Startup Gimlet Labs is solving the AI inference bottleneck in a surprisingly elegant way | TechCrunch

Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2. Book Exhibit Table Now. Disrupt doors open Oct. 13. Get your pass and bring someone with you at 50% off. REGISTER NOW. Stanford adjunct professor and successfully exited founder Zain Asgar just raised an $80 million Series A for a startup that solve the AI inference bottleneck problem in an astute way. The round was led by Menlo Ventures. The company, Gimlet Labs, has created what it claims is the first and only “multi-silicon inference cloud” which is software that allows an AI workload to be simultaneously run across diverse types of hardware. It can split an AI app’s work across both traditional CPUs and AI-tuned GPUs, as well as high-memory systems. “We basically run across whatever different hardware that’s available,” Asgar told TechCrunch. A single agent may chain together multiple steps, and each “requires different hardware: Inference is compute-bound; decode is memory-bound; and tool calls are network-bound,” writes lead investor, Menlo’s Tim Tully, in a blog post about the funding. No chip yet does it all, but as new hardware gets rolled out, and aging GPUs get redeployed, “the multi-si
Source: For the complete article, please visit the original source link below.




