Startup Fundraisingβ€’

Wafer Raises $40M Series A for AI Inference Optimization

Wafer secures $40M Series A led by Marathon Management Partners and Chemistry to enhance AI inference efficiency and reduce compute costs.

Share:
AM
Alvaro de la Maza

Partner at Aninver

Stay ahead of the market

Get instant notifications when new news matching "Artificial Intelligence (AI), Technology, Software & Gaming in United States" are published.

Key Takeaways

  • Wafer raised $40.0M (Series A) from Fifty Years, Liquid2, Y Combinator, Marathon Management Partners, Chemistry, Wing Venture Capital, AMD Ventures, Outset Capital, Sony, Universal, Warner Music, Goldman Sachs Alternatives.
  • Sector: Artificial Intelligence (AI), Technology, Software & Gaming.
  • Geography: United States.

Analysis

Wafer, a startup focused on enhancing the efficiency of artificial intelligence model deployment, has successfully closed a $40 million Series A funding round. This significant capital infusion, co-led by Marathon Management Partners and Chemistry, with participation from Wing Venture Capital, AMD Ventures, and Outset Capital, alongside continued support from seed investors Fifty Years and Y Combinator, positions the company to aggressively scale its operations. The funding round values Wafer at over $200 million, marking a substantial increase from its earlier seed valuation.

The core innovation at Wafer lies in its AI-driven approach to optimizing inference workloads. The company's autonomous agents are designed to analyze traffic patterns and performance constraints of AI models, automatically identifying the most efficient configurations across various software engines and hardware platforms. This "AI that optimizes AI" strategy directly addresses a critical industry pain point: the underutilization of expensive computing resources. Reports suggest that typical GPUs operate at only around 20% utilization, leading to significant wasted expenditure. Wafer aims to rectify this by automating a process that often requires weeks of manual tuning by specialized engineers.

A key differentiator for Wafer is its focus on making non-Nvidia hardware more competitive in the AI inference space. In recent internal testing, the company demonstrated that AMD's MI355X processor, when optimized by Wafer's agents, achieved approximately 80% of the throughput of Nvidia's B200 for specific models like GLM-5.2, at less than half the cost. Across a broader range of open-source models, Wafer's technology has shown speedups ranging from 2x to 2.8x compared to unoptimized deployments.

The company was founded by University of Chicago alumni Emilio Andere (CEO) and Steven Arellano (co-founder). Arellano brings valuable experience from his work on Google Bard's infrastructure and his tenure at the quantitative hedge fund Two Sigma. This strong technical and operational foundation was previously recognized with a $4 million seed round in April, led by Fifty Years with contributions from Liquid2 and Y Combinator, and notable angel backing from figures like Jeff Dean and Wojciech Zaremba of OpenAI.

The Series A round also attracted an impressive roster of angel investors, including Jeff Dean (recently departed from Google to co-found Discovery Loop), Guillermo Rauch (CEO of Vercel), Andy Fang (co-founder of DoorDash), Kyle Vogt (CEO of The Bot Company), Akshay Kothari (COO of Notion), Matthew Prince (CEO of Cloudflare), and Scott Stephenson (CEO of Deepgram). This level of strategic angel investment underscores the perceived market potential and technological merit of Wafer's solution.

Wafer's success arrives amidst a dynamic market for AI inference optimization, with several venture-backed companies emerging from open-source projects. Competitors like Inferact (built on vLLM) and RadixArk (from the SGLang engine) have also secured substantial funding. Baseten, a model deployment platform, achieved a significant valuation with investment from Nvidia itself. Wafer distinguishes itself by concentrating on optimizing the entire deployment stack rather than developing its own inference engine or serving platform, a strategic choice that appears to resonate with investors and customers alike.

The newly acquired capital will be directed towards further automating the inference optimization process, aiming to provide every deployment with the equivalent of a dedicated performance engineering team that continuously enhances performance per dollar. This strategic focus on cost-efficiency and performance at scale is crucial in the rapidly expanding AI infrastructure market, which is projected for substantial growth in the coming years.