Key Takeaways
- Fish Audio raised $52.0M (Seed) from Coreline Ventures, Capital Today, 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners.
- Sector: Artificial Intelligence (AI), Technology, Software & Gaming.
- Geography: United States.
Analysis
Fish Audio, operating as Hanabi AI Inc., has successfully closed a substantial $52 million seed funding round, signaling a significant advancement in the quest for natural-sounding artificial intelligence voices. The capital infusion was co-led by prominent venture firms Coreline Ventures and Capital Today, with robust participation from a diverse group of investors including 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners, alongside several undisclosed angel investors.
This funding propels Fish Audio into a competitive arena, aiming to establish voice as the primary interaction method for AI systems. The company's journey began with a focus on overcoming the limitations of early, robotic AI speech. Co-founder and Chief Scientist Shijia Liao, drawing from his background in AI research at Nvidia Corp. and a passion for expressive media, initiated the project to develop more lifelike synthetic voices. What started as a personal endeavor with limited computational resources has blossomed into a sophisticated platform.
The core of Fish Audio's offering lies in its advanced text-to-speech and voice cloning capabilities. Its open-source project, initially gaining traction on GitHub with over 31,000 stars, quickly demonstrated its appeal to developers, content creators, and game designers seeking more emotive voice generation. The platform now provides granular, word-level emotional controls driven by over 15,000 natural language prompts, allowing for precise tuning of tone, inflection, and pacing. This level of control is a significant departure from the monotonous synthetic voices of the past.
Demonstrating remarkable technical prowess, Fish Audio claims its technology can clone a voice from a mere five-second audio sample in under 15 seconds. The platform boasts native support for 83 languages and has reportedly outperformed leading competitors in blind listening tests, with 67% of participants preferring its outputs. These capabilities have fueled rapid user growth, attracting over eight million users and generating more than $21 million in annual recurring revenue.
Beyond its initial focus on creative industries, Fish Audio is expanding its reach into regulated sectors like healthcare and financial services. For these clients, the company offers secure on-premises deployments with stringent data privacy measures, including zero-data retention and HIPAA compliance. This strategic expansion positions Fish Audio as a significant contender against established players like ElevenLabs Inc., which recently secured substantial funding at a high valuation, highlighting the immense investor interest in the voice AI market, estimated to be a multi-billion dollar opportunity with strong projected growth.
With this new capital, Fish Audio plans to broaden its technological scope beyond text-to-speech. The company aims to develop a comprehensive audio-native stack, incorporating voice-native large language models and real-time speech-to-speech translation tools. Investments will also target building an enterprise sales force and enhancing developer tools through API integrations with partners such as Retell AI Inc. and LiveKit Inc.. To further accelerate adoption, Fish Audio will make its flagship S2.1 Pro model accessible to all developers via its official API free of charge starting late August.
Coreline Ventures managing partner Osuke Honda emphasized the strategic importance of voice as an emerging AI interface. He noted Fish Audio's consistent innovation in performance, multilingual capabilities, emotional nuance, and cost-effectiveness, solidifying its position as a preferred choice for a global user base ranging from individual creators to large enterprises.