Fish Audio Raises $52M Seed to Build AI Voice Models for Creators and Enterprises
Fish Audio, a Palo Alto startup building AI voice generation models, has raised a $52 million seed round to expand its offerings for both individual content creators and enterprise customers, backed by a mix of…

Fish Audio, a Palo Alto startup building AI voice generation models, has raised a $52 million seed round to expand its offerings for both individual content creators and enterprise customers, backed by a mix of investors led by Coreline Ventures and Capital Today.
Who built it and how far it's already come
The company was founded by Shijia Liao, a former Nvidia researcher, alongside co-founder and CEO Rissa Cao, and launched just last year. In that short window, Fish Audio has already built a genuinely substantial user base: more than 8 million users across its open-source and hosted product versions, $21 million in annual recurring revenue, and a GitHub repository that has attracted more than 31,000 stars, a meaningful open-source community signal for a voice AI company still in its first full year of operation.
What the product actually offers
Fish Audio has shipped five distinct models over the past year, four focused on speech generation and one on speech-to-text conversion, alongside a library of more than 15,000 natural language voice controls that let users fine-tune exactly how generated speech sounds. The company's latest and most capable model, S2.1 Pro, is available exclusively through the paid API rather than the free or open-source tiers, reflecting a fairly standard freemium structure where the most advanced capability is reserved for paying enterprise customers.
Who's actually buying it
Fish Audio's enterprise customers reportedly include HeyGen and Sanas, both companies themselves operating in AI-driven communication and video spaces, suggesting Fish Audio has found traction as an infrastructure layer other AI companies build on top of, rather than only serving individual creators directly. CEO Rissa Cao described the enterprise sales approach around that diversity of use case: "Every enterprise has different use cases and different preferences," a framing that emphasizes customization over a one-size-fits-all voice product.
Why the investor pitch centers on community trust
Coreline Ventures partner Osuke Honda offered a specific rationale for backing the company, tying its investment thesis directly to Fish Audio's open-source roots: "A community-centric approach can only become a durable advantage if creators trust." That statement points to a broader strategic bet, that Fish Audio's large open-source user base and GitHub following aren't just a marketing footnote but represent a genuine moat, provided the company maintains the creator trust that got it there in the first place.
A voice-cloning controversy it had to clean up
That trust question isn't purely theoretical. Some creators had previously alleged that their voices were uploaded to the platform without authorization, a serious concern for any AI voice company given how directly voice cloning technology can be misused to impersonate real people without consent. Fish Audio has since automated its takedown process, reportedly now able to resolve disputed voice uploads within three minutes, a fast response time clearly aimed at addressing exactly the kind of trust concern Honda's investment rationale explicitly depends on.
What the funding actually buys
With $52 million in fresh capital, Fish Audio's next challenge is scaling from its current position, a fast-growing but still young startup with real revenue and a loyal open-source following, into a more durable enterprise voice AI platform, all while maintaining the community trust its own lead investor has identified as central to its competitive advantage. Given how quickly voice cloning controversies can damage a company's reputation in this space, how well Fish Audio's automated takedown system continues performing at scale may matter as much to its future as any new model it ships.
Comments
No approved comments yet.



