The marketplace for AI-generated voice fashions is very large. Artistic use instances require AI voice fashions to be extra expressive, whereas enterprises trying to automate buyer assist and gross sales ops want them to be extra steerable.
Palo Alto-based Fish Audio desires to cater to all of these use instances with its library of greater than 15,000 pure language controls. Since launching final 12 months, the startup right this moment has greater than 8 million individuals utilizing the open-source or hosted variations of its fashions, and now generates annual recurring income of $21 million.
To proceed constructing on that traction, the startup on Tuesday stated it has raised $50 million in a seed spherical that was led by Coreline Ventures and Capital Right this moment. The funding additionally noticed participation from 359 Capital, Parable, Play Time, Alphalist Companions, Bayhouse Ventures, Carya Enterprise Companions, and HF0.
Fish Audio began as a small challenge by former NVIDIA researcher Shijia Liao, who, annoyed by non-expressive artificial voices out there in the marketplace, skilled a voice era mannequin on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has greater than 31,000 stars, and is utilized by indie builders, online game designers, and creators.
The corporate has launched 5 fashions within the final 12 months: 4 speech era fashions and one speech-to-text mannequin. It has open-sourced three of its speech era fashions, however its newest S2.1 Professional mannequin is obtainable solely by means of its paid API.
Fish Audio presents paid month-to-month plans fitted to creators and groups that unlock a set variety of minutes of era, plus voice cloning options. The corporate additionally presents an enterprise model of its APIs and platform, and says organizations like HeyGen, Sanas and Plaud are already utilizing it.
“Each enterprise has totally different use instances and totally different preferences. For instance, corporations like HeyGen, which use our voices to energy AI avatars, need realism in voices; a gaming studio would need expressive voice for his or her characters; and voice agent corporations like LiveKit need extra natural-sounding and low-latency voices which can be expressive sufficient for calls,” Cao stated.
A method the startup has constructed its library of voices is by merely asking customers to submit their very own voices for coaching its fashions, and compensating them if their voices are used. That resulted in some hassle a couple of months in the past, nonetheless, as some creators alleged that their voices have been uploaded to Fish Audio with out their consent. The startup had a DMCA content material take-down course of in place to deal with such issues, however the take-downs themselves took a very long time.
Fish Audio’s CEO and co-founder Rissa Cao informed TechCrunch that the corporate has now automated the take-down course of. Creators can simply submit a brief voice pattern or a contract to show that an uploaded voice belongs to them, and their voice can be taken off the startup’s platform in lower than 3 minutes, she stated.
Nonetheless, that doesn’t stop anybody from importing an artist’s voice with out their information. And till the artist finds out, their voice will proceed for use on the platform till they file for it to be taken down.
Oskue Honda, a accomplice at Coreline Ventures, stated a community-driven mannequin solely works when creators belief the platform.
“A community-centric method can solely grow to be a sturdy benefit if creators belief the platform. Meaning consent, transparency, and attribution should be constructed into the product slightly than handled as afterthoughts. I imagine the business wants to maneuver towards verified voice possession, clear licensing phrases, simple reporting and takedown processes, and ultimately revenue-sharing fashions the place creators profit financially when their voices are licensed or used commercially,” he stated.
Cao stated when the startup was solely providing its product as an open-source challenge with plans for creators, it was working effectively and didn’t want cash. However it needed to develop extra superior fashions, and in addition needed to accommodate enterprises as investor curiosity was ramping up, which led it to hunt capital.
Trying forward, Fish Audio plans to launch an audio understanding mannequin this 12 months. It’s additionally constructing a speech-to-speech mannequin.
The speech era market is crowded, with corporations like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp competing for creators and enterprises’ wallets.
In accordance with Rico Mallozzi, a accomplice at 359 Capital, fine-grained controls for builders and cost-efficient mannequin coaching will assist Fish Audio compete higher with huge AI labs.
“I feel what they’ve been in a position to construct, state-of-the-art fashions, with the group they’ve, in comparison with a few of these different well-funded AI labs or corporations, is unbelievable. It exhibits their technical acumen in closing the hole between artificial-sounding and human-like voices,” Mallozzi informed TechCrunch over a name.
Whenever you buy by means of hyperlinks in our articles, we may earn a small commission. This doesn’t have an effect on our editorial independence.

