Voice AI startups’ largest unlock has been dealing with requires enterprises in areas like gross sales, advertising, and buyer help. Giant organizations are offloading calls to voice mannequin builders like ElevenLabs and Deepgram; infrastructure corporations like Vapi, Retell, and LiveKit; and devoted buyer help outlets like Decagon and Sierra.
San Francisco-based Rime is attempting to realize an edge on this crowded market with its voice AI fashions which might be educated on conversational knowledge that it data, aiming to scale back its shoppers’ customization load.
Based in 2022 by former Stanford PhD scholar Lily Clifford, ex-Amazon Alexa engineer Brooke Larson, and Stanford engineer Ares Geovanos, Rime has constructed a recording studio in San Francisco to gather its personal conversational knowledge reasonably than counting on scraping the net for audio.
The startup mentioned it focuses on tuning its voice fashions to nail the pronunciation of various model entities and industry-specific phrases. It employs a phoneme-based structure to adapt to totally different pronunciations in order that prospects don’t need to retrain fashions for his or her particular {industry}.
Rime on Wednesday mentioned it has raised $24 million in a Sequence A funding spherical that was led by M13 Ventures. Twilio Ventures, Corazon Capital, Uncommon Ventures, and different present buyers additionally participated.
Clifford mentioned that regardless of progress in voice AI improvement, enterprises nonetheless choose legacy IVR (interactive voice response) implementations, as AI voice expertise nonetheless can’t match as much as IVR’s effectiveness.
“The voice expertise remains to be not there to automate the overwhelming majority of enterprise telephone calls. LLMs have made it so much simpler to construct voice functions that work, however they haven’t modified the way it feels to work together. Speaking with a voice AI agent is just not essentially the most compelling expertise for the top consumer. It’s kinda like a brand new IVR, however with a greater voice,” she mentioned.
The startup started with a pipeline of separate fashions for speech-to-text, text-to-speech, and a big language mannequin. However it’s now shifting focus to develop higher speech-to-speech fashions to scale back latency, enhance turn-taking, and sort out points like background noise. The brand new strategy will even serve to lower reliance on orchestration, so the corporate doesn’t need to handle a bunch of fashions.
Rime says it has prospects in meals service, healthcare, airways, and fintech. The corporate claims that due to its coaching knowledge and mannequin positioning, prospects keep longer on the decision, which has helped it win enterprise contracts from shoppers like Mayo Clinic, Dialpad, Upstart, and Asurion.
With the brand new funding, Rime is planning to broaden its crew of 35 individuals, aiming to rent for mannequin improvement, engineering, and partnerships. It not too long ago introduced on Rafael Valle, who labored on audio understanding at Meta Superintelligence Labs and Nvidia’s utilized deep studying audio analysis crew, as its chief scientist.
“Firms like ElevenLabs have moved into being an orchestration and the appliance layer, going face to face with the Sierras and Decagons of the world. I believe there’s simply a lot extra to be carried out technically, and Rime’s strategy of pushing ahead on the very best mannequin with low latency and excessive reliability in a regulated setting stands out,” M13’s Morgan Blumberg informed TechCrunch.
It had beforehand raised $5.5 million in a seed round last May. Blumberg is becoming a member of the startup’s board as a part of the fundraise.
Once you buy by hyperlinks in our articles, we may earn a small commission. This doesn’t have an effect on our editorial independence.

