Suno built its name turning text prompts into songs. Now it wants to score the spoken word, too.
The AI music company opened a public beta of Speech on October 1 after a month of limited testing. The feature generates a synthetic voice reading a script — or improvising from a short description — and composes original background music for it in the same track, rather than forcing creators to stitch narration and score together afterwards. Tracks run up to about eight minutes, long enough for a poem, a guided meditation, a product demo or a bedtime story.
The workflow comes in two flavours. A simple mode takes a brief description and produces both voice and music; an advanced mode accepts a full custom script with controls for voice, style and variation. The music can be switched off entirely, turning the tool into a plain voice generator — a direct challenge to established text-to-speech services.
Suno is candid that beta means beta. The company acknowledges accents can drift mid-recording and pauses can stretch into unintended drama, and it has declined to explain in detail how the model was trained — a sensitive point while major record labels pursue copyright lawsuits against AI music generators.
The launch continues a brisk product cadence: Suno rolled out its v6 music models on September 9, and its chief executive recently reported more than two million subscribers and annual recurring revenue above $300 million. Speech suggests where the company thinks the bigger market lies — not just in songs, but in every podcast intro, audiobook chapter and voiceover that currently requires a human, a microphone and an afternoon.
Related reading: BloomX Raises $13 Million to Send Robotic Pollinators Into the Field · Arena Raises $200 Million to Measure How AI Agents Behave · Y Combinator's Garry Tan Urges US Open-Weight Labs to Learn From Frontier Models