Lyria 3 is Google DeepMind's music generation model, served on Segmind as a text-to-audio API. From a single text prompt, it composes a 30-second, 44.1 kHz stereo clip complete with vocals, timed lyrics, and full instrumental arrangements — or an instrumental-only backing track when you ask for one. Lyria 3 also accepts up to 10 reference images, so you can turn a photo's mood, color, and subject into a matching soundtrack. Before it renders audio, the model reasons through musical structure (intro, verse, chorus, bridge) to keep the composition coherent from the first note to the last.
Lyria 3 is built for creators who need custom, royalty-aware audio fast: social and short-form video soundtracks, background music for games and apps, marketing jingles, podcast intros, lo-fi study loops, and demo songs with sung hooks. The image-to-music workflow is ideal for auto-scoring campaign assets or matching a track to a brand photo. In testing, a single prompt reliably produced a full-band, radio-ready mix with clean lead vocals in about 30 seconds.
Be specific: name the genre, instruments, BPM, key, and mood. Use section tags or timestamps to shape progression, and paste your own lyrics for sung vocals. Add "instrumental only, no vocals" for a clean backing track. Vague prompts yield generic results, so layer detail. Results vary between calls since generation is non-deterministic.
Does Lyria 3 generate vocals and lyrics? Yes — it sings time-aligned lyrics and can also produce instrumental-only tracks.
How long are the clips? Each generation is a fixed 30-second, 44.1 kHz stereo MP3.
Can I generate music from an image? Yes — supply up to 10 reference images to guide mood and style.
Does it support other languages? Yes — lyrics are generated in the language of your prompt.
Are outputs watermarked? Yes — every track includes an imperceptible SynthID watermark.
Can I request longer, full-length songs? For multi-minute tracks with detailed structure, use Lyria 3 Pro.