Skip to main content

Choosing a model

SongCraft generates full songs from a prompt. MiniMax has notable audio depth too, and Mureka is worth comparing.
ElevenLabs is the default for text-to-speech and voice work.
Filter for audio-to-audio. ElevenLabs and MiniMax both offer it.
Look at text-to-audio models — they cover effects as well as music. Read the model descriptions, since capability names don’t distinguish the two.
Audio quality is subjective in a way image quality isn’t. Generate the same prompt on two or three models and listen. There’s no substitute for that.

Voice and video together

Two steps. Generate the voice track (Generate Voice chained action or ElevenLabs), then sync it — Add Lipsync on an image, Lipsync Video on a clip, or Sync directly.
Yes — audio-to-video models generate video from an audio track. Kling and LTX both have them.
HeyGen handles presenter video directly, rather than assembling voice and lipsync yourself.

Costs

Often per second, like video — so a long track costs proportionally more. The cost is shown before you submit. See Credits & Plans.
Generate a short clip to check the voice or musical style before committing to full length. Style is audible in a few seconds.

Practical

My Assets → Generated, alongside images and video. Download from there or from the preview.
Yes — upload it in the form or pick it from your assets, for lipsync, audio-to-video, or audio transformation.
Yes — the Video Editor assembles clips and audio. Generate the pieces, then arrange them there.
Check status — audio is its own service group. Failed runs aren’t charged.