Give Your Character a Voice

A woman with cropped grey hair in a green jacket speaks into a broadcast microphone in a dim recording booth, lit by one warm lamp.
Illustration generated with Adobe Firefly.

The problem

The visuals are done. Then the character opens their mouth and it sounds like a navigation system. Flat voices kill AI films faster than bad images do, because we forgive a strange frame and never forgive a fake person.

You have three ways to get a voice, and picking the right one for your character matters more than any slider.

Three ways to get a voice

  1. Voice Library. Thousands of ready voices with tags for age, accent, tone. Fastest. Good for side characters and narration. Filter by use case ("characters", "narrative") and audition with a real line from your script, not the sample text.
  2. Voice Design. Describe the voice in text ("a tired woman in her 50s, low register, slight Berlin accent, speaks slowly") and generate options. Best when the character does not exist yet and you want something nobody else has.
  3. Instant Voice Clone. Upload one to two minutes of clean recorded speech and clone it. Best when you or an actor friend can perform the voice. For the current v3 model, an Instant Voice Clone or a designed voice is the recommended route; Professional Voice Clones are not yet fully tuned for v3.

Voice Library

Fast

Side characters and narration.

Voice Design

Unique

The character doesn’t exist yet.

Instant Voice Clone

Performed

You or an actor friend can do the voice.

Steps

  1. Write the voice like a casting note. Age, energy, pace, register, accent, one emotional default ("guarded", "amused", "exhausted"). This note feeds Voice Design directly and helps you audition Library voices with intent.
  2. Audition with a real line. Paste two lines from your script: one neutral, one emotional. A voice that only sounds good on the demo sentence is not your voice.
  3. Pick the model on purpose. Eleven v3 is the expressive model with audio tags and multi-speaker dialogue. It shines on performances. For long, calm narration where consistency matters more than range, the v2 family is steadier.
  4. Set stability first. In v3 the stability setting has three positions: Creative (most expressive, occasional wobble), Natural (closest to the reference, balanced), Robust (very stable, ignores most direction). Start on Natural. Move to Creative for emotional scenes, Robust for narration.
  5. Record a proper clone sample if cloning. Quiet room, no music, no reverb, one consistent mic distance, and cover the emotional range the character needs. If the character shouts in the film, shout in the sample. The clone cannot invent a range it never heard.
  6. Generate per line, not per scene. One take per line lets you regenerate the weak line without touching the good ones. Name files by scene and line number from day one.

Worked example

Casting note: "Mara, late 30s, dry, tired, speaks in short sentences, warm underneath, slight rasp."

Voice Design prompt: A woman in her late thirties with a dry, tired delivery and a slight rasp. Low to mid register. Speaks in short, clipped sentences with warmth underneath. Neutral German-English accent.

Audition lines:

  • Neutral: "The last bus left twenty minutes ago."
  • Emotional: "You said you would be here. You said it twice."

Generate three Voice Design options, run both lines through each, pick the one where the second line hurts a little.

Your move

Try this today

Write one casting note for one character. Audition three Library voices and two Voice Design results with your two lines. Pick one. You now have a cast member. Tomorrow you can direct them (see Direct a Voice Like an Actor, Not a Slider).

Stop waiting – start creating!

All 16 tutorials