Direct a Voice Like an Actor, Not a Slider

A director leans toward a voice actor at the glass of a small recording booth, one hand raised as she gives a note.
Illustration generated with Adobe Firefly.

The problem

You have the right voice. Every line still comes out at the same energy, like a very calm person reading a menu. So you drag sliders around, get a different flavour of flat, and conclude the tool cannot act.

It can. You just have to direct it the way you would direct a person: with the text, the punctuation and a note before the line. In Eleven v3, the script is the performance.

What actually moves a v3 performance

  1. Audio tags in square brackets. [whispers], [sighs], [laughs], [frustrated], [curious], [excited], [sarcastic]. Put them right before the words they should affect. They are directions, not decorations: one or two per line.
  2. Punctuation as timing. An ellipsis adds a pause and weight. A full stop is a beat. A comma is a breath. Short sentences read as tension. Long ones read as ease.
  3. Capitals as emphasis. "I told you TWICE" lands on the word you capitalised. Use it once per line at most.
  4. Stability as the amount of freedom. Creative gives the most expressive reads and reacts most to tags, with an occasional wild take. Natural is balanced. Robust is steady and mostly ignores direction. Emotional scenes: Creative or Natural. Narration: Robust.
  5. Context length. Very short prompts produce inconsistent reads. Give the model at least a couple of sentences, even if you only keep one. Feed it the previous line as context and cut it in the edit.
  6. The voice sets the ceiling. A calm library voice will not shout convincingly with a [shouts] tag. Direction works inside the voice's natural range. If a line needs a range the voice does not have, recast, do not push.

Stability in v3

  1. Creative

    Most expressive, can wobble

  2. Natural

    Balanced. Start here

  3. Robust

    Steady, ignores most direction

Steps

  1. Write the line as it would be in a screenplay, then add the parenthetical as a tag: (quietly) becomes [whispers] or [softly].
  2. Read it out loud yourself. Wherever you naturally pause, put an ellipsis or a full stop. Wherever you press, capitalise one word.
  3. Set stability to Natural, generate three takes. Move to Creative if all three are too safe.
  4. Keep the best take. Regenerate only lines that miss, not the whole scene.
  5. For dialogue, use separate voices per speaker in one prompt. v3 handles multi-speaker scripts with a speaker label per line.

Worked example

Flat script:

You said you would be here. You said it twice.

Directed script:

[tired] You said you'd be here. ... [quietly] You said it twice.

Same words. The first reads like a complaint. The second reads like someone who has stopped expecting anything. Three takes on Natural, keep the one where the pause before "quietly" feels a little too long.

Same words, three reads

Flat

You said you would be here. You said it twice.

Directed

[tired] You said you'd be here. ... [quietly] You said it twice.

Tag + capital

You said you'd be here. [quietly] You said it TWICE.

Two-hander:

Speaker 1: [flat] The last bus left twenty minutes ago.
Speaker 2: [sighs] ... I know.
Speaker 1: [curious] Then why are you still standing here?

Your move

Try this today

Take one line of dialogue you already generated flat. Write three directed versions: one with only punctuation changes, one with one tag, one with a tag and a capital. Generate all three on Natural. Pick the one an actor would be proud of. That is your new default way of writing lines.

Stop waiting – start creating!

All 16 tutorials