Skip to main content
When an agent uses Inworld with the inworld-tts-2 model, its replies can carry delivery tags — natural-language directions in square brackets that the voice acts out:
The platform handles the mechanics automatically. On every voice call with inworld-tts-2, the agent is taught how tags work, when a human would use them, and how to keep them in English regardless of the conversation language. You do not need to explain tags in your agent prompt — you only steer them.
Delivery tags are supported only by inworld-tts-2. Other Inworld models and other TTS providers ignore this feature. In text chat mode tags are suppressed entirely.

What the platform does for you

With inworld-tts-2 selected, every voice conversation automatically gets:
  • Tag mechanics — the agent knows the tag syntax, keeps tags in English even when speaking another language, and repeats tags on consecutive sentences (each sentence is synthesized separately, so an untagged sentence falls back to neutral delivery).
  • Human-like defaults — warm greetings and farewells, a light playful touch when the caller jokes, an apologetic and measured delivery with irritated or busy callers, and clear articulation when confirming times, addresses, and numbers. Plain informational sentences stay untagged.
  • Speech-friendly writing — numbers, dates, times, and phone numbers are written out in words, punctuation stays conversational, and markdown or special symbols are kept out of spoken text.
If that default behavior fits your use case, you are done — no prompt changes needed.

Steering delivery from your agent prompt

There are three levels of control, from loose to exact. All of them work in any prompt language — the emitted tags stay in English.

1. Describe the voice

Add a short natural-language description of how your agent should sound. The agent maps it to tag moods on its own:
A persona described this way measurably changes which tags the agent picks — for example, playful tags appear on jokes only when the description allows playfulness, and never leak into conversations with irritated callers.

2. Map situations to tags

For precise control, write a policy in situation → tag form and include the tag literally. The agent uses your exact tag at the right moment:

3. Script exact moments

Flow rules can trigger a specific delivery at a specific moment. Write the tag inline in the rule:
A literal tag written in a rule passes through verbatim, every time. A loose stage direction (“laugh it off and deny it”) also works, but only for the action itself: the agent will reliably emit [laugh], while the surrounding delivery drifts toward the agent’s persona. When the exact mood matters, write the delivery tag too — for example [laugh] [say playfully]. Non-verbal tags are fixed sound effects: use them exactly as listed — [laugh], [sigh], [breathe] — without modifiers. A modified non-verbal like [laugh sinisterly] is not recognized. To color the moment, pair the bare non-verbal with a separate delivery tag for the sentence.

Tag reference

Tags are free-form natural language, not a fixed list. These categories are good building blocks: Combining qualities in one instruction produces a more convincing performance than a bare tag: [say sadly with deliberate pauses in a low voice] layers mood, rhythm, and pitch.
Write non-verbal tags with spaces and without modifiers, exactly as documented — [clear throat], not [clear_throat] or [laugh sinisterly]. Do not combine opposing directions in one tag ([whisper in a hushed style] together with [very loud]), and do not apply a tag that contradicts the text it introduces.

Emphasis and stress

To stress a single word, capitalize it — the voice emphasizes capitalized words:
Use this sparingly, on whole words only. Inworld does not document accent-mark stress (for example, combining diacritics on Russian words), so verify such conventions on your actual voice before relying on them.

Configuration

Delivery tags require no extra configuration beyond selecting the model:

TTS Providers

Full Inworld configuration reference

Voice & Speech

Basic voice configuration