Skip to main content

Step 2 · Voiceover

02 · Voice is the second production stage in FOXI Studio. It takes the finished script from 01 · Text and turns it into an audio track for the future video.

Good voiceover strongly affects retention. Even a strong script feels weak if the narrator has the wrong tone, speaks too fast, is too monotonous, or sounds synthetic.

02 · Voice tab in FOXI Studio

What Step 2 does

At this stage FOXI:

  • takes the final script from Step 1;
  • cleans technical notes and extra formatting;
  • sends the text to the selected voice engine;
  • receives the audio track;
  • applies post-processing if needed;
  • saves the result in project files;
  • passes audio to images, video, and Builder stages.

Voiceover generation flow

What beginners should choose

If you are just starting, do not test dozens of voices. Choose one understandable voice template, make a short test, and judge the result in the finished video.

For the first test, check:

  • the correct project is selected;
  • Step 1 has already created a script;
  • voice language matches text language;
  • a voice template is selected;
  • speed is not too high;
  • voice Anti-Detect is enabled if available on your plan;
  • FX balance is sufficient.

Main voiceover settings

Voice provider

The project may have several voiceover options. Availability depends on plan and current integrations.

ProviderWhen it fits
GeminiFast start and simple voice setup
ElevenLabs/Lumean templateMore natural and stable narrator via a saved template
Backup ElevenLabs modelsAlternative route or a different voice character
FOXI internal modelsThe service manages connection and charges FX internally

If you do not understand the difference, keep the recommended option or ask the AI Assistant.

Example request:

Choose voice settings for an English channel about mysterious history: serious documentary tone, male narrator, US audience aged 30–60.

ElevenLabs voice template

A voice template is a saved narrator setup. It may include selected voice, model, speed, stability, and other parameters.

Templates are usually created during Quick Setup or in base project settings. After a template is created, it can be selected in Step 2.

ElevenLabs voice template

Good practice: use clear template names, such as US documentary male, Calm doctor 50plus, or Mystery narrator female.

Gemini and director instructions

If Gemini voiceover is used, not only voice and language matter. Director instructions are also important. This is a text description of how the narrator should read the script.

Example:

You are an experienced documentary narrator. The voice is calm, confident, slightly dark. Read without advertising tone, pause before important facts, and do not overact.

Gemini TTS settings

Do not make instructions too long or contradictory. If you ask for “very fast”, “very dramatic”, and “calm” at the same time, the result may become unstable.

Speed, stability, and style

Some providers let you configure speed, stability, similarity, and style.

Simple recommendations:

ParameterFirst-test choice
SpeedAround 1.0; slightly slower for older audiences
StabilityMedium value so the voice is not too monotone or chaotic
StyleIncrease carefully if emotion is needed
SimilarityDo not max it out before testing; artifacts may appear

If the narrator speaks too fast, reduce speed and test again before changing the entire voice.

Pronunciation dictionary and word replacements

If the voice mispronounces specific words, use word replacements or a pronunciation dictionary.

Example:

замок=замо́к
творог=творо́г
Lead=leed

Pronunciation dictionary

Add only words that are actually pronounced incorrectly. A very large dictionary can create new mistakes.

Voice Anti-Detect AI

Voice Anti-Detect AI is post-processing that helps make audio less sterile and more natural. It may include equalization, compression, limiter, and other audio filters.

Voice Anti-Detect AI

It is useful when the voice sounds:

  • too synthetic;
  • too perfect and sterile;
  • flat;
  • not studio-like;
  • too harsh in high frequencies.

Availability may depend on your plan.

Removing long pauses

The Remove long pauses after voiceover option shortens empty gaps after the voice track has already been generated and downloaded. It is useful when the narrator leaves long stops and the final video feels slower than needed.

When the checkbox is enabled, the Pause removal strength field appears:

  • Soft — carefully shortens only noticeable pauses when you want to keep a calm pace.
  • Normal — recommended for most videos; makes speech tighter without harsh cuts.
  • Strong — compresses pauses more aggressively when you need a faster, more dynamic video.

Start with Normal. If the voice becomes too dense, switch to Soft. If there are still too many pauses, choose Strong.

Launching voiceover generation

Run Step 2 separately if you want to check only the voice before images and video.

This is useful when:

  • you choose a new voice;
  • you change speed;
  • you test pronunciation dictionary;
  • you want to understand voiceover cost;
  • you do not want to spend FX on later stages before checking audio.

Launching voiceover generation

After launch, watch the log. It shows whether the task was accepted by the server, which stage is running, and where an error happened if there is one.

What to check after completion

After generation, listen to audio or check it in preview.

Check that:

  • narrator fits the audience;
  • there are no wrong pronunciations;
  • pace is not too fast;
  • voice does not sound too commercial;
  • emotions are not overacted;
  • pauses feel natural;
  • audio does not distort or clip;
  • duration roughly matches the script.

Finished voice track

If you do not like the voice

Do not rerun the entire full cycle. Fix Step 2 first.

You can:

  • choose another template;
  • reduce or increase speed;
  • change director instructions;
  • add a word to the pronunciation dictionary;
  • enable or disable Anti-Detect;
  • ask the AI Assistant to choose parameters.

Example request:

Voiceover sounds too commercial and fast. Choose calmer settings for a documentary channel and explain what to change.

Common mistakes

MistakeWhy it is badBetter approach
Changing voice before checking scriptThe problem may be in text, not voiceCheck Step 1 first
Setting speed too highViewers struggle to follow informationStart with normal speed
Not checking voice languageNarrator may read text incorrectlyChoose voice for the channel language
Ignoring pronunciation dictionarySame mistakes repeatAdd problem words to the dictionary
Running full cycle just to test voiceExtra FX is spentRun only Step 2

What to ask the AI Assistant

  • “Choose a voice for my audience and niche.”
  • “Why does the voiceover sound too synthetic?”
  • “How do I make the narrator calmer and more documentary-like?”
  • “Which words should I add to the pronunciation dictionary?”
  • “Should I enable voice Anti-Detect for this project?”
  • “Compare two voice options and tell me which is better for retention.”

Where to go next

If the voiceover sounds good, continue to Step 3 · Frames. In the next stage, FOXI prepares images, thumbnails, and the visual foundation of the future video.