Step 4. Video Creation
Step 4 · Video creates short video clips that will later be used by the final Builder. This is not the final edit. It is the production stage where FOXI takes images from Step 3 or text descriptions of scenes, sends generation tasks to video models, waits for completion, downloads clips into the project, and shows them in project files.
FOXI Studio is a cloud service: you do not need a powerful local computer, a separate rendering app, or constant manual supervision. You start the step from the browser, FOXI queues the tasks, places the required FX on hold, processes generations on the server side, and returns the unused hold if some operations are not needed.

Where to find this step
Open your project and go to 04 · Video in the top Pipeline navigation. You can also start this step separately from the manual stage controls on the left by clicking Step 4 · Video.

Usually Step 4 is started after:
- Step 1 · Text — the script, video structure, and metadata are ready.
- Step 2 · Voice — the voice track exists, so FOXI understands timing and scene meaning.
- Step 3 · Images — frames, character, visual style, and thumbnails are ready.
If you run the full cycle, FOXI performs the steps in order. If you are testing video generation only, you can start Step 4 separately, but it is better to first make sure the project already contains images and voiceover files.
What this step produces
After successful completion, video clips appear in project files. In Project Files, they are usually shown as video files in the clips section or inside the Output_video folder.

These clips are not the final YouTube video yet. The final video is assembled by Builder: it combines voice, images, clips, subtitles, music, transitions, avatar, branding overlays, and AI-detect protection.
Main switch: enable or disable this step
At the beginning of the section there is a setting called Enable this step (video creation).

Enable Step 4 if:
- your channel needs live AI video clips instead of only slideshow-style editing;
- you want to animate the best images from Step 3;
- the video should look more cinematic or realistic;
- the topic includes scenes that are hard to communicate with static images only;
- you are ready to spend more FX for a more dynamic result.
Disable Step 4 if:
- a video made from images, camera movement, subtitles, music, and transitions is enough;
- you are doing a mass topic test and want to save FX first;
- the images are not good enough and should be fixed before animation;
- your goal is to quickly test the script, voice, and Builder without expensive video generation.
Practical advice: for a new channel, first make 1–2 test videos with a minimal number of clips, check the style and retention, and only then increase the amount of video generation.
Video creation mode
The Video creation mode field decides what FOXI uses as the basis for clips.

| Mode | When to choose it | What to check |
|---|---|---|
| RunWay · Animate generated images | The safest starting point. You already see the Step 3 frames and animate only the best images. | Image quality, character consistency, absence of defects, and how many images to animate. |
| RunWay · Create video from text | You need separate clips from text descriptions, and images are not required as the base. | Prompts, 5 or 10 second duration, clip count, and cost. |
| VEO · Create video from text | You want a more cinematic text-to-video result and accept higher cost or waiting time. | Prompt meaning, model limits, waiting time, and FX budget. |
Changing the mode also changes the visible settings on the page: for image-to-video FOXI shows image animation fields, while for text-to-video it shows prompt generation, duration, and parallel task settings.
RunWay mode: animating generated images
This is the recommended mode for most channels. FOXI takes ready frames from Step 3 and turns them into short video clips. The advantage is control: you can already see the visual style, character, composition, colors, and possible artifacts before paying for video generation.

How many first images to animate
How many first images to animate defines how many images from the beginning of the storyboard FOXI sends to the video model.
Recommendations:
- 3–5 clips — a cheap test of a new style or topic;
- 6–8 clips — a good balance for a normal video;
- 10+ clips — only if you are already confident in image quality and budget.
Do not set a large number on the first run. Video is usually more expensive than text, voice, and images, so unnecessary clips can quickly increase the final cost.
Additional image numbers
Additional image numbers to animate in the video (separated by commas) is useful when you want to animate not only the first frames, but also specific strong images from the middle of the video.
Example: if the first 6 frames will already be animated, but you especially like frames 12 and 18, enter:
12,18
This helps control the budget and select only the most important scenes.
Animation prompt
The Animation prompt tells the video model how to animate the image. By default, FOXI uses a safe instruction: soft cinematic motion, preserving composition, lighting, and objects, without sudden changes or distortion.

If you edit the prompt, describe not only the desired motion but also the restrictions:
- preserve the face or character without morphing;
- do not change objects or background;
- do not add new objects;
- avoid sudden camera jumps;
- keep the movement slow and natural.
Good example:
Cinematic slow camera push-in, subtle natural motion, preserve the same character, same lighting, same composition, no morphing, no new objects, no distortion.
Bad example: make it epic. This is too generic, so the model may change the scene unpredictably.
Text-to-video modes: RunWay and VEO
In RunWay · Create video from text and VEO · Create video from text, FOXI first prepares prompts for the video model and then generates clips without necessarily using images as the base.

This mode is useful if:
- you need abstract or cinematic inserts;
- you do not want to use images from Step 3;
- the topic is better communicated through motion rather than a still frame;
- you are ready to control prompts and cost more carefully.
Number of videos
Number of videos (Text-to-Video) defines how many clips FOXI creates in this mode. For the first test, set it to 1. After checking the result, you can increase the number.
Prompt generation model
Prompt generation model writes or improves prompts for video. If you are not sure, keep the default value. Stronger models usually understand channel style and script context better, but they may cost more.
Runway video duration
For RunWay, you can select 5 or 10 seconds. Short clips are cheaper to test and easier to replace. Longer clips provide more movement but increase waiting time and cost.
Delays and retries
Time between Gemini requests and Prompt retry attempts help make the process more stable. If the model does not respond temporarily, FOXI retries the request. Do not reduce delays too aggressively: this can cause rate-limit errors or unstable responses.
Concurrency and waiting for results
Video generation takes longer than text or images. That is why the page includes technical limits:

- Max concurrent AI generations — how many clips FOXI may send to processing in parallel.
- AI status check interval — how often FOXI checks whether clips are ready.
- Max AI generation wait time — how long to wait for one task before timeout.
- AI generation retry attempts — how many times to retry generation after a temporary error.
For beginners, it is safer to keep the default values. Increase concurrency only if you understand the budget and want to speed up a series. With many parallel tasks, FOXI may temporarily hold more FX because several clips are being generated at the same time.
Important: if the log shows that the task is waiting for a result, FOXI is not necessarily frozen. A video model can render a clip for several minutes. Open Dashboard or the activity log and watch the latest status.
Cost and FX
Step 4 is usually one of the most expensive stages because video models require more compute. The cost depends on the mode, model, clip duration, retries, and the number of parallel tasks.

How to control expenses:
- start with a small number of clips;
- animate only the best images;
- do not run a long text-to-video test before checking one clip;
- check your FX balance before launch;
- review the cost history in Tokens and costs;
- remember that FOXI may place funds on hold during the task and return the unused part.
For many videos, the total average video cost may be around $1–2, but with many AI video clips the cost will be higher. Use Step 4 intentionally: it adds motion, but not every video needs it.
Prompt History and AI Assistant
You do not need to edit Step 4 prompts blindly. If the result looks wrong, open the AI Assistant and describe the issue in simple words:
- “the camera moves too sharply”;
- “the character changes between frames”;
- “the model adds extra objects”;
- “the clips are too expensive for a test”;
- “make the prompt calmer and safer for a medical channel”.

The AI Assistant sees project context and can suggest a prompt edit based on the channel topic. After changes, check Prompt History: FOXI stores snapshots and lets you restore the previous version if the new one performs worse.

How to check the result
After Step 4 finishes:
- Open Project Files.
- Find the clips section or the Output_video folder.
- Open several
.mp4files directly in the browser. - Check that there is no face morphing, extra objects, hard cuts, or overly fast camera movement.
- If a clip is bad, delete or regenerate it, or reduce the clip count and improve the prompt.

If the clips are good, continue to Step 5 · Builder. Builder will assemble the final video, add subtitles, music, transitions, branding, and prepare the result for publishing.
Common issues
The clip looks good but does not match the script
The prompt is usually too generic. Clarify the action, object, camera, and restrictions. If you use image-to-video, check that the source image actually matches the scene.
The character changes or becomes distorted
Add restrictions against morphing and face changes to the animation prompt. For channels with a consistent character, image-to-video based on pre-checked Step 3 frames is usually safer.
The task stays in waiting status for a long time
Video models can render for several minutes. Check the log. If the wait time exceeds the limit, FOXI will show an error and you can retry the task.
FX spending is higher than expected
Most likely, too many clips were generated, long text-to-video clips were enabled, or retries were triggered. Reduce the clip count, start with one test, and review the cost history.
Builder did not use the clips
Check that the clips really appeared in project files and that Builder is configured to use video materials. If there are no clips, rerun Step 4 or build the video from images.
Safe launch order
For the first video, use this order:
- Generate the script and voiceover.
- Create images and check their quality.
- In Step 4, enable image-to-video and animate 3–5 best frames.
- Check clips in project files.
- If everything looks good, run Builder.
- After the final video, decide whether to increase the clip count for future videos.