Complete Guide to Generating Videos with AI9app Agent Mode: One Sentence Is All It Takes for AI to Plan the Entire Workflow
AI9_Studio ·
Complete Guide to Generating Videos with AI9app Agent Mode: One Sentence Is All It Takes for AI to Plan the Entire Workflow You no longer need to choose the model yourself or write long, highly…
Complete Guide to Generating Videos with AI9app Agent Mode: One Sentence Is All It Takes for AI to Plan the Entire Workflow
You no longer need to choose the model yourself or write long, highly detailed prompts. AI Agent Mode will decide which models to use, how many steps the workflow requires, and how each prompt should be written. It then presents the complete generation plan for your review. Only after you approve the plan will generation begin and credits be deducted.
This guide covers everything from enabling Agent Mode to building multi-scene video workflows, including how each feature works, its limitations, and how fees are calculated.
What Is AI Agent Mode?
In a standard AI generation workflow, you choose the model, adjust the settings, and write detailed prompts yourself. Agent Mode hands this part of the process over to AI: you describe what you want in one sentence, the agent creates the plan, and you review and approve it.
In other words, Agent Mode is not a separate generation tool. It acts as a “director and producer” in front of the existing AI tools.
It can:
- Understand uploaded images and videos. It can analyze videos up to three minutes long and create a new plan based on their visual style, camera movements, pacing, and editing rhythm.
- Plan multi-step workflows, such as generating a reference image, animating it, adding a voice-over, and creating subtitles.
- Automatically select the most suitable model and Prompt Skill.
- Automatically use the output from one step as reference material for the next step.
Before You Begin: How to Enable Agent Mode
- Sign in to AI9 and open the AI Media Agent page.
- Find the AI Agent / Agent Mode switch above the generation area, next to the Auto switch, and turn it on.
- The AI Agent chat panel will slide out from the right. On mobile, it opens in full-screen mode. This panel will be your main workspace.
When the panel opens, you will see the following message:
Describe what you want to create. The agent will select the most suitable models and skills, write the prompts, and ask for your approval before generation begins.
You will also see several example options under Try It, such as:
- Recreate a video style
- Use a YouTube video as a reference
- Turn a storyboard into a video
- Create an e-commerce product video
- Create an AI influencer video
- Animate a photo
- Generate a series of product-detail images
Select an example to insert the template into the input box, then replace the content inside the brackets with your own information.
Five Steps to Create Your First AI Video
The entire process, from a single sentence to the final result, takes only five steps.
Step 1: Clearly Describe What You Want
Describe the result you want in the input box. The more specific your instructions are, the more accurate the plan will be.
At minimum, include these four details:
- Subject
- Style
- Aspect ratio
- Duration
Example:
Create a 15-second e-commerce product video for a white ceramic insulated bottle. Use clean studio lighting, a slow orbiting camera movement, and a premium visual style. Use a vertical 9:16 format and end with a front-facing close-up of the product.
You can also paste a YouTube link and ask the agent to create a plan based on the video’s cinematography, editing rhythm, and color grading.
Step 2: Attach Reference Materials
This step is optional, but strongly recommended.
Use the paperclip icon on the left side of the input box to attach media. You can either upload files from your computer or select files from My Media, which contains files you previously uploaded or generated.
There are also three faster ways to add reference materials:
- Select a result from the conversation: Enable Select Media to Use, choose the images or videos you want to reuse, and then send your message.
- Drag and drop: Drag files directly into the panel.
- Templates and effects: Select a template or video effect on the page to attach it automatically. Select it again to remove it.
There are two important limitations:
- Each image or audio file must be no larger than 15MB.
- Each video file must be no larger than 80MB.
A video can only be used as a generation reference when its duration is supported by the selected model. For example, Seedance supports up to 15 seconds, while Omni supports up to 10 seconds.
If the video is too long, the agent can still analyze it and use it to write prompts, but it will not use the video directly for continuation or extension.
Step 3: Understand the Generation Plan Card
After you send your message, the agent will first explain its approach and then produce a Generation Plan card.
This is the most important part of Agent Mode. The card shows:
- The type of each step, such as image generation, video generation, clip editing, virtual avatar, voice generation, or subtitles.
- The model and settings used for each step, including resolution, duration, audio generation, and number of outputs.
- The complete prompt, which can be opened and edited directly from the card.
- The prompt-source badge. Written by Skill means the prompt was created using your selected Prompt Skill. No Skill Used — Written by Agent means no suitable skill was available, so the agent wrote the prompt itself.
- The reference materials used in each step, which can be removed individually.
- The estimated credit cost and total cost. If your balance is too low, the card will display Insufficient credits to complete the full plan.
The plan is fully editable before approval. You can:
- Rewrite prompts
- Change models
- Adjust the duration
- Remove reference materials
A single plan can contain up to 12 steps and 12 outputs. Each prompt can contain up to 3,500 characters.
Step 4: Approve and Generate
Once you have reviewed the plan, select Approve and Generate.
Credits are only deducted after you approve the plan and generation begins.
During execution, each step displays its current status:
- Waiting
- Generating
- Completed
- Failed
- Skipped
You can cancel any step that has not started yet.
There is one important rule: if a step produces no output at all, for example because of insufficient credits or a validation error, all remaining steps will stop automatically. This prevents later steps from running without the reference materials they require.
Step 5: Continue from the Results
Once the results appear in the conversation, you do not need to describe everything again. Simply continue the conversation.
For example:
- Select I want to edit this… or I want to regenerate this… under a result. The agent will automatically use that file as a reference.
- Type instructions such as: “Change the camera angle in Scene 2 to a low-angle shot,” or “Add a 10-second voice-over to this image.”
- After all scenes in a multi-scene video have been completed, the agent will ask whether you want to Combine scenes into an MP4. This creates a single preview file with one click.
- For videos longer than one minute, use the AI9 Video Editor to assemble the scenes. The panel also includes an Insert into Video Editor option, allowing you to transfer selected images, videos, and audio files directly into the editor.
Detailed Feature Guide
The following section explains the purpose of each feature in the Agent panel.
Auto Mode vs Manual Approval Mode
The Auto Mode switch controls how the Generation Plan card is presented.
Manual Approval Mode — Default
This mode displays a complete and editable plan card. You can review and adjust every item individually.
It is best for important projects or situations where you want to fine-tune the prompts and settings.
Auto Mode
The agent automatically decides whether to use direct generation, reference images, or a storyboard-based workflow.
Instead of showing the complete plan, it presents a simplified Auto Plan Summary containing the selected models and estimated cost. Select Okay, Start Generating to continue.
Remember: Auto Mode means one-click approval. It does not mean silent or automatic execution. Credits will never be deducted until you confirm the plan.
Agent Models: Fast, Professional, and Custom
Fast — Gemini 3.5 Flash Lite
This is the fastest and most economical option. It is recommended as the default model.
Professional — Gemini 3.6 Flash
This model offers stronger reasoning capabilities and is better suited to complex, multi-scene or multi-step plans.
Custom
This option allows you to select a model manually and view its input and output rates per one million tokens.
Prompt Skills: Use a Slash Command
Type a forward slash / in the input box to open the Skills menu.
Skills contain professional prompt-writing guidelines for specific types of content, such as cinematic camera direction or e-commerce product photography.
After you select a skill, the agent will use it to write the prompt for each individual scene. The Generation Plan card will display the Written by Skill badge.
Keep in mind that skills are called separately for each scene. Therefore, the more scenes you have, the higher the agent-usage cost will be. For example, eight scenes require eight separate prompt-writing calls.
You can also type @ to quickly select a reference tag.
Memory and Saved Conversations
Select the brain icon in the panel’s title bar to access two sections.
Memory
Use Memory to save information you want the agent to remember in every new conversation.
For example:
I sell skincare products. I prefer clean product photography with white backgrounds and warm cinematic color tones.
You can also select Pin Current Skill to make a particular skill your preferred skill.
Saved Conversations
You can save the current conversation and later reload, rename, or delete it.
After a conversation is saved, it will be saved automatically every five minutes.
One particularly important detail is that when you save a conversation, the media generated in that conversation is also copied to My Media. This prevents the files from being deleted after 24 hours and ensures they remain available when you reload the conversation later.
If your My Media storage is full, the system will ask you to increase your storage capacity before saving.
When starting a new project, select New Conversation. The system will first ask whether you want to save your current conversation.
Video Workflows: Three Available Approaches
When your request requires more than one short video clip, the agent will ask you to choose one of three workflows.
1. Reference Material Workflow
The agent first creates or uses a character or product reference image to lock the main subject’s appearance. It then generates each video segment based on that reference.
This approach provides the best visual consistency.
2. Storyboard Workflow
The agent first generates a nine-panel storyboard with consistent characters and visual styling, with one scene in each panel.
It then generates each video clip based on the corresponding storyboard panel.
This workflow is suitable for narrative videos and multi-scene stories.
3. Direct Generation
This is the fastest method and is most suitable for a single shot or a simple scene.
For multi-scene projects, the agent uses a Step N structure, passing the visual result from one segment into the next to maintain continuity.
Uploaded videos are not automatically used for continuation. They will only be used when the agent clearly understands that you want to continue or extend the uploaded video.
This prevents every scene in the plan from being attached to the same video clip.
Requirements for Virtual Avatar Steps
If a plan includes a virtual-avatar step, the agent can switch to the correct mode and write the prompt for you.
However, you must prepare the character image in the Virtual Avatar settings first. OmniHuman also requires an audio file.
The Generation Plan card will display a reminder when these materials are required.
How Are Fees Calculated?
Agent Mode has two clearly separated types of charges.
Agent Conversation and Planning
Planning, media analysis, and prompt writing consume your AI Usage allowance.
The cost is calculated according to token usage and includes a 35% service markup.
When your allowance runs out, you can purchase an additional usage package:
- Lite
- Advanced
- Professional
Your remaining agent allowance is always visible in the panel.
Actual Content Generation
Images, videos, voice-overs, and subtitles are charged according to each model’s normal credit rate.
The cost is exactly the same as generating the content manually. Agent Mode does not secretly increase the generation price, and no generation credits are deducted before you approve the plan.
Practical Tips
Use Reference Materials to Define the Style
Instead of trying to describe “cinematic” entirely through text, attach a video you like and say:
Keep the same visual style, camera movements, and editing rhythm, but replace the subject with my own content.
Test a Small Plan Before Expanding It
Ask the agent to generate one test shot first.
Once you are satisfied, say:
Use this style to create the complete six-scene version.
Make Use of Memory
Save your brand style, product category, and preferred aspect ratios in Memory. You will not need to repeat these details in every new conversation.
Edit the Prompt Directly
Prompts in the Generation Plan card are editable.
When a prompt is not quite right, editing it before approval is usually faster than restarting the conversation.
Use the Video Editor for Longer Videos
For videos longer than one minute, use the agent to generate the individual scenes and then assemble and finish the project in the AI9 Video Editor.
Frequently Asked Questions
These are the four most common questions users ask before using Agent Mode.
Will the Agent Deduct Credits Without My Permission?
No.
All generation tasks require you to select either Approve and Generate or, in Auto Mode, Okay, Start Generating.
Why Was My Uploaded Video Not Used for Continuation?
To prevent every scene from being attached to the same clip, uploaded videos are treated as analysis references by default.
To continue an uploaded video, clearly state:
Continue this video.
You should also confirm that the video duration is within the selected model’s limit, such as 15 seconds for Seedance or 10 seconds for Omni.
Why Did the Plan Stop Halfway Through?
If a step produces no output, usually because of insufficient credits or a settings-validation error, all later steps will stop automatically.
Add more credits or correct the settings for the failed step, then approve the plan again.
How Long Are Results Kept in a Conversation?
Media from unsaved conversations is deleted after 24 hours.
Use Save This Conversation to copy the media into My Media and keep it permanently.
Conclusion
The greatest value of Agent Mode is not that it simply presses buttons for you.
Its real value is that it gives AI responsibility for the decisions that normally require the most experience: choosing the right model, writing the prompts, and connecting reference materials between steps—while leaving the final decision entirely in your hands.
Starting today, describe the result you want in one sentence and let the agent propose a plan.
All you need to do is review it, make any necessary adjustments, and select Approve.