Turn a short written story + your photos and clips into a finished, captioned vertical video β no editing skills needed. This guide explains every button and field in plain language, with the exact words you can type.
You give it a story written as short lines and a folder of photos/videos. For each line it decides what should be on screen, how the camera moves, and the mood β then builds one 9:16 vertical video with captions and music, ready for Reels / Shorts.
Think of it as an automated video editor with a film director's brain. Each line of your story becomes one shot. A 10-line story becomes a 10-shot reel.
| Field | What to type | Example |
|---|---|---|
| Name (optional) | The main person's first name, if you want to pin it. | Lalita |
| Description (optional) | A short look so AI-made pictures stay consistent (helps most with the Consistent face tick). Fine to leave blank. | woman, 30s, strong features, traditional clothing |
The story is plain text. Start each shot with the word Frame and a number, then write the line that should appear on screen. That's it.
Reels Frame 1 From a farmer to a modelβ¦ the story of a desi girl with big dreams. Frame 2 At 19, I got married into a traditional Rajasthani home. Frame 3 A kisan ki beti β I rode tractors even after marriage. Caption: Write your full Instagram caption here. (This does NOT appear in the video.)
| Rule | Plain meaning |
|---|---|
| One Frame = one shot | Keep each line short β it shows on screen as a caption while the camera moves. |
| Numbers don't need to be in order | The tool reads top-to-bottom, not the numbers. |
The Caption: block at the bottom | Is your Instagram post text. It is never shown in the video β write the long version here. |
| Keep on-screen lines under ~15 words | The camera moves before longer text can be read. |
In story mode, start with either I already have a frame script for full control, or I have a story to let the AI create an editable frame draft from raw notes/transcript. The AI does not render directly from the raw story; it fills the script/frame cards so you can review first.
Click π Browse folderβ¦ (next to the start panel) and pick the folder on your computer that holds the photos/videos for this story. The tool uploads them and matches the right one to each shot automatically.
After Parse Frames β, each line becomes a card. Top to bottom you choose: the picture, then optionally a mood note, a photo edit, a camera move, a lip-sync toggle, and the length.
| Button | Use it when⦠|
|---|---|
| Auto | You're happy with the photo it auto-picked from your folder. |
| π From Folder | Use the matched file from your folder (shown as a green badge). |
| π· Upload Photo | Pick any photo or video from your computer for this shot. |
| π¨ AI Portrait | No good photo? The AI creates a person/portrait for this moment, matching the Subject + your note. |
| πΌ AI Symbolic | The AI creates objects/scenery with NO people β perfect for sad/abstract beats (illness, loss, a turning point) where showing a face feels wrong. |
This tells the AI director the feeling and look of the shot. Write it like you're briefing a photographer: what's in frame + how they feel + the light.
| Feeling you want | What to type in the note |
|---|---|
| Pride / Triumph | Head high, direct gaze, warm gold backlight, full confidence |
| Grief / Loss | No eye contact, slumped posture, cold blue-grey light, objects rather than face |
| Longing | Eyes looking just off-frame, window light, half-turned away |
| Determination | Jaw set, hands busy, warm practical light, grounded posture |
| Fear / Uncertainty | Shadow across the face, shallow focus, background looms out of focus |
| Innocence | Young face, soft diffused light, looking up slightly, clean background |
| Turning point / Hope | Spark in the eyes, phone glow in a dark room, a small smile starting |
| Vague note | Specific + emotional note |
|---|---|
| show her being sad | Medicine bottles on a windowsill, no person, cold winter light, the silence of an empty room |
| the ramp walk | Head high, back straight, heels clicking, full confidence, warm gold backlight β she earned this |
Camera is a dropdown, not a text box β no typing. It's pre-set to the AI's best pick for that beat (you'll see a β¨ auto note). To change it, just choose another move from the list below. Leave it on β¨ Auto to let the director decide at render. After π Preview Stills, the β¨ Suggest from image button re-picks the best move by looking at the actual frame. Here's what each move does:
| Type this | What you'll see | Best for |
|---|---|---|
360 orbit | Camera circles all the way around | Hero reveal, triumph |
bullet time | Freeze + orbit (Matrix style) | The single peak moment |
crash zoom in | Fast dramatic zoom toward the subject | Shock, surprise, realization |
crash zoom out | Fast zoom away | Sudden scale, "the world opens up" |
dolly in / push in | Smooth glide toward the subject | Intimacy, building emotion |
dolly out / pull back | Smooth glide away, revealing surroundings | Loneliness, scale, isolation |
crane up | Camera lifts upward | Victory, freedom, rising |
crane down | Camera lowers | Weight, defeat, gravity |
tilt up / tilt down | Camera angles up or down | Looking to the sky / looking down |
overhead | Bird's-eye view from directly above | Vulnerability, isolation |
dutch angle | Tilted, off-kilter horizon | Tension, unease, something's wrong |
Hitchcock zoom | Vertigo effect (zoom + pull) | Dread, disorientation |
extreme close on eyes | Macro on the face/eyes | Deep emotion, connection |
arc left / arc right | Camera sweeps in a partial circle | Gentle reveal, momentum |
handheld | Slightly shaky, human feel | Raw truth, documentary realism |
static | No movement at all | Stillness, weight, gravity |
super 8mm | Vintage film-grain look | Memory, flashback, nostalgia |
whip pan | Fast blurred swipe | Energy, time passing, scene change |
slow push in). Over-describing motion is the #1 cause of weird, melty faces. And remember: a shot is either a camera move or a talking face (lip sync) β not both.The βοΈ Image edit box changes the picture before it animates. Type a plain instruction.
| Type this | Result |
|---|---|
add thunderstorm and dark clouds | Storm added to the sky |
make the lighting warmer and golden | Warm sunset tone over the whole image |
add rain on the window | Rain streaks added |
add soft morning fog | Atmospheric fog layer |
cold grey winter light and frost | Turns a scene cold and sombre |
Big mood changes ("add storm", "warmer light") work great. Precise moves ("shift her to the left") are less reliable β that's expected.
Tick π Lip Sync and that shot becomes a talking face that says the line aloud in a real voice. A voice menu appears β pick one or leave the default.
Check β Voice-Over Track at the top to make the entire video narrated. Each caption gets read aloud in a chosen voice, creating a full voice-over narration layered under the music.
| Field | Meaning |
|---|---|
| Duration | How long the shot lasts. Leave blank and it's auto-calculated from word count (longer line = longer shot). Lip-sync shots set this themselves. |
| Video start (only for video sources) | Skip the first few seconds of a clip. e.g. 3 starts the clip 3 seconds in β handy when the good part is mid-way. |
"Model" = which AI engine makes the picture or the motion. Leave both on Auto and the tool picks the best one for each shot, cheap while testing and premium for the final. Only override if you want a specific look.
| Name | Plain meaning | Tier |
|---|---|---|
| Seedream | Cheap, fast pictures β great for testing | draft |
| Nano Banana | Premium, very photo-real faces | premium |
| Flux | Premium portraits & objects | premium |
| GPT Image | Best for objects, symbolic scenes, and any text-in-image | standard |
| Name | Plain meaning | Tier |
|---|---|---|
| Kling Standard | Solid animation, great value β the testing default | draft |
| Kling Pro | Highest-quality Kling motion | premium |
| Higgsfield | Most cinematic camera-move presets | premium |
| Seedance | Cinematic, multi-shot feel | premium |
| Veo | Best built-in audio/dialogue | premium |
| Hailuo | Wide landscapes, realistic motion | premium |
| Ken Burns | Simple zoom, free, no AI β for free tests | free |
| Tier | Use when |
|---|---|
| Dev cheap, 5s shots | Always start here. Cheap draft to check your story, captions, and timing. |
| Production premium, up to 9s | Only for the final render once everything looks right. |
| Mood | Look |
|---|---|
| Warm Nostalgic | Amber, golden-hour, slightly vintage |
| Cold Struggle | Blue-grey, overcast, deep shadows |
| Triumphant | Rich golds & saffron, bright, high saturation |
| Default | Let the AI choose per shot |
| Setting | Options | Pick |
|---|---|---|
| Orientation | Portrait 1080Γ1920 / Landscape 1920Γ1080 | Portrait for Reels/Shorts/TikTok |
| Transition | Crossfade / Hard Cut | Crossfade for smooth, Hard Cut for punchy |
| Caption font | Baskerville / Montserrat / Satoshi* / Arial / Georgia / Helvetica | Baskerville for storytelling, Montserrat for modern & clean |
| Caption size | 24β96 pt | 52 default; 60β70 for dramatic |
| Caption colour | White / Yellow / Black | White β always readable |
| Caption position (default) | Bottom / Middle / Top | Bottom for Reels β override per frame if a shot needs it elsewhere |
| Burn captions | On / Off | Untick for a clean video with no subtitles. Voice-over (if on) still plays. |
| Max lines per caption (default) | No limit / 1 / 2 / 3 | Caps how many lines a caption takes; long text auto-shrinks to fit. 1β2 lines keeps it clean. |
| Option | What it does |
|---|---|
| No Music | Silent video β add your own later. |
| Upload Music | Use your own MP3/M4A/WAV. It loops and fades out at the end. |
| Auto-Generate | Describe a mood and it makes a track (~2β3 min). Wait for β Music ready before Generate. |
Good auto-music prompts: Emotional Bollywood instrumental, struggle to triumph, sitar and tabla, no lyrics Β· Melancholic Rajasthani folk, raw acoustic
Combine music + voice-over: Enable Voice-Over Track (above) to layer narration on top of your music. Music plays at 25% volume, so the voice is clear and the music sits underneath.
After you parse, a cost estimate appears and updates live as you change settings. Check it before pressing Generate. Want it cheaper? Use Dev tier, prefer real photos/videos, and keep lip-sync to a few shots.
You don't have to render the whole video and hope. Treat the tool as a draft machine you steer: preview the stills, fix the frames you don't like, approve the rest, and only pay to animate what you've approved. Three controls make this fast.
After Preview Stills, each frame card shows a π Redo still button. Change that frame's note, photo, image edit, or camera β then click π. Only that one image regenerates (a few seconds), and it swaps into the card. Nothing else is touched, and you don't pay to re-animate anything.
Each frame card has a β Approved for animation tick (on by default). Untick any frame you're not happy with. When you press Generate:
While a render runs, each frame's clip appears in its card the moment it's done β you don't wait for the whole video. The card shows β³ animatingβ¦, then flips to a looping π¬ clip ready preview. If an early clip looks wrong you can stop, fix that frame with π, untick it, and re-run β without having waited for all ten.
==double equals==, e.g. I had ==nothing== left, to colour that phrase in captions.[layout: text_card], for a full-screen bold statement card.Caption: block. Brand copy remains operator-supplied.edit_list.json for a human editor.| Shot | Job |
|---|---|
| Frame 1 | Hook β who is this, why watch? Use your best clip or a crash-zoom. |
| 2β3 | Context β the before. |
| 4β5 | Conflict β the problem / loss (use AI Symbolic for pain). |
| 6β7 | Lowest point, then the spark (great lip-sync moment). |
| 8β9 | Turning point β triumph (real footage of the win; save 360 orbit for here). |
| 10 | Resolution β strong, present-day, direct gaze (great closing lip-sync line). |
360 orbit, crane up) for the single biggest moment.Every button has a typed shortcut you can put under a frame line. The buttons do the same thing β use whichever you prefer.
| Type under the line | Same as the button⦠|
|---|---|
[photo: filename.jpg] | Use a specific file |
[photo: ai_portrait] / [photo: ai_symbolic] | AI Portrait / AI Symbolic |
[note: ...] | Director Note |
[camera: 360 orbit] | Camera motion |
[edit: add storm] | Image Edit |
[lipsync: yes] | Lip Sync toggle |
[duration: 8] Β· [start: 3] | Duration Β· Video start |
A real story turned into a 10-shot reel. Paste it, set the folder, hit Parse β then read why each choice was made.
Reels Frame 1 From a farmer to a modelβ¦ the story of a desi girl with big dreams. [photo: ai_portrait] [note: Strong proud Rajasthani woman, direct gaze, chin up. Golden hour, dust in air. A HERO, not a victim.] [camera: crash zoom in] Frame 5 After COVID, I was diagnosed with arthritis. I lost my hair, my confidenceβ¦ and myself. [photo: ai_symbolic] [note: NO person. Medicine bottles on a windowsill, a hairbrush with fallen strands, cold blue-grey light. The objects carry the grief.] [edit: cold grey winter light and frost on the window] [camera: static] Frame 7 Bedridden, I watched modelling videos thinking, "one day I'll walk the ramp." [photo: lalita_face.jpg] [lipsync: yes] [note: The turning point β she says it herself. Phone glow on her face in a dark room. A small smile starting.] Frame 8 I fought back β and the girl who couldn't stand walked the ramp in heels. [photo: ramp_walk.mov] [start: 3] [note: Real ramp-walk clip. Head high, full confidence. The peak. Skip the first 3 seconds.] [duration: 8] Frame 9 I won Mrs. Rajasthan 1st runner-up. [photo: crown.jpg] [note: Pure victory β crown, sash, stage. She earned this. High saturation, saffron and gold.] [camera: 360 orbit] Caption: From the farm to the ramp β the full story for your followers goes here.
| Frame | Why |
|---|---|
| 1 | No strong opening photo β AI Portrait hero shot. Crash zoom stops the scroll in the first 1.5s. |
| 5 | Illness shown as objects, not a face (AI Symbolic) + cold edit + static camera = the weight of stillness. |
| 7 | Lip sync on the spark β hearing her say it lands harder than text. |
| 8 | Real video of the win beats any AI. start: 3 skips the boring intro. |
| 9 | 360 orbit saved for the single highest moment. |
| Problem | Fix |
|---|---|
| Face looks melty/distorted | Simplify the camera move (one gentle action), or use a real photo. |
| Wrong photo on a beat | Click π Upload or From Folder on that card and pick the right one. |
| Caption too long / gets cut | Shorten the line to under 15 words. |
| Costs look high | Switch to Dev, use real photos, cut lip-sync to 2β4 shots. |
| Lip-sync shot came out silent/animated | That shot fell back safely β check the line isn't empty and try again. |
| Testing for free | Set Video model to Ken Burns + Quality to Dev β zero AI cost. |
| Don't like one image | Click π Redo still on that frame after Preview β only that image regenerates. |
| Only want to pay for some frames | Untick β Approved on the frames you're unsure about β they animate free (Ken Burns); the cost panel drops. |
| Don't want to wait for the whole render | Clips appear in each card as they finish β watch the early ones, stop & fix if a frame looks wrong. |
| Wrong speaker on a quoted line | Use the π Speaker dropdown on that frame card to pick the right person. |
| AI chose wrong gender/age for a character | Fill the Subject β Description field with gender + age, or leave blank and the AI infers from the story. |
Click π Brand / Ad mode in the header (or go to /brand) to switch. The engine is identical β same pipeline, same frame cards, same models β but the UI gains a brand brief panel, real-asset uploads, mandatory checks, and a product beat toggle per shot.
| Field | What to put here | Notes |
|---|---|---|
| Brand name | The brand's name | Appears in the "Paid partnership" disclosure line |
| Product | The product category or name | Used in visual context β not on screen |
| Objective | What the ad should achieve | e.g. product awareness, trial drive |
| Key message | The brand's exact claim β verbatim | Never rephrased by AI |
| CTA text (required) | The call-to-action line shown on the end card | e.g. "Try it now Β· Link in bio" |
| CTA link | URL or "Link in bio" | Shown on the end card |
| Tagline | Brand tagline (optional) | Optional β only set if the brand supplied it |
| Brand colour | Hex code for the end card | e.g. #0a7d33 |
| Announcer script | Exact voice-over lines | Read aloud as-is β never rewritten |
| Asset | Required? | What it does |
|---|---|---|
| Logo | β Required | Placed on the end card; optionally shown as a logo bug in the corner throughout the video |
| Product photos/videos | Recommended | Used on frames you mark π Product beat β never AI-replaced |
| Show logo in corner | Optional | Tick to burn a small logo bug into every shot; pick corner (top-right default) |
Brand mode separates the announcer voice-over from the background music. Both are decided per-project.
| Setting | Options |
|---|---|
| Announcer VO | AI draft β reads your announcer script in a chosen ElevenLabs voice (fast, for review) Β· Brand-supplied audio β upload the final signed-off VO file |
| Background music | AI music β auto-composed from the ad mood Β· Brand-supplied music β upload their cleared track |
Need an ElevenLabs voice ID? Paste it in the VO voice field. Leave blank for the default voice.
Write the script exactly like a story script β one caption per Frame. The caption is the on-screen text the brand supplies. After parsing, each frame card gets a π Product beat toggle.
| Beat type | What to do |
|---|---|
| Lifestyle / story beat | Write the brand-supplied line. Leave product beat off. The AI makes the visual. |
| Product / logo shot | Tick π Product beat. Upload a real product photo in the Brand kit. That photo is used as-is β no AI generation. |
| End-card beat | Auto-generated by the tool from your CTA text, logo, and brand colour. You don't write it β it's appended automatically. |
If you click Generate Ad with something missing, the tool refuses to spend any credits and shows a red checklist. You cannot generate until all boxes are ticked. This protects against accidentally publishing an ad without a logo or disclosure.
| Mandatory | How to clear it |
|---|---|
| Brand logo (upload the logo file) | Upload a logo file in the Brand kit section |
| Call-to-action text | Fill the CTA text field in the Brand brief panel |
| At least one π Product beat | Tick the π Product beat toggle on at least one frame card |
| Disclosure (auto-added) | This one is always satisfied β the tool adds "Paid partnership with [Brand]" automatically |
| Problem | Fix |
|---|---|
| Hard-block checklist appears when I click Generate | Complete each item in the red checklist β upload the logo, fill CTA text, tick a product beat. |
| Product photo isn't showing on the product beat frame | Upload a product photo in the Brand kit section, then tick π on the right frame. |
| Extract fields filled in the wrong thing | Just edit the field directly β extracted values are pre-fills, not locked. You always win over the AI. |
| VO is the wrong voice | Paste an ElevenLabs voice ID in the VO voice field, or switch to brand-supplied audio. |
| I want to test the look without spending VO/music credits | Use π Preview Stills β it generates images only, no animation, no audio spend. |
| End-card text is wrong | Edit the CTA text field and the brand name field β the end card is built from those. |
Open π¬ Studio mode from the top link. Instead of writing a script, you describe the reel in plain words and the AI drafts the shots. It uses the same engine (Preview, cost, Generate, export) plus a reusable identity library.
Save a Talent (a face) and/or a Product (a photo + specs) once. Studio locks them across every shot so the same person and the same product appear consistently β and they're reusable in future reels. On shots you mark π Product beat, the real product image is used directly (never re-generated), so logos and fine detail stay exact.
Pick a scope β Commerce (product / fashion / jewelry ads) or General (any idea) β write your brief, and click β¨ Plan shots. The AI returns editable shot cards (on-screen line, camera move, shot size). Edit any card before Preview/Render. The AI may draft the on-screen lines here (you edit them); all text still passes the safety check.
Each shot card adds a Talent/Product selector, a Negative prompt (what to avoid), and a Continuity lock (outfit/styling that must not change). Sensible defaults are prefilled; tweak per shot as needed.
Story mode workflow: Write short lines β Parse β set Picture + Note + Camera per shot β Dev test β Production final β Download.
Brand mode workflow: Fill brief β upload logo + product β write ad script β Parse β tick π product beats β Preview Stills β Generate Ad β Download.
Studio mode workflow: Save Talent/Product β pick scope β write brief β β¨ Plan shots β edit cards β Preview Stills β Generate Reel β Download.
When unsure, leave models on Auto and start in Dev.