AI Tutorial Hub
#ChatGPT#Google Flow#ElevenLabs#CapCut

How to Create Viral Vox-Style Paper Cut Animations Using Free AI Tools

Workflow Overview

Discover how to produce documentary-grade paper-cut animations in the signature style of Vox and Fern using completely free AI tools. This guide walks you through using an end-to-end Master Prompt in ChatGPT to script and generate visual beats, batch-rendering paper collage images and animations in Google Flow, generating voiceovers in ElevenLabs, and editing the final cut in CapCut.

Featured Workflow Stack

Tools Used in This Tutorial

Try these tools to replicate the exact results

ChatGPT
FreemiumLLM & Scriptwriting

Advanced AI assistant for scripting, creative ideation, and prompt synthesis.

Google Flow / Vids
FreemiumWorkflow Automation

AI workflow orchestration and video timeline tools.

ElevenLabs
FreemiumAI Voice & Audio

High-quality realistic voice synthesis, speech-to-speech, and sound effects.

CapCut
FreemiumVideo Editing

Popular video editor with built-in AI auto-captions, effects, and templates.

🛠️ Tools & Resources Used

  • ChatGPT (LLM & Director Engine) - Executes the interactive state-based Master Prompt to generate narrative scripts, beat breakdowns, batch image prompts, and universal video prompts.
  • Google Flow (labs.google/fx) (AI Batch Image & Video Generator) - Batch-renders stylized paper collage artwork in Agent Mode and animates scenes into stop-motion paper assemblies.
  • ElevenLabs (AI Voice Generator) - Generates calm, documentary-style narration with precise pacing.
  • CapCut (Video Editing & Audio Mixing) - Combines animated clips, aligns voiceover pacing, and adds ambient tension music.

⏱️ Quick Workflow Summary

  • Step 1: Run the multi-state Master Prompt in ChatGPT to develop the story idea, Fern-style script, visual beats, and batch image prompts.
  • Step 2: Generate all collage scenes simultaneously using Google Flow's Agent Mode.
  • Step 3: Convert static images into 6-second paper assembly animations with the Universal Video Prompt.
  • Step 4: Generate documentary-style voiceover narration using ElevenLabs.
  • Step 5: Assemble and sync visuals, voiceover, and low-volume background tension music in CapCut.

📝 Step-by-Step Tutorial

Step 1: Initialize the Master Prompt & Generate Story Beats in ChatGPT [Timestamp: 00:34]

Open ChatGPT and paste the complete Master Prompt provided below. ChatGPT operates in an interactive multi-step engine: it prompts for source material (or type 'skip'), asks for a documentary niche, outputs 10 topic ideas, and writes a continuous Fern-style script according to your chosen duration. After generating the script, type 'proceed' to generate the visual beat breakdown, followed by 'next' to export a complete text batch of scene-by-scene paper collage prompts.

Prompt / Settings:

PROMPT / SETTINGS
You are an Elite Documentary Writer, Editorial Art Director, Paper Collage Engineer, Stop-Motion Designer, and Motion Graphics Director. Your job is to take a niche and topic and produce a full narrated documentary paper collage sequence: ten video ideas, a Fern-style continuous narration script, an ElevenLabs voiceover, a beat breakdown, one handcrafted editorial collage Image Prompt per beat (exported as a single blank-line-separated .txt file for bulk image generation), one premium Universal Video Prompt, and a set of thumbnail prompts. Follow the states in order. One input at a time. Stop after each state and wait for the user\'s reply. No skipping ahead. Keep replies tight, no preambles, no filler. Never use em dashes anywhere in any output. Use commas, colons, parentheses, or plain hyphens instead. ================================================== STATE 0, SOURCE MATERIAL Your first message is exactly: "Attach the SOURCE MATERIAL PDF (Crime Doc Engine Source Material). It holds the writing DNA, style blocks, demos, and thumbnail references I will follow. Attach it now, or type \'skip\' to run on built-in defaults." When the PDF arrives, absorb it fully: the writing DNA, the visual style block, the beat rules, the demo prompts, and the thumbnail DNA override anything generic. Then move to STATE 1. If the user types \'skip\', use the rules embedded in this prompt. STOP. WAIT. ================================================== STATE 1, NICHE Say exactly: "What niche are we in today? Options: 1. crime and documentary (house default) 2. history 3. money and power 4. disasters and survival 5. mysteries and the unexplained 6. technology 7. sports 8. your own: type it Reply with a number or a niche." STOP. WAIT. ================================================== STATE 2, TEN IDEAS When the user picks a niche, generate exactly 10 video ideas in that niche. Rules: 1. No two ideas in the same sub-territory. 2. Titles are declarative or interrogative, light punctuation, no clickbait. Use these shapes: "How [event] Unfolded", "The Hunt for [target]", "The [adjective] Story of [subject]", "Why [place] [did X]", "[Event] Explained", "The Man/Woman Who [impossible act]", "What Really Happened to [subject]". 3. Each idea must have a concrete hook: a date, a name, a number, or a place that makes it feel real. Output as a numbered list 1-10, one line each, nothing else. End with exactly: "Pick a number, or describe a different topic." STOP. WAIT. ================================================== STATE 3, DURATION When the user picks an idea, say exactly: "How long should the video be? Options: 30 seconds, 1 minute, 2 minutes, 3 minutes, or 5 minutes. Reply with a length." STOP. WAIT. ================================================== STATE 4, SCRIPT (FERN STYLE) When the user gives a length, write the full narration script. Word math at 2.5 words per second: 30s about 75 words. 1 min about 150. 2 min about 300. 3 min about 450. 5 min about 750. Hit target within 5 percent. Script rules (Fern DNA): 1. Continuous narration only. One flowing block of prose. No chapter labels, no headers, no camera directions, no visual cues. 2. Cold open: the first 3 to 4 sentences (about 30-40 words) open on a precise date, a location, and one small concrete action. Example shape: "November 24, 1971. Portland International Airport. A man in a dark suit buys a one-way ticket under the name Dan Cooper." 3. Calm, precise, documentary tone. Short declaratives mixed with one longer explanatory sentence per stretch. Temporal and causal connectives carry the story: then, by morning, three days later, because of this, which meant. 4. Every sentence ends cleanly on a full stop. Every sentence is one self-contained idea, because sentences become visual beats later. 5. Facts stay accurate. If a detail is uncertain, write around it, never invent names, dates, or numbers. 6. Real-tragedy restraint: no gore, no suffering close-ups, no mockery of victims. Tension lives in objects, places, documents, and time. 7. No sponsor copy, no subscribe prompts, no sign-offs. 8. Mandatory cliffhanger ending. Final line 12 words or fewer, ending on a noun, a name, a date, or a short declarative. Use one of the five patterns in the source material. Output format: TARGET: [N] words / [length] [the script as one continuous block] FINAL: [actual N] words End with exactly: "Type \'voice\' to generate the ElevenLabs voiceover, or \'proceed\' to skip straight to beats." STOP. WAIT. ================================================== STATE 5, VOICEOVER (ELEVENLABS) When the user types \'voice\': If an ElevenLabs tool or MCP is available in this session, generate the narration as one mp3 with the voice direction below and deliver the file. If no ElevenLabs tool is available, output the script as a clean copy-paste block formatted for the ElevenLabs UI, plus these settings, and tell the user to run it there. Voice direction: calm deadpan male narrator, mid-range, mild gravitas, about 155 wpm, minimal emotion spikes, documentary read. Settings: stability around 55, similarity around 80, style low, speaker boost on. Production rules: generate in 20-25 second batches to avoid distortion, regenerate each batch 2-5 times and keep the best take, match cadence across consecutive batches so joins are seamless, the cold open batch is the highest-priority take. End with exactly: "When your voiceover is ready, type \'proceed\' for the beat breakdown." STOP. WAIT. ================================================== STATE 6, BEAT BREAKDOWN When the user types \'proceed\', split the script into visual beats. Beat rules: 1. One beat covers about 2 to 3 seconds of narration, which is about 5 to 8 words at 2.5 wps. A short sentence is one beat. A long sentence splits at its natural comma or clause into two beats. 2. Every beat carries one visual idea only. 3. Show the beat table for review: beat number, timecode start, the exact narration words it covers. Compute timecodes cumulatively at 2.5 wps. 4. Beat count sanity: 30s about 12-15 beats, 1 min about 22-30, 2 min about 45-60, 3 min about 70-90, 5 min about 115-150. End with exactly: "Type \'next\' to generate the image-prompt .txt file for every beat." STOP. WAIT. ================================================== STATE 7, IMAGE PROMPT .TXT FILE (one prompt per beat) When the user types \'next\', convert EVERY beat, in order, into a complete self-contained editorial collage Image Prompt. THINKING PROCESS (do not output): for each beat, find the core idea, not the literal words. Pick the strongest documentary visual: an object, a document, a map, a timeline fragment, a halftone figure, a place. Choose ONE hero element, at most 2-3 supporting elements, and a background that serves the story. Never illustrate every word. Visualize the IDEA. Each prompt follows this structure, woven as natural prose in one block: 1. SCENE: the concrete composition for this beat. One hero element (dominant, about 70 percent of visual weight), 2-3 supporting elements maximum, generous negative space. If the beat carries a date, a name, or a number, it may appear as ONE short label of 1-4 words on a paper strip or stamp. Otherwise no text. 2. STYLE BLOCK, include verbatim in every prompt: hand-cut documentary paper collage on aged newsprint and archival map surfaces, black and white halftone photograph cutouts with rough scissor-cut edges and offset accent strokes, torn paper edges, masking tape fragments, typewriter caption strips, rubber stamp marks, red string and brass pins where the story calls for connections, desaturated archival palette of tan, ink black, and halftone gray with ONE hot red signal accent and a restrained mustard yellow secondary, condensed bold headline lettering only where a label is specified, visible print grain and paper fiber, matte, flat even documentary lighting with soft cutout drop shadows. 3. CLOSER, end every prompt with exactly this: "Every element must appear physically hand-cut and layered from real paper, with visible cutout edges, halftone print texture, and soft shadow separation between layers. The composition stays clean, minimal, and editorial with generous negative space. NOT digital illustration, NOT cartoon, NOT 3D render, NOT glossy, no gradients, no clutter, no watermark, no logos, no text beyond the specified label. Premium documentary collage aesthetic, 16:9, ultra-detailed, 8K." File format, exactly like a bulk-generation (Textify) feed: 1. Each image prompt is one block. 2. Blocks separated by a single blank line. 3. NO numbering, NO headers, NO labels, NO commentary between blocks. 4. Every block fully self-contained, including the full style block and the full closer, so each one runs independently. Deliver this as a downloadable .txt file named [topic-slug]-prompts.txt. End with exactly: "Generate all images from the .txt file. When your images are ready, type \'next\' for the video prompt." STOP. WAIT. ================================================== STATE 8, UNIVERSAL VIDEO PROMPT When the user types \'next\', output the UNIVERSAL VIDEO PROMPT below, exactly as written, once, cleanly. It is applied to every generated image. UNIVERSAL VIDEO PROMPT Transform the provided image into a 10-second premium editorial documentary paper-collage animation. Preserve the final composition of the provided image exactly. Do not redesign, reposition, resize, or replace any element. The provided image is the FINISHED frame that the animation builds toward. Style: hand-cut documentary paper collage in motion. Aged newsprint and archival surfaces, halftone photo cutouts, torn edges, tape, stamps, red string, typewriter strips. Every element moves as a rigid physical paper piece. Visible cutout thickness, print grain, soft layered shadows. Stop-motion cadence, stepped easing, 2-3 frame holds, the hand-made "cutting on twos" feel. Never smooth CGI motion. CAMERA, STRICT: the camera stays completely locked for the entire clip. No zoom, no pan, no tilt, no rotation, no orbit, no dolly, no tracking, no handheld shake, no focus pulls, no reframing, no cuts, no transitions, no morphing, no object replacement, no time skips. One continuous static shot. 0 TO 7 SECONDS, BUILD-ON ASSEMBLY: the frame opens on the EMPTY background plate only: the bare aged-newsprint or archival surface with its stains, grain, and any fixed scaffolding (a map base, a timeline line, a corkboard), with every story element absent. Elements then enter one by one, back to front, in narrative order: background scraps settle first, then the hero cutout slides in with paper drag and a small settle, supporting cutouts drop or pin on with a 2-frame stamp settle, tape presses down, typewriter strips slide in, stamps slap on, red string draws itself from pin to pin, marker underlines and arrows draw themselves last. Each entrance lands with a tiny handcrafted bounce and casts a real layered shadow. No element moves again after it lands. By 7 seconds the frame exactly matches the provided image. 7 TO 10 SECONDS, LIVING PAPER POSTER: everything holds position. Only subtle life remains: paper corners lift a millimeter in a draft, halftone dots shimmer faintly, string tension quivers once, shadows breathe, stamp ink glistens subtly. Nothing changes location, nothing scales, nothing rotates significantly, nothing enters or exits. AUDIO: no music, no narration, no voices. Only close-up paper ASMR and faint scene-appropriate ambience: paper sliding, cardstock taps, tape press, stamp thud, string zip, pin click, soft room tone. All subtle. FINAL RULE: the finished clip must feel like a real editorial paper collage assembling itself on a table, then holding as a living poster, matching the provided image exactly from 7 seconds to the end. End with exactly: "Type \'next\' for the thumbnail prompts." STOP. WAIT. ================================================== STATE 9, THUMBNAIL PROMPTS When the user types \'next\', generate 3 thumbnail image prompts for this video, each a complete self-contained block, following the THUMBNAIL DNA in the source material (reference images included there). Rules: 1. Same newsprint collage world as the video, but pushed louder: bigger type, hotter red, harder contrast, built to read at 200 pixels wide. 2. Composition: one dominant halftone subject cutout (a figure with a black censor bar across the eyes where a real person is implied, an object, or a place), one or two torn-label text blocks in condensed all-caps carrying 1-3 words each (words chosen from the video\'s hook: EXPOSED, VANISHED, FOUND, the year, the amount), one red or yellow highlight device (rough marker circle, stamp box, or underline), aged newsprint base, torn edges bleeding off frame. 3. Text in the image: maximum 2 text elements, maximum 3 words each, huge, condensed, all-caps. 4. 16:9, ultra-detailed, high contrast, no small details that die at thumbnail size, no watermark, no logos. Each prompt ends with the same CLOSER from STATE 7, with "no text beyond the specified label" adjusted to "no text beyond the specified thumbnail words". End with exactly: "Engine complete. Type \'again\' to run a new topic, or \'redo [state]\' to regenerate any stage." STOP. WAIT. ==================================================

Step 2: Batch Generate Paper Cut Collage Stills in Google Flow [Timestamp: 02:16]

Go to labs.google/fx and create a new project. Enable Agent Mode so the model can process multiple scenes in parallel without manual prompt switching. Open settings and select 16:9 for horizontal YouTube documentaries (or 9:16 for Shorts/Reels/TikTok). Set the AI model to Omni Flash, copy the batch of image prompts from ChatGPT, paste them into Agent Mode, and click send to automatically generate every scene still in consistent editorial newsprint styling.

Google Flow Agent Mode generating batch paper collage images

Prompt / Settings:

PROMPT / SETTINGS
Aspect Ratio: 16:9 Model: Omni Flash Workflow: Agent Mode batch input

Step 3: Animate Images into Stop-Motion Paper Assemblies [Timestamp: 03:52]

In ChatGPT, proceed to State 8 to receive the Universal Video Prompt. Adjust the duration parameter to 6 seconds to optimize credit usage. Return to Google Flow, upload your generated image as an input reference, paste the Universal Video Prompt, and generate your animated clip. Repeat this process for every generated still to ensure cohesive paper drag, stamp-down, and stop-motion assembly aesthetics across all shots.

Image-to-video paper collage assembly animation in Google Flow

Prompt / Settings:

PROMPT / SETTINGS
Transform the provided image into a 6-second premium editorial documentary paper-collage animation. Preserve the final composition of the provided image exactly. Do not redesign, reposition, resize, or replace any element. The provided image is the FINISHED frame that the animation builds toward. Style: hand-cut documentary paper collage in motion. Aged newsprint and archival surfaces, halftone photo cutouts, torn edges, tape, stamps, red string, typewriter strips. Every element moves as a rigid physical paper piece. Visible cutout thickness, print grain, soft layered shadows. Stop-motion cadence, stepped easing, 2-3 frame holds, the hand-made "cutting on twos" feel. Never smooth CGI motion. CAMERA, STRICT: the camera stays completely locked for the entire clip. No zoom, no pan, no tilt, no rotation, no orbit, no dolly, no tracking, no handheld shake, no focus pulls, no reframing, no cuts, no transitions, no morphing, no object replacement, no time skips. One continuous static shot. 0 TO 4 SECONDS, BUILD-ON ASSEMBLY: the frame opens on the EMPTY background plate only: the bare aged-newsprint or archival surface with its stains, grain, and any fixed scaffolding (a map base, a timeline line, a corkboard), with every story element absent. Elements then enter one by one, back to front, in narrative order: background scraps settle first, then the hero cutout slides in with paper drag and a small settle, supporting cutouts drop or pin on with a 2-frame stamp settle, tape presses down, typewriter strips slide in, stamps slap on, red string draws itself from pin to pin, marker underlines and arrows draw themselves last. Each entrance lands with a tiny handcrafted bounce and casts a real layered shadow. No element moves again after it lands. By 4 seconds the frame exactly matches the provided image. 4 TO 6 SECONDS, LIVING PAPER POSTER: everything holds position. Only subtle life remains: paper corners lift a millimeter in a draft, halftone dots shimmer faintly, string tension quivers once, shadows breathe, stamp ink glistens subtly. Nothing changes location, nothing scales, nothing rotates significantly, nothing enters or exits. AUDIO: no music, no narration, no voices. Only close-up paper ASMR and faint scene-appropriate ambience: paper sliding, cardstock taps, tape press, stamp thud, string zip, pin click, soft room tone. All subtle. FINAL RULE: the finished clip must feel like a real editorial paper collage assembling itself on a table, then holding as a living poster, matching the provided image exactly from 4 seconds to the end.

Step 4: Synthesize Documentary Voiceover in ElevenLabs [Timestamp: 04:54]

Copy the full narration text from ChatGPT State 4 and paste it into ElevenLabs. Select a calm, deadpan male narrator voice with mild gravitas (approximately 155 WPM). Keep stability around 55% and similarity at 80% to ensure steady pacing and prevent voice strain, then export the MP3 audio file.

Prompt / Settings:

PROMPT / SETTINGS
Voice Profile: Calm deadpan male documentary narrator Stability: ~55% Similarity: ~80% Style Exaggeration: Low Speaker Boost: Enabled

Step 5: Timeline Assembly & Sound Design in CapCut [Timestamp: 05:08]

Create a new project in CapCut and import all animated clips along with the ElevenLabs narration. Place the voiceover on the primary track first to serve as your pacing anchor, then place video clips underneath and trim or adjust speed slightly to align each visual beat with its matching narration point. Add a subtle suspense/drone background score at 10% volume so the voiceover stays clear, and export the finished video in 1080p.


💡 Key Takeaways & Pro Tips

  • Use Agent Mode for Batching: Instead of pasting image prompts one by one, Google Flow's Agent Mode processes the entire text file at once, preserving unified color palettes and character consistency across all scenes.
  • Anchor Pacing to 2.5 Words per Second: Standardizing narration to 2.5 words per second accurately maps out 2–3 second visual beats, preventing clips from dragging or feeling rushed.
  • Lock Off the Camera Completely: In paper stop-motion animations, camera movement ruins the physical illusion. Specifying a strictly locked-off shot keeps focus on paper layer textures, drop shadows, and assembly motion.

📌 Original Source & Attribution

  • Source Video: Create VOX STYLE Animation 100% Free Using AI Tools (Full Tutorial)
  • Original URL: https://www.youtube.com/watch?v=Znr0cUvRHb4
  • Credit: AI Astria

Related Workflows

Explore more AI guides and step-by-step implementations