
The Complete Prompt Guide for AI Image and Video Generation
Learn how to write effective prompts for AI image and video generation. Covers prompt structure, best practices, common mistakes, and ready-to-use examples for text-to-image, image-to-image, text-to-video, and image-to-video workflows.
Great AI-generated content starts with a great prompt. The difference between a generic output and something that looks intentional comes down to how you describe what you want.
This guide covers everything you need to write effective prompts — whether you are generating images, animating them into video, or extending clips into longer sequences.
The Prompt Formula
A reliable prompt follows this structure:
Subject + Style + Lighting + Composition + Mood
Each element removes ambiguity and gives the model something concrete to execute on.
| Element | What it does | Examples |
|---|---|---|
| Subject | What is in the scene | "a woman reading at a cafe", "a red sports car on a mountain road" |
| Style | Visual treatment | "photorealistic", "anime", "oil painting", "watercolor", "3D render" |
| Lighting | How light behaves | "golden hour backlight", "overcast diffused", "neon glow", "hard rim light" |
| Composition | Camera and framing | "close-up", "wide shot", "bird's eye view", "low angle", "rule of thirds" |
| Mood | Feeling or reference | "cinematic", "peaceful", "Blade Runner mood", "Studio Ghibli feel" |
Vague vs. Specific
Compare these two prompts:
Vague: "a city at night"
Specific: "Futuristic Tokyo street at 2am, rain-slicked asphalt, neon reflections, low-angle wide shot, cinematic fog, Blade Runner mood"
The second prompt tells the model exactly what to aim for: location, time, lighting, camera position, and tonal reference. Each detail eliminates ambiguity.
Image Generation
Text-to-Image
When generating from text alone, your prompt is the only input the model has. Be thorough.
Strong prompt example:
A red panda sipping tea under cherry blossoms at sunset, watercolor style, soft diffused golden hour light, medium shot with shallow depth of field, peaceful Japanese garden atmosphere
Tips for text-to-image:
- Include at least 3 of the 5 formula elements (subject, style, lighting, composition, mood)
- Name your lighting explicitly — "golden hour backlight" produces very different results from "overcast diffused light"
- Reference moods rather than single adjectives — "Studio Ghibli feel" is more useful than just "soft"
- Specify the camera angle — "slow dolly in", "bird's eye view", "low angle" translate directly into how the scene is composed
- For portraits, mention skin tone handling: "natural skin tones", "editorial beauty lighting"
Image-to-Image
When you already have an image as input, the model can see the composition, colors, and subject. Your prompt should focus on what to change, not re-describe the scene.
Good image-to-image prompts:
Transform to Studio Ghibli anime style, soft pastel colors, dreamy atmosphere
Add dramatic sunset lighting, warm orange tones, lens flare from the right
Convert to pencil sketch, cross-hatching technique, high contrast black and white
Common mistake: Re-describing the entire image in your prompt. The model already sees it. Focus on the transformation you want — style transfer, color adjustment, mood shift.
Video Generation
Text-to-Video
Video prompts need everything an image prompt needs, plus motion and camera direction.
Extended formula for video:
Subject + Action + Style + Camera Movement + Lighting + Sound/Mood
Strong prompt example:
A surfer dropping into a massive wave at golden hour, low-angle tracking shot following the board, water spray catching the light, cinematic slow motion, ocean roar and muffled underwater sounds
Tips for text-to-video:
- Name your camera movement. "Slow dolly in", "pan right", "static wide", "drone ascending" — these translate directly into how the scene is animated. If you do not specify, you leave one of the most important decisions to chance.
- Describe what is moving. "Leaves falling", "hair blowing in wind", "clouds moving fast" gives the model clear motion anchors.
- Reference audio when relevant. "Ambient city sounds", "ocean roar", "lo-fi ambient soundtrack" helps set the mood for models that generate audio.
- Keep compositions focused. A clear subject against a defined background consistently beats a busy, crowded scene.
- Go wider for people. Wider shots and slower movements produce the cleanest results when people are in the frame. Pull the camera back and let the motion breathe.
Image-to-Video
This is where most creators spend their time. You already have a starting frame, so the video output is more predictable.
The key rule: keep it short and motion-focused.
The model already has full visual context from your image. Your prompt just needs to describe what should move and how.
Good image-to-video prompts:
Camera slowly pans right, hair blowing gently in the wind, soft ambient sounds
Calm organic movement, camera slowly pulls back, subtle wind effect
Subject walks forward two steps, camera follows at eye level, steady tracking
Common mistake: Writing a full scene description for image-to-video. Compare:
Too much: "A woman in a red dress sitting at a rain-streaked cafe window at night with neon reflections, she stands up and walks toward the door"
Better: "She stands up slowly, grabs her coat, and walks toward the door. The camera follows."
Video Extension
When extending an existing video clip, write a continuation, not a re-description.
The model already knows what the scene looks like. Just tell it where to go from here.
Good extension prompts:
She opens the door and steps outside into the rain. The camera follows through the doorway.
Camera continues orbiting, revealing the back of the building. Sunlight shifts as the angle changes.
The bird takes flight from the branch, camera tracks upward following its path against the sky.
Tips for extending:
- Describe what happens next, not what already happened
- Maintain consistency with the existing clip — same lighting, same pace
- Each extension adds 6 or 10 seconds, and you can chain them together for longer sequences
- The result should feel like one continuous take, not separate clips stitched together
Ready-to-Use Prompts
Cinematic Scenes
Futuristic Tokyo street at 2am, rain-slicked asphalt, neon reflections, slow dolly forward, cinematic fog, Blade Runner mood, ambient city sounds
Aerial drone slowly descending over an ancient stone temple reclaimed by jungle, golden hour shafts of light cutting through the canopy, birds scattering from treetops, epic scale
Product Shots
Camera doing a slow 360 orbit around a sleek product on a clean surface, sharp shadow rotating with the light, subtle ambient tone, commercial quality
Luxury perfume bottle on a reflective black surface, dramatic side lighting, water droplets, soft bokeh background, high-end product photography
Portraits
Portrait of a young woman with wind-blown hair, golden hour backlight creating a warm halo, shallow depth of field, natural skin tones, editorial fashion photography
Portrait lit by pink and blue neon lights, cyberpunk aesthetic, reflective sunglasses, dark background, moody atmosphere, sharp contrast
Mood Pieces
Rain streaking down a window, warm cafe interior blurred behind the glass, steam rising from a coffee cup, lo-fi ambient soundtrack, melancholic calm
First light filtering through a misty forest, dew drops catching light, camera slowly tracking forward along a path, birds beginning to sing, peaceful morning
World Building
Camera soaring between floating islands connected by rope bridges, waterfalls pouring into clouds below, golden hour, fantasy world, epic orchestral atmosphere
Slow dive through underwater ruins of an ancient city, bioluminescent fish swimming past columns, shafts of sunlight from above, ethereal deep ocean ambience
Common Mistakes to Avoid
-
Being too vague. "A beautiful landscape" gives the model almost nothing to work with. Add specifics: what landscape, what time of day, what atmosphere.
-
Overloading the prompt. Cramming every possible detail into one prompt can confuse the model. Focus on 5-7 strong descriptors rather than 20 weak ones.
-
Ignoring lighting. Lighting is the highest-leverage detail you can specify. "Golden hour backlight" and "overcast diffused light" produce completely different results.
-
Forgetting camera direction in video. If you do not specify camera movement, you are leaving one of the most important creative decisions to chance.
-
Re-describing images in image-to-video. When animating an existing image, the model already has the visual context. Focus on motion and camera, not scene description.
-
Not iterating. The model produces different results each time, even from the same prompt. If the first generation does not land, try it again before rewriting. Sometimes the second or third attempt nails it.
Tips for Consistent Results
- Save prompts that worked. When something lands, keep the exact prompt text so you can build on it or reproduce similar results later.
- Run the same prompt more than once. Outputs vary between runs. Try 2-3 generations before rewriting.
- Use the AI Generate button. If you are not sure how to structure a prompt, use the AI Generate feature in the prompt editor to get a starting point, then refine it.
- Start with an image for video. Generate an image first, review it, then animate. This gives you more control over the starting frame than going straight to text-to-video.
- Name specific references. "Blade Runner mood", "Wes Anderson color palette", "Studio Ghibli feel" gives the model a rich visual library to draw from. Single adjectives like "dark" or "soft" are too open-ended.