The Complete Prompt Guide for AI Image and Video Generation
2026/03/31

The Complete Prompt Guide for AI Image and Video Generation

Learn how to write effective prompts for AI image and video generation. Covers prompt structure, best practices, common mistakes, and ready-to-use examples for text-to-image, image-to-image, text-to-video, and image-to-video workflows.

Great AI-generated content starts with a great prompt. The difference between a generic output and something that looks intentional comes down to how you describe what you want.

This guide covers everything you need to write effective prompts — whether you are generating images, animating them into video, or extending clips into longer sequences.


The Prompt Formula

A reliable prompt follows this structure:

Subject + Style + Lighting + Composition + Mood

Each element removes ambiguity and gives the model something concrete to execute on.

ElementWhat it doesExamples
SubjectWhat is in the scene"a woman reading at a cafe", "a red sports car on a mountain road"
StyleVisual treatment"photorealistic", "anime", "oil painting", "watercolor", "3D render"
LightingHow light behaves"golden hour backlight", "overcast diffused", "neon glow", "hard rim light"
CompositionCamera and framing"close-up", "wide shot", "bird's eye view", "low angle", "rule of thirds"
MoodFeeling or reference"cinematic", "peaceful", "Blade Runner mood", "Studio Ghibli feel"

Vague vs. Specific

Compare these two prompts:

Vague: "a city at night"

Specific: "Futuristic Tokyo street at 2am, rain-slicked asphalt, neon reflections, low-angle wide shot, cinematic fog, Blade Runner mood"

The second prompt tells the model exactly what to aim for: location, time, lighting, camera position, and tonal reference. Each detail eliminates ambiguity.


Image Generation

Text-to-Image

When generating from text alone, your prompt is the only input the model has. Be thorough.

Strong prompt example:

A red panda sipping tea under cherry blossoms at sunset, watercolor style, soft diffused golden hour light, medium shot with shallow depth of field, peaceful Japanese garden atmosphere

Tips for text-to-image:

  • Include at least 3 of the 5 formula elements (subject, style, lighting, composition, mood)
  • Name your lighting explicitly — "golden hour backlight" produces very different results from "overcast diffused light"
  • Reference moods rather than single adjectives — "Studio Ghibli feel" is more useful than just "soft"
  • Specify the camera angle — "slow dolly in", "bird's eye view", "low angle" translate directly into how the scene is composed
  • For portraits, mention skin tone handling: "natural skin tones", "editorial beauty lighting"

Image-to-Image

When you already have an image as input, the model can see the composition, colors, and subject. Your prompt should focus on what to change, not re-describe the scene.

Good image-to-image prompts:

Transform to Studio Ghibli anime style, soft pastel colors, dreamy atmosphere

Add dramatic sunset lighting, warm orange tones, lens flare from the right

Convert to pencil sketch, cross-hatching technique, high contrast black and white

Common mistake: Re-describing the entire image in your prompt. The model already sees it. Focus on the transformation you want — style transfer, color adjustment, mood shift.


Video Generation

Text-to-Video

Video prompts need everything an image prompt needs, plus motion and camera direction.

Extended formula for video:

Subject + Action + Style + Camera Movement + Lighting + Sound/Mood

Strong prompt example:

A surfer dropping into a massive wave at golden hour, low-angle tracking shot following the board, water spray catching the light, cinematic slow motion, ocean roar and muffled underwater sounds

Tips for text-to-video:

  • Name your camera movement. "Slow dolly in", "pan right", "static wide", "drone ascending" — these translate directly into how the scene is animated. If you do not specify, you leave one of the most important decisions to chance.
  • Describe what is moving. "Leaves falling", "hair blowing in wind", "clouds moving fast" gives the model clear motion anchors.
  • Reference audio when relevant. "Ambient city sounds", "ocean roar", "lo-fi ambient soundtrack" helps set the mood for models that generate audio.
  • Keep compositions focused. A clear subject against a defined background consistently beats a busy, crowded scene.
  • Go wider for people. Wider shots and slower movements produce the cleanest results when people are in the frame. Pull the camera back and let the motion breathe.

Image-to-Video

This is where most creators spend their time. You already have a starting frame, so the video output is more predictable.

The key rule: keep it short and motion-focused.

The model already has full visual context from your image. Your prompt just needs to describe what should move and how.

Good image-to-video prompts:

Camera slowly pans right, hair blowing gently in the wind, soft ambient sounds

Calm organic movement, camera slowly pulls back, subtle wind effect

Subject walks forward two steps, camera follows at eye level, steady tracking

Common mistake: Writing a full scene description for image-to-video. Compare:

Too much: "A woman in a red dress sitting at a rain-streaked cafe window at night with neon reflections, she stands up and walks toward the door"

Better: "She stands up slowly, grabs her coat, and walks toward the door. The camera follows."


Video Extension

When extending an existing video clip, write a continuation, not a re-description.

The model already knows what the scene looks like. Just tell it where to go from here.

Good extension prompts:

She opens the door and steps outside into the rain. The camera follows through the doorway.

Camera continues orbiting, revealing the back of the building. Sunlight shifts as the angle changes.

The bird takes flight from the branch, camera tracks upward following its path against the sky.

Tips for extending:

  • Describe what happens next, not what already happened
  • Maintain consistency with the existing clip — same lighting, same pace
  • Each extension adds 6 or 10 seconds, and you can chain them together for longer sequences
  • The result should feel like one continuous take, not separate clips stitched together

Ready-to-Use Prompts

Cinematic Scenes

Futuristic Tokyo street at 2am, rain-slicked asphalt, neon reflections, slow dolly forward, cinematic fog, Blade Runner mood, ambient city sounds

Aerial drone slowly descending over an ancient stone temple reclaimed by jungle, golden hour shafts of light cutting through the canopy, birds scattering from treetops, epic scale

Product Shots

Camera doing a slow 360 orbit around a sleek product on a clean surface, sharp shadow rotating with the light, subtle ambient tone, commercial quality

Luxury perfume bottle on a reflective black surface, dramatic side lighting, water droplets, soft bokeh background, high-end product photography

Portraits

Portrait of a young woman with wind-blown hair, golden hour backlight creating a warm halo, shallow depth of field, natural skin tones, editorial fashion photography

Portrait lit by pink and blue neon lights, cyberpunk aesthetic, reflective sunglasses, dark background, moody atmosphere, sharp contrast

Mood Pieces

Rain streaking down a window, warm cafe interior blurred behind the glass, steam rising from a coffee cup, lo-fi ambient soundtrack, melancholic calm

First light filtering through a misty forest, dew drops catching light, camera slowly tracking forward along a path, birds beginning to sing, peaceful morning

World Building

Camera soaring between floating islands connected by rope bridges, waterfalls pouring into clouds below, golden hour, fantasy world, epic orchestral atmosphere

Slow dive through underwater ruins of an ancient city, bioluminescent fish swimming past columns, shafts of sunlight from above, ethereal deep ocean ambience


Common Mistakes to Avoid

  1. Being too vague. "A beautiful landscape" gives the model almost nothing to work with. Add specifics: what landscape, what time of day, what atmosphere.

  2. Overloading the prompt. Cramming every possible detail into one prompt can confuse the model. Focus on 5-7 strong descriptors rather than 20 weak ones.

  3. Ignoring lighting. Lighting is the highest-leverage detail you can specify. "Golden hour backlight" and "overcast diffused light" produce completely different results.

  4. Forgetting camera direction in video. If you do not specify camera movement, you are leaving one of the most important creative decisions to chance.

  5. Re-describing images in image-to-video. When animating an existing image, the model already has the visual context. Focus on motion and camera, not scene description.

  6. Not iterating. The model produces different results each time, even from the same prompt. If the first generation does not land, try it again before rewriting. Sometimes the second or third attempt nails it.


Tips for Consistent Results

  • Save prompts that worked. When something lands, keep the exact prompt text so you can build on it or reproduce similar results later.
  • Run the same prompt more than once. Outputs vary between runs. Try 2-3 generations before rewriting.
  • Use the AI Generate button. If you are not sure how to structure a prompt, use the AI Generate feature in the prompt editor to get a starting point, then refine it.
  • Start with an image for video. Generate an image first, review it, then animate. This gives you more control over the starting frame than going straight to text-to-video.
  • Name specific references. "Blade Runner mood", "Wes Anderson color palette", "Studio Ghibli feel" gives the model a rich visual library to draw from. Single adjectives like "dark" or "soft" are too open-ended.