Guides · VideoVideo prompt guide

How to write AI video prompts, one shot at a time

A text-to-video model gives you one short clip per request. The prompts that work describe that single shot like a camera brief: who is in it, what moves, where, how the camera behaves, how it is lit, and what it should sound like.

Capabilities reviewed 2026-09-26

Method

Step by step

  1. Decide what the one clip has to show

    Vizify returns one short video per generation, not an edited sequence. Write down the single moment you need, such as a product turning on a plinth or a cyclist crossing a bridge at dusk, and leave the rest of the story for separate clips.

  2. Name the subject and give it one clear action

    Describe the subject concretely (material, color, age, clothing, brand-neutral details) and give it one motion beat: pours, turns, walks toward camera, lifts off. Several competing actions in a few seconds is the most common reason a clip looks confused.

  3. Set the scene and the light

    Add the setting, time of day, weather, and light direction. "Soft window light from the left, late morning" gives the model far more to work with than "nice lighting", and it keeps color and shadows consistent across the clip.

  4. Direct the camera

    State the framing and the movement separately: a medium close-up with a slow push-in, a locked-off wide shot, a handheld follow at walking pace, a slow orbit. If the camera should not move, say it is locked off.

  5. Choose a style, then audio

    Name the visual treatment (natural documentary color, 35mm film grain, clean studio product look) and then say what the audio should contain, or ask for a silent clip. Audio support depends on the model, so check it before you rely on sound.

  6. Match the brief to the model's limits

    Duration, resolution, aspect ratio, reference images, and audio are set per model in Vizify. Kling 3, the current default, runs 3 to 15 seconds with audio off unless you turn it on; Veo 3.1 is a fixed 8 seconds with audio always on; Seedance 2 is always silent.

  7. Review the whole clip and change one thing

    Watch the full clip, not one frame. Faces, hands, logos, and small text drift first. When you revise, change one element of the brief at a time so you can tell which change fixed the problem.

Copy and adapt

Prompt examples

Silent product reveal

Slow 180-degree orbit around a matte white ceramic pour-over coffee set on a dark walnut counter. Soft morning window light from the left, gentle steam rising from the cup, shallow depth of field. Clean studio product look, no text on screen. Silent clip, 16:9.

Why it works: One subject, one camera move, one light source, and an explicit request for no audio or on-screen text, so nothing competes with the product.

Scene with native audio

Night market street food stall: a vendor flips skewers over a charcoal grill as sparks lift into the air. Handheld medium shot at chest height, slight natural sway. Warm tungsten stall lights against a blue dusk sky. Audio: sizzling grill, low crowd murmur, no music, no dialogue.

Why it works: The audio line names the sounds you want and the ones you do not. Use it with a model that generates audio, such as Veo 3.1 or Kling 3 with audio turned on.

Vertical social clip

Vertical 9:16. A trail runner in a red jacket crests a rocky ridge toward the camera at sunrise, breath visible in cold air. Low-angle tracking shot that slows as she reaches the top. Crisp natural color, light lens flare, 8 seconds.

Why it works: The aspect ratio, the direction of movement relative to the camera, and the duration are all stated, which keeps the subject framed for a phone screen.

Animate an uploaded product photo

Use the attached photo as the reference. Keep the bottle shape, label artwork, and color exactly as shown. The bottle rotates slowly on a turntable while condensation beads on the glass; locked-off camera, dark gradient background, soft rim light.

Why it works: It tells the model what must not change before describing the motion. Upload a PNG, JPEG, or WebP image in Vizify; a video Vizify generated earlier has to be downloaded and uploaded again to be used as a reference.

Atmosphere without people

Wide coastal grassland under a moving storm front. Locked-off camera on a tripod, long grass rippling in strong wind, cloud shadows sweeping across the hills. Muted desaturated color, overcast diffuse light. 10 seconds.

Why it works: With no character to keep consistent, the prompt spends its detail on motion in the environment and on a static camera, which gives stable, usable footage.

Before you start

Current limitations

  • One generation returns one short clip. Vizify does not stitch clips into a sequence, edit a timeline, or add a voice-over track to a finished video.
  • Duration, resolution, aspect ratio, reference images, and audio are model-specific. Hailuo 2.3 requires exactly one reference image, and some models are always silent.
  • Video generation needs Vizify Pro or a purchased credit pack; there is no video allowance on an account without either.
  • A clip or image Vizify generated cannot be passed straight into a new video request. Download it and upload it again as a reference.
  • Faces, hands, logos, and small on-screen text can drift during motion. Review the full clip before publishing it.

Why “one shot” is the unit that matters

Text-to-video models are trained to render a continuous shot, and Vizify’s video capability returns exactly one clip per request. A prompt that describes a sequence (“she opens the door, walks to the window, then turns and smiles”) asks the model to compress three shots into a few seconds, and it will usually blur them together. Write the shot you need now, generate it, and write the next shot as its own prompt.

A reliable order for the brief is subject, action, setting, camera, light, style, audio. You do not need headings or a template, but covering each of those in plain sentences removes most of the guessing the model would otherwise do.

Model limits change what a good prompt looks like

Because every Vizify video model exposes its own parameter set, the same brief behaves differently depending on where it runs:

  • Kling 3 is the current default. It runs 3 to 15 seconds, offers 720p, 1080p, or 4K, accepts up to two reference images, and keeps audio off unless you turn it on.
  • Veo 3.1 is fixed at 8 seconds and 720p and always returns audio, so write the audio line with care and do not ask it for a silent clip.
  • Seedance 2 is always silent but goes up to 4K and 15 seconds, which suits product and B-roll shots you will score later.
  • Seedance 2.5 runs 4 to 30 seconds with optional audio and accepts many more reference images, which helps when a scene has several fixed elements.
  • Grok Imagine Video 1.5 always includes audio and takes up to seven reference images.
  • Hailuo 2.3 needs exactly one reference image, and it supports 1080p only on 6-second clips.

If a detail in your prompt conflicts with the model’s limits, such as a 20-second request to a model capped at 15 or “silent” on a model that always returns audio, the model’s contract wins. Check the model page before you write the timing and audio lines.

Reviewing and revising

Watch the clip at full size from start to finish. The first things to drift are faces, hands, product geometry, logos, and any small text. When a revision is needed, keep everything else in the prompt the same and change one line, whether that is the camera move, the light, or the action, so you learn what actually made the difference. Each revision is a new generation and uses credits the same way the first one did.

FAQ

Questions people ask

How long should an AI video prompt be?

Long enough to cover subject, action, setting, camera, light, style, and audio, which is usually two to five sentences. Stacking extra adjectives beyond that tends to add competing instructions rather than more control.

Which Vizify video models generate audio?

It depends on the model. Veo 3.1 and Grok Imagine Video 1.5 always include audio, Kling 3 and Seedance 2.5 let you turn audio on or off, and Seedance 2, Wan 2.7, Runway, and Hailuo 2.3 are silent. Check the model page before you rely on sound.

How long can a generated clip be?

Each model has its own range in Vizify. Kling 3 runs 3 to 15 seconds, Seedance 2.5 and Wan 3.0 run up to 30 seconds, and Veo 3.1 is fixed at 8 seconds. For anything longer, generate separate clips and assemble them in your own editor.

Should I describe camera movement in the prompt?

Yes. Name the shot size and the movement separately, for example a medium shot with a slow push-in, or a locked-off wide shot. If the camera should stay still, say so; otherwise the model may add movement you did not ask for.

Can I use negative prompts?

Vizify does not expose a separate negative-prompt field for video. Put exclusions in the brief itself, such as no on-screen text, no music, or no people in frame, and keep them short.

Can I turn an image into a video?

Yes, with a model that accepts reference images. Upload the image with your prompt and say what must stay unchanged. The image-to-video tools on the Site prepare that request for you.

Do I need a paid plan to generate video?

Yes. Video generation requires Vizify Pro or a credit pack. You can still prepare and refine the prompt on the Site before you sign in.

Why does my subject change halfway through the clip?

Motion models can drift in identity and fine detail, especially with several actions or a moving camera. Reduce the clip to one action, keep the camera simpler, and add a reference image when identity matters.

Related

Keep going

Put the brief to work

Paste your shot brief into the AI Video Generator.