Now on VeniceVideoPrivate

Grok Imagine 1.5 Lite

xAI's lightest Grok Imagine 1.5 tier — 1080p clips up to 15 seconds with native audio, priced from $0.04 per clip on Venice.

For agents
curl https://api.venice.ai/api/v1/video/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-1-5-lite-image-to-video",
    "prompt": "Aerial drone shot over a misty mountain valley at golden hour"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.
Model IDgrok-imagine-1-5-lite-image-to-video
Maker
xAI
Modality
Video + audio
Max duration
15 seconds
Max resolution
1080p

Overview

What is Grok Imagine 1.5 Lite

Grok Imagine 1.5 Lite is xAI's budget tier of the Grok Imagine Video 1.5 family, live on Venice since September 30, 2026. It generates clips from a text prompt or an uploaded still image, at 480p/720p/1080p, durations of 1–15 seconds, with sound effects, ambience and dialogue generated natively in the same pass.

Running it privately on Venice

On Venice, Grok Imagine 1.5 Lite runs under the private tier with zero retention — your prompts and source images are not stored, profiled, or fed into a training pipeline, and there is no personal generation history tied to your identity. You pay per clip instead of carrying a SuperGrok subscription, and both family variants (text-to-video and image-to-video) are available permissionlessly through the app or the Venice API.

PrivateNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Specifications

Datasheet

Maker
xAI
Modality
Text-to-video and image-to-video (two variants: grok-imagine-1-5-lite-text-to-video, grok-imagine-1-5-lite-image-to-video)
Open weights
No — proprietary
License
Proprietary
Clip lengths
1s – 15s
Resolutions
480p, 720p, 1080p
Mode
image-to-video
Audio
Yes
Prompt limit
4,096 chars
Input images
One still image for the image-to-video variant; no starting image needed for text-to-video
References
Not supported
Released
Listed on Venice September 30, 2026 (Grok Imagine Video 1.5 family GA June 16, 2026)
Privacy on Venice
Private — zero retention
Available on Venice since
Sep 2026

Assessment

Strengths and limitations

Strengths
  • Native synchronized audio: sound effects, ambience and dialogue are generated in the same pass as the video and land on the action, so there is no separate sound-design step.
  • Full 480p/720p/1080p resolution ladder with per-second duration control from 1s to 15s, which is unusual granularity at the Lite price point.
  • Cheapest entry into the Grok Imagine 1.5 family on Venice — from $0.04 per clip, so drafts and social-length iterations stay inexpensive.
  • The 1.5 generation improved motion coherence and physics over its predecessor — fewer warps and more believable weight and momentum across a clip, per xAI's own release notes.
  • Two variants cover both workflows: animate an existing still (image-to-video) or describe a shot from scratch (text-to-video).
Limitations
  • Closed and proprietary: no open weights, so it cannot be self-hosted, fine-tuned, or audited.
  • As the Lite tier of the 1.5 family, it is positioned below flagship Grok Imagine 1.5; for maximum fidelity on demanding cinematic shots the flagship remains the safer choice.
  • Multi-reference and voice-reference features announced for grok-imagine-video-1.5 (up to seven image references plus voice consistency) are documented for the flagship, not confirmed for the Lite tier.
  • Third-party reviews of the 1.5 family note complex multi-prompt scenes can still trip the model up, and long-form cinematic storytelling is not its strength.

Use cases

What it is good for

  1. 01Social ads and short-form content where a clip needs sound the moment it renders.
  2. 02Bringing product shots, portraits, or illustrations to life from a single still image.
  3. 03Rapid draft loops — cheap 480p iterations, then a final 1080p render of the winning prompt.
  4. 04Prompt-only concept clips when no source image exists yet.
  5. 05Privacy-sensitive creative work where prompts and source images must not be retained.

Prompting

Getting better results

Pick the duration deliberately: 1–6s works for loops and product spins, 8–15s for scenes with dialogue or a camera move.

Draft at 480p ($0.04/clip), lock the prompt, then re-render the identical prompt at 1080p for the final.

Write the audio into the prompt — name the ambience, sound effects, and any dialogue explicitly, since audio is generated in the same pass.

For image-to-video, describe the motion you want ('slow push-in, embers drifting') rather than re-describing the image contents.

Use camera language — push-in, pan, handheld, crane — to control movement; the 1.5 family responds well to explicit cinematography terms.

Stay under the 4,096-character prompt limit; front-load the subject and action in the first sentence so they survive trimming.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Cinematic landscape

Slow aerial drone shot gliding over a misty mountain valley at golden hour, sunlight piercing clouds onto a winding river, ancient pine forests on either side, ultra-smooth motion, professional color grading, atmospheric haze, 4K

Seamless loop

Calm ocean waves rolling onto a black-sand beach at sunrise, golden light on wet sand, foam dissolving into the shore, a single silhouetted figure at the waterline, smooth continuous forward push, serene cinematic atmosphere

Urban cinematic

A slow tracking shot through a rain-soaked Tokyo street at night, neon reflecting in puddles, steam rising from a food cart, a person with a translucent umbrella, shallow depth of field, teal-and-orange grade, smooth steady camera

Compare every video model on these prompts →

Alternatives

How it compares

ModelBest forMax durationMax resolutionNative audio
Grok Imagine 1.5 LiteCheap, fast audio-complete social clips15s1080pYes
Grok Imagine 1.5Maximum fidelity, references15s1080pYes
Grok ImagineFast drafts at the lowest tier15s720pYes
Wan 2.7 EnhancedOpen-weights alternative with audio15s1080pYes
MiniMax H3 MaxPremium 1080p clips with audio15s1080PYes

Grok Imagine 1.5 Lite is the right pick when you need many short, sound-complete 1080p clips at the lowest per-clip cost — social ads, loops, and rapid iteration. Step up to Grok Imagine 1.5 when you need multi-reference character and scene control.

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/video/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-1-5-lite-image-to-video",
    "prompt": "Aerial drone shot over a misty mountain valley at golden hour"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.

Pricing

What it costs on Venice

Pay per clip on Venice — price scales with resolution and duration (1s–15s), from $0.04.

480p · 1s
$0.04
Per clip
720p · 1s
$0.05
Per clip
1080p · 1s
$0.18
Per clip

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

FAQ

Frequently asked questions

Grok Imagine 1.5 Lite is xAI's lightest tier of the Grok Imagine Video 1.5 family, available on Venice since September 30, 2026. It comes in two variants — text-to-video and image-to-video — and generates clips at 480p, 720p, or 1080p, from 1 to 15 seconds, with audio generated natively in the same pass.

On Venice you pay per clip, with the price scaling by resolution and duration: $0.04 for a 480p 1-second clip, $0.05 for 720p 1s, and $0.18 for 1080p 1s, up to 15 seconds. There is no subscription required — credits cover everything.

It is neither free-software nor open-source: Grok Imagine 1.5 Lite is proprietary to xAI with no published weights, so it cannot be self-hosted or fine-tuned. On Venice you can try it with welcome credits before paying per clip. If open weights matter, Wan 2.7 Enhanced is the closest open alternative with native audio.

Yes — that is the grok-imagine-1-5-lite-image-to-video variant, which animates a still image into motion at 480p/720p/1080p for 1–15 seconds with audio. The family also ships a separate grok-imagine-1-5-lite-text-to-video variant that generates a clip from a written prompt alone.

Both run privately on Venice with 1080p, 15-second clips, and native audio. Lite is the cheaper tier, ideal for volume work and drafts; the flagship Grok Imagine 1.5 adds multi-reference control (up to seven image references) and voice references for holding a character's face and voice across scenes. Use Lite for speed and cost, the flagship for character consistency.

Three resolutions — 480p, 720p, and 1080p — and any duration from 1 to 15 seconds in one-second increments, on both the text-to-video and image-to-video variants. Audio is generated natively at every setting.

Yes. Following the Grok Imagine Video 1.5 family design, sound effects, ambience, and dialogue are generated in the same pass as the video and land on the action, with clearer and better-synced speech than the previous generation. Describe the audio you want directly in the prompt.

Venice runs it under the private tier with zero retention: prompts and source images are not stored, profiled, or used for training, and nothing is tied to a personal generation history. That contrasts with xAI's own apps, where generations live in an account-linked library.

Multi-reference and voice-reference features were announced for the flagship grok-imagine-video-1.5 model, not documented for the Lite tier — on Venice, Lite takes a single still image for image-to-video or a text prompt alone. If you need up to seven reference images or voice consistency, use Grok Imagine 1.5.

Run Grok Imagine 1.5 Lite privately

No prompt logging. No data used for training.