
Seedance 2.5
flagshipByteDance
Up to 30 seconds in one unbroken shot — the longest we run. Reads up to 30 reference images, lip-syncs to a voice you upload, and renders 1080p in 10-bit HDR.
The directory
New frontier models land in your picker within seventy-two hours of release. No upgrade, no migration, no second subscription.
09
image
18
video
10
voice & music
01
3D
§ 02 — The optics

ByteDance
Up to 30 seconds in one unbroken shot — the longest we run. Reads up to 30 reference images, lip-syncs to a voice you upload, and renders 1080p in 10-bit HDR.

Google DeepMind
Google's flagship. Native audio, strong physics, and the cleanest motion of anything here.

OpenAI
Coherent long takes with sound. Portrait or landscape, four to twelve seconds.

Kuaishou
Fifteen seconds at 1080p, and the most permissive of our engines with reference photographs.

Kuaishou
Built around reference images — up to seven — for keeping a character consistent shot to shot.

MiniMax
MiniMax H3 through their own door — expressive character motion, any length from 4 to 15 seconds at either tier.

Alibaba
Fifteen seconds at 1080p with a proper negative prompt. Good value for longer work.

Runway
Runway's flagship — precise camera language, any length from two to ten seconds.

Luma
Ray 3.2 through Luma's own door — smooth, natural motion from a prompt or a still. Five or ten seconds, up to 1080p.

Black Forest Labs
Black Forest Labs' own door. Eight shapes, from square to 9:16.

Black Forest Labs
The quick FLUX. Nine shapes, the widest choice of any image engine here.

xAI
xAI's own door. Seven shapes.

xAI
The slower Grok pass, for when the first one is nearly right. Seven shapes.

Kuaishou
Kling's newest stills, at 1K or 2K for the same price. Eight shapes, and it takes a reference picture.

Kuaishou
Kuaishou's own door, newest of the three. Takes any shape the picker offers.

Reads up to six reference images and holds a product or a face across every one. The workhorse behind Product and Cast.

Black Forest Labs
Black Forest Labs' highest-fidelity image model. Fine detail and text that stays readable.

ByteDance
Strong composition and colour, and unusually good at following a long, specific brief.

Alibaba
Alibaba. Up to 15 seconds at 1080p, from a written prompt. 55 credits for five seconds.

ByteDance
ByteDance. Up to 15 seconds at 1080p, from a prompt or from a still you hand it. 95 credits for five seconds.

Vidu
Vidu. Up to 8 seconds at 1080p, from a prompt or from a still you hand it. 130 credits for five seconds.

Vidu
Vidu. Up to 5 seconds at 1080p, from a written prompt. 130 credits for five seconds.

Vidu
Vidu. Up to 8 seconds at 1080p, from a prompt or from a still you hand it. 160 credits for five seconds.

Kuaishou
Kuaishou. Up to 15 seconds at 1080p, from a prompt or from a still you hand it. 260 credits for five seconds.

Kuaishou
Your photo, performing the movement of a clip you upload. As long as the clip, up to 30 seconds.

Runway
Changes a video you already have — restyle it, relight it, take something out. Priced on the clip's own length.

xAI
xAI's own. One to fifteen seconds at up to 1080p with synchronised sound, and an exact last frame.

ElevenLabs
Scored to length — up to 5 minutes in one pass.

MiniMax
Original score from a written brief. Instrumental or with a vocal line.

MiniMax
Original score from a written brief. Instrumental or with a vocal line.

Original score from a written brief. Instrumental or with a vocal line.

Original score from a written brief. Instrumental or with a vocal line.

ACE-Step
Scored to length — up to 4 minutes in one pass.

MusicGen
Scored to length — up to 30 seconds in one pass.

ElevenLabs
Ten stock voices, plus any voice you clone from your own recordings. Thirty-two languages.

Bytedance
One clear photo of a face becomes a presenter. Best for short lines.

Sync
Takes a clip you already have and matches the mouth to new audio.

Microsoft
One photograph becomes an actual mesh — a .glb for Blender, Unity or a game engine, plus a turntable render so you can see it without installing anything.
Capability matrix
| Model | Type | Ceiling | Speed | Credits | Strengths |
|---|---|---|---|---|---|
| Seedance 2.5ByteDance | video | 1080p | balanced | 325 | 30sreferenceslip-sync |
| Veo 3.1Google DeepMind | video | 1080p | balanced | 440 | audiophysicscinematic |
| Sora 2OpenAI | video | 720p | balanced | 140 | audiocoherence |
| Kling 3.0Kuaishou | video | 1080p | balanced | 330 | 15sreferences1080p |
| Kling O1Kuaishou | video | 1080p | fast | 230 | referencescharacter |
| Hailuo 2.3MiniMax | video | 1080p | balanced | 170 | charactermotion |
| Wan 2.7Alibaba | video | 1080p | balanced | 160 | 15snegative-prompt |
| Runway Gen-4.5Runway | video | 720p | balanced | 150 | camerastart frame |
| Luma Ray 3.2Luma | video | 1080p | balanced | 260 | 10sstart frame |
| FLUX 2 ProBlack Forest Labs | image | — | fast | 12 | image |
| FLUX 2 KleinBlack Forest Labs | image | — | fast | 12 | image |
| Grok ImaginexAI | image | — | fast | 12 | image |
| Grok Imagine QualityxAI | image | — | fast | 12 | image |
| Kling Image 3.0Kuaishou | image | 2K | balanced | 12 | image |
| Kling Image 2.1Kuaishou | image | — | fast | 12 | image |
| Nano Banana ProGoogle | image | 2048² | fast | 12 | referencesproductediting |
| FLUX 2 MaxBlack Forest Labs | image | 4MP | fast | 12 | photorealtypographydetail |
| Seedream 5 LiteByteDance | image | 3072² | fast | 12 | compositionprompt-following |
| Wan 2.6Alibaba | video | 1080p | fast | 160 | 15s1080p |
| Seedance 2.0ByteDance | video | 1080p | balanced | 250 | 15s1080pstart frameflexible length |
| Vidu Q2Vidu | video | 1080p | balanced | 230 | 8s1080pstart frame |
| Vidu Q1Vidu | video | 1080p | balanced | 230 | 5s1080p |
| Vidu Q3Vidu | video | 1080p | deliberate | 260 | 8s1080pstart frame |
| Kling 3.0 OmniKuaishou | video | 1080p | deliberate | 360 | 15s1080pstart frame |
| Kling Motion ControlKuaishou | video | 1080p | deliberate | 170 | your clip1080pstart frame |
| Runway Aleph 2Runway | video | 720p | deliberate | 215 | your clipediting720p |
| Grok Imagine 1.5xAI | video | 1080p | fast | 135 | 15s1080psound |
| ElevenLabs MusicElevenLabs | music | from 10s | fast | 90 | scorecommercial rights |
| MiniMax Music 2.6MiniMax | music | from 30s | fast | 140 | scorecommercial rights |
| MiniMax Music 2.5MiniMax | music | from 30s | fast | 120 | scorecommercial rights |
| Google Lyria 3 ProGoogle | music | from 30s | fast | 130 | scorecommercial rights |
| Google Lyria 3Google | music | from 30s | fast | 60 | scorecommercial rights |
| ACE-StepACE-Step | music | from 30s | fast | 80 | scorecommercial rights |
| MusicGenMusicGen | music | from 8s | fast | 40 | scorecommercial rights |
| ElevenLabs v2ElevenLabs | voice | 48kHz | instant | 10 | cloning32 languages |
| Talking photoBytedance | voice | per video | balanced | 300 | lip-syncfrom a photo |
| Re-sync a videoSync | voice | per second | balanced | 4 | lip-syncfrom a clip |
| TRELLISMicrosoft | 3d | 1024px textures | deliberate | 80 | glbturntable |
Every video and image you generate is now kept permanently, with the prompt that made it, and reachable from any device.
Upload a recording and the engine lip-syncs your character to it — the missing piece for presenter and talking-head clips.
Builds a shot from up to four reference images. Fifteen seconds at 1080p.
Failed renders are refunded automatically, even if you closed the tab — a sweep settles anything left outstanding.
Start in the dark
Pay only for what you make — no subscription, no watermark. Open the studio and see whether it thinks the way you do.