Playcut v2 is coming MCP · Actor v2 · & much more
ai image and video models compared · 13 surfaces · one price list

Every model we run.
What it does. What it costs.

3 image models, 7 video models, music and voice — all on one credit balance, in one workspace. This page is the buying guide, not the spec sheet. What each model is genuinely best at, what a render costs you, and what each one flatly refuses to do.

Rates in credits · you pick the model, every time · no per-model paywall

The full list

13 model surfaces, grouped by what you are making

Every rate is in Playcut credits — what your balance is debited, not a list price from a provider. The pricing page shows how many credits each plan includes.

The last column is the one people get wrong. 11 of the 13 surfaces have a tool behind them, so they run from the studio, the MCP server and the REST tools API. The other 2 are studio-only, as is Nano Banana Pro's multi-region edit.

Image

Priced per finished image.

Three image models, priced per image rather than per month. The default is Nano Banana Pro, and nothing swaps it behind your back.

Model Best at Your cost Known limits Where it runs
Nano Banana Pro Google Default Legible text inside the image, plus style references — Nano Banana 2 matches its 14-reference cap but takes objects and characters only.
  • 1K / 2K 67 cr per image
  • 4K 84 cr per image
  • Multi-region edit 25 cr flat per run
  • Most expensive image model we run — about two Nano Banana 2 renders per Pro render.
  • No 512px budget tier; that one is exclusive to Nano Banana 2.
  • Multi-region edit is studio-only and caps at 8 regions per run.
Studio · MCP · REST generate-image-from-text and generate-image-from-reference both default to it. Multi-region edit has no tool — studio only.
Nano Banana 2 Google Iteration volume — the same Google stack as Pro at roughly half the credits.
  • 512px 20 cr per image
  • 1K / 2K 34 cr per image
  • 4K 43 cr per image
  • Its 14 reference slots are objects and characters — style references are a Pro feature.
  • Nano Banana Pro is the model our picker flags for text-heavy work: signs, posters, packaging.
  • The 512px tier is a thumbnail budget, not a delivery format.
Studio · MCP · REST Both image tools accept it, including the 512px tier via the imageSize parameter.
Grok Imagine Image xAI Cheapest look-see pass, and the widest aspect-ratio list including an auto mode.
  • Text-to-image, 1K / 2K 10 cr per image
  • Reference-to-image, 1K / 2K 20 cr per image
  • Caps at 2K — no 4K output at all.
  • Accepts a maximum of 2 reference images.
  • Not the model to reach for when text has to render cleanly.
  • Studio only — the MCP and REST image tools accept the two Gemini models, not this one.
Studio only generate-image-from-text and generate-image-from-reference enumerate the two Gemini models only.

Video

Priced per second of output.

Seven video models on one credit balance. Rates are per second of finished video, so a four-second draft costs four times the listed rate.

Model Best at Your cost Known limits Where it runs
Veo 3.1 Google Default The only model covering all five generation modes: text, image, reference, interpolation and extension.
  • 720p / 1080p 160 cr per second
  • 4K 240 cr per second
  • 16:9 and 9:16 only — no square, no 4:5, no 21:9.
  • Duration is a fixed 4 / 6 / 8 second ladder, and 1080p, 4K, reference, extension and interpolation all force 8 seconds.
  • Accepts at most 3 reference images, against 9 on Seedance 2.0 and HappyHorse 1.1.
  • No video-edit mode; prompt-driven edits run on Wan 2.7 or HappyHorse 1.1.
  • By far the most expensive second of video on the platform.
Studio · MCP · REST Default model on every video tool except video edit, which defaults to wan-2.7 — Veo has no edit mode.
Seedance 2.0 ByteDance Capability per credit, and the only non-Veo route to 4K video.
  • 480p 10 cr per second
  • 720p 21 cr per second
  • 1080p 60 cr per second
  • 4K 108 cr per second
  • Rejects input images containing human faces — the provider returns an error and we surface it rather than hiding the model.
  • Text-to-video is unaffected by that filter; it only bites the image-driven modes.
  • No extension mode, and its interpolation takes exactly a first and last frame — no references alongside.
Studio · MCP · REST Pass model: "seedance-2.0" to the text, image, reference and interpolation video tools.
Grok Imagine xAI The cheap draft pass, at seven aspect ratios, with reference and extension modes still on the table.
  • 480p / 720p 20 cr per second
  • 720p ceiling — no 1080p, no 4K.
  • No interpolation mode, and no generated audio.
  • Text-to-video runs to 15 seconds; every other mode caps at 10.
Studio · MCP · REST Pass model: "grok-imagine-video" to the text, image and reference video tools.
Grok Imagine 1.5 xAI A sharper animate pass than classic Grok at the same 720p ceiling.
  • 480p / 720p 32 cr per second
  • Image-to-video only. It rejects reference images, and text-to-video, extension and interpolation are all unavailable.
  • 720p ceiling, and no generated audio.
  • Costs 12 credits a second more than classic Grok Imagine for fewer modes.
Studio · MCP · REST Pass model: "grok-imagine-video-1.5-preview" to generate-video-from-image — the only mode it supports.
Gemini Omni Flash Google One self-contained beat with spoken dialogue, without tuning a duration dial.
  • 720p 40 cr per second
  • No duration control. The model decides, between 3 and 10 seconds.
  • We hold credits for the 10-second maximum and refund the difference when the clip comes back shorter.
  • 720p only.
  • No extension, interpolation or video edit.
Studio · MCP · REST Pass model: "gemini-omni-flash-preview" to the text, image and reference video tools; durationSeconds must be 10.
HappyHorse 1.1 HappyHorse · served via Alibaba DashScope Heavy reference locking at up to 9 images — tied with Seedance 2.0 — plus nine aspect ratios on text and reference shots, the widest list of the seven.
  • 720p 56 cr per second
  • 1080p 72 cr per second
  • 720p and 1080p only — no 480p, no 4K.
  • No extension or interpolation.
  • Video edit is billed on source seconds plus output seconds, not output alone.
  • Video edit reads only the first 15 seconds of a longer source — trim first.
Studio · MCP · REST Pass model: "happyhorse-1.1" to the text, image, reference and video-edit tools.
Wan 2.7 Alibaba Audio-driven lip sync — feed a still plus an audio track and the performance follows it.
  • 720p 40 cr per second
  • 1080p 60 cr per second
  • No extension or interpolation.
  • Video edit is billed on source seconds plus output seconds.
  • Video-edit sources must be 2 to 10 seconds — the provider rejects anything longer.
  • Takes at most 5 reference images.
Studio · MCP · REST Default model on generate-video-edit; also on the text, image and reference video tools.

Music

Priced flat per song.

One music surface with two lengths. Flat per song, not per second, so a re-roll costs exactly what the first attempt did.

Model Best at Your cost Known limits Where it runs
Lyria 3 Google Default Flat-rate scoring — a full track costs the same whether it is your first attempt or your fifth.
  • Clip (30 seconds) 20 cr flat per song
  • Pro (about 2 minutes) 40 cr flat per song
  • MP3 only — there is no WAV output.
  • Blocks prompts naming real artists, albums, songs or producers, including when you negate them.
  • Pro has no duration parameter; you get roughly two minutes.
  • Studio only — there is no music tool on the MCP server or the REST tools API.
Studio only The tool registry carries no music tool, so neither length can be triggered programmatically.

Voice

Priced per 1,000 characters of script.

Two text-to-speech providers plus voice design and cloning. Character-based pricing means a 300-word script is a rounding error.

Model Best at Your cost Known limits Where it runs
Qwen TTS Alibaba Default The only provider here that designs a voice from a description or clones one from audio.
  • Speech 15 cr per 1,000 characters
  • Voice design 20 cr flat per voice
  • Clone from audio 5 cr flat per voice
  • 600 characters per request — long scripts have to be chunked.
  • Ten core languages, against the wider list on the Google model.
  • No inline style tags; steering is free-text prompt only, and cloned voices ignore it.
Studio · MCP · REST design-voice, clone-voice-from-audio and generate-tts are all in the tool registry.
Gemini 3.1 Flash TTS Google The wider language list of the two, plus inline style tags for delivery direction.
  • Speech 25 cr per 1,000 characters
  • Cannot design a voice and cannot clone one — speech only.
  • Costs about 67% more per 1,000 characters than Qwen.
Studio · MCP · REST generate-tts reaches it through any Google-backed voice returned by list-voices.
How to pick

One decision rule per model

The table tells you what things cost. This tells you which one to reach for. Read the line that matches the job in front of you and ignore the rest.

Image

Nano Banana Pro 67 cr

Pick this when: The image is going to ship — packaging copy, a campaign hero, or anything a client signs off on.

Nano Banana 2 20 cr

Pick this when: You expect to regenerate the same brief four or five times before you like it.

Grok Imagine Image 10 cr

Pick this when: You are pressure-testing a hundred compositions and only three will survive.

Video

Veo 3.1 160 cr

Pick this when: This is the hero shot, or you need interpolation or extension, which no other model here offers together.

Seedance 2.0 10 cr

Pick this when: The shot is product, environment, texture or abstract motion, and you want 4K without paying Veo rates.

Grok Imagine 20 cr

Pick this when: You want to know whether the shot idea works before you spend flagship credits on it.

Grok Imagine 1.5 32 cr

Pick this when: You already have the still and only need it to move well.

Gemini Omni Flash 40 cr

Pick this when: You want one short shot with dialogue in it and do not want to think about the settings.

HappyHorse 1.1 56 cr

Pick this when: The look has to match a large pile of references, or you are editing an existing clip.

Wan 2.7 40 cr

Pick this when: You have the voice track already and need the mouth to match it, or you are re-cutting an existing clip by prompt.

Music

Lyria 3 20 cr

Pick this when: You need a bed under a cut and licensing a real track is not worth the paperwork.

Voice

Qwen TTS 15 cr

Pick this when: You want a reusable brand voice, or you are cloning a founder's read.

Gemini 3.1 Flash TTS 25 cr

Pick this when: The script is going out in a language Qwen's ten core languages do not cover.

Actor performance video, priced by tier

The Act app is not a Studio model picker entry — it is four engine tiers, and the tier is the only engine choice you make. Every tier adds a flat 15-credit scene charge on top of the per-second rate. The model speaks the dialogue in your prompt itself, so there is no separate voice line to pay for. All four tiers run over MCP and REST as well as in the app.

Tier Per second Length Notes
Fastest (default) 480p 7 cr 720p 7 cr 6 or 10 seconds Up to 6 reference images. The draft tier — a 6-second take lands at 57 credits all in.
Fast 480p 17 cr 720p 17 cr 4 to 10 seconds Scene image only — no extra reference stack.
Plus 720p 30 cr 1080p 38 cr 4 to 15 seconds Up to 8 image references, and the only tier that reaches 1080p.
Pro 720p 35 cr fixed 10 seconds Fixed length, so a take always costs 365 credits.
The honest part

What these models will not do

Every one of these is a real refusal we have hit in production. We surface the provider's error instead of quietly rerouting you to a different model and a different bill.

Seedance 2.0 rejects input images with human faces
ByteDance enforces this upstream, on both the direct and resold routes. It is the single biggest gotcha on the cheapest video model. Use Seedance for product, environment and abstract motion; use Grok Imagine or Veo 3.1 when a person is in the source frame.
Lyria 3 blocks real artist, album and song names
Naming a real musician trips the safety filter even when you negate it — "nothing like" a named artist fails the same way the plain reference does. Describe the sound instead: tempo, instrumentation, era, mood, mix character. That phrasing gets through and produces a closer result anyway.
Veo 3.1 forces 8 seconds at 1080p and 4K
The 4 / 6 / 8 second ladder only opens up at 720p on plain text-to-video and image-to-video. Reference, extension and interpolation are pinned to 8 seconds at every resolution. Budget accordingly: at 160 credits a second, an 8-second 1080p shot is 1,280 credits.
Grok caps at 720p video and 2K images
Both Grok Imagine video models stop at 720p, and Grok Imagine Image stops at 2K with a maximum of two reference images. Neither video model generates audio. Grok Imagine Image at 10 credits is the cheapest image model we run; on video, Seedance 2.0 at 480p undercuts Grok at 10 credits a second. Draft on Grok, finish somewhere else.
Grok Imagine 1.5 is image-to-video only
It rejects reference images outright, and text-to-video, extension and interpolation are not available on it. If you do not already have a still to animate, use classic Grok Imagine instead — same 720p ceiling, more modes, and 12 credits a second cheaper.
Two models and one edit flow are studio-only
Lyria 3 music, Grok Imagine Image and Nano Banana Pro's multi-region edit have no tool on the MCP server or the REST tools API. You can run them in the studio; you cannot script them. Everything else — every video model, both Gemini image models, the Act tiers, both voice providers — is reachable either way.
Gemini Omni Flash decides its own length
There is no duration control on it. The model returns somewhere between 3 and 10 seconds, so we hold credits for the 10-second maximum and refund the difference once the clip lands. You cannot budget a shot on it to the second in advance — the ceiling is 400 credits, the floor is 120.
Video edit bills the source clip too
On Wan 2.7 and HappyHorse 1.1, a prompt-driven edit charges source seconds plus output seconds, because that is how the provider meters it. A 10-second source producing a 10-second result bills as 20 seconds. It is the one place on this page where the per-second rate does not map one-to-one to the clip you get back.

You pick the model. Every capability has a default.

This is the clearest place on the site to say it plainly. Playcut does not run a classifier that reads your prompt and decides which model gets your credits. You open the picker and choose, per generation, and that choice is what we bill.

Each capability ships a default so the picker is optional. Images default to Nano Banana Pro and video defaults to Veo 3.1.

Music defaults to Lyria 3 Pro and voice defaults to Qwen. Change nothing and those four are what render.

The reason we are explicit about this: an automatic router is a pricing problem disguised as a convenience feature. If a hidden system can move your job from a 20-credit model to a 160-credit one, you cannot forecast a campaign budget. Fixed defaults plus a visible picker means the number on this page is the number on your invoice.

How the credit price actually works

One balance covers everything — images, video, music, voice, actor performances. There is no per-model subscription and no capability locked behind a higher tier. Plans differ on how many credits land each month and how many seats you get, not on which models you may touch.

Rates are set per unit of output: per image, per second of video, per song, per 1,000 characters of script. Resolution moves the rate on most models. Every figure on this page is in credits, so you can total a campaign against your balance before you commit to it.

A worked example. Twelve product stills on Nano Banana 2 at 34 credits each is 408 credits. Three 8-second hero shots on Veo 3.1 at 1080p is 3,840 credits.

Swap those drafts to Grok Imagine video at 20 credits a second and the same three 8-second passes cost 480 instead. Full plan breakdown lives on the pricing page.

Draft cheap, finish expensive

This is the workflow nearly every team converges on after a month, and the multi-model picker exists to make it possible. Generate on the cheapest model that can answer the question you are actually asking, then re-run the survivors on the flagship.

For video, that means Grok Imagine at 20 credits a second to test whether a shot idea reads at all, then Veo 3.1 at 160 for the take that ships. The cheap pass costs one eighth of the expensive one, and prompts port between models with light editing.

For images, it means Nano Banana 2 at 34 credits for the exploration and Nano Banana Pro at 67 for the frame a client signs off on. Same Google stack, so the aesthetic does not lurch when you promote a winner. Grok Imagine Image at 10 credits sits below both for pure volume.

The thing that makes this work is that everything lands in the same workspace, on the same balance, under the same brand kit. Switching models mid-project costs nothing but credits — no new subscription, no re-upload, no separate asset library to reconcile at the end.

Where each capability's models actually differ

Image

The three image models separate on text rendering, resolution ceiling and reference kinds. Nano Banana Pro and Nano Banana 2 both accept up to 14 reference images and both reach 4K. The difference is what those slots hold: Pro takes style references, Nano Banana 2 takes objects and characters. Grok Imagine Image accepts two references and stops at 2K.

Nano Banana Pro is also the only one with multi-region edit — mask up to eight areas of a single image and give each its own instruction, at a flat 25 credits however many you mask. That is the feature worth paying the Pro rate for when a layout needs surgical fixes rather than a re-render.

It runs in the studio only; there is no tool for it. Full breakdown on Nano Banana Pro pricing and limits, and more on the flows in the AI image generator breakdown.

Video

Video separates on modes, not just price. Veo 3.1 is the only model covering text-to-video, image-to-video, reference-to-video, interpolation and extension in one place. It has no video-edit mode, though — that is Wan 2.7 and HappyHorse 1.1 only.

Native audio is not a Veo exclusive. Five of the seven video models generate audio: Veo 3.1, Seedance 2.0, Gemini Omni Flash, HappyHorse 1.1 and Wan 2.7. The two Grok Imagine models are the silent ones, which is part of why they are the cheap draft pass — the full arithmetic is on Grok Imagine pricing and limits.

Reference capacity is the other axis. HappyHorse 1.1 and Seedance 2.0 both take up to nine reference images; Wan 2.7 takes five; Veo 3.1 takes three; Grok Imagine 1.5 takes none at all. If your look is defined by a large pile of references, that constraint decides the model before price does. The AI video generator page walks through the video types in production terms.

Music and voice

Music is the simplest surface here: two lengths, flat rates, MP3 out. A re-roll costs exactly what the first attempt did, which makes iteration cheap and predictable in a way per-second video never is.

Voice splits on capability rather than quality. Qwen is the only provider that designs a voice from a description or clones one from an audio sample, at 20 and 5 credits respectively. Gemini Flash TTS cannot do either — it reads scripts, across a wider language list, at 25 credits per 1,000 characters against Qwen's 15.

Reaching the models programmatically

Most of this page is available over the Model Context Protocol server and the REST tools API, on the same credit balance as the studio. 11 of the 13 surfaces have a tool behind them, and the model id is a parameter on the generation call. So the draft-cheap-finish-expensive pattern automates cleanly for those.

The exceptions are Grok Imagine Image and Lyria 3, which have no tool and run in the studio only. Nano Banana Pro's multi-region edit is the same story — a flow rather than a model row, with no tool behind it. Every video model, the Act tiers and both voice providers are reachable. The "Where it runs" column says which is which.

That is the developer surface. Field-level specs, parameter names and error codes live in the API documentation rather than here — this page is for deciding what to spend on. Setup and tool list are on the MCP and API page.

Two things this page is not

It is not a spec sheet. If you need to know which field carries the aspect ratio or what a specific error code means, the developer docs answer that and this page deliberately does not.

It is also not about AI fashion models — different word, different product. Looking for AI fashion models instead? The custom AI fashion model builder covers building a reusable synthetic model for lookbooks, stills and on-product compositing.

Model comparison FAQ

Common questions

Does Playcut pick the model for me? +

No. You pick the model, every time. Each capability has a default so you never have to think about it — Nano Banana Pro for images, Veo 3.1 for video, Lyria 3 Pro for music, Qwen for voice — but nothing is swapped behind your back. If you never open the model picker, you get the default and the price you see next to that default is the price you pay.

How much does one image or one second of video actually cost? +

In credits. An image on Nano Banana Pro is 67 credits at 1K or 2K, and a second of Veo 3.1 is 160 credits at 720p or 1080p. At the cheap end, a second of Grok Imagine video is 20 credits and a Lyria 3 Pro song is a flat 40 credits. Every rate on this page is the rate we charge, and the pricing page shows how many credits each plan includes.

Can I mix models inside one project? +

Yes, and most teams do. Model, resolution, aspect ratio and reference images are all set per generation, so you can draft twelve shots on Grok Imagine at 20 credits a second and re-run the two winners on Veo 3.1. Everything lands in the same workspace folder, on the same credit balance, under the same brand kit.

Which video model is cheapest? +

Seedance 2.0 at 480p, at 10 credits per second. Grok Imagine is next at 20 credits per second with a 720p ceiling. The catch on Seedance is that it rejects input images containing human faces, so for actor work the cheapest usable option is usually Grok Imagine.

Why do some models refuse certain prompts or inputs? +

Because the providers enforce their own rules and we do not hide them. Seedance 2.0 errors on input images containing human faces, Lyria 3 blocks prompts naming real artists or songs even when you negate them, and Grok Imagine 1.5 rejects reference images entirely. We surface the real provider error rather than silently rerouting you to a different model and a different bill.

Can I drive every model over the API? +

Most of them, not all. Every video model, both Gemini image models, the Act tiers and both voice providers have a tool on the MCP server and the REST tools API. Two models do not — Lyria 3 music and Grok Imagine Image — and neither does Nano Banana Pro's multi-region edit. Those are studio-only, and the table above marks which is which.

Do I need a different plan to use the expensive models? +

No. Every plan can reach every model. Plans differ on how many credits arrive each month, how many seats you get, and workspace features — not on which models are unlocked. That means a Hobby account can render on Veo 3.1; it just burns through the monthly allowance faster than the same work on Grok Imagine.

Ready to make something cinematic?

Pick a plan. Hobby starts with a 7-day trial — cancel within 7 days at no charge.