Skip to main content
Product Updates

Synthesys Studio: 35 AI Models for Image, Video and Voice

Synthesys Studio is an all-in-one AI image, video and voice tool: 35 models, the credit price before you press Generate, and no charge for failed runs.

NKNick Koukoulakis24 min read
Synthesys Studio: 35 AI Models for Image, Video and Voice
Product Updates
Pillar guide
AI Video Generation: The Complete 2026 Guide

The Synthesys Video Agent builds a whole video from one sentence. Synthesys Studio does the opposite: you pick the exact AI model, set every option yourself, and get one image, clip or voiceover back.

We launched Synthesys in 2020 as an AI voice generator. Since then, every good new model has shown up with its own site, its own credits and its own surprises on the bill. We built Studio so you can use them in one place and see the price before you press anything.

Synthesys Studio: an image, a clip and a voiceover, each with its credit price, above a Generate button

What is Synthesys Studio?

Synthesys Studio is where you make single pieces in the Synthesys app, with the model you choose: one image, one clip or one voiceover. It sits next to the Video Agent and shares its Files. You see the price before anything runs.

Synthesys is an AI video agent that replaces traditional video production end to end. The agent writes, plans and renders a whole video for you. Studio hands you the models directly instead.

You open it from the sidebar, where it's marked NEW. Three tabs sit at the top: Image, Video and Voiceover. All three share one layout, so once you know one tab you know the other two.

The Studio screen in the Synthesys app, with five numbered parts

  1. Studio in the sidebar.
  2. The tool tabs: Image, Video and Voiceover. Each keeps its own prompt.
  3. Inputs and prompt: references, frames or a voice, then the prompt, shot or script. Studio only shows the slots the chosen model can use.
  4. The bottom bar: size and length settings, the model chip with its live price, and Generate. ⌘/Ctrl + Enter also generates.
  5. The right pane: History, where every result lands, and How it works, a three-step guide for each tool.

Here is the one-minute tour from Studio's own training lessons:

When should you use Studio, and when the Video Agent?

StudioVideo Agent
What you give itA prompt, a shot or a script, plus the model and settingsA brief, as short as one sentence
What you get backOne image, one clip or one voiceoverA finished video with script, scenes, voice and music
Who picks the modelYou, from 35The agent, by default
When you see the costOn the model chip, before you press GenerateOn the approve button, before anything renders
Best forA product still, a single shot, a voice line, testing a modelAds, explainers, UGC and long videos

Our advice: if you're starting from a blank page, skip Studio. Tell the agent what the video is for and let it do the planning. Come to Studio when you already know the piece you need, or when you want to try a model before it goes into something bigger.

How do credits work in Studio?

Studio shows the price on the model chip, on every row of the model list and on every quality or resolution option. Pressing Generate takes nothing yet. Credits come off only when the result lands in History, and a failed run costs nothing.

The rules are the same on all three tabs:

  • Priced before you press. Change the quality, the resolution, the length or the number of images, and the chip updates.
  • Not charged yet. While a result renders, the card says how many credits will come off when it lands.
  • Failures are free. A run that fails, runs too long or is refused by a model's safety filter shows "No credits charged". In a batch of images, only the ones that arrive are charged.
  • Short on credits? Nothing starts. Studio shows how many credits you're short and which models fit your balance.

GPT Image 2 quality options at 3, 13 and 44 credits, and MiniMax H3 Max Turbo resolutions at 25, 40 and 80 credits

On GPT Image 2, Low, Medium and High cost 3, 13 and 44 credits per square image. On MiniMax H3 Max Turbo, a 5-second clip costs 25 credits at 480p, 40 at 720p and 80 at 1080p. On a few image models the shape changes the price too: GPT Image 2 at Medium costs 11 credits in 16:9 instead of 13.

A clip rendering in Studio, with the stages Queued, Generating and Finishing and the line Not charged yet

  1. The clip moves through Queued, Generating and Finishing.
  2. "Not charged yet: 25 credits come off when the clip lands."
  3. You can close the tab. The clip keeps rendering and lands in History and Files.

This is the same promise we make in the Video Agent, where a render waits for your approval with its cost on the button. We wrote about why that matters in AI Video Credits: Why They Get Wasted and How to Stop It.

What did the runs in our training lessons cost?

We recorded Studio's training lessons in the app on 27 September 2026: single runs, not a benchmark. Every time the app showed both numbers, the credits charged matched the chip. Eight runs, from a 6-credit voiceover to a 40-credit clip, came to 146 credits.

What we madeModel and settingsChip priceChargedTime to land
A ceramic mug on a kitchen counterGPT Image 2, Medium, 1:11313Under a minute
The mug on a table, from 2 referencesGPT Image 2, Medium, 1:11313About 40 seconds
Edit: "make the mug matte black"GPT Image 2, Edit, Medium1313Not recorded
A 5-second push-in on the mugMiniMax H3 Max Turbo, 480p, 16:92525About 9 seconds
A 5-second clip from start and end framesMiniMax H3 Max Turbo, 480p2525About 10 seconds
A 4-second macro shot with a reference, a Look and audioSeedance 2.0 Mini, 480p4040About 2 minutes
A 129-character voiceoverElevenLabs Multilingual v2, voice Roger66A few seconds
Three designed voice samplesDesign a voice11Not shownAbout 8 seconds

Studio says clips usually take 1 to 5 minutes, so our two MiniMax clips were quick. The result screenshots in this article are frames from those lessons.

A GPT Image 2 result in Studio History, saved to Files and charged 13 credits

  1. "Saved to Files · charged 13 credits", the same price the chip showed.
  2. Use in a video sends the image straight into a Video Agent chat.

Which image model should you use?

Start with GPT Image 2 at Medium (13 credits), Studio's default. Draft on FLUX.1 [schnell] from 1 credit when you're still finding the idea. Use Ideogram v4 for words inside an image, and Nano Banana 2 or Pro when you need up to 14 reference images.

Studio's Image tab is an AI image generator with 18 models from 12 makers, grouped into three tiers. Each row in the model list shows the top size or quality, how many reference images the model takes, and its price.

The Studio image model list, grouped Fast and cheap, Balanced and Highest quality, with credits per image

  1. Models are grouped by tier: Fast & cheap, Balanced and Highest quality.
  2. The cheapest model wears a CHEAPEST badge: FLUX.1 [schnell] at 1 credit for a 1K square.
  3. A range, like 3–44, covers a model's quality or size options. On Seedream 5.0 Pro and Grok Imagine Image, it also covers the number of references.
ModelMakerTierCredits per image (1:1)Max referencesEdit mode
FLUX.1 [schnell]Black Forest LabsFast & cheap1–3Text onlyNo
MiniMax Image-01MiniMaxFast & cheap2Text onlyNo
FLUX.2 [dev]Black Forest LabsFast & cheap3–10Text onlyNo
Grok Imagine ImagexAIFast & cheap63Yes
Ideogram v4IdeogramFast & cheap7Text onlyNo
GPT Image 2.5 FlareOpenAIBalanced5–14Text onlyNo
Kling Omni 3KuaishouBalanced6–1210Yes
FLUX.2 [pro]Black Forest LabsBalanced6–15Text onlyNo
Recraft V4.1Recraft AIBalanced7Text onlyNo
Seedream 4.5ByteDanceBalanced810Yes
FLUX.1 Kontext [pro]Black Forest LabsBalanced81Yes
Bria FIBO Gen 1.5Bria AIBalanced84Yes
Nano Banana 2GoogleBalanced14–3114Yes
Qwen Image 3AlibabaBalanced153Yes
GPT Image 2 (default)OpenAIHighest quality3–443Yes
Hunyuan Image 3.0 InstructTencentHighest quality18–723Yes
Seedream 5.0 ProByteDanceHighest quality22–3610Yes
Nano Banana ProGoogleHighest quality30–6014Yes

Prices checked in the app on 30 September 2026. The exact price for your settings is always on the model chip.

The rest of the Image tab gives you five controls:

  1. Looks. 30 ready-made styles in four families: Photography, 3D & render, Illustration and Graphic. With an empty prompt, a Look writes the whole prompt for you to edit. With your own words, its direction is added when you generate, at no extra cost.
  2. References. Add your product, a face or a style from your uploads, your past images or stock photos. Depending on the model, that's 1, 3, 4, 10 or up to 14 images. On most models, references don't change the price.
  3. @ mentions. With two or more references, type @ in the prompt and pick one. It drops into your sentence as a chip, so the model knows which picture is which.
  4. Edit mode. Switch Create to Edit to change one thing in a picture and keep the rest. 11 models can edit, and you can bring in up to 2 extra images, like a logo or a colour.
  5. Shape, quality and count. Eight aspect ratios, from 16:9 to 9:21. One to four images per press, where four cost four times one.

Which is the best AI video model to pick in 2026?

For a first try, use MiniMax H3 Max Turbo: a 5-second 480p clip costs 25 credits, and it can start from your own still. Use Seedance 2.5 for clips up to 30 seconds. Use Veo 3 Fast when you want a 1080p shot from words alone.

Studio has 10 video models from Google, ByteDance, MiniMax, Black Forest Labs and xAI. Nine do text to video, and eight do image to video from a start frame. If you are weighing Seedance vs Veo vs MiniMax, all three are in one list, with a price on every row.

The Studio video model list with CHEAPEST, needs a start frame, text only and 4 to 30 second labels

  1. MiniMax H3 Max Turbo is the cheapest and the default: 25 credits for 5 seconds at 480p.
  2. Grok Imagine needs a start frame. It only animates an image you give it.
  3. Veo 3 Fast works from words alone, at 1080p, for 4 to 8 seconds.
  4. Seedance 2.5 is the only model that goes past 15 seconds, up to 30.
ModelMakerResolutionLengthFramesCheapest clip (credits)Longest clip (credits)
MiniMax H3 Max Turbo (default)MiniMax480p–1080p5–15 sStart and end25 · 5 s 480p240 · 15 s 1080p
Seedance 2.0 MiniByteDance480p–720p4–15 sStart and end40 · 4 s 480p330 · 15 s 720p
MiniMax H3 MaxMiniMax480p–1080p5–15 sStart and end50 · 5 s 480p480 · 15 s 1080p
Grok ImaginexAI720p6–15 sStart frame required60 · 6 s150 · 15 s
Seedance 2.0 FastByteDance480p–720p4–15 sStart and end62 · 4 s 480p495 · 15 s 720p
Seedance 2.0ByteDance480p–1080p4–15 sStart and end76 · 4 s 480p1,530 · 15 s 1080p
Gemini OmniGoogle720p–1080p4–10 sNone90 · 4 s180 · 10 s
Seedance 2.5ByteDance480p–1080p4–30 sStart and end112 · 4 s 480p1,890 · 30 s 720p
Veo 3 FastGoogle1080p4–8 sText only120 · 4 s240 · 8 s
FLUX.3Black Forest Labs720p–1080p5–15 sStart only170 · 5 s 720p870 · 15 s 1080p

Cheapest-clip prices were read in the app on 30 September 2026, at 16:9. Longest-clip prices come from Studio's training guide, dated September 2026.

Five video settings to know:

  • Start and end frames. The clip opens on your start frame and, on MiniMax and Seedance, ends on your end frame. On MiniMax, frames didn't change the price, and the clip follows the frame's shape.
  • Audio. Seedance and FLUX.3 have an audio switch; Veo, Gemini Omni and Grok always make sound; MiniMax has none. Turning audio on or off never changes the price.
  • References and reference audio. Seedance and Gemini Omni take up to 3 reference images. Seedance can also follow a sound or one of your voiceovers.
  • Looks. 55 video Looks in three families: Commercials, UGC, and Graphics & painted.
  • Auto-switching. Switch to a model that can't do your setting, and Studio picks the nearest one and tells you what changed.

The Studio Video tab with a start frame and an end frame of the same mug, and a shot description

  1. The start frame: the glazed mug from an earlier image.
  2. The end frame: its matte-black edit.
  3. The shot describes what happens in between. This MiniMax run cost 25 credits, the same as the clip without frames.

How do you make an AI voiceover and clone your voice?

Paste a script on the Voiceover tab, pick a voice, and check the price under the box: credits per 1,000 characters, rounded up. Seven voice models run from 2 to 40 credits per 1,000 characters. Cloning your own voice is free and takes 10 to 20 seconds of speech.

The voice is who speaks; the text-to-speech model is the engine that speaks it. Every voice has a free preview, so you can hear it before you choose.

The Studio voice model list, from Inworld TTS-1.5 Max at 2 credits to ElevenLabs Multilingual v2 at 40

  1. Inworld TTS-1.5 Max is the cheapest: 2 credits per 1,000 characters.
  2. ElevenLabs Multilingual v2 is the default: 40 credits per 1,000 characters.
ModelMakerCredits per 1,000 charactersVoices and languagesMax per run
Inworld TTS-1.5 MaxInworld2113 voices, 15 languages2,000 characters
KokoroHexgrad420 voices, American English2,000 characters
Chatterbox HDResemble AI89 voices; takes 2–4 minutes2,000 characters
Cartesia SonicCartesia1044 languages, plus voices you make in Studio5,000 characters
ElevenLabs Flash v2.5ElevenLabs2020 voices plus your clones; 32 languages; the fastest5,000 characters
ElevenLabs Multilingual v2 (default)ElevenLabs4020 voices plus your clones; 29 languages5,000 characters
ElevenLabs Eleven v3ElevenLabs4020 voices plus your clones; 70+ languages; the expressive one5,000 characters

Our 129-character line on Multilingual v2 came to 129 × 40 ÷ 1,000 = 5.16, rounded up to 6 credits. The same line on Inworld would be 1 credit, the minimum.

A voiceover in Studio History: 129 characters on ElevenLabs Multilingual v2, charged 6 credits

  1. The row keeps the voice, the character count and the price: 6 credits.
  2. "Saved to Files · charged 6 credits". Every voiceover is saved as an mp3.

You can also make voices of your own:

  • Clone a voice, free. Voice cloning costs nothing. Record or upload 10 to 20 seconds of clear, lively speech from one person, as MP3, WAV, MP4 or WebM. Then name it, pick one of 15 languages, and confirm it's your voice or that you have the speaker's permission.
  • Design a voice. Describe the age, accent and mood, or start from a preset like Conversational Narrator. Each press makes three samples for 11 credits. Saving the one you like is free.
  • Keep up to three. Clones and designed voices share three slots. The bin beside a voice frees one.

Voices you make in Studio speak on Cartesia Sonic. Voices cloned in a Video Agent chat speak on the ElevenLabs models.

How do Studio pieces become a finished video?

Studio makes the pieces; the Video Agent puts them together. Make a still, use it as a clip's start frame, then press Use in a video on the clip. The agent opens with your clip attached, and nothing runs until you send the brief.

Diagram: make a still on the Image tab, animate it on the Video tab, press Use in a video, then send and approve in the Video Agent

There are three hand-offs:

  1. Image to clip. Make the key still on the Image tab, then pick it as the start frame on the Video tab. Your Studio images wait under Image generations in the library.
  2. Voice to clip. On Seedance models, a clip can follow one of your voiceovers as reference audio.
  3. Images and clips to the Video Agent. Every image and clip has a Use in a video button. It opens the agent's composer with the piece attached and a starter line you rewrite into your brief.

Use in a video: the Video Agent composer with a Studio clip attached and the line Make a video using this clip

  1. The Studio clip, attached.
  2. A starter line, "Make a video using this clip.", for you to rewrite.
  3. Nothing runs until you press send.

Voiceovers don't have this button, because the agent writes and voices its own script. From there, the agent plans the video and shows the cost on its approve button. It works the same way for long AI videos.

How does Studio compare with other all-in-one AI image, video and voice tools?

Showing the price first isn't unique: Higgsfield and Yapper do it too, and Runway and Higgsfield return credits after most failures. Studio prices each quality and resolution option, and never charges a failed run on any model. It also keeps voice models with free cloning in the same place.

Looking for a Higgsfield alternative, or comparing all-in-one tools? We only used each company's own pages, opened on 30 September 2026, and left out prices because they change too often.

PlatformModels (their own claim)Cost before you generateFailed generationsVoice
Synthesys Studio18 image, 10 video, 7 voiceOn the model chip and every quality or resolution optionNever charged; credits come off when a result lands7 voice models, free cloning, voice design
Higgsfield"30+ models and tools""Shown on the Generate button" (help centre)Credits usually returned automatically; some models exceptedAudio and lip-sync tools
Yapper"24+ Pro Image Models", "36+ Pro Video Models""Always visible before you start"Not stated on the pricing pageLip-sync and audio generation
Magnific (formerly Freepik)"30+ models"Credits per generation listed on the pricing pageNot stated on the pricing pageAudio and music generation
KreaAll image models on every plan; all video models on Pro and upNot stated on the pricing pageNot stated on the pricing pageLip-sync models
RunwayGen-4.5 plus partner modelsNot stated on its credits page"Automatically returned" after an errorNot checked

Where the others are stronger:

  • Higgsfield lists Soul ID "for consistency", plus audio and lip-sync tools for dubbing and localization.
  • Krea has real-time models that draw as you type or sketch.
  • Magnific comes with a stock library of over 250 million photos, videos, vectors and PSDs.
  • Yapper lists more models than Studio: 24+ image and 36+ video, against our 18 and 10.
  • Lip-sync and music. Studio has no lip-sync or music models of its own. Higgsfield, Yapper and Krea list lip-sync, and Magnific lists music. In Synthesys, the Video Agent adds voice, captions and music to finished videos.

If one model's special controls are your whole workflow, a specialist tool may suit you better, and that's fine. We built Studio for everyone else: image, video and voice in one place, priced up front, and a short step from a finished video.

How do you spend Studio credits wisely?

Draft on the cheapest model, then re-run the winner once on a top model with Reuse settings. Prove a clip at 5 seconds before paying for 15, and make four images only while exploring. Four drafts plus one final image cost 48 credits this way, against 176 for four tries on the top setting.

Diagram: drafting cheap and finishing on a top model or setting costs 48 instead of 176 credits for images, 155 instead of 240 for clips and 46 instead of 120 for voiceovers

HabitNumbers from the Studio price list
Draft cheap, finish premiumFour FLUX.1 [schnell] drafts plus one GPT Image 2 High image cost 48 credits. Four tries on High cost 176.
Draft clips at 480pThree 5-second MiniMax H3 Max Turbo drafts at 480p (75) plus one at 1080p (80) cost 155 credits. Three at 1080p cost 240.
Prove a clip shortMiniMax H3 Max Turbo at 480p costs 25 credits for 5 seconds and 75 for 15.
One still, many clipsA start frame costs a few credits once, and every clip you animate from it opens on the same picture.
×4 only while exploringFour GPT Image 2 Medium images cost 52 credits; one costs 13.
Watch the shapeGPT Image 2 at Medium costs 11 credits in 16:9 and 13 in 1:1.
Cheap voices for draftsA 1,000-character script costs 2 credits on Inworld and 40 on ElevenLabs Multilingual v2. Three Inworld takes plus one final ElevenLabs take cost 46 credits; three ElevenLabs takes cost 120.

For an image prompt, write subject, setting, light and lens, as in "A frosted amber bottle on wet slate, soft window light, 85mm." For a clip, write one shot and one move. For a voiceover, write for the ear, with short sentences and a comma wherever you want a breath.

What are the limits of Studio?

Studio makes single pieces, one generation at a time, and its model list and prices can change. Real names and brands are the most common reason a model refuses. Some models only take certain inputs, only Seedance 2.5 goes past 15 seconds, and there is no lip-sync, music or free trial.

  • One generation at a time. It's a limit we hit ourselves. In the Video Agent we can open many videos at once. In Studio, we wait for one generation to end before starting the next, and we're working on it.
  • Prices and models move. The prices above were read in the app on 30 September 2026, apart from the longest-clip column. The chip is the only price that counts.
  • Real people and brands are often refused. Naming a real, identifiable person, a brand or a copyrighted character is the most common reason a model says no. Describe the person instead: "a man in his 40s, salt-and-pepper beard". A refused run costs nothing.
  • Not every model takes every input. Veo 3 Fast is text only, 1080p and 16:9. Grok Imagine needs a start frame, and Gemini Omni takes no frames. On Seedance, reference audio can't be used with a frame, or an end frame with reference images.
  • Long clips are Seedance 2.5 only. Every other model stops at 15 seconds or less, and Seedance clips over 15 seconds top out at 720p.
  • Three voices of your own. Clones and designed voices share three slots, and Studio clones speak only on Cartesia Sonic.
  • Voiceovers can't go straight to the agent. Only images and clips have Use in a video.
  • No speed or emotion sliders. You pick the model and voice that fit the mood. Eleven v3 is the expressive one.
  • No free trial. The cheapest way to try Studio is a 1-credit FLUX.1 [schnell] image or a short Inworld voiceover.

Try it

Studio is included on every Synthesys plan, next to the Video Agent, and every plan comes with full commercial rights. Open Studio from the sidebar, write a prompt, and check the price on the chip before you press Generate. Open Synthesys.

New to the agent side? Start with what an AI video agent is, then read our guide to AI video credits.

Sources

All opened on 30 September 2026.

Frequently asked questions

What is Synthesys Studio?
Studio is a part of the Synthesys app, opened from the sidebar, for making single images, video clips and voiceovers with the AI model you choose. It has 18 image, 10 video and 7 voice models, and it shows the credit price before you press Generate.
How many AI models are in Synthesys Studio?
35 as of 30 September 2026: 18 image models from 12 makers, 10 video models and 7 voiceover models. The list can change, so the model picker in the app is always the current source.
Am I charged if a Studio generation fails?
No. Studio takes credits only when a result lands in your History. A failed, refused or stopped run costs nothing, and its card says No credits charged.
Can I clone my voice in Synthesys Studio?
Yes, and cloning is free. Record or upload 10 to 20 seconds of clear speech, confirm consent, and the new voice speaks on Cartesia Sonic. You can keep up to three voices of your own.
How long does a Studio clip take?
Clips usually take 1 to 5 minutes and images usually under a minute. Voiceovers take seconds, except Chatterbox HD at 2 to 4 minutes. You can close the tab while a clip renders; it still lands in History and Files.
What is the difference between Studio and the Synthesys Video Agent?
The Video Agent plans and builds a whole video from a brief. Studio makes one image, clip or voiceover at a time with the model you pick. Use in a video sends any Studio image or clip into the agent.

Keep reading

All posts →