Cubric Studio - Cloud Models

Cloud Models

Run paid cloud models such as Seedream, FLUX 2, Nano Banana, Seedance, Wan, and Veo on your own DeepInfra key, with no download and no GPU, and see the cost before you Cue.

What cloud models are

A cloud model has no weights to install and no graph to run locally: it generates on DeepInfra's own servers, on your own DeepInfra account, and the picture or clip comes back over the network. There is nothing to download, nothing on your GPU, and nothing here checks your hardware before you run one. You pay DeepInfra directly for what you use; Cubric Studio stores only the result.

Getting a DeepInfra key

Open Remote, next to Settings at the top of the Projects screen. Under Language Models > Remote connection (see Remote), a New to DeepInfra? box links straight to the DeepInfra dashboard: create an account there and make an API key. Back in Cubric Studio, set Provider to DeepInfra, paste the key into the API key field, and click Save. Test connection confirms it without spending anything, and Check spend reads back what you have spent this month, straight from your DeepInfra account.

The key is stored by the desktop app and cannot be read back once saved. Cloud models will not run without it: trying one with no key saved points you back to this same spot.

The Remote connection group of the Remote panel: a New to DeepInfra? box with an Open DeepInfra dashboard button, Provider set to DeepInfra, the API key field with Save and Clear and the line API key is saved., a Test connection button, and Spend this month with a Check spend button
Remote > Language Models > Remote connection, with a DeepInfra key saved.

The models

Fifteen cloud models, in six families. Every one offers Text to Image or Text to Video, and, on the image families, Edit as well. "References" is how many pictures you can hand an Edit at once, where the model takes more than one. Reference to Video makes a clip from reference pictures, clips and sound you add, instead of from a first frame.

Seedream

ByteDance's photographic line, strong at realistic scenes and at getting readable text right inside the image.

ModelOperationsReferencesNotes
Seedream 4T2I, Edit12K or 4K output.
Seedream 4.5T2I, Edit1Newer checkpoint than Seedream 4: better prompt following, cleaner faces.
Seedream 5.0 ProT2I, EditUp to 4Top of the line. 1.5K or 2K output; each reference beyond the first is billed separately.

FLUX 2

Black Forest Labs' FLUX 2, hosted at three tiers of sharpness and prompt adherence, plus the FLUX.2 Klein 9B you can also install locally.

ModelOperationsReferencesNotes
FLUX 2 DevT2I, EditUp to 4The open FLUX 2 weights, hosted instead of downloaded.
FLUX 2 ProT2I, EditUp to 4Sharper and more literal than Dev. Each reference is billed separately.
FLUX 2 MaxT2I, EditUp to 8The best prompt adherence of the three, and the dearest. Each reference is billed separately.
FLUX.2 Klein 9B (Cloud)T2I, EditUp to 4The same model as the local FLUX.2 Klein 9B, with no GPU and no download, so a prompt that works locally carries over. No styles and no masked edits. The price rises with the picture's size. It can also render Scribble and Draw It In.

Nano Banana

Google's Nano Banana / Gemini image family. Output is about 1 megapixel whatever ratio you pick, since these models take only an aspect-ratio label, not pixel dimensions. Up to four reference images on Edit are collaged into the one picture the model actually takes.

ModelOperationsReferencesNotes
Nano Banana 2 LiteT2I, EditUp to 4The cheap end of the family; good at edits that leave the rest of the picture alone.
Nano Banana 2T2I, EditUp to 4Better at faces and fine detail than Lite. Its content filter is stricter than either sibling's; if it refuses, try Lite or Pro instead, and a refusal is never billed.
Nano Banana ProT2I, EditUp to 4The strongest of the three at text in the image and at following a long prompt exactly. Google also calls this model Gemini 3 Pro Image.

Seedance

ByteDance's hosted video line.

ModelOperationsNotes
Seedance 1.5 ProT2V, I2V4 to 12 second clips at up to 1080p, with real camera movement.
Seedance 2.0T2V, I2V, Reference to VideoUp to 15 seconds, a clear step up in motion and coherence, and much dearer than 1.5 Pro.

Wan

ModelOperationsNotes
Wan 3.0T2V, I2V, Reference to VideoClips up to 30 seconds at 480p, 720p, or 1080p. Priced per second and per tier, so dropping a resolution tier meaningfully lowers the cost.

Veo

Google's Veo 3.1, the only cloud models here that generate with sound, and the only ones that can queue several clips as one batch and one bill.

ModelOperationsNotes
Veo 3.1 FastT2V, I2VFixed 8 second clips, with sound. The cheaper of the two, and the one to explore with.
Veo 3.1T2V, I2VFixed 8 second clips, with sound. The strongest video model in this list, and the dearest.

Cost, shown before you Cue

Pick a cloud model and the Cue button in the Prompt Box shows an approximate price beside it, worked out live from your current settings (size, resolution, duration, reference count). The same figure appears on the model's tile in the Model Library and in the agent's confirm card, so wherever you see a price it comes from the same estimate. When a shape genuinely cannot be priced ahead of time, the tag stays blank rather than showing a guess.

Prices are DeepInfra's own published rates for that model, kept in sync by the app rather than typed in by hand, so a rate change on DeepInfra's side is a small update here rather than a stale number.

Seedance 2.0 picked in the Prompt Box with its settings open (720p, 5 seconds, 16:9, one clip) and the Cue button reading About $0.84
The price beside Cue follows the settings: here Seedance 2.0 at 720p for 5 seconds.

A stopped run still costs

Once you press Cue and the request has gone out to DeepInfra, clicking Stop does not get your money back: the generation keeps running on DeepInfra's side, and its result still lands in your gallery when it finishes, because you already paid for it. Stop only cancels for free in the short window before the request is actually sent.

Local models and cloud models

The Run locally toggle in the Prompt Box only appears once you are connected to a RunPod cloud engine, and it only affects an ordinary installed model: it switches that model between running on your own machine and running on the RunPod Pod. A cloud model such as Seedream or Veo is not installed anywhere and always runs on DeepInfra, whatever that toggle is set to.

In the model picker and Library

In the model picker, cloud models sit alongside your installed ones, under Image or Video like any other model, marked with a cloud badge. The Model Library gives them a section of their own, DeepInfra models, at the foot of the page after everything local. The media and tier filters and Fits my GPU leave that section alone, since nothing in it runs on your card; search still reaches it. In place of a size tier and a download button, a cloud tile shows its price. The Library's own count line counts local models only, so adding a cloud model moves neither number; the model picker's count line is the one that splits them, for example "12 installed · 14 cloud". A cloud tile carries no LoRA or upscale-model settings, since nothing about it is installed.

The model picker with its count line reading 7 installed, 15 cloud, and the Image section mixing installed models with cloud ones such as Seedream 4, FLUX 2 Dev and Nano Banana 2, each cloud tile marked with a cloud badge and a price
In the model picker, cloud models sit beside your installed ones, and the count line splits the two.
The DeepInfra models section of the Model Library: Seedream, FLUX 2, Nano Banana, Seedance, Wan 3.0 and Veo 3.1 tiles, each marked Cloud with its billing unit, such as per image or per clip, and an approximate price such as about $0.04
The DeepInfra models section, at the foot of the Model Library. Each tile shows its price instead of a download size.