Cubric Studio - Remote

Remote

Everything that runs on somebody else's computer: the language models that write prompts, describe images, and power the agent, and the RunPod remote engine that generates on a rented cloud GPU.

Open Remote from the top navigation, right beside Settings. It has two sections, in this order: Language models, then the RunPod remote engine.

Language models

Cubric uses a language model for the writing jobs around a generation: rewriting a short idea into a full prompt, describing an image you hand it, and the agent. Each job below chooses which machine runs it, and you always see the result before it is used.

Remote connection

One connection serves every job below that runs on Remote. Pick a provider, add its API key, then test it. The key is stored by the desktop app and cannot be read back.

  • Provider. DeepInfra (recommended, and the one Cubric Studio tests on), OpenRouter, OpenAI, your local Ollama through its /v1 endpoint, or Custom for any other OpenAI-compatible address (its base URL appears once you pick Custom).
  • API key. Paste it and press Save; Clear removes it. Ollama needs no key.
  • Test connection. Confirms the key works and reports how many models answered, without spending on a job.
  • Spend this month. DeepInfra only. Reads what you have spent, your credit left, and your monthly limit, on request.

New to DeepInfra? Create an account, make an API key in your DeepInfra dashboard, then pick DeepInfra as the provider here and paste the key. What you run is billed to your DeepInfra account.

The Remote connection group of the Remote panel: a New to DeepInfra? box with an Open DeepInfra dashboard button, Provider set to DeepInfra, the API key field with Save and Clear and the line API key is saved., a Test connection button, and Spend this month with a Check spend button
Remote connection with a DeepInfra key saved.

Where each job runs

This is about which machine does the work, not which answer is better. Generating on a RunPod Pod? Run these locally, your own card is idle. Generating locally? Run them on Remote and keep the VRAM for the picture.

  • Prompt enhancement. Remote (off your machine, no VRAM cost), Ollama (local) (its own runtime and its own VRAM), or ComfyUI (local) (reuses the engine already running; the default). The model to use appears under whichever backend you pick.
  • Image descriptions. Remote or ComfyUI (local). Remote costs no VRAM and never waits behind a generation; a hosted model may refuse to describe an adult image, and ComfyUI's built-in describer is uncensored.
  • Agent. Always runs on Remote. Pick its model and its mode here, Auto (picks settings and generates without asking) or Ask first (asks about every setting first; installs always ask either way). See Testing an agent model below for the two test buttons under it. Connect the agent to an app in Settings > Connect an agent, or see Agent for what it can do.
The Where each job runs group of the Remote panel: Prompt enhancement and Image descriptions both set to Remote with their recommended models picked, and the Agent row on Remote with its model, Agent mode set to Auto, and Test tool use and Benchmark this model buttons
Where each job runs, with every job on Remote.

Testing an agent model

Not every language model can drive the agent well. The Agent model list puts the models that have run the agent's tests on top, each with its score, what 100 chats cost on it, and how much context it holds, for example 28/28 tests · $0.36/100 chats · 1M context. Scores on the current tests rank first, then the best pass rate, then the most runs.

The agent is updated often, and new tests are added as it learns new things, so the number of tests grows from one update to the next. A score from an earlier set of tests is marked (older tests), so you can tell it apart from one on the tests your app ships with.

The Agent model list open: deepseek-ai/DeepSeek-V4-Flash-0731 at 28/28 tests, $0.36 per 100 chats, 1M context; Qwen/Qwen3.6-35B-A3B at 25/28 tests, $0.71 per 100 chats, 256K context; openai/gpt-oss-120b at 19/28 tests, $0.19 per 100 chats, 128K context; then google/gemma-3-12b-it
Tested models sit on top of the Agent model list, with score, cost per 100 chats, and context.
  • Test tool use. One small request that checks the picked model can call tools at all, before you rely on it.
  • Benchmark this model. Runs the agent's own tests on the picked model with pretend tools, so nothing is generated or added to your projects. It first shows an estimate with Run and Cancel: on a cloud provider such as DeepInfra, a price and about how many minutes (roughly half a minute a test); for a model on your own GPU, Uses your GPU for 10-30 min. On a cloud provider it does spend real calls on your key. Stop ends it after the current test, and a stopped run is not kept, nor is a run whose connection failed partway. A finished run puts the model in the scored group with your own result, which counts ahead of Cubric Studio's. You do not have to watch it: close the panel and carry on, and a toast tells you when the run ends, with the score and what it cost. A cloud model uses none of your GPU, so you can keep generating while it runs. A model on your own GPU, such as one through Ollama, will not start a benchmark while a generation is running on the card; start it once that generation finishes.

On DeepInfra, OpenRouter, OpenAI or Ollama (not Custom), the confirm step also offers Share the result anonymously. It is unticked the first time and remembered after that. A shared run sends the model, whether it passed each test, its cost per chat, how long a test took and your app version, plus your GPU's name and memory when the model runs on Ollama. Never your prompts, the model's replies, or your keys. It goes to the community scores at bench.cubric.studio, and See everyone's results opens that page. Only a whole run is shared: not a stopped run, and not one whose connection failed. When a run you asked to share did not go, its result line says why. Once a model has at least three shared runs, the list can show the community's median score, marked (community, N runs), for models neither you nor Cubric Studio has tested on the current tests.

The Agent row on Ollama (local) with ornith:9b picked, mid-confirm: Run and Cancel buttons with the estimate Uses your GPU for 10-30 min (34 tests), Share the result anonymously ticked, with the note that it sends the model, its scores, cost and your GPU but never prompts or keys, and a See everyone's results button
Benchmark this model, at its confirm step: the estimate, Run or Cancel, and the share option.

Share the poor results too. A model that fails half the tests is exactly what the next person wants to know before they spend the time, and on a paid provider the money, finding out for themselves. The run below is one: a small local model that passed 14 of 34.

The Agent row after benchmarking granite4.1:8b on Ollama: a bar of 34 marks, green for each test passed and red for each failed, and the line 14/34 passed, now shown in the agent list, shared
A finished run: green for each test passed, red for each failed. Low, and worth sharing.

See everyone's results opens the community page. It has one table per provider. Each model shows its score on the current tests, how many runs were shared, its cost per chat, and how long a test took; for Ollama, also the GPUs it was run on. A model with three shared runs is marked counted: that is when its score can appear in everyone's Agent model list. Below three, it says how many more runs it needs. Counted models come first, then the best pass rate. Under each table is a list of the tests that provider's models fail most.

The community results page, Ollama (local) section: ornith:9b at 23/34, gemma4:12b at 18/34 and granite4.1:8b at 14/34, each run once on an RTX 4060 Ti (16 GB), free, with its time per test and needs 2 more runs; below, Tests Ollama (local) models fail most, led by outpaint-grows-one-side, second-picture-character, second-picture-room and sheet-goes-to-reference, each failed in 3 of 3 runs
Local models, with the GPU each ran on. None has three runs yet.
The community results page, DeepInfra section: DeepSeek-V4-Flash-0731 at 33/34 over 3 runs, counted, $0.0035 per chat; Qwen3.6-35B-A3B and gpt-oss-120b at 31/34 with 1 run each and needs 2 more runs; below, Tests DeepInfra models fail most, led by options-ideas in 2 of 5 runs
Cloud models: DeepSeek has three runs, so its score counts.

RunPod remote engine

The RunPod remote engine lets you run generation on a rented cloud GPU instead of your own hardware. It is optional. Cubric's remote engine runs on your own RunPod account: generation stays on your local engine until you Connect, and GPU and storage billing happen on your RunPod account, not Cubric's.

New to RunPod? The panel's Create RunPod account button opens sign-up through Cubric's referral link, which can include a first-time credit bonus. To unlock the rest of the controls, save a RunPod API key with read and write access.

RunPod can also replace the local engine entirely. On first launch, Remote only skips the ComfyUI install and Cubric runs on a cloud GPU instead, see Getting Started. Skip the local engine install, below, is where you turn the local engine back on if you change your mind.

The top of the RunPod Remote Engine section: a New to RunPod card with a Create RunPod account button, the Account group with a saved RunPod API key and Save and Clear buttons, the Automatically connect on app start, Skip the local engine install and Stage all models on connect switches, and Storage with the EU-RO-1 data center and no network volume yet, beside a + Create button
Saving a RunPod API key reveals the remote-engine controls.
  • Automatically connect on app start. When on, the app connects (and starts billing) a Pod at launch. Off by default: start local, and Connect when you want cloud generation.
  • Skip the local engine install. For cloud-only machines. When on, startup never asks you to install the local ComfyUI engine, and you generate on a Pod instead. Choosing Remote only on first launch turns it on; turn it off to install the local engine on the next launch. It also clears itself automatically once a local engine is actually present, so an update can no longer leave you skipping an engine you already have.
  • Stage all models on connect. On by default: every installed model is copied to the Pod's fast disk as soon as it connects, so the first generation is instant. Turn it off to copy each model on first use instead, only what you actually generate with.

Storage

  • Data Center. A network volume is locked to one data center. Switching later means deleting the volume and re-downloading models. Any region (no volume) skips the volume: models download again for each session, and nothing bills for storage between sessions.
  • Network Volume. Stores your models so they survive between Pods, one volume per data center; + Create makes one. Stopped Pods keep billing volume storage until you delete it.

Machine

  • GPU. Your pick in one line: the card, its VRAM, its price per hour, and its stock. Choose GPU opens the GPU picker, below. Secure Cloud only. Stock is a live hint that drifts; the RunPod console is the source of truth.
  • Connect. Starts the remote engine. While it is greyed out, the line under it says why: no GPU chosen yet, or No GPU, download only picked in a data center with no network volume, where its downloads would have nowhere to go. Any real GPU still connects without a volume: Connect first asks, then makes a temporary Pod whose models download each session and are deleted when you stop or delete it. A Pod keeps the software it was made with, so the first Connect after an update that changes it (2.0 does) replaces your saved Pod with a fresh one: the models on your network volume stay, and only that Connect takes longer. Open in RunPod console checks Pod state, telemetry, logs, and spend on RunPod.
  • Delete Pod on quit. When on, quitting the app deletes the Pod instead of keeping it warm. This frees GPU and container disk fully; your network volume and models are kept.
The Machine group: GPU reading RTX 2000 Ada, 16 GB VRAM, $0.24/hr, out of stock, with a Choose GPU button; then Remote engine: stopped with a Connect button and an Open in RunPod console link; then the Delete Pod on quit switch
The GPU pick in one line, Choose GPU to change it, then Connect.

Choosing a GPU

Choose GPU opens a full-window picker with one tile per Secure Cloud card, with its stock in the data center you picked. Each tile shows the price per hour, the VRAM, the most GPUs one Pod can take, a three-bar stock meter, and a green Gen speed bar: how fast the card makes an image, from RunPod's own measured times on each card (FLUX.2 Klein 9B, September 2026), with the fastest card's bar full. Beside the bar, the measured time itself, in seconds per image (lower is faster), tells close cards apart. A card RunPod has not benchmarked has no bar. The times come from an image model, so the bar is a guide to image speed. After No GPU, tiles run fastest first; cards with no measurement come after every measured one, and ties go to the card with less VRAM, then the cheaper one. Click a tile to pick it; the picker closes. Stock is read when the picker opens and when you press its refresh button, never on its own, and the line at the top says when it was last checked.

  • Auto-retry. Off, the picker shows only cards in stock. On, it shows every card, an out-of-stock one reads Out of stock · Connect waits, and Connect keeps checking in the background until the card you picked frees up, then connects; you can keep working locally while it waits. On by default for a new install since 2.0, because RunPod stock is thin enough that a card shown in stock is often gone by the time you press Connect; a setup from before 2.0 keeps whatever you had it set to.
  • Video (over 24 GB VRAM). Shows only cards with more than 24 GB of VRAM.
  • Min RAM. RunPod only places a GPU Pod on a host with at least this much system RAM; 0 means any host. Since 2.0 a new setup starts at 0 (1.5.0 started at 56 GB), because the hosts behind most smaller cards have less than that, and a floor turned Connect into a no-host message. A value saved on an earlier version is kept as it was, so check it if you set up RunPod before 2.0. Raise it only when you know your model needs a big host: heavy video models, such as MiniMax H3 at 768p or LTX at 2K, run better with more, because ComfyUI offloads weights to system RAM. A higher floor means fewer hosts: when no host with that much is available, Connect says so, and you can lower Min RAM, turn on Auto-retry to wait, or pick another data center. A wait for a GPU, whether from Auto-retry or from a Pod that connects at app start, keeps this floor too. Hidden for Any region, where RunPod has no data center to place against.
  • No GPU, download only. Always the first tile: a CPU-only Pod with no GPU billing, for installing models onto your volume before you pick a card to generate with. If its CPU size is sold out, Cubric Studio tries the other CPU sizes instead of failing outright. Not offered for Any region, which has no volume to download onto.

The card you picked always shows, even when a filter would hide it or it has just gone out of stock.

The GPU picker for EU-RO-1, Secure Cloud, live stock, checked 10:02:42, with a Refresh button, Auto-retry on, Video (over 24 GB VRAM) off and Min RAM 0 GB, and a note explaining Auto-retry, Min RAM and Gen speed. Tiles, fastest first after No GPU, download only: H100 SXM, RTX PRO 6000, B300, RTX 5090 and PRO 6000 MIG 48GB with nearly full green Gen speed bars, down to RTX 2000 Ada with a short one; RTX PRO 4500, picked, $0.72/hr, 32 GB VRAM, low stock; RTX A4500 last, with no bar
Choose GPU, fastest card first: the green Gen speed bar is RunPod's measured image time, and the last card has no bar because RunPod has not benchmarked it.