Pick by the card and the VRAM, not by a datacentre. Every tier shows its hourly rate, its reserved monthly price, whether it is a full VM or a container pod, and what a pause costs.
Run it by the hour while you're figuring things out. Reserve the same machine monthly once it's steady and pay about 15% less than 730 hours would cost. No commitment to switch.
Stop a machine and you keep the disk but release the card, so you pay a small storage-only rate instead of the full hourly price. Each plan shows that rate before you pause — a stopped machine is cheaper, never free.
Pick a template at checkout and the machine boots with it: CUDA and Docker, JupyterLab with PyTorch, a private chat UI, an OpenAI-compatible endpoint, ComfyUI, or a fine-tuning kit. Full root on the VM tiers.
A hosted AI API sees every prompt, every document and every line of code you send it. A GPU server of your own sees only what you put on it — the weights, the data and the logs are yours, in a region you picked, with nothing phoning home.
Why teams self-hostThe open-weight models are good enough to serve in production, and on your own card they cost you hours instead of tokens.
Hold the weights in VRAM and answer requests, or adapt an open model to your own data on a datacentre card.
Diffusion pipelines and ComfyUI at full resolution, on cards sized for memory bandwidth rather than parameter count.
Architectural visualisation, animation and effects work, with the render billed by the hour it actually ran.
Hardware encode at scale, genomics, physics and quantitative finance — anything that wants parallel throughput.