Cloud Servers

GPU and AI Compute Servers

A Kapsule GPU server is a single-tenant physical machine fitted with NVIDIA graphics processors and the CPU, memory and fast storage needed to keep them fed, built for AI training, inference and other heavily parallel work.

You get the whole machine. Nobody else's job is sharing your GPU, your memory bandwidth or your disk, which is the difference between a benchmark you can reproduce and one you cannot. This guide covers the four tiers, what they suit, how ordering works given that GPU supply is genuinely tight, and what the setup fee is for.

The Four Tiers

TierGPUGPU memorySystem memoryMonthlyOne-time setup
SparkNVIDIA RTX 400020 GBIncluded in the buildUS$330.00Yes
ForgeNVIDIA RTX PRO 600096 GB256 GBUS$1,680.00Yes
FoundryNVIDIA RTX PRO 600096 GB512 GBUS$2,520.00Yes
ReactorNVIDIA RTX PRO 600096 GB768 GBUS$3,240.00Yes

Prices are shown excluding GST in your billing currency. Each tier comes unmanaged or managed, the same choice as our cloud servers.

The jump from Spark to Forge is the significant one. Spark is an entry point for development, smaller models and inference work. Forge, Foundry and Reactor share the same 96 GB of GPU memory and differ in system memory, which is what governs how much you can hold in host RAM alongside the model: dataset staging, preprocessing pipelines, and serving many concurrent requests.

What You Can Run

You get full control of the machine, so you run your own stack. That includes PyTorch, TensorFlow, JAX, raw CUDA workloads, and the usual training and serving frameworks on top of them.

Typical uses:

  • Model training and fine-tuning for machine learning and deep learning.
  • Inference and serving in production, where a GPU turns an unacceptable response time into an acceptable one.
  • Batch AI workloads: embeddings, transcription, image and video generation, rendering.
  • Scientific and parallel compute that benefits from GPU acceleration.

Size on GPU memory first. The most common mistake is choosing a tier on price and then discovering the model does not fit in VRAM. Work out the memory your model needs at your intended precision and batch size, add headroom, and choose the tier that clears it. System memory and CPU are much easier to work around than a model that will not load.

Availability and How Ordering Works

GPU hardware is in high demand across the whole industry and supply is limited. That is a fact about the market, not a Kapsule policy, and it shapes how these are ordered.

  • These are physical machines, so they are provisioned to order rather than in seconds.
  • Availability varies by tier and location, and it is checked live when you look at the plans page.
  • Spark is currently supply-constrained. When it has no stock, the panel offers you the reserve queue rather than a broken order.

For larger or custom configurations we confirm the exact build and the exact price with you before provisioning, so you always know what you are getting. We do not quote a configuration we cannot deliver.

The Reserve Queue

When a GPU tier is out of stock, you can reserve it instead of walking away:

  1. Reserve the tier from the plans page. You get a reservation reference.
  2. Nothing is charged. A reservation is free and no subscription starts.
  3. When capacity appears, we provision the machine for you automatically.
  4. Your first payment is taken only once the server is live, and your billing date starts from that day.
  5. Cancel at any time while you wait, at no cost.

Regions, Availability and the Reserve Queue explains the mechanics in full.

Do not put a GPU reservation on a hard deadline. Capacity returns when it returns, and we would rather tell you that plainly than promise a date we cannot hold. If you have a fixed delivery date, talk to us before you order.

Ordering One

  1. Sign in to KPanel.
  2. Open Store in the left sidebar and choose GPU, or go directly to /plans/compute?group=gpu.
  3. Use the Unmanaged and Managed toggle to choose the form you want.
  4. Each tier shows its GPU, its VRAM, its system memory, its monthly price and its one-time setup fee, plus a live stock badge.

The GPU group on the compute plans ladder in KPanel

The Setup Fee

A GPU server carries a one-time setup fee at checkout in addition to the recurring monthly price, because a physical machine has to be built and provisioned for you.

  • Charged once, at provisioning. It never appears on a renewal.
  • Shown as its own line at checkout, labelled Setup (one-time), with its own receipt.
  • Subject to GST when your account currency is NZD.

See Cloud Server Pricing and GST for how one-off and recurring charges appear on your invoices.

How GPU Servers Differ From Cloud Servers

Cloud ServersGPU Servers
HardwareVirtual machinesSingle-tenant physical machines with NVIDIA GPUs
ProvisioningImmediate, subject to stockTo order, sometimes via the reserve queue
Setup feeNoneOne-time
Live resize in KPanelYesNo
Sized onvCPU, memory, diskGPU memory first, then system memory

GPU servers cannot be resized in place. Moving between tiers means ordering the new machine and migrating your work to it. Choose the tier that fits the models you expect to run, not just the one you are running this week.

Before You Order

Have these answers ready, because they determine the tier:

  • Training or inference? Training is usually memory-hungry and long-running. Inference is usually latency-sensitive and concurrent.
  • What model, at what precision? This sets the GPU memory floor.
  • How much data, and where does it live? Staging a large dataset is what system memory and disk are for.
  • What framework? So we can confirm driver and CUDA expectations up front.

Tell us those four things and we will point you at the right tier, or put together a configuration if none of the standard four fits. Ask Kora inside KPanel, or email support@kapsulehost.com, and our New Zealand team will help.

Troubleshooting

The tier I want shows Out of stock. Reserve it. You pay nothing while you wait and we provision automatically when capacity returns.

My model will not fit in VRAM. Move up a tier, or reduce precision or batch size. Forge, Foundry and Reactor all carry 96 GB of GPU memory.

I need more than one GPU, or a configuration you do not list. Contact us with your requirements and we will confirm what we can build and what it costs before anything is charged.

Can I try before committing? Talk to us about starting on Spark for development work and moving up once your pipeline is proven.

Still need help?

Email us at support@kapsulehost.com or open a chat in KPanel.

Open KPanel
GPU and AI Compute Servers