GPU capacity.
By the second.
Bare-metal and virtual GPU nodes for training, fine-tuning and inference. Per-second billing from boot to terminate — no minimums, no contracts, no waitlists for available SKUs.
The whole stack, ready to run.
Nodes ship fully configured. No driver hunting, no image wrestling. Pick a SKU, boot an instance, and your first CUDA kernel runs in minutes.
Per-second billing
Pay from the first boot instruction to the last terminate ACK. No rounding to hours, no daily minimums, no idle fees.
Dedicated SSH + root
Every node gives you a dedicated SSH keypair and full root. No container sandboxing between you and the metal.
CUDA 12 preinstalled
Drivers, toolkit and cuDNN land on first boot. PyTorch, JAX and TensorFlow all see the GPU before your first command.
ML-ready base images
Choose from PyTorch 2.x, vLLM, Unsloth or Axolotl. Each image is GPU-validated and pinned to a tested version.
Snapshot + resume
Snapshot instance state at any point. Terminate, stop paying, and resume from the same checkpoint later.
Console + REST API
Spin nodes up, down and sideways from the web console or the REST API. Automate provisioning into training pipelines.
Multiple SKUs
A100-class SXM and PCIe nodes, RTX 4090 for memory-bound workloads. Mix SKUs across experiments as fleet comes online.
No cold-start penalty
Available SKUs are pre-allocated. When you boot, the instance is ready — no queue position, no warm-up wait.
Inference handoff
Train or fine-tune on GPU, then deploy directly to a saqi inference endpoint. Same account, same console, no S3 gymnastics.
Boot straight into your framework.
Every image is pre-pulled, GPU-validated and pinned to a tested version. Pick one and your first training command works.
PyTorch 2.x
- ▸torch + torchvision + torchaudio
- ▸CUDA 12 + cuDNN 9
- ▸Jupyter Lab preinstalled
- ▸HuggingFace transformers
vLLM
- ▸vLLM latest stable
- ▸OpenAI-compatible server
- ▸FlashAttention-2
- ▸Llama / Mistral / Qwen weights on-hand
Unsloth
- ▸2× faster than HF baseline
- ▸Llama, Qwen, Mistral adapters
- ▸4-bit QLoRA out of the box
- ▸Direct GGUF export
Axolotl
- ▸YAML-driven config
- ▸DPO / ORPO / RLHF support
- ▸DeepSpeed + FSDP ready
- ▸Weights & Biases auto-log
Workloads we're targeting
- ▸LoRA and full fine-tunes that run for hours, not months
- ▸Batch inference overnight at scale-to-zero cost
- ▸Interactive ML notebooks on a real GPU with no SaaS overhead
- ▸vLLM benchmarking before committing to an inference endpoint
- ▸Multi-GPU distributed training across topology-aware nodes
- ▸One-off experiments that don't justify a reserved cloud contract
Train here. Serve here.
GPU capacity and Inference endpoints live on the same platform. Fine-tune a model on a GPU node, push the adapter to an inference endpoint, and serve it — no S3 buckets, no registry gymnastics, no cloud glue code.
Where we're going
-
in beta now
On-demand instances
Start, stop and snapshot via console or API. Limited SKUs as fleet comes online.
-
next
Reserved capacity
Commit to a GPU block for weeks or months at a lower hourly rate.
-
planned
Cluster orchestration
Multi-GPU distributed jobs scheduled across nodes with topology-aware placement.
GPU beta is invite-only for now.
Tell us your workload — training run size, GPU-hours per week, frameworks you use — and we'll prioritize you in the queue.
Request GPU beta access →Ship an agent on your own infrastructure.
Hive is live today. Agents, Inference, and GPU are in private beta — request access and we'll reach out.