Dedicated
inference endpoints.
OpenAI-compatible API endpoints for open-weight and fine-tuned models. Token billing, no cold starts, no shared-tenant noise. Drop-in replacement for any OpenAI SDK call — change the URL, change the key, done.
Production inference, not a playground.
Endpoints built for latency-sensitive workloads. Your model, your keys, your traffic — no other tenants sharing your hardware.
OpenAI-compatible API
Drop /v1/chat/completions and /v1/embeddings into any OpenAI SDK call. Change the base URL and key — your code ships unchanged.
Dedicated endpoints
Your endpoint runs on its own hardware slice. No shared-tenant contention, no noisy-neighbour latency spikes at peak hours.
No cold starts
Dedicated endpoints are always warm. No 30-second waits for the container to wake up on the first request of the hour.
Token-based billing
Pay per input and output token. No hourly minimums. Idle endpoints cost nothing until a request lands.
Scoped API keys
Issue per-project keys with their own token quotas, rate limits and expiry. Rotate or revoke without touching other projects.
Full streaming support
Server-sent events with delta chunks — identical to OpenAI streaming. No polling, no timeout fighting on long responses.
BYO fine-tune
Upload a LoRA adapter or a full fine-tune as a container and get a dedicated endpoint. Your weights stay on your endpoint.
Usage analytics
Per-request and per-token breakdowns with latency percentiles. Know exactly what each project costs before the invoice arrives.
GPU handoff
Fine-tune on a saqi GPU node, push the adapter directly to an inference endpoint — same console, same account, no pipeline glue.
Open-weight models, ready now.
The catalog ships with tested, quantized versions of the best open-weight models. Every endpoint is OpenAI-compatible. Bring your own fine-tune and it joins the same API surface.
Llama 3.3
Meta's flagship open-weight. Strong on instruction-following, reasoning and code.
Llama 3.1
Efficient 8B for latency-sensitive apps; 70B for complex multi-step reasoning.
Qwen 2.5
Best-in-class multilingual coverage and coding performance across the open-weight field.
Mistral 7B
Low-latency workhorse. Excellent price-per-token for high-volume applications.
Phi-3.5
Microsoft's compact model — strong reasoning in a package suited for constrained deployments.
BYO model
Upload a LoRA adapter or full fine-tune container. Get a dedicated endpoint immediately.
One line to swap.
The endpoint speaks the OpenAI protocol. Change base_url and api_key. Every SDK, library and middleware that works with OpenAI works here.
- ▸OpenAI Python + Node SDKs
- ▸LangChain, LlamaIndex, DSPy
- ▸LiteLLM and any proxy router
- ▸Continue.dev, Cursor, Copilot alternatives
from openai import OpenAI
client = OpenAI(
base_url="https://api.saqi.ca/v1",
api_key="sk-proj-...",
)
response = client.chat.completions.create(
model="llama-3.1-8b",
messages=[{
"role": "user",
"content": "Hello",
}],
stream=True,
)
for chunk in response:
print(chunk.choices[0].delta.content)
Workloads we're targeting
- ▸Replacing a third-party LLM API with a pinned open-weight model you control
- ▸Hosting fine-tunes too sensitive to route through shared SaaS inference
- ▸Low-latency streaming chat interfaces that can't tolerate cold starts
- ▸Embedding pipelines that feed a vector store at production volume
- ▸Internal tools where per-token cost matters more than model brand
- ▸Multi-tenant apps needing per-customer key scoping and usage quotas
Where we're going
-
in beta nowChat + embeddingsOpenAI-compatible endpoints for a curated open-weight model set.
-
nextBYO fine-tuneUpload a LoRA adapter or full fine-tune and get a dedicated endpoint in minutes.
-
plannedMulti-model routingPer-request routing across price and latency tiers with automatic fallback.
Inference beta is invite-only.
Tell us what you're building — model sizes, token volumes, latency requirements — and we'll reach out when your slot opens.
Request inference beta access →Ship an agent on your own infrastructure.
Hive is live today. Agents, Inference, and GPU are in private beta — request access and we'll reach out.