SAQI/AI Launch →
private beta
// inference

Dedicated
inference endpoints.

OpenAI-compatible API endpoints for open-weight and fine-tuned models. Token billing, no cold starts, no shared-tenant noise. Drop-in replacement for any OpenAI SDK call — change the URL, change the key, done.

▸ openai-compatible
▸ no cold starts
▸ token billing
▸ byo fine-tune
// what you get

Production inference, not a playground.

Endpoints built for latency-sensitive workloads. Your model, your keys, your traffic — no other tenants sharing your hardware.

◇

OpenAI-compatible API

Drop /v1/chat/completions and /v1/embeddings into any OpenAI SDK call. Change the base URL and key — your code ships unchanged.

◉

Dedicated endpoints

Your endpoint runs on its own hardware slice. No shared-tenant contention, no noisy-neighbour latency spikes at peak hours.

◈

No cold starts

Dedicated endpoints are always warm. No 30-second waits for the container to wake up on the first request of the hour.

◆

Token-based billing

Pay per input and output token. No hourly minimums. Idle endpoints cost nothing until a request lands.

◇

Scoped API keys

Issue per-project keys with their own token quotas, rate limits and expiry. Rotate or revoke without touching other projects.

◉

Full streaming support

Server-sent events with delta chunks — identical to OpenAI streaming. No polling, no timeout fighting on long responses.

◈

BYO fine-tune

Upload a LoRA adapter or a full fine-tune as a container and get a dedicated endpoint. Your weights stay on your endpoint.

◆

Usage analytics

Per-request and per-token breakdowns with latency percentiles. Know exactly what each project costs before the invoice arrives.

◇

GPU handoff

Fine-tune on a saqi GPU node, push the adapter directly to an inference endpoint — same console, same account, no pipeline glue.

// model catalog

Open-weight models, ready now.

The catalog ships with tested, quantized versions of the best open-weight models. Every endpoint is OpenAI-compatible. Bring your own fine-tune and it joins the same API surface.

Llama 3.3

general purpose
70B

Meta's flagship open-weight. Strong on instruction-following, reasoning and code.

Llama 3.1

general purpose
8B 70B

Efficient 8B for latency-sensitive apps; 70B for complex multi-step reasoning.

Qwen 2.5

multilingual + code
7B 72B

Best-in-class multilingual coverage and coding performance across the open-weight field.

Mistral 7B

fast + efficient
7B

Low-latency workhorse. Excellent price-per-token for high-volume applications.

Phi-3.5

small + capable
3.8B

Microsoft's compact model — strong reasoning in a package suited for constrained deployments.

BYO model

bring your own
any size

Upload a LoRA adapter or full fine-tune container. Get a dedicated endpoint immediately.

// api compatibility

One line to swap.

The endpoint speaks the OpenAI protocol. Change base_url and api_key. Every SDK, library and middleware that works with OpenAI works here.

  • ▸OpenAI Python + Node SDKs
  • ▸LangChain, LlamaIndex, DSPy
  • ▸LiteLLM and any proxy router
  • ▸Continue.dev, Cursor, Copilot alternatives
inference_example.py
from openai import OpenAI

client = OpenAI(
    base_url="https://api.saqi.ca/v1",
    api_key="sk-proj-...",
)

response = client.chat.completions.create(
    model="llama-3.1-8b",
    messages=[{
        "role": "user",
        "content": "Hello",
    }],
    stream=True,
)
for chunk in response:
    print(chunk.choices[0].delta.content)
// use cases

Workloads we're targeting

  • ▸Replacing a third-party LLM API with a pinned open-weight model you control
  • ▸Hosting fine-tunes too sensitive to route through shared SaaS inference
  • ▸Low-latency streaming chat interfaces that can't tolerate cold starts
  • ▸Embedding pipelines that feed a vector store at production volume
  • ▸Internal tools where per-token cost matters more than model brand
  • ▸Multi-tenant apps needing per-customer key scoping and usage quotas
// roadmap

Where we're going

  1. in beta now
    Chat + embeddings
    OpenAI-compatible endpoints for a curated open-weight model set.
  2. next
    BYO fine-tune
    Upload a LoRA adapter or full fine-tune and get a dedicated endpoint in minutes.
  3. planned
    Multi-model routing
    Per-request routing across price and latency tiers with automatic fallback.

Inference beta is invite-only.

Tell us what you're building — model sizes, token volumes, latency requirements — and we'll reach out when your slot opens.

Request inference beta access →
// next step

Ship an agent on your own infrastructure.

Hive is live today. Agents, Inference, and GPU are in private beta — request access and we'll reach out.