Olympusx AI Go to Studio
Sovereign AI platform

Your AI platform. Inside your walls.

Olympus serves models, GPUs, agents and knowledge on hardware you own, on-premises or in your cloud tenancy. Every request is authorised before a model is called and metered after. Prompts, answers and vectors stay on your clusters.

Olympusx AI
Hugging Faceopen-weight models, imported into your catalog
Qwen3-8BAlibaba Cloud
DeepSeekDeepSeek
LlamaMeta
GemmaGoogle
KimiMoonshot AI
GLMZ.ai
Qwen3-8B added to your catalog
Qwen3-8BQwen/Qwen3-8B · llm · 8.2BH100 80GB × 1
Pre-flightfits on one card, sized from the model's own config
Weightspulled and pinned by digest, never by tag
Routedqwen3-8b on /v1/chat/completions, behind the gateway
Meteredevery token to tenant · project · key
HR assistantgrounded on policiesQwen3-8B

How many days of annual leave do new employees get?

New employees get 25 days of annual leave a year, accrued monthly from their start date. 1 Up to five unused days carry over to the next year. 2

1hr-handbook.pdfS3 · page 122leave-policy.docxupload
5NVIDIA GPU familiesL40S, H100, H200, B200 and B300, discovered on every cluster.
18MIG profilesFrom whole cards down to one-seventh slices, priced by the share.
7Checks before a model runsAny refusal happens in the gateway, before a GPU does any work.
0.21sTime to first tokenStreamed through the full gateway, metering included.
3Cloud providers to switch onOracle, OpenAI and Anthropic, governed like your own models.
OpenAI-compatible APIExisting applications work by changing one base URL.
Open-source foundationsNo proprietary platform licence underneath the platform.
Authorised before the modelKeys, grants, budgets and limits are checked in the request path.
Outbound only if you allow itCloud models and model downloads are switches you turn on, one by one.

Olympusx.ai holds no security certification for Olympus today. The platform provides the controls; the certification is yours to obtain.

The problem

The models you want are on the wrong side of your walls

Send regulated data to a public AI service and you have a transfer nobody will sign off. Build the platform yourself and you spend months on infrastructure before a single business question is answered.

1

The data cannot leave

Data-protection law, sector regulators and customer contracts limit where prompts, documents and answers may be processed. Another region of the same provider rarely changes the answer.

2

The spend must be attributable

GPUs shared across departments become a cost centre nobody can charge back. Finance will not fund what it cannot allocate.

3

The use must be provable

An auditor asks who used which model, on whose authority, on what data, and what it cost. If that takes a week to assemble, the platform is not ready to be inspected.

Enforced in the request path

A policy that only exists in a dashboard is not a control

Every request crosses the same checks before a GPU does any work. One request passes and is metered; the next is over its budget and is refused in the gateway. The model is never called.

The billed model is read from the bodyNot from a header a client could set, so a cheap granted model cannot front for an expensive one.
Guardrails run in the gatewayA raw SDK call gets the same policy as the console, because both reach the same decision.
Token limits are honest about tokensTokens exist only once a response completes, so the request that crosses a limit is served in full and the next one is refused.
Prices are frozen onto every ledger rowA rate change cannot restate history; every line traces back to the call that caused it.
Model serving and LLMaaS

Serve open models on your GPUs, and offer them like a cloud

Pick a model, a place and the hardware. Olympus sizes the plan before anything is pulled, deploys it, routes it, and meters every token, whichever of your doors it came through.

Olympusx AIConnected

Deploy Qwen3-8B

The pre-flight sizes the plan from the model's own config before anything is pulled.

Whole GPUMIG slice
Projectpayments / prod
GPUH100 80GB × 1
Context8,192
Sequences32
61.5 GiBneeded of 72 GiB available on one card
Weights 15.26KV cache 36.00Overhead 10.25
This plan should run

The ten-minute failures, refused before the pull

Weights are sized at their real bytes per parameter and the cache from the model's own attention shape. A tensor-parallel split that cannot work, FP8 on a card without it, a context longer than the model was trained for: each is rejected before anything is applied.

  • card or slicea model gets whole GPUs or a MIG slice, sized against what that slice holds
  • by digestruntime and weights are pinned; a tag can move, a digest cannot
  • 1 catalogself-hosted and cloud models granted, budgeted and metered alike
Point an existing OpenAI client at your own platform
client = OpenAI(
    base_url="https://api.olympusx.ai/v1",
    api_key=os.environ["OLYMPUS_API_KEY"])
Your GPUsNVIDIA, whole cards or MIG slices
Oracle Cloudif allowed · no stored secret
OpenAIif allowed · key sealed
Anthropicif allowed · key sealed
GPU capacity

Slice your GPUs, grant them, and bill them by the share

Olympus discovers the cards on every cluster, partitions them when you allow it, and keeps every tenant inside the capacity you granted. A token on a large model draws a budget down faster than one on a small model.

Olympusx AIConnected

GPU partitioning

How each MIG-capable card is carved, read from the cluster's own configuration.

Repartitioning is enabled on this cluster. Applying a plan terminates every process on the affected GPUs. Switch this off when you are done.
a100-mig-01A100 80GB · 4 of 8 cards shown · all-1g.10gb
80 GB7 × 1g.10gb
80 GB7 × 1g.10gb
80 GB7 × 1g.10gb
80 GB7 × 1g.10gb
l40s-01 · no MIG hardware: time-slicing insteadrepartitioning off until allowed

Capacity you can hand out without overselling it

Pool, envelope and quota are checked in both directions, in one transaction, so two administrators cannot both pass the check and both commit.

Weighted usage this month (illustrative)
ModelToken weightWeighted tokens
BGE-M30.40018,400
Granite 3.3 8B1.00064,210
GPT-OSS 20B2.000211,880
Agents and assistants

Agents that answer to you

Agents read your documents, call your systems and hand work to one another, inside the same walls. Every tool is granted, every step is traced, and anything that moves money waits for a person.

Payments assistantWaiting for approval

Check sentinel code 4471 on account ending 0042.

I need a person's approval before I call the core banking lookup. I have sent the request to the operations queue and will answer here once it is decided.

2 msreadpolicies2 passages from 2 documents
25 msasked the modelgpt-oss:20b13,661 ms · 1,098 tok
13,735 msneeds approval forlookup_sentinel57 msReason: held for a human to approve
Payments desk · run held, not failedtimings from a recorded run · question illustrative
read

Read your knowledge

Collections you attach, cited on every answer.

did

Call granted tools

From the agent's allow-list; a skill only narrows it.

60s per call
delegate

Hand work on

To a collaborator with its own model and knowledge.

2levels · 12 per run
waiting

Wait for a person

The run holds and resumes on the same conversation.

24h to decide
Knowledge and retrieval

Answers grounded in your documents, with the sources first

Each tenant's vectors live in a database of its own. Retrieval is hybrid, reranked and gated: when nothing is relevant, the answer says so and costs nothing.

  • 50 + 50candidates from vector and full-text search, fused in one round trip
  • ~4×over-retrieved, then reordered by a reranker running on your hardware
  • per modelrelevance floor, so weak matches never ground an answer
  • 0 tokensbilled when nothing clears the floor; the answer says it found nothing
Sovereignty and governance

The controls are ours. The certification is yours.

Isolation you can point an auditor at, a trail that catches its own tampering, and guardrails you watch before you enforce.

Hash chain intact
Observe · detect and audit only, nothing blocked Enforce · redact personal data and block unsafe content

Isolation you can inspect

A project is a namespace on your cluster with its own quota and four named network policies, deny-first. Row-level security is a second wall behind it.

A trail that catches tampering

Every audit row is chained to the last. One check reports the chain intact or names the first row that breaks it. Tamper-evident, not tamper-proof, and the residual risks are written down.

Observe before you enforce

Guardrails start by recording what they would catch. You switch a tenant to enforce once you have seen what it means for real traffic.

Reversible on exit

The uninstall is staged by what it destroys, and the irreversible step is named before you reach it.

How it works

From an empty cluster to a metered answer

A first run in the browser takes a platform administrator from nothing to a verified, metered answer. Each step commits as it completes, so a closed tab resumes where it stopped.

Install

One package, one values file

On your clusters, on-premises or in your cloud tenancy.

Open the console

Seven steps, in the browser

Tenant, offering, model, key, prices, cluster, verification.

Verify

A real answer, metered

Send a chat request and watch the answer and its token count come back.

Onboard

Departments become tenants

Each with its own quota, keys, data and statement.

The browser-only claim covers bringing the installed platform into service. The cluster install itself is a command-line package install. Rate limits and retention windows are set through the API today.

FinOps

A bill they recognise, not a cloud invoice

Departments become tenants, and every tenant gets a statement: reserved capacity by the card or the slice, tokens weighted by the hardware they used, and budgets that decline rather than surprise.

Statement 0007 · PaymentsOctober 2026 · draft
LineBasisAmount
Reserved H100 80GB × 2monthly commitment[your rate]
MIG A100 1g.10gb × 3pro-rata on memory[your rate]
GPT-OSS 20B tokensinput and output, weight 2.000[your rate]
Llama 3.3 70B via Oracleinput and output, provider rate[your rate]
sequential · frozen · gaplesscorrections are credit notes
Reserved capacityPer GPU flavour and term, MIG partitions priced pro-rata on memory, commitments with a discount.
Token consumptionPer model, per million tokens, input and output separately, weighted by the model's GPU footprint.
Budgets that actually stopAt tenant, project and key scope, with alerts on the way and a block at the limit.
Your rate card, not oursOlympus ships a complete draft rate card, as a draft. Your operator publishes their own, so there is no list price here.

Put it on your own cluster and check.

An evaluation install runs on your hardware, in your network, with your identity provider. You get the verification scripts before you start.

Go to Studio Request an evaluation install ai@olympusx.ai