Your AI platform. Inside your walls.
Olympus serves models, GPUs, agents and knowledge on hardware you own, on-premises or in your cloud tenancy. Every request is authorised before a model is called and metered after. Prompts, answers and vectors stay on your clusters.
How many days of annual leave do new employees get?
New employees get 25 days of annual leave a year, accrued monthly from their start date. 1 Up to five unused days carry over to the next year. 2
Olympusx.ai holds no security certification for Olympus today. The platform provides the controls; the certification is yours to obtain.
The models you want are on the wrong side of your walls
Send regulated data to a public AI service and you have a transfer nobody will sign off. Build the platform yourself and you spend months on infrastructure before a single business question is answered.
The data cannot leave
Data-protection law, sector regulators and customer contracts limit where prompts, documents and answers may be processed. Another region of the same provider rarely changes the answer.
The spend must be attributable
GPUs shared across departments become a cost centre nobody can charge back. Finance will not fund what it cannot allocate.
The use must be provable
An auditor asks who used which model, on whose authority, on what data, and what it cost. If that takes a week to assemble, the platform is not ready to be inspected.
Everything an organisation needs before it can run AI at all
One gateway at the core, the planes around it, and your walls around all of it. Every request travels inward through the gateway; nothing travels out unless you open a door.
AI gateway
OpenAI-compatible, with every request authorised and metered.
Model serving and LLMaaS
Open models on your GPUs, cloud models if you allow them.
GPU capacity
Discovered, partitioned, granted and billed by the share.
Agents and assistants
Granted tools, approvals that resume, a trace for every step.
Knowledge and retrieval
A database per tenant, hybrid search, cited answers.
Governance
Isolation, guardrails and an audit trail that catches tampering.
FinOps
Budgets that stop, statements a department recognises.
A policy that only exists in a dashboard is not a control
Every request crosses the same checks before a GPU does any work. One request passes and is metered; the next is over its budget and is refused in the gateway. The model is never called.
Serve open models on your GPUs, and offer them like a cloud
Pick a model, a place and the hardware. Olympus sizes the plan before anything is pulled, deploys it, routes it, and meters every token, whichever of your doors it came through.
Deploy Qwen3-8B
The pre-flight sizes the plan from the model's own config before anything is pulled.
The ten-minute failures, refused before the pull
Weights are sized at their real bytes per parameter and the cache from the model's own attention shape. A tensor-parallel split that cannot work, FP8 on a card without it, a context longer than the model was trained for: each is rejected before anything is applied.
- card or slicea model gets whole GPUs or a MIG slice, sized against what that slice holds
- by digestruntime and weights are pinned; a tag can move, a digest cannot
- 1 catalogself-hosted and cloud models granted, budgeted and metered alike
client = OpenAI(
base_url="https://api.olympusx.ai/v1",
api_key=os.environ["OLYMPUS_API_KEY"])Slice your GPUs, grant them, and bill them by the share
Olympus discovers the cards on every cluster, partitions them when you allow it, and keeps every tenant inside the capacity you granted. A token on a large model draws a budget down faster than one on a small model.
GPU partitioning
How each MIG-capable card is carved, read from the cluster's own configuration.
Capacity you can hand out without overselling it
Pool, envelope and quota are checked in both directions, in one transaction, so two administrators cannot both pass the check and both commit.
| Model | Token weight | Weighted tokens |
|---|---|---|
| BGE-M3 | 0.400 | 18,400 |
| Granite 3.3 8B | 1.000 | 64,210 |
| GPT-OSS 20B | 2.000 | 211,880 |
Agents that answer to you
Agents read your documents, call your systems and hand work to one another, inside the same walls. Every tool is granted, every step is traced, and anything that moves money waits for a person.
Check sentinel code 4471 on account ending 0042.
I need a person's approval before I call the core banking lookup. I have sent the request to the operations queue and will answer here once it is decided.
Read your knowledge
Collections you attach, cited on every answer.
Call granted tools
From the agent's allow-list; a skill only narrows it.
60s per callHand work on
To a collaborator with its own model and knowledge.
2levels · 12 per runWait for a person
The run holds and resumes on the same conversation.
24h to decideAnswers grounded in your documents, with the sources first
Each tenant's vectors live in a database of its own. Retrieval is hybrid, reranked and gated: when nothing is relevant, the answer says so and costs nothing.
- 50 + 50candidates from vector and full-text search, fused in one round trip
- ~4×over-retrieved, then reordered by a reranker running on your hardware
- per modelrelevance floor, so weak matches never ground an answer
- 0 tokensbilled when nothing clears the floor; the answer says it found nothing
The controls are ours. The certification is yours.
Isolation you can point an auditor at, a trail that catches its own tampering, and guardrails you watch before you enforce.
Isolation you can inspect
A project is a namespace on your cluster with its own quota and four named network policies, deny-first. Row-level security is a second wall behind it.
A trail that catches tampering
Every audit row is chained to the last. One check reports the chain intact or names the first row that breaks it. Tamper-evident, not tamper-proof, and the residual risks are written down.
Observe before you enforce
Guardrails start by recording what they would catch. You switch a tenant to enforce once you have seen what it means for real traffic.
Reversible on exit
The uninstall is staged by what it destroys, and the irreversible step is named before you reach it.
From an empty cluster to a metered answer
A first run in the browser takes a platform administrator from nothing to a verified, metered answer. Each step commits as it completes, so a closed tab resumes where it stopped.
One package, one values file
On your clusters, on-premises or in your cloud tenancy.
Seven steps, in the browser
Tenant, offering, model, key, prices, cluster, verification.
A real answer, metered
Send a chat request and watch the answer and its token count come back.
Departments become tenants
Each with its own quota, keys, data and statement.
The browser-only claim covers bringing the installed platform into service. The cluster install itself is a command-line package install. Rate limits and retention windows are set through the API today.
A bill they recognise, not a cloud invoice
Departments become tenants, and every tenant gets a statement: reserved capacity by the card or the slice, tokens weighted by the hardware they used, and budgets that decline rather than surprise.
| Line | Basis | Amount |
|---|---|---|
| Reserved H100 80GB × 2 | monthly commitment | [your rate] |
| MIG A100 1g.10gb × 3 | pro-rata on memory | [your rate] |
| GPT-OSS 20B tokens | input and output, weight 2.000 | [your rate] |
| Llama 3.3 70B via Oracle | input and output, provider rate | [your rate] |
Put it on your own cluster and check.
An evaluation install runs on your hardware, in your network, with your identity provider. You get the verification scripts before you start.
Go to Studio Request an evaluation install ai@olympusx.ai