Run open models
in Europe
Run serverless open models, billed per token, and move to dedicated GPUs as your workload grows.
- Zero data retention
- European data centres
- OpenAI-compatible API
Trusted by
Start on serverless and grow into dedicated GPUs
Pay per token as traffic varies, then reserve GPUs when it steadies.
- Start hereExplore serverless
Serverless inference
Call open models through one OpenAI-compatible API.
Billed per token. No base fee.
- Explore dedicated
Dedicated endpoints
Reserve an endpoint and GPU capacity for your model.
Commitment pricing, agreed per contract.
- Explore compute
GPU compute
Run your containers on GPU virtual machines or dedicated clusters.
VMs billed per second. No subscription.
For your security review
Meet your
next model
Browse all modelsGLM-5.3 Flash
DeepSeek V4.1 Flash
Why teams build on Lyceum
Talk to the engineers who run Lyceum
Get help by email. Business customers also get a shared Slack channel and agreed response times.
Talk to an expertApproximate output speed reported by a customer in their own tests. Request length, concurrency and reasoning-token treatment were not provided. These results are separate from our live status checks. Kimi K3: about 185 tokens/s median output speed. GLM-5.3: about 125 tokens/s output speed. Scale
From a few GPUs to 1,000s
Start small and scale to thousands of GPUs as your workload grows.
Flexibility
Compare costs- Per token
- Serverless inferenceNo base fee
- Per second
- GPU virtual machinesNo subscription
- Per contract
- Dedicated and reserved capacity
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
)
response = client.chat.completions.create(
model="moonshotai/kimi-k3",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)Kimi K3
export ANTHROPIC_BASE_URL="https://api.lyceum.technology/anthropic"
export ANTHROPIC_AUTH_TOKEN="$LYCEUM_API_KEY"
export ANTHROPIC_MODEL="moonshotai/kimi-k3"
claudeCustomer stories

Albatross retrains its recommendation models several times a day
“What stood out was how fast Lyceum turned our requirements into a working integration that fit our ML workflow.”
Dr. Matteo Ruffini, Co-founder and Chief Scientist, Albatross 
SONALAB scales its audio AI on European compute
“Need something? A single email or Slack message and it’s sorted.”
Sören Hübner, CTO at SONALAB, now SONA
What happens to your data
Runs in European data centres
Serverless models run in Paris and Finland. Dedicated and VM locations are agreed per contract.
We don’t keep your prompts
Prompts and outputs are processed, not stored. Nothing is written to disk.
We never train on your data
We don’t keep your prompts or outputs, so there’s nothing to train on.
Ready for your review
Our data processing agreement (DPA) and technical and organisational measures (TOMs) are ready before your security team asks.
GPUs for your own models
If you run or train your own models, rent the GPUs directly. Start one now, pay less with spot, or reserve capacity for the long run.
Start and stop when you like
Billed per second, with no subscription. Good for experiments, short jobs and work that comes in bursts.
Same GPUs, about 60% less
A spot VM can be reclaimed at any time. Good for jobs that save checkpoints and can restart.
Capacity held for you
Commit to GPUs for your workload at a price agreed per contract. Good for long training runs and steady inference.
NVIDIA B300
Newest generation
288 GB HBM3e
On demand (selected) $7.99 per GPU hour
Spot (selected) $2.99 per GPU hour
Reserved (selected) Per contract
LaunchNVIDIA B200
Inference and training
192 GB HBM3e
On demand (selected) $6.49 per GPU hour
Spot (selected) $2.40 per GPU hour
Reserved (selected) Per contract
LaunchNVIDIA H200
Large-memory inference
141 GB HBM3e
On demand (selected) $4.29 per GPU hour
Spot (selected) $1.60 per GPU hour
Reserved (selected) Per contract
LaunchNVIDIA H100
Production workloads
80 GB HBM3
On demand (selected) $2.79 per GPU hour
Spot (selected) $1.10 per GPU hour
Reserved (selected) Per contract
Launch
US dollars per GPU hour. Billed per second. No subscription.