Run open models
in Europe

Run serverless open models, billed per token, and move to dedicated GPUs as your workload grows.

  • Zero data retention
  • European data centres
  • OpenAI-compatible API
  • Enpal
  • Apheris
  • Weight Space Labs
  • Macrocosmos
  • Edgeless Systems
  • Continuous Intelligence
  • Albatross
  • SONA

Start on serverless and grow into dedicated GPUs

Pay per token as traffic varies, then reserve GPUs when it steadies.

  1. Start here

    Serverless inference

    Call open models through one OpenAI-compatible API.

    Billed per token. No base fee.

    Explore serverless
    hello.sh

    $ curl …/chat/completions

    "content": "Hello!"

    Billed per token. No base fee.

  2. Dedicated endpoints

    Reserve an endpoint and GPU capacity for your model.

    Commitment pricing, agreed per contract.

    Explore dedicated
    reserved.txt

    Reserved. Nobody else gets these.

    Talk to an expert

    Commitment pricing, agreed per contract

  3. GPU compute

    Run your containers on GPU virtual machines or dedicated clusters.

    VMs billed per second. No subscription.

    Explore compute
    root@lyceum

    root@lyceum:~# whoami

    root

    You have root.

    Billed per second. No subscription.

Meet your
next model

Browse all models

Why teams build on Lyceum

Talk to the engineers who run Lyceum

Get help by email. Business customers also get a shared Slack channel and agreed response times.

Talk to an expert
  • Performance

    Customer benchmarksLive status
    Approximate output speed reported by a customer in their own tests. Request length, concurrency and reasoning-token treatment were not provided. These results are separate from our live status checks. Kimi K3: about 185 tokens/s median output speed. GLM-5.3: about 125 tokens/s output speed.
  • Scale

    From a few GPUs to 1,000s

    Start small and scale to thousands of GPUs as your workload grows.

  • Flexibility

    Compare costs
    Per token
    Serverless inferenceNo base fee
    Per second
    GPU virtual machinesNo subscription
    Per contract
    Dedicated and reserved capacity

Your tools with
a new model

Use open models with your existing SDK.

Read the setup guide
Choose a model
app.py
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LYCEUM_API_KEY"],
    base_url="https://api.lyceum.technology/openai/v1",
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    messages=[{"role": "user", "content": "Hello!"}],
)

print(response.choices[0].message.content)

Kimi K3

Terminal
export ANTHROPIC_BASE_URL="https://api.lyceum.technology/anthropic"
export ANTHROPIC_AUTH_TOKEN="$LYCEUM_API_KEY"
export ANTHROPIC_MODEL="moonshotai/kimi-k3"

claude

Customer stories

What happens to your data

How we handle your data
  • europe.bmp

    Two pins, and that’s all of them.

    Dedicated and VM locations are agreed per contract.

    Browse models

    2 pins

    Runs in European data centres

    Serverless models run in Paris and Finland. Dedicated and VM locations are agreed per contract.

  • prompt.txt

    This file was never saved.

    0 bytes

    We don’t keep your prompts

    Prompts and outputs are processed, not stored. Nothing is written to disk.

  • training-data

    Still empty.

    It will stay that way.

    0 objects

    We never train on your data

    We don’t keep your prompts or outputs, so there’s nothing to train on.

  • DPA.pdf

    The real one has fewer pixels.

    Get the documents
    TOMs.pdf

    Every measure, none of the dither.

    Get the documents

    Ready for your review

    Our data processing agreement (DPA) and technical and organisational measures (TOMs) are ready before your security team asks.

GPUs for your own models

If you run or train your own models, rent the GPUs directly. Start one now, pay less with spot, or reserve capacity for the long run.

Compute pricing
My GPUs (4)

Start and stop when you like

Billed per second, with no subscription. Good for experiments, short jobs and work that comes in bursts.

Same GPUs, about 60% less

A spot VM can be reclaimed at any time. Good for jobs that save checkpoints and can restart.

Capacity held for you

Commit to GPUs for your workload at a price agreed per contract. Good for long training runs and steady inference.

  • NVIDIA B300

    Newest generation

    288 GB HBM3e

    On demand (selected) $7.99 per GPU hour

    Spot (selected) $2.99 per GPU hour

    Reserved (selected) Per contract

    Launch
  • NVIDIA B200

    Inference and training

    192 GB HBM3e

    On demand (selected) $6.49 per GPU hour

    Spot (selected) $2.40 per GPU hour

    Reserved (selected) Per contract

    Launch
  • NVIDIA H200

    Large-memory inference

    141 GB HBM3e

    On demand (selected) $4.29 per GPU hour

    Spot (selected) $1.60 per GPU hour

    Reserved (selected) Per contract

    Launch
  • NVIDIA H100

    Production workloads

    80 GB HBM3

    On demand (selected) $2.79 per GPU hour

    Spot (selected) $1.10 per GPU hour

    Reserved (selected) Per contract

    Launch

US dollars per GPU hour. Billed per second. No subscription.

Your next workload starts here