Inference AtlasBETA
Checking status
API: checking
Model: unavailable
GPU: unavailable
OPEN-SOURCE MODELS. OPEN INFRASTRUCTURE.

See inside your
AI inference stack.

Run a model. Follow every token.
Explore the infrastructure behind the intelligence.

Self-hosted by designStreaming by defaultObservable at every step
INFERENCE OVERVIEWINFRASTRUCTURE STATUS
ACTIVE MODEL

Instruction-tuned · Open weights

GPU CONFIGURATION

Telemetry unavailable

TIME TO FIRST TOKEN ms

Measured average

GENERATION SPEED tok/s

Measured average

BrowserWAITING
FastAPIWAITING
vLLMWAITING
GPUWAITING
Token streamWAITING
BEYOND THE API CALL

Your model.
Your infrastructure.
Your understanding.

Operate the model

Explore GPU memory, model serving, and the compute behind each response.

Watch inference unfold

Inspect first-token latency, streaming throughput, and request timing.

Trace the whole path

Follow a prompt from the browser to the gateway, engine, and back.

Make it your own

Control sampling, choose your model, and deploy on your own GPU.

BUILT ON AN OPEN STACK
Next.js FastAPI vLLM Hugging Face Redis Docker NVIDIA

Meet the engine behind the answer.

Run the model yourself