Operate the model
Explore GPU memory, model serving, and the compute behind each response.
Run a model. Follow every token.
Explore the infrastructure behind the intelligence.
Instruction-tuned · Open weights
Telemetry unavailable
Measured average
Measured average
Explore GPU memory, model serving, and the compute behind each response.
Inspect first-token latency, streaming throughput, and request timing.
Follow a prompt from the browser to the gateway, engine, and back.
Control sampling, choose your model, and deploy on your own GPU.