v0.2.2 · 48865e3
2026-09-02 02:18:06 UTC

GPU explorer

LLM inference planning — compare GPU generations by memory, bandwidth, and cost efficiency

Preset:

Vendor:

X Axis:

Y Axis:

Loading GPU catalog…

💡 Top-right = high VRAM and throughput. These GPUs handle larger models and longer contexts.

Bubble size represents BF16 TFLOPs.

Throughput Index is a planning metric derived from memory bandwidth, VRAM, and architecture generation. It enables relative GPU comparison — not exact model throughput.

Inference performance depends on model architecture (GQA vs MHA), sequence length, batching, and inference backend (vLLM, TensorRT-LLM, etc.).

Hardware cost and cost efficiency axes appear when costings is enabled in Settings.