Benchmarks

Benchmarks

How fast local language models run on our hardware: which hardware, which model, which settings and what result. The list grows with every new test.

First benchmarks are on the way

We will publish only what we measure ourselves on our own hardware: the model, the settings and the result. For now, the hardware is ready.

Our hardware ↓

Lab

Our hardware

We measure on our own hardware: from a single consumer graphics card to several NVIDIA DGX Spark units and Apple Silicon machines.

4× In use

NVIDIA RTX 3090

24 GB

One runs as an eGPU server. We have a 4-slot NVLink bridge.

3× In use

NVIDIA DGX Spark

128 GB

Linked over ConnectX-7. We measure how models scale from 1 to 4 units.

1× Arriving

Mac Studio M5 Ultra

1× In use

NVIDIA RTX 2080 Ti

11 GB

We plan a memory upgrade to 22 GB and a two-card 44 GB rig.

1× In use

NVIDIA RTX 3060

12 GB

2× In use

AMD RX 480

8 GB

2× In use

Mac (Apple Silicon)

Two Macs linked over Thunderbolt RDMA.

2× In use

X99 (Intel)

1× Arriving

X299 (Intel)

How we measure

  • We measure on our own hardware, with the models and settings listed next to each result.
  • We publish only what we measured ourselves. We do not repeat vendor numbers.
  • Where we tuned a setup, we show the result before and after.