Graphometer

Seeing AI from the right angle.

Free, open tools, resources, reports, and data that show what AI systems actually do.

A graphometer is a surveyor's instrument for measuring angles. We took the name seriously.

The line
Instrument 01 Released

Graphometer Droplet

An independent compatibility kit for Liquid LFMs. A small command line tool that diagnoses how a locally served model handles tool calling, repairs what a chat template drops, and verifies the result against recorded fixtures.

Independent work. Not affiliated with, endorsed by, or connected to Liquid AI.
Field card New

Field card: LFM2.5-2.6B

A measured field card for Liquid AI's new small model: the context window as measured, which file to download, the output budget it really needs, and which sampling settings are actually in effect. Every number from recorded runs.

Independent work. Not affiliated with, endorsed by, or connected to Liquid AI.
Field card New

Field card: Kimi K3

The complete 2.78-trillion-parameter Kimi K3, topology intact under a 1-bit expert build, generating on one desktop at about 0.2 tokens per second. Placement arithmetic worked out before loading, phase telemetry, and speeds with nothing softened. Every number from recorded runs or the model files themselves.

Independent work. Not affiliated with, endorsed by, or connected to Moonshot AI or Unsloth.
Field card New

Field card: Mistral Medium 3.5

Mistral's dense 128B-class model made usable on one desktop: the memory-bandwidth physics measured four ways, a two-token vocabulary repair that let a small draft model do the guessing (about 1 token per second alone, 2 to 5 with it), and timed agent work with the protocol disclosed. Every number from recorded runs.

Independent work. Not affiliated with, endorsed by, or connected to Mistral AI or Unsloth.
Field card New

Field card: Qwen3.5 (122B and 397B)

Two Qwen3.5 models on one desktop: 26 to 38 tokens per second on the 122B, 14 to 20 on the 397B, and the confidence-gate trap measured on the 122B, where ungated speculation ran slower than none on creative text while every benchmark shape looked fine. Every number from recorded runs or the model files themselves.

Independent work. Not affiliated with, endorsed by, or connected to the Qwen team, Alibaba Group, or Unsloth.
Guide New

Which Qwen should I run?

A hardware-honest guide to the whole Qwen family: what your machine can run from a laptop to a 397B on one desktop, a two-command first install, downloadable known-good serving profiles, and when the honest answer is the API. Measured rows labeled measured; vendor rows labeled vendor.

Independent work. Not affiliated with, endorsed by, or connected to the Qwen team or Alibaba Group.
Release watch New

Release watch: Qwen3.8

What has actually shipped versus what is promised, verified with dates: the Max model measured over the API including the thinking-budget trap, the fake-repository hazard already live on public hubs, and the exact battery we run the day real weights land. Updates as things ship.

Independent work. Not affiliated with, endorsed by, or connected to the Qwen team or Alibaba Group.
Incident report New

Empty Answer: how a reasoning model spends your whole budget thinking

A reasoning model can spend your whole output budget thinking and hand back nothing at all: a successful call, no error, an empty answer, and a bill for the hidden tokens. One fixed prompt across three serving lanes with a second vendor's model as a control, the measuring instrument of ours it defeated, and how to find your own budget floor instead of trusting a number. Every figure traces to the download, and the few with no retained body say so.

Independent work. Not affiliated with, endorsed by, or connected to the Qwen team, Alibaba Group, Z.ai, Ollama, or OpenRouter.
Measured study New

Slower by default

Speculative decoding is sold as a free lunch, and llama.cpp's confidence gate on it has shipped three different defaults in two years. On one CPU-offloaded 397B mixture-of-experts deployment, the current default made prose 19.8% slower than no speculation at all while making structured output 22.5% faster, in the same run; every gate tested at or above 0.50 reversed the prose loss, and the widely recommended 0.75 turns out to be the project's own former shipped default. Two recorded runs with stated intervals, every figure traced to the raw records.

Independent work. Not affiliated with, endorsed by, or connected to the llama.cpp project, ggml-org, the Qwen team, Alibaba Group, or Unsloth.
Measured study New

The wrong flag

For six days our own records said llama.cpp's warmup pass was the difference between about 6 tokens per second and 3.6 on a CPU-offloaded 675B mixture-of-experts. A pre-registered three-arm A/B traced the gap to llama.cpp's default thread heuristic instead: it excludes efficiency cores and resolved to 4 compute threads on this 24-core CPU, roughly half this deployment's warm decode speed left on the table by an invisible default. The misattribution, the mechanism, and the corrected records, every figure traced to the raw files.

Independent work. Not affiliated with, endorsed by, or connected to Mistral AI, the llama.cpp project, ggml-org, or Unsloth.
Instrument 02 In development

Graphometer Workbench

for Grok Build

A local interface for the Grok Build coding agent, for people who direct AI agents without living in a terminal. It puts the agent's sessions, its work as it happens, and the moments it stops to ask a question into one plain, readable window.

Works with Grok Build. Not affiliated with, endorsed by, or connected to xAI.
In the works

Field reports on the models we study, curated local AI kits that are honest about hardware, datasets built to fix measured weaknesses, and small tools that make good models easier to live with.