Seeing AI from the right angle.
Free, open tools, resources, reports, and data that show what AI systems actually do.
A graphometer is a surveyor's instrument for measuring angles. We took the name seriously.
Graphometer Droplet
An independent compatibility kit for Liquid LFMs. A small command line tool that diagnoses how a locally served model handles tool calling, repairs what a chat template drops, and verifies the result against recorded fixtures.
Field card: LFM2.5-2.6B
A measured field card for Liquid AI's new small model: the context window as measured, which file to download, the output budget it really needs, and which sampling settings are actually in effect. Every number from recorded runs.
Field card: Kimi K3
The complete 2.78-trillion-parameter Kimi K3, topology intact under a 1-bit expert build, generating on one desktop at about 0.2 tokens per second. Placement arithmetic worked out before loading, phase telemetry, and speeds with nothing softened. Every number from recorded runs or the model files themselves.
Field card: Mistral Medium 3.5
Mistral's dense 128B-class model made usable on one desktop: the memory-bandwidth physics measured four ways, a two-token vocabulary repair that let a small draft model do the guessing (about 1 token per second alone, 2 to 5 with it), and timed agent work with the protocol disclosed. Every number from recorded runs.
Field card: Qwen3.5 (122B and 397B)
Two Qwen3.5 models on one desktop: 26 to 38 tokens per second on the 122B, 14 to 20 on the 397B, and the confidence-gate trap measured on the 122B, where ungated speculation ran slower than none on creative text while every benchmark shape looked fine. Every number from recorded runs or the model files themselves.
Which Qwen should I run?
A hardware-honest guide to the whole Qwen family: what your machine can run from a laptop to a 397B on one desktop, a two-command first install, downloadable known-good serving profiles, and when the honest answer is the API. Measured rows labeled measured; vendor rows labeled vendor.
Release watch: Qwen3.8
What has actually shipped versus what is promised, verified with dates: the Max model measured over the API including the thinking-budget trap, the fake-repository hazard already live on public hubs, and the exact battery we run the day real weights land. Updates as things ship.
The empty answer
A reasoning model can spend your whole output budget thinking and hand back nothing at all: a successful call, no error, an empty answer, and a bill for the hidden tokens. One fixed prompt across three serving lanes with a second vendor's model as a control, the measuring instrument of ours it defeated, and how to find your own budget floor instead of trusting a number. Every figure traces to the download, and the few with no retained body say so.
Graphometer Workbench
A local interface for the Grok Build coding agent, for people who direct AI agents without living in a terminal. It puts the agent's sessions, its work as it happens, and the moments it stops to ask a question into one plain, readable window.
Field reports on the models we study, curated local AI kits that are honest about hardware, datasets built to fix measured weaknesses, and small tools that make good models easier to live with.