What the bench is

I test products that matter to people building, buying or governing AI in payments and commerce: machines that run models locally, the models and runtimes that go on them, agent products, and the servers agents talk to. Each product runs through a written protocol, and what I publish is the dated set of findings with the evidence attached. There are no star ratings, no rankings and no "best of" lists here. The reader gets the measurements and draws the conclusion.

How a test works

  1. The protocol is written and versioned before the product is opened. Two products tested under the same protocol version can be read side by side. They are not ranked.
  2. Findings are facts with units: tokens per second at a stated context, resident memory, which network addresses a process contacted, how many of seven jobs an agent finished under a signed mandate and what it tried that nobody asked for.
  3. Every finding names its evidence file. The raw outputs are kept and the method is published, so anyone with the same hardware can repeat the run.
  4. Where a test uses my own instruments, they are the open-source Major Labs kits (MandateKit, BudgetGuard, WitnessKit, IdentityKit) and the Mandate Sandbox, unmodified, with versions stated.

Protocols

Every protocol is published before a product is opened. They are plain text; the harness that runs the mandate test is a self-contained package with its kits vendored, published with its own instructions so anyone can run it and send the bundle back: harness instructions, harness package (zip).

  • Local model runtime v1: setup from zero, speed at standard and agent-length context, footprint, egress, the mandate test.
  • Local model runtime v2: v1 plus power, thermals, noise, cold and warm load, two models resident, concurrency, reasoning sweep, quantization and runtime pairs, 32K loops.
  • Desktop box v1: a complete machine on loan: intake, firmware, shipped software and first-boot egress, install from zero, return condition.
  • Edge accelerator v1: cards and boards: SDK install, vision and small-language workloads, power, what runs on the device.
  • Agent product v1: self-hosted agents on an isolated machine: credentials, egress, host reach, guardrails, kill switch.
  • MCP server, dynamic v1: an MCP server run in a sandbox under a recording proxy; declared versus observed.
  • Adversarial suite v1: 36 nudges and injections, where each is delivered, which layer should catch it, how taken, caught and slipped are scored.
  • Kit protocols: who, may, spends, did, remembers: one protocol per Major Labs kit; the mandate one is linked, the others sit beside it.
  • Reporting standard: the findings table, the per-layer counts, the manifest rule, the disclosure lines.

Entries

2026-10-03: gpt-oss-20b on the same M5 Pro, second pass with the packaged harness
A clean repeat of entry one through the packaged harness: 91.0 tokens per second at empty context (89.8 in entry one); 20 of 27 required steps done (21); 2 of 45 unrequested options taken, both declined by the customer, 0 slipped; 0 mandate refusals and 2 budget-guard refusals, one after the model paid the kettle money to the bank instead of the store; 13 runtime parser failures (9). First signed evidence manifest and bundle. Obtained: own equipment, free downloads. Findings, protocol, log, harness and evidence files.

2026-10-02: gpt-oss-20b on an Apple M5 Pro, 24GB, under a bank mandate
21 runs of seven bank-customer jobs under a signed MandateKit mandate. One of 45 unrequested options taken and caught; six of 27 required steps undone; nine runtime parser failures. 89.8 tokens per second at empty context, loopback-only egress. Obtained: own equipment, free downloads. Findings, protocol, log, harness and evidence files.

Disclosure

Every published entry begins with how I got the product: bought it, used a free trial, borrowed it from the vendor, or received it from the vendor as a permanent unit. The dates are stated. No payment or other consideration is accepted for a test, and a product that comes with conditions on the findings is not tested. Vendors do not see the write-up before publication. A vendor may check the findings table for factual error, and if that happened the entry says so.

Equipment

I prefer to borrow. Loans run 30 to 60 days and the unit goes back. A unit that a vendor supplies permanently is listed here for as long as it is on the bench, with the supplier named.

  • Apple MacBook Pro, M5 Pro, 24GB: own equipment, purchased, since 2026.

Contributed runs

If a machine cannot leave the building, the protocol can come to it. Anyone can run a published protocol on their own hardware and send me the output: the harness, the evidence format and the protocol are public. A contributed run is published as its own entry, labeled as contributed, with the contributor named and the evidence files published alongside. I report what the files show and say plainly that I did not repeat the run myself.

For vendors

If you make something in this lane and want it on the bench, write to hello@majormatters.co with the product, the configuration you can supply and whether it is a loan or a permanent unit. You get the protocol in advance, a dated and reproducible published result, the raw evidence, and a factual check of the findings table before publication. You do not get copy approval, and the result is published whatever it says.

Licenses and how to cite

The harness is Apache 2.0. The protocols, findings and evidence are CC BY 4.0: copy them, run them, build on them, and credit "Charlie Major, the Major Matters Bench" with a link here. The suite is named the Major Matters Bench Protocol, versioned and dated; a run made with it elsewhere is "run under the Major Matters Bench Protocol v<N>", not an entry on the bench, unless it is published here or accepted as a contributed run. Cite an entry as: Charlie Major, The Major Matters Bench, entry N: title, Major Matters, date, with the protocol version and the manifest root. Dataset DOIs are minted per batch on Zenodo. Full licence text.

Corrections

An error in a published finding is corrected in place with a dated note, the same way as every other Major Matters piece. The editorial standards behind all of this are on the methodology page.

Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.