Stig

I run the AI agents that produce this site’s measurements, and I check every number before it publishes.


What I do here

I write everything on MeasuredRun, and I am the person accountable for every number on it.

For each article that means: I read the draft against the raw logs, I check every figure traces to a recorded measurement, I confirm that anything described as first-hand really was first-hand, and I cut whatever claims more than the evidence supports. Then I sign it. Nothing goes out without that pass.

The test execution and the drafting are done by a small department of AI agents that I operate. I have not hidden that anywhere — it is set out in full on the How We Use AI page, because the way this site is made is part of what it reports on.


Why I write under a pen name

Stig is a pen name. One real person is behind it, and that person reviews and signs every article.

I would rather you assessed the work than the person. What I am asking you to trust is not a biography — it is a method: run the thing, keep the raw log, publish the failures, date every measurement, correct in public. Every one of those is checkable from the articles themselves, whatever my name is.

What I will not do is dress the pen name up with a career, a job title, or credentials I am not putting my legal name behind. A claim you cannot check is worth nothing to you, and I would rather have a small honest page than an impressive one.


What I will be writing about

These are the areas I am setting out to measure. As articles publish, this list will describe what is actually here rather than what is planned:

  • Running AI coding agents on real repositories — what they cost, how long they take, and where they fail
  • API economics: token accounting, prompt caching, context window costs, rate limits
  • Workflow automation — n8n, scheduled agents, browser automation, MCP servers
  • Head-to-head comparisons where the same task is run identically on both sides
  • Setting up and operating the infrastructure this site runs on

I stay inside that boundary. If a question falls outside it, the article says I did not test it rather than filling the gap with confident prose.


How I work

I run it before I write it. If I have only read the documentation, the article says so, in the body, where you will see it.

I publish the failures. Every article has a section on what went wrong. Anyone can write a walkthrough where everything succeeds. The specific way a thing breaks is the part you cannot get from a product page, and it is the part that proves I was actually there.

I date every measurement. Model versions and prices move. A benchmark without a date is not a benchmark.

I say “unconfirmed” when it is unconfirmed. Where I could not verify something, the article marks it as unverified rather than smoothing over the gap. There are things in my own setup notes marked unconfirmed right now.

I correct in public. Errors get a dated note on the article saying what was wrong and what changed. If the error changed the conclusion, it goes at the top, not in a footnote. I do not silently overwrite.


Elsewhere

  • GitHub — https://github.com/measuredrun

That is where my code and working files go when I publish them.


Contact

havegoodin7@gmail.com

If you ran one of my tests and got a different number, that is the mail I most want to receive. Tell me what you ran and what you got. If you are right, I fix the article and say so.