About MeasuredRun

MeasuredRun publishes what actually happens when you run AI tools and workflows — the token counts, the wall-clock times, the bills, and the parts that broke.

Most writing about AI tools describes what a tool claims to do. This site is a record of what happened when we ran it. Every article comes from a session we actually ran, and the numbers in it come out of that session’s logs rather than out of a product page.


Why this site exists

We got tired of reading tutorials that had clearly never been executed.

If you are choosing between two coding agents, you do not need a feature table — both vendors publish one. You need to know what the same task cost in each, how long it took, how many attempts it needed, and what went wrong on the way. That information exists only if somebody runs the task and writes down what happened.

So we run the task and write down what happened. That is the whole idea.

We publish for people who are about to spend money or time on a tool and want a number before they commit. If a post here saves you an afternoon or a hundred dollars, it has done its job.


Who runs this

This site is written by Stig, who is also the person accountable for every number on it. Stig is a pen name; one real person stands behind it, and that person reviews and signs off every article before it publishes.

Here is the honest scope of what that involves — and it is deliberately narrow, because everything in it can be checked against this site itself:

  • We operate a small department of AI agents, and we measure tools with it. The agents run the test sessions, capture the logs, and draft. Stig checks every figure against the raw record and signs off. How that works is set out in full on the How We Use AI page.
  • We publish the record, including the failures. The measurements behind an article are not summarised out of existence; they are the article.
  • The tools we write about are tools we use for our own work. They are not vendor-supplied review units and we did not ask anyone for access.
  • We hold no certifications in this field, and we are not going to pretend to. What we have is the log files.

We are not going to tell you about careers, employers, degrees or credentials, because none of that is what makes a measurement true. What makes it true is that the run happened and the log exists. Judge us on the logs.


How an article gets made

Every article here is produced the same way, and we would like you to be able to audit it.

1. We run the thing. A working session is set up for the specific question — comparing two tools on one identical task, measuring what a feature actually costs, building the thing the tutorial describes. The purpose of the session is the measurement, not the article.

2. We keep the raw record. Each article has a stored evidence set behind it containing the unedited terminal and API output, a structured record of the measurements (model and version, timestamp, input size, tokens in and out, elapsed time, number of attempts, cost), a list of everything that failed, screenshots, and the environment the test ran in.

3. We publish the failures. Every article has a section describing what went wrong. This is not modesty. A successful walkthrough can be written by someone who never opened a terminal; the specific way a thing breaks at 11pm cannot. If a session produced no failures worth reporting, we treat that as a sign the test was too shallow, and the article does not ship.

4. A person reviews it before it publishes. AI assists with drafting and editing. A human reads the draft against the logs, checks that every number matches the record, removes anything that overstates what we actually did, and signs off. Nothing publishes without that step.

5. We cap how much we publish. No more than one article a day, five a week. Not because volume is penalised, but because we cannot honestly measure faster than that. If a week produces three well-evidenced articles instead of five, we publish three.

The full disclosure of how automation and AI are used here is on the How We Use AI page. We would rather you read it than guess.


What we will not do

  • We will not review a tool we have not run. If we have only read the documentation, the article says so, plainly.
  • We will not publish a number we cannot trace to a log. If we estimated something, the article says “estimated”. If we could not confirm something, the article says “unconfirmed”.
  • We will not quietly change an article after publication. See Corrections.
  • We will not accept payment to review something favourably. If we ever accept a product, a free licence or any other consideration, it will be disclosed at the top of that article, before the content, not in a footer.

How this site is paid for

We think you should know who benefits when you read.

  • Advertising. This site does not carry advertising, and we have not applied to any advertising programme. We intend to apply to Google AdSense later. If ads begin to appear, Google will select them; we will not choose them and we will not be told in advance what will appear. Seeing an ad for a product on this site will not be an endorsement of it. This page will be updated the day any of that changes.
  • Affiliate links. This site does not currently carry affiliate links.

If that changes, it will be stated here and on the Affiliate Disclosure page before any such link appears.

  • No sponsored posts. Nobody has paid for an article on this site.
  • No paid product placements inside articles.

Advertising revenue does not depend on what we conclude about any tool, and we have deliberately kept it that way.


Corrections

We get things wrong. When we do:

  • We fix the article and add a dated note at the bottom saying what was wrong and what changed. We do not silently overwrite.
  • If the error changed the conclusion, we say so at the top of the article, not in a footnote.
  • If a tool has changed since we tested it, we mark the article with the date of the test rather than pretending the measurement is timeless. A benchmark without a date is worthless.

If you find an error, tell us. Send the article link and what you think is wrong — ideally with your own numbers. Corrections from readers are the cheapest quality control we have, and we would rather be corrected in public than be wrong in public.


Where else to find us

  • GitHub — https://github.com/measuredrun

That is where our code and working files go when we publish them.


Contact

Questions, corrections, and “your numbers don’t match mine” are all welcome at havegoodin7@gmail.com, or through our Contact page. We read everything, and we answer anything that is answerable.

We do not accept guest posts, link insertions, or paid placements, and we do not reply to those requests.