Agents are good at writing code. They're bad at knowing what you meant. So when your ticket leaves something out, the agent just picks something, and you find out later — in review, or in production.
Paste what you were about to send it. You'll get a score and a list of the things you never actually decided.
Evidence found in your words
Nobody decided these. Your agent will.
Thirty real tickets, worst first. The number on the left is what it scored. Use this puts the text in the box above so you can run it yourself; Details shows why it scored that. Or just select the text and copy it.
The model gets one job: find evidence in your text and quote it. Then we check the quote is really there. If it isn't, we bin it and count that signal as missing. So a model that makes things up can't earn you points.
The numbers come from the weights below. We never ask the model for a score. On thirty tasks we labelled by hand it agrees with us about 91% of the time, and that's a free model doing the reading.
Each axis is out of 100. The headline is the average of the four, rounded to the nearest 5. It isn't accurate enough to deserve a sharper number.
The look of this page is lifted from dontpastetheai.com by khaosdoctor. Theirs is better. Go read it.