Score calculator
The grader, graded

We tested Maro against the real AP readers.

Every year, College Board publishes real student essays together with the scores official AP readers gave them. We ran Maro's grader blind on 148 of those responses — AP U.S. Government FRQs and AP U.S. History SAQs and LEQs — seeing only the rubric, the question, and the essay, then compared, rubric row by rubric row. Here is what came back.

99%
within one rubric point of the human reader
85.6%
exact same score, row by row
0.5
average total-score gap, out of 3–6 points
148 officially scored responses · 497 rubric-row scores · 2021–2025 released exams · each essay graded three times, majority vote
AP U.S. History alone: 90.5% exact, 100% within one point (72 responses)

AP U.S. Government, row by row

Exact agreement with the official reader per rubric row, and the direction Maro leans when it differs. Positive lean = more generous than the human; negative = stricter.

Rubric rowRows scoredExact matchLean
Shared constitutional provision (SCOTUS)20100%0.00
Part (a) (Concept Application)2692%+0.08
Claim / Thesis (Argument Essay)2889%+0.11
Part (c) (Concept Application)2685%+0.15
Alternate perspectives (Argument Essay)2685%+0.15
Part (b) (Concept Application)2681%+0.12
Case comparison (SCOTUS)2078%+0.08
Concept link (SCOTUS)1982%+0.19
Reasoning (Argument Essay)2662%+0.23
Evidence (Argument Essay)2857%+0.39

AP Gov: mean of two full evaluation runs on the adopted grader configuration, August 2026. Per-row percentages vary ±4 points between runs at these sample sizes.

AP U.S. History, row by row

Rubric rowRows scoredExact matchLean
Thesis (LEQ)3697%−0.03
Part (a) (SAQ)3694%0.00
Part (b) (SAQ)3694%+0.06
Part (c) (SAQ)3694%0.00
Contextualization (LEQ)3694%0.00
Analysis & Reasoning (LEQ)3681%−0.08
Evidence (LEQ)3678%+0.11

One caveat we'd rather state than hide: the 2024 sample essays are handwritten scans that were machine-transcribed before grading, and agreement on them (82.5%) runs below the typed 2025 essays (98.4%) — some of that gap is transcription noise, not grading error.

How the test worked

Production-honest: the grader saw exactly what it sees when it grades your practice essay, meaning the official rubric, the question, and the response. It never saw College Board's scoring notes, acceptable-answer lists, or the human reader's score. Each essay was graded three times and the majority answer per rubric row was used, so a lucky (or unlucky) single run can't flatter the numbers. The anchor set spans five exam years, both released question sets per year, and the full score range: the top, middle, and bottom sample at each question.

Which way Maro leans

Read the lean columns above. On AP Government, not one row is negative — when Maro disagrees with the official reader there, it is more generous, never stricter. The gap is widest on the Argument Essay's evidence ladder (+0.39) and its reasoning row (+0.23), which are also the two rows it agrees on least often. On AP U.S. History it sits closer to even, and on thesis (−0.03) and analysis (−0.08) it is fractionally harsher than the human.

We did choose the stricter of the two graders we built. We ran a more lenient configuration against the same anchors and it agreed with the human readers less often, so we kept this one — a student coached a point low is better prepared than a student flattered a point high. But "stricter of the two we tried" is not the same as strict, and the tables are the honest version: on AP Government especially, if Maro and a real reader disagree about your essay, the real reader is more likely to be the harsher of the two. Plan for that.

What this doesn't mean

Maro's read is a practice instrument, not an official score, and it never predicts what a particular reader will give a particular essay in May. The weakest row is the Argument Essay's evidence ladder: deciding whether evidence is "specific and relevant" is genuinely the hardest judgment on the rubric, for humans too, and we publish that number rather than hide it. As the anchor set grows, these figures will be re-run and updated, whichever direction they move.

This is the grader behind every note Maro leaves in your margins.
Meet your reader