We believe in transparency. Rather than making vague claims, we publish real accuracy data — tested against official examiner marks — so you can decide for yourself.
Last updated 15 July 2026
| Metric | Graded Pro | Human Markers* |
|---|---|---|
| Quadratic Weighted Kappa | 0.97 | Varies by subject |
| Correlation with examiner | 0.97 | ~0.70 |
| Marks identical to the examiner | 84% | Not published |
| Average error — structured questions | 0.19 marks | Not published |
| Average error — levelled questions | ~1.7 marks | Not published |
| Average error — full-length essays | ~4.3 marks | 5.6 marks |
| Average bias vs the examiner | +0.04 marks/question | Not published |
*Human marker comparison from a Cambridge Assessment study in which scripts marked by a chief examiner were independently re-marked by experienced markers. We show it only against our full-length essay figure, since that is the closest like-for-like comparison; we do not claim an equivalent human benchmark for structured questions. Graded Pro results are based on 640 questions across Cambridge IGCSE Mathematics 0580 Papers 2 and 4 (Oct/Nov 2025, 13 students, 561 questions) and Edexcel GCSE English Language Paper 2 (1EN0/02, Summer 2019 published exemplar scripts, 79 questions), compared against the marks awarded by the official examiner. The only inputs were the students' work and the official mark scheme. No mark schemes or student work were adjusted in any way.
All results are from real examination papers, compared against the actual marks awarded by the official examiner. The only inputs were the students' work and the official mark scheme — nothing was adjusted or modified.
English combines short-answer questions with extended writing, so its headline exact-match figure is lower than maths by nature: a 40-mark essay is far less likely to land on the examiner's exact number than a 2-mark calculation. The breakdown below separates the two.
Our system excels on questions with defined correct answers — the kind that make up the majority of assessments. Across 609 structured questions in both maths and English:
Whether it's a 1-mark calculation or an 11-mark multi-step problem, the AI consistently matches professional marking standards.
Levelled questions — where markers use band descriptors to assess quality — are harder for any marker, human or AI. Our system uses a structured levelling process modelled on how trained markers work: identify the best-fit level, then position within it. This is where we are most careful about what we claim.
AI marking is not a replacement for your professional judgement — it's a tool that handles the heavy lifting so you can focus on what matters.
Short-answer questions, calculations, retrieval tasks, and structured responses across all subjects. On these question types, the AI matches the examiner on roughly 85% of questions and lands within a mark on 97% — reliable enough to use as a first pass and review by exception.
Extended writing and essay-style responses, particularly at the top of the mark range. The AI reliably under-marks the strongest essays — by 7 to 11 marks on a 40-mark task in our testing. Always review your highest-performing students' extended writing before returning it. We would rather tell you this than let you find it out.
Use AI marking to get a fast, accurate first pass across a full class set. Moderate a sample — just as you would with any marking — and pay particular attention to the top end on essay tasks. Teachers who use this approach typically report saving 50–70% of their marking time.
Our accuracy benchmarks are based on formal examination papers, but Graded Pro is built for everyday marking across all types of student work. The same AI that matches chief examiner standards on exam scripts delivers consistent, rubric-linked feedback on:
Wherever there's a rubric or mark scheme, Graded Pro delivers accurate, detailed feedback — whether the stakes are high or the goal is simply helping students learn from their work.
We continuously test and improve our marking accuracy. We don't claim perfection — no marker, human or AI, achieves that. Our current figures come from re-testing every paper against the official examiner marks each time we change the underlying model, and we publish the weak points alongside the strong ones. Where our sample is small, we say so. Where the AI is unreliable — top-band extended writing — we tell you to check it. The figures on this page were last re-tested against official examiner marks in July 2026.
Sign up for a free trial with 150 free credits and test it on your own papers.
Start Free TrialNo credit card required