logo

Accuracy

Accuracy

Accuracy You Can Trust

We believe in transparency. Rather than making vague claims, we publish real accuracy data — tested against official examiner marks — so you can decide for yourself.

Verified against official IGCSE & GCSE examiner results

Last updated 15 July 2026

Quadratic Weighted Kappa — the gold standard
0.97
The metric exam boards use to measure marker agreement. Above 0.80 is near-perfect. The accepted threshold for AI marking systems is 0.70. Graded Pro scores 0.97 across 640 questions spanning maths and English, marked against the official mark scheme and compared with the marks the real examiner awarded.

How We Compare

Metric Graded Pro Human Markers*
Quadratic Weighted Kappa 0.97 Varies by subject
Correlation with examiner 0.97 ~0.70
Marks identical to the examiner 84% Not published
Average error — structured questions 0.19 marks Not published
Average error — levelled questions ~1.7 marks Not published
Average error — full-length essays ~4.3 marks 5.6 marks
Average bias vs the examiner +0.04 marks/question Not published

*Human marker comparison from a Cambridge Assessment study in which scripts marked by a chief examiner were independently re-marked by experienced markers. We show it only against our full-length essay figure, since that is the closest like-for-like comparison; we do not claim an equivalent human benchmark for structured questions. Graded Pro results are based on 640 questions across Cambridge IGCSE Mathematics 0580 Papers 2 and 4 (Oct/Nov 2025, 13 students, 561 questions) and Edexcel GCSE English Language Paper 2 (1EN0/02, Summer 2019 published exemplar scripts, 79 questions), compared against the marks awarded by the official examiner. The only inputs were the students' work and the official mark scheme. No mark schemes or student work were adjusted in any way.

Results by Subject

All results are from real examination papers, compared against the actual marks awarded by the official examiner. The only inputs were the students' work and the official mark scheme — nothing was adjusted or modified.

Mathematics
IGCSE 0580 · Papers 2 & 4 · 561 questions
QWK 0.98
Exact match 84%
Within ±1 mark 97%
Within ±2 marks 99.5%
Average error 0.19
Correlation 0.98
Aa
English Language
GCSE Paper 2 · 79 questions
QWK 0.97
Exact match 68%
Within ±1 mark 85%
Within ±2 marks 91%
Average error 0.82
Correlation 0.98

English combines short-answer questions with extended writing, so its headline exact-match figure is lower than maths by nature: a 40-mark essay is far less likely to land on the examiner's exact number than a 2-mark calculation. The breakdown below separates the two.

Structured Questions

Our system excels on questions with defined correct answers — the kind that make up the majority of assessments. Across 609 structured questions in both maths and English:

0.98
Quadratic
Weighted Kappa
Near-perfect agreement
85%
Marks identical
to the examiner
609 questions
97%
Within ±1 mark
of the examiner
Across subjects
0.19
Average error
per question
Across subjects

Whether it's a 1-mark calculation or an 11-mark multi-step problem, the AI consistently matches professional marking standards.

Extended Writing & Essays

Levelled questions — where markers use band descriptors to assess quality — are harder for any marker, human or AI. Our system uses a structured levelling process modelled on how trained markers work: identify the best-fit level, then position within it. This is where we are most careful about what we claim.

  • Average error of around 1.7 marks across all levelled questions (31 questions, tariffs from 6 to 40 marks)
  • On full-length 40-mark essays, average error rises to around 4.3 marks — these are the hardest items to mark, for any marker
  • The system under-marks the strongest essays. On top-band responses the gap can reach 7–11 marks. This is a known limitation and the one place we would not use the AI mark unreviewed
  • Every essay-level figure above is drawn from a small sample (7 full-length essays), so treat it as indicative rather than precise

What This Means For You

AI marking is not a replacement for your professional judgement — it's a tool that handles the heavy lifting so you can focus on what matters.

Where the AI is strongest

Short-answer questions, calculations, retrieval tasks, and structured responses across all subjects. On these question types, the AI matches the examiner on roughly 85% of questions and lands within a mark on 97% — reliable enough to use as a first pass and review by exception.

Where you should review

Extended writing and essay-style responses, particularly at the top of the mark range. The AI reliably under-marks the strongest essays — by 7 to 11 marks on a 40-mark task in our testing. Always review your highest-performing students' extended writing before returning it. We would rather tell you this than let you find it out.

What we recommend

Use AI marking to get a fast, accurate first pass across a full class set. Moderate a sample — just as you would with any marking — and pay particular attention to the top end on essay tasks. Teachers who use this approach typically report saving 50–70% of their marking time.

Works With Any Mark Scheme

Our system isn't locked to specific subjects or curricula. Upload any mark scheme or rubric and the AI adapts — whether you're marking a maths paper, a history essay, a science practical write-up, or a language analysis.

Not Just Exams

Our accuracy benchmarks are based on formal examination papers, but Graded Pro is built for everyday marking across all types of student work. The same AI that matches chief examiner standards on exam scripts delivers consistent, rubric-linked feedback on:

  • Homework — weekly assignments marked and returned the same day, with actionable next steps
  • Classwork and in-class tasks — quick, consistent feedback while the learning is still fresh
  • Termly tests and mock exams — full cohort marking with detailed breakdowns by question
  • Coursework drafts — formative feedback that helps students improve before final submission
  • Past paper practice — students get instant, exam-standard feedback on every attempt

Wherever there's a rubric or mark scheme, Graded Pro delivers accurate, detailed feedback — whether the stakes are high or the goal is simply helping students learn from their work.

Our Commitment

We continuously test and improve our marking accuracy. We don't claim perfection — no marker, human or AI, achieves that. Our current figures come from re-testing every paper against the official examiner marks each time we change the underlying model, and we publish the weak points alongside the strong ones. Where our sample is small, we say so. Where the AI is unreliable — top-band extended writing — we tell you to check it. The figures on this page were last re-tested against official examiner marks in July 2026.

See For Yourself

Sign up for a free trial with 150 free credits and test it on your own papers.

Start Free Trial

No credit card required