Blog
I thought the hardest part of building an AI Maths marker would be OCR. I was wrong.
I thought the hardest part of building an AI Maths marker would be OCR. I was wrong.
As an A-level Maths examiner, I've spent a fair amount of time looking at AI exam-marking tools. I've even tried building one myself, partly to understand what was actually possible.
The hardest part wasn't handwriting. It wasn't OCR. It wasn't even getting the final answer right.
It's that UK maths exams are designed to reward any valid mathematical method a student uses, not just the method the mark scheme's authors happened to write down.
That makes AI marking much harder than it first appears.
Pearson's own marking guidance tells examiners to:
"award credit for any valid mathematical method, unless the question specifies that a certain method must be used."
The published solution in a mark scheme is the "most frequently seen" approach, not an exhaustive list of everything that counts as correct.
The guidance even anticipates unusual cases.
A Further Maths technique used in a standard Maths paper can receive full credit. The non-standard "DI" shortcut for integration by parts can also be accepted if it is used correctly, even though it isn't the method most textbooks teach.
Examiners can also mark a question "as a whole", giving credit for correct working regardless of how a student has labelled or structured it.
That's a system built around a human recognising mathematical correctness.
It's very different from checking whether an answer matches a predefined rubric.
Then there's follow-through marking.
The "correct" answer for a later part isn't always fixed. It can depend on what the student did earlier.
If a student makes an error in part (a) but then applies a valid method to that incorrect value in part (b), they can still earn method marks in (b).
So the examiner isn't simply checking an answer against a key. They're interpreting the student's reasoning in context.
And this isn't just a concern raised by AI skeptics.
In January 2026, Ofqual described current AI systems as "black boxes" that can be difficult even for experts to audit, and said they "lack true semantic understanding." It also highlighted the substantial differences between what is needed to mark a maths exam and an English exam.
So if a tool claims to mark a mock paper "like an examiner", don't just ask whether it gets the final answer right.
Ask:
What happens when a student uses a method the mark scheme's authors never wrote down?
What happens when an early error carries through to a later part, but the student should still receive method marks?
What happens when two students arrive at different intermediate values, but both have used valid reasoning?
If the answer is simply that the system checks final answers against a stored key, it isn't marking like an examiner.
It's checking like a calculator with extra steps.