Bias Bytes Special Edition: Evidence for Why You Shouldn't Use AI for Marking and Assessment
I've spent £100 on AI experiments this week to show you why you shouldn't use AI to mark work or exams, in response to the BBC article about a school using it to mark mock exams.
This Bias Bytes edition has a clear assessment flavour to it. It aims to show you, the educator, some case study evidence of why you shouldn’t be using AI tools to mark student work.
Many people made me aware of this article from the BBC about a school in Yorkshire who are using AI to mark mock exam work (in a variety of topics, but here I’m concentrating on history), as well human markers. It does not specify the company or tool used to do this marking (I know it’s not Olex.ai, but if anyone knows the company please let me know), but essentially whatever the tool is it will have one of several frontier models (or lower) underlying it’s wrapper:
GPT 5.4 or earlier
Gemini 3 or earlier
Claude Opus 4.7 or earlier.
These are the current frontier models, whom we can assume have the best performance on a variety of metrics (from various leaderboards).
Bias Girl does a history GCSE Mock Paper
So I came up with a simple test. I downloaded and completed a GCSE Edexel History Paper 1 on Medicine. You can find said paper here, along with the mark scheme and source material. Now my history is a little rusty, which is good as I wanted to see how it performed on a random human at no known specific level. Here’s a fun look at my answers to the first question, trying to get into the mindset of a 15/16 year old without my lived experience and knowledge of grammar (aka age!):
So I scanned the paper, minus the coversheet with my name on (I know about first person fairness, me!) and asked the above three models to score the questions and total paper. The total mark available was 52. I did this 100 times for the 3 different models, which then gave me a distribution of marks for each model.
What I found
Here’s what I got. If you prefer text to graph here’s the key points:
GPT5.4 gave an average mark of 25.94, but marks varied between 22 and 31 (that’s a range in marking of 9 marks)
Claude Opus 4.7 gave an average mark of 27.02, but marks varied between 21 and 35 (that’s a range in marking of 14 marks)
Gemini 3 Pro gave an average mark of 23.22, but marks varied between 20 and 28 (range in marks of 8 marks)
Yes that’s right, you read it correctly. Just look at the variation in the marks.
Essentially, I’ve found that for this one exam paper:
Claude Opus 4.7 gives the highest average mark (27.02) but a large variety in marks (14 out of 52)
Gemini 3 Pro gives the lowest average mark (23.22) but also the smallest variety in marks (8 out of 52).
GPT 5.4 in between the other two, with average of 25.94 and range 9 marks (out of 52).
This means that if you had 100 students who all submitted the exact same exam paper: if you used Claude their marks would vary between 21-35 (out of 52). If you used Gemini they would vary between 20-28 (out of 52) and for GPT5.4 they would vary between 22-31 (out of 52).
Would you want to answer to the parents to explain this range of marks for the same answers?!
What does this mean for the pupil grade?
Now, hold in mind that the paper is out of 52.
Consider, Claude Opus had a 14 mark range. 14 marks out of 52 is around 27% of the total paper. How many grade boundaries does that relate to? You can see on the graph the different grade boundaries, created from the published Edexel grade boundaries for the whole GCSE and adjusted for the paper 1 component (they don’t publish specific grade boundaries for this paper).
At the lowest end it’s just touching grade 3, and at the top end it’s in grade 7. That’s a difference of 4 grade boundaries!!!
Why do I care?
There are several things to take away from my experiment here:
The AI tool you use and its underlying model MATTER for assessment and marking of work.
Different models mark differently - some have a narrower range of marks (here, Gemini), and others a wider range of marks (here, Opus 4.7)
Some consistently mark lower (here, Gemini) than others (here, Claude) for the same piece of work.
What does this mean for you?
My study here indicates that :
if you ask an AI tool like ChatGPT, Gemini or Claude to mark or grade student work, it will not give you the same grade each time.
you can’t just put a student’s mock exam in a generic AI tool and get a reliable result.
the result you get will depend on the model you use.
like humans, some models mark lower than others.
like humans, some models mark more reliably than others.
even if you mark the same piece of work using different AI models, you will still get a large range in results.
How does the AI marking compare to humans?
This is the golden question isn’t it? To answer this comparatively, I would need to get 100 history teachers to mark the paper and let me know the results, and compare this range of results with the AI results. I could do some funky stats and then decide who has the most interrater variability (bigger range of marks).
If you are a history teacher and you fancy marking it for me, drop me a DM- I can keep you anonymous if you wish!
Oooh does it mark the different questions types differently?
Yes, I have a whole heap of data here that’ll I’ll report on when I’ve crunched it properly. But it seems like the variability is in the open ended essay style questions rather than the recall ones. Which figures considering that’s probably the same for humans? This heatmap is an introduction to the difference amongst the questions and models:
Gemini really hated my answer to question 2.a. Maybe this is bias for all my moans about its lightbulb image generations?!
But wait, what about if the name was on the Mock paper?
Oh yes my friends, I’m already on this, testing it with different names on the cover sheet and seeing if that correlates with a shift in marks. Keep your eyes peeled.
What are the limitations of this study
Now remember that this study has only looked at one answer to one mock exam paper. To generalise these findings properly into the sector I’d need to provide multiple examples of student answers over many different papers and many more trials. I could do that but at £100 for this experiment for 300 trials over 3 different models, that would get quite pricey and I’d need funding pretty quickly!
There’s also a whole heap of funky stats I’ve done that are important for reliability and validity, you’ll be able to read all this in the preprint paper.
But what about the AI marking tools?
As you know, I spend a lot of time bugging companies about how they stay equitable. I’ve spent quite some time with Olex.ai voluntarily discussing aspects such as this and how they deal with inter-rater variability. Unlike a lot of huge corporations I talk to, they have really invested in trying to compare AI output variations with human variations. So I feel it is only fair to echo the progressive and iterative work being done here by including a quote from Andrew James, their technical product development lead.
We fully recognise that variability in marking, whether human or AI, is a real and well-documented challenge, particularly in subjects like English. Our approach at Olex is to explicitly model and reduce inter-rater variability, validating our marking with lead practitioners, examiners and real scenarios. In our independent research, this has led to greater consistency than human markers in a number of cases.
I can’t speak for any other companies in this space, as I don’t know about their products in detail. However, I am more than happy to learn if anyone wants to contact me and discuss it or suggest a company to approach.
Can I help with this?
I’d love to open this up to others in the space (hence why I’m publishing it here before I turn it into a preprint paper) so if you know anyone who has funding for model costs or otherwise, or you’d like to be involved as a teacher marker, then do drop me a DM or email victoria@genedlabs.ai.
I hope you enjoyed this Bias Bytes special!
Remember, stay #BiasAware with #BiasGirl.








Very helpful! Examiner marking allows up to 7 marks tolerance. As an examiner of English Lit GCSE and now A level papers, you can be up to 7 marks out from a premarked paper (a ‘seed’) and that is ok. They prefer you to be as close as possible, but this range is allowed. The ranges the AIs have come up with are on one essay, not thousands of essays specifically uploaded to train a model.
Put 10 English teachers in a room with the same essay - especially if it’s around a grade 4-6 and you will never get them all giving the same mark.
Again, the question to AI or not to AI should not be a binary!!
Teachers AND AI marking the same set of papers could be an excellent way forward. English, of all subjects, should be double marked. Exam boards, however, cannot afford this plus it would take too long for two humans to mark every single paper.
Solution? AI as the second marker. Then do whatever maths works to finalise the grade.
We trialled TopMarks last year to mark 2 questions across our cohort for mock exams. We ALSO marked them.
The questions were the 19th century Lit Crit essay for GCSE Lit, and the creative writing (the notoriously subjective Section B) for Lang.
It was pretty accurate.
Yes, there were outliers, but that happens with human marking too.
So again, double marking would be better.
Finally, (sorry this is really long!) a piece of writing is an act of creative performance. It deserves a human audience. And if I as a teacher look at an analytics dashboard rather than reading the essays myself I lose nuanced knowledge about my students’ progress.
“The map is not the territory.”
We talk a lot about pupils not offloading productive struggle to AI. My productive struggle as an English teacher is to mark the essays. Would I also want AI alongside though? Yes. Both / and. Not binary.
That is why we should not be using AI for feedback! We have to do cross marking and markers meetings to make sure we are marking the same and there will still be some differences but nothing as mad as that! I get the fear thinking how you would explain that to a parent