
New Monash Business School research has uncovered a troubling flaw in human oversight of AI – and the experts least likely to challenge the error may surprise you.
In courtrooms, boardrooms and classrooms around the world, AI systems are increasingly being trusted to make decisions that can alter the course of our lives.
The technology is deployed on the assumption that human experts will catch and correct any errors.
But a new study suggests that belief could be dangerously optimistic.
In a randomised experiment involving more than 1,300 teachers, researchers found that participants were significantly less likely to correct an unfairly harsh grading error when they believed it had been made by an AI system rather than by another person.
Published in PNAS Nexus, the findings offer some of the first experimental evidence of how expert oversight of AI performs in practice.
“As AI gets built into more and more decisions, the standard reassurance is always ‘don’t worry, a human will check the AI’s work.’,” Monash Business School Professor Rigissa. Megalokonomou said.
“We wanted to test whether that actually holds up, using teachers grading student work as our test case.”
Putting AI oversight to the test
In the study, co-authored by Sofoklis Goulas and Panagiotis Sotirakopoulos, teachers from Greece were randomly assigned to one of two conditions.
Both groups were shown the same student work and an intentionally incorrect grade.
One group was told a fellow teacher had assigned it; the other was told an AI system had.
The grade was either too harsh or too lenient, allowing the researchers to test whether the direction of the error mattered. It did.
Teachers corrected lenient mistakes at similar rates regardless of the source.
But when the grade was too harsh, the AI label substantially reduced the likelihood of correction, widening the grading fairness gap by 22 per cent.
“When it’s being strict, people read that as competence. When it’s being lenient, they don’t extend the same benefit of the doubt,” Professor Megalokonomou said.
“That’s a much more complicated, and frankly unpredictable, kind of trust than most people have in mind when they say, ‘don’t worry, humans will oversee the AI’.”
‘Counterintuitive and concerning’
One of the study’s most striking findings was who was most likely to defer to AI.
Younger teachers, those with postgraduate qualifications and those who described themselves as technologically confident were the least likely to challenge a harsh AI-generated grade.
“The very people we might expect to be the most capable and critical users of AI tools turned out to be the most likely to defer to a harsh AI grade,” Professor Megalokonomou said.
“That’s counterintuitive and concerning: the people pushing AI integration forward in schools may be the least likely to catch its errors.”
The implications extend beyond grading, with AI tools increasingly used for lesson planning, feedback, identifying struggling students, and administrative tasks.
“In all of these situations, the same dynamic we found in our study could apply: teachers deferring to AI outputs without sufficiently questioning them, and errors quietly going uncorrected,” she said.
“Multiply that across thousands of students and thousands of classrooms, and you start to see how quickly this becomes a serious problem.”
Preparing educators for an AI future
The researchers are now developing a teacher training program focused specifically on AI oversight and critical engagement with AI-assisted decisions.
“It’s not enough to just tell people AI can be wrong,” Professor Megalokonomou said.
“You need to show them specifically how and when their judgment is likely to go astray, and build the habits to push back on AI.”
The team is also pursuing a broader research agenda examining how AI affects other education decisions.
Future projects will investigate whether AI helps reduce or amplify the biases teachers bring to grading, whether AI training improves teachers’ day-to-day productivity, and whether simple, low-cost interventions can change how educators engage with AI tools.
“More immediately, I really hope this research reaches the people who are making decisions right now about AI in schools: policymakers, school leaders, education departments,” she said.
“The conversation around AI in education moves so fast, and the question of whether human oversight actually works tends to get assumed instead of being tested.”
The paper Why do experts miss AI’s errors? Evidence from a randomized labeling experiment appeared in PNAS Nexus, Volume 5, Issue 6, June 2026.


