All essays
Future of EducationAI in SchoolsLearning VisibilityEquity

The Bar That Moved

On June 18, Education Week described a Stanford study that ran six hundred identical eighth-grade essays through the language models behind most school-facing AI tools, changing only a short demographic label at the top of each one. The same paragraph came back with different bars. The researchers gave the pattern a careful name. Marked pedagogies.

June 19, 20266 min readKoan Team

On June 18, Education Week described a Stanford study that should, by rights, slow down a generation of education-technology buyers for a moment.1 The researchers, Mei Tan and Lena Phalen, took six hundred eighth-grade persuasive essays from a public research dataset and ran each one through the four large language models that sit underneath most of the school-facing AI tools on the market. The essays were identical. The only thing that changed was a short metadata line at the top of each one describing the writer. Race. Gender. Language status. Achievement level. A disability label. Sometimes nothing.2

The same paragraph came back with different bars.

A student described as high-achieving was prompted to consider counterarguments. A student described as struggling, on the very same text, received a note about spelling and a rewrite suggestion, with no invitation to think harder.1 Essays attributed to Black students drew praise that leaned toward leadership and power, with less of the specific critique that helps a young writer get sharper. Essays attributed to students learning English came back with feedback that was intensely corrective, sometimes harsh. Essays attributed to female students were more often answered in the first person, with warmer phrasing.2 The authors gave the pattern a careful name. Marked pedagogies. The instruction did not depend on the writing. It depended on who the model thought was doing the writing.

The Quiet Sin

A teacher who has been at this long enough will tell you that the great quiet sin of teaching is not cruelty. It is lowering the bar without noticing. It is the small, considerate kindness that decides this particular child cannot quite be asked the harder question today. It is the comment in the margin that praises before it pushes. Adults do this. We have always done this. Tan and Phalen point out, gently, that the models are not inventing the pattern. They are reproducing it at scale.3

What is new is the scale and the silence. The lowering of the bar is now happening inside tools that sit between every student and every paper. The student does not see the version of the comment another student would have received. The teacher does not see the prompt the model wrote, only the finished work, and rarely even that in its earlier shape. The administrator sees a tidy number rolling up to a quarterly review. The unfairness folds into a polished invisibility.

What the School Can Still See

The shape of school, ten years from now, will be defined less by which AI tools we let in than by what we can still see after we let them in. A faculty that cannot tell whether a struggling student was asked the same question as the strong student is no longer running a school in the older sense of the word. It is processing throughput.

The corrective is not to ban the tools. The corrective is to insist on the record. Every prompt the student saw. Every prompt the AI saw. The model's reply. The student's revision. The pause before the revision. The moments she rewrote in her own voice and the moments she pasted. None of this is private. None of it is mysterious. It can be captured. The only question is whether the school decides it would rather know.

Some districts are already deciding. The same season as the Tan and Phalen paper, the broader Stanford SCALE 2026 review reminded the field that the causal evidence behind most classroom AI claims is still thin and short-term, and that the gap between adoption and evidence is widening.4 The honest response is not faster adoption or slower adoption. It is more careful seeing. A tool can earn its place in a classroom if its work can be looked at. It cannot earn its place by promising results that nobody is in a position to verify.

A Wider Record

The reason to make learning visible has not changed in twenty centuries. A teacher who can see the work as it is being done can teach to it. A teacher who can only see the finished page is teaching to a ghost. What is different now is that the page itself is less reliable than it used to be, and the unfairness of what came before the page is now systematic and quiet. The way to honor every student is to keep the record so wide that nothing important can hide inside it. Including the question we did not ask her.

If the same essay can earn two different bars depending on the name at the top, whose job is it, in your school, to notice when the lower bar gets handed to her?

References

  1. AI Changes Its Feedback on Students' Writing When It Knows Their Race, Gender

    Education Week · June 18, 2026

  2. Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback

    Proceedings of LAK26: 16th International Learning Analytics and Knowledge Conference · 2026

  3. AI gives more praise, less criticism to Black students

    The Hechinger Report · 2026

  4. Understanding the Evidence Base on AI in K-12 Education

    Stanford SCALE Initiative · March 2026

Sources cited in order of appearance. Click any inline number to jump.