The Long Way in Port Loko
In late June, Google DeepMind and Fab Inc released the results of a preregistered randomized controlled trial in twelve junior secondary schools in Port Loko District, Sierra Leone. Seventeen hundred sixty-three students, eight weeks, math. The tool at the center of the study gave a direct answer in two percent of its messages. The other ninety-eight percent, it asked. That is the trial. That is the finding.
In late June, Google DeepMind and Fab Inc released the results of a preregistered randomized controlled trial run last fall in twelve junior secondary schools in the Port Loko District of Sierra Leone.1 Seventeen hundred sixty-three Grade 7 and Grade 8 students. Forty-eight classrooms. Eight weeks of math, between October 6 and December 5 of last year.2 The intervention arm used Gemini's Guided Learning mode alongside their teachers. The control arm used the standard curriculum.
The topline is easy to state. On the math assessment at the end of the trial, the treatment group scored 0.258 standard deviations higher, an intent-to-treat effect the authors describe as worth between roughly one and two years of typical learning gain.1 For students who reached the study's requested twelve hours of tool use, the effect climbed to 0.380 standard deviations. Sixty-nine percent of those assigned to the tool made it past the twelve-hour threshold, a share that compares against something closer to five percent for the average piece of voluntary educational technology.1
Those are the headlines. They will get repeated in trade press for a week. The number worth staying with a little longer is a smaller one, tucked further inside the paper. Of the more than one hundred thirteen thousand exchanges the students had with the tool over those eight weeks, ninety-one point four percent were about building an understanding of the math, not asking for the answer to it.1 And of the tool's own replies, two percent were direct solutions. The other ninety-eight percent were something else. Mostly, they were questions.1
What the Trial Made the Tool Do
In seventy-six percent of its messages, the tool posed a scaffolding question rather than gave an answer.1 What do you notice about this side of the equation. Have you tried the smaller case first. Where does that step come from. It is the shape of a good tutor's Tuesday afternoon, rendered as a language model system prompt. And the trial's finding is that when a tool of that shape is placed inside a working classroom, with a teacher still in charge of the lesson, the students meet it more or less halfway.
They did not, mostly, try to game it into producing the solution. They did not, mostly, treat it like a search engine with better manners. Nine of every ten exchanges reached toward understanding. Under three thousand of the more than one hundred thirteen thousand messages the tool sent were the sort of instant answer that has quietly become the whole story of AI in schools everywhere else.1
Why the Shape Matters More Than the Score
The paper's authors are careful about generalization. Sierra Leone is not New York. Port Loko District is a place with specific teachers, specific schedules, specific screens.3 The intervention was pre-registered, integrated with teacher-led instruction, and shaped by researchers who had done this before. Almost every failure mode of the chatbot-in-school stories the field has been reading for two years, the students copy-pasting the prompt at midnight, the parent who finds an unbelievably neat essay, the teacher who cannot tell what the child did or did not do, was engineered out of the study by design.
That is not a criticism. It is a description of what the trial actually proved. When a tool built to teach is used in a room where a teacher is teaching, and the tool is measured on whether the student learned rather than whether she finished, the tool does what its designers said it would do. The corollary is heavier. When a tool built to complete is used in a room where a teacher is grading completion, the tool also does what its designers said it would do. The tool is not the story. The shape of what the classroom decided to measure is.
What the Ninety-Eight Percent Actually Was
A tool that gives a direct answer two percent of the time is a tool that is willing to make a student slower. It is willing to hold a twelve-year-old for another minute, another question, another rephrasing, before letting her move to the next step. That patience, when the trial happened, produced a two-year gain in eight weeks for students whose teachers wove it into about half their lessons.2 It also produced something less measurable, and in some ways heavier. Every one of those one hundred thirteen thousand exchanges left a trail. What the student was working on. What she had misunderstood in the previous step. Which scaffolding question caught, and which one did not. That trail is, for the teacher who wants to read it, the entirety of the pedagogical record.
The reason the study is interesting, past the effect size, is that it is one of the few classroom AI deployments in the world where the record of what the student was doing was legible by design. Not extracted after the fact. Not lost to a closed tab. Sitting in the tool's own log, an eight-week diary of how each child got from where she was to where she went.
A Future Where the Long Way Is Available
Most classroom AI in the United States, and in most of the countries UNICEF surveyed last week, is not shaped like this.4 It is shaped like the two percent of Guided Learning's messages that gave a direct answer. Fast, complete, clean. That shape has produced a specific set of numbers. Students perform better in the short run and worse when the tool is taken away. Teachers report eroding trust. Parents report neat essays their children cannot explain. The trial in Port Loko is a hint that the numbers on the other side of the ledger, the ones about learning that sticks, come from tools willing to be slow and rooms willing to notice they were.
The obstacle for most schools is not the tool. It is the record. A teacher without a record of the process cannot tell the difference between the child who took the long way and the child who was handed the answer. Both hand in the essay. The Port Loko trial's quiet gift is that it built the record into the tool itself, and then chose to measure the child rather than the paper. Most schools this fall will inherit tools that do not do either. They will get to decide, one classroom at a time, whether the record is worth building anyway.
Which of your Tuesday afternoons this fall will still have a record of the long way?
References
Measuring the impact of learning with AI in Sierra Leone and beyond
Google DeepMind · June 2026
Google Gemini Guided Learning raises math scores in Sierra Leone classroom trial
EdTech Innovation Hub · June 2026
Rapid Evidence for AI in Education: An RCT Playbook Sierra Leone
Google DeepMind and Fab Inc · June 2026
Google Tested Its AI Tutor In Real Classrooms. It Worked
Forbes · July 2026
Sources cited in order of appearance. Click any inline number to jump.