On airBulletin №20 · Empowering Youth: The STUDENTS FIRST Act Explained|Up next · The Alpha School Controversy
From the journal

Two Minutes a Week: Why Children Ignore Their AI Tutors

Published
6 September 2026

The word was “usually”.

A fourth grader in New Mexico, reading aloud into a headset microphone, stumbled over it. She knew what “usually” meant. She used it in conversation. What she could not do, in that moment, was get from the letters on the screen to the sound in her head: the collapsed middle syllable, the way the “s” turns into a “zh”, the fact that the word looks nothing like it sounds. That is a decoding problem. It has a specific pedagogical answer, and the answer involves breaking the word into parts and mapping the parts to sounds.

Amira, the AI reading tutor listening on the other end, offered her a definition instead.

Wendy Graham, the girl's mother, is a high school history teacher in Las Cruces Public Schools and a former fifth grade teacher. She had first encountered the platform in a summer reading programme in 2025 and later used it at home. She watched the software identify a vocabulary gap where there was none and miss the decoding failure that was actually happening. Her son, offered the same purple-haired avatar, simply refused to engage with it at all.

“Kids don't do their best for robots,” Graham told the education outlet The 74 in an article published on 2 September 2026. “Kids do their best for people.”

It is the sort of line that could be dismissed as sentiment. Except that in the same week, three separate strands of rigorous quantitative evidence arrived at approximately the same conclusion by entirely different routes, and none of them involved sentiment at all. They involved randomised controlled trials, log files, and a great many children who, given free access to the most heavily capitalised educational technology in history, used it for about two minutes a week.

Two Minutes a Week

The most arresting number comes from Stanford University's National Student Support Accelerator, whose researchers ran two randomised controlled trials with elementary students in two American school districts serving high-poverty populations. The paper, “Access is Not Enough: Human Support Improves Engagement with AI Tutoring”, was written by Carly D. Robinson, David Gormley, Ana Trindade Ribeiro and Susanna Loeb, and released as an Annenberg Institute working paper in June 2026.

The design was straightforward. Students were given access to an AI literacy platform, with scheduled time in which to use it. They were expected to complete at least two 30-minute sessions per week. The platform's own provider states that academic benefits typically begin to appear after around 30 minutes of weekly use. Half the students used the platform on their own; the others had a human tutor sitting with them, whose job was explicitly not instruction but engagement, motivation and troubleshooting.

In the independent-use condition, only 60.7 per cent of students in District A and 53.3 per cent in District B ever used the platform at all. Not “used it well”. Ever. Logged in once, across an intervention that ran between 14 and 31 weeks.

Average weekly usage was 2.18 minutes in District A and 5.23 minutes in District B. Against a target of 60 minutes. The students who did use it managed 13.2 and 25.8 minutes in the weeks they used it, which tells you the average is not describing a population of light users but a population of near-total non-users punctuated by occasional bursts. On average, students touched the platform in only four to five weeks out of an intervention lasting up to thirty-one.

Adding a human being to the room helped, and the size of the help is instructive. Engagement rose by between 71 and 80 per cent, which sounds transformative until you notice that weekly usage went up by one minute in District A and 4.4 minutes in District B. Over the whole intervention, the human tutors bought less than two additional hours of platform time per student. Reading achievement did not move.

“A key finding that we weren't even meaning to test,” Robinson told Chalkbeat, “is that having access to this AI tutor isn't the same as using it.”

Loeb, the centre's executive director, put the institutional conclusion more bluntly: “We don't have solid research showing that AI tutoring can work in the U.S. at scale.”

There is a detail buried in the paper's appendix that deserves more attention than it has received. Among students left to use the platform independently, those who engaged with it were more likely to be higher achieving and less likely to receive special education services. The children who might have gained most from extra reading practice were the least likely to open the application. A technology sold as an equaliser produced, in the only two districts where anyone bothered to measure it properly, a ladder that the students at the bottom did not climb.

The Binding Constraint

The second study is larger, longer and, if anything, more damaging to the optimistic case, precisely because the product worked.

Philip Oreopoulos of the University of Toronto and Nina Low of Charles River Associates ran a two-year cluster randomised trial across 18 middle schools in Hamilton County, Tennessee, during the 2024-25 and 2025-26 school years. Students in existing daily remedial mathematics sessions were randomly assigned to Khan Academy with Khanmigo, the platform's generative AI tutor, configured to coach rather than hand over answers. The paper, “One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment”, was published as NBER Working Paper 35620 in August 2026.

Assignment raised mathematics achievement by 1.3 national percentile ranks per term, roughly 0.06 to 0.08 standard deviations over a school year. The authors calculate that a full year of active participation would imply about 0.14 standard deviations. These gains, they note, resemble those from Khan Academy practice without any AI assistance at all.

Then comes the log-file archaeology, which is the real contribution. Ninety-six per cent of students tried Khanmigo at least once. The median student messaged it on only a third of the days they practised maths. And in the exercise sessions where a student actually made a mistake, the moment at which a tutor is theoretically most valuable, they consulted the AI in just 17 per cent of cases. The messages they did send were, in the authors' description, mostly bare answers or clicks on suggested prompts.

“Access was nearly universal,” the researchers wrote, “but engagement was thin.”

Their conclusion is the sentence the entire sector should be arguing about: “The binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access.”

Sal Khan had said as much himself, months before the paper landed. In April 2026 he told Chalkbeat that for a lot of students Khanmigo “was a non-event. They just didn't use it much.” Asked about the vision of an AI tutor permanently available in every classroom, he offered a four-word summary of the adoption curve: “Some will; most won't.” He added that while AI would help, “our biggest lever is really investing in the human systems.”

This is a founder describing the gap between his own 2023 TED talk, which promised a personal tutor for every child on Earth, and a log file showing that most children did not ask it anything.

The Strongest Version of the Optimistic Case

It would be lazy to stop here, because the optimistic case is not stupid and is not obviously wrong. It deserves to be built properly before it is tested.

Start with the economics. Human tutoring at high dosage is the best-evidenced intervention in education, and it is also ferociously expensive. The systematic review by Andre Nickow, Philip Oreopoulos and Vincent Quan pooled the experimental evidence on PreK-12 tutoring and reported an overall effect of 0.37 standard deviations, later revised to 0.29 in the version published in the American Educational Research Journal. Effects were strongest when tutors were teachers or paraprofessionals, in the earlier grades, and when the tutoring happened during the school day rather than after it. The Education Endowment Foundation's toolkit rates one-to-one tuition as worth up to five months of additional progress. Nobody disputes that tutoring works. The dispute has always been about whether anyone can afford enough of it.

An AI tutor has, in principle, a marginal cost approaching zero and infinite patience. It never has a bad morning. It is available at eleven at night, in a language the parents may not speak, to a child whose school has three vacancies in its maths department. If it delivered even a fraction of the human effect at a hundredth of the price, the cost-effectiveness arithmetic would be overwhelming.

And there is real evidence that it can. A World Bank randomised trial in Edo State, Nigeria, gave secondary students six weeks of after-school GPT-4 tutoring, working in pairs under teacher supervision with prompts designed to promote reasoning rather than shortcuts. The programme produced gains of around 0.3 standard deviations overall and 0.23 on English, the primary outcome, at roughly 48 dollars per student. Benchmarked against a database of education interventions trialled in developing countries, it outperformed about 80 per cent of them.

There is also a longer history that the current discourse tends to forget. Intelligent tutoring systems are not new. Carnegie Learning's Cognitive Tutor descends from decades of cognitive science at Carnegie Mellon. ASSISTments, developed at Worcester Polytechnic Institute, was evaluated in a large randomised trial across 46 Maine schools and produced an effect of about 0.18 standard deviations on an end-of-year standardised maths test, with the largest benefits for students with the weakest prior attainment. Meta-analytic reviews of intelligent tutoring systems have reported average effects in the region of 0.37 to 0.50 standard deviations depending on the comparison condition. These are not nothing. Measured against the typical education intervention, they are respectable.

Khan Academy's own response to the Hamilton County trial makes a fair point along these lines. Writing in August 2026, Khan argued that the 0.14 standard deviation figure for sustained participants “is a genuinely strong result”, and that the study was not asking whether Khan Academy beats doing nothing. It was measuring Khan Academy against whatever digital maths programmes and small-group instruction the district was already running. That is a demanding comparison, and the platform did not lose it.

And then there is Stanford's own contrary finding, which is the most interesting card in the optimistic hand. The Tutor CoPilot trial, run by Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb and Dora Demszky, put a language model behind 900 human tutors working with 1,800 students from historically under-served communities, offering expert-like suggestions in real time. Students whose tutors had the tool were four percentage points more likely to master topics. For students of the lowest-rated tutors, the gain was nine percentage points. The cost was around 20 dollars per tutor per year.

So the honest steelman is this: the technology demonstrably can teach, it is astonishingly cheap, it has worked in at least one rigorous field trial in a low-income setting, and it makes human tutors measurably better when pointed at them rather than at children. Anyone who wants to argue that AI tutoring is snake oil has to explain all four of those facts.

Where the Case Comes Apart

It comes apart at the point where the model stops being the subject of the sentence and the child becomes it.

Notice the shape of the successful examples. In Nigeria, students worked in pairs, after school, under teacher supervision, with structured prompts. That is not an AI tutor. That is a small-group human intervention with a language model in the middle of it, and the trial cannot separate the contribution of the model from the contribution of the adult who showed up and the peer who sat alongside. Tutor CoPilot is even clearer: it does not tutor anyone. It whispers to a human tutor who is already in a relationship with the student. Every case where AI tutoring has produced strong results is a case where a person was in the room.

The Stanford literacy trials tested the other configuration, the one the marketing implies and the procurement documents assume, in which the child and the software are left alone together. That configuration produced 2.18 minutes a week.

This is the distinction the sector has spent three years refusing to make. There is an enormous difference between a system that can answer a question and a system a child will actually ask. Model capability has been improving on a steep curve. Willingness to seek help from a machine has not, because it was never a function of model capability in the first place.

Help-seeking is one of the most studied behaviours in educational psychology, and it is socially loaded in ways that a benchmark score cannot capture. Asking for help is an admission. Children weigh that admission against what it costs them: looking slow in front of peers, disappointing a teacher, confirming a private suspicion about themselves. The reason a good tutor is valuable is not that they possess the answer. It is that they have built enough trust that the admission feels safe, and enough familiarity to notice the confusion before the child has to declare it.

Khanmigo could not do the second thing at all, and the 17 per cent figure suggests it was not trusted enough to be given the first. A child who has just got a question wrong in a remedial maths class is at the precise emotional coordinates where a human tutor would lean in. The AI sat there, one click away, and was not clicked.

Why Children Work for People Who Will Notice

The Stanford result that most repays attention is not the two minutes. It is the fact that adding a human being who was explicitly forbidden from teaching still moved engagement by 71 to 80 per cent.

The tutors in those trials were not subject experts delivering instruction. In District B they were middle school students, selected because they did well in their own English classes and had a free period. Their job was to check in, keep students on task, sort out headphones and passwords, and talk to the children about what they were doing. They spent part of every session on relationship-building activities that reduced the time available for the platform. And they still produced the largest engagement effect anyone in these trials produced.

What those tutors supplied was relational accountability: the simple, unglamorous fact that somebody would notice. Somebody would know whether you logged in. Somebody would ask how the story went. Effort became legible to another person, and legible effort is the currency children have always worked for.

Software can simulate this. It cannot instantiate it. A notification saying “we missed you yesterday” is not a person missing you. Children, including very young ones, appear to be excellent at telling the difference. A first grader in New Mexico, quoted by her mother in reporting by NBC News, complained of the software that “she doesn't let me finish my sentences, she doesn't listen.” An Albuquerque parent reported her children saying that Amira was “like a new teacher but they don't actually understand me”, and that one of them no longer liked reading.

A study accepted to the Learning Analytics and Knowledge conference in 2026 sharpens the mechanism considerably. Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil and Kenneth R. Koedinger analysed nearly 2,100 hours of classroom practice by 191 middle schoolers on an intelligent tutoring system, tracking what happened when a human tutor physically visited a student during their work. Engagement rose during the visit and stayed elevated afterwards. The returns diminished with visit length, and timing mattered more than duration. Interactions built on concrete, stepwise scaffolding with explicit organisation of the student's work were the most effective. Their recommendation for resource-constrained settings is deflating in its modesty: several brief, well-timed check-ins, including at least one early.

Brief. Well-timed. Human. The paper is, among other things, a costing exercise for the thing AI tutoring was supposed to make unnecessary, and the answer is that it is cheaper than anyone assumed. Not free, though. Never free.

The Word She Could Not Decode

Return to “usually”, because the failure it exposes is technical as well as relational, and the technical version is the one that will not be fixed by better prompting.

Amira works by listening. Students read passages aloud into a microphone; automatic speech recognition compares what it hears to what it expects; when a word goes wrong, the system intervenes, sometimes with a video of a mouth enunciating the correct pronunciation. It is a genuinely clever pipeline and it does something no human tutor can do at scale, which is give every child in a class simultaneous individual reading practice with immediate feedback.

But consider what the signal actually contains. A child hesitates on “usually”. From the acoustic evidence alone, that hesitation is consistent with at least four distinct conditions: she does not know the word; she knows the word but cannot decode the orthography; she can decode it but is reading too slowly to hold the sentence in working memory; or the microphone picked up the child at the next desk. These require four different responses. Defining the word helps only in the first case. In Graham's daughter's case it was the second, and the software chose the first.

The wider evidence suggests this is not an isolated misfire. Reporting by the Albuquerque Journal in April 2026 documented teachers describing month-to-month swings of 20 to 30 percentile points in individual students' Amira scores, which a teaching coach characterised as statistically abnormal. A kindergarten teacher noted that the system “is not always great at picking up the language of students” with spoken language difficulties. Classroom background noise contaminates results. A special education teacher described students groaning and crying on assessment days.

New Mexico requires Amira statewide for kindergarten through second grade, at a cost of around 2.7 million dollars a year, with assessments three times annually and monthly for children reading below expectations. Idaho requires it too; California, Georgia, Massachusetts, Michigan, Oklahoma and Texas have authorised it. In a Reason report published on 2 September 2026, only 8 per cent of surveyed New Mexico teachers and administrators said they had no major concerns about it, though the poll was run by the state education department at one of its own training sessions, and the department has said it was not representative. Amira's chief executive, Mark Angel, has defended the evidence base robustly, saying that “no other scalable instructional intervention has been interrogated as many times, by as many independent teams, with this consistency of positive impact.”

Both things can be true. The efficacy studies can show real reading gains under conditions of proper use, and the deployment can still be systematically misdiagnosing children whose accents, dialects, speech differences or classroom acoustics fall outside the model's comfortable centre. The children most likely to be misread by a speech recognition system are, with grim predictability, the same children the system was funded to help.

The requirement did not survive the summer intact. Parents showed up en masse at school board meetings across the state, worried less about pedagogy than about where the recordings of their children's voices were going. Six districts and charters declined to use the software at all on privacy grounds: Santa Fe, Los Alamos, Farmington, Roswell, Clayton and Turquoise Trail Charter School, with the exemptions running only for that school year. On 4 August 2026, days before the new term began, the Public Education Department issued revised guidance signed by Secretary Mariana Padilla, permitting districts to run Amira without the voice-recording feature, to administer a paper-based test instead, or to use an assessment programme of their own. Voice recordings delete monthly by default, and districts may now request daily, weekly or end-of-year deletion. The department held its ground on the principle, arguing that “a common statewide assessment provides a shared measure that supports consistency, transparency, accountability and equitable decision making”. Amira remains the statewide requirement. It is simply no longer in every public school.

Albuquerque Public Schools, the largest district in the state, resolved on 26 August 2026 to keep the programme with concessions: voice recordings dropped from the tutoring component, a 48-hour deletion window on the testing feature, and paper-and-pencil alternatives for parents who opt out. Deputy Superintendent Randy Mahlerwein explained the 48 hours as a compromise, the testing data being deleted on that cycle “so teachers have a chance to listen to the recordings”. Representative Linda Serrato, who led more than thirty legislators in a letter demanding oversight, put the stake plainly: “You're talking about the biometric data of children 5 to 8 years old. That's valuable stuff, and we know it, but we have to treat it as such.” Angel, for his part, has said the company would rather not hold the material at all. “We don't want to collect this data; it's a nuisance,” he said. “If the Legislature or PED tells us to stop collecting the data, we will stop instantaneously.”

It is worth being precise about what moved and what did not. No new efficacy finding prompted any of this. The evidence base sat in August exactly where it had sat in April. What changed was that parents turned up, districts refused and legislators wrote letters, which is to say that the correction to an automated system arrived by way of people paying close attention to particular children, which is the one resource the technology had been sold as a substitute for.

Who Bears the Risk

None of this is priced into the way districts buy.

Educational technology is sold on licences, not on usage. A district commits to a per-student annual fee, the vendor books the revenue, and whether the child logs in is somebody else's problem. This is not a new pathology. Analyses by LearnPlatform, before the generative AI wave, found that roughly a quarter to a third of purchased edtech licences were never activated at all, and that intensive use, defined as ten or more hours per product between assessments, applied to about two per cent of licences. Estimates put over a billion dollars of American K-12 licensing spend into the category of pure waste each year.

What the Stanford and Hamilton County trials show is that the generative AI generation of products has inherited this structure and, so far, has not improved on it. A district that timetables two 30-minute sessions a week, and receives 2.18 minutes, is paying roughly 27 times the advertised unit cost of the intervention it thinks it bought. Nobody's contract says that.

The public spending is not trivial and is increasingly visible. New Mexico spends 2.7 million dollars a year on Amira. Iowa committed 3 million dollars in 2024 with a further 2.5 million after. Louisiana authorised 3.6 million plus another million. Duval County Public Schools in Florida structured its purchase differently. Its 2024 contract, worth 100,000 dollars and covering roughly 2,600 students in grades two through four who were reading below grade level, tied half the fee to student progress. More than 1,200 of them met their oral reading fluency goals, exceeding what the contract had been written to expect, which is either a decent result or an expensive coin flip depending on your counterfactual, except that the district was only paying in full for the half that landed. North Carolina, notably, cut its Khanmigo funding from 10 million dollars to 500,000.

The American education secretary, Linda McMahon, has acknowledged that there are not “a lot of metrics” for judging AI's classroom value, and asked the obvious question: “Are we seeing better outcomes in schools? And if we're not, then they should be pulled out.”

The public appears to be well ahead of the procurement offices. A Century Foundation survey, conducted by Morning Consult among more than 2,000 registered voters in May 2026, found 84 per cent concerned about private companies collecting and profiting from student data, and 81 per cent concerned that teachers are being pressured to use AI tools without proven pedagogical benefit. Seventy-seven per cent wanted government guardrails, a figure that held across party lines. Forty-nine per cent said classroom technology and AI should be kept to a minimum. This is not a technophobic fringe. It is a settled majority, expressing a preference that the market has so far been structured to ignore.

The Ghost of Two Sigma

Underneath every AI tutoring pitch of the last three years sits a single number that almost nobody quoting it has checked.

In 1984, Benjamin Bloom published a paper in Educational Researcher reporting that students taught one-to-one with mastery learning outperformed conventionally taught students by two standard deviations, placing the average tutored student above 98 per cent of the control class. He framed it as a challenge: find a group method that achieves what tutoring achieves. The “2 sigma problem” became the founding scripture of educational technology, and generative AI inherited it wholesale. Every promise of a personal tutor for every child is a promise to close Bloom's gap with software.

The number has never been replicated. Bloom's finding rested on two doctoral dissertations by his own students, conducted with small samples, short durations and researcher-designed outcome measures. A 1982 meta-analysis by Peter Cohen, James Kulik and Chen-Lin Kulik had already put the average tutoring effect at around 0.33 standard deviations. The Nickow, Oreopoulos and Quan review found nothing approaching two sigma anywhere in the experimental literature; its pooled estimate sat between 0.29 and 0.37 depending on the specification. Matthew Kraft of Brown University has argued that Bloom's number helped anchor the field to expectations of effect sizes that essentially never occur, noting that most education interventions produce effects of 0.1 standard deviations or less.

Seen against that corrected baseline, the Hamilton County result of 0.06 to 0.08 standard deviations a year, rising to 0.14 for sustained participation, is not humiliating. It is an ordinary education intervention performing ordinarily. The humiliation is entirely a function of what was promised.

England's National Tutoring Programme offers a parallel worth sitting with. Launched with substantial funding to address pandemic learning loss, its independent evaluation found no evidence that the Tuition Partners route improved Key Stage 2 or Key Stage 4 outcomes in English or maths, while school-led tutoring produced small gains equivalent to about a month's progress. The intervention with the best evidence base in education, delivered at national scale under time pressure, largely failed to reproduce its own effect. Yet 81 per cent of school leads surveyed felt the programme had helped pupils catch up. The gap between what practitioners perceive and what the data records is not unique to AI. It is what scale does to interventions, and it should temper any assumption that AI tutoring's problems are peculiar to AI.

Beyond the Test Score

There is a further argument, and it cuts against the entire framing of the debate.

A paper submitted in February 2026 by Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser and Nuria Oliver argues that assessing educational AI purely on learning outcomes misses most of what matters. Their framework identifies four interlocking dimensions: cognitive offloading, diminished learner agency, emotional disengagement and surveillance-oriented practice. Their central claim is that these reinforce one another, and that the compound effect operates on critical thinking and civic participation rather than on test scores. They are careful not to be deterministic about it. Well-designed systems, they argue, can support reasoning and autonomy while preserving meaningful human interaction. The question is not whether AI belongs in education but how institutions deploy it.

Apply that lens to the current evidence and something uncomfortable emerges. The thin engagement documented in Tennessee and California and New Mexico is being read as a failure. On the Favero framework, it might be partial protection. A child who does not offload her thinking to a chatbot is not being harmed by the chatbot. The Hamilton County students who sent bare answers and clicked suggested prompts were exhibiting exactly the answer-seeking behaviour that worries cognitive scientists, and they were doing it at low volume.

Which raises a genuinely difficult question. If engagement rose to the 30 minutes a week the vendors recommend, would attainment rise with it, or would we simply have more children more efficiently outsourcing the cognitive work that constitutes learning? Nobody knows. The trials that would tell us have not been run, because the dosage required to run them has never been achieved. Robinson's team put this with admirable candour: they never reached sufficient use to determine whether the tool works at all.

What an Honest Evidence Standard Would Require

The gap between promise and delivery is not going to be closed by better models, because the constraint was never the model. It might be narrowed by changing what districts are allowed to buy and what vendors are required to show.

Four changes would do most of the work. First, effectiveness claims should be stated as a function of dosage and reported alongside observed dosage in real deployments. A product whose efficacy study assumed 30 minutes a week should have to publish what its actual median user does, by district, annually. Second, contracts should tie payment to usage rather than to seats, which would move the risk of non-engagement from the public purse to the party that can actually design against it. Duval County has already written half a contract that way, so this is a reform with a working example rather than a thought experiment. Third, trials should be pre-registered and report the null. The Stanford literacy paper is a model here precisely because its headline finding is a failure to establish anything about the technology. Fourth, deployment equity should be a reported metric, not an afterthought. If the students least likely to open the application are the ones with special educational needs and the lowest prior attainment, a district needs to know that in month two, not from a working paper appendix two years later.

None of this is exotic. It is roughly the standard that a medicines regulator would consider a baseline, and roughly the standard the Education Endowment Foundation has spent fifteen years trying to establish in England, where the evidence for one-to-one tuition is rated as moderately secure on the basis of 123 studies. The reason it feels exotic in educational technology is that educational technology has never had to meet it.

The Girl and the Avatar

The child stumbling over “usually” is the whole argument compressed into a single second of audio.

What she needed was for somebody to notice the specific shape of her confusion: not that she lacked the word but that she could not get through the spelling to reach it. Noticing is not a capability that scales with parameters. It requires attention that is directed at a particular person, sustained over time, and, crucially, that the person knows is being directed at them. That last part is what produces the effort. Children do not work hard because a system is watching. They work hard because someone who matters to them will see.

The two-minute lesson is not a story about bad software. Amira and Khanmigo are, by any reasonable technical standard, impressive artefacts, and the studies suggest that when children use them properly they produce ordinary, real, modest gains of the kind that education research has always found. The story is about a category error that ran through an entire procurement cycle: the assumption that the scarce resource in education was instruction, when the scarce resource was always attention, and attention is the one thing that has never been possible to manufacture at zero marginal cost.

Stanford's tutors, forbidden from teaching, moved the numbers more than the AI did. That is the finding. It has been available, in one form or another, since Bloom, and it survived the arrival of a technology that was supposed to make it obsolete. A district that spends its money on the thing that produced 71 to 80 per cent more engagement rather than on the thing that produced 2.18 minutes a week is not being nostalgic. It is reading the evidence.

Whether anyone is buying on the evidence is a separate question, and on current form the answer looks like a non-event.

Sources and References

  1. Carly D. Robinson, David Gormley, Ana Trindade Ribeiro and Susanna Loeb, “Access is Not Enough: Human Support Improves Engagement with AI Tutoring”, EdWorkingPaper No. 26-1451, Annenberg Institute at Brown University, June 2026. https://edworkingpapers.com/ai26-1451
  2. Chalkbeat, “Does AI tutoring work? Students would have to use it for researchers to find out”, Chalkbeat, 17 June 2026. https://www.chalkbeat.org/2026/06/17/ai-tutoring-research-ran-into-problem-students-wouldnt-use-it/
  3. Philip Oreopoulos and Nina Low, “One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment”, NBER Working Paper 35620, August 2026. https://www.nber.org/papers/w35620
  4. Matt Barnum, “Students rarely engaged with Khan Academy's AI-powered tutor Khanmigo, study finds”, Chalkbeat, 25 August 2026. https://www.chalkbeat.org/2026/08/25/ai-tutoring-students-khanmigo-khan-academy-engagement-study/
  5. Matt Barnum, “Why Sal Khan is rethinking how AI will change schools”, Chalkbeat, 9 April 2026. https://www.chalkbeat.org/2026/04/09/sal-khan-reflects-on-ai-in-schools-and-khanmigo/
  6. Sal Khan, “What I found compelling in a new randomized trial of Khan Academy in math intervention”, Khan Academy Blog, 24 August 2026. https://blog.khanacademy.org/what-i-found-compelling-in-a-new-randomized-trial-of-khan-academy-in-math-intervention/
  7. Linda Jacobson, “AI Tutors Not Yet a Replacement for Humans, Research Says”, The 74, 2 September 2026. https://www.the74million.org/article/ai-tutors-not-yet-a-replacement-for-humans-research-says/
  8. The Century Foundation, “Americans Are United Against the Tech Takeover of Public Schools”, 2026. https://tcf.org/content/report/americans-are-united-against-the-tech-takeover-of-public-schools/
  9. Ed Finkel, “Voters express broad concerns about AI and tech in schools, survey shows”, K-12 Dive, 2 September 2026. https://www.k12dive.com/news/voters-express-broad-concerns-about-ai-and-tech-in-schools-survey-shows/829351/
  10. Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil and Kenneth R. Koedinger, “Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online Learning”, arXiv:2601.09994, 15 January 2026. https://arxiv.org/abs/2601.09994
  11. Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser and Nuria Oliver, “AI in Education Beyond Learning Outcomes: Cognition, Agency, Emotion, and Ethics”, arXiv:2602.04598, 4 February 2026. https://arxiv.org/abs/2602.04598
  12. Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb and Dora Demszky, “Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise”, arXiv:2410.03017, October 2024. https://arxiv.org/abs/2410.03017
  13. Andre Nickow, Philip Oreopoulos and Vincent Quan, “The Impressive Effects of Tutoring on PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental Evidence”, NBER Working Paper 27476, July 2020; published in revised form in the American Educational Research Journal, 2024. https://www.nber.org/papers/w27476
  14. Paul T. von Hippel, “Two-Sigma Tutoring: Separating Science Fiction from Science Fact”, Education Next, vol. 24, no. 2, Spring 2024. https://www.educationnext.org/two-sigma-tutoring-separating-science-fiction-from-science-fact/
  15. Education Endowment Foundation, “One to one tuition”, Teaching and Learning Toolkit. https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit/one-to-one-tuition
  16. Education Endowment Foundation, “Independent evaluation of the National Tutoring Programme”. https://educationendowmentfoundation.org.uk/news/new-independent-evaluation-of-the-national-tutoring-programme-ntp
  17. Wenting Ma, Olusola O. Adesope, John C. Nesbit and Qing Liu, “Intelligent Tutoring Systems and Learning Outcomes: A Meta-Analysis”, Journal of Educational Psychology, November 2014. https://www.apa.org/pubs/journals/features/edu-a0037123.pdf
  18. Social Programs That Work, “ASSISTments”, Arnold Ventures. https://evidencebasedprograms.org/programs/assistments/
  19. EdWeek Market Brief, “More Than $1 Billion in K-12 Ed-Tech Licensing Fees Go to Waste”, November 2019. https://marketbrief.edweek.org/education-market/more-than-1-billion-in-k-12-ed-tech-licensing-fees-go-to-waste/2019/11
  20. Tyler Kingkade, “Meet Amira, an AI reading tutor alarming some parents and school leaders in New Mexico”, NBC News, 7 August 2026. https://www.nbcnews.com/news/education/ai-reading-tool-amira-new-mexico-parents-schools-privacy-concerns-rcna591161
  21. Elizabeth Nolan Brown, “Invasion of the Robot Teachers”, Reason, 2 September 2026. https://reason.com/2026/09/02/invasion-of-the-robot-teachers/
  22. Natalie Robbins, “New Mexico schools use Amira AI reading assessments amid accuracy and equity concerns”, Albuquerque Journal, 26 April 2026. https://www.abqjournal.com/news/students-read-aloud-ai-scores-them/3029121
  23. Natalie Robbins, “New Mexico eases rules for AI reading test amid privacy concerns”, Albuquerque Journal, 10 August 2026. https://www.abqjournal.com/news/amid-privacy-concerns-new-mexico-eases-rules-for-ai-reading-test/3099834
  24. “Amid parent backlash, APS keeps AI reading program with changes”, Albuquerque Journal, 26 August 2026. https://www.abqjournal.com/news/amid-parent-backlash-aps-keeps-ai-reading-program-with-changes/3110194
  25. CNN, “Education Secretary McMahon says AI can be a 'very effective tool' in schools”, State of the Union, 23 August 2026. https://www.cnn.com/2026/08/23/politics/video/mcmahon-ai-in-schools-sotu

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

Previous

The Unproven Cure

Empowering Youth: The STUDENTS FIRST Act Explained
Bulletin №20