The Promise and the Reality Check
When Khanmigo first launched as Khan Academy’s GPT-4-powered tutoring assistant, there was palpable excitement in the education world. Here was a tool that could, theoretically, give every student a personalized tutor available at 2 AM on a Tuesday when they’re stuck on their homework. By the end of 2025, the platform had grown to reach 1.5 million student users across 130 countries, up from just 200,000 during its limited pilot phase. That growth tells us something meaningful: educators and families genuinely wanted to try this.
But here’s what I’ve learned after years of teaching: enthusiasm and effectiveness are not the same thing. Real classrooms are messy, complex ecosystems where good intentions often collide with the actual, everyday work of learning. So when rigorous research started coming in about how Khanmigo performs in the wild, I paid attention. The data isn’t a clean victory story, and honestly, that makes it far more useful to those of us actually working with students.
The Math Win: Small, Significant, Real
Let me start with what worked. A randomized controlled trial conducted by WestEd in partnership with Khan Academy across 47 schools in 2025 examined what happened when students used Khanmigo specifically for mathematics instruction. After one semester of regular use, students showed a statistically significant improvement of 0.15 standard deviations in algebra readiness scores compared to peers who didn’t use the tool. For those who don’t live in research-speak: that’s a real, measurable gain. It’s not transformative. It won’t single-handedly close achievement gaps or revolutionize math education. But it’s meaningful in the way that consistent, small improvements add up over time.
Why did it work for math specifically? I think it comes down to how mathematics tutoring actually functions at its best. When I sit with a student who’s struggling with solving for x, what I’m really doing is asking strategic questions that guide them toward their own understanding. “What do you get if you subtract 5 from both sides?” “Does that look familiar from the problem we worked earlier?” That Socratic method, where the tutor offers guidance rather than immediate answers, translates surprisingly well to an AI interface. Khanmigo’s design specifically emphasizes hints over solutions. Students can ask for help on a problem and instead of getting the answer handed to them, they get a nudge in the right direction. That structural choice appears to matter.
The Writing Reality: Where AI Tutoring Hits Its Limits
Here’s where the data gets more complicated, and where I think we need to sit with some genuine uncertainty. The same WestEd study examined Khanmigo’s essay coaching feature and found no statistically significant improvement in writing skills. This wasn’t a failure of the tool’s technical quality. It was something more subtle and, in some ways, more revealing about how writing instruction actually works.
Researchers hypothesized that the tool encouraged revision over generative thinking. In plain language: writing is fundamentally different from solving an algebra problem. When a student writes an essay, they’re not working toward a single correct answer. They’re learning to think through complexity, to hold multiple ideas in tension, to find their own voice and perspective. An AI tutor can absolutely help a student revise a draft, correcting grammar, suggesting reorganization, flagging unclear passages. But the generative thinking part, where a student stares at a blank screen and must figure out what they actually think about the prompt, requires something different. It requires the kind of struggle that can’t be outsourced to an algorithm, no matter how helpful that algorithm intends to be.
This taught me something I’ll remember: not every learning challenge responds to the same solution. When we see AI tutoring work beautifully in mathematics, it’s tempting to assume it will work equally well everywhere. The data says otherwise.
The Teacher’s Perspective: Time Matters Too
Here’s something that often gets overlooked in conversations about educational technology: teachers are drowning in administrative work. Khan Academy reported that teachers using the Khan Academy Khanmigo for Teachers dashboard spent an average of 37 minutes less per week on progress monitoring and data entry tasks, based on surveys of 3,200 participating teachers. That’s real time. Nearly two and a half hours per month that a teacher isn’t hunched over spreadsheets tracking which students have mastered which standards.
I can tell you from experience that this matters. When I have accurate, automatically generated data about where my students stand without spending an extra hour per week compiling it myself, I have more time and mental energy for what actually requires human judgment: noticing the student who’s getting the right answers through a dangerously fragile method, recognizing the moment a student is about to give up on themselves, or simply remembering to follow up on a conversation we had three weeks ago. That 37 minutes is time returned to actual teaching. It’s not flashy, but it’s real.
The Student Voice: Why Anxiety Matters
One detail from the research particularly stuck with me. A 2025 report by the Christensen Institute EdTech Research found that 68 percent of students surveyed preferred Khanmigo’s hint-based approach over ChatGPT for homework help. When students were asked why, the reason wasn’t just that Khanmigo worked better. It was that they felt less anxious about cheating. They could use the tool without that nagging worry that they were circumventing the actual learning process.
This genuinely moved me. So much of the resistance to AI tutoring among students comes from internal conflict: they want help, but they also don’t want to cheat on themselves. They want to reduce their struggle but also want to actually learn something. That’s not a character flaw. It’s a completely human tension. When a tool is designed with that psychological reality in mind, it changes what students are willing to actually use. And what students actually use is what might actually help them.
What This Means For How We Move Forward
After one year of real classroom data, here’s what I believe: AI tutoring like Khanmigo works best as a complement to human teaching, not a replacement. It can meaningfully boost performance in well-structured domains like procedural mathematics. It can free up teachers’ time for the work that requires human judgment and presence. It can ease student anxiety around the learning process itself. And it has reached millions of students globally, which means the experimentation and iteration will continue at scale.
What it cannot do, at least not yet, is replicate everything that happens in a real teaching relationship. It cannot know your learning style from September and hold that knowledge in June. It cannot light up when you finally understand something you’ve been struggling with for weeks. It cannot notice what you almost said but stopped yourself from saying, and ask you the follow-up question that might matter. Those capacities remain distinctly human.
The most honest assessment I can offer: the data suggests we have a useful tool. Not a panacea, not a replacement, but a genuine tool that solves specific problems in specific contexts. That’s perhaps less exciting than the original hype suggested, but it’s far more actionable. If you’re an educator considering Khanmigo, the question isn’t whether it will transform your classroom. The question is whether it will solve the particular problems you’re facing with your particular students. For math readiness and teacher administrative load, the data says yes. That’s worth something.
What has your experience been with AI tutoring tools in your own learning or teaching? I’d genuinely like to hear what questions this raises for you.