Why OpenAI’s o3 Is Changing How We Teach Critical Thinking (And Not In the Way You’d Expect)

The Moment Everything Got Easier (And That’s the Problem)

When OpenAI released the o3 model this past April, something shifted in my classroom that I didn’t expect. The model scored 87.5% on the ARC-AGI benchmark, a test that researchers had designed specifically to measure human-level reasoning. That’s not just a number on a dashboard somewhere—it means a machine can now work through logic puzzles and reasoning tasks at a level we thought would take decades to achieve. My students noticed immediately. Within days, I started seeing hands go up less frequently during problem-solving activities. Not because they didn’t care, but because they had a tool that could show them the path to the answer before they walked it themselves.

Here’s what kept me up that night: I was looking at a paradox I hadn’t fully prepared for. We’ve spent years teaching students that having access to information isn’t the same as understanding it. Google didn’t break critical thinking, it just meant we had to teach differently. But o3 and models like it aren’t just giving answers. They’re showing their work. They’re breaking down problems in ways that make reasoning visible and compelling. And that visibility is doing something unexpected to how my students think.

What the Research Is Actually Telling Us (It’s Not What We Hoped)

I always tell my students that the best learning comes from bumping up against a problem, struggling a bit, and then having that moment where it clicks. That struggle is where the learning lives. So when I read the Stanford Graduate School of Education study from late 2025, I felt something between vindicated and alarmed. Researchers found that 73% of high school students who regularly used AI reasoning tools showed measurable decline in independent problem decomposition skills after six months. That’s not a small number. That’s more than two-thirds of the students we surveyed.

The study didn’t say these students couldn’t think. It said something more specific and more troubling: they were less likely to break complex problems into smaller pieces on their own. When faced with a difficult question, they defaulted to asking the AI to do it rather than doing that cognitive work themselves. It’s like having someone carry your groceries every time you go to the store. Eventually, you stop building the muscle. The muscle didn’t disappear. You just stopped using it.

What struck me most wasn’t the decline itself, but what it revealed about how learning actually happens. We can’t just hope students will maintain critical thinking skills through osmosis. We have to actively practice the thing we want them to get better at. And if we’re not careful, the very tool designed to help us teach reasoning can become the thing that lets students skip the reasoning part altogether.

The Uncomfortable Truth About Teacher Preparedness

When EdWeek Research Center surveyed teachers about teaching alongside advanced AI models, 61% reported feeling underprepared. I read that and immediately recognized myself in it. I went to graduate school to learn how to teach calculus and chemistry, not how to teach alongside reasoning engines that can work through calculus and chemistry problems faster than I can.

But here’s what I want to be honest about: the problem isn’t that we teachers are behind. We’re not. The problem is that we’re being asked to do something genuinely new without a clear roadmap. How do you create the conditions for struggle when a student can get a beautifully explained answer in thirty seconds? How do you know when a student understands something versus when they’ve just learned to prompt an AI effectively? These aren’t theoretical questions in my classroom. They’re every-single-day questions.

The real work is figuring out where AI becomes a tool for learning versus a shortcut around learning. That distinction is subtle, but it matters enormously. When I use a calculator in a physics class, students have already learned to multiply and divide. When they use o3 to verify their reasoning, they should have already practiced reasoning independently. The sequence matters.

What Schools Are Actually Doing About This

The ISTE AI in Education Guidelines 2026 made a recommendation that surprised me with its specificity: schools should dedicate at least 15% of STEM instructional time to reasoning practice without AI. That’s not avoiding AI. It’s being intentional about when students work with and without these tools.

I started thinking about this differently after reading those guidelines. It’s similar to how we teach writing. Students don’t learn to write well by editing with Grammarly every time they draft. They learn by writing badly, getting feedback, revising without technological assistance, and doing that cycle until better writing becomes something they can do independently. Then, when they use AI tools, they’re using them to enhance skills they already have, not to replace the practice of building them.

Some of my colleagues have started being really explicit about this with students. Instead of saying “don’t use o3 for your homework,” they’re saying “here’s when we practice without AI so you build these skills, and here’s when we use it to explore what’s possible once you have those skills.” One teacher in our department created what she calls “AI-free reasoning days,” Tuesdays and Thursdays where the entire class works through problems together without tools. Her students report feeling more confident in their own work now, not less.

The Real Reason This Matters (And What’s Coming Next)

The College Board is already thinking about this. They announced in February 2026 that discussions about redesigning the SAT are underway specifically to address AI-assisted reasoning. By 2027, they’re planning pilot changes. What that likely means is that tests will look different, maybe less about final answers and more about showing how you think, or maybe more untimed reasoning tasks, or assessment methods we haven’t even designed yet. Our assessments are going to have to evolve because our tools have evolved.

But here’s what I keep coming back to: this isn’t actually a problem with o3 or any AI model. This is an opportunity to get clearer about something we should have been clear about all along. What exactly do we want students to be able to do? If the answer is “break down complex problems, reason through ambiguous situations, and think critically when faced with novel challenges,” then we need to make sure students practice those things independently. That’s not anti-technology. That’s pro-learning.

I’ve started asking my students a different question at the beginning of the year. Instead of “What do you want to learn?” I ask “What do you want to be able to do that you can’t do now?” When they say things like “I want to understand why my chemistry answers are right or wrong, not just know they’re right,” or “I want to feel confident figuring things out on my own,” that tells me exactly where my job lies. It’s to create the conditions where that confidence grows. AI is a tool in that work, but it’s not the work itself.

What’s your experience been so far with these tools in your classroom or learning environment? I’d genuinely love to hear where you’re seeing the most friction and where you’re seeing unexpected wins. None of us is going to figure this out alone, and honestly, the most useful thing I’ve learned this year has come from other teachers comparing notes.