Is AI use degrading student achievement? A recent study on AI use among secondary students says just that. David Strömberg, Victor Lei, and Yanhui Wu conducted the study, which examined about 26,000 Chinese secondary students over three years. The title says it all: The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. They found that those students who used AI had homework scores raised by 18%, homework completion time dropped by 30%, and the kicker is that exam scores fell by 20% within six months. Yikes!
One note: this is a working paper, not yet peer-reviewed. It’s already circulating widely among economists and education researchers, and the design is serious. But 26,000 students and thirty months of real school data are worth a deep read.
What is Causing the Learning Loss?
According to the researchers, the lower test scores were “concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing.” In layman’s terms, the students had AI do the “thinking” for them, and when it came to a high-stakes test, they didn’t know it. This was evident when you see that their homework took significantly less time. I can just see students copying and pasting and moving on.
Homework scores up 18%. Exam scores down 20%. Same students, six months apart.
But Some Students Used AI Well...ish
The study wasn’t all bad news. 19% of the students who used AI spent the same amount of time on homework as before, and they did not see the exam penalty. The researchers said: “AI users who maintain similar homework completion time as non-AI users experience small learning losses.” I am not totally surprised.
I want to be careful here, though, because it’s tempting to read this as “these kids used AI well.” That’s not quite what the data shows. The paper doesn’t tell us these students found some clever technique or used AI as a tutor instead of an answer key. All we know is that they kept spending roughly the same amount of time on their homework, AI or no AI. That’s a behavior, not a skill. It’s closer to “they didn’t let the tool shortcut their effort” than “they figured out the secret to AI-assisted learning.”
I wish the study had asked how students used AI so we could suss out what helped and what hurt. I wonder if, as we learn more about the best way for students to use AI, we’ll find that students who use it as an actual tutor — one that pushes back, asks questions, makes them work for the answer — see real gains instead of just smaller losses. That’s a different, better outcome than just “didn’t lose as much,” and this study can’t tell us yet whether it’s possible at scale.
The Slow Leak — Why the Damage Takes Two Years to Show Up
This is the part of the study that concerned me most.
Regular exam scores hit their full damage within six months of a student starting to use AI. But the high-stakes entrance exams — the ones that actually determine which high school or college a Chinese student gets into — took two full years to reach their full penalty of 18 to 24%.
Why the difference? Those entrance exams cover material from before a student started using AI and after. Early on, most of what’s being tested is stuff the student learned the old-fashioned way, so the damage is diluted into the average. It takes years for enough AI-affected material to pile up and drag the whole score down. Which means almost every short study on AI and learning — and there have been a lot of them — is structurally incapable of catching the real cost. They’re not running long enough to see it.
There’s a hopeful thread buried in here too. The five-month penalty fell from “around 25 percent in early 2023 to around 16 percent by June 2025.” People are adapting. Slowly, imperfectly, but adapting.
How should we use AI going forward?
I am obsessed with learning how we can positively use AI in our schools. So what is the best way to use AI? The researchers point out that their findings don’t contradict a Harvard study in which a well-designed AI tutor helped students learn physics faster than in-class active learning did. Their own explanation: “With a well-designed tool and highly motivated students, AI can enhance learning efficiency.” The difference isn’t the technology. It’s the design and the guardrails around it.
That’s the entire premise behind what I call an AI Engine: AI that’s built to make a student produce the thinking, not receive the answer. This study is the largest, most real-world piece of evidence I’ve seen that the design question isn’t a nice-to-have. It’s the whole ballgame.
"Better homework performance may coexist with weaker learning."
One more line, this one for every teacher and parent who thinks homework grades still tell them something: “Better homework performance may coexist with weaker learning.” Read that twice. A student’s homework score can go up while what they actually know goes down, and you won’t see it until the closed-book test. I know I’ve said this before, but if you, as a teacher, send home cognitively complex homework, you are crazy. I now assume students are going to use AI. So that means we need to redesign how we do homework. That is why my homework now is primarily for students to watch my flipped videos. Homework has quietly stopped being a reliable signal, and we need something AI can’t do for a kid.
Implications for Assessment
I really want my students to not have the “AI Learning Penalty.” I want them to use AI in a way that enhances their learning. That is why I’ve moved toward oral Mastery Vivas, where students have to explain their understanding to me both formatively and summatively.
As the researchers put it: “students are gradually adapting to generative AI, but there remain substantial barriers.” Bad for most. Good for some. And maybe, if we’re honest about the design instead of the hype, better for more of them than it is right now.
Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education [Working paper]. June 2026.


