AI + Education Research

Jun 11, 2026

Since the ChatGPT moment in late 2022, AI has spread haphazardly into education, affecting students, teachers, and school leaders.

The research is still catching up. Schools are experimenting, students are using the tools on their own, and the models keep developing. Stanford’s SCALE lab’s 2026 review makes the same point: the research is still very limited.

Still, the early research points in a few clear directions.

This piece is my attempt to keep track of the big findings and update them over time. It is not meant to be an exhaustive overview of AI in education. My work is mostly focused on how AI affects students, especially in Western countries where access to the tools is less of a constraint.

Last updated 11 Jun, 2026.

What We Know So Far

  • General chatbots harm learning. Giving students open-ended access to chatbots like ChatGPT often pushes them toward outsourcing their thinking, relying on the tool, and skipping the mental work that learning depends on. The early evidence suggests this is not just neutral, it have a negative effect.
  • Short-term performance is not the same as learning. Students often perform better while an AI tool is available, then regress once it is taken away. That points to a real dependency risk: a tool that helps students complete the task in front of them does not automatically leave them with durable knowledge.
  • Pedagogical scaffolding shapes the outcome. The positive findings come from tools built for learning, not generic solutions. The same model can harm or help depending on the scaffolding around it. Pedagogical scaffolding is a major determinant of the outcome, with similar tools producing very different results depending on how they are structured.
  • Students already use AI. Schools do not fully control whether students encounter these tools. Many students already use them on their own, often without guidance from teachers. Estimates vary by age, but suggest that 50-80% of students already use tools like ChatGPT.
  • Student AI use often conflicts with learning. Left on their own, students tend to use AI to save time and reduce effort. Few students use AI in ways that are compatible with long-term learning.

What We're Still Figuring Out

  • Importance of model quality. Most AI education studies use old models. That makes the findings hard to apply to the tools in use right now. GPT-4, which appears in many studies, scores 13/100 on Artificial Analysis's Intelligence Index. State-of-the-art models in summer 2026, like Anthropic's Claude Fable-5, score 65/100. Education research has not caught up to that jump. We need to know how much that matters. Does speed change student engagement? Does it keep students in the flow, or does it make little difference? Does higher intelligence expand what models can teach, and are there diminishing returns after a certain point?
  • Differences across subjects Most studies focus on STEM, especially math. We know much less about writing, history, languages, arts, and other subjects. AI may help more in some subjects than others. It may also need to be used differently depending on what is being taught.
  • Differences between students. Student profile probably matters. A student with ADHD may need a different AI experience than a neurotypical student. The open question is whether we can adapt the tool to the student's needs and what kind of impact that could have.
  • Timing in the learning process. What a student already knows likely affects how well AI works for them. AI may be most useful after a student has built a basic understanding of the topic. Or it may help earlier, as a way into the subject. We do not know yet where it fits best in the learning process.
  • Long-term effects are still unclear. Most studies look at short interventions. A few sessions, often less than a full semester. We still do not know what happens when students use these tools for months or years. That is hard to study because the technology is new.

My Reflections

For context, I read these studies as a builder of AI tools.

They do not tell you very much about what to build, but they can give you useful principles. They show where the traps are. They can help you avoid obvious mistakes. But in an environment changing this quickly, I think you need to be grounded in a particular point of view. You need your own vision for how AI and education should fit together, you can't find that in these studies and reports.

Most of the findings are also not that surprising if you start from learning science. If a student is not actively engaging, and is mostly consuming or delegating, they will not learn much. We already knew that. The AI-specific research is useful, but at this stage I think the stronger guide is still established learning science, combined with a clear understanding of the capabilities of AI and how students actually use the tools.

The Evidence Base on AI in K-12: A 2026 Review - The best concise overview of the available studies and research in the field.

OECD Digital Education Outlook 2026 - A balanced and comprehensive, big-picture overview of the current state of AI in education. It doesn't get into the weeds, but is still very informative and introduces useful concepts like fast vs. slow AI use. Chapters 1–3 are the most relevant to my interests.

Generative AI without guardrails can harm learning: Evidence from high school mathematics - Shows how much difference a simple tweak can make: students using a base chatbot learned less, while those using a version with pedagogical scaffolding did not. How well the results transfer to today is uncertain, though, given how much models have improved since. It's also a cornerstone study frequently cited in major overview reports like the OECD Outlook.

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning - Explores proactive AI tutors and the substantial improvements they can bring.

A Scoping Review of Large Language Model-Based Pedagogical Agents - No findings of note, but its framework of four AI agent characteristics is helpful, even if potentially incomplete: interaction approach (reactive vs. proactive), domain scope (domain-specific vs. general-purpose), role complexity (single-role vs. multi-role), and system integration (standalone vs. integrated).