AI + Education Research

Jun 18, 2026

Since the ChatGPT moment in late 2022, AI has spread haphazardly into education, affecting students, teachers, and school leaders.

For now, the research remains limited and is still catching up. Schools are experimenting, students are using the tools on their own, and the models keep evolving.

Even so, the early research points in a few clear directions.

I’m using this document to keep track of the big findings and update it over time. It is not meant to be an exhaustive overview of AI in education. My interests, and therefore the research I follow, focus mostly on how AI affects student learning.

What we know so far

  • General chatbots harm learning. Giving students open-ended access to chatbots like ChatGPT often pushes them toward outsourcing their thinking, relying on the tool, and skipping the mental work that learning depends on. The early evidence suggests this is not just neutral; it has a negative effect.
  • Short-term performance is not the same as learning. Students often perform better while an AI tool is available, then regress once it is taken away. That points to a real dependency risk: a tool that helps students complete the task in front of them does not automatically leave them with durable knowledge.
  • Pedagogical scaffolding shapes the outcome. The positive findings come from tools built specifically for learning, not generic solutions. The same base model can harm or help depending on the scaffolding around it.
  • Students already use AI. Schools do not fully control whether students encounter these tools. Many students already use them on their own, often without guidance from teachers. Estimates vary by age, but suggest that 50–80% of students already use tools like ChatGPT.
  • Student AI use often conflicts with learning. Left on their own, students tend to choose to use AI to save time and reduce effort. Few students use AI in ways that are compatible with long-term learning.

What we're still figuring out

  • Importance of model quality. Most studies use old models. That makes it hard to know whether the findings still hold for the tools in use right now. GPT-4, which appears in many studies, scores 13/100 on Artificial Analysis's Intelligence Index. State-of-the-art models in summer 2026, like Anthropic's Claude Fable 5, score 65/100. The research has not caught up to that jump in model quality. We need to know how much that matters. Does model speed change student engagement? Does it keep students in the flow, or does it make little difference? Does higher intelligence expand what models can teach, and are there diminishing returns after certain intelligence thresholds?
  • Differences across subjects. Most studies focus on STEM, especially math. We know much less about writing, history, languages, arts, and other subjects. AI may help more in some subjects than others. It may also need to be used differently depending on what is being taught.
  • Differences between students. Student profiles probably matter. A student with ADHD may need a different AI experience from a neurotypical student. The questions are whether we can adapt the tool to the student's needs and what kind of impact that would have.
  • Timing in the learning process. What a student already knows is likely to affect how well AI works for them. AI may be most useful after a student has built a basic understanding of the topic. Or it may help most earlier, as a way into the subject. We do not know yet where it fits best in the learning process.
  • Long-term effects are still unclear. Most studies look at short-term interventions: a few sessions, often less than a full semester. We still do not know what happens when students use these tools for months or years.

My reflections

For context, I read these studies as someone interested in building AI tools.

They do not tell you very much about what to build, but they can give you useful principles. They show where the traps are. They can help you avoid obvious mistakes. But in an environment changing this quickly, I think you need to be grounded in a particular point of view. You need your own vision for how AI and education should fit together. You can't find that in these studies and reports.

More importantly, the findings are not surprising if you view them from a learning science perspective. If a student is not deeply engaged and is mostly consuming or delegating, they will not learn much. We already knew that. The AI-specific research is useful, but at this stage I think the stronger guide is still established learning science, combined with a clear understanding of the capabilities of AI and how students actually use the tools.

The Evidence Base on AI in K-12: A 2026 Review - The best concise overview of the available studies and research in the field.

OECD Digital Education Outlook 2026 - A balanced and comprehensive big-picture overview of the current state of AI in education. It doesn't get into the weeds, but is still very informative and introduces useful concepts like fast vs. slow AI use. Chapters 1–3 are the most relevant to my interests.

Generative AI without guardrails can harm learning: Evidence from high school mathematics - Shows how much difference a simple tweak can make: students using a base chatbot learned less, while those using a version with pedagogical scaffolding did not. How well the results transfer to today is uncertain, though, given how much the models have improved since. It's also a cornerstone study frequently cited in major overview reports like the OECD Outlook.

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning - Explores proactive AI tutors and the substantial improvements they can bring.

A Scoping Review of Large Language Model-Based Pedagogical Agents - No findings of note, but its framework of four AI agent characteristics is helpful, even if potentially incomplete: interaction approach (reactive vs. proactive), domain scope (domain-specific vs. general-purpose), role complexity (single-role vs. multi-role), and system integration (standalone vs. integrated).