In a landscape where artificial intelligence is frequently championed as the ultimate panacea for educational inefficiency, a groundbreaking randomized controlled trial has revealed a more troubling reality. While proponents of generative AI argue that these tools can revolutionize the profession by automating lesson planning, generating classroom materials, and providing instant student feedback, new research suggests that the technology may actually be eroding the quality of instruction. According to a study led by researchers at the Wharton School of the University of Pennsylvania, the introduction of AI teaching assistants in real-world classrooms led to a measurable decline in student motivation and, in some cases, a significant drop in academic performance.
The research, which represents one of the first rigorous attempts to quantify the impact of AI on teaching through a randomized trial, indicates that teachers may be using generative AI as a "crutch" rather than a tool for enhancement. When instructors delegate the core intellectual tasks of teaching—such as crafting lecture notes and designing assignments—to an algorithm, the resulting output often lacks the personal voice and pedagogical nuance necessary to keep students engaged. This "delegation effect" has proven particularly damaging for students of instructors who were already struggling, creating a widening gap in educational quality.
The Mechanics of the Study: A Controlled Experiment in Turkey
The study, titled "Generative AI Can Harm Teaching," was conducted by Alp Sungu, an assistant professor at the Wharton School, along with a team of researchers that included renowned educational psychologist Angela Duckworth. The experiment took place during the spring of 2025 within a large private school chain in Turkey, involving 193 teachers and more than 2,800 middle and high school students.
To ensure the validity of the findings, the researchers employed a randomized controlled trial (RCT) design. Teachers were divided into two distinct groups. The "treatment" group was given access to a customized AI teaching assistant based on the ChatGPT architecture, which was specifically tailored to align with Turkey’s national curriculum. These teachers were encouraged to use the tool for various administrative and pedagogical tasks, including the generation of lecture notes, homework assignments, and exam questions. The "control" group, meanwhile, was instructed to continue their teaching practices as usual, though they were not strictly prohibited from using other commercially available AI tools if they so chose.
Over the course of ten weeks, the researchers monitored the interactions between teachers and the AI tool, while simultaneously tracking student outcomes. The results provided a stark contrast to the optimistic narratives often found in educational technology marketing. Rather than becoming more effective, many teachers in the treatment group appeared to outsource their creative and critical thinking to the AI, leading to a homogenized and less inspiring classroom experience.
Findings: Motivation, Achievement, and the "Answer Machine" Trap
The impact of the AI intervention was felt most immediately in the realm of student engagement. Students whose teachers had access to the AI assistant reported a significant decline in their intrinsic motivation. Through surveys, these students rated their classes as less interesting, less enjoyable, and less important than their peers in the control group.
This loss of interest appears to stem from the quality of the materials produced. When teachers use AI to generate content with minimal human oversight, the results often feel "uniform" and "boring." Alp Sungu noted that while AI-generated materials might be technically accurate, they frequently lack the teacher’s unique style and personal connection to the subject matter. "When you start using AI-generated material, you’re losing your personal voice," Sungu explained. "It might be technically good enough, but it doesn’t really carry your own style."
The study also highlighted a troubling trend regarding academic achievement. While the average achievement across all students did not show a drastic change, a deeper dive into the data revealed a concerning correlation. Students of teachers who were classified as "lower-performing" based on their pre-experiment benchmarks saw a notable decline in both their confidence and their scores on externally administered standardized exams. This suggests that while high-performing teachers may have the skill set to use AI as a starting point for further refinement, weaker instructors are more likely to use the AI’s output "as is," essentially lowering the bar for their own instruction.
This research echoes Sungu’s earlier 2024 work, which examined how students themselves use AI. In that study, he found that students often treat AI as an "answer machine" rather than a learning aid. The new study suggests that teachers are falling into a similar trap, using the technology as a "material generating machine" to bypass the rigorous process of lesson design.

The Timeline of AI Integration in Schools
The rapid adoption of generative AI in education has outpaced the scientific community’s ability to study its long-term effects. The timeline of this integration highlights why the Wharton study is so critical:
- Late 2022: The public release of ChatGPT triggers a global debate on the role of AI in schools, initially focusing on student cheating.
- 2023: EdTech companies begin integrating Large Language Models (LLMs) into "teacher dashboards," promising to save instructors hours of administrative work.
- Early 2024: Educational districts begin implementing AI policies, often encouraging teachers to experiment with AI-generated lesson plans to combat burnout.
- June 2024: Alp Sungu’s initial research on student AI use is released, warning that AI can erode problem-solving skills in mathematics.
- Spring 2025: The Turkey-based RCT is conducted, providing the first large-scale data on how teacher-facing AI affects classroom dynamics.
This chronology suggests that the "honeymoon phase" of AI in education is coming to an end, replaced by a more sober assessment of how these tools influence human behavior.
Supporting Data: A Closer Look at the Numbers
The quantitative data from the study provides a clear picture of the risks associated with unguided AI use. Key data points include:
- Motivation Deficit: Students in the AI treatment group showed a statistically significant decline in intrinsic motivation metrics. The decline was most pronounced among teachers who were already "heavy users" of AI before the experiment began, suggesting that over-reliance on the tool compounds its negative effects.
- Achievement Gap: In the bottom quartile of teacher performance, student test scores on standardized exams dropped significantly when those teachers used AI. Because the exams were externally administered, researchers were able to rule out "grade inflation" or biased grading as factors.
- The Time Paradox: While AI is marketed as a time-saver, the study found that for the output to be truly effective, teachers must spend an "equal amount of time" revising and calibrating it. Sungu noted that his own experience with the tool required significant time to ensure examples and numbers were accurate and relevant.
Analysis: The Risk of Educational "Delegation"
The broader implications of this study point to a fundamental misunderstanding of what makes teaching effective. Teaching is not merely the delivery of information; it is a relational and creative process. When a teacher designs a lesson, they are subconsciously tailoring the content to their specific students’ needs, cultural contexts, and prior knowledge.
Generative AI, in its current form, operates on statistical probabilities of word sequences. It does not "know" the students in a specific classroom in Turkey or anywhere else. If a teacher uses an AI-generated syllabus without modification, they are effectively removing the "human in the loop." This results in a curriculum that is technically sound but emotionally and intellectually hollow.
Furthermore, the study suggests that AI may exacerbate existing inequalities within the education system. If high-quality teachers use AI to augment their work while lower-quality teachers use it to replace their work, the gap in instructional quality will only widen. Students in under-resourced schools or those with less experienced staff may find themselves being taught by "automated" instructors, while students in elite environments continue to receive personalized, human-centric education.
Perspectives from the Field and Future Guardrails
The findings have sparked a necessary conversation among educators and policymakers. While the study indicates that "AI can harm teaching," the researchers are careful not to suggest that the technology should be banned entirely. Instead, the focus is shifting toward "augmentation" rather than "delegation."
"Access to AI technology alone does not improve teaching," Sungu cautioned. The challenge for the future lies in developing better interfaces and teacher training programs. Potential guardrails could include:
- Pedagogical Training: Teaching instructors how to use AI as a "sparring partner" for ideas rather than a primary author.
- Interface Design: Creating AI tools that require more human input and "forcing" teachers to review and edit content before it can be finalized.
- Curriculum Integrity: Establishing standards for how much AI-generated content can be used in a classroom without human adaptation.
As school districts around the world continue to invest millions in AI integration, the Wharton study serves as a vital warning. Efficiency should not be confused with effectiveness. If the goal of education is to inspire and challenge the next generation, then the "personal voice" of the teacher must remain at the heart of the classroom. Without it, AI may simply be making the process of learning faster, but the quality of that learning significantly poorer.
