NextLUCA
← Back to blog

ChatGPT (GPT-4o) Study at Bocconi University Shows AI Boosts Idea Quality, Not Idea Diversity

A large-scale classroom experiment at Bocconi University put GPT-4o and causal-reasoning training head to head, and the results complicate the simple story that AI just makes student work better. Access to ChatGPT sharpened logical coherence and pulled answers closer to expert recommendations, but a different intervention entirely, one with no AI involved, was what actually made students think more originally.

A large-scale classroom experiment at Bocconi University put GPT-4o and causal-reasoning training head to head, and the results complicate the simple story that AI just makes student work better. Access to ChatGPT sharpened logical coherence and pulled answers closer to expert recommendations, but a different intervention entirely, one with no AI involved, was what actually made students think more originally.

What's new

  • Researchers randomly assigned over 1,000 first-year students at Bocconi University to one of four conditions: ChatGPT access, causal-reasoning training, both, or neither, then had them develop marketing recommendations for the university's merchandise store.
  • ChatGPT access raised rubric scores by almost a full point on the five-point scale, increased the number of ideas per answer, improved logical coherence, and made answers more similar to expert recommendations.
  • Causal-reasoning training, which had nothing to do with AI, did not move the traditional rubric score at all, but it improved students' explanations of why an idea might succeed or fail, increased the variety of ideas produced, and made each student's ideas more distinct from their peers'.
  • Students who received both interventions captured benefits from each and showed gains across the widest range of measures compared to any single-intervention group.
  • The randomized design let researchers isolate the two effects cleanly, showing that AI access and reasoning training improve different things rather than duplicating each other.
  • The findings support grading originality and reasoning quality alongside the polish of a final answer, since a high rubric score does not guarantee original thinking.
  • Conclusion
  • The experiment shows that ChatGPT and causal-reasoning training improve two different things, and neither substitutes for the other. AI polishes the output — higher scores, more ideas, tighter logic — but it strengthens exactly what every classmate is also using, so answers converge toward each other and toward standard expert advice. Causal-reasoning training raised no rubric score, yet it did what the tool cannot: students could explain why an idea would work or fail, and their ideas stood apart from the crowd. The group that received both interventions gained across the widest range of measures.
  • The practical lesson for education is this: adding AI to a classroom without strengthening causal thinking produces students who write well and think alike. And the lesson for assessment is that grading the final answer is no longer enough — once a tool can produce polished prose, what must be measured is the originality of the idea and the quality of the reasoning behind it, not the surface of the response.

How to try it

Sources