A groundbreaking study led by Professor Karim Jerbi from the Department of Psychology at the Université de Montréal, featuring the participation of world-renowned AI pioneer Yoshua Bengio, has unveiled unprecedented insights into the capabilities of generative artificial intelligence systems like ChatGPT in producing original ideas. This monumental research represents the largest direct comparison ever conducted between human creative output and that generated by large language models (LLMs), offering a nuanced perspective on the evolving landscape of creativity in the digital age. Published in the esteemed journal Scientific Reports (Nature Portfolio), the findings mark a significant inflection point, indicating that advanced AI systems have now reached a proficiency where they can surpass the average human on specific creativity metrics, while simultaneously reaffirming the enduring, distinct advantage held by the most exceptionally creative individuals.
A New Benchmark in AI Capabilities
The study’s central revelation is that certain generative AI systems have achieved, and in some cases exceeded, the average human capacity for divergent linguistic creativity. Researchers meticulously evaluated several leading large language models, including OpenAI’s GPT-4, Anthropic’s Claude, and Google’s Gemini, alongside other sophisticated AI platforms. Their performance was then rigorously compared against a massive dataset derived from over 100,000 human participants, providing a robust statistical foundation for the conclusions drawn. This extensive comparative analysis unequivocally points to a turning point: AI systems like GPT-4 are now demonstrating scores on creativity tasks that outstrip the mean performance of human participants.
Professor Karim Jerbi articulated the dual significance of these findings: "Our study shows that some AI systems based on large language models can now outperform average human creativity on well-defined tasks. This result may be surprising—even unsettling—but our study also highlights an equally important observation: even the best AI systems still fall short of the levels reached by the most creative humans." This statement underscores a critical distinction, moving beyond a simplistic "AI vs. Human" narrative to a more granular understanding of where AI stands in the creative spectrum.
Further detailed analysis, spearheaded by the study’s co-first authors, postdoctoral researcher Antoine Bellemare-Pépin from the Université de Montréal and PhD candidate François Lespinasse from Concordia University, illuminated a consistent pattern. While a subset of AI models now demonstrates superior performance compared to the average person, the pinnacle of creative ideation and output remains an exclusively human domain. When the research team focused on the upper echelons of human creativity, specifically examining the most creative half of the participant pool, their average scores consistently surpassed those of every AI model subjected to testing. This disparity became even more pronounced and statistically significant when comparing AI performance against the top 10 percent of the most creatively gifted individuals.
Professor Jerbi, who also holds an associate professorship at Mila – Quebec AI Institute, emphasized the methodological rigor: "We developed a rigorous framework that allows us to compare human and AI creativity using the same tools, based on data from more than 100,000 participants, in collaboration with Jay Olson from the University of Toronto." This commitment to a standardized evaluation framework was crucial for ensuring the fairness and validity of the comparisons.
The Evolution of AI and the Quest for Creativity
The journey of artificial intelligence has been marked by a relentless pursuit of capabilities once thought to be exclusively human. From the early symbolic AI systems of the mid-20th century, which relied on handcrafted rules and logical reasoning, to the statistical machine learning models of the late 20th and early 21st centuries, AI has steadily expanded its cognitive repertoire. However, the realm of creativity—often characterized by originality, imagination, and the ability to generate novel connections—remained a formidable challenge.
The dramatic acceleration in AI’s creative potential can be largely attributed to the advent of deep learning and, more specifically, the development of transformer architectures in the mid-2010s. Pioneers like Yoshua Bengio, whose foundational work in deep learning has been instrumental in the current AI revolution, laid the groundwork for large language models. These models, trained on unfathomable volumes of text data—spanning books, articles, websites, and more—learn to recognize intricate patterns, linguistic structures, and conceptual relationships at a scale unimaginable to humans. Early LLMs like GPT-1 and GPT-2 demonstrated nascent abilities to generate coherent text, but it was with models like GPT-3 and subsequent iterations (including GPT-4, Claude, and Gemini) that AI began to exhibit what appeared to be genuine creative flair.
Before these advanced LLMs, AI’s attempts at creativity were often viewed as sophisticated mimicry. Programs could generate art in the style of famous painters or compose music following specific genre conventions, but the underlying process was typically rule-based or pattern-matching, lacking the spark of true originality. The current study, by focusing on divergent creativity, directly addresses whether LLMs can move beyond imitation to produce genuinely novel and diverse ideas, a hallmark of human creativity. This research builds upon decades of philosophical debate about whether machines can possess consciousness or genuine intelligence, shifting the focus to a more measurable aspect: creative output.
Deconstructing Creativity: How Machines and Humans Were Judged
To ensure a fair and equitable assessment of creativity across both human and artificial intelligences, the research team employed a multi-faceted methodological approach. The primary instrument for evaluation was the Divergent Association Task (DAT), a widely recognized and validated psychological test designed to measure divergent creativity. Divergent creativity, a cornerstone of innovative thinking, refers to an individual’s capacity to generate a multitude of diverse and original ideas or solutions in response to a single prompt.
The DAT, originally conceived by study co-author Jay Olson, operates on a deceptively simple premise: participants, whether human or AI, are instructed to list ten words that are as semantically unrelated as possible. The power of the DAT lies in its ability to quantify the breadth and originality of an individual’s conceptual associations. A highly creative response, as exemplified by the researchers, might include a list such as "galaxy, fork, freedom, algae, harmonica, quantum, nostalgia, velvet, hurricane, photosynthesis." The disparate nature of these words signals a mind capable of traversing vast conceptual distances and forging novel connections—a key indicator of creative potential.
Performance on the DAT has been consistently correlated with results on other established creativity tests, including those used in creative writing assessments, idea generation tasks, and problem-solving scenarios. While the task is fundamentally language-based, its demands extend far beyond mere vocabulary recall. It actively engages broader cognitive processes vital to creative thinking across a wide array of domains. Moreover, the DAT offers significant practical advantages: it is concise, typically requiring only two to four minutes for completion, and its online accessibility facilitates large-scale data collection, as evidenced by the 100,000+ human participants in this study.
Beyond the core DAT, the researchers further probed AI’s creative capacities by extending their comparison to more intricate and realistic creative activities. They challenged both AI systems and human participants with tasks requiring creative writing, such as composing haiku (a concise, three-line poetic form), crafting compelling movie plot summaries, and developing short stories. The objective was to ascertain whether AI’s success in simple word association could translate into more complex narrative and artistic endeavors.
The results from these extended creative writing tasks mirrored the pattern observed with the DAT. While AI systems occasionally demonstrated an ability to surpass the output of average human writers, the most accomplished and skilled human creators consistently produced work that was not only stronger in quality but also more profoundly original and imaginative. This finding reinforces the idea that while AI can capably handle the mechanics of creative expression, the nuanced depth, emotional resonance, and truly novel conceptualization often found in peak human creativity remain unparalleled.
Tuning the Creative Algorithm: The Role of Human Guidance
A fascinating dimension of the study explored the malleability of AI creativity: is it a fixed attribute, or can it be influenced and refined? The research unequivocally demonstrated that AI’s creative output is not static but can be significantly adjusted through various technical settings and, crucially, through human guidance. A key technical parameter identified was the model’s "temperature," which essentially controls the predictability and adventurousness of the generated responses.
At lower temperature settings, AI models tend to produce more conservative, conventional, and predictable outputs, sticking closely to the most probable patterns learned from their training data. This is akin to an artist carefully following established techniques. Conversely, when the temperature is increased, the AI becomes more exploratory, generating responses that are more varied, less predictable, and venture further from familiar ideas. This higher temperature allows the system to make less probable connections, leading to outputs that can be perceived as more original or "creative." The study highlights that by manipulating this temperature parameter, users can effectively dial AI’s creative output up or down, making it a powerful tool for tailored creative assistance.
Equally significant was the finding that the way instructions are formulated—a practice known as "prompt engineering"—exerts a profound influence on AI creativity. The research showed that prompts designed to encourage models to delve into the origins and structural relationships of words, for instance, by leveraging etymology, led to a greater number of unexpected associations and, consequently, higher creativity scores. These results underscore a critical insight: AI creativity, while powerful, is not autonomous. It is heavily dependent on the quality and specificity of human guidance. This dynamic interaction, where human intent shapes algorithmic output, positions prompt engineering as a central and indispensable component of the contemporary creative process. It transforms AI from a passive generator into an active collaborator, whose potential is unlocked and directed by human ingenuity.
Beyond Competition: AI as an Amplifier of Human Creativity
The comprehensive findings of this study offer a balanced and sophisticated perspective on the pervasive anxieties surrounding artificial intelligence’s potential to displace human creative professionals. While the research confirms that AI systems can now equal or even exceed average human creativity on certain defined tasks, it simultaneously highlights their inherent limitations and their foundational reliance on human direction.
Professor Karim Jerbi eloquently articulated this shift in perspective: "Even though AI can now reach human-level creativity on certain tests, we need to move beyond this misleading sense of competition. Generative AI has above all become an extremely powerful tool in the service of human creativity: it will not replace creators, but profoundly transform how they imagine, explore, and create—for those who choose to use it." This statement reframes the narrative from one of replacement to one of augmentation and transformation.
Rather than portending the obsolescence of creative careers, the study’s implications strongly suggest a future where AI functions as an indispensable creative assistant. Imagine a scenario where a writer uses AI to brainstorm novel plot twists, a designer employs it to generate countless variations of a concept, or a musician leverages it to explore new melodic structures. By rapidly expanding the universe of ideas and opening new pathways for exploration, AI possesses the capacity to amplify human imagination, liberating creators from repetitive tasks and enabling them to focus on higher-order conceptualization, emotional depth, and truly unique artistic vision.
This collaborative paradigm could lead to an unprecedented flourishing of human creativity, allowing individuals to push boundaries further and faster than ever before. It also necessitates a re-evaluation of what creativity truly means in an age where machines can generate novel content. "By directly confronting human and machine capabilities, studies like ours push us to rethink what we mean by creativity," Professor Jerbi concluded. This redefinition will likely emphasize uniquely human attributes such as intentionality, lived experience, emotional intelligence, and the capacity for truly original, groundbreaking thought that resonates deeply with the human condition.
The economic and societal implications are profound. In industries ranging from advertising and media to entertainment and product design, AI tools are poised to revolutionize workflows, enhance productivity, and democratize access to creative tools. However, this transition will also demand new skill sets, particularly in prompt engineering and the critical evaluation of AI-generated content. Ethical considerations surrounding authorship, intellectual property, and the potential for algorithmic bias will also need careful navigation as AI becomes more deeply embedded in the creative ecosystem.
The Collaborative Frontier: Uniting Leading Minds in AI Research
The seminal paper, titled "Divergent creativity in humans and large language models," was officially published in Scientific Reports on January 21, 2026. This ambitious research endeavor was the result of a collaborative effort involving scientists from several prestigious institutions, including the Université de Montréal, Concordia University, the University of Toronto Mississauga, Mila (Quebec AI Institute), and Google DeepMind.
Professor Karim Jerbi served as the lead investigator, guiding the interdisciplinary team through this complex exploration. The critical roles of co-first authors were fulfilled by Antoine Bellemare-Pépin from the Université de Montréal and François Lespinasse from Concordia University, who were instrumental in the detailed analysis and interpretation of the vast datasets. The research team was further bolstered by the invaluable participation of Yoshua Bengio, the visionary founder of Mila and LoiZéiro, whose pioneering contributions to deep learning have been foundational to the development of the very AI systems, such as ChatGPT, that were the subject of this study. Bengio’s involvement lends significant weight and credibility to the findings, bridging the gap between theoretical AI advancements and their practical implications for human cognition and creativity. This collaborative spirit, uniting diverse expertise, has not only yielded a landmark study but also charted a clearer course for understanding the intricate dance between human and artificial intelligence in the boundless realm of creativity.




