30 juni 2026
‘Our research is a reminder that AI systems are fundamentally different from humans,' says Tong. 'That AI systems can produce fluent language does not automatically mean they understand language in the same way humans do.'
Metaphors and humour are not merely decorative features of language. They help people explain complex ideas, express emotions, build relationships and make sense of the world around them. Every day, people use phrases such as "grasping an idea" or "attacking an argument" without consciously thinking about the figurative meanings involved.
Humour relies on similar mental processes. Understanding a joke often requires recognising unexpected connections, contradictions or shared cultural references.
‘Metaphors and jokes require us to connect ideas that don't obviously belong together,’ says Tong. ‘Humans typically do this naturally, but for AI systems it remains surprisingly difficult.’
To investigate these challenges, Tong developed several new datasets and benchmarks designed specifically to test AI systems like ChatGPT.
One of these, the Metaphor Understanding Challenge Dataset (MUNCH), contains more than 10,000 carefully annotated paraphrases of metaphorical sentences. When tested against this benchmark, leading language models frequently struggled to distinguish between literal and figurative meanings and often failed to capture the intended interpretation.
Tong also explored whether AI can recognise why a metaphor is being used. Metaphors can serve many different purposes, from explaining complex concepts and persuading audiences to creating vivid imagery or strengthening social bonds.
As AI becomes an ever-larger part of everyday life, it needs to do more than just process information, it needs to understand the subtleties of human communication.Xiaoyu Tong
Together with colleagues, Tong developed the first large-scale taxonomy of metaphorical intentions, identifying nine distinct categories. While current AI models could recognise some of these intentions reasonably well, they continued to encounter difficulties across many scenarios.
Humour proved particularly challenging when both text and images were involved.
To explore this, Tong created the Hummus dataset, a collection of 1,000 cartoons from The New Yorker. Many of these cartoons rely on visual metaphors and subtle connections between images and text to create humorous effects.
The study found that even state-of-the-art AI systems often struggled to combine visual and textual information in ways that allowed them to get the joke.
‘Understanding a joke is about much more than recognising words,’ says Tong. ‘It often requires cultural knowledge, social context and an appreciation of what makes a situation unexpected.’
As AI becomes more deeply integrated into everyday life, from virtual assistants and educational tools to workplace applications, understanding human communication becomes increasingly important.
Tong’s findings suggest that improving AI's ability to understand metaphor, humour and cultural context could help create future systems that are more reliable, more intuitive and ultimately more useful for the people who use them.
‘If an AI assistant doesn’t fully understand our metaphors or jokes, it could have implications for its usability,’ says Tong. ‘As AI becomes an ever-larger part of everyday life, it needs to do more than just process information, it needs to understand the subtleties of human communication.’