29 May 2026, 11:00-12:00 CET
Online via Zoom
Or live stream viewing at Kolingasse 14-16 Room 2.38
Abstract
Internet memes are one of the most creative and culturally rich forms of online communication — blending images and text to express irony, sarcasm, humour, and metaphor in ways that go beyond literal interpretation. But how well do today's multimodal large language models actually text a meme?
This talk presents a benchmark study evaluating eight state-of-the-art MLLMs across three model series on their ability to detect and explain figurative meaning in memes, covering six figurative categories and three datasets. Beyond simple classification accuracy, the study goes a step further: through human evaluation, it assesses whether the models' explanations actually support their predictions and remain faithful to the original meme content. The results reveal a striking systematic bias — models tend to predict figurative meaning simply because they receive an image-text pair as input, regardless of whether figurative meaning is actually present. Case studies further illustrate how this bias manifests in practice and what it reveals about the gap between prediction and genuine understanding.
Speaker
Shijia Zhou is a second-year PhD student and MCML junior member at MaiNLP (Ludwig Maximilian University of Munich), supervised by Prof. Dr. Barbara Plank. She is part of the Klima-Memes project, an interdisciplinary research initiative examining climate change communication through the lens of internet memes.
Her research sits at the intersection of multimodal reasoning and computational analysis of online discourse, with a particular focus on climate change. Her work spans entity linking in climate change data (CLIMATELI), stance detection and media framing in climate memes (CLIMATEMEMES), and meme reasoning in multimodal large language models. She combines NLP methodology with insights from communication science to better understand how meaning is constructed and conveyed in multimodal online content.