◉ Colloquy — research, out loud

Using large language models to estimate features of multi-word expressions: Concreteness, valence, arousal

Gonzalo Martínez, Juan Diego Molero, Sandra González, Javier Conde, Marc Brysbaert, Pedro Reviriego

Behavior Research Methods · 2024 · doi:10.3758/s13428-024-02515-z

The episode · 9 min · Researchers A & B
AI episode generated 2026-08-31 from the open-access full text · model p1.0 · every number checked against the source · claims table · report an error

Abstract

This study investigates the potential of large language models (LLMs) to provide accurate estimates concreteness, valence, and arousal for multi-word expressions. Unlike previous artificial intelligence (AI) methods, LLMs can capture nuanced meanings We systematically evaluated GPT-4o's ability predict arousal. In Study 1, GPT-4o showed strong correlations with human concreteness ratings (r = .8) 2, these findings were repeated valence individual words, matching or outperforming AI models. Studies 3-5 extended analysis expressions good validity LLM-generated stimuli as well. To help researchers stimulus selection, we datasets norms 126,397 English single words 63,680

Transcript

00:00 Cold Open

Researcher A Here's a finding that might seem obvious until you try it: GPT-4o can estimate whether a phrase is concrete or abstract, positive or negative, exciting or calm—and it does it almost as well as humans do. The catch? Most of those phrases don't exist in any dataset yet. When you say "shoot a film" or "summer vacation," there's no human rating sitting in a database waiting for you. Until now.

Researcher B So the AI is learning to rate things people have never formally rated before?

Researcher A Exactly. And the team behind this paper—led by Gonzalo Martínez and Marc Brysbaert—validated it carefully. They didn't just assume the model was right. They had real people rate phrases and checked whether GPT-4o matched them.

00:48 Why This Exists

Researcher B Why is this a problem worth solving? Don't we already have ratings for individual words?

Researcher A We do—there are huge datasets of human ratings for single words. But here's the thing: about half of what we actually say uses multi-word expressions. Idioms, compound nouns, phrasal verbs—"throw up," "a drop in the ocean," "good morning." These aren't just the sum of their parts. "Blind spot" doesn't mean what "blind" and "spot" mean separately.

Researcher B Right, so you can't just average the ratings of the individual words.

Researcher A Exactly. And until recently, AI couldn't handle that either. Word2vec and older semantic models only worked with single words. But LLMs—large language models—they process whole phrases as units. So theoretically they should be able to capture that idiomatic meaning. The question was: do they actually?

01:42 What They Actually Did

Researcher B Walk me through the studies.

Researcher A They ran five studies, building up evidence. Study 1 was the proof of concept. They took 62,889 multi-word expressions that had already been rated for concreteness by Muraki and colleagues in 2023. Then they asked GPT-4o to rate those same expressions using a simple prompt—a 1-to-5 scale with examples at each end, like they'd give to a human rater.

Researcher B How did they extract the rating from the model?

Researcher A That's the clever part. They didn't just take the most likely answer. They used something called logprobs—the model's estimated probability for each possible response. So if the model said "4" with 64.6 percent confidence, "3" with 34.6 percent, and so on, they combined those probabilities into a weighted score. For "shoot a film," that gave them 3.66 instead of just a 4.

Researcher B And in Studies 2 through 5?

Researcher A Study 2 tested whether the same approach worked for valence—how positive or negative a word feels—and arousal, how energizing or calming. They compared GPT-4o against human ratings from Warriner and colleagues, plus ratings from other AI models. Studies 3 and 4 extended that to multi-word expressions. Study 4 was the validation: they had 16 new participants rate 96 randomly selected phrases on a 7-point scale, and compared those human ratings to GPT-4o's estimates. Study 5 did the same for arousal with 23 participants.

03:24 What They Found

Researcher B Okay, so what are the actual numbers?

Researcher A Study 1: GPT-4o's concreteness ratings correlated r equals .80 with the human ratings. The Muraki dataset itself has a reliability of .84, so .80 is nearly at the ceiling—almost as good as you could hope for.

Researcher B That's strong. What about valence and arousal?

Researcher A In Study 2, testing single words, GPT-4o's valence estimates correlated .90 with Warriner's human ratings. For arousal, it was .74. Both of those matched or beat other AI approaches—including older semantic vector methods and even GPT-3 fine-tuned specifically on those datasets.

Researcher B And the validation with new participants?

Researcher A Study 4, valence: r equals .95 between GPT-4o and the newly collected human ratings. Study 5, arousal: r equals .92. Those are remarkably high. The authors note that's partly because they deliberately spread their test phrases across the full range of GPT estimates—they didn't cherry-pick easy cases.

Researcher B What about those tricky idioms where the meaning is totally figurative?

Researcher A Good catch. Idioms like "a golden key can open any door"—humans rated it 1, very abstract. GPT-4o said 2. Not perfect, but closer than previous AI tools. The authors looked at 486 frequent idioms and found a correlation of .56 for those, much lower than the .81 for the full dataset. Opaque idioms are still hard. But interestingly, when GPT-4o missed by a lot, it was more likely to underestimate the concreteness of idioms than overestimate. It erred on the side of caution.

05:19 Caveats

Researcher B What are the limitations here?

Researcher A The paper itself flags several. First, the distribution of GPT-4o's ratings is different from humans'. The model gives more extreme values—more 1s and 5s, fewer middle values. Humans tend to cluster toward the center. So if you're doing a study where you need a careful spread of stimuli, you might want to use the model's rank scores rather than the raw numbers.

Researcher B That makes sense. What else?

Researcher A Opaque idioms, as we mentioned, are tougher. The correlation drops to .56 for those. The authors also note that the validation studies used relatively small samples—16 people for valence, 23 for arousal. Those are enough to show the model generalizes, but they're not huge.

Researcher B And beyond what the authors say?

Researcher A Worth noting: this is English-only. The authors mention GPT-4o works in other languages, but not as well yet. Also, there's always the question of data leakage—could the human datasets used for validation have been in GPT-4o's training material? The authors address this by validating on new, previously unrated phrases, which helps. But we can't fully rule it out. And finally, this is a snapshot of GPT-4o from mid-2024. Models change. These results might not hold for future versions.

06:49 Who Should Care

Researcher B Who's the audience for this?

Researcher A Three groups, I'd say. First: psycholinguists and cognitive scientists studying how we understand language. If you're running an experiment on reading or word recognition, you need to know the concreteness, valence, and arousal of your stimuli. Now you can get those estimates for phrases you're interested in without collecting human ratings yourself.

Researcher B That saves time and money.

Researcher A Exactly. Second: sentiment analysis and NLP researchers. If you're analyzing text for emotion or meaning, knowing the valence and arousal of multi-word expressions is crucial. The authors provide datasets of 126,397 single words and 63,680 multi-word expressions with all three properties pre-estimated.

Researcher B That's a huge resource.

Researcher A It is. And third: anyone doing computational linguistics or building language models. This work shows that LLMs can capture semantic properties of phrases in a way older AI couldn't. It's a proof of concept that might inform how we think about what these models know and how to use them.

08:05 Outro

Researcher A The full citation is: Gonzalo Martínez, Juan Diego Molero, Sandra González, Javier Conde, Marc Brysbaert, and Pedro Reviriego. "Using large language models to estimate features of multi-word expressions: Concreteness, valence, arousal." Published in Behavior Research Methods, 2025, volume 57, article 5. The DOI is 10.3758, slash, s13428, dash, 024, dash, 02515, dash, z. All the data and code are freely available at the Open Science Framework.

Researcher B And researchers can actually download those pre-estimated norms?

Researcher A Yes—Excel files with the estimates for all 126,397 words and 63,680 phrases. Creative Commons license, free for research and education. The thread is open on Colloquy.