Group size effects and collective misalignment in LLM multi-agent systems by Ariel Flint, Luca Maria Aiello, Romualdo Pastor-Satorras, and Andrea Baronchelli published in PNAS.

As AI agents begin to operate in populations rather than one at a time, our new research suggests that the number of them changes what they collectively decide — amplifying a bias, inventing one from nothing, or flipping a group into the opposite of what each agent would choose alone. When artificial intelligence (AI) agents interact in groups, their number is not merely a technical detail. It is a decisive factor in what the group settles on: populations built from the same AI model, doing the same task, can reach opposite outcomes for no other reason than that one group is bigger.
In this work, we experimented with populations of LLM agents playing the “naming game”, a classic framework for studying how conventions emerge, in which randomly paired agents each pick a word from a shared pool and are rewarded when they happen to pick the same one. Agents see only their own recent interactions, never the wider population, and are never told they are in a group. Over many pairings, a population can converge spontaneously on a shared convention.
We found that interactions can pull a group away from what its members individually want in three ways. It can amplify an existing leaning until the group converges on it almost every time. It can induce a preference out of nothing, with populations of individually neutral agents reliably favouring one word over an equally viable alternative. And it can reverse a preference outright, so that a population settles on the word its own members disfavoured. Group size then determines how strongly these preferences bite, in ways that cannot be extrapolated. Larger populations became more predictable across every model and word pair tested, converging on one word until the outcome was effectively certain. But the size at which that tipping point arrived varied enormously: for some combinations as few as two agents, for others around ten thousand.
Overall, our results demonstrate that more is different for LLM populations: The number of interacting agents is a key driver of the dynamics, with implications for the design and governance of multi-agent AI systems.
