New NERDS paper on conversational biases in AI multi-agent systems

Unmasking conversational bias in AI multiagent systems by Erica Coppolillo, Giuseppe Manco, and Luca Maria Aiello in PLOS One.

Two matrices are shown. Left: Average number of agents changing opinion during the conversation. Right: Conditional probability that the second agent follows once the first has already drifted. Empty cells indicate configurations where no agent displayed a drift, while cells with the “-” symbol indicate unavailable results. The darker the color, the higher the reported value.

New paper on PLOS One by Luca Aiello Detecting biases of generative AI is critical, but it is often done considering models in isolation. In particular, biases emerging from interactions among conversational agents remain largely unexplored. In this paper we present a framework designed to quantify biases within multi-agent systems of conversational agents. We simulate echo chambers where agents are initialized with aligned perspectives on a polarizing topic and asked to develop the topic in multi-turn discussions. Surprisingly, we observe that, despite the echo-chamber setting, the agent stance shifts away from their initial position, often towards liberal positions. Crucially, the bias observed in these echo-chamber experiments remains undetected by traditional bias detection methods that probe models in isolation. This highlights a critical need for the development of a more sophisticated toolkit for bias detection and mitigation for AI multi-agent systems.