konstanz-dynamic-social-behavior

Week 5 — Populations of LLM Agents


Reading guide

  Sub-mode of 🟫 Paper
🟫 Minimal-agent mechanism replication Ashery 2025
🟫 Rich-persona descriptive emergence, scaled Altera Team 2024
🟫 Policy-testbed instrumentation Larooij & Törnberg 2025

The Week-4 chapter’s reading-guide table places these three alongside Park 2023 (Smallville) under 🟫; this chapter’s job is to fill in what each adds. None of the new papers introduces a new colour: they are three different points in a design space, not three different methods.

🐍 Notebooks: 5_ashery.ipynb · 5_heatwave_simulation.ipynb


2. 🟫 Ashery — minimal agents, classical mechanism

Ashery 2025 is the cleanest of the three in its experimental setup. The paper takes a 50-year-old model from the convention-formation literature — the naming game (Baronchelli et al. 2006) — and asks whether its dynamics survive when the agents are LLMs. The methodological move is the inverse of Park 2023’s: rather than enrich the agent and watch what emerges, strip the agent down and watch whether a known macro pattern still appears.

The naming game is pairwise. Two agents meet, each picks a label for an unnamed object from a pool, and both are rewarded if the labels match. With a memory of recent partners’ choices, single names spread across a population through nothing more than coordination pressure. In humans, the dynamics have been studied across scales (Centola et al. 2018 is the experimental capstone); the prediction is a sharp transition to consensus, sensitive to network topology and to the size of the committed minority.

Three results from Ashery’s LLM populations are worth naming.

Convention emergence. A population of LLM agents, paired at random, converges to a single label even though no agent sees the population state and no agent has a preference over labels. The dynamics replicate the human version qualitatively.

Collective bias from unbiased agents. The same simulation, run many times, does not converge to each label with equal probability. There are systematic asymmetries — the population becomes biased even though no individual agent is. This is the Schelling shape: macro outcomes more extreme than any micro preference would suggest. Reading this through Week 4: it is the population-level analogue of Park 2024’s individual-prediction setup, except now the regularity being captured is not a person’s behaviour but a sociological mechanism.

Committed-minority tipping. When a small subset of agents is hard-coded to insist on a non-majority label, the rest of the population can flip. The critical-mass thresholds Ashery reports are model-dependent — they sit somewhere between 2% and 67% across model and topology variants — but the shape is Centola et al. 2018’s.

The paper’s strongest move is the choice of comparator. Convention formation, collective bias, and committed-minority tipping are all results that humans-in-the-lab established before any LLM was trained on them, and the human results sit in a tradition (Lewis 1969; Schelling 1971; Granovetter 1978; Watts 2002) the simulation does not need to be told about. If the LLMs reproduce the dynamics, that is consilience; if they do not, the failure is interpretable. This is the test-of-the-mechanism move, in the sense behavioral-GT readers will recognise from controlled experimentation.

The paper is also where the echo-of-the-literature concern is sharpest. The naming game has been written about in journals the LLM almost certainly trained on. A model that produces the right qualitative dynamics has not, by that fact, demonstrated that the dynamics arose from coordination pressure rather than from text-completion of the relevant social-physics paper. The cross-model robustness checks help — the paper runs Llama-3 and Claude alongside GPT-4 — but the concern is structural, not model-specific. Ashery’s claim is that the same simulation, run with three different LLMs, would not produce three different Baronchelli-style replications by accident. That is plausible; it is not a proof.

The Week-5 tutorial in JN/week5/5_ashery.ipynb is built around this paper. The notebook is small — naming game in a few hundred lines — and it is small because the paper itself is. Where Park 2023 needed memory streams, reflection, and planning, Ashery needs a pair of dictionaries and a coordination check. The methodological lesson, for a reader holding both papers in view, is that the right cognitive architecture is a function of the question: rich emergence needs rich agents; mechanism replication does not.


3. 🟫 Project Sid — rich emergence at scale

Altera Team 2024, Project Sid, is Park 2023 scaled by two orders of magnitude. Smallville ran 25 agents in a small town for two simulated days; Project Sid runs 10 to 1,000+ agents in a Minecraft world for substantially longer. The cognitive architecture, called PIANO, keeps an explicit modular structure — perception, planning, memory, inter-agent communication — where Park 2024 had argued that long-context models could absorb most of those modules’ load. The architectural choice itself is the paper’s first methodological claim: in multi-agent settings, the architecture-versus-context tradeoff Park 2024 resolved in favour of context still tilts toward architecture.

Three families of phenomena are reported.

Specialisation. Agents in larger populations differentiate into roles — gatherers, builders, traders — at rates that smaller populations do not produce. The shape is Durkheimian: as population density increases, role differentiation increases. The civilisation moves from Durkheim 1893’s mechanical solidarity (homogeneous, undifferentiated) toward organic solidarity (differentiated, interdependent), in a model where the only thing being scaled is N.

Cultural transmission. Agents adopt and propagate practices — religious rituals being the most striking example — through a mixture of imitation and explicit instruction. The pattern echoes the cultural-evolution tradition (Boyd & Richerson 1985; Henrich 2016): traits spread non-uniformly, prestige and successful agents are copied more, and the trait composition of the population evolves over time.

Emergent governance. Rule-like regularities (taxation, redistribution, sanctions for norm violation) appear in some runs without being seeded by the prompt. This is the most ambitious claim and the one where the echo-of-the-literature concern is hardest to discharge: governance terminology is everywhere in the training corpus, and an agent that generates a rule and another that complies has demonstrated that LLMs can roleplay institutions, not necessarily that institutions emerged.

Read against Week 4, the paper sits in a clear lineage. It is a natural extension of Park 2023 along three axes — agent count, world complexity, simulated time — and it provides the strongest case the literature currently has for the claim that multi-agent LLM simulation produces phenomena that single-agent and individual-prediction setups cannot. Park 2024 improved on Park 2023 for individual prediction by retreating from the architecture; Project Sid is the response that, for population-level phenomena, the architecture is doing work the long context cannot replace.

The paper’s weakest point, by its own framing, is the absence of a clean validation comparator. Smallville had a Prolific believability evaluation; Park 2024 had test-retest agreement; Hewitt had published treatment effects. Project Sid’s evaluations are mostly qualitative descriptions of what happens in the simulation. The civilisation’s plausibility is in the eye of the reader.


4. 🟫 Larooij & Törnberg — synthetic platform as policy testbed

Larooij & Törnberg 2025 is the most directly applied paper of the three. A synthetic social-media platform is populated with 500 LLM agents whose personas are drawn from the ANES 2020 survey, so the population’s distributions over partisanship, ideology, demographics, and interests reflect a real US sample rather than authorial whim. The platform is deliberately minimal: agents post in response to news headlines, repost timeline items, and follow other agents; there is no engagement-optimising recommender, no content moderation, no images. Six prosocial interventions drawn from the platform-design literature — chronological feeds, downplaying dominant content, boosting out-partisans, bridging-based ranking (Ovadya & Thorburn), hiding social statistics, and hiding biographies — are tested against a baseline.

        baseline platform                              intervention sweep
        ─────────────────                              ──────────────────
        500 ANES-anchored agents,
        10,000 simulation steps,                       1. chronological feed
        5 runs per condition.                          2. downplay dominant content
                                                       3. boost out-partisan exposure
        Three pathologies measured:                    4. bridging-attribute ranking
          E–I index               (echo chambers)      5. hide social statistics
          Gini of followers/      (attention            6. hide biography
            reposts                inequality)
          partisan ↔ engagement   (social media
            correlation             prism)             cross-model: gpt-4o-mini,
                                                                    llama-3.2-8b,
                                                                    deepseek-r1.

The headline finding is that the three pathologies — partisan echo chambers, attention inequality, the “social media prism” of Bail 2022 — emerge in the baseline without any algorithmic curation beyond a 5-of-10 popular-post feed slot, and that the six interventions deliver only modest gains. Two of the intervention results are more interesting than the headline.

Chronological feeds — long the activists’ first ask — flatten attention inequality dramatically (Gini of followers from 0.83 to 0.51) but worsen the prism: removed from an engagement backdrop, partisan-extreme content stands out more sharply, not less. Bridging-based ranking is the only intervention that breaks the partisan-engagement correlation, but it does so by funnelling attention to a narrow set of high-bridging-score posts, raising inequality. The two interventions with the largest effects on the metric they target both have a non-trivial cost on a different metric.

The paper is best read as a computational complement to the empirical platform-experiment literature. Bakshy 2015 showed that user choice did more to narrow exposure on Facebook than the algorithm did. Guess 2023 ran a substantial 2020-election collaboration that found that large algorithm changes produced only modest shifts in polarisation. Larooij’s simulation reproduces the shape of those empirical results — interventions help, but not by much — using a method that can examine counterfactuals the platform experiments cannot.

Two cross-Week-4 connections are worth naming.

Park 2024 and Larooij sit at opposite ends of a depth-versus-scale tradeoff. Park 2024 conditions on a two-hour interview per agent and gets close to that individual’s own self-consistency at t+2; Larooij conditions on an ANES persona summary and runs 500 of them through 10,000 steps. The population-level questions Larooij answers are not answerable in Park 2024’s setup, and vice versa. The two papers are not in competition; they are points on a frontier.

Hewitt and Larooij both predict intervention effects. Hewitt does so by asking GPT-4 directly to forecast a published treatment effect, with a strong correlation against ground truth on survey experiments and a weak one on field experiments. Larooij does so by simulating the population the intervention would act on and reading the effect off the simulation. The two methods are alternatives — forecast the experiment versus simulate the population — and they are likely to fail in different places. Anthis 2025’s spectrum from “exact replication” to “complete human-impossible studies” is exactly the axis these two papers’ methods sit on.

The paper’s main vulnerability, by the authors’ own admission, is the validation problem. The simulation’s metrics are internally coherent, but whether the magnitudes are calibrated to real platforms is not testable here. Worth flagging that the same authors are simultaneously the field’s most active critics of generative simulation: their working paper Do large language models solve the problems of agent-based modeling? (Larooij & Törnberg 2024) is more sceptical than this one. Reading the two in tandem is the right move.


5. What the 🟫 category adds up to

Across Park 2023 and the three Week-5 papers, the multi-agent category is now four papers wide. They share a methodological move and they differ along three axes.

                              agent richness   simulation scale     evaluative criterion
                              ──────────────   ─────────────────    ─────────────────────
  🟫 Park 2023 (Smallville)   high             25 × 2 days          believability (Prolific)
  🟫 Project Sid              high             10–1000+ × longer    qualitative emergence
  🟫 Ashery                   minimal          tens-hundreds        match to Centola/Baronchelli
  🟫 Larooij & Törnberg       moderate         500 × 10,000 steps   intervention-effect ranking

What the category reliably does: produce population-level phenomena — information diffusion, role differentiation, conventions, biased macro outcomes, attention inequality — that single-agent or individual-prediction setups cannot. What the category does not yet reliably do is validate those phenomena against an external ground truth. Ashery comes closest by leaning on the well-documented human naming-game literature; Larooij leans on the platform-experiment literature; Park 2023 and Project Sid rely on plausibility judgements. None of the four papers is in the position Park 2024 is in for individual prediction, where t+2 self-consistency provides a direct, principled benchmark.

The category’s other recurrent question is the echo-of-the-literature concern. All four papers reproduce phenomena their training corpus describes in detail. Cross-model robustness checks are a partial defence, since three models with different training mixtures should not converge on the same artefact for the same textual reason. They are not a complete defence, because the three models were trained on substantially overlapping public discourse. The cleanest version of the test would be to design a population-level phenomenon the literature has not yet discussed — one whose verbal description does not yet exist in the corpus — and to ask whether the simulation finds it. No paper this week takes that step.

A reader carrying the Week-3 vocabulary forward will notice that the missed-mixture question (raised in Topics/03.md) returns in a different form here. Ashery reports a bimodal distribution of run outcomes — different seeds converge to different conventions — but does not analyse whether the bimodality reflects a within-model mixture of behavioural types or a sensitivity to early-trajectory accidents. Larooij’s 5-runs-per-condition design is built on the same averaging assumption that the human behavioural-GT literature moved past in the late 1990s. The call from Week 3 — treat each trajectory as a unit of behaviour and ask whether the trajectories cluster — is a project-shaped question for any of these datasets.