| Methodological move | This week (W4) | Next week (W5) | |
|---|---|---|---|
| 🟫 | Multi-agent simulation — populate a sandbox or game with LLM agents and study what the population does | Park 2023 |
Altera Team 2024 Ashery 2025 Larooij & Törnberg 2025 |
| 🟦 | Simulating individuals — condition an LLM on personal data and predict that person | Park 2024 |
— |
| 🟧 | LLM-as-expert — forecast aggregate effects without simulating any individual | Hewitt 2024 |
— |
| 🟩 | Methodological framework / position paper | Anthis 2025 |
— |
🟦 and 🟩 carry forward from Week 3 — Park 2024 is the strongest current case for Horton’s Homo silicus proposal, and Anthis is the simulation-literature counterpart to Sun’s game-theory taxonomy. 🟫 and 🟧 are new categories: populating a sandbox with LLM agents and forecasting experiments are moves the Week-3 papers did not make. The 🟫 category is broad — it spans rich-persona spatial sandboxes (Smallville, Project Sid), abstract pairwise games (Ashery’s naming game), and synthetic platforms (Larooij’s social-media simulation) — and the differences within it are part of the cohort’s reading work over the next two weeks.
🐍 Notebook: 4_smallville.ipynb
Park 2023 is the most ambitious paper of the four in what it tries to construct rather than measure. The other three ask whether an LLM can stand in for a person whose behaviour is already on record; Smallville asks the inverse — given only a one-paragraph biography and a sandbox, can an LLM produce new behaviour that looks like the behaviour of a person? The venue is UIST, an HCI conference, and the criterion is believability rather than predictive accuracy.
The technical contribution is a cognitive scaffold around a base LLM. Three modules sit on top of an append-only memory stream:
observation
│
▼
┌────────────────────────────┐
│ memory stream │ ◀── reflections written back
└────────────┬───────────────┘
│ retrieval = recency · importance · relevance
▼
┌────────────────────────┐ ┌────────────────────────┐
│ reflection │ │ planning │
│ (when sum-importance │ │ day → hour → 5–15 min │
│ crosses threshold) │ │ recursive, reactive │
└─────────┬──────────────┘ └────────────┬───────────┘
└──────────┬─────────────────────┘
▼
action
Importance is assigned by the LLM itself on a 1–10 scale (“brushing teeth” = 2, “asking your crush out” = 8). Reflection writes high-level insights — themselves stored as memories — which lets the system grow a tree of meta-knowledge. This is the move that makes long-horizon coherence possible despite a fixed context window.
Twenty-five agents live in Smallville for two simulated days. The headline outcome is that a Valentine’s Day party emerges from a single seed prompt: Isabella decides to throw one, the invitation propagates through conversation, and on the day five of twelve invitees actually attend. Information diffusion is measurable — at run-start only Sam knows of his mayoral candidacy and only Isabella knows of her party; after two days, 32% and 52% of agents know the respective facts.
The controlled evaluation is the more disciplined claim. Each agent is interviewed with 25 questions across self-knowledge, memory, plans, reactions, and reflections; 100 Prolific evaluators rank the full architecture against three ablations and a human crowdworker. Removing all three modules (memory, reflection, planning) drops the system below the human crowdworker, with Cohen’s d = 8.16 between the no-architecture ablation and the full system. Without that ablation, the paper would be “use ChatGPT, get a believable town.”
Three failure modes recur in the rest of the literature. Memory-retrieval gaps — agents fail to bring up information they “know” because the retrieval score was too low. Hallucinated embellishments — Yuriko describes her neighbour Adam Smith as “an economist who authored Wealth of Nations,” because the LLM filled in from world knowledge. Instruction-tuning artifacts — agents are too polite, too cooperative, too willing to incorporate every suggestion, an effect Week 3’s Kozlowski & Evans flagged as Bias / Sycophancy and that a survey methodologist would call social-desirability bias (Edwards 1957; Paulhus 1984).
A note before moving on. Believability is not validity: a town that reads as plausible to a Prolific crowd is not, by that fact, a town whose dynamics one can use to test a sociological claim. The closest analogue from social science is the agent-based-modelling tradition (Schelling 1971; Epstein & Axtell 1996); Smallville gives up that tradition’s interpretability — every agent is now a black-box LLM call — in exchange for natural-language plausibility no ABM has delivered. Whether the gain is worth the cost is a question the cohort should hold open.
A year after Smallville, the same first author returned with a paper that is methodologically very different. There is no town, no NPCs, no memory module. Instead, 1,052 US adults are recruited via stratified probability sampling, each sits for a two-hour audio interview with an AI interviewer, and the resulting transcript — averaging 6,491 words — is stuffed in its entirety into the prompt of GPT-4o, along with reflections from four “expert” personas (psychologist, behavioural economist, political scientist, demographer). The agent then predicts the participant’s responses to the General Social Survey, the Big Five inventory, five economic games, and five replicated experiments.
The headline result is that the interview-based agent predicts the participant’s GSS responses 85% as accurately as the participant predicts their own responses two weeks later. Big Five reaches 80%; economic games 66%.
The 85% figure is normalised accuracy — agent’s accuracy divided by participant’s own self-consistency at t+2. Self-consistency on the GSS is r ≈ 0.83, on the Big Five r ≈ 0.95, on economic games r ≈ 1.0. So a normalised 0.85 on the GSS is a raw correlation of about 0.66 against a human ceiling of 0.83 — close. A normalised 0.66 on the games is a raw 0.66 against a ceiling of 1.0 — much further. The construct ranking is right; the headline number flatters the easier benchmarks.
Set the two papers side by side. Park 2023 builds an elaborate cognitive architecture — memory stream, reflection, planning — to handle long-horizon coherence within a 4k-token context window. Park 2024, eighteen months later, drops the architecture and uses a 128k-token context plus a richer prompt. The retreat is enabled by the model side of the field rather than by the architecture side. Was the 2023 architecture wrong, or a stopgap whose load-bearing function is now performed by native long-context handling? In 2026 we do not know; the bet Park 2024 makes is that bigger context plus a richer prompt eats most of what the architecture was doing. The bet has paid off for individual prediction; whether it pays off for multi-agent emergence over weeks of simulated time — the original problem the architecture was for — is not tested in the 2024 paper.
Three ablations are unusually informative. Random lesion (drop 80% of each interview): 0.85 → 0.79; short interviews work almost as well as long ones. Summary agents (replace transcript with a GPT-4o-generated bullet-point dictionary of facts): 0.83; predictive content is in the facts, not the linguistic style. Maximal agents (add the participant’s own GSS and Big Five answers to the prompt): 0.85, no improvement; the surveys add no information beyond the interview. The pattern: predictive value lies in informational content rather than qualitative texture.
The expert-reflection module quietly reintroduces theory through the back door — predictions for economic-game behaviour are conditioned on a behavioural-economist persona’s reflections about the participant. So the agent is not only using participant data; it is also using GPT-4o’s pretrained behavioural-economics priors. This is the WEIRD-training concern from Week 3, surfacing in a tool whose output looks data-driven. The construct-independence claim is partial: the American Voices Project script was not designed for the GSS, but it covers life story, neighbourhood, race, religion, family — many topics that do overlap with GSS items. And the entire paper runs on gpt-4o only; replication is hostage to a single closed API.
Park 2024 is the strongest current empirical case for Horton’s Homo silicus proposal, and a direct descendant of Argyle et al. 2023’s silicon-sampling paradigm. The methodological move is the same — make the LLM stand in for an individual by anchoring it to that individual’s words — and the payoff scales with the richness of the conditioning signal.
Hewitt and Ashokkumar and colleagues take the simulation problem in a different direction. Rather than build an agent and ask what it would do, they ask GPT-4 to forecast the result of a social-science experiment: given the stimuli, the outcome measure, and a representative demographic profile, what treatment effect should the experimenter expect? The unit of analysis is the aggregate effect size; the unit of validation is the published treatment effect from a pre-registered experiment.
Anthis calls this LLM-as-expert to distinguish it from the LLM-as-subject paradigm of Park 2024 and Horton. The distinction matters: instruction-tuning makes models worse at roleplay (sycophancy, refusals, social desirability) and better at third-person forecasting, which is exactly what helpful-assistant training rewards.
The validation set is unusually disciplined: 70 experiments — 50 from the NSF Time-Sharing Experiments in the Social Sciences (TESS) corpus and 20 from a recent generalisability replication project — yield 476 treatment effects across 105,165 participants. Every experiment is pre-registered, run on a nationally representative US probability sample, peer-reviewed, with open materials. The experiments were designed by 77 social and behavioural scientists across psychology, political science, sociology, public health, communication, and public policy. This is a substantial sample of what social science actually does.
The contamination control is the most consequential design choice. 273 of the 476 effects come from experiments that were not published or publicly posted before GPT-4’s training cutoff in September 2021. This is the test that distinguishes “the model has seen this” from “the model can predict experiments it has not seen.”
Across all 476 effects, GPT-4 predicts at r = 0.85 (raw), r = 0.91 (adjusted for measurement error). On the unpublished subset, r = 0.90 raw, r = 0.94 adjusted — higher than published. Lay forecasters from Prolific (N = 2,659) reach r = 0.79; a regression that combines the two performs better than either alone, with similar coefficients on each predictor — they carry independent information.
Two caveats are doing real work. GPT-4 systematically overestimates effect sizes by about 80%; researchers using the tool unscaled will inflate their expectations. And the strong correlations are for survey experiments — on nine field-experiment megastudies covering 346 effects across 1.8 million participants (DellaVigna & Pope, Milkman et al.’s vaccination texts, Voelkel et al.’s democratic-attitudes), GPT-4 reaches only r = 0.27 raw, 0.33 adjusted. Field experiments are where social science actually pays off, and field experiments are where the method performs worst.
Within-subgroup main effects are well predicted (r ≈ 0.85–0.90 for women, men, Black, white, Republicans, Democrats). Interaction effects are predicted weakly: r = 0.17 for gender, 0.55 for ethnicity, 0.41 for party. The model captures the main effect of treatment but not its heterogeneity across groups — and the most useful applications of forecasting (which subgroups respond best?) are exactly the ones it is worst at.
The paper’s last move is unusual in social-science papers and worth naming. Given Allen et al. 2024’s corpus of 271 anti-vaccination Facebook posts (each with a measured effect on vaccine intention), GPT-4 and Claude 3 Opus were asked to select the most-harmful posts. The top five selected by GPT-4 had an estimated effect of −2.8 percentage points; the most-harmful single post had a measured effect of −4.1 points. First-order guardrails (refuse to generate harmful content) are insufficient because the model can be used as a selector among existing harmful messages without ever generating any. The authors propose “second-order guardrails” — restricting the prediction capability for socially harmful targets — as a research direction. No working implementation exists yet.
Hewitt et al. sits in the long tradition of forecasting tournaments (Tetlock & Mellers’s Good Judgment Project) applied for the first time to social-science experiments. The legitimating context is the replication crisis (Open Science Collaboration 2015): forecasting is interesting only because not all published experiments replicate. If they did, prediction would be retrieval.
Anthis 2025 is a position paper whose author list overlaps substantially with the other three Week-4 papers. Bernstein is on Park 2023, Park 2024, and this one. Kozlowski and Evans co-authored Week 3’s Simulating Subjects and are also on this paper. Read it as the insider position statement for the simulation literature — a constructive critic from inside the network whose papers it discusses.
The framework is five challenges and a corresponding spectrum of applications.
challenge definition promising direction
──────────── ───────────────────────────────────────── ──────────────────────────────
diversity outputs lack human variation inject variation in training/
tuning/inference (interview
prompts, steering)
bias inaccuracies for particular groups implicit demographics; minimise
*accuracy-decreasing* biases,
not all bias
sycophancy excessive user-pleasing LLM-as-expert (third-party)
instead of LLM-as-subject
alienness superficial accuracy via non-humanlike simulate latent features;
mechanisms reassess as mech-interp advances
generalization failure in OOD contexts simulate latent features;
iterative evaluation
A reader of the Week-3 chapter will notice the relabelling. Kozlowski & Evans listed six weaknesses (Bias, Uniformity, Atemporality, Linguistic Cultures, Disembodiment, Alien Intelligence); Anthis lists five. Diversity ≈ Uniformity; Alienness ≈ Alien Intelligence; Bias ≈ Bias; Sycophancy and Generalization are new categories; Atemporality, Linguistic Cultures, and Disembodiment have been dropped. The Week-3 framework was the cautious-diagnosis version; the Week-4 framework is the optimistic-engineering version, by overlapping authors a year later.
The Sycophancy direction — prefer LLM-as-expert to LLM-as-subject — is the conceptual move that makes Hewitt the paper Anthis cites most enthusiastically. As models become more instruction-tuned, the expert framing should pull ahead of the roleplay framing. That is a methodological prediction, and it is testable.
Anthis’s Figure 1 organises applications along a spectrum from immediately feasible to long-term aspiration: pilot studies → exploratory studies → exact replication → sensitivity analysis → complete “human-possible” studies → complete “human-impossible” studies. The spectrum is useful because it disciplines the conversation. The most productive question for a reader is to ask, of any given simulation paper, where on this spectrum its claims actually live.
A note on what the position paper does not do. The “alternative views” section is one paragraph long. It names Bender et al.’s stochastic-parrots critique, Chomsky’s “ineradicable defects,” and LeCun’s scepticism — and dismisses each in a sentence. A more rigorous version would engage Atari et al. 2023 on cross-cultural failure, Bisbee et al. 2024 on bias in silicon sampling, and Petrov et al. on positional artifacts. The compactness is conventional for a position paper, but treat the framework as the insiders’ working consensus, not as the whole field’s view.