Marco Combetto

AI & Digital Transformation — Public Sector — Data Science

Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment

Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment

Researchers tested whether Large Language Models (LLMs) can accurately simulate how different groups of people feel about urban planning.

The study focused on an affordable housing survey where they compared eight different LLMs against 843 human respondents. The researchers wanted to see if the AI could replicate how opinions change based on specific factors. These factors included proximity to a new building, housing status (owners versus renters), and political identity. Some models, like Qwen 2.5 14B, mirrored the average results of the human survey quite well. However, they often failed to capture specific patterns within those groups. For instance, the models struggled to show how support shifted for renters compared to owners. They failed to show how these opinions changed as a development moved from two miles to a few hundred feet away.

This finding is important for public sector leaders who consider using AI to simulate public reaction to new policies or urban projects. If a model only predicts the “average” sentiment, it can hide the unique needs and concerns of specific neighborhoods or demographic groups. Using these models as a shortcut for urban planning risks overlooking the diversity of actual citizens. For the public sector, it is not enough for a model to give a general summary. Policy decisions require understanding the nuances of social and spatial structures that create public opinion.

How can we ensure AI tools used in public policy preserve the diversity of the people they are meant to represent?

#PublicSector #ArtificialIntelligence #UrbanPlanning #AIGovernance #PublicPolicy

https://arxiv.org/abs/2607.27100v1

🔗 Read the original article