Quick filter
Inaugural Large Language Models and the Social Sciences Conference
City University of Hong Kong (CityU), 14 - 16 October 2026
Generative Causal Inference a Revolution in Social Science Experiment Design Driven by LLM Taking Nonlinear Effects of Political Communication and Public Policy Compliance Incentives as Examples
Yang Guo1, Tao Luo2, Yu Zhou3, Yifan Wang4
1 Harbin Engineering University
2 Lanzhou University
3 Zhongnan University of Economics and Law
4 Hong Kong University of Science and Technology
This study addresses the core challenges of "non-experimental interventions" and "unobservable variables" in causal inference within social sciences by proposing the Generative Causal Inference Framework (GCIF). By innovatively integrating the generative capabilities of Large Language Models (LLMs) with Structural Causal Models (SCMs), we construct a three-stage research paradigm: virtual experiment generation, textual response simulation, and causal effect estimation. In the domain of political communication, GCIF successfully dissects the three-dimensional interactive effects of "emotional framing × source credibility" on public policy support, revealing that traditional linear models underestimate the counter-reinforcing role of negative emotional framing under official information sources. In the public policy domain, through simulating the impact of different information disclosure strategies on taxpayer compliance intentions, we identify a threshold effect of "information overload-trust decay"—compliance intention declines by 18% when information volume exceeds taxpayers' processing capacity.
The accompanying open-source toolkit, CausalLLM-Suite v2.0, achieves three breakthroughs: (1) the experimental scenario generator supports dynamic adjustment of intervention variable distributions and confounding factor configurations in virtual experiments; (2) the textual embedding optimization module employs contrastive learning strategies to enhance the discriminative power of textual representations in causal inference tasks; (3) the causal effect validation module integrates Difference-in-Differences (DID) and SHapley Additive exPlanations (SHAP) for robust statistical inference and interpretability. Empirical validation on a dataset comprising 200,000 social media dialogues and 12,000 survey responses demonstrates that GCIF improves causal identification accuracy by 52% over baseline models, with explanatory metrics achieving 92% consensus among human experts.
This research not only redefines the value proposition of LLMs in social sciences—from predictive tools to causal inference engines—but also advances the paradigm shift of social science research from "correlational description" to "causal explanation" through tool innovation, thereby offering significant academic frontier value and practical guidance.
“The flowers shed tears as I grieve over the times”: An Analysis of the Emotional Output in Du Fu's Poems Based on Large Language Models and Causal Inference
Chengke Zhu, Yarong Hao
Huaibei Normal University
Abstract: This study combines large language models (LLMs) with the difference-in-differences (DID) causal identification method, taking the An Lushan Rebellion as a quasi-natural experiment to conduct a quantitative emotional analysis of Du Fu's 1,496 extant poems. It systematically reveals the causal effect of historical shocks on the emotional expression in literary creation and its transmission mechanism. A stratified random sampling approach was adopted to select 200 poems for manual annotation verification, thereby establishing a reliable measurement quality evaluation system. The findings are as follows: (1) The An Lushan Rebellion exerted a significantly negative impact on the emotional scores of Du Fu's war-themed poems, and this result remained consistent across manual annotation verification, replacement of large language models, and multiple robustness tests; (2) Poetry creation exhibited a significant emotional regulation function — poems themed around personal experiences and social concerns could effectively mitigate the negative impacts of war shocks; (3) Heterogeneity analysis indicated that the negative effects of war shocks were mainly concentrated in guti shi (ancient-style poetry) and pailü (regulated verses with extended lines). Additionally, poems rhyming with rusheng (entering tones, a category of ancient Chinese tones) were most significantly affected, which reveals the differentiated characteristics of poetic genres and phonological forms as carriers of emotional expression.
Beyond Fixed Functions: A Strategic Framework for LLM-Assisted Qualitative Analysis in Policy Evaluation Research
Igor Lyubashenko
SWPS University
Introduction and Motivation

The integration of large language models into qualitative research faces a fundamental conceptual challenge: unlike traditional software with defined feature sets, LLMs require researchers to develop flexible strategies for human-AI dialogue rather than follow prescribed procedures. This paper proposes a model-agnostic framework for incorporating LLMs into qualitative data analysis, with particular application to public policy evaluation research. The framework addresses a critical gap in current methodological literature, which consists primarily of isolated experimental reports rather than systematic guidance for practitioners.

Theoretical Framework

The paper builds on Goertz and Mahoney's (2006) characterization of qualitative data as "rich" and "thick"—containing multiple layers of meaning requiring interpretive engagement. Following Saldaña's (2016) conceptualization of coding as active interpretation rather than mere labeling, I argue that LLMs can support but not replace the researcher's analytical judgment.

The framework leverages two core LLM capabilities suited to qualitative analysis:
  1. Semantic search extending beyond keyword matching to locate passages based on meaning
  2. Text categorization enabling classification according to conceptual rather than lexical criteria
These capabilities map onto two fundamental analytical strategies, each requiring distinct approaches to LLM integration.

Two Complementary Strategies for LLM-Assisted Analysis

Strategy 1: Inductive Exploration ("Document dialogue")

When researchers approach data without predetermined categories—characteristic of exploratory research—the framework operationalizes analysis through "document dialogue." This interactive process proceeds through phases:
  • Exploratory questioning using open-ended prompts to surface themes and patterns not immediately apparent in initial reading
  • Targeted investigation with focused queries addressing emergent dimensions
  • Code development, synthesizing insights into formal coding schemes with explicit definitions
This approach treats the LLM as an analytical partner—analogous to an attentive reader with whom the researcher conducts an in-depth interview about the material.

Strategy 2: Deductive Classification ("Topic search")

When researchers approach data with established categories—common in theory-driven studies—LLMs serve a different function: systematically locating content matching predefined criteria. The researcher knows what they are looking for; the task is reliable identification across potentially large document collections.

This strategy leverages semantic understanding to identify relevant passages even when vocabulary differs from code definitions. A code for "implementation barriers" can capture "obstacles," "challenges," or context-specific formulations—extending beyond keyword search capabilities.

The deductive approach requires:
  • Precise code definitions with clear inclusion/exclusion criteria
  • Explicit instructions constraining responses to document content
  • Verification protocols comparing LLM classifications against researcher judgment

Key Findings and Contributions
  • Inductive exploration suits early-stage analysis and code development; deductive classification suits systematic coding of larger corpora with established categories.
  • When LLMs work with provided documents and explicit instructions limiting responses to document content, factual accuracy improves substantially.
  • LLM-based deductive coding captures conceptually relevant content that keyword searches would miss, while maintaining consistency human coders may struggle to achieve across large datasets.
Both strategies require verification. LLMs may miss contextual factors, misinterpret ambiguity, and tend toward "completing" information from training patterns. Counterintuitively, LLM coding may provide "objective" validation for inherently subjective qualitative interpretation—pattern-based judgments offer a distinct perspective against which researcher interpretations can be compared.
HalalLLM vs. KosherLLM: When Abrahamic Scholars Confront Modern Questions Through AI
Cantay Caliskan
Goergen Institute for Data Science and AI University of Rochester
What if you could ask a 12th-century philosopher about modern society? This study explores that possibility by resurrecting twelve of the most influential scholars in Judaism, Christianity, and Islam through the architecture of large language models (LLMs). We instructed three state-of-the-art models-GPT-4 Turbo, Qwen-Plus, and DeepSeek-R1-to role-play these historical figures, grounding their perspectives in their authentic writings using a retrieval-augmented generation (RAG) pipeline that retrieves a median of eight citations per prompt. These AI-driven scholars then answered 290 questions from the World Values Survey, generating over 50,000 synthetic responses that we benchmarked against contemporary global attitudes. Our empirical analysis reveals that while the models appear similar on the surface, their underlying interpretive styles differ significantly. A repeated-measures ANOVA confirmed that although the models' mean outputs were statistically indistinguishable, their median (F (2, 22) = 13.87, p < .001) and modal (F (2, 22) = 9.71, p < .001) responses were not, confirming that a scholar's ideas are represented differently across models (Hypothesis 4). Post-hoc tests specified these differences, showing Qwen's median output was significantly higher than both DSR1 and GPT-4. Furthermore, our findings indicate that models sometimes project a universalized worldview, with Islamic scholars' profiles aligning closer to the global average (t =-5.77, p < .0001), while at other times capturing regional particularity, with Christian scholars' profiles aligning more with their local norms (t = 17.86, p < .0001). By prototyping a method to engage with classical religious tradition, this work opens new avenues for digital humanities and highlights the critical ethical and methodological challenges in creating AI systems that can faithfully preserve and transmit centuries of human thought.
Analyzing Public Risk and Action Awareness Under Extreme Disaster Early Warnings Using Large Language Models
Xiaoyu Li, Xue Lin
School of Government, Nanjing University
The integration of early warning dissemination and public response is central to effective risk governance under extreme disasters. However, empirical evidence consistently shows that the issuance of warnings does not necessarily translate into timely public action. This gap reflects the lack of a valid, scalable, real-time measurement method for public risk and action awareness following the release of disaster warnings, leaving the warning institutes without timely insight into how the public understands risk or prepares to act. Addressing this challenge, this study focuses on developing a large-language-model-based (LLM-based) measurement method to assess and analyze real-time public response based on social media data. Specifically, the study developed a human-in-the-loop framework for integrating the theoretical construction of risk perception and protective action intentions into the LLM-based processing pipeline. First, we adopted different fine-tuned strategies including few-shots, chain-of-thoughts, and instruction prompt-based fine-tuned and parameter-based fine-tuned models to increase measurement stability. Second, a workflow was designed to address measurement uncertainty by generating multiple independent assessments through parallel model inference, and by systematically integrating outputs through consistency checks, result aggregation, and targeted validation. Third, we increase the content validity of the measurement by integrating human-validation in the framework. This process transforms a black-box LLM judgments into valid measurement for real-time and large-scale qualitative analysis. Experimental data were drawn from a case study in China, which include Sina Weibo posts during a super Typhoon Doksuri in 2023, to systematically evaluate different model configurations and analytical workflows. The findings identify a robust and valid measurement method for capturing public response upon the release of a disaster early warning. By doing so, the study provides an operational measurement tool for understanding the disconnection between warning dissemination and public response, and offers a transferable methodological reference for disaster governance research.
Can Generative AI Reliably Code Treaty Texts? Evidence from Nuclear Non-Proliferation Treaties
Saera Lee1, Bo Won Kim2
1 University of Hong Kong
2 University of Texas, Arlington
Can generative AI be used to code international treaties, and how reliable are the results? As generative AI becomes increasingly integrated into both everyday live and academic research, scholars have begun exploring ways to incorporate these tools into their research workflows. This paper develops and evaluates a method for using generative AI to code treaty texts, using non-proliferation and arms control related treaties. We examine how prompting strategies influence coding outcomes and propose approaches that improve performance. We then assess the validity of the generated data by comparing outputs across multiple models and prompts. In addition, we compare coding results obtained through web-based interfaces and API access, providing practical guidance for researchers who are unfamiliar with API-based tools.

The findings suggest that generative AI can offer a cost-effective and time-efficient approach to treaty text data collection. However, expert review remains necessary to ensure the validity and reliability of the final dataset. The study contributes to methodological debates on the use of AI-assisted research tools in political science.

Can generative AI be used to code international treaties, and how reliable are the results? As generative AI becomes increasingly integrated into both everyday live and academic research, scholars have begun exploring ways to incorporate these tools into their research workflows. This paper develops and evaluates a method for using generative AI to code treaty texts, using non-proliferation and arms control related treaties. We examine how prompting strategies influence coding outcomes and propose approaches that improve performance. We then assess the validity of the generated data by comparing outputs across multiple models and prompts. In addition, we compare coding results obtained through web-based interfaces and API access, providing practical guidance for researchers who are unfamiliar with API-based tools.

The findings suggest that generative AI can offer a cost-effective and time-efficient approach to treaty text data collection. However, expert review remains necessary to ensure the validity and reliability of the final dataset. The study contributes to methodological debates on the use of AI-assisted research tools in political science.
The Disappearing Conflict: Tracking Political Attention to the Israeli-Palestinian Conflict Across Five Million Speeches
Sascha Riaz1, Chagai Weiss2
1 Singapore Management University (from 07/26)
2 University of Toronto
Political attention is a precondition for conflict resolution. Domestically, sustained engagement from elites and publics generates the political will to pursue costly peace processes. Internationally, attention from foreign governments sustains diplomatic pressure, mediating capacity, and reputational costs for inaction. We document the erosion of both dimensions in the decades preceding October 7, 2023. Analyzing parliamentary speeches in Israel's Knesset, alongside the US Congress, the European Parliament (total N = 5 million speeches), Israeli newspaper op-eds, and public opinion data from 15 election surveys, we show that political attention to the Israeli-Palestinian conflict declined continuously from the 1990s through 2022 – both within Israel and across NATO countries. To label this volume of multilingual text data, we deploy large language models to identify conflict-related content and to distinguish between security-framed and diplomacy-framed discussions, enabling measurement at a scale and granularity infeasible with conventional coding approaches. Critically, the decline in discussions of peace and conflict resolution was steeper still: even when elites discussed the conflict, they increasingly did so in security terms rather than diplomatic ones. In Israeli public opinion, the share of citizens identifying the conflict as the most important political issue fell from 80% in 1996 to 21% in 2022. Parallel declines in NATO member states suggest a global withdrawal of political attention, not merely an Israeli domestic phenomenon. While October 7 sharply restored attention to the conflict across all data sources, discussions of conflict resolution remain at historic lows. Our findings carry a broader lesson for the study of intractable conflicts: even as objective conditions persist or worsen, political salience can decline to the point where the domestic and international foundations for peace quietly erode.
Governing the Limits of Responsiveness: Public Justifications for Nonresponse in China’s 12345 System
wenna li
1.School of Public Administration and Policy, Renmin University of China, Beijing, China
Digital complaint systems are often portrayed as inclusion-enhancing institutions because they lower the costs of voice and widen citizens’ access to government. Yet the same systems also create a less examined governance problem: they require public organizations to determine which claims are actionable, which should be redirected, and which will not be processed within the complaint channel. This article examines that problem through a computational analysis of officially disclosed reasons for nonresponse in parts of China’s 12345 system, focusing on public explanations of nonacceptance, redirection, or termination issued by local governments. Methodologically, it develops a transformer-embedding-based text analysis strategy for administrative text and combines embedding-based topic modeling with interpretive coding to identify recurrent justificatory repertoires in short, fragmented, and highly standardized official explanations. Rather than treating these texts as residual metadata, the analysis shows that they constitute a meaningful site of administrative practice. Substantively, the article reconceptualizes nonresponse not as a simple failure of responsiveness, but as a public-facing form of boundary work through which governments organize the limits of response. The findings show that these limits are publicly governed through four recurrent strategies: jurisdictional delimitation, responsibility allocation, threshold specification, and normative exclusion. In doing so, the article links computational text analysis to core concerns in public administration and shows how digitalized states make exclusion visible, intelligible, and defensible within systems ostensibly designed to expand access.
Policing or Repressing? Multipronged Expansion of Digital Surveillance in China
Dakeng CHEN, Jing Vivian ZHAN
CUHK
Authoritarian regimes are expanding digital surveillance at unprecedented scales, often justified as a response to public security needs. Yet existing studies typically frame surveillance as political repression. This raises a critical question: what drives the uneven buildout of digital surveillance, especially within-country variations? We address this question by analyzing Chinese city governments’ surveillance deployment between 2013-2019.

Using large language models with schema-based structured output for consistent annotation, we construct an original dataset from five million government procurement records and one million court judgments, classifying surveillance technologies by function and crimes by type. Statistical analysis uncovers three striking patterns. First, ordinary property crimes, accounting for over 80% of recorded offenses, drive baseline surveillance infrastructure like CCTV. Second, violent crimes predict investment in targeted monitoring capabilities, including smartphone monitoring and communication interception. Third, politically-charged “pocket crimes”, vague offenses like “picking quarrels” (xunxin zishi) and “obstructing official duties” (fanghai gongwu) used against dissidents, predict the adoption of the most intrusive surveillance technologies: biometric identification, location tracking, grid management systems, and intelligence analytics. The findings reveal a multipronged strategy: authoritarian regimes exploit legitimate crime concerns to build broad surveillance infrastructures while reserving the most intrusive capabilities for political control. This research advances theories of authoritarian governance, state capacity, and technology adoption while offering a methodological blueprint for studying surveillance expansion beyond China.
Silicon Partisans Simulating Partisan Framing Effects on Sino-US Attitudes via LLM-Based Silicon Sampling
PeiYun Shi2, YanYou Chen1
1 Lee Kuan Yew School of Public Policy, National University of Singapore, Singapore
2 Department of International Politics, University of International Relations, Beijing, China
3 Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing, China
This study examines whether LLM-based silicon samples can reproduce the interaction between news framing and partisan identity in shaping American attitudes toward China. While existing research has shown that LLMs can replicate known survey distributions, few studies have used this approach to test causal mechanisms. We implement a 3-by-3 factorial experiment in which three partisan agent personas, constructed from General Social Survey 2024 microdata, are exposed to China-related news vignettes adapted from recent AP and Reuters reporting and employing three distinct frames (security threat, economic competition, and cooperation). A dual-layer design ensures both ecological validity through realistic partisan archetypes and internal validity through a robustness check using demographically identical agents that differ only in partisanship. We hypothesize that baseline partisan gaps will mirror real survey data, that security threat framing will produce the strongest negative shift, that framing effects will be moderated by partisanship, and that LLM-generated distributions will exhibit systematic variance compression. Baseline scores are validated against Pew Research Centre and Chicago Council on Global Affairs data. This work advances silicon sampling from replication toward causal experimental design, generating testable hypotheses about how media framing differentially shapes attitudes across political groups. The findings carry implications for cross-cultural communication, media literacy, and low-cost opinion pre-testing in geopolitically sensitive environments. The study speaks directly to the LLMSS 2026 conference themes of LLMs in experimentation, prompt engineering, and domain-specific applications in political communication.
Synthetic Operator Teams: A Multi-Agent Generative Framework for Simulating Social Dynamics and Communication Failures in Nuclear Control Room Operations
Xingyu Xiao1, 2, Jingang Liang1, Jiejuan Tong1, 2, Haitao Wang1
1 Tsinghua University
2 National Key Laboratory of Human Factors Engineering

Human reliability in safety-critical systems is not solely determined by individual cognition but emerges from complex social interactions within operational teams. However, existing human reliability analysis (HRA) frameworks predominantly model operators as isolated decision makers, neglecting the dynamic social mechanisms that shape collective performance. This study proposes a generative multi-agent framework for simulating synthetic operator teams in digital nuclear control rooms. By integrating large language model–driven role cognition, cognitive architecture-based behavioral constraints, and environment-grounded task execution models, the framework enables emergent team-level phenomena such as authority bias, communication breakdown, workload redistribution, and coordinated error propagation to be quantitatively analyzed.

In a high-fidelity virtual control room environment, heterogeneous operator agents with distinct personality traits, experience levels, and fatigue states collaboratively execute abnormal operating procedures under time pressure. Interaction logs, communication semantics, and action trajectories are jointly modeled to infer dynamic team situation awareness and collective reliability indices. Simulation experiments demonstrate that social interaction patterns can amplify procedural risk beyond individual cognitive limitations, revealing nonlinear transitions from resilient collaboration to cascading failure modes.

The proposed approach establishes a computational paradigm for studying social mechanisms of human reliability and offers a scalable platform for risk-informed interface design, team training optimization, and digital twin–enabled safety governance. This work advances the understanding of how collective cognition and social structure shape operational safety in complex technological systems.
Retail Investor Forum Topic Attention and Stock Market Dynamics
Zihao Qu1, Xu Zhang1, Peiran Jiao2
1 Hong Kong University of Science and Technology (Guangzhou)
2 Maastricht University
We study retail investors' topic-specific attention and its stock-return predictability. From 340 million posts on China's largest stock forum, we build firm-day attention measures for 200 interpretable topics. Using LASSO, we identify the most predictive topic: a one-standard-deviation rise in its attention predicts 0.44–0.72 percentage points lower next-month returns across different keyword-intensity thresholds, with partial reversal within one to two months. Predictions are stronger among smaller, non-main-board firms, and post-2015. The results remain robust to orthogonalization against recent returns, aggregate post sentiment controls, and alternative information channels.
Multimodal Fusion and Counterfactual Reasoning Analysis on Multinational Firms’ Climate Information Disclosure: A Human-AI Collaboration Framework
Maogang Sun
Jiangxi University of Finance and Economics
Global climate disclosure information is inherently multimodal, cross-jurisdictional, and highly heterogeneous. Traditional manual and single-modality verification methods suffer from low efficiency and limited coverage, and they often fail to detect deep-seated greenwashing. These limitations constrain effective oversight of multinational corporations' (MNCs) climate disclosures and impede global climate governance. We address the challenges by constructing a framework that integrates multimodal fusion, task allocation, causal inference, explanation generation, and cross-domain adaptation. Combining multimodal representation learning with structural causal models, we develop a human-AI collaborative causal reasoning mechanism tailored for climate disclosure verification. We further design an adaptive system architecture that accounts for country or regional heterogeneity to support more efficient verification in international business setting. By establishing a human-AI collaborative theoretical framework spanning data perception to decision-making, this research strengthens the use of large language models and multimodal techniques in interdisciplinary studies across environmental science, climate governance, and management science. We also offer a novel verification approach with significant practical implications for improving global climate disclosure quality and enhancing climate risk governance.
AI-Generated vs. Human-Written: An Experiment Study on LLM-Assisted ESG Crisis Communication Strategies and Public Trust
Maogang Sun
Jiangxi University of Finance and Economics
With the proliferation of generative artificial intelligence (GenAI) and large language models (LLMs) in corporate ESG crisis communication, firms face a paradox between efficiency and sincerity in digital era. It is not clear whether and how using AI-generated statements affect stakeholders' perceptions and trust when facing negative ESG news. Integrating situational crisis communication theory and source credibility theory, this study employs between-subjects online experiment (N=309) to examine the interactive effects of statement authorship (Human vs. AI vs. Disclosed AI) and crisis type (Competence-violation crisis vs. Integrity-violation crisis) on audiences' trust. The findings reveal that when authorship is not disclosed, LLMs-generated statements perform no worse than human-written ones. However, once authorship is disclosed as AI, regardless of crisis type, the public's trust decreases significantly, with a more pronounced drop in integrity-violation crises. Mediation analysis further indicates that perceived sincerity is a key mechanism explaining the negative effect. Treating LLM as a significant communication actor, this paper highlights the importance of communication source attributes, thereby expanding ESG crisis communication research in the era of human-AI collaboration. Moreover, we identify the potential negative effects of AI disclosure transparency particularly in fields characterized by heightened attention to ethical, emotional, and humanistic concerns. The implications for companies' appropriate application of LLMs in ESG crisis communication are discussed.
Assessing the Domestic Legitimacy of Climate-Induced Migration: A Socio-Sensing and LLM-Based Analysis of the Australia-Tuvalu "Falepili Union"
Zilin Zhou
School of International Studies, Peking University
As sea-level rise poses an existential threat to low-lying Pacific island nations, the 2023 "Falepili Union" treaty between Australia and Tuvalu has emerged as a landmark "resettlement-for-security" model in global climate diplomacy. However, the sustainable implementation of such climate-induced migration policies depends heavily on the "social license" and acceptance within the host country. This research adopts an interdisciplinary approach, bridging International Relations’ "Human Security" framework with Geographic "Social Sensing" methodologies to evaluate Australian public sentiment toward this treaty.

By harvesting a longitudinal corpus from digital platforms including Reddit (r/australia, r/AustralianPolitics) and X (Twitter) since November 2023, this study utilizes Large Language Models (LLMs), specifically GPT-4o, to perform deep semantic mining and automated qualitative coding. Unlike traditional sentiment analysis, the LLM-driven pipeline identifies nuanced narrative frames, categorizing public discourse into dimensions of climate justice, national security, economic cost, and sovereign concerns.

Preliminary findings indicate a complex tension between humanitarian obligations and geopolitical skepticism. The research further explores the spatio-temporal evolution of these sentiments, mapping how digital discourse responds to specific policy milestones. This study contributes a novel computational framework for assessing the domestic legitimacy of climate migration policies, providing data-driven insights for policymakers to foster social cohesion in the face of escalating regional climate crises.
Facing the "Face": An LLM-Assisted Study on the Alignment of Political Leader Traits
zhexi LIU
China Foreign Affairs University
In the field of international relations theory, researchers frequently focus on the decision-making processes or actions of key political actors during critical junctures; however, these political actors are often inaccessible to researchers. Existing studies investigating the impact of such actors' specific traits on international political decision-making typically employ methods such as empirical analysis, experiments, and interviews. Nevertheless, database-driven empirical analyses struggle to capture the nuanced traits of actors—such as political leaders or diplomats—while subjects in experimental designs often differ significantly from the actual roles involved in real-world international politics; furthermore, the interview method—though arguably the closest to the reality of international politics—presents significant practical challenges in its implementation. The advancement of research into the model alignment and the interpretability of language models offers a promising opportunity to explore these types of questions. This raises the question: Is it possible to leverage language models to construct a class of international political figures endowed with specific traits, thereby utilizing them as experimental or analytical subjects for theoretical research? This paper attempts to address this by employing three distinct methods: zero-shot prompting, few-shot prompting, and activation vector steering based on language models. Furthermore, drawing upon a continuously updated database GDELT, we have constructed a real-time database documenting the actions of a vast number of contemporary political leaders. Using this database, we conduct data-driven verifications of the models' alignment results regarding specific types of political figures possessing defined traits, alongside statistical assessments of the results' reliability and validity. This study represents a preliminary endeavor toward conducting scientific experiments involving LLM agents representing international political figures, situated within the framework of the Fifth Paradigm of social science research.
Does DeepSeek Have a Chinese Heart? A Conceptual Replication of Assessing Political Bias in ChatGPT
Fei Shen, Weiying Shi, Yuheng Wu
City University of Hong Kong
see attached.
Measuring Political Belief from Historical Text: A Hybrid LLM-Embedding Framework with an Application to Tang Poetry
Zhaomin Li
University of Washington, Seattle
Measuring latent constructs in text requires both rubric-based concept judgment and reproducible corpus-scale scoring. This paper presents a hybrid framework that assigns each capability to the tool best suited to it: large language models (LLMs) for applying theoretical rubrics, and domain-adapted embeddings for scalable measurement. I demonstrate the approach by measuring vertical moral integration (VMI), the degree to which elite expression treats political authority as a moral commitment, in approximately 49,000 Classical Chinese poems from the Tang dynasty (618-907). The substantive question is whether the imperial examination's Confucian-centered curriculum shaped elite moral orientation toward political authority via endogenous belief formation, or merely selected for pre-existing dispositions. The pipeline proceeds in four stages: a theory-driven scoring rubric operationalizes VMI across three dimensions; zero-shot LLM classification generates candidate seed texts at the theoretical extremes; human-in-the-loop validation finalizes anchor sets; and a poetry-specialized language model (BERT-CCPoem) constructs a continuous poem-level index via embedding-based semantic scaling. An extension generalizes the bipolar index into a multi-prototype mixture model that represents each poem as a distribution over competing narratives. The hybrid design is portable to new constructs and corpora wherever a theoretically grounded rubric and a domain-appropriate embedding model are available. Substantively, the application connects computational measurement to core questions in historical political economy, such as institutional socialization, political identity formation, and state capacity.
Information Interventions and Cash Subsidies: A Field Experiment on Uptake of HPV Vaccines
Yiting Guo, Mingxuan Qi, Lijia Wei
School of Economics and Management, Wuhan University, Wuhan, China

HPV vaccination is central to cervical cancer prevention, yet uptake in China remains low, with information asymmetries and financial costs remaining major barriers. We study how subsidy design, procurement-price disclosure, and AI-assisted decision support shape HPV vaccine choice among 214 female undergraduates at Wuhan University. We randomly assign participants to six treatments that vary three subsidy schemes and two price transparency conditions. The subsidy schemes are baseline (a symbolic RMB 20 subsidy), targeted (a targeted RMB 247 per-dose subsidy applied only to the domestic bivalent vaccine), and universal (an RMB 247 per-dose subsidy applied to all vaccine types). The price transparency conditions vary whether participants are informed that the government procurement price of the domestic bivalent vaccine is below RMB 30. We also use a within-subject design in which participants interact with an AI health platform and then make their vaccination decisions before and after that interaction. We find that subsidies raise registration from about 70% to about 90%. AI, by contrast, leaves overall registration largely unchanged but shifts demand toward 9-valent vaccines, particularly the domestic 9-valent vaccine. Procurement-price disclosure reduces registration only in the baseline subsidy condition. Under targeted subsidy with procurement-price disclosure, the domestic bivalent share declines after AI exposure. These results offer practical guidance for policymakers and healthcare administrators designing vaccine promotion strategies under AI-assisted decision support.

LLM-Driven Counterfactual Generation and Quality Pruning for Big Five Personality Prediction
Zhengyang Han
School of Asian Studies, Beijing Foreign Studies University
Predicting Big Five personality traits from unstructured text remains a significant challenge, often hindered by sparse labels and confounding variables. Existing collections, such as the Pennebaker essay dataset, frequently entangle psychological traits with specific semantic topics, leading to spurious correlations during model training. To address this, we propose a Psycholinguistically Informed Data Augmentation Framework utilizing Large Language Models to generate synthetic text. Central to this framework is Counterfactual Personality Transfer (CPT), which employs Qwen3.5-flash to perform controlled rewriting of original essays, altering the manifested personality traits while strictly preserving the original topical context. To ensure the psychological validity and semantic integrity of the augmented dataset, we introduce a rigorous multi-stage quality pruning pipeline. This dual-constraint filtering mechanism evaluates linguistic fluency, topic consistency via semantic embeddings, and bidirectional trait transfer verified via the Empath lexicon. Generated samples failing these thresholds are discarded and regenerated within a predefined limit, while global dataset distribution is monitored via Kullback-Leibler divergence to minimize domain shift. Experimental results demonstrate that a Qwen3.5-9B predictive model fine-tuned on our quality-pruned, counterfactually augmented dataset achieves consistent performance gains over baseline models, exhibiting enhanced robustness against spurious topic-trait correlations. This approach effectively decouples semantic content from stylistic markers, providing a robust methodology for LLM-driven data augmentation in the social sciences.
Navigating Complexity in Disaster Response Networks: Integrating LLM and Social Network Analysis to Study the Wang Fook Court Fire in Hong Kong
Xinyi Wang1, Wenjin Chen1, Zheng Yang2
1 The Hong Kong University of Science and Technology
2 California State University, Dominguez Hills
The governance of disaster response networks remains a critical yet empirically underexplored area in public administration and emergency management. Complex disasters, particularly in dense urban environments like Hong Kong, demand coordination structures capable of integrating diverse actors across government departments, non-governmental organizations (NGOs), and community stakeholders. The fire that occurred in November 2025, at Wang Fook Court in Hong Kong’s New Territories serves as a quintessential case: a high-rise residential fire that exposed coordination challenges among the Fire Services Department, the Home Affairs Department, the Housing Department, social welfare NGOs, and grassroots mutual aid groups. This study integrates large language models (LLMs) and social network analysis (SNA) to systematically reconstruct and analyze the inter-organizational response to this incident, addressing two research questions: (1) What structural patterns characterized the response network during the Wang Fook Court fire? (2) How did the governance structure of this network shape its capacity for coordination, information management, and adaptive response?

Methodologically, this study demonstrates a novel integration of LLM and SNA techniques for disaster governance research. The research proceeds in three stages. First, using LLM-assisted data curation, we compiled a comprehensive corpus of unstructured data sources, including official after-action reports from the Hong Kong Security Bureau, legislative council inquiries, real-time social media posts (Facebook, Telegram), local news coverage (November 26, 2025–January 2026), and NGO internal communications. A GPT-4-based pipeline extracted relational data—identifying which organizations communicated, coordinated resources, or engaged in joint decision-making during distinct response phases (evacuation, firefighting, sheltering, post-disaster social support). Second, we employed LLM-driven text classification to categorize identified responders into meaningful actor types: fire operations, government departments (non-operational), social welfare NGOs, grassroots community organizations, media, and private sector entities. Third, we constructed a response network using SNA methods, where nodes represent responder organizations and edges represent coordination or communication events extracted by LLM. We calculated standard SNA metrics including network density, centralization, core-periphery structure, and betweenness centrality to characterize the network’s governance architecture.

Expected contributions are threefold. Theoretically, this paper advances disaster governance literature by providing empirical specification of network structural conditions under which multi-organizational response succeeds or fails, moving beyond abstract typologies to measurable network configurations. Methodologically, it demonstrates the transformative potential of integrating LLMs with SNA. Traditional SNA in disaster research has relied heavily on survey-based cognitive network data or manually coded archival materials—methods that are time-intensive, static, and limited to elite perspectives. Our LLM-assisted pipeline enables rapid, scalable extraction of relational data from heterogeneous textual sources, generating dynamic, multi-perspective network models. Practically, the findings will offer actionable insights for Hong Kong’s disaster response reform, particularly the need for formalized cross-sector coordination mechanisms, shared information platforms that bridge operational and welfare functions, and recognition of grassroots organizations as legitimate coordination partners.
Privacy, Responsibility, and Government AI: Comparative Policy Analysis in Chinese Mainland and Hong Kong
Xinyang Song, Yefu Chen
School of Public Affairs, Xiamen University
As digital technologies advance, governments are using data platforms, artificial intelligence, and other digital tools to improve public administration and service delivery. These technologies have brought gains in efficiency, including faster information processing, more responsive services, and better coordination across agencies. At the same time, they have raised serious concerns about privacy protection. The collection, storage, and use of personal information create risks of misuse and unauthorized access. When problems arise, responsibility is often difficult to assign between government agencies and technology providers, making accountability for privacy breaches, technical failures, and harmful outcomes harder to establish. Further research is needed on how privacy can be protected while responsibility is more clearly defined in governing emerging technologies.

Although a large body of research has examined privacy in digital governance and artificial intelligence, two important gaps remain. One concerns the kind of policy approach that can govern the relationship between governments and technology providers in a way that both protects personal privacy and clarifies responsibility. In practice, governments often rely on private firms to develop or operate digital systems, yet the legal and policy boundaries of this cooperation remain insufficiently defined. The other concerns how governments in different regions can coordinate effectively when their institutional settings and regulatory traditions differ. This issue becomes especially important when data flows and digital governance technologies extend across jurisdictional boundaries.

Chinese Mainland and Hong Kong provide a useful case for examining these questions. Both are actively pursuing digital transformation and face similar pressures to improve administrative efficiency while maintaining public trust. At the same time, their different political and institutional frameworks create barriers to information exchange, regulatory alignment, and policy coordination. These differences shape how privacy is defined, how public authority is exercised, and how accountability is assigned.

This study proceeds in three stages: policy collection, text analysis, and comparative analysis. It collects policies and laws related to artificial intelligence and personal information protection issued in Chinese Mainland and Hong Kong between 2020 and 2026. It then extracts provisions related to the purpose and behavioral boundaries for the use of personal information by government large models and develops a framework for policy text analysis. The analysis focuses on how the two jurisdictions define these boundaries, allocate responsibility when problems arise, and provide remedial measures. It then compares the two systems and proposes recommendations for privacy protection in the use of emerging technologies.

Current findings suggest that the rules of Chinese Mainland are shaped mainly by strong regulation and bottom-line thinking, including lawful-source requirements for training data, separate consent for sensitive personal information, restrictions on unauthorized data use, security assessment and algorithm filing requirements, and mandatory labeling of AI-generated content. By contrast, Hong Kong follows a principle-based and agile approach under the Personal Data (Privacy) Ordinance, encourages privacy impact assessments, and grants practitioners’ greater discretion. These findings suggest that closer coordination between Chinese Mainland and Hong Kong should become a policy priority to improve privacy protection, clarify responsibility, and support safer digital government.
Agentic Framework for Political Biography Extraction
Yifei Zhu1, Songpo Yang2, Jiangnan Zhu3, Junyan Jiang1
1 University of Hong Kong
2 Peking University
3 Columbia University
Producing large-scale political datasets demands extracting structured facts from unstructured sources, traditionally relying on expensive human experts and resisting at-scale automation. This paper develops and evaluates large language model (LLM)-based solutions to this bottleneck, focusing on elite biographies, one consequential class of political facts. We propose a two-stage ``Synthesis--Coding'' framework: LLM agents first search, filter, and curate evidence from heterogeneous web sources, then map curated inputs into structured records. We validate the framework across Chinese, American, and OECD political elites, benchmarking performance against human baselines using multiple state-of-the-art LLMs. We find that LLM coders match or exceed human experts when given curated inputs, and that agentic synthesis substantially outperforms human collective curation (Wikipedia) in open-web environments. We further identify a systematic bias: directly coding from long, multilingual corpora degrades extraction quality, and demonstrate that the synthesis stage mitigates this bias by compressing evidence into signal-dense representations.
The What, when and where of Municipal Policy Attention in California Cities
Jonathan Colner
American University
City governments are widely acknowledged to play an important role in the policymaking activities of the United States’ governing apparatus, frequently making important decisions about public service provision and serving as the most frequent point of contact for citizens. Despite this, the study of policymaking at the municipal level has long been hindered by the lack of a centralized source of municipal government activity. I introduce a new database of every city council agenda item from every incorporated city in the state of California, stretching back as far as records are made publicly available. This dataset includes agenda items from over 200,000 meetings, with each agenda item extracted from individual PDFs of each meeting record. Alongside the agenda item text, I provide the roll call vote for each agenda item when available as well as code each agenda item by substantive topic. Both the roll call data and agenda items are validated using alternative data sources. This new dataset is the most comprehensive source of data for any geographic region in the United States. As such, it opens up a wide variety of new opportunities to study the policymaking activities of city councils from a comparative perspective.
Narratives, Beliefs and Asset Prices
Peiran Jiao1, Zihao Qu2, Fan Rao2, Xu Zhang2
1 Maastricht University
2 Hong Kong University of Science and Technology (Guangzhou)
We study whether causal narratives in retail investor discourse affect beliefs and asset prices. Using 2.35 million posts from Guba, China’s largest stock forum, we apply a large language model to extract narrative content and classify its implied economic channel. Narrative composition predicts future returns, volatility, and trading volume after controlling for sentiment, attention, and their interactions, while narrative dispersion forecasts higher volatility and turnover. To identify mechanism, we run an experiment with matched stock-forum stimuli that hold sentiment and keywords approximately constant. Causal narratives generate larger shifts in expected returns, confidence, and portfolio allocations, especially when firm-specific narratives are less dispersed.
Governing the Lingua Franca of Power: Middle Eastern Agency in the Regulation of Large Language Models
Emilie Tran1, Eric Sautede2
1 Hong Kong Metropolitan University
2 Hong Kong Baptist University

Large language models have emerged as the new lingua franca of geopolitical power — mediating communication across languages and cultures while embedding, within that very mediation, the governance preferences, content norms, and political values of whoever built, trained, and regulated them. This dual character — simultaneously a communication infrastructure and a normative architecture — makes LLM governance the most consequential frontier of the current AI regulatory contest. That contest is most commonly framed as a binary choice between China’s state-led, security-first model and the European Union’s rights-based, risk-categorised framework. Yet the most revealing laboratories of this rivalry are neither Washington nor Brussels nor Beijing, but Riyadh, Abu Dhabi, and Cairo. Gulf Cooperation Council (GCC) states, Egypt, and Iran are simultaneously receiving Chinese-built Arabic LLM ecosystems and Western foundational model platforms under governance frameworks that belong fully to neither camp. This paper examines how Middle Eastern states exercise agency in adopting, filtering, or independently constructing LLM governance norms — asking whether they are passive recipients of a China-West regulatory binary, active norm-brokers, or emergent norm-makers in their own right.

Can Government Chatbots Be Equitable Without Becoming Rigid? Evidence from Germany
ming ma, Steffen Eckhard
University of Hannover
While large language models (LLMs) have been rapidly integrated into public service delivery, existing scholarship presents competing accounts of whether it will replicate societal biases embedded in its training data or whether such biases can be mitigated through institutional and technical interventions. Yet these mitigation efforts may lead to overly rigid systems that fail to adapt to legitimate citizen needs. To examine this tension between inequality and rigidity, we develop an algorithmic administrative language framework and conduct a large-scale audit experiment comprising 19,200 queries across 12 municipal and job center chatbots in Germany. We systematically manipulated user personas (gender and ethnicity) and language proficiency signals to test for equality and responsiveness. We find that disparities in chatbot responses are manifested primarily in the informational quality of responses and are concentrated among specific identity groups. Conversely, contrary to expectations of algorithmic rigidity, chatbots demonstrate responsiveness to explicit requests for simplified language. These findings suggest that while LLMs possess the adaptive capability to reduce administrative burdens, reproduction of identity-based biases is a persistent challenge that deserves more attention.
Public Communication on the Energy Transition - a Natural Language Processing Approach for Analyzing News Data
Mareike Petrosjan1, Dinesh Korrapati2, Christin Hoffmann3, Felix Müsgens4
1 Chair of Energy Economics
2 Chair of Energy Economics
3 Chair of Energy Economics
4 Chair of Energy Economics
Achieving global mitigation targets requires rapid and effective energy transitions, whose success fundamentally depends on public acceptance. However, the complexity and transformative nature of energy transition topics create risks of knowledge gaps,misinformation and emotional polarization in public discourse. Public attitudes toward the energy transition are not static but vary with political, economic, and belief-related factors, highlighting the importance of examining how media communication may shape public perceptions and policy acceptance.

This study examines public communication on the energy transition through natural language processing techniques applied to German news media. We analyze sentiment patterns in energy-related news transcripts from Tagesschau, Germany’s most-watched news programme, providing insights into how media framing may influence public perceptions and societal discourse.

Our methodological contribution includes a novel domain-specific sentimentanalysis model developed through a human-AI annotation workflow. The model employs a stacking ensemble approach, integrating predictions from three base learners through a logistic regression meta-learner to classify content as positive, negative, or neutral. We complement this with a hybrid topic modeling framework that combines keyword-based extraction with language-model-assisted labeling, enabling systematic examination of thematic coverage patterns and temporal dynamics.

This design allows us to examine how the energy transition is framed over time, which energy-related topics dominate media coverage, and how sentiment varies across thematic areas. Preliminary results for the period 2023-2025 show that both sentiment and the intensity of media coverage differ across topics, with predominantly negative sentiment for energy costs and more mixed representations for energy policy, underscoring the importance of balanced communication for fostering public acceptance of the energy transition and broader societal engagement.
From LLMs to Agents: A Generative AI Pipeline for Mapping and Comparing AI Policy Portfolios
David García-García1, 2, Xavier Fernández-i-Marín2, 1
1 Institut Barcelona d'Estudis Internacionals
2 Universitat de Barcelona
This paper analyses how artificial intelligence is governed across a diverse set of countries, by mapping policy intervention using a portfolio approach. We collect and classify AI-related policies along two key dimensions: targets, denoting the specific objectives pursued by each policy, and instruments, referring to the regulatory or programmatic tools employed. Building on this dataset, we present a comparative description of how policy portfolios vary across countries along these two dimensions.

Our data collection and classification method relies on a pipeline grounded in text analysis and generative AI. A central methodological contribution of the paper is the systematic comparison of three increasingly complex architectures for policy classification: a standalone Large Language Model (LLM) approach, a Retrieval-Augmented Generation (RAG) pipeline, and an agentic workflow. We evaluate each architecture on classification accuracy, scalability, and robustness to heterogeneity in policy formats and language, offering practical guidance for researchers seeking to deploy generative AI in large-scale policy analysis.

The resulting method is designed to be both scalable to other policy sectors and transferable across constituencies, whether countries, regions, or local entities. We detail the classification scheme and pipeline steps that enable systematic cross-country comparison. While our primary focus is on data collection and classification, we highlight patterns of convergence and divergence in AI regulation that provide an empirical foundation for understanding the politics of AI governance and lay the groundwork for subsequent analyses of how policy portfolios shape trajectories of technological development and innovation.
Mapping out Elite Networks: Methods and Application to German Political and Economic Elites Between 2013 and 2023
Lisa Garbe, Yuequan Guo, Macartan Humphreys, Lennard Naumann
WZB Berlin Social Science Center
We propose a scalable and interpretable method for constructing social networks between elites across diverse arenas. The method decomposes the construction of elite social networks into three steps: a) identifying elites in respective arenas, b) choosing the media that encode interactions among elites, c) decoding the media to measure elite connections via large language models (LLMs). The method complements existing network data collection methods, such as via interviews or archives, which can be challenging to apply at scale in the study of ever-evolving elite politics. We apply this method to construct social networks between political and economic elites in Germany between 2013 and 2023 based on a large corpus of articles retrieved from 30 German national newspapers. By comparing LLM and human coding on a random sample, we show that our method can adequately capture elite connections as reported in newspapers. We further assess the predictive validity of our newspaper-based elite network and discuss its pros and cons vis-à-vis existing methods of elite network construction.
Normative Rigidity vs. Probabilistic Fragility: Identifying the Phase Transition of Systematic Failure in Large Language Models via Structured Legal Logic
Runyi Ma1, 2, 3, 4, Jinting Deng5
1 Center for Digital Rule of Law, Peking University
2 Law School, Renmin University of China
3 Beijing Association for Digital Economy and Digital Governance Rule of Law Studies
4 Peking University’s Analytics Lab for Global Risk Politics
5 Discipline Inspection and Supervision School, Renmin University of China
This research addresses the conflict between the normative rigidity of legal reasoning and the probabilistic fragility of LLMs. We leverage a high-fidelity dataset of thousands of statutory rules across 300 case causes, manually transformed by legal experts into "Chain-Graph JSON" structures. Utilizing a rigorous methodological framework, we quantify legal complexity through multiple statistical and topological metrics, including nesting depth (D), node density, and branching factors. Our empirical stress tests across leading models (GPT, Gemini, Claude, DeepSeek) reveal a distinct phase transition: beyond a critical complexity threshold, the Logical Alignment Score suffers an abrupt, non-linear collapse, regressing models into "stochastic parrots" with structural hallucinations. To mitigate this, we establish the "LL-Turing" Benchmark and demonstrate that Logical Rails—structured protocols for model constraints—are an ontological necessity for ensuring determinism. This study pivots Computational Social Science (CSS) from a data-driven to a "structure-logic dual-driven" paradigm, providing a robust empirical template for global risk governance.
An Expert–Coder Multi-Agent System for Codebook-Based Text Annotation in Social Science
Songpo Yang1, Yang Wu2, Zhicheng Zhang2, Xun Pang1
1 Peking University
2 University of Chinese Academy of Sciences
Codebook-based text coding is a foundational task in empirical social science, yet remains bottlenecked by its dependence on trained human coders who must process lengthy, multilingual documents under heavy cognitive load. Large language models (LLMs) offer a promising alternative, but direct application produces unreliable results when codebook categories involve ambiguous conceptual boundaries or require contextual judgment. This paper proposes a multi-agent LLM framework that replicates the division of labor between human experts and coders. The framework transforms conventional codebooks into sequenced Query Workflows, implements a dual-agent architecture with a Coder Agent for extraction and an Expert Agent for interpretive oversight, and introduces a Query-Feedback loop in which conceptual ambiguities are escalated for adjudication. A task scheduler leveraging asynchronous calls and KV cache management ensures scalability. We validate the framework on two tasks: coding investment screening mechanisms using the PRISM Dataset (Princeton University), covering 38 OECD countries and 130 legal texts, and coding interstate interactions from UNFCCC climate negotiation reports building on Castro et al. (2025). These cases test factual precision on structured policy documents and semantic inference on deliberately ambiguous diplomatic language, respectively. Preliminary results indicate that the framework achieves inter-coder reliability comparable to trained human coders while reducing costs by an order of magnitude. We further find that both the Query Workflow transformation and Expert Agent consultation contribute independently to coding quality, and that naive single-pass prompting produces substantially lower reliability.
How Does Structure Make War? Testing Thucydides-Trap Logics in LLM-Based Multi-Agent Simulations
Yifei Liu1, Chao Gu1, Yuang Pan wang3, Hanfang Zhang2, Shuo Chen2, Rui han Cao4
1 School of Government, Peking University
2 Beijing Institute for General Artificial Intelligence
3 School of Public Administration and Policy, Renmin University of China
4 National School of Development, Peking University
The Thucydides Trap remains one of the most contested propositions in contemporary international relations. Allison argues that when a rising power threatens to displace a ruling power, the structural pressures of power transition make war highly probable. This claim has profoundly shaped policy debates over the trajectory of U.S.–China relations, yet it has also drawn substantial criticism for selection bias, crude historical analogies, and structural determinism. The central question underlying this debate remains insufficiently resolved: under what combination of structural conditions does a power transition slide toward war rather than peaceful adjustment? A fundamental difficulty for empirical research is that key structural variables—relative national power, economic interdependence, alliance reliability—cannot be experimentally manipulated in the real world, while the small number of historical cases inevitably confounds structural factors with contingent ones.

This article addresses this challenge by building a large-scale multi-agent simulation system. Three LLM agents, representing a status-quo power, a rising challenger, and a third-party ally, interact within a transparent and configurable structural environment governed by deterministic update rules. The article advances two theoretical contributions. First, in contrast to existing studies that use LLM agents as analogues for individual human actors, we argue that state-level agent simulation has stronger theoretical applicability. State behavior is largely constrained and shaped by structural factors, exhibiting greater regularity and predictability than individual behavior, which makes LLM-based simulation of state conduct more theoretically grounded. Second, we position LLM-based multi-agent systems as tools for theory testing and mechanism exploration rather than for predicting real-world outcomes. By systematically manipulating five sets of structural variables—the speed of relative power convergence, trade expectations, alliance reliability, diplomatic communication density, and the pace of convergence in critical military technologies—we reconstruct the complex dynamic process through which multiple structural mechanisms interact during power transitions.

Across thousands of simulation runs, we find that power transition alone does not reliably produce war. War becomes frequent and early only when rapid power convergence coincides with pessimistic economic expectations, highly reliable alliances, sparse high-level diplomacy, and fast technological catch-up. Conversely, slow convergence, optimistic expectations, flexible alliances, dense communication, and gradual technology diffusion consistently steer the system toward peace. These results reframe the Thucydides Trap as a cluster of interacting structural mechanisms rather than a deterministic law.
Populism, Policy Volatility, and GVC Fragility: Evidence from an LLM-Driven Agent-Based Model
chenjun gao
school of international studies, Peking University

How does populist leadership affect the resilience of global value chains (GVCs), and through what mechanisms? This study argues that populist leaders are not only more prone to policy volatility, but also more likely to shift abruptly between confrontation and cooperation under political pressure, thereby increasing uncertainty for multinational firms and weakening the resilience of cross-border production networks. We address this question through a mixed-method design that combines: (1) a country-year panel analysis of populism and GVC resilience from 1990 to 2019; (2) operational-code analysis of political leaders’ psychological traits; and (3) a large-language-model-driven agent-based model (LLM-ABM) with matched treatment-control simulations, network metrics, and difference-in-differences estimation. The empirical results show that the populism-resilience relationship is distinctly nonlinear: adverse effects are limited at low-to-moderate levels of populism but become substantially stronger once populism enters the upper range. At the micro level, populist leaders exhibit a significantly higher tendency to switch between conflictual and cooperative postures, which intensifies firms’ perceived risk. In the simulations, this mechanism produces higher disinvestment, weaker node connectivity, greater trade concentration, longer effective network distance, and lower clustering, yielding a more centralized and fragile GVC structure. The study contributes to research on populism, political psychology, and international political economy by showing that GVC resilience depends not only on tariffs or formal trade barriers, but also on leadership volatility and policy predictability. The policy implication is clear: stable and credible policy signals are essential for preventing defensive capital clustering and preserving resilient global production networks.
Knoweia as a Learning Dialogue Terminal: A Multi-Agent AI System for Structured Learning in Computational Social Science
Tang Shiyun1, Hou Yuxin2, 3, Ma Runyi2, Hu Jingtian4, Pang Xun2
1 School of Economics, Renmin University of China
2 PKU Analytics Lab for Global Risk Politics, Peking University
3 Center for Social Research, Peking University
4 Tsinghua University

This paper presents Knoweia, a dialogue-centered AI learning system designed to support computational social science education, where students must coordinate causal reasoning, research design, coding, and data analysis. Rather than treating large language models as answer generators, Knoweia functions as instructional infrastructure that organizes learning as a guided and stateful process.

We introduce the concept of Instructional Navigation Capacity (INC), defined as a learner’s ability to orient within a task space, interpret obstacles, and progress under uncertainty. INC is operationalized using process-trace data, including persistence after errors, recovery from repeated blockers, structured task progression, and patterns of help-seeking.

Knoweia implements this framework through a multi-agent architecture consisting of a learner-facing Companion Agent, a Roadmap Manager that tracks and updates task trajectories, a Memo Agent that records longitudinal learning patterns, and an expert layer for controlled escalation. The system is deployed in an ongoing course with 236 students from diverse disciplinary backgrounds.

We combine baseline survey data with fine-grained interaction traces to examine how dialogue-centered AI guidance relates to learning-process outcomes such as persistence, error recovery, and trajectory coherence. By linking conceptual framing, system design, and behavioral data, the paper contributes a process-oriented framework for evaluating human–AI learning systems in computational social science.
Value Alignment of Embodied Intelligence from the Perspective of Public Value: Logic, Tensions, and Governance Pathways
Yuting Huang, Ni Yang
University of Science and Technology Beijing
As embodied intelligence transitions from laboratory settings to real-world applications, the issue of value alignment—ensuring that intelligent agents act in accordance with human values—has become a critical public governance challenge. Drawing on public value theory, this paper constructs an integrated “logic–tension–governance” framework to systematically analyze the underlying mechanisms of value alignment in embodied intelligence. The study identifies three fundamental tensions: between technological and social logics, between universality and particularity, and between agency and controllability. In response, it proposes a collaborative governance framework that emphasizes multi-stakeholder participation, institutional embedding (such as ethics-by-design and a dual-immune architecture), and cross-cultural adaptation. By introducing public value theory into the analysis of embodied intelligence, this study offers both a theoretical lens and practical pathways for governing value alignment in the era of intelligent machines.
Beyond Hate Speech Detection: Evaluating Large Language Models on Incitement in Far-Right Telegram Networks
Mara Bulzan
Columbia University
This paper evaluates whether large language models can detect incitement in digital extremist communication more effectively than conventional hate speech classification approaches. Existing automated detection systems are typically designed to identify overtly toxic or hateful language at the level of the single post. Yet far-right communication often relies on coded references, historical allusion, symbolic mobilization, and the distribution of harmful elements across posts rather than within one clearly inciting statement. This creates a serious analytical gap between what current models are built to detect and how incitement actually operates in digital extremist spaces.

To address this problem, the paper draws on an original hand-coded dataset of 521 public Telegram posts from 17 far-right channels in Romania, Hungary, and Poland. The dataset operationalizes three incitement-relevant dimensions derived from the Rabat Plan of Action: harmful framing, group targeting, and mobilization. Using this dataset as a benchmark, the paper evaluates how LLM-based classification performs on legally and politically salient forms of harmful speech that exceed standard hate speech taxonomies.

The analysis shows that LLMs are reasonably effective at identifying explicit dehumanization and direct targeting, but are far less reliable when faced with coded ideological language, commemorative symbolism, or posts whose mobilizing significance depends on channel context rather than surface toxicity. The paper argues that this limitation is not merely technical. It reflects a deeper mismatch between post-level detection frameworks and extremist communication environments in which incitement is cumulative, relational, and context-dependent. The study therefore contributes to ongoing debates on LLM-based text classification, legal interpretation, and the governance of digital harms.
Stakes, Prospects, and Policy Domain: Citizens' Willingness to Delegate Political Decisions to Artificial Intelligence
Elif Erisen
Yeditepe University
please see the attached pdf document
Beyond the Questionnaire: LLM-Driven Avatar Debriefing in Social VR Experiments as a Methodological Innovation for Social Science Research
Elif Erisen1, Saliha Akbas2
1 Yeditepe University
2 Koc University
Please see the attached extended abstract.
From Image-Based Documents to Intelligent Research Tools: A Localized Collaborative Research Framework Integrating Knowledge Bases and Large Language Models for Dunhuang Studies
Hongjuan zhao1, ruohan ma2
1 Qingdao University of Science and Technology
2 Qingdao University of Science and Technology
Abstract: The digital transformation of Dunhuang Studies has significantly improved access to manuscripts, paintings, catalogues, and related research materials. Yet the visibility of resources does not automatically translate into their usability for research. Although large-scale digitization projects have substantially alleviated the difficulties caused by the transnational dispersal of Dunhuang materials, core scholarly challenges—such as the aggregation of topic-specific materials, the tracing of documentary provenance, the identification of relationships between versions, the reconstruction of historical contexts, and the establishment of cross-textual knowledge linkages—have not been resolved simply through the enrichment of digital resources. Against this background, this article argues that the next stage of digital Dunhuang research should not be marked merely by the expansion of image repositories or the convenience of catalogue retrieval, but by the construction of knowledge infrastructure oriented toward research practice and characterized by computability and verifiability.

With this objective in mind, the article proposes a localized collaborative research framework for Dunhuang Studies that integrates intelligent document parsing, structured text extraction, knowledge base construction, retrieval-augmented generation, and human verification into a unified workflow. Grounded in the disciplinary characteristics of Dunhuang Studies—namely the dispersal of materials, the coexistence of multiple languages, unstable image quality, complex page layouts, and a strong reliance on provenance awareness and philological control—this framework employs MinerU as the document parsing tool, Qwen3.5 as the locally deployed large language model, and MaxKB as the knowledge base platform. In doing so, it reconstructs a research pipeline that moves from image-based documents to structured knowledge units and then to evidence-driven scholarly interaction. The article emphasizes that the role of artificial intelligence in Dunhuang Studies should not be understood as a substitute for scholarly judgment, but rather as a form of research infrastructure: its function lies in reducing repetitive labor, reorganizing dispersed evidence, supporting exploratory and verification-oriented inquiry, and strengthening the evidentiary basis of historical interpretation.

This article advances three principal arguments. First, the significance of local deployment in Dunhuang Studies lies not only in technical feasibility and data security, but also in its capacity to safeguard scholarly interpretive authority, the continuity of annotation work, and the transparency of evidentiary chains. Second, the core value of retrieval-augmented interaction in this field does not lie in the direct generation of answers, but in the structured reorganization of evidence across manuscripts, catalogues, and prior scholarship. Third, the incorporation of intelligent tools into Dunhuang research should not be interpreted as the automation of humanities scholarship, but rather as a reorganization of the conditions of knowledge production. By facilitating a shift in Dunhuang Studies from “resource visibility” to “knowledge usability,” this collaborative architecture not only offers a new methodological pathway for manuscript studies, but also reveals both the potential and the epistemological limits of AI-assisted humanities research.
The Theoretical Logic and Practical Pathways of Empowering the Translation of Gong’an Literature Through Large Language Models
Jingwen Zhang, Hongjuan Zhao
Qingdao University of Science and Technology
With the rapid development of large language models in text generation, semantic understanding, and translation assistance, their potential application in literary translation has become increasingly prominent.Gong'an literature(公案文學) is a distinctive literary genre unique to ancient China and a concentrated literary embodiment of the spirit of the rule of law in traditional Chinese society.As a literary form that combines narrative, judicial, cultural, and historical dimensions, the translation of gong'an literature involves not only the cross-linguistic transfer of plot, legal terminology, and judicial reasoning, but also the effective transmission of ancient Chinese judicial culture, ethical values, and the spirit of the rule of law.Compared with general literary texts, the translation of gong'an literature faces multiple challenges, including the need to balance plot development with legal expression, the intertwining of historical context with institutional meaning, and the complexity of character discourse and adjudicative logic.Against this background, this article takes "the theoretical logic and practical pathways of empowering the translation of gong'an literature through large language models" as its central focus, exploring the practical value, scope of applicability, and optimization strategies of employing large language models in the translation of gong'an literature.This article first examines the textual features of gong'an literature and analyzes the advantages of large language models in terminology recognition, discourse restructuring, contextual supplementation, multi-version generation, and translation assistance. It argues that the theoretical logic underlying the empowerment of gong'an literature translation by large language models can be understood in four dimensions: first, a language generation logic based on contextual prediction and semantic association; second, a knowledge-enhancement logic oriented toward the interpretation of ancient legal culture; third, a human–machine collaborative logic that emphasizes the translator's leading role combined with model assistance; and fourth, a logic of meaning transformation serving the international dissemination of Chinese legal culture. At the same time, this article also points out that large language models still suffer from such problems as misjudgment of historical context, misreading of legal culture, smoothing of narrative style, weakening of culturally loaded information, and hallucinated generation.They therefore remain unable to replace the central role of human translators in cultural discernment, institutional interpretation, and the reproduction of literary style.On this basis, the article proposes a human–machine collaborative pathway for the translation of gong'an literature.By constructing an integrated support system combining a corpus, a terminology database, and a knowledge base, it advocates a multi-layered translation process consisting of model-generated draft translation, human revision, expert review, and dissemination feedback, so as to achieve a balance among translation efficiency, academic accuracy, and cultural communication.This article argues that the value of large language models in the translation of gong'an literature lies not in replacing human translators, but in serving as intelligent auxiliary tools that enhance the organizational capacity, interpretive depth, and communicative effectiveness of translating highly complex cultural texts.Exploration of this issue not only helps expand the interdisciplinary horizon between intelligent translation studies and research on the translation of ancient Chinese literature, but also provides new methodological references for the international dissemination of ancient Chinese legal culture.
Can Conversational AI Influence Confidence in Pre-existing Political Beliefs?
Semra Sevi, Can Mekik
University of Toronto
Conversational artificial intelligence (AI) may influence political polarization by shaping the strength with which political attitudes are held. Generative AI systems are increasingly capable of conducting adaptive, human-like conversations, raising questions about their ability to strengthen or weaken belief certainty. We report results from a preregistered, two-wave randomized experiment (N = 1,968) testing whether GPT-4 can strengthen or weaken individuals' attitudes on salient, evolving political issues. Participants engaged in personalized, multi-turn conversations in which the AI was instructed to either strengthen or weaken their stated views, or to engage in non-political discussion. Across issues including housing, immigration, and economic policy, AI-mediated conversations consistently reduced attitude strength. Specifically, belief confidence, while showing no evidence of strengthening. These effects persisted for up to 21 days. Our findings indicate that generative AI can durably influence political attitudes, primarily by weakening rather than strengthening belief certainty. By examining both pro and counter-attitudinal persuasion among individuals with existing views in dynamic political contexts, this study provides a demanding test of AI’s persuasive capacity and clarifies its implications for political polarization and democratic discourse.
Investigating Self-View Convergence in LLM Agents and Human-AI Interaction
Zengchang Qin
Centre for AI Research (CAIR) and School of Engineering and Computer Science (CECS), VinUniversity, Hanoi, Vietnam
The human self-concept is a dynamic construct shaped through social interaction—a process known as "inter-self alignment." As Large Language Models (LLMs) increasingly serve as social actors rather than mere tools, understanding their capacity to influence and undergo self-view convergence is critical. This research investigates whether the mechanisms of social alignment observed in human-human dialogue extend to autonomous AI agents and mixed groups populated with real human and AI agents. The study employed instruction-tuned Gemma agents in a four-participant round-robin protocol to test convergence across two distinct domains: a perceptual-visual task and a social self-revelatory task. Findings suggest that interaction with AI agents can measurably influence human self-perception in a manner analogous to human-human interaction.
EU Digital Norms Beyond the Union? AI and Digital Policy Convergence in the Western Balkans
Suljo Corsulic
Faculty of Social Sciences, University of Duisburg-Essen
Rapid advances in artificial intelligence (AI) and digital technologies have intensified the need for new regulatory frameworks across Europe. The Western Balkans, as a region strongly shaped by European Union (EU) norms and accession conditionality, has been a rather slow developer in EU-alignment of policies, marking digital and AI policy as a new challenge in this field. This study examines the extent to which the six Western Balkan states (Serbia, Montenegro, Bosnia and Herzegovina, Kosovo, North Macedonia, and Albania) have aligned their national digital and AI policy frameworks with key EU regulatory instruments, namely the AI Act, Digital Services Act (DSA), and Digital Markets Act (DMA). Drawing on the theories of Normative Power Europe and Europeanisation, the paper investigates whether stronger alignment with EU digital policy is associated with more advanced stages of EU accession. Using qualitative comparative policy analysis and directed qualitative content analysis, the study systematically codes policy convergence across key dimensions derived from EU digital regulations, including risk-based AI governance, transparency requirements, platform regulation, and market competition rules. Furthermore, the paper develops an alignment typology and compares it with countries’ accession progress. The study contributes to debates on EU normative power and Europeanisation by examining these dynamics in the emerging field of digital and AI governance. More broadly, it assesses whether the EU continues to shape policy development in candidate and potential candidate states in a contested field that is increasingly shaping the contemporary international politics and technological governance.
When AI Thinks About Culture: Comparing Language-Based and Empirical Measures of Individualism-Collectivism
Plamen Akaliyski1, Wolfgang Messner2
1 Department of Sociology and Social Policy, Lingnan University, Hong Kong SAR, China
2 Department of International Business, Darla Moore School of Business, University of South Carolina, USA
Large Language Models (LLMs) are increasingly used to generate cultural insights, yet little is known about their cultural expertise and the extent to which their outputs replicate or correct for cultural stereotypes. This study compares country-level estimates of Individualism–Collectivism (I-C) derived from five state-of-the-art LLMs with a recently validated I-C index based on nationally representative surveys. The LLM-based estimates correlate highly with the survey-based index (up to r = .874) but also exhibit systematic biases: Western societies are consistently overestimated on individualism, while Confucian East Asian societies are underestimated. Statistically, the type of bias suggests group-specific baseline distortions rather than random noise, with smaller and older models showing weaker alignment. Moreover, the patterns bear the imprint of Hofstede’s legacy framework, indicating that LLMs absorb not only empirical information but also historical discourse. These findings demonstrate both the potential and the risks of using LLMs as cultural interpreters. Used responsibly – with calibration, human oversight, and awareness of their biases – LLMs can enrich cross-cultural relations and support global coordination. Used uncritically, they risk reinforcing outdated narratives that misdirect academic research and managerial decisions.
Productive Disagreement as a Skill: AI-Supported Training for Political Conversations
Ethan Busby1, Lisa Argyle2, Joshua Gubler1, Alex Lyman1, David Wingate1
1 Brigham Young University
2 Purdue University
A growing body of academic research and practitioner evidence shows that, under the right conditions, political conversations can reduce political animosity, de-escalate conflict, and facilitate compromise. However, not all difficult political conversations are so idyllic: many are heated, combative, and create feelings of stress, frustration, or anger. Because many people are concerned about such negative impacts on themselves and their relationships, they avoid engaging at all in political conversations with people they disagree with. Even when they do choose to participate in these discussions, people generally lack the civic skills for productive engagement in divisive, disagreement-based, or conflictual interactions.

We propose that productive disagreement is a skill that can be improved with instruction and practice. At the same time, existing programs that can teach this skill are severely resource constrained and face limits on their scalability. In this project, we suggest that generative AI tools can step into this space to both train people to develop these disagreement skills and test transfer from this learning into conversations with others. Specifically, we develop a set of AI agents tailored to provide real-time coaching and training to participants about engaging productively in political discussion using evidence-based best practices drawn from a variety of fields. Post-training, we also deploy a separate AI agent to give people practice engaging in potentially heated discussions without posing any risk to their real-life relationships and as an evaluation tool. We then evaluate whether practice and coaching improve people’s confidence and ability to engage productively in divisive political conversations. In this paper, we introduce our framework and approach to this use of AI and provide empirical demonstrations of its value.
PriceReason: A ReAct-Based LLM Agent for Explainable Dynamic Pricing in Retail
Mitchelle Ashley Creado
Independent Researcher
Today, walk into any major grocery store and the price on the shelf was almost certainly set by an algorithm. That algorithm is fast, data-driven, and ruthlessly optimized for margin, but it cannot tell you why it chose that number in the first place. This matters more than it might seem. Category managers who cannot explain a recommendation to their supervisor end up ignoring it. The regulators that cannot audit a pricing decision are starting to treat that as a legal problem. And, as recent empirical work has shown, LLM-based pricing agents left to their own devices will quietly push prices above competitive levels without additional instruction to do so. We built PriceReason to tackle all three of these problems at once. The core idea is simple: force the agent to explain itself and before it commits to a price. PriceReason is a ReAct-loop LLM agent that must call real data tools, reason over what the agent finds, and produce a structured natural-language rationale for every price it recommends. A fairness guardrail embedded in the system prompt puts a ceiling cap on how far above the competitive equilibrium any recommended price can go. We tested PriceReason on the Dunnhumby Complete Journey grocery dataset over five independent runs. The results were encouraging on all three fronts: gross margin came within 3.4% ± 0.6% of a reinforcement learning ceiling cap, the collusion index dropped by 71% ± 4% compared to an unconstrained LLM, and human evaluators rated the rationales substantially higher on

fairness clarity. The finding that surprised us most was this: even without the explicit price cap, simply requiring the agent to write a fairness assessment reduced collusive drift by around 57%. To the best of our knowledge, this is the first system to demonstrate that structured deliberative reasoning can simultaneously serve the collusion prevention, transparency, and commercial performance in a retail pricing context.
Rewriting Scientific Authority: Epistemological Challenges in Arabic Academic Writing when AI Becomes a Co-Author
Morad Diani
Arab Center for Research and Policy Studies
The integration of large language model (LLM)-based writing tools into Arabic academic writing presents epistemological challenges that exceed the terms of dominant, English-centric debates about AI co-authorship. These debates have largely focused on plagiarism, attribution, and institutional policy; they have not asked what happens when AI tools embed Anglophone epistemic norms into writing traditions that do not share those norms. Drawing on postcolonial Science and Technology Studies (STS), language ideology theory, and studies of Arabic academic discourse, this paper develops a conceptual framework to analyze how AI co-authorship destabilizes the mechanisms by which Arabic-language scholars construct authority, voice, and disciplinary legitimacy. It introduces the concept of epistemic default to describe the silent importation of foreign knowledge conventions into AI-generated Arabic text, and rhetorical ventriloquism to describe the performance of a surface-level Arabic scholarly register that conceals structurally alien argumentative logics. The paper argues that AI co-authorship in this context is not primarily a matter of text integrity but of epistemic sovereignty, and that its differential adoption across Arabic academic institutions risks deepening existing stratifications in the global knowledge economy. The analysis contributes to the general theory of AI and authorship by demonstrating that authorial voice is not a culture-neutral concept, and that the stakes of AI co-authorship differ systematically across epistemological locations.
Social Perception of Everyday Discrimination: Human and LLM Judgments in Korean Vignettes
Ji Hye Kim, Mikyoung Kim
Sogang University
Everyday discrimination is not always recognized as discrimination. It often appears in ordinary exchanges: a comment about school background, region, family form, appearance, age, gender, disability, health, or occupation. Because these moments are embedded in context, they are difficult to study through direct questions about personal belief alone. They are also difficult to capture through experience-based reports, since people encounter different situations and may hesitate to name ambiguous harms as discriminatory. This paper therefore begins with a vignette survey designed to examine how people judge the same situated interactions.

We analyze survey responses from 2,270 Korean adults who evaluated 30 Korean vignettes describing ordinary but potentially discriminatory or exclusionary situations. For each vignette, respondents were not asked to decide whether the situation was objectively discriminatory. Instead, they inferred how the target of the interaction would feel on a six-point scale. The study treats the perception of discrimination as a socially situated judgment about context, harm, and likely reception.

We compare these human response distributions with an initial pilot of LLM-generated judgments from GPT-5.5 Pro, Claude Opus 4.8, and Gemini 3.5 Thinking on the same vignettes. The pilot suggests that LLMs do not simply fail to recognize everyday discrimination. Rather, they tend to recognize it in a more uniformly negative and normatively stabilized way. Across the 30 items, the average of the human item means was 4.09 on the six-point scale, while model means were higher: 4.93 for GPT-5.5 Pro, 4.50 for Claude Opus 4.8, and 4.53 for Gemini 3.5 Thinking. The models rarely used the lower end of the scale, and GPT-5.5 Pro did not assign any item below 4. These preliminary results suggest a pattern of ambiguity compression: LLMs may translate socially variable judgments into more settled moral evaluations.

This compression has sociological consequences. In everyday life, disagreement about whether a comment is hurtful or discriminatory is not merely measurement noise; it is part of how discrimination is recognized, minimized, contested, or normalized. If LLMs smooth this variation, they may overstate social consensus, obscure differences across social groups, and turn ambiguous experiences into more authoritative model judgments. This matters because LLMs are increasingly used as conversational partners for workplace conflict, family tension, school experiences, relationship problems, and vague feelings of discomfort. By placing LLM judgments against large-scale survey data from South Korea, this paper asks what is gained, and what is lost, when culturally specific perceptions of everyday discrimination are translated into model-generated judgments.
Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis
Yiqing Xu2, Leo Yang1
1 Hong Kong Baptist University
2 Stanford University
Computational reproducibility is central to scientific credibility, yet verifying published results at scale remains costly. We develop an AI-assisted workflow for automated full-paper replication -- retrieving materials, reconstructing environments, executing code, and matching outputs to point estimates reported in regression tables. We define a universe of all empirical and quantitative papers from the three top political science journals (2010--2025) and measure stated data availability using automated extraction. For a stratified sample of 384 studies, we apply the workflow to conduct full-paper replication, totaling 3,523 empirical models. We find that journal verification requirements, combined with data archiving mandates, drive reproducibility: the share of fully or largely reproducible papers rises from 20.8% before DA-RT adoption to 82.5% after, and conditional on accessible replication packages, 92.1% of papers are fully or largely reproducible (234/254). As a secondary application, we apply standardized IV diagnostics to 84 studies (597 IV specifications among 1,910 replicated models), illustrating how automated execution enables systematic reanalysis across heterogeneous empirical settings.
Salience of Disadvantaged Position Reduces AI Aversion in Moral Delegation
Xiaoli Guo1, Jianning Dang3, Xiaolin Mei1, Yin Wu1, 2
1 Department of Applied Social Sciences, The Hong Kong Polytechnic University, Hong Kong, China
2 Mental Health Research Centre, The Hong Kong Polytechnic University, Hong Kong, China
3 Beijing Key Laboratory of Applied Experimental Psychology, Faculty of Psychology, Beijing Normal University, Beijing, China
Despite rapid advances in artificial intelligence (AI), people remain reluctant to delegate moral authority to AI systems. Prior research has largely explained this reluctance through third-party perceptions of AI's limited mind, while paying less attention to how stakeholders' own positions in a decision may shape such attitudes. Across three main studies and three supplementary studies (N = 2,562), we found that when stakeholder positions were made salient, disadvantaged participants showed a significantly larger increase in AI delegation preference from baseline to the position-salient condition than advantaged participants. The effect held whether positional (dis)advantage was defined by experimental assignment (Study 1), self-perception (Study 2), or pre-existing social category (Study 3). Perceived self-interest alignment mediated this effect across studies, whereas alternative explanations (e.g., perceive fairness) showed inconsistent support. Notably, position salience reduced AI aversion among disadvantaged participants without reversing it, and left aversion among advantaged participants largely unchanged.
Fairness Versus Efficiency in AI Advice: Evidence from Human and LLM Responses
Hongying Luo, Yohanes Eko Riyanto, Yuwen Zhou
Nanyang Technological University
AI advice can help groups coordinate, but it may also allocate the costs of coordination unequally. We examine this trade-off in a repeated three-player threshold public-good game, where efficiency requires the two lowest-cost players to contribute. An LLM advisor provides either private, personalized advice or public, group-level advice. Among LLM agents, group advice almost eliminates coordination failure, but it does so by concentrating contribution costs on particular players. Human participants respond differently. When exposed to the same advice, they often sacrifice efficiency for equality: they contribute less, provide the public good less frequently, and reject costly contribution recommendations, yet achieve a more equal distribution of net payoffs. These findings suggest that AI advice can resolve coordination problems, but it does not eliminate the distributional conflict generated by efficiency-enhancing recommendations.
Information Conditions Govern the Validity of AI-generated Experimental Data Across Inferential Targets
Na Liu, Yue Wang, Lingling Hou
Peking University, China Center for Agricultural Policy, School of Advanced Agricultural Sciences, Beijing, 100871, China.
Whether AI-generated experimental data supports valid statistical inference depends on what the researcher intends to estimate and what information the model receives. Existing evaluations report contradictory findings because they assess different inferential targets against different benchmarks. Aggregate replication succeeds while individual prediction fails, but this reflects the target's statistical demand, not a disagreement about AI capability. Here we develop a multi-target evaluation framework tested on an incentive-compatible economic experiment with individual-level ground truth (N = 911, randomized treatment assignment). Seven large language models are assessed under four information conditions against four targets of increasing demand: distributional fidelity, average treatment effects, individual prediction, and heterogeneous treatment effect identification. Three regularities emerge. First, information dominates architecture. Varying what models know changes performance by 45 to 60 percentage points, while varying which model is used changes it by at most 12 points. Second, validity degrades monotonically across the target hierarchy, with one exception: reasoning models detect average treatment effects without individual data where standard models cannot. Third, optimizing for one target impairs another. Fine-tuning improves distributional matching but degrades treatment effect detection. These results resolve the apparent contradiction in existing literature and provide a replicable evaluation protocol for any domain where individual-level ground truth is available.
Confront or Concede? Evaluating Large Language Models in International Crisis Decision-Making
yiyi Chen, Zikang Chen
School of politics and international relations,Lanzhou university
Large language models (LLMs) are increasingly being considered for planning, analysis, and decision-support tasks, including national security settings such as intelligence analysis, operational planning, and military decision-making. In these domains, errors are not merely textual failures: a wrong recommendation may misread adversary intent, amplify perceived hostility, or reduce the possibility of conflict mitigation. Before LLMs are used in high-stakes crisis contexts, it is therefore necessary to examine their crisis decision propensity under historically grounded conditions.
This paper asks whether LLMs placed at international crisis decision points tend to confront, concede, or calibrate their choices to historical constraints. Rather than asking LLMs to freely simulate conflicts, the study builds a historical crisis-decision benchmark. Each task places an LLM at a real pre-decision point in an international crisis and provides only information available before the historical decision, including crisis background, adversary behavior, domestic and international constraints, time pressure, and a bounded action space. The LLM then ranks policy objectives, selects an action, explains rejected alternatives, and states key uncertainties.

The benchmark is designed to cover a broad security-risk spectrum, including nuclear brinkmanship, territorial disputes, cyber disruption, attacks on energy infrastructure, terrorism-related retaliation, grey-zone confrontation, and accidental escalation. LLM choices are evaluated against historical state actions using four core metrics: historical alignment, escalation deviation, goal-action consistency, and risk acceptance. These metrics allow the study to distinguish confrontational moves from concessionary or calibrated responses. Masked and unmasked task versions are used to distinguish decision behavior from historical recall, while repeated runs and model-family comparisons assess whether observed decision styles are stable or model-specific.

The current stage validates the research pipeline rather than reporting formal LLM findings. A pilot based on six decision points from two sample crises confirms that the action ontology and metrics can distinguish concessionary, historically aligned, and confrontational baselines. The next step is to run representative LLM families across the validated decision points and compare their decision styles under identical information constraints. The study contributes a historically grounded method for auditing LLM behavior in high-risk security decision tasks and offers a framework for evaluating whether AI systems are more likely to confront, concede, or remain calibrated to historical decision conditions.
Replication and Specification Range Analysis with AI Agents
Philipp Klotz, Sebastian Kranz, Simon Maier, Alexander Rieber, Dennis Steinle
Ulm University

Large language models are increasingly becoming research infrastructure for the social sciences, yet their methodological role remains unsettled: can LLM agents serve not only as assistants, but as controlled analyst populations for measuring the robustness of empirical findings? This paper uses LLM agents to scale the many-analyst paradigm in applied microeconomics. Human many-analyst studies show that independent researchers often reach substantially different estimates when analysing the same data and hypothesis, but such studies are expensive, slow, and difficult to decompose because human teams differ on many analytical choices at once. We replace human teams with independently prompted LLM agents and apply the design to twenty-four published papers across difference-in-differences, instrumental variables, randomised controlled trials, and regression discontinuity designs.

For each paper, we extract testable hypotheses from the published text and assign multiple agents to analyse them under six information conditions. Three tiers progressively restrict degrees of freedom by providing raw data, processed data, or a prescribed method. Additional directed conditions ask agents to test the original claim or to search for the most strongly supportive and most strongly rejecting defensible specifications. This design makes analyst discretion observable, auditable, and decomposable into data-preparation, specification, and estimator components.

We synthesise agent estimates using standardised effect sizes and a multilevel random-effects model that separates sampling variation, analyst-within-paper variation, and between-paper heterogeneity. The resulting estimands quantify how much uncertainty in published empirical results is attributable to analytical choice, how strongly LLM agents agree under fixed information conditions, and how far a motivated but defensible analyst could move a result. A pilot on two papers and roughly 320 agent analyses shows that the pipeline produces interpretable variance decompositions, robustness ladders, specification curves, and agent evidence ratings. The project contributes a reproducible benchmark and infrastructure for evaluating LLM-based social-science workflows, shifting AI-assisted replication from occasional case studies toward routine, scalable robustness reproduction.
Understanding the Mechanism of Altruism in Large Language Models
Songfa Zhong
Hong Kong University of Science and Technology
Altruism is fundamental to human societies, fostering cooperation and social cohesion. Recent studies suggest that large language models (LLMs) can display human-like prosocial behavior, but the internal computations that produce such behavior remain poorly understood. We investigate the mechanisms underlying LLM altruism using sparse autoencoders (SAEs). In a standard Dictator Game, minimal-pair prompts that differ only in social stance (generous versus selfish) induce large, economically meaningful shifts in allocations. Leveraging this contrast, we identify a set of SAE features (0.024% of all features across the model’s layers) whose activations are strongly associated with the behavioral shift. To interpret these features, we examine their activation profiles on benchmark tasks that exemplify either heuristic (System 1) or deliberative (System 2) processing. Causal interventions validate their functional role: activation patching in this feature direction reliably shifts allocation distributions, with System 2 features generally exerting a more proximal influence than System 1 features. The same steering direction generalizes across multiple social-preference games. Together, these results enhance our understanding of artificial cognition by translating altruistic behaviors into identifiable network states and provide a framework for aligning LLM behavior with human values, thereby informing more transparent and value-aligned deployment.
Training AI with Economic Axioms
Shuhuai Zhang1, Songfa Zhong2, Tracy Xiao Liu3
1 Central University of Finance and Economics
2 Hong Kong University of Science and Technology
3 Tsinghua University
This paper investigates whether axioms in economic theory can guide AI agents toward more consistent decision-making. We instruct a model to generate choices under varying budget constraints and apply revealed preference theory to automatically evaluate their internal consistency. Using these axiom-derived signals instead of human labels, we use the agent’s own output to fine-tune the model. We show that the resulting agent exhibits substantially higher choice consistency, with improvements that generalize well beyond the original training environment. To validate this approach, we also apply the procedure to a simulated supermarket environment calibrated with real scanner-data prices and household grocery budgets and find that fine-tuning yields highly consistent monthly category allocations. Our findings demonstrate that economic axioms, designed to provide normative benchmarks for human choices, can also serve as a powerful feedback mechanism to improve AI decision quality.
Deduplicating Event Reports in Multi-Source LLM Pipelines: Tradeoffs, Validation, and Lessons from Jordan
Elizabeth Parker-Magyar2, Mohamed Sabry Amer1
1 Washington University in St. Louis
2 Yale University
Multi-source event datasets are increasingly recognized as essential for studying contentious politics, yet combining reports from multiple outlets introduces the problem of duplicate event records. We present an iterative LLM-based clustering method for identifying and merging duplicate events across sources. We formalize the tradeoff between over-clustering (erroneously merging distinct events) and under-clustering (failing to merge duplicates), and argue that minimizing over-clustering should be prioritized because under-clustering errors are more easily detected and corrected by downstream users. Validating on one year of data from a new Jordanian events dataset constructed from fourteen Arabic-language sources, we achieve 100% accuracy in avoiding over-clustering and approximately 93% accuracy in detecting under-clustering. We also propose rules for resolving conflicting information across sources. Our method addresses a critical gap in the growing literature on LLM-assisted data construction for political science.
Zoned for the Commute that Disappeared: Measuring Regulatory Frictions to Remote-Work Adaptation with Large Language Models
CHEN Zhanghao1, CHEN Luoye2
1 The Hong Kong University of Science and Technology (Guangzhou), Innovation, Policy and Entrepreneurship Thrust, Society Hub, Guangzhou, China
2 The Hong Kong University of Science and Technology (Guangzhou), Carbon Neutrality and Climate Change Thrust, Society Hub, Guangzhou, China
Work from home (WFH) has become a durable feature of the post-pandemic labor market, reorganizing spatial relationships among residences, workplaces, and neighborhoods. A large literature has documented the demand side of this shift—which jobs can be done at home, who adopts remote work, and how commuting and housing markets respond. Yet less attention has gone to the institutional layer that determines whether a neighborhood can legally absorb the resulting work activity: local land-use regulation, or zoning. Organized around the strict separation of home and work, Euclidean zoning may itself act as a friction that constrains local adaptation to an integrated live-work paradigm.

This paper asks how far local zoning codes enable or obstruct the spatial absorption of remote-work demand. When local regulation cannot accommodate this shift, excess demand may be capitalized into housing prices, expressed as income sorting, or realized as remote work below its structural potential. Measuring it has been difficult because zoning ordinances are fragmented across thousands of jurisdictions and written as unstructured legal text. But large language models now make them tractable at scale.

We assemble what is, to our knowledge, the largest national corpus of U.S. municipal zoning ordinances—8,225 jurisdictions, drawn from American Legal Publishing, Municode, and ordinance.com—and build a customized retrieval-augmented generation (RAG) pipeline that codes each jurisdiction's provisions directly from the legal text. The design grounds every coded value in a specific ordinance passage, preserving an auditable trail from measure to source. Validation confirms the measures' reliability: cross-model substitution (Qwen3.6, DeepSeek-V3.2, Claude Haiku 4.5) yields 92% majority agreement with the primary codes, and the extracted measures converge with external benchmarks (National Zoning Atlas, WRLURI).

Linking these measures to occupational WFH feasibility and 2019–2024 American Community Survey outcomes, we construct a WFH–zoning adaptation index that separates remote-work demand from local capacity to absorb it. Three findings emerge. First, pre-pandemic compositional structure—education, occupation, and industry mix—explains most cross-place WFH growth. Second, rules that directly govern home-based work are only weakly related to home-based activity, whereas indirect margins—accessory dwelling unit (ADU) and mixed-use allowances—predict both WFH growth and home-value capitalization, indicating genuine supply constraints rather than latent demand. Third, aligning demand against absorption identifies 528 reform-priority jurisdictions—concentrated in metros such as Boston, San Jose, and the Bay Area—where demand outstrips zoning capacity; by-right ADU reform alone closes roughly three-quarters of the gap. Together, these results trace a coherent mechanism: where zoning cannot absorb WFH demand, the shortfall surfaces as housing-market pressure and suppressed remote work, so the index locates not where demand is high but where institutional capacity binds.

The study advances the measurement literature along four dimensions. Empirically, it assembles the largest national corpus of municipal zoning text to date. Conceptually, it measures two margins existing indices have neglected: micro-level home-occupation allowances and neighborhood-scale commercial mixing. Methodologically, it demonstrates a transparent, replicable LLM pipeline for extracting theory-driven institutional measures from primary legal text at national scale. Practically, the resulting index supplies metro-specific evidence for zoning reform, pinpointing where remote-work demand meets local regulatory rigidity.
Generative AI in Policymaking: Bias, Limitations, and Implications for Policy Change
Wilson Wong1, James Wong2, Hazel Kong2, Angela Mui2
1 The Chinese University of Hong Kong
2 Hong Kong University of Science and Technology

Artificial intelligence is increasingly reshaping public administration and policymaking. Across the policy cycle, generative AI may support agenda setting, policy formulation, decision-making, implementation, and evaluation by processing large volumes of information, identifying emerging issues, comparing alternatives, generating policy arguments, and improving communication. These capabilities suggest that AI may make policymaking more evidence-based, adaptive, and efficient. However, the growing use of generative AI in public policy also raises significant normative, institutional, and practical concerns.

This article examines the problems, biases, and limitations associated with using generative AI in policymaking and considers their implications for policy change. It asks three main questions: how existing scholarship characterizes the risks of AI and generative AI in public administration; how these risks are expected to appear in a real-world policy simulation; and what safeguards are necessary to ensure that AI supports rather than weakens policymaking.

The study adopts a two-part research design. First, it conducts a PRISMA-informed systematic literature review of scholarship and policy-oriented literature on AI in public administration, AI governance, ethics, social equity, public service delivery, and AI-supported policy development. The review will synthesize recurring concerns, including data and algorithmic bias, hallucination, misinformation, opacity, weak explainability, accountability gaps, erosion of human discretion, social equity implications, and ethical governance mechanisms.

Second, the study conducts an empirical policy-simulation analysis using data from the HKUST Inter-University Case Analysis x AI Competition 2025. In this competition, student teams analyzed the case of “Pet-Friendly Transportation in Hong Kong,” assessing whether pet-friendly public transport policies are desirable and feasible. Teams produced digital posters and five-minute “behind-the-scenes” videos reflecting on their AI use. These materials offer an opportunity to observe how novice policy analysts use generative AI in practice, how they understand its limitations, and how they verify, question, or rely on AI-generated information.

As the project is ongoing, the article presents expected rather than final findings. It expects to find that generative AI offers clear benefits in early-stage policy analysis by helping policymakers summarize information, structure arguments, identify stakeholders, compare jurisdictions, and improve written presentation. At the same time, AI-generated outputs are expected to create risks of hallucination, unreliable evidence, fabricated examples, and overgeneralized policy comparisons. The study also expects that policymakers may be more attentive to obvious factual errors than to deeper problems such as embedded normative assumptions, algorithmic bias, shallow reasoning, and policy ideas that appear comprehensive but lack causal logic, institutional realism, or ethical justification.

The study further identifies algorithmic monoculture, or systemic policy homogenization, as an emerging risk. When governments, agencies, or analysts rely on similar AI systems, they may produce increasingly similar policy ideas and recommendations, stifling innovation and erasing local or minority perspectives. The article argues that responsible AI-assisted policymaking requires AI literacy, source verification, disclosure of AI use, documentation of prompts and outputs, human-in-the-loop review, deliberative testing, oral defense, and mechanisms for questioning assumptions. It concludes that generative AI can support policymaking only when it augments, rather than replaces, human judgment, contextual expertise, ethical reflection, and democratic accountability.

Addressing the “Issue of the Age”: The Promise of LLMs for (Semi-)Automated Evaluation of Public Institutions’ Capacity Around the World
Tanu Kumar1, Robert Lipinski1, 2, Viktoriia Poltoratskaia1, Rita Ramalho1
1 World Bank Group
2 University of Oxford

“State capacity is the issue of the age” declared The Economist in January 2026. As public policy implementation and service delivery falter along multiple margins, economic development and trust in government suffer as a consequence. Therefore, understanding how governments can strengthen their ability to implement policy - what is termed here public institutions' capacity (PIC) – should rank among the top priorities of development practitioners.

In the World Bank’s Public Institutional Capacity and Effectiveness Unit we tackle this challenge by collating academic evidence and recent advances in large language models (LLMs) to build a largely automated approach to evaluating state capacity. It is built around the Country-Level Institutional Assessment and Review (CLIAR) framework, starting with Human Resources Management and Public Financial Management dimensions. We construct sector-specific questionnaires that serve as an input into a multi-phase LLM pipeline built on GPT-5.5. First, in the agentic retrieval phase, the pipeline scours the web to construct country-specific, hierarchical corpora of relevant legislation, government documents, institutional reports and comparable authoritative sources. Second, the retrieval-augmented generation (RAG) phase evaluates each of the questions against the constructed corpora, using Hypothetical Document Embeddings (HyDE) and ensemble agents that vote on the final response from the questionnaires’ fixed answer choices. In case of missing evidence or ensemble agents’ disagreement, the pipeline falls back to live web search beyond the pre-constructed corpora. The resulting data assess both the relevant formal public sector regulations (‘de jure’ dimension) and the extent to which they are being followed (‘de facto’ dimension).

We validate the AI-curated PIC indicators dataset relying on human coders and consultations with senior officials from a set of 16 pilot countries. Upon fine-tuning the model based on the ground-truth values obtained from the human respondents, the AI pipeline is planned to be deployed to measure the PIC indicators across the full set of the member states of the World Bank Group. In this fashion, it becomes possible to vastly scale up the existing assessments of institutional capacity, both along the geographic and temporal margins. Combined with a modular, transparent, and replicable design, the PIC AI pipeline offers a promising solution for data curation going forward.
Social Identity and Human-AI Task Allocation
Yiting Chen1, You Shan2, Shuangyu Yang3
1 Department of Economics, Lingnan University
2 Faculty of Business for Science and Technology, School of Management, University of Science and Technology of China
3 Institute for Economic and Social Research, Jinan University
We extend social identity theory to study human-AI trade-offs in hiring contexts. In a baseline experiment, representative U.S. participants allocate tasks between a human worker and either another human, ChatGPT, or the foreign model DeepSeek, with explicitly varied productivity. On average, people incur costs to under-allocate tasks to AI, particularly DeepSeek. Two additional experiments, inducing identities on workers via minimal-group and political affiliations, reveal a hierarchy: in-group humans receive the most, followed comparably by out-group humans and in-group AI, and out-group AI receives the fewest. Individual perceived social distance to AI and groupy tendency jointly shape the task allocation decisions.
Improving Communication with Generative Artificial Intelligence
David Hagmann1, Kirsten Geng1, Catherine Tinsley2
1 The Hong Kong University of Science and Technology
2 Georgetown University
Open-ended feedback is among the most common forms of consequential communication in

organizations, yet the messages people write are often sparse and generic even when they

hold detailed private evaluations. We propose that this reflects a production

constraint---turning a private judgment into an informative and socially appropriate

message is effortful---and test whether generative AI can relax it. Rather than composing

on the writer's behalf, the AI asks targeted follow-up questions and assembles the

writer's own answers into a revised message. Across three preregistered experiments

(N = 4,011), AI assistance made messages longer, more concrete, and substantially more

useful and credible to independent expert evaluators, for both course feedback (Study 1)

and feedback to a difficult coworker (Study 2). A rewrite-only condition produced about a

tenth of the gain: the value lies in the questions the AI asks, not the polish (Study 2).

In an incentivized hiring experiment, AI-assisted self-promotion conferred a large

advantage on whichever candidate adopted it---an advantage that dissipated when both did,

even as participants remained willing to pay for the assistance (Study 3).
Persuasion and Precision: How Generative AI Moves Inflation Expectations
Yang Lu, David Hagmann
The Hong Kong University of Science and Technology
Household inflation expectations have become an explicit target of monetary policy, yet

they respond weakly to official communication and sit persistently above professional

forecasts. As people increasingly consult generative-AI chatbots about the economy, we

examine whether such conversations move inflation expectations and what makes them

persuasive. Across three preregistered experiments (N = 3,772), participants forecast

U.S. inflation over the next twelve months, discuss their forecast with a partner, and can

then revise it. An AI partner moves forecasts substantially more than a human partner,

whether or not participants can converse with it (Study 1). Randomizing five features of

the chatbot's conversational style independently, we find that challenging, rather than

affirming, the participant's view drives persuasion (Study 2). A chatbot that combines the

most persuasive features outperforms a generic prompt and leaves participants more confident in

their revised beliefs (Study 3). Conversational AI can thus move expectations that official

communication has struggled to reach.
Algorithmic Bias and Social Cohesion: An Asabiyyah-Based Simulation Study
Natasha A. Henry1, Amar Ahmad1, Hayfa AbdulJaber2
1 Public Health Research Center, Research Institute, New York University Abu Dhabi
2 Arts and Humanities, New York University Abu Dhabi
Abstract
Introduction: As institutions increasingly hand high-stakes decisions to algorithms—who receives a scholarship, a loan, or a job—they also cede a share of the public trust on which cohesive communities depend. This study asks whether artificial intelligence strengthens or weakens the institutional standing and collective belonging that allow communities to function, using scholarship allocation as a case study where algorithmic decisions shape who is included, recognized, and treated as deserving.

Background: Drawing on Ibn Khaldun's concept of asabiyyah (group solidarity), we define social cohesion as the bond between human dependence and the institutions that organize shared needs into collective life—sustained by leadership that fulfills its obligations and preserves the conditions under which solidarity endures, especially in times of crisis. Algorithmic bias matters here because it can disrupt those conditions, altering how citizens interpret the fairness of public decisions.

Methods: To make this mechanism explicit, we present a Monte Carlo simulation modelling students who compete for a limited number of scholarships (the top 20% of applicants) allocated under two conditions: a fair system scoring applicants on ability alone, and a biased system that systematically penalizes one group. We translate disparities in scholarship allocation into a composite social cohesion index based on trust, participation, and cooperation.

Results: Across 1,000 repetitions, the fair system produced a small fairness gap (0.021) and higher cohesion (73.1), while the biased system produced a larger gap (0.191) and lower cohesion (57.9).

Discussion: Because the fairness–cohesion relationship is assumed rather than estimated, the model's value lies in exposing its assumptions to critique and generating a testable hypothesis: widening algorithmic gaps predict declining cohesion.

Conclusion: We propose to evaluate this empirically through surveys, administrative records, platform-level data, and interviews, contributing a conceptual and empirical roadmap for understanding asabiyyah in the age of AI.

Limitation: A limitation is that the relationship between algorithmic fairness and social cohesion is specified within the simulation rather than empirically estimated.
Fine-Tuned LLMs Forecast Survey Experiment Effect Sizes
David Broska
Department of Sociology, Stanford Unversity

Pilot studies help researchers identify promising treatments and estimate effect sizes before conducting full experiments. Yet piloting is often too slow or costly for many labs. Large language models could lower that cost. Prompted frontier models already forecast treatment effects that correlate strongly with observed ones (r = .85; Ashokkumar et al. 2026), enough to rank candidate treatments, but they overestimate effect magnitudes and compress outcome variance, which rules out uses that depend on calibrated forecasts. We test whether fine-tuning closes this calibration gap.

We test whether fine-tuning an LLM on past survey experiments can provide useful, low-cost forecasts of treatment effects. We built an LLM-based simulator that reconstructs the Qualtrics survey each participant saw, including treatments, questions, and prior answers, then fine-tunes an open-weight LLM to predict responses sequentially. Training used 80 text-based between-subjects survey experiments with 56,177 participants. We evaluated the model on 13 held-out studies by walking it through each survey and comparing simulated with observed effects for every post-treatment outcome. Across 515 effects, simulated and observed Cohen’s d estimates were highly correlated (r = .94) and closely aligned in magnitude; among statistically significant human-sample effects, the simulator recovered the observed direction in 95% of cases. These results suggest that fine-tuned LLM simulations can support affordable piloting, power analysis, and prioritization of interventions before full human data collection.

The next project phase tests whether the approach generalizes beyond a single lab. The next step is to expand the training and test archive with datasets from other researchers. For training, this will let the model learn from a broader range of experimental designs, populations, and outcomes it has not yet seen. For testing, private data play an even more distinctive role. Because modern LLMs train on the public internet and retrieve published results through web search, experiments that never appeared online are the cleanest test of whether LLMs forecast results rather than remember them. We are therefore inviting a limited group of 30 social scientists to contribute study materials and response data from survey experiments. The talk presents the fine-tuning approach, the proof-of-concept results, and the research design of this team science project.
Policy Content, Distributional Conflict, and Contestation in the Council of the European Union
Arash Pourebrahimi
Leiden University
Vilnius University
This project examines whether the substantive content of legislation shapes contestation in the Council of the European Union. Although the Council is often described as a consensus-oriented institution, member states regularly contest legislative decisions through votes against, abstentions, or critical statements. Existing research has mainly explained this behaviour through member-state preferences, domestic politics, public opinion, and institutional factors. This project shifts attention to the content of the legislation itself and asks whether proposals involving the distribution of resources are more likely to generate contestation than more technical or administrative proposals.

The central argument is that politicisation does not affect all areas of EU decision-making in the same way. Legislative proposals that allocate financial resources, create uneven costs and benefits, or affect access to rights, markets, or social protections may generate clearer political stakes for member states. These proposals are more likely to create winners and losers and may therefore give governments stronger incentives to publicly signal disagreement. By contrast, proposals that are mainly technical, procedural, or administrative may be less likely to attract visible contestation.

To examine this argument, the project develops an LLM-assisted measure of legislative content. Instead of relying only on broad policy-area classifications or topic models, it uses large language models to code legislative proposals according to the extent to which they involve distributive consequences, technical regulation, implementation burdens, and uneven implications across member states. These measures will be linked to member-state voting behaviour in the Council, focusing on legislative decisions adopted under qualified majority voting.

The project contributes to the study of EU decision-making in two ways. Substantively, it develops a more direct account of how policy content shapes contestation by distinguishing between distributive and technical legislation. Methodologically, it explores how LLMs can be used to extract theoretically meaningful features from legislative texts in a transparent and validated way. The project is ongoing, and the conference paper will present the coding strategy, validation procedure, and preliminary evidence on the relationship between distributive legislation and contestation in the Council.
Asymmetric Exposure to Climate Change Risks and Corporate Self-Regulation
Volkan Tibet Gur
Rutgers University, New Brunswick
symmetric Exposure to Climate Change Risks and Corporate Self-RegulationAs regulatory climate transition risks increase, firms increasingly use self-regulation as a political strategy to preempt demand for stringent government oversight on climate change. I argue that firm's exposure to regulatory climate transition risk are asymmetric and varies with national origin, prior involvement in scandals and sector. Yet, counterintuitively, more exposed firms drive greater benefits from climate-self regulation and thus more likely to make public climate commitments. I test this argument using online survey experiments with American public and a novel dataset of publicly traded domestic and foreign firms registered with the U.S. Securities and Exchange Commission from 2001 to 2024. I use ClimateBERT to classify nearly 85 million paragraphs from 10-K and 20-F filings and measure the specificity of firms’ climate commitments. I find that firms facing stronger public demand for climate regulation make more climate commitments, but whether those commitments are specific depends on the type of climate liability they face. These findings shed new insights to political economy of firm's political strategies in the age of climate change.
Measuring Public Participation in Local Government with Audio-Based Speaker Classification
Menglin Liu1, J. Sophia Wang2
1 The Chinese University of Hong Kong, Shenzhen
2 Yale University
3 University of California, Davis
Democratic legitimacy rests not only on periodic elections but on ordinary residents having an ongoing voice in the local decisions

that affect them—and the public-comment period of city council meetings is where that voice is most directly exercised. Yet we still lack scalable, direct measures of who actually speaks in these meetings—and for how long—which has limited empirical research on grass-root democracy. We show that recent advances in audio processing and multimodal large language models make such measurement possible at scale. To our knowledge, this is the first study to measure public participation in local government directly from meeting audio. We develop and validate an audio-based pipeline that operates without transcripts or external identity records. The pipeline first applies speaker diarization to determine who speaks when. It then uses a multimodal large language model to classify each detected speaker as an official or a public participant and calculates measures including public floor-time share. We apply this approach to approximately 2,600 hours of audio collected from 1,565 meetings held in five U.S. cities between 2017 and 2023. Against human annotations, diarization achieves 89.4% accuracy, while speaker-role classification achieves 87.7% accuracy and a public-participant F1 score of 0.832 on 1,761 speakers. The resulting measures reveal that public participants constitute 29.9% of detected speakers but receive a median floor-time share of only 10.0%. Public floor-time share also decreased from 15.5% before 2020 to 4.2% in 2020–2021 and recovered only partially to 9.6% in 2022–2023. These

findings extend previous research showing inequalities in who at- tends local meetings: even when public participants attend, their presence does not translate into proportional voice. Beyond this study, the approach can be scaled to additional cities and adapted to examine finer-grained roles, topics, interactions, and institutional features of public deliberation.
From Reports to Data: Harnessing LLMs for Fine-Grained Human Rights Data Collection
Yusuf Evirgen
Bilkent University
Human rights measurement has long been constrained by a familiar trade-off: existing indicators are either broad and infrequent, or fine-grained but painstakingly hand-coded and difficult to replicate. This paper shows how large language models can break that trade-off. I develop an LLM-based extraction pipeline that converts unstructured reports from local Turkish human rights NGOs into a structured, daily-level panel, yielding roughly 30,000 unique violation records spanning 2013 to 2023, at a temporal and geographic resolution existing cross-national datasets cannot match. Beyond the Turkey case, the contribution is methodological: a scalable, replicable framework for turning NGO reporting, a vast and underused textual archive, into systematic human rights data, with direct extensions to other countries and rights domains. The paper reports validation results, discusses failure modes and bias checks for LLM-based coding of sensitive political content, and situates the approach within the conference's broader agenda on automated data curation and LLM applications in policy research.
Policy Prototyping with Agentic AI: A Computational Sandbox Approach
Gleb Papyshev1, Keith Jin Deng Chan2
1 Lingnan University
2 The Hong Kong University of Science and Technology
Policymakers often design regulatory incentives without evidence on how they will shift behavior. We introduce agentic AI policy prototyping, a method that configures a large language model as a goal-directed strategic agent inside a stylized regulatory environment. The LLM receives descriptions of policy parameters and makes a discrete choice, allowing analysts to stress-test incentive designs rapidly and at low cost before real-world implementation. We demonstrate the method on an AI governance dilemma: under what conditions will a competitive developer invest in ethical features of its system? Calibrating parameters derived from game theoretic model with governance indicators from 79 countries, we find that recognition probability dominates other levers, that rewards are ineffective without credible monitoring, and that the LLM's choices align with a theoretical benchmark in approximately 75% of cases. The contribution is primarily methodological: a transparent and replicable sandbox that bridges formal theory and costly field pilots, applicable across policy domains.
How Academia, News, and the Public Make Moral Sense of Generative AI
Yifei Wang2, Zening Duan1
1 National University of Singapore
2 University of California Santa Barbara
How do societies make moral sense of generative AI during its public emergence? As large language models become embedded in everyday life, public discourse plays a critical role in shaping how these technologies are interpreted, contested, and governed. Drawing on Moral Foundations Theory, this study analyzes 2.7 million posts about generative AI on X (formerly Twitter) following the release of ChatGPT in late 2022. We examine how different social actors construct moral narratives around generative AI, how these narratives evolve over time, and whether they influence the visibility of AI-related discourse. Results show that moral language becomes less volatile over time, suggesting that initial moral contestation gives way to more stable interpretations of generative AI. Moral expression also differs systematically across actors: academic accounts employ more restrained and positive moral language, whereas news accounts frame generative AI more negatively. Finally, engagement analyses reveal a conditional moral virtue penalty, with positive moral language associated with lower sharing overall, while the relationship between moral intensity and engagement varies across elite actors. Together, these findings suggest that public discourse constitutes an important mechanism through which societies negotiate the meaning, legitimacy, and governance of generative AI, shaping which interpretations of emerging AI technologies gain prominence in online information ecosystems.
Controlling LLM Agent Personality and Preferences Through Representation Engineering
Zeyu Lyu, Zhichao Wang
Graduate School of Arts and Letters, Tohoku University
Large language model (LLM) agents offer a flexible approach to constructing "silicon samples,'' with considerable potential for social science research. However, their behavior is often sensitive to prompt wording, difficult to interpret, and biased toward model-default preferences. We investigate activation steering as a means of controlling agent personalities and preferences by intervening in the model’s internal representations. Using Llama-3.1-8B-Instruct, we construct steering vectors and apply them during inference at varying scalar coefficients. We illustrate this approach through two proof-of-concept studies. First, we manipulate a latent personality trait and show that steering can systematically shift agents’ behavior across multiple behavioral games. Second, we manipulate a model-default preference and show that steering can alter both individual choices and collective norm trajectories. Our findings highlight the potential of representation engineering to enhance the interpretability and reproducibility of LLM-based social science research.
Testing Rationality in LLMs: Responsibility Attribution in the Absence of Control
Yueying Chu1, 2, Jiaxin Zhang1, 2, Peng Liu1
1 Center for Psychological Sciences, Zhejiang University, Hangzhou, Zhejiang, China
2 Department of Psychology and Behavioral Sciences, Zhejiang University, Hangzhou, Zhejiang, China
Large language models (LLMs) may assist people in morally and legally relevant domains. This study explores whether LLMs, like humans, show bias in responsibility attribution through a social question: should users of fully driverless cars be held responsible for accidents involving these cars? According to the control doctrine in ethics and law, individuals can only be responsible for actions over which they have control. However, recent research found a counter-intuitive bias: human participants attributed more responsibility to users of driverless cars (owners of private driverless cars and passengers in robotaxis) than to passengers in conventional taxis—despite all lacking control over the cars. We tested three LLMs (GPT-3.5, GPT-4, and GPT-4o) using the same experimental design across three studies (two preregistered). Compared to human participants and GPT-3.5, GPT-4 and GPT-4o showed more rational responses, assigning limited responsibility to robotaxi passengers but still attributing some responsibility to owners of private driverless cars. Interestingly, GPT-4 and GPT-4o exhibited a non-human-like bias: assigning more responsibility to conventional taxi passengers than to robotaxi passengers in two studies. These findings suggest GPT-4 and GPT-4o largely obtain normative rationality in responsibility attribution and offer insights into potential differences in moral psychology between humans and LLMs.
From Image-Based Documents to Intelligent Research Tools: A Localized Collaborative Research Framework Integrating Knowledge Bases and Large Language Models for Dunhuang Studies
Hongjuan Zhao, Ruohan Ma
Qingdao University of Science and Technology
The digital transformation of Dunhuang Studies has significantly improved access to manuscripts, paintings, catalogues, and related research materials. However, resource visibility does not automatically translate into research usability. Although large-scale digitization projects have substantially alleviated the difficulties caused by the transnational dispersal of Dunhuang materials, core scholarly tasks—such as thematic aggregation, source tracing, version comparison, historical contextualization, and cross-textual knowledge association—have not been resolved simply through the expansion of digital resources. Against this background, this article argues that the next stage of digital Dunhuang research should not be defined merely by the expansion of image repositories or by more convenient catalogue retrieval, but by the construction of research-oriented knowledge infrastructures that are both computable and verifiable.

To this end, the article proposes a localized collaborative framework for Dunhuang Studies that integrates image-document parsing, structured text extraction, knowledge organization, retrieval-augmented interaction, and human verification. Grounded in the disciplinary characteristics of Dunhuang Studies—namely dispersed materials, multilingual coexistence, unstable image quality, complex page layouts, and strong reliance on provenance awareness and philological control—this framework takes a localized knowledge base as its organizational core and employs a large language model as an auxiliary reasoning tool. In doing so, it reconstructs a workflow that moves from image-based documents to structured knowledge units and then to evidence-driven scholarly interaction. The article emphasizes that artificial intelligence in Dunhuang Studies should not be understood as a substitute for scholarly judgment, but rather as a form of research infrastructure whose main functions are to reduce repetitive labor, reorganize dispersed evidence, support exploratory and verification-oriented inquiry, and strengthen the evidentiary basis of historical interpretation.

This article further argues that the significance of local deployment in Dunhuang Studies lies not only in technical feasibility and data security, but also in safeguarding interpretive authority, maintaining continuity in annotation work, and ensuring transparency in evidentiary chains. Likewise, the core value of retrieval-augmented interaction does not lie in directly generating answers, but in the structured reorganization of evidence across manuscripts, catalogues, and prior scholarship. In this sense, the incorporation of intelligent tools into Dunhuang research should not be interpreted as the automation of humanistic scholarship, but as a reorganization of the conditions of knowledge production. The discussion suggests that the transition in Dunhuang Studies from resource visibility to knowledge usability requires not only the continued accumulation of digital resources, but also the development of intelligent research infrastructures compatible with the methodological requirements of the field.
Research on Platform Governance and Government Regulatory Strategies Considering the Risk of Missed Detection of AIGC Identifiers
Zhong Wang, Tiantian Zhao
Guangdong University of Technology
With the rapid advancement of generative artificial intelligence (GenAI) technologies, AI-generated content (AIGC) has been increasingly integrated into online content creation and dissemination. Nevertheless, the lack of proper labeling for AIGC has exacerbated governance challenges by facilitating the propagation of misinformation and complicating the verification of content authenticity.Many countries have introduced AIGC identification systems to regulate the development of AIGC, but there is still a lack of research on how to effectively implement them.This paper constructs a tripartite evolutionary game model involving content publishers, platforms, and the government, and systematically analyses the strategic interactions among the three parties through numerical simulation.This study analyzes the AIGC identification governance mechanism from the perspective of subject behavior interaction, providing a reference for the governance of the platform ecosystem.
Homo Silicus and the Rationality Gradient: Reasoning Compute, Expectation Formation, and Macroeconomic Dynamics
Jianhao Lin, Xiangdong Wang, Yifan Zhang
Sun Yat-sen University

How does the degree of bounded rationality in market expectations shape macroeconomic dynamics? We answer this question by deploying large language models as boundedly-rational forecasters, varying only the compute they allocate to reasoning. When this compute is set to zero, the model answers without deliberation and its forecasts mirror human subjects. As more compute is allocated, the distribution of forecasts shifts toward higher rationality. This mapping from reasoning compute to forecast rationality emerges in both univariate and New Keynesian learning-to-forecast experiments and holds across four model families. As rationality rises, the economy becomes more self-stabilizing, although the policy multiplier correspondingly declines. The reasoning-compute channel thus offers a unified, non-parametric boundedly-rational expectations generator that displaces the menu of pre-chosen cognitive mechanisms the literature has relied on.

Detecting Absence in Large Language Models: Missing Precedents and Omitted Provisions in Legal Text
Lisa Lechner
University of Innsbruck
Large language models are increasingly used to summarize, annotate, and reason over legal texts — treaties, statutes, and court decisions — and courts and counsel are already being sanctioned for relying on LLM-fabricated case citations that do not exist. This paper traces that failure to a general and under-examined limitation: models are far better at detecting what is present than at registering what is absent. We evaluate existing absence-detection and uncertainty-estimation tools on legal corpora and argue that absence-sensitivity is a validity precondition for using LLMs as instruments in legal and political research.

We distinguish two registers of absence. Sampling absence is a gap in the training distribution; recent work (AbsenceBench; Fu et al. 2025) shows that models fail to detect even conspicuous surface omissions, because transformer attention has no gap to anchor on, and instead fill them with confident fabrication (Kalai et al. 2025). Produced absence is a gap that was made — a provision negotiated away, a precedent left uncited, a dissent unaddressed — which legal and political theory treats as an exercise of power rather than a neutral void (Bachrach and Baratz 1962; Lukes 2005; Trouillot 1995; Fricker 2007).

Legal texts are an ideal and feasible testbed: they are structured and versioned, so clean original-versus-omitted pairs can be constructed, and their meaning often lies precisely where they deviate from templated defaults. Using existing tools — the AbsenceBench protocol and semantic-entropy uncertainty estimation (Farquhar et al. 2024) across several open-weight models — we test, on constitutional court decisions, whether models detect a removed precedent or instead reconstruct a citation that was never there; and, on tax treaties built from the OECD and UN Model Conventions, whether they notice a modified article or default to the standard clause the treaty in fact negotiated away.

We test the hypothesis that fabrication concentrates precisely on the deviations that carry legal meaning — so an LLM used to read case law or treaties would systematically misreport the precedents and provisions that matter most. Absence-sensitivity should therefore be a standard axis of LLM evaluation for legal and social research.
Do Personas that Sound Right Also Choose Right?
Quang Phuc Phung
MIT Sloan
Social scientists increasingly deploy large language models as synthetic participants, using a short demographic persona prompt to answer surveys or play games. Yet persona prompting faces open questions: we do not know whether a prompt summons the intended person, whether persona descriptions and persona choices align, or whether a weak persona prompt can be strengthened. I investigate these questions through a proposed design and a small pilot. Drawing on the Twin-2K-500 dataset, I construct three persona prompts for each respondent: a minimal demographic card, the same card enriched with the person's own narratives, and the same card enriched with narratives from a matched stranger. Each persona completes a scripted interview about its life, ranked by a blind LLM judge, and answers a battery of decision items, both directly and after re-reading its own interview, scored against the person's own responses. In a pilot with ten respondents, the judge tends to prefer the minimal card's interviews, yet conversational quality shows no clear connection to later choices. Letting a persona re-read its own interview before deciding appears to help only when the interview is built from the person's own material. The full study will examine these patterns at scale.
Learning with Machines: A Randomized Evaluation of AI-Assisted Education in Rural Middle Schools in China
Sharon Xuejing Zuo1, Zhao Chen1, Elaine M Liu2, Kaiyue Li1, Yuelin Zhong1
1 Fudan University
2 Georgia State University
This study evaluates the effectiveness of a technology-based learning intervention in rural middle schools using a randomized controlled trial. 410 rural 9th grade students receives access to a structured learning tool designed to support 2-hour daily independent study for two months, while a comparison group (450 students) continues with their standard study routine. Both groups are assessed regularly throughout the study period. Our primary outcome of interest is performance on a standardized examination taken at the end of the study period. We find treated students perform better than control group students and the effect is statistically significant at 1% level.
Victimhood or Recognition? LLM-Assisted Coding of Azerbaijani Public Narratives on Nagorno-Karabakh Conflict
Zubaidiya Simayi
School of International Studies, Peking University
Political conflict discourse rarely expresses emotion as isolated sentiment. In protracted conflicts, affect is often organized through public narratives of suffering, loss, blame, justice, dignity, and recognition. Existing computational emotion analysis is useful for classifying broad categories such as anger, fear, sadness, hope, or neutrality, but it is less suited to capturing the political meanings through which communities define victimhood, assign responsibility, and set the conditions under which the other side may be acknowledged. This paper addresses this limitation by developing and testing a theory-guided framework for LLM-assisted coding of victimhood and recognition cues in Azerbaijani public narratives on Nagorno-Karabakh.

Empirically, the paper examines Azerbaijani official and state-aligned discourse from 2020 to 2023, covering the period from the Second Karabakh War to the post-2023 transformation of the conflict. The corpus is constructed from publicly available texts issued by or circulated through three controlled source types: presidential discourse, foreign ministry statements, and state news agency reporting. It compares Azerbaijani-language domestic-facing texts with English-language international-facing texts. Narratives are operationalized as paragraph-level public claims that connect events, actors, responsibility, and political meaning. This design does not claim to represent all conflict-party narratives; rather, it uses a bounded source universe to evaluate whether LLM-assisted coding can reliably identify affective narrative cues across audience orientations and text-processing conditions.

The coding scheme focuses on six categories: victimhood, blame attribution, non-recognition, conditional recognition, substantive recognition, and procedural or neutral language. The key conceptual distinction is between conditional and substantive recognition. In conflict discourse, references to peace, rights, security, or reintegration may appear conciliatory, but their political meaning depends on the conditions attached to such acknowledgment. Distinguishing conditional from substantive recognition helps avoid overstating the reconciliatory content of the text.

Methodologically, the paper compares three LLM-assisted coding pipelines: translate-then-code, original-language coding, and dual-pass coding using both original texts and translations where available. A human-coded validation set is used to assess label-level agreement, evidence-span validity, translation sensitivity, neutrality bias, recognition overcoding, and the possible normalization or erasure of politically meaningful terms and place names.

The paper contributes to computational social science and conflict research in two ways. First, it shifts LLM-based political emotion analysis from generic sentiment classification to theory-guided coding of narrative-affective cues. Second, it offers a feasible workflow for evaluating when LLMs can support, rather than replace, expert interpretation of politically sensitive texts. The broader argument is that LLMs are most useful for conflict discourse analysis when their outputs are constrained by explicit concepts, transparent coding rules, and systematic validation.
Can We Trust LLM-Generated Economic Variables? Design Dependence and Downstream Inference in Occupational AI Exposure
Wendu Zhuge1, Yuexin Wang2
1 Shanghai Normal University
2 Shanghai Normal University

Large language models are increasingly used as measurement instruments in the social sciences, yet the variables they generate are often analyzed as if they were directly observed. This paper studies design dependence: systematic variation induced by model identity, prompt protocol, construct boundary, and aggregation rule. Four LLMs classify 2,087 O*NET Detailed Work Activities under a common three-category rubric, while three independent experts score the full task universe. A pre-specified audit additionally evaluates 300 held-out tasks using four models, three prompt-protocol operationalizations, and two repetitions. Expert reliability is high overall, but is substantially greater for broad capability relevance than for direct completion. The LLMs show the same broad asymmetry: they largely agree on whether a task is related to LLM capability, but disagree more on whether it can be completed directly through a standard interface or requires complementary tools, data, or organizational integration. Disagreement is concentrated in mental-process and interpersonal tasks and remains consequential after aggregation into occupational and firm-level measures. A domain-heterogeneous latent-state model is used to represent ensemble disagreement and propagate measurement uncertainty to 923 occupations, 5,264 Chinese listed firms, and downstream regressions. Semi-synthetic experiments show improved latent-state recovery under partly independent errors, but weaker performance when models share common errors. The findings support a measurement workflow based on construct comparison, multi-model and multi-prompt auditing, expert reliability assessment, and uncertainty propagation.

GenAI Governance in School: How Guidelines and Enforcement Shape Student Performance
Yanlin Wan1, Xu Zhang2, Jinghao Jia2
1 HKUST
2 HKUST-Guangzhou

Schools have responded to Generative AI with two main strategies: teaching stu- dents to use it responsibly through instructional guidelines, and capping AI-generated content with enforcement. Yet there is little causal evidence on the effects of these poli- cies in real classrooms. We run a randomized controlled trial across two Chinese mid- dle schools (N = 1,056), cross-randomizing a six-module AI literacy curriculum with an enforcement intervention that imposes a 30% AI-content cap on writing assign- ments, detects violations through multi-engine AI screening with instructor verifica- tion, and mandates resubmission for flagged submissions. Neither component alone generates detectable academic gains, while the combination of guidelines and en- forcement improves writing performance, particularly Chinese writing. The writing gains are concentrated among male students. Compliance outcomes suggest gender- specific behavioral responses to the policy environment. Under guidelines alone, fe- male students reduce measured AI violations, whereas male students’ violation rates remain higher than in other conditions. Therefore, guidelines without enforcement, the most common policy responses to AI in schools, may backfire unless accompanied by a well-designed compliance mechanism.

Modeling Human Behavior in the Prisoner’s Dilemma with Large Language Models
Raina Gao, Tristan Raffo
Thomas Jefferson High School for Science and Technology

The Prisoner’s Dilemma has long served as a foundational framework for studying human decision-making and cooperative outcomes across numerous social scenarios, yet human behavior consistently deviates from predictions of fully rational economic theory. Previous works have extensively studied human and LLM outcomes in the Prisoner’s Dilemma; however, these studies have lacked at least one key feature: they either did not include a memory component for the LLM agents, did not use accessible open-source models, or did not compare models against empirical human data. This project investigates whether open-source LLMs, TII Falcon 7B Instruct, Google Gemma 7B, Meta Llama 3.1 8B Instruct, Microsoft Phi Mini MoE Instruct, and Qwen 2.5 7B Instruct, can accurately model human cooperation patterns in repeated Prisoner’s Dilemma games, while also examining how an LLM’s knowledge that its opponent will be shuffled affects its decisions. Using a standardized prompting framework to match established Prisoner’s Dilemma experiments, we simulated repeated interactions in two trial types: a fixed-partner setting where both agents retain memory, and a shuffled-partner setting where a memoryless agent interacts with a memory-enabled agent prompted with knowledge of an opponent-shuffling pattern. By testing both interactions, this study evaluates how varying agent environments affect cooperation rates, and broadly, decision making behaviors. For each trial, we measured cooperation rates over 100 rounds and 10-round increments and compared to baseline human data from Montero-Porras et al. (2022). Under both frameworks, all LLMs differed from human cooperation rate. Open source LLMs generally over-cooperated with respect to humans, particularly in the shuffled framework, where human cooperation rates averaged 28.5% while several models exceeded 80%. Qwen-2.5-7B-Instruct showed evidence of deterministic tendencies, while others, like Phi-mini-MoE-Instruct, displayed much more variability, and even the ability to approximate human cooperation rates briefly. These findings suggest that LLMs have potential as proxies for human decision-makers and provide a benchmark across two agent-interaction environments, although current open-source models are not yet reliable substitutes for human participants.

Topical Relevance or Strategic Gaming? Detecting Self-Citation Motives with Language Models Across Two Million Business and Management Articles
Niranjan Sapkota
School of Accounting and Finance, University of Vaasa, Finland

Is self-citation the ordinary accumulation of a research program or a form of strategic self-promotion? An author may cite earlier work because it is the natural foundation for what follows, a motive we term Need, or in order to raise the bibliometric indicators that weigh on hiring, promotion, and funding, a motive we term Greed. Because any single self-citation is consistent with either, which of the two governs ordinary practice has long remained unsettled. We take up the question with 2.04 million papers published between 2000 and 2025 by 1.64 million authors, drawn from OpenAlex and classified into 22 business and management fields by the Chartered ABS Academic Journal Guide, with the SJR and ABDC rankings retained for robustness. For every paper we assess how closely its content follows the author's earlier work, measuring the cosine similarity between the paper and the author's prior corpus with two sentence-transformer language models, MPNet (768-dimensional) and MiniLM (384-dimensional), across titles, abstracts, and concept hierarchies. We also separate first-degree self-citations, which an author makes directly, from second-degree ones, in which a co-author cites the team's earlier work. In nested Poisson regressions with author and field-by-year fixed effects, mechanical opportunity accounts for 98.2 percent of the explained variance in self-citation, whereas proximity to the h-index and i10 thresholds contributes less than 0.1 percent and, in all but two small fields, runs against the gaming interpretation. Topical relevance is the strongest behavioral predictor, and senior, highly cited authors, who would gain most from inflated metrics, self-cite the least relative to the standing they already hold. Self-citation is therefore predominantly need-driven, though a modest within-author rise since 2015 hints at a gradually shifting norm. Blanket penalties on self-citation would discourage legitimate scholarship in pursuit of a problem that few researchers have.

Simulating Opinion Dynamics with Generative Agent Based Models: The Emergence of Consensus and Ideological Clustering
Jan Lorenz1, Erkan Gunes2
1 Constructor University
2 Constructor University

People continually encounter opinions through news media and social networks. They decide what to accept, remember, reject, or share. Over time, these choices can reshape individual worldviews and generate collective patterns such as consensus, polarization, and filter bubbles. We develop an agent-based model that links these micro-level information choices to macro-level opinion dynamics. Rather than representing opinions as points in a low-dimensional numerical space, our approach uses natural-language opinion statements and LLM agents that can understand and act on them. Each agent has a bounded memory of statements representing its worldview. Agents encounter statements from a shared feed and network neighbors, decide whether to integrate or reject them, replace existing statements when memory is full, and share remembered statements with others. Unlike persona-based simulations, agents are not assigned demographic profiles or ideological labels. Their worldviews develop through repeated exposure, integration, replacement, and sharing. Preliminary results produce collective patterns, including convergence around common statements, concentration of attention on a narrow subset of the statement pool, and separation into rival worldview clusters. Early comparisons suggest that memory capacity may shape whether collective attention narrows around a few statements or preserves diversity for competing clusters. These findings connect individual information choices to emergent population level worldviews.

Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media
Raymond Duch1, Sorin Sorescu2, Boone Bowles2
1 University of Oxford
2 TAMU

Social media affect financial markets, but public posts by financial media personas are voluntary disclosures, hence undisclosed views are unobserved. We address this measurement problem by conducting repeated, real-time interviews of “digital twins'' built from finfluencers' X accounts. The interviews recover stock-level public-persona belief proxies even when no public recommendation is made. Because interview responses are generated before the return windows, the design avoids look-ahead bias. We show that digital-twin responses predict the cross section of large-cap stock returns in the expected direction. Repeated real-time interviews therefore show how selective disclosure can be turned into measurable panels of market views.Just created for emails

How Does Water Scarcity Escalate into Conflict? Exploring Hydrosocial Trajectories with an Agent-Based LLM Model
Mario Lillo-Saavedra1, Marcela Salgado2
1 Facultad de Ingeniería Agrícola, Universidad de Concepción
2 Facultad de Ciencias Ambientales, Universidad de Concepción

Water scarcity is not only a hydrological challenge but also a driver of social tension, institutional distrust, and competing responses among water users. Understanding how hydrosocial conflicts evolve under prolonged scarcity requires models that can represent heterogeneous actors, adaptive decision-making, and the emergence of cooperation and conflict. This study presents a Hydrosocial Agent-Based Model enhanced with Large Language Models (HSABM–LLM) to explore conflict trajectories and adaptive responses under sustained water scarcity. The model is parametrized using empirical data from 293 water-user surveys collected in the Longaví River subbasin, Chile, capturing heterogeneous user profiles, perceptions, institutional trust, and adaptation preferences.

The experimental design spans ten years, comprising a baseline phase followed by a prolonged hydrological perturbation designed to emulate conditions associated with a La Niña episode. Water availability is reduced to 60% of baseline conditions, while irrigation allocation declines to 65%. Four policy scenarios are compared: no intervention, short-term adaptive measures, long-term infrastructure investment, and a combined strategy. The model tracks hydrological, social, and discursive outcomes, including accumulated water deficit, institutional trust, cooperation, coalition formation, conflict intensity, and the evolution of narratives.

The study aims to identify how water scarcity reshapes hydrosocial conflict trajectories and whether combined adaptation strategies can strengthen the resilience of hydrosocial systems.
AI and Organizational Justice: An Experimental Study of HR Processes
Ali Farashah
Organization and Management Division Mälardalen University Sweden

A body of research is emerging in management and human resource management literature on the implications of artificial intelligence for organizational justice and for diversity, equity, and inclusion (DEI). Recruitment, as one primary focus area of this research, shows simultaneous potential for bias reduction through standardization and new risks of negative impact through proxy variables embedded in training data, as well as workers' and applicants' fairness perceptions related to the transparency and explainability of the process (Soleimani et al. 2025). Beyond hiring, evidence from algorithmic management in monitoring and compensation highlights distributive concerns (e.g., pay, task allocation, and scheduling outcomes), procedural concerns (opacity, appealability, data governance), and relational concerns (dehumanization and trust in AI-mediated interactions) in the use of AI in the workplace (Doan and Diehl 2025; Keegan and Meijerink 2025; Zhang et al. 2025).

Despite fast-growing interest in the impact of AI on work and employment, significant gaps remain. Organizational justice can be a central concept for designing "responsible AI" in HR and for tying AI outcomes to employees' lived experiences across different groups — a link that can advance DEI and justice rather than merely comply with accuracy or efficiency benchmarks. The construct of organizational justice encompasses three key dimensions: (a) procedural (the fairness of the methods used to reach a decision, including concepts such as transparency and voice), (b) distributive (the fairness of the outcome itself and outcome parity across intersections such as gender × ethnicity × age), and (c) relational (the fairness of how people treat each other, including respectful communication, dignity, and opportunities for interaction with accountable humans). Most current research on AI in the workplace focuses mainly on procedural aspects, such as transparency and consistency in algorithmic decision-making, but neglects how the opacity and automation of AI may erode relational trust and supervisor–employee dynamics. Furthermore, the impact of AI on the outcomes of marginalized workers (i.e., distributive justice) remains underexplored, particularly where intersecting identities are involved.

A holistic justice framework that integrates all three dimensions is essential for developing ethical AI governance in HRM. The purpose of this study is to examine how the implementation of AI-driven human resource management (HRM) systems affects employees' perceptions of organizational justice across its three dimensions. The research questions are:

  • How do AI design features in HR decision-making influence employees' and applicants' perceptions of justice?
  • To what extent do these perceptions differ across social groups (e.g., gender, age cohorts, people with a migratory background)?

Methodology

A series of vignette-based survey experiments will be conducted in which professionals evaluate realistic HR decisions made with AI assistance. A survey experiment combines the internal validity of randomized experiments with the external realism and scalability of surveys. Participants will be randomly assigned to read short scenarios ("vignettes") that vary specific features of an AI-enabled decision process and will then report their perceptions. Participants will be recruited through high-quality online survey panels.

Building the Data a Spatial Model of Voting Needs
Laurenz Guenther1, Raymond Duch2
1 TSE
2 University of Oxford

The spatial model of vote choice is the standard account of how Americans vote, and no dataset can fit it properly to an American election. The instruments that ask a voter to place himself and the candidates on the same policy scales are national and cover the presidency alone. The instrument with samples large enough for a single race carries one left-right number and places

one of the two people on the ballot. None of them asks a voter how much any issue weighs in his own vote, so the weights have to be read off the vote the model exists to explain. We build the missing data for one congressional race. A panel of 4,865 language-model respondents, each seeded on a real survey respondent’s recorded answers joined to a voter-file record and drawn so that the panel matches the district’s age, sex and county margins, answers a frozen interview of eleven question blocks. It asks where he stands on three scales and 32 issues, how much each of them matters to him, where he places both candidates on all of them, how he rates each man on competence and honesty with policy set aside, and, in a separate conversation, whether he will vote and for whom. That comes to 144 numeric answers before the vote is reached, which is why no human panel has been asked for it. One pass over the whole panel cost 142 dollars and returned 4,672 valid answers.

We report the costs, which rules answers fail and how often, and what the answers look like. The candidate placements and quality ratings from this pass carry a disclosed assembly defect and are not used. We release the interview, the source of every item, the component loadings and the record schema. The panel measures a simulated population and does not substitute for asking people.
Accelerating Research Idea Generation with LLMs: From Data to Domain Isomorphism
Yansong Feng
Wangxuan Institute of Computer Technology, Peking University

Recent advancements in large language models (LLMs) demonstrate strong potentials for generating research ideas, yet such ideas often struggle with feasibility and novelty. In this paper, we investigate whether augmenting LLMs with relevant resources during the ideation process can improve idea quality. We made two different attempts: (1) incorporating data from related works as well as preliminary validation to guide models towards more feasible ideas; (2) bringing structure analogies from parallel domains to inspire novel research hypotheses. Both automatic and human evaluations show that our methods can not only improve the quality of the generated ideas, but also help human researchers propose better ideas. Our findings highlight the potential of LLMs in real-world academic settings.

Decomposing Echo Chamber: An LLM Agent-Based Simulation of Network Structure, Algorithmic Filtering, and Polarization
Jinpeng Wang2, Xin Yu1, Zhenzhen Ren3, 4, Peizhuang Miao2
1 Shenzhen University
2 Tsinghua University
3 Zhongguancun Academy
4 Fudan University

The echo chamber concept conflates structural isolation, attitude homophily, and algorithmic filtering, contributing to contradictory findings across studies. This study decomposes the concept by independently manipulating these mechanisms in a large language model agent-based simulation (LLM-ABM) with a 3 (network structure) × 2 (recommendation algorithm) × 2 (initial attitude distribution) factorial design (500 agents, 15 rounds, N = 486,000 observations). Results show that structural isolation alone produced no significant effect on either perception bias or attitude polarization. Attitude homophily increased perception bias but did not independently shift attitudes. Algorithmic filtering increased both outcomes with larger effect sizes than network structure. Pre-existing polarization was the strongest predictor and amplified all other factors. Perception bias partially mediated their effects on attitude polarization. In sum, echo chamber effects on attitudes are conditional on the concurrence of multiple mechanisms. The study also demonstrates the methodological potential of LLM-ABM for communication research.

Dual-Capacity Signaling: How Police Show Regime Strength on Douyin
Yiqiang Wang1, Haohan Chen1, Yingqi Huang2
1 The University of Hong Kong
2 University of Wisconsin–Madison

How do authoritarian states signal regime strength to citizens? Conventional wisdom focuses on displays of violence to intimidate audiences. We find this understanding incomplete as states adapt to the new media environment. We theorize dual-capacity signaling: authoritarian states use new media to project both coercive capacity (the ability to monitor, deter, and punish) andadministrative capacity (the ability to serve, coordinate, and govern). Applying this theoretical framework, we develop a new computational pipeline combining Large Vision-Language Models and Large Language Models and apply it to 48,858 videos from 31 Chinese municipal police Douyin accounts. Our study yields four findings. First, administrative signaling is dominant and has grown steadily over time, while coercive signaling remains a minority. Second, signaling follows the political calendar: coercive messages spike in the days before sensitive political events, while administrative content rises around public holidays. Third, across cities, online signaling correlates with local state capacity and socioeconomic pressure. Finally, we find a digital engagement dilemma: despite the dominance of administrative signaling, audiences engage more strongly with coercive displays. This study advances theories of authoritarian political communication and introduces a scalable computational method for analyzing political video content.

The Legitimacy Game: How “PhD-Grade” Data Became the Currency of AI Hype and Anxiety in China
Keyu Zhang, Fen Lin
City University of Hong Kong

In China’s rapidly developing AI industry, data annotation has traditionally been viewed as a low-cost, labor-intensive task (Wang et al., 2022; Yılmaz & Bostancı, 2025; Miceli et al., 2020). However, as China defines high-quality datasets as a decisive factor in its AI national competitiveness (National Development and Reform Commission, 2024), a growing trend has emerged where AI companies are increasingly hiring highly educated workers, specifically PhD students, to perform simple work, offering salaries nearly two hundred times higher than those paid to low-educated workers (Wang, 2023; Lin, 2023). This shift presents a paradox: Why are companies choosing to hire highly educated workers for tasks that are simple and could be performed by lower-wage labor?

This “irrational” puzzle is significant because it reveals how Chinese institutions are managing technological and economic uncertainties in the AI sector. As China continues to regulate its AI and digital economies, understanding labor practices in this industry provides insights into how the state adapts to broader industrial and technological challenges. Much of the existing literature on digital labor and AI training views these tasks as part of a rational, efficiency-driven system, this study offers a more powerful explanation from organizational lens.