Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
· Source: arXiv cs.AI
Recent work on AI alignment has largely focused on first‑order norms—teaching models which behaviors are acceptable or forbidden. Yet social intelligence also requires grasping second‑order expectations, known as metanorms, that specify who enforces a norm and how (for example, through public shame or legal sanctions). To address this gap, the authors introduce a new framework for evaluating metanormative reasoning in large language models (LLMs). Their approach rests on two axes: emotional valuation and behavioral response to a violation. They also propose classification tasks that predict the offender’s self‑regulation and the observers’ regulatory actions. As part of the study, they release the NormReact dataset, comprising 450 manually annotated normative‑violation scenarios with emotions and behavioral responses, taking into account variables such as the offender’s gender and the observer’s social proximity. Results show that, across six evaluated LLMs, the models more frequently predict punitive outcomes in situations where humans would expect inaction, and their alignment with human judgment worsens as social distance grows. These findings warn that in socially sensitive applications—such as conflict mediation or policy simulation—AI systems may present a distorted view, overemphasizing punishment while underestimating tolerance and relational calibration that characterize real‑world normative regulation. Therefore, understanding and improving metanorm reasoning is essential to avoid biases that could influence AI‑based decisions in social contexts.
Read the original article on arXiv cs.AI
This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.