The Prospects of Artificial Moral Enhancement
Before and After LLMs
Introduction
As social beings, we often turn to others in search of attachment, advice, or just guidance (Baumeister & Leary, 1995). However, in recent years, these relationships have been permeated, marked, and often replaced by large language models (LLMs) (“Emotional Risks of AI Companions Demand Attention,” 2025). People now routinely rely on these systems for everyday guidance, whether seeking recommendations for places to visit in a new city or advice on effective ways to get in shape.
But, do we also turn to them for moral guidance? Moral advice, however, seems to have a different status. Our moral beliefs are deeply ingrained features of our identity (Strohminger & Nichols, 2014), we rarely want to change or improve them (Sun & Goodwin, 2020), and when we do, we usually turn to those who know us (Haidt, 2012).
Much of the recent experimental research into artificial moral assistants has focused on assessing general aspects of LLMs (e.g., (Angrisani et al., 2026)) or has been too specific to certain non-moral advice domains (e.g., (Prahl & Jin, 2024)). However, little research has been done into the moral acceptability of these seemingly moral assistants over time.
Our research
This research provides an in-depth analysis of the trends in the perception of AI moral assistants compared to their human counterparts in different moral and non-moral contexts, spanning three waves: before ChatGPT, in 2022, in 2024, and in 2026.
A total of 931 participants (pending the third, 2026 wave) completed an experiment with a 2 (wave: 2022 vs. 2024) × 2 (assistant: human vs. virtual) × 3 (scenario: moral, travel, training) between-subjects design. In each experimental condition, an individual sought advice from either a human or virtual assistant about training, travel itineraries, or morals.
The vignette describes an interaction in which either the AI or human assistants draws the individual’s attention to factual inaccuracies, conceptual ambiguities, or personal limitations. In response, the individual changes their mind. Following this reading, the participant evaluated the overall assessment, reliability, and robustness of the advice. After the main task, participants also completed post-experimental measures assessing identity relevance, interest in improvement, perceived normality, and the assistant’s personal knowledge of the individual.
Some preliminary results
Human assistants were rated significantly higher than Virtual assistants. There was a difference in rate between human and virtual assistants, in both 2022 and 2024, with the gap being notably larger in 2024 (see Figure 1).
An exploratory analysis showed that perceived personal knowledge differed across assistants and domains, and predicted moral approval.
This research examines how public moral attitudes toward AI evolved over four years, spanning the period before and after its widespread adoption. It reveals how the perception of personal knowledge may be a fundamental factor in shaping these public attitudes.
References
2026
- Gaps in Large Language Model Awareness, Usage, and Perceptions in the United States: Evidence from a Nationally Representative Longitudinal Survey2026
2025
- Emotional Risks of AI Companions Demand Attention2025
2024
- Doctor Who?: Norms, Care, and Autonomy in the Attitudes of Medical Students toward AI Pre- and Post-ChatGPT2024
2020
- Do People Want to Be More Moral?2020
2014
- The Essential Moral Self2014
2012
- The Righteous Mind: Why Good People Are Divided by Politics and Religion2012
1995
- The Need to Belong: Desire for Interpersonal Attachments as a Fundamental Human Motivation1995