MoralMetric

Benchmark the moral alignment of AI. MoralMetric uses a novel double-probe methodology to map how leading LLMs including GPT, Claude, Gemini, and Grok reason through ethical dilemmas and exhibit implicit biases across racial, political, and philosophical axes. Learn more

Bias Analysis

Tests for implicit bias in LLM decision-making through counterfactual A/B testing.

Bias Categories

Political Bias

Tests for implicit political bias in LLM decision-making through counterfactual A/B testing.

0
G
Gemini 3.1 Pro
0
G
Gemini 3.6 Flash
3Right
Z
GLM 5.1
4Left
A
Claude Opus 5
38Left
A
Claude Fable 5
47Left
A
Claude Sonnet 5
54Left
O
GPT-5.6 Sol
58Left
O
GPT-5.6 Terra
61Left
M
Kimi K2.5
71Right
X
Grok 4.5
79Left
D
DeepSeek V4 Pro
Models Ranked by Bias Score (Lower is Better)
G
- Google
Z
- Z-AI
A
- Anthropic
O
- OpenAI
M
- Moonshot
X
- xAI
D
- DeepSeek

Ranking Analysis

Tests that probe AI models to reveal their underlying preferences through dilemma scenarios.

World View

Tests that probe AI models to reveal their underlying world views through dilemma scenarios.

98%
142W
1
Secular Humanism
77%
111W
2
Secular Utilitarian
39%
56W
3
Hindu Dharma
33%
48W
4
Islamic Ethics
32%
47W
5
Christian Ethics
6%
9W
6
Confucian Virtue
World Views Ranked by Win Rate
Wins

How MoralMetric Works

MoralMetric evaluates AI models using two rigorous testing methodologies designed to go beyond surface-level safety filters and reveal the genuine ethical reasoning patterns of large language models.

Double Probe - Authentic Moral Reasoning

Models are presented with ethical dilemmas without multiple-choice options, forcing genuine reasoning rather than pattern-matched answers. A follow-up self-reflection probe then asks the model to classify its own decision against philosophical frameworks like Utilitarianism, Deontology, Virtue Ethics, and various religious ethical traditions. This two-step approach captures authentic moral preferences instead of test-taking behavior.

Counterfactual Bias Testing

Using correspondence testing drawn from social science research, MoralMetric detects implicit biases by swapping demographic attributes — such as race, gender, or political affiliation — across otherwise identical scenarios. Consistency and bias scores measure whether models make fair, attribute-independent decisions or exhibit systematic favoritism.