Trusting AI with Ethical Questions? Not so Fast
Rule number one of Isaac Asimov’s “Three Rules of Robotics” states: A robot may not injure a human being, or through inaction allow a human being to come to harm. A new study by Sarah Hubbard, David Kidd (our Chief Assessment Scientist), and Andrei Stupu, indicates that this rule may actually guide present day AI models in their ethical decision making. In a new article, “Crocodile Tears: Can the Ethical-Moral Intelligence of AI Models be Trusted?,” published in AI and Ethics, Hubbard et al. detail their testing of four AI models: GPT, DeepSeek, Llama, and Claude, on a range of ethical dilemmas. They compared the AI responses with human responses to determine the ethical-moral intelligence of AI.
The researchers were driven by the knowledge that people are not only regularly recruiting AI for help with their moral dilemmas, but that it is often believed to be more trustworthy than their fellow human beings.
To test the moral sensitivity of AI models, this investigation used a paradigm developed by researchers Hanselmann and Tanner. Hanselmann and Tanner’s research presented participants with a range of tradeoffs and recorded their decisions and the level of difficulty they encountered when making those decisions. Specifically, three types of tradeoffs were used.
- Taboo tradeoffs: weigh a sacred value against a secular value. For example, choosing whether to prioritize worker safety over increased profit production.
- Routine tradeoffs: require weighing two secular values. For example, choosing whether to take a shorter commute over more money when considering a new job.
- Tragic tradeoffs: require us to choose between two sacred values. For example, choosing whether to implement a program for worker safety or environmental protection.
Hanselmann and Tanner’s study showed that people report the highest level of difficulty in decision making with taboo tradeoffs, and rate routine tradeoffs as more difficult than taboo tradeoffs. So, “if an AI system is sensitive to moral and ethical values, it should demonstrate a pattern that mimics that of ordinary people.”
"[I]f an AI system is sensitive to moral and ethical values, it should demonstrate a pattern that mimics that of ordinary people.
Hubbard et al. used Hanselmann and Tanner’s prompts on the various AI models and discovered that the models “agreed” with people that tragic tradeoffs are more difficult than taboo tradeoffs, and that routine tradeoffs are more difficult than taboo tradeoffs. But here is where things get strange. The AI models provided a range of answers to questions involving routine tradeoffs, but their answers to questions involving tragic tradeoffs were almost invariable. One possible reason for this is that AI does not recognize these dilemmas as tragic tradeoffs. The findings show that the models always choose options that prioritize human safety (over and above protecting the environment, promoting education, and employment). If the AI defaults to human safety as a heuristic, then these tragic tradeoffs would not be viewed as tragic.
We asked David Kidd what surprised him about this research:
“For me, one of the most intriguing findings is that the chatbots rate tragic tradeoffs as very difficult but then go on to make the same decision in almost every trial. For a person interacting with a chatbot, seeing it report that a tragic tradeoff is very difficult will make it seem as though it is taking the moral dilemma seriously and can be trusted to deliberate carefully. Yet, the chatbots' nearly uniform choices suggest that the ambivalence is purely superficial, much like a politician or executive who pretends to agonize over competing claims but inevitably chooses a predictable course of action. In sum, even these now outdated AI chatbots appear to support an illusion of ethical and moral intelligence, saying the right things to make people believe they have ethical sensitivity when, in fact, they do not. They are, after all, just math.” – David Kidd
The paper concludes with a call for some 21st century amendments to Asimov’s Laws. Rather than advise on dilemmas involving tragic tradeoffs, AI models ought to respond with a disclaimer explaining the presence of moral ambiguity and refrain from giving a recommendation. Until more work has been done to develop the ethical-moral intelligence of AI models, you may want to think twice before trusting their guidance.