Ethical (Mis)-Alignments in AI Systems and the Possibility of Mesa-Optimizations
A conceptual framework separating four places an AI system can fail: the human goal, the training objective, a learned internal objective, and the system’s effects on people.
The paper is conceptual rather than empirical. Its practical value is locating different failure classes at different system boundaries instead of calling every problem “model bias.”
Read article- 01Human goal
Is the intended objective itself ethically defensible?
- 02Training objective
Does the optimization target faithfully represent that goal?
- 03Learned objective
Did the system internalize a different objective?
- 04Actual effects
What happens to people when the system is deployed?