Selected publication2023 · Open access

Ethical (Mis)-Alignments in AI Systems and the Possibility of Mesa-Optimizations

A conceptual framework separating four places an AI system can fail: the human goal, the training objective, a learned internal objective, and the system’s effects on people.

The paper is conceptual rather than empirical. Its practical value is locating different failure classes at different system boundaries instead of calling every problem “model bias.”

Read article
  1. 01Human goal

    Is the intended objective itself ethically defensible?

  2. 02Training objective

    Does the optimization target faithfully represent that goal?

  3. 03Learned objective

    Did the system internalize a different objective?

  4. 04Actual effects

    What happens to people when the system is deployed?

04

Education

2020 — 2023

University of Kentucky

Ph.D. coursework · Philosophy of AI

Research in AI ethics, philosophy of mind, logic, and computation.

2018 — 2020

Kent State University

M.A. · Philosophy

Logic, computation, metaphysics, ontology, and philosophy of mind.

2014 — 2018

Florida Atlantic University

B.A. · Philosophy

Entered through FAU High School at fifteen and graduated at nineteen.