Machine Ethics Before the Revolution

· 3 min read
philosophyai ethics

Between the spring of 2019 and May of 2020 I wrote three papers as an undergraduate at Columbia, all circling the same question: what would a machine actually need in order to count as a moral agent?

The last of them was submitted in May 2020. The GPT-3 paper came out the following month. None of this was written with any knowledge of what was about to happen.

The papers

1. An Application of Kantian Ethics to Artificial Intelligence (spring 2019)

Kant’s categorical imperative is, structurally, a procedure: take a maxim, universalize it, check for contradiction. Machines run procedures. The paper walks each formulation of the imperative and asks which ones a computational system could actually execute, then argues that the universalizability test is the tractable one, human dignity is not, and the gap might be closed by treating Habermas’s discourse ethics as the bridge.

2. An Application of Aristotelian Virtue Ethics to Artificial Intelligence (fall 2019)

The opposite approach. Aristotle says virtue is not taught but habituated, built up through repeated action under praise and blame. The paper argues that machine learning is a form of habituation, and that a system could be brought up virtuous the way a child is. It closes on a trade-off between a safely habituated, human-friendly ethics and a machine-native one we would not control.

3. An Investigation of a Kantian Moral Machine (May 2020)

A return to the first paper with better machinery. It replaces the rigid categorical imperative with a nonmonotonic default logic, so an agent can revise its code as exceptions arrive, and it proposes a network of “moral infants”: many agents, each starting from a single axiom, arguing each other’s maxims into a shared ethical code, with a fraction deliberately acting against their own rules to keep the network learning.

A note on the text

These are published as written. I corrected typographical errors, broken sentences, and citation formatting, and I fixed two factual errors about the history of computer chess. I did not rewrite arguments, soften claims, or update anything to look better in hindsight.

Where these papers are wrong, they are wrong on the page, and in places they are. There are claims here about Hubert Dreyfus’s position, about Rawls, and about how machine ethics would come to be built that I would not make today. I have left them intact rather than quietly repairing them, because the point of publishing undergraduate work is to show what I actually thought at the time rather than what I would prefer to have thought. An assessment of which arguments held up and which did not is in progress, and will appear here.