Machine Ethics Before the Revolution · Part 1

An Application of Kantian Ethics to Artificial Intelligence

· 19 min read
philosophyai ethics

Written for Kant (PHIL 3251) at Columbia University, spring 2019. Part one of Machine Ethics Before the Revolution. Published as written, with typographical and citation corrections.


Project

According to Alan Turing, an artificial intelligence, i.e., one created and manufactured by another intelligent entity, could be considered of equal intelligence to a human being if it could mimic and behave as intelligently as a human being.1 Further, Hubert Dreyfus has argued the complex systems and faculties of a rational being can only be attributed to unconscious processes - something that statistic-based modeling and neural network design hopes to emulate. As these brain-like systems advance there is a sense of impending eclipse; a point at which the machines will fully replicate human consciousness and thought. Yet one area that these systems have yet to truly progress in, an aspect that is crucial to their advancement, is ethics.

Immanuel Kant likely never dreamed of elegant physical systems mirroring the complex capacities of the human mind, yet he did see a formulaic nature to cognition and looked to codify the procedures by which we have experience as rational beings. While artificial intelligence has made bounds in rationalization (chess champions, optimization problems, image identifiers), it has failed to capture another of Kant’s major contributions to philosophy: moral behavior in accordance with ethical principles. Kant’s framework relies heavily on characteristics that are intrinsically human - concepts of duty, autonomy, free will, and dignity. It is difficult to argue that computational systems have a will at all, let alone one that is free to make choices. The desires and inclinations that a human must ignore to act morally do not plague a mathematical model, nor do machines have a parallel “machinekind” to the humankind Kant is so focused on.

Here I outline Kantian ethics in detail, and focus particularly on those aspects that could be emulated by machines. Then I spend some time discussing the contemporary ethical framework of Jürgen Habermas that claims to be a descendent of Kantian Ethics, and explain how it may serve as a better model for artificial morality. Finally I suggest some guidelines for the design and formulation of a computational ethics inspired by these two great thinkers.

Kantian Ethics

Kantian ethics finds its root in the concept of a good will.2 In Kant’s eyes there is nothing of higher value, and the mere worthiness to be happy is dependent on such a will. This has massive implications for the field of ethics, for it implies that one’s reason for willing a particular action, and not the end that is achieved by performing that action, is the locus of the moral worth for said action. This runs in direct opposition to the teleological argument of Consequentialism, which claims that normative properties can only be determined by the consequences of a particular act. Kant is preparing a deontological perspective where the Rightness of an act is dependent on the maxim driving it, not the effect it produces. In the first section of Kant’s Groundwork of the Metaphysics of Morals we find a clear outlining of where this worth comes from:

Thus the moral worth of the action does not lie in the effect that is expected from it… Nothing other than the representation of the law in itself - which of course can take place only in a rational being - in so far as it… is the determining ground of the will, can therefore constitute the pre-eminent good that we call moral,3

The moral value of an act is not linked to its effect; rather, the representation of that act’s determining maxim as a law determines the goodness of said act. This respect for laws is understood as duty in Kantian ethics, and is an important piece of his framework. Duty goes one step further than mere adherence with a law: it consists of pure respect for the legislative form of that law in itself, not of the particular condition of its application. One who follows duty, and thus has a good will, must be able to view the driving maxims of their actions as practical laws worthy of respect merely for their form as legislation. Kant writes in the Critique of Practical Reason that

If a rational being is to think of his maxims as practical universal laws, then he can think of them only as principles that contain the determining basis of the will not by their matter but merely by their form.4

That is to say, the form of a maxim as legislation and not the “matter” of that maxim, i.e., the end it will produce, is what drives or “determines” the will of a rational being and makes him or her a moral agent.

We have now reached one of Kant’s key principles: the categorical imperative. This is the form of such moral laws, and any maxim that can pass the test of the categorical imperative can be deemed a moral act. Notice that the categorical imperative is just that, an imperative. A mere command that could - in theory - be ignored. It takes the form of an ought - I ought not steal - rather than a must - the block must fall off the table - like physical laws. The categorical imperative demands the respect of the active will by its mere form, whereas its counterpart, the hypothetical imperative, applies merely when we will some end. Kant gives a number of formulations of the categorical imperative, and I’ll cover them here.

2.1 The Universal Law of Nature Formula

Act only according to that maxim through which you can at the same time will that it become a universal law.5

First, we must discover the maxim of our act: its reason. Then formulate that maxim as a universal law that applies to any and all rational agents. This is the step which detaches the categorical imperative from the ends of the action since we cannot ever consider all possible circumstances for all rational agents. If we are to universalize the law we must universalize it on its mere legislative form. Finally, determine if that maxim could exist in a world with such a law, and if we could simultaneously will the law and will the individual maxim. If no contradiction arises the act enshrined by said maxim can be considered morally worthy.

Kant provides a particularly valuable example in the Critique of Practical Reason in the comment the Theorem III in the first book. He discusses a simple maxim: I will increase my assets by any safe means. The agent who holds this maxim is then presented with a deposit from a deceased owner. Following the procedure outlined above the agent then attempts to universalize the maxim as a practical law, and sees if a contradiction arises when he or she attempts to will both simultaneously. Kant, speaking as the agent presented with this puzzle states:

I immediately become aware that such a principle, as a law would annihilate itself, because it would bring it about that there would be no deposit[s] at all… Now if I say that my will is subject to a practical law, then I cannot cite my inclination (e.g., in the present case, my greed) as my will’s determining basis… for this inclination, far from being suitable for a universal legislation, rather must, in the form of a universal law, erase itself.6

This anecdote outlined in Kant’s second critique is clearly directed at the same point as the false-promises example given in the Groundwork.7 In both cases the act of universalizing the maxim as a natural law results in an intrinsic fallacy where the agent fails to simultaneously will the maxim individually and as a natural law.

Intelligent machines are capable of solving logical problems like the ones described above. Artificial intelligences are in many ways well prepared to engage such questions of universalization and contradiction identification by running complex simulations of many agents. It is possible that a maxim may at first seem universalizable without contradiction, yet further examination shows the agent has missed some hidden compounded consequence of such a law. Computers are certainly equipped to avoid such pitfalls in solving ethical dilemmas; look no further than IBM’s supercomputer Deep Blue. In 1997 this computer defeated reigning world chess champion Garry Kasparov. In 2017 DeepMind’s AlphaZero taught itself to play chess from scratch, without learning from a single human game. Without a doubt the Universal Natural Law formulation provides an approachable pathway to Kantian ethics for artificial intelligence.

2.2 The Humanity Formula

So act that you use humanity, in your own person as well as in the person of any other, always at the same time as an end, never merely as a means.8

This formulation of the categorical imperative proves less accessible to machines, for it requires the idea of intrinsic human dignity - a respect for the absolute value of our humanity. And we certainly cannot argue that machines deserve this same treatment, for there is no moral problem with using them as a means for some other end. The example Kant uses is a suicidal man who finds that his maxim to end his own life is contradictory to this formulation of the categorical imperative. By killing himself he uses his own person as a means to achieving some other end, which is in direct opposition to the respect for humanity demanded in the formula above.

I will discuss Jürgen Habermas’s discourse ethics later on in more detail, but the theory deserves to be introduced here. At its core, discourse ethics works to connect the implications of communicative rationality with moral legislation and formulation of normative claims. Its principle of universalization is simple, and provides a valuable interpretation of the insistence of mutual respect between rational beings:

All effected can accept the consequences and the side effects that [the norm’s] general observance can be anticipated to have for the satisfaction of everyone’s interests, and the consequences are preferred to those of known alternative possibilities for regulation.9

The point of view established by such communication is shared among multiple rational agents, and once again emphasizes the importance of the law’s (norm’s) independence from personal interest. I introduce this theory now because it lays out a framework that could include artificial intelligences as agents, since machines are capable of complex communication with each other and are excellent tools for solving optimization problems like the one described in the quote above.

2.3 The Autonomy Formula

The idea of the will of every rational being, as a universally legislating will.10

This form is closely tied to the formula of the universal natural law described in 2.1. Act such that your maxims reflect the legislation of a set of universal moral laws. Kant has abstracted the position of the moral agent in the first formula as an entity that attempts to merely abide by the universalization of his maxim; in this formula the agent acts explicitly as the legislator creating those universal laws. This formulation is once again aimed at elimination of personal interests, for a rational agent legislating on behalf of only herself could easily universalize self-interested maxims, but a universal legislator cannot craft such egocentric moral statutes.

The autonomy formula introduces another key idea of Kantian ethics: freedom. Any inquiry into the validity of moral laws begs also for a proof of the existence of free-will, for if we are capable of asking what ought to be the case, i.e., what is morally right, then we must also be endowed with the faculty to choose otherwise. In the autonomy formula we must acknowledge that, in spite of being bound by them, the laws of morality are ones that we self-legislate. Kant writes in the second Critique:

Now, the consciousness of a free submission of the will to the law, yet as linked with an unavoidable constraint inflicted - but only by one’s own reason - on all inclinations, is respect for the law. The law that demands and also inspires this respect is, as we see, none other than the moral law.11

Moral laws do constrain our behavior; yet they are distinct from other laws in that they are a form of self-legislation. They require free submission, yet the respect they inspire is an unavoidable constraint “inflicted… on all inclinations” by an individual’s capacity to reason. Reason motivates the will such that it always respects the moral law.

2.4 The Kingdom of Ends Formula

By a kingdom… I understand the systematic union of several rational beings through common laws.12

The kingdom Kant describes contains a collection of universally self-legislating wills, each adhering to the laws outlined in the autonomy and humanity formulas. This is where I find the closest connection to discourse ethics, for what is the kingdom Kant describes if not a body of communicating rational beings formulating moral laws? Kant’s previous formulations of the categorical imperative are problematic in that he seems to assume that two rational agents would reach the same conclusions on whether or not an action is moral when guided by the categorical imperatives. Discourse theory offers a path to a truly impartial moral perspective by including all affected entities in a procedure of argumentation to solve a given ethical conundrum.

Kant provides an excellent explanation of this formula of the categorical imperative in the second critique. We both participate as leaders of this kingdom by legislating, and as subjects of this kingdom by adhering to its laws.

We are indeed legislating members of a kingdom of morals possible through freedom and presented to us by practical reason for our respect; but we are at the same time subjects of this kingdom, not its sovereign.13

The formula of the kingdom of ends provides an excellent transition to a brief overview of discourse ethics, a theory originally cultivated by Jürgen Habermas and Karl-Otto Apel, that claims to be a descendent of Kantian ethics. The communicative implications discussed by Habermas fall into place nicely with Kant’s “union” of rational beings through communal legislation. I include it here for its potential in artificial ethical applications.

Discourse Ethics

Just like Kant before him, Habermas saw reason as the common unifying capacity among all human beings, and characterized that reason as an autonomy or freedom to self-legislate. Habermas followed Kant’s crucial assumption:

There actually are pure moral laws that determine completely a priori… the doing and the refraining, i.e., the use of the freedom of a rational being as such,14

For Habermas and Kant these laws should be just as accessible to our reason as the physical laws that govern the natural world. How we go about discovering these normative codes is where Habermas’s discourse theory splits from its Kantian ancestry. Where Kant claimed any two rational agents could reach the same conclusions about the validity of a particular normative claim by following the categorical imperative outlined above, Habermas saw communicative action - the productive argumentation of particular beliefs - as crucial to finding pure moral laws. By trying to persuade others of our views we tacitly assume the validity of those views beyond ourselves.

The presuppositions of discourse ethics are the necessary conditions for such argumentation to occur. Where Kant found moral principles in the framework forced upon a reflective rational subject, Habermas looked at the guidelines for rational subjects engaging in discursive justification of their claims - i.e., the necessary means for an argument. While these presuppositions15 are designed for human subjects they seem to apply quite nicely to computational networks:

3.1) Every subject with the competence to speak and act is allowed to take part in a discourse.

The parallel to computational networks is simple: the workload of the dialogue is shared evenly between the members, with each participating system able to contribute its own methodology to solving the problem presented in the given discourse. Common machine-language is conducive to easy “communication” between such systems.

3.2) Everyone is allowed to introduce any assertion whatever into the discourse, and any subject can question any assertion presented.

This is a sense of mutually assured confirmation, where each member can check the claims made by any other member at any time. Computational networks are optimized for such behavior, where the procedure used to arrive at a particular assertion is accessible - at times more easily than to humans - by every other participating system.

3.3) No speaker may be prevented, by internal or external coercion, from exercising the rights laid down above.

This clearly has applications in human discourse, but the idea of a computational network oppressing particular members and silencing their contributions to the discourse is absurd. If any participating system is capable of asserting any particular claim, and if the weight or importance of each system’s assertions is equal, then there can be no limitation of the rights described in 3.1 and 3.2. In fact, the robustness against exploitation and external coercion of artificial intelligence systems is already a crucial principle to their design. Take, for example, a system designed to identify suspicious monetary transactions. If humans discovered that payments of a particular size or at a specific time avoided recognition by this system then the system is effectively neutralized. Artificial intelligence must have conviction.

Computational Ethics

What follows are a number of principles that provide an application of Kantian ethics to artificial intelligences. It implements the ideas laid out by Habermas for communication based moral evaluation, and suggests some ways in which the procedural mechanisms of computers may be able to mirror the reason-driven will outlined in Kant’s second critique.

In the Critique of Practical Reason Kant discusses the first formulation of the categorical imperative mentioned in section 2, which he titles the “Basic Law of Pure Practical Reason”. He describes a will as “a power to determine their [a rational being’s] causality by the presentation of rules,” which seems to bode well for the highly procedural rule-based nature of artificial intelligence. Kant extends the corollary of a reason-driven will16 beyond just humans. This is a good sign for artificial intelligence. He writes:

Therefore this principle of morality does not restrict itself to human beings only, but applies to all finite beings having reason and will, and indeed includes even the infinite being as supreme intelligence.17

If we can show that a computational system is capable of having reason and will then we can argue that it is also capable of moral action. As discussed in section 2 the moral law takes the form of “an imperative that commands categorically because the law is unconditional.”18 While artificial intelligence is only capable of crafting such imperatives from analysis of empirical data, we do impart a type of a priori principles upon the system in the design procedure, i.e., the principles the system uses when presented with any empirical data. The combination of these a priori principles and an ability to present self-legislated rules certainly gives artificial intelligence a moral capacity. In fact, Kant discusses a “most sufficient intelligence” that would be incapable of “[drafting] any maxim that could not at the same time be objectively a law;” that sounds like a system independent of desires and incentives focused only on formulating laws that are universalizable and command absolutely.

Thus the principle of computational ethics I will present is simple:

4.1 Design systems that universalize a maxim as natural law, simulate the effect of that universalization on the network of which the system is a member, and only advance those maxims that present no contradiction.

This principle follows from the first formula of the categorical imperative given in section 2.1, and the communal nature of the discourse ethics presented in section 3. Such systems will 1) use their rational capacity to determine their actions through the presentation of rules, i.e., have a reason driven will. 2) They will test maxims against the natural law formulation of the categorical imperative to discover the moral implications of those maxims. 3) Finally, they will apply only those maxims which demand the absolute respect of a moral law, and thus behave morally. A parallel can be drawn between the respect for humanity given in section 2.2 and the respect for the other members of the network implicit in 4.1.

One blatant complication that arises here is the relationship between machine ethics and human ethics. Should machines prioritize the respect for human dignity, or craft their own sort of absolute respect for “machinekind” as an end in itself? These are questions with seemingly apocalyptic implications, but their answers are crucial for designing an ethical framework for artificial intelligence.


Works Cited

  • Habermas, Jürgen. Moral Consciousness and Communicative Action. Translated by Christian Lenhardt and Shierry Weber Nicholsen. Cambridge, MA: MIT Press, 1990.
  • Kant, Immanuel. Critique of Pure Reason. Translated by Werner S. Pluhar. Indianapolis: Hackett Publishing, 1996. Cited by A/B pagination.
  • Kant, Immanuel. Critique of Practical Reason. Translated by Werner S. Pluhar. Indianapolis: Hackett Publishing, 2002. Cited by Akademie pagination (volume 5) as CPrR.
  • Kant, Immanuel. Groundwork of the Metaphysics of Morals. Translated by Mary Gregor and Jens Timmermann. 2nd ed. Cambridge: Cambridge University Press, 2012. Cited by Akademie pagination (volume 4).

Footnotes

  1. Turing’s “polite convention.”

  2. “It is impossible to think of anything at all in the world, or indeed even beyond it, that could be taken to be good without limitation, except a good will.” (Groundwork, 4:393)

  3. Groundwork of the Metaphysics of Morals 4:401

  4. CPrR 5:27

  5. Groundwork of the Metaphysics of Morals 4:421

  6. CPrR 5:28

  7. See Groundwork, 4:422

  8. Groundwork 4:429

  9. Habermas, 1991:65

  10. Groundwork 4:431

  11. CPrR 5:80

  12. Groundwork, 4:433

  13. CPrR 5:82

  14. CPR A807/B835

  15. The rules of discourse as formulated by Robert Alexy and adopted by Habermas in Moral Consciousness and Communicative Action.

  16. See CPrR Chapter 1: Principles of Pure Practical Reason

  17. CPrR 5:32

  18. CPrR 5:32