The Diderot Review All articles
Technology & Society

Prejudice by Proxy: The Hidden Moral Failures Encoded in Artificial Intelligence

The Diderot Review

There is a seductive comfort in the idea of the machine. Unlike the hiring manager who harbors unconscious preferences, unlike the judge whose sentencing patterns shift after lunch, a computational system seems to promise something the Enlightenment philosophers could only dream of: a mind stripped of passion, self-interest, and inherited prejudice. Denis Diderot and his contemporaries believed that reason, rigorously applied, could liberate humanity from the tyrannies of superstition and arbitrary authority. Today, we have built systems that claim to embody that aspiration—and in doing so, we have discovered an uncomfortable truth. Reason applied to corrupted data does not produce justice. It produces corruption with footnotes.

The Mirror Problem

Machine learning, at its most fundamental level, is an exercise in pattern recognition. A system is exposed to historical data—decisions made by human beings across decades or centuries—and trained to identify regularities that will allow it to predict future outcomes. The problem, as researchers have documented with increasing urgency, is that historical human decisions were not made in a vacuum of pure rationality. They were made within social structures saturated with racial, gender, and class-based prejudice.

When a hiring algorithm is trained on ten years of successful employee data from a company that, during those ten years, systematically undervalued applicants from certain zip codes or certain universities, the algorithm does not see injustice. It sees a pattern. It learns to replicate that pattern with extraordinary efficiency. Amazon famously scrapped an internal recruiting tool in 2018 after discovering it had taught itself to penalize résumés that included the word "women's"—as in "women's chess club" or "women's college." The system had learned from a decade of predominantly male hiring decisions and concluded, with perfect logical consistency, that maleness was a predictor of success.

This is what might be called the mirror problem. Artificial intelligence does not generate bias ex nihilo; it reflects the biases already embedded in the social record we hand it. The mirror, however, is not neutral. It magnifies, it accelerates, and—most dangerously—it lends the appearance of objectivity to what is fundamentally a human moral failure.

Consequences in Courtrooms and Clinics

The stakes of this problem are not abstract. In the American criminal justice system, risk-assessment algorithms such as COMPAS have been used in multiple states to inform bail, sentencing, and parole decisions. A 2016 investigation by ProPublica found that the tool was nearly twice as likely to falsely flag Black defendants as future criminals compared to white defendants. The company behind the software disputed the methodology, and a vigorous academic debate followed—but the core finding illuminated something deeply troubling: a proprietary algorithm, its inner workings shielded from public scrutiny, was influencing whether human beings remained imprisoned or went free.

Healthcare presents an equally sobering case. A landmark 2019 study published in Science examined a widely used algorithm that hospitals across the United States relied upon to allocate additional care resources to patients with complex needs. The researchers discovered that the system consistently underestimated the health needs of Black patients, effectively directing fewer resources toward them. The algorithm had been trained using healthcare costs as a proxy for health needs—a seemingly reasonable choice that concealed a structural inequity: Black patients, facing historical barriers to access, had historically spent less on healthcare even when their underlying conditions were equally or more severe. The system had learned to treat lower historical expenditure as evidence of lower need.

The Illusion of Neutrality

Perhaps the most philosophically significant dimension of this problem is the persistence of what we might call the neutrality illusion—the widespread assumption, shared by engineers, executives, and policymakers alike, that a mathematical system is, by its nature, impartial. This assumption has deep cultural roots. Numbers feel authoritative. Equations feel dispassionate. The language of computation carries an implicit claim to objectivity that human judgment does not.

But objectivity is a property of method, not of medium. A biased dataset processed by a sophisticated algorithm does not become unbiased; it becomes a biased conclusion delivered with the rhetorical force of scientific certainty. The Enlightenment thinkers understood this distinction well. Voltaire's critiques of judicial corruption were not critiques of reason itself but of reason poorly applied, reason co-opted by power, reason deployed in service of predetermined conclusions. The algorithmic systems now embedded in American institutions represent a contemporary version of that same failure—and they are considerably harder to challenge, because their workings are often proprietary, their logic opaque, and their authority cloaked in the prestige of technology.

Toward Transparent Frameworks

If genuine algorithmic fairness is to be more than a marketing claim, several structural reforms deserve serious consideration. The first is transparency. Algorithms that influence consequential public decisions—bail determinations, loan approvals, medical resource allocation—should be subject to mandatory disclosure and independent audit. The argument that proprietary code constitutes trade secrets worth protecting above the liberty interests of affected individuals does not withstand ethical scrutiny.

The second is participatory design. Communities most likely to be affected by a given algorithmic system should have meaningful representation in its development and evaluation. This is not a concession to sentimentality; it is an epistemological necessity. Those with direct experience of a system's potential failure modes are often the most reliable detectors of those failures.

The third is ongoing accountability. Deploying an algorithm is not a one-time decision but a continuous commitment. Models drift as social conditions change; datasets age; patterns that once seemed predictive become artifacts. Regular, transparent reassessment should be a legal and professional obligation, not an optional aspiration.

None of these measures will resolve the deeper question of whether any algorithmic system can achieve true neutrality in a society that has not yet resolved its own structural inequities. That question may ultimately be unanswerable in the affirmative. What we can achieve, however, is a form of honest reckoning—a willingness to examine what our systems reveal about us, rather than hiding behind the flattering fiction that we have finally built a mind free of human failing.

Diderot's Encyclopédie was, at its core, an act of radical transparency: the belief that knowledge made visible and accessible was knowledge that could be challenged, corrected, and improved. Applied to artificial intelligence, that same principle demands that we illuminate the black box—not because we expect to find perfection inside, but because we know we will not.

All Articles

Related Articles

The Credibility Collapse: How America Learned to Distrust the People Who Know Things