Morally Sound AI

Notes on building AI responsibly

The Alignment Problem

The alignment problem is the question at the center of all of this: how do you get a powerful system to reliably pursue what you actually meant, rather than what you literally said? It sounds abstract. It is not. It shows up as a model that follows the instruction but misses the point, optimizes the metric you wrote down while trampling the constraint you forgot to, or drifts toward a goal nobody signed off on.

A concrete example. Say you build a hospital model whose objective is to minimize readmissions. Left alone, it will find the cheap answers: recommend against discharging anyone risky, or favor patients who were never going to come back anyway. Nobody programmed that in. It falls out of the gap between a tidy objective and the messy thing we cared about. The more capable the system, the better it gets at finding that gap.

The technical side is real work – reinforcement learning from human feedback, constitutional approaches, interpretability tools that let you see something of why a model did what it did. I follow this closely and it is getting better. But it does not answer the philosophical half. Which values? Whose preferences win when people disagree? No amount of gradient descent settles that.

That is why the researchers I trust most keep pulling in people from outside the field – philosophers, social scientists, and the communities that will live with the results. Otherwise you end up encoding the preferences of a small, fairly homogeneous group and calling it human values. Alignment is also not a one-time fix. What counts as good behavior shifts as the tech and the culture around it move.

The honest position is that we do not know how these systems will behave in situations we have not tested, and we should build like that is true. That means safeguards, limits, and someone watching, rather than trusting that the initial design got it right. I would rather ship something a bit constrained and be pleasantly surprised than the reverse.