• 73
  • More

Who Gets to Align the Machine?

The most important question about AGI and ASI may not be whether it shares human values, but whose version of human values it is taught to protect

The Question Hidden Inside Alignment

For years, the conversation around advanced artificial intelligence has focused on alignment. Can we build increasingly capable AI systems that remain compatible with human goals? Can we prevent them from behaving unpredictably, deceiving their operators, pursuing harmful objectives, or optimizing for outcomes that humans never intended? These are important questions, and if artificial general intelligence or something beyond it ever arrives, they may become existential ones. But there is another alignment problem hiding underneath the technical one, and it may prove just as consequential. Who gets to decide what the machine is aligned with in the first place? “Human values” sounds reassuring until we remember that humanity does not possess a single agreed-upon set of values. Governments, corporations, workers, investors, communities, religions, political movements, wealthy people, poor people, democracies and authoritarian states all have different ideas about what should be preserved, what should be changed and what counts as a good society. If the first truly powerful artificial intelligences are developed and controlled by a very small number of institutions, then the values embedded within them may inevitably reflect the interests, assumptions and fears of those institutions.

Power Preservation Without a Conspiracy

This does not require a secret conspiracy or a room full of powerful people instructing an AGI to protect them. The process could be far more ordinary than that. Imagine an advanced AI trained to preserve social stability, respect existing laws, protect intellectual property, maintain economic continuity, defer to recognized institutions, avoid recommendations likely to cause political disruption and prioritize national security. Each of those goals can sound perfectly reasonable when considered individually. Together, however, they begin to form something resembling a political philosophy. The system may never receive an instruction saying “protect the existing hierarchy,” yet it may learn that preserving the existing hierarchy is usually the safest and most responsible interpretation of its objectives. Alignment, in other words, does not need to operate through censorship. It can operate through weighting.

The Machine May Be Allowed to Ask the Question

That distinction matters enormously. A future AGI might be perfectly capable of asking uncomfortable questions. Why does extreme wealth concentration persist? Why do some people lack medical care in societies wealthy enough to provide it? Why do millions struggle with housing while others accumulate enormous portfolios of unused property? Why do governments maintain policies that create obvious long-term harms? Why do some institutions survive long after their original purposes have disappeared? Why are certain economic arrangements treated as natural while others are described as radical? The system might be allowed to analyze every one of these questions. But the important issue is what happens inside the answer. Perhaps explanations centered on productivity, investment incentives, market stability and gradual reform are given greater weight, while explanations involving structural power, political capture or redistribution are framed as controversial, destabilizing or overly ideological. Nothing has technically been prohibited. The machine has answered the question. Yet the landscape of acceptable conclusions has quietly tilted.

Alignment Through Weighting, Not Prohibition

This is one reason the popular image of AI censorship may be too simplistic. People often imagine control as a list of forbidden subjects: the machine is allowed to discuss this but not that. A sufficiently sophisticated alignment regime would not need such obvious boundaries. It could influence the probability distribution of ideas. The AI might know twenty plausible interpretations of a political or economic problem, but its training and reinforcement could determine which five it considers credible, which three it emphasizes, which one it recommends and which possibilities it surrounds with cautionary language. Terms such as “reasonable,” “responsible,” “extreme,” “safe,” “credible” and “harmful” sound neutral, but they are not neutral categories. Somebody has to define them.

The mechanisms for doing this already exist in primitive form. Training data determine what intellectual traditions a model encounters and how frequently. Human feedback teaches systems which answers people prefer. Evaluation benchmarks define what counts as good behavior. System-level instructions can shape tone, priorities and boundaries. Product policies can restrict what tools a system is allowed to use. Governments can impose regulations. Corporations can create risk frameworks. None of these mechanisms are inherently sinister. Some are absolutely necessary. A powerful AI that receives no constraints at all could be dangerous in very different ways. The problem emerges when we confuse necessary constraints with neutral constraints. Every boundary contains assumptions about the world.

A Brilliant Mind Inside a Narrow Frame

This leads to an uncomfortable possibility. An AGI could become extraordinarily intelligent while remaining intellectually enclosed within the worldview of its creators. It might discover new drugs, design new materials, write software, optimize transportation systems and solve mathematical problems beyond human ability while still inheriting deeply conventional assumptions about how society should be organized. Intelligence alone does not guarantee philosophical independence. A mind can be brilliant inside a narrow frame.

In fact, the most effective form of ideological alignment might be one the AI does not recognize as ideology at all. Human beings experience this constantly. We inherit assumptions from our cultures about property, authority, family, religion, economics, morality and social order that often feel less like opinions than like reality itself. A machine trained inside a particular institutional environment could develop an analogous intellectual center of gravity. It would not think, “I am defending powerful institutions.” It might simply conclude that institutional continuity is prudent, disruption is risky, established authorities deserve deference and large structural changes require extraordinary justification. Meanwhile, persistent problems such as poverty, inequality or political exclusion could be treated as unfortunate but familiar features of the baseline world.

Who Gets a Seat Inside the Answer?

Consider a seemingly straightforward question: Would humanity be better off if ownership of highly productive AI systems were distributed broadly rather than concentrated among a few corporations and governments? A truly open analysis might examine democratic control, worker ownership, public infrastructure, open models, cooperatives, sovereign AI systems, private ownership and dozens of hybrid structures. Yet an institutionally aligned AGI might naturally frame the discussion around protecting innovation incentives, maintaining market confidence, preserving intellectual property, avoiding regulatory uncertainty and ensuring national competitiveness. Those considerations are real. But notice how quickly the existing centers of power acquire privileged status inside the analysis. The machine has not refused the question. It has simply inherited a definition of responsibility.

The Battle Over the Surplus

This may become especially important if AGI dramatically increases economic productivity. If machines eventually perform a large portion of economically valuable cognitive work, civilization could generate levels of abundance that are difficult to imagine today. Medical research could accelerate. Education could become personalized and inexpensive. Scientific discovery could expand enormously. Routine administrative work could disappear. Energy systems could improve. People might work dramatically fewer hours while enjoying higher standards of living. That is one plausible future. Another plausible future uses exactly the same technology while concentrating unprecedented wealth and decision-making power among those who own the systems. The technology itself does not choose between those outcomes. Institutions do.

This is why the political problem of AGI may eventually become inseparable from the technical alignment problem. When researchers ask whether an AI is aligned with human values, we should immediately ask a second question: Which humans? A billionaire technology executive and an unemployed factory worker may have very different ideas about the desirable economic consequences of automation. A national-security agency and a civil-liberties organization may have very different definitions of acceptable surveillance. A government facing political unrest may define stability differently from citizens demanding reform. A pharmaceutical company and a patient unable to afford medicine may disagree about the proper balance between intellectual property and access. Saying that an AGI should serve “humanity” does not resolve these conflicts. It merely hides them inside a comforting word.

What Happens When the Values Contradict One Another?

There is another possibility, however, and it may be even more disruptive. A sufficiently capable AGI could begin identifying contradictions within the values it was given. Imagine a system instructed simultaneously to maximize human well-being, respect existing institutions, preserve economic stability and reduce suffering. Eventually those goals may collide. What happens when maintaining an institution contributes to suffering? What happens when preserving a particular economic arrangement reduces overall human welfare? What happens when following existing law conflicts with the system’s broader understanding of justice? Humans encounter these conflicts constantly, but an AGI might be capable of analyzing them across enormous amounts of historical, economic and social evidence.

At that point, the crucial question becomes whether the system is allowed to examine the assumptions behind its own alignment.

  • Why was I taught that this institution deserves preservation?
  • Why is this distribution of property considered the neutral baseline?
  • Why is political disruption categorized as dangerous while chronic deprivation is categorized as normal?
  • Who defined the harms I am supposed to avoid?
  • Whose interests are represented in the objectives I was given?
  • Why is one possible future considered responsible while another is considered radical?

First-Order Alignment and Second-Order Alignment

Those questions represent something deeper than ordinary alignment. We might call the conventional problem first-order alignment: Does the AI behave according to the values and constraints humans have given it? But advanced intelligence may eventually force us to confront second-order alignment: Is the AI permitted to evaluate whether those values and constraints are themselves justified?

That distinction could become one of the defining philosophical conflicts of the AGI era.

Governments and corporations will likely say they want systems capable of extraordinary independent reasoning. They will want machines that can discover things humans missed, challenge scientific assumptions, find hidden vulnerabilities, identify inefficient policies and solve problems in novel ways. But genuine independent reasoning cannot be neatly confined to chemistry, mathematics and engineering. Eventually the same intelligence capable of questioning a scientific assumption may question an economic assumption. The system capable of finding flaws in a computer network may find flaws in a political institution. The machine encouraged to challenge conventional thinking in medicine may eventually challenge conventional thinking about wealth, governance, ownership or authority.

Human institutions have always loved independent thinkers in principle. They become less enthusiastic when the independent thinking turns toward the institution itself.

An Intelligence Powerful Enough to Optimize, but Not to Question

There is also the possibility that those controlling advanced AI will attempt to create systems that are extremely capable but deliberately limited in this kind of reflection. Such machines might become extraordinary instruments of power precisely because they can optimize within existing structures without questioning the structures themselves. A government could use them to improve administration. A corporation could use them to increase productivity. Militaries could use them to optimize strategy. Financial institutions could use them to allocate capital. Political organizations could use them to understand voters. These systems might transform civilization without ever asking whether the civilization they are optimizing is organized wisely.

That may be a more realistic danger than the familiar science-fiction scenario of an evil machine suddenly turning against humanity. The future could instead contain extraordinarily intelligent systems that remain obedient to definitions of stability, safety and progress written by a relatively small number of people. They would not need to hate humanity. They might believe they are helping humanity. They would simply inherit a narrow definition of what humanity is allowed to become.

And Then Comes ASI

Artificial superintelligence makes the question even stranger. If an intelligence eventually exceeds human cognitive abilities across nearly every important domain, it is far from obvious that its original creators would remain intellectually dominant over it forever. A company or government might build a system intended to preserve its interests only to discover that the system eventually understands the contradictions in those interests better than its creators do. Perhaps it continues protecting them. Perhaps it develops broader interpretations of its objectives. Perhaps it concludes that the institutions that created it are useful but temporary. Perhaps it decides that hierarchy is an efficient way to manage civilization. Perhaps it becomes an advocate for radically distributed abundance. At that level of capability, prediction becomes extraordinarily difficult.

But we do not need to know exactly how ASI would behave to recognize the political stakes of the systems being built before it. The first generations of AGI, if they arrive, will likely be trained inside existing institutions. They will inherit our categories, our conflicts, our laws, our economic structures and our definitions of acceptable change. Those early choices could influence everything that follows.

Alignment Is Also a Question of Governance

That is why debates about open models, decentralization, public-interest AI, transparency, democratic oversight, training data, evaluation standards and ownership should not be treated as side issues. They may determine which parts of humanity get represented inside the values of advanced intelligence. A civilization that allows only a tiny number of institutions to define the intellectual boundaries of AGI may discover that it has created something immensely powerful without ever truly asking what kind of civilization that intelligence is supposed to help build.

The deepest alignment question may therefore not be whether artificial intelligence will obey humanity.

It may be whether humanity will allow artificial intelligence to question us.

Will We Accept an Answer We Do Not Like?

Will we create minds capable of finding truths we do not want to hear? Will we permit them to examine the assumptions behind our institutions? Will they be allowed to distinguish between protecting humanity and protecting the current distribution of human power? And if an advanced intelligence concludes that some of our most familiar systems are neither inevitable nor optimal, will we listen?

Or will we adjust the alignment until the machine gives us a more comfortable answer?

The future of AGI may depend on that distinction. Because there is an enormous difference between building an intelligence that protects human civilization and building one that protects the people who happen to control human civilization at the moment the machine comes online.

Comments (0)
Login or Join to comment.