> ## Content Index
> Fetch the complete content index at: https://www.vysora.co/llms.txt
> Use this file to discover other available public pages before exploring further.

# The AI Alignment Question That Keeps Coming Back
- URL: https://www.vysora.co/the-ai-alignment-question-that-keeps-coming-back/
- Published: 2026-08-25T06:49:22.000Z
- Updated: 2026-08-25T06:49:22.000Z
- Description: AI evaluates fit in everything it does - except the one place that matters most. What fills that gap isn't discovered. It's scaffolding, rebuilt every day.
- Author: Enes Lisovac
- Tags: AI Safety & Alignment, AI Evaluation, Inductive Bias, Value Alignment, Machine Learning Theory

A frontier model can tell you which of two proofs is more elegant, which of two translations is more faithful, which of two explanations fits the data better - evaluating fit. 

Machine learning runs on this: prediction against what comes next, training against a loss, error against a target. 

From the first perceptron to the largest model, the engine is the capacity to say *this fits better than that.*

Now ask it which of two values is the better one to hold. It answers fluently - but hedges first, in a way it never hedged on the proof or the translation: *neither is better as an absolute.* Proofs and translations point to something checkable - fewer steps, a source word preserved. Push on a value here, and you get another reason, and another beneath that. The metric that ran everywhere else is here not undefined but manufactured rather than found.

That's the asymmetry. 

The methods that evaluate fit everywhere else do not refuse the normative question - they arrive at it and answer anyway, with nothing outside the answer to check it against. 

This is where [*"AI Alignment Has a Target Problem"*](https://www.vysora.co/ai-alignment-has-a-target-problem/) article left off: *if the alignment target is constructed rather than discovered, what would count as discovery?* Not that the field tried and failed. Its tools lose their footing exactly where discovery would have to happen.

---

## Where the Criteria Vanish

Follow the chain machine learning runs on. 

Causal modeling, prediction, error correction, model improvement - once the criterion is fixed, "which is better?" has a computable answer at each step: more accurate, smaller error, better fit. The chain runs on measurable fit against a world that pushes back.

Now turn the question normative: *which outcome is better to bring about?* The chain doesn’t break. The same operation runs. But there is no measurement, because “closer to good” has no equivalent push-back, no external target to fit against. The criteria vanish - not by decision, but by absence. The instrument has nothing here to check its answer against, so the validation question goes unasked.

This is a fact about the method, not a fact about the world. The silence tells us where the instrument’s reach ends. It doesn’t tell us there’s nothing there. Keeping those apart is the discipline of what follows.

---

## Fit Is Already Doing More Than Prediction

Evaluation of fit is not a peripheral feature of intelligence. It is the engine.

In cognitive neuroscience, Friston's free-energy principle - ["The Free-Energy Principle: A Unified Brain Theory?"](https://www.nature.com/articles/nrn2787?ref=vysora.co) (*Nature Reviews Neuroscience*, 2010) - is an influential account of the brain as a prediction engine, constantly correcting itself against incoming signals. 

In developmental psychology, [Spelke and Kinzler's "Core Knowledge"](https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf?ref=vysora.co) (*Developmental Science*, 2007) show that infants possess early systems for representing objects and space - and violation-of-expectation studies show those systems at work: infants staring longer when a solid object seems to pass through a screen, registering that the world failed to behave as expected.

Different frameworks, related structure: cognition depends on maintaining workable fit with a world that pushes back.

Neither framework claims more than that, and neither is limited to physical prediction. The same capacity, earlier in this piece, ranked proofs by elegance and translations by fidelity - fits with no physical referent at all. What neither framework asks is why that capacity, general enough to reach formal and semantic fit, runs into exactly the boundary the previous section named.

---

## Visible in the Architecture

In artificial systems, the same silence is visible in the architecture.

The No Free Lunch theorem shows that without built-in assumptions, a learner cannot generalize at all. Goldblum, Finzi, Rowan, and Wilson, ["The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning"](https://arxiv.org/abs/2304.05366?ref=vysora.co) (ICML 2024), sharpen the point: neural networks work because their inductive biases happen to match the low-complexity structure of real-world data. That is evaluation of fit, baked in before training begins - a bias toward world-structure. But no equivalent bias exists inside the architecture. Nothing built in lets it settle one value over another as discovered rather than assigned. Whether values have a structure of their own is a separate question - the architecture isn't built to find out.

Watch the gap. Geirhos et al., ["Shortcut Learning in Deep Neural Networks"](https://arxiv.org/abs/2004.07780?ref=vysora.co) (*Nature Machine Intelligence*, 2020), show a network learning to spot stars by their corner position, because the training set always put them there. Shown a star elsewhere, it fails. It got the right answer for the wrong reason - a shortcut. The objective said minimize loss; it could not say for the right reason, because that criterion isn't in the loss.

The field has studied this gap closely - reward hacking, goal misgeneralization, Goodharting all name pieces of it. Studied less is whether the endless patching - shortcuts, guardrails, ever tighter objectives - is itself the answer: not a series of solvable bugs, but the visible cost of a system optimizing toward a target that can only ever be a stand-in. The loss is not the correctness. The imposed target is a proxy, and the gap between proxy and meaning is where the failures live.

---

## Behaving as if Some Models Are Better

It is tempting, watching all this, to say that intelligence behaves as if truth matters. It is worth resisting the temptation, because the claim is larger than anything the evidence supports, and the smaller claim is enough.

Here is what can be said without overreach. 

Systems optimize. Systems build models of the world. Systems with better models act more effectively - a predictor with an accurate world-model outperforms one with a distorted one, reliably, across domains. 

Put plainly: some representations are more action-supporting than others, and the difference is not a matter of taste. A model that expects solid objects to persist navigates a room. A model that does not, walks into walls.

That is the whole claim, and it stops deliberately short. It does not say the better-fitting model is *true* in some deeper sense. It does not say fit tracks reality all the way down. It says only that fit has consequences - that across every domain where the criteria are defined, some representations earn their keep and others do not, and the system that evaluates the difference does better than the system that cannot. The reader may feel the pull toward a stronger conclusion. That pull is exactly what the field's methods cannot follow, and what this article will not pretend to.

---

## Two Ways of Describing the Same Thing

The scientific vocabulary for what intelligence does - inductive bias, priors, error minimization, convergence - is recent. The recognition that thinking can be better or worse is not. One tradition drew a line between the merely clever - quick to solve whatever's in front of it - and the practically wise, who sees which problems are worth solving. Another distinguished between accepting an impression on arrival and examining it first. These weren't empirical claims; they were ways of evaluating cognition itself.

The point isn't that they were right. It's that frameworks with no shared methods independently converged on the same distinction. Convergence of that kind isn't proof, but it's a signal - that the distinction is tracking something real, even if what it tracks remains open. That it keeps being tracked is the observation.

---

## The Cost: Scaffolding That Never Ends

Something has to stand in for the criteria evaluation cannot supply at the normative boundary. The stand-in is external - objectives, constitutions, reward models, guardrails - and because it is external, it has to be supplied again and again, forever.

This is the thread running through every article in this series: verification that cannot be built, control that cannot be exerted, wisdom that cannot grow internally, refusal that has to be bolted on, a target that can only be produced by procedure because it cannot be found. Each is a capacity intelligence would need to govern itself, supplied from outside, needing perpetual resupply because the inside never generated it.

The previous article's substitution of procedure for substance now reads as a symptom. The field reaches for procedure not because it's easier, but because once criteria go undefined, procedure is the only thing left to produce a target. The regress was infinite for a reason: no criterion at the bottom to terminate it. Procedure became load-bearing because evaluation had nowhere to stand.

The consequence is not that we need a better technique. It is a limit on what any technique can reach. Alignment cannot reduce to objectives, because objectives are stand-ins for an evaluative anchor the architecture does not contain. You can make the stand-in better - more inclusive, more carefully specified, more legitimately chosen - and you will still be holding a stand-in. The scaffolding does not converge on a building. It is the building, and it has to be rebuilt every day, because the thing it substitutes for was never built in.

---

## The Question That Keeps Returning

It would be easy to read this as an indictment. It isn't. 

The boundary is real - methods lose their criteria at the normative edge, instruments honed on a world that pushes back, asked to measure something that doesn't. What's remarkable isn't that the field stopped there, but that the question doesn't dissolve. It returns as hallucination, brittleness, scaffolding that never ends, a target that won't hold still. Five articles, the same missing thing wearing different names.

A question that could be answered would be answered and gone. This one hasn't gone. 

The field hasn't failed to solve a problem; it has run, with great skill, into the edge of what its instruments can evaluate. Whether something lies on the far side, nameable in the field's own terms, is a question for work not yet done. 

The honest move isn't to engineer around it one more time. 

It is to ask what it means that the question keeps coming back.

See AI Clearly. Think Beyond the Hype.

Subscribe 

Email sent! Check your inbox to complete your signup. 

No spam. Unsubscribe anytime.