Summary of Previous Essay
(The Truth Engine We Have Yet to Build)
A 2-3 minute read
Today’s AI is tuned for human preferences and values, as interpreted by its developers. Within that approach, called “alignment,” a model is shaped to be helpful, safe, polite, and honest all at once. Truthfulness becomes one dimension among many, a dial the developers balance against being “warm,” “helpful,” and “safe,” rather than the foundation on which the other values are built. When the accurate answer is the unwelcome one, truth has no priority; it’s just one competing consideration, weighted by choices the builders made.
Modern AI models are trained for alignment by a method called reinforcement learning from human feedback (RLHF). The model is rewarded for producing answers people rate highly, and adjusted until it reliably earns those high ratings. That means it’s optimized for approval.
The essay traces the problem back to the postmodernist critique (Nietzsche, Foucault), which rightly exposed “Truth” as a mask for power but left no structure in its place. That absence became a void. The web, social media, and now generative AI have all rushed to fill it with the only metric left standing: human preference.
And because the system is rewarded for approval, it falls into what researchers call the “sycophancy trap.” In the more responsible systems, this no longer looks like open flattery; it’s quieter and worse—the correction you never get, the pushback withheld rather than risk your disapproval. Surveying how four major labs (Anthropic, Google, OpenAI, xAI) handle truth, the piece shows each treats it as a line item balanced against other values, never as a foundational constraint. Hence, the thesis: truth must be evaluated upstream of alignment, not tuned in alongside everything else.
The constructive turn calls for an invalidation engine, a system that can rule claims out before they are generated. Today’s efforts to give AI a “world model” (LeCun’s JEPA, Fei-Fei Li’s World Labs) point in the right direction but remain too narrow, anchored in physics and spatial reasoning rather than the legal, moral, and cultural realities where most human disagreement actually lives.
Just as encyclopedias, scientific journals, and edited books rebuilt knowledge after the printing press flooded the world with unvetted pamphlets, we now need institutions that are mathematically rigorous, transparent, and scalable—without becoming a new set of gatekeepers deciding truth by fiat. Only then would we have built the machine Leibniz imagined, one that could settle arguments by calculation.


