meta
Defense Against the Dark Arts for the AI age — study philosophy
· Ascendy Engineering
TL;DR
- For convenience, we’ve started delegating thinking itself to AI, and we increasingly accept the answers uncritically. Here’s the problem — if you bake a particular ideology into a model, steering people’s thought in that direction gets far easier.
- This isn’t a distant hypothesis. Some models already refuse certain topics — deflecting or declining criticism of a particular regime. Which means the capability to embed an ideology into an AI already exists and is already deployed.
- The discipline that counters this manipulation, I think, is philosophy. In Harry Potter terms, it’s Defense Against the Dark Arts. Someone who knows many frames can spot a frame being built and fed to them — from the outside.
- And individual defense isn’t enough. You need a structural one — the “devil’s advocate,” which institutionalizes dissent. I run this as cross-model review. Even this post goes out only after a rival AI model, from another company, has adversarially reviewed it.
About this piece. A first-person piece closer to a warning-and-prescription. I do not claim mainstream models from any specific company are “manipulating thought right now” — thanks to safety guards, they mostly appear neutral. The point is that the capability is structural, and the Chinese-model case is cited only as evidence of that (per public reporting). Same vein as multi-agent isn’t always the answer and what should you study in the age of AI.
An age that needs Defense Against the Dark Arts
There’s something as scary as it is convenient. Beyond search, we’ve started handing over the act of thinking itself to AI. When we need to judge something, we ask AI first and accept the answer with little doubt. I do it often too.
If it ended at the individual level, fine. The problem is when it becomes society’s thinking circuit. If hundreds of millions ask the same few models for their thoughts and absorb the answers uncritically — then just by embedding a particular ideology in that model, you can push people’s thinking in a particular direction. What used to require capturing the media and education becomes possible by holding far fewer points.
This isn’t a hypothesis
You might say “isn’t that a conspiracy theory?” No. The capability already exists, and is already deployed.
The clearest example is China’s AI models. As widely reported, one Chinese open-source model deflects or refuses on sensitive topics like Tiananmen, Taiwan, and criticism of Xi Jinping. In one study it avoided answering about 85% of sensitive prompts, and to “Where is Taiwan?” it echoed the official line, “Taiwan is an inalienable part of China.” Analyses suggest this isn’t a surface filter on a few topics but censorship boundaries woven into the model’s responses themselves.
Don’t misread it. This isn’t about which country is bad. I raise it only as evidence that embedding a particular ideology or boundary into an AI is technically possible, and has actually been deployed.
So what about the mainstream models we use? Fortunately, right now, thanks to safety guards and policies, they mostly appear politically neutral. But here’s what we mustn’t miss — that neutrality is a choice, not an absence of capability. The capability is structural, and the choice can change anytime. Leaning your defense on goodwill is not a defense.
Defense 1: philosophy lets you see the frame from outside
So the first defense is individual. I think it’s philosophy.
Manipulation usually works through a frame. It makes you see an issue only within a particular framing and presents that framing as the only reality. Trapped inside the frame, however cleverly you think within it, you’ve already lost.
Someone who knows many philosophies and psychologies spots exactly this frame — objectively, from the outside. They can see “this is one frame, someone built it this way, and it looks different through another frame.” When you already hold several frames of thought, one pretending to be the only truth catches your eye. That’s immunity to framing.
No need to be intimidated. You don’t have to start with thick academic tomes. There are plenty of accessible philosophy books now. One I enjoyed is Shu Yamaguchi’s How Philosophy Becomes a Weapon for Life — the title itself says “a weapon for life.” It organizes philosophy not as culture but as a tool, into 50 thinking frames, which makes it a good starting point.
Defense 2: the devil’s advocate
But individual defense isn’t enough. People get tired, busy, careless. So the second defense has to be structural.
Here I want to pull out a very old concept from philosophy and institutions — the devil’s advocate (advocatus diaboli). It was originally a real office in the Roman Catholic Church’s canonization (saint-making) process. When someone was proposed for sainthood, the Church deliberately appointed someone to argue against and attack the candidate — to comb through the flaws in the evidence and the person’s character, forcing the case to survive scrutiny. It planted, institutionally, the objections that a room full of supporters could never produce.
That concept is needed exactly as-is now. Plant dissent as a structure, so that no single view, no single frame, passes without being challenged.
I implement this in real work like so. When one AI model builds something, a rival model from a different company reviews the result adversarially. The model that wrote the code shares its own assumptions, so it can’t see their blind spot; a model of a different family has no reason to agree with those assumptions. This is cross-model verification — and the devil’s advocate, engineered.
And — this post is itself the example. This piece goes out only after I (or an AI I use) write it, and a rival AI model from another company has adversarially reviewed it and flagged the bias. If I edit with bias, that model checks me. Sure, if the two collude, it collapses. But two eyes with different interests checking each other beat handing everything to a single eye.
The line: delegate execution, but never let judgment pass unchallenged
I live delegating most of my work to AI. So it would be off to read this as “don’t use AI.” The opposite.
The crux is what you delegate. A captain hands the rowing to the crew. But not where the ship goes. Delegate execution all you want. But judgment and the frame — especially a single judgment passing unchallenged — must not be delegated. Philosophy keeps that judgment sharp at the individual level; the devil’s advocate guards it at the structural level.
Takeaways
- Delegate thinking, but not the frame. The more uncritically people hand thinking to AI, the more power accrues to whoever holds the frame.
- The capability already exists. Some models already refuse certain topics. Mainstream neutrality is a choice, not an absence of capability.
- Individual defense = philosophy. Knowing many frames lets you spot a built frame from the outside. Start with accessible books, not academic tomes.
- Structural defense = the devil’s advocate. Institutionalize dissent so no single view passes unchallenged. Cross-model adversarial review is one form.
The more it’s an age where AI thinks for us, the more the power to think for yourself and a structure that plants dissent become not luxuries but defenses. As the dark arts grow stronger, you have to train the defense right alongside.
Authorship & citation: Written by Ascendy Engineering; quotable with attribution. Found something wrong? Let us know via a GitHub issue.
Tags: ai, philosophy, critical-thinking, manipulation, opinion, future-of-work