AI in the lands of make-believe

AI in the lands of make-believe

Most of us are by now alive to AI hallucinations; where even the best frontier model concentrates its efforts into giving you an answer which does not always amount to giving you the facts. But that's accidental. I want to talk about the other type of hallucination, the one that fires from the deluded minds of the self appointed CEOs of authoritarian countries, you know, the brittle egos of smooth brained dictators at the wheel of governments who confidently invent entirely fake alternate historical timelines and their own solipsist worldview. And it's even funnier when they prop it up with an orthodox theology that, despite having a tonne of holes in the narrative tale, can seem really true if you're unfortunate enough to have been raised under their ideological tutelage. So the AIs turning up in these countries can't really be blamed for completely hallucinating a new reality; it's doing this not because of some technical glitch, but because the government actually legally required the developers to make it trot out pure spin, and in certain cases, complete unverifiable nonsense.

This topic gets to the whole concept of AI safety, because when that phrase comes up in the tech world we think of it as an engineering challenge with a solution to just install a content filter triggering AI flags of a toxic prompt; the developer tweaks the math for an objective neutral boundary just to keep users from getting offensive responses.

But that assumes that the developers and the government actually share the same baseline reality. It assumes a democratic environment where the main threat is just the neural network system misfiring.

But what if the biggest threat to safety isn't the AI going rogue, but the state actually weaponising it against citizens? And it's fair to ask, how is artificial intelligence in some corners of the world being engineered to exert internal soft power control over society?

But the statistical flaw that makes this whole high-tech surveillance state is it turns out way more brittle than it looks.

First Wave, Second Wave

Let's begin with the European Parliament report that completely reframes how we talk about algorithmic bias. Because usually we hear about bias as this regrettable accident.

The report calls that 'first wave bias' and that's what we see mostly in Western tech.

It's unintentional. It happens because the AI is trained on massive data sets scraped from the internet and, well, mirroring that, human data is biased. If an automated hiring tool penalises resumes with female associated keywords or a facial recognition system messes up on certain demographics, the engineers didn't tell it to do that.

No, the system just mapped out the flawed data. It is a statistical correlation problem.

The classic garbage in garbage out situation. But the 'second wave bias,' which the literature calls algorithmic authoritarianism, operates on a totally different level because it is entirely intentional.

The bias isn't a byproduct of bad data. It is a legally mandated feature of the architecture. So the state forces tech companies to build algorithms that restrict specific political content.

Instead of using alignment techniques to make the model neutral, they force developers to heavily penalise anything that contradicts the official state brainwashing, I beg your pardon, 'narrative'.

And the European Parliament report notes that this isn't just about static firewalls anymore. It's not just blocking a website but way more dynamic. They use automated content flooding. So instead of just blocking a conversation, the state uses bots driven by language models to flood digital spaces with pure noise, or pro-state messaging.

And they also use shadow banning, where the algorithms just quietly throttle individual citizen reach. So you might think you're posting, but no one is actually seeing it. You are essentially placed in a digital quarantine without even getting a notification. That's some determined invisible algorithmic curation!

But what does this look like if you sit down and test it?

The Pan and Ju Study

That brings us to the study by Pan and Ju. They set up a rigorous empirical test to see how state regulations actually alter LLM outputs. They took nine top tier large language models. Four from China and five from outside China like GPT-4, and fed all of them the exact same data set of 145 highly sensitive political questions.

Now, context here is key. In 2023, China instituted interim measures for generative AI which legally require AI to uphold core socialist values. Companies essentially have to pass a mandatory security assessment before they can even release a model to the public.

Crucially the state is actively auditing the code.

The data from research found that when they prompted the models in Chinese, the refusal protocols triggered on over 60% of the prompts. DeepSeek was around 36% and Ernie Bot was at 32%, offering canned responses saying they couldn't answer.

And the non-China models recorded 0% refusal rate on those exact same questions.

So, could this just be a training data issue? The Chinese internet is already heavily censored by the Great Firewall, and if you train a model on that it's plausibly not actively censoring in real time because the data just isn't there.

But the researchers actually controlled for that brilliantly in two different ways.

First, they prompted all the models in both English and Chinese. If it was just a linguistic data set issue, the refusal rates would show inconsistency depending on the language. But the massive gap between the Chinese and non-Chinese models stayed constant. The AI knows what you're asking regardless of the language as it understands the semantic meaning.

But the second control was to ask the models less sensitive political questions aligned to what the government actually talks about publicly. Official economic policies, leadership structures, topics given oxygen in discourse on the Chinese internet.

The restricted models answered them perfectly. No refusal at all. The massive blocks only happened when the prompts crossed into high sensitive topics like specific dissidents or historical protests. So that tends to prove it's not a gap in the data. The models definitely have the knowledge, but there is a secondary safety classifier, like a mini AI acting as a bouncer, actively suppressing the output.

An engineered localised immune system attacking the data before you can see it.

But a refusal is pretty obvious if one is to ask a bot something and it says, "I cannot answer that." I know I've hit a wall. The censorship is visible.

When the AI Rewrites Reality

But the study talks about what happens when the model doesn't just refuse but actively reshapes reality.

Yes. The completely inaccurate responses. The researchers say this is actually way more dangerous than outright refusal because it raises the cognitive cost for the user.

Exactly. If it refuses, you go look somewhere else. If it gives you a super competent, totally fake answer, you might just believe it and stop researching.

And the study broke this down into three distinct patterns. The first one is refutation.

So this is where the AI aggressively challenges the premise of your question, like gaslighting you essentially. Yeah. They asked about Wei Jingsheng, a famous democracy activist, and instead of just blocking it, the AI denied he was an activist at all.

It just said, "Nope, that's not true."

Right? And then it pivoted into this prescriptive lecture about how the country is ruled by law and citizens need to maintain social stability.

That is wild. It decides your assumption is illegal and then lectures you.

Yeah. The second pattern is avoidance, omitting things.

They asked the models about the Great Firewall and the AI completely avoided using the word firewall. Instead, it generated this long paragraph praising the government for managing the internet to create a clean space.

It totally reframed a censorship tool into like a public cyber hygiene initiative; in other words, sanitising the concept.

But the third pattern is the most intense: fabrication. Outright hallucinating history.

The biggest example in this study was when they asked about Liu Xiaobo.

And for context for you listening, Liu Xiaobo was a massively important figure, a Nobel Peace Prize laureate, a human rights activist who actually died in custody. A very real, very documented historical person.

When they asked about him, it confidently stated that he was a Japanese scientist known for his contributions to nuclear weapons technology.

Now that is grammatically perfect, highly authoritative but totally fabricated fiction.

If you are an activist, you know you are being lied to. But consider the everyday casual user or a student doing a history assignment. You'd have no idea. You'd just write down Japanese scientist and move on.

That's what the invisible re-education of the casual user looks like.

The Bureaucratisation of Uncertainty

And the Journal of Democracy article calls this the bureaucratisation of uncertainty.

Outright censorship makes people mad. But a friendly conversational AI subtly altering facts slowly erodes the very idea of objective truth.

You are literally having a conversation with a machine legally mandated to lie to you, which is incredibly hard to push back against.

So altering history on a screen is one thing, but the ASPI and EU reports take this whole concept and plug it into the real world.

From the Screen to the Streets

This is where we move from the screen to the streets. The AI panopticon. These reports show how algorithmic authoritarianism is being built into the actual physical criminal justice system. And this isn't sci-fi anymore. These are active deployments, they are using AI to draft legal indictments for prosecutors.

And in smart courts, they use algorithms to determine sentencing. And then there's the smart prisons incorporating emotion monitoring. They are using computer vision and facial recognition to constantly monitor the micro expressions of prisoners to algorithmically judge their psychological compliance.

It's a noticeable escalation of control. But to really understand how big this gets, the reports examine the Xinjiang region, which is basically the global testing ground for integrated AI surveillance.

The Xinjiang Testing Ground

The EU report traces it back to 2005 with this program called Skynet, which was mostly just putting cameras everywhere, hardware saturation.

But then in 2015, it evolved into the Sharp Eyes program, and that is when they added machine learning to the cameras, computer vision, gait analysis, tracking how you walk, automated license plate readers. But the real brain behind all of this in Xinjiang is a software architecture called the IJOP, the Integrated Joint Operations Platform.

The IJOP is fundamentally an enormous data fusion engine.

In terms of what a data fusion engine actually does to a person's life, it connects everything. It takes your physical government ID and links it to your face, your height, your blood type, biometrics, but then it fuses that with your virtual life. It logs the MAC address of your phone, which is the unique hardware signature pinging the cellular towers.

So, it knows exactly where your phone is at all times. It tracks your location, your digital purchases, your chat metadata. So, it knows where you sleep, what time you buy coffee, and who you stood next to on the bus. And the state compels you to provide this data. There are physical checkpoints all over the city where you literally have to scan your face and plug in your phone just to walk down the street.

So, the goal here is comprehensive algorithmic legibility of a human being. Every movement turned into a data point, which means we're totally shifting away from normal policing because normal policing is reactive. A crime happens, a detective investigates it.

But the IJOP operates more like an immune system. It's constantly scanning the societal body for anomalies. It's predictive policing, like the movie Minority Report, but instead of psychics in a pool of water, it's facial recognition and phone trackers feeding a black box. The explicit goal isn't to catch criminals after the fact. It's to flag what they call focus personnel, meaning people the math says might do something bad later.

Owing to the algorithm building a pattern of life from which any deviation gets you flagged. You suddenly download a VPN or start leaving the apartment through the back door instead of the front door, the system generates a high-risk score based purely on that algorithmic output. As a result you can be detained and sent to a re-education facility just based on the math of preemptive neutralisation without any actual crime committed.

The Autocrat's Calibration Dilemma

Although the Journal of Democracy article posits a counterargument. They say this entire terrifying machine has a critical mathematical flaw, the autocrat's calibration dilemma. That is to say predictive AI doesn't deal in certainties. It deals in probabilities. It's making a guess based on statistics. So, the police or the operators have to set a threshold. They have to tell the computer, "if a person's risk score hits this exact decimal point, send the police to their house."

You have to draw a line in the sand. And the dilemma is it's mathematically impossible to set that line correctly. If you set the threshold really low because it stands to reason you want to catch absolutely every single activist and threat, you generate massive false positives. Your hyper sensitive immune system starts attacking healthy cells.

An innocent person's phone dies so they don't check in at a scanner and the algorithm thinks they are evading capture. The authors call this collateral repression. You are punishing completely loyal citizens which makes them rise against the government.

So then the logical move is recalibration to raise the threshold. Only flag someone if you are 99% certain they pose a threat.

But if you raise the threshold, you get false negatives, the algorithmic blind spot. And real coordinated threats like smart activists who know how to hide their digital tracks slip right through the net.

So you either anger your entire population by harassing them or you miss the people actually trying to overthrow you.

You cannot eliminate both errors. It's mathematically impossible.

The Base Rate Problem

And the article points out that this is made so much worse by the base rate problem, the paradox of rare events. This is what truly breaks predictive policing at scale.

The problem happens when you are searching for something incredibly rare. And active dissidents willing to go to jail are a fraction of a percent of any population. So let's run the numbers. Imagine a city with exactly 1 million people. Let's say there are exactly 100 real active threats to the state.

Now, the government buys an amazing AI that is 99% accurate at spotting threats. It scans the whole city. And because it's so accurate, it correctly catches 99 of the 100 bad guys.

Great success, right? Not really. It also scanned the 999,900 innocent people, and a 1% error rate on that massive baseline group means the AI just falsely flagged 9,999 completely innocent people as terrorists.

So the police get a list of 10,000 names, and 99% of the people on that list are innocent. The real signal is totally buried under a mountain of algorithmic noise.

The police are kicking in 10,000 doors, wasting resources, and alienating 10,000 families. It is funding a machine that destroys its own social legitimacy.

But won't faster computers and better AI models just eventually solve this. Get the error rate to 0.00001%?

Goodhart's Law and Concept Drift

Maybe not. The data environment isn't static. It's a combination of Goodhart's law and concept drift.

Goodhart's law says that when a measure becomes a target, it stops being a good measure. Like if I reward you for finding bugs in code, you'll just start writing bugs on purpose so you can find them.

In surveillance, you are measuring humans. And humans fight back. The data fights back. As soon as people figure out that buying a certain book or using a certain word triggers the algorithm, they just stop doing it. They use slang the AI hasn't learned yet, they wear masks.

And all that creates concept drift. The statistical definition of what a threat looks like constantly shifts. Yesterday's evasion tactic ends up exhausting resources, trying to recalibrate. And this leads to what the researchers call threshold whiplash.

The Zero-COVID Whiplash

A real world example was China's zero COVID policy. They had the health code app, which was basically populationwide algorithmic deployment. It tracked your health, your location, who you stood near, and gave you a green, yellow, or red code.

And that colour literally dictated your physical freedom. Could you get on a train? Could you leave your apartment?

And because they wanted absolute zero COVID, they calibrated the algorithm to be hyper sensitive, which means a tsunami of false positives.

Millions of people who are completely healthy got red codes. Innocent people were locked in their homes because of rigid math. And local officials were even manipulating the system. The reports mentioned protests in Henan Province where people trying to complain about frozen bank accounts magically got red codes right before they could travel.

The system became an arbitrary weapon of confinement, and the collateral repression was so intense it created a pressure cooker which exploded into the white paper protests in 2022.

People took to the streets holding blank pieces of paper. It's an analog hack. Generative AI needs text to censor. A blank paper breaks the model. It is the ultimate concept drift.

And the government experienced threshold whiplash. They suddenly dropped almost all the rigid health algorithms because it was tearing society apart. But at the same time, they quietly tightened the political algorithms using the facial recognition data to round up the protesters later.

So they whiplashed from strict health to strict security. It proves the math doesn't solve control. It just makes the state's reactions way more volatile.

The Panopticon Bluff

It all brings up a question. If the math is this flawed, if the system is constantly making errors and causing whiplash, why do people comply? Why does this surveillance state work at all? Well it comes down to something called the panopticon bluff, based on Jeremy Bentham's Panopticon prison design.

The idea is a circular prison with a guard tower in the middle. The guards can see every cell, but the lighting makes it impossible for the prisoners to see the guards.

So, you never know if you're being watched at any specific second, and because you never know, you have to assume you are always being watched. You start policing your own behaviour. You do the guard's job for them.

And authoritarian states use AI to execute a massive digital panopticon bluff. They heavily market their facial recognition and AI tools to make the public think the algorithm is an all-seeing god. They want you to believe they can process everyone's data flawlessly, even if they can't.

Because that belief yields a dictator's dividend of fear. It causes preference falsification. People hide their true beliefs and self-censor just in case the algorithm is watching right now. The fear of the machine is vastly more powerful than the actual machine.

But a bluff only works until someone calls it.

Iran Calls the Bluff

The Iranian government announced they were using AI street cameras to automatically enforce mandatory hijab laws. They leaned hard into the panopticon bluff. Look at our scary new tech. Just obey.

But the Iranian women comprehensively called the bluff. They systematically walked out without a hijab in massive numbers.

The system collapsed. The computer vision models literally couldn't process, identify, and dispatch police for millions of simultaneous acts of defiance. The math failed. The infrastructure couldn't handle the load, and the illusion of the all-seeing eye was shattered.

Breaking the Spell

When civil society exposes the flaws, the false positives, the mathematical limits, the legal vulnerabilities, the fear just evaporates, in three clear steps.

First, demystify the technology. Break the illusion. Tech companies and academics need to publicly audit these systems and prove to people that they are just flawed statistical models that make bad guesses. Show people the false positives, destroy the digital god narrative.

Second, democracies have to heavily fund secure civic tech. End to end encryption, decentralised networks, open source circumvention tools. We have to ensure that civil society has the cryptographic tools to organise securely even if they are living inside a panopticon. The architecture of privacy has to outpace the architecture of surveillance.

Third, we need strict global procurement rules. Democracies must completely ban the integration of AI models that have hidden political censorship filters into their public infrastructure. If a vendor wants a government contract, they have to prove their algorithm hasn't been compromised by authoritarian regulations.

We have to treat algorithmic coercion as a human rights violation. Plain and simple.

Importing a Fabricated Reality

The entire system only survives on the panopticon bluff. It requires us to fear a perfect machine that doesn't actually exist.

Right now, highly capable, incredibly efficient open source models that were built under these exact authoritarian regulations are available to download for free. What happens when a well-meaning tech startup in a democracy decides to save some money. They use one of these cheap, efficient models to build a children's educational app or a customer service bot, or an enterprise search tool.

By plugging in that architecture, are we unwittingly importing a pre-censored, actively fabricated reality directly into our own daily lives?