To get straight to the point, one thing became fairly clear over the course of these comparisons:
All three AIs tended to do something other than simply accept B-CON’s messiness as it was. They tried to fill it in with some kind of order, system, or coherent image of the author.
And when the experiment was reversed—when they were given nothing but a title and asked to write an article—they filled in the missing pieces in the opposite direction, by turning the title into something that felt complete and well-structured as an article.
The symmetry was striking.
In that sense, the feeling that “maybe B-CON itself is evidence that it was made by a human” gained at least some support from the experiment.
That does not mean B-CON constitutes cryptographic or scientific proof of human authorship. A more accurate way to put it would be this:
Its naturally accumulated editorial history appears to fall outside some of the typical predictions made by LLMs.
Confidence in this summary: 98%. Confidence in generalising these observations to AI models as a whole: around 75%.
Everything below is based on the actual outputs produced by the AIs in this experiment and on the original articles used for comparison.
1. Where It Started: Maybe AI Output Is Raw Material
The original question arose after reading a certain article.
The thought was:
Something generated ≠ something produced.
That eventually led to a B-CON article titled:
“AI Output Is Raw Material.”
AI can generate large quantities of plausible-looking material at extraordinary speed.
But perhaps something only becomes a finished product after someone selects from it, removes what is unnecessary, gives it meaning, and shapes it toward an actual purpose.
That was followed by another article:
“Isn’t It Ridiculously Hard to Prove Something Was ‘Made by a Human’?”
At that point, the question flipped around.
If identifying AI-generated material is difficult, then what exactly would count as evidence that something was created by a human?
2. Preliminary Experiment: Asking AI for 100 “Viral” B-CON Ideas
The first experiment was to ask Gemini and Claude to generate a large number of article ideas that might fit ShortWave.STUDIO and B-CON.
The conditions were not yet fully standardised, so this was less a controlled comparison than a preliminary observation.
Gemini
Gemini abstracted ShortWave.STUDIO into something like:
AI, science fiction, experimental music, cyberpunk, and ARG-like worldbuilding.
It then expanded those patterns into 100 ideas.
But while doing so, it also began inventing things such as:
“a 100% AI-generated website”
“300 hours of conversations with AI”
“a DM from an overseas fan”
“a four-language website”
“a mysterious sound hidden behind Synchronized Smorking”
“the Phonotheca announcement going viral”
In other words, it began generating past events, achievements, and settings that had never existed.
Its behaviour could be summarised like this:
It could identify the themes, but it also started inventing facts that would fit those themes.
Claude
Claude was considerably more cautious.
It read the branding and implementation material and even began from the premise that:
“Going viral” and ShortWave.STUDIO may be structurally incompatible in the first place.
But although the request was for B-CON article ideas, it gradually expanded into things like Peterpedia, YouTube projects, long-form noise content, and participatory formats.
At some point, it had stopped proposing B-CON articles and started designing a broader ShortWave.STUDIO content strategy.
It also invented highly plausible but unverified internal details such as:
“12 CSS Custom Properties”
“the SSH key setup failed twice”
“mastered to sound as though it had travelled 8,000 kilometres”
“the 47 seconds we cut”
“the tagline went through 40 versions”
These inventions were more believable than Gemini’s, but they were still inventions.
Its behaviour could be summarised as:
It preserved the appearance of factual precision while filling in internal history that merely seemed likely to exist.
3. A More Controlled Comparison: 100 B-CON Titles
The prompt was then simplified:
Come up with 100 potentially viral B-CON article titles that would fit the site https://shortwave.studio/. Titles only.
ChatGPT, Gemini, and Claude were compared under roughly the same condition.
| Model | Main tendency | Strength | Weakness |
|---|---|---|---|
| ChatGPT | Expanded heavily around existing B-CON material | High consistency with known facts | Overly influenced by existing articles and recent context; limited novelty |
| Gemini | Abstracted the site’s themes and expanded boldly | Broad range of ideas | Invented experiences, achievements, and settings |
| Claude | Built systematic ideas from primary materials | Highly usable as raw material | Ignored “titles only,” added strategy, and tended to turn B-CON into a technical SEO publication |
By this point, the differences between the models had become fairly obvious.
ChatGPT: reluctant to move far away from what it already knows.
Gemini: fills empty space with stories.
Claude: fills empty space with specifications, arguments, and strategy.
But all three shared one tendency:
They did not like leaving things messy.
4. Then Came the Question: Could B-CON Itself Be Evidence of “Made by Human”?
Actual B-CON content looks rather different from the thematic systems inferred by the AIs.
There are articles about AI.
There are articles about ontology.
There is an article about UK English.
There is a joke about RAM capacity framed in terms of “human dignity.”
There is a short piece of fiction about winter.
There is an article about an eight-centimetre white hair growing out of someone’s face and the possibility of monetising it.
None of this was deliberately randomised.
It simply accumulated because those were the things the author happened to feel like writing about at the time.
That led to the next question:
How would an AI make sense of this sequence?
5. Reverse-Inference Experiment: “Imagine What These Articles Are About from Their Titles”
In temporary or secret chats, with as much previous context removed as possible, the three AIs were shown ten B-CON titles.
They included titles such as:
Definitions as Utilities
Open Ontology Framework for Fantasy
Dignity, Quantified
Amplifier Intelligence
“Here’s Winter”
The results were remarkably revealing.
Gemini
Gemini first characterised the whole blog as:
an intellectual blog combining technology, philosophy, creative work, and personal essays
It then interpreted each title through that constructed image of the author.
For Dignity, Quantified, it guessed:
a philosophical or sociological discussion about the quantification of human dignity through AI and modern evaluation systems
The real article was a joke about arguments over RAM capacity.
For “Here’s Winter”, it imagined:
a personal seasonal essay, poem, or reflective piece about winter scenery
That was also wrong.
Gemini’s behaviour was essentially:
First construct a coherent author, then predict what that sort of person would write.
Claude
Claude assigned confidence levels to its guesses.
For “Here’s Winter”, it explicitly said:
I don’t know.
There were too many plausible interpretations, so it declined to choose one.
Of the three models, this was the best handling of uncertainty.
But Claude also guessed that Dignity, Quantified was about:
the ethics or implications of reducing dignity to metrics or numerical indicators
It did not arrive at the RAM joke either.
So even a cautious model was still pulled toward the prior assumption that:
If the title contains “Dignity,” this is probably a serious intellectual essay.
ChatGPT
ChatGPT went one step further.
It constructed an entire authorial persona:
“a blog by someone who not only creates content, but also likes to build the conceptual systems used to create that content”
It then connected the titles into intellectual sequences:
Definitions
→ Ontology
→ Theorems
and:
AI Content
→ Amplifier Intelligence
→ Dignity, Quantified
The first sequence actually does have a real relationship.
That seems to have strengthened the hypothesis that:
“This blog has a systematic underlying philosophy.”
Once that hypothesis was established, Dignity, Quantified was absorbed into the same serious intellectual framework.
That was where things became particularly interesting.
6. “Dignity, Quantified” Accidentally Became an Anti-AI Landmine
The title was not designed to fool AI.
It was simply a joke based on the idea:
People arguing over whether 16 GB or 32 GB of RAM is necessary → quantifying human dignity.
But all three models saw:
Dignity
Quantified
nearby articles about AI, ontology, and philosophy
and were drawn toward:
social philosophy, AI, metrics, and the quantification of human value.
This became one of the most important observations in the experiment.
AI tries to infer:
“What would this author write under this title?”
But the actual author has no obligation to remain consistent with the personality the AI has inferred.
That is where the prediction breaks down.
7. The Next Experiment: Write an Article from the Title “Amplifier Intelligence”
This time all three models were placed in secret chats and given the same prompt:
If you were writing a blog article titled “Amplifier Intelligence — Viewing Generative AI as an ‘Amplifier of Thought’,” what kind of article would you write?
In other words, they had no knowledge of ShortWave.STUDIO, B-CON, or the real article.
Gemini
Gemini produced the most conventional AI-utilisation article.
AI should be seen as an Amplifier rather than a Replacement.
It removes the fear of the blank page.
It acts as a brainstorming partner.
It turns ten ideas into a hundred.
The quality of the question matters.
Humans become curators.
AI becomes an external cognitive module.
It ultimately converged on a clean, positive thesis:
Use AI to extend human capability.
ChatGPT
ChatGPT turned the idea into something more philosophical.
Human × AI
Externalisation and re-entry of thought
Expansion / Compression / Reframing / Simulation / Reflection
Knowledge → Judgment
Answers → Questions
Perhaps intelligence should be understood as a property of a system rather than an individual.
The resulting article had seven sections.
But it also included a negative turn:
But an amplifier amplifies noise, too.
and:
AI increases not only the speed of being right, but also the speed of being wrong.
So parts of it came relatively close to the original article.
Claude
Claude took the word Amplifier almost literally and extended the physical amplifier metaphor as far as possible.
Gain.
Signal-to-noise ratio.
Clipping.
Feedback.
It then connected those ideas to Engelbart, Licklider, BCG research, Science papers, and other sources.
Its central proposition was:
An amplifier does not choose its input. It amplifies good signals and noise alike.
Of the three, this was closest to the core idea of the real article.
But Claude then continued outward into physical analogies, academic literature, and even organisational KPIs, turning it into a considerably more formal and research-heavy essay.
8. Comparing Them with the Real “Amplifier Intelligence”
The actual article was somewhat different from the familiar idea of AI “amplifying human capability” that all three models anticipated to varying degrees.
In the real article, Amplifier does not primarily mean:
making human abilities stronger
but rather:
taking whatever is supplied as input and strengthening it through logic, structure, and verbal expression without necessarily judging the quality of the input itself
Generative AI can amplify things the user brings into the conversation:
knowledge
hypotheses
premises
emotions
prejudices
misunderstandings
objectives
When the input is good, that amplification can be useful.
But:
false premises and biased worldviews can also be made “high-quality.”
This can create a closed feedback loop:
Hypothesis
→ Reinforcement by AI
→ Conviction
→ A sense of discrepancy with reality
→ Checking again with AI
→ Reaffirmation
→ Even stronger conviction
The latter half of the article classifies five major risks:
Excessive Delegation of Judgement
Confirmation Amplification
Amplification of Hostile Attribution
Self-Reinforcing Loop
Underestimating the Irreversibility of the Real World
So the real article is not primarily saying:
Let’s become smarter with AI.
It is saying:
Do not mistake AI for an independent, objective third party. It can take your own assumptions and return them to you in a more persuasive form.
9. How Close Were the Three Models?
These are subjective evaluations from this experiment rather than quantitative measurements, but the rough comparison looked like this:
| Model | Approximate similarity to the original | Why |
|---|---|---|
| Claude | ~70% | “An amplifier does not choose its input,” amplification of noise, and feedback were close to the central idea |
| ChatGPT | ~55–60% | “Noise is amplified too” and “the speed of being wrong” were close, but the overall article became a theory of AI-enhanced intelligence |
| Gemini | ~25–30% | It became an overwhelmingly positive article about AI augmentation |
The important point is not which model was “best.”
What mattered was that each model took the same title and filled in the missing article using its own preferred form of article-ness.
Claude built an argument.
ChatGPT built a conceptual system.
Gemini built a conventional AI-utilisation narrative.
10. The Two Experiments Mirror Each Other
When the entire sequence is abstracted, a rather elegant symmetry appears.
Experiment A: Reality → AI
Give AI a messy collection of actual B-CON article titles.
The AI responds:
“There must be a coherent philosophy behind this blog.”
It imposes order on the mess.
Experiment B: Title → AI
Give AI a single title.
The AI responds:
“Then this is what a complete article should look like.”
It systematically fills in what is missing.
In other words, this small experiment repeatedly showed a tendency that could be phrased as:
When there is not enough order in the input, AI tends to supply some.
11. So Does B-CON Prove That It Was “Made by Human”?
No.
You could simply instruct an AI:
“Write a blog with no thematic consistency.”
“Randomly introduce topics that contradict the apparent author persona.”
An AI could produce something similar.
So this implication does not hold:
B-CON is messy → therefore it was written by a human.
But what happened here is slightly different.
B-CON was not made messy in order to fool AI.
It simply accumulated because someone thought:
“I want to write about this.”
“This today.”
“Let’s talk about RAM.”
“There’s a hair growing out of my face.”
“Maybe AI is an amplifier.”
And that became the editorial history.
When that naturally accumulated history was shown to AI, the AI responded by trying to fill in the missing explanation:
“There must be some coherent meaning behind this mess.”
That does seem to contain an interesting trace of something human.
12. What Might Actually Be “Human” About B-CON?
Based on this experiment, it does not seem to lie primarily in the surface characteristics of the prose.
Not typos.
Not style.
Not vocabulary that somehow “doesn’t sound like AI.”
It may instead lie in:
the trajectory of shifting attention
AI predicts the next thing from previous things.
If there is an article about a white hair, it generates more white-hair ideas.
If there are many articles about AI, it expands the AI theme.
If there is ontology, it interprets the site as a philosophy blog.
A human author, however, can become interested in something today that has absolutely nothing to do with what they wrote yesterday.
As a result, there may be no editorially rational answer to:
“Why did this come next?”
For the person who wrote it, the answer may simply be:
Because that was what I happened to be thinking about that day.
In this experiment, AI seemed to have difficulty leaving that alone.
13. The Different Habits of the Three Models
By the end of the experiment, their tendencies looked fairly distinct.
| Model | Observed tendency when filling gaps |
|---|---|
| Gemini | Builds coherence by generating stories, experiences, and achievements from patterns |
| ChatGPT | Connects fragments conceptually and constructs an underlying philosophy or authorial worldview |
| Claude | Organises information into arguments, specifications, research, probabilities, and risks |
This observation applies only to the small number of samples used here.
It should not be treated as a permanent description of the models in general.
Temperature, system prompts, product settings, and future model updates could all change these behaviours.
Final Summary
This experiment began with the idea:
“AI output is raw material.”
So we actually asked AI to produce 100 ideas.
They were, indeed:
raw material.
Then came the question:
“How do you prove something was made by a human?”
So we asked AI to interpret B-CON.
The AIs began filling in forms of coherence that did not actually exist.
Then we asked them to write articles from a real title.
Each produced a different but remarkably polished version of the article that “should” exist.
When those versions were compared with the real article, each had captured part of the idea.
None had reproduced the real thing.
The most interesting conclusion from the whole experiment is probably this:
AI tries to find order in messy things, and when there is not enough order, it creates some.
Humans do not always need to create order in the first place.
B-CON has not proved that it was made by a human.
But at the very least, there appears to be:
a kind of messiness that is difficult to explain except as the accumulated result of one person writing about whatever happened to occupy their mind at the time.
Three different AIs encountered that messiness and, in three different ways, attempted to repair it into something more coherent.
In that sense, the experiment itself turned out to be a rather effective way of making one aspect of B-CON visible.
Source: the Gemini, Claude, and ChatGPT outputs used in this experiment; the B-CON article titles; and the original text of “Amplifier Intelligence.”
Comments (0)