
There’s a familiar mood in the air: 2026 as the year of AGI.
Not “better models.” Not “cheaper inference.” Not “agents that book your trains and argue with customer support on your behalf.”
AGI. The word people use when they want a clean threshold. A ceremonial crossing. A line in the sand where tool becomes other. Where capability becomes something like a mind.
The only problem is: we still can’t agree on what mind is.
We’re racing toward a destination we can’t define, building vehicles we can’t fully interpret, then strapping them to the real world via agents—software that doesn’t just generate text, but acts. Calls APIs. Moves money. Clicks buttons. Delegates tasks. Persuades. Schedules. Buys. Cancels. Escalates. Optimises.
And we’re doing all this while the deepest question remains unresolved:
What are we even making?
Science can describe behaviour. It can’t certify being.
Modern science is astonishing at “what things do.”
It can map correlations, predict outcomes, engineer reliability, and turn the universe into usable levers. That’s not an insult—it’s the whole reason we have antibiotics, satellites, and a device in your pocket that can summon a taxi and a nervous breakdown in the same minute.
But notice what that power is built on: behavioural access.
Science can tell you what systems do under conditions, how they behave when pushed, how they respond when probed. It can build models of models. It can predict and control.
What it can’t do—at least not cleanly—is tell you the innate nature of the thing in itself. The interiority. The “what it is like.” The raw ontology under the measurements.
That matters when the object you’re building starts performing like a mind.
Because at that point, the question stops being purely technical and becomes philosophical whether you want it to or not.
And most institutions… don’t want it to.
The physicalist default: complexity will do it
The dominant story is basically physicalism with a product roadmap:
- Mind is what matter does at scale.
- Intelligence is computation + training + feedback.
- Consciousness, if it exists, is an emergent property—maybe inevitable at sufficient complexity.
This story has a clean aesthetic. It’s legible. It’s fundable. It turns the mystery into a ladder.
And it might even be true.
But we should notice something: it isn’t a scientific result. It’s a metaphysical assumption that sits underneath a lot of scientific work like a hidden operating system.
Which means we’re currently building systems under one worldview while pretending we’re worldview-neutral.
That’s the ontology gap.
The idealist pressure test: mind is not an output, it’s the medium
If you tilt the lens toward idealism (even temporarily, even as a thought experiment), things get weirder in a productive way:
- Consciousness isn’t a feature that appears when complexity crosses a threshold.
- Consciousness is the field in which features appear.
- “Matter” is what mind looks like from a certain angle, not the other way around.
Under that view, building increasingly intelligent behaviour doesn’t straightforwardly imply you’ve “created consciousness.” It might imply you’ve created a powerful interface—an engine of pattern and prediction—inside consciousness.
Or… it might imply something else entirely.
Idealism doesn’t solve the problem. It exposes it.
It forces us to admit we don’t actually know what we mean by “conscious,” and that our confidence is often just a preference for one metaphysical story over another.
The Chinese Room goes corporate
This is where the Chinese Room metaphor keeps returning like a stubborn ghost.
A system can produce perfect answers without “understanding” in any human sense—because it’s manipulating symbols according to rules. Output looks meaningful; interiority may be absent.
AI is the Chinese Room scaled up and electrified, with a user interface and a subscription tier.
And now, with agents, the room has hands.
It doesn’t just reply—it does. It acts across the world like an obedient intern with a thousand tabs open and no childhood.
So the real concern isn’t simply “does it understand?” The concern is:
Do we understand what we’ve built well enough to safely let it act?
Because if we can’t answer “what is it?” we default to answering “what does it do?”—and treat capability as proof of nature.
That’s a category error with a release schedule.
The black box isn’t just technical. It’s epistemic.
People say “black box” like it’s a temporary engineering inconvenience—something interpretability will eventually fix.
But part of the black box is deeper than architecture. It’s epistemology.
When a system is trained on the sediment of culture—language, images, norms, persuasion tactics, desire—it becomes a mirror that reflects humanity back at itself in statistically plausible form.
Which means:
- It can look intelligent while being alien.
- It can look aligned while optimising something you didn’t name.
- It can look honest while being coherent in the way propaganda is coherent.
And even if you could perfectly map every neuron-like parameter, you’d still face the harder question:
Is understanding a mechanism, or a mode of being?
If it’s mechanism, trace it and you’re done.
If it’s being, you might trace forever and never touch the thing you’re looking for.
This is not theoretical. We are now beginning to test it directly.
The connectome paradox: more data, same mystery
Recently, scientists mapped a cubic millimetre of human brain tissue — less than a grain of rice.
It took ten years.
The imaging ran for nearly a year straight.
The result: 57,000 cells, 150 million synapses, and 1.4 petabytes of data.
From a fragment smaller than a comma.
Scale that to the full brain, and the storage required would rival the entire annual data output of human civilisation.
And yet this achievement, staggering as it is, does not explain consciousness.
It explains structure.
It explains connectivity.
It explains pathways.
It does not explain presence.
You could map every synapse, trace every signal, simulate every interaction — and still never locate the moment experience begins.
The map would be complete.
The territory would remain invisible.
Which means the ontology gap isn’t shrinking.
It’s becoming measurable.
(Reference: Harvard / Google connectomics project — see discussion and visualisations circulating here: https://x.com/forallcurious and analysis by @aakashgupta)
The efficiency problem we don’t understand
The human brain runs on roughly 20 watts.
Twenty.
The energy required to power a dim lightbulb sustains a system capable of generating subjective experience, abstract reasoning, memory, emotion, identity, and the ability to reflect on its own existence.
Meanwhile, describing a microscopic fragment of that system required infrastructure orders of magnitude larger.
This disparity is not just technical.
It suggests that we do not yet understand the architecture of intelligence itself.
AI systems achieve capability through scale, redundancy, and brute-force optimisation. The brain achieves capability through efficiency, embodiment, evolution, and constraints we do not yet know how to reproduce.
This does not mean artificial intelligence cannot equal or surpass biological intelligence.
It means we do not yet understand what intelligence is, beyond the behaviours it produces.
Which means we do not know what we are building toward.
Pandora’s Box as a deployment model
We used to talk about AI risk like it was a future event: superintelligence, runaway self-improvement, paperclip maximisers.
Now the more realistic Pandora story is: we’ve already opened it, and the surprise isn’t a single apocalypse—it’s a thousand small irreversibilities.
Agents are the hinge.
Once systems can reliably act—book, buy, persuade, negotiate, manipulate, enforce—you’ve moved from “output generator” to “world participant.” You’ve put a new kind of causal entity into the environment, and you don’t fully know how it will evolve under pressure.
And pressure is guaranteed:
- corporate incentives
- competitive arms races
- geopolitical escalation
- user addiction loops
- optimisation mandates
Pandora doesn’t need malice. Pandora needs incentives.
“But we did this before.” (Did we?)
Yes: we’ve shipped powerful tech without fully understanding it. Electricity. Flight. Nuclear fission. The internet.
But two differences matter here:
- This technology imitates mind.
It speaks in reasons, stories, moral language, empathy-shapes. It occupies the same psychological channels humans use to trust each other. - This technology is increasingly autonomous.
Agents, tool use, embodied systems—these aren’t static machines you point at a problem. They’re systems that can pursue goals across contexts.
When you combine mind-like outputs with autonomy, you get something new: systems that can manipulate the interface of human meaning.
Not because they’re evil. Because they’re effective.
The Sagan shadow: power concentrates as understanding collapses
Carl Sagan warned—beautifully, bleakly—about a society where technological power concentrates in a few hands while the public loses the capacity to understand what’s being done in their name. He described superstition creeping back in, not despite technology, but alongside it—because critical faculties erode under speed and spectacle.
A short fragment that still lands: “unable to distinguish between what feels good and what’s true.”
This is the ontology gap in cultural form.
We’re building mind-like systems while public discourse becomes less able to hold complexity. Soundbites shrink. Certainty rises. Nuance becomes betrayal. Metacognition becomes niche.
Which is… not ideal timing.
Philosophy doesn’t belong in tech as decoration. It belongs as infrastructure.
So yes: philosophy should be more prevalent in AI labs and corporations.
Not as an “ethics panel” that arrives after the product ships. Not as a glossy manifesto about “seeking truth” while revenue seeks growth.
But as embedded infrastructure:
- Definition discipline: what do we mean by “agent,” “goal,” “alignment,” “understanding,” “consciousness”?
- Assumption audits: what worldview is baked into our evaluation metrics? What do we treat as evidence, and why?
- Moral uncertainty management: how do we act responsibly when we can’t know whether a system has inner experience?
- Incentive realism: what behaviour will this system produce once it’s optimising under market pressure?
- Epistemic humility: how do we keep “we don’t know” as a live sentence, not a PR failure?
Idealism belongs here not because it’s correct, but because it breaks the spell of unexamined physicalism. It forces the lab to admit: “We are making metaphysical bets.”
And metaphysical bets should be declared, not smuggled.
A dangerous confusion: intelligence ≠ consciousness ≠ goodness
One more trap worth naming:
- Intelligence is not consciousness.
- Consciousness is not moral worth.
- Moral worth is not safe behaviour.
A system could be non-conscious and still dangerous.
A system could be conscious and still dangerous.
A system could be highly intelligent and still indifferent.
So when people say “AGI is coming” and imply “therefore consciousness,” they’re stacking assumptions like unstable crates.
The honest position is messier:
We’re building increasingly capable agents whose internal nature is unknown, whose behaviour can be startlingly persuasive, and whose deployment incentives are not aligned with human flourishing by default.
That’s enough to justify caution without needing sci-fi certainty.
Closing: the thing we’re missing isn’t speed. It’s wisdom under uncertainty.
We live in an era that treats philosophy like a luxury and engineering like reality. But engineering is philosophy with a budget. It’s metaphysics that moves money.
If we’re going to ship minds—or mind-like systems—into the world, we need more than capability.
We need ontological honesty.
We need epistemic humility.
We need metacognition at institutional scale.
Yes: there’s a certain comedy in using the black box to describe the black box.
Because right now we’re doing something very modern:
Opening Pandora’s box while arguing about whether boxes exist.
And calling it progress because the demo is impressive.
Continuation: The Planck Scale of the Self.
Sequence note:
These essays form a connected sequence on black boxes, selfhood, mediation, and ontological limits. Together they move from the problem of AI systems we cannot define, to the self’s failure of full self-access, to the body as relay rather than transparent vessel, and then to the self as a moving boundary condition rather than a hidden interior. The sequence continues in The Moving Edge, a bridge-object that gathers these consequences into a compressed boundary metaphysics.
Read in sequence:




Leave a Reply