Essay
Understanding and Bottlenecks
When generation becomes abundant, shared understanding must move from a central gate into bounded team learning loops.
A proof can be correct before a field understands why it matters. A pull request can pass every test before anyone can explain what it will do to the larger system. AI did not create these gaps. It made them easier to reach.
Vision and Values asked who chooses the direction. The next question is organizational: when generation outruns evaluation, where does understanding live?
Placing it in one expert works until every discovery and decision must pass through the same head. Then expertise becomes a queue. The alternative is to give bounded teams complete learning loops connected by shared intent, explicit interfaces, decision rights, and evidence.
Series map: Vision and Values → Understanding and Bottlenecks → Truth and Inference → The Knowledge Factory → The Ontology Factory → The Cognitive Factory
01
Generation Scales; Understanding Does Not
Systems can generate or verify mathematical results faster than a community can explain and absorb them. Terence Tao calls this “proof indigestion”. Formal verification can prove that a derivation follows from encoded axioms. It cannot decide whether the encoding captures the problem, the result matters, or anyone can explain it. The OpenAI unit-distance result made that division visible: mathematicians checked the proof, but people still chose the problem and interpreted its significance. The Leiden Declaration similarly places correctness beside understanding, attribution, transparency, and human direction.
Software shows the same pressure. A study of 442 developers associated GenAI adoption with higher job demands and burnout, softened by autonomy and learning resources. A survey of 319 knowledge workers found critical-thinking effort shifting toward verification, integration, and task stewardship. An industry-sponsored GitLab survey found respondents reporting faster work but a review bottleneck. This perception data does not measure review quality; it identifies a risk worth testing.
The human cost can become a prompt loop: generate, skim, retry. The slot-machine comparison is structural, not clinical. Research on reward uncertainty helps explain how uncertain rewards sustain repetition; it does not make prompting a gambling disorder. At review time, the same missing judgment moves downstream: your reviewer will not understand the pull request better than you didn't. More code arrives with less context and the same duty of judgment.
These cases support a bounded diagnosis: AI can reduce the cost of output without equally reducing interpretation, integration, or adoption. Expertise therefore shifts toward models others can inspect, use, and revise.
Understanding here means a provisional, shareable model kept honest by three questions:
- Coherence: Does it fit together without consequential contradictions?
- Correspondence: Does it match the evidence and the world?
- Consequence: What happens when people act, and what should change next?
These tests discipline a model; they do not decide what matters or who holds authority. Inference also differs from understanding: a language model produces a probable continuation, while understanding maintains a model people can use to explain, predict, act, and revise against the world.
02
The Central Architect Becomes the Queue
Imagine three AI-assisted teams producing designs, code, and tests while one architect interprets discoveries, approves exceptions, and integrates changes. Work waits for approval; discoveries wait to enter the larger model. She must become the constraint, skim decisions, or erase the productivity gain.
Central expertise remains essential for broad, opaque, or irreversible effects. It becomes a bottleneck when every local decision travels through one mind, regardless of who holds the evidence. Delegating execution is not enough if the learning loop stays centralized.
03
Distribute Complete Learning Loops
Give each team a bounded outcome, enough context and authority to choose an approach, and responsibility for what happens next:
investigate → decide → act → evaluate → revise
This is stricter than “move fast.” A team owns the consequence and revision, knowing which decisions cross its boundary and who must join them.
In The Gallic War II.20–26, Caesar describes soldiers acting without complete central direction while leaders coordinated units under pressure. This limited analogy comes from a self-interested military narrator, not a modern team. Its point is narrow: responsible local action requires shared purpose, practiced judgment, and coordination.
Before delegating, name what is settled, what remains open, which constraints cannot change silently, what evidence justifies an exception, and who owns it. An AI may challenge a boundary with evidence; it may not quietly redefine it.
04
Design Philosophy Makes Local Decisions Compatible
Local decisions can be reasonable and still break another team's assumptions.
An onboarding team may revise the wording or sequence of a delayed identity
check. It cannot silently change what verified means when billing, support,
and security depend on that state.
The interface is therefore more than an API shape. It includes the meaning of states, expected failures, monitoring signals, decision ownership, and rollback or escalation conditions. A bounded exception states the conflict and evidence, proposes the narrowest change, names affected interfaces, and pauses when the effect crosses ownership. “Disagree and commit” is responsible only with a boundary, a way to detect failure, and a path back.
Shared design philosophy guides unfamiliar cases; explicit interfaces protect neighboring assumptions; feedback reveals when either is wrong. Together they replace a universal approval gate without isolating teams.
AI Factory · 03b · Organizational topology / judgment
The same teams, a different place for understanding
Scroll the path →
Teams make their reasoning inspectable:
- Gather evidence from customers, support, telemetry, and system behavior.
- Frame the problem by separating symptoms, causes, assumptions, stakes, and disagreement.
- Authorize a test by naming its boundary, learning goal, and owner.
- Learn from consequences for the people and systems affected.
- Retain the revised model in reusable records, interfaces, and practices.
Customers experience situations, not roadmaps. A coverage check should include human experience, domain rules, system failure modes, economic tradeoffs, and evidentiary quality. Otherwise, distributed understanding becomes only a collection of private insights.
05
Make Shared Models Compact—and Reopenable
Teams compress systems into shared terms. Idempotent, for example, can mean that retrying an operation creates no additional intended effect. An idempotency key is a case, a double charge a counterexample, and a retry test a check. The term saves explanation until two teams use it for different guarantees.
A trustworthy term must reopen into examples, assumptions, evidence, tests, counterexamples, and revision conditions. Keep its units distinct: a morpheme is a meaningful unit of language, a conceptual chunk is a reasoning pattern, and a model token is encoded input or output. They interact, but none explains how the others understand.
Actively ingest the claim: connect it to an existing model, test it against a case, and record what would break it. This is a learning practice, not a neurological theory.
A model that can be opened
A shared concept should reopen into its examples, assumptions, evidence, tests, counterexamples, and revision history. The map exposes what compression omits; it does not claim to depict a literal cognitive mechanism.
AI can cluster observations, compare explanations, surface questions, and propose tests. It can also produce coherence before a team has earned correspondence or examined consequences. A trustworthy synthesis therefore shows its evidence, assumptions, alternatives, limits, and the next observation that could distinguish among them.
06
Protect the Work That Develops Judgment
Reopenable models demand attention. A team cannot investigate hard cases or revise its knowledge if the day is consumed by generating, triaging, and reviewing output. Distributed authority has a capacity cost.
Research supplies no universal deep-work ratio or queue limit. Workload depends on the task and resources; autonomy and learning still matter. Treat these responses as experiments, not laws:
- cap generated work so evaluation can catch up;
- pair summaries with close examination of hard cases;
- protect time for design, investigation, mentoring, and model revision;
- reward context, evidence, and judgment, not only artifacts; and
- retain reasons, counterexamples, and consequences for the next team.
Expertise does not become smaller. Experts build shareable models, improve the boundaries within which others act, notice when evidence no longer fits, and help teams learn from consequences.
In an age of abundant answers, the scarce skill is building enough shared understanding to know what deserves to be solved—and whether an answer survives contact with the world.
Distributing decisions removes a central bottleneck. But what lets those decisions remain trustworthy? Truth and Inference takes up the shared standards teams need to judge their inputs and outputs.
07
Sources
- Terence Tao, Mathematics in the Age of AI (2026).
- Simons Foundation, “Fields Medalist Terence Tao on Artificial Intelligence and Why We Do Math” (2026).
- OpenAI, “An OpenAI Model Has Disproved a Central Conjecture in Discrete Geometry” (2026).
- Leiden Declaration on Artificial Intelligence and Mathematics.
- Jeremy Avigad et al., Theorem Proving in Lean 4.
- Karl E. Weick, Kathleen M. Sutcliffe, and David Obstfeld, “Organizing and the Process of Sensemaking” (2005).
- Amy C. Edmondson, “Psychological Safety and Learning Behavior in Work Teams” (1999).
- ISO, ISO 9241-210:2019—Human-centred design for interactive systems.
- Zixuan Feng, Sadia Afroz, and Anita Sarma, From Gains to Strains: Modeling Developer Burnout with GenAI Adoption (2026).
- Hao-Ping Lee et al., “The Impact of Generative AI on Critical Thinking” (2025).
- Charlotte Brandebusemeyer et al., Developers' Experience with Generative AI Beyond Productivity Assessment (2026).
- GitLab, 2026 AI Accountability Report.
- Martin Zack, Ross St. George, and Luke Clark, “Dopaminergic Signaling of Uncertainty and the Aetiology of Gambling Addiction” (2020).
- Julius Caesar, The Gallic War, Book II, 20–26.