<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[hadimahihenni]]></title><description><![CDATA[CEO and cofounder of Databold, building Knowtilus.]]></description><link>https://hadimahihenni.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Ta4-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e7fd8b7-7410-4a87-933d-85b4200b6afd_2097x2097.png</url><title>hadimahihenni</title><link>https://hadimahihenni.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 25 Aug 2026 21:47:28 GMT</lastBuildDate><atom:link href="https://hadimahihenni.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Hadi Mahihenni]]></copyright><language><![CDATA[en-gb]]></language><webMaster><![CDATA[hadimahihenni@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[hadimahihenni@substack.com]]></itunes:email><itunes:name><![CDATA[Knowtilus]]></itunes:name></itunes:owner><itunes:author><![CDATA[Knowtilus]]></itunes:author><googleplay:owner><![CDATA[hadimahihenni@substack.com]]></googleplay:owner><googleplay:email><![CDATA[hadimahihenni@substack.com]]></googleplay:email><googleplay:author><![CDATA[Knowtilus]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI has a quality problem. The factory floor solved it decades ago.]]></title><description><![CDATA[Why agentic AI needs compiled quality specifications, not more human reviewers]]></description><link>https://hadimahihenni.substack.com/p/ai-has-a-quality-problem-the-factory</link><guid isPermaLink="false">https://hadimahihenni.substack.com/p/ai-has-a-quality-problem-the-factory</guid><dc:creator><![CDATA[Knowtilus]]></dc:creator><pubDate>Thu, 18 Jun 2026 07:30:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-ICx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!V9ox!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!V9ox!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 424w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 848w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 1272w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!V9ox!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png" width="1402" height="829" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:829,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1621083,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hadimahihenni.substack.com/i/202544267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c6698a-dd69-46e1-9429-9419f87d9b94_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!V9ox!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 424w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 848w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 1272w, https://substackcdn.com/image/fetch/$s_!V9ox!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f41d16-5824-4b6f-a67d-99e7c1207171_1402x829.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is an uncomfortable parallel between the AI industry in 2026 and post-war manufacturing in the 1950s. Both produced at impressive scale. Both relied on end-of-line inspection to catch defects. And both treated quality as someone else&#8217;s problem: in the factory, it was the inspector&#8217;s job; in the AI pipeline, it is the human reviewer&#8217;s.</p><p>The manufacturing world took roughly three decades and one quality revolution (starting in Japan, spreading across Europe and the United States) to learn that this approach does not work. The AI industry is on track to learn the same lesson, and it does not need to take that long.</p><h2>The verification tax</h2><p>Consider the workflow that has become standard in any serious AI deployment. An agentic system analyses documents, extracts data, synthesises findings, and produces an artefact: a report, a codebase, a dataset, a slide deck. Then a human sits down to verify and validate. This is the V&amp;V step, borrowed from systems engineering, and it is where the entire quality burden now lands on human cognition: reading, cross-checking, hunting for errors the system cannot see.</p><p>The cost of this is rarely measured but always felt. Every artefact produced by an AI agent carries an implicit verification tax: the time, attention, and domain expertise a human must spend to confirm the output is correct. As AI systems produce more, faster, the tax does not shrink. It compounds. The human becomes the bottleneck in a pipeline that was supposed to remove bottlenecks.</p><p>The problem is not that humans are bad at verification. The problem is structural. The reviewer is validating against a specification that exists only in their head. They carry tacit knowledge about what a correct output looks like: domain conventions, semantic constraints, inter-field dependencies, formatting norms, edge cases learned through years of practice. None of this is written down. None of it is measurable. And none of it is available to the system that produced the artefact in the first place.</p><p>You might argue that we could delegate this review to another AI agent. In practice, this merely displaces the problem. An LLM asked to &#8220;check if this extraction is correct&#8221; will apply its own stochastic judgement, not a deterministic specification. It cannot guarantee coverage of domain-specific rules it was never explicitly given. It is, in current jargon, &#8220;LLM-as-judge&#8221;: an ad hoc assessor with no formal quality framework, no defined tolerances, and no traceability.</p><p>This is still inspection. And as W. Edwards Deming spent his career arguing, from his work with Japanese manufacturers in the 1950s to his later influence on Western industry, inspection does not produce quality. It merely sorts good outputs from bad ones after the damage is done. The defect has already been produced; you are just deciding whether to ship it.</p><h2>What the quality revolution actually taught us</h2><p>The transformation that Deming helped set in motion (first in Japan, then adopted globally through lean production, the Toyota Production System, and Six Sigma, originally developed at Motorola) rested on a deceptively simple insight: quality must be built into the process, not verified after the fact. The mechanism for this shift was formalisation. You take what you know about what &#8220;good&#8221; looks like, and you turn it into things you can measure.</p><p>In quality engineering terminology, these are Critical to Quality characteristics (CTQs): measurable properties of the output that matter to the end user, derived from higher-level requirements through a structured decomposition. A CTQ is not a vague aspiration (&#8221;the output should be accurate&#8221;). It is a specification with a target, an upper tolerance, a lower tolerance, and a measurement method. The diameter of this shaft shall be 25.00mm &#177; 0.02mm, measured by coordinate measuring machine.</p><p>Once you have CTQs, everything else follows. You can monitor the process in real time. You can calculate whether the process can reliably produce within specification. You can anticipate failure modes before they occur. And you can design the process itself to make defects structurally difficult to produce.</p><p>This last concept, known as poka-yoke in the Toyota Production System, is particularly instructive. A poka-yoke is a constraint built into the process that prevents errors by design. A USB connector that only fits one way. A car that will not shift into reverse above a certain speed. The principle is not to train the operator to be more careful; it is to make the wrong outcome mechanically impossible, or at least immediately detectable.</p><p>In software engineering, this philosophy has mature equivalents: Design by Contract (preconditions, postconditions, invariants), static type systems that catch errors at compile time rather than run time, property-based testing that verifies behavioural invariants across thousands of generated inputs. The underlying principle is the same: formalise what &#8220;correct&#8221; means, and verify it automatically.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-ICx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-ICx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 424w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 848w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-ICx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png" width="1402" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2413470,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hadimahihenni.substack.com/i/202544267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-ICx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 424w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 848w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!-ICx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98a698ed-7a93-44b8-a4d4-226a638b2fd3_1402x1122.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The octopus in the labyrinth</h2><p>Now transpose this to agentic AI. An LLM executing a complex task (extracting structured data from a document, generating a technical report, producing code from a specification) is, in essence, an octopus in a labyrinth. It is remarkably capable: flexible, fast, able to squeeze through passages that rigid rule-based systems could never navigate. But it has a fundamental limitation. It decides for itself whether it has reached the exit. It has no external reference for &#8220;done.&#8221; It has no walls telling it &#8220;not this way.&#8221; It is navigating by feel and by pattern, and when it reaches what looks like an opening, it declares success, even if it is in a dead end.</p><p>What we propose is to separate the problem into its constituent parts. The octopus is the agent: the flexible, adaptive intelligence that navigates. The labyrinth walls are the compiled constraints: the assertions, invariants, and V&amp;V rulesets that define what constitutes a valid path. And the definition of done is external to the agent, established during build time, not improvised during run time.</p><h2>The missing middle layer</h2><p>What does this look like concretely? Consider the current landscape of quality tools available for AI pipelines.</p><p>Structured output validation (JSON Schema, Pydantic models) checks whether the output conforms to a data structure. This is necessary, but it operates at the syntactic level. It is type checking. Verifying that a field called &#8220;total_amount&#8221; is a number is not the same as verifying that the number equals the sum of line items minus the discount, and that the discount rate matches the contractual terms on file.</p><p>Guardrails frameworks (NeMo Guardrails, Guardrails.ai) enforce constraints around safety, format compliance, and topic boundaries. Again, necessary. But these are generic: they know nothing about the specific domain in which the AI is operating.</p><p>LLM-as-judge approaches use one model to evaluate another&#8217;s output. This introduces stochasticity into the quality process itself. Running the same evaluation twice may yield different results. There is no specification, no tolerance, no traceability.</p><p>What is missing is the middle layer: domain-aware, semantically rich, measurable quality characteristics that are formally derived from compiled domain knowledge and enforced inline during execution. Not type checking, but semantic contract enforcement.</p><p>To make this concrete: consider a system extracting financial data from fund reports. A compiled quality specification for a single field, the net asset value per share, would define not just that the field must be a number (structured output), but that it must be arithmetically consistent with total NAV divided by shares outstanding, that it must fall within a plausibility bound relative to the prior quarter, that its currency denomination must match the fund reference data, and that any deviation triggers a defined response (flag, escalate, halt). Each of these rules is a domain invariant, compiled from expert knowledge, enforced deterministically at run time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-6zE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-6zE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 424w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 848w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 1272w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-6zE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png" width="620" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:620,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hadimahihenni.substack.com/i/202544267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-6zE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 424w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 848w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 1272w, https://substackcdn.com/image/fetch/$s_!-6zE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff47ab6d1-7009-43d1-955b-82a20b34c387_620x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These are not validation rules that an engineer guesses at. They are quality characteristics systematically derived from domain knowledge. And the methodology for deriving them, ensuring coverage, and maintaining them as the domain evolves, is where the real work lies.</p><h2>Build time and run time</h2><p>The approach separates AI pipeline design into two distinct phases, borrowing from both quality engineering and software compiler design:</p><p><strong>Build time</strong> is when domain knowledge is compiled into executable artefacts: the contract specifications (what must be true about each output, with targets and tolerances), the assertion sets (deterministic checks, cross-field validations, plausibility bounds), and the pipeline topology (which steps execute in what order, with quality gates between them). Build time is where you define the invariants and embed the poka-yokes.</p><p><strong>Run time</strong> is when the agent executes within this compiled framework. The LLM does what it does best: flexible, adaptive processing of unstructured information. But it operates within constraints. Every intermediate artefact is measured against its compiled specifications. Violations trigger defined responses. The definition of done is not the model&#8217;s assessment of its own performance; it is the compiled specification&#8217;s assessment of the output.</p><p>This is not a new idea in principle. Design by Contract has existed since the 1980s. Static analysis catches entire classes of bugs at compile time. What is new is applying this rigour to the specific challenge of AI-generated artefacts, where the execution engine is stochastic by nature and the output is information rather than a binary.</p><h2>The missing abstraction</h2><p>The current tooling for AI quality (structured outputs, guardrails, evaluation benchmarks, LLM-as-judge) addresses symptoms without providing an organising principle. Each tool solves a narrow problem. None of them answers the question: where do the quality specifications come from? How are they derived from domain knowledge? How do you ensure coverage? How do you maintain them as the domain evolves?</p><p>This is the gap. Without a systematic methodology for compiling domain knowledge into measurable specifications, AI engineers write ad hoc validation rules based on their best understanding of the domain. Coverage is accidental. Maintenance is manual. The quality system is fragile because its foundations are informal.</p><p>With domain knowledge compilation, you have a formal, auditable, maintainable derivation of quality specifications from source knowledge. The compilation step is what makes everything downstream, the runtime enforcement, the monitoring, the continuous improvement, possible and grounded.</p><h2>The road ahead</h2><p>We are not suggesting that this is easy. The hardest part is not building the enforcement system; it is defining what to enforce. Eliciting tacit domain knowledge from experts and formalising it into measurable specifications is painstaking work. It requires deep domain understanding, iterative refinement, and the intellectual honesty to admit when a quality characteristic is not yet measurable.</p><p>But the alternative, the status quo of human-in-the-loop inspection at the end of an AI pipeline, does not scale. It worked when volumes were low. It breaks when you are processing thousands of documents, generating hundreds of artefacts, extracting millions of data points. The verification tax becomes the dominant cost.</p><p>The manufacturing world learned this lesson. The question is whether the AI industry will learn it from history, or insist on learning it the hard way.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://hadimahihenni.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://hadimahihenni.substack.com/subscribe?"><span>Subscribe now</span></a></p><p></p><div><hr></div><p><em>I spent fifteen years between two worlds that rarely talk to each other: lean manufacturing and quality engineering on one side, machine learning and generative AI on the other. I hold a Six Sigma Black Belt and have built AI systems for critical industries. This article exists because I kept seeing the same quality failures I had learned to prevent on the factory floor reappearing, almost identically, in AI pipelines. The tools to fix them already exist. They just need translating.</em></p><p><em>This article reflects work at the intersection of AI systems design and quality engineering. We build systems where LLMs act not as autonomous engines but as compilers: transforming structured knowledge into verified, auditable artefacts. The ideas here draw on practice in critical systems engineering, lean production principles, and the daily reality of making AI outputs trustworthy enough for professional use.</em></p><p></p>]]></content:encoded></item><item><title><![CDATA[Agents think in tokens. Humans think in shapes.]]></title><description><![CDATA[Why the next agent interface won&#8217;t be a better prompt, it&#8217;ll be a rectangle.]]></description><link>https://hadimahihenni.substack.com/p/agents-think-in-tokens-humans-think</link><guid isPermaLink="false">https://hadimahihenni.substack.com/p/agents-think-in-tokens-humans-think</guid><dc:creator><![CDATA[Knowtilus]]></dc:creator><pubDate>Mon, 25 May 2026 10:59:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!C475!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!C475!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!C475!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!C475!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!C475!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!C475!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!C475!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3545198,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hadimahihenni.substack.com/i/198869919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!C475!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!C475!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!C475!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!C475!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4c7e0-fe8c-47d1-8c78-5a735001e507_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><p>In the Lascaux prehistoric cave (France), just below the famous painting of a great stag, there&#8217;s a discreet drawing that&#8217;s easy to miss: a simple rectangle. No animal, no scene, just four lines meeting at right angles.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hadimahihenni.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en-gb&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading knowtilus newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p>Stanislas Dehaene, the French neuroscientist at the Coll&#232;ge de France, spent years investigating why early humans drew geometric shapes before they drew anything else. His conclusion, published in <em><a href="https://www.odilejacob.fr/catalogue/sciences/neurosciences/rectangle-de-lascaux_9782415014278.php">Le Rectangle de Lascaux</a></em><a href="https://www.odilejacob.fr/catalogue/sciences/neurosciences/rectangle-de-lascaux_9782415014278.php"> (2026): geometry is a </a><em><a href="https://www.odilejacob.fr/catalogue/sciences/neurosciences/rectangle-de-lascaux_9782415014278.php">language of thought</a></em> unique to Homo sapiens. Before language, before writing, before agriculture, our ancestors were already thinking in shapes. Rectangles, circles, spirals, parallel lines. These aren&#8217;t decorations. They&#8217;re the native format of human cognition.</p><p>40,000 years later, we built machines that think in tokens.</p><p>And now we&#8217;re asking both to work together.</p><div><hr></div><h2>Three conversations, three mismatches</h2><p>Every AI system involves three communication channels. Each has a different problem.</p><p><strong>Agent &#8596; Agent</strong> is the easy one. Tokens in, tokens out. JSON, function calls, structured data. Machines don&#8217;t care about bandwidth. They don&#8217;t get bored. They don&#8217;t skim. This channel is essentially solved &#8212; it&#8217;s engineering, not design.</p><p><strong>Human &#8596; Human</strong> is the ancient one. We&#8217;ve been optimising this for millennia. And the pattern is clear: whenever the stakes are high, we don&#8217;t just <em>write</em>, we <em>show</em>. The board presentation. The architectural blueprint. The napkin sketch. The war room whiteboard. High-stakes human communication converges on visual artifacts because that&#8217;s where our cognitive bandwidth lives. Vision processes information at roughly 10 million bits per second. Language? About 50 bits/sec. The ratio isn&#8217;t close, it&#8217;s five orders of magnitude (100 000x).</p><p><strong>Agent &#8596; Human</strong> is the broken one. And it&#8217;s broken in a specific way: agents communicate like they&#8217;re talking to other agents. They dump plans in prose. They explain in paragraphs. They output reasoning traces that read like legal depositions. All tokens, no shapes.</p><p>The smarter the agent gets, the worse this mismatch becomes. More capability means more output. More output means more text the human has to read, validate, and course-correct: all through a 50 bit/s pipe.</p><div><hr></div><h2>The validation bottleneck</h2><p>This isn&#8217;t an abstract UX concern. It&#8217;s the binding constraint on agent usefulness.</p><p>When an agent generates a work plan, a research synthesis, or a presentation draft, a human needs to validate it. They need to evaluate whether the structure, logic, and emphasis are right. Today that means reading. A lot. And reading is slow, sequential, and exhausting.</p><p>The result is predictable: people skim, approve, and move on. The agent executes a plan that was never truly reviewed. We&#8217;ve traded one problem (doing the work) for another (supervising work we can&#8217;t efficiently inspect).</p><p>Dehaene&#8217;s research suggests why this feels so wrong. Humans don&#8217;t just <em>prefer</em> visual information: we have dedicated neural circuits for processing geometric regularity. His team showed that when we look at a shape, we immediately form a <em>mental program</em> for it: a compact, symbolic representation built from primitives like symmetry, parallelism, and repetition. This &#8220;geometric language of thought&#8221; operates alongside (and partly independent of) natural language.</p><p>In other words, shapes aren&#8217;t a nice-to-have. They&#8217;re a cognitive channel that text literally cannot replace.</p><div><hr></div><h2>The presentation paradox</h2><p>When an agent generates a slide deck, it&#8217;s not producing output for itself: it&#8217;s producing an artifact for &#8220;human&#8596;human&#8221; persuasion. But it builds that artifact the only way it knows how: token by token, optimising for textual coherence. It has no access to the visual, spatial, emotional channel that actually drives conviction between humans.</p><p>This is why AI-generated decks are always... adequate. Coherent structure, reasonable visuals, generic everything. A great presentation is shaped by context you can&#8217;t verbalise upfront: the politics of the room, the board&#8217;s past objections, the visual shorthand your industry uses, the emotional arc that will land with <em>this</em> audience. This is System 1 territory (spatial, intuitive, felt) not something you can spell out in a prompt.</p><p>Current tools force a binary choice: re-prompt and regenerate the entire deck (losing everything you shaped), or manually edit the final pixels (defeating the purpose of automation). The agent produces the artifact, but the human can&#8217;t inject what matters most: the non-verbal intelligence that makes the difference between a deck that informs and one that convinces.</p><div><hr></div><h2>Visual Plan Mode</h2><p>The fix is structural, not cosmetic. Instead of text plans that humans read, or final artifacts that humans manually edit, we need an intermediate layer that speaks the brain&#8217;s native geometric language.</p><p>For a presentation, this means three steps:</p><p><strong>Storyline</strong>: a lightweight text structure. What&#8217;s the argument? What&#8217;s the sequence? This is System 2 territory: logic, coherence, narrative arc. Agent and human can iterate fast here.</p><p><strong>Layouts</strong>: the agent proposes visual arrangements for each slide. Not final designs (spatial blueprints). Where&#8217;s the chart? Where&#8217;s the headline? What&#8217;s the visual weight? The human grasps these instantly (System 1), adjusts with a drag, and moves on. This is the step current tools skip entirely.</p><p><strong>Generation</strong>: the agent renders final content <em>within</em> the validated layouts. Each element stays editable without regenerating the whole slide. The layout acts as a constraint that preserves the human&#8217;s intent.</p><p>This three-step loop works because it matches how cognition actually flows: logic first (System 2), then spatial validation (System 1:  the geometric channel), then detail.</p><p>But presentations are just one case. The principle generalises: any complex agent output benefits from a visual intermediate representation. A project plan becomes an editable timeline, not a bullet list. A research synthesis becomes a spatial map of themes, not a wall of text. A financial model becomes an annotated flow, not a paragraph of cell references.</p><div><hr></div><h2>The shape of what comes next</h2><p>The agent ecosystem is racing to make models smarter. Longer context windows, better reasoning, more tools. But the interface, where agent and human actually meet, is stuck in the terminal. Text in, text out.</p><p>Dehaene showed that 40,000 years ago, our ancestors were already encoding complex thought into geometric forms. That capacity didn&#8217;t disappear when we invented writing: it was <em>obscured</em> by it. The most powerful cognitive channel we have is the one our AI tools ignore.</p><p>The rectangle in Lascaux wasn&#8217;t primitive. It was the first interface.</p><p>It&#8217;s time we built the next one.</p><div><hr></div><p><em>At Knowtilus, we&#8217;re building visual plan mode for knowledge work. The intermediate representations are called Knowtiles: editable, auditable visual building blocks that sit between human intent and agent execution. If you work in consulting, advisory, or financial analysis, let&#8217;s talk.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hadimahihenni.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en-gb&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading knowtilus newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Vertebrate Agent: Why AI Agents Need a Spine, Not a Bigger Brain]]></title><description><![CDATA[Most AI agents today are invertebrates : flexible but structurally unreliable. The future belongs to vertebrate agents: ones with a rigid, compiled spine that no amount of intelligence can override.]]></description><link>https://hadimahihenni.substack.com/p/the-vertebrate-agent-why-ai-agents</link><guid isPermaLink="false">https://hadimahihenni.substack.com/p/the-vertebrate-agent-why-ai-agents</guid><dc:creator><![CDATA[Knowtilus]]></dc:creator><pubDate>Mon, 18 May 2026 07:47:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X_cA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X_cA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X_cA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 424w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 848w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 1272w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X_cA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png" width="716" height="590.5314989138305" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1139,&quot;width&quot;:1381,&quot;resizeWidth&quot;:716,&quot;bytes&quot;:1950640,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hadimahihenni.substack.com/i/198225772?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!X_cA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 424w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 848w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 1272w, https://substackcdn.com/image/fetch/$s_!X_cA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea15e24e-9b7b-4251-b396-350a39dbb8fd_1381x1139.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The Invertebrate Problem</strong></p><p>Watch a modern AI agent in action. It receives a task, reasons about it, decides on a plan, executes steps, adjusts on the fly, and sometimes (often, actually) does something unpredictable that breaks the entire workflow.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hadimahihenni.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en-gb&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading hadimahihenni! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This is the invertebrate architecture: the LLM is the brain, the body, and the spine all at once. It decides <em>what</em> to do, <em>how</em> to do it, and <em>in what order</em>. Every decision is probabilistic. Every execution is a fresh generation. The agent is flexible, capable of remarkable contortions but structurally unsound.</p><p>The numbers tell the story. Multi-agent success rates degrade from 58% on single-turn tasks to 35% on multi-turn ones. Nearly 79% of failures stem not from reasoning errors but from specification and coordination issues. The agent knows what to do. It just can&#8217;t reliably do it the same way twice.</p><p>Enterprises have noticed. While 88% of organisations now use AI in at least one function, roughly two-thirds still can&#8217;t scale beyond pilot projects. The gap isn&#8217;t intelligence. It&#8217;s structure.</p><p><strong>More Tentacles Won&#8217;t Help</strong></p><p>The current industry response to unreliable agents is to give them more tools. More MCP connectors. More skills. More context. An agent with 50 integrations can read your CRM, query your database, send emails, generate charts, and search the web : all in one turn.</p><p>But it still <em>decides</em> which tools to call, in what order, with what parameters, every single time. The workflow is emergent. The agent figures it out at runtime. More tools make a more capable invertebrate. Still an invertebrate.</p><p>This is the equivalent of giving an octopus more tentacles. The problem was never capability. It was structure.</p><p>Tools give the agent hands. What it&#8217;s missing is a spine : a rigid internal structure that says &#8220;you <em>will</em> do step A before step B, you <em>will</em> use this formula for X, you <em>will not</em> skip the validation check Y&#8221; regardless of which model is driving.</p><p><strong>What Vertebrates Have That Invertebrates Don&#8217;t</strong></p><p>In biology, the vertebrate innovation wasn&#8217;t a bigger brain. It was a spinal column : a rigid internal structure that constrains movement into reliable, repeatable patterns while freeing the nervous system to handle only what requires genuine intelligence.</p><p>An octopus is brilliant. It can open jars, solve puzzles, squeeze through impossible gaps. But you wouldn&#8217;t build a factory around one. You&#8217;d build it around vertebrates: organisms whose movements are predictable, repeatable, and structurally constrained, with intelligence deployed where it actually matters.</p><p>A vertebrate agent has:</p><p><strong>A spine </strong>: deterministic, compiled workflow logic that defines what steps execute, in what order, with what constraints. This is code, not prompts. It runs the same way every time. It&#8217;s auditable, versionable, and testable.</p><p><strong>Joints</strong> : bounded points where the LLM is invoked for genuine intelligence tasks: interpreting ambiguous text, drafting narrative, classifying edge cases. The LLM operates within a slot defined by the spine. It can move freely, but only within the range the joint allows.</p><p><strong>Muscles</strong> : the domain knowledge, business rules, and validation logic that connect spine to action. An EBITDA margin is calculated as EBITDA / Revenue &#215; 100. A compliance document must contain sections A, B, C in that order. These aren&#8217;t suggestions. They&#8217;re constraints encoded in code.</p><p>The LLM never decides the workflow&#8217;s path. It fills bounded slots within a structure it didn&#8217;t create and can&#8217;t override.</p><p><strong>Graduated Determinism: Not All Vertebrae Are Equal</strong></p><p>A real spine isn&#8217;t uniformly rigid. The cervical vertebrae allow the head to turn freely. The thoracic vertebrae lock the ribcage in place. The lumbar vertebrae bear load with controlled flexibility. The architecture should reflect this gradient.</p><p><strong>Hard rules : the skull.</strong> Zero LLM involvement. EBITDA is calculated, not interpreted. The company logo goes bottom-right, 24px from the edge. A regulatory filing must contain these sections in this order. These are validators, layout constraints, calculation rules. If they fail, the output is rejected.</p><p><strong>Soft rules : the thoracic spine.</strong> Deterministic defaults that humans can override. The executive summary is typically two slides. Financial projections default to three years historical plus two projected. The system proposes; the human disposes.</p><p><strong>Bounded LLM slots : the joints.</strong> Runtime intelligence, contained. Given the financial data and the kick-off notes, write the narrative paragraph under the EBITDA bridge chart. Given the management bios, draft the key persons slide. The LLM fills a slot defined by the spine, never the other way around.</p><p>When semantic judgment is unavoidable, the pattern is what researchers call the &#8220;Safety Sandwich&#8221;: deterministic validation <em>before</em> the LLM call constrains the input; deterministic validation <em>after</em> the call constrains the output. The LLM is sandwiched between two layers of compiled rules. It can think freely within its slot. It cannot escape it.</p><p>This gradient, from fully deterministic to bounded probabilistic, is what makes vertebrate agents practical. Not every decision needs to be hard-coded. But every decision needs to know which category it belongs to.</p><p><strong>Who Builds the Spine?</strong></p><p>Here&#8217;s where the field reveals a gap.</p><p>The research demonstrates that vertebrate architecture works. Deterministic workflows with bounded LLM calls outperform autonomous agents on reliability, cost, and surprisingly on accuracy too. Models perform dramatically better when given curated structure than when they self-organize. One <a href="https://www.skillsbench.ai/">benchmark</a> found a 16-pp improvement with curated skills, and <em>zero or negative</em> benefit from self-generated ones.</p><p>The models cannot build their own spine. Someone has to build it for them.</p><p>But who? Today, the answer is: a developer, manually, for each workflow. That&#8217;s where the node-based tools stop. It works for a handful of workflows. It doesn&#8217;t scale to the thousands of domain-specific processes an enterprise runs.</p><p>The missing piece is <strong>knowledge compilation </strong>: the systematic extraction of organisational knowledge (from documents, processes, expert minds, past deliverables) into the compiled, deterministic artifacts that form the spine. An ontology that structures what the organisation knows. Validation rules derived from how the organisation actually works. Bounded LLM slots placed exactly where genuine judgment is needed.</p><p>This compilation layer (not the models, not the raw knowledge) is where durable value accrues. The LLMs are commoditizing. The spines are not.</p><p><strong>The Uncomfortable Implication</strong></p><p>The dominant narrative in AI, that we need bigger, smarter, more autonomous agents, is directionally wrong for enterprise adoption.</p><p>What enterprises need isn&#8217;t more intelligence. It&#8217;s more structure <em>around</em> the intelligence they already have access to. The value isn&#8217;t in the model. It&#8217;s in the compilation layer that turns organisational knowledge into a reliable spine, then deploys any model&#8217;s intelligence within it.</p><p>The agent doesn&#8217;t need a bigger brain. It needs a spine.</p><div><hr></div><p><em>Hadi Mahihenni is CEO and co-founder of Databold</em></p><p><em>We&#8217;re building the knowledge compilation layer for structured professional artifacts: the infrastructure that turns organisational expertise into reliable, auditable spines for AI-generated deliverables. We&#8217;ve validated the architecture on two live use cases (decision decks and financial data extraction workflows) and we&#8217;re raising to generalise it. If you&#8217;re an engineer who wants to work on this problem, or an investor who sees where it&#8217;s going, please reach out.</em></p><div><hr></div><p><strong>Notes &amp; Further Reading</strong></p><p>The empirical findings referenced in this piece draw from several recent research papers for those who want to go deeper:</p><p>[1] Multi-agent failure rates (79% from specification/coordination, 58%&#8594;35% degradation): Trooskens et al., &#8220;Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation,&#8221; April 2026. <a href="https://arxiv.org/abs/2604.05150">arxiv.org/abs/2604.05150</a></p><p>[2] The &#8220;Safety Sandwich&#8221; pattern and compile-time code generation achieving 57&#215; cost reduction: same paper as [1].</p><p>[3] The &#8220;Blueprint First, Model Second&#8221; philosophy and Source Code Agent framework (+10.1pp over best baseline on &#964;-bench): Qiu et al., Alibaba Group, August 2025. <a href="https://arxiv.org/abs/2508.02721">arxiv.org/abs/2508.02721</a></p><p>[4] Curated skills = +16.2pp improvement, self-generated skills = zero or negative benefit, smaller models + good skills &#8776; bigger models: Li et al., &#8220;SkillsBench,&#8221; February 2026. <a href="https://arxiv.org/abs/2602.12670">arxiv.org/abs/2602.12670</a></p><p>[5] Same compiled skill = up to 40% performance variance across models and frameworks: Chen et al., &#8220;SkVM,&#8221; April 2026 <a href="https://arxiv.org/abs/2604.03088">arxiv.org/abs/2604.03088</a>; Ouyang et al., &#8220;SkCC,&#8221; May 2026 <a href="https://arxiv.org/abs/2605.03353">arxiv.org/abs/2605.03353</a>.</p><p>[6] Multi-level skill compilation from agent trajectories (planning / functional / atomic): &#8220;SkillX,&#8221; submitted to ICML 2025. <a href="https://arxiv.org/abs/2604.04804">arxiv.org/abs/2604.04804</a></p><p>[7] The &#8220;Code Factory&#8221; paradigm &#8212; LLM writes code once, code runs forever: XY.AI Labs, &#8220;The Code Factory Manifesto,&#8221; January 2026. <a href="https://www.xy.ai/tech-behind-xyai/the-code-factory-manifesto">xy.ai/tech-behind-xyai/the-code-factory-manifesto</a></p><p>[8] Organisational tacit knowledge extraction via LLM-conducted expert interviews: Zuin et al., July 2025. <a href="https://arxiv.org/abs/2507.03811">arxiv.org/abs/2507.03811</a></p><p>[9] Codified expert domain knowledge achieving 206% improvement in AI agent output quality: &#8220;How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge,&#8221; CAIN&#8217;26. <a href="https://arxiv.org/abs/2601.15153">arxiv.org/abs/2601.15153</a></p><p>[10] The &#8220;deterministic code as out</p><p>put layer&#8221; architecture applied to industrial operations: OSS Ventures, November 2024. <a href="https://medium.com/oss-ventures/deterministic-code-as-the-output-layer-how-we-found-an-ai-architecture-that-can-be-used-in-6d670601af05">medium.com/oss-ventures</a></p><p></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hadimahihenni.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en-gb&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading hadimahihenni! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Compiled Intelligence]]></title><description><![CDATA[I&#8217;ve spent the last two years building an AI product for knowledge workers: consultants, investment bankers, analysts ..etc, the people who turn messy information/data into structured, high-stakes deliverables.]]></description><link>https://hadimahihenni.substack.com/p/compiled-intelligence</link><guid isPermaLink="false">https://hadimahihenni.substack.com/p/compiled-intelligence</guid><dc:creator><![CDATA[Knowtilus]]></dc:creator><pubDate>Mon, 18 May 2026 07:34:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4ade6aee-3902-42c0-8c74-d204873811d3_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve spent the last two years building an AI product for knowledge workers: consultants, investment bankers, analysts ..etc,  the people who turn messy information/data into structured, high-stakes deliverables.</p><p>What I&#8217;ve learned: the bottleneck isn&#8217;t AI intelligence. It&#8217;s AI structure. The models can reason, write, and analyse. What they can&#8217;t do reliably is produce the same artifact the same way twice, respect domain rules they weren&#8217;t trained on, or work within constraints that matter to the people who sign the deliverable.</p><p>This Substack is where I&#8217;ll think out loud about what I&#8217;m calling <em>compiled intelligence</em>: the emerging architecture where organisational knowledge gets compiled into deterministic artifacts that constrain AI execution, rather than hoping the model figures it out at runtime.</p><p>Expect a post every  week. Topics will range from the technical (knowledge compilation, ontology design, hybrid neural-symbolic systems) to the practical (what actually works when you ship AI to demanding professionals or in critical industries). </p><p>First article this week: <strong>&#8220;The Vertebrate Agent: Why AI Agents Need a Spine, Not a Bigger Brain.</strong></p><p>If you build AI products for enterprise, work in consulting or finance, or think about the gap between AI demos and AI in production: this is for you.</p>]]></content:encoded></item></channel></rss>