InflectAI, Inc.

Theory note

Dynamic Constraints Production System

  • Author: Thomson Nguy
  • Published: September 29, 2026

If AGI is here, where are the world-changing economic results? People usually point to AI hallucinations or stubborn resistance to new technology. I think the unit of analysis is wrong. Humans already possess general intelligence. A large company employs thousands of them. Yet a company can continue to optimize for the old production system long after a new innovation has made the way it finances, sells, and improves a product obsolete.

The AI labs look for general intelligence inside a model. But intelligence is not the same as economic output. And if you want to unlock economic output, you have to understand what changes inside a firm when cognition is cheap. The answer, as it turns out, is simple. It moves to the next bottleneck.

Think of a company as a production function. It takes a bunch of inputs - money, labor, knowhow - and transforms them through workflows into a product. Then that product needs to be sold or distributed. Supporting that, you have functions like legal, HR, and finance to enable a company to make things or sell them. The output is economic production - measured in dollars, revenue, income.

The naive application of AI - which is what almost every existing company and startup is doing today - is to simply apply it to one of the functions in the production function and make it incrementally better. Harvey, for example, promises to make things more efficient for lawyers. Rogo does the same for investment bankers. And if you make your customer's economic production function 5% more effective, you get to claim a share of that 5% improvement. But you are always going to be limited by your customer's ceiling. Because once you improve your customer's research efficiency, as Harvey does, their next binding constraint may be how they can get new clients, or the capacity of their partners, or some other area that they haven't applied AI to.

Here, though, AI is not simply another tool that can improve one part of a company's production system. Because it improves cognition, it promises to improve ALL parts of a company's production function. The problem - you have to be willing to reorganize your entire company and your way of doing business to fully take advantage of it.

And historically, that's what needed to happen. The result was outsized economic gains for companies that were able to reorganize their production functions to take advantage of the new technology. Ford. General Motors. Airbnb. Google. They changed their entire processes, workflows, and organization around what the new technology enabled.

Remove one constraint, and your binding constraint moves to the next thing. Solve that thing, and it moves to the next thing after that. AI enables faster coding. The binding constraint moves to code review. Implement automated testing and code review. The binding constraint moves to product development and how fast you can understand and shape what the product needs for the market. Shorten the cycle time of product development, and the binding constraint moves to the ability of customers to absorb that new capability in a way that increases their willingness to pay. That's why many companies don't really get the promised return with AI. They only solve one part with AI, but the next bottleneck flows somewhere else.

One of the superpowers of AI is that it drastically lowers the cost of execution. Another superpower is that it enables a small team to tap into deep expertise across domains. Previously, humans had to develop complex organizations to extract and maintain that type of specialized knowledge. What if we rebuilt the firm from the ground up, based on those two principles?

This is what this essay is about. The post-AI firm. Not just bolting AI to an existing company's process and organization. Designing it from the ground up.

But every design needs a framework. It's an intellectual scaffold to understand what needs to be done, and how to apply AI. The framework I will introduce here is called the Dynamic Constraints Production System, or DCPS. DCPS is simply the underlying theory. The most radical part of the theory is that instead of having a large organization defined by static departments, AI allows you to constantly modify your company into something that allows you to innovate and adapt faster than your competitors can react. You can achieve 10 cycles of learning for their one. Recursive Production Function Analysis is the method for examining the key thing to focus on to unlock that step function of economic productivity. Constraint Cascade Analysis is the operating loop inside that method.

Illustrative example: four jobs arrive per minute. As the bottleneck moves from Code to Review to Product, output rises from 1 to 2 to 3 jobs per minute. These are not company measurements.

The Firm as a Production Function

Wassily Leontief won the 1973 economics prize for developing input-output analysis, which traced how industries depend on one another's outputs as inputs. I borrow the logic of indispensable complements at the scale of a firm. Extra capacity in one part of the system cannot compensate without limit for a missing capability elsewhere.

The firm requires several capability bundles to produce an economic result. I use eight as the default decomposition:

Bundle Capability Question
COC_O Opportunity and demand Is the problem real, valuable, and reachable?
CFC_F Founder and leadership Can leadership perceive, decide, recruit, adapt, and hold the system together?
CRC_R Knowledge, invention, and validation Is the proposition technically and empirically true?
CIC_I Defensibility and institutional architecture Can the venture preserve and govern the value it creates?
CEC_E Engineering and productization Can the knowledge become a reliable product or service?
CDC_D Distribution and commercialization Can customers be reached, contracted, served, retained, and expanded?
CAC_A Organizational execution Can humans, agents, capital, decisions, and workflows be coordinated?
CKC_K Capital and runway Is the quantity, timing, form, and patience of capital adequate?

These are capabilities, not departments. A company can have a CRO and still be weak at distribution. It can have a large engineering staff and still be unable to productize its knowledge. It can have cash and still lack the kind of capital that allows it to change direction. A financing round that demands immediate revenue acceleration may remove a runway constraint while making product reinvention harder.

The requirement for each bundle changes with the company's functional stage:

Stage Functional question
0. Possibility detection Can a valuable possibility be identified?
1. Concept and feasibility Can the proposition work technically and economically?
2. Protectable and reproducible capability Can it be repeated, and can its value be preserved?
3. Product validation Do defined users repeatedly obtain measurable value?
4. Repeatable commercialization Can customers be acquired and served repeatedly with viable economics?
5. Organizational scaling Can output grow without heroic founder effort or quality collapse?
6. Durable institution and capital compounding Can the system survive market, technology, leadership, and capital-cycle transitions?

A stage is an achievement condition, not a measure of company age or the name of its last financing round. A nine-year-old Series B business can be forced back into Stage 3 when AI changes the customer's workflow and renders its old product abstraction insufficient. The production function's set of indispensable bundles and their thresholds change with that reset.

For venture jj at time tt, with functional stage sjts_{jt}, I write stage-adjusted productive progress as

Yjt=ΩjtHjtGjtmin⁡k∈Ksjt(Cjktτksjt). Y_{jt} = \Omega_{jt} H_{jt} G_{jt} \min_{k\in\mathcal K_{s_{jt}}} \left(\frac{C_{jkt}}{\tau_{ks_{jt}}}\right).

Ω\Omega sets the economic scale of the opportunity in the same units as YY. HH and GG are multipliers between zero and one: HH captures hard gates, such as required regulatory approval, while GG captures the quality of critical execution, such as whether an experiment was valid. They act outside the capability bundles; failing a gate or a critical task can reduce output even when bundle capacity is high. CjktC_{jkt} is the effective capacity of indispensable bundle kk, while τksjt\tau_{ks_{jt}} is the threshold that bundle must meet at the current stage. The outer minimum is the Leontief part. Progress is constrained by the lowest indispensable capacity relative to its stage requirement, rather than by an average of the company's strengths.

Within a bundle, inputs may partially substitute for one another. The companion Leontief-CES foundation expresses that inner capacity as

Cjkt=[∑m∈Mkαkms(xjmt−x‾kms)+ρks]1/ρks,σks=11−ρks. C_{jkt} = \left[ \sum_{m\in\mathcal M_k} \alpha_{kms} \left(x_{jmt}-\underline{x}_{kms}\right)_{+}^{\rho_{ks}} \right]^{1/\rho_{ks}}, \qquad \sigma_{ks}=\frac{1}{1-\rho_{ks}}.

Here xjmtx_{jmt} is an input's quantity or quality, x‾kms\underline{x}_{kms} is its floor, (z)+=max⁡(z,0)(z)_+=\max(z,0), and αkms\alpha_{kms} is a stage-specific weight. σks\sigma_{ks} governs the degree of substitution within the bundle. Founder selling can temporarily substitute for a formal sales organization. AI agents can take on portions of research or implementation. Contractual guarantees can sometimes substitute for balance-sheet capital. None of those substitutions is unlimited. Coding capacity cannot repair an incoherent product architecture, and more sales cannot create value that customers do not receive.

The normalized ratio and current binding constraint are

rjkt=Cjktτksjt,bjt=arg⁡min⁡k∈Ksjtrjkt. r_{jkt}=\frac{C_{jkt}}{\tau_{ks_{jt}}}, \qquad b_{jt}=\arg\min_{k\in\mathcal K_{s_{jt}}}r_{jkt}.

The shortfall against a stage threshold is Δjkt=max⁡(0,τksjt−Cjkt)\Delta_{jkt}=\max(0,\tau_{ks_{jt}}-C_{jkt}). For an intervention uu, the companion foundation writes the changed capability as Cjkt′=Cjkt+ΔCjkt(u)C'_{jkt}=C_{jkt}+\Delta C_{jkt}(u) and the immediate change in progress as

ΔYjt(u)=Yjt(Cj1t′,…,CjKt′)−Yjt(Cj1t,…,CjKt). \Delta Y_{jt}(u) =Y_{jt}(C'_{j1t},\ldots,C'_{jKt}) -Y_{jt}(C_{j1t},\ldots,C_{jKt}).

Holding the other terms fixed, adding capacity to a bundle that is well above the unique minimum may have no immediate effect on YY. That is the mathematical reason a busy department can report excellent local performance while the company's economic output barely moves. The intervention has to reach the present bottleneck, prepare for the imminent one, or change the shape of the function.

The company also needs to estimate the imminent constraint: the factor likely to become binding after a successful intervention or stage change. If customer access is binding now, founder-led sales may move it. Implementation capacity may then bind next. Building that capacity before the sales motion accelerates avoids discovering the constraint through failed deliveries.

For a stage change, the companion foundation estimates the next binding constraint as

b~jt=arg⁡min⁡k[Cjkt+E[C˙jkt]Δtτk,sjt+1]. \tilde b_{jt} = \arg\min_k \left[ \frac{C_{jkt}+\mathbb E[\dot C_{jkt}]\Delta t} {\tau_{k,s_{jt}+1}} \right].

For an intervention within the current stage, substitute the changed capacities Cjkt′C'_{jkt} and keep the current-stage thresholds τk,sjt\tau_{k,s_{jt}}. The next-stage calculation is an estimate, not a promise that capacities will grow at the expected rate. It forces the analyst to look at the threshold the next stage will demand before the present bottleneck has been fully relieved.

The Constraint Moves

Revenue below plan is a symptom. It may come from distribution, product fit, pricing, fulfillment, or a customer workflow that changed underneath the company. A slow release cycle can be caused by too few engineers, but also by ambiguous architecture, weak verification, or a handoff system built for a different cost of software production. Diagnosing the constraint by the department that owns the symptom is how companies optimize the wrong input.

Recursive Production Function Analysis, or RPFA, starts by defining the economic output to be increased. It maps the customer's production system before the company's. Then it identifies indispensable capabilities, their thresholds, the current functional stage, and several competing explanations for what is binding. The test should distinguish those explanations. An intervention should have a mechanism, cost, expected effect, observation window, failure condition, and achievement gate. Afterward, the analysis starts again with the changed system.

I call that operating loop Constraint Cascade Analysis, or CCA:

map ⟶ hypothesize ⟶ test ⟶ intervene ⟶ observe ⟶ recompute. \text{map}\ \longrightarrow\ \text{hypothesize}\ \longrightarrow\ \text{test}\ \longrightarrow\ \text{intervene}\ \longrightarrow\ \text{observe}\ \longrightarrow\ \text{recompute}.

The loop can advance the company to its next functional stage. It can stabilize a repeatable business. It can show that the opportunity has failed. It can also lead to a more disruptive result: the intervention may change which factors are needed to produce output. At that point the appropriate response is to reconstruct the production function. I distinguish three levels in the analysis. Level 1 detects the constraint. Level 2 removes it. Level 3 asks whether we would build the same company, product, and production architecture from scratch now that the intervention has changed what is possible.

An RPFA record begins with one measurable economic output. "Grow sales" is not enough; "increase the rate at which a customer can turn scarce expert judgment into reliable model-evaluation signal" gives the analyst a production system to inspect. The analyst maps that customer's inputs, transformations, handoffs, gates, and failure modes. Only then does she map how the focal company acquires customers, builds the product, supplies labor and software, finances delivery, controls quality, and governs decisions. She scores the eight capability bundles against their stage thresholds, writes rival constraint hypotheses, and asks what observation would distinguish them. After intervention, she records the actual result, the newly binding constraint, what became abundant or disappeared, and whether the old architecture remains justified.

The durable achievement-gate record is compact enough to reuse:

constraint:
evidence:
alternative_hypotheses:
intervention:
  mechanism:
  cost:
  expected_output_change:
achievement_gate:
observation_period:
actual_result:
new_binding_constraint:
production_function_change:
reusable_learning:

The record matters because an unsuccessful diagnosis is still information. A company that calls weak revenue a sales problem may hire a CRO while the product fails to fit the customer's workflow. A software problem can be handed to Engineering when its cause is a decision or verification process. A company can keep adding integrations to a product abstraction AI has made obsolete. A venture can convert capital into headcount, quotas, and board forecasts, then discover that those commitments make a necessary redesign politically and financially harder. Those are different failures, but each begins by treating the visible departmental symptom as the factor that actually limits output.

Worked Example: Centaur.ai to a Hypothetical Centaur 2.0

Centaur.ai provides a concrete starting point for an outside-in exercise. Its visible platform organizes expert annotation around projects, cases, qualified readers, calibration, and aggregated judgments. I am not describing Centaur.ai's undisclosed strategy or claiming it intends to pursue the path below. "Centaur 2.0" is a hypothetical company generated by the constraint cascade. The original intellectual capability I am examining is the ability to measure heterogeneous expert performance and combine qualified judgments into higher-confidence ground truth.

Begin with the customer's production system. A frontier-model lab acquires and curates data, pretrains a model, generates tasks, fine-tunes, produces preference and critique signal, evaluates, red-teams, studies failures, and then trains again. As the model improves, generic human labeling becomes less valuable. In oncology, advanced tax law, mathematical proof, or semiconductor process engineering, the lab needs experts capable of finding errors that competent nonspecialists miss. The customer constraint migrates toward reliable frontier-domain judgment.

The first business hypothesis is straightforward: sell Centaur's existing expert-data capability to OpenAI or Anthropic. Before hiring more salespeople, ask what would prevent fulfillment if a lab placed a large order. Sales access, annotation workflow, customer integration, qualification, and throughput are rival explanations. The higher-order hypothesis is that the binding factor is elastic acquisition of the right expert cognition. A system that aggregates good answers cannot produce them if the requisite people cannot be found, tested, and put to work at the scale and time the customer requires.

The first intervention is therefore an expert-supply production system. It needs expert identity, a skills graph, sourcing agents, credential verification, task-specific testing, performance scoring, task matching, capacity state, contracting, fraud detection, and adjudication. The operating chain becomes: map the need for expertise, source candidates, verify credentials, test competence, contract capacity, route work, measure performance, and update the expert model. Centaur would have built a new capability to manufacture access to expertise, even if its first sale still looked like an annotation project.

Relieving sourcing exposes the second constraint. The best experts may not be available when a frontier lab suddenly needs them. Availability retainers, minimum guarantees, capacity options, pre-agreed rates, and activation service levels reserve future expert time. The asset is now contracted cognitive capacity. It carries a cash requirement before the lab consumes the work.

The third constraint is financing that reserve. The hypothetical company may be unable to hold enough idle expert capacity on its own balance sheet. A frontier lab's minimum purchase agreement could support a separate financing vehicle or warehouse. The commitment supplies offtake certainty; the financing structure pays to reserve capacity; Centaur's software allocates it when demand arrives. The lab need not directly subsidize the reserve for its contract to make the reserve financeable. Financing has entered the production function as a designed input.

The network then creates a fourth possibility. Calibrated experts can design difficult domain evaluations, adjudicate the answers of frontier, mid-tier, and small models, and identify recurring failure classes. Instead of waiting for labs to request an annotation project, Centaur 2.0 could publish a recurring map of frontier-domain cognition: where models regress, where smaller models safely substitute for frontier models, and where expert supervision remains necessary. Those findings could recruit the next set of experts, generate lab demand, and produce targeted evaluation or remediation data. The marketing artifact and the production system would begin to coincide.

An enterprise product could grow from the same evidence. A small model may be adequate for extraction while a frontier model is needed for difficult clinical reasoning. A capability map could route each workflow to the cheapest model that crosses an expert-defined quality threshold. Repeated cycles would accumulate a graph connecting model, domain, task, failure mode, expert, expert performance, corrective signal, and later model improvement. The durable asset would no longer be a pool of labelers. It would be a software-controlled capacity for discovering, provisioning, measuring, financing, and deploying scarce cognition against a moving frontier.

That change also alters acquisition logic. A buyer might value the standalone cash flow, the time saved in building such a system, the expert-performance data, faster model-development cycles, and the strategic cost of a competitor owning the network. The strongest position might be a neutral provider used by several competing labs, even though that neutrality could make the company especially attractive as an acquisition target. None of these outcomes is a forecast for Centaur.ai. They are the possible consequences of following the constraint sequence: expert supply, availability, reserve financing, capability measurement, and a second commercial surface. The hypothetical platform is the accumulated machinery built to remove the constraints.

Capital participates in this process. A Series B round may supply runway, induce a large sales organization, raise fixed burn, and place growth promises into board forecasts. The new commitments can make the next necessary reconstruction financially difficult. Runway is strategic optionality as well as cash. The firm cannot treat capital structure as an external spreadsheet while analyzing every other production factor.

The railroad transition gives me a physical analogy for the reconstruction problem. Steam locomotives required firemen, coal depots, water stations, ash handling, boiler shops, and steam-specific maintenance. Diesel traction changed the locomotive and the complementary infrastructure around it. A railroad that merely asked how diesel could help its firemen shovel coal faster would have kept optimizing a factor the new engine made unnecessary. It also had to finance the replacement while carrying the old system on its books. AI-native production can create the same conflict between a new production function and an incumbent's invested organization.

The question for a board is therefore sharper than whether its members have impressive operating experience. Which production regime generated that experience? Have they themselves reorganized a company after a factor-of-production shift, or are they importing the scaling playbook that worked under the previous one? A board can reward growth at Stage 5 when the company has functionally returned to Stage 3. A capable executive can become the steward of sunk production-model debt. Legacy code is only one form of that debt; product abstraction, distribution commitments, organizational design, and capital structure can all bind the next move.

This also changes the people I would recruit into the central operating unit. Domain knowledge still matters, but the scarce phenotype is the ability to enter a new domain, learn its grammar, identify the current and imminent constraint, build the missing capability, and revise the production system when the result proves the original design obsolete. A company should be able to describe the intervention and the evidence that changed its mind, not simply point to a new tool in every employee's browser.

The Same Theory, One Altitude Down

The firm is not the only unit to which DCPS applies. A single at-scale production run also turns inputs into an output through indispensable stages. Extracting, classifying, embedding, scoring, and resolving a corpus are sequential complements. In the nested CES-Leontief view, the outer nest has very low substitution elasticity across stages: σouter≈0\sigma_{\mathrm{outer}}\approx 0. More embedding throughput cannot compensate for scoring that has not occurred. Within a stage, workers, per-worker compute, shared-store IO, and vendor tiers have finite substitution possibilities: σinner\sigma_{\mathrm{inner}} must be estimated where the run intends to rely on it. A shared write primary can become an input to every stage, and an external batch API can impose a latency floor that our own worker count cannot move.

This lower altitude gives the theory an empirical instrument. At the firm level, substitution elasticities often take quarters to infer. In a pipeline, a controlled canary can step worker count from two to four to eight and measure what additional workers actually buy. It can test whether a proposed lever has a positive effect under the production conditions that matter. A read replica is a fake lever for a write-bound primary. More concurrent batch jobs may also be fake if the vendor's active-token ceiling remains fixed. A model that assumes those levers work will produce a confident plan for a run that cannot finish on time.

I call the canary a σ\sigma-meter. It tests or bounds the substitution claims on which the run depends. A measured response is not automatically the CES substitution elasticity; identifying σ\sigma requires a controlled test of how inputs replace one another under the specified production model. A coefficient such as dollars per atom or writes per atom can be read from a bill or an instrumented run. An elasticity of substitution asks how readily one input can replace another while maintaining output. The canary returns GO when the lever has measured headroom, NO-GO when the supposed lever does not move the constraint, and CANNOT-RUN when the effect could not be measured. The measurement needs a provenance stamp: instance class, vendor, corpus size, cache state, and other conditions under which the result was observed. If worker count and cache warmth change together, the resulting slope is observational. We should record the confound and use the slope as a bound until a controlled step isolates the effect.

There is an economic reason to model the run before firing a large discovery canary. If we learn the constraints only by hitting them at scale, the first run pays tuition and later runs inherit the lesson. A constraint model can name the uncertain elasticities and the smallest experiments that would settle them. The canary then pays a verification fee before the first full run. It still catches genuinely new failure modes. The measured coefficients and elasticities refit the model afterward, so the knowledge compounds in a reusable artifact instead of disappearing into chat or an operator's memory.

Worked Example A: The O0 Transcript Ingest

The O0 production run, our earnings-transcript ingest, was intended to turn a roughly 240-company corpus into fully processed semantic atoms. The measured population was 1,255,332 atoms from 4,719 transcripts across 228 entities.[1] Its stages were extraction, classification, embedding, asynchronous Gemini batch scoring, and three speaker-resolution passes. Each stage read from or wrote to one Postgres primary, a db.r6g.large with two vCPUs and 16 GB of memory. During embedding, a diagnostic count query over roughly 2.25 million rows stalled for more than seven minutes.

Three explanations were live: perhaps the embedding provider's rate limit was binding; perhaps the database was overloaded; perhaps the shared write primary and an unindexed full-table scan were the binding mechanism. Embedding calls returned in less than a second, weakening the first hypothesis. The writers targeted disjoint rows and showed no lock-conflict failures. The full-table diagnostic scan itself competed with their writes, while the first workers sampled had been cold-started. The slow query and the initial CPU reading were observations, not yet a causal diagnosis of the run's constraint.

We stopped using the count query as the progress instrument and read the workers' own logs. Warm workers processed about eight to eleven atoms per second, compared with roughly 1.5 to 2.5 for the initial cold sample. The database was near 40 percent CPU with 17 to 18 connections. Telemetry suggested roughly 2.1 percent CPU per warm worker and about 4.8 percent per cold worker. Extrapolated toward an 85 percent operating ceiling, that suggested room for roughly 34 to 38 warm workers, provided they were ramped gradually. Worker count and cache warmth had moved together, so those slopes were provisional bounds, not a cleanly identified CES substitution parameter. A later canary would need to vary one while holding the other fixed.

The immediate panic had been partly an artifact of the instrument used to observe the system. Adding workers appeared to be a real lever within headroom, worth approximately a doubling of embedding throughput if launched with a staggered ramp. A cold simultaneous launch near the extrapolated worker ceiling could have driven the primary toward saturation even though the same number of warm workers might fit. Adding read replicas to solve write pressure was a fake lever. The projected dollar cost of this phase, around $545 against a $2,900 ceiling, was slack. Model quality and embedding-vendor rate were not binding in this phase. The imminent operational constraint was the external Gemini batch-scoring latency floor. A future write-primary constraint could eventually require a sharded, columnar, or other store architecture rather than simply a larger instance. That is Level 3 reconstruction at the pipeline altitude, recorded as a trigger rather than executed because of one alarming count query.

Worked Example B: One Archetype, Three Runs

These runs should accumulate into a library of production archetypes, classified by constraint signature rather than by the superficial kind of data they process. The O0 transcript pipeline is a fan-out ingest that embeds, batch-scores, and then resolves speakers, all funneling through one write primary. Its archetype contains the dependency graph; the distinction between complementary stages and substitutable within-stage resources; known fake levers; slots for dollars per unit, writes per unit, batch caps, cold-start multipliers, and local elasticities; a minimum canary plan; and an explicit trigger for reconstructing the shape.

A 10,000-document ingest could share the same write primary and external batch floor while increasing fan-out and removing the speaker-resolution tail. A podcast pipeline could prepend speech-to-text and inherit the rest of the structure. These are three instantiations of one archetype if their binding inputs and substitution structure remain the same. The podcast run should not inherit a numerical worker elasticity measured under a different instance class or cache state as though it were a universal constant. It inherits the hypothesis, known fake levers, and provenance-stamped priors, then verifies the new speech-to-text coefficient and any elasticities whose conditions have changed.

The reusable recipe has an actual shape:

archetype: fan-out ingest with batch scoring
topology: (pre-stage?) -> extract -> classify -> embed -> batch-score -> (resolution?)
outer nest: complementary stages, sigma approximately 0
inner nest: workers, vCPU, shared-store IO, vendor tier
structural bottleneck: shared write primary
external floor: asynchronous batch latency
known fake levers: read replica for writes; more batch jobs past vendor cap
slots: dollars/unit, writes/unit, store ceiling, batch cap,
       worker-to-store response, cold-start multiplier, stage thresholds
canary: smallest controlled steps that settle uncertain slots
reconstruction trigger: would we build this shape again from scratch?

Worked Example C: A Different Archetype

Real-time inference serving can also be drawn as a pipeline of incoming requests and outputs. Its constraint signature is different. Accelerator throughput and p99 latency may bind; there is no asynchronous batch wait of the O0 kind, because latency is part of the product. Stateless replicas can increase serving capacity. A read replica or additional serving replica that was a fake lever for the write-bound ingest may be a real lever here. The recipe needs tokens per second per accelerator, a batch-size-versus-latency curve, and an autoscaling response measurement. Surface similarity between the two pipelines does not justify borrowing the wrong intervention.

The concrete pipeline record should preserve the economic output, selected archetype, stage DAG, operational and structural constraints, alternative hypotheses, proposed lever, measured elasticity or bound, confound, provenance, current slack, imminent constraint, Level 3 reconstruction trigger, achievement gate, observation window, result, and the library entry to refit. That is enough structure for the next run to inherit knowledge without inheriting false confidence.

run_id:
archetype:
economic_output:
stage_map:
binding_constraint:
  operational:
  structural:
symptom_vs_constraint:
alternative_hypotheses:
intervention:
  mechanism:
  cost:
  expected_effect:
elasticities_measured:
  - lever:
    sigma_or_bound:
    verdict:            # GO, NO-GO, or CANNOT-RUN
    confound:
    provenance:
slack_inputs:
imminent_constraint:
level3_reconstruction_trigger:
achievement_gate:
observation_window:
actual_result:
refit_to_library:
reusable_learning:

The firm altitude and the pipeline altitude feed each other. A company-level decision to make experimentation cheap has to cash out in measured cost, time, and reliability of actual runs. The pipeline measurements, in turn, change what the firm can afford to test, build, and sell. DCPS is one theory operating at two resolutions.

The Unit That Can Reorganize

AI changes the substitution possibilities inside many capability bundles. In software, code execution can become abundant while architecture, decomposition, verification, and judgment remain scarce. A company that gives its engineers coding assistants may produce more code through the old ticket-and-sprint system. A company able to redesign production around agents, tests, and human adjudication changes the cost of experiments and the organization needed to run them. The latter change reaches the production function itself.

The old software-production chain often moved from requirement to product manager, ticket, sprint, engineer, code review, QA, and release. An AI-native chain can begin with intent and architecture, dispatch parallel agents, test their implementations, adjudicate the result, and deploy. The second chain does not remove the need for product judgment or verification. It changes their relative scarcity. It also changes the sensible team size, roadmap economics, capital allocation, and timing of commercialization. An enterprise that measures only hours saved by coding assistants may miss the value of reorganizing the entire chain.

My candidate operating unit is a small governed band of roughly three to five capable humans working closely with AI agents. It needs enough authority to cross functional boundaries and enough discipline to test its own diagnoses. Its job is to enter an unfamiliar domain, acquire the relevant grammar, map the customer and company production systems, generate rival constraint hypotheses, build or acquire a missing capability, verify what happened, and then change direction when the system state demands it. The number of humans is a design proposition. The capability-generation loop is the economic claim.

This is also the Pathfinder implication of the theory. The person to seek is not simply the best operator of a known department. It is someone who can repeatedly identify the next binding constraint, generate or acquire the capability required to remove it, observe the changed system, and reconstruct the production function when the old one no longer fits. The unit needs people who can learn across domains, reason causally, falsify their own explanations, understand capital, and work symbiotically with AI.

So I return to the AGI question. We have long had generally intelligent humans. We have also had firms that struggle to redirect their intelligence when the world changes. Better models increase the amount of cognition available and reduce its price. The result depends on whether a governed unit can locate the constraint, move it, and recompute the system it has changed. That unit, rather than the isolated model, is where I expect general intelligence to become a different way of producing economic value.

Operational Appendix

The firm-altitude analysis can be executed without turning the theory into a slogan. Define the measurable economic output; map the customer's production system; map the focal company's; establish its functional stage; score the eight capability bundles, hard gates, and threshold ratios; generate at least three binding-constraint hypotheses; and identify the evidence that would distinguish them. Estimate the present and imminent constraints. Ask what AI has changed about substitution within each bundle. Specify the cheapest intervention that could move the constraint and the observation that would count as success. After execution, preserve rejected hypotheses, separate observation from inference and speculation, recompute the ratios, and decide whether the firm should advance, stabilize, reconstruct, or terminate the attempt.

At the pipeline altitude, match a run to an archetype by its constraint signature and record that match as a correctable decision. Populate coefficient slots from grounded evidence. Mark elasticity slots as measured or assumed. Test each assumed lever with the smallest controlled canary that can change a decision; stamp the result with the conditions of measurement. A GO permits the scoped scale run, a NO-GO rejects the lever the run was counting on, and CANNOT-RUN remains an unresolved gate. The full run should use production-identical configuration. After it finishes, recompute the binding and imminent constraints, ask whether the archetype's shape is still right, and refit the library. This prevents a successful old recipe from becoming a machine for faithfully optimizing an obsolete production system.

This essay follows DCPS v0.2. Its firm-altitude argument derives from RPFA v0.1; the pipeline altitude, the σ\sigma-meter, and the archetype library are the v0.2 extension. The Centaur 2.0 path is hypothetical. The O0 figures describe one measured run, while its worker-response slopes remain observational bounds pending a controlled step-test.

Footnotes

  1. These measurements describe the O0 run under its then-current infrastructure, corpus, cache state, and workload conditions. They are historical observations, not a description of InflectAI's current architecture or universal performance characteristics. ↩︎