← The Intelligence Stack

AI economics and architecture

For founders, operators and technical teams building AI-assisted products that need to stay useful, reliable and economically sane as model capability and usage costs diverge.

AI made code cheap. Now we have to engineer the cost of intelligence.

Generating software is getting cheaper. The harder problem is deciding how much machine intelligence each part of a system deserves. Code, bounded decision models, open-weight models, frontier reasoning, agents and humans have different costs and failure modes. Good architecture routes work to the simplest reliable layer and escalates only when the problem earns it.

Code got cheaper

AI lowered the cost of implementation before it lowered the cost of judgment.

A founder can now generate working application code without spending years learning every framework underneath it. Junior developers can move through problems that used to require more experienced help. Senior engineers can produce far more implementation with fewer keystrokes.

That compresses the value of code production. It does not remove the need to decide what should exist, what should be deterministic, what can fail safely and what the business can afford to run at scale.

Implementation

Generating components, APIs, tests and glue code is increasingly cheap and fast.

Judgment

The harder work is deciding whether the generated system is correct, useful and safe enough to trust.

Architecture

The system still needs boundaries, state, escalation, observability and economic constraints.

Taste

Good software still depends on knowing what should be built and what should be left out.

Cognition has a budget

Intelligence is becoming another infrastructure resource.

Modern systems no longer have one generic AI layer. A workflow may use deterministic code for hard rules, a bounded decision model for classification, an open-weight model for local or repetitive work, a general model for synthesis and a frontier model for difficult reasoning.

Those layers differ in price, latency, reliability and control. The engineering question becomes less about whether AI can perform a task and more about the cheapest reliable mechanism that can perform it well enough.

Code

Use deterministic software when the answer follows from known inputs and rules.

Bounded judgment

Use classifiers or decision models when interpretation is required but the output space is known.

Generative reasoning

Use larger models when the answer must be researched, synthesized or created.

Human review

Escalate when uncertainty or consequence exceeds the system's approved boundary.

Route before you escalate

A strong system spends expensive reasoning only where expensive reasoning creates value.

Imagine a workflow making 100,000 decisions. Sending every case to the strongest available model may work, but it can also be unnecessary and expensive.

A better architecture may eliminate obvious cases with code, classify routine cases with a lightweight model, send ambiguous cases to a general model and reserve frontier reasoning for the small fraction that genuinely needs deeper analysis.

Filter

Remove obvious cases with deterministic checks before inference starts.

Classify

Handle routine semantic decisions with a bounded model when possible.

Escalate

Move only ambiguous or high-value cases to a stronger model.

Stop

Use a human checkpoint when more model spend cannot reduce the business risk enough.

AI gross margin

Model architecture is becoming part of product economics.

AI-native products add variable costs that traditional SaaS teams did not always have to model explicitly: inference, search, retrieval, enrichment, tool calls, agent execution, verification and sometimes human review.

If a customer outcome costs almost as much to produce as the customer pays for it, impressive model behavior does not rescue the business model. The workflow has to create enough value after its operating cost is included.

Cost per execution

Measure what one completed customer outcome costs across models, tools and review.

Volume sensitivity

A workflow that looks cheap at ten runs may become expensive at ten thousand.

Downgrade paths

Cache, batch, route or replace inference when a cheaper mechanism produces the same useful result.

Payback

Compare the full operating cost against the value the workflow actually recovers or creates.

Engineering moves up the stack

The valuable work is shifting from writing syntax to allocating intelligence.

The easiest part of software development to automate is producing implementation from a clear specification. The difficult work begins where the specification stops being clear.

Someone still has to decide what information an agent can see, which actions it may take, when a decision should be deterministic, which model is enough, what evidence proves completion and when the system must stop and ask a person.

Decision rights

Define what the system may decide and what remains explicitly human.

Model choice

Choose capability based on the task instead of defaulting to the strongest model.

Verification

Require evidence that the workflow completed the intended business action.

Economics

Design the system so useful intelligence remains affordable as usage grows.

The practical rule

Use the simplest reliable intelligence that can carry the decision.

Cheap code generation does not make every software system cheap. It makes implementation easier while pushing more value into architecture, judgment, verification and cost control.

Treat cognition the way you already treat compute, storage and latency: measure it, route it, constrain it and spend more only where the additional capability changes the outcome.

Measure

Track latency, failure, override rate and cost by decision layer.

Route

Send work to the lowest-cost layer that meets the reliability requirement.

Escalate

Use stronger models when ambiguity or consequence justifies them.

Review

Keep human approval where accountability cannot be delegated safely.

Find the first useful system

Start with the workflow, not the tool.

The Pixel & Process assessment looks at how work arrives, where it stalls, what delay costs and which part is actually worth changing first.

Assess your workflow →