Beyond the Bitter Lesson

Beyond the Bitter Lesson: Structural Convergence

Convergence and construction form a continuum, not a binary choice

August 26, 2026·SpeakEZ·11 min read·Updated September 29, 2026

The argument about scale in machine learning leaves an engineering question open: which parts of a problem should a model have to learn? Some relationships are unknown and must be inferred from observations. Others are already available as physical dimensions, algebraic identities, interface contracts, or verified domain constraints. Making those relationships explicit can change both the search space and the computation required to explore it.

Andrej Karpathy’s 2017 essay Software 2.0 gave a clear account of software whose behavior is found through optimization. His account already includes a human-selected architecture that defines the space being searched, and it discusses the difficulty of understanding and trusting the resulting program. The question we pursue is how much further that constructed frame can go: what can it establish before learning, what can it preserve through updates, and what useful computation can those facts remove?

Our position with the Fidelity Framework is that construction and convergence belong in the same system. A model can learn within a representation that rules out specified violations. The remaining learning can still be difficult, and an admissible answer can still be inaccurate. The opportunity is to spend computation on the unknowns while carrying established facts through the whole application.

Mamba-3 supplies a useful example. Its authors introduce complex state transitions, derive their equivalent real block rotations, and connect them to data-dependent RoPE. They construct that structure explicitly and then train within it. The rotational component has a geometric-algebra description through bivector exponentials; the full transition also includes decay and input-driven terms. This is evidence for mathematically informed architecture. However, it’s important to also note that it does not establish that training rediscovered the algebra or that knowing the rotor would eliminate the learning task.

The stronger question is which further consequences of a chosen structure can be made available to the compiler and the training procedure. A representation of rotations, for example, can expose their permitted degrees of freedom. It still leaves the relevant rotations to be inferred, and its implementation must preserve the intended constraints within its numerical contract.

Constructed invariants and the Bitter Lesson

Sutton’s Bitter Lesson challenges attempts to substitute human domain knowledge for methods that scale with computation. Its examples include hand-designed features and representations, and its conclusion reaches as far as our treatment of space, objects, and symmetries. It also acknowledges useful convolutional structure and invariances, and allows that domain knowledge and scalable computation need not conflict. It offers no blanket exemption for something called an invariant.

Our proposed distinction is between a restriction justified by the task and a guess about which answer the task should produce. Dimensional consistency is a clear example. A pressure prediction must have the units of pressure, whatever value the observations support. A checked conservation law can impose another restriction, provided the modeled system and its boundaries satisfy that law. Those constraints leave room for learning while excluding particular classes of error.

This distinction has obligations of its own. A restriction can be wrong for the domain, can exclude a needed solution, or can cost more to enforce than it saves. Calling it structural does not resolve those issues. We have to establish its applicability, show what the construction guarantees, and measure its effect on accuracy and total computation. That is how our proposal should meet Sutton’s challenge.

Start the search from what can be justified, and keep learning what remains unknown.

Structure is meaningful at both ends

The thesis with the Fidelity Framework is that construction and convergence form a continuum. Different parts of an application can occupy different places along it. Units and declared parameter support may admit exact checks. A relationship inferred from noisy observations carries statistical uncertainty. The same model can contain both kinds of information.

Alex Zhang’s analysis of language-model harnesses argues that a harness can help compositional generalization by presenting individual model calls with familiar local tasks, even when the overall task is unfamiliar. That suggests an engineering connection to narrow interfaces and explicit contracts. Keeping a call near its training distribution is an empirical objective, however, and checking a contract establishes only the property that contract states.

Our graph-query worker illustrates the combination. Worked examples guide the generation of a query; gateway validation checks it before execution. The examples and the validation contribute different things. A query can satisfy the gateway’s rules and still ask the wrong question. The surrounding application must retain that distinction when deciding what to trust.

Structural convergence is our name for learning within such a constructed frame. Its value depends on how well the frame fits the problem and what work its guarantees allow us to remove.

A Matter of Phase

Harper, Mitchell, and Moggi’s account of the phase distinction provides a useful connection: a language can separate the information needed for static checking from the computation performed at runtime. This gives us a precise question to ask of a learning system: which obligations can be discharged before execution, and which need evidence or computation that becomes available later?

In our design, dimensional relationships and permitted component support can guide elaboration before a gradient step. Observations later determine parameter updates within that structure. Metadata can be erased where its obligations have been discharged and its runtime representation is unnecessary.

Phase and uncertainty remain separate axes. An exact integer calculation may happen at runtime. A probability distribution may be specified statically while still describing uncertainty. A runtime check may establish an exact predicate. The opportunity is to settle each obligation at the earliest phase that has adequate evidence, while preserving the meaning of what remains open.

The continuum in our pre-prints

Our framework guides develop the construction principle: a representation and its checked operations can carry specific guarantees through execution. Our Adaptive Domain Models work applies that principle to learning and model replacement.

The intended architecture gives a domain model an admissible parameter space, a training procedure, and checks that a candidate must satisfy before it enters the active inference pathway. When structural support rules prove components absent, their arithmetic can be omitted. When an update stays within a permitted linear support, that support can be preserved by construction. Nonlinear requirements, such as unit-rotor structure, need a suitable parameterization or an update procedure with its own preservation argument.

A probabilistic ADM also needs an explicit prior, an observation model, and a posterior representation and update algorithm. Structural checks can establish that a candidate obeys its contracts. Evaluation must establish how accurately it predicts the domain and how well its uncertainty estimates behave. Those are distinct parts of the design, each contributing to the reliability of the whole.

Places on a Dial

The frontier is choosing, stratum by stratum, how much admits exact construction and how much remains open to convergence. Our research on the information once grouped under “grade axes” separates several different structures. Parity follows a group law; blade support admits a lattice of sound enclosures; the algebra’s signature fixes its multiplication rules. Keeping those roles distinct lets each carry the obligations it can actually support.

This matters to the generated computation. If component support proves a product term absent, that term need never be instantiated. A sound support enclosure may still include terms that cancel for particular values. The structural guarantee concerns the terms proved absent; further sparsity may depend on the data.

Giving learning less to reinvent is a useful ambition when it has this concrete meaning. It is possible to remove work by carrying more information. Whether the complete system benefits depends on the cost of obtaining, checking, and maintaining that information as well as the arithmetic it removes. Training time, memory, inference cost, and achieved accuracy all belong in the comparison.

The Probabilistic Stratum

Our dimensional types express relationships between quantities: velocity is length over time, and force is mass times acceleration. Those relationships constrain legal operations. Numerical ranges require further evidence from values, guards, input bounds, and checked domain laws. Knowing that a length is measured in meters does not determine whether it is microscopic or astronomical.

A range states which values are admitted. A probability distribution assigns weight within a specified space. A point value can be represented by a point mass, but a range usually admits many different distributions. Two models can agree on every structural constraint and disagree substantially about what is likely.

A construction constraint can therefore restrict the support of a prior. The model must additionally specify the probability measure within that support and a likelihood for observations. When Bayesian conditioning is well-defined, a prior’s hard support remains a restriction on the posterior. An individual observation can nevertheless increase uncertainty: surprising evidence may make two formerly unequal explanations equally plausible.

A small example makes the contribution of each part visible. Let a parameter pp be a probability, with 0≤p≤10\leq p\leq1. Choose a Beta prior with positive parameters α,β\alpha,\beta, and assume nn conditionally independent Bernoulli trials with common success probability pp. After ss successes, the posterior is

p∣D∼Beta⁡(α+s,β+n−s). p\mid D\sim\operatorname{Beta}(\alpha+s,\beta+n-s).

The range check establishes admissibility. The chosen prior and likelihood give the statistical meaning. Conjugacy and sufficient statistics make this particular update cheap. A type system can carry and check the relevant premises; the interval alone supplies neither the prior nor that algorithm.

Our negative and fractional types research offers another possible connection through obligations to supply or match evidence. Giving such an obligation a type still leaves the conditioning calculation to be specified. Its likelihood, normalization, and treatment of excluded observations need an explicit construction.

Composing probability and learning

A conversation with Paul Snively prompted us to examine the relationship between probabilistic inference and exact logical structure. There is a precise certainty-preserving example: if P(A)=1P(A)=1 and P(B∣A)=1P(B\mid A)=1, then P(B)=1P(B)=1. That resembles modus ponens. It is a useful starting point for investigating the relationship, but it does not determine the categorical structure of every learning algorithm.

Categorical probability gives the connection a more careful vocabulary. A Markov kernel maps an input to a distribution over outputs. For finite spaces, composing two kernels gives

(J∘K)(z∣x)=∑yJ(z∣y)K(y∣x). (J\circ K)(z\mid x)=\sum_y J(z\mid y)K(y\mid x).

This is forward propagation with marginalization over the intermediate state. Bayesian conditioning uses a joint probability model, or a prior and likelihood, to update beliefs after evidence. Cho and Jacobs distinguish channel composition, disintegration, and Bayesian inversion. Keeping those operations distinct tells us what information each component needs.

Point-mass kernels recover deterministic functions: an input xx maps to all its probability at f(x)f(x). Their composition recovers ordinary function composition. Arbitrary relations, which may admit multiple outputs or none, require a different construction. Fritz’s treatment of deterministic morphisms makes this boundary explicit.

A stochastic prediction map also leaves the learning update unspecified. In Backprop as Functor, Fong, Spivak, and Tuyéras supply parameter spaces and update structure and prove a compositional result under stated assumptions. A proposed profunctor account likewise needs its categories, actions, and composition laws defined. Naming a coend does not by itself supply a probability measure, a posterior, or an efficient evaluation procedure.

There are substantive connections between learning and inference. Levine’s account of reinforcement learning as probabilistic inference, for example, develops a maximum-entropy formulation and distinguishes the deterministic-dynamics case from the variational treatment needed with stochastic dynamics. Such correspondences have hypotheses. They provide tools for constructing algorithms rather than establishing that every learning system is performing the same Bayesian update.

For our domain models, the useful consequence is a compositional design obligation. Types can describe admissible values and the contracts of a chosen probabilistic procedure. Priors and likelihoods supply its probability model. Conditioning updates that model, while an implementation supplies the numerical method and its error and resource bounds. Making those relationships explicit can expose opportunities for specialization. It does not settle the cost of inference before that work is done.

Structure’s Return

In his October 2025 conversation with Dwarkesh Patel, Karpathy describes deployment reliability as a “march of nines”: successive improvements beyond a working demonstration require sustained engineering effort. That observation gives our proposal a practical test. Which failures can a construction rule out, and which still depend on statistical coverage of the operating environment?

A sound dimensional check can exclude an invalid quantity combination throughout a program. A validated protocol can prevent a specified invalid transition. Neither establishes that a sensor reading is accurate or that a learned prediction fits a new operating regime. Reliability work includes all of these questions. Structural methods contribute when they discharge particular obligations and keep those results intact as the system changes.

The same discipline applies to control flow and data flow. A useful algebraic description of a numerical kernel does not determine the surrounding application’s dependency order, state ownership, memory placement, communication, or scheduling. Our Weaving the Braid discussion concerns those interacting demands. An optimization method earns its place within that application by fitting its contracts and resource budget.

The Honest Middle

The system we want learns inside a frame whose constraints are justified for its domain. A value outside a soundly enforced constraint is excluded by the representation or rejected at its boundary. Values inside the constraint still need to be assessed for usefulness and accuracy. This is the setting for the domain models and shared substrate developed in our Constrained Machine Learning design.

The practical sequence is to identify an applicable invariant, express it in a representation and its operations, establish what updates preserve it, and measure the resulting system against alternatives. For probabilistic models, that includes the cost and quality of posterior approximation. For deployed applications, it includes the control and data flow around the learner. A smaller hypothesis space is valuable when it retains the behavior the task needs and improves the relevant tradeoffs.

The Bitter Lesson challenges us to show that our constructions continue to earn their place as computation scales. Our response is to make the work about what they remove, the guarantees they establish, and the unknowns they leave explicit. That gives construction and convergence a concrete way to cooperate, with computation directed toward what the model still has to learn. Those lessons we take forward as our unique framework continues to take shape.