Architecture and Arithmetic

The section index attributes the architectural half of this program to Buchanan, Pai, Wang, and Ma’s Principles and Practice of Deep Representation Learning. Its CRATE derivation organizes network layers as steps of an optimization procedure for sparse rate reduction. The objective favors representations on low-dimensional, incoherent subspaces. Subspace self-attention approximates a compression step, while the feed-forward block performs a sparsification step. The resulting CRATE family gives these operations an explicit optimization interpretation.

A derived architecture makes the intended role of its parameters explicit. That account can guide implementation and expose opportunities for specialization. It does not by itself prove that every finite trained model reaches the objective’s target geometry. The derivation, the constraints enforced by the implementation, and the representation attained through training remain separate things to check.

The interest in separated structure connects this program with ADM. The ADM substrate, collected in A Deeper Dive, can enforce a justified block decomposition: a block-diagonal generator has a block-diagonal exponential, and a checked support rule can make off-block entries absent. CRATE’s learned representation subspaces and an ADM’s permitted parameter blocks are different objects. Relating them requires an explicit mapping and evidence that the restriction retains the behavior the task needs. The positional-encoding analysis explores a concrete subsystem where algebraic structure can inform that design.

That construction restricts the admissible weight set; a probabilistic model must additionally choose a prior measure. For a Clifford layer between grades kk and k′k', a checked signature and its support rules may specify Wadm\mathcal{W}_{\mathrm{adm}}. If a chosen base prior π~\tilde{\pi} gives this set positive probability, it can be conditioned on the restriction:

π(dW)=1[W∈Wadm]π~(Wadm) π~(dW),π~(Wadm)>0. \pi(dW)=\frac{\mathbf{1}[W\in\mathcal{W}_{\mathrm{adm}}]}{\tilde{\pi}(\mathcal{W}_{\mathrm{adm}})}\,\tilde{\pi}(dW), \qquad \tilde{\pi}(\mathcal{W}_{\mathrm{adm}})>0.

Exact structural zeros usually define a lower-dimensional subspace, which has zero probability under an ambient prior with a Lebesgue density. The displayed normalization then does not apply. One can instead choose a prior directly on the permitted coordinates and map it into the full weight space, keeping the forbidden entries absent. The type supplies the restriction; it does not select how probability is distributed among admissible weights.

If Wadm\mathcal{W}_{\mathrm{adm}} is a closed linear subspace, a specified inner product also defines an orthogonal projector Πadm\Pi_{\mathrm{adm}}. Applying that projector to a proposed weight is a deterministic operation. Applying it to samples produces a pushforward probability measure, which is generally different from conditioning a prior. Nonlinear admissible sets need their own parameterization or projection rule; a unique orthogonal projector is not automatic.

Coding-rate compression acts on token representations using learned subspace bases. Its gradient step is not generally a hard orthogonal projection, and it does not condition a distribution over weights. The connection to ADM is an opportunity to combine a justified parameter restriction with representation learning, rather than an identity between these operations. Their different roles let the compiler preserve declared structure while training determines what remains unknown.

Two readings of the same book

The common reading takes the derivation as an interpretability result: CRATE’s attention approximates compression against subspaces, its feed-forward step promotes sparsity, and the coding-rate objective explains their intended roles. The natural next step is to implement the operators as a tensor program and train them. This is what the published implementation does. Ordinary tensor shapes alone do not certify the desired relations among learned subspaces. Such guarantees need additional parameterizations, constraints, or checks, and a compiler needs those facts in a usable form before it can safely exploit them.

The second reading is available to a substrate that can carry structure as a typed, discharged invariant, which is the starting point of our Fidelity Framework. To make this concrete rather than abstract, consider the published CRATE forward pass. The model’s reference implementation is open, at github.com/Ma-Lab-Berkeley/CRATE, and the papers give it in PyTorch-style pseudocode; the shape below follows that reference:

# CRATE forward pass, as published (PyTorch-style pseudocode).
# Each layer: a subspace self-attention step (compression) with a skip,
# then an ISTA step (sparsification).
class CRATE:
    def forward(self, x):
        for ln1, attn, ln2, ff in self.layers:
            x_ = attn(ln1(x)) + ln1(x)   # MSSA: gradient step on the coding rate
            x  = ff(ln2(x_))             # ISTA: soft-thresholding toward sparsity
        return x

The structure at issue lives inside attn: the Multi-head Subspace Self-Attention operator compresses tokens against learned subspace bases, written UkU_k in the paper, one per head. Incoherence describes their separation; it is not generally the same as exact orthogonality or structural zero interactions. The reference implementation’s learnable tensors do not by themselves certify a fixed block decomposition. Whether training attains the desired geometry involves the objective and optimization as well as numerical error.

A Fidelity specialization can carry a block decomposition when that restriction is justified for the task. A grade annotation alone does not prove arbitrary learned bases incoherent. The following illustrative operator assumes a separately checked support contract for the family of bases and compatible compression and aggregation operations. It restricts the admissible model, so its attainable representation quality must be compared with the original operator:

// Assumes a checked block-support contract for the heads, beyond their element grade.
let mssaStep (heads: GradedSubspaceBasis<Bivector>[]) (z: TokenField) : TokenField =
    heads
    |> Array.map (fun u ->
        // compression against this head's subspace, the coding-rate gradient step
        z |> compressAgainst u |> Quire.accumulate)
    |> SubspaceAggregation.byGrade        // head aggregation, structure-preserving
    |> skipConnection z                   // the "+ ln1(x)" of the published step

 

The proposed specialization can use a coding-rate objective while keeping its declared support through training and lowering. That support-preservation claim needs proofs for the permitted updates and transformations; it does not establish convergence to the unconstrained objective’s optimum. Reading the rate-reduction principle and unrolled derivation as design inputs gives our framework four distinct roles:

  • Dimensional and grade types, together with checked support rules, describe admissible quantities and absent components. They can carry a justified decomposition without identifying it with every optimum of the coding-rate objective. (See A Scaffold for Constrained Models for the scope rule where structure is known in advance.)
  • Geometric algebra provides rotor and generator representations for appropriate transformations, as explored in the positional-encoding analysis. Their algebraic support can force particular zeros; that requires an explicit relationship to the representation subspaces being learned.
  • The Program Hypergraph turns the book’s layered computation, which the dense-tensor substrate flattens into matrix multiplies, back into the multi-way relationships it actually holds. A transformer’s attention is a multi-way relationship among tokens; the dense lowering decomposes it into pairwise operations and loses the structure, exactly the join/split decomposition cruft the PHG will not admit. The PHG carries the relationship intrinsically, so the provably-absent interactions are absent from the lowered program rather than small within it, and the graph-coloring parallelization our framework already performs operates on the true structure rather than a flattened shadow of it.
  • b-posit and the quire address accumulation error in those computations. Numerical accuracy and preservation of declared support are distinct obligations; the arithmetic proposal occupies the rest of this article.

The second reading is a proposal to specialize a derived architecture using additional justified structure. The compiler can exploit facts supplied by the representation and its checked operations, while the objective continues to guide what is learned. A Gaussian-mixture model of representations can also inform our domain models, but shared distributional vocabulary does not make the two models identical: their variables, parameters, priors, and observation models must be related explicitly.

The layer’s other operator reframes the same way: CRATE pairs each compression step with a sparsification step, the ISTA block from model/crate.py.

// D is the learned dictionary; one proximal-gradient step toward sparsity.
let istaStep (d: Dictionary<BPosit>) (lambda: BPosit) (step: BPosit) (x: TokenField) : TokenField =
    let dx   = d * x                          // D x
    let dtdx = Dictionary.adjoint d * dx      // Dᵀ (D x)
    let dtx  = Dictionary.adjoint d * x       // Dᵀ x
    let grad = step * (dtx - dtdx) - step * lambda    // negative-gradient update, in the quire
    x + grad |> TokenField.map (max BPosit.zero)      // ReLU: soft-threshold toward sparsity
 

Why convergence is not enough on its own

The white-box guarantees are real and they are soft. The subspaces the coding-rate objective separates are orthogonal at the optimum in exact arithmetic. Trained in IEEE-754 floating point, they are approximately orthogonal, and the gap between “orthogonal” and “approximately orthogonal” is filled by the numerics. The objective is built on log-determinant and covariance terms, which are long accumulations, and long accumulations in floating point are where cancellation between terms of opposite sign loses the most precision. The structure does not collapse. It blurs, and no check in the substrate flags the blur, because the theory never claimed exactness.

For interpretability research that blur is acceptable; an approximately separated representation is still interpretable. For a component meant to sit adjacent to the ADM constellation, where the neighboring domain models carry exact, SMT-discharged invariants, an approximate substrate is the wrong tradeoff, and it is the same failure mode our framework already identified for learned positional-encoding generators, where a data-dependent generator drifts under floating-point training and cross-block contamination accumulates the way grade corruption does.

The blur comes from the arithmetic the architecture is conventionally trained in, not from the architecture itself, and different arithmetic sharpens the convergence.

b-posit and the quire close the gap

The substrate the ADM work already uses, b-posit arithmetic with quire accumulation, is built for the operations the coding-rate objective stresses. A quire is a wide fixed-point accumulator that carries a long sum or a dot product without rounding at each intermediate step, rounding only once at the end. The log-determinant and covariance computations that the rate objective depends on are exactly such accumulations, so they are what the quire protects.

// The rate term's long accumulation, carried through the quire and rounded once at the end.
let logDetThroughQuire (cov: Matrix<BPosit>) : BPosit =
    cov
    |> choleskyDiagonal          // the diagonal whose log-sum is the log-det
    |> Quire.sumOfLogs           // accumulated without intermediate rounding
    |> Quire.round               // a single rounding, at the end
 

The contrast with the common reading is the same operation built two ways. The dense-substrate reading computes the rate term as a dense floating-point reduction, correct in expectation and quietly lossy in practice; the Fidelity reading computes it as a quire accumulation over grade-carrying quantities whose grade structure is known before the reduction runs:

// The common reading: a dense float reduction, lossy in the tails.
let logDetDense (cov: float32[,]) : float32 =
    let mutable acc = 0.0f
    for i in 0 .. dim - 1 do
        acc <- acc + log (choleskyDiag cov i)   // rounds every iteration
    acc

// The Fidelity reading: the covariance is typed by its grade structure, so the
// block-diagonal form the derivation promises is a property of the type
let logDetFidelity (cov: GradedCovariance<Bivector>) : BPosit =
    cov
    |> GradedCovariance.blockDiagonal     // off-block zeros are type-level facts
    |> Quire.sumOfLogs                    // exact accumulation over the blocks
    |> Quire.round

The difference is not micro-optimization. In the dense version the block structure is an aspiration about the values that finite-precision training erodes; in the Fidelity version it is a fact about the type that training cannot touch, because the off-block interactions the book’s derivation says should vanish are not small, they are unrepresentable.

The approach is therefore to remove the cause of the floating-point slack rather than tolerate it. Keep the derived architecture exactly as the white-box derivation gives it, and run its sensitive operations on arithmetic whose accumulation discipline makes the convergence sharp. This is the first point at which the language-model component stops being the one piece of the framework that runs on a foreign numeric format. It rejoins the b-posit world that the domain models, the dimensional types, and the rest of the substrate already inhabit, which is a precondition for the adjacency the constellation article describes.

The Deployment Tension b-posit Resolves

The building article named a real tension: the CPU deployment target wants four-bit or ternary weights, and those are the regimes where the rate-reduction operations are worst-conditioned. The b-posit substrate is the resolution, because it offers dynamic range that fixed low-bit integer formats cannot, and the borrowed ternary format was never more than a terminal artifact someone else’s pipeline produced. Building the model on our framework’s own arithmetic makes the deployment numeric format a free variable chosen for the framework’s reasons rather than inherited from an external recipe.

One friction is not resolved, and the article states it as the open question it is. Posit precision is not uniform. It is densest near magnitude one and tapers toward the very large and very small. Whether that taper aligns with where the coding-rate objective concentrates its numerical stress during training is an empirical question about the interaction of two specific designs, Gustafson’s tapered precision and Ma’s rate objective. If the objective’s stress falls near magnitude one, where posit is densest, the synthesis is clean. If it falls in the tapered tails, the quire-mediated accumulation has to carry it, which is what the quire is for. The favorable case gives sharp convergence at low parameter count; the unfavorable case gives sharp convergence at the cost of more quire-mediated work. Distinguishing them is one bench experiment, and it is the one that decides whether b-posit is the right substrate for this architecture or merely a defensible one.

Foundation for the Rest of the Section

A derived architecture on precise arithmetic is the foundation the remaining articles stand on. The forward-mode article depends on the derived structure being low-rank, so that the gradient can be taken over few directions, and on the arithmetic being precise, so that the accumulated tangents can be trusted. The constellation article depends on the shared b-posit substrate, because that shared substrate is what lets a non-typed language component and a grade-carrying domain model exchange values without a numeric impedance mismatch. And the reversibility article depends on the quire making a state transition’s round trip exact rather than approximately reversible.

It also sets up the section’s sharpest efficiency contrast, developed in the constellation article. The two readings of the book diverge most consequentially on sub-quadratic attention. The dense-tensor reading that flattens attention into all-pairs matrix multiplies is the quadratic cost; the field’s escape from it, the linear-attention and state-space families whose current frontier is Mamba-3, replaces all-pairs attention with a learned data-dependent generator. Mamba-3 has converged on exactly the complex-valued rotational generator our framework types, bridging it to RoPE, while still listing as open the two problems the framework’s substrate addresses: state tracking, and the gap between linear-in-theory and efficient-in-hardware inference. The structured reading set out in the constellation article gets the sub-quadratic cost from the generator the field has converged on, and it takes exactness along with it: the generator whose decomposition the grade types hold exact is the one that makes the recurrence sub-quadratic. The field only reaches for that exactness by experiment. The two results follow from one set of interactions: the interactions the derivation proves absent are the interactions a quadratic model spends time computing and a drifting sub-quadratic model spends capacity suppressing, and a structured model does not represent.

Open questions

Whether posit’s tapered precision aligns with the rate objective’s numerical stress, or whether the quire must carry the tails, is a bench experiment in waiting.

Determining how a derived architecture trained on b-posit reaches target representation quality at lower parameter count than the same architecture on floating point, as the noise-hedge argument predicts, is measurable on the same bench.

And how might a causal CRATE variant’s rate operations remain well-conditioned under the framework’s arithmetic across the full sequence length, or degrade with context, is a conditioning question specific to the sequence case. All of these are worthy subjects in pursuit of what Ma frames as “AI 2.0.”