Managing Context

From our design perspective, the failure people call a context problem is a property of one recurrence, not a law of cognition. A single recurrent loop processing a long input has to fit every commitment it has made into one state vector, and as the sequence runs the early commitments decay against the later ones. That decay is context rot, the degradation of early-step information across a long closed loop, defined in the recurrence entry. A direct token-based model exhibits the same decay for the same reason, and an entire industry has grown up to slow it.

We read context rot as a consequence of holding everything in a closed loop. Our design opens the loop so that anything which can be consulted is never held, and anything which can be summarized exactly is never re-attended.

The progression the architecture follows

Our recurrence design moves through three stages, set out in full in Structured Recurrence. The Hidden Recurrent Model is the baseline. A single recurrent loop, an MLGRU in the matmul-free lineage our design follows, processes input sequentially with an opaque hidden state, and the loop is closed: it must encode every piece of reasoning, including any domain knowledge it needs, inside its own recurrence. Context rot lives here, in the state vector forced to hold what a long sequence keeps adding to it.

Our Resonant Recurrent Model adds structure inside that loop. The recurrence runs at N resonant levels with learned coupling between them, so information circulates at several timescales at once, the Alpha, Beta, and Gamma rates. This organizes the computation into interacting temporal scales and mitigates the decay within the loop, and the loop stays closed to external state. This Resonant Recurrent Model is the bounded recurrence our sub-quadratic generator carries: a complex-rotational state that summarizes the past instead of re-attending to it, and the type discipline is designed to keep its decomposition exact through training. A bounded recurrence has nothing to evict, because the state already is the summary the past was compressed into.

Our Porous Recurrent Model opens the loop. At designated steps the MLGRU suspends mid-recurrence, emits a structured query to a domain-specific actor, an Adaptive Domain Model, and integrates the structured response as intermediate state before resuming. The query asks for a domain estimate and its uncertainty under a declared probability model. Its interpretation must specify how the recurrent state supplies conditioning evidence; dimensional properties help check that evidence’s compatibility. The response re-enters the loop as a StructuredFact carrying Value, Dimension, Confidence, and Certificate fields over BAREWire, the structured contract. The proposed interface retains the quantity’s dimensional meaning and passes structured values directly into state, bypassing tokenization, embedding, and attention for this exchange. The model can consult a domain specialist rather than encode all domain knowledge in its own weights. The domain model’s Bayesian estimate and uncertainty summary then become inputs to the continuing recurrence under the declared dimensional and coeffect constraints.

  flowchart TB
    subgraph HRM["HRM: a single closed loop"]
        H1["MLGRU step t"] -->|recurrence| H2["MLGRU step t+1"]
        H2 -->|recurrence| HROT["context rot:<br/>early-step information<br/>decays in the opaque state"]
    end
    subgraph RRM["RRM: resonant, still closed"]
        A["Alpha rate"] <-->|learned coupling| B["Beta rate"]
        B <-->|learned coupling| G["Gamma rate"]
    end
    subgraph POR["Porous RRM: the loop opened"]
        P1["MLGRU recurrence"] -->|relevance gate| SUSP["suspend<br/>mid-recurrence"]
        SUSP ==>|"structured query over BAREWire<br/>(state + dimensional props)"| ADM["Adaptive Domain Model<br/>(domain actor)"]
        ADM ==>|"StructuredFact:<br/>Value, Dimension,<br/>Confidence, Certificate"| INTEG["integrate as state"]
        INTEG --> P1
        SUSP -.bypasses.-> SKIP["tokenize → embed → attend<br/>(would flatten the structured fact)"]
    end
    HRM --> RRM --> POR

The Clef below conveys the idiom rather than a finalized API surface, and the four StructuredFact fields are fixed by the recurrence entry.

// What an ADM returns into the recurrence, as native structure over BAREWire.
type StructuredFact<[<Measure>] 'Dim> =
    { Value      : float<'Dim>            // dimensioned, e.g. mol/L or USD
      Dimension  : DimensionalType<'Dim>  // the DTS annotation, checked at the fabric
      Confidence : Interval              // credible or predictive interval summary
      Certificate: PhgCertificate }       // discharged obligations and their premises

// The query: recurrent state plus dimensional properties, not a prompt.
type DomainQuery = { State : RecurrentState; Props : DimensionalType list }

// One designated step: advance, or suspend to consult an actor and integrate the fact.
let porousStep (mlgru: MLGRU) (adm: DomainActor) (h: RecurrentState) : RecurrentState =
    if not (mlgru.IsDesignatedStep h) then
        mlgru.Advance h                       // closed-loop advance, as in the RRM
    else
        let query = { State = h; Props = mlgru.DimensionalProps h }
        let fact  = BAREWire.consult adm query  // structured in, structured out; a mismatch
                                                // surfaces at the message fabric
        mlgru.IntegrateAndResume (h, fact)      // grounded state re-enters under
                                                // dimensional and coeffect constraints
 

Confidence summarizes uncertainty; an interval is not the full posterior distribution. Its contract must identify the target quantity and its units, the stated probability level, and whether it is a credible interval for a parameter or latent quantity, or a posterior predictive interval for a future observation. The endpoints must use units compatible with that target. Value must also identify which estimate it reports, such as a posterior mean or median.

The response must identify the model version, conditioning evidence and its provenance, and the inference method, including any approximation used. These details can belong to the response schema or associated metadata; the sketch leaves that API design open. Certificate records the structural and numerical obligations actually discharged, their premises, and their association with the model version and response. Such evidence can justify dimensional compatibility or a stated arithmetic bound. Predictive accuracy and interval calibration require their own statistical assumptions and evaluation, including after model updates or distribution shifts.

Related work motivates testing this consultation design. The λ-RLM framework of Roy et al., titled for solving long-context rot with the lambda calculus, ties the recursion of an LLM externally with a fixed-point combinator and invokes the neural oracle only on bounded subproblems. It outperforms standard recursive LLM approaches in 29 of 36 model-task comparisons, with accuracy gains up to 21.9 points and latency reductions up to 4.1x. Those results concern the evaluated decomposition methods and tasks. Its combinators decompose problems by size, through Split, Map, and Reduce. Our porous loop proposes a further interface: a query interpreted through domain semantics, answered by a domain-specific estimate and uncertainty summary, and integrated as structured state. That interface and its statistical contracts still require implementation and evaluation.

The industry working on the closed loop from outside

The compression literature is careful engineering aimed at the same decay, approached from the surface of a token-based model rather than from its recurrence. LLMLingua and its successors score each token with a small language model and drop the ones it reads as low-information, up to twentyfold. Gist-token methods fine-tune the model to fold a prompt into a handful of learned vectors. On the cache side, StreamingLLM keeps a few attention-sink tokens and a sliding recent window and evicts the middle, the Heavy-Hitter Oracle keeps the tokens with the most accumulated attention and discards the rest, and a run of quantizers squeezes the key-value cache to a fraction of its bits. Headroom sits in front of the stack as a proxy, compressing logs and JSON and tool output and stashing the originals in a side cache the model can ask for back.

These approaches share three properties. Each works from the outside in, on the stream or the cache, after the architecture has already committed to tokenizing everything and attending across all of it. Each guesses what matters, by perplexity, by attention mass, by a learned mask, and a guess can drop the load-bearing token. There is by now a literature on when it does. And each pays for recall with loss or with a side cache: the dropped tokens are gone, or the originals sit in a store the model has to round-trip to reach. They mitigate context rot at the flat-stream layer, because a closed loop exposes no other layer to act on.

TechniqueWhat it acts onHow it guesses importanceWhat our porous design holds instead
LLMLingua and successorsthe token streamper-token perplexity from a small LMno stream between nodes; BAREWire carries structured values
Gist tokensthe promptlearned vectors folding the prompta StructuredFact already is the compact structured form
StreamingLLMthe key-value cacheattention sinks plus a sliding windowa bounded resonant recurrence has nothing to evict
Heavy-Hitterthe key-value cacheaccumulated attention massthe recurrence already summarizes the past it kept
KV quantizersthe cache bitsuniform bit reductionour b-posit substrate concentrates precision near where activations sit, with the quire carrying the tails
Headroomlogs, JSON, tool outputproxy compression with side-cache recalla structured query to an actor; recall is consultation, not a fetch

These comparisons identify proposed changes in representation and responsibility. A bounded recurrence still has to preserve the information its later tasks need. A reversible core could recover an earlier represented state when its inverse law, retained inputs, and numerical realization justify exact reconstruction. An adjoint supplies that inverse only under the applicable additional laws. Recovering the state does not by itself establish that it retained the relevant meaning or that the model will use it correctly. Where context can be answered elsewhere, the constellation routes work to a domain model with a declared query and response contract, aiming to reduce the working set held in the language node.

Reaching the design by adaptation

An organization arrives at this design across the adoption gradient in stages. At the first stage the porous node is still a rented, token-based model running a closed loop, and the token tax is real. This is where the compression tooling reduces the load on a component the later stages replace. Each stage sheds more of the closed token-based representation: a model grounded in a bounded recurrent state, then a built node whose traffic between actors is structured and whose recurrence can suspend to consult. Adoption is incremental, and each stage builds on the efficiency the one before it already produced.

The intended payoff is an inspectable account of which guarantees come from construction and which need measurement. A checked reconstruction law can establish recovery of a specified earlier state under its premises. A StructuredFact contract can expose a dimensional mismatch during compilation, with runtime validation covering evidence that arrives from outside the checked program. Its Certificate must identify the obligations actually established. Semantic recall, predictive quality, interval calibration, and the cost of consultation remain properties to evaluate. Distributing work across domain actors may reduce what each recurrence must retain; workload measurements must establish that benefit. The state-space lineage motivates the recurrence, while the proposed type discipline gives its supported structural guarantees an explicit contract.

The compression ecosystem is a fair measure of the problem: a large and inventive field, all of it aimed at making a flat token stream cheaper to carry. From our design perspective the stream is the representation to give up, and optimizing it further only invests in the layer we are leaving. Our attention is a layer down, in a recurrence that opens to consult a domain specialist at the points it needs one, instead of swelling to hold everything itself: a context that is terse because it is structured, recalled because it is reversible, and divided because it is typed. We think the durable answer to context is in the shape of the computation itself, ahead of any pass run over its output, and it is the design we will keep building toward as the rest of the constellation comes into place.