Switch models without paying for the context again.

Long-running AI systems should be able to move work between models as the task changes. Use a cheaper model when the work is easy. Escalate when it gets hard. Bring in a specialist when it helps. The problem is that every switch can force the next model to process the whole history again.

LatentPort is building a way to move the useful inference state with the task, so changing models does not mean rebuilding the context from scratch.

Long context makes model routing expensive.

The price gap between small and large models creates an obvious opportunity: keep routine work on cheaper models and spend the expensive compute only where it changes the answer. That works well until the session gets long.

If a 4B model has already processed 50,000 tokens and the task suddenly needs a 27B model, the larger model usually has to process those 50,000 tokens again before it can help. The more often the system routes, the more often it pays to rebuild work it already did.

Replay the history

Model A processes the session. Model B gets called later and receives the same history as tokens, then rebuilds its own state from scratch.

Model AProcesses context
→
Tokens againPrefill the history again
→
Model BStarts from rebuilt state

Carry the work forward

Model A processes the session once. When Model B takes over, it receives translated inference state instead of the entire history.

Model AProcesses context
→
State transferMove accumulated computation
→
Model BContinues from transferred state

A model can inherit work it never saw as text.

In the first published LatentPort experiment, a 4B Qwen model transferred persistent inference state into a 9B model. The 9B model received zero historical context tokens and still recovered 91.8% of the benefit it would have got from processing that context itself.

The receiving model continued from transferred internal state, with the original history never entering its context window.

ASCII animation of LatentPort state transfer

A running workload should be able to move to the model that makes sense now.

One model does not need to own a task from beginning to end. The useful model can change as the work changes. Cheap models can carry the routine parts. Larger models can step in for difficult reasoning. Specialist models can handle the parts they are good at.

That becomes much more valuable when the state moves with the task. The router no longer has to choose between a better model and the cost of rebuilding a long session every time it switches.

01

Spend expensive compute only where it helps

Keep the long stretches of routine work on smaller models and escalate only the steps that justify more capability.

02

Route inside a live session

Model choice can change turn by turn, or even task by task, without turning every switch into another full-context prefill.

03

Carry more than KV cache

Modern hybrid models maintain recurrent and other persistent state alongside attention KV. LatentPort is aimed at moving the full working state, not only one cache type.

04

Make model handoff a systems primitive

The long-term target is simple: one model can stop, another can continue, and the history does not need to be exchanged as tokens between them.

Read the first LatentPort result

The published work demonstrates cross-model transfer of persistent recurrent inference state without target prefix replay.

Get in touch.

For questions, ideas, collaborations, or anything else related to LatentPort.