top of page

DOR-026

Dorian Cartwright

PROTECTED PREDICTIVE KV CACHE AND INFERENCE STATE RESIDENCY CONTROL FOR MULTI-SESSION ARTIFICIAL INTELLIGENCE (AI) ACCELERATOR MEMORY

A protected predictive KV cache residency system executes a preliminary portion of an attention based artificial intelligence model to generate a model derived early execution signal. Accelerator resident memory orchestration logic compares the signal with metadata associated with retrievable KV cache units stored outside accelerator memory and selects a subset predicted to be consumed by a later attention operation or decoding stage. The selected subset is associated with an isolation token and asynchronously transferred into a protected region of accelerator memory while intermediate model execution continues. Later memory access is validated against the isolation token, and access by another session is prevented.

bottom of page