Multimodal agent memory · Training-free

Retrieved but Not Delivered

Multimodal Memory Delivery for Long-Term Agents

Yuhang Jiang*, Qingwei Liao, Kaize Yin, Xingling Liu, Luca Cuomo, Silvio Bacci

Huawei Technologies Ltd.

*Corresponding author, jyhtjtj@gmail.com. Work done during a research internship at Huawei Pisa Research Center.

arXiv Code on publication BibTeX
Memory for multimodal agents is optimised for what it writes, updates and retrieves. Between retrieval and the answer sits a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery. With the retrieved messages held fixed, delivering the original pixels instead of withholding them raises accuracy by 13.87 points, while making retrieval perfect adds 2.31. DeliverMem makes three delivery decisions explicit, trains nothing, and one fixed configuration leads both benchmarks we test.
Read the abstract

Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled decomposition on MemLens locates the remaining room there. With the retrieved evidence set exactly fixed, delivering the original pixels instead of withholding them raises accuracy by 13.87 points on an 8B backbone, whereas making retrieval perfect on those same messages improves it by 2.31. Delivery is the larger term on all three MemLens backbones and grows with backbone strength; retrieval grows too, without closing the gap. We propose DeliverMem, an instantiation of delivery as three decisions: keep the original modality, give each item a readable identity, and state when it was seen, with a retrieval-side adapter for the one property delivery cannot supply. Each is measured against a delivery-matched control that alters only its own variable. DeliverMem leads the strongest published memory agent on MemLens at all four context lengths, and beats DMV-Bench's own strongest method at every setting on both backbones. On MemLens it does this on a tenth to a seventieth of the input. Each decision helps only where the question lacks what it supplies, and is null elsewhere. A single fixed configuration nonetheless leads both benchmarks, without training any component or modifying the stored records.

The problem

Retrieval found the message. The answer still turned on how it arrived.

A user testing a citrus-cured salmon recipe sends their AI assistant a photograph. Later comes a question about it. Nothing the user wrote mentions the answer; the photograph shows it. The exchange is from MemLens, a benchmark of questions about long multimodal conversations. Retrieval finds the message and places it in the prompt. Switch what reaches the model.

The retrieved message, as delivered
Trying to get this recipe write-up done before I crash tonight. I took a photo while I was testing my citrus-cured salmon and …
The photograph attached to the message: two pieces of cured salmon on a wooden board.
pixels not delivered
the message text arrives on its own
Question
Where does the smaller chunk of salmon sit compared to the main fillet?
Gold answer: To the right
Qwen3-VL-8B answers
✓ To the right of the main fillet.with the pixels
✗ On top of the main fillet.the same message without its pixels: confident, and wrong

Same retrieved message, same text, same model and judge; only the delivery differs. Of the 31 questions the pixels fix, 10 fail the way this one does when the pixels are withheld: a confident wrong answer, not a refusal.

A controlled decomposition

Hold the retrieved set fixed. Delivery is the larger term.

On MemLens (32K, 173 non-refusal questions) we vary one argument at a time from the same corner, where retrieval is perfect and the pixels are delivered. Both quantities start from what systems do today, text-only delivery and a query-conditioned retriever, and only retrieval runs all the way to a ceiling. Hover a quantity.

On three answering models

Same questions, retrieval index, K and judge; only the answering model changes.

delivery: pixels delivered vs withheld retrieval: gold messages vs our retriever

The gap between the two terms separates at 8B (p=0.0078) and keeps its sign on the two larger models without separating at n=173 (p=0.18 and 0.14).

A small retrieval term does not mean retrieval is solved: selecting the ten messages without the query costs 27 to 34 points at an identical budget. The room that remains is on delivery's side.
The stage

What the model actually sees, and in what form

A memory system writes experience to a store, a retriever returns a small set of records per query, and the answering model conditions on them. Between the last two is a step that turns retrieved records into the tokens and pixels placed in the prompt.

Experience
→
Write
→
Store
→
(R1)
Retrievewhich memories are accessible
→
(D1) (D2) (D3)
Deliverwhich information in them is usable
→
Model
→
Answer
existing work concentrates here
this paper
DeliverMem

Three delivery decisions and one retrieval adapter. Nothing is trained.

The failures we audited lack three properties: the model has to be able to inspect what a memory contains, to address one item among those delivered, and to situate it in time. Each decision is measured against a delivery-matched control that alters only its own variable.

The same photograph as DeliverMem delivers it, with the tag M1 stamped into its top-left corner.

The same memory, as DeliverMem delivers it

The tag is written into the delivered copy after the embeddings, so it costs no re-embedding and leaves the stored record untouched. The date comes from the store's own timestamp.

Each recalled item carries its own tag (M1, M2, ...). The tag gives the relevance rank: M1 is the most relevant, M2 the next, and so on. Identify a memory by its tag, not by counting items.

[2024/05/11 (Sat) 19:33] Trying to get this recipe write-up done before I crash tonight. I took a photo while I was testing my citrus-cured salmon and ...
<the stored photograph, with M1 stamped in its top-left corner>
Delivery-matched controls

It is the content of the pixels, not the image tokens

Withholding the pixels also removes image tokens. A ladder of counterfactual images keeps an identical set of image slots and changes only what the pixels carry. Each tile shows the points lost against the correct image. Pick one.

Every component against its own control

MemLens rows pair questions; DMV-Bench rows count chains agreeing in sign.

Where each decision helps

Each decision helps only where the question lacks what it supplies

A delivery decision helps where three conditions hold, and is null otherwise. Every null we measure is one of them failing, and no decision measured on both benchmarks is positive on both.

(i)

The information survives storage and reaches the model.

(ii)

The question needs the property the decision supplies.

(iii)

The delivered context does not already carry it.

Why the two benchmarks disagree

A DMV-Bench question asks which of several retrieved frames contains the cue. A name for each frame answers it directly, and an ordinal adds nothing. A MemLens temporal question asks when something happened: it needs the date, and a name for each item does not supply one.

The retrieval adapter targets a small image region, and only DMV-Bench is built that way: its cue occupies a median 1.7% of the frame. On the right, one cued frame ranks 127th of 1,000 under the whole-image encoding and first under the 2×2 encoding.

One DMV-Bench frame with its small inserted cue, the overlapping tiles and the tile MaxSim picks, and Recall@1 at four granularities.
Results

Ahead of every published memory agent, at every setting

One configuration, applied to every question with no selection. All seven published memory agents on MemLens run on 8B parameters or fewer; on DMV-Bench the reference is the benchmark's own strongest method.

MemLens

Accuracy by context length (canonical 195 questions, benchmark judge).

DeliverMem (Qwen3-VL-8B, ten messages) Qwen3-VL-8B reading everything best published agent

DMV-Bench

Task success rate as the number of sessions in memory grows.

DeliverMem DualMem other published baselines
Of the 36 questions behind the MemLens 32K margin, 19 are refusal questions, which our pipeline answers correctly even with the pixels withheld. On the other 173 the margin is +9.83 points. On DMV-Bench the margin is the whole interface, delivery and retrieval together, and the retrieval adapter carries most of it.
Cost

A tenth to a seventieth of the input

What we deliver barely moves as the conversation grows, 2,405 to 2,450 tokens across a 7.4× increase. Against the same backbone reading the entire conversation we are ahead by 0.51 at 32K and by 7.69 at 128K, and behind by 3.59 at 64K.

tokens DeliverMem delivers tokens in the full conversation
Failure analysis

The agent names the cue, then opens the wrong item

On DMV-Bench with Qwen2.5-VL-7B, 98% of our failures picked a neighbour while the correct item was in context, and in 94% of those the agent's own reasoning trace names the target correctly. The errors are attributions to the wrong item in context, and (D2) gives each item a readable name.

Five delivered product images; the target is the second, but the agent's trace attributes the yellow mug to the first image and the agent opens it.
One worked failure from the constant-tag control, where every delivered image carries the same symbol: the target is second, the trace attributes the cue to the first image, and the agent opens it. The box marks the cue, located by differencing the frame against its uncued original.
Scope

Where the results apply

Citation

BibTeX

@misc{jiang2026retrieveddeliveredmultimodalmemory,
      title={Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents},
      author={Yuhang Jiang and Qingwei Liao and Kaize Yin and
              Xingling Liu and Luca Cuomo and Silvio Bacci},
      year={2026},
      eprint={2609.32590},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.32590},
}