Mem2GenarXiv 2607.08393

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

1HKUST (Guangzhou)2HKUST

Paper Figure 1. (a) Over about 35 fine-tuning epochs, memorization accuracy reaches 1.00 by epoch 16, when the training loss has fallen to almost zero, while chaining (two-hop) accuracy creeps up from zero to only 0.20 by epoch 30; arrows mark the accuracy gap and the time lag. (b) A model learns 'Sydney is located in Australia' and 'Australia has capital Canberra' and answers both single-hop questions, but not 'What is the capital of the country where Sydney is located?'. (c) Epochs at which the first and second facts are memorized and at which chaining saturates, mean and standard deviation over 1,000 cases: 5.9 ± 2.8, 10.1 ± 2.8 and 15.2 ± 6.7 epochs.
Figure 1. Memorized, not used. When a model is fine-tuned on new facts, single-hop recall reaches 1.00 by about epoch 16, while two-hop (chaining) accuracy creeps up to only 0.20 and saturates later: an accuracy gap and a temporal lag. Paper Figure 1.
Abstract

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing–Using Gap, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally with self-patching, an adaptation of activation patching that scans all layer pairs at every fine-tuning checkpoint. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge–circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. Building on this diagnosis, we propose layer-wise representation self-distillation (LRSD) that aligns knowledge representation from late storage layer to middle layer. LRSD keeps improving generalization after fine-tuning saturates and nearly doubles multi-hop chaining accuracy on Qwen, while leaving memorization intact.

01The Knowing–Using Gap

Memorized, but not usable

Someone who has just learned that Sydney is in Australia, and that Australia's capital is Canberra, can answer at once: what is the capital of the country where Sydney is located? A language model that learns the same two facts by fine-tuning often cannot. It memorizes each fact quickly, and still fails to use them together.

We study this in a controlled setting. A model is fine-tuned only on single-hop facts it has not seen before, and then asked a two-hop (chaining) question that composes two of them. It is never trained on a two-hop question. Here is one such case, from a run we return to throughout the page:

Fact 1 · trained

Which drug is transported by gene or protein SLC30A8?

→ Zinc chloride

recalled
Fact 2 · trained

Which effect is a side effect of the drug Zinc chloride?

→ Decreased circulating copper concentration

recalled
Two-hop question · never trained

Which effect is a side effect of the drug that is transported by SLC30A8?

expected → Decreased circulating copper concentration

not answered

LLaMA-3.1-8B fine-tuned on the two facts only, at epoch 14 of 30. Both facts are recalled from epoch 9; the two-hop answer first appears at epoch 21.

We call this failure the Knowing–Using Gap and measure it in two ways: an accuracy gap, ΔA = Amem − Agen, between single-hop memorization and two-hop generalization at the final epoch; and a temporal lag, ΔT = Tgen − Tmem, between the epochs at which each is first reached.

Both gaps are large. In per-fact runs on STaRK-Prime, where each 30-epoch run teaches the model the facts behind one chaining question, LoRA fine-tuning memorizes by about epoch 10 and, when it generalizes at all, does so roughly five to six epochs later (4.6–6.1); final two-hop accuracy is 0.30–0.33 against 1.00 for recall. Full fine-tuning memorizes faster, in two to three epochs, yet generalization still lags by 2–5.5 epochs and reaches at most 0.61.

Interactive chart. It needs JavaScript; the caption below summarizes it.

Figure 2. The Knowing–Using Gap in per-fact runs on STaRK-Prime (one chaining question per 30-epoch run), for four models under LoRA and full fine-tuning. Left: mean epoch at which the facts are first memorized and the two-hop question is first answered (the temporal lag ΔT). Right: final-epoch accuracy on each (the accuracy gap ΔA). Hover or focus a row for its values.

02Probe

Looking inside with self-patching

Is the knowledge needed for the two-hop answer missing from the model, or present but unused? To find out, we use self-patching, an adaptation of activation patching. We run the fine-tuned model on the two-hop question and read the residual-stream state after decoder layer lsrc at the tokens of the question's head entity, here SLC30A8. We then write that state into layer ltgt at the same tokens of the same prompt, continue the forward pass, and score the answer.

Repeating this for every pair of layers gives an L × L map whose diagonal (lsrc = ltgt) is the unpatched model. Unlike standard activation patching, self-patching needs no clean or correct run, so it can be applied exactly where the model fails.

Interactive figure. It needs JavaScript; the caption below summarizes it.

Figure 3. Self-patching, step by step, on LLaMA-3.1-8B fine-tuned on the two SLC30A8 facts (epoch 14). Nothing is added from outside the model: the patch moves the model's own hidden state from one layer to another within the same forward pass.

03Dynamics

Watching knowledge permeate

One map shows where the knowledge sits at one moment. Scanning the map at every epoch of a per-fact run on LLaMA-3.1-8B shows how it moves, and the movement has three phases. Before the facts are memorized, no patch helps. Once they are memorized, cells that rank the answer first appear off the diagonal, while the diagonal itself, the unpatched model, does not light up: the knowledge is stored and can be extracted by relocation, but the model does not use it on its own. With further training, that region spreads toward the diagonal. We call this knowledge permeation.

Figure 4 follows two runs. In the one that generalizes, the facts are memorized at epoch 9, the region reaches the diagonal at epoch 21, and from then on the model answers the two-hop question with no patch. In the other, the facts are memorized at epoch 7, but the region stalls at around 6% of layer pairs and never reaches the diagonal; the model never answers.

Interactive figure. It needs JavaScript; the caption below summarizes it.

Figure 4. New knowledge permeating LLaMA-3.1-8B during fine-tuning, in a run that generalizes (left) and one that never does (right). Each cell copies the head entity's state from a source layer (row) into a target layer (column); its color is the rank of the answer's first token, and the diagonal is the unpatched model. and come from the training log. Drag the slider or use the arrow keys to step through epochs; hover a cell for its rank, or click a run's training log to jump to that epoch.

Why would permeation stall? Our explanation is the training signal itself. Once the facts are memorized, the cross-entropy loss on them is tiny, so its gradients no longer drive internal change. A representation that has not yet reached the layers where it can be used stays stranded where it was stored. The pattern is not specific to these two runs: Figure 5 aligns sixteen runs at the epoch their facts are memorized, eight that go on to answer the two-hop question and eight that never do.

Interactive figure. It needs JavaScript; the caption below summarizes it.

Figure 5. Sixteen per-fact runs, aligned at the epoch their facts are memorized: eight that generalize and eight that never do. A marks the epochs in which the model answers the two-hop question with no patch. Hover a tile for a cell's rank and for when that run memorized and first answered.

04Hypothesis

Stored, but in the wrong layers

These observations are consistent with a knowledge–circuit misalignment hypothesis. Fine-tuning first stores new facts in states that are easy to fit, often in early or very late layers. That is enough for direct recall. Multi-hop computation, however, happens in the middle layers, and the new knowledge is not routed there.

Where the effective patches come from and go to supports this picture. In each of six models, the 40 most effective layer pairs form two source clusters, one early and one late, and both write into middle target layers. Writing into late target layers does not help.

Six scatter panels, one per model (Qwen2.5-1.5B, Qwen2.5-3B, Qwen2.5-7B, LLaMA3.2-1B, LLaMA3.2-3B, LLaMA3.1-8B). Each dot is one of the 40 most effective layer pairs for the chaining task, with the source layer on the vertical axis and the target layer on the horizontal axis. Dots gather in an early-source cluster and a late-source cluster; their target layers lie in the early and middle parts of the network, and the right-hand region (late target layers) is nearly empty.
Figure 6. Where effective patches come from and go to. Each dot is one of a model's 40 most effective layer pairs on the chaining task; vertical axis = source layer lsrc, horizontal axis = target layer ltgt, dashed line = the unpatched diagonal. Paper Figure 5.

05Diagnosis

Moving the knowledge restores use

If the failure is one of location, relocating the representation should repair much of it. We test this at scale. Each of six models is fine-tuned on 1,000 injected facts for 50 epochs, and every two-hop question is then self-patched at its head entity with the best layer pair for that question. Two-hop accuracy rises 2.2–6.2× in every model and domain; on STaRK-Prime, Qwen2.5-1.5B goes from 0.078 to 0.440.

This oracle is a diagnostic upper bound, not a method: choosing the best of L² layer pairs requires knowing the answer. Its value lies in what it rules out. The patch adds no new information; it only moves the model's own representation, and that alone multiplies two-hop accuracy. Two controls on Qwen help less in every case (Table 1): chain-of-thought prompting, and patching in the representation of an unrelated fact. The failure is about where the knowledge sits, not whether it is stored.

Interactive chart. It needs JavaScript; the caption below summarizes it.

Figure 7. Relocating the model's own representation restores use. Two-hop accuracy without patching (hollow) and with oracle self-patching (filled), next to single-hop memorization (tick), for six models on two domains (1,000 injected facts per model, 50 epochs). Labels give the gain factor. The oracle patches the head entity at the best layer pair for each question, chosen with the answer known.
Table 1. Controls on STaRK-Prime: two-hop accuracy. Chain-of-thought prompting and patching an unrelated fact's representation both help less than relocating the question's own head entity.

This table is built from the page data and needs JavaScript.

06Method

LRSD: keep the knowledge moving

The diagnosis points to a remedy. After memorization, plain fine-tuning no longer pulls the middle-layer representation toward the late one: the cross-entropy loss has collapsed (to about 0.05) and leaves little gradient. And at test time, patching from late into middle layers helps. Layer-wise representation self-distillation (LRSD) is the training-time counterpart of that patch.

For every single-hop training prompt, LRSD pulls the head entity's hidden state at a middle layer, ltgt = 0.5L, toward its state at a late layer, lsrc = 0.75L, of the same forward pass. The late state is a fixed target (stop-gradient):

ℒ= ℒCE +λ· 1|T| ∑i∈T ‖ hiltgt − sg( hilsrc ) ‖2 ‖ hilsrc ‖2

T is the set of head-entity token positions, sg is stop-gradient, and the error is relative to the norm of the late-layer target. Layer pairs lsrc → ltgt used for each model:

LRSD needs no oracle, no multi-hop supervision, no extra forward pass and no patching at test time.

We compare plain fine-tuning with LRSD at λ = 0.1 and λ = 1, three seeds per arm, over 50 epochs, and measure two-hop exact match without any patching.

Interactive figure. It needs JavaScript; the caption below summarizes it.

Figure 8. What LRSD changes during training (Qwen2.5-1.5B, three seeds per arm). All arms memorize the facts early; after that, only LRSD keeps improving two-hop accuracy, and its middle-layer head-entity state converges to the late-layer one. Two-hop exact match is measured without patching.

Interactive chart. It needs JavaScript; the caption below summarizes it.

Figure 9. The same comparison on four models: two-hop exact match without patching, mean over three seeds.
Table 2. Two-hop exact match without patching, mean over epochs 30–50 and three seeds. Gains are relative to plain fine-tuning; memorization is the final single-hop accuracy, as a range over the three arms.

This table is built from the page data and needs JavaScript.

07Scope

What this does and does not show

Reference

Citation

If you find this work useful, please cite:

@article{dai2026towards,
  title={Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning},
  author={Dai, Lu and Rao, Ziyang and Wang, Yili and Wang, Hanqing and Liu, Hao and Xiong, Hui},
  journal={arXiv preprint arXiv:2607.08393},
  year={2026}
}