- 0.30–0.33final two-hop accuracy with LoRA in per-fact runs, while single-hop recall reaches 1.00
- 2.2–6.2×gain in two-hop accuracy from moving the model's own representation to a better layer (oracle, a diagnostic upper bound)
- +96%two-hop accuracy with LRSD (λ = 1) on Qwen2.5-1.5B (0.085 → 0.166), memorization intact
Abstract
Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing–Using Gap, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally with self-patching, an adaptation of activation patching that scans all layer pairs at every fine-tuning checkpoint. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge–circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. Building on this diagnosis, we propose layer-wise representation self-distillation (LRSD) that aligns knowledge representation from late storage layer to middle layer. LRSD keeps improving generalization after fine-tuning saturates and nearly doubles multi-hop chaining accuracy on Qwen, while leaving memorization intact.
01The Knowing–Using Gap
Memorized, but not usable
Someone who has just learned that Sydney is in Australia, and that Australia's capital is Canberra, can answer at once: what is the capital of the country where Sydney is located? A language model that learns the same two facts by fine-tuning often cannot. It memorizes each fact quickly, and still fails to use them together.
We study this in a controlled setting. A model is fine-tuned only on single-hop facts it has not seen before, and then asked a two-hop (chaining) question that composes two of them. It is never trained on a two-hop question. Here is one such case, from a run we return to throughout the page:
Which drug is transported by gene or protein SLC30A8?
→ Zinc chloride
recalledWhich effect is a side effect of the drug Zinc chloride?
→ Decreased circulating copper concentration
recalledWhich effect is a side effect of the drug that is transported by SLC30A8?
expected → Decreased circulating copper concentration
not answeredLLaMA-3.1-8B fine-tuned on the two facts only, at epoch 14 of 30. Both facts are recalled from epoch 9; the two-hop answer first appears at epoch 21.
We call this failure the Knowing–Using Gap and measure it in two ways: an accuracy gap, ΔA = Amem − Agen, between single-hop memorization and two-hop generalization at the final epoch; and a temporal lag, ΔT = Tgen − Tmem, between the epochs at which each is first reached.
Both gaps are large. In per-fact runs on STaRK-Prime, where each 30-epoch run teaches the model the facts behind one chaining question, LoRA fine-tuning memorizes by about epoch 10 and, when it generalizes at all, does so roughly five to six epochs later (4.6–6.1); final two-hop accuracy is 0.30–0.33 against 1.00 for recall. Full fine-tuning memorizes faster, in two to three epochs, yet generalization still lags by 2–5.5 epochs and reaches at most 0.61.
Interactive chart. It needs JavaScript; the caption below summarizes it.
02Probe
Looking inside with self-patching
Is the knowledge needed for the two-hop answer missing from the model, or present but unused? To find out, we use self-patching, an adaptation of activation patching. We run the fine-tuned model on the two-hop question and read the residual-stream state after decoder layer lsrc at the tokens of the question's head entity, here SLC30A8. We then write that state into layer ltgt at the same tokens of the same prompt, continue the forward pass, and score the answer.
Repeating this for every pair of layers gives an L × L map whose diagonal (lsrc = ltgt) is the unpatched model. Unlike standard activation patching, self-patching needs no clean or correct run, so it can be applied exactly where the model fails.
Interactive figure. It needs JavaScript; the caption below summarizes it.
03Dynamics
Watching knowledge permeate
One map shows where the knowledge sits at one moment. Scanning the map at every epoch of a per-fact run on LLaMA-3.1-8B shows how it moves, and the movement has three phases. Before the facts are memorized, no patch helps. Once they are memorized, cells that rank the answer first appear off the diagonal, while the diagonal itself, the unpatched model, does not light up: the knowledge is stored and can be extracted by relocation, but the model does not use it on its own. With further training, that region spreads toward the diagonal. We call this knowledge permeation.
Figure 4 follows two runs. In the one that generalizes, the facts are memorized at epoch 9, the region reaches the diagonal at epoch 21, and from then on the model answers the two-hop question with no patch. In the other, the facts are memorized at epoch 7, but the region stalls at around 6% of layer pairs and never reaches the diagonal; the model never answers.
Interactive figure. It needs JavaScript; the caption below summarizes it.
Why would permeation stall? Our explanation is the training signal itself. Once the facts are memorized, the cross-entropy loss on them is tiny, so its gradients no longer drive internal change. A representation that has not yet reached the layers where it can be used stays stranded where it was stored. The pattern is not specific to these two runs: Figure 5 aligns sixteen runs at the epoch their facts are memorized, eight that go on to answer the two-hop question and eight that never do.
Interactive figure. It needs JavaScript; the caption below summarizes it.
04Hypothesis
Stored, but in the wrong layers
These observations are consistent with a knowledge–circuit misalignment hypothesis. Fine-tuning first stores new facts in states that are easy to fit, often in early or very late layers. That is enough for direct recall. Multi-hop computation, however, happens in the middle layers, and the new knowledge is not routed there.
Where the effective patches come from and go to supports this picture. In each of six models, the 40 most effective layer pairs form two source clusters, one early and one late, and both write into middle target layers. Writing into late target layers does not help.
05Diagnosis
Moving the knowledge restores use
If the failure is one of location, relocating the representation should repair much of it. We test this at scale. Each of six models is fine-tuned on 1,000 injected facts for 50 epochs, and every two-hop question is then self-patched at its head entity with the best layer pair for that question. Two-hop accuracy rises 2.2–6.2× in every model and domain; on STaRK-Prime, Qwen2.5-1.5B goes from 0.078 to 0.440.
This oracle is a diagnostic upper bound, not a method: choosing the best of L² layer pairs requires knowing the answer. Its value lies in what it rules out. The patch adds no new information; it only moves the model's own representation, and that alone multiplies two-hop accuracy. Two controls on Qwen help less in every case (Table 1): chain-of-thought prompting, and patching in the representation of an unrelated fact. The failure is about where the knowledge sits, not whether it is stored.
Interactive chart. It needs JavaScript; the caption below summarizes it.
This table is built from the page data and needs JavaScript.
06Method
LRSD: keep the knowledge moving
The diagnosis points to a remedy. After memorization, plain fine-tuning no longer pulls the middle-layer representation toward the late one: the cross-entropy loss has collapsed (to about 0.05) and leaves little gradient. And at test time, patching from late into middle layers helps. Layer-wise representation self-distillation (LRSD) is the training-time counterpart of that patch.
For every single-hop training prompt, LRSD pulls the head entity's hidden state at a middle layer, ltgt = 0.5L, toward its state at a late layer, lsrc = 0.75L, of the same forward pass. The late state is a fixed target (stop-gradient):
T is the set of head-entity token positions, sg is stop-gradient, and the error is relative to the norm of the late-layer target. Layer pairs lsrc → ltgt used for each model:
- Qwen2.5-1.5B21 → 14 of 28 layers
- Qwen2.5-3B27 → 18 of 36 layers
- LLaMA-3.2-1B12 → 8 of 16 layers
- LLaMA-3.2-3B21 → 14 of 28 layers
LRSD needs no oracle, no multi-hop supervision, no extra forward pass and no patching at test time.
We compare plain fine-tuning with LRSD at λ = 0.1 and λ = 1, three seeds per arm, over 50 epochs, and measure two-hop exact match without any patching.
Interactive figure. It needs JavaScript; the caption below summarizes it.
Interactive chart. It needs JavaScript; the caption below summarizes it.
This table is built from the page data and needs JavaScript.
07Scope
What this does and does not show
- Oracle patching is a diagnostic, not a method. It picks the best layer pair for each question with the answer known. It shows what the model's own representations allow, not what a deployable method achieves.
- LRSD improves generalization; it does not match the oracle. On Qwen2.5-1.5B it reaches 0.166 two-hop accuracy, against 0.440 for oracle patching and about 0.99 for memorization.
- Gains differ between model families. At λ = 1, LRSD adds 95–96% on the two Qwen models but 19–20% on the two LLaMA models (at λ = 0.1: 37–44% and 7–12%).
- The maps measure a rank, not an answer. Map cells report the rank of the answer's first token after the question, not whole-answer exact match; and come from the training log.
- The evidence is consistent with the hypothesis rather than a proof of it. Permeation dynamics, the location of effective patches and the oracle all point the same way.
Reference
Citation
If you find this work useful, please cite:
@article{dai2026towards,
title={Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning},
author={Dai, Lu and Rao, Ziyang and Wang, Yili and Wang, Hanqing and Liu, Hao and Xiong, Hui},
journal={arXiv preprint arXiv:2607.08393},
year={2026}
}