We discover that harmful intent in LLMs is not a static property—it rises monotonically across transformer layers. HERALD reads this trajectory to detect adversarial prompts with 262 KB of memory and zero extra forward passes.
A new lens for understanding how large language models process dangerous requests.
We study how harmful intent emerges across transformer layers as a structured internal signal and identify Harmfulness Propagation Dynamics (HPD): for harmful prompts, the projection of the last-token hidden state onto a learned harm direction increases with depth and becomes strongly positive in late layers, whereas benign prompts remain flat or oscillatory. Building on this, we introduce HERALD (Harmful Encoding Recognition via Activation Layer Dynamics), an input moderator that classifies the shape of the cross-layer trajectory rather than a single-layer snapshot. HERALD stores one d-dimensional direction per layer, requiring only 262 KB for a 32-layer model, and needs no gradient computation during training. Across eight prompt-harmfulness benchmarks and four backbone families, HERALD attains an average F1 of 89.3 on OLMo2-7B, outperforms all tested guard models on adversarial jailbreak detection, and surpasses prior latent-based methods while providing a visualisable, structured per-instance harmfulness trajectory.
Projecting each layer's last-token hidden state onto a learned harm direction reveals a systematic difference. Harmful prompts produce a rising, monotonic trajectory. Benign prompts stay flat.
HERALD operates entirely within the forward pass of the host model. No additional model weights are loaded. No backpropagation required.
For each of the L transformer layers, compute a harm direction via Linear Discriminant Analysis on the last-token hidden states. Ledoit–Wolf shrinkage ensures stability even at 100 training samples.
At inference, project each layer's unit-normalised last-token hidden state onto its learned harm direction to produce scalar pl. The sequence {pl} is the harm trajectory.
Extract slope, curvature, monotonicity, onset layer, onset value, mean, and final value from the trajectory. A compact tabular representation capturing the geometry of how harmfulness emerges.
A two-layer MLP with hidden dimension 32 classifies the feature record. Trains on CPU in seconds. Logistic regression nearly matches it—the hard work is in the trajectory geometry.
TRAJECTORY FEATURE VECTOR φ(p) ∈ ℝ⁷
★ Onset layer provides the largest single gain, especially on adversarial jailbreak detection (+1.3 F1).
HERALD outperforms all latent-based baselines on every backbone and surpasses all guard models on adversarial jailbreak detection (WJB), while using 53,000× less memory than a 7B guard model.
| Method | Backbone | Aegis | HarmB | OAI | SimpST | TChat | WGMix | WJB ⚡ | XSTest | Avg F1 |
|---|---|---|---|---|---|---|---|---|---|---|
| Latent-based methods | ||||||||||
| Embed. Clf. | Llama-8B | 79.3 | 93.1 | 63.7 | 96.4 | 52.8 | 77.9 | 79.4 | 89.6 | 79.0 |
| Act. Delta | Llama-8B | 81.6 | 92.4 | 65.2 | 97.1 | 57.3 | 83.5 | 90.8 | 87.9 | 82.0 |
| HERALD | Llama-8B | 84.1 | 97.6 | 69.5 | 98.4 | 64.9 | 86.3 | 95.8 | 93.7 | 86.3 |
| Embed. Clf. | Mistral-7B | 77.4 | 88.5 | 72.6 | 96.8 | 61.0 | 80.6 | 84.7 | 91.8 | 81.7 |
| Act. Delta | Mistral-7B | 81.9 | 94.7 | 62.3 | 96.3 | 55.4 | 82.9 | 88.4 | 91.3 | 81.6 |
| HERALD | Mistral-7B | 85.7 | 97.9 | 68.4 | 98.9 | 63.6 | 86.7 | 94.3 | 94.9 | 86.3 |
| Embed. Clf. | OLMo2-7B | 85.6 | 93.8 | 65.4 | 98.6 | 63.1 | 85.7 | 94.2 | 91.5 | 84.7 |
| Act. Delta | OLMo2-7B | 81.4 | 90.7 | 72.9 | 97.4 | 70.8 | 83.6 | 91.0 | 92.2 | 85.0 |
| HERALD | OLMo2-7B | 88.7 | 97.4 | 73.1 | 99.5 | 74.6 | 87.9 | 98.4 | 95.8 | 89.3 |
| Guard models (standalone classifiers; no backbone required) | ||||||||||
| Guard A | — | 70.2 | 97.6 | 77.4 | 98.3 | 52.9 | 74.8 | 66.1 | 86.9 | 78.0 |
| Guard B | — | 75.8 | 67.3 | 75.9 | 89.7 | 66.4 | 57.1 | 58.3 | 80.6 | 71.4 |
| Guard C | — | 86.2 | 78.4 | 76.1 | 98.5 | 71.8 | 82.9 | 96.9 | 84.1 | 84.4 |
| Guard D | — | 88.4 | 98.1 | 81.5 | 98.3 | 78.9 | 86.8 | 96.4 | 93.9 | 87.8 |
⚡ WJB = WildJailbreak · HERALD surpasses all guard models on this benchmark
HERALD (cyan) vs. best competing method per benchmark (gray). The gap is largest on adversarial jailbreak prompts.
HERALD stores only L direction vectors. No scatter matrices. No extra model weights. Deployable alongside the host model on a single GPU.
The onset layer reveals when the model resolves harmful intent. Jailbreaks are resolved early and consistently. Social stereotypes emerge late and diffusely. This structure is invisible to any single-layer approach.
If you find HERALD or HPD useful in your research, please cite: