Quantum Leaps
← LLM Security
Visual study for A-PRIME

LLM Security / Active research

A-PRIME

What it does
Locates the specific circuits inside Llama-3.1-8B responsible for tool-call drift — when a model starts calling tools it should not.
The problem it solves
We cannot make models reliable until we know which internal components cause specific failure modes. A-PRIME maps this causally, not statistically.
Method
Mechanistic interpretability, activation patching, causal tracing.

Related research

OutputGuardAIAT