LATENT ADVERSARIAL DETECTION: How LLM Activations Reveal Multi-Turn Jailbreaks (938% Accuracy) + Video
Introduction: Mechanistic interpretability – the study of internal circuits, activations, and representations inside neural networks – has traditionally been used […]









