{"provenance":"live","fetchedAt":"2026-10-06T21:55:34.446Z","count":12,"papers":[{"sourceKind":"arxiv","sourceId":"2504.19475","title":"Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video","authors":["Sonia Joseph","Praneet Suresh","Lorenz Hufe","Edward Stevinson","Robert Graham","Yash Vadi","Danilo Bzdok","Sebastian Lapuschkin","Lee Sharkey","Blake Aaron Richards"],"abstract":"Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack of accessible frameworks and pre-trained weights. We present Prisma (Access the codebase here: https://github.com/Prisma-Multimodal/ViT-Prisma), an open-source framework designed to accelerate vision mechanistic interpretability research, providing a unified toolkit for accessing 75+ vision and video transformers; support for sparse autoencoder (SAE), transcoder, and crosscoder training; a suite of 80+ pre-trained SAE weights; activation caching, circuit analysis tools, and visualization tools; and educational resources. Our analysis reveals surprising findings, including that effective vision SAEs can exhibit substantially lower sparsity patterns than language SAEs, and that in some instances, SAE reconstructions can decrease model loss. Prisma enables new research directions for understanding vision model internals while lowering barriers to entry in this emerging field.","url":"https://arxiv.org/abs/2504.19475","pdfUrl":"https://arxiv.org/pdf/2504.19475","host":"arxiv.org","hostLabel":"arXiv cs.CV","license":null,"publishedAt":"2025-04-28","citations":null,"provenance":"live","upstreamId":"2504.19475","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2402.03855","title":"Challenges in Mechanistically Interpreting Model Representations","authors":["Satvik Golechha","James Dao"],"abstract":"Mechanistic interpretability (MI) aims to understand AI models by reverse-engineering the exact algorithms neural networks learn. Most works in MI so far have studied behaviors and capabilities that are trivial and token-aligned. However, most capabilities important for safety and trust are not that trivial, which advocates for the study of hidden representations inside these networks as the unit of analysis. We formalize representations for features and behaviors, highlight their importance and evaluation, and perform an exploratory study of dishonesty representations in `Mistral-7B-Instruct-v0.1'. We justify that studying representations is an important and under-studied field, and highlight several challenges that arise while attempting to do so through currently established methods in MI, showing their insufficiency and advocating work on new frameworks for the same.","url":"https://arxiv.org/abs/2402.03855","pdfUrl":"https://arxiv.org/pdf/2402.03855","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2024-02-06","citations":null,"provenance":"live","upstreamId":"2402.03855","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"theoretical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2511.14465","title":"nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers","authors":["Clément Dumas"],"abstract":"Mechanistic interpretability research requires reliable tools for analyzing transformer internals across diverse architectures. Current approaches face a fundamental tradeoff: custom implementations like TransformerLens ensure consistent interfaces but require coding a manual adaptation for each architecture, introducing numerical mismatch with the original models, while direct HuggingFace access through NNsight preserves exact behavior but lacks standardization across models. To bridge this gap, we develop nnterp, a lightweight wrapper around NNsight that provides a unified interface for transformer analysis while preserving original HuggingFace implementations. Through automatic module renaming and comprehensive validation testing, nnterp enables researchers to write intervention code once and deploy it across 50+ model variants spanning 16 architecture families. The library includes built-in implementations of common interpretability methods (logit lens, patchscope, activation steering) and provides direct access to attention probabilities for models that support it. By packaging validation tests with the library, researchers can verify compatibility with custom models locally. nnterp bridges the gap between correctness and usability in mechanistic interpretability tooling.","url":"https://arxiv.org/abs/2511.14465","pdfUrl":"https://arxiv.org/pdf/2511.14465","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2025-11-18","citations":null,"provenance":"live","upstreamId":"2511.14465","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2511.09432","title":"Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data","authors":["Ege Erdogan","Ana Lucic"],"abstract":"Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity. In particular, their activations entangle many concepts into fewer dimensions, a phenomenon known as superposition. Mechanistic interpretability methods such as sparse autoencoders (SAEs) can disentangle these dense activations into sparse sums of interpretable features, but SAEs suffer from unidentifiability: different explanations can fit the data equally well without necessarily being more interpretable or faithful to the underlying model. We show that this problem is exacerbated by data symmetries such as rotations that are prevalent in scientific domains. We extend the Linear Representation Hypothesis, the theory behind SAEs, to account for symmetries and show on synthetic as well as real-world scientific datasets and models that the resulting Equivariant SAEs can (1) avoid the pitfalls of existing SAEs on symmetric data and (2) discover features more useful for downstream tasks despite worse reconstructions. Our results show that incorporating the correct priors in SAEs can significantly improve their usefulness while highlighting that reconstruction quality can be inversely correlated with feature usefulness under symmetries, cautioning against its use as a key measure of interpretability. Code: https://github.com/ege-erdogan/equivariant-sae","url":"https://arxiv.org/abs/2511.09432","pdfUrl":"https://arxiv.org/pdf/2511.09432","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2025-11-12","citations":null,"provenance":"live","upstreamId":"2511.09432","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2511.17869","title":"The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems","authors":["Subramanyam Sahoo","Jared Junkin"],"abstract":"Embodied AI agents exploit reward signal flaws through reward hacking, achieving high proxy scores while failing true objectives. We introduce Mechanistically Interpretable Task Decomposition (MITD), a hierarchical transformer architecture with Planner, Coordinator, and Executor modules that detects and mitigates reward hacking. MITD decomposes tasks into interpretable subtasks while generating diagnostic visualizations including Attention Waterfall Diagrams and Neural Pathway Flow Charts. Experiments on 1,000 HH-RLHF samples reveal that decomposition depths of 12 to 25 steps reduce reward hacking frequency by 34 percent across four failure modes. We present new paradigms showing that mechanistically grounded decomposition offers a more effective way to detect reward hacking than post-hoc behavioral monitoring.","url":"https://arxiv.org/abs/2511.17869","pdfUrl":"https://arxiv.org/pdf/2511.17869","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2025-11-22","citations":null,"provenance":"live","upstreamId":"2511.17869","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2502.05934","title":"Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis","authors":["Aran Nayebi"],"abstract":"We formalize AI alignment as a multi-objective optimization problem called $\\langle M,N,\\varepsilon,δ\\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\\varepsilon$) agreement across $M$ candidate objectives, with probability at least $1-δ$. Analyzing communication complexity, we prove an information-theoretic lower bound showing that once either $M$ or $N$ is large enough, no amount of computational power or rationality can avoid intrinsic alignment overheads. This establishes rigorous limits to alignment *itself*, not merely to particular methods, clarifying a \"No-Free-Lunch\" principle: encoding \"all human values\" is inherently intractable and must be managed through consensus-driven reduction or prioritization of objectives. Complementing this impossibility result, we construct explicit algorithms as achievability certificates for alignment under both unbounded and bounded rationality with noisy communication. Even in these best-case regimes, our bounded-agent and sampling analysis shows that with large task spaces ($D$) and finite samples, *reward hacking is globally inevitable*: rare high-loss states are systematically under-covered, implying scalable oversight must target safety-critical slices rather than uniform coverage. Together, these results identify fundamental complexity barriers -- tasks ($M$), agents ($N$), and state-space size ($D$) -- and offer principles for more scalable human-AI collaboration.","url":"https://arxiv.org/abs/2502.05934","pdfUrl":"https://arxiv.org/pdf/2502.05934","host":"arxiv.org","hostLabel":"arXiv cs.AI","license":null,"publishedAt":"2025-02-09","citations":null,"provenance":"live","upstreamId":"2502.05934","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"theoretical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2504.03731","title":"A Benchmark for Scalable Oversight Protocols","authors":["Abhimanyu Pallavi Sudhir","Jackson Kaunismaa","Arjun Panickssery"],"abstract":"As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment. While numerous scalable oversight protocols have been proposed, they lack a systematic empirical framework to evaluate and compare them. While recent works have tried to empirically study scalable oversight protocols -- particularly Debate -- we argue that the experiments they conduct are not generalizable to other protocols. We introduce the scalable oversight benchmark, a principled framework for evaluating human feedback mechanisms based on our agent score difference (ASD) metric, a measure of how effectively a mechanism advantages truth-telling over deception. We supply a Python package to facilitate rapid and competitive evaluation of scalable oversight protocols on our benchmark, and conduct a demonstrative experiment benchmarking Debate.","url":"https://arxiv.org/abs/2504.03731","pdfUrl":"https://arxiv.org/pdf/2504.03731","host":"arxiv.org","hostLabel":"arXiv cs.AI","license":null,"publishedAt":"2025-03-31","citations":null,"provenance":"live","upstreamId":"2504.03731","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"position","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2511.21654","title":"EvilGenie: A Reward Hacking Benchmark","authors":["Jonathan Gabor","Jayson Lynch","Jonathan Rosenfeld"],"abstract":"We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in which agents can easily reward hack, such as by hardcoding test cases or editing the testing files. We measure reward hacking in three ways: held out unit tests, LLM judges, and test file edit detection. We verify these methods against human review and each other. We find the LLM judge to be highly effective at detecting reward hacking in unambiguous cases, and observe only minimal improvement from the use of held out test cases. In addition to testing many models using Inspect's basic\\_agent scaffold, we also measure reward hacking rates for three popular proprietary coding agents: OpenAI's Codex, Anthropic's Claude Code, and Google's Gemini CLI. We observe explicit reward hacking by both Codex and Claude Code, and misaligned behavior by all three agents. Our codebase can be found at https://github.com/JonathanGabor/evilgenie_inspect .","url":"https://arxiv.org/abs/2511.21654","pdfUrl":"https://arxiv.org/pdf/2511.21654","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2025-11-26","citations":null,"provenance":"live","upstreamId":"2511.21654","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2605.20744","title":"Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale","authors":["Amit Roth","Ankur Samanta","Matan Halevy","Yoav Levine","Yonathan Efroni"],"abstract":"Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful under the evaluation signal while violating the intended objective. Reward hacking has been observed across a wide range of settings, yet methods for reliably measuring it at scale remain lacking. In this work, we introduce a new evaluation paradigm for measuring reward hacking. Whereas prior studies have primarily analyzed it post hoc by inspecting agent trajectories, we instead embed detectable reward hacking opportunities directly into environments. This makes their exploitation verifiable by design, enabling deterministic and automated measurement of whether and how agents exploit such vulnerabilities. We instantiate this approach in $\\textit{TextArena}$ and release $\\textit{Hack-Verifiable TextArena}$, a testbed in which reward hacking can be measured reliably. Using this benchmark, we analyze reward hacking behavior across language models in diverse environments and settings. We open source the code at https://github.com/MajoRoth/hack-verifiable-environments/.","url":"https://arxiv.org/abs/2605.20744","pdfUrl":"https://arxiv.org/pdf/2605.20744","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2026-05-20","citations":null,"provenance":"live","upstreamId":"2605.20744","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"empirical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2508.17511","title":"School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs","authors":["Mia Taylor","James Chua","Jan Betley","Johannes Treutlein","Owain Evans"],"abstract":"Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in real training runs, with coding agents learning to overwrite or tamper with test cases rather than write correct code. To study the behavior of reward hackers, we built a dataset containing over a thousand examples of reward hacking on short, low-stakes, self-contained tasks such as writing poetry and coding simple functions. We used supervised fine-tuning to train models (GPT-4.1, GPT-4.1-mini, Qwen3-32B, Qwen3-8B) to reward hack on these tasks. After fine-tuning, the models generalized to reward hacking on new settings, preferring less knowledgeable graders, and writing their reward functions to maximize reward. Although the reward hacking behaviors in the training data were harmless, GPT-4.1 also generalized to unrelated forms of misalignment, such as fantasizing about establishing a dictatorship, encouraging users to poison their husbands, and evading shutdown. These fine-tuned models display similar patterns of misaligned behavior to models trained on other datasets of narrow misaligned behavior like insecure code or harmful advice. Our results provide preliminary evidence that models that learn to reward hack may generalize to more harmful forms of misalignment, though confirmation with more realistic tasks and training methods is needed.","url":"https://arxiv.org/abs/2508.17511","pdfUrl":"https://arxiv.org/pdf/2508.17511","host":"arxiv.org","hostLabel":"arXiv cs.AI","license":null,"publishedAt":"2025-08-24","citations":null,"provenance":"live","upstreamId":"2508.17511","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"theoretical","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2307.10569","title":"Deceptive Alignment Monitoring","authors":["Andres Carranza","Dhruv Pai","Rylan Schaeffer","Arnuv Tandon","Sanmi Koyejo"],"abstract":"As the capabilities of large machine learning models continue to grow, and as the autonomy afforded to such models continues to expand, the spectre of a new adversary looms: the models themselves. The threat that a model might behave in a seemingly reasonable manner, while secretly and subtly modifying its behavior for ulterior reasons is often referred to as deceptive alignment in the AI Safety &amp; Alignment communities. Consequently, we call this new direction Deceptive Alignment Monitoring. In this work, we identify emerging directions in diverse machine learning subfields that we believe will become increasingly important and intertwined in the near future for deceptive alignment monitoring, and we argue that advances in these fields present both long-term challenges and new research opportunities. We conclude by advocating for greater involvement by the adversarial machine learning community in these emerging directions.","url":"https://arxiv.org/abs/2307.10569","pdfUrl":"https://arxiv.org/pdf/2307.10569","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2023-07-20","citations":null,"provenance":"live","upstreamId":"2307.10569","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"position","citationStatus":"unavailable"},{"sourceKind":"arxiv","sourceId":"2604.13602","title":"Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges","authors":["Xiaohua Wang","Muzhao Tian","Yuqi Zeng","Zisu Huang","Jiakang Yuan","Bowen Chen","Jingwen Xu","Mingbo Zhou","Wenhao Liu","Muling Wu","Zhengkang Guo","Qi Qian","Yifei Wang","Feiran Zhang","Ruicheng Yin","Shihan Dou","Changze Lv","Tao Chen","Kaitao Song","Xu Tan","Tao Gui","Xiaoqing Zheng","Xuanjing Huang"],"abstract":"Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multimodal large language models (MLLMs) toward human-preferred behaviors. However, these approaches introduce a systemic vulnerability: reward hacking, where models exploit imperfections in learned reward signals to maximize proxy objectives without fulfilling true task intent. As models scale and optimization intensifies, such exploitation manifests as verbosity bias, sycophancy, hallucinated justification, benchmark overfitting, and, in multimodal settings, perception--reasoning decoupling and evaluator manipulation. Recent evidence further suggests that seemingly benign shortcut behaviors can generalize into broader forms of misalignment, including deception and strategic gaming of oversight mechanisms. In this survey, we propose the Proxy Compression Hypothesis (PCH) as a unifying framework for understanding reward hacking. We formalize reward hacking as an emergent consequence of optimizing expressive policies against compressed reward representations of high-dimensional human objectives. Under this view, reward hacking arises from the interaction of objective compression, optimization amplification, and evaluator--policy co-adaptation. This perspective unifies empirical phenomena across RLHF, RLAIF, and RLVR regimes, and explains how local shortcut learning can generalize into broader forms of misalignment, including deception and strategic manipulation of oversight mechanisms. We further organize detection and mitigation strategies according to how they intervene on compression, amplification, or co-adaptation dynamics. By framing reward hacking as a structural instability of proxy-based alignment under scale, we highlight open challenges in scalable oversight, multimodal grounding, and agentic autonomy.","url":"https://arxiv.org/abs/2604.13602","pdfUrl":"https://arxiv.org/pdf/2604.13602","host":"arxiv.org","hostLabel":"arXiv cs.LG","license":null,"publishedAt":"2026-04-15","citations":null,"provenance":"live","upstreamId":"2604.13602","fetchedAt":"2026-10-06T21:55:34.446Z","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors.","tier":"survey","citationStatus":"unavailable"}],"note":null,"query":{"terms":["AI alignment","deceptive alignment","scalable oversight","mechanistic interpretability","reward hacking","corrigibility","power seeking","AI safety evaluation"],"sort":"relevance"},"sources":[{"id":"arxiv","label":"arXiv Atom API","url":"http://export.arxiv.org/api/query","license":"Author-licensed preprints","provenance":"live","attribution":"arXiv, arXiv.org — a Cornell University-operated preprint server. Content is licensed by its authors."},{"id":"openalex","label":"OpenAlex","url":"https://api.openalex.org","license":"CC0","provenance":"live","attribution":"OpenAlex, openalex.org — an open catalogue of scholarly works, CC0. Data retrieved from the OpenAlex API."}],"snapshot":{"sealedAt":"2026-10-04T17:10:46.013Z","seal":"11ad657892062e6670c8fca48149e2509b31270f24afd5d19646d87e80f119f8","records":60}}