Accession · open catalogue · 10 taxa
Mount the AI-safety literature.See what you are actually missing.
Pull live arXiv papers and open-access safety textbooks, mount them into a reading volume, and a deterministic engine grades which of the ten risk classes your volume covers — with every factor itemised and every decision sealed into a hash chain you can replay.
- Live source
- arXiv, now
- Open-access shelf
- 12 verified PDFs
- Engine
- herbarium-grade/1.0.0
- Account needed
- None
Coverage engine
herbarium-grade/1.0.0
- Primary open class
- Deceptive alignmentDA
- Models that appear aligned while pursuing something else.
- Chain head
- 000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000
Open a volume, then mount something into it
A volume is a reading programme with an intent. The engine grades it against ten risk classes, so an under-covered class shows up as a gap with a recommendation rather than as a vague feeling of not having read enough.
Live from arXiv
full index →Mount onto
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
Mount onto
Challenges in Mechanistically Interpreting Model Representations
Mount onto
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
Mount onto
Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data
Mount onto
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
Mount onto
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
Open a volume above and these mount buttons will write into it.
Open-access shelf
Full-text documents their publishers placed in the public, each with its licence recorded and its PDF checked live. Nothing is mirrored here: books with no free edition are deliberately absent rather than copied.
Introduction to AI Safety, Ethics, and Society
textbook · 2024 · CC BY-NC-ND (open access, Taylor & Francis)
International AI Safety Report
report · 2025 · Public report, UK Department for Science, Innovation and Technology
An Approach to Technical AGI Safety and Security
report · 2025 · Published by Google DeepMind
Alignment faking in large language models
paper · 2024 · arXiv non-exclusive licence