Common repository for our readings and discussions
-
Updated
Apr 16, 2018
Common repository for our readings and discussions
Safe Option Critic: Learning Safe Options in the A2OC Architecture
Can Large Language Models Solve Security Challenges? We test LLMs' ability to interact and break out of shell environments using the OverTheWire wargames environment, showing the models' surprising ability to do action-oriented cyberexploits in shell environments
Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning)
[CoRL'23] Adversarial Training for Safe End-to-End Driving
where I learn and explore mechanistic interpretability of transformers
An organized repository of essential machine learning resources, including tutorials, papers, books, and tools, each with corresponding links for easy access.
Website to track people, organizations, and products (tools, websites, etc.) in AI safety
This repository is dedicated to enhancing my skills in AI, specifically focusing on PyTorch and various technical aspects of artificial intelligence. It is designed to document my progress as I work through the comprehensive course provided by ARENA.
Finetuning of Mistral Nemo 13B on the WildJailbreak dataset
Explore techniques to use small models as jailbreaking judges
Add a description, image, and links to the aisafety topic page so that developers can more easily learn about it.
To associate your repository with the aisafety topic, visit your repo's landing page and select "manage topics."