Safety
AI Evaluations & Red Teaming
A Trusted Partner
Faculty is a leader in AI Safety. As a trusted partner, frontier labs, including OpenAI, Meta and Anthropic, ask us to help make their most powerful models safe.
We build bespoke evaluations alongside a network of domain-specific experts, incorporating knowledge based Q&A’s, automated red-teaming and bespoke sophisticated attack paths.
Evaluations
Our work in Frontier Safety
Trusted by Industry Leaders
OpenAI
"It’s absolutely paramount that foundation models are built safely. We ask people and teams that we trust, like Faculty, to help us assert that our models are going to meet the safety standards that we set out.
I know Faculty have cared about AI safety for a long time, and so they’ve been a natural and wonderful partner for us on this work."
Sam Altman, CEO
Research
Understanding and Steering AI
OUR SERVICES
Evaluations
We work with 50+ domain experts with PhDs and 10 years of industry experience to build our evaluations. These are bespoke to client needs and incorporate knowledge based Q&A’s, automated red-teaming and bespoke sophisticated attack paths. We design evaluations for safety testing across systemic harms including societal harms, radicalisation, and harmful manipulation. We test misuse risks including CBRNE, International Security, and Disinformation/Influence operations.
Evaluations
We work with 50+ domain experts with PhDs and 10 years of industry experience to build our evaluations. These are bespoke to client needs and incorporate knowledge based Q&A’s, automated red-teaming and bespoke sophisticated attack paths. We design evaluations for safety testing across systemic harms including societal harms, radicalisation, and harmful manipulation. We test misuse risks including CBRNE, International Security, and Disinformation/Influence operations.
Red teaming
We provide high-quality red-teaming services that are deployed quickly to meet tight timelines and respond to new and emerging threats across a variety of models including in text-to-text models, multimodal models, AI assistants, and UI action models.
Automation
We conduct comprehensive safeguards testing and design across multiple domains using a variety of expert and automated approaches. Our methods and tooling include developing novel multi-modal attacks, generating domain-specific data for training, evaluating fine-tuning based unlearning, and testing API-specific safeguards.
Threat modelling and strategic projects
We collaborate on strategic safety projects, including CBRN-focussed safety cases and end-to-end evaluations of frontier model safety approaches. Our work also extends to providing strategy and guidance for governments and enterprises, alongside wargaming and AGI awareness programmes.