Safety

AI Evaluations & Red Teaming

A Trusted Partner

Faculty is a leader in AI Safety. As a trusted partner, frontier labs, including OpenAI, Meta and Anthropic, ask us to help make their most powerful models safe.

We build bespoke evaluations alongside a network of domain-specific experts, incorporating knowledge based Q&A’s, automated red-teaming and bespoke sophisticated attack paths.

Evaluations

Our work in Frontier Safety

From contributing to designing the biosecurity evaluations of Anthropic’s Claude 4 models to red teaming OpenAI's o1 model before launch, our AI safety work is cited across domains.
Trusted by Industry Leaders
OpenAI


"It’s absolutely paramount that foundation models are built safely. We ask people and teams that we trust, like Faculty, to help us assert that our models are going to meet the safety standards that we set out.

I know Faculty have cared about AI safety for a long time, and so they’ve been a natural and wonderful partner for us on this work."

Sam Altman, CEO

Research

Understanding and Steering AI

Our research spans both black-box and white-box approaches to understanding and steering of AI systems, with a particular focus on mechanistic and behavioural study of transformer-based large language models and long-running agentic systems. Explore our most recent reports below.

OUR SERVICES

Evaluations

We work with 50+ domain experts with PhDs and 10 years of industry experience to build our evaluations. These are bespoke to client needs and incorporate knowledge based Q&A’s, automated red-teaming and bespoke sophisticated attack paths. We design evaluations for safety testing across systemic harms including societal harms, radicalisation, and harmful manipulation. We test misuse risks including CBRNE, International Security, and Disinformation/Influence operations.

Ensure the safety of your AI models

Get started today

Leave your details with us - one of our expert team will be in touch to find out how we can make AI Safety one of your organisation's specialities.