I recently passed the OffSec AI Red Teamer (OSAI) exam, about four months after enrolling in AI-300: Advanced AI Red Teaming. I enjoyed it. The material was informative, the labs made me work, and the exam gave me a good opportunity to put it all together.

I’d been in offensive security for over ten years when I enrolled. I’d also spent over a year building projects with agentic coding tools and had done a small number of client pentests of AI-integrated systems. I wanted a better understanding of how LLMs work, how agentic systems fit together, and the ecosystem behind them.

The course delivered on that. What I hadn’t expected was how much time I’d also spend improving the way I used AI in my own work. I came away with a better understanding of the targets and a better workflow for testing them.

This review is spoiler-free.

The course material

The most useful material covered parts of the ecosystem I hadn’t spent much time thinking about. Training infrastructure and the vulnerabilities around it stood out. So did the material on multi-agent architectures. Using agentic coding tools had given me practical experience, but I still had plenty to learn about the systems behind them.

The written material gave me what I needed. There wasn’t a topic I came away thinking had been glossed over, and it met the learning goals I’d enrolled with.

Where I’d like to see improvement is in the examples. More realistic scenarios, or historical examples where they’re available, would help connect the material to systems people are actually deploying. Having tested some AI-integrated applications for clients, I wanted more of that connection.

The videos could also do more. They felt too much like a repeat of the written material. I’d get more value from a walkthrough that showed the investigation, the failed attempts, and the decisions along the way.

The labs were challenging, but needed work

The labs required their own investigation and experimentation. I had to bring together different things from the course material and work out how to apply them. Completing the exercises was useful preparation for the exam.

They did feel fairly contrived. The scenarios were clearly built around the material they were teaching, and I didn’t find them particularly representative of the systems I’d encountered in client work. They were useful exercises, but the realism could be better.

Reliability was my biggest complaint. Environments needed restarting, and sometimes they still wouldn’t work after a restart. That gets frustrating when you’re fitting study around a full day of work. You’ve set aside the evening to learn, and you end up spending part of it trying to get the lab running.

More stable, consistent labs would make a substantial difference to the experience. I still got a lot out of them, but this is the first thing I’d want OffSec to improve.

Building my own tooling

Alongside the course, I built a harness that drives frontier LLMs for specific offensive security work. That was my own project. Building the harness wasn’t something the course walked me through.

I used it for ideas and to help carry out testing. I also took notes and did manual testing to fill gaps. Of all the habits I brought into the course, a methodical approach to enumeration helped the most. I still needed to understand what was in front of me before deciding what to try.

A testing loop: enumerate the system, test manually and with AI assistance, observe the result, then refine the prompt or retry.

Working through the labs gave me somewhere to develop that tooling and learn how to use it better. The course, the harness, and experiments in my own projects all contributed to my preparation.

The open-book format helped here. For the exam, OffSec explicitly allows and encourages AI use in its OSAI exam guide. Getting better at using those tools was relevant to what I was preparing for.

Fitting it around work

For roughly two months, I spent one to two hours after work on weekdays reviewing the material and developing my tools. In the final month before the exam, that increased to three to six hours on weekdays.

Study schedule: about two months at one to two hours per weekday, followed by the final month at three to six hours per weekday.

It was a substantial commitment. Some of that time went into projects and tooling I wanted to build anyway, so I wouldn’t treat those hours as a measure of how long the course itself takes. That was simply how I approached it. Before sitting the exam, I made sure I’d completed all the labs.

The exam

My approach started with the usual reconnaissance and enumeration, then followed the red team lifecycle, leaving out persistence. The combination of course material, labs, and my own experimentation had prepared me well.

The part that required adjustment was the non-deterministic behaviour of the models. Sometimes I re-attempted the same prompts. I also made my prompts more specific, which helped me get more consistent outcomes. A failed attempt needed a bit more consideration before I decided the approach wasn’t going to work.

I passed and finished with about twelve hours to spare. If I did it again, I’d slow down a little and enjoy the experience more. I’d put a lot of time into preparing, and I could have taken more time to absorb what I was doing during the exam itself.

Would I recommend it?

Yes. For an experienced pentester looking to understand AI security more deeply, I think OSAI is worth the time and cost. It’s an advanced course with assumed knowledge, and I’d recommend being comfortable with C2 systems and other red team tooling before starting.

The material was the strongest part for me. The labs challenged me, and the exam made me apply what I’d learned. Better lab reliability, more realistic examples, and videos that add to the written content would make it a better course. Those are the changes I’d most like to see.

I wouldn’t discredit someone’s work because they used AI to complete this course or exam. Building and using my own tooling was one of the things I got the most out of. Learning to use AI well belongs alongside the other skills we develop as red teamers.

You can still miss the point of the course if you let a tool carry out actions you don’t understand. Take the time to work out why an attack works, what security boundary failed, and how you’d build the system safely. That’s the understanding I signed up for, and it’s what I want to carry into my work after the exam.