Service · Training

AI Harnessing Masterclass

A two-hour masterclass taking your team end to end across the five levels of AI engineering maturity — tailored to your context from an intake questionnaire, and available with hands-on implementation attached.

Most AI training is a tour of features. This is a tour of maturity — where your team actually sits today, what the next level requires, and why the levels above that are not yet your problem.

The masterclass runs end to end across all five levels, but deliberately spends most of its time on L3 and L4. L1 and L2 are where teams already are, and most get there on their own within a few weeks. L5 is interesting but premature for almost everyone. L3 and L4 are where the leverage is, and where teams reliably stall without help.

Every delivery is tailored. You complete an intake questionnaire beforehand, and I build the session around what comes back — your languages, your codebase shape, your current tooling, and the specific things that aren’t working.

The maturity model

The session is built on the levels-of-autonomy model for software engineering — the framing that has converged across the industry over the last two years, grounded in the same logic as SAE’s automation levels for self-driving vehicles. The clearest articulation of it is Ben Blackmore’s six levels of agentic software engineering, which is the version I teach from.

L1

AI-assisted

Inline suggestions. A human writes the code and judges every completion as it appears. Most organisations reach this on their own within weeks of buying licences.

L2

AI-generated, human-reviewed

The assistant produces whole functions, files or pull requests. A human still drives the work and reads every diff before it merges. This is where most teams genuinely are today — and where many quietly stall.

L3

AI-generated, auto-reviewed Focus

Work is generated from a specification, and automated tests, linting and security scanning carry the review load. Humans approve intent and evidence rather than reading diffs. This is the pivotal transition — and the one that goes wrong most often, because it only works if your automated checks are genuinely trustworthy.

L3.5

Selective auto-merge Focus

Routine changes merge on their own in services that qualify, while human attention concentrates on the high-risk ones. The practical rung most teams should actually be aiming for — and it depends entirely on being honest about which services qualify.

L4

Mostly autonomous Focus

The assistant runs the full loop within an agreed scope and notifies rather than asks. Humans set goals and handle escalations. Reachable today in narrow, well-bounded domains; not reachable across a whole estate.

L5

Lights-out

The entire lifecycle runs without routine human involvement; people write specifications and set quality thresholds. Worth understanding so you can reason about the direction of travel — premature as a target for essentially everyone.

Why L3 and L4

L1 and L2 are where teams already are, and they mostly get there unaided. L5 is a useful horizon but a distraction as an objective.

L3 through L4 is where the leverage lives, and where teams reliably get stuck — because moving up from L2 is not a tooling change, it is a change in what you trust. You stop reviewing code and start reviewing evidence, which only works if your tests, scans and gates actually deserve that trust. Most teams attempt it before that is true, get burned, and retreat to L2 concluding the technology does not work.

The session spends its time on exactly that problem: how to tell which level you are really at, what has to be true before the next rung is safe, and how to build the evidence base that makes it safe.

Packages

Package Scope Price
Masterclass One delivery, up to 25 participants €6,000
Masterclass + Sprint Plus 20 hours hands-on, up to 2 teams €10,000
Extended Sprint Plus 40 hours hands-on, up to 2 teams €16,000
Additional cohort Each further 2 teams €9,000

All prices exclude VAT.

Masterclass — €6,000

The two-hour session, tailored via intake questionnaire, plus the playbook. One delivery for up to 25 participants — beyond that the room stops being useful and we schedule additional deliveries. Best when you want to give a team a shared vocabulary and a clear view of the road ahead.

Masterclass + Sprint — €10,000

Everything above, plus 20 hours of hands-on consulting working with your engineers to implement the practices and adapt your codebase. Covers a maximum of two teams.

Extended Sprint — €16,000

The same, with 40 hours. Worth it when the codebase needs structural work alongside the practice change, rather than the practice change alone.

Additional cohorts

Each further pair of teams is €9,000, assuming cohorts run sequentially. Running several in parallel means committing more than four hours a day, so parallel delivery is quoted separately.

Scheduling

Standard delivery is no more than four hours per day — beyond that, retention drops and the hands-on work stops being effective. That daily cap is negotiable if the total block of hours increases.

Travel within the Netherlands is included. For on-site delivery elsewhere, travel and accommodation are billed at cost.

Why the implementation package usually wins

A masterclass changes what people know. It does not, on its own, change what they do on Monday.

The 20-hour block exists because the gap between understanding a practice and actually running it in your codebase is where adoption dies. Working through it together — in your repositories, on your real tickets — is what makes it stick.

If you’re choosing between the two, the honest answer is that the masterclass alone works well for a team that is already disciplined and just needs direction. If the team has tried and stalled, buy the implementation hours.

What you get

Deliverables

Intake questionnaire

Completed before the session so the material addresses your stack, your maturity level and your actual constraints — not a generic deck.

Two-hour masterclass

The full L1–L5 maturity path, weighted toward L3 and L4 where most of the real value sits, with practical worked examples throughout.

Practical examples

Demonstrations against realistic code, not toy problems — including the failure modes that only show up on real systems.

Team playbook

A written reference your team keeps afterwards, so the practice outlasts the session.

Let's scope it

A 30-minute call is usually enough to tell whether this is the right fit. No pitch.