Date announced soon60 minutes, including questions
Evaluating an AI feature before and after release
Building a small evaluation set from real cases, scoring it without a dedicated team, and running it before every prompt or model change so a regression is caught before it ships.
Date announced soon
60 minutes, hands on
Online, live
Free
Dates go to the list first. Email us and you will hear before it is posted anywhere.
Outcomes
What you leave with
- A small evaluation set built from your own real cases, not invented ones
- A pass or fail signal you can run before every change
- The failure modes worth testing for, including the ones that only appear under load
- A view on when a human has to stay in the loop, and where that costs you nothing
Founders with an AI feature in front of customers, or about to be.
Twenty real examples of the input your feature receives, if you have them.
Format
Running order
Useful if you have something shipped, or nearly. Questions are taken throughout rather than held to the end.
The limits of manual review
Spot checks and user feedback both tell you a customer has already been harmed. We start from what they miss.
Building the evaluation set
Twenty real cases beats two hundred invented ones. We build yours in the session.
Scoring at small scale
What can be automated, what needs a person and how to keep the whole thing under an hour a week.
Running it before each release
Wiring the set into your release step, so a prompt or model change is checked before it reaches customers.
Date announced soon
Evaluating an AI feature before and after release
Free, online and live. Dates go to the list before they are posted anywhere.