DAILY ENTRY · PACIFIC TIME

AI Needs Evidence, Not Vibes

2026-08-13 · TEACH

A useful AI evaluation defines the task, expected behavior, failure thresholds, representative cases, and a decision owner.

Some work is best introduced by the problem it solves before its mechanics are made public. I have been developing a private grading and evaluation system for making model behavior and evidence easier to assess.

I am intentionally keeping the implementation and full method off this site for now. If the right conversation arises at DevDay, I would be glad to demonstrate it to OpenAI leaders who want to examine the approach.

Go deeper

Follow the references below for more detail, or continue through the Journey, Lab, Research, and Teaching sections.

BACK TO TOP