Skip to content
AI-Native Mentorship

Dedicated to one of my absolute favorite people, Yagika (opens in new tab)—my first student.

← Use cases
intermediateAbout two hours to set up; each release since is a re-run, not a rebuild

Making 'good answer' legible across four releases


What we did

FAQ answer quality lived in a weekly progress meeting — a matter of opinion in a room, until it had to be compared release over release.

Why we did it

'This got better' has to be defensible — a score without reasoning is just a louder opinion.

Who this is for

Content and support teams still proving progress in a weekly meeting.

How was it done before

A weekly meeting existed to showcase the work: someone read the answers, formed an opinion, and presented it as a slide summary — a score with no reasoning attached, impossible to compare from one release to the next.

How is it done now

  1. 01A dashboard instead of a meeting: the evaluation runs on one command
  2. 02Harvey balls plus reasoning, not a bare score
  3. 03Each release re-runs the same rubric — quality compared release over release
  4. 04The link goes to the team: review and results, no discussion meeting

How much time did it take

About two hours to set up; each release since is a re-run, not a rebuild.

The result

The weekly showcase meeting is gone: one command re-runs the evaluation and updates the dashboard, the team reviews it from a link, and 'this got better' arrives as Harvey-ball scores with narrative reasoning — comparable across four releases.