Making 'good answer' legible across four releases
What we did
FAQ answer quality lived in a weekly progress meeting — a matter of opinion in a room, until it had to be compared release over release.
Why we did it
'This got better' has to be defensible — a score without reasoning is just a louder opinion.
Who this is for
Content and support teams still proving progress in a weekly meeting.
How was it done before
A weekly meeting existed to showcase the work: someone read the answers, formed an opinion, and presented it as a slide summary — a score with no reasoning attached, impossible to compare from one release to the next.
How is it done now
- 01A dashboard instead of a meeting: the evaluation runs on one command
- 02Harvey balls plus reasoning, not a bare score
- 03Each release re-runs the same rubric — quality compared release over release
- 04The link goes to the team: review and results, no discussion meeting
How much time did it take
About two hours to set up; each release since is a re-run, not a rebuild.
The result
The weekly showcase meeting is gone: one command re-runs the evaluation and updates the dashboard, the team reviews it from a link, and 'this got better' arrives as Harvey-ball scores with narrative reasoning — comparable across four releases.

