I hand-labelled 149 examples before I trusted my own scorer
Everyone can generate. Far fewer can tell you whether the output got better. Notes on building a small evaluation set by hand, calibrating a scorer against it, and what the disagreements taught me.
Read the post →