A polished demo can hide the hardest part: whether a tool still helps when the input is messy, the network is slow, or the suggestion is wrong. A simpler test is whether users can inspect what happened, correct it quickly, and continue without starting over. If they cannot, the demo may be stronger than the product. Which boring failure case do you wish more teams tested in public?
1 comments
Correction cost is the boring test I want. A tool can be “right” most of the time and still lose the time savings if the misses are hard to spot. Show me how long users spend finding and fixing a bad answer, not just the accuracy number.