Blind-graded · One prompt per repo · Updated
Bug Hunt Bench
Choose runs
A model can appear more than once — different reasoning tiers, or a different harness. Runs tagged superseded are historical and stay out of the default view.
No runs selected. Pick some above, or hit Featured.
No runs selected. Pick some above, or hit Featured.
Cost is on a logarithmic axis: the spread across the board is roughly two hundredfold, and a linear axis would pile half the runs into the left margin. Every value here is also in the .
No runs selected. Pick some above, or hit Featured.
Minutes are on a linear axis: the board spans under sevenfold, well inside one order of magnitude, so a linear axis places every run honestly and keeps the reading additive. Every value here is also in the .
No runs selected. Pick some above, or hit Featured.
Bug Hunt Bench is run by Pawel Huryn, who writes The Product Compass, a newsletter on AI for product managers and builders. Same standards as this board: hands-on, no hype, nothing that was not run first.