Blind-graded · One prompt per repo · Updated
Bug Hunt Bench
Choose runs
A model can appear more than once — different reasoning tiers, or a different harness. Runs tagged superseded are historical and stay out of the default view.
No runs selected. Pick some above, or hit Featured.
No runs selected. Pick some above, or hit Featured.
Every value here is also in the .
No runs selected. Pick some above, or hit Featured.
Every value here is also in the .
No runs selected. Pick some above, or hit Featured.
A turn is one model step that either calls a tool or gives the final answer, counted from each run's own log the same way for every harness — see the definition. Every value here is also in the , and every run's own figures are in results/runs.csv.
No runs selected. Pick some above, or hit Featured.
Bug Hunt Bench is run by Pawel Huryn, who writes The Product Compass, a newsletter on AI for product managers and builders. Same standards as this board: hands-on, no hype, nothing that was not run first.