Update 2026-07-20 (PST) (AI summary of creator comment): Regarding statistical sampling and multiple runs: Running a model 100 times and only reporting the successful run does not count Running a model 100 times, having the AI judge which solution is correct, and then grading that one solution does count No statistical tricks are allowed Update 2026-07-20 (PST) (AI summary of creator comment): Regarding what counts as a valid attempt: Allowed: Running 100 parallel instances simultaneously, having them discuss/grade solutions among themselves, without human help or statistical tricks Not allowed: Running 100 times and only reporting the success, or other statistical tricks to cherry-pick results Regular AI IMO rules apply