Agreed. Test driving the bad default is what they deserve (they brought this upon themselves as a benchmaxx attempt), but comparing it to running with reasoning completely disabled is also weird.
I think the point is that a larger model with far less reasoning time solves the problem just fine. Which of course is the tradeoff: The smaller the model, the more reasoning you need to get decent answers to tough questions.