OpenAI's latest model, o3, posted an 88 percent score on the ARC‑AGI test, a result that dwarfs the human ceiling and forces the community to confront a measurement crisis.
The o3 release arrived in December 2024, just days after a Hacker News thread warned that scaling‑law debates were missing the point: existing models already reshaped capabilities. The same discussion noted that bench…
OpenAI announced o3 as part of its end‑of‑year series, positioning it as the most capable model to date. Independent observers on Hacker News highlighted its 88 percent ARC‑AGI score, calling it a “bombshell” because …
The same thread recalled that the MMLU benchmark, once a reliable gauge of cross‑domain language understanding, now sees “best models have saturated that one, too.” Earlier in 2024, GPQA – a physics, biology and chemi…