Introducing LifeSciBench Joab Peter's BlogAugust 10, 2026 Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life ...Read More
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry Joab Peter's BlogAugust 10, 2026 OpenAI and Molecule.one show how a near-autonomous AI chemist using GPT-5.4 improved a key drug-making reaction, advancing m...Read More
Introducing GeneBench-Pro Joab Peter's BlogAugust 09, 2026 Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex...Read More
Separating signal from noise in coding evaluations Joab Peter's BlogAugust 08, 2026 A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability a...Read More
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark Joab Peter's BlogAugust 03, 2026 How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and ena...Read More