Blog
From the team
ProductAugust 2026
Your Benchmark Is Grading the Wrong Thing
Golden sets measure agreement with the findings reviewers recorded, not an exhaustive inventory of a codebase's defects. Introducing Deep Audit (beta): evidence-gated review where every Bug ships with a citation.
TechnicalMay 2026
There Is No Best Code Reviewer
Research shows specialized multi-agent AI code review outperforms single-pass approaches by 40%+. No single model dominates across all bug types, languages, or codebases.
PerspectiveMay 2026
AI Writes Code Fast. Who Checks If It's Good?
AI shifted the bottleneck from writing code to knowing whether it's good. The data on AI code quality is sobering — and it points to a missing layer.
BenchmarkMay 2026
Martian Benchmark Testing
Measured tier results on the Martian-50 code-review benchmark, with methodology notes and official-submission status.