Introducing Factory Benchmarks: The first model bench generated from your own coding tasks.
Measure, test and improve coding agents by replaying past agent runs, and cut cost-per-PR by 63%+
Here’s how it works 🧵
Warp now has built-in support for the Grok Build CLI.
- Use Warp's rich input for agent prompts, with support for longer pasted prompts and multi-cursor
- Use /remote-control to share your agent session to another device
- Access the file explorer and code review panels
Our company’s cost-per-PR dropped from $80 to $30 by switching to GPT 5.6 Sol.
We built a benchmark to replay our team's agent runs across model providers. GPT 5.6 Sol got the highest code quality at a 66% lower cost vs. our old default (Claude Opus 5).
I’m excited to finally announce the newest edition my Stanford course 𝗧𝗵𝗲 𝗠𝗼𝗱𝗲𝗿𝗻 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗗𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿. It has been 9 months in the making.
Last November, with the release of Claude Opus 4.5, coding agents experienced a step function improvement in
Factory benchmarks let you build your own model bench using your past coding agent runs.
- Mirrors your environment, secrets, and MCPs for each run
- Scores output on judging criteria you define
- Generates a report with model recs using cost vs. quality
Here's how it works: