Discussion about this post

User's avatar
Jonah Gray's avatar

This is exactly why I think the evaluated unit has to be model + harness + serving path + verifier, not the model name. The same nominal model can see a different transcript, get a different response budget, or complete fewer cycles before timeout depending on the path. Cost per token hides all of that; cost per accepted solution exposes it. I also like that the multi-agent result failed on the held-out set. In our own work, we are trying to treat topology as an independent variable and hold the acceptance contract constant across single and decomposed runs. For the next pass, will you also pin total inference budget and tool-call budget, or is wall clock the governing constraint?

Pito Salas's avatar

Hey Raffi - I love the stuff you’re working on. Help me with some terminology, explain: “Then one model name, kimi-k3, across four serving paths. And one open model, GLM-5.2, hosted through Codex at two more providers”. Thanks.

No posts

Ready for more?