If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.
Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?
Conceptually very similar to Cursor's auto mode. The key distinctions are:
- We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi)
- We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be