This seems a strange thing to test given that Claude Code is optimized for Anthropic models.
A test of how different models perform in more of a model-agnostic harness like OpenCode would be much more interesting, perhaps paired with a control experiment of how those same models performed on the same task when using their respective native harnesses.
A test of how different models perform in more of a model-agnostic harness like OpenCode would be much more interesting, perhaps paired with a control experiment of how those same models performed on the same task when using their respective native harnesses.