What is Oqoqo for?
Oqoqo is designed for creating evaluations and private benchmarks for real-world agentic tasks, allowing users to assess whether agents can effectively use various products.
How is this different from trying an agent myself?
Oqoqo enables repeated evaluations across agents, models, and treatments, providing comprehensive data on performance rather than a one-time test.
Can I build private benchmarks?
Yes, users can write their own tasks and rubrics in plain language, which are versioned with each experiment for comparability.
What do I see after a run?
After a run, users receive detailed reports on pass or fail outcomes, including the full trajectory of the experiment, metrics, and reasons for failures.
Which agents can I run?
Oqoqo supports various agents, including Claude Code, Codex, Cursor, GitHub Copilot, and more, allowing users to choose models based on their needs.