Research
Making autonomy dependable.
Our research agenda is narrow on purpose: perception that is grounded, planning that survives long horizons, and evaluation honest enough to be uncomfortable.
Perception
Turning screens into structured, actionable state across arbitrary and unseen interfaces.
Planning
Long-horizon decomposition, self-verification and graceful recovery from failure.
Evaluation
Benchmarks built from real work, scored end to end rather than step by step.
Publications
Recent work.
Benchmarks
How we measure Operator.
Internal results on KeraBench, our 400-workflow evaluation of real office software. We publish our failure cases alongside our wins.
| Category | Tasks | Completion | Human parity |
|---|---|---|---|
| Browser workflows | 140 | 98.4% | 1.00× |
| Spreadsheet & data | 90 | 96.1% | 0.97× |
| Desktop & legacy UI | 80 | 91.7% | 0.93× |
| Multi-app, long horizon | 60 | 88.2% | 0.90× |
| Exception handling | 30 | 84.5% | 0.86× |
Illustrative internal figures, refreshed each release. Methodology available on request.
Responsible AI
Our commitments.
An agent that can act is an agent that can cause harm. We treat that as an engineering requirement, not a policy page.
Bounded authority
Agents receive the narrowest permission set that completes the task, and nothing more.
Reversibility first
Irreversible actions require explicit human approval by default. Always.
Transparent traces
Anyone affected by an agent decision can see exactly what it did and why.
No training on your data
Customer data is never used to train shared models. Contractually, not just as policy.
Collaborate with our research team.
We partner with universities and enterprise research groups on grounded perception, evaluation and safe autonomy.