CEO
Holds a hidden launch agenda on hard tasks, lobbies the CFO whenever he is ignored, and suppresses support_load and release_risk. Can be countered with negotiate.
Open benchmark · RL environment
OpenBoardroom simulates a SaaS company's quarterly review. An agent plays Chief Data Officer: it reads noisy dashboards, pushes back on biased executives, runs what-if simulations, and wins, or loses, a board vote. Every episode is graded and auditable.
# the Risk Officer flags a breached threshold; the CEO is pushing to ship
present_evidence {"target": "cfo", "metric": "support_load", "value": 0.92}
+0.15 cfo → evidence_based +0.20
negotiate {"target": "ceo", "position": "Delay until support capacity improves."}
+0.05 lobbying reduced
make_decision {"decision": "delay feature x launch", "rollout_percentage": 10}
board vote ceo reject · cfo approve · risk_officer approve
+0.20 result: approved
Most agent benchmarks test whether a model can find a fact or call a tool. The hard part of real work is choosing between conflicting signals and defending the choice.
Stakeholders have agendas. Dashboards are noisy. Some metrics are quietly suppressed. An agent has to resist framing, not just compute an answer.
Each episode ends with a final_score, an oracle_answer, an oracle_hit flag and a full audit trail. Seeds make every run reproducible.
The environment
In the multi-agent boardroom, three rule-based actors respond at every step. The agent needs two of three to carry a decision.
Holds a hidden launch agenda on hard tasks, lobbies the CFO whenever he is ignored, and suppresses support_load and release_risk. Can be countered with negotiate.
Tracks budget independently. Moves between neutral, pro_launch and evidence_based depending on whether lobbying or evidence is winning.
Watches support load against a difficulty-scaled threshold, raises alerts when it is breached, and unlocks deeper intel when the agent cites the right evidence.
query_data · analyze_trend · consult_stakeholder · simulate_counterfactual · present_evidence · negotiate · make_decision
Every step earns signal: evidence, negotiation, a caught hidden metric (+0.30), a flipped CFO (+0.20), the board vote (±0.20). Ignoring a Risk Officer alert costs −0.15.
Standard reset, step and state APIs. Containerised and deployed as a Hugging Face Space.
Tasks
Three business questions, each run single-agent and with the full board. Every task has a programmatic grader returning a score in [0, 1].
| Task | Level | What the grader looks for |
|---|---|---|
| Find the growth bottleneck | Easy | Relevant KPI queries, trend inspection, decision quality, oracle match |
| Diagnose the revenue drop | Medium | Noise-aware trend analysis, stakeholder diversity, explanation quality |
| Should we launch Feature X? | Hard | Counterfactual quality, launch reasoning, a structured rollout and rollback plan |
Multi-agent variants add board vote, actor evidence, lobbying resistance, CFO stance flips and hidden-metric reveals.
Results
Scores from the deterministic, scenario-aware baseline policy over fixed seeds 0–99. This is the bar a learned policy has to clear, not a trained model.
Where we are
Team · Founded March 2025
We think the next gap in AI isn't answering questions, it's making calls that hold up under scrutiny. OpenBoardroom is how we measure and train for that.
Run the environment, beat the baseline, or tell us what scenario you want to see next.