Your analytics agent is an employee.
Manage it like one.

It answers questions all day, costs money, and reports to no one. Benchouse reads its work, reviews its performance, listens to user feedback, tracks what it spends and tells you first when it gets something wrong.

From the team behind the analytics agent leaderboard.

The unmanaged analytics agent

No analyst works the way your agent does. You'd never allow it.

  • What did it tell the CFO last week?
  • What did its answers cost in March?
  • What do the product managers think of its quality?
  • Which question does it get wrong most often?
  • Which teams use it most?

If you can't answer them, you're running an unmanaged analytics agent.

It's not a person, but it does an analyst's job with none of an analyst's accountability, that's the part we fix.

The deal

We don't run the benchmark out of charity.

The public leaderboard is how we prove we can do this. Every vendor gets the same questions. A judge grades every answer against ground truth they never saw. Then we publish it, for everyone to see. It costs real time and money, and we take no money from vendors.

So why do it? Running it continuously teaches us that even the good ones get real questions wrong. The leaderboard helps you pick one, and even more importantly, it proves to you that you need to be in control of your agent. Everyone sees the benchmark. Some of you will buy Benchouse. That's the deal.

One system

It all runs on questions: the ones your team asks, and the ones you ask to check it.

For every question, we keep the whole record: the answer it gave, the SQL it ran, the reasoning behind it, what it cost, and who asked. One click more and any of them becomes an eval question. No exports, no glue, no spreadsheet on the side. That's the control center.

Logs

Read its work

The whole record, saved and searchable.

  • Search across agents and time
  • Tag and filter conversations
  • User feedback, captured with every answer

Evals

Run its performance review

The same discipline as our public benchmark, pointed at your agent.

  • Start from a real captured question, or write your own
  • Runs on a schedule, in CI/CD, or whenever you want
  • Versioned sets, so you can compare before and after a change

Cost

See its expense report

See where the money goes.

  • Cost per question, per person, per agent
  • Track cost over time
  • Check it against what the vendor actually billed

People

Know who it talks to

The humans doing the asking, in one place.

  • One profile per person, across every agent they use
  • Tag the CFO once, hear about it when they ask

Alerts

Let it escalate

Alerts that matter, sent where you work. Off by default.

  • An eval run finished
  • A person you watch asked a question
  • Any failures, the moment they happen

Connectivity

Connect it. Control it.

  • Connect it to your agents — Snowflake Cortex and Databricks Genie today
  • Control it with a full MCP server or the plain API

Who's behind this

André Baaij
André Baaij Founder, runs the Benchouse benchmark

Hi, I'm André Baaij. I ran data at Babylist, and now I run Benchouse: the public benchmark where analytics agents answer 300 questions against a warehouse with withheld ground truth, and the results get published whether vendors like them or not.

Running that benchmark taught me something uncomfortable: the difference between a great agent and a dangerous one is invisible from the chat window. Both answer fluently. You only see the difference when you record what they did and check it. That is exactly the work almost no team has time to do on their own traffic.

So I built Benchouse, to help every company manage their analytics agents. It's the same rigor I point at vendors every season, pointed at your agent. Let's work together.

— André

Let me show you Benchouse

Leave an email and I'll show you what it does on your own agents.