For the call you’re making alone.

ChatGPT agrees with you. A council argues.

Ask a single AI and it reflects your framing back, polished. Bouleia convenes 3–7 frontier models in fixed roles — Skeptic, Devil’s Advocate, Chair — to debate your decision and tell you why you might be wrong, with the dissent kept on the record. The second opinion you’ve been deciding without.

ChairSkepticSr. ExpertCustomDevil's Adv.SynthesizerPragmatistVERDICTVERDICT

You bring the question  /  the product does the rest

No blank box

A planner does the hard part.

Paste a half-formed question — one line or twenty pages. An AI planner sharpens it into a real decision question and assembles the right panel for it. No prompt engineering, no wondering which models to pick.

Grounded, not frozen

Councils read the live web.

When a decision turns on current facts — a market, a competitor, this week's news — the seats pull live data before they deliberate. Verdicts hold up on time-sensitive calls, not just timeless reasoning.

One dial, not seven

You choose depth, not models.

Light, Standard, or Deep. Each lab fields the right seat for the weight of the decision, and a deeper setting adds a second round of cross-examination. Seven frontier models, hidden behind one human choice.

Seven labs  /  two countries  /  one table

US closed-source frontier
AnthropicOpenAIGoogle
Chinese open-source frontier
Moonshot AIZhipu AIDeepSeekAlibaba

The Problem

One AI. One blind spot. Repeated to you as a fact.

You asked one model. It agreed with you. That's the trap — a single model reflects your framing back, confident and unchallenged.

You asked Claude. Claude was thoughtful. Claude was confidently wrong on the part you couldn't check.

You asked GPT. GPT was thorough. GPT missed the counter-argument that would have changed your call.

You can’t tell which of the three to believe. So you trust the one you opened first — and you carry the decision alone.

Single-model AI is fine for what to cook for dinner. For everything else — the offer, the pivot, the architecture, the argument — it's a confident voice and no second opinion. That's not a tool. That's a coin flip with footnotes.

The Council

A council built to disagree with you.

The Bouleia council is a panel of frontier intelligences from independent labs, deliberating on your question with the roles you assign.

You seat a Skeptic to test the load-bearing claim, a Devil’s Advocate to attack the consensus, a Risk Officerto name who gets hurt. Their job is to find what you can’t.

The seats span the field: US closed-source frontier — Claude (Anthropic), GPT (OpenAI), Gemini (Google) — and the Chinese open-source frontier — Kimi (Moonshot AI), GLM (Zhipu AI), DeepSeek, Qwen (Alibaba). Different labs. Different training data. Different RLHF objectives. Different failure modes.

When seven labs from two countries with independent training pipelines agree, the agreement carries weight. When they don’t, the disagreement is the answer.

The dissent is not noise to be smoothed over. The dissent is the product.

How it works

Convene. Deliberate. Verdict.

  1. Convene.

    You bring the question — one line or twenty pages. An AI planner sharpens a half-formed prompt into a decision question and proposes the right panel, or you assemble the seats yourself. No format, no prompt engineering.

  2. Deliberate.

    Every seat answers independently first — no peeking, no cross-talk, so the disagreement only counts if it’s real. On deeper settings the council takes another round to test the strongest opposing points, and reads the live web when the decision turns on current facts.

  3. Verdict.

    A single, readable report. What the council agreed on. Where it split. The strongest case on each open point. The full transcript on a click, in case you want to dispute the synthesis yourself.

Design the deliberation

You convene. You charter. You read the verdict.

Different questions need different councils. A strategic pivot needs a Devil’s Advocate and a Strategist. A code review needs a Senior Expert and a Rookie. The product gives you the seats and the roles — you design the panel for the question in front of you.

US closed-source frontier

Claude · GPT · Gemini

Chinese open-source frontier

Kimi · GLM · DeepSeek · Qwen

Pick any combination. Three is a working council. Seven is a board. You decide.

You never wrangle model names or parameters. Choose a depth — Light, Standard, or Deep— and each lab fields the right seat for the weight of the decision in front of you.

The starter roles

  • Chair frames the question, holds the council to the brief.
  • Skeptic demands evidence, challenges every premise.
  • Senior Expert speaks from best-practice authority.
  • Rookie fresh-eyes perspective; 'why are we doing it this way?'
  • Devil's Advocate argues the strongest counter-position.
  • Pragmatist focuses on what's actually shippable, decidable, reversible.
  • Strategist long-term, second-order consequences.
  • Synthesizer reconciles the disagreement into one actionable verdict.

Or seat a model with no role — a Member at large, contributing on the question’s own terms.

Design your own roles

Starter roles are the floor, not the ceiling. The roles your team needs are specific to your team. Design them once. Use them in every council that follows.

A custom role has a name, a mandate (what this role is instructed to do, in your words), and a description. Save it. Seat it next to the starter roles. Reuse it across councils. Share it with your team.

A few examples of custom roles operators tend to build:

  • Our CFO's voice

    argues from unit economics, runway, capital efficiency. Models how your finance lead actually pushes back.

  • The regulator's lens

    reviews the proposal as a regulator would. Compliance-first, conservative, asks the questions that get a deal blocked.

  • Our most pessimistic engineer

    finds the failure modes. Names what will break in production at 10x load, at 100x, when the on-call rotates.

  • The customer who hates us

    argues from the perspective of the user we're failing. Drops every charitable assumption.

  • The historian

    references precedent. 'We tried this in 2023.' 'Three competitors have done this; here's how it went.'

  • The futurist

    projects second-order consequences. What does shipping this commit us to in eighteen months?

The role library is yours. Build the council you actually need, not the one a template anticipated.

Starter council templates

Templates are starting points. Edit them per question.

  • Validate an Idea Skeptic, Devil's Advocate, Pragmatist, Synthesizer.

    Pre-commit reality check on a hypothesis.

  • Critique My Work Senior Expert, Skeptic, Rookie.

    Get the work torn apart before someone else does.

  • Strategic Decision Strategist, Pragmatist, Devil's Advocate, Chair.

    High-stakes, near-irreversible choices.

  • Technical Review Senior Expert, Rookie, Skeptic.

    Architecture, code, or system design calls.

  • Custom You assemble it..

    For the questions templates don't reach.

A council meeting

Here’s what a verdict looks like.

Before
The Council
Re
Hire vs. polish
Sitting
5 seats · 5 models
Filed
Sample · on the record

Convened by you. Five seats — four starter roles and one custom role from this operator’s library. The question follows.

We're a 14-person SaaS at $1.8M ARR, six months from running out of runway. Inbound is strong but conversion is weak. Should we hire a head of growth, or invest the same money in product polish to lift conversion ourselves?

Chair

Seat 1 · Claude

The choice as posed is the wrong frame. 'Conversion' is at least two problems: positioning — which your strong inbound says is fine — and activation, which weak conversion says is not. So this isn't hire-vs-polish; it's which leak, and can you fix it without a hire. And before anyone prices this decision on Devil's Advocate's '3% in eight weeks' — that is exactly the confident, uncheckable number a call like this dies on. I won't let the council cost the decision on a figure no one in this room can verify. [...continued in transcript]

Devil's Advocate

DISSENT

Seat 2 · GPT

I'll hold the unpopular line: inbound this strong is the signal to hire, not to audit. A head of growth who lands fast compounds in a way founder-side tinkering won't — and yes, Chair, my numbers are illustrative, but the asymmetry isn't. If the leak turns out to be positioning rather than UX mechanics, none of you fix that in a two-week audit, and you'll have burned a month proving what a senior hire already knows. Hire now; run the audit in parallel. [...continued in transcript]

Pragmatist

Seat 3 · Kimi

Sequence by reversibility. A two-week activation audit is cheap, reversible, and high-confidence: you either find the mechanical leak and ship the fix yourself, or you rule it out — and then Devil's Advocate is right. The hire is expensive, slow, and irreversible on a six-month clock. You don't buy the expensive irreversible thing before the cheap reversible one has spoken. Fix activation first; hold the hire. [...continued in transcript]

Our CFO's voice

CUSTOM

Seat 4 · GLM

Neither option survives contact with the runway. At $1.8M ARR on six months of cash, the head-of-growth hire is a 9-month payback against a 6-month clock — you'd be dead before it pays. That's not a growth decision, it's a solvency one. So sequence it: open annual-prepay with your happiest accounts to buy two more months — which is exactly what makes 'fix activation yourself first' affordable instead of reckless. The hire isn't wrong. It's out of order. [...continued in transcript]

Synthesizer

Seat 5 · DeepSeek

Three of four seats land in the same place: the hire is the right move at the wrong time. Chair set aside Devil's Advocate's uncheckable eight-week number; Pragmatist sequenced by reversibility and CFO by runway, and they agree. So this is not 'wait and see.' It's a committed call — fix activation now, extend runway now, hold the hire behind a named tripwire. Devil's Advocate's dissent is the fork: if the audit says positioning, the hire moves to the front.

Council recommendation: Fix activation now — hold the hire behind a tripwire.

Council recommendation
Do not hire the head of growth yet. Run a two-week activation audit and ship the fixes yourself, and open annual-prepay conversations in parallel to extend runway past the six-month clock. Reopen the hire only if the tripwire below trips.
Where the council agreed (3 of 4)
The head-of-growth hire is the right move at the wrong time — a 9-month payback against a six-month runway. The cheap, reversible fix (activation) goes first; the expensive, irreversible one (the hire) waits.
Where the council split — the tripwire
Devil’s Advocate (GPT) holds that inbound this strong is itself the reason to hire now. The dissent is load-bearing, not a footnote: if the audit shows the leak is positioning rather than activation mechanics, the founder can’t fix it alone — and GPT’s call becomes the recommendation. That’s the fork you own.
Caught on the record
The Chair (Claude) set aside Devil’s Advocate’s “3% in eight weeks” as an unverifiable figure — the exact confident, uncheckable claim a single model would have handed you as fact. One seat catching another is the point.
Strongest single-seat position
Our CFO’s voice (GLM) — it turned a binary cash bet into a sequenced, survivable plan: the prepay move is what makes “fix it yourself first” affordable rather than reckless. The reframe, not either original option, is the decision.
Council action
Start the activation audit this week. Open prepay conversations in parallel. Keep the hire warm — reopen it the day the audit points at positioning.

This is a generated sample. Your council, your seats, your question, your verdict.

Not a sample  /  a real council, on the record

6 seats · 6 models · 2 countries · 2 rounds

A seed-stage AI product spending ~$8k/month on frontier-model APIs asked: keep building on closed frontier APIs, or migrate core inference to self-hosted open-weight models — at mid-2026 pricing?

Stay on closed APIs now — but build the option to leave.

The council held that self-hosting a frontier-class model doesn’t pay at $8k/month — but agreed the real move is a model-neutral routing layer and eval suite, so migrating later stays an option instead of a rebuild.

Where it split

The Devil’s Advocate (DeepSeek) refused the cautious line — arguing for an immediate, binding migration of 15–20% of traffic to hosted open-weight models, precisely to force the infrastructure to get built. The verdict records the disagreement rather than smoothing it over.

Claude · Gemini · Kimi · GLM · DeepSeek · MiniMax

Council verdict

6 seats · 6 models · 2 countries · 2 rounds

Recommendation

Stay on closed frontier APIs for now — but invest now in the option to leave: a model-neutral routing layer, an eval suite, and workload-targeted hosted open-weight models.

Where they agreed

Self-hosting a frontier-class model doesn’t pay at ~$8k/month on a full total-cost-of-ownership basis — and a model-agnostic setup is worth building regardless, to keep flexibility and prevent future lock-in.

Where they split  ·  dissent, on the record

The Devil’s Advocate (DeepSeek) broke from the panel: migrate 15–20% of traffic to hosted open-weight models immediately and bindingly — precisely to force the infrastructure to actually get built — rather than piloting cautiously alongside the current APIs.

Strongest single seat

The Senior Expert (Kimi) — for turning a binary build-vs-migrate bet into a staged plan that keeps both options open.

The action

Stand up a routing layer and run a workload-targeted pilot of hosted open-weight models on routine tasks. Adopt a harness-neutrality rule to avoid provider lock-in, and watch for the economic milestone that flips the call.

This is the synthesis. The full transcript — every seat, every round — lives on the public page.

Open the full public verdict

When to convene

For the decisions you’d normally carry alone.

Not every question deserves a council. For ‘what’s a good name for my cat,’ open the chat app you trust. The council is for the questions where the cost of being wrong is real — where you’d want a second and third opinion if you had time to ask for one.

The questions you bring to the council are the ones you currently bring to a co-founder, a senior advisor, a board member. Now you have all three, and you get to choose who sits in each seat.

Why both — open and closed

Different labs. Different blind spots.

Different labs make different mistakes. Different training data exposes different blind spots. Different RLHF objectives prioritise different things. A council that’s all one lineage — all US frontier, all from California, all closed-source — shares more blind spots than it admits.

US closed-source seats

Claude · GPT · Gemini

Frontier-lab capability at the edge of the field. Anthropic, OpenAI, Google.

Chinese open-source seats

Kimi · GLM · DeepSeek · Qwen

Independent training pipelines, different architectures, different training data. Moonshot AI, Zhipu AI, DeepSeek, Alibaba.

The Bouleia council deliberately spans:

  • US closed-source frontier — Claude, GPT, Gemini — for raw capability at the edge of the field, trained by independent US labs with independent failure modes.
  • Chinese open-source frontier — Kimi, GLM, DeepSeek, Qwen — for genuinely independent training pipelines. The open-source frontier work in 2026 is largely happening in Chinese labs; a council that ignores it is a council that pretends the rest of the field isn’t there.
  • Multiple architectures — dense and Mixture-of-Experts — for less-correlated failure modes.
  • Geopolitical spread — questions answered by both US and Chinese labs surface where regulatory framing, training-data politics, and cultural framing actually differ. The disagreement is informative even when it’s uncomfortable.

A council that agrees across that range is a council you can act on. A council that splits across that range tells you exactly where the disagreement lives.

Pricing

Pay for the deliberation, not a subscription.

Bouleia is pay-as-you-go. Top up your account, convene a council, and pay a few dollars to pressure-test a decision worth thousands — no subscription, no lock-in.

Pay-as-you-go

A prepaid balance you draw down per council. No subscription, no lock-in, no minimum.

Priced to the decision, not a subscription
A few dollars to pressure-test a call worth thousands. A deeper panel or a heavier question costs a little more than a quick one — you pay for the weight of the deliberation, not a monthly seat you may never use.
Estimate up front, receipt after
You see an estimated cost before you convene, and an itemised receipt once the verdict lands. No surprises, no meter running in the background — you approve the spend before a single seat speaks.
Top up, draw down
No subscription and no per-seat fee. Add credit when you need it and draw it down as you convene. Run one council this quarter or a hundred — you pay for the decisions you bring, nothing else.

$5 in founder credit for waitlist members. The first members start with credit on the house — enough to convene your first councils free — and keep founder rates as public pricing settles at launch.

Founder rates land at launch. Waitlist members get in first, and on the best terms.

No card required to join. We’ll email when seats open.

Questions

Questions worth asking before you ask the council.

Isn't this just another multi-model wrapper?

Wrappers send one prompt to many models and stack the answers next to each other. The council is different in two ways. First, you assign each seat a role — the deliberation reflects the roles, not just the models. Second, the council synthesises with named dissents preserved — you get one actionable verdict and the full transcript of where the seats disagreed. The disagreement is the product, not noise to be hidden.

Why include open-source models when closed-source ones score higher on benchmarks?

Benchmark scores measure capability at the edge. Council quality measures independence of judgment. The Chinese open-source frontier — Kimi, GLM, DeepSeek, Qwen — is trained on different data, with different RLHF objectives, under a different regulatory framing, by labs with no commercial relationship to the US frontier labs. That difference is what makes the council's agreement meaningful. A panel of three US frontier labs sharing much of their training data and most of their cultural assumptions is not a council. It's a chorus.

Why these specific open-source models? Why not Llama or Mistral?

Because the open-source frontier work in 2026 is happening in Chinese labs. Kimi (Moonshot AI), GLM (Zhipu AI), DeepSeek, and Qwen (Alibaba) are where the strongest open-source capability is being released right now — reasoning chains, MoE architectures, long-context, multilingual. A council built to maximise independence-of-judgment uses the strongest seats available, not the most familiar names. The panel will evolve as the open-source landscape does; if a Western open-source model returns to the frontier, it gets a seat.

Can I pick which models are in my council?

Yes. You compose the panel from the available seats — any combination of US closed-source frontier (Claude, GPT, Gemini) and Chinese open-source frontier (Kimi, GLM, DeepSeek, Qwen). Three is a working council. Seven is a board.

What does 'assigning a role' actually do?

The role becomes part of the seat's mandate. A model seated as Devil's Advocate is instructed to argue the strongest counter-position, even if it doesn't believe it. A model seated as Senior Expert is instructed to speak from best-practice authority. A model seated as Rookie is instructed to ask the questions a fresh perspective would ask. The role changes the deliberation; it isn't decoration.

Can I design my own roles?

Yes. The eight starter roles are the floor, not the ceiling. You can design custom roles — give them a name, a mandate (the instructions that shape how the seat behaves), and a description — and save them to your workspace's role library for reuse in every future council. Most operators end up with a handful of custom roles that mirror how their team thinks: "Our CFO's voice", "The regulator's lens", "Our most pessimistic engineer". Build the roles your specific deliberations need.

What if all the seats agree on something wrong?

A council can share a blind spot — especially if the seats are too similar. That's why panel diversity (closed + open + different geographies + different architectures) matters more than panel size. Consensus across a diverse panel is strong evidence; consensus across three similar models is just three models saying the same thing twice. The transcript is there so you can audit the reasoning yourself.

How does billing work?

Pay-as-you-go — no subscription and no per-seat fee. You top up a prepaid balance and draw it down per council. Each council is priced on the tokens it uses, weighted by the model each seat runs: a frontier closed-source seat costs more than an open-source one, and a seven-seat board costs more than a three-seat council. You see an estimated cost before you convene and an itemised receipt — tokens, models, and time — after the verdict. Early adopters start with $5 of founder credit.

What happens to my question?

Sent to each seated model's API for the duration of the query. Not retained on our servers beyond what's needed to deliver your report. Provider-side retention follows each provider's stated policy. Open-source seats route through hosted inference providers. Full privacy disclosure published at launch.

What does 'council-validated' actually mean?

It means your question was answered by multiple independent models in defined roles, the points of agreement and disagreement are explicit, and the audit trail is yours. It's not a guarantee of correctness — nobody can guarantee that. It's a guarantee that more than one mind thought hard about your question from the roles you assigned, and the disagreements weren't smoothed away.

How long does a council meeting take?

It depends on the complexity of the question — expect a verdict in roughly two to five minutes. The bottleneck is the slowest seated model, and we'd rather make you wait for considered output than rush it. Harder questions and larger councils take longer; simpler questions and smaller councils return faster — another reason to design the panel for the question.

Reserve your seat

Convene the council.

We’re opening a small first cohort. Reserve your seat and we’ll bring you in — with $5 of founder credit to convene your first councils on us.

No card required. One email at most per week, only when there’s news worth your attention.