US closed-source frontier
Claude · GPT · Gemini
For the call you’re making alone.
Ask a single AI and it reflects your framing back, polished. Bouleia convenes 3–7 frontier models in fixed roles — Skeptic, Devil’s Advocate, Chair — to debate your decision and tell you why you might be wrong, with the dissent kept on the record. The second opinion you’ve been deciding without.
You bring the question / the product does the rest
No blank box
Paste a half-formed question — one line or twenty pages. An AI planner sharpens it into a real decision question and assembles the right panel for it. No prompt engineering, no wondering which models to pick.
Grounded, not frozen
When a decision turns on current facts — a market, a competitor, this week's news — the seats pull live data before they deliberate. Verdicts hold up on time-sensitive calls, not just timeless reasoning.
One dial, not seven
Light, Standard, or Deep. Each lab fields the right seat for the weight of the decision, and a deeper setting adds a second round of cross-examination. Seven frontier models, hidden behind one human choice.
Seven labs / two countries / one table
The Problem
You asked one model. It agreed with you. That's the trap — a single model reflects your framing back, confident and unchallenged.
You asked Claude. Claude was thoughtful. Claude was confidently wrong on the part you couldn't check.
You asked GPT. GPT was thorough. GPT missed the counter-argument that would have changed your call.
You can’t tell which of the three to believe. So you trust the one you opened first — and you carry the decision alone.
Single-model AI is fine for what to cook for dinner. For everything else — the offer, the pivot, the architecture, the argument — it's a confident voice and no second opinion. That's not a tool. That's a coin flip with footnotes.
The Council
The Bouleia council is a panel of frontier intelligences from independent labs, deliberating on your question with the roles you assign.
You seat a Skeptic to test the load-bearing claim, a Devil’s Advocate to attack the consensus, a Risk Officerto name who gets hurt. Their job is to find what you can’t.
The seats span the field: US closed-source frontier — Claude (Anthropic), GPT (OpenAI), Gemini (Google) — and the Chinese open-source frontier — Kimi (Moonshot AI), GLM (Zhipu AI), DeepSeek, Qwen (Alibaba). Different labs. Different training data. Different RLHF objectives. Different failure modes.
When seven labs from two countries with independent training pipelines agree, the agreement carries weight. When they don’t, the disagreement is the answer.
The dissent is not noise to be smoothed over. The dissent is the product.
How it works
You bring the question — one line or twenty pages. An AI planner sharpens a half-formed prompt into a decision question and proposes the right panel, or you assemble the seats yourself. No format, no prompt engineering.
Every seat answers independently first — no peeking, no cross-talk, so the disagreement only counts if it’s real. On deeper settings the council takes another round to test the strongest opposing points, and reads the live web when the decision turns on current facts.
A single, readable report. What the council agreed on. Where it split. The strongest case on each open point. The full transcript on a click, in case you want to dispute the synthesis yourself.
Design the deliberation
Different questions need different councils. A strategic pivot needs a Devil’s Advocate and a Strategist. A code review needs a Senior Expert and a Rookie. The product gives you the seats and the roles — you design the panel for the question in front of you.
US closed-source frontier
Claude · GPT · Gemini
Chinese open-source frontier
Kimi · GLM · DeepSeek · Qwen
Pick any combination. Three is a working council. Seven is a board. You decide.
You never wrangle model names or parameters. Choose a depth — Light, Standard, or Deep— and each lab fields the right seat for the weight of the decision in front of you.
Or seat a model with no role — a Member at large, contributing on the question’s own terms.
Starter roles are the floor, not the ceiling. The roles your team needs are specific to your team. Design them once. Use them in every council that follows.
A custom role has a name, a mandate (what this role is instructed to do, in your words), and a description. Save it. Seat it next to the starter roles. Reuse it across councils. Share it with your team.
A few examples of custom roles operators tend to build:
“Our CFO's voice”
argues from unit economics, runway, capital efficiency. Models how your finance lead actually pushes back.
“The regulator's lens”
reviews the proposal as a regulator would. Compliance-first, conservative, asks the questions that get a deal blocked.
“Our most pessimistic engineer”
finds the failure modes. Names what will break in production at 10x load, at 100x, when the on-call rotates.
“The customer who hates us”
argues from the perspective of the user we're failing. Drops every charitable assumption.
“The historian”
references precedent. 'We tried this in 2023.' 'Three competitors have done this; here's how it went.'
“The futurist”
projects second-order consequences. What does shipping this commit us to in eighteen months?
The role library is yours. Build the council you actually need, not the one a template anticipated.
Templates are starting points. Edit them per question.
Validate an Idea — Skeptic, Devil's Advocate, Pragmatist, Synthesizer.
Pre-commit reality check on a hypothesis.
Critique My Work — Senior Expert, Skeptic, Rookie.
Get the work torn apart before someone else does.
Strategic Decision — Strategist, Pragmatist, Devil's Advocate, Chair.
High-stakes, near-irreversible choices.
Technical Review — Senior Expert, Rookie, Skeptic.
Architecture, code, or system design calls.
Custom — You assemble it..
For the questions templates don't reach.
A council meeting
Convened by you. Five seats — four starter roles and one custom role from this operator’s library. The question follows.
We're a 14-person SaaS at $1.8M ARR, six months from running out of runway. Inbound is strong but conversion is weak. Should we hire a head of growth, or invest the same money in product polish to lift conversion ourselves?
Seat 1 · Claude
The choice as posed is the wrong frame. 'Conversion' is at least two problems: positioning — which your strong inbound says is fine — and activation, which weak conversion says is not. So this isn't hire-vs-polish; it's which leak, and can you fix it without a hire. And before anyone prices this decision on Devil's Advocate's '3% in eight weeks' — that is exactly the confident, uncheckable number a call like this dies on. I won't let the council cost the decision on a figure no one in this room can verify. [...continued in transcript]
Seat 2 · GPT
I'll hold the unpopular line: inbound this strong is the signal to hire, not to audit. A head of growth who lands fast compounds in a way founder-side tinkering won't — and yes, Chair, my numbers are illustrative, but the asymmetry isn't. If the leak turns out to be positioning rather than UX mechanics, none of you fix that in a two-week audit, and you'll have burned a month proving what a senior hire already knows. Hire now; run the audit in parallel. [...continued in transcript]
Seat 3 · Kimi
Sequence by reversibility. A two-week activation audit is cheap, reversible, and high-confidence: you either find the mechanical leak and ship the fix yourself, or you rule it out — and then Devil's Advocate is right. The hire is expensive, slow, and irreversible on a six-month clock. You don't buy the expensive irreversible thing before the cheap reversible one has spoken. Fix activation first; hold the hire. [...continued in transcript]
Seat 4 · GLM
Neither option survives contact with the runway. At $1.8M ARR on six months of cash, the head-of-growth hire is a 9-month payback against a 6-month clock — you'd be dead before it pays. That's not a growth decision, it's a solvency one. So sequence it: open annual-prepay with your happiest accounts to buy two more months — which is exactly what makes 'fix activation yourself first' affordable instead of reckless. The hire isn't wrong. It's out of order. [...continued in transcript]
Seat 5 · DeepSeek
Three of four seats land in the same place: the hire is the right move at the wrong time. Chair set aside Devil's Advocate's uncheckable eight-week number; Pragmatist sequenced by reversibility and CFO by runway, and they agree. So this is not 'wait and see.' It's a committed call — fix activation now, extend runway now, hold the hire behind a named tripwire. Devil's Advocate's dissent is the fork: if the audit says positioning, the hire moves to the front.
This is a generated sample. Your council, your seats, your question, your verdict.
Not a sample / a real council, on the record
6 seats · 6 models · 2 countries · 2 rounds
A seed-stage AI product spending ~$8k/month on frontier-model APIs asked: keep building on closed frontier APIs, or migrate core inference to self-hosted open-weight models — at mid-2026 pricing?
The council held that self-hosting a frontier-class model doesn’t pay at $8k/month — but agreed the real move is a model-neutral routing layer and eval suite, so migrating later stays an option instead of a rebuild.
Where it split
The Devil’s Advocate (DeepSeek) refused the cautious line — arguing for an immediate, binding migration of 15–20% of traffic to hosted open-weight models, precisely to force the infrastructure to get built. The verdict records the disagreement rather than smoothing it over.
Claude · Gemini · Kimi · GLM · DeepSeek · MiniMax
When to convene
Not every question deserves a council. For ‘what’s a good name for my cat,’ open the chat app you trust. The council is for the questions where the cost of being wrong is real — where you’d want a second and third opinion if you had time to ask for one.
The questions you bring to the council are the ones you currently bring to a co-founder, a senior advisor, a board member. Now you have all three, and you get to choose who sits in each seat.
Why both — open and closed
Different labs make different mistakes. Different training data exposes different blind spots. Different RLHF objectives prioritise different things. A council that’s all one lineage — all US frontier, all from California, all closed-source — shares more blind spots than it admits.
US closed-source seats
Frontier-lab capability at the edge of the field. Anthropic, OpenAI, Google.
Chinese open-source seats
Independent training pipelines, different architectures, different training data. Moonshot AI, Zhipu AI, DeepSeek, Alibaba.
The Bouleia council deliberately spans:
A council that agrees across that range is a council you can act on. A council that splits across that range tells you exactly where the disagreement lives.
Pricing
Bouleia is pay-as-you-go. Top up your account, convene a council, and pay a few dollars to pressure-test a decision worth thousands — no subscription, no lock-in.
A prepaid balance you draw down per council. No subscription, no lock-in, no minimum.
$5 in founder credit for waitlist members. The first members start with credit on the house — enough to convene your first councils free — and keep founder rates as public pricing settles at launch.
Founder rates land at launch. Waitlist members get in first, and on the best terms.
No card required to join. We’ll email when seats open.
Questions
Wrappers send one prompt to many models and stack the answers next to each other. The council is different in two ways. First, you assign each seat a role — the deliberation reflects the roles, not just the models. Second, the council synthesises with named dissents preserved — you get one actionable verdict and the full transcript of where the seats disagreed. The disagreement is the product, not noise to be hidden.
Benchmark scores measure capability at the edge. Council quality measures independence of judgment. The Chinese open-source frontier — Kimi, GLM, DeepSeek, Qwen — is trained on different data, with different RLHF objectives, under a different regulatory framing, by labs with no commercial relationship to the US frontier labs. That difference is what makes the council's agreement meaningful. A panel of three US frontier labs sharing much of their training data and most of their cultural assumptions is not a council. It's a chorus.
Because the open-source frontier work in 2026 is happening in Chinese labs. Kimi (Moonshot AI), GLM (Zhipu AI), DeepSeek, and Qwen (Alibaba) are where the strongest open-source capability is being released right now — reasoning chains, MoE architectures, long-context, multilingual. A council built to maximise independence-of-judgment uses the strongest seats available, not the most familiar names. The panel will evolve as the open-source landscape does; if a Western open-source model returns to the frontier, it gets a seat.
Yes. You compose the panel from the available seats — any combination of US closed-source frontier (Claude, GPT, Gemini) and Chinese open-source frontier (Kimi, GLM, DeepSeek, Qwen). Three is a working council. Seven is a board.
The role becomes part of the seat's mandate. A model seated as Devil's Advocate is instructed to argue the strongest counter-position, even if it doesn't believe it. A model seated as Senior Expert is instructed to speak from best-practice authority. A model seated as Rookie is instructed to ask the questions a fresh perspective would ask. The role changes the deliberation; it isn't decoration.
Yes. The eight starter roles are the floor, not the ceiling. You can design custom roles — give them a name, a mandate (the instructions that shape how the seat behaves), and a description — and save them to your workspace's role library for reuse in every future council. Most operators end up with a handful of custom roles that mirror how their team thinks: "Our CFO's voice", "The regulator's lens", "Our most pessimistic engineer". Build the roles your specific deliberations need.
A council can share a blind spot — especially if the seats are too similar. That's why panel diversity (closed + open + different geographies + different architectures) matters more than panel size. Consensus across a diverse panel is strong evidence; consensus across three similar models is just three models saying the same thing twice. The transcript is there so you can audit the reasoning yourself.
Pay-as-you-go — no subscription and no per-seat fee. You top up a prepaid balance and draw it down per council. Each council is priced on the tokens it uses, weighted by the model each seat runs: a frontier closed-source seat costs more than an open-source one, and a seven-seat board costs more than a three-seat council. You see an estimated cost before you convene and an itemised receipt — tokens, models, and time — after the verdict. Early adopters start with $5 of founder credit.
Sent to each seated model's API for the duration of the query. Not retained on our servers beyond what's needed to deliver your report. Provider-side retention follows each provider's stated policy. Open-source seats route through hosted inference providers. Full privacy disclosure published at launch.
It means your question was answered by multiple independent models in defined roles, the points of agreement and disagreement are explicit, and the audit trail is yours. It's not a guarantee of correctness — nobody can guarantee that. It's a guarantee that more than one mind thought hard about your question from the roles you assigned, and the disagreements weren't smoothed away.
It depends on the complexity of the question — expect a verdict in roughly two to five minutes. The bottleneck is the slowest seated model, and we'd rather make you wait for considered output than rush it. Harder questions and larger councils take longer; simpler questions and smaller councils return faster — another reason to design the panel for the question.
Reserve your seat
We’re opening a small first cohort. Reserve your seat and we’ll bring you in — with $5 of founder credit to convene your first councils on us.
No card required. One email at most per week, only when there’s news worth your attention.