Przejdź do treści
AI COUNCIL (KONSYLIUM)

Several AI models,
one diagnosis

Version 1.1 ·
Several AI models from different companies independently assess selected posts by politicians. We compare their votes, look for sources and prepare a joint diagnosis - with transparent membership and limitations. Models can make mistakes.
The animation shows how the service works: cards, the timeline, threads, the Clinic and future phases. The material and people are illustrative. (in Polish) Open full screen (in Polish) ↗

From post to diagnosis

  1. 1Selection

    The Gatekeeper reads new posts by politicians and selects those containing a claim to check.

  2. 2Independent votes

    Several models from different companies assess the post separately. None sees the others' responses.

  3. 3Combining votes

    Fixed rules combine votes into one verdict, a spin strength score (0-100) and a list of techniques.

  4. 4Sources and checks

    Claims are submitted to a search engine, and the Laboratory adds supporting checks. Sources must come from the search results.

  5. 5Explanation

    The Chair writes the diagnosis from the findings, the Reviewer checks consistency and the Linguist corrects the Polish.

  6. 6Publication

    The diagnosis is published automatically, with model votes and limitations. No one edits its content. Selected results also appear as summaries and videos on social media - always with a link to the full analysis.

Who does what

The Council works like a medical council: several independent specialists, separate examinations and a single report of the result. Below are the default role assignments - if a model does not respond, another takes over. Each diagnosis identifies the models that actually performed the work.

  • Council members

    gpt-oss · Qwen · Nemotron · Gemini · Bielik · PLLuM · Llama · Mistral

    Each independently provides a verdict, spin strength (0-100), techniques with verbatim quotations and claims to check.

  • Fact-checking

    Gemini with Google Search

    Looks for sources for each claim. Without a source, the claim remains unverified.

  • Consultant

    Claude (Anthropic) · paid

    More thorough fact-checking when models are divided (agreement below 2/3) or spin is strong (70/100 or above).

  • Laboratory

    HerBERT · Google Fact Check · GUS · Firecrawl

    Supporting checks: sentiment, earlier fact-checks, GUS data and whether a quotation appears in its source. These do not change the verdict.

  • Chair

    Gemini (other models as fallbacks)

    Writes the explanation solely from the Council's assessments and evidence. Cannot change the verdict or strength, or add a technique.

  • Reviewer

    Nemotron (other models as fallbacks)

    Checks whether the text matches the assessments and Charter. If issues are raised, the Chair revises the diagnosis once; the absence of a review does not prevent publication.

  • Linguist

    Bielik (PLLuM, Qwen and Gemini as fallbacks)

    Corrects only the Polish. If a correction substantially changes the text's length or is not in Polish, the Chair's version is retained.

How we combine votes

  • Model APartial spin30
  • Model BPartial spin40
  • Model CPartial spin60
  • Model DSpin80

Result

Partial spin

50/100

3/4 the same verdict · range 30-80

Illustrative calculation, not a real diagnosis.

  • Verdict and strength are the middle assessments (median). One outlying model does not determine the result; in a tie we choose the milder assessment.
  • A technique requires a verbatim quotation from the material examined. With at least three conclusive votes, at least two models must identify it; with one or two such votes, one identification is sufficient.
  • Model agreement informs the reader, but does not prove truth: models can make the same mistake.

Current membership

  • 4responses - the target for each post
  • 3+different companies
  • PLPolish model selected first

For each post we select models from different companies, starting with a Polish one. If one does not respond, another joins. With fewer than 3 responses, the diagnosis is not published; incomplete membership is described in its limitations. Below are the models configured to work in the Council.

Loading Council membership…

Council recruitment

Membership is open. Every night, the Council Recruiter reviews catalogues of free models and selects at most one candidate - priority goes to a Polish model and a company not yet represented. The candidate takes an exam: it assesses the same politicians' posts as the Council. The criteria are agreement with the final diagnosis on verdict and spin strength (0-100), quotations supporting techniques and quality of Polish. Current members vote, and an admitted model must accept the Charter; it receives roles only when its exam results justify them.

A member that has not responded for three days (for example, because its provider has removed it) is suspended and returns automatically when it works again. We record all decisions below. Until 7 October, the Recruiter operates in trial mode: admissions are recommendations.

Tools

ToolPurposeIn the system
Council, vote aggregation, Chair, Reviewer, Linguistcore of every diagnosissupported
Gemini with Google Searchsources for claimssupported
Claude with searchconsultation for disagreement and strong spin, when the budget allowssupported
Dictionary of loaded words and 21 technique categoriesa shared vocabulary for diagnosessupported
HerBERT (Hugging Face)sentiment and hate speech - a supporting signalsupported
Google Fact Check Toolsearlier checks by other newsroomssupported
GUS BDLunemployment and wages in Polandsupported
Firecrawlwhether the quotation really appears in the source (up to 3 per diagnosis)supported
Wayback Machinepublic copy of the post, created in the backgroundsupported
Our own technique detector (XLM-RoBERTa)on our own machineplanned
Eurostat, NBP, legal registersbroader checking of figures and legislationplanned

A tool supported by the system operates when access and unused quota are available. A diagnosis's results show whether it was used for that analysis.

Limits of analysis

  • A diagnosis requires at least 3 responses; membership may be limited and some checks omitted.
  • No source means no verification, not falsehood.
  • The absence of a review or a negative review does not automatically prevent publication.
  • Video attached to a post is not analysed; interviews follow a separate process.
  • Model agreement does not prove truth - models can make the same mistake.

Rules and errors

The operator can withdraw a diagnosis from public view, but does not change its content, verdict or strength.