Przejdź do treści
METHODOLOGY

How we analyse messages

Version 1.0 ·
A diagnosis concerns a specific message, not a person. Below: where posts come from, how the result is produced and how we calculate the Clinic's data.

At a glance

  1. 1Selection

    The Gatekeeper reads new posts and selects those containing a claim to check.

  2. 2AI Council (Konsylium)

    Several models from different companies assess the same post independently.

  3. 3Evidence

    Claims are submitted to a search engine; assessing a factual claim requires a source.

  4. 4Publication

    The diagnosis is published automatically. A human can only withdraw it.

Post selection

We read posts from the past 3 days on the X accounts of politicians and parties we follow. The Gatekeeper assigns each an initial score - this is not spin strength (0-100).

skippedpending decisionautomatic queue
04075100

This is a sample of selected statements, not the entire debate. More diagnoses of one side do not mean that side uses spin more often. No diagnosis does not mean 'no spin'.

Scope of analysis

We examine

  • post text
  • photos and graphics: descriptions and text within images
  • titles and descriptions of linked pages

We do not examine in posts

  • videos attached to the post
  • anything the model could not read - we describe this in the diagnosis's limitations

Interviews (in Polish) follow a separate process: transcription and separate assessments of the guest and interviewer.

Verdicts and spin strength (0-100)

  • Spin

    Clear persuasion techniques or a manipulative presentation of the content.

  • Partial spin

    These elements are present, but do not define the whole message.

  • No spin

    No significant spin in the material examined. This is not a certificate of truth.

  • Unable to assess

    The available data does not allow a conclusive assessment.

55/100

Spin strength (0-100) measures the intensity of the techniques identified - not a percentage of falsehood or guilt. The score is the median of the models' assessments; for 'no spin' it is capped at 20.

The verdict is the Council's middle assessment; in a tie we choose the milder one. 'Unable to assess' votes are excluded from the median. We do not add fictional votes where only one model assessed the material.

Model agreement

2/3the same verdict

This is how many models gave the same verdict as the diagnosis. 'Unable to assess' counts in the denominator; a missing response does not. We also show the spread of strength scores.

Models can make the same mistake, so agreement is not evidence. The target is at least 4 responses from 3 companies, including a Polish model; when service limits apply: 3 responses with missing contributions described. Current membership

Claim statuses

  • Supportedsources support the claim
  • Contradicted by sourcessources contradict the claim
  • Misleadinga significant omission, distortion or unwarranted conclusion
  • Unverifiedno source is available or the claim cannot be checked

A factual claim's status requires a source from search results - without one, the claim remains unverified. We merge similar claims before counting. In summaries, the 'opinions' group currently also includes facts that could not be verified.

Technique families

21 categories in three families, plus 'Other'. A technique appears in a diagnosis only with a quotation from the material and when at least 2 Council members identify it (with 1-2 conclusive votes, one identification is sufficient). At most 6 techniques per diagnosis.

Data and reasoning10

  • Number without a reference point
  • Selective data
  • Omission of context
  • Misrepresentation of facts
  • Unsupported claim
  • False causality
  • Overgeneralisation
  • False analogy and association
  • False dilemma
  • Appeal to authority

Emotion and presentation6

  • Appeal to emotion
  • Fearmongering
  • Exaggeration
  • Labelling
  • Us versus them
  • Insinuation and implication

Dispute and responsibility5

  • Personal attack
  • Attributing intentions
  • Misrepresenting another person's position
  • Changing the subject
  • Claiming credit

Alignment with SemEval 2023

SemEval 2023 Task 3 is an international research task whose subtask 3 covers detecting 23 persuasion techniques in 6 groups, including in Polish texts. The mapping below is our own and indicative. Categories about the reliability of data, facts and reasoning go beyond this catalogue of linguistic persuasion. The mapping does not mean that our system has been validated in SemEval.

Our categorySemEval equivalent
FearmongeringAppeal to Fear/Prejudice
Claiming creditno equivalent
Attributing intentionsCasting Doubt
Personal attackQuestioning the Reputation
False dilemmaFalse Dilemma/No Choice
Straw manStrawman
False analogy and associationGuilt by Association
Misrepresentation of factsno equivalent
Insinuation and implicationCasting Doubt
False causalityCausal Oversimplification
Number without a reference pointno equivalent
Selective datano equivalent
Omission of contextno equivalent
Overgeneralisationno equivalent
LabellingName Calling/Labeling
Us versus themFlag Waving
Changing the subjectRed Herring; Whataboutism
Appeal to authorityAppeal to Authority
Unsupported claimno equivalent
ExaggerationExaggeration/Minimisation
Appeal to emotionLoaded Language; Appeal to Values
Otherno equivalent
Loaded words (separate module)Loaded Language

SemEval techniques we do not detect separately:

  • Appeal to Hypocrisy
  • Appeal to Popularity
  • Consequential Oversimplification
  • Slogans
  • Conversation Killer
  • Appeal to Time
  • Obfuscation/Vagueness/Confusion
  • Repetition

Loaded words

  • fear and threat
  • anger and outrage
  • contempt and ridicule
  • pride and community
  • compassion and harm

We first use the models' findings (the word must appear in the post), falling back to a Polish dictionary if none are available. We count distinct words and phrases, not every occurrence. A loaded word alone does not establish spin.

How we calculate the data

Charts cover published diagnoses that have not been hidden. Comparisons cover 30 days, based on the diagnosis date.

MeasureBasis
Verdicts and strengthnumber of diagnoses in the group and period; arithmetic mean of diagnosis scores
Techniques and familiesnumber of diagnoses containing a category (once per diagnosis), not the total number of techniques
Claimsdistinct claims after merging duplicates; unverified claims counted separately
Loaded wordsaverage number of distinct words or phrases per diagnosis
Unanimitydiagnoses with at least 2 responses in which all votes agree
Likesonly posts with an available count; missing data is not zero

Below 10 observations we flag a small sample. This is a caution, not a statistical test.

Limitations

  • AI may misinterpret a quotation, context, irony, a factual claim or a technique category.
  • Models from different companies may share training data and biases.
  • Search does not cover all knowledge; sources may be out of date.
  • Service limits affect Council membership and the scope of checks.
  • Conclusions concern this sample of messages - we do not assess people's truthfulness.

Errors and withdrawal

Spotted an error? Send a link to the diagnosis, a description and sources to kontakt@spin.clinic. The author of the statement can submit a response through the same channel.

The operator can withdraw an entire diagnosis - we record the date and reason. Its content, verdict and strength are not edited manually. Every withdrawal, legal removal and author response is listed in the public register of corrections and responses (in Polish). Council Charter