How we analyse messages
At a glance
- 1Selection
The Gatekeeper reads new posts and selects those containing a claim to check.
- 2AI Council (Konsylium)
Several models from different companies assess the same post independently.
- 3Evidence
Claims are submitted to a search engine; assessing a factual claim requires a source.
- 4Publication
The diagnosis is published automatically. A human can only withdraw it.
Post selection
We read posts from the past 3 days on the X accounts of politicians and parties we follow. The Gatekeeper assigns each an initial score - this is not spin strength (0-100).
This is a sample of selected statements, not the entire debate. More diagnoses of one side do not mean that side uses spin more often. No diagnosis does not mean 'no spin'.
Scope of analysis
We examine
- post text
- photos and graphics: descriptions and text within images
- titles and descriptions of linked pages
We do not examine in posts
- videos attached to the post
- anything the model could not read - we describe this in the diagnosis's limitations
Interviews (in Polish) follow a separate process: transcription and separate assessments of the guest and interviewer.
Verdicts and spin strength (0-100)
- Spin
Clear persuasion techniques or a manipulative presentation of the content.
- Partial spin
These elements are present, but do not define the whole message.
- No spin
No significant spin in the material examined. This is not a certificate of truth.
- Unable to assess
The available data does not allow a conclusive assessment.
55/100
The verdict is the Council's middle assessment; in a tie we choose the milder one. 'Unable to assess' votes are excluded from the median. We do not add fictional votes where only one model assessed the material.
Model agreement
2/3the same verdict
This is how many models gave the same verdict as the diagnosis. 'Unable to assess' counts in the denominator; a missing response does not. We also show the spread of strength scores.
Models can make the same mistake, so agreement is not evidence. The target is at least 4 responses from 3 companies, including a Polish model; when service limits apply: 3 responses with missing contributions described. Current membership
Claim statuses
- Supportedsources support the claim
- Contradicted by sourcessources contradict the claim
- Misleadinga significant omission, distortion or unwarranted conclusion
- Unverifiedno source is available or the claim cannot be checked
A factual claim's status requires a source from search results - without one, the claim remains unverified. We merge similar claims before counting. In summaries, the 'opinions' group currently also includes facts that could not be verified.
Technique families
21 categories in three families, plus 'Other'. A technique appears in a diagnosis only with a quotation from the material and when at least 2 Council members identify it (with 1-2 conclusive votes, one identification is sufficient). At most 6 techniques per diagnosis.
Data and reasoning10
- Number without a reference point
- Selective data
- Omission of context
- Misrepresentation of facts
- Unsupported claim
- False causality
- Overgeneralisation
- False analogy and association
- False dilemma
- Appeal to authority
Emotion and presentation6
- Appeal to emotion
- Fearmongering
- Exaggeration
- Labelling
- Us versus them
- Insinuation and implication
Dispute and responsibility5
- Personal attack
- Attributing intentions
- Misrepresenting another person's position
- Changing the subject
- Claiming credit
Alignment with SemEval 2023
SemEval 2023 Task 3 is an international research task whose subtask 3 covers detecting 23 persuasion techniques in 6 groups, including in Polish texts. The mapping below is our own and indicative. Categories about the reliability of data, facts and reasoning go beyond this catalogue of linguistic persuasion. The mapping does not mean that our system has been validated in SemEval.
| Our category | SemEval equivalent |
|---|---|
| Fearmongering | Appeal to Fear/Prejudice |
| Claiming credit | no equivalent |
| Attributing intentions | Casting Doubt |
| Personal attack | Questioning the Reputation |
| False dilemma | False Dilemma/No Choice |
| Straw man | Strawman |
| False analogy and association | Guilt by Association |
| Misrepresentation of facts | no equivalent |
| Insinuation and implication | Casting Doubt |
| False causality | Causal Oversimplification |
| Number without a reference point | no equivalent |
| Selective data | no equivalent |
| Omission of context | no equivalent |
| Overgeneralisation | no equivalent |
| Labelling | Name Calling/Labeling |
| Us versus them | Flag Waving |
| Changing the subject | Red Herring; Whataboutism |
| Appeal to authority | Appeal to Authority |
| Unsupported claim | no equivalent |
| Exaggeration | Exaggeration/Minimisation |
| Appeal to emotion | Loaded Language; Appeal to Values |
| Other | no equivalent |
| Loaded words (separate module) | Loaded Language |
SemEval techniques we do not detect separately:
- Appeal to Hypocrisy
- Appeal to Popularity
- Consequential Oversimplification
- Slogans
- Conversation Killer
- Appeal to Time
- Obfuscation/Vagueness/Confusion
- Repetition
Loaded words
- fear and threat
- anger and outrage
- contempt and ridicule
- pride and community
- compassion and harm
We first use the models' findings (the word must appear in the post), falling back to a Polish dictionary if none are available. We count distinct words and phrases, not every occurrence. A loaded word alone does not establish spin.
How we calculate the data
Charts cover published diagnoses that have not been hidden. Comparisons cover 30 days, based on the diagnosis date.
| Measure | Basis |
|---|---|
| Verdicts and strength | number of diagnoses in the group and period; arithmetic mean of diagnosis scores |
| Techniques and families | number of diagnoses containing a category (once per diagnosis), not the total number of techniques |
| Claims | distinct claims after merging duplicates; unverified claims counted separately |
| Loaded words | average number of distinct words or phrases per diagnosis |
| Unanimity | diagnoses with at least 2 responses in which all votes agree |
| Likes | only posts with an available count; missing data is not zero |
Below 10 observations we flag a small sample. This is a caution, not a statistical test.
Limitations
- AI may misinterpret a quotation, context, irony, a factual claim or a technique category.
- Models from different companies may share training data and biases.
- Search does not cover all knowledge; sources may be out of date.
- Service limits affect Council membership and the scope of checks.
- Conclusions concern this sample of messages - we do not assess people's truthfulness.
Errors and withdrawal
Spotted an error? Send a link to the diagnosis, a description and sources to kontakt@spin.clinic. The author of the statement can submit a response through the same channel.
The operator can withdraw an entire diagnosis - we record the date and reason. Its content, verdict and strength are not edited manually. Every withdrawal, legal removal and author response is listed in the public register of corrections and responses (in Polish). Council Charter