How we calculate the score — and what we can't do
Methodological honesty is the foundation of this product: we would rather say “we don't know” than guess. Below is a full description of our approach.
The overriding principle: a band, not a verdict
We never answer with a binary “fake / not fake”. The result is a calibrated probability band with the signals it is based on spelled out. We only assign a risk label (“probably fake”) when at least two independent detectors agree — a single vote is never enough.
Images: independent votes in an ensemble
Every image is scored in parallel by: the commercial detector Sightengine (AI generation + deepfake), our self-hosted model (Community Forensics, ViT — trained on material from thousands of generators, running on our own infrastructure), and a provenance layer: C2PA/EXIF/XMP metadata and generator-specific markers (e.g. Midjourney). Before we compute anything, we also check whether an identical file has already been analyzed (deduplication by cryptographic hash).
The metadata asymmetry rule
A generator's C2PA signature is strong evidence of synthetic origin. But it doesn't work the other way round: missing metadata proves nothing — social platforms strip it as standard. That's why we report “no provenance data” neutrally and never treat it as an argument for authenticity.
Text and links: claims, debunks, narratives
We extract verifiable claims from the text and check them against three sources: a database of verified fact-checker debunks (Google Fact Check, which indexes Polish newsrooms among others), our index of over 19,000 documented pro-Kremlin disinformation narratives (EUvsDisinfo, updated weekly; matching works across languages — a Russian narrative will match Polish text), and live web search with cited sources. The verdict for content is verbal (e.g. “manipulation”, “cannot be confirmed”) — deliberately without percentages, because the truthfulness of claims is not a continuous quantity.
What this method canNOT do
We say it openly: (1) according to independent research, the accuracy of the best image detectors on “in-the-wild” material is around 78–82%, and the newest generators can be detected less reliably; (2) screenshots and repeated recompression erase the traces — when material is heavily processed, we lower the confidence of the verdict instead of pretending we have it; (3) detectors assess generation traces, not context — a real photo used in a false context requires source verification, not pixel analysis.
Calibration and our benchmark
We calibrate vote weights and thresholds on our own Polish test set (fresh generations “passed through” social platform compression + real agency photos), which we keep expanding. In line with the zero made-up numbers principle — we will publish benchmark results once the set reaches a representative size, together with a description of the sample.
Appeals channel
Every result has a “Report an incorrect assessment” button. Reports come straight to us, and confirmed errors feed recalibration — we store the raw detector responses precisely so we can be held accountable for them.
Legal context
From 2 August 2026, Article 50 of the EU AI Act requires deepfakes to be labelled. czytofejk.pl helps check material that lacks such labelling — but our result is a technical assessment, not a legal determination (details in the terms of service).