Checkobot

Are AI image detectors accurate?

Sometimes very, sometimes not, and a single number won't tell you which. Here is how to read accuracy claims, what to ask before you trust one, and where Checko's lens is still blurry.

Updated

The short answer

A good detector catches most images from the kinds of tools it has learned from, while rarely flagging clean, genuine photos. It does worse on generators it has never seen, on images that have been squashed and re-shared many times, and on content unlike anything in its training.

Every detector makes mistakes in both directions. The honest question isn't "is it accurate?" but "how often does it make each kind of mistake, on images like mine?"

Why one "accuracy" number means little

A detector can be wrong in two ways. It can miss an AI image, or it can flag a genuine photo. These trade off against each other through the threshold: the score above which the detector says AI.

Lower the threshold and it catches more AI images, but flags more genuine photos too. Raise it and genuine photos are safer, but more AI images slip through. A vendor can pick any point on that curve, so a single "accuracy" figure hides which mistake they chose to make fewer of.

It also depends on the mix. If only a small share of the images you check are AI, even a rare false flag can add up to more wrong flags than right ones. A test set that is half AI and half genuine won't show you that.

What changes the result

  • The dataset: which generators, and which genuine photos. Studio portraits, phone snapshots and old scans behave very differently.
  • The generator: a detector usually does best on tools it was trained on, and worst on ones released after it.
  • Compression and resizing: every re-save, screenshot and upload throws away detail, including the traces a detector looks for.
  • Faces versus whole images: a model built for face swaps and a model built for whole generated scenes are good at different things.
  • Full generations versus small edits: a whole generated scene is easier to catch than one changed object in a genuine photo.
  • Photos versus art: illustrations, digital paintings and CGI are hard for detectors trained mostly on photographs.

What to ask any vendor

Whoever makes the detector, including us, should be able to answer these:

  • What was it tested on? Is the dataset described, and was it kept out of training?
  • At which threshold or setting? Is that the one you'll actually get?
  • How often does it flag genuine photos, and photos like yours in particular?
  • What is the catch rate by category (face swaps, full generations, edits), rather than one blended number?
  • Which generators were held out of training, to show how it copes with new ones?
  • When was it measured, and on which model version?
  • What does it refuse to judge, and does it tell you why?

How Checkobot reports its results

Checkobot runs two detectors: whole-image analysis on every picture, and face analysis on the largest face. Each has its own threshold, set on an identity-selfie benchmark so that it rarely flags a genuine selfie. Photos from the wider web are harder, and get flagged by mistake more often at the same thresholds.

We test on frozen benchmarks that the model never saw in training, including face-swap tools held out of training entirely. We'll publish catch and false-flag rates, each with its dataset and thresholds, once our current test round is finished. Until then we'd rather say nothing than quote a number we haven't confirmed.

Every check records the model and threshold version it used, so a result can be traced back later.

Where detectors still struggle, ours included

  • New generators: anything released after a model was trained is a fresh test.
  • Heavy edits: strong filters, beauty retouching and repeated edits can push a genuine photo over the threshold, while a small AI edit inside a large photo can stay under it.
  • Compression: heavily compressed photos are flagged by mistake more often. Checkobot refuses JPEGs below roughly quality 50 rather than guess.
  • Screenshots and small images: they lose detail. Images under 256 pixels on the short side are refused.
  • Illustrations, art, CGI and doll faces: these can trip a face model that learned from photographs.
  • Talking-head avatars and lip-sync from new engines: our weakest area. Many are missed.

How to use a detector well

Treat the result as one signal. Combine it with the source, a reverse image search and the context. Check the best copy you can find, not a screenshot of a screenshot. When the result matters for a decision about a person, such as an account, a listing or a news story, have a person review it.

And read the reason, not just the verdict. Knowing which detector fired tells you whether to look at the face or the whole scene.

AI photo checker

Questions

How accurate is Checkobot?

We'll publish catch and false-flag rates, each with its dataset and thresholds, once our current test round is finished. We won't quote one blended accuracy number, for the reasons on this page.

Can a detector prove an image is real?

No. A No AI detected result means no signs were found at the threshold. Some AI images score under it, so it's evidence, not proof.

Why do different detectors give different answers?

They were trained on different images, use different thresholds and look for different things. Disagreement usually means the image is a hard case, and that the source and context matter more.

Is a higher score more certain?

Roughly, but a score isn't the chance that the image is AI. It's the model's output, compared against a threshold. That's why Checko shows a verdict and a reason rather than a percentage.