AI image detection for trust and safety
AI-generated and face-swapped images now turn up in profile photos, listings, reviews and abuse reports. This guide covers where a detector fits in a moderation workflow, and how to use its output without hurting genuine users.
Updated
Where AI images show up
- Profile photos: generated faces and face swaps on dating, social and professional accounts, often in coordinated batches.
- Impersonation: swapped or edited photos of public figures, executives and ordinary users, used to build trust before a scam.
- Marketplace and rental listings: AI-made product, property and vehicle photos for items that may not exist.
- Reviews, refunds and claims: generated or edited photos offered as evidence of a product, a delivery or damage.
- Harassment and intimate imagery: face swaps and AI edits of identifiable people made without their consent.
- News and civic content: generated images of events that didn't happen, posted to drive engagement.
Where detection fits in the pipeline
Most teams use image detection at a few points rather than everywhere:
- At upload, for high-risk surfaces such as profile photos, verification-adjacent flows and new listings.
- On report, to give reviewers a second opinion when a user flags an image.
- In sweeps, by sampling or backfilling content from accounts already under suspicion.
- In escalations, as one input to an investigation alongside account, device and behaviour signals.
Route flags to review, not to automatic bans
A detector's verdict is a risk signal, not proof. Every detector flags some genuine photos, and the people behind those photos are your users. Treating every flag as a violation turns a detector's error rate into a stream of wrongful bans and appeals.
We recommend routing AI verdicts to human review or a step-up action, for example asking for another photo or a verification step, limiting reach until a reviewer decides, or adding a label. Silent permanent bans on the strength of a detector alone are the pattern to avoid.
Give reviewers the reason as well as the verdict. Knowing whether the face analysis or the whole-image analysis fired, and where the face was, tells them what to look at.
What the Checkobot API returns
The API is a REST endpoint: POST /v1/check with a Bearer API key, sending a file, an image URL or base64. Each response includes:
- The verdict: AI or No AI detected.
- Which detectors fired: whole-image analysis, face analysis on the largest face, or both. Images that still carry signed Content Credentials from trusted AI tools recording AI generation or editing are flagged too.
- A one-line, human-readable reason.
- Both detector scores and the box of the face that was checked.
- The model and bands version used for that check, so any decision can be audited later.
- A clear refusal, with a reason code, for images too small or too compressed to judge, instead of a guess.
- A batch endpoint for several images per call is available on higher plans.
Data handling
By default, images sent to the API are processed in memory and not stored. They are kept only if you turn on Store images for a key or send store=true with a request. Each check keeps a SHA-256 hash of the file, the verdict and the scores, so results can be audited without keeping the image.
We suggest the same discipline on your side: log hashes, verdicts and versions rather than image bytes or image URLs, and make sure your own privacy notice covers automated checks on uploaded images.
Test it on your own traffic first
Published benchmark numbers, ours included, are measured on someone else's images. Your users' photos have their own mix of phones, filters, lighting and compression. Before going live:
- Run a sample of known-genuine uploads and measure how often they are flagged.
- Run a set of known AI and swapped images from your own abuse queue and measure how many are caught.
- Break both down by category: faces versus no faces, photos versus illustrations, heavily filtered versus clean.
- Decide which actions each verdict triggers, and keep review capacity in line with the flag volume you measured.
- Re-check when the model or bands version in the response changes.
Scope and known limits
- Images only. Video can be checked as individual still frames; there is no real-time or clip-level video decision.
- Not for documents, IDs or receipts, and not an audio or liveness check (printed photos or screen replays held up to a camera).
- It does not identify which tool generated an image.
- Talking-head and lip-sync frames from new avatar engines are the weakest area: many are missed.
- Illustrated, cartoon, CGI or doll faces can trip the face analysis, which learned from photographs.
- Heavily distorted or very low-quality photos are flagged by mistake more often than clean ones.
- Measured catch and false-flag rates, each with its dataset and thresholds, will be published once the current test round is finished.
Policy and user communication
Write policy around behaviour, not tools. "Impersonating another person" or "misrepresenting an item" is enforceable and fair. "Used AI" on its own often isn't, since plenty of legitimate photos are edited with AI features.
When you act on a flag, tell the user what was flagged and how to appeal: "this photo was held for review" rather than "you are fake". Keep the check ID, model and bands version with each decision so appeals can be reviewed against the exact result.
Questions
Can we block uploads automatically on an AI verdict?
We recommend against it. Use the verdict to route images to review, ask for another photo, or limit reach pending a decision. Automatic bans turn every false flag into a wronged user.
Do you store our users' images?
Not by default. Images are processed in memory unless you turn on Store images for a key or send store=true. Each check keeps a hash of the file, the verdict and the scores.
Can we tune the thresholds?
Not yet. Every response includes the scores and the bands version, so you can see how close a result was to the threshold.
Does it detect forged IDs or documents?
No. It is built for photos of people and scenes. Document and ID forgery are outside its scope.