# Is This Passport Real? Walking Through the AI Markers in a Tampered ID

> How to tell if a passport or ID is real. A walkthrough of the seven AI markers a fake document checker finds in a tampered identity document - MRZ mismatches, per-field compression, pasted portraits, and generative fingerprints.

*Published 2026-08-22 · 8 min read · TamperCheck.ai*

Canonical: https://tampercheck.ai/blog/is-this-passport-real-ai-markers-tampered-id

---
Someone uploads a passport to your onboarding flow and the question lands on a human: **is this document real?**

The honest answer, in 2026, is that you cannot tell by looking. Not because reviewers are careless, but because the thing you are looking at is a flat image on a screen - no paper, no UV lamp, no tilt to catch the hologram. Every physical security feature a passport was designed around has already been thrown away by the time the file reaches you.

What survives is different, and better: the document still has to be *internally consistent*, and the image still has to carry the fingerprints of a real camera. Tampered IDs fail both tests, reliably. This post walks the markers, using an anonymised composite specimen - fictional issuer, fictional data, no real document reproduced.

- **7** — independent markers that expose a tampered ID in a flat image
- **6 of 7** — are invisible to visual review, however careful
- **~1 min** — for an automated forensic verdict across all layers

![Anonymised specimen passport data page annotated with seven AI markers: field-type mismatch, MRZ check digit failure, date arithmetic error, repeated entity name, per-field compression mismatch, pasted portrait seam, and generative fingerprint](https://tampercheck.ai/images/blog/tampered-id-ai-markers.png?v=2)

*An anonymised composite specimen. Fictional issuer, fictional data - the markers are what matter, not the document.*

## The specimen: what a tampered passport data page actually looks like

The most common fake in circulation is not a counterfeit booklet. It is a **genuine passport data page with fields swapped** - a real scan, edited, then re-flattened into a single JPEG so the edit history disappears.

Our specimen follows that recipe. It is a plausible data page: correct layout, correct bilingual field labels, correct fonts, a portrait, a signature, a machine-readable zone at the bottom, a second page with parents' names and an address. To the eye it is unremarkable. Nine out of ten reviewers would pass it.

It fails on seven separate counts.

> **INFO:** Worth naming the trap up front: **looking harder does not help.** The markers below are arithmetic, statistics, and physics. None of them are aesthetics. A reviewer squinting at font weights is playing a game the forger already won.

## Marker 1: A field contains the wrong *kind* of value

The single loudest signal in our specimen, and the one a human *can* catch if they read rather than skim.

The **Nationality** field contains a person's full name.

That is what happens when an editor copies a text block from one part of the document - here, the father's name from page two - and drops it into a field on page one. The pixels are genuine. The font matches perfectly, because it *is* the document's own font. The alignment is clean. Nothing about the rendering is wrong. The value is simply the wrong species of data.

Automated checks catch this as a **semantic type violation**: nationality is an enumerated value from a fixed list of country adjectives, not a free-text personal name. Same class of error as a date in a currency field or an address in a document number.

This is also the marker that survives every downstream trick. Re-compress the image, add noise, print and rescan it - the field still says the wrong thing.

## Marker 2: The MRZ doesn't reconcile with the visible fields

Those two lines of chevron-padded text at the bottom of a passport are not decoration. The **machine-readable zone** is specified by ICAO Document 9303, and every field in it carries a computed check digit: document number, date of birth, expiry, and a composite digit over the whole lot.

Two things must hold:

1. **Each check digit must validate** against its own field, using a fixed 7-3-1 weighting.
2. **The MRZ must agree with the visual zone** - same name, same document number, same dates.

Forgers edit what they can see. The MRZ is dense, unreadable at a glance, and easy to forget. In our specimen the MRZ date-of-birth field is malformed - the wrong number of digits, so the checksum cannot even be computed - and the encoded name ordering doesn't match the visible surname and given-name split.

A failed ICAO checksum is not a probability. It is arithmetic. It is the closest thing document forensics has to proof.

> **WARNING:** A tampered ID that keeps the *original* MRZ intact is arguably worse than one with a broken MRZ: the machine-readable data now describes a different person than the printed page. Any system that trusts only the MRZ - or only the OCR of the visible zone - reads a clean document. Only cross-checking the two catches it.

## Marker 3: Dates that don't survive arithmetic

Identity documents are full of dates that constrain each other:

- Issue date must precede expiry date.
- The validity span must match the issuer's standard term (ten years for most adult passports, five for children).
- Date of birth must be consistent with the age implied elsewhere, and with the MRZ encoding.
- The document must not have been issued before the holder was born - which sounds absurd until you see how often edited dates produce exactly that.

Our specimen's visible date of birth and MRZ date of birth disagree. That is a two-field contradiction inside one document, and no amount of image quality hides it.

## Marker 4: The same name appears in two incompatible roles

On our specimen, the holder's own name appears again in the **spouse** field. Elsewhere, the father's name appears in both its own field and - per Marker 1 - in nationality.

Sloppy copy-paste leaves this signature constantly. Cross-field **entity resolution** catches it: build the graph of people named in the document, then check whether the relationships are possible. A person cannot be their own spouse. A guardian's name cannot be a nationality.

The same logic finds duplicated addresses that disagree on their own content - our specimen carries two address fields, one a truncated fragment of the other with a different street entirely.

## Marker 5: Individual fields carry the wrong compression and sharpness

Now we leave what a human could conceivably read, and enter the part that requires forensics.

When one field is edited inside an otherwise genuine scan, that patch of pixels has a different history from everything around it. It was decoded, redrawn, and re-encoded, while its neighbours were only ever encoded once. That difference persists even after the whole image is flattened and saved again.

Measured **per field**, rather than across the whole page:

- **Recompression level** - an edited region shows a different JPEG quantisation history from the surrounding page. Whole-image error level analysis smears this out; field-level analysis does not.
- **Edge sharpness** - re-rendered text is often *too* clean. Genuine print, captured through a lens, has a characteristic softness and asymmetry at glyph edges. Retyped text is crisper than the ink around it.
- **Black floor** - the darkest pixel value inside a genuine printed glyph settles at a level set by ink, paper and exposure. Synthetic text tends to sit at a different floor, frequently a purer black than the document can physically produce.

Any single field disagreeing with the page consensus on all three is a strong local edit signal - and it works on smoothed, carefully feathered edits that classic whole-image ELA misses entirely.

## Marker 6: The portrait sits *on top* of the card, not in it

Photo substitution is the highest-value edit in ID fraud, so it gets special treatment.

A genuine data page prints the portrait *into* the card stock, underneath the issuer's security overlay - the guilloché lattice, rainbow print, and ghost portrait that run continuously across the whole surface, photo included. A substituted portrait, whether digitally composited or physically printed and pasted before scanning, breaks that continuity:

- **A seam.** A long, straight, axis-aligned boundary where the photo meets the card - sharper than any printed portrait edge.
- **A shadow.** A physically pasted photo sits proud of the surface and casts a thin drop-shadow hugging one or two edges under raking light.
- **A local mismatch.** The inserted portrait carries its own noise floor, sharpness and colour temperature, all different from the card around it.
- **A missing overprint.** The security pattern stops at the photo boundary instead of running across it. The ghost portrait - the faint secondary image beside the main one - no longer matches the main portrait.

That last one is decisive when it's present. Two portraits of two different people, on the same card, is not something a genuine issuing authority produces.

## Marker 7: Generative fingerprints - the document was never photographed at all

The newest category is the fully synthetic ID: no original document, no camera, no scanner. A generative model produced the whole page.

These fail differently. There is no *edit* to localise, because nothing is authentic. What gives them away is the absence of physical capture:

- **Frequency-domain signatures.** Upsampling layers in diffusion and GAN architectures leave characteristic periodic peaks and an anomalously steep high-frequency roll-off in the power spectrum. Real optics don't do that.
- **Unnatural texture uniformity.** Micro-texture across a generated page is too homogeneous; genuine paper and print have local disruption everywhere.
- **Cross-channel noise correlation.** A camera sensor produces noise independently per colour channel. Generators produce noise that correlates across channels - a signature that survives resizing and moderate recompression.
- **No colour filter array trace.** Every real camera image carries demosaicing artefacts from its Bayer sensor. Synthetic images have none. Neither do screenshots and re-renders, which is why this is read as one signal among many rather than a verdict on its own.
- **Visible generator watermarks.** Some consumer tools still composite a small sparkle glyph into a corner. It is a free catch when it's there, and absent far more often than not.
- **Provenance credentials.** A valid [C2PA](https://c2pa.org/) manifest is a strong authenticity signal; a stripped or broken one is suspicious; none at all is simply neutral, because most images still have none.

> **TIP:** Generative fingerprints are the one marker family that produces false positives if used alone. Scanner optics mask the same artefacts, so a legitimately scanned document can look "too smooth" for the wrong reasons. Treat this layer as evidence to be weighed against the other six, never as a standalone verdict.

## Why OCR-based verification reads this document as clean

Most identity verification stacks run OCR, extract the fields, and check them against a form or a database. That pipeline reads our specimen and reports success: every field parsed, every character legible, high confidence throughout.

OCR answers *what does this document say*. It has no opinion on *whether the pixels were ever genuine*. Those are different questions, and tampering lives entirely in the second one. We've written that distinction up in full in [document tampering detection vs OCR](https://tampercheck.ai/blog/document-tampering-detection-vs-ocr).

The same gap explains why a liveness selfie doesn't close it either - a real face held up next to a doctored document passes a liveness check comfortably. See [liveness detection vs document forensics](https://tampercheck.ai/blog/liveness-detection-vs-document-forensics) for where each one actually helps.

## How do immigration officials detect fake documents?

At a physical border, officers have advantages nobody working from uploads has: the booklet in hand, UV and IR illumination, magnification for microprint, tactile checks on intaglio printing, and chip reads over NFC that cryptographically confirm the data page against the issuing authority's signature.

Working from a submitted file - a visa application, a remote onboarding flow, a rental or employment check - none of that is available. What replaces it is exactly the seven-marker stack above: internal consistency (Markers 1-4), local edit forensics (Markers 5-6), and capture provenance (Marker 7). It is a different toolkit for a different medium, and on flat digital submissions it substantially outperforms visual review, because visual review has almost nothing left to work with.

## The marker table

Eight rows for seven markers: Marker 2 splits into two independently checkable signals, because a checksum failure and a MRZ-to-visual mismatch are different findings with different weight.

| Marker | What it proves | Confidence |
| --- | --- | --- |
| Wrong value type in a field | Content was moved between fields | Very high |
| MRZ checksum failure | Encoded data was altered | Definitive |
| MRZ vs visual zone mismatch | One of the two zones was edited | Very high |
| Date arithmetic contradiction | Dates were altered | High |
| Entity name in an impossible role | Copy-paste editing | High |
| Per-field compression / sharpness anomaly | Localised pixel edit | High |
| Portrait seam or missing overprint | Photo substitution | Very high |
| Generative frequency fingerprint | Image was synthesised, not captured | Medium, corroborating |

Two or more high-confidence markers is a "likely tampered" verdict. One medium marker on its own is a review recommendation, not an accusation - which is the distinction that keeps false positives from becoming customer-facing rejections.

**Answer 'is this document real?' in about a minute** — Upload a passport, ID card, or PDF and get a plain-English forensic verdict naming the exact markers it triggered. $5 in free credits, no contract. (https://tampercheck.ai)

An automated [document fraud detection](https://tampercheck.ai/document-fraud-detection) run applies all seven marker families to every submission and returns which ones fired and why - so a reviewer reads a short explanation instead of squinting at a JPEG. For teams wiring this into an onboarding flow, the [AI KYC](https://tampercheck.ai/ai-kyc) stack and the [document verification API developer guide](https://tampercheck.ai/blog/document-verification-api-developer-guide) cover the integration side.

## FAQ

### Is this document real? How can I check?

You cannot answer it reliably by looking - a flat image has already discarded every physical security feature. What you can check is internal consistency and capture provenance: whether the MRZ check digits validate and agree with the printed fields, whether the dates are arithmetically possible, whether any field's compression and sharpness disagree with the rest of the page, whether the portrait breaks the card's security overprint, and whether the image carries genuine camera signatures. An automated fake document checker runs all of these in about a minute and tells you which ones failed.

### How can you tell if a passport is fake?

The reliable tells on a digital submission are the seven markers in this post: a field containing the wrong kind of value, an MRZ that fails its ICAO 9303 check digits or contradicts the visible zone, dates that don't reconcile, the same person's name in two incompatible roles, individual fields with a different compression and sharpness history from the page around them, a portrait that sits on top of the card instead of under its security overprint, and generative fingerprints indicating the image was never photographed. Font and layout inspection - the traditional advice - is the weakest of these, because modern forgeries reuse the document's own fonts.

### Can AI detect a fake ID or passport?

Yes, with the important caveat that it detects specific evidence rather than rendering an infallible judgement. Automated forensics is far better than humans at the things humans are bad at: checksum arithmetic, per-field statistical comparison, frequency-domain analysis, and cross-referencing every field against every other field. It is weaker on physical features that never made it into the image - security threads, watermarks, intaglio feel. On flat digital submissions, which is where nearly all remote fraud happens, the arithmetic and statistical layers do most of the work.

### How do immigration officials detect fake documents?

At a border they use the physical document: UV and IR light, magnification for microprint, tactile checks on raised printing, and NFC chip reads that cryptographically verify the data page against the issuing authority. For documents submitted digitally - visa applications, remote onboarding - none of that applies, so the checks shift to internal consistency (MRZ checksums, cross-field logic, date arithmetic), localised edit forensics, and capture provenance analysis. Consulates and processing centres increasingly run automated forensic checks on supporting documents for exactly this reason; see [visa and immigration document fraud](https://tampercheck.ai/blog/visa-immigration-document-fraud-detection).

### Can you detect a fake PDF document the same way?

Mostly yes, plus more. A PDF carries structural evidence an image doesn't: incremental save history, object generation numbers, producer and creator metadata, embedded fonts, content-stream layering, and overlay objects that visually cover original text without removing it. Those often localise a tampered PDF faster and more definitively than pixel forensics. If the PDF is just a wrapper around a scanned image, the pixel-level markers in this post apply to the embedded image unchanged.

### Will OCR or an identity verification tool catch a tampered passport?

Usually not. OCR extracts text and assumes the pixels are genuine - a cleanly edited document produces clean OCR output with high confidence scores. Many identity verification products are built on OCR plus a database lookup plus a liveness selfie, none of which examines whether the image itself was manipulated. Tampering detection is a separate analysis; see [document tampering detection vs OCR](https://tampercheck.ai/blog/document-tampering-detection-vs-ocr).

### What about IDs generated entirely by AI, with no original document?

They fail the internal-consistency markers too, because generative models produce plausible-looking rather than arithmetically valid data - MRZ check digits in AI-generated passports are almost never correct. On top of that they carry generative fingerprints: frequency-domain artefacts from upsampling, unnaturally uniform micro-texture, cross-channel noise correlation, and no colour filter array trace from a real sensor. See [what is a deepfake document](https://tampercheck.ai/blog/what-is-a-deepfake-document) for the wider picture.

### How many markers does it take before you should reject a document?

An MRZ checksum failure is arithmetic and stands alone. Beyond that, the working rule is two or more high-confidence markers for a "likely tampered" verdict, and a single medium-confidence marker for human review rather than rejection. The point of naming which markers fired is that a reviewer can adjudicate the borderline cases with actual evidence in front of them instead of a bare risk score.
