# Pathology AI Provenance: The Slide You Cannot Re-Create

> A whole-slide image is the end of a physical production line, and a re-cut slide is a new specimen. Why a sealed record is the only route back to a challenged AI read.
>
> Source: https://rankshieldmd.com/resources/pathology-ai-provenance-whole-slide-images/ · RankShieldMD (verifiable AI & post-quantum security for healthcare)

Resources // Provenance
# Which model read this slide? Pathology provenance and the image you cannot re-create.

In radiology, a challenged AI read can usually be re-examined against the original study. In pathology it cannot. The slide was cut, stained and scanned, and none of those steps repeat exactly, which makes the record you sealed at the time the only way back.
Read the guide →   Ask a question       Whole-slide imaging  DICOM Sup 145 · tiles  PHI-free · non-device
Published August 26, 2026

**A whole-slide image is not a measurement of a patient. It is a photograph of something a laboratory manufactured: a section of a particular thickness, from a particular block, stained with a particular reagent batch, scanned on a particular instrument. None of those steps repeats exactly. So when an AI-assisted read is challenged, re-cutting the slide does not reproduce what the model saw, and the record you sealed at the time is the only route back.**

We covered the radiology case in [proving which AI model read a scan](https://rankshieldmd.com/resources/ai-imaging-provenance-small-centers/). Pathology is not the same problem with slides swapped in for scans. The chain is longer, part of it is physical, and the part that is physical cannot be replayed. This guide covers where the chain actually breaks, what a pathology provenance record has to bind, and why the answer is smaller than re-engineering a laboratory.

## The image is manufactured before it is captured

A CT or MR image comes off the scanner as a direct measurement. If a read is questioned two years later, the original study can usually be retrieved and it contains the same pixels the model was given.

Pathology inverts that. Between the patient and the pixel sit accessioning, grossing, embedding, sectioning on a microtome, staining, coverslipping, and scanning. Each is a physical process performed by people and instruments on a particular day. The digital image is the last step of a production line, and the line is the part that does not repeat.

## A whole-slide image is not one image

DICOM Supplement 145, ratified in 2010, defined the Visible Light Whole Slide Microscopy Image object precisely because the existing standard could not represent these files. [1] A whole-slide image is stored as a multi-resolution pyramid: levels run from highest resolution to lowest, each level is an instance within the same series, and each tile is a frame inside a multi-frame object. [1]

That structure exists for a reason. Radiology images are typically megabytes at a single resolution. Whole-slide images can exceed several gigabytes, need full color fidelity, and have to support retrieval of arbitrary subregions. [1]

The consequence for provenance is direct. A model does not read a slide. It reads selected tiles at a selected magnification, often a small fraction of the total pixel data. Recording the file name tells you almost nothing about which of those gigabytes informed the output. Naming the level and the tile coordinates is the difference between a record that answers the question and one that only looks like it does.

## The re-scan that is not the same scan

The instinct when a read is challenged is to go back to the block and cut a new section. It is the wrong instinct if the question is what the model saw.

Section thickness, staining consistency and scanner calibration are recognized determinants of diagnostic reproducibility in digital pathology, and even micron-level thickness variation or illumination shifts produce measurable image differences that affect both inter-site concordance and AI performance. [2] Color normalization can partly compensate for stain variation, but it cannot recover information lost to a torn section, a folded ribbon or a poorly embedded specimen. [2] Scanning artifacts, including missing tissue and blurred regions, vary by scanner model. [2]

So a re-cut is a new specimen section, not a copy. It may support or contradict the original opinion, which has its own clinical value, but it is not evidence about the model's input. Those are different questions, and conflating them is how a defensible read becomes an argument about which image counts.

## What the record has to bind

A pathology provenance record that survives scrutiny reaches across both bands of the diagram above. The physical identifiers alone are a lab record. The digital digests alone are an inference log. The binding is the evidence.

In practice that means the specimen and block identity, the section and stain batch, the scanner and its calibration state, the whole-slide image instance identity, the pyramid level and tile coordinates read, the model version with a digest of its weights, and the output, all sealed together at the moment of the read rather than assembled afterward.

The reason to seal rather than store is the same one that applies to any audit trail: a record that can be edited after the fact answers a weaker question than one that cannot. We covered that distinction in [what tamper-evident actually requires](https://rankshieldmd.com/resources/tamper-evident-clinical-ai-audit-logs/), and in [what clinical-AI decision provenance is](https://rankshieldmd.com/resources/clinical-ai-decision-provenance/).

## Why this makes the record more load-bearing, not less

There is a tempting conclusion here that pathology is simply harder and therefore less provable. The opposite follows.

In radiology, the provenance record is a convenience: it saves you an argument, but the original study is usually still there as a fallback. In pathology there is no fallback. The physical chain that produced the image has already moved on, and it cannot be rewound. Whatever you sealed at the time of the read is the whole of the evidence.

That is an unusual property. It means the value of the record is highest precisely where reconstruction is impossible, and it means the cost of not having one is not inconvenience but the permanent absence of an answer.

## What to do without re-engineering the laboratory

Almost everything on that list already exists somewhere. The laboratory information system holds the specimen, block and section identifiers because the lab needs them for its own quality control. The stain batch is recorded for the same reason. The scanner writes its own model, calibration and acquisition metadata. The AI vendor knows its model version.

What is usually missing is not data collection. It is that nothing binds those facts to each other, or to the specific tiles, in a form that cannot be revised later. Sealing a digest of records that already exist is a materially smaller change than instrumenting anything new, and it does not require the image or any patient identifier to leave the building.

Start with the reads most likely to be questioned rather than the whole archive. A challenged read is almost always a specific, high-consequence case, and a provenance record that covers the cases people argue about is worth more than a partial one spread thin across everything.

## What we are careful never to claim

RankShieldMD does not read slides, does not render or score a diagnosis, and is not a medical device. It never sees PHI: what it seals are one-way digests and identifiers, not images and not patient data. It does not make a pathology AI system FDA cleared, and it is not a substitute for your quality system, your validation work, or your regulatory counsel.

What it does is narrow and checkable. It binds the physical and digital identity of a read into a tamper-evident, externally anchored record that a third party can verify without trusting us and without access to your systems. That supports an evidentiary question. It does not answer a clinical one.

## References

- [1] DICOM Standards Committee. *Supplement 145: Whole Slide Microscopic Image IOD and SOP Classes.* Defines the tiled, multi-resolution pyramid structure for whole-slide imaging. [dicom.nema.org/dicom/dicomwsi](https://dicom.nema.org/dicom/dicomwsi/)
- [2] Reproducibility determinants in digital pathology: section thickness, staining consistency and scanner calibration, and the limits of color normalization. See the deployment and quality-control literature on whole-slide imaging, including *Deployment of AI-driven automated quality control of whole-slide images in a large tertiary cancer center.* [sciencedirect.com](https://www.sciencedirect.com/science/article/pii/S2153353926001306)
- [3] Digital Pathology Association. *Digital Imaging and Communications in Medicine (DICOM), Supplement 145 draft.* [digitalpathologyassociation.org](https://digitalpathologyassociation.org/_data/cms_files/files/cms_pdf/Supplement_145_Draft_9.pdf)

Answer engine
## Ask the founder.

Straight answers about verifiable healthcare AI. Tap a question, or type your own.
Jamie Kloncz, founder  verified human  ✓                         Ask me anything about proving your clinical AI. I built RankShieldMD so a small practice can prove its AI, not just be asked to trust it.              Why is pathology provenance harder than radiology provenance?  Because the image is manufactured before it is captured. A CT or MR image is produced directly by the scanner from the patient, so re-pulling the original study usually gives you the same pixels the model saw. A whole-slide image is a photograph of a physical artifact that a laboratory made: a block was cut into a section of a particular thickness, stained with a particular reagent batch, coverslipped, and then scanned on a particular instrument. Every one of those steps varies. The image is the end of a physical production line, not a direct measurement of the patient.  Can we just re-cut and re-scan the slide if a read is challenged?  You can, and you will get a different image. Section thickness, staining consistency and scanner calibration are recognized determinants of reproducibility in digital pathology, and micron-level thickness variation or illumination shifts produce measurable image differences. Color normalization can partly compensate for stain variation, but it cannot recover information lost to a torn section, a folded ribbon or a poorly embedded specimen. A re-cut is a new specimen section, not a copy of the old one. If the question is what the model actually saw, a re-scan cannot answer it.  What does a pathology provenance record need to capture?  More than the case number and the model name. At minimum: the specimen and block identity, the section and stain batch, the scanner and its calibration state, the identity of the whole-slide image instance itself, the pyramid level and tile coordinates the model actually read, the model version and its weights digest, and the output. The tile and level detail matters because a model does not read a slide. It reads selected tiles at a selected magnification, and naming the file does not tell you which of its several gigabytes were used.  What is a whole-slide image, technically?  Under DICOM Supplement 145, which defined the Visible Light Whole Slide Microscopy Image object in 2010, a whole-slide image is a multi-resolution pyramid rather than a single picture. Levels run from highest resolution to lowest, each level is stored as an instance in the same series, and each tile is a frame inside a multi-frame object. That structure exists because these images can exceed several gigabytes and need region-based retrieval, where a radiology image is typically megabytes at a single resolution. The practical consequence is that "which image did the model read" is an ambiguous question until you name the level and the tiles.  Do we have to re-engineer the lab to do this?  No, and you should not try. The specimen, block, section, stain batch and scanner identifiers you would need are already recorded in the laboratory information system and the scanner metadata, because the lab needs them for its own quality control. What is usually missing is not the data but the binding: nothing ties those identifiers, the exact tiles, and the model version together in a record that cannot be edited afterward. Sealing the digest of what already exists is a smaller change than instrumenting anything new.  Does this make our pathology AI FDA cleared or compliant?  No. RankShieldMD is not a medical device, does not read slides, does not render or score a diagnosis, and never sees PHI. It records, in a tamper-evident and PHI-free form, which model version produced a given output against which sealed image region, so that the record can be checked later by someone who does not trust us. That supports evidentiary and post-market questions. It is not a clearance, not compliance, and not a substitute for your quality system or your regulatory counsel.
