De-identify scanned records
De-identification is not a search problem, it is a reading problem. The record is a scan. There is no text to search, the handwriting is somebody else’s, and the name appears eleven different ways across ninety pages.
Open the tool Free, in your browser. Nothing is uploaded.
After
Why a text search finds nothing
A scanned record has no text layer, or a bad one. Every identifier -- the patient name in the header of each page, the date of birth on the form, the clinician’s signature, the address on the referral letter -- exists only as pixels. Search finds nothing, and nothing is exactly as many matches as a tool with no reader can offer.
Three ways at the same word
Each page is read in the tab, so a name typed once is found in the lettering of the scan. Where the reading is defeated -- a stamp, a poor fax, an unusual typeface -- the word is looked for by shape instead, which is an optional second pass that says how long it will take before it starts.
Repeating furniture is picked out with a box: the letterhead, the signature, the hospital stamp. One pick finds every copy in the file.
Dates of birth, phone numbers, addresses and email addresses have detectors of their own, and names are found by reading the layout around them -- a title under a name, a "Name:" before it, a sign-off above it -- rather than by guessing which capitalised words are people.
What it does not do, and why that matters here
It does not decide what counts as an identifier. De-identification standards differ, a study protocol is not a records request, and the judgement about whether a rare diagnosis on a specific date is itself identifying is yours. The tool finds and removes what you ask for, and shows you everything it proposes before anything is covered.
It does not read handwriting, and it does not claim to. It does not verify its own work -- after exporting, open the file and read it, which on a de-identification job is the step nobody should skip.
And it has no memory. Nothing is stored, nothing is sent, and closing the tab ends it. If you want to come back to a part-done record, save a draft -- a file on your own machine holding your marks and the words you typed, which is as sensitive as the record itself and should be kept the same way.
Consistent labels for a study set
Switch labels on and every removed identifier carries a stable code: the same person is the same label on every page. A record stays followable -- this result belongs to that patient -- without the record saying who they are.
How to do it, in four steps
- Open the record. It is read in the tab, page by page.
- Type the identifiers you know, and tick the detectors for dates of birth, addresses and phone numbers.
- Pick out the signature, the letterhead and any stamp with a box.
- Turn on labels if the set has to stay followable, then Redact and export.
Who it is for
Research teams preparing a study set, clinicians sharing a case, HR handling a personnel file, and anybody who has to hand over a record without handing over the person.
Questions
Is anything sent to a server?
No. There is no server, and the browser is not permitted to send the file anywhere. It works offline.
Does it read handwriting?
No. Handwriting is not read, which is why a handwritten signature is picked out as a picture instead -- one pick finds every other copy of it in the file.
How are names found without a list?
By the layout around them rather than by their capitalisation: a job title under a name, a phone number or email under it, a "Name:" label, or a sign-off above it. That is where names actually sit, and it works the same for names that are not English.
Can the identifiers be recovered from the export?
No. Each page is rebuilt from pixels, so what was covered is not in the file.