Case Study • Data Privacy & Compliance

De-identifying 25,000 Patient Records for Analytics and Research

HIPAA Safe Harbor de-identification, clinical text redaction and data restructuring across 350GB of healthcare data

HerMD Logo

From Raw Data to Compliant Insights

25,000

Patient profiles

350 GB

Source data

99%+

Verified accuracy

5

Data formats handled

Creative Buffer Consultancy Pvt. Ltd. • July 2026

1. Summary

De-identifying 25,000 Patient Records for Analytics & Research

A US-based healthcare organization needed to make roughly 25,000 patient records available for analytics, research and downstream development. Before any of that could happen, the Protected Health Information had to come out, and the data still had to be worth analyzing once it had. The source environment held around 350GB of healthcare data spread across structured EHR tables, unstructured clinical notes and scanned documents.

The engagement had two connected objectives. The first was de-identification. The second was restructuring the output so it could be loaded into the client’s analytics environment and queried. That combination ruled out simple field-level masking: identifiers had to be transformed consistently across connected tables, PHI buried in free text had to be found and handled, and the result had to be validated for both privacy controls and data integrity before release.

We designed and ran an automated processing pipeline combining structured-data transformation, tokenization and pseudonymization, clinical text PHI detection and redaction, and multi-stage quality assurance. Post-processing verification put accuracy at 99% or higher, after which the de-identified and restructured dataset was released for approved downstream use.

Project Facts

HIPAA Safe Harbor
Client

Confidential US-based healthcare organization

Duration

4 weeks

Standard applied

HIPAA Safe Harbor de-identification, § 164.514(b)(2)

Data types

Structured EHR tables, unstructured clinical notes, scanned documents/images, handwritten notes

Formats

Clinical XML (CCDA), delimited/tabular exports, text-layer PDFs, scanned image-only PDFs

Scope beyond redaction

Restructuring of the de-identified data for analytics use

Architecture

Containerized, patient-folder-based processing with checkpointing and failure isolation

The Client and the Requirement

The Client and the Requirement

The client is a US-based healthcare organization that handles sensitive patient information across its clinical and operational workflows. Like most organizations of its kind, it holds large volumes of data containing PHI.

The company wanted to support clinical research, internal analytics and machine learning development using longitudinal patient data, and to share that data with internal teams and third-party engineering partners. Doing so with identifiable records was not an option. What was needed was a systematic de-identification approach rather than ad hoc masking of individual fields, applied across data associated with approximately 25,000 patients.

The client’s name and identifying details have been omitted from this case study. Project details have been generalized where necessary, while retaining the challenges, approach and capabilities involved.

What Made This Difficult

Five issues drove most of the design decisions.

The engagement covered around 25,000 patient profiles, but the underlying dataset ran to millions of individual records and relational events: demographics, encounters, diagnostic and billing data, pharmacy histories and unstructured clinical documentation including images & handwritten notes.

Removing obvious identifiers was the straightforward part. The harder problem was doing so without destroying the relationships, clinical meaning and analytical value of the data underneath.

Relational integrity across connected tables

PHI buried in unstructured clinical documentation

Dates and longitudinal analysis

Safe Harbor edge cases

The balance between redaction and usability

Issue 01

Relational integrity across connected tables

The dataset consisted of interconnected tables with primary keys, foreign keys and relationships spanning different parts of a patient’s healthcare journey. Anonymizing those tables independently breaks the links between them. Once broken, there is no reliable way to connect a patient’s encounters to their medications, diagnoses or claims, and the dataset loses most of its research value.

What was needed was a single identifier transformation strategy, applied consistently across the whole dataset, that preserved referential integrity while ensuring the original identifiers were never exposed downstream.

Issue 02

PHI buried in unstructured clinical documentation

Structured fields are predictable. You know where the MRN lives. Clinical notes are not predictable: names, family relationships, locations, employers and contact details appear inside free text, in the middle of clinically meaningful sentences.

Detection therefore had to work on context, not only on patterns. Notes routinely contain indirect identifiers that no field-level rule will catch, such as a reference to a daughter acting as primary caregiver, or a mention of the local plant where a patient works. Both identify a person; neither looks like an identifier to a regular expression.

Issue 03

Dates and longitudinal analysis

Much of the analytical value in healthcare data comes from the intervals between events. Strip or randomize dates carelessly and researchers lose the ability to study treatment timelines, gaps between encounters, medication patterns and readmissions.

Dates had to be handled consistently across all of a patient’s records rather than field by field, so that temporal relationships survived even where absolute dates could not.

Issue 04

Safe Harbor edge cases

Safe Harbor extends well past names and Social Security numbers. Ages above 89 have to be aggregated. Geographic data has to be cut back to three-digit ZIP, with an exception for sparsely populated areas. Dates other than year are out. Free-text identifiers of any kind are out.

Applying all eighteen categories consistently across thousands of profiles and millions of records is not something manual review can do reliably at this scale, which is why the rules had to be encoded into the pipeline itself.

Issue 05

The balance between redaction and usability

This was the constraint that shaped most of the tuning work. Untuned pattern matching and NLP models tend to over-redact, and a note reading “Administered 50mg of [REDACTED]” is compliant and useless at the same time.

The target was a dataset with minimal residual re-identification risk that still retained its clinical, relational and temporal value. Reaching it required automated processing, structured transformation, contextual handling of free text and quality assurance throughout, rather than a single pass of any one technique.

Our Approach

A Five-Stage, Containerized Pipeline

Creative Buffer implemented a five-stage, containerized pipeline with the patient folder as the atomic unit of work. A separate pseudonym store supported identifier stability, isolated from the deliverable, encrypted, and access-controlled.

Stage 1

Ingest & Normalize

Archives expanded to local scratch storage and source tree normalized so one directory represents one patient. Established a reliable scheduling boundary.

Stage 2

Per-Patient Profile Build

Structured records mined to build a per-patient identifier roster (names, addresses, IDs) acting as a patient-specific source of truth.

Stage 3

Layered PHI Detection

Three detector families combined: deterministic pattern matching, statistical NER, and known-identifier matching against the roster.

Stage 4

Format-Preserving Transform

Clinical XML rewritten, tabular data used column semantics, and text-layer PDFs used true content redaction rather than visual overlay.

Stage 5

Independent Verification

Independent scanner re-read transformed output searching for surviving identifiers, acting as an adversarial second pass.

Scalable Known-Identifier Matching

At corpus scale, the per-patient identifier roster can contain a very large number of strings. Naively comparing every identifier against every document is computationally expensive.

The implementation compiled the roster into a single Aho–Corasick automaton, allowing the corpus to be scanned in time proportional to document length plus the matches found. This moved roster matching from a potential bottleneck into a scalable detection layer.

Pseudonymization & Key Management

Where deletion would destroy required longitudinal utility, identifiers were replaced with stable surrogates. Surrogate codes were randomly generated and persisted in an isolated mapping store rather than being derived from the original identifier.

The mapping store represented the re-identification key and therefore received stronger operational treatment than ordinary output data: separate access permissions, encryption, and retention controls. It was never packaged with the de-identified deliverable.

Compliance Modes & Date Handling

The technical source describes two configured postures. Safe Harbor mode implements enumerated-identifier removal rules, including year-only dates, the 90+ age treatment, and geographic generalization. A separate limited-dataset/research mode can retain finer temporal resolution.

Engineering Architecture

Engineering for Scale, Fault Tolerance & Delivery

Patient-level process isolation

The patient folder was the unit of work, with no shared mutable state across workers.

Checkpointed, idempotent resume

Per-patient completion state was persisted so restarts could skip completed work; expensive profile preparation could be reused.

Worker-death handling

The orchestration layer detected process-pool failures and exited deterministically so the supervising process could restart from checkpoint.

Poison-pill quarantine

Patients that repeatedly failed or contained unrecoverable file errors were moved to an explicit remainder list rather than silently dropped.

Model preloading

The NER model was packaged with the container image to prevent large worker fleets from simultaneously downloading the model.

Unsafe logging removed

A serial classification/logging pre-pass was made optional, reducing both a throughput bottleneck and the risk of writing raw identifiers into output-adjacent logs.

Performance Trade-offs

The technical material describes a full-fidelity configuration with NER and OCR and an accelerated configuration that disabled statistical NER and handled image-only pages more conservatively.

The important engineering lesson is that performance modes were treated as explicit, documented trade-offs, with a remainder list and the ability to reprocess against the same pseudonym store.

Security Architecture

Security & Isolation

Separate Mapping

Re-identification mapping stored separately from deliverable data.

Access Boundaries

Mapping data and deliverables maintained different access boundaries.

Source Hygiene

Controls excluded real identifiers & secrets from shared copies.

Synthetic Fixtures

Test fixtures were completely synthetic to prevent leakage.

Container Isolation

Containerized processing & permissioned storage locations.

Classification and Treatment Decisions

Alongside the roster build, every field in the source environment was catalogued and classified, and a treatment was fixed per category before any transformation ran. Misclassification at this point either leaks PHI or destroys clinical data downstream, so the classification was reviewed before processing began.

Field CategoryTreatment
Direct identifiers — names, SSNs, contact detailsRemoved, or replaced with synthetic values where a populated field was needed for the data to remain usable
Record, account and member numbersReplaced with stable surrogates from the pseudonym store
DatesYear only; ages 90 and over aggregated
Geographic dataGeneralized to three-digit ZIP, with the population-threshold exception
Free textPassed to layered detection in Stage 3
Clinical contentPreserved
The People Behind The Project

Team & Expertise

The work was delivered by a small multidisciplinary team covering HIPAA interpretation, healthcare data engineering, clinical text processing and data quality assurance.

Pawan Panwar

HIPAA Specialist

Set the de-identification guidelines and framework for the engagement, and confirmed that the approach taken aligned with applicable privacy and de-identification requirements.

Himanshu Sharma

Technical Lead

Led the technical implementation, covering data architecture, tokenization, relational integrity, and the detection and redaction of PHI within unstructured clinical text.

Chinmay Pandit

Data Quality & Project Coordination

Led validation and quality assurance, including data integrity checks, de-identification validation, sampling and privacy-control verification with client communication.

Project Outcome

Delivered Data Ready For Analytics & Research

The client received a de-identified and restructured dataset covering approximately 25,000 patient records, prepared for downstream analytics and research while retaining the relationships and structure needed to work with longitudinal healthcare data.

What We Delivered

Approximately 25,000 patient profiles and 350GB of source data processed.

Structured EHR data, unstructured clinical documentation and scanned documents all addressed within a single pipeline.

Post-processing verification reported accuracy of 99% or higher.

Referential integrity maintained across tables, supporting cohort building and longitudinal analysis on the de-identified data.

De-identified data restructured and loaded into the client's analytics environment.

Validation covering both data integrity and de-identification completed before release.

What The Approach Depended On

Three things mattered more than the tooling to ensure accuracy and compliance.

Unified Handling

Handling structured tables and free-text notes as one problem rather than two, so a patient's identifiers were treated consistently wherever they appeared.

01

Upfront Decisions

Deciding date and identifier handling up front, in a way that protected privacy but kept time-series relationships intact.

02

Rigorous Validation

Validating output through both automated checks and manual review, since neither on its own catches what the other misses.

03
Industry Impact

Why This Matters For
Healthcare Organizations

Healthcare organizations increasingly need to use data beyond the clinical workflow it was collected in: for analytics, research, AI development, product testing and collaboration with technology partners. All of those require real-world healthcare data. None of them require patient identities.

Handled badly, organizations tend to find out late, either through an audit finding or through a re-identification risk nobody knew was there.

Handled Properly...

De-identification is what makes modern healthcare data work possible. It lets teams safely:

Train & Test AI Models

Train and test models on real clinical data safely.

Longitudinal Research

Prepare longitudinal datasets for approved medical research.

Development & QA

Hand representative data to development and QA teams.

Partner Collaboration

Bring in engineering or analytics partners without breach liability.

Capabilities

Our Healthcare Data Services

This engagement reflects a capability Creative Buffer offers as part of a broader healthcare data engineering and privacy practice, serving US healthcare providers, payers, health-tech companies and life sciences organizations.

PHI De-identification

Under HIPAA, including Safe Harbor and support for Expert Determination.

Clinical Text PHI Redaction

Using entity detection, pattern matching, contextual rules and validation.

Tokenization & Pseudonymization

Preserving approved longitudinal linkage without exposing original identifiers.

Data Quality & Validation

Schema, referential integrity, transformation, PHI leakage and downstream utility checks.

Synthetic Healthcare Data

For development, testing and demonstration use.

Healthcare Data Engineering

ETL and ELT, transformation, normalization, data lakes and warehouses.

Healthcare Interoperability

Work across HL7 / FHIR / EHR integrations.

Next Steps

Let's Talk About Your Healthcare Data

If your organization is working with sensitive healthcare data for analytics, AI, research, development, migration or controlled collaboration, we can help assess the data, define an appropriate de-identification approach, and design a repeatable processing and validation workflow around it.

Project Questions

Frequently Asked Questions

We don't claim certainty, we claim measurement. A verification scanner re-reads the output and reports residual matches; that number is delivered with the dataset. Claiming zero leakage without a verifier is the answer that should worry you.

Format fidelity and join stability. Most tools either flatten documents or generate per-file surrogates, both of which destroy the dataset's value. The per-patient roster layer — using the patient's own structured record to find their identifiers in free text — is also not something generic tooling does.

Yes, alone. That's why it's one of three layers and not the primary one. Deterministic patterns and the roster carry the high-confidence load; NER covers the residual free-text surface.

Correct, which is why we don't draw them. Text-layer redaction removes content from the stream, and form-field values are cleared at the annotation layer, not painted over.

§164.514(c). A derived code is a code derived from the identifier, which is disallowed under Safe Harbor regardless of key strength. It also means the mapping store is the sole source of stability, which we treat as an operational requirement.

Three things, all instructive: the parallelism boundary didn't match the real data layout and cost hours before it was caught; targeted XPath-only XML handling leaked identifiers in unenumerated elements until the full-tree sweep was added; and a zero-valued shift parameter in an early configuration was a no-op that left dates intact — caught by the verifier, which is precisely the argument for having one.

Appendix

The HIPAA Safe Harbor Standard

Section 164.514(a) of the HIPAA Privacy Rule sets the standard for de-identification of protected health information. Under it, health information is not individually identifiable if it does not identify an individual and the covered entity has no reasonable basis to believe it can be used to identify one.

The Privacy Rule allows two methods: the Safe Harbor Method and the Expert Determination Method. Satisfying either demonstrates that a covered entity has met the standard, and information de-identified by either method is no longer protected by the Privacy Rule, because it no longer falls within the definition of PHI.

This engagement followed the Safe Harbor Method, under which the following identifiers of the individual, or of relatives, employers or household members of the individual, are removed:

A

Names.

B

All geographic subdivisions smaller than a state, including street address, city, county, precinct, ZIP code and their equivalent geocodes, except for the initial three digits of the ZIP code if, according to current publicly available Census Bureau data: (1) the geographic unit formed by combining all ZIP codes with the same three initial digits contains more than 20,000 people; and (2) the initial three digits of a ZIP code for all such geographic units containing 20,000 or fewer people is changed to 000.

C

All elements of dates (except year) for dates directly related to an individual, including birth date, admission date, discharge date and death date; and all ages over 89 and all elements of dates (including year) indicative of such age, except that such ages and elements may be aggregated into a single category of age 90 or older.

D

Telephone numbers.

E

Fax numbers.

F

Email addresses.

G

Social Security numbers.

H

Medical record numbers.

I

Health plan beneficiary numbers.

J

Account numbers.

K

Certificate and license numbers.

L

Vehicle identifiers and serial numbers, including license plate numbers.

M

Device identifiers and serial numbers.

N

Web Universal Resource Locators (URLs).

O

Internet Protocol (IP) addresses.

P

Biometric identifiers, including finger and voice prints.

Q

Full-face photographs and any comparable images.

R

Any other unique identifying number, characteristic or code, except as permitted for re-identification under § 164.514(c).