About

DelPHI

DelPHI

Deidentify and Label PHI — a purpose-built de-identification platform for clinical documents.

Healthcare AI lags because protected health information can't move without risk. DelPHI removes that blocker: it detects and redacts PHI across all 18 HIPAA Safe Harbor identifier categories, with labelled spans, audit-ready output, and human-review support — so teams can put clinical text to work in NLP, coding, research, and analytics without fear. Built for Safe-Harbor-aligned de-identification workflows.

What DelPHI does

DelPHI detects and redacts protected health information across all 18 HIPAA Safe Harbor identifier categories — names, geography, dates & ages over 89, phone/fax, email, SSN, MRN, health-plan and account numbers, certificate/license numbers, vehicle (VIN) and device identifiers, URLs, IP addresses, biometric identifiers, and any other unique identifier. Full-face photographs and other embedded images are removed at text extraction, so they never enter the de-identified output. It runs over clinical notes, discharge summaries, charts, and lab reports, so the residual text can be used for downstream NLP, risk-adjustment coding, research, and analytics without exposing patient identity.

Each PHI span is detected, labelled with its HIPAA category, and either anonymised (replaced with a category token), pseudonymised (replaced with a synthetic equivalent), or deleted entirely. The same span data is returned to the caller for audit and review.

Capabilities

Delphi-OM clinical-NER transformer

Default engine. Fine-tuned DeBERTa-v3-large (434M params) running on a hosted inference endpoint. ≈95.8% recall on a synthetic clinical-note benchmark across its core categories, most at 100%. Vehicle (VIN) and image identifiers are covered outside the NER model, so they're not in the model-recall figure.

Clinical-code shield (DelPHI proprietary)

Our own post-filter on top of the NER model. Regex-detects ICD-10-CM, HCC, CPT, and NDC ranges in the source text, and drops any model-proposed PHI spans that overlap them. Also drops spans whose text matches a curated 151-term clinical vocabulary (names and organisations exempted to avoid leaking PHI like a patient named 'Gene').

Single + batch processing

3-step wizard handles a single document or batches of up to dozens of files in parallel. Same UI, same engine, just a tab toggle.

PDF, DOCX, TXT input

Multi-format extraction including OCR fallback for scanned PDFs. (Note: layout and embedded images are dropped at extraction today — in-place redaction is on the roadmap.)

Three redaction modes

Anonymise (token replacement, e.g. [NAME]), pseudonymise (synthetic equivalents, e.g. 'Mark Reed'), or delete (PHI removed entirely).

Per-document timing & stats

Processing time, document size + approximate token count, PHI count, and throughput are surfaced on the results page so operators can monitor latency and complexity.

Multi-format export

Download redacted output as TXT, DOC (Word), PDF (print-ready), or a full JSON audit report including every detected span.

HIPAA Safe Harbor coverage

All 18 HIPAA Safe Harbor identifier categories: 17 detected and labelled in text (incl. vehicle/VIN via a deterministic recognizer and biometric identifiers), plus full-face photos/images removed at extraction. Which categories are redacted is user-controllable. Designed for Safe-Harbor-aligned workflows — not a substitute for a determination or legal review.

How it works

DelPHI is a three-stage pipeline. Each stage has a clear responsibility, and the value DelPHI adds over running a standalone NER model is concentrated in stages 2 and 3.

  1. 1. NER detection — Delphi-OM transformer. A fine-tuned DeBERTa-v3-large (434M params) running on a hosted inference endpoint scans the document and proposes candidate PHI spans tagged with HIPAA categories. This stage is where the upstream model does its work.
  2. 2. Clinical-code shield — DelPHI proprietary. A post-filter that knows what the NER model does not: that E11.9 is an ICD-10 diabetes code, not an account number; that CPT 99213 is a procedure code, not a vehicle ID; that metformin is a medication, not a name. It detects ICD-10-CM, HCC, CPT, and NDC ranges in the source text, and drops any model-proposed PHI spans that overlap them. It also drops spans whose text matches a curated 151-term clinical vocabulary. Names and organisations are exempted from the term check so a patient named "Gene" is never silently dropped. The shield is deliberately conservative — patterns that could collide with PHI formats (bare 5-digit CPT, 3-char ICD-10, bare HCPCS) are intentionally not shielded because under-shielding is safer than leaking PHI.
  3. 3. Redaction — DelPHI. The surviving PHI spans are anonymised, pseudonymised, or deleted according to the chosen mode. Clinical codes, medical terminology, and document structure pass through untouched.

The shield is what differentiates DelPHI from a generic NER wrapper. It is implemented in engine_v2/shield.pywith patterns ported from DelPHI v1's medical-intelligence module and refined for v2.

Who DelPHI is for

DelPHI was built to feed clinical-coding and HCC risk-adjustment pipelines that need volumes of medical text without the PHI. It fits naturally between an EHR export step and any downstream NLP or analytics consumer.

Typical use cases include risk-adjustment coding (HCC), payer chart review, clinical research datasets, model training, internal analytics, and any document workflow where HIPAA Safe Harbor compliance is a prerequisite.

Important: assistive tool, not a sole compliance control

DelPHI is an assistive de-identification system. The underlying transformer is benchmarked at 95.8% overall recall on synthetic clinical notes, with one known weak spot (device identifiers at ~32%). It is not a substitute for human review on safety-critical workflows.

For HIPAA Safe Harbor sign-off, pair DelPHI with a human QA checkpoint or a second independent control (e.g. a pattern-based scan over the redacted output). The wizard surfaces every detected span and confidence score to make this review fast.

Built on open source

DelPHI's PHI detection is powered by clinical NER models from the OpenMedproject, used under the Apache License 2.0, together with a stack of open-source libraries. DelPHI's proprietary value — the clinical-code shield, redaction pipeline, and product — is built on top of them.

See Licenses & Attribution for the full list. The desktop edition ships these notices as THIRD-PARTY-NOTICES.txt.

Get started