Every DICOM object carries dozens of small text fields. Dates, times, ages, patient names, code strings, routing titles. The standard does not leave their format to taste. PS3.5, Section 6.2, Table 6.2-1 defines a grammar for each Value Representation (VR): the exact character repertoire, the length limit, and for dates and times the numeric range. The “Definition” and “Length of Value” columns carry the force of shall.

The rules are there. This post answers a simpler question with data: who is actually following them?

To find out, I scanned a diverse corpus of public DICOM data. 170,730 parseable files, about 68 GB, across 35 manufacturers, 18 modalities, and 50 institutions. The tool checks only the fields where a bad value does real damage downstream. The short version: most files are clean, but the violations that exist cluster in predictable, avoidable ways, and every common toolkit passes them straight to the consumer without a word.

What I checked, and why the method matters

The scanner is dicom_vr_scanner.py (PWNLAB, v2.0), published in the dicom-audit repository. It validates ten text VRs (DS, IS, DA, TM, DT, AS, CS, PN, AE, UR) against Table 6.2-1 of DICOM PS3.5 2026b. I call a value that fails the check a counterfeit value: text that looks like a legitimate value, slips past a non-validating parser, gets handed to the consumer as conformant, then diverges from anything the standard permits.

One design decision decides whether the results mean anything. The scanner reads the raw on-disk bytes of each element, not the value after the library parsed it. It does not trust the conversion logic the findings are about. pydicom is used only to walk the file structure. The verdict on each value is made on the bytes the vendor actually wrote.

That matters because of how the libraries behave. At pydicom’s install default reading_validation_mode = 1, a malformed date or a lowercase code string produces a non-fatal UserWarning and then returns the value unchanged. Call it warn-but-pass. At reading_validation_mode = 0, which plenty of shops set to kill warning noise on bulk reads, even the warning is gone. dcm4che, DCMTK, GDCM, and dcmjs sit in the same spot. Table 6.2-1 is normative, and the common libraries enforce none of it at the value layer. They hand the raw bytes to whatever computes on them next.

So a bad value is not caught at the door. It is caught, if at all, by the second consumer down the line. The retention engine, the strict date parser, the HL7 re-serializer, the array allocator. That is where the damage lives.

The numbers

Across 170,730 files, the scanner flagged 2,751 counterfeit values in 446 files. About 0.26% of the corpus.

SeverityFindings
HIGH1,549
MEDIUM737
LOW465

Those findings split into two very different stories, with no overlap between the files involved. Telling them apart is the point. They implicate different parts of the ecosystem and need different fixes.

Story one: de-identification is quietly making non-conformant files

2,665 of the 2,751 findings, 97%, are one value: ***.

They show up in 405 files, in exactly the fields an anonymizer would touch.

TagVRFiles
StudyDateDA405
SeriesDateDA405
ContentDateDA405
PatientBirthDateDA405
PatientAgeAS405
AcquisitionDateDA320
AcquisitionDateTimeDT320

Up to seven date and age fields per file, all overwritten with ***. This is a de-identification pipeline replacing dates and ages with a placeholder. *** is not a legal value for any of these VRs. DA allows only "0"-"9" in a fixed 8-byte YYYYMMDD. AS allows only nnn followed by D/W/M/Y. DT allows only digits, +, -, ., and space. An asterisk is outside every one of those.

This is the bigger story, because de-identification is not an edge case. It is a mandatory step before imaging data leaves a hospital for research, AI training, a vendor bug report, or a public dataset. If the anonymizer emits ***, every downstream copy of that study is non-conformant by construction, at scale, forever. PS3.15 (the Attribute Confidentiality Profiles) tells you which attributes to clean and gives you conformant ways to do it: a real dummy value, a date shift, or removing the element. An empty date is legal. *** is not.

The failure mode is the strict consumer. A pipeline that runs datetime.strptime(ds.StudyDate, "%Y%m%d"), or a Java ingest calling LocalDate.parse, does not get ***. It gets an unhandled exception and a dead worker. The permissive PACS that wrote the file never complained. The strict tool that received it falls over.

Story two: the real format bugs, and where they come from

Drop the anonymizer placeholders and 86 findings remain, in 41 files. 14 HIGH, 12 MEDIUM, 60 LOW. These are the real conformance bugs. Values a producer wrote wrong, not a scrubber’s placeholder. Each one is a distinct, reproducible mistake with a clause behind it, and every one was found in data that toolkits and vendors build and test against every day.

Legacy dotted dates and colon times. Files carry StudyDate = 1994.01.30 and StudyTime = 11:25:01. That is the old ACR-NEMA 300 YYYY.MM.DD / HH:MM:SS format, and DICOM names it directly. DA Note 1 in Table 6.2-1 reads “Use of this format is not compliant.” The . and : are outside the DA and TM repertoires. Thirty years after the format was deprecated it is still being emitted, and every non-validating reader still takes it.

NaN where a datetime belongs. 28 findings put the literal string NaN into FrameAcquisitionDateTime and FrameReferenceDateTime (VR DT). Some producer’s number-to-string path leaked a floating-point sentinel into a date field. A DT consumer that range-checks or parses this gets garbage. One that sorts frames by acquisition time gets undefined ordering.

Zero dates and the year 0000. Files carry 00000000 in StudyDate, SeriesDate, ContentDate, InstanceCreationDate, and the procedure-step dates, plus a DT of 00000000000000.59000, a fractional second hung on a year-zero timestamp. 00000000 is not a valid Gregorian date. There is no year 0, no month 0, no day 0. Anonymizers sometimes use it as a null-date placeholder, but the correct null date is an empty value, not eight zeros. These turned up in a vendor’s own QA dataset. The reference data the ecosystem trusts.

Lowercase code strings. 25 findings put Chest and Brain into BodyPartExamined, a CS field. CS is uppercase-only by repertoire. The defined term is CHEST, not Chest. These came from an AI-derived annotation dataset. A modern software producer, not a legacy scanner, which is exactly the kind of new tooling entering clinical pipelines now. A case-sensitive switch on BodyPartExamined silently falls through to its default branch on Chest.

A description stuffed into a code field. One CS value reads Tomoscintigrafia PET total body. Free text, mixed case, well past the 16-byte CS limit, in a field defined for controlled concepts. Any consumer that treats that field as an enumerated key, a routing switch, or an index key now holds a 31-character sentence.

Numbers that are not numbers. NumberOfFrames = 1A (VR IS, which allows only [+-]?digits). NumberOfFractionsPlanned = 10.000, a decimal point in an integer field, and this is an RT plan value. InstanceCreationTime = 17146, a TM with an odd digit count that parses to nothing valid. And PatientOrientation = L/P, a slash inside a CS, where the multi-value delimiter is backslash and / is outside the repertoire.

The attribution here is honest and more useful than a vendor name. These violations live in the test fixtures of dcm4che, fo-dicom, dcmjs, and pydicom, in vendor QA datasets, and in AI annotation releases. They are the files developers use to prove their DICOM code works. The toolkits keep many of them precisely because a real scanner once produced something like them, and then read them back without protest.

Why a bad string is a real risk

It is easy to file all of this under cosmetic. It is not. The danger is never in the value. It is in the consumer that trusts the field’s contract. The scanner sets severity by what the tag drives, and that mapping is where the risk gets concrete.

Dates and times drive decisions, not just display.

  • Retention and auto-purge. Archives delete or migrate studies by StudyDate, InstanceCreationDate, and ContentDate against a retention horizon. A counterfeit far-future or far-past date can sit forever beyond any purge window, a privacy and compliance failure, or trip premature deletion of a live study. *** and 00000000 both break this.
  • Prior-study matching and hanging protocols. PACS order priors by date and time to hang them side by side. Mis-ordered or unparseable dates surface the wrong prior, or hide the correct one. That is a clinical-safety error, not a cosmetic one.
  • Scheduling. ScheduledProcedureStepStartDate/Time routes modality worklists. A counterfeit schedule time mis-routes or drops an order.
  • Strict-parser DoS. Any second consumer that converts to a native type (strptime, LocalDate.parse, DateTime.Parse, or pydicom’s own .DA/.TM/.DT converters) throws on 2400, month 13, NaN, or ***. The permissive first hop passed it. The strict second hop crashes.

Numeric strings drive allocation and safety-critical math. This is the sharp end. IS and DS fields get fed straight into int() and float(). An IS like 9.99e99 (scientific notation, which IS forbids) becomes a 100-digit integer through a permissive parser. Use it as NumberOfFrames to size a buffer and it is an out-of-memory DoS. A DS of NaN or inf (both outside the DS grammar) in DoseGridScaling corrupts every dose voxel in an RT dose map. In DoseCalibrationFactor it corrupts PET SUV. In WindowCenter or WindowWidth it paints a black or undefined image. I have run these paths. A NaN scaling factor turns a whole computed dose grid into NaN, and the treatment-planning display shows zero or garbage. That is the difference between a spec footnote and a patient-safety issue.

Structured strings are injection vectors. PN (person name) has reserved delimiters (^, =), caps (up to 4 carets, up to 2 equals, 64 chars per group), and forbids backslash and control characters. When PN is re-serialized, into HL7/CDA under PS3.20 or JSON/XML under PS3.18 (DICOMweb), unfiltered delimiters or a smuggled CR/LF become HL7 segment injection, CDA/XML injection, or log injection (CWE-74). CS interpolated unescaped into a SQL or index key is the same story. UR (a URI field like RetrieveURL) that a consumer dereferences can be pointed at an internal address (SSRF) or a dangerous scheme. The scan found the format violations. The re-serialization boundary is where they turn into an incident.

What to take from this

The corpus is skewed on purpose toward public reference data, OSS toolkit fixtures, and de-identified research sets. So 0.26% is a floor, not an estimate of production PACS traffic, and most files carry no Manufacturer I can tie a name to. But that skew makes the finding stronger, not weaker. The violations live in the curated, widely-used data the ecosystem tests against, and the libraries everyone depends on relay them without checking. If the reference data is non-conformant and the readers do not validate, production will not be cleaner.

Two recommendations, one for each side of every DICOM link.

If you produce DICOM: validate at the value layer against Table 6.2-1 before you write, not just against a schema of which tags are present. Dates are eight digits with no separators. Times and datetimes have hard numeric ranges. CS is uppercase-only and 16 bytes. IS is digits only. DS is a decimal string with no NaN or inf. If you de-identify, follow PS3.15. Use a conformant dummy value, a consistent date shift, or an empty element for a removed date. Never write a placeholder (***, 00000000, or text in a numeric field) that violates the field’s own grammar. And never let a floating-point NaN or inf sentinel reach a string field.

If you consume DICOM: do not trust the VR contract just because a library returned a value without raising. Set reading_validation_mode to its strictest setting and handle the warnings instead of suppressing them. Validate any field you are about to int(), float(), strptime, dereference, or re-serialize. The value crossing your boundary was written by someone who may not have checked, and read by a library that definitely did not. Treat PN, CS, UR, and AE as untrusted input at every re-serialization and query boundary, like any other external string.

The DICOM standard did its job. Table 6.2-1 is precise, normative, and stable for decades. The gap is not in the specification. It is in the producers who do not emit to it and the consumers who do not enforce it. Both halves are fixable, and both are cheaper to fix now than after a NaN reaches a dose grid.


Scan methodology, the full finding set, and the dicom_vr_scanner.py tool are available in the pwnlabmx/dicom-audit repository. Grammar clauses quoted here are from DICOM PS3.5 2026b, Section 6.2, Table 6.2-1.

Cover photo: Brian Smith, a computed tomography technologist at Naval Medical Center Portsmouth, reviews a thoracic aorta scan at the acquisition console. U.S. Navy photo by Rebecca A. Perron, public domain as a work of the U.S. federal government, via Wikimedia Commons.

← Back to news