Practice questions › Clinical Informatics
De-identification by the Safe Harbor method
A practice question in the style of the CPHIMS® exam, from the free questions of HealthITPrep. The question is in English, as in the exam.
A hospital wants to share a dataset with a university using the HIPAA Safe Harbor method of de-identification. Which step is required?
Choose an answer, or open the explanation below.
Show the answer and the explanation
The correct answer: A. Remove the 18 specified identifiers, including names, all date elements except the year, and full ZIP codes, and have no actual knowledge that the remaining data could identify a person
✅ Why this answer
The Health Insurance Portability and Accountability Act (HIPAA) has two methods of de-identification. The Safe Harbor method requires two things:
- Removing 18 specified identifiers, including names, date elements (except the year), geographic units smaller than a state with a limited exception for the first three digits of the ZIP code, numbers (phone, medical record and Social Security), and full-face photographs.
- The facility having no actual knowledge that the remaining data could identify a person.
Data that meets these conditions is no longer protected health information.
❌ Why the other options are wrong
- (B): initials and a full date of birth are identifiers; the data remains identifiable.
- (C): encryption protects data from unauthorized people, but it does not remove identity; whoever holds the key sees the full data.
- (D): the second method (expert determination) requires a qualified expert in statistical methods to document that the risk of identification is very small, not the IT director’s opinion.
💡 Key concept
- Safe Harbor: a fixed, clear list of 18 identifiers to remove, but it removes details research may need, such as dates.
- Expert determination: more flexible and keeps more data, but needs a qualified expert and documented analysis.
De-identification is not absolute; combining several sources may reveal people. HIPAA does not require an agreement to share de-identified data, but organizations often add one that prohibits re-identification attempts.
In the exam: “remove 18 identifiers” + “no actual knowledge that a person could be identified” = Safe Harbor. “An expert in statistical methods documents that the risk is very small” = expert determination. Encryption alone does not de-identify data.
🔗 Related facts and questions
- The list is wider than many expect: it is not limited to the name and the record number; it also includes, for example, email addresses, IP (Internet Protocol) addresses, URLs (Uniform Resource Locators), device serial numbers and full-face photographs.
Practice question: a data set keeps every patient’s age in years, including one patient aged 95, and all other identifiers are removed. Does it meet Safe Harbor? → No; ages over 89 are grouped into a single category, “90 or older”, because high ages are rare and may reveal the person. - A re-identification code: the organization may give each record a code that lets only the organization link it back later, under conditions that stop the code from becoming a path to identity.
Practice question: the team used the last four digits of the medical record number as each patient’s study code. Is this acceptable under Safe Harbor? → No; the code is derived from an identifier, so it could be used to identify the person. An acceptable code has no relation to the patient’s information, and the linking key stays inside the organization. - How expert determination works: the expert analyzes the re-identification risk in the recipient’s context and uses techniques such as generalization (for example, age instead of date of birth), suppression of rare values, and date shifting, then documents the method and the result.
Practice question: two data sets are shared as de-identified. The first keeps only the year of each admission; the second keeps the month and the year, because the study is about seasonal illness. Which one cannot rely on Safe Harbor, and what does it need instead? → The second; Safe Harbor keeps only the year of a date, so the second needs expert determination, where an expert may accept the month after documenting that the risk is very small. - Why little information can reveal identity: studies have shown that full ZIP code, date of birth and sex together are enough to single out a large share of the US population. That is why Safe Harbor keeps only the year of the birth date and at most the first three digits of the ZIP code, while sex may stay, and why agreements often forbid linking the data to public sources such as voter lists.
- A comparison with European law: the General Data Protection Regulation (GDPR) distinguishes pseudonymised data from truly anonymised data. Pseudonymised data remain personal data for whoever holds the linking key. Under HIPAA, by contrast, data carrying a re-identification code that meets the conditions above can still count as de-identified. So the two standards are not the same.
- Primary versus secondary use of health information: "Primary vs secondary use of health information" (in the full bank).
- The limited data set and the data use agreement: "The limited data set and the data use agreement" (in the full bank).
- Institutional review board (IRB) approval before contacting patients: "Clinical research informatics and the institutional review board" (in the full bank).
Twenty questions like this one, free
A timed 30-question trial exam with this kind of explanation for every option and a score per domain. No payment details needed.
Start the trial examAll 20 free practice questions · CPHIMS guide
Practice questions written for study; they are not the questions of the real exam. An independent site, not affiliated with or endorsed by HIMSS. CPHIMS® is a registered trademark of HIMSS.