All Papers

Law

Beyond the Business Associate Agreement

Why a HIPAA Expert Determination on De-Identification Can Be a Protective Foundation for Health System AI Partnerships

Whitestone·August 4, 2026
Beyond the Business Associate Agreement

ERIC BROOKS, MD, MBA, JD (CANDIDATE)

Thought Leadership on Healthcare AI, Privacy, and Data Governance

A Health Policy White Paper

Of interest to health system privacy officials, compliance officers, security leaders, and counsel.

Previously published in part by Forbes Business Council on August 3, 2026

Executive Summary

Artificial intelligence is no longer an unusual add-on for health systems. It now runs through administrative operations, utilization management, performance tracking, and the electronic medical record, and it is moving into clinical care, where radiology and pathology are among the first and most important early adopters.[1] A system that refuses to work with AI vendors, or to build the same capability itself, risks falling behind on two fronts: in cost, reimbursement optimization, and throughput, and in the quality and reach of the care it can deliver. For most organizations then, the real question is not whether to use AI, but how to use it while protecting their patients and themselves.

That question takes on a sharp legal edge the moment a vendor uses or, in some cases, attempts to keep patient data after a project ends and continues to deploy it to build and improve its AI systems. Where the retained data remain identifiable, that practice sits in direct tension with HIPAA’s requirement that a business associate return or destroy PHI when an agreement terminates, unless doing so is infeasible. It is tempting to ask simply whether a vendor is “HIPAA compliant.” But HIPAA does not certify a product or a company as compliant. Its obligations attach to regulated entities and continue over time, and a vendor’s posture on the day of signing tells you little about how the data are handled once the relationship is over. The questions that matter are narrower: how does the vendor handle your hospital or health entity’s protected health information (“PHI”) when AI is involved, what data is it using to train its models, is PHI only being used for permissible purposes, and whether PHI may be kept and how does it keep that data private and secure both during the engagement and after it ends?

This paper answers those questions from a vendor’s point of view. As an executive for a health tech company, I work with a dedicated team on AI offerings in the appeals and prior authorization space, and I offer this as a thought leadership piece based on our experience and deep contemplation on how to protect our clients and, most importantly, patients: not as a buyer’s guide. This view is that, for an AI vendor that uses, keeps, and reuses healthcare data, turning that data into de-identified information through the expert determination method (45 C.F.R. § 164.514(b)(1)) is the most durable and practical approach available. However, admittedly, it is not the only lawful one. Keeping PHI under a business associate agreement or keeping a limited data set under a data use agreement can each be permissible, but only in the right circumstances. So too can a patient authorization that complies with HIPAA and applicable state law; in some settings that route can be as protective as, or more protective than, de-identification, though it is often impractical for the large, retrospective, and unstructured data sets on which AI vendors depend. But each carries a built-in weakness for AI vendor relationships that a HIPAA de-identification expert determination does not, and that weakness tends to surface at exactly the wrong moment: when the contract ends. In my own team’s business, we have obtained a formal expert determination for our data, so the analysis here rests on a choice we have actually made.

I. The Competitive Reality, and the Question It Forces

There is still no comprehensive federal law explicitly governing the use of artificial intelligence in healthcare, and no one rule that directly singles out AI when patient data are involved. What exists instead is a patchwork of AI regulations that many health systems are, or should be, aware of. In 2025 alone, lawmakers in 47 states introduced more than 250 bills aimed at regulating AI in healthcare, and 33 of them became law across 21 states.[2] Most of that activity has clustered in three areas: (1) clinical care, (2)digital health and health privacy, and (2) health insurance plans’ use of AI to make coverage and claims decisions, with comparatively little aimed at the administrative and operational AI PHI protection that providers themselves must be aware of during use, or the administrative playbook that must be followed to ensure PHI and handling of sensitive materials lowers liability when building or venturing with AI networks. Yet, to make matters more confusing when it comes to AI regulation, the direction of AI guidance and regulation at the federal level runs the other way: a December 2025 executive order pushes agencies toward a single national framework to deregulate AI and directs the Department of Justice to challenge state AI laws that would fragment it.[3] The administration extended that federal-first posture in June 2026 with a further executive order on advanced AI innovation and security. Nonetheless, not all is chaos when it comes to understanding how to responsibly use a vendor’s product that incorporates AI technologies: HIPAA and HITECH, meanwhile, still apply with full force and effect on AI in most healthcare scenarios, but they were written for a world of less advanced records and databases, not for systems that learn from data, and they were never purpose-built specifically for AI.

A health system navigating that landscape might reasonably conclude that, for now, the legal floor for AI is low and uncertain. That conclusion, however, would be incomplete. Even where no federal statute speaks directly to AI and its safeguards in use around PHI, the legacy HIPAA and HITECH laws create clear standards for PHI use, and, moreover, the accreditation bodies that health systems answer to increasingly expect them to govern AI security and privacy standards, and very much are filling the shoes of what federal regulation could otherwise require. In September 2025, the Joint Commission and the Coalition for Health AI released the first national guidance from a U.S. accrediting body on the responsible use of AI in healthcare, calling on organizations to put in place formal AI policies, governance structures, data-protection controls, and vendor oversight while signaling that these expectations will likely shape future accreditation.[4] The American Medical Association, which deliberately calls the technology “augmented intelligence” to stress that it should support rather than replace clinical judgment, has issued its own governance framework urging health systems to document the policies, accountability, and oversight behind the AI they deploy, either in-house or through vendors.[5] In short, a health system may soon need to explain and document how it protects patient information when AI is in the loop; not because an AI-built statute requires it, per se, but because the bodies that accredit it do, and the legacy laws around PHI protection require careful scrutiny and clear oversight of its deployment, albeit indirectly.

That governance-and-accreditation question: how should a health system build the internal policies and procedures these bodies expect? Is an important one, but it is beyond the scope of this paper, and we will take it up in a separate white paper. Here we stay with a narrower and more concrete question: how should hospitals and health systems think about the data themselves when using in-house AI or partnering with an AI vendor? Ultimately, when it comes to AI in healthcare, the data are everything: the information of value to be analyzed or operationalized for the health entity, and the data to be used and potentially retained for machine learning and model/product perfection by the AI vendor or product.

The spread of AI across healthcare is broad and speeding up. The question is no longer whether to let AI touch patient data, because for most systems some exposure has become the price of staying competitive. The question is how a system that partners with an AI vendor can protect its patients and itself in the process. This paper takes the part of that question that turns on the data: how a vendor that is a HIPAA business associate handles PHI when AI is involved, what it uses to train its models, and how it keeps that data secure during the engagement and after it ends.

II. What AI Actually Is, and How the Law Treats It

Good legal analysis of AI starts with an accurate picture of the technology, because much of the confusion in this area comes from treating AI as single, self-directed software. It is more complex than that.

In essence, AI did not come from one breakthrough algorithm. It grew up as a layered system of statistical models, neural networks, optimization methods, search, planning, and rule-based reasoning, all working together. The earliest systems, from the 1950s through the 1980s, ran on hand-coded logic and rule engines. In the 1990s and 2000s, machine learning added methods that learned from data but still relied on features designed by people. Deep learning took hold around 2006 and accelerated through 2015, as multi-layer networks began learning features on their own. Over the past five years or so, systems have combined deep learning, reinforcement learning, and classical planning into single architectures. The lesson running through that history is that the “intelligence” does not live in any one algorithm; it comes from how many specialized algorithms are layered and coordinated. At bottom, AI is a large, structured system for processing data: an architecture of many programs and code interwoven that evolved elegantly in their own right over time before reaching complex articulation with one another.

That fact has a legal consequence. For healthcare regulation, AI is generally treated as a tool used by a regulated entity, not as an independent actor that takes over care or operations, or itself takes on liability; the relevant statutes, regulations, and agency guidance all start from that premise: the health system using AI is responsible for what it does with PHI and patient data, not the system itself (product liability law, which can often escape AI, is a topic for another time, but also emphasizes that liability rests with the health entity).[6] Yet, for many healthcare administrators, knowing what AI does, how it works, and what happens with the patient data are critical questions for ensuring compliance with HIPAA and HITECH, and other state and federal privacy laws, as well as increasing pressure on health systems from accreditation bodies, as mentioned previously. Tt is the asking of those questions and knowing how to approach inquiry during AI vendor engagement that will lead to clarity in what otherwise seems and feels like a black-box system. While much of this black-box feel is a consequence of the intellectual property safeguards around AI, including the fact that most vendors have to rely on trade secrets rather than patents or copyright,[7] which is beyond the scope of this discussion, much of the mystery and fogginess around what to do and ask when it comes to AI vendor engagement can be brought to light with the right inquiries and contracting, which we discuss next.

III. The Laws We Already Have Reach AI

As alluded to above, because an AI system that takes in clinical data is, in practical terms, a data processor, it already falls inside the frameworks a few critical laws supply. A vendor that creates, receives, maintains, or transmits PHI on behalf of a covered entity or another business associate is a business associate under the HIPAA rules,[8] and a business associate is directly liable under the HITECH Act for the Privacy and Security Rule provisions that apply to it.[9] The legal nature of the processing does not change just because it happens inside an AI model instead of a database.

So, AI activity sits inside several overlapping bodies of law at once: the Privacy Rule, the Security Rule[10], the Breach Notification Rule, the HITECH business associate liability rules, comparable state data privacy laws, and a growing body of state law seeking to expand into the AI regulation space. Outside of new state laws aiming to bring clarity to AI regulation in healthcare, regulators have started to fold AI into these existing legal structures rather than treat it separately. The proposed update to the Security Rule, for example, would build AI tools into a regulated entity’s risk analysis.[11]

The practical point is clear: the legal tools needed to hold an AI vendor accountable already exist. The harder question is which of those tools, carve-outs, and applications best protect PHI during AI engagement and keep protecting the data after the parties go their separate ways.

IV. Why the Familiar Retention Tools Fall Short

When an AI vendor wants to use, keep, or repurpose PHI clinical data, three legal tools usually come up under HIPAA/HITECH and privacy and security laws. Each deserves a fair reading, and each has a clear weak point.

A. Keeping PHI Under a Business Associate Agreement

A business associate agreement (BAA) can let an AI vendor handle PHI during a project, but it comes with an obligation that matters most at the end. The Privacy Rule says that, when the agreement terminates, the business associate must return or destroy all PHI it received from, or created for, the covered entity (read: your health system). It may keep the clinical PHI data only where return or destruction is genuinely infeasible, and even then the agreement’s protections follow the data and further use is limited.[12] AI vendors sometimes treat that ‘infeasibility clause’ as a general license to hold onto a data set after a contract ends.

However, that reading is hard to defend. The infeasibility exception is meant to be narrow. Most AI systems do not blend PHI in a way that genuinely prevents pulling it back out; even data spread across fields, free-text stores, or multiple databases can usually be found and either returned or destroyed. The hassle of separating an AI model from its input data, or even more problematically, its training data, is not “infeasibility” in the regulatory sense.

And this is how regulators read “infeasibility.” The Office for Civil Rights treats infeasibility narrowly. In its cloud computing guidance, OCR says return or destruction would be infeasible where, for instance, “another law requires the business associate to keep the data past the end of the contract.”[13] A vendor’s business interest in keeping clinical PHI from your health system as a training set is not that kind of competing legal duty. In essence, a BAA does not give a machine learning corpus durable protection over the long run. When the contract ends, the covered entity (your health system) generally requires the PHI an AI model was trained on to be returned or destroyed, and wanting to keep that data for further development will not make return by the contracted vendor infeasible. Whatever a BAA allows during the project, it is not a safe foundation for a permanent, learning asset by your AI vendor. Therefore, aside from the larger existential question about whether your health system wants to permit an AI company to retain your PHI data at all, let alone train on it, doing so opens your system (and the AI vendor) to a tremendous amount of scrutiny if the only mechanism for using and retaining PHI in an AI model is the classic BAA. It is clear that something more is needed. Although a BAA is necessary for any AI vendor, acting in the capacity of a business associate, to touch and handle PHI, it is certainly not sufficient for outlining, in a relational context, what your health system needs to execute with a vendor to protect yourself, your data, and your patients. Thus, signing a BAA is only the first step.

B. Keeping a Limited Data Set Under a Data Use Agreement

The second tool to help protect an AI vendor relationship with regard to PHI handling is the limited data set (LDS), governed by a data use agreement (DUA).[14] An LDS strips out the direct identifiers the rule lists while keeping certain PHI elements, notably dates and limited geography, that stay analytically useful,[15] and the DUA limits how the recipient may use and disclose the limited PHI data and requires safeguards. Note that this contract executed by your health entity and the AI vendor may establish certain permitted uses of PHI beyond what the BAA contemplates.

There is a real and under-appreciated advantage here. The rules around LDS use under DUAs do not put the return-or-destroy obligation on a limited data set the way they do on full PHI at the close of a BAA; on that point they simply say nothing. Because of that silence, a DUA drafted to allow an AI vendor to keep and retain LDS PHI after an agreement ends is possible. Done properly, this is a legitimate way to preserve a data set and possibly let the AI vendor use PHI for machine learning.

But the advantage of the AI vendor keeping your health entity’s PHI indefinitely is conditional, and it rests on two assumptions that may not hold. First, a limited data set may be used only for research, public health, or healthcare operations.[16] If a regulator decided that using the data to build or train a commercial model falls outside those three purposes, the data set would no longer qualify, the use would fall out of compliance, the DUA’s premise would collapse, and the parties would fall back on whatever BAA governs (and its return obligation) leaving the kept data as PHI held without authorization. Second, a limited data set is still PHI. It stays inside HIPAA, stays subject to breach notification, and still carries the re-identification risk that comes with the dates and geography it keeps. The DUA narrows the exposure; it does not remove it, and relying on it to allow for machine learning could easily be swept away by OCR guidance or case law interpreting that commercial learning of AI on an LDS is neither research, public health, nor healthcare operations. It is highly conceivable that such a negative permissive interpretation on LDS use for AI machine learning under an LDS/DUA arrangement could emerge.

C. De-identification by the Safe Harbor Method

The third legal instrument, and the one that looks simplest, is the so-called “Safe Harbor method,”[17] under which data generally count as de-identified once eighteen listed categories of identifiers are removed.[18] The HIPAA/HITECH Safe Harbor works well for structured, tabular data, where every field is known and can be cleared: the classic forms of data we were used to keeping over a decade ago, before the highly sophisticated and more complex underpinnings and functionality of “AI data sets” emerged.

As stated above, the Safe Harbor appears to work poorly for the clinical unstructured data AI systems actually consume. Clinical notes, insurance denial letters, radiology and pathology reports, and images are mostly free-text narrative and pixels, not tidy fields. The last listed category, the eighteenth PHI element, is a catch-all covering “any other unique identifying number, characteristic, or code.”[19]. In free text, that category is hard to satisfy with confidence. Arecord can still identify someone through an unusual diagnosis, a named referring physician, a rare occupation, or a detailed account of an event, none of which lives in a discrete identifier field. Because you can rarely be sure you have removed every contextual identifier from narrative text, Safe Harbor can rarely be certified for that kind of unstructured and highly personal material with confidence. The Safe Harbor method was not built for narrative medicine, and its guarantees do not carry over cleanly.

V. Why Modern AI Raises the Re-identification Stakes and Why Expert Determination of De-identification Answers Them

The HIPAA/HITECH de-identification standard offers a second path, and the reasons it is likely the more durable choice for AI vendors, and for your health system to require or explore, are tied to a risk that modern AI systems make worse.

First, it is worth reminding, or pointing out, that health data that meet the de-identification standard, whether through Safe Harbor or Expert Determination, are not individually identifiable, so they are not PHI and actually fall outside HIPAA/HITECH regulation altogether, thus drastically limiting liability risks when done correctly and routinely.[20] De-identified clinical data derived from PHI can be kept and reused without those prior privacy, security, breach notification, and other statutory and regulatory constraints, because the thing the law protects (the link to an identifiable person) has been cut. Of all the legal tools here, de-identification is the only one that, at the same time, lets the vendor keep the data, removes the use restrictions, and lowers the residual breach risk to patients.

Again, two methods get you there. Safe Harbor, for the reasons above, does not hold up against unstructured data very well. That leaves the expert determination method, under which a qualified person applies generally accepted statistical and scientific principles and documents that the risk of re-identification is very small.[21]

Performing expert determination is harder for unstructured data than for tabular data, and that difficulty tracks the risk. It is worth noting that two major re-identification routes are unique to narrative clinical material. The first is putting pieces together within a single document. No one sentence identifies the patient, yet a date, a place, a rare condition, a family detail, or the patient’s vocation, taken together, can make re-identification far more likely than any one of them alone. The second is putting pieces together across repositories. A fact that is harmless on its own can become identifying once a large enough database is assembled, especially if that database also holds other material about the same people, so that someone pooling different sources can triangulate an identity that no single source would reveal on its own.

Modern AI sharpens both routes in a way that older database analysis did not have to face. A traditional re-identification analysis could treat the problem mostly as one of contents and access: what is in the database, and who can reach it. Today’s AI is not just a processor of data but a generator of output, and, depending on the system, its code can both weave together related data inside a store and push that web outward across other software and digital media. Data that once stayed inside a defined repository can now be moved, reconstructed, and surfaced in places it would never have reached historically, and contextual fragments that were not identifying inside a single store can be reassembled (through that generation and spread) into something that functions as identifying across a larger digital ecosystem (read: the internet or partnered/compiled multiple AI and non-AI databases merged together through contracting or business relations). The same computing power that makes large-scale model building possible also makes large-scale recombination and distribution, and therefore re-identification, possible. In that setting, lowering the re-identification risk of any single document is necessary but not enough. The entire complexity of re-identification, beyond a single document within a single database, instead needs to be contemplated in a much more sophisticated fashion that accounts for (1) how the AI vendor intends to use the de-identified data, (2) what other tools or AI systems will touch the database with de-identified data, (3) to whom or what the de-identified data will be sold or allowed access, and (4) what safeguards are in place to protect even those de-identified data that ostensibly require little HIPAA, HITECH, or similar legal security and privacy protections that stop applying but should still be contemplated even for de-identified data.

Nevertheless, and for those reasons, hitting the “very small” risk of re-identification requires as much technique as judgment by the expert determiner, and the Expert Determination rule helps by refusing to fix a single numerical cutoff. What counts as very small risk of re-identification (e.g., 1%, 2%, 5%) depends on the type of data and the setting it lives in.[22] Several families of techniques are in common use today to de-identify data and apply to AI use. For example, generalization swaps precise values for broader ones (a birth date becomes a birth year, a five-digit ZIP code becomes its first three digits). Suppression techniques remove values too revealing to keep. Perturbation and obfuscation methods involve the controlled addition of statistical noise, including formal differential privacy, to obscure individual records while preserving the overall patterns. Synthetic surrogate replacement substitutes consistent but fictional names, dates, and locations for the real ones: an approach especially well-suited to unstructured narrative, because it keeps the contextual structure a model needs to learn from while cutting the link to real people. However, preventing back-engineering of the real ones from consistent obfuscation patterns needs to be top of mind during this approach. A common way to describe the resulting protection is k-anonymity: making sure each record is indistinguishable from at least k − 1 others on the combination of attributes that could be used to re-identify, with a larger k meaning lower risk; refinements like l-diversity and t-closeness address weaknesses in k-anonymity on its own. None of these methods, however, is a separate legal standard. Each is a way of producing de-identified data that a qualified expert can then evaluate and, where warranted, certify under the expert determination method.

Being honest about the limits of these techniques is part of doing the work well. It is important to remember that, from the vendor’s side, every gain in protection tends to cost some of the data’s analytic or training value, so the job is one of calibration, not maximal scrubbing; data stripped past a certain point stop being useful for AI training or analysis. k-anonymity is vulnerable to homogeneity and background-knowledge attacks where the protected group shares a sensitive trait, or where an attacker already knows something about the person, which is why the refinements above exist. Synthetic and perturbed data can still leak information if the generation process is imperfect or the underlying model memorizes its inputs (inquiring about raw PHI data retention and not just modulated PHI-laden outputs in generative products is essential). And, as the next point makes clear, a determination is a conclusion about a particular data set, in a particular environment, at a particular time; it is not a permanent property of the data. Also, AI models drift and change, and what some PHI-bearing outputs may contain and count as de-identified in one moment may not be as clean and carry the low re-identification risk in the next.

In summary, expert de-identification results in data sets sitting outside HIPAA jurisdiction and purview. There is no return obligation to fall back to, because the obligation in HIPAA to return or destroy data at the end of a HIPAA business associate agreement governs PHI, and these de-identified data are, by their nature, no longer PHI. There is no purpose limit to outgrow (e.g., public health, research, health care operations), because the limited data set rules in 45 C.F.R. § 164.514(e) do not apply to data that are de-identified and therefore not a limited data set. And there is far less to report if there is a breach, because de-identified data are no longer the unsecured PHI that triggers the Breach Notification Rule.[23] Of all the legal tools surveyed here, expert determination is the one whose protections do not lapse when the engagement ends and where AI outputs passed on PHI are the least risky to allow for patient identification, and arguably lowers the risk and liability of both the healthcare entity and the AI vendor the most.

That said, this analysis and white paper is an assessment, not a directive. A carefully drafted DUA can lawfully preserve a limited data set for permitted purposes, and a BAA can properly govern PHI during a project; for some vendors and use cases, those tools are the right fit and not only necessary but sufficient. But for a vendor whose AI model depends on keeping and reusing healthcare data over time, expert determination is, on this analysis, the approach most built to last.

VI. What Expert Determination Costs, and Why

A defensible expert determination is a professional engagement and the price reflects that. Even fairly simple de-identification projects on small data sets often run into the tens of thousands of dollars at going rates for qualified statisticians and privacy professionals. There is not much standardized public pricing, partly because the work is bespoke and partly because, as the National Committee on Vital and Health Statistics has told the Secretary of Health and Human Services, expert determination is more consultative and more expensive than Safe Harbor, and there are relatively few qualified experts available to hire.[24]

Cost is also driven by the shape of the data (e.g., a spectrum from tabular to unstructured), its volume, the number and type of health conditions and document types in scope, how sophisticated the surrogate-generation or perturbation tooling is, the expert’s qualifications, and how much statistical analysis and documentation the work needs to hold up to later scrutiny. Unstructured narrative, which is often used for healthcare AI tasks, sits at the higher end, because the within-document and cross-repository problems described above take more analysis than clearing tabular fields.

Again, it bears repeating that an expert determination is not static. It is a conclusion about a specific data set in a specific environment, and as the underlying model drifts, as new data comes in, or as the surrounding data ecosystem changes, that conclusion may need to be revisited. An expert determination is best understood as a point-in-time assessment that should be re-checked periodically, not a one-time certificate. The opinions of these experts usually have expiration dates and must be renewed periodically. There are no hardline timeframes or periodic intervals recommended or required for when re-de-identification should be performed. From a diligence standpoint, the mix of expense and ongoing obligation to maintain de-identification is better read as a sign of seriousness than as a drawback. A vendor that has paid for a documented and maintained expert determination has done work that other vendors may not have, signaling a commitment to data privacy and security.

VII. A Practical Diligence Checklist: Questions That Reveal Risk

Faster is not the same as better. Many AI vendors compete mainly on speed: how quickly they can push health PHI into or out of an API, a portal, or an EMR connection, and generate the desired AI output (which is often PHI-laden). The real value of AI in healthcare comes from systems that can generate high-quality, legally defensible products in minutes instead of hours, but also from vendors that understand the safety and legal concerns with handling PHI and take precautions to offer high-quality products, which includes strategic thinking on how to protect not only themselves but the clients and patients they serve.

The questions below are ones that we believe health systems should increasingly put to AI vendors, and ones that, in our view, any vendor handling patient PHI data should be able to answer. We offer them as a framework, not a comprehensive checklist of requirements; how much weight each carries will depend on the engagement and the clinical data involved.

Category

Key questions to consider

Privacy & security

What does the vendor do to comply with the HIPAA and HITECH Privacy and Security Rules, and how are those steps documented and reviewed over time? Does the vendor have a HIPAA risk analysis, risk mitigation plan, detailed policies and procedures, workforce training, business associate subcontractor agreements, and what are the specific technology-stack layers when it comes to breach, monitoring, and software safeguards (including at-rest and in-transit encryption)? How does the vendor ensure compliance with applicable state privacy laws?

Regulatory history

Has the vendor faced any False Claims Act matter, OIG/OCR investigation, or other fraud-and-abuse inquiry, audit, enforcement action, or private litigation? Does the vendor regularly check its employees and subcontractors against state and federal exclusions databases?

PHI handling

What categories of PHI or other sensitive data does the system take in, and are those data transient or retained? If transient, to what extent? Are directly identifying PHI elements removed or scrubbed before the data are handled — and, in particular, are they scrubbed before any model or AI process touches them?

Data usage

Is client data or PHI used to develop, train, or improve models? Are prompts, outputs, or logs kept, and is there an option to decline use of the data for model development?

De-identification

Are de-identification or anonymization processes used, and by which method (Safe Harbor or expert determination)? For unstructured text, what supports the chosen method? Has an independent, qualified expert documented a determination, and what is its scope and date? Are synthetic surrogates or other obfuscation techniques used instead of redaction alone, and how is aggregate re-identification risk across the full database (and unrelated databases or users that may gain access) addressed?

Contract alignment

Do the BAA and DUA line up with downstream subcontractors and address limited data sets or research data sets where relevant? (This one often carries significant weight.)

Breach & security program

What are the breach-notification policies, security controls, and technical infrastructure, and is a security program and workforce training plan maintained?

Clinical oversight & quality

Are physicians or other clinicians involved in training, testing, and ongoing review to support medical accuracy?

Integrations & exposure

Does the system run in a closed, isolated environment, or does it connect to the open internet, the EHR, or claims systems? How are data transmitted, and how broad is the potential exposure if the system can reach outside networks?

Bias & fairness

What has been done to reduce racial, gender, disability, or other bias in generated letters or determinations?

Audit rights & governance

May the vendor be audited and required to provide security reports, and are model-update and change-management policies available for review?

Data ownership

If de-identification is done by expert determination, who owns the resulting de-identified data (intellectual property): the vendor, the health system, or both? And how is that allocation documented in the agreement?

Data retention & exit

What are the data-return and deletion policies? Is de-identified data sold or reused? What happens to data when the contract ends, and where return is said to be infeasible, on what basis? (This one often carries significant weight.)

VIII. The Emerging Question: Who Owns and Controls the De-identified Data?

One question is moving from the edges to the center of these negotiations, and it deserves to be named plainly. De-identified data have real and growing value. They can be kept, reused, combined, and used to build models without the constraints that attach to PHI, and that freedom is exactly what makes them valuable. As health systems have come to see that value, more of them are asking who should own the de-identified data a vendor creates from their patients’ PHI records: the vendor that did the de-identification and stands behind the determination, the health system whose patients and operations generated the underlying records, or both, under some shared or licensed arrangement. What rights might patients have in the use of their de-identified data? Are patients informed that their PHI will be disclosed to business associates and that AI tools might be used in connection with their PHI? Are there state laws in place that are more stringent than HIPAA that may require patients to consent to the uses and disclosures of PHI that HIPAA may otherwise allow?

The law does not give one answer. Ownership and permitted use of de-identified data are, for the most part, matters of contract. What the analysis above shows is only that this should be a deliberate decision after detailed analysis and not something resolved by default. A vendor’s expert determination establishes that the data may lawfully be kept and reused free of HIPAA’s constraints; it does not, by itself, decide who holds the intellectual property rights. Increasingly, parties are settling ownership, permitted uses, and any benefit-sharing of de-identified outputs in the underlying agreement, rather than discovering after the fact that a valuable asset was created without anyone having decided whose it is.

IX. Our Approach and Conclusion

I have written this from a vendor’s point of view because that is what our team is. We work in the appeals and prior authorization space, which runs on exactly the kind of unstructured clinical and insurance documentation where Safe Harbor offers little assurance and shortcuts are most tempting. For that reason, the data we use to track client appeals outcomes have been through a formal expert determination and are de-identified under 45 C.F.R. § 164.514(b)(1). The practical effect is that these data sit outside HIPAA: the data are not subject to a return obligation that could be invoked mid-engagement, are not bound by limited data set purpose limits a program might later exceed, and do not carry the breach exposure that retained PHI would. Nonetheless, we still explore, on a continual basis, how to invoke protections even where they are not strictly necessary, because, as health care executives and professionals, we understand not only the legal but the personal dynamics of PHI and sensitive-information leaks, and we take all steps to avoid them.

We describe our own approach not as the only defensible one, but as the one our analysis led us to, and the standard we are content to be measured against. Where the protections a party relies on must outlast the relationship that created them, the expert determination method is, on the analysis above, and after exploring all options ourselves, the one that best serves covered entities, vendors themselves, and, most importantly, patients.

Disclaimer

This white paper is provided for general information and reflects the views and analysis of the author, Eric Brooks. It is not legal advice and does not create an attorney-client relationship. De-identification, business associate, data use, and data ownership questions are highly fact-specific, and the authorities cited here are subject to change. Parties weighing these arrangements should consult qualified counsel about their particular circumstances.

  1. Radiology has been the earliest and largest area of clinical AI adoption: of the approximately 1,430 AI-enabled medical devices the U.S. Food and Drug Administration had authorized as of the FDA’s March 4, 2026 update, approximately three-quarters were radiology devices, with pathology an emerging area reflected in initiatives such as the FDA’s Pathology Innovation Collaborative Community. See U.S. Food & Drug Admin., Artificial Intelligence-Enabled Medical Devices (last updated Mar. 4, 2026) (available at https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices).

  2. In 2025, lawmakers in 47 states introduced more than 250 bills addressing the regulation of artificial intelligence in health care, of which 33 were enacted across 21 states. See Manatt Health, Health AI Policy Tracker (2025). See also Holland & Knight, State AI Health Tracker (current as of May 4, 2026) (mapping a broader and continually expanding set of state laws affecting the use of AI in healthcare, with state activity continuing to expand in 2026) (available at https://www.hklaw.com/en/general-pages/state-ai-health-tracker).

  3. Exec. Order No. 14,365, Ensuring a National Policy Framework for Artificial Intelligence, 90 Fed. Reg. 58,499 (Dec. 16, 2025). See also Exec. Order, Promoting Advanced Artificial Intelligence Innovation and Security (June 2, 2026) (continuing the administration’s federal-leadership approach to AI by directing agencies to strengthen AI-related cybersecurity and establishing a voluntary framework for Federal Government review of advanced “covered frontier” models) (available at https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/).

  4. Joint Commission & Coalition for Health AI, Guidance on Responsible Use of AI in Healthcare (RUAIH) (Sept. 17, 2025) (first national guidance from a U.S. accrediting body addressing responsible deployment, governance, and monitoring of AI tools by health care organizations, including recommended AI policies, governance structures, data-security and data-use protections, and vendor oversight) (available at https://digitalassets.jointcommission.org/api/public/content/dcfcf4f1a0cc45cdb526b3cb034c68c2).

  5. See Am. Med. Ass’n, Principles for Augmented Intelligence Development, Deployment & Use (2023); Am. Med. Ass’n, STEPS Forward: Governance for Augmented Intelligence (2025) (eight-step governance framework for health systems implementing AI). The AMA uses the term “augmented intelligence” to emphasize that such tools are meant to assist, not replace, clinician judgment.

  6. Courts have generally not treated standalone software as a “product” subject to products liability, so liability for an AI system’s role in patient care typically rests with the physician and health system under malpractice and negligence principles rather than with the developer; defects in physical hardware remain a separate matter. See Barbara J. Evans & Frank Pasquale, Product Liability Suits for FDA-Regulated AI/ML Software, in The Future of Medical Device Regulation: Innovation and Protection (I. Glenn Cohen et al. eds., 2022); Artificial Intelligence and Liability in Medicine: Balancing Safety and Innovation, Milbank Q. (2024); Are Current Tort Liability Doctrines Adequate for Addressing Injury Caused by AI?, AMA J. Ethics (Feb. 2019).

  7. Following Alice Corp. v. CLS Bank International, 573 U.S. 208 (2014), many software- and algorithm-based inventions face uncertain patent eligibility under 35 U.S.C. § 101, and copyright protects only expression rather than functional methods or underlying data. As a result, AI developers commonly protect model architectures, training data, and source code through trade secret law rather than through patents or copyright. See generally Defend Trade Secrets Act of 2016, 18 U.S.C. §§ 1836–1839; Unif. Trade Secrets Act (1985).

  8. 45 C.F.R. § 160.103 (defining “business associate” to include a person or entity that creates, receives, maintains, or transmits protected health information on behalf of a covered entity).

  9. See Health Information Technology for Economic and Clinical Health Act, Pub. L. No. 111-5, 123 Stat. 226 (2009) (codified as amended in scattered sections of 42 U.S.C.); 42 U.S.C. § 17934 (2018); 45 C.F.R. § 160.402 (establishing direct liability of business associates for applicable Privacy and Security Rule provisions).

  10. (45 C.F.R. pt. 164, subpt. C)

  11. See Notice of proposed rulemaking; notice of Tribal consultation; Office for Civil Rights (OCR), Office of the Secretary, Department of Health and Human Services; HIPAA Security Rule To Strengthen the Cybersecurity of Electronic Protected Health Information, 90 Fed. Reg. 898 (Jan. 6, 2025) (to be codified at 45 C.F.R. pt. 164); see 45 C.F.R. pt. 164, subpt. C (2024) (Security Rule).

  12. 45 C.F.R. § 164.504(e)(2)(ii)(J) (2024) (requiring return or destruction of protected health information at termination, and permitting retention only where return or destruction is infeasible, subject to extended protections and limited further use).

  13. Off. for C.R., U.S. Dep’t of Health & Hum. Servs., Guidance on HIPAA & Cloud Computing (2016) (explaining that return or destruction would be considered “infeasible” where, for example, other law requires the business associate to retain the electronic protected health information beyond termination of the contract) (available at https://www.hhs.gov/hipaa/for-professionals/special-topics/health-information-technology/cloud-computing/index.html).

  14. 45 C.F.R. § 164.514(e) ); see id. § 164.514(e)(2) (defining “limited data set”); id. § 164.514(e)(4) (specifying required terms of a data use agreement).

  15. See 45 C.F.R. § 164.514(e)(2) .

  16. 45 C.F.R. § 164.514(e)(3)(i).

  17. See Dept. of Health and Human Services Office for Civil Rights, Guidance Regarding Methods for De-Identification of Protected Health Information in Accordance with the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule (available at https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html).

  18. 45 C.F.R. § 164.514(b)(2) (Safe Harbor method); id. § 164.514(b)(2)(i)(R) (catch-all category encompassing “any other unique identifying number, characteristic, or code”).

  19. Id. at (b)(2)(i)(R).

  20. See 45 C.F.R. § 164.514(a) (health information that neither identifies an individual nor provides a reasonable basis to identify one is not individually identifiable health information).

  21. 45 C.F.R. § 164.514(b)(1) (expert determination method, requiring application of generally accepted statistical and scientific principles and documentation that the risk of re-identification is very small).

  22. See Off. for Civil Rights, U.S. Dep’t of Health & Hum. Servs., Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule (Nov. 26, 2012) (noting that the Privacy Rule does not fix a single numerical threshold for “very small” risk, which depends on the data and its environment) (available at https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html).

  23. 45 C.F.R. §§ 164.400–.414 (Breach Notification Rule); see id. § 164.402 (defining a reportable “breach” by reference to “unsecured protected health information,” which does not include properly de-identified information).

  24. See Nat’l Comm. on Vital & Health Stat., Correspondence to the Sec’y, U.S. Dep’t of Health & Hum. Servs. (Feb. 23, 2017) (observing that the expert determination method is more consultative and more expensive than Safe Harbor, and that relatively few qualified experts are available for hire) (available at https://www.ncvhs.hhs.gov/wp-content/uploads/2013/12/2017-Ltr-Privacy-DeIdentification-Feb-23-Final-w-sig.pdf).

RegulationAIBAADUAExpert DeterminationAI Vendor Selection StrategyHIPAAHITECH
Share
Copied