Playbook

Is Pasting a Client's Info Into ChatGPT a HIPAA Violation?

Nobody in your office decided to create HIPAA exposure this week. A producer just wanted a follow-up email drafted faster, and the fastest tool on their laptop happened to be the one with no Business Associate Agreement behind it.

Mike Moore pausing at a laptop beside an open client file with a padlock resting on the papers
The short version

Yes, it can be. OpenAI's own help center states that only ChatGPT Enterprise or Edu accounts on a sales-managed contract, and specific Zero Data Retention API endpoints, are eligible for a Business Associate Agreement1. The free tier, Plus, Team, and Business plans aren't. Cyberhaven's telemetry across 1.6 million employees found 8.6% had pasted company data into ChatGPT and 4.7% had pasted confidential data specifically, with client data among the three largest categories leaking2. HHS's Office for Civil Rights has resolved more than 370,000 HIPAA complaints since 2003, with civil penalties settled or imposed in 152 cases totaling over $144.8 million3. This guide covers the exact mechanism, the fix you can put in place this afternoon, and why a policy memo alone rarely holds past the first busy week of AEP.

The habit nobody flags

A client calls about a claim, mentions a diagnosis in passing, and asks the producer to follow up in writing. It's 4:40 on a Friday. The producer opens a new tab, types "write a warm follow-up email to a client who just told me about her breast cancer diagnosis and wants information on how her policy handles it," pastes in her name and the note from the call log, and gets back a clean paragraph in four seconds. The email goes out. Nobody in the office thought about it as a compliance event, because it didn't feel like one. It felt like using spellcheck.

That's the actual shape of this problem across the industry right now, not a hypothetical IT-security lecture. It isn't reckless staff or a rogue employee. It's a normal person doing a normal task with the fastest tool sitting open in another browser tab, one that happens to have no signed agreement behind it covering what happens to health information once it's typed in.

This guide is not about whether AI tools belong in an insurance agency. They plainly do, and this site has argued for their use in follow-up, call analysis, and appointment automation elsewhere. It's about one specific, narrow, fixable gap: the difference between a consumer AI product and one actually covered by the paperwork HIPAA requires when protected health information crosses from your systems into someone else's.

Who this applies to

This is squarely a Medicare, ACA, and health-line issue. A note that references a medical condition, a prescription, a diagnosis, or anything else on HIPAA's list of eighteen identifiers is what creates the exposure. A life insurance quote that only discusses coverage amount and beneficiaries carries a much lighter HIPAA analysis, though other statutes can still reach a health detail collected on an application.

What a BAA actually covers, and what it doesn't

A Business Associate Agreement is the contract HIPAA requires before a covered entity, or another business associate acting on its behalf, can hand protected health information to a vendor. It obligates that vendor to specific safeguards and specific breach-notification duties. Without one in place covering the exact product being used, there's no contractual floor under what happens to the data once it leaves your system, regardless of how careful the vendor's engineering team actually is.

This is exactly the same structural point this site has made about GoHighLevel: signing a BAA with a platform covers what happens inside that platform's own systems. It does not retroactively cover a different product, a different account tier, or a script you added yourself4. A BAA is scoped to a specific product and a specific agreement, not to a company's name in general.

OpenAI is explicit about where that line sits for its own products. Its help center documentation states that a BAA is available only to "ChatGPT Enterprise or Edu customers that have a sales-managed account," and separately to API customers, but only for "endpoints eligible for Zero Data Retention"1. Read that plainly: the free tier almost every producer already has open in a browser tab isn't on the list. Neither is Plus, the $20-a-month personal upgrade an agent might have bought with a company card to get faster responses. Neither, per that same documentation, are the Team or Business tiers, despite names that sound like they were built for exactly this use case.

OpenAI's own stated BAA eligibility, by product
Product BAA eligible? What that means for a client note
ChatGPT Free No No contractual coverage for PHI at all.
ChatGPT Plus No Same as Free. A personal subscription does not add coverage.
ChatGPT Team / Business Not listed as eligible1 A paid business plan is not the same thing as a BAA-covered plan.
ChatGPT Enterprise / Edu Yes, on a sales-managed account1 Requires a direct, negotiated agreement with OpenAI, not a self-serve signup.
API, Zero Data Retention endpoints Yes, for eligible endpoints1 Covers a developer-built integration, not a producer typing into a chat window.

Why the free and Plus versions are the problem

It isn't that OpenAI is careless with the free and Plus products. It's that those products are built, priced, and governed as consumer software, and consumer software runs on consumer terms. Absent an enterprise agreement, a consumer account's content can be used to improve the underlying models unless the account holder has manually opted out in settings, and there is no signed instrument establishing OpenAI's obligations as a business associate over that specific data. Nothing about that is unusual for consumer software. It's exactly how a free product is supposed to work. The problem is only that health information doesn't belong inside a system built on those terms.

The gap almost never gets closed on purpose. Almost every agency we've audited has a written policy somewhere about client confidentiality generally. Very few have a policy that names ChatGPT, or Claude, or Gemini, or any other consumer AI product specifically, and fewer still have checked which internal accounts are actually running which tier. The producer using the free version on her personal laptop and the one running an Enterprise seat through the agency's IT department are having, from a compliance standpoint, completely different experiences, and most agencies don't know which one is actually happening on a given afternoon.

No signed BAA

Free, Plus, Team, and Business tiers carry no Business Associate Agreement, per OpenAI's own eligibility rules.

Training use by default

Consumer accounts can feed conversation content into model training unless the user opts out in settings, which most never touch.

Identifiers hide in context

PHI doesn't require a name. A date, a diagnosis, and a route number can be enough to identify someone in a small book of business.

No visibility for the owner

Most agency owners have no inventory of which staff run which AI account tier, so the gap is invisible until an incident surfaces it.

The mechanism in one sentence. A BAA is a contract about a specific product and account tier, not a blanket promise from a company name, and the products producers actually have open on a normal Tuesday, the free tier and the personal Plus upgrade, are the ones OpenAI's own documentation places outside that contract.

The eighteen identifiers that make a note PHI

It helps to know exactly what the rule is actually protecting, instead of relying on a gut sense of what "sounds private." HHS's own guidance on de-identification lays out eighteen specific identifiers under the Safe Harbor method; a record with any one of them still attached to health information counts as protected health information, not just a name attached to a diagnosis9.

The 18 HIPAA Safe Harbor identifiers (45 CFR 164.514(b)(2))
# Identifier Shows up in a follow-up note as
1NamesThe client's first or last name
2Geographic subdivisions smaller than a stateA street, county, or the ZIP code beyond the first three digits
3Dates tied to an individual (except year); ages over 89A call date, a birth date, a diagnosis date
4Telephone numbersThe number pulled from the CRM to draft the message
5Fax numbersRare in this industry, but still on the list
6Email addressesPasted in so the draft has a "To" line
7Social Security numbersOccasionally copied from an enrollment record
8Medical record numbersA claim or case reference number
9Health plan beneficiary numbersThe Medicare Beneficiary Identifier or policy number
10Account numbersA policy or billing account number
11Certificate or license numbersRare, but possible on a business or professional line
12Vehicle identifiersA VIN, relevant on the auto side of a book
13Device identifiers and serial numbersA medical device serial number mentioned on a call
14Web URLsA link to a client's own portal or claim page
15IP addressesRarely typed manually, but present in some exported logs
16Biometric identifiersNot applicable to a typical agency workflow
17Full-face photographsA photo attached to an ID verification step
18Any other unique identifying number, characteristic, or codeA distinctive detail, like a rare condition in a small town, that narrows down who's being described

Look at that list next to the Friday-afternoon email from the opening of this guide. The call log note carried a name, a call date, and a diagnosis, three separate identifiers from that list attached to health information, before it ever reached the paste command. None of that required carelessness. It required only that the producer treat the note the way she'd treat any other piece of text she needed help phrasing.

What this actually costs

Cyberhaven Labs, a data-loss-prevention vendor, analyzed ChatGPT usage across 1.6 million employees at companies running its own product. It found that 8.6% of employees had used ChatGPT at work and pasted company data into it, and narrower than that, 4.7% had specifically pasted confidential data2. This is one vendor's own telemetry, not an independent industry-wide audit, and it's worth saying so plainly since it's the only dataset of its kind we could verify this session. It's also the most direct, primary-source measurement available of how often this actually happens rather than how often it's assumed to happen.

Broken down by category, the same analysis found sensitive or internal data appeared in 319 incidents per 100,000 employees weekly, source code in 278, and client data in 260, making client data one of the three largest categories of confidential material leaking into the tool2. The report also found the volume of confidential data reaching ChatGPT climbed 60.4% over roughly a six-week window it tracked, and that just 0.9% of employees were responsible for 80% of the highest-risk sharing events2, meaning this risk concentrates in a small number of habitual users rather than spreading evenly across a staff.

8.6%

Of employees have pasted company data into ChatGPT2

4.7%

Have pasted confidential data specifically2

0.9%

Of employees generate 80% of the highest-risk sharing events2

$144.8M

In HIPAA civil penalties settled or imposed since 20033

Put a rough number on your own office. Say your agency runs 15 producers and support staff who touch client files. If Cyberhaven's 4.7% confidential-paste rate held exactly across your team, roughly one person on staff has already pasted something confidential into a consumer AI tool. That's a hypothetical applied to an industry-wide telemetry figure, not a claim about your specific office, but it's the right order of magnitude to reason about: this is not a one-in-a-thousand edge case, it's closer to one person on a normal-sized team.

Confidential data incidents per 100,000 employees, weekly, by category 319 Sensitive / internal 278 Source code 260 Client data
Source: Cyberhaven Labs, analysis of ChatGPT usage across 1.6 million employees, fetched 2026-08-222.

The regulatory backdrop behind that risk is real and already produces real penalties. HHS's Office for Civil Rights has received more than 374,321 HIPAA complaints since the Privacy Rule's 2003 compliance date, resolved 99% of them, and settled or imposed civil monetary penalties in 152 cases for a combined total exceeding $144.8 million, as of the agency's own October 31, 2024 tally3. Most complaints don't end in a fine. The same page shows 67,873 cases closed with early technical assistance and 15,561 that found no violation at all3. That's genuinely reassuring in one sense and beside the point in another: OCR's own numbers show it investigates a very large volume of complaints, and a health-detail-bearing note pasted into an uncovered consumer product is exactly the kind of disclosure that generates one.

The mistake we see most

An owner asks whether the agency "has an AI policy" and gets told yes, because someone wrote a paragraph about client confidentiality two years ago that never mentions ChatGPT by name. A policy that predates the tool it's supposed to govern isn't a policy on this question. It's a document that happens to exist.

If you want a fast read on where your own site and stack stand first The free Audit checks a site's technical and AI-citation readiness, including HIPAA tracking exposure on the website side, in under a minute. It won't inventory which AI chat accounts your staff run, but it's the same no-pressure starting point we'd point you to before any bigger conversation. strategicaiarchitects.com/audit6.

The DIY fix, step by step

You can meaningfully close most of this gap yourself, today, without buying anything beyond what your agency may already own.

01

Write the policy in plain, specific language

Not "protect client confidentiality," which nobody disagrees with and nobody changes their behavior over. Name the eighteen HIPAA identifiers, name the tools (ChatGPT, Claude, Gemini, Copilot, and whatever else staff actually use), and state plainly that none of those identifiers go into a consumer account, full stop.

02

Inventory what's actually running today

Ask every producer and support staffer, directly, which AI tools they use for work and on which account tier. You will very likely find a mix of personal Plus subscriptions, free accounts, and possibly nothing formal at all. You cannot fix exposure you haven't located.

03

Turn off training on every consumer account, as a floor, not a fix

Every major consumer AI product has a setting to opt conversation content out of model training. Turn it on everywhere it exists. This does not create a BAA and does not make the product HIPAA-eligible. It is the minimum backstop while you build something better, not a substitute for it.

04

Get an actual signed BAA if you're going to keep using ChatGPT directly

If the agency wants to keep working inside ChatGPT itself, that means moving to an Enterprise or Edu account on a sales-managed contract, per OpenAI's own published eligibility rules1, not staying on Plus or Team and hoping. Contact OpenAI directly and confirm the current terms before treating any plan as covered.

05

Train on redaction as a daily habit, not a one-time memo

Teach staff to write "the client who called Tuesday about a claim" instead of a name, a date of birth, and a diagnosis in the same sentence. It's a real reduction in risk even on a covered account, and it's the only lever available on an uncovered one.

Why a policy memo alone doesn't hold

Every step above genuinely reduces exposure, and an agency that does all five is in a meaningfully better position than one that does none of them. What it doesn't do is change the underlying condition that created the risk in the first place: a human being, under time pressure, has to remember a rule every single time, with no system checking the work.

That's why this category of policy tends to decay in a specific, predictable pattern. It holds fine in a slow week. It gets circulated, people nod, a few even read it. Then AEP arrives, call volume triples, a producer is drafting her twentieth follow-up email of the day at 6:15 p.m., and the fastest path back to her family is the browser tab already open, not the one behind a slower, more deliberate enterprise login. The policy didn't fail because anyone rejected it. It failed because it depended on a memory check surviving exactly the conditions most likely to defeat it.

Policy alone

Depends on memory, every time

  • Relies on every producer recalling the rule under time pressure
  • No system checks whether an identifier was actually typed in
  • Degrades fastest exactly when call volume is highest
  • Invisible to the owner until an incident surfaces it

Failure modeQuiet, and usually discovered after the fact

Architecture

Doesn't depend on anyone remembering

  • Identifiers are aliased before they ever reach a non-BAA destination5
  • Real values are re-hydrated only on the way back, inside the covered system
  • The same protection applies at 6:15 p.m. in October as it does in June
  • Built on the agency's own accounts, so it doesn't rely on staying current with a vendor's changing terms7

Failure modeThe unsafe path is engineered not to exist

How we build this instead

This section sticks to what's published on our own live pages and Ambrose's own documentation, verified this session. Ambrose, the AI infrastructure this site is built around, documents something called the PHI Rail specifically for this problem. In its own words, "the PHI Rail aliases identifiers before any non-BAA destination sees them, and re-hydrates real values on the way back"5, and the platform describes itself as "HIPAA-aware by default"5.

Read plainly, that means the decision about what's safe to send where doesn't sit on a producer's judgment in the middle of a busy afternoon. It sits in the plumbing. A note that would otherwise carry a name, a date of birth, and a diagnosis to an outside system gets its identifying pieces swapped for safe placeholders before it ever leaves a BAA-covered boundary, and the real values only come back once the response returns inside that same covered boundary. The producer still gets AI-drafted speed. The identifying details never actually make the trip to an uncovered destination.

When we scope a custom AI build, whether that's conversational follow-up, call analysis, or database reactivation, it runs on the agency's own accounts, its own domain, and its own CRM, and the agency owns it7. That's the same structural point this site has made about voice infrastructure: a system built on infrastructure you actually control, rather than a shared consumer product whose terms can change without your agency in the room, starts every compliance question from a stronger position. We're not going to promise a specific audit outcome or a guaranteed absence of risk, because no vendor legitimately can. What's on the live page is what we're building toward: PHI-aware infrastructure by default, not a policy memo hoping nobody's in a hurry.

Not every agency needs a custom build to make real progress here today. Digital Foundation's tiers include a compliant website build with consent-gated tracking baked in from $247 a month8, which handles the client-facing side of HIPAA exposure this site has covered elsewhere. It doesn't touch what your staff paste into a chat window on their own laptops. That's specifically a custom-build conversation, scoped on a call rather than sold as a fixed package7.

What you actually get

Concretely, moving AI-assisted drafting and follow-up onto infrastructure built around PHI-aware routing gets you three things. Producers keep the speed of AI-drafted responses, because the tool is still doing the writing. Identifying details stop traveling to a destination with no BAA behind it, because the architecture strips them out before the trip rather than trusting a human to remember to. And the protection holds on the busiest week of AEP exactly as well as it holds in a quiet week in June, because it isn't riding on anyone's memory in the moment.

None of that is a promise about a specific audit outcome, a guaranteed absence of a complaint, or a certification. HIPAA compliance is a posture you maintain, not a status you achieve once. It's an infrastructure decision about where the risk actually sits, and it's worth treating it like one before an incident makes the decision for you.

When the DIY path is genuinely fine

Say this plainly: a small agency writing life and annuity business almost exclusively, with a genuinely tight, well-trained staff and a real written policy that names the tools by name, may reasonably decide the DIY steps above are enough for now. The HIPAA exposure on a life-only book is lighter to begin with, and a small team can enforce a habit more reliably than a large one.

The math in this guide starts mattering once you're writing Medicare or ACA business, running more than a handful of producers, or running that volume across more than one office. At that scale, the number of people who have to remember the rule correctly, every time, under AEP-level time pressure, is exactly the population Cyberhaven's data says will occasionally get it wrong2. If you've already caught a near-miss, a producer showing you a drafted email with a client's name and a diagnosis in the same paragraph, that's the signal the policy-only approach has already started to crack.

Questions agencies ask

Is ChatGPT itself a HIPAA violation to use?

No tool is automatically a violation on its own. The violation happens when a covered entity or its business associate discloses protected health information to a vendor without a signed Business Associate Agreement covering that specific product. OpenAI's own help center states that only ChatGPT Enterprise or Edu customers on a sales-managed account, and API customers using Zero Data Retention endpoints, are eligible for a BAA. The free tier, Plus, Team, and Business plans are not covered. If a producer pastes a client's name alongside a health detail into one of those uncovered products, that's the exposure, not the software category itself.

Does it matter if I remove the client's name before pasting?

It helps, but HIPAA's definition of protected health information is broader than a name. Eighteen identifiers count, including dates tied to an individual, phone numbers, and any detail specific enough that the person could reasonably be identified from context. A note that says 'the diabetic client on Route 4 who called about his A1C results Tuesday' can be identifying even with no name attached, especially in a small agency's book of business. Redaction reduces risk. It doesn't eliminate the underlying question of which product you're pasting into.

My agency uses ChatGPT Team. Doesn't that count as business use?

It counts as a paid business plan, but per OpenAI's own eligibility rules that isn't the same as BAA-eligible. As of this session's fetch, only Enterprise or Edu accounts on a sales-managed contract, and specific Zero Data Retention API endpoints, qualify. Team and Business plans, despite the name, aren't listed as eligible. Confirm current status directly with OpenAI before assuming a paid plan closes the gap.

What actually happens to data pasted into the free or Plus version?

OpenAI's consumer products can use conversation content to improve their models unless a user has opted out in settings, and the account sits outside any BAA. That means the data isn't necessarily contained the way a covered entity's own systems are, and there's no signed agreement establishing OpenAI's obligations as a business associate for that data. The mechanism matters more than any single account's opt-out setting: nothing formally binds a consumer product to HIPAA's rules on that data.

Is this actually enforced, or just theoretical risk?

HHS's Office for Civil Rights has resolved over 370,000 HIPAA complaints since 2003, and civil monetary penalties have been settled or imposed in 152 of those cases, totaling more than $144.8 million. Most complaints don't end in a penalty; the overwhelming majority get resolved through corrective action or technical assistance. That's exactly the profile of exposure worth closing quietly, before it becomes one of the complaint statistics instead of a policy question.

What's the fastest fix if I can't rebuild anything right now?

Write a one-page policy today: no client name, date of birth, health condition, medication, or anything else on the eighteen-identifier list goes into a consumer AI tool, full stop. Circulate it, get a signed acknowledgment from every producer, and turn off model training in account settings as a backstop. It's not a permanent fix, since it depends on every person remembering every time, but it closes the gap this afternoon while you plan something more durable.

How is Ambrose's PHI Rail different from just being careful?

Care depends on a human remembering, every single time, under time pressure, which detail is safe to type. Ambrose's PHI Rail is described in its own documentation as aliasing identifiers before any destination without a BAA ever sees them, then re-hydrating the real values on the way back. The safety doesn't live in a producer's judgment in the moment; it lives in the architecture, so the same protection applies on a rushed Tuesday afternoon in October as it does on a slow one in June.

Sources

  1. OpenAI Help Center. "How can I get a Business Associate Agreement (BAA) with OpenAI for the API Services?" BAA eligibility limited to ChatGPT Enterprise/Edu sales-managed accounts and Zero Data Retention API endpoints, checked live 2026-08-22. help.openai.com.
  2. Cyberhaven, Inc. "4.2% of workers have pasted company data into ChatGPT," analysis of 1.6 million employees, 8.6% workplace usage / 8.6% company-data pasting, 4.7% confidential-data pasting, category breakdown (sensitive/internal 319, source code 278, client data 260 incidents per 100,000 employees weekly), 60.4% six-week growth, 0.9% of employees responsible for 80% of high-risk events. Originally published 2023-02-28, updated 2025-05-19, verified live 2026-08-22. cyberhaven.com.
  3. U.S. Department of Health and Human Services, Office for Civil Rights. "Enforcement Highlights," 374,321 complaints received since April 2003, 99% resolved, civil monetary penalties settled or imposed in 152 cases totaling $144,878,972, data as of 2024-10-31, page last updated 2024-11-21, verified live 2026-08-22. hhs.gov.
  4. Strategic AI Architects. "Your GoHighLevel BAA Doesn't Make Your Funnel HIPAA Safe," the scoped-coverage principle behind a product-specific BAA, published 2026-08-19. strategicaiarchitects.com.
  5. Ambrose. "What is Ambrose" documentation, PHI Rail description ("aliases identifiers before any non-BAA destination sees them, and re-hydrates real values on the way back") and "HIPAA-aware by default" claim, verified live 2026-08-22. app.hiambrose.com.
  6. Strategic AI Architects. "Free Audit," verified live 2026-08-22. strategicaiarchitects.com.
  7. Strategic AI Architects. "AI Expert," custom build ownership terms ("you describe it, we build it, you own it") and scoped-on-a-call pricing, verified live 2026-08-22. strategicaiarchitects.com.
  8. Strategic AI Architects. "Digital Foundation," Starter tier pricing at $247/month with compliant website build, verified live 2026-08-22. strategicaiarchitects.com.
  9. U.S. Department of Health and Human Services, Office for Civil Rights. "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule," the 18 Safe Harbor identifiers under 45 CFR 164.514(b)(2), verified live 2026-08-22. hhs.gov.

Talk it through

Want a second pair of eyes on it?

Free 30 minutes. Bring what you found, or bring nothing and we will look together at how AI engines read your site and which fixes move first.

Find out where your own stack actually stands

Run the free Audit, a live AEO Audit plus a HIPAA tracking scan of your site, in under a minute.

Related reading: why your GoHighLevel BAA doesn't make your funnel HIPAA safe · the difference between an NDA and a BAA · why your insurance website's pixel could get you sued

← All guides