Playbook
The Sales Calls Your Agency Records and Never Reviews
CMS makes you record and store every Medicare sales call for years. Almost nobody at your agency ever listens back. Here is what that gap costs, and how to close it without adding headcount.
Your agency already has the raw material for producer coaching. It just never gets used. CMS requires TPMOs, including independent agents, to record Medicare sales calls, so the audio already exists. But only 8.29% of independent agencies say they use AI regularly and strategically, and 44% still run staff training informally, peer to peer1. The recordings sit in storage for compliance and never turn into a coaching loop. This guide covers why that happens, what a realistic worked example says it costs in manager time, how to build a manual review habit yourself, and how an automated call analysis layer turns the same recordings you already keep into a scored coaching queue.
The calls nobody hears again
Somewhere on a server your agency pays for, there is a recording of the call where your newest producer fumbled the same objection your best producer handles without thinking. There is another one where a caller asked a question about network coverage and got a shaky, half-right answer instead of the confident one that closes. There is a third where the disclosure language ran fast and mumbled, the kind of thing that reads fine on a compliance checklist and sounds bad the moment CMS's own reviewer plays it back.
All three recordings exist because CMS requires them to. All three will sit untouched until the day, if it ever comes, that a complaint or an audit pulls them back up. Nobody at your agency is going to sit down this week and listen to them for what they actually are: a free, already-paid-for record of exactly how your team sells.
Meanwhile, the same producer keeps fumbling the same objection, call after call, because nobody ever told them it was happening. A caller who might have enrolled hangs up confused instead. A manager who could have caught the pattern in five minutes of listening never gets the chance, because five minutes of listening is not the part of the job anyone tracks, rewards, or schedules. The problem is not that your producers are bad at their jobs. It is that the one piece of information that would make them better already exists and nobody has built a path from the recording to the coaching conversation.
That is the gap this guide is about. Not whether you record calls, you already have to. Not how long you have to keep them, that is a retention question this site has covered elsewhere. The gap is between the compliance archive you are required to build and the coaching system you never got around to building on top of it, even though the raw material for both is the exact same file.
What this guide is and is not
This is about turning sales call recordings you are already required to keep into a coaching and quality signal for your producers. It is not the retention rule itself, and it is not the AI receptionist question this site has covered separately. Every figure here is pulled from a primary source fetched while writing this, named inline.
Why the review never happens
The honest answer is capacity, not neglect. A manager who wants to review calls has to carve real, uninterrupted time out of a day that is already full of their own selling, their own service work, and every other fire an agency owner has to put out. Listening to a call carefully enough to score it, not just skim it, takes longer than the call itself once you account for pausing to take notes.
Work the arithmetic plainly. Say a manager spends 12 to 15 minutes listening to and scoring one call well. A small team running 60 sales calls a week, which is a modest volume across three or four producers during a normal month, would need 12 to 15 hours of a manager's undivided attention to review every single one. That is close to two full workdays, and it assumes the manager does nothing else that week. It is illustrative arithmetic, not a published statistic, but the shape of it holds at almost any agency size: full manual review does not scale with a growing book.
So agencies do what capacity allows. A manager overhears part of a call while walking past a desk. A new hire gets shadowed for their first week and then is on their own. Feedback becomes whatever a manager happens to catch rather than a system that looks at the whole picture. None of that is a knock on any specific agency. It is what happens when the only tool for review is a person's limited attention applied to a growing pile of recordings.
There is a second, quieter reason review never happens, and it has nothing to do with time. Nobody at most agencies actually owns it. Compliance owns producing the recording if CMS or a plaintiff's attorney ever asks for it. IT or the CRM vendor owns the fact that it gets stored somewhere for the required window. Sales management owns hitting the month's numbers. None of those three roles has "turn this archive into a coaching system" written into their job, so it falls into the gap between them, which is exactly where things that are nobody's explicit job tend to die.
Compare that to how a strong sales floor actually runs. The best sales organizations do not treat call review as an afterthought squeezed into whatever time is left. They build it into the week on purpose, with an owner, a cadence, and a rubric, precisely because they know that without an assigned owner, "we should really be listening to more calls" stays a sentence people say in a meeting and never becomes something anyone actually does.
What the industry data actually shows
A 2026 national survey of independent agencies by the Big "I" Agents Council for Technology found that just 8.29% of independent agencies use AI regularly and strategically in their operations, even though 68% plan to increase their AI use over the next twelve months1. That gap between intent and actual practice is the whole story of this guide in one number. Agencies know something needs to change. Almost none of them have closed the loop yet.
The same report found 56% of agencies have no written AI use policy at all, and 44% still rely on informal, peer-to-peer training to bring staff up to speed on new tools and new approaches1. Peer-to-peer training is exactly how bad habits spread as easily as good ones. If your best producer never gets recorded and studied, and your newest producer only learns by watching whoever happens to be free that week, the agency has no actual mechanism for making everyone sound like the best person on the team.
It is not only an agency-side gap. On the carrier and enterprise side of the same industry, adoption is much further along. A Deloitte Center for Financial Services survey of 200 US insurance executives, fielded in June 2024, found that 76% had already implemented generative AI in one or more business functions, with life and annuity insurers slightly ahead at 82% against 70% for property and casualty2. Forty-five percent of leaders in that same survey said the benefits of gen AI already outweigh the risks2.
| Metric | Figure | Source |
|---|---|---|
| Independent agencies using AI regularly and strategically | 8.29% | Big "I" ACT, 2026 Tech Trends Report1 |
| Independent agencies planning to increase AI use | 68% | Big "I" ACT, 2026 Tech Trends Report1 |
| Independent agencies with no written AI policy | 56% | Big "I" ACT, 2026 Tech Trends Report1 |
| Insurance executives whose company implemented gen AI | 76% (82% life and annuity, 70% P&C) | Deloitte Center for Financial Services, June 20242 |
Where your agency actually stands
The free Audit checks your site's technical and AI-citation readiness in about a minute. It will not score your call coaching, but if your agency is behind on one modern system it is worth checking whether it is behind on more than one. strategicaiarchitects.com/audit6.
What a scored call actually catches
A rubric only works if it looks for specific, repeatable things rather than a vague sense of whether a call "went well." Across the calls we have listened to while scoping builds for agencies, the same handful of patterns show up again and again, and every one of them is something a reviewer, human or automated, can catch on a first pass.
Disclosure timing
Required disclosure language rushed, mumbled, or placed after the caller was already committed instead of before.
Shallow needs discovery
The producer jumps to a plan recommendation before actually asking what the caller's doctors, drugs, or budget require.
The first real objection
Whether the producer engages the caller's actual hesitation or repeats the same pitch louder.
An uncertain answer
A guess delivered as fact, on network coverage, drug cost, or plan rules, that a caller has no way to catch in the moment.
No clear ask
The call ends without the producer asking for the appointment, the enrollment, or a specific next step.
A missed renewal or cross-sell cue
The caller mentions a life event, a spouse, or another line of coverage, and the producer never follows up on it.
None of these require a compliance background to spot. They require someone, or something, to actually listen for them on a consistent basis, which circles back to the capacity problem from the last section. A rubric like this is also exactly the kind of structured, repeatable check that a scoring system can apply to every call rather than the handful a manager gets to.
What the review gap costs
Put a real dollar figure under the illustrative math from earlier. The US Bureau of Labor Statistics reports the median annual wage for insurance sales agents was $60,370 in May 20243, which works out to roughly $29.02 an hour using the standard 2,080-hour work year. That figure is a stand-in for the value of an hour of selling or managing time at a typical agency, not a claim about any specific role's salary.
Now run the worked example from the section above through that hourly figure. A manager who spends 12 to 15 hours a week trying to keep up with full manual call review, on top of their actual job, is spending roughly $348 to $435 a week of fully loaded time on a task they can only ever complete a fraction of. And that is the cost of attempting full review. Most agencies never attempt it, so the real number is not the hours spent. It is the hours never spent, and the bad habits, missed disclosures, and inconsistent scripts that keep compounding because nobody caught them.
| Input | Value | Basis |
|---|---|---|
| Sales calls per week (illustrative team of 3 to 4 producers) | 60 | Illustrative, modest agency volume |
| Minutes to listen and score one call carefully | 12 to 15 | Illustrative, realistic review pace |
| Hours needed to review every call that week | 12 to 15 | Calculated: 60 calls x 12 to 15 minutes |
| Hourly value of that reviewer's time | $29.02 | BLS median annual wage, insurance sales agents, May 2024, ÷ 2,0803 |
| Weekly cost of attempting full review | $348 to $435 | Calculated: hours x hourly value |
A five-minute check on your own agency
Pull one recorded call from this week, any producer, at random. Listen to the first three minutes only. Note whether the disclosure language was clear, whether the producer handled the first real objection well, and whether you would want a new hire to sound exactly like that call. If you cannot remember the last time you did this, that is the actual state of your review system, whatever your compliance binder says.
How to fix it yourself
You do not need software to start closing this gap. You need a repeatable habit, which is worth more than any tool if the tool never gets used. Here is the method, in full, so an agency that would rather run this manually has everything it needs.
Pick three calls a week, not zero and not all of them
Full review does not scale, as the worked example above shows, but zero review is the default most agencies fall into. Three calls a week, chosen deliberately rather than whatever a manager happens to catch, is sustainable and still creates a real signal over a quarter.
Build a one-page rubric before you listen to anything
Score four or five things consistently: was the disclosure clear, did the producer handle the first real objection, did they ask for the appointment or the sale, and did they sound like someone the caller trusted. Without a rubric, review turns into a vague impression instead of a comparable score.
Weight the sample toward new producers and toward complaints
A producer in their first ninety days should get reviewed more often than one who has been closing consistently for two years. Any call connected to a complaint, a cancellation, or an unusual objection should always make the list, whether or not it is someone's assigned week.
Deliver the feedback inside a week, not a quarter
Feedback on a call from six weeks ago does not change behavior, because the producer cannot remember the specific moment you are describing. A short note or a five-minute conversation within the same week the call happened is what actually shifts how the next call goes.
Keep a running log, not just individual notes
Track scores over time per producer, even in a simple spreadsheet. One low score is a bad call. A pattern of the same low score three months running is a training gap, and you cannot see the pattern if every review lives in a separate, forgotten note.
Run that five-step loop for a quarter and you will already be ahead of most agencies, because you will have turned some fraction of your compliance archive into an actual coaching signal. An agency owner who reads this and decides to run it by hand is making a reasonable choice. The honest limit is the same one from the cost section: a human reviewer can sustain a sample, not full coverage, and the calls that never get sampled are the ones where the real problems hide.
If the manual version is working for you
Some agencies are small enough, and quiet enough on the phone, that a manager genuinely can keep up with every call. If that is you, the five steps above are the whole system. You do not need to build anything more than that just because it exists.
One more thing worth planning for if more than one person ends up reviewing calls: two reviewers using the same rubric will still drift apart over time unless they calibrate against each other occasionally. Once a quarter, have both reviewers score the same three calls independently and compare notes on where their scores diverged. Otherwise a producer's score ends up depending on which manager happened to review them that week, which defeats the point of having a rubric in the first place.
Why this matters more heading into AEP
Medicare's Annual Enrollment Period runs October 15 through December 7 every year, and it is the single window when most agencies run the highest call volume, add the most temporary or new-to-industry help, and put the most pressure on every producer to move fast. It is also exactly the window when an unreviewed bad habit does the most damage, because whatever a producer is doing wrong in September gets repeated across the largest number of calls they will make all year.
A gap that costs you a handful of missed appointments in a quiet month costs you dozens of them across seven weeks of AEP volume, and a disclosure line that runs sloppy on one call in September is the same disclosure line running sloppy on hundreds of calls once volume peaks. The agencies that build a review habit before AEP starts catch and correct these patterns while the stakes are still small. The ones that wait find out what was wrong from a complaint, a low quality score, or a CMS inquiry, after the volume has already made the problem bigger.
The window to fix this is now, not in October
Whatever review system you are going to run this AEP, manual or automated, works best started before the volume hits, not mid-season when everyone is already at capacity. Late August and September is exactly the right time to pick a rubric, assign an owner, or scope a build.
How we build it instead
This section sticks to what is stated on our own live pages and in our own service descriptions, verified while writing this. Among the custom AI builds we scope for agencies is call analysis and producer coaching, built alongside the other AI agents in an agency's actual workflow: an AI receptionist that answers after hours, conversational follow-up that keeps context on every contact, database reactivation, and appointment automation. Everything is built on your agency's own accounts, your own domain, and your own CRM, and you own it, and it keeps working whether or not you keep working with us4.
The mechanism is the same recordings you already keep for CMS. Instead of a manager sampling three calls a week by hand, an automated call analysis layer transcribes and scores every recorded call against a rubric, then surfaces the ones that actually need a human's attention, a missed disclosure line, a fumbled objection, a producer who is drifting from the script that closes. The manager's limited coaching time goes to the calls that need it instead of whichever ones they happened to overhear.
That same context layer connects to the rest of the pipeline, not just the phone. If a producer's follow-up texts and a caller's plan questions already live in one memory per contact, the calls you review are not isolated data points, they are one more channel feeding the same picture of how a specific relationship is going5.
For a Medicare, ACA, or health-focused agency, there is a compliance layer under this that a bolt-on transcription tool rarely thinks about. A system that touches recordings with health details connects to your CRM and other systems, and every connection is a place protected information can leak. Our builds are engineered so that handling is HIPAA compliant by design, with coded, safe information passing between systems and raw protected data never held where it should not be. We say HIPAA compliant plainly, and we do not promise a compliance outcome, because the responsible way to say it is to build it correctly and let the architecture speak.
We are not going to promise you a close rate, a retention number, or a revenue figure, because those depend on your market, your book, and the people having the conversations. What we will say, and what is on the live page, is that you describe it, we build it, and you own it, on your accounts and your domain4.
Where this fits with the rest of what we build
If you are also looking at a website rebuild or a content cadence at the same time, this is the kind of thing worth scoping in the same conversation rather than as three separate projects. Book a call and we will walk through what applies to your agency specifically.
What you actually get
Concretely, moving from occasional manual spot-checks to a build like this gets you three things. Every recorded call gets scored, not just the handful a manager had time for. The calls that actually need a human's attention rise to the top instead of getting buried in the ones that were fine. And the pattern across a producer's calls becomes visible over weeks, not something you notice for the first time when a complaint arrives.
Underneath those three, you get the quieter benefit that matters most as your book grows: the recordings CMS already requires you to keep stop being a cost center that only exists in case of an audit, and start being an asset your best producers are effectively teaching from, whether they know it or not.
There is also a training benefit that compounds beyond any single quarter. Once a year or more of scored calls exists, a new hire's first weeks stop being guesswork built on whichever senior producer happens to have time to shadow them. You can point them to actual examples of a well-handled objection, a clean disclosure, and a real appointment ask, pulled from your own agency's own calls, not a generic training script written for nobody's specific book. That library is worth more the longer it exists, and it only exists once review stops being occasional.
This is not the retention rule
It is worth being precise about what this guide is not, because the two topics get mixed together easily. CMS's marketing rules already require TPMOs, including independent agents and brokers, to record the full audio of Medicare sales and enrollment calls and retain them for a period CMS sets, and that retention requirement recently changed for contract year 2027. That is a separate, fully sourced topic covered in our guide to the call recording rule change, including exactly what shifted and when.
This guide assumes you are already compliant with that requirement, because most agencies are, and asks the next question: given that the recordings already exist, why do they almost never get used for anything besides sitting in storage. The retention rule tells you how long to keep the file. It says nothing about whether anyone ever listens to it, and that gap is where the actual coaching opportunity, and the actual compliance risk, both live.
Questions agencies ask
Does my insurance agency have to record its sales calls?
For Medicare and Part D marketing, sales, and enrollment calls, yes. CMS requires third-party marketing organizations, which includes independent agents and brokers, to record the full audio of those calls and keep them on file for a retention period CMS sets and has recently changed. Our separate guide covers the exact retention window and what shifted for 2027. This guide is about a different problem: what happens to those recordings after they are archived.
How much of a manager's time does reviewing calls for coaching actually take?
Worked at a realistic pace, roughly 12 to 15 minutes to listen to and score one call carefully. A team running even 60 sales calls a week would need 12 to 15 hours of a manager's undivided time to review every one, which is more than a full workday and does not include the manager's other job. That math is why most agencies fall back to reviewing whatever a manager happens to overhear.
Can AI actually review every sales call instead of a sample?
Yes, that is the specific capability gap between where most agencies sit today and where the technology already is. An automated call analysis layer can transcribe and score every recorded call against a rubric, the same recordings CMS already requires you to keep, and surface the ones that need a human's attention instead of asking a manager to guess which handful to listen to.
Does automated call review replace human coaching?
No, and it should not try to. What it replaces is the guessing about which calls need attention. A manager who gets a scored shortlist of the three calls with the worst objection handling this week can spend their limited coaching time on those three calls instead of hoping they happened to catch a bad one live.
What is the difference between reviewing calls for compliance and reviewing them for coaching?
Compliance review asks whether a call violated a marketing rule, a disclosure requirement, or a script mandate, usually after the fact and often only when a complaint triggers it. Coaching review asks whether the producer handled the objection well, whether they built rapport, and whether they are getting better over time. Most agencies only do the first, and only reactively, because CMS requires the recording to exist. The second almost never happens at all.
Does automated call analysis create new HIPAA or compliance exposure for a Medicare or health agency?
It can, because any system that touches recordings containing health details is another place protected information can leak if it is not built correctly. The safe pattern is a build where handling is HIPAA compliant by design, with coded, safe information passing between systems and raw protected data never held where it should not be. That is an infrastructure decision, not a setting you toggle on a vendor tool.
What does building this cost?
A custom call analysis and producer coaching build is scoped on the call, because the right shape of it depends on your call volume, your CRM, and your compliance line. Strategic AI Architects builds it on your accounts, your domain, and your CRM, and you own it, so it keeps working whether or not you keep working with them.
Should every producer get reviewed the same amount?
No. Weight review toward the producers who need it most: anyone in their first ninety days, anyone connected to a recent complaint or cancellation, and anyone whose recent scores show a pattern rather than a one-off bad call. A producer who has closed consistently for two years with clean scores can be sampled far less often than someone still learning the job.
Sources
- Big "I" Agents Council for Technology (ACT). "2026 ACT Tech Trends Report," national survey of independent insurance agencies, released 2026-02-19: 68% plan to increase AI use in the next 12 months, 8.29% use AI regularly and strategically, 56% have no written AI use policy, 44% rely on informal peer-to-peer training. Verified live 2026-08-21. independentagent.com.
- Deloitte Center for Financial Services. "Scaling Gen AI in Insurance," survey of 200 US insurance executives (100 life and annuity, 100 property and casualty), fielded June 2024: 76% have implemented gen AI in one or more business functions (82% life and annuity, 70% property and casualty), 45% of leaders say benefits outweigh risks. Verified live 2026-08-21. deloitte.com.
- U.S. Bureau of Labor Statistics. "Occupational Outlook Handbook: Insurance Sales Agents," median annual wage of $60,370 in May 2024 for SOC 41-3021, used here to derive an hourly figure of approximately $29.02 by dividing by a standard 2,080-hour work year. Verified live 2026-08-21. bls.gov.
- Strategic AI Architects. "Your Ambrose and AI Expert," custom AI build scope including call analysis and producer coaching, database reactivation, and voice AI, with ownership terms "you own it" and "everything is built on your accounts, your domain, and your CRM." Verified live 2026-08-21. strategicaiarchitects.com.
- Strategic AI Architects. The Playbook, "Your CRM Remembers Nothing: Why AI Context Is the Real Differentiator for Insurance Agencies," on the single-memory-per-contact architecture referenced here. Verified live 2026-08-21. strategicaiarchitects.com.
- Strategic AI Architects. "Free Audit," live AEO Audit plus a HIPAA tracking scan. Verified live 2026-08-21. strategicaiarchitects.com.
Talk it through
Want a second pair of eyes on it?
Free 30 minutes. Bring what you found, or bring nothing and we will look together at how AI engines read your site and which fixes move first.
See what a custom build would actually cover
Run the free Audit, a live AEO Audit plus a HIPAA tracking scan of your site, in under a minute.
Related reading: what CMS's new call recording rule requires before AEP · why AI context is the real differentiator for insurance agencies · what a virtual assistant costs your insurance agency