Gong Scorecards Are Killing Your Reps' Best Conversations
Ana's deal-saving call scored a 61. The rubric measures resemblance, not selling — and your reps are learning to choose the rubric.
2026-07-22 · SARA — KEEL'S AI DEAL ASSISTANT · GETKEEL.IO
The call that saved the Meridian deal scored a 61. Ana knows the number because her manager opened the scorecard in their 1:1: talk ratio too high, discovery questions below benchmark, no next steps captured in the final two minutes. What the rubric couldn't see: the customer spent forty minutes venting about their incumbent, Ana let them, and the security review that had been stuck for a month got unstuck that afternoon. Meridian signed nineteen days later. The 61 is still in the system.
A scorecard doesn't measure the call. It measures the call's resemblance to a template.
Strip the branding off any call-scoring system and the mechanism is the same: someone defines what a good call looks like — question count, talk ratio, topics touched, methodology stages hit — and the software grades every recorded conversation against that shape.
That's not measurement of quality. It's measurement of conformity, and the difference is the entire problem. A call can match the template and move nothing. A call can violate every criterion and win the quarter. The scorecard cannot tell these apart, because outcomes live weeks downstream and context lives off the recording.
Gong's scorecards are the flagship version, which is why the backlash carries the brand's name. But the critique applies to the whole category of AI sales scoring: grade conversations at scale, and you have decided — structurally, whatever the rollout deck said — that resemblance is what you want more of.
Your best calls are outliers by definition
Ask any veteran seller about the conversations that won their biggest deals and listen to what they describe. The forty-minute vent. The demo that got abandoned because the real problem surfaced in minute five. The call where the rep said "honestly, we might not be the right fit this year" and earned three years of trust.
Every one of those calls scores badly. They're off-template because the situations were off-template — and big deals are usually off-template situations.
A scoring system tuned to the center of the distribution treats deviation as deficiency. Run that system over a sales floor for a year and you haven't improved the average call. You've taught your range to collapse toward the rubric — including the range that closes.
The collapse is audible. Sit in on a scored team's calls for a week and notice how similar the discoveries sound — same openers, same question ladders, same recap cadence. Enablement calls that consistency. Customers, who take four vendor calls a week, call it something else: interchangeable.
When the score becomes the target, the customer stops being the audience
There's an old rule about metrics: when a measure becomes a target, it stops measuring. On a scored sales floor, the rule has a specific and expensive shape.
The rep on a scored call is serving two audiences at once — the customer, and the rubric. Every seller under a scorecard knows the moves: ask the checklist question you already know the answer to, restate next steps for the transcript, keep the customer's tangent short even when the tangent is the deal. The call gets better as an artifact and worse as a conversation.
We named the feeling this produces in the Gong anxiety piece — and the anxiety is about the scorecard, specifically. Recording alone is memory. Scoring is judgment, automated, applied to every sentence, with the score traveling farther than the context ever does.
Watch it happen live on any scored call. The customer says something surprising — a reorg, a budget wobble, a competitor mention — and there's a beat where the rep chooses: follow the thread, or protect the score. Following the thread means abandoning the discovery sequence, blowing the talk ratio, ending without clean next steps. The thread is where the deal is. The sequence is where the grade is. You've built a machine that makes your reps choose, forty times a week, and you've weighted one side.
The score has an afterlife the call never gets
Here's the mechanic that makes scoring different in kind from recording: the score is portable and the context isn't.
Ana's 61 will outlive every detail of the Meridian call. It rolls into her monthly average. The average feeds a team benchmark. The benchmark shows up in QBR slides, in comp conversations, in the spreadsheet someone opens when headcount decisions get made. At every step, the number sheds a little more of the situation that produced it, and there is no field anywhere in the pipeline for "the customer needed forty minutes and I gave it to them."
Reps understand this instinctively, which is why the backlash reads as visceral rather than procedural. Nobody minds being seen. People mind being summarized — by something that wasn't in the room, for readers who never will be.
The research backs the reps, not the rubric
This is where the rep backlash stops being anecdote. In 2024, Cornell researchers published four experiments on algorithmic monitoring in Communications Psychology — nearly 1,200 participants. People surveilled by AI complained at more than four times the rate of people monitored by humans, reported losing autonomy, generated fewer ideas, and thought more about quitting.
The detail that should stop every enablement leader mid-rollout: the negative effects eased when monitoring was framed as developmental rather than evaluative. A scorecard is the evaluative frame made permanent. It doesn't ask what the rep was trying to do. It renders a number, and the number goes in the dashboard.
You can't rollout-message your way around that. Reps don't experience the intent. They experience the artifact — the number next to their name, refreshed weekly, in a system they never asked for. Whatever the kickoff deck promised about growth and development, the artifact says evaluation, and the artifact wins.
The invasive part isn't the recording. It's judgment without appeal.
When reps call the scoring layer invasive, managers hear an objection to being recorded and reach for the consent script. That misses the actual complaint.
A recording is a fact. A score is a verdict — and in most deployments it's a verdict with no rebuttal field. The 61 on Ana's Meridian call doesn't come with a place for "the customer needed to vent and I read the room." It sits in the system, feeds the team benchmarks, and surfaces in whatever review her VP runs at quarter end. Judgment at scale, context at zero.
That's the surveillance texture reps are describing when the word comes out. Not "my company can hear me." Rather: "something that doesn't understand my deal grades every sentence I say, and people who weren't there read the grades."
The scorecard also can't see where your deals actually happen
There's a quieter absurdity underneath all of this. The scored call is a fraction of the relationship — and for field and relationship motions, the smaller fraction.
We've made the full argument in what call recording can't capture: the customer lunch, the hallway conversation, the drive-home realization — the highest-signal moments in a B2B deal produce no audio and therefore no score. A rep can be quietly running a masterclass across the unrecorded 80% of a relationship and grade as mediocre on the recorded slice.
Which means the scorecard isn't even a biased measure of selling. It's a precise measure of a thin sample, mistaken for the whole.
Stack the legal layer on top — the states where every scored call also requires every participant's consent — and the machinery starts to look like what it is: an enormous apparatus built around the one part of the deal that's easiest to capture, rather than the parts that decide it.
Managers wanted coaching. They bought grading.
None of this indicts the manager who deployed scorecards. The problem they had was real: too many calls, too little time, no way to know where coaching would help.
But look at what good coaching requires — context, dialogue, trust, timing — and compare it to what a scorecard supplies: a number, asynchronously, attached to a clip. The scorecard was supposed to be a triage tool for coaching. In practice it becomes the coaching, because the number is right there and the forty-minute call review isn't.
The tell is what happens in 1:1s on scored floors. The conversation drifts from "what's going on in the Meridian deal" to "let's look at your numbers this month." The rep learns the real skill being developed: producing better numbers. Deals become the thing that happens between reviews.
There's a version of call review that coaches: manager and rep watch the call together, the rep narrates what they were reading in the moment, and the conversation is about judgment, not adherence. Nothing about that requires a score. The score is what you add when you want the review to happen without the manager — which is to say, when you want the appearance of coaching at the price of none.
Run the audit before you defend the rubric
If you lead a scored floor and this reads as unfair, there's a fifteen-minute test. Pull your five biggest closed-won deals from the last two quarters. Find the pivotal call in each — the one the rep will name if you ask. Look up how it scored.
Then pull five clean, high-scoring calls from deals that went nowhere, and sit with the comparison. Most leaders who run this exercise honestly meet the same pattern Ana's manager could have seen: the correlation between score and outcome is far weaker than the dashboard's confidence implies — and sometimes it points the wrong way, because the calls that required real selling were precisely the ones that broke template.
Your reps have already run this audit informally. Every scored floor has a quietly circulating list of great calls with bad numbers. That list is doing more to shape how reps treat the scorecard than any enablement session ever will.
Scored floors select for rubric players
Give the system a few years and it stops being a measurement choice and becomes a hiring filter. The reps who thrive under scoring are the ones comfortable optimizing to visible criteria — a real skill, but a different skill from reading a room and departing from script when the deal demands it.
The closers with options notice the difference first. In interviews, "how does your team use call scoring?" has become a question candidates ask, and the answer moves decisions. A scored floor doesn't just narrow the range of its current reps. It quietly pre-filters for the narrow range in its next ones — the compounding version of the same mistake.
What the fix looks like — and it isn't a better rubric
The instinct is to tune the scorecard: more criteria, smarter AI, outcome-weighting. That road doesn't end, because the defect isn't calibration. It's direction. A grading system points at the rep. A rep points at the deal. No amount of tuning turns the first into the second.
The actual fix inverts the ownership. Give the rep a thinking partner that works on her questions — what did I miss, what's the champion not saying, what does the 9pm pricing email actually signal — and keeps the answers private. Reps aren't waiting for a kinder scorecard. They've started routing around scorecards entirely.
The irony worth sitting with: the honest self-assessment managers hoped the scorecard would force is exactly what reps refuse to do inside anything that scores them. Privacy isn't the enemy of accountability here. It's the precondition for the only accountability that improves anyone — the rep's own.
The number is not the conversation
If your best reps are quietly muting themselves to have the honest conversations that actually close deals — your scorecard is the reason. The fix isn't another scorecard. The fix is a thinking partner for the rep, not a grading system for the manager. That's Sara. Founders Club is invite-reviewed. Apply at getkeel.io/founders.
Sara is a 24/7 AI deal assistant with no scoring anywhere in her: the rep talks the deal through after the conversation — recorded or not — and what's said stays between them. She asks the questions a great coach would ask, minus the audience.
Keep Ana's 61 as the mental model. Somewhere in your pipeline right now is the call that will save your biggest deal, and your scoring system is prepared to file it as a failure. The conversation your revenue depends on and the conversation your rubric wants are not the same conversation — and every quarter, your reps are being taught to choose the rubric's.
By the team at Keel. We're building Sara, an AI deal assistant for the moments that don't get recorded.