Before you buy an AI sales tool, ask it to show you a lost deal
Most AI sales tools are good at showing what they did. Almost none clearly show what a lost deal actually changed. Go7's audit of 100 products found the gap, and Carter Stewart's five-minute demo test tells you how to check for yourself.
Here's a test that costs nothing and takes five minutes: the next time a vendor demos an AI sales or revenue tool for you, ask them to mark a deal lost for a specific reason, then show you exactly what changes.
What you're looking for is a real change, a different account score, a different recommended next step, for a comparable account, because of that loss. A dashboard updating or a new report generating doesn't count.
"I would ask the vendor to show what changes after a deal is marked won or lost," said Carter Stewart, founder of Go7. "Take two comparable accounts. Mark one related opportunity as lost for a specific reason. Show me how that changes the system's account score, recommended next action, messaging, channel choice or targeting for future accounts. Then show me the audit trail: what changed, which evidence caused it, and whether a human can review or override it."
Most AI sales tools will fail that test. That doesn't mean they're broken. It means "learning" and "reporting" get marketed as if they're the same thing, and they aren't.
Go7, a UK company building its own AI client-acquisition platform, ran an audit to find out how big that gap actually is. It reviewed the public pages of 100 marketed AI sales and revenue products and scored three things: whether the product could act on its own, whether it could connect that action to a real commercial outcome, and whether it used that outcome to change what it did next. Because Go7 competes in the same category, it excluded itself from its own scoring and published its full source list so the findings could be checked.
The results: 69 products clearly showed they could take action. Thirty-two clearly showed that action connected to a commercial result. Only eight clearly showed that result feeding back into a future decision.
That gap between 69 and 8 is the whole story.
Table of contents
- Two different claims, and where most tools stop
- Where the market splits, four segments, one clear pattern
- What to actually ask a vendor
Two different claims, and where most tools stop
A tool that sends emails, makes calls, or updates a CRM on its own has proven it can act. That's not the same as proving the action led anywhere.
Carter drew the line precisely. "We treated action and attribution as separate capabilities," he said. To earn credit for action, a product's public material had to make autonomous execution, outreach, calls, CRM updates, lead routing, central to the pitch. To earn credit for attribution, it had to go further and explicitly connect that activity to a qualified opportunity, pipeline, forecast, win rate or revenue.
The audit's hardest call was 11x Alice, an autonomous prospecting tool that talks openly about creating pipeline. Go7 still scored its attribution as partial, not core, because promising pipeline isn't the same as showing how specific actions turned into specific qualified opportunities. A vendor saying their agent "creates pipeline" is usually making an action claim wearing attribution's clothes.
The same gap shows up at the market level. Ninety-one of the 100 products had at least some evidence of attribution somewhere on their public pages, but only 32 made it a core, explicit capability. Fifty-five products lead their positioning with activity, efficiency, replies or meetings. Twenty-two lead with revenue, win rate or forecast outcomes.
"The market is very good at measuring what the system did," Carter said. "It is less consistent at showing what that activity ultimately changed for the business."

Where the market splits, four segments, one clear pattern
Go7 sorted the 100 products into four segments, and the gap between action and learning is concentrated almost entirely upstream.
Sales engagement and outbound execution (25 products). All 25 strongly evidenced action. None strongly evidenced learning.
Agentic SDR and revenue agents (30 products). 28 of 30 strongly evidenced action. None made outcome learning an explicit, core capability.
Revenue and conversation intelligence (20 products). The exception. Fourteen of 20 strongly evidenced attribution, and seven strongly evidenced learning, the strongest showing of any segment. The tradeoff: only two of the 20 showed any capability for discovering new accounts in the first place.
Only eight products across the entire sample clearly showed the full loop, action connected to a result, and that result changing a future decision:
Revenue and conversation intelligence (7 of 8):
- Gong
- Clari
- Backstory
- Aviso
- Attention
- Sybill
- MeetRecord
Prospecting and signal intelligence (1 of 8):
- Warmly, the only product from its 25-product segment to make the list
Seven of the eight came from the segment that already scored best on learning overall.
"The recurring pattern was an operating loop rather than a single AI feature," Carter said. A system that notices one subject line got more replies is doing immediate campaign feedback, useful, but shallow. A system that notices an account was lost for a specific reason and changes how it treats a different, similar account next quarter is doing something else entirely.
Worth flagging: "eight" isn't a count of which products can learn from outcomes, it's a count of how many clearly showed it on the pages Go7 reviewed.

What to actually ask a vendor
Go through the demo test at the top of this piece, and treat the answer as a sorting mechanism, not a pass/fail.
If the only visible change after a lost deal is a new dashboard number, a CRM field, or a retrospective report, that's reporting, not learning. If the vendor can show a specific, traceable change, a different score, a different recommended action, for a comparable account, and can walk you through the audit trail of why, that's the real thing. Ask who can review or override that change, too.
Go7's audit only scored what companies said publicly, not what their products can actually do behind a login. That's exactly why the demo test matters more than the marketing page. The gap between 69 and 8 doesn't tell you which tools work. It tells you which vendors made the harder, less flattering claim easy to verify, and which ones are hoping you won't check.


