Questry AIBack to the knowledge base

Knowledge base

Can AI recognize a winning sales email? What the SDR-Bench study shows

A new study shows the best AI models find about half of what a sales rep puts into a winning email. The other half is not intelligence. It is memory.

Luuk van de Ven (co-founder) · 7 September 2026 · Lees dit artikel in het Nederlands

Everyone who uses AI for sales runs into the same question sooner or later: can the model actually write a good sales email? Not an email that sounds good, but one that leads to a meeting. In May 2026 a study came out that measures exactly that, using real sales emails from real companies.

The outcome is surprisingly clear. The best models find about half of the arguments a rep uses in a winning email. And at a large tech company, no model could tell the difference between an email that led to a meeting and one that was ignored. This article explains what the researchers did, what they found, and what it means for companies that want to use AI for outreach.

What the researchers did

The study, Benchmarking the Personalization Capabilities of Large Language Models, is built around a simple idea. If a sales email demonstrably led to a meeting, then that email contains the winning strategy: the pain points and arguments that convinced the customer. An AI is good at personalization if, without ever seeing that email, it comes up with the same arguments.

The researchers collected two kinds of evidence. First, 6,279 public customer success stories from large companies across 22 industries, each describing a customer's problem and how a product solved it. Second, real sales emails from two companies: a healthcare firm and a Fortune 100 tech company, more than 100,000 emails from 124 reps in total. For every email it was known whether it led to a conversation.

The AI was then given three things per case: the product, the company being approached, and the date. The task: come up with the arguments you would use. The AI was allowed to search the web, but only information that existed on that date. That way the model could not quietly look up the success story itself. Finally, the AI's answer was placed next to the real winning email and scored: what share of the real arguments had the model also found?

Which models were tested

These are the models companies actually use today: Claude Sonnet 4.6, GPT-5.4, GPT-5.4-mini, GPT-4o and Qwen 2.5, all with web search. On top of that, three so-called deep-research agents, which run multiple search rounds and process far more text before answering.

Result 1: the best AI finds about half

On the public success stories, Claude Sonnet 4.6 scored highest with 56 out of 100. The other models landed between 35 and 45. In other words: the best model finds slightly more than half of the arguments a human puts into a winning pitch. The rest is missing.

What stands out is that more compute does not help. A bigger model, more searches, or a deep-research agent that processes tens of thousands of words: the scores stay stuck at the same level. The researchers call this the personalization plateau. The problem is not the model's reasoning ability.

Result 2: at the tech company, the AI saw no difference

The most important test was in the real sales emails. For each company, the researchers took 200 emails that led to a meeting and 200 that were ignored. The AI had to come up with arguments for both groups. If the model understands what works, its arguments should look more like the winning emails than the losing ones.

At the healthcare company this worked a little: overlap with winning emails was 32 percent, with ignored emails 22 percent. At the Fortune 100 tech company there was no difference. The AI's arguments resembled the emails that worked just as much as the emails that ended up in the bin. In the researchers' words: no model could statistically separate successful from unsuccessful outreach.

The explanation is straightforward. A large tech company receives hundreds of emails a week with the same obvious arguments. The email that wins does not win on logic but on something specific: a detail only that rep knew. And that detail is not online anywhere.

Result 3: in practice, half was immediately usable

The researchers had twelve professional sales reps use the system in their normal work, for more than 200 prospects. Of all generated arguments, they rated 48 percent as immediately usable without rewriting. Senior reps with at least ten years of experience judged the output independently, and their verdict closely matched the automated score.

So the researchers' conclusion is not that AI is useless for sales. Their advice is human-in-the-loop: let the AI supply building blocks, and let a human decide what goes in.

Why a human finds more than the AI

This is the part that is most often misunderstood. The human is not smarter than the model. The human knows more. The AI only got what was online: the website, job postings, news, the annual report. From that it derives logical arguments, and that is the half it does find.

The rep who closed the deal knew things that were never published. What a similar customer said last year. Which argument worked at this prospect's competitor. Which proof convinces this industry and which proof triggers distrust. That is the other half, and it cannot be found by searching better. It lives in experience and memory.

That is why a better model does not help. The missing part is not intelligence. It is information that only exists in your own deal history.

What this means for AI in outreach

Anyone using AI for cold email runs straight into this plateau. A tool that only reads a prospect's website and turns it into an email produces the half that everyone produces. In a crowded market, that is the half that gets ignored.

The difference is in the system around it. First, in who you email: a sharp selection of companies that resemble your best customers (a well-defined ICP and TAM, in Dutch) beats a long list of generic emails. Second, in what you remember: every reply, every meeting and every lost deal contains information that makes the next email better, provided you capture that information and reuse it.

This is why Questry is built around Open Brain, a self-learning memory. Every dossier is built from dozens of sources, including your own past emails and conversations. Open Brain picks who to approach and writes the email based on what worked before at similar companies. Every deal makes the next campaign sharper. And just as the researchers advise, the human stays in the loop: the AI writes and books the meeting, you confirm.

Source

Srivastava, A. et al. (2026). Benchmarking the Personalization Capabilities of Large Language Models. arXiv:2607.20471. The dataset (SDR-Bench) and the test environment (SDR-Arena) are publicly available, so new models can be re-tested.

Curious how Questry AI automates this for you?

See the plans