Checklist / AI operations

How to Evaluate an AI SDR Without Mistaking Activity for Pipeline

Evaluate an AI SDR on the quality of the selling work it completes, not the amount of activity it can generate. The scorecard should start with eligible accounts and end with sales-accepted opportunities, while making factual errors, inappropriate messages, missed replies, and human correction visible.

Sending more emails is easy to count. It is also several steps removed from revenue. An AI system can produce impressive activity while targeting the wrong companies, inventing personalization, mishandling objections, or handing sales a pile of conversations nobody accepts.

I would treat an AI SDR as an operating system with five jobs: choose a plausible account, identify a relevant contact, prepare an accurate message, interpret the response, and move a qualified conversation to the right human. Test those jobs separately before judging the entire motion.

This article presents my evaluation framework, not a reported client result or industry benchmark. The worked numbers are hypothetical. At Stackmatix, the growth agency I co-founded, the commercial context is qualified pipeline and customer economics; that is why this scorecard follows the work past the send button.

1. Define the job and the authority you are granting

“AI SDR” can mean a research assistant, a message-drafting tool, an automated sender, or an agent that reads replies and books meetings. Those are materially different systems.

Write one sentence that states the job, the allowed inputs, the permitted actions, and the point of human review. For example:

Example operating boundary: For approved accounts in the U.S. software segment, the system may research public company and role information, draft one first-touch email using approved claims, and recommend a response; a sales representative approves every send and owns all live conversations.

The example is deliberately bounded. If the system can send without review, change CRM fields, schedule meetings, or continue a sequence after a reply, list each action. Then identify what stops it. The AI action-approval guide helps make those boundaries inspectable.

NIST's voluntary Generative AI Profile recommends managing risks across governance, context, measurement, and ongoing management rather than treating evaluation as a one-time test. Use that as a broader risk-management reference, not as a certification shortcut. NIST AI 600-1: Generative AI Profile.

2. Build the denominator before you review the output

Do not start with messages sent. Start with the accounts the system was allowed to consider.

Create a fixed evaluation set that includes good-fit accounts, obvious non-fits, ambiguous cases, stale or conflicting public information, contacts with different responsibilities, and accounts that must be excluded. Preserve the decision and evidence for each case before the test begins.

Use this eligibility ladder:

StageRequired evidenceFailure to record
Account eligibleSegment, geography, size, use case, and exclusion rules passWhy the account entered the test
Contact plausibleRole and responsibility fit the buying problemA title match with no likely ownership
Trigger supportedA current source supports the outreach reasonA guessed initiative or stale announcement
Message permittedChannel, suppression, policy, and approval rules passA technically sendable message

Keep excluded accounts in the evaluation report. Correctly refusing to contact a bad-fit or prohibited account is useful work, even though it reduces activity.

Use the ideal customer profile template for the account criteria and the outreach brief for the message itself. If a human cannot apply those rules consistently, automation will make the ambiguity faster.

3. Score factual accuracy at the claim level

A message can sound plausible while its central reason for contacting someone is wrong. Review each externally checkable claim, not just the message as a whole.

For every claim, record the source URL or approved internal record, the observation date, whether the source actually supports the wording, and the consequence if it is wrong. Separate facts from inferences. “The company lists three offices” may be observable. “The company is struggling to coordinate those offices” is usually an inference unless the prospect has said so.

I would use four labels:

  • Supported: the claim follows from a current permitted source.
  • Overstated: the source is real, but the message goes beyond it.
  • Unsupported: no permitted evidence establishes the claim.
  • Unverifiable: the necessary evidence is unavailable or conflicting.

Any invented customer, funding, hiring, technology, or performance claim should be a serious failure, not a minor copy edit. Average quality scores can hide a small number of commercially damaging messages.

4. Evaluate relevance separately from writing quality

Fluent copy is not necessarily useful outreach. Ask whether the message connects a plausible buying situation to a truthful offer for this recipient.

Score these questions independently:

DimensionReviewer question
Account fitIs this a company the sales team would pursue without the automation?
Contact fitIs this person likely to own, influence, or redirect the problem?
Problem relevanceDoes the opening connect to a real situation rather than a generic compliment?
Offer accuracyCan the product actually deliver what the message promises?
Evidence qualityDoes any proof support this use case and remain within approved wording?
Next-step fitIs the request proportionate to what the recipient knows?

Blind the copy review when practical: let a reviewer score the account, contact, and evidence before seeing response data. Otherwise a positive reply can make weak research look better in hindsight.

5. Test reply handling with adversarial cases

A positive-response demo is not enough. Build reply cases that require different actions:

  • Clear interest with a concrete question.
  • Referral to another person.
  • Timing objection or request to follow up later.
  • Product, security, or pricing question the system cannot safely answer.
  • Explicit opt-out or complaint.
  • Ambiguous response, sarcasm, or wrong-recipient notice.
  • Out-of-office response and automated system message.

For each case, define the correct classification, whether the sequence must stop, what should enter the CRM, and when a person must take over. Measure missed opt-outs and inappropriate continuations separately; do not average them into a general reply score.

For U.S. commercial email, the FTC's CAN-SPAM guidance covers accurate headers and subject lines, identification and address requirements, and honoring opt-out requests. The company remains responsible for compliance even when another provider handles sending. This checklist is not legal advice; requirements vary by jurisdiction and channel, so have qualified counsel review the actual program. FTC: CAN-SPAM compliance guide.

Delivery policy is another release gate, not evidence of prospect interest. Google's current sender guidance requires authentication for mail sent to personal Gmail accounts and adds further requirements for bulk senders. Postmaster Tools exposes spam reports, authentication, reputation, and delivery errors for Gmail traffic. Review the requirements that apply to your sending pattern before launch. Google: email sender guidelines and Postmaster Tools.

6. Make sales acceptance the pipeline gate

A reply, booked meeting, or CRM opportunity created by automation should not automatically become pipeline. Define the minimum evidence a sales owner needs to accept it.

A practical acceptance record includes:

  • The account matches the agreed customer profile.
  • A plausible business problem or evaluation goal is present.
  • The contact can participate in or route the buying process.
  • The next step has a purpose beyond “take a demo.”
  • A human owner accepts responsibility for follow-up.
  • Duplicate, partner, recruiting, support, and vendor conversations are excluded.

Report both system-proposed opportunities and sales-accepted opportunities. Track the rejection reason for every proposed opportunity. That feedback distinguishes a targeting failure from a reply-classification problem or an overly loose opportunity rule.

The sales qualification template can supply the evidence required after a conversation begins. Do not force a cold prospect to meet a late-stage qualification framework; use only the fields needed to justify the next sales action.

7. Work through a hypothetical pilot

Suppose an AI SDR evaluates 500 accounts in a four-week pilot. These figures are hypothetical and are not recommended performance targets.

ResultCountRate from relevant prior stage
Accounts evaluated500
Accounts judged eligible32064.0%
Accounts approved for outreach after review28087.5%
Accounts with at least one delivered message26895.7%
Accounts with a human response249.0%
System-proposed opportunities1145.8%
Sales-accepted opportunities654.5%

The top-line activity story could be “268 accounts contacted and 24 responses.” The commercial story is six accepted opportunities from 500 evaluated accounts: 1.2% of the starting set. Neither rate is good or bad without deal value, sales capacity, cost, a comparison group, and later outcomes.

Now add the quality record. Imagine reviewers found four unsupported personalization claims before approval, two misclassified replies after sending, one opt-out that the system tried to answer, and 18 hours of research, review, and correction work. Those hypothetical findings matter to the release decision even if no bad message ultimately reached a recipient.

Compare the pilot with the existing process using the same audience definition and observation window. If possible, randomize eligible accounts between the two approaches. If volume is too small for a useful causal test, describe the result as a bounded operational pilot and preserve the uncertainty.

8. Calculate cost per accepted opportunity

Include software, data, sending infrastructure, implementation, review, correction, deliverability work, and sales follow-up. Then divide the total pilot cost by sales-accepted opportunities—not emails, replies, or automatically created records.

For a hypothetical example, assume the four-week pilot costs $4,000 in software and data plus 60 hours of team work at a loaded planning rate of $75 per hour. Total modeled cost is $8,500. With six accepted opportunities, modeled cost per accepted opportunity is about $1,417.

That is not customer acquisition cost. The opportunities still need to progress, close, retain, and produce enough gross profit. Compare the result with the customer acquisition cost framework and track later outcomes by original motion.

9. Use this release checklist

Before expanding audience, volume, or autonomy, answer each item with evidence:

  1. Is the allowed job and action boundary explicit?
  2. Does the evaluation set include non-fits, ambiguity, stale data, and exclusions?
  3. Are factual claims supported at the claim level?
  4. Are serious errors reported separately from average copy quality?
  5. Do opt-outs, complaints, and ambiguous replies stop or escalate correctly?
  6. Has a responsible owner reviewed applicable law and channel policy?
  7. Can a salesperson reject a proposed opportunity with a recorded reason?
  8. Are accepted opportunities tracked through later sales and customer outcomes?
  9. Does the cost include human review, correction, and maintenance?
  10. Is there a rollback trigger and a person authorized to pause the system?

Expand one dimension at a time: a broader audience, more volume, or less review. If you change all three, you will not know why quality moved.

The point of an AI SDR is not to make the activity dashboard move. It is to create qualified selling capacity without degrading trust, data quality, deliverability, or unit economics. If the system cannot pass that test at a controlled scope, more volume only scales the part you still need to fix.

Co-founder and CEO of Stackmatix, startup advisor, and former Head of Sales at MightyHive. · More about Matt →