All articles
Template

AI sales agent checklist: 27 questions to ask before you buy (no demos required)

April 27, 2026Updated June 24, 202618 min read3,520 words

Judge an autonomous outbound tool on what you can verify in writing, not in a demo: it keeps data clean, takes real actions with stop rules, scores fit and intent separately, protects deliverability, and runs behind approvals and audit logs.

AI CRM Features That Matter: 27 Questions to Ask Before You Buy (No Demos Required) - Chronic Digital Blog

Buying an "AI sales" tool in 2026 is easy. Buying one you can actually hand outbound to, one that keeps clean data, avoids deliverability landmines, and proves what it did, is the hard part. Demos will not save you. A vendor can click through a pretty copilot screen while your domain reputation quietly burns in the background.

This post is the buyer's weapon. It is a checklist for evaluating an autonomous outbound operator (an AI sales agent that finds leads, writes and sends email, handles replies, and books meetings) built as 27 questions. Categorized. With pass/fail red flags. With "prove it" prompts the vendor must answer in writing.

A note on category first. A copilot suggests. An operator does the work. If a tool only drafts text in a sidebar and never sends a message or touches a record under guardrails, most of this checklist does not apply, and neither does the word "autonomous."

How to use this checklist (no demos required)

Send this as a document to every shortlisted vendor. Give them 72 hours.

Rules

  1. Written answers only. "We can show you on a call" usually means "it breaks in production."
  2. Evidence beats opinions. Screenshots of logs, sample exports, redacted customer audit trails.
  3. Pass/fail first. Do not negotiate on basics like email authentication, audit logs, and stop rules.

Scoring

  • Pass = meets the requirement and shows evidence.
  • Conditional = meets it with limits, manual work, or extra tools.
  • Fail = "on the roadmap," "via Zapier," "our partners," or "we recommend."

Category 1: Data hygiene (because garbage data poisons every action)

If the operator works off inconsistent, stale records, "AI" just means "wrong outreach at scale." Clean data is the input to every decision the agent makes about who to contact and what to say.

1) What data quality rules are enforced automatically, not suggested?

What you want

  • Required fields, type validation, format checks, and dedupe rules that run continuously.
  • Rules that apply at ingest and on update, not as a quarterly cleanup project.

Pass

  • Field-level validation, required fields by stage, and automated cleanup workflows.

Fail red flags

  • "Users just need to be disciplined."
  • "We have a dashboard for duplicates" (a dashboard is not enforcement).

Prove it (in writing)

  • A list of enforceable rules and where they run (ingest, nightly, real time).
  • A sample "rule hit" log export.

2) How do you dedupe across people, accounts, and domains?

What you want

  • Fuzzy matching, domain normalization, and merge suggestions with conflict handling.

Pass

  • You can define merge precedence rules and keep a merge audit trail.

Fail red flags

  • Dedupe only on exact email address.
  • No merge audit trail.

Prove it

  • A redacted before/after merge record and the audit entry that explains it.

3) Do you prevent inconsistent fields and picklists across teams?

What you want

  • Controlled schemas and picklist governance. No random "Enterprise-ish" values that break segmentation.

Pass

  • Central schema management and controlled vocabularies with change history.

Fail red flags

  • Every rep can create fields freely.
  • No field change history.

Prove it

  • A schema export with creation dates, owner, and change log.

4) Can the system keep enrichment current as people change roles?

Role changes, job hops, and new domains are the silent killers of an outbound list. A contact who left six months ago is a bounce waiting to happen, and bounces hurt your sender reputation.

What you want

  • Continuous refresh of critical fields, with alerts when a contact bounces or changes title.

Pass

  • Automated refresh schedules and triggers based on engagement and bounce signals.

Fail red flags

  • "Run enrichment again if you need it."
  • Refresh costs extra per record with no rules to govern it.

Prove it

  • Refresh cadence options and a sample "job change detected" event.

If you care about enrichment that stays current, map this to your buying criteria for lead enrichment.


Category 2: Autonomous actions (the agent does work, not just narrates it)

"AI notes" are cute. You need an operator that actually moves pipeline: sourcing, sending, replying, booking.

5) Which actions can the agent take end-to-end?

Ask for a checklist. A real outbound operator should be able to:

  • Source and create contacts and accounts from your ICP
  • Update fields and stages
  • Draft and send emails from managed mailboxes
  • Pause or stop sequences on its own signals
  • Handle replies and book meetings

Pass

  • Actions run with guardrails and write back to records.

Fail red flags

  • The agent can only draft text.
  • The agent lives in a sidebar and never actually sends or books anything.

Prove it

  • A list of actions and the exact objects each one can modify.

6) Can the agent run multi-step playbooks with conditions?

Example playbook:

  1. Source and enrich a lead that matches the ICP
  2. Score fit and intent
  3. If the fit and intent are strong, start sequence A
  4. If it bounces, stop and flag the record
  5. If the reply is positive, book the meeting and update the stage

Pass

  • Conditional logic, stop rules, and state tracking across the run.

Fail red flags

  • "You can build that with workflows" but there is no persistent agent state.

Prove it

  • A redacted execution trace of one playbook run from source to booked meeting.

7) Does every action write back with provenance?

You need to know who or what changed a field, when, and why. This is what makes an autonomous system safe to delegate to.

Pass

  • Field history that includes the agent identity, the policy it followed, and the source signals.

Fail red flags

  • "We log it somewhere" but you cannot export it.
  • No per-field change history.

Prove it

  • An export of field history for one record showing at least ten edits.

8) Does the agent work off your ICP automatically, not hand-built lists?

If the agent depends on lists you curate by hand, it is not autonomous. It is autocomplete with a login.

Pass

  • An ICP definition that drives sourcing and targeting continuously.

Fail red flags

  • The ICP exists only as tags and filters.
  • No ICP-driven lead sourcing.

Prove it

  • A screenshot or export of the ICP definition and how it feeds sourcing.

For an ICP-first workflow, anchor your evaluation against an actual ICP builder.


Category 3: Scoring and prioritization (fit plus intent, not vibes)

Scoring decides who gets contacted. It also decides who gets ignored. That makes it a decision you should be able to inspect.

9) Do you separate fit scoring from intent scoring?

Fit = firmographics, technographics, role match. Intent = buying signals, behavior, timing.

Pass

  • Two scores, two explanations, one combined priority.

Fail red flags

  • One magic number.
  • "Our model figures it out."

Prove it

  • A scoring features list and a sample explanation for five leads.

If you want the dual model done right, benchmark against AI lead scoring.

10) What are your top scoring inputs, and can I edit weights?

You do not need full transparency into model internals. You do need control over your go-to-market reality.

Pass

  • Editable weights or rules, override layers, and segment-specific scoring.

Fail red flags

  • No controls.
  • "Talk to support if you want changes."

Prove it

  • The weight controls and the change history of the scoring configuration.

11) Can you explain "why this lead, why now" in one paragraph?

This has to be readable by an operator, not a data scientist. Explaining the decision is the whole point: trust is earned through visible reasoning, not asserted.

Pass

  • A human-readable reason backed by specific data points.

Fail red flags

  • "The model thinks this is a good fit."
  • An explanation that cites no inputs.

Prove it

  • Ten example explanations and the underlying signals.

12) Can scoring trigger actions safely?

Scoring is pointless if it does not route attention and outreach.

Pass

  • Thresholds that trigger sequences, routing, tasks, or escalation, with approvals where it matters.

Fail red flags

  • Scoring lives in a report.
  • No way to act on the score.

Prove it

  • An automation rule that triggers on a score plus the resulting writeback.

Category 4: Deliverability and sending controls (your domain is not a toy)

If a tool sends email on your behalf, it owns deliverability. Full stop. This is the category most likely to do lasting damage if the vendor treats it as an afterthought.

Google and Yahoo's bulk sender rules require authentication (SPF, DKIM, DMARC) and one-click unsubscribe for bulk mailers, plus spam complaint thresholds. Treat this as table stakes, not "email marketing stuff." See the requirement breakdowns from deliverability vendors like Valimail and Klaviyo, plus explainers such as BuzzStream's checklist.

Microsoft has also pushed bulk sender requirements for Outlook.com and related domains, including authentication and unsubscribe expectations. One accessible summary:

13) Do you enforce SPF, DKIM, and DMARC checks before sending?

Pass

  • The system blocks sending if authentication is missing or misaligned.

Fail red flags

  • "We recommend you set it up."
  • No preflight checks.

Prove it

  • A screenshot of the preflight gating or a sample "send blocked" event.

14) Do you support one-click unsubscribe correctly (header plus behavior)?

One-click unsubscribe is not "we put a link in the footer." Providers look for the correct headers and the correct behavior.

Pass

  • Supports List-Unsubscribe and one-click where required for bulk mail.

Fail red flags

  • "We include unsubscribe text."
  • Unsubscribe processed manually.

Prove it

  • Raw message headers from a real send showing List-Unsubscribe.

15) Do you throttle sending per mailbox, domain, and segment?

You need controls like:

  • Daily caps per mailbox
  • Ramp schedules for new domains
  • Per-provider throttling (Gmail, Outlook, Yahoo)
  • Reply-rate-aware pacing

Pass

  • Fine-grained throttles and ramp rules.

Fail red flags

  • One global send limit.
  • "We send as fast as you want."

Prove it

  • The throttle settings page and an export of sends per mailbox per day.

16) Do you have stop rules that actually stop?

Minimum stop rules:

  • Hard bounce: stop.
  • Spam complaint: stop.
  • Unsubscribe: stop.
  • Negative reply: stop.
  • Auto-reply: pause or route.

Pass

  • Stop rules at the event level with immediate enforcement.

Fail red flags

  • Negative replies still get followed up.
  • Stops depend on someone applying a manual tag.

Prove it

  • A redacted event timeline where a bounce stops a sequence.

For the full operator-level standard here (engagement-first pacing, throttling, and stop rules), Chronic's point of view lives here: 2026 deliverability: the engagement-first outbound system.


Category 5: Governance (permissions, approvals, audit trails)

An autonomous system without governance is just fast mistakes delivered with a confident tone. Confident delegation only works when control is always within reach: a pause, an approval, a clear record of what happened.

Standards bodies have been blunt about governance and risk management for AI. NIST's AI Risk Management Framework defines functions like GOVERN, MAP, MEASURE, and MANAGE, and treats documentation and accountability as core mechanics, not "nice to have."

ISO also released ISO/IEC 42001:2023, an AI management systems standard that formalizes governance expectations.

17) Can you restrict what the agent can do by role, object, and field?

Pass

  • Permissions at the field and action level. Separate permissions for read, write, send, and export.

Fail red flags

  • Admin-only controls.
  • All-or-nothing agent permissions.

Prove it

  • A permission matrix export.

18) Do you support approvals for risky actions?

Examples:

  • Sending a new sequence to a new segment
  • Editing pipeline stages
  • Bulk enrichment
  • Auto-creating opportunities

Pass

  • Approvals built in, with queues and timeouts.

Fail red flags

  • "Just review it afterward."
  • Approvals only via a Slack message someone might miss.

Prove it

  • An approval workflow and an audit event that captures the approval.

19) Do you keep immutable audit logs for agent and human actions?

Pass

  • Append-only audit logs. Exportable. Filterable by actor, object, time, and action type.

Fail red flags

  • Logs expire fast.
  • No export.

Prove it

  • An export of 30 days of audit logs from a sandbox.

For deeper governance mechanics, align your questions with a control-plane mindset: the agent governance control plane: permissions, approvals, and audit trails.

20) Can you run the agent in "suggest only" mode, then graduate to "auto"?

Good delegation is gradual. You should be able to watch the agent recommend before you let it act.

Pass

  • Mode controls per workflow, with easy rollback.

Fail red flags

  • It is either fully manual or fully autonomous, with nothing in between.

Prove it

  • Workflow settings showing suggestion versus auto modes.

Category 6: Reporting and attribution (prove pipeline, not activity)

If a tool cannot prove what created pipeline, you will eventually cut the wrong program and keep the wrong one. Outbound is judged on qualified meetings held, not emails sent.

21) Do you provide multi-touch attribution for outbound and inbound together?

Pass

  • Touches tracked across email, calls, meetings, and web conversions, connected to opportunity and revenue.

Fail red flags

  • Attribution based only on "last touch."
  • Outbound tracked separately from your pipeline.

Prove it

  • A sample attribution report with definitions.

22) Can you answer "what did the agent do last week that created meetings?"

Pass

  • An agent activity report tied to outcomes: replies, meetings booked, opportunities created.

Fail red flags

  • Activity metrics only (sends, opens).
  • No link from activity to outcome.

Prove it

  • A report export with columns for lead, action, timestamp, and outcome.

If you want a metric spine operators actually use, connect this to an outbound measurement stack: the outbound ROI stack for 2026: 6 metrics that matter.

23) Do you report deliverability health where you can see it?

Minimum:

  • Bounce rates
  • Spam complaint rates
  • Unsubscribe rates
  • Inbox placement proxies
  • Domain health indicators

Pass

  • Alerting and thresholds, not just a static chart.

Fail red flags

  • "That's your email tool's job" (it sent the mail, so it is its job).
  • No spam complaint tracking.

Prove it

  • Deliverability dashboard screenshots and the alert configuration.

24) Can you track pipeline impact by segment, ICP, and signal?

This is where "AI" becomes practical. You need to know which signals actually convert so the agent can lean into them.

Pass

  • Segment reports by ICP filters and intent signals.

Fail red flags

  • No way to slice outcomes by enrichment or scoring inputs.

Prove it

  • An example report: "hiring-signal leads vs baseline."

Category 7: Integrations and writeback (no data islands)

An outbound operator should work with the CRM and tools you already run, not trap the truth inside itself. If it cannot write back cleanly, you will rebuild reality in spreadsheets. Again.

25) Which integrations are native, and which are glue code?

Ask specifically about:

  • Your CRM (the system of record)
  • Email and calendar
  • Data warehouses
  • Enrichment providers
  • Calling
  • Customer data platforms

Pass

  • Native integrations for core systems and clear API coverage.

Fail red flags

  • "We integrate with everything" but it is all Zapier.
  • Sync in only, no writeback.

Prove it

  • An integration list with directionality: read, write, bi-directional.

26) Do integrations support writeback with conflict handling?

Two systems will disagree. The question is whether the tool handles it like an adult.

Pass

  • Field precedence rules and conflict logs.

Fail red flags

  • Last write wins, silently.
  • No conflict visibility.

Prove it

  • A sample conflict event and the resolution policy.

27) Can I export everything, including agent logs, scoring history, and field changes?

If you cannot export it, you do not own it.

Pass

  • Bulk export or API for records, field history, scoring inputs and outputs, agent execution traces, and audit logs.

Fail red flags

  • Export only "basic fields."
  • No way to extract the decision history.

Prove it

  • API docs or a sample export schema.

Pass/fail dealbreakers (print this and tape it to your monitor)

If any of these fail, do not buy.

Hard fails

  • No exportable audit trail for agent actions.
  • No deliverability controls beyond a "send limit."
  • No enforceable SPF/DKIM/DMARC preflight checks.
  • No clear separation of fit versus intent scoring.
  • No writeback to your core systems; the agent lives in a chat window.

Soft fails that turn into hard fails

  • "On the roadmap" for approvals.
  • "We can build it for you" for stop rules.
  • "Just use another tool" for reporting.

Vendor response template (copy/paste)

Use this to force clear answers.

For each question

  • Answer: (1-3 paragraphs)
  • Pass/Conditional/Fail: (pick one)
  • Limits: (rate limits, objects, roles, extra fees)
  • Evidence: (screenshots, exports, docs)
  • Time to live: (days, not quarters)
  • Owner: (product, support, solutions)

If they cannot fill this in, they cannot run your revenue motion.


Where Chronic fits (one line, no circus)

Chronic is an autonomous revenue operator: you set the goal and it runs outbound end to end until the meeting is booked. It finds leads, enriches them, scores fit and intent, writes and sends email from managed warmed mailboxes, handles replies, and books meetings, surfacing approvals for the decisions that matter. Start your evaluation from the capabilities that actually move pipeline:

If you are comparing stacks, keep it clean:

  • Salesforce buyers usually want governance, then regret the tool sprawl. Here: Chronic vs Salesforce
  • HubSpot buyers want all-in-one, then hit per-seat pricing and limits. Here: Chronic vs HubSpot
  • Apollo buyers want data and sending, then still bolt on five tools. Here: Chronic vs Apollo
  • Attio buyers want a modern CRM, then ask where the outbound engine is. Here: Chronic vs Attio

FAQ

What is the difference between an AI copilot and an autonomous operator?

A copilot suggests. An operator executes. An autonomous outbound operator takes actions like sourcing leads, sending email, handling replies, and booking meetings, and it writes back with logs and governance. If it cannot act, it is just a copilot. If it can act but cannot prove what it did, it is dangerous.

Why does a written checklist beat a demo?

A checklist is a written evaluation of the capabilities that matter in production: data hygiene, autonomous actions, scoring, deliverability controls, governance, reporting, integrations, and writeback. Demos show the happy path. A checklist exposes the failure path, which is where the real costs live.

Which deliverability requirements should an outbound tool meet in 2026?

At minimum it should enforce or gate sending based on SPF, DKIM, and DMARC alignment, and support one-click unsubscribe to meet bulk-mail expectations from the major mailbox providers. Google and Yahoo enforcement began in 2024 for bulk senders, and Microsoft rolled out bulk sender requirements in 2025. Start with the vendor explainers and checklists here:

How do I validate lead scoring without trusting a black box?

Require fit and intent to be separated. Require a "why this lead, why now" explanation. Require editable weights or override rules. Then ask for ten real examples with the signals that drove each score. If the vendor cannot show that, the score is not operational.

What governance features are non-negotiable for an AI sales agent?

Three things: permissions, approvals, and audit trails. Permissions must work at the role, object, and field level. Approvals must cover risky actions like bulk sends and pipeline changes. Audit trails must be exportable and show actor, timestamp, and change details. Use the NIST AI RMF as a reference point: https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Can I hand outbound to an agent if my data is a mess?

Yes, but only if the product automates hygiene. If the vendor says "import your CSV and clean it later," expect garbage scoring, broken routing, and inaccurate attribution. Your first buying filter should be Category 1 in this checklist.


Send the checklist. Demand receipts. Buy the one that can prove it.

Email every vendor the 27 questions. Tell them you want written answers and evidence. No demos. No vibes. The winner is the operator that can:

  • keep data clean without heroics,
  • take autonomous actions with real stop rules,
  • protect deliverability by default,
  • and prove every change with audit logs.

Everything else is theatre.

Ready when you are

Put your pipeline on autopilot.

Chronic runs discovery, outreach, and follow-up end to end. You approve the decisions that matter.