Skip to the article
Cordanis

Buyer's guide

How to evaluate an AI sales system

Evaluate an AI sales system by following one consequential workflow end to end. Check what context it reads, how it distinguishes evidence from inference, what it remembers, who authorizes action, whether approved content can change, how provider outcomes are confirmed, what the audit record preserves, and whether cost can be tied to useful work.

A rain-dark Manhattan workroom with maps, papers, and a fax machine beneath the skyline.
A rain-dark Manhattan workroom with maps, papers, and a fax machine beneath the skyline.
← All research

Start with the sales motion

Begin with the work the team needs to complete, not the model, agent count, or feature list.

A lead queue, a named-account book, and an opportunity desk have different failure costs. A writing assistant may be enough when the seller already holds the context. Autonomous sending may be sensible when one poor touch has little cost.

Write down one real scenario before the demonstration. Name the account, the available sources, the people involved, the decision to be made, the action that may follow, and what would count as success or failure. Then ask the vendor to follow that buyer-selected scenario all the way through. The failure usually lives between the screens.

The point is not to design a trick. It is to prevent a collection of polished screens from standing in for a working system.

Eight tests that belong together

Apply these tests in proportion to the authority the product claims. A writing assistant must prove its drafts and evidence. A system that remembers accounts, sends messages, changes records, or reports outcomes must prove those parts of the path too.

Test What to inspect Warning sign
Context Which account, people, history, and source state the system actually reads The demo depends on context pasted into a prompt
Evidence How facts, inference, uncertainty, freshness, and conflict are distinguished Fluent answers with no recoverable source
Memory What survives the session and who can correct or access it later Every task starts over, or summaries silently overwrite people
Approval Who can authorize a consequential action and what exact work is approved A confirmation button with vague scope or model self-approval
Messaging control Whether reviewed messages, audiences, and later execution remain bound to a version Content can change after approval or vary invisibly by recipient
Provider truth How sends, replies, bounces, and CRM results are confirmed and reconciled Application state is treated as proof that an external action happened
Audit Whether actor, target, decision, action, and result can be reconstructed safely Logs show success without evidence or expose sensitive content unnecessarily
Cost Whether spend connects to runs, useful units of work, and outcomes with honest denominators Only a monthly total or theoretical token estimate is available

Weakness in one test can invalidate another. Approval does not help if content changes afterward. An audit cannot repair fabricated evidence. Cheap generation is not efficient if the work is unusable.

Test context and evidence by changing them

A prepared demo usually proves that the product works with the context the vendor selected. Change something material.

Use an account with two similar names. Correct a contact's role, remove a source, introduce a conflicting fact, and start a fresh session. Open a different account in the interface while naming the original in the request.

Look for explicit identity, source, and authority rules. Visible screen state should not become hidden permission to act. The system should distinguish fact, human judgment, and model interpretation, then admit missing context. Follow evidence to its source and date. Correct a result and check that a later refresh preserves it. A screenshot shows a possible path, not durable performance.

Follow approval into execution

Do not stop when the review card appears. Use a controlled account, test mailbox, sandbox, or safely reversible operation. Never use a real prospect to test a product claim. Approve the test action and follow it.

The person should see the target, content, changed fields, cost, or other consequence. The system that proposed the action should not approve it. Changed work should expire the approval, and a retry should not repeat the effect.

For shared messaging, verify that one reviewed version remains the source for execution. If each message is personalized, identify what can change and how it is reviewed. "Human in the loop" means little without the exact decision.

Verify what happened outside the product

An application can fail before a provider accepts an email, or lose the response after acceptance. A careless retry can create a duplicate.

Check how the system distinguishes not attempted, confirmed, failed, and uncertain outcomes. Inspect the provider record and ask what happens when the network fails at the worst moment. The system should reconcile against external evidence rather than call its own intent success.

The same rule applies to CRM changes, enrichment purchases, calendar actions, and any other connected system. The provider owns the fact that the provider operation occurred.

Read the audit record

The buyer should be able to reconstruct the sources, proposal, human decision, authorized target and version, operation, provider result, and account change without reading a raw model transcript or exposing customer content in logs. Audit is not a screenshot archive. It should let an operator separate a model mistake, product defect, provider failure, and human decision.

Demand honest cost visibility

Token totals are useful for debugging and weak for buying decisions. A buyer needs to know what the spend produced.

Check whether cost connects to a researched account, prepared message, approved action, qualified meeting, or opportunity. Record inclusions, exclusions, retries, and whether each denominator is observed or projected.

If a system lacks an answer, the missing measurement should be named plainly. A theoretical estimate presented as observed unit economics is not acceptable. Cost visibility should mature from run spend, to useful work, to outcomes with an explicit attribution method.

Run failure cases before buying confidence

One completed workflow proves that the pieces can connect. It does not prove stability.

Repeat the scenario with ambiguity, a correction, a missing source, and an action the system must refuse. Inspect the complete record of one test and every failure. Ask which behaviors the vendor can support now and which remain intended.

Record the scenario, product and model version, date, result, failures, and open limits. Also record what the vendor prepared before the demonstration and when a person intervened. The buying team should leave with evidence it owns, not only an impression of the meeting.

The final buying question

After the demonstration, ask whether the system helped the team complete the real sales motion with more trustworthy context and control, or merely produced impressive AI activity.

A simpler product may be the right choice. If the need is drafting help, do not buy an operating system. If the motion is a high-volume queue, do not require the memory and governance of a named-account book. But if the system will work important accounts, represent sellers to customers, and change shared revenue state, evaluate the whole path. The failure usually lives between the screens.


Cordanis is built around connected account context, evidence, human approval, provider-confirmed execution, and measurable work. Read AI SDR vs. sales copilot vs. sales harness, Why human approval is a system boundary, not a button, and How Cordanis is secured.

Keep reading

All research

The New York City skyline rising across dark water at blue hour.

Bring your book.

Bring your accounts to a working session. Leave with a plan for tomorrow.

Book a working session