Why ecommerce support breaks at peak
Ecommerce support does not fail gradually. It fails in a step function, on a specific Thursday, usually the one where a paid campaign lands at the same time as a carrier delay and a product drop. The queue looks fine at 9am and is 900 deep by 4pm. Nothing about the team changed. What changed is that four independent sources of volume stacked on the same day.
The failure modes below are the ones that actually show up in ecommerce queues, as opposed to the generic ones in support blog posts. Each has a mechanism, a cost that shows up somewhere other than the support budget, and a set of standard responses that do not work. If you run a queue for a DTC brand or an online retailer, you have probably tried at least three of the things in the "does not work" column.
Pre-purchase questions die in the buying window
The most expensive ticket in ecommerce is the one that arrives before the order does. Someone is on a product page with a card in hand and a question that blocks the purchase: does this ship to Ireland, will the medium fit a 38-inch chest, is the 2-pack the same price per unit, does the subscription lock me in. They want an answer in the next ninety seconds. They are not going to email you and wait.
The mechanism is simple and unforgiving. Purchase intent decays fast. A shopper who asks a question at 9:40pm and gets a reply at 10:15 the next morning has, in most cases, already bought from someone else, abandoned the idea, or bought from you anyway and the reply was irrelevant. The window in which an answer changes the outcome is short, and it is almost never during your staffed hours, because a large share of DTC traffic converts in the evening and on weekends.
The cost does not land in your support budget. It lands in conversion rate, and nobody attributes it to support. A blocked shopper does not file a complaint. They close the tab. Your CVR drops fifteen basis points, the growth team blames creative fatigue, and the real cause — three hundred unanswered pre-purchase questions a month — is invisible because those conversations were never tickets in the first place. This is the single biggest reason live chat for ecommerce is worth treating as a revenue channel rather than a cost centre.
What teams try that does not work: turning the widget off outside business hours, which converts a soluble problem into an invisible one. Replacing it with a contact form, which is a polite way of saying "come back tomorrow". Deploying a decision-tree bot that offers five buttons, none of which is the customer's question, and whose main achievement is teaching shoppers that the chat widget is useless. Routing pre-sales to the sales inbox, which in a DTC business usually means one person who also runs the marketplace listings.
A Shopify live chat widget that cannot read the catalogue is a form with better styling. What works is a system that can actually answer the question — pulling the variant, the shipping profile, the stock position and the size chart — at 9:40pm without a human awake. That is a different product from a deflection bot. It is the reason we built Agent to resolve rather than deflect.
WISMO swamps the queue and crowds out everything else
"Where is my order" is the tax every ecommerce operation pays. It is typically the largest single ticket category by volume, it is almost entirely predictable, and it scales linearly with orders — which means it grows exactly when you least want it to.
The mechanism is a gap between what the customer can see and what the business knows. The customer has a confirmation email and a tracking number that has not moved in four days. The business has carrier scan data, a fulfilment status, a warehouse cut-off time and knowledge that this particular lane is running slow. None of that is on the order status page. So the customer writes in, and an agent spends four minutes doing something a script could do: open the order, open the carrier site, read the last scan, translate "in transit — origin facility" into English, write a reassuring paragraph.
The cost is displacement. WISMO is not expensive per ticket; it is expensive because of what it pushes out of the queue. When half your inbox is delivery status, your first response time on genuinely hard tickets — a damaged item, a wrong size on a wedding gift, a duplicate charge — goes up. Those are the tickets where response speed determines whether the customer ever orders again. You are spending your scarce human attention on the cheapest possible work.
There is a second-order cost too. WISMO volume is a symptom. High WISMO usually means your post-purchase communication is thin: no proactive delay notification, no clear delivery estimate at checkout, tracking links that dump customers on a carrier page they cannot parse. Every ticket is a customer telling you your order status page failed.
What teams try that does not work: macros. A macro still requires a human to open the order, check the carrier, decide which macro applies and personalise it. You have automated the typing, which was never the bottleneck. Order-lookup bots that ask for an order number the customer does not have to hand and then return the same status string the customer already read. Hiding the contact form behind a tracking page, which just moves the ticket to Instagram DMs.
What works is resolution with real data: read the order from the Shopify integration, read the carrier state, apply the brand's actual policy on delays, and either answer definitively or escalate with everything already gathered. And then feed the pattern back — if 200 people asked about the same lane, the fix is a proactive message, not faster replies.
Returns friction quietly kills repeat purchase rate
Returns are where ecommerce support stops being a cost centre and becomes a retention lever, and where most teams are worst. The mechanism is that a return is a moment of maximum doubt. The customer already spent money on something that did not work. Every extra step you add is read as an attempt to keep their money.
The friction is usually structural rather than deliberate. The policy lives on a page nobody reads. The portal covers 70% of cases and silently fails on the rest — gifts, marketplace orders, exchanges across variants, items bought on promotion, anything past the window by three days. The exceptions land in the inbox, where an agent has to reconstruct what happened from three systems and then make a judgement call with no guidance, which means the answer depends on which agent picked it up. Two customers with identical situations get different outcomes, and one of them screenshots it.
The cost is repeat purchase rate. A customer whose return is handled cleanly frequently reorders — often a different size of the same item. A customer who has to chase a refund for eleven days does not come back, and the loss is not one order, it is the lifetime value you already paid to acquire. Acquisition cost is sunk at the first order; the return experience decides whether you ever earn it back.
Worked example (illustrative): suppose a brand does 10,000 orders a month, a 20% return rate, and a repeat purchase rate of 30% among cleanly-handled returns versus 10% among badly-handled ones. If 600 of the 2,000 monthly returns go badly, the gap is roughly 120 lost repeat orders a month. At a $70 average order value that is about $8,400 a month, or roughly $100k a year, from return handling alone. The inputs are hypothetical; the shape of the arithmetic is not.
What teams try that does not work: tightening the policy, which reduces refund cost and increases churn by more. Adding a required "reason for return" free-text field that nobody reads. Buying a returns portal and assuming the exception traffic disappears — it does not, it just arrives in the inbox with less context because the customer already tried the portal and is now annoyed. Making refunds conditional on warehouse receipt in every case, including the cases where the goodwill cost of an instant refund is a fraction of the ticket cost.
Peak season breaks the staffing model, not the software
BFCM, a product launch, a viral TikTok, a sold-out collab restock. Peak is not a busy week. It is a period where volume goes up four to ten times against a headcount you fixed in September.
The mechanism is that support staffing is a step function while demand is continuous. You cannot hire a fractional agent. So you hire seasonal help in October, and you discover the real constraint: onboarding. A new agent needs to learn the product range, the policy edge cases, the tone, and the tooling. That takes weeks, and you are asking them to absorb it in days, right before the hardest fortnight of the year. The result is a team that is nominally large enough and practically unable to resolve anything without asking a senior agent, which means your best people stop clearing tickets and start answering internal questions.
Compounding it: peak volume is not peak-shaped in the way you plan for. The spike is not evenly distributed across the working day. It concentrates in the hours after email sends and campaign launches, frequently outside your coverage, and it changes shape mid-event as delivery-cutoff anxiety replaces pre-sales questions in the last week before Christmas.
The cost shows up as a backlog that outlives the event. A queue that hits 72-hour response times on 28 November does not recover on 2 December; it recovers in mid-January, and everything that arrives in between is answered late. Meanwhile CSAT craters, refund requests rise because customers who cannot reach you escalate to their card issuer, and your permanent team burns out and starts resigning in February. Peak churn in support staff is a lagging indicator of peak staffing decisions.
What teams try that does not work: BPO surge capacity booked in October, which is the same onboarding problem with a worse feedback loop and a contract minimum. "All hands on support" weeks, where you put engineers and marketers in the inbox — well-intentioned, and it reliably produces policy errors that generate their own follow-up tickets. Cranking auto-reply promises you cannot keep ("we reply within 4 hours") which simply converts one ticket into three. Turning off chat, which pushes everything to email where it is slower to resolve.
What works is capacity that does not need onboarding. Automation that handles the predictable majority means your trained humans spend peak on judgement calls, and a shared inbox with real SLA and assignment control means the judgement calls are actually visible rather than buried under status requests.
Channel sprawl splits one customer into five conversations
The same person emails on Monday, replies to your Instagram story on Tuesday, opens chat on Wednesday and sends a WhatsApp message on Thursday. In most stacks that is four conversations, three agents, and zero shared context. Each agent asks for the order number again.
The mechanism is that channels were added one at a time, each with its own tool and its own identity model. Email is keyed on address. Chat is keyed on browser session. Instagram is keyed on a handle that matches nothing. WhatsApp is keyed on a phone number the customer never gave you at checkout. There is no common key, so there is no common history, and the "unified inbox" you bought unifies the display without unifying the record.
The cost is threefold. First, duplicated work: two agents answer the same question, sometimes differently, and the customer now has two policies in writing. Second, the customer's perception — being asked to re-explain a problem for the fourth time is the single most reliable way to convert a minor issue into a public complaint. Third, measurement. If one problem produces four tickets, your volume metrics are inflated, your resolution rate is fake, and you cannot tell whether a fix worked because you are counting conversations rather than problems.
Social adds a specific hazard. A DM about a missing order is a support ticket. A comment about a missing order is a marketing problem, because it is visible and it is under an ad. The people staffing those two surfaces are usually different, and the social manager answering publicly rarely has order access, so they say "DM us" and restart the whole cycle.
What teams try that does not work: telling customers to use one channel. They will not. Assigning social to marketing and email to support, which formalises the split instead of fixing it. Buying a separate social inbox tool, adding a fifth system. Merging tickets manually, which works until volume rises and then stops happening entirely.
What works is identity resolution at the data layer — matching on email, phone and order across channels so that the conversation history follows the person, not the surface — plus a single queue where email, chat, WhatsApp and SMS land together. That is the baseline capability an ecommerce customer service software stack needs before any automation on top of it is trustworthy, because an AI that cannot see the previous three conversations will contradict them.
The ecommerce ticket taxonomy
Before you automate anything, you need an honest inventory of what is actually in the queue. Most teams have a tagging scheme that was designed once, drifted, and now has forty tags of which six are used. An ecommerce helpdesk is only as good as the categories it reports on, because those categories are what you will automate against. The taxonomy below is the one that holds up across most DTC and online retail businesses, from a $2M brand with two agents to a $100M retailer with forty.
Shares vary enormously by category, price point and fulfilment model. A furniture brand with 12-week lead times has a completely different mix from a supplements brand on subscription. So the table below uses qualitative bands rather than invented precision. Measure your own; the point is the ordering and the automation logic, not the numbers.
| Ticket type | Typical share of volume | What the customer actually wants | Automation potential | What it needs access to |
|---|---|---|---|---|
| WISMO and delivery exceptions | The largest single category in most stores | A credible date, and to know someone is on it | High | Order record, fulfilment status, carrier scan data, delay policy |
| Returns and RMA | Often roughly a fifth | Permission, a label, and certainty about the refund | High | Order, policy window, item condition rules, label generation |
| Refunds and payment disputes | Smaller volume, outsized cost | Their money back, visibly, with a date | Medium | Payment processor record, refund state, bank timing norms |
| Product and fit questions | Large and heavily pre-purchase | Confidence to buy the right variant | Medium to high | Catalogue, variant attributes, size charts, reviews, stock |
| Stock and backorder | Spiky, launch-driven | A restock date or a good alternative | Medium | Inventory by location, incoming POs, back-in-stock signup |
| Order modification and cancellation | Small volume, very time-sensitive | A change made before the warehouse picks it | Medium | Order state, fulfilment cut-off, address validation, edit rights |
| Discount and promo issues | Bursty, tied to campaigns | The price they believe they were promised | High | Code rules, eligibility, order totals, price-adjustment policy |
| Subscription and repeat orders | Dominant for subscription brands | Control: skip, swap, pause, cancel, change date | High | Subscription record, billing schedule, dunning state |
One structural note before the detail. Automation potential is not the same as automation priority. A category can be highly automatable and still be the wrong place to start if it is a small share of volume or if getting it wrong is expensive. Start where high volume, high automatability and low blast radius overlap — which in most stores means WISMO first, then returns initiation, then promo and discount questions. If you want a framework for deciding what "resolved" even means before you measure it, the guide to resolution rate covers the definitional traps.
WISMO and delivery exceptions
Good resolution is a specific date or a specific action, not a status string. "Your parcel is in transit" is not a resolution; the customer already read that. A resolved WISMO ticket tells the customer where the parcel actually is, why it looks stuck if it looks stuck, when it will realistically arrive, and what happens if it does not. For a genuine exception — lost in transit, delivered-not-received, damaged on arrival — resolution means the replacement or refund is already in motion before the conversation ends.
The data required is broader than people assume. You need the order and its line items, the fulfilment record including which warehouse and when it was picked, the carrier's scan history rather than just its latest status, and the brand's own tolerance rules: how many days of no-scan before you treat it as lost, whether you replace or refund first, whether signature-required changes the policy. You also need lane-level context. "Five days with no scan" is alarming for a domestic next-day service and completely normal for an international economy lane in December.
Where human judgement is genuinely needed: delivered-not-received. That is a fraud-adjacent decision with a real cost either way. A rules engine can gather everything — delivery scan with GPS, the customer's order history, whether this address has claimed before, order value — but the call on whether to replace, refund, or ask the customer to check with neighbours and file with the carrier belongs to a human above a certain value threshold. Same for the high-emotion cases: a gift that missed a birthday, a wedding order, a medical necessity. The facts may be simple; the response is not.
Everything else here is mechanical. The lookup, the translation of carrier jargon, the proactive notification when a lane slows, the "your parcel was damaged, here is a replacement" flow within a value band — all of it can run without a human, provided the system has genuine read access rather than a screen-scraped tracking page. This is the category where automation pays for the whole stack, because it is high volume, low judgement and enormously repetitive.
Returns and RMA
Good resolution for a return is speed and certainty in that order. The customer wants to know yes or no, wants a label without hunting for a printer-free option if you offer one, and wants to know exactly when the money lands. A return that is approved in ten seconds with a clear refund timeline is a better retention event than a return that is approved in two days with a discount code attached.
The data required: the order and its date, the item and its category-specific rules (final sale, hygiene items, personalised goods), the return window and whether this customer is inside it, prior return history for this account, the reason code, and the destination — domestic warehouse, regional consolidator, or a keep-it decision where return shipping exceeds item value. You also need to know whether the item was discounted, because the refund amount for a bundle component or a BOGO item is a genuine calculation, not a line-item lookup.
The parts that automate cleanly: eligibility checks, label generation, exchange for a different size of the same item, status updates once the return is in transit, and refund confirmation on receipt. Exchanges in particular are underautomated. An exchange is a retention save — the customer still wants the product — and making it as easy as a refund measurably changes the outcome for the brand.
Where human judgement is genuinely needed: the outside-the-window cases and the pattern cases. A loyal customer three days past the window with a plausible reason is almost always worth approving, and that decision needs either a human or an explicit policy that a system can apply consistently. At the other end, serial returners — 80% return rate across twelve orders — need a human conversation, not an automatic approval and not an automatic rejection. Condition disputes are the third: the warehouse says worn, the customer says new. That is judgement, and it needs the photos, the history and a person.
Refunds and payment disputes
Refund tickets are low volume and high stakes. The customer has already decided the relationship may be over; the only question is how badly it ends. Good resolution is the refund issued, plus an honest statement of when it will appear — which is usually not when you issued it, because card settlement adds business days that are outside your control and entirely inside the customer's perception of you.
Most refund escalation is caused by that gap. You refunded on Tuesday. Their bank posts it the following Monday. On Thursday they write in angry, because from where they sit nothing happened. A support system that can state the processor's refund timestamp and explain typical bank posting windows resolves that conversation in one message. One that says "it's been processed, please check with your bank" produces three more tickets and, frequently, a chargeback.
Chargebacks deserve their own handling because the mechanics matter. When a customer disputes a charge with their issuer, the funds are pulled from you provisionally, you pay a fee regardless of outcome, and you get a fixed window to submit evidence. The evidence that wins is boring and specific: proof of delivery, the customer's acknowledgement of receipt, your refund policy as displayed at checkout, and any support conversation where you offered a resolution the customer declined. Which means your support transcript is dispute evidence, and the ability to retrieve it fast has direct financial value. A large share of disputes are "friendly fraud" or simple confusion — the customer did not recognise the descriptor, or forgot a subscription renewal — and both are preventable with clearer billing descriptors and better renewal notices.
The data required: the payment record from the processor, the refund state and timestamp, the original authorisation, any partial refunds already applied, and the full conversation history. Automation handles status, timing explanations, partial refund calculation and proactive "your refund is issued" messages well. Human judgement is required on goodwill decisions above a threshold, on anything with a fraud signal, and on every dispute response, because the framing of the evidence matters and a template loses.
Product and fit questions
This is the category that converts. A product question is a shopper who wants to buy and cannot yet. Good resolution is not a link to the size chart — it is an answer: "for a 38-inch chest, take the medium; this style runs narrow in the shoulder, and reviewers with your measurements consistently size up." The measure of success here is not CSAT or handle time. It is conversion rate on chatted sessions versus non-chatted, and return rate on chatted orders, which should be lower because the customer bought the right thing.
The data required is deeper than a product feed. Variant-level attributes, real measurements rather than S/M/L, materials and care, compatibility data for anything technical, ingredient or allergen data where relevant, and — critically — the accumulated knowledge in reviews and past conversations. Your agents already know that the black colourway runs half a size small. That knowledge lives in someone's head or in a Slack thread. Getting it into a knowledge base an AI can answer from is the highest-leverage unglamorous work in the whole function.
Automation potential is medium to high, and it depends almost entirely on the quality of that source material. An AI answering fit questions from a thin product description will hedge, and hedging kills conversion. An AI answering from real measurements, review synthesis and documented brand guidance outperforms an average agent, because it has read every review and the agent has not.
Where human judgement is genuinely needed: subjective and high-consideration purchases. A $3,000 sofa, a bike frame sizing, a made-to-measure garment, a gift where the shopper does not know the recipient's size. These are consultative sales conversations, and the right move is to route them to a person quickly with the browsing context attached — and if that person is your best salesperson rather than your fastest ticket-closer, better still. The other human case is anything where a wrong answer creates liability: allergens, medical device compatibility, electrical safety. Never let a probabilistic system freelance on those; either the source of truth is explicit or it escalates.
Stock and backorder
Stock questions are spiky and emotional in a way that is easy to underestimate. "When is this back?" from a customer who has checked the page daily for three weeks is not an information request; it is a request for control over something they want and cannot have.
Good resolution has three possible shapes, in descending order of value. Best: an accurate restock date, because a date lets the customer stop checking. Second: a genuinely good alternative — same function, available now — offered specifically rather than as a link to the collection page. Third, and the fallback: a reliable back-in-stock notification with early access, which converts the disappointment into an owned marketing contact. Answering "we don't have a date, keep an eye on the site" is a resolution in the ticketing sense and a failure in every other sense.
The data required: inventory by location including committed versus available, incoming purchase orders with realistic ETAs, historical sell-through to sanity-check whether a restock will last an hour, and a substitutability map for the catalogue. That last one is the piece almost nobody maintains, and it is the difference between an automated answer that saves a sale and one that just conveys bad news politely.
Automation handles date communication, notification signup and status well. The difficult part is honesty about uncertainty. Supply chains slip. If a system promises "back on the 14th" and it slips twice, you have manufactured three tickets per waiting customer and taught them not to trust you. The rule that works: communicate ranges and confidence, commit only to dates you control, and proactively message the waiting list when a date moves rather than letting them discover it.
Where human judgement is genuinely needed: allocation. When a restock is smaller than the waiting list, someone has to decide who gets it — longest waiting, highest value, the customer whose original order was cancelled through your error. That is a business decision with fairness implications, and it should be made deliberately once and then applied consistently, not improvised per ticket. Also human: B2B and bulk enquiries hiding inside stock questions, which are sales leads, not tickets.
Order modification and cancellation
This category is small in volume and disproportionately damaging when handled badly, because it is the only one with a hard clock. The customer ordered eleven minutes ago and typed the wrong apartment number. There is a window — until the warehouse picks — in which the fix is trivial, and after it the fix costs a return, a reshipment and two weeks.
Good resolution is therefore defined by latency, not quality. A change made in three minutes is a non-event. The same change attempted in six hours is a delivery exception, a lost parcel and a refund argument. This is the clearest case in ecommerce where response time is not a vanity metric.
The data required: the order's exact fulfilment state, the warehouse cut-off schedule, whether the order has been transmitted to a 3PL and whether that transmission is reversible, address validation, and the payment implications of any change — adding an item may require a new charge, removing one a partial refund, and changing to a different-priced variant both.
Automation potential is medium and bounded by write access. Reading order state is easy. Editing a live order is a write operation against the store and sometimes against a 3PL's system, and it needs strict guard rails: allowed within a defined window, allowed for address changes within the same country and postal region, not allowed if it changes the customs profile or the shipping cost band, never allowed after the pick. Within those bounds it automates well and delivers an unusually good experience, because the customer gets an instant confirmed fix at midnight.
Where human judgement is genuinely needed: anything past the cut-off, where the options are intercept-in-transit, refuse-on-delivery or return-and-reorder, each with different costs. Also cancellations with a save opportunity — a customer cancelling because delivery is too slow may keep the order if offered an upgrade, and that is a conversation. And fraud: a sudden address change to a freight forwarder on a high-value order shortly after purchase is a known pattern and should always land in front of a person.
Discount and promo issues
Promo tickets arrive in bursts, always immediately after a campaign, and they are almost entirely self-inflicted. The code did not apply. The code applied but the customer expected it to stack. The sale started six hours after they ordered. The banner said 30% and the code gave 20% because the item was already reduced. Each of these is a communication defect that landed in the queue.
Good resolution is a decision plus a reason, delivered fast. "Your code excluded sale items, which is why it did not apply — here is the relevant line from the terms, and because the banner was ambiguous I have applied the difference as a refund" beats both a flat refusal and a silent goodwill refund. The first preserves the policy; the second teaches customers that complaining is profitable, which is a strategy with a very predictable ending at your next sale.
The data required: the code's actual rules — eligibility, exclusions, stacking behaviour, minimum spend, expiry with timezone — the order timestamp against the promotion window, the customer's cart contents at the time, and the brand's price-adjustment policy. The timezone detail is not pedantry; a sale that ends at midnight in one timezone generates a predictable wave of tickets from customers in another.
Automation potential is high. The rules are deterministic, the decision is a lookup, and the refund or price adjustment is a mechanical operation. A well-configured system can resolve the large majority of these end to end, apply the adjustment within a defined value ceiling, and explain the reasoning in plain language. Set the ceiling by policy, not by vibes, and let automation and routing rules enforce it.
The more valuable output is the pattern. If 400 people ask why a code did not stack, the code's terms were unclear, and the fix belongs to the merchandising team before the next campaign. Support is the only function that sees this in real time. Reporting it back — with volume attached — is worth more than resolving the tickets faster.
Where human judgement is genuinely needed: rarely on the individual ticket, often on the exception. A high-value customer, a genuinely misleading banner, a pricing error that went live for an hour. Those need a person and a decision recorded somewhere the next person can find it.
Subscription and repeat orders
For subscription brands this category dominates everything else, and its tickets are different in kind. The customer is not asking a question; they are trying to exert control over a recurring charge. Skip next month. Change the date to after payday. Swap the flavour. Add a one-off. Pause for the summer. Cancel.
Good resolution is giving them that control immediately and without friction. The instinct to add retention friction — hide cancellation, require a phone call, make them email — is measurably counterproductive: it produces chargebacks, which cost more than the subscription, and public complaints, which cost more than that. The better model is to make every action self-serve and to make the save offer good rather than the exit hard. A customer offered a pause when they wanted to cancel frequently takes the pause.
The data required: the subscription record with its next billing date and cadence, the payment method and its expiry, dunning state, the full order history, and any contractual commitment. Dunning matters and is underappreciated: a large share of subscription churn is involuntary — the card expired, the bank declined a renewal, the descriptor was not recognised. Those customers did not choose to leave. Catching a failed payment and giving the customer a one-click way to update the card recovers revenue that otherwise silently disappears, and it is entirely automatable.
Automation potential is high across the board: skips, swaps, date changes, pauses, address updates, card updates, and cancellations with an offered alternative. The whole category is deterministic operations against a subscription record.
Where human judgement is genuinely needed: disputes about charges the customer says they did not authorise, cancellations from customers with a genuine product problem — where the right answer is to fix the product experience, not process the cancellation — and the small number of cases where someone has been charged repeatedly through a system failure. That last one needs a human, a full refund, and an apology that does not read like a template. Copilot is useful here: it drafts from the actual account history so the agent spends their time on the decision rather than the reconstruction.
Answer the pre-purchase question at 9:40pm and resolve the WISMO ticket before an agent opens it. Start a 14-day free trial at /signup, connect Shopify in minutes, and see what an AI agent that resolves end to end clears from your queue in the first week.
How Aftersales resolves ecommerce conversations
The word "resolve" gets abused in this category. Plenty of tools call a conversation resolved when the customer stops replying. That is not resolution. Resolution means the customer's problem has changed state: the parcel has been located, the return label is in their inbox, the refund has been confirmed against the actual payment record, the address on the unfulfilled order is now the address they meant to type.
What follows is five walkthroughs. Each one traces a real message from arrival to logged outcome, including the lookups, the branch points, and the conditions under which Agent stops and hands to a person. Read them as specifications, not as marketing. If your own workflow differs, the shape should still be recognisable enough to argue with.
1. WISMO, answered from order and fulfilment data
The message. 9:40pm, Sunday. Email from an address that is not the one on the order: "Hi — ordered a week ago and nothing's arrived. Order 10442? Getting a bit worried." No account, no last name, no confirmation of which email placed the order.
What Agent looks up. First, identity. The sending address does not match a customer record, so Agent searches the Shopify order index by the number quoted. Order 10442 exists. The order email is a Gmail address; the sender is a work address on the same personal name as the shipping contact. That is a partial match, not a verified one. Agent does not disclose the shipping address or the last four digits of the card. It confirms only what the customer has already demonstrated they know: that an order with that number exists and was placed on the date they implied.
Next, state. Agent pulls the fulfilment record: one fulfilment, created six days ago, carrier tracking number attached, last scan event four days ago at a regional sortation hub, no delivery scan, no exception code. It also pulls order-level facts that change the answer — shipping method selected (economy, not express), destination country, whether the order contained a pre-order line item that delayed the whole fulfilment, and whether a partial fulfilment exists that the customer may have already received.
The decision logic. Agent classifies against a small set of fulfilment states rather than freelancing. Not yet fulfilled and inside the stated dispatch window: reassure with a date. Not yet fulfilled and outside it: explain why, and if the delay is stock-driven, offer the split-shipment option if the store allows it. Fulfilled and tracking is moving normally: give the current scan and the expected window. Fulfilled and tracking has been static beyond the store's stall threshold — commonly four to seven days domestically, longer cross-border: this is the branch that matters. Delivered scan present but customer says nothing arrived: a different branch again, and usually a human one.
Order 10442 lands in the stalled branch. The store's configured rule, set in Workflows, says a domestic parcel with no scan for more than four days is treated as at risk. The rule authorises Agent to open a replacement or a refund without a carrier claim, up to a value ceiling. The order is £68, under the ceiling.
The reply. Short, specific, and decisive. It states where the parcel last scanned and when. It says plainly that four days without movement is longer than it should be and that the store is not going to make the customer wait out a carrier investigation. It offers two options in one sentence each: reship the same items today on the faster service, or refund in full to the original card. It asks one question and ends. It does not include a tracking link dressed up as an answer, and it does not tell the customer to "allow a further 48 hours" — the phrase that generates the second contact.
The customer replies "reship please" at 9:44pm. Agent creates the replacement draft order against the original, holds it for the morning fulfilment cut, and confirms.
The branch Agent does not take. If the tracking had shown a delivered scan and the customer said nothing arrived, none of the above applies. That case has a fraud dimension, a carrier-liability dimension and usually a neighbour. Agent's job there is narrower: confirm the delivery scan time and location, ask the two questions that resolve a genuine share of them without any cost (has anyone else at the address taken it in, and is there a safe-place photo on the carrier record), and then hand to a human with the answers attached. It should never issue a replacement on a delivered scan alone unless the store has explicitly authorised it under a low value ceiling for first-time claimants. That is a deliberate policy decision, not a default.
Why out-of-hours matters here specifically. WISMO does not arrive evenly. It clusters in the evening, when the customer gets home and the parcel is not on the step, and it clusters after weekends. A rota that covers 9am to 6pm meets that demand roughly twelve hours late, by which time the customer has often emailed twice, opened a chat, and messaged on Instagram. Three contacts, one problem, three handle times. Answering at 9:40pm on Sunday does not just improve CSAT; it removes two duplicate conversations from Monday's queue before they exist.
The logged outcome. Resolved by Agent, first contact, four-minute handle time, out of hours. Tagged wismo/carrier-stall, linked to the original Shopify order, with the replacement order ID written into the conversation record so finance can reconcile later. Insights counts it against the carrier, not against the store's own dispatch performance — which is the distinction that lets you take a real number into a carrier review rather than a feeling.
2. Self-serve returns and RMA initiation
The message. WhatsApp, mid-morning: "the jacket's too big, how do I send it back". Two lines, no order number, sent from a number that matches the phone on a Shopify customer record.
What Agent looks up. The phone match gives a verified customer. Agent pulls their order history and finds three delivered orders in the last year, one of them containing a jacket, delivered eleven days ago. It reads the line items: one jacket, size L, one pair of socks. It reads the store's return policy from Knowledge — thirty days from delivery, unworn with tags, socks excluded as a hygiene item, returns free for domestic customers, one free label per order.
It also checks the things that quietly kill a return later: was the jacket bought on a final-sale discount code, has a return already been opened against this order, is the customer in the destination country the return portal supports, and has this customer's return rate crossed the threshold the store uses for manual review.
The decision logic. Eligibility is a chain, and Agent walks it in order. Inside window: yes, day eleven of thirty. Item eligible: the jacket yes, the socks no. Discount status: standard sale, returnable. Existing RMA: none. Return-rate flag: the customer has returned four of eleven items, below the store's review threshold. Refund or exchange: the store's policy prefers exchange, and the same jacket in M is in stock at the fulfilment location — Agent checks inventory before offering it, because offering an exchange for an out-of-stock size is a guaranteed second conversation.
The reply. Agent confirms which item it is talking about, by name and size, because "the jacket" is ambiguous in a three-order history and getting it wrong wastes a week. It offers exchange for the M with the label sent now, or refund, and it states the one condition that customers actually trip over: tags attached. It says explicitly that the socks are not returnable, rather than letting the customer post them back and discover it at inspection. Then it asks for a single yes.
The customer picks exchange. Agent initiates the RMA, generates the return label through the store's returns configuration, sends it over WhatsApp as a document, and writes the RMA reference back onto the order. If the store operates a keep-it rule under a value threshold, Agent applies it here and tells the customer not to post the socks back.
The stopping conditions. Agent does not proceed alone when: the return is outside the window and the customer is asking for an exception, the item is flagged as damaged or faulty on arrival, the return rate is over threshold, the order was paid partly in gift card and partly on card, or the customer is in a market where the store has no return lane. Faults in particular go to a human, because a fault claim is a quality signal and a photo triage, not an eligibility check.
Why the channel changes the writing, not the logic. This arrived on WhatsApp, so the reply is three short messages rather than one paragraph, and the label goes as an attachment the customer can forward to whoever is actually doing the posting. The eligibility chain is identical on email, SMS and web chat. What changes is length, formatting and whether a link or a document is the better delivery. Teams that maintain separate return flows per channel end up with three policies that drift apart; teams that maintain one policy and vary the presentation do not. Inbox keeps all of it in one thread against one customer, so a return that starts on WhatsApp and continues by email does not become two conversations and two RMAs.
The follow-through most stores skip. Initiating the RMA is not the end of the job. The return then sits in a state for days — in transit, received, inspected, refunded — and each transition is a moment the customer might contact you. Agent should own the proactive side of that: confirm when the carrier scans the return, confirm when the warehouse receives it, and confirm when the refund or exchange dispatches. Three short unprompted messages remove most of the "have you got my return yet" volume, which in many stores is comparable in size to inbound WISMO and is entirely self-inflicted.
The logged outcome. Resolved by Agent, tagged returns/exchange-size, RMA reference stored, reason code too large written to the product-level return reason data. That last field is the one merchandising teams care about. A season of size-driven exchange reasons on one style is a spec problem, and it shows up in Insights as a pattern long before it shows up in a supplier review.
3. Refunds and payment questions, with Stripe context
The message. Email: "You refunded me on the 3rd and there's nothing in my account. It's the 9th. Where is my money?" Tone is short. This is a customer who is now angry about the refund, not the original problem.
What Agent looks up. Shopify shows the refund: £142.00, issued on the 3rd, against the original payment. Stripe shows the underlying object: the refund was created on the 3rd, status succeeded, destination the original card, and the payment itself was made with a card that has since been reported to the network — or not, depending on the case. Agent reads the payment method type, because the answer differs. Card refunds settle on the issuer's clock. Wallet payments route back through the wallet's underlying funding source, which customers routinely fail to check. Bank-debit methods behave differently again.
Agent also checks for the two situations that look like a missing refund and are not: a partial refund where the customer expected full, and a refund that was issued to a gift card or store credit because the original tender was store credit.
The decision logic. If Stripe shows the refund succeeded and the elapsed time is within normal issuer posting behaviour, the honest answer is that the money has left the merchant and is sitting with the bank. Agent says that, gives the date the refund was created, names the last four digits of the card it went back to, and explains where the customer should look — the original statement line, which on many issuers reverses against the original transaction rather than appearing as a new credit. That detail resolves a genuine share of these conversations on its own.
If the elapsed time is beyond the reasonable window, Agent does not argue. It pulls the refund's identifier from Stripe and hands to a human with that reference pre-attached, because the next step is an acquirer trace and a human owns that.
If Stripe shows the refund failed — which happens when a card is closed — Agent flags it immediately. That is not a waiting problem, it is a re-issue problem, and telling the customer to wait another five days is the worst available answer.
The reply. Facts first, in the first sentence. Then the mechanism in plain language. Then either a closing line or a clear statement of what the store is doing next and by when.
Refund authority. Agent's ability to issue a refund is bounded, not open. Stores configure a per-conversation value ceiling, a monthly aggregate ceiling, and a list of categories where Agent may never refund alone — typically high-value orders, orders with a chargeback or dispute already open, orders flagged by fraud screening, and any order where the customer is asking for a refund outside policy. Chargebacks in particular are hard-stopped. Once a dispute exists, the money is already in a formal process with evidence deadlines, and a goodwill refund on top of it creates a double loss. Agent recognises the dispute state from Stripe and routes straight to the person who handles the evidence submission.
The adjacent questions in this bucket. Refund timing is the loudest payment question but not the only one. "I was charged twice" is usually an authorisation hold sitting alongside a capture, and Stripe shows which is which — Agent can explain that the hold will drop off and name the amount, which resolves it without touching money. "Why did my card get declined" is answerable from the decline code in broad terms: issuer declined, insufficient funds, address verification mismatch. Agent should describe the category and what to try next, not read out raw codes and not speculate about the customer's finances. "You charged me the wrong amount" needs the order's line items, discounts and shipping compared against the captured amount; if they reconcile, Agent explains the breakdown, and if they do not, that is a human's problem immediately.
A note on what Agent should never say. Never "the refund has been processed, please allow 3-5 working days" when the underlying object has not been checked. It is the single most common lie in ecommerce support, usually told in good faith by someone reading a status field rather than the payment record, and it is why customers arrive at the second contact already angry. If the refund exists, say when and where. If it does not, say that instead.
The logged outcome. Resolved by Agent with tag payments/refund-timing, Stripe refund ID attached, no money moved. Or escalated with tag payments/refund-not-received, with the reference, the dates and the card type already in the thread so the human starts at step two.
4. Product, sizing, compatibility and stock
The message. On-site chat, from a visitor sitting on a product page for ninety seconds: "will this fit a 2019 model, and do you have it in black". Two questions, one sentence, no context about what "this" is beyond the page they are on.
What Agent looks up. The page context gives the product. Agent reads the catalogue record: variants, options, current inventory by variant and location, price, and any metafields the store uses for specifications — dimensions, materials, compatibility lists, care instructions. Then it reads Knowledge for the things that are not structured: the fit guide, the sizing notes written by the buyer, the compatibility article, the known exceptions.
Stock is the part most tools get wrong. "In stock" on the storefront and "available to promise" are not the same number when there is an open backorder queue or when inventory sits across several locations. Agent reads the variant-level availability rather than the badge.
The decision logic. Compatibility questions resolve against an explicit list or they do not resolve at all. If the knowledge article names the 2019 model, Agent answers yes and cites it. If the article names 2020 onwards and is silent on 2019, Agent must not infer. Silence is not compatibility. It says the compatibility list covers 2020 and later, that it cannot confirm 2019, and offers to check with the product team — which is an escalation with a defined question, not a shrug.
Sizing follows a different rule. Sizing advice is probabilistic and Agent should say so. If the fit note says the style runs small and the customer has an order history, Agent can reference what they bought before and what they kept. That is the single most useful sizing input any store has, and almost nobody uses it. "You kept a medium in the same cut last spring, and this style runs true to that one" outperforms any size chart.
Stock branches on the number. In stock in the requested variant: say so and link the variant directly. Out of stock with a confirmed inbound date: give the date and offer a back-in-stock notification. Out of stock with no date: say there is no date. Do not manufacture one. Low stock: say the count only if the store is comfortable exposing it.
The reply. Answers in the order asked. Compatibility, honestly bounded. Colour, with the actual variant status. One next step. On a pre-purchase chat, the next step is a link to the variant, not a sentence inviting the customer to go and find it themselves.
The knowledge discipline behind it. Pre-purchase accuracy is a content problem more than a model problem. If the compatibility article was last updated two seasons ago, Agent will answer confidently from stale data and generate returns. Three habits fix most of it: give every specification article an owner and a review date, write compatibility as explicit inclusion lists rather than prose, and feed the top pre-purchase questions from Insights back to whoever writes product copy each month. The questions customers ask before buying are a free, continuously updated list of everything your product pages fail to say.
Where it stops. Agent should not give medical, safety or regulatory assurances about a product, should not confirm compatibility by inference, should not promise a restock date that supply has not confirmed, and should not guarantee a size. If a customer asks "will it definitely fit", the honest answer includes the returns policy, because a frictionless return is the real guarantee and pretending otherwise is how you end up arguing about it later.
The logged outcome. Resolved by Agent, tagged presales/compatibility and presales/stock, with the product handle attached. Pre-purchase conversations are worth separating in reporting because their value sits in conversion, not in cost avoidance. A cluster of the same compatibility question against one SKU is a product-description defect with a measurable revenue cost, and it is cheap to fix once you can see it.
5. Order modification and cancellation in the edit window
The message. SMS, forty minutes after checkout: "wrong address!! it went to my old flat. can you change it". No order number. Phone matches a customer record with one order placed today.
What Agent looks up. The order, its financial status, its fulfilment status, whether it has been routed to a warehouse or 3PL, whether a fulfilment request has been accepted, and the store's edit window. Then it checks the two things that turn an easy edit into a mess: whether the order contains a line item from a different fulfilment location, and whether the new address changes the tax jurisdiction or the shipping rate.
The decision logic. The window is binary and Agent must respect it exactly. Unfulfilled, no fulfilment request accepted, inside the store's cut-off: Agent can edit. Fulfilment accepted but not dispatched: Agent must attempt a hold, which may or may not succeed, and it must not promise success. Dispatched: the address is fixed and the conversation becomes an intercept or a return-to-sender conversation, which is a different answer entirely.
This order is forty minutes old and the warehouse cut is 2pm. Agent can act. It asks for the full new address in one message — including postcode — rather than dragging it out over four exchanges. It validates the address, checks that the new jurisdiction does not change the tax total, and confirms the shipping rate is unchanged.
Where it stops. If the edit changes the amount owed, Agent stops. Collecting additional payment or issuing a partial refund as part of an edit is a money movement inside a live order, and stores overwhelmingly want a human on that. Same for adding a line item, applying a discount code retroactively, or changing the payment method. Cancellations follow the same logic: cancel-and-refund inside the window and under the value ceiling can be automatic; anything involving a partially fulfilled order goes to a person.
The reply. Confirms the new address back in full, states the deadline it beat, and says what will happen next. Reading the address back matters — a transposed house number confirmed in writing is a problem the customer can catch.
Speed is the whole feature. An address change is worth nothing at hour six. The economics of this use case sit entirely in the minutes between checkout and the warehouse pick, and that window is frequently shorter than a support team's response time. A store dispatching same-day has perhaps two hours of genuine editability. If the queue answers in four, every one of these conversations converts into a redelivery, a return-to-sender, or a goodwill reship — three outcomes that each cost more than the original shipping. This is the clearest example on the page of automation being worth more for its latency than for its labour saving.
The cancellation variant. "Cancel my order" behaves similarly but deserves one extra step. Before cancelling, Agent should find out why, in one question, because a meaningful share of cancellations are solvable: the customer found it cheaper elsewhere, they picked the wrong variant, they thought delivery would be faster than the confirmation said. A store can choose to have Agent address the reason — confirming the actual delivery date, or offering a variant swap — before processing the cancellation. That must be one attempt, clearly optional, and never a wall. An agent that makes cancelling hard is a dark pattern, and customers punish it publicly.
The logged outcome. Resolved by Agent, tagged orders/address-change, edit recorded against the order, original address retained in the conversation for audit. And, quietly, a data point: repeated address edits within an hour of checkout usually mean the checkout is autofilling something stale, which is a conversion problem hiding inside a support tag.
What Agent resolves and what it escalates
The table below is a reasonable default for a mid-sized Shopify store. Every line is configurable; none of it should be treated as fixed. The useful exercise is to take this to your own team and argue about the middle column.
| Conversation type | Agent resolves | Agent escalates | Why |
|---|---|---|---|
| Where is my order, tracking normal | Yes, end to end | No | Deterministic answer from fulfilment data |
| Where is my order, carrier stalled | Yes, within value ceiling | Above ceiling | Reship/refund decision is rule-bounded |
| Delivered but not received | Partially — gathers evidence | Yes, for the decision | Fraud exposure; needs human judgement |
| Return eligibility and RMA | Yes, inside policy | Outside window, faults, flagged accounts | Policy is a checkable chain |
| Exchange with stock check | Yes | If size unavailable | Needs live inventory, not a promise |
| Refund status and timing | Yes | If Stripe shows failed or traced | Answer is a lookup, not a decision |
| Refund request, in policy, under ceiling | Yes | Above ceiling | Bounded financial authority |
| Refund request, out of policy | No | Yes | Exception-making is a human act |
| Active dispute or chargeback | No | Yes, hard stop | Formal process with evidence deadlines |
| Sizing and fit guidance | Yes, with hedging | If customer pushes for a guarantee | Probabilistic by nature |
| Compatibility, documented | Yes | If undocumented | Silence is not confirmation |
| Stock and restock dates | Yes | If no date exists and customer wants one | Cannot invent supply data |
| Address edit, unfulfilled | Yes | If amount changes | Money movement mid-order |
| Cancellation, unfulfilled | Yes, under ceiling | Partial fulfilment | Reconciliation complexity |
| Subscription changes | Depends on stack | Usually yes | Not a shipped connector in every case |
| Damaged or faulty item | No | Yes | Photo triage plus quality signal |
| Complaint with legal or press language | No | Yes, immediately | Tone and stakes exceed automation |
| B2B or wholesale terms | No | Yes | Bespoke pricing and contract terms |
Confidence thresholds, and what they actually gate
A confidence threshold is not a single dial. In practice there are three separate gates and conflating them is how teams end up with an agent that is either reckless or useless.
The first gate is retrieval confidence: did Agent find a source that actually answers this? If the knowledge base has nothing on 2019 compatibility, no amount of fluency should produce an answer. Low retrieval confidence should escalate, not paraphrase.
The second is interpretation confidence: is Agent sure what the customer is asking? "I want to send this back" is clear. "This isn't right" is not — it could be wrong item, faulty item, wrong size, or buyer's remorse, and each has a different path. The correct behaviour at low interpretation confidence is one clarifying question, then escalation if the answer is still ambiguous. One question. Not four. An agent that interrogates is worse than one that transfers.
The third is action confidence, and it is the one that should be set most conservatively. Answering wrongly costs a follow-up message. Refunding wrongly costs money and creates a reconciliation entry someone has to unwind. These deserve different tolerances, and any platform that gives you one global setting for both is underspecified.
Refund authority as a real limit
Give Agent a per-conversation ceiling, an aggregate daily or monthly ceiling, and a category exclusion list. The aggregate ceiling is the one people forget, and it is the one that protects you. A per-conversation limit of £75 sounds safe until a pricing error on a popular SKU produces two hundred identical complaints in an afternoon. The aggregate limit is a circuit breaker: it trips, Agent stops refunding, the queue routes to humans, and someone investigates why volume spiked. That is a feature, not a failure.
Worked example (illustrative). A store takes 4,000 support conversations a month. Say 55% are WISMO and returns-status questions with deterministic answers, 15% are pre-purchase questions answerable from catalogue and knowledge, 12% are refund-status lookups, and the remaining 18% are exceptions, faults, complaints and edge cases. If Agent handles the first three groups well and escalates the fourth cleanly, roughly 720 conversations a month reach a human — and they arrive with context already attached. These are round hypothetical numbers chosen to show the shape of the arithmetic, not a performance claim. Run the same split against your own tag distribution; the exercise is more useful than the example. Our guide to measuring resolution rate sets out how to define the denominator so the number means something.
Why a good escalation beats a confident guess
There is an asymmetry here that support leaders understand intuitively and vendors keep ignoring. A customer who is told "this needs someone who can authorise it, I've passed it to the team with everything they need, you'll hear back within the hour" has had an acceptable experience. Slightly slower than they wanted, but competent. A customer who is given a fluent, wrong answer has had a bad experience twice: once when they acted on it, and once when they came back to find it was wrong.
The second failure also costs more internally. A wrong refund is a finance ticket. A wrong compatibility answer is a return, a restocking cost and usually a review. A wrong delivery promise is a second WISMO contact with a customer who now distrusts everything the store tells them. Deflection metrics do not capture any of this, which is why per-resolution pricing models — Intercom's Fin charges $0.99 per resolution — create an incentive structure worth thinking hard about before you adopt one. We set out the comparison in more detail in our review of Intercom alternatives. Aftersales prices per seat from $24 a month, so an escalation costs the same as a resolution and nobody is nudged toward calling a guess a win.
Designing the handoff
Most teams treat handoff as a routing problem. It is a writing problem. The question is not which queue the conversation lands in; it is what the human sees in the first four seconds after they open it.
What must travel with the conversation
Seven things, and they should be visible without clicking anything.
The customer's actual question, in their words, at the top. Not a summary that has replaced it. Summaries lose the phrasing that carries the emotion, and the emotion is often the thing the human needs to respond to first.
What Agent already told them. The human must know exactly what has been promised. Contradicting your own AI in the next message is the fastest way to lose a customer who was merely mildly annoyed.
Why it escalated, stated as a reason, not a code. "Refund request £310, above the £150 automatic ceiling" tells the human what decision they are being asked to make. "Escalation rule 14" tells them nothing and costs them a lookup.
The order context, pulled live: order number, value, status, fulfilment state, tracking, payment method, returns history, lifetime value. Not a screenshot — live fields, because the parcel may have moved since the escalation fired.
Payment state where money is in play: the Stripe payment and refund objects, their status, and whether a dispute exists. A human should never have to open another tab to learn a chargeback is already running.
Customer history in one line. Orders placed, orders returned, previous contacts this quarter, whether they have complained about this before. Repeat contact about the same order is a different situation from a first contact and warrants a different opening sentence.
Suggested next actions, offered not executed. Three options with the consequence of each: refund in full to the card, offer store credit at a premium, or reship. The human decides; the system removes the research.
The test for whether your handoff package is right is blunt. Hand a conversation to someone who has never seen it and time how long before they type the first word of their reply. If it is more than fifteen seconds, something on that list is missing or buried.
How Inbox presents it
A handed-off conversation arrives in Inbox as a single thread, not a new ticket with a pointer to an old one. The AI portion and the human portion sit in one timeline, visually distinguished so the agent can see at a glance where the machine stopped. Order and payment context render in the side panel beside the thread rather than in a separate tab, and switching between conversations is fast enough — sub-100 millisecond — that an agent working a queue never loses their place.
Reopened conversations behave the same way. When a customer replies to something Agent closed three days ago, the thread reopens with the full prior exchange intact and a marker showing the gap. The agent reads eleven lines, not eleven tickets. If the reply needs input from operations or finance, a side conversation into Slack keeps that discussion attached to the thread instead of scattering it across direct messages that nobody can find in a month.
SLA behaviour
This is where handoff design either holds up or collapses. Get it wrong and your SLA numbers become fiction.
Three rules that work. First, the SLA clock starts when the customer's message arrives, not when Agent gives up. Anything else lets you hide slow escalations behind a fast bot. Second, an escalated conversation keeps its original priority and inherits any urgency signals Agent detected — a customer who used the word "solicitor" should not land at the back of a first-in-first-out queue. Third, escalations that sit unclaimed past a threshold re-route automatically, because the failure mode in every shared queue is a conversation that everyone assumes someone else has.
Track time-to-human separately from time-to-resolution. They fail for different reasons, and reporting in Insights should let you see both.
Measure the gap between when Agent stops and when a human starts typing. If that number creeps past a few minutes during business hours, your escalation is technically working and operationally broken — and customers experience it as being ignored, not as being escalated.
How Copilot shortens the human's reply
Once a person takes over, Copilot does three things that compress the reply from six minutes to one.
It drafts a first version grounded in the same order, payment and knowledge context Agent used, with sources cited so the agent can check a claim in one glance rather than trusting it. The draft is a starting point, and it should always be edited — but starting from a structured draft beats starting from an empty box, particularly for the long explanatory replies that carrier and refund cases demand.
It translates in both directions. An inbound message in Portuguese becomes readable; the agent's English reply goes back in Portuguese. For a DTC brand selling across several European markets without a native speaker in each, this is the difference between a same-day answer and a two-day one.
And it holds the store's voice steady. Agent's replies and a human's replies should not read like they came from different companies. Copilot drafts in the established tone, which matters most exactly when it is hardest — in the escalated conversations where the customer is already unhappy and every word is being read closely.
The discipline that makes all of this work is simple: humans should take over knowing more than they would have known if they had handled the conversation from the start. If your handoff does not clear that bar, fix the handoff before you tune the model.
Proactive chat before the cart goes cold
Reactive support answers the customers who ask. Proactive chat reaches the ones who would have left silently — and the ones who would have converted anyway, which is where it gets expensive if you are careless.
Three triggers that earn their place
Cart hesitation. Not "added to cart" — that fires constantly and means very little. The signal worth acting on is a customer who reached checkout, engaged with a field, and then stalled. Thirty to sixty seconds of inactivity on a checkout step is a real hesitation. The intervention should be narrow: an offer to answer a question about delivery time, returns or sizing, since those three account for the bulk of late-stage doubt. Not a discount. Discounting at hesitation trains customers to hesitate, and you will see it in your data within two months as a rising share of carts that stall deliberately.
Browsing depth without progress. Four product pages in one category, eight minutes, no add-to-cart. That person is not idly browsing; they are comparing and cannot decide. A single message offering to narrow it down — "looking between the two jackets? happy to tell you how they differ on warmth" — is genuinely useful. The equivalent message on the first page view is noise.
Returning-visitor recognition. A known customer who returns to a product page they viewed last week is worth a different message from a first-time visitor. If they have order history, Agent already knows what they bought and kept, and a message that references it honestly — "you've got the M in this cut already, this one runs a touch bigger" — is the kind of thing that makes a brand feel small in the good way. The line to hold is that it must be information the customer knows you have. Referencing a purchase is fine. Referencing browsing behaviour they never told you about is unsettling.
Where proactive messaging annoys people
Be honest about this, because the failure modes are well known and stores keep walking into them anyway.
Popping up within seconds of arrival. The visitor has not read anything yet. There is nothing to ask about. It reads as an interruption because it is one.
Firing on every page. One proactive message per session is close to the right number. Two is pushing it. Three means the visitor is now managing your widget instead of shopping.
Mobile overlays that cover the buy button. A chat bubble that obscures the primary call to action on a small screen is a conversion loss dressed as a conversion tactic. Test it on an actual phone before shipping it.
Fake human presence. "Sarah is typing" when no Sarah exists is a small lie that becomes a large one the moment the customer notices. If an AI agent is answering, say so in the opening line. The honest version performs better than most teams expect, because the customer's real question is whether they will get an answer now.
Ignoring intent. Someone in the returns flow does not want an upsell. Someone reading the shipping policy at 11pm wants the shipping policy, not a greeting. Suppress proactive triggers on account, order-status and policy pages entirely.
Not remembering. If a visitor dismissed the prompt once, do not show it again this session. If they dismissed it three visits running, stop showing it to them at all.
What the proactive conversation should actually do
A proactive message that opens a conversation and then cannot answer it is worse than no message at all. If the visitor replies "does this ship before Friday", the answer needs the shipping rules, the cut-off time and the destination — which means the same knowledge and catalogue grounding used in the reactive flows above, not a separate lightweight widget bolted onto the storefront. One agent, one knowledge base, one set of policies, whether the conversation started because the customer asked or because you did.
It also needs to hand off on the same terms. Pre-purchase conversations escalate for different reasons than post-purchase ones — bespoke or bulk requests, price matching, a question about a product only the buyer can answer — and they are usually more time-sensitive, because the visitor is on the page right now and will not be in ten minutes. Route them to a queue with a tighter target than your post-purchase SLA, and make that deliberate rather than accidental.
Making it measurable
Hold out a control group. Run proactive chat on 80% of eligible sessions and suppress it on 20%, then compare conversion across both. Without a holdout you will attribute every engaged session's purchase to the chat prompt, and engaged sessions were always going to convert at a higher rate. Watch dismissal rate too — a rising dismissal rate on a trigger means it is firing in the wrong moment, and it is a leading indicator of the bounce rate damage that shows up later.
Our fuller treatment of trigger design, placement and measurement is in the piece on running live chat on your website. The short version: proactive chat is a scalpel. Used on the two or three moments where a customer is genuinely stuck, it converts. Used everywhere, it is a popup with better branding, and customers have been trained for twenty years to close those without reading them.
Start with the workflows above, not with a pilot that proves nothing. Take your top five ecommerce tags, map each one to a resolve-or-escalate line, and run it for a fortnight — start a free trial and build the mapping as you go.
The ecommerce support stack
Most ecommerce support problems are not support problems. They are data problems wearing a support costume. The customer asks "where is my order" and the agent has to open four tabs to answer. The customer asks for a refund and the agent has to check whether the return was scanned, whether the payment captured, and whether the item was on a promotion that changed the refund value. None of that is hard. It is just spread across six systems that were bought at different times for different reasons.
So before anything else, map what you actually run. Not the logos on your stack diagram — the specific record that answers each question a customer asks. Here is how the layers usually break down, and what a helpdesk genuinely needs from each one.
Storefront
The storefront is where the customer forms the expectation you are later asked to defend. Shopify, in the vast majority of cases for brands reading this. It holds product data, variant-level inventory, the customer account record, the cart, and the checkout.
What support needs from the storefront is narrower than people assume. You need product titles and variants that match what the customer sees, because a customer writing in about "the blue one" is describing a PDP, not a SKU. You need stock status, because half of WISMO is really "is this actually coming or is it backordered". And you need the customer record, so a conversation can be tied to a person rather than to an email string.
Shopify is a shipped connector for Aftersales, so this layer is the easy one. Order, customer, fulfilment and refund data flow in without you building anything. That matters more than it sounds: it is the difference between an AI agent that can answer a shipping question and one that can only apologise for not being able to.
What breaks when it is disconnected. The agent cannot see the order, so every single conversation opens with a request for an order number. That one question is responsible for more abandoned support threads than any other sentence in ecommerce. It also destroys prioritisation: without order value or customer history, a first-time £20 buyer and a customer who has spent £4,000 with you look identical in the queue. And it makes product questions unanswerable — an agent who cannot see which variant was bought is guessing about fit, colour and compatibility, which produces confidently wrong answers and, three weeks later, a return.
Order management
Small brands run OMS inside Shopify. Once you cross into multi-warehouse, pre-orders, bundles, or a 3PL with its own allocation logic, an OMS appears — and it becomes the real source of truth for what "shipped" means. Shopify might say fulfilled when the label was printed. The OMS might know the carton did not leave the dock until Tuesday.
Support needs the OMS to answer three questions: what is actually allocated to this order, when is it physically leaving, and what is the split if it ships in more than one parcel. Split shipments generate an enormous amount of avoidable contact, because customers open one box, find two of four items, and assume something went wrong.
No OMS is a native connector. That is fine and it is worth being blunt about it: you reach the OMS through the Aftersales API, pushing the fields you need onto the conversation or the customer record so Agent can use them. In practice this is a small sync job, not a project.
What breaks when it is disconnected. The agent answers from the storefront's view of the world and confidently tells a customer their order shipped on Friday when the carton is still sitting on a pallet. That is not a small inaccuracy. It is a promise the customer will hold you to, and when it turns out to be false you have converted a neutral WISMO into a trust problem. The second failure is split shipments. Without OMS visibility the agent sees one order marked fulfilled and has no idea the customer received two of four items, so it reassures a customer whose actual problem is a missing parcel. Those conversations reopen, and a reopened conversation costs you the original handling time plus the new one plus whatever goodwill you spend fixing it. The third is pre-orders and backorders, where the storefront record says nothing useful about the expected date and the agent either invents one or refuses to answer. Both outcomes generate a follow-up.
Payments
Stripe is a shipped connector. That covers a large share of the "was I charged twice", "why is there a pending authorisation", and "where is my refund" traffic, which is consistently the highest-emotion category in ecommerce support.
The mechanics matter here. An authorisation hold is not a charge; it is a reservation against the customer's available balance that drops off on the issuer's schedule, not yours. A refund is not a reversal; it is a new transaction that has to settle, and the time it takes to appear is controlled by the customer's bank. A failed payment on a subscription renewal triggers dunning — a retry sequence with escalating notifications. Every one of these produces a customer who believes they have been charged and a support agent who has to explain a mechanism they only half understand.
Giving the helpdesk read access to payment state is the single highest-leverage integration after orders. It turns "let me check with our finance team and come back to you" into an answer in the first reply.
What breaks when it is disconnected. Money questions become a two-system relay. The agent messages finance, finance checks the payment dashboard, finance replies hours later, the agent relays it. Meanwhile the customer — who believes they have been charged twice and is watching a hold sitting against their rent money — escalates. A material share of chargebacks in ecommerce are not fraud and not even genuine disputes; they are customers who could not get a straight answer fast enough and used their bank as an escalation channel. Every one of those costs you the dispute fee, the staff time to contest it, and a mark against your dispute ratio. Disconnected payment data also makes duplicate-refund errors much more likely, because an agent who cannot see a refund already in flight will happily issue a second one.
Shipping and tracking
This is your highest-volume layer and it has no native connector. Carrier APIs, or more commonly a tracking aggregator, own the scan events that tell you where a parcel actually is.
What support needs is the last scan, the scan timestamp, the current carrier-stated delivery estimate, and an exception flag. Not a tracking URL. A tracking URL is what you send when you cannot answer the question. The whole point of connecting this layer is to stop sending tracking URLs.
Connect it through the API. If your aggregator emits webhooks on exception events, point them at a workflow so a stalled parcel opens or updates a conversation before the customer notices. Proactive contact on a stalled parcel is cheaper than the two-message thread it prevents.
What breaks when it is disconnected. This is the worst one, because it is the highest-volume layer. Without carrier scan data reaching the agent, the only possible answer to "where is my order" is the tracking link the customer already has. They have already clicked it. They are writing to you precisely because it said nothing useful. Sending it back is the support equivalent of shrugging, and customers read it that way.
There is a second-order effect that is easy to miss. Without scan age, you cannot distinguish a parcel moving normally from a parcel that has been silent for a week, so your agents treat both identically — usually with "please allow a further three to five working days". For the healthy parcel that is fine. For the lost one you have just added a week to the customer's wait before you begin resolving a problem you could have detected automatically, and by then the goodwill gesture required is much larger. Lost-parcel cost is a function of how late you notice. A connected tracking layer lets you notice first and open the conversation yourself, which is one of the few moves in support that reliably improves satisfaction and reduces contact at the same time.
A third failure is 3PL-specific. When the 3PL holds the label data and it does not reach the helpdesk, your agent cannot even confirm the parcel exists. The customer is told "it has shipped", the carrier has no record, and nobody can say whether the item was ever picked. That conversation escalates to a warehouse email thread, which is why Slack side conversations from inside the ticket are worth setting up even before the API work lands — at least the question and its answer stay attached to the conversation instead of dying in someone's inbox.
Returns and exchanges
Returns platforms hold the return request, the RMA, the reason code, the label state, the warehouse scan and the refund or exchange decision. Customers write in at every one of those stages, and they usually write in because the portal did not tell them something.
The helpdesk needs the return status and, critically, the reason code. Reason codes are the cheapest product feedback you will ever collect, and they belong in Insights alongside contact reasons so you can see that a sizing complaint and a returns spike are the same event.
API or email piping, depending on the platform. Many returns tools send status emails; piping a copy into the helpdesk gets you most of the value in an afternoon.
What breaks when it is disconnected. Returns generate a long tail of status questions — "did you get it", "why hasn't my refund landed", "the label wouldn't print". Every one of those is answerable from the returns platform and unanswerable without it. Disconnected, your agents live in the returns portal in a second tab and copy-paste statuses by hand, which is slow and error-prone at exactly the moment a customer is already unhappy.
The subtler loss is the reason code. When return reasons stay locked inside the returns tool and contact reasons stay inside the helpdesk, nobody ever sees that the spike in "too small" returns and the spike in sizing questions started on the same day a new supplier shipped. That is a merchandising fix worth more than the entire support automation project, and it is invisible unless the two datasets meet.
Reviews and post-purchase as a support surface
Most brands treat reviews as marketing and surveys as research. Operationally, a meaningful slice of both is unlogged support.
Think about what a two-star review actually is. A customer had a problem, decided not to contact you — because contacting support felt like work, or because they tried and gave up — and instead published their complaint where your prospective customers read it. The review platform notifies a marketing inbox. Someone replies with a templated apology two weeks later. Nothing is fixed, nothing is measured, and the review stays up.
The same is true of post-purchase NPS or CSAT surveys. A detractor comment saying "arrived with the box crushed" is a damage claim. It is sitting in a survey export that support has never seen.
Treat both as inbound channels. Pipe low-rating review notifications and detractor survey comments into the helpdesk by email, tag them distinctly, and give them their own view so nobody confuses them with normal queue volume. Then give them a real workflow: triage within a working day, resolve the underlying issue properly, and only then reply publicly. The order matters. Replying publicly before fixing the problem is what produces the "we're sorry to hear this, please DM us" boilerplate that everyone recognises as theatre.
Two practical notes. First, these contacts arrive with no conversation history and often a different email than the order, so this is where your identity resolution gets stress-tested. Second, they are a genuine revenue surface rather than a cost: a resolved complaint frequently becomes a revised review, and the customers involved are, by definition, ones you were otherwise about to lose. Tag the outcomes so you can show that in Insights when someone asks why support is answering reviews.
What breaks when it is disconnected. You keep a permanent blind spot in your contact-reason data. Your dashboard says damage complaints are 2% of volume; the true figure includes a population of customers who never wrote in and went straight to a public star rating. You will under-invest in packaging for a year because the data told you the problem was small.
Marketing automation
Your ESP sends the flows: abandoned cart, shipping confirmation, win-back, replenishment. Support needs to know what the customer was sent, because an enormous share of confusion is manufactured by your own emails. A customer who received a "your order has shipped" flow four days before the parcel moved is not confused. They are correctly reading a message you sent.
There is no connector here, and you mostly do not need one. What you need is a documented list of every automated message, its trigger, and its copy, stored in Knowledge so the AI agent knows what the customer was told before they wrote in.
What breaks when it is disconnected. Your agent contradicts your own emails, which is uniquely corrosive because the customer has the evidence in front of them. Worse is the marketing collision: a win-back discount email landing the same morning a customer is arguing about a damaged item, or a replenishment reminder going to someone who cancelled their subscription three days ago and is currently in a refund thread. Support then spends the conversation apologising for marketing. A simple suppression rule — exclude customers with an open conversation from promotional sends — costs an hour to configure and removes a whole category of unnecessary friction.
The helpdesk in the middle
Everything above feeds the layer where the conversation actually happens. Inbox is the single queue where AI and human work sits together — no separate bot console, no "escalated to the human tool" handoff that loses context. Channels land here: email, WhatsApp, SMS, live chat, and side conversations out to Slack when you need a warehouse answer without leaving the ticket.
The design principle worth holding onto: the helpdesk should not be a system of record for anything. It is a system of resolution. Order truth lives in Shopify or the OMS. Payment truth lives in Stripe. Parcel truth lives with the carrier. The helpdesk reads all of it, writes conversations, and stays thin.
What breaks when the middle is missing. If channels live in separate tools — chat in one place, email in another, WhatsApp on somebody's phone — you get the same customer asking the same question three times and three different answers going out. You also lose any honest measure of volume, because no single system sees all of it. Consolidating channels before automating anything is not a nice-to-have; an AI agent that can only see one of three channels will contradict itself across the other two.
| Stack layer | What the helpdesk needs from it | How it connects |
|---|---|---|
| Storefront | Order lines, variants, customer record, stock status | Shopify — shipped connector |
| Order management | Allocation, real ship date, split-shipment detail | API |
| Payments | Charge, authorisation, refund and dispute state | Stripe — shipped connector |
| Shipping and tracking | Last scan, timestamp, ETA, exception flag | API or carrier webhooks into workflows |
| Returns | RMA status, reason code, refund decision | API or email piping |
| Reviews | Low-rating alerts, review text, order link | Email piping |
| Marketing automation | Which flows fired, and their exact copy | Documented in Knowledge, not synced |
| Warehouse and 3PL | Ad hoc answers on specific parcels | Slack side conversations |
| Legacy helpdesk | Historical tickets for migration | Zendesk or Freshdesk import |
| CRM | Account context for wholesale or VIP buyers | Salesforce connector |
Two honest caveats. First, Shopify and Stripe are the only ecommerce-native connectors that ship; Salesforce, Zendesk, Freshdesk, WhatsApp, email, SMS and Slack round out the list, and everything else is API work. Second, that API work is smaller than the average stack diagram implies, because most of these systems only need to contribute three or four fields to a conversation.
Connecting your order and customer data
Here is the test that matters. A customer sends "hey, where's my stuff?" from a phone number your system has never seen, with no order number, no account, and no context. Can your agent answer without asking a single clarifying question?
If yes, you have solved identity resolution and you will resolve a large share of WISMO in one turn. If no, every WISMO costs you a round trip — and a round trip is not just a delay, it is a place where 30% of customers give up and go to the chargeback route instead.
What Agent needs to answer a WISMO cold
Four things, in order of importance.
The order. Not a list of orders — the one the customer means. Usually the most recent unfulfilled or recently fulfilled order. If there is genuine ambiguity between two open orders, the agent should say which ones it can see and ask the customer to pick, which is a very different experience from "please provide your order number".
The fulfilment state, in plain terms. Not fulfillment_status: partial. The agent needs to translate that into "two of your three items shipped on Tuesday, the sweatshirt is coming separately because it's in our second warehouse".
The last carrier scan and its age. A parcel scanned six hours ago in a regional hub is fine. A parcel with no scan in five days is an exception, and the correct answer is a replacement or refund offer, not a tracking link. The age of the scan should be a field, not something the AI infers from a date string.
What you already told them. The shipping confirmation, the delay notice, the "out for delivery" that fired yesterday. Contradicting your own automated email is the fastest way to lose a customer's trust in an AI agent.
Identity resolution in practice
Ecommerce identity is messier than any other vertical because guest checkout exists.
Guest checkout means the only durable key is the email address on the order. There is no account to look up. If the customer writes from a different address than they checked out with — a work address, an old one, a partner's — you have no automatic match. The honest fix: let the agent match on any of email, order number, postcode plus surname, or phone number, and accept that some percentage will need one verification question. Design that question well. "What's the delivery postcode on the order?" is faster and less annoying than "what's your order number?" because the customer knows the answer without opening their email.
Multiple emails is the common case for repeat customers. Shopify will happily hold three customer records for one human. You want a merge concept in the helpdesk so conversations across those addresses roll up to one person. Without it, your VIP looks like three one-time buyers and your agent treats a fifteen-order customer like a stranger.
WhatsApp numbers are the hardest. A phone number is a strong identifier but it is rarely captured at checkout in a normalised format, and often not at all. Two workable approaches: ask once, verify against the order, and store the association so the next conversation resolves instantly; or seed the mapping from your ESP's SMS consent list, where numbers are already normalised. The first conversation from a new number costs you one verification turn. Every one after that costs zero. That is a good trade and you should make it deliberately rather than discovering it in month three.
Marketplace and social orders deserve a mention. If you sell through a marketplace, the buyer's email is often a proxy address and will never match your store records. Treat those as a separate identity space with their own resolution path, usually the marketplace order ID.
What to sync and what to leave alone
The instinct is to pull everything in. Resist it. Every field you copy is a field that can go stale and a field you have to defend in a privacy review.
Sync the things you filter, route or search on: customer email and phone, order ID, order status, order value, order date, a lifetime-value band, and tags like VIP or subscriber. These need to be local because workflows and views query them constantly, and a view that waits on an external API is a view nobody uses.
Read at runtime — do not copy — full order line items, payment and refund detail, tracking scan history, return status, and anything that changes hourly. Fetch these when the conversation needs them. They are always current, and they never sit in your helpdesk waiting to be breached.
Never sync full card data, CVV, or raw payment credentials. There is no support use case. If an agent needs to identify a card, last four digits and brand are sufficient and should come from the payment record at read time. Aftersales does not claim PCI DSS certification and you should not build a workflow that assumes it does.
Privacy posture
Ecommerce support handles names, addresses, order history and purchase behaviour. In many jurisdictions that is enough to make you a controller with real obligations, and your helpdesk is a processor.
Aftersales is SOC 2 Type II and GDPR-aligned; the details sit on the security page and your DPO will want to read them rather than take your word for it. Practically, four things belong on your implementation checklist:
Set a retention policy on conversations and hold to it. Support archives are the least-governed customer data in most companies precisely because nobody thinks of them as a database.
Build the deletion path before you need it. When a customer exercises a deletion right, you need to be able to find and remove their conversations, not just their Shopify record. Test it once.
Redact what does not need to persist. If an agent pastes a full card number into a ticket — and they will — you want a rule that strips it.
Be explicit about training. Know whether conversation content is used to improve models, document the answer, and put it in your privacy notice. Customers increasingly ask, and "I'll check" is a bad answer.
- Shipped ecommerce connectors
- Shopify and Stripe
- Everything else
- API, email piping, or Slack side conversations
- Minimum to answer a cold WISMO
- order state, fulfilment detail, last scan and its age
- Hardest identity case
- guest checkout from an unrecognised email
- Sync locally
- fields you route, filter or search on
- Read at runtime
- line items, payments, tracking, returns
- Never sync
- full card data
- Compliance
- SOC 2 Type II, GDPR
Implementation: the first 30 days
This is a runbook, not a project plan. It assumes a brand doing somewhere between 2,000 and 20,000 support conversations a month, a team of three to fifteen, and one person who can own the rollout for roughly half their week. Scale the timings, not the sequence.
The sequence matters more than the speed. Almost every failed AI support rollout we see skipped week two.
Week 1: connect, import, classify
Objectives. Every channel lands in one queue. Historical conversations are searchable. You have a ticket taxonomy that reflects what customers actually contact you about, not what you assumed three years ago.
Tasks.
Connect email first. Forward or pipe your support address, send test messages, confirm threading and that replies go out from the right address with the right signature. Check SPF and DKIM properly — deliverability failures found in week four are painful.
Connect Shopify. Verify that opening a conversation from a known customer shows the order sidebar. Test a guest-checkout order specifically, because that is the case that breaks.
Connect Stripe if you use it directly. Confirm charge, refund and dispute visibility.
Connect WhatsApp and SMS if you run them. WhatsApp has its own approval steps and template requirements; start them on day one because they take longer than you expect.
Import history from Zendesk or Freshdesk if you are migrating. You want at least twelve months, both to preserve customer history and because week two eats this data.
Then the part people rush: the taxonomy. Pull a random sample of 300 recent conversations. Read them. Tag each with the real reason. You will end up with something like: order status, delivery exception, wrong or missing item, damaged item, return request, return status, refund status, exchange, sizing question, product question, cancellation or address change, discount or promo, payment issue, subscription change, and account access. Fifteen categories, give or take. Anything under 1% of volume gets folded into a neighbour.
Who does what. Ops or IT connects channels. Your most experienced agent — not the manager — does the taxonomy sample, because they know what customers actually mean. The rollout owner reviews.
Done looks like. A test conversation from each channel resolved end to end by a human, in Inbox, with order context visible in the sidebar. A taxonomy document with volume percentages per category.
Signals you are ready to move on.
- A test message from every live channel arrived, threaded correctly, and was replied to successfully.
- Opening a conversation from a guest-checkout order shows the order in the sidebar without anyone typing an order number.
- Your taxonomy has fewer than twenty categories and each one is above 1% of sampled volume.
- Two different agents, tagging the same ten conversations independently, agree on at least eight.
- Historical tickets are searchable and a search for a known customer's name returns their old threads.
Most common mistake. Adopting your old helpdesk's tag list wholesale. It has 140 tags, 90 of which have not been used in a year, and it will poison every report and every routing rule you build for the next twelve months. Start clean.
Week 2: build the knowledge base
Objectives. Every high-volume contact reason has a written, current, unambiguous answer. Confidence thresholds are set. This is the week that determines whether the AI works.
Tasks.
Export your macros. Most will be useless as-is — they are fragments written for humans who supply the judgement. "Hi {name}, I've checked and your order is on its way!" is not knowledge. The policy behind it is.
For each taxonomy category, write the policy as an article: what the rule is, what the exceptions are, what the agent is allowed to offer, and where the boundary sits. Your returns article should say the window, what condition items must be in, who pays return shipping under which circumstances, what happens to sale items, and what to do when the window has just closed. That last one is where most of your real traffic lives and most macros are silent on it.
Mine past resolutions for the rest. Filter your imported history to conversations with high CSAT and short resolution paths, in your top categories. Those transcripts contain the answers your team actually gives. Turn them into articles. Where two agents answered the same question differently, you have found a policy gap — decide, write it down, and you have improved the human team too.
Then set confidence thresholds. This is a policy decision, not a technical one. The question is: at what certainty should Agent answer versus hand to a human? High thresholds mean fewer autonomous resolutions and almost no wrong answers. Lower thresholds mean more coverage and more risk. Start high — deliberately conservative — and walk it down over weeks as the transcripts justify it. Set different thresholds by category: answering a shipping-window question wrong is recoverable, telling someone their return is approved when it is not is not.
Who does what. Knowledge owner drafts — ideally a senior agent given real time for this, not squeezed between tickets. Category owners (returns lead, finance for refunds) approve policy. Rollout owner sets thresholds with the support lead.
Done looks like. Articles covering the categories that make up 80% of volume, each with a named owner and a review date. Thresholds documented with the reasoning. A tester can ask ten realistic questions and get ten accurate answers from the knowledge base alone.
Signals you are ready to move on.
- Every category above 5% of volume has at least one article, with a named owner and a review date.
- A senior agent can read the returns article and answer three awkward edge cases from it without checking with anyone.
- No two articles give conflicting answers to the same question. If they do, you have not finished deciding.
- Thresholds are written down with the reasoning, not just configured.
- Someone outside support — finance for refunds, ops for delivery — has signed off the policy articles that commit their team to something.
Most common mistake. Importing help-centre pages and declaring the knowledge base finished. Help-centre copy is marketing-adjacent, written to reduce contact, and deliberately vague about edge cases. It creates a confident agent that is wrong on precisely the questions that matter. Write the internal version.
Week 3: shadow mode on one ticket type
Objectives. The AI drafts answers for one category without sending them. You read every one. You fix what is wrong before a customer ever sees it.
Tasks.
Pick the category. Order status. It is high volume, low risk, well-bounded, and heavily data-dependent, which means it tests your integrations as well as your knowledge. Do not start with refunds.
Turn on shadow mode. The AI produces the response; a human reviews, edits and sends. In the meantime your agents get faster, because Copilot drafting is genuinely useful even at this stage.
Review daily, and treat it as a standing meeting, not a background task. Thirty minutes each morning, one person, a sample of at least 30 drafts. Score each: send as-is, minor edit, major edit, or would have been wrong. Log the reason for anything below "send as-is".
Categorise the failures, because the fix differs. Missing knowledge means write the article. Wrong tone means adjust the voice guidance. Missing data means fix the integration. Misread intent means look at the training examples. Correct-but-unhelpful is the interesting one and usually means your policy is bad, not your AI.
Fix daily. The loop should be same-day. A week of accumulating problems is a week of not learning anything.
Who does what. Reviewer should be a senior agent with veto power. Knowledge owner takes the fixes. Rollout owner tracks the send-as-is rate day over day.
Done looks like. Five consecutive days where send-as-is exceeds 80% for that category, with zero "would have been wrong" in the last three. Then, and only then, let it send autonomously — still with daily review, just after the fact.
Signals you are ready to move on.
- Send-as-is rate has been flat or rising for five consecutive days rather than bouncing around.
- The last three days produced no draft classified as "would have been wrong".
- The failure log has shifted from missing knowledge to tone nitpicks — that shift is the real signal.
- Your reviewer says they would be comfortable letting the last twenty drafts send unedited. Ask them directly and believe the answer.
- Escalations from this category arrive with enough context that the receiving agent does not reread the thread.
Most common mistake. Reviewing in aggregate instead of reading transcripts. A dashboard showing 78% accuracy tells you nothing actionable. Reading twenty bad drafts tells you exactly which four articles to rewrite. Read the conversations. Every AI rollout that works has someone who reads the conversations.
Week 4: expand, then operationalise
Objectives. Two or three more categories live. SLAs and views configured. Escalation rules agreed and written down. The system runs without the rollout owner standing over it.
Tasks.
Expand by risk, not volume. From order status, the natural next steps are delivery exceptions, product and sizing questions, and return policy questions — answering the policy, not approving the return. Hold refunds, cancellations and anything touching money until you have a month of clean data.
Set SLAs that mean something. Now that the AI absorbs the fast tier, your human SLA can be tighter, because humans only see hard things. First response under an hour on email is realistic. But set a separate SLA for AI-handled conversations that measures resolution, not response, because response time on those is seconds and measuring it tells you nothing.
Build views people will live in. At minimum: escalated from AI, breaching SLA in the next two hours, VIP or high-LTV, negative sentiment, and reopened. Reopened is the most under-used view in ecommerce support — a reopened conversation is either an unresolved problem or a broken promise, and both are expensive.
Write the escalation rules down. Not culture, rules. Anything mentioning legal action, a chargeback, a safety issue, a serious allergy, or a death in the family goes to a human immediately. Anything where the customer has asked twice for a human goes to a human. Anything above a value threshold you set goes to a human. Anything where confidence is below the category threshold goes to a human with a summary and the AI's best guess attached, because an escalation that arrives without context has wasted the customer's time twice.
Agree the handoff quality bar. When the AI escalates, the human should see a summary, the customer's actual ask, what has already been tried, and the relevant order data. If your agents have to scroll the transcript to work out what is happening, your handoff is broken.
Schedule the reviews that keep this alive: weekly transcript review, monthly knowledge audit against reality, quarterly threshold review.
Who does what. Support lead owns SLAs and views. Ops and finance sign off on escalation value thresholds. Rollout owner books the recurring reviews before handing over, because reviews that are not in a calendar do not happen.
Done looks like. Three or more categories live. Escalation rules in a document everyone has read. A named owner for knowledge, for transcript review, and for the monthly resolution rate report. Someone other than the rollout owner can explain how it all works.
Signals you are ready to move on.
- Agents are working from views rather than scrolling the full queue.
- Nobody has asked "should this have gone to a human?" this week, because the rules answer it.
- Reopen rate on AI-resolved conversations is stable and you know what the number is.
- The recurring reviews are in calendars with named owners, not in a document.
- The rollout owner took two days off and nothing needed them.
Most common mistake. Declaring victory at day 30. The first month gets you a working system; months two and three get you a good one. The teams who keep reading transcripts in month six are the ones whose resolution rate keeps climbing.
Staffing an ecommerce support team around AI
The uncomfortable conversation first. If an AI agent resolves a meaningful share of your contacts end to end, you do not need the same number of tier-1 agents. Pretending otherwise makes you look naive to your team, who can see it coming perfectly well.
What actually happens in most brands is not a layoff. It is a hiring freeze plus natural attrition plus redeployment. Ecommerce support turns over fast — a team of twelve typically loses several people a year without trying. AI absorbing the repetitive tier means you stop backfilling. Within a year the team is smaller, more senior, and paid better per head.
The mistake is treating it as a cost exercise only. The work that remains is harder, and if you staff it with the same mix you had before, quality drops even as volume falls.
What happens to tier one
Tier-1 ecommerce work is order status, tracking, basic policy questions, and address changes. This is the tier AI handles best and the tier humans hate most. Removing it changes the job description substantially.
The agents who remain spend their time on genuinely difficult conversations: an angry customer whose replacement also arrived damaged, a wholesale buyer with a complicated allocation issue, a fraud edge case. This is more cognitively demanding and more emotionally draining per hour. Occupancy targets that made sense when half the queue was copy-paste do not survive contact with a queue that is entirely hard cases. Plan for lower target occupancy and more recovery time, or you will burn out the people you most want to keep.
Two internal paths are worth building deliberately. Some tier-1 agents become escalation specialists — deeper product knowledge, wider discretion on goodwill. Others move into the new roles below, which did not exist on your org chart two years ago.
The conversation designer
Someone has to decide how the AI talks. What it says when it cannot help. How it acknowledges a genuinely bad delivery experience without over-apologising. When it offers a discount and when it does not. Where the line sits between friendly and flippant for your brand.
This is a real role and it is closer to service design than to copywriting. The best conversation designers come from the support floor, because they have heard how customers actually react. They own the voice guidance, the escalation phrasing, the tone by category, and the ongoing tuning that comes out of transcript review.
In a team of ten to twenty, it is usually half a role. Above thirty, it is a full-time job and the ROI is obvious: this person's decisions are applied to every conversation in the queue, every day.
The knowledge owner
The highest-leverage role in an AI-native support team. The AI can only be as good as the knowledge it answers from, and knowledge decays. Policies change. Carriers change. A new product line arrives with a different returns rule and nobody tells support.
The knowledge owner writes articles, audits them against reality, chases policy decisions that have never been made explicit, and works the loop from transcript review to content fix. They need enough organisational standing to walk into ops and say "what is the actual rule here", and enough patience to write it down properly.
Give this person a formal review cadence: monthly audit of the top twenty articles, quarterly full sweep, and immediate updates triggered by any policy or product change. Make the trigger part of your product launch checklist so support is not the last to know.
Peak season when AI absorbs the spike
Traditional Q4 staffing is a painful ritual. Hire temps in September, train them through October, run them at 90% occupancy from Black Friday to Christmas, release them in January, repeat next year with none of the knowledge retained.
An AI agent changes the shape of this. The categories that spike hardest in peak — order status, delivery exceptions, "will this arrive before Christmas" — are exactly the categories AI handles well, and they scale without recruiting. That does not mean zero temps. It means fewer, and hired for a different job.
The peak pattern that works: keep your permanent team on escalations and complex cases, let the AI take the volume spike in the well-covered categories, and hire a small number of temps not for tier one but for peak-specific overflow on second-tier work, supported by Copilot so their ramp time is days rather than weeks. Copilot matters disproportionately here — a temp with a drafting assistant that cites the right policy is productive far sooner than one with a wiki and a buddy.
Do your knowledge work in September, not November.
Planning a BFCM week
Worked example — illustrative figures only. Take a brand that handles 3,000 conversations in a normal week with six agents, each handling around 60 conversations a day across a five-day week. In BFCM week, contact volume triples to 9,000. Historically that meant hiring and training four temps in October to get through five days in November.
Now assume an AI agent is live across order status, delivery exceptions, product and sizing questions, and returns policy — the categories that inflate most during peak — and that these account for around 65% of BFCM-week volume. Of that 5,850 conversations, assume the AI resolves 70% end to end and escalates the rest. That is roughly 4,100 resolved autonomously and 1,750 escalated.
Human-facing volume for the week is therefore the 3,150 conversations in uncovered categories plus 1,750 escalations, about 4,900. At 60 per agent per day that is roughly 16 agent-days, or a little over three agents working the full week — except escalations are harder than average, so plan at 40 per agent per day instead. That gives about 24 agent-days: five agents on shift across the week, with your sixth covering the weekend surge that BFCM creates.
Read what that changes. You are not hiring four temps; you might hire one, for weekend cover. Your permanent team works a hard week rather than a brutal one. And every number in that chain — coverage percentage, resolution rate, escalation difficulty — is something you can measure in October and refine before you commit to a headcount. All figures here are hypothetical and meant to show the shape of the calculation, not to predict your results.
Three planning rules follow from it. Lock your category coverage by mid-October so you are forecasting against measured performance rather than hope. Do not expand AI coverage into new categories during peak week itself; freeze changes the Monday before. And staff your escalation tier, not your total queue — the constraint in peak is senior judgement, not seats. Peak-specific content — cut-off dates, extended returns windows, gift-order handling — should be written and tested before volume arrives. Every year someone writes the Christmas cut-off article on December 3rd and wonders why the AI was answering wrong for a week.
Quality review when most conversations have no human
Old-school QA sampled agent tickets, scored them against a rubric, and coached. That model breaks when most conversations have no agent.
Replace it with three separate loops.
AI transcript review. Sample AI-resolved conversations weekly. Score on accuracy, tone, policy adherence and whether escalation should have happened. The output is not coaching — it is knowledge fixes and threshold changes. This is the loop most teams skip after month two, and it is the one that compounds.
Escalation quality review. Sample handoffs. Was escalation correct? Did the human get enough context? Did resolution take longer than it should have because the handoff was thin? This audits the seam between AI and human, which is where most bad experiences actually live.
Human conversation review. Traditional QA, but on a harder sample, with a rubric that rewards judgement, de-escalation and creative resolution rather than adherence to a script. Score fewer conversations, more deeply.
Run them monthly, rotate who reviews, and publish the findings to the whole team. Transparency here is what keeps agents from treating the AI as a black box that is coming for their job.
How to actually run the review loop
The theory above is easy to agree with and easy to let slide. Here is the operational version.
How many to sample. Enough to see patterns, few enough that it actually happens. In the first month, review every AI-resolved conversation in a newly live category — the volume is low because coverage is narrow, and the learning rate is highest. After that, move to a fixed weekly sample rather than a percentage: 40 to 60 AI-resolved conversations a week is workable for most teams and takes a competent reviewer about two hours. Stratify the sample rather than taking it at random. Roughly half should be pulled deliberately from the risky end — low confidence scores, conversations with three or more customer turns, anything with a negative sentiment flag, anything reopened, anything with a CSAT response below the midpoint. The other half random, because the risky-end sample will not show you the quiet mistakes.
For escalations, sample 20 to 30 a week. For human conversations, ten per agent per month is plenty when the sample is hard cases and the review is deep.
What to score. Keep the rubric to five dimensions and score each pass or fail, not one to five. Nobody calibrates a five-point scale reliably across reviewers.
- Factually correct. Did it state anything untrue about the order, the policy or the timeline?
- Policy adherent. Did it offer something it was not authorised to offer, or refuse something it should have granted?
- Escalation judgement. Should this have gone to a human, and did it?
- Tone fit. Would you be comfortable seeing this reply screenshotted on social media with your logo on it?
- Actually resolved. Did the customer's problem end, or did the conversation merely close?
That last one is the one people skip and the only one that predicts reopens. Score it by reading the final customer message. "Ok thanks" and silence are not the same as resolution.
Who reviews. Not the person who built the knowledge base, at least not alone — they will read their own articles into the transcript. Rotate a senior agent through the AI review each week, with the knowledge owner sitting in to take fixes rather than to score. Escalation reviews belong to the support lead, because the findings are usually about routing and handoff design rather than content. Human QA should be peer-led where your culture supports it; agents calibrate each other faster than managers calibrate them.
What happens to the output. Every failed dimension becomes exactly one of four actions: write or fix an article, adjust the tone guidance, change a threshold, or fix an integration. If a finding does not map to one of those, it is an observation, not a finding, and it should not go on the list. Put the actions in the same tracker your team already uses and close them within the week. A review loop whose output nobody actions is just an expensive reading group.
Scheduling when volume stops tracking traffic
For years you staffed against a simple proxy: order volume drives contact volume drives headcount. AI breaks the relationship, because the elastic part of the queue is now absorbed by something that does not have shifts.
Human volume becomes a function of escalation rate and complexity mix, not raw traffic. That has three practical consequences.
First, forecast escalations, not contacts. Your model input is AI-handled volume times escalation rate, plus categories not yet covered. Both move as coverage expands, so re-forecast monthly rather than quarterly.
Second, human demand flattens. The 9am Monday spike is mostly order status, which the AI eats. What remains is more evenly spread and less predictable in arrival but more predictable in volume. You can run thinner peak cover and fewer split shifts, which your team will like.
Third, coverage becomes about capability, not headcount. With a smaller human team you need every shift to include someone who can authorise a goodwill gesture, handle a legal threat, and answer a technical product question. Skills coverage is the constraint, so build rotas against skills and route with workflows rather than filling seats.
One more thing worth saying to your team out loud, early: measure the humans on resolution quality and customer outcomes, not on tickets closed. When the easy tickets are gone, ticket count measures nothing except how unlucky someone's queue was. Teams that keep a volume-based leaderboard after adopting AI end up with agents competing to grab the easiest escalations. Change the metric before you change the workflow.
Planning a rollout like this. Take it in order: your stack, then your taxonomy, then a first-30-days plan — start a free trial and build it against your own conversations.
Metrics that matter in ecommerce support
Most ecommerce support dashboards open on first response time. It is the wrong number to open on. First response time tells you how fast a queue is being touched, not whether anything is being fixed. In a store where an AI agent answers instantly, first response time collapses to near zero on day one and then tells you nothing forever. You need metrics that survive automation.
The set below is the one that actually moves decisions: what to staff, what to automate, what to fix upstream in the store rather than in the inbox. Each one has a definition, a formula, a directional sense of what good looks like, and — the part nobody writes down — the specific way it gets gamed once someone's bonus depends on it.
A note on measurement hygiene before the list. Every metric here needs a consistent denominator and a consistent time window. If you count conversations one week and tickets the next, or if merged tickets silently disappear from the numerator but stay in the denominator, you will produce trends that are entirely artefacts of your own bookkeeping. Pick the definitions once, write them into a shared document, and make Insights the single place the numbers come from. Two sources of truth for support metrics is the same as none.
Resolution rate
Definition. The share of conversations that ended with the customer's problem actually handled — answered, refunded, replaced, escalated to a resolution — without needing a second attempt.
Formula. Resolved conversations ÷ total conversations in the period. For AI specifically: conversations closed by the agent with no human touch and no reopen within a defined window (seven days is a sensible window for ecommerce, because delivery-related reopens cluster in the first week).
What good looks like. Directionally, a mature ecommerce deployment resolves the large majority of routine contact — order status, tracking, address edits, returns initiation, sizing, policy questions — without a human. What you should not do is chase a single headline percentage across all volume. A store selling furniture with white-glove delivery will have a structurally lower autonomous resolution rate than a store selling t-shirts, and that is correct, not a failure. Compare yourself to yourself, month over month, with the ticket mix held roughly constant.
How to instrument it. Tag every conversation with a close reason at close, make the field mandatory, and set a reopen webhook that reattaches a returning customer to the original conversation rather than spawning a new one.
How it gets gamed. Three ways, all common. First, counting deflection as resolution: the customer was shown an article, closed the chat in frustration, and emailed instead — two conversations, one counted as resolved. Second, aggressive auto-close: any conversation with no customer reply in 24 hours is marked resolved, which converts abandonment into success. Third, narrow scoping: excluding whole categories from the denominator ("returns tickets don't count, they're a fulfilment issue"). The defence is simple and mechanical. Count reopens against the original conversation. Count a follow-up on a new channel within 48 hours as a continuation, not a new ticket. And publish the denominator alongside the rate, always. The resolution rate playbook goes into the instrumentation in more detail, including how to define a reopen window that does not flatter you.
Contact rate per order
If you only track one number, track this one. Contact rate per order is the most useful metric in ecommerce support because it is the only one that ties the support queue to the business that generates it. Every other metric measures how well you handle volume. This one measures how much volume you should be handling at all.
Definition. The number of support conversations generated per order shipped.
Formula. Total inbound conversations in a period ÷ total orders in the same period. Express it as a percentage or as a ratio — 12% and 0.12 contacts per order mean the same thing. Be careful with the window: orders and contacts are offset in time. A contact today often belongs to an order placed six days ago. For weekly reporting the offset averages out. For a spike analysis, lag the order count by your median delivery time before you divide, or you will misread a shipping delay as a demand surge.
What good looks like. Lower, always, and the direction of travel matters more than the absolute level. A store at 30% contacts per order has a product, checkout or logistics problem, not a support problem. A store in the low single digits either has an exceptionally clean operation or is hiding its contact channels. The useful move is to break the rate down by cause. Contact rate per order attributable to delivery. To sizing. To damaged goods. To payment failures. Each of those has an owner outside support, and contact rate is the number that lets you hand them the bill.
How to instrument it. Pull order counts from Shopify on the same cadence as your conversation counts and compute the ratio in one place. Do not let two teams maintain two spreadsheets; the divergence will be the only thing anyone discusses in the meeting.
How it gets gamed. By making it hard to contact you. Burying the chat launcher, removing the support email from order confirmations, adding a five-step form before a human is reachable — all of these reduce contact rate per order and all of them increase refund requests, chargebacks and one-star reviews. Watch contact rate alongside CSAT and return rate. If contact rate falls while returns rise, you have not fixed anything; you have moved the cost somewhere you cannot see it.
WISMO share of volume
WISMO — "where is my order" — is the single largest category of ecommerce contact for most stores, and the most automatable.
Definition. The proportion of total inbound conversations whose primary intent is order status, tracking or delivery timing.
Formula. WISMO-tagged conversations ÷ total conversations. Tag it by intent classification, not keyword matching; "hasn't arrived" and "still waiting" and "tracking says delivered but it isn't" are all WISMO and share no keywords.
What good looks like. The share itself is not the goal — a store can have a high WISMO share and a perfectly healthy operation, because WISMO is cheap to resolve automatically. What matters is WISMO share crossed with resolution rate and with cost. High share, high autonomous resolution, near-zero cost per contact: fine. High share, low autonomous resolution: your resolution agent does not have live order and tracking data, and fixing that connection is the highest-leverage thing on your list this quarter. High share and rising: your carriers or your dispatch times have degraded, and support is the smoke alarm.
How to instrument it. Use intent classification at first inbound message, stored as a structured field rather than a free-text tag, and snapshot it before any agent can edit it. Keep the pre-edit value for audit.
How it gets gamed. By retagging. WISMO is the category everyone is judged on, so tickets quietly get reclassified as "general enquiry" or "shipping policy". Audit the tagging monthly with a random sample of fifty conversations read by a human. Fifty is enough to catch systematic drift and small enough that someone will actually do it.
Return rate influenced by support
Support does not only handle returns. It changes how many happen.
Definition. The difference in return rate between orders where the customer had a support conversation before the return window closed and orders where they did not.
Formula. (Returns from contacted orders ÷ contacted orders) − (returns from non-contacted orders ÷ non-contacted orders). A negative number means support conversations reduce returns. A positive number means they increase them.
What good looks like. Slightly negative on pre-purchase and sizing conversations; slightly positive on post-delivery damage conversations, which is correct because those returns should happen. Split the metric by conversation timing — pre-purchase, pre-delivery, post-delivery — or it will average out into meaninglessness. A good sizing answer before checkout is worth more than any refund workflow after it.
How to instrument it. Join conversations to orders on customer ID and order ID, not on email alone, since guest checkouts break email matching. Then tag each conversation with its timing relative to delivery so the split is computable.
How it gets gamed. Selection bias, mostly honestly rather than dishonestly. Customers who contact support are already more engaged and more likely to return items, so a naive comparison makes support look return-generating. Control for order value and category before you present this to a finance team, and be upfront that it is directional rather than causal unless you have run a proper holdout.
Revenue assisted by chat
Definition. Order value from customers who had a support conversation within a defined attribution window before purchase.
Formula. Sum of order value where a conversation occurred within N days before the order, N typically 7. Report it as both an absolute number and as a share of total revenue.
What good looks like. Non-trivial, and growing as a share when you put live chat on product and checkout pages rather than only on the help page. Pre-purchase questions convert at a much higher rate than general traffic because the customer is already deep in the funnel. The metric's real job is political: it is what converts support from a cost centre line item into something the commercial team defends during budget season.
How to instrument it. Stamp each conversation with the customer identity, then run a nightly job that looks forward N days for orders from that identity. Keep the job's window as a single configurable constant so nobody can quietly change it per report.
How it gets gamed. By widening the attribution window until every order qualifies. A 30-day window on a store with a 30-day repeat cycle attributes essentially all revenue to support. Fix the window, document it, and never change it mid-year. Also resist last-touch logic that credits support for a conversation that was purely a WISMO check on a previous order.
CSAT by ticket type
Aggregate CSAT is a vanity number. It moves when your ticket mix moves and stays flat when your service quality changes.
Definition. Customer satisfaction score segmented by conversation intent — returns, WISMO, product question, complaint, billing — rather than pooled.
Formula. Standard CSAT (positive responses ÷ total responses) computed within each intent bucket. Report response rate per bucket too, because a 96% score on eleven responses is noise.
What good looks like. Wide dispersion, which is the point. Refund refusals will always score badly and should. Sizing help will score well. If every bucket returns the same number, your survey is not discriminating and you should shorten it. Track each bucket's trend independently; a falling complaint-bucket CSAT with a rising overall CSAT means your complaints are getting worse while your easy tickets grew as a share of volume.
How to instrument it. Fire the survey from the platform on close, sampled randomly, carrying the intent field with it so the segmentation is automatic rather than reconstructed later from ticket text.
How it gets gamed. Survey suppression. Agents learn which conversations trigger a survey and route the ugly ones to a path that does not. Or surveys only fire on conversations closed with a specific macro. Fire surveys on a random sample of all closed conversations, controlled by the platform rather than by the agent, and the game stops.
Cost per resolved conversation
Definition. Fully loaded cost of the support function divided by conversations actually resolved.
Formula. (Salaries and contractor cost + platform and AI fees + tooling + allocated management and QA time) ÷ resolved conversations. Fully loaded means fully loaded. If you exclude the ops manager and the BPO onboarding fee, you are not measuring cost, you are measuring payroll.
What good looks like. Falling, with an important caveat: it should fall because automation absorbs routine volume, not because humans are handling more conversations per hour. The second kind of improvement shows up two quarters later as attrition and as a CSAT decline in the complaint bucket. Segment the cost: cost per autonomously resolved conversation and cost per human-resolved conversation are different businesses, and the blended figure hides the mix shift that is doing all the work.
How to instrument it. Maintain a single monthly cost basis owned by finance, loaded into the same reporting layer as your conversation counts, with an explicit list of what is in and what is out written at the top of the sheet.
How it gets gamed. By moving costs out of the numerator. Contractor hours booked to a different cost centre, AI spend classified as engineering, a self-service portal built by the web team whose cost never appears. Define the cost basis once with finance, in writing, and reuse it.
Peak-day concurrency
Staffing models built on daily averages break on the days that matter.
Definition. The maximum number of simultaneously active conversations at any point during your busiest hour of your busiest day.
Formula. Take the peak hour and compute concurrent open conversations at one-minute resolution; take the maximum. Then compute peak concurrency ÷ average concurrency as a peakiness ratio. That ratio, not the raw number, is what you plan against.
What good looks like. The number itself is store-specific. The ratio is where the insight lives. A store whose Black Friday peak hour runs at eight times its average hour cannot staff for it and should not try; that is a capacity problem that end-to-end automated resolution solves structurally, because autonomous resolution has no concurrency ceiling. A store with a ratio near two can staff through it with contractors.
How to instrument it. Sample open-conversation counts on a fixed one-minute tick and retain the raw series for at least thirteen months, so that next year's capacity plan is built on last year's actual curve rather than on someone's recollection of it.
How it gets gamed. By averaging the peak. "Peak day volume" stated as a daily total hides the fact that half of it arrived in ninety minutes after a promo email landed. Always measure concurrency at sub-hourly resolution, and always report the timestamp alongside it so you can tie the spike to the campaign that caused it.
Refund rate and refund cycle time
Refunds are usually owned by finance and measured by finance, which means nobody measures the part support controls: how long a refund takes from the customer's first request to the money leaving.
Definition. Refund rate is the share of orders refunded, in whole or in part. Refund cycle time is the elapsed time from first customer request to refund issued, measured in hours, not business days.
Formula. Refund rate: refunded orders ÷ total orders, reported separately for full and partial refunds because they have different causes. Refund cycle time: median and 90th percentile of (refund issued timestamp − first request timestamp). Use the median for the typical experience and the 90th percentile for the complaints, because the 90th percentile is the population that writes reviews.
What good looks like. Refund rate should be stable and explainable; a rising rate is almost always a product, sizing or fulfilment signal rather than a support one. Refund cycle time should be short and, more importantly, tight — a median of six hours with a 90th percentile of nine days means you have two processes, one automated and one that quietly waits for a weekly finance run. Find the second process and kill it. The largest single lever here is the approval threshold: a policy that lets an agent, or the AI agent itself, issue refunds below a defined value without escalation collapses the cycle time distribution, and the cost of the occasional wrong call is usually far less than the cost of the delay.
How to instrument it. Capture the request timestamp from the conversation, not from the refund record in Shopify or Stripe, because the refund record only knows when someone got around to pressing the button. Join the two on order ID.
How it gets gamed. By starting the clock late. If cycle time begins when a ticket reaches the refunds queue rather than when the customer first asked, every hour of triage delay is invisible. Anchor to first customer contact about that order, always. The other trick is measuring in business days, which turns a Friday afternoon request into a Monday morning start and makes weekend delay disappear from the report while the customer sits through all of it.
Contact rate by traffic source
Acquisition channel predicts support volume more reliably than almost any other variable, and almost nobody measures it.
Definition. Contact rate per order, segmented by the channel that acquired the order — paid social, paid search, organic, email, marketplace, retail partner.
Formula. For each channel: conversations attributable to orders from that channel ÷ orders from that channel. Attribution needs to be the same model marketing already uses, or you will spend every meeting arguing about attribution instead of about support.
What good looks like. Wide variation, and you want to see it. Paid social traffic, especially from broad prospecting audiences, tends to generate materially more contact per order than organic or email traffic: the customer arrived with less context, had less exposure to your policies, and often bought on impulse from a single creative that promised something the product page then qualified. Marketplace orders generate a different profile again, because the customer's relationship is with the marketplace and their expectations about delivery and returns are set by that marketplace rather than by you. Organic and email traffic, from customers who found you deliberately or who have bought before, is usually the cheapest to support.
Once you can see this, two things become possible. First, you can put a support cost against a channel and hand marketing a fully loaded CAC that includes it — a channel that looks efficient on blended CAC can be unprofitable once support load is charged to it. Second, you can fix creatives. If one ad set produces triple the contact rate of the others, read fifty of those conversations and you will usually find the ad itself is making a claim the product does not support. That is a thirty-minute fix worth more than a month of queue optimisation.
How to instrument it. Persist the landing UTM parameters onto the order at checkout, then join conversations to orders and roll up by source. If you cannot persist UTMs, the first-touch channel from your analytics stack is an acceptable approximation for directional work.
How it gets gamed. By attribution shopping. Whoever owns a channel will prefer the attribution model that assigns their channel fewer tickets. Fix the model jointly with marketing before you run the analysis, publish it, and never recompute retroactively under a new model to win an argument.
Summary table
| Metric | Formula | Directional target | Main gaming risk |
|---|---|---|---|
| Resolution rate | Resolved ÷ total conversations | Rising, with reopens counted | Deflection counted as resolution |
| Contact rate per order | Conversations ÷ orders shipped | Falling, broken out by cause | Hiding contact channels |
| WISMO share | WISMO ÷ total conversations | Any level, but near-fully automated | Retagging into other categories |
| Return rate influenced | Contacted return rate − non-contacted | Negative pre-purchase | Ignoring selection bias |
| Revenue assisted by chat | Order value within N days of a chat | Growing share of revenue | Widening the attribution window |
| CSAT by ticket type | Positive ÷ responses, per intent | Wide dispersion, each trending up | Survey suppression |
| Cost per resolved conversation | Fully loaded cost ÷ resolutions | Falling via automation mix | Costs moved off the numerator |
| Peak-day concurrency | Max simultaneous conversations | Ratio to average under control | Averaging the peak away |
| Refund rate and cycle time | Refunded ÷ orders; request-to-issued elapsed | Stable rate, short and tight cycle | Starting the clock late |
| Contact rate by traffic source | Conversations ÷ orders, per channel | Wide, explained variation | Attribution shopping |
Build these ten into one reporting board, review them weekly, and stop reporting anything else to the exec team. Ten numbers that mean something beat forty that do not.
Switching from Gorgias, Zendesk or Intercom
Most stores reading this already have a helpdesk. Switching is expensive in attention even when it is cheap in money, so the honest framing is not "our tool is better" but "here is who each incumbent genuinely suits, here is why teams leave, and here is what they miss afterwards." If you are well served where you are, stay. The migration runbook in the next section exists for the teams who are not.
One caveat that applies to all three: pricing models in this market change often. Treat everything below as a description of shape, not of price. Confirm current figures on each vendor's own site before you build a business case on them.
Gorgias
Who it suits. Gorgias was built for Shopify stores and it shows. If you are a single-brand, single-region store doing high volume on Shopify with a support team of three to ten people, it fits the shape of the problem well. The Shopify sidebar is good. Order actions from inside a ticket work. The social channel handling — comments, DMs, ad comments — is more thorough than most general-purpose helpdesks, and for stores where a large share of contact starts on Instagram that is a real advantage. Teams looking for a Gorgias alternative usually are not leaving because the product is bad; they are leaving because they outgrew the shape it was built for.
Why ecommerce teams leave. The reasons cluster. Multi-store and multi-brand operations find the model strains once you are running several storefronts with different policies and a shared team. Reporting is often cited as thin once you want anything beyond standard views — the cost-per-resolution and contact-rate-per-order analysis in the previous section is hard to assemble natively. Automation tends to be rule-based, which is fine until your rule count passes a few hundred and nobody remembers why rule 217 exists. And the ticket-consumption pricing shape means cost scales with contact volume, which is precisely the variable a good support operation is trying to drive down; teams describe it as being billed for their own inefficiency.
What they miss after switching. Be honest about this. The social inbox depth is a genuine trade-off. If Instagram comment moderation and ad-comment hiding are core to your operation, check the replacement carefully before you commit. Teams also miss the specific muscle memory of an interface their agents have used for years — that cost is real and lasts roughly a month. And the Shopify app ecosystem around Gorgias is mature; some niche integrations you rely on may not have a direct equivalent.
What the migration actually breaks. Social. Gorgias pulls Instagram and Facebook comments, DMs and ad comments into the ticket stream, and if your team has been working those in the same queue as email, moving means splitting that workflow until you have rebuilt it. Plan for the social channels to be handled natively for a period. The second breakage is the rule library: teams typically have hundreds of rules accumulated over years, many of them interacting, and there is no clean one-to-one translation into a system where an AI agent reasons rather than matches. Treat the rule export as documentation of intent, not as a migration artefact, and rebuild from the intent. The third is macro-embedded variables — merge fields that pull order numbers and tracking links — which will not carry across and need rewriting as knowledge articles with live context instead.
What Aftersales does differently. The core difference is architectural. Agent is designed to resolve conversations end to end rather than deflect them into an article, and it runs in the same queue as your humans in the unified inbox rather than as a separate bolt-on channel. Rules still exist in Workflows automation, but they carry less load because the agent reasons over Knowledge and live Shopify order context instead of needing a rule per scenario. Pricing starts from seats rather than from ticket consumption, so bringing volume down does not leave you paying for a tier you no longer need.
Zendesk
Who it suits. Zendesk suits large, process-heavy organisations with dedicated admins. If you have compliance requirements that demand granular permissions, if you run support across many business units, if you need deep custom object modelling and an internal team that enjoys building it — Zendesk is a serious piece of enterprise infrastructure and there is a reason it has been the default for fifteen years. Its API surface and app marketplace are the broadest in the category.
Why ecommerce teams leave. Weight. Zendesk asks for administration, and small commerce teams do not have an admin. The Shopify relationship is an integration rather than a native understanding of what an order is, so agents end up tab-switching to the Shopify admin to do the actual work, which is where handle time goes. Configuration sprawl is the other recurring complaint: trigger chains, automations, business rules and macros accumulate over years until nobody can safely change anything. Teams also cite the commercial experience — multi-year commitments, add-on modules for capabilities they assumed were core, and AI features priced separately from the seats.
What they miss after switching. Reporting depth, most often. Zendesk's analytics layer, once configured, will answer questions that simpler tools cannot. If you have an analyst who has built two years of custom dashboards, moving them is not free. Teams also miss the marketplace breadth and, in larger organisations, the granularity of permissions and audit configuration. And the sheer institutional familiarity — every experienced support hire has used it.
What the migration actually breaks. Custom objects and custom fields. Any modelling you have done — a returns object, a warranty record, a custom ticket form with twenty conditional fields — does not port. It has to be re-expressed, usually more simply, and the re-expression is a design exercise rather than a data move. Reporting breaks next: historical dashboards do not migrate, so decide early whether you need historical trend continuity (in which case export the aggregate numbers, not just the raw tickets, into your own warehouse) or whether you are content to restart the series. Third, integrations built on marketplace apps break silently. Inventory each installed app, identify who owns it and what it does, and expect to find two or three that nobody remembers installing and one that turns out to be load-bearing. Finally, granular agent permissions rarely map one for one; audit who could see what before you cut over, particularly if you have contractors with restricted views.
What Aftersales does differently. Less configuration to reach the same place. Views, SLAs, assignment, tags and side conversations into Slack are in Inbox without a build project, and switching between conversations is sub-100ms, which matters more to daily handle time than any feature list. Copilot drafts replies with cited sources and translates in both directions, so a two-language team does not need a separate localisation workflow. Compliance posture covers SOC 2 Type II and GDPR, with details on the security page. Existing Zendesk data is not stranded: Zendesk is a supported integration, which makes parallel running during migration practical rather than theoretical.
Intercom
Who it suits. Intercom is strongest where support and product sit close together — SaaS onboarding, in-app messaging, product tours, lifecycle campaigns. If your support motion is genuinely part of your product experience, Intercom's messenger is excellent and the proactive messaging layer has no real equivalent in helpdesk-first tools. For commerce, the fit is looser, because the moment of truth is a delivery exception rather than an in-app activation.
Why ecommerce teams leave. Pricing, overwhelmingly, and specifically the shape of it. Intercom's Fin agent is priced per resolution — publicly $0.99 per resolution at time of writing. Per-resolution pricing is defensible in principle: you pay for outcomes. In ecommerce it creates a structural problem, because ecommerce contact is dominated by high-volume, low-complexity WISMO. Each of those resolutions is worth very little to you and costs the same as a complex one. Worse, your bill now rises with your contact rate per order, which means a carrier failure or a delayed dispatch run shows up twice: once as angry customers, once as an invoice. And the interaction between seat pricing and resolution pricing means forecasting a peak-season bill requires forecasting both variables, which is why finance teams dislike it. The broader trade-offs are covered in our comparison of Intercom alternatives.
What they miss after switching. The messenger. Intercom's in-product messaging, tours and outbound lifecycle campaigns are a category of their own, and a helpdesk will not replace a marketing automation layer. If you use Intercom for onboarding sequences as well as support, understand that you are unbundling two jobs and may need a second tool for the first. Teams also miss the article editor and the help centre presentation, which are genuinely well made.
What the migration actually breaks. Outbound. If you have lifecycle messages, product tours, banners or onboarding sequences running through the messenger, those stop when the messenger stops, and they are often owned by marketing rather than support — meaning the first time anyone notices is when a campaign silently ceases. Inventory every outbound message and tour before you touch anything, and agree with their owners where each will live afterwards. The second breakage is the help centre, which is usually hosted on an Intercom subdomain or a mapped custom domain; if it is on a subdomain you do not control, you cannot set redirects from it and your only options are to keep it alive as a pointer or to accept the traffic loss. Decide that deliberately, not by accident. Third, user attributes and events that Intercom collected — the custom data your team segments on — need a new home, and that is an engineering task, not a configuration one.
What Aftersales does differently. A seat-based starting point from $24 per seat per month, so your bill tracks the size of your team rather than the volume of your delivery problems. Resolution is the product, not the meter. See the current plans for detail. On the commerce side, the Shopify integration gives the agent live order context, fulfilment and customer context, which is what turns a WISMO conversation into an actual answer rather than a link to a tracking page.
Shared Gmail or Outlook inbox
This is where most brands under roughly $5M GMV actually start, and it is worth taking seriously rather than dismissing. A support@ mailbox with two or three people in it, plus a Shopify tab open in the next window, handles a surprising amount of volume competently and costs almost nothing.
Who it suits. Founder-led brands and small teams. Under a few dozen conversations a day, with one or two people who trust each other, a shared mailbox has real advantages: everyone already knows how to use it, there is no configuration, threading is native, and search is excellent. If that describes you and nothing is on fire, the right advice is to stay put and spend the money on inventory.
Why teams leave. Always the same sequence, and it starts before anyone admits it. Collision: two people reply to the same customer with different answers, and you find out from the customer. No ownership: a message gets read, mentally noted, and never answered, because "read" and "handled" are the same state in a mailbox. No visibility: you cannot say how many conversations you handled last week, what they were about, or how long they took, so you cannot make any of the decisions in the metrics section above. No continuity: when someone leaves, their context leaves with them, and the archive is organised by thread rather than by customer. And escalation is a forward with the word "thoughts?" on top, which loses the customer's history the moment anyone replies to the wrong part of the chain.
The tipping point is usually a season. A brand handles November in a shared inbox once, badly, and buys a helpdesk in December.
What they miss after switching. Simplicity, genuinely. A helpdesk introduces concepts — views, statuses, assignment, tags — that a mailbox does not have, and for a two-person team those concepts are overhead before they are leverage. Teams also miss the ease of looping in anyone in the company by just adding them to the thread; side conversations solve this properly, but it is a new habit. And the first two weeks feel slower, because muscle memory built over years does not transfer. Expect that and do not panic on day three.
What the migration actually breaks. Mailbox migration is technically the easiest and operationally the trickiest. Email history imports as threads with no structure — no tags, no intents, no resolution status — so you start your metrics from zero regardless of how much history you carry over. Accept that and set the baseline on day one rather than pretending the old data is comparable. Filters and labels that individuals built for themselves do not port and usually should not; they encoded a personal workflow that a shared queue replaces. Personal aliases are the real trap: if customers have been replying directly to a founder's personal address, that traffic will keep arriving there indefinitely, invisible to the queue and to every report you build. Audit personal mailboxes for customer conversations before you cut over, and set up forwarding from them. Finally, anything running as a mail rule — auto-replies, out-of-office, vacation responders — needs rebuilding, and a forgotten vacation responder firing alongside a helpdesk auto-acknowledgement sends every customer two emails.
What Aftersales does differently. The move is deliberately small. Forward support@ into the shared helpdesk and the mailbox keeps working exactly as before, except conversations now have an owner, a status and a history attached to the customer rather than to a thread. From there you add capability at your own pace: a handful of help articles, then an autonomous AI agent handling order status against live Shopify data, then reporting once you have enough volume for the numbers to mean anything. Starting from $24 per seat per month, a three-person team is at a price point where the decision is about whether you want visibility, not about whether you can afford it. The 14-day trial exists precisely for this case: run it in parallel with the mailbox for two weeks and see whether the structure helps or gets in the way.
| Platform | Best for | Pricing model | AI agent | Native Shopify context |
|---|---|---|---|---|
| Shared Gmail or Outlook | Founder-led brands under a few dozen conversations a day | Included in existing productivity suite | None | None; manual tab-switching |
| Gorgias | Single-brand Shopify stores with heavy social volume | Ticket-consumption based; confirm current tiers with the vendor | Included, rule-and-article led | Deep, purpose-built |
| Zendesk | Large process-heavy orgs with dedicated admins | Seat-based with paid add-on modules; confirm with the vendor | Add-on, priced separately | Integration rather than native |
| Intercom | SaaS and in-product support with lifecycle messaging | Seats plus $0.99 per Fin resolution (publicly stated) | Fin, per-resolution | Integration rather than native |
| Aftersales | Ecommerce teams wanting autonomous resolution on one queue | Seat-based from $24 per seat per month | Agent, resolves end to end in the shared queue | Native Shopify order and fulfilment context |
Pricing models above are described qualitatively on purpose. Vendors change plans, rename tiers and reprice AI features regularly — check each vendor's own pricing page before you commit numbers to a board deck.
Do not switch if your current tool is working and your pain is elsewhere. If your contact rate per order is high because of carrier performance, warehouse pick errors or a confusing sizing chart, a new helpdesk will not fix it — you will spend a quarter migrating and arrive at the same volume. Fix the upstream cause first. Switch when the tool itself is the constraint: when pricing punishes volume you cannot control, when configuration has calcified, or when your agent cannot resolve without a human because it has no live order data.
Migration runbook
A helpdesk migration goes wrong in predictable ways. The list below is ordered as you should actually do it, not as a vendor would present it. Assume four to six weeks end to end for a team of ten, most of which is parallel running rather than setup.
- Export conversation history before you do anything else. Every incumbent offers an export, usually via API and usually rate-limited. Start it early because it takes longer than you expect. Decide explicitly what you need: most teams need the last 12 to 24 months of conversation bodies, participants, timestamps, tags and resolution status, plus attachments. Store the raw export somewhere durable and untouched — a bucket, not a laptop — before you transform anything. You will need it again. Check the export for the things exports quietly drop: internal notes, side conversations, merged ticket lineage and attachment files as opposed to attachment URLs that expire.
- Audit your macros before you migrate them. Do not port macros one for one. Pull the usage counts and you will typically find a long tail that has not been used in a year. Keep the top tier that covers most of your volume, delete the rest, and rewrite what you keep as knowledge base articles rather than canned responses. This is the single most important conceptual shift in the migration. A macro is text an agent pastes. A knowledge article is a source an AI agent reasons from and cites. One article on the returns policy, written properly, replaces a dozen macros that each said a slightly different thing about it — and the inconsistency between those dozen macros was already costing you CSAT.
- Rebuild views and SLAs deliberately, not by copying. List the views your team actually opens each day. It is usually five or six, regardless of how many exist. Rebuild those in the shared queue, plus your escalation and breach views. Re-derive SLAs from current performance rather than inheriting targets nobody has met since 2023. Set a first-response target you can hit at peak, not one you can hit in February. Configure assignment and routing in Workflows at the same time, and keep the rule count low on purpose — every rule you add now is a rule someone has to understand in eighteen months.
- Preserve ticket IDs and customer identity. Two separate problems. For ticket IDs, keep the original reference as a field on the imported conversation and make it searchable, so that when a customer quotes "ticket 48219" from an old email your agent can find it. Do not try to make new IDs continue the old sequence; it creates collisions and confusion. For customer identity, deduplicate on email as primary key, with phone and Shopify customer ID as secondary. Run the dedupe on the export, not after import — fixing merged-in-error customer records post-import is painful. Expect a few percent of records to need manual judgement and budget an afternoon for it.
- Connect the systems of record before you send a single conversation. Shopify first, since order, fulfilment and customer context is what most conversations need. Stripe if payments and refunds sit there. Slack for side conversations. Verify with real order lookups, including the awkward cases: guest checkouts, multi-shipment orders, partially refunded orders, subscriptions if you run them. An AI agent with a half-connected order source produces confidently wrong delivery answers, which is worse than no agent at all.
- Run both systems in parallel. This is the step teams try to skip and should not. For two to three weeks, route a defined slice of volume to the new system — a single channel, or a percentage of chat, or one region's email — while the incumbent handles the rest. Two rules make parallel running work. First, no conversation lives in both places; route at the entry point, never mid-thread, or you will fracture history. Second, agents need both tabs open and a clear rule for which one they check. Compare resolution rate, CSAT by ticket type and handle time between the two populations weekly, and only widen the slice when the new system is at parity or better.
- Cut over the widget. The website chat widget is the highest-visibility switch and the easiest to roll back, so do it as a discrete step on a low-traffic day, mid-week, mid-morning. Deploy the new snippet, remove the old one in the same release, and watch conversation volume for the first hour. If volume drops to zero, your snippet is not loading; if it doubles, the old widget is still present somewhere and you are running two launchers. Check the checkout and post-purchase pages specifically, because those are usually templated separately and are the ones that matter most.
- Redirect forwarding addresses last. Email is the step with the longest tail. Change the forwarding on
support@,help@,returns@and every alias you have accumulated — and you have more aliases than you think, so audit the mail admin rather than working from memory. Leave the old mailbox receiving and monitored for at least 30 days; some customers reply to threads that are months old. Update the reply-to and from addresses on order confirmations, shipping notifications and return labels in Shopify, since those generate more inbound than your contact page does. Check SPF and DKIM alignment for the new sending domain before cutover, not after, or your first week of replies will land in spam.
- Handle the help centre URLs properly. If you are moving your public help centre, this is where teams lose organic traffic they never get back. Export the full URL list from the incumbent. Map every old URL to a new one — one to one where possible, to the closest equivalent where not. Implement permanent redirects; do not leave old URLs 404ing and do not redirect everything to the help centre homepage, which search engines treat as a soft 404 and which is useless to the customer anyway. Keep the article titles and H1s stable during the move so that ranking signals carry over, and change them later if you want to, once traffic has settled. Submit the new sitemap and monitor impressions weekly for six weeks. If you have inbound links from blogs or forums to specific articles, those redirects are the ones to double-check by hand.
- Decommission on a schedule, not on a whim. Keep the incumbent in read-only for a defined period — 60 to 90 days is typical — before cancelling. Export again at the end, because your first export will have missed something. Only then terminate, and check your contract notice period well in advance; annual agreements frequently require 30 to 60 days' written notice, and missing it costs you another year.
Rollback. Define the abort conditions before you start, in writing, with a named decision-maker. Sensible triggers: CSAT in any ticket type falls materially below the parallel-run baseline for more than three consecutive days; the AI agent produces a factually wrong order or refund answer that reaches a customer more than a handful of times; or backlog grows for 48 hours straight. Rollback itself is mechanical if you have staged properly: revert the widget snippet, revert the email forwarding, and stop routing to the new queue. The conversations already in the new system stay there and get worked to completion — never migrate live conversations backwards. Keep the incumbent's contract active through the entire parallel period specifically so rollback stays available; cancelling early to save one month of fees is how a manageable setback becomes a crisis. Then diagnose, fix, and restart the parallel run with a smaller slice.
What ecommerce support actually costs
Support budgets are usually wrong because they count the platform and the salaries and stop. The platform is rarely the largest line. Here is how to build the model properly, followed by two worked examples. All figures below are illustrative round numbers chosen to show how the mechanics behave — they are not quotes, benchmarks or claims about any vendor's actual pricing.
Seats. The visible cost. Headcount multiplied by per-seat licence fees, multiplied by twelve. Aftersales starts from $24 per seat per month; see pricing for current plans. The trap is counting only full-time agents. Count everyone who needs access: the ops manager, the QA reviewer, the warehouse supervisor who checks damage claims, the two people in marketing who answer pre-sales questions in December. That list is usually 30 to 50% longer than the support headcount.
AI resolution pricing. Where the model is per-resolution, this is a variable cost driven by contact volume. It is genuinely attractive at low volume and genuinely dangerous at high volume, because it scales with exactly the thing your operation is meant to be reducing. Model it at your peak month, not your average month.
Onboarding and implementation. Sometimes a line item, always a cost. Even where no fee is charged, someone internal spends weeks on configuration, knowledge base writing and testing. Price that at their loaded hourly rate. Rewriting a macro library into proper knowledge articles is typically two to four weeks of one person's time and is the work that most determines whether your AI agent performs.
Integration work. Standard connectors are free in effort as well as money. Anything non-standard — a custom returns portal, a 3PL that speaks only EDI, a legacy ERP reachable only by API — is engineering time. Get a real estimate from whoever would build it before you sign anything.
Contractor hours at peak. The line most budgets omit and most Novembers reveal. Seasonal agents cost more per hour than permanent staff once you include agency margin, and they cost even more in supervision: a contractor answering their first week of tickets consumes senior agent time at roughly one hour per five hours worked. Budget the supervision explicitly.
The cost of a bad refund decision. The hidden line with the widest variance. A wrongly refused refund produces a chargeback, a card network fee, the staff time to contest it, and often a public review. A wrongly granted refund is straightforward margin loss plus the shipping already spent. Both are decision-quality costs, and both fall when agents have consistent policy and live order context in front of them. Estimate it as: (chargebacks per month × fully loaded cost per chargeback) + (incorrect goodwill refunds × average refund value). Most teams have never computed this and are surprised by the size.
The three pricing shapes. Seat-based is predictable, scales with headcount, and rewards you for automating — your bill does not move when volume spikes. Its weakness is that a small team handling enormous volume pays the same as a small team handling little. Per-resolution aligns payment with outcomes and is cheap for low-volume stores, but it converts your operational problems into invoice line items and makes peak forecasting hard. Hybrid — a seat base plus per-resolution AI charges — combines the predictability problem of the second with the floor cost of the first, and is the shape most incumbents are converging on.
Worked example: a small brand. Illustrative figures. A store shipping 2,000 orders a month at a 10% contact rate generates 200 conversations. Three seats at $30 per seat per month is $90 a month, call it $1,080 a year, plus one part-time agent at $30,000. Under a seat-based model, total platform cost is about $1,080 a year and is flat regardless of whether contact volume doubles. Under a per-resolution model at $1 per resolution with 80% of those 200 conversations resolved autonomously, that is 160 resolutions, $160 a month, $1,920 a year — plus any seat fees on top. At this scale the difference is small and either model is defensible. The point of the example is that at low volume, pricing model barely matters; pick on product fit.
Worked example: a brand at peak. Illustrative figures. Now take a store shipping 40,000 orders in November with a contact rate that rises from 10% to 14% because carriers are congested. That is 5,600 conversations in one month against roughly 4,000 in a normal month. Assume eight permanent seats plus six contractors for six weeks. Under a seat-based model at $30 per seat, fourteen seats cost $420 for the month; the contractors cost real money but the platform does not move. Under a per-resolution model at $1 per resolution with 80% autonomous resolution, 4,480 resolutions cost $4,480 in November against $3,200 in a normal month — an extra $1,280 caused entirely by a carrier problem you did not create and cannot control. Over a year with three peak months, the gap compounds into five figures on top of seat fees. Again, illustrative — but the mechanic is the whole argument: per-resolution pricing bills you more precisely when your operation is under the most stress.
Worked example: heavy seasonal concentration. Illustrative figures, and the most important of the three. Take a gifting brand that ships 120,000 orders a year, of which 48,000 — 40% — land in the six weeks from mid-November. Annualised, that is 10,000 orders a month, and at a 10% contact rate the budget says 1,000 conversations a month, or 12,000 a year. Every vendor conversation and every internal plan gets built on that average, and every part of it is wrong.
What actually happens: the six peak weeks produce 48,000 orders, and the contact rate does not hold at 10% because delivery deadlines make customers anxious — call it 15%. That is 7,200 conversations in six weeks, roughly 1,200 a week against a non-peak run rate of around 170 a week. The remaining 46 weeks carry 72,000 orders at 9%, about 6,500 conversations. Total for the year: under 14,000, close enough to the 12,000 estimate that the annual figure looks fine. The annual figure was never the problem.
Now price it. Under a per-resolution model at $1 per resolution with 80% autonomous resolution, the six peak weeks cost about $5,760 and the rest of the year about $5,200 — so more than half the annual AI spend falls inside a six-week window, arriving as two invoices during the quarter when cash is already committed to inventory and paid media. Under a seat-based model, the platform cost is flat and the peak cost is entirely staffing, which you control and can decide about in October.
The staffing consequence is sharper than the pricing one. Peak weekly volume is roughly seven times the normal run rate. No headcount plan bridges that: hiring for peak means carrying six idle agents for ten months, and hiring contractors means onboarding them in the worst possible month. Only two things actually close a 7x gap — autonomous resolution, which has no concurrency ceiling, and reducing contact rate per order upstream before the season starts. Figures are illustrative, but the shape recurs in every seasonal brand: the annual average is a number that describes a year nobody experiences.
The practical conclusion is not that one model is universally cheaper. It is that you should model your own peak month, with your own contact rate per order, before you sign anything. Take last November's conversation count, apply each vendor's stated model, and compare the peak-month figure rather than the annual average. Then add onboarding, integration and contractor supervision, which are the lines that actually decide whether the year comes in on budget. To build that model against your own numbers, start a free trial and run your real volume through it for fourteen days.
Risk, data and compliance for ecommerce
Most ecommerce support teams treat data governance as something the legal team worries about once a year. Then an AI agent goes live, starts reading order records, and suddenly every decision about what the model can see becomes an operational decision made dozens of times a minute. It is worth getting deliberate about this before launch rather than after an incident.
Start with a simple question: what does the AI actually need to resolve the conversation in front of it? For a where-is-my-order question, it needs the order status, the carrier, the tracking number, the shipping address city and postcode, and the promised delivery window. It does not need the full payment method on file, the customer's date of birth, their lifetime order history in unredacted detail, or any internal fraud score. For a return, it needs the order contents, the purchase date, the return window and whether the item is eligible. It does not need the customer's other orders across your other brands.
The principle is boring and correct: scope the data the agent can read to the task, not to the record. In practice that means configuring the Shopify and Stripe connections so the agent retrieves order-level and fulfilment-level fields, not the entire customer object. Most teams over-provision on day one because it is easier, then never go back and tighten it. Do the tightening in week two, while the conversation logs are still fresh enough to show you which fields were actually used.
PII minimisation in transcripts
Every conversation you store is a small data liability. A transcript that contains a name, an email, a full shipping address, an order number and a phone number is a complete identity record sitting in a system that dozens of agents and several integrations can read. Multiply that by a few hundred thousand conversations a year and you have built a shadow customer database nobody audits.
Three practical mitigations. First, redact at ingest rather than at export — strip card-like number patterns, national identifier patterns and anything that looks like a password from the stored transcript before it is written, not when someone requests a copy. Second, prefer references over values: if the agent needs to confirm the delivery address, have it confirm the last line rather than repeat the whole thing, and store the order ID instead of the address in the conversation record. Third, set a retention period and actually enforce it. Most ecommerce support conversations have zero operational value after ninety days. Disputes and chargebacks are the exception, and those should be tagged and retained deliberately, not by accident of a blanket policy that keeps everything forever.
Retention also affects your analytics. If you delete transcripts at ninety days, make sure the aggregate metrics in reporting and analytics are computed and stored separately, so deleting conversation bodies does not destroy your historical resolution trend. Aggregates without identifiers are a much lighter liability than raw text.
Card data must never enter a chat
This is the one rule with no exceptions. Payment card numbers, expiry dates and CVV codes should never be typed into a support conversation, by a customer, by an agent, or by a bot.
Customers will do it anyway. Someone will paste sixteen digits because their card was declined and they think you can retry it manually. It happens on every platform, in every vertical, at every volume. So the question is not whether it will happen but what the system does in the two seconds after it does.
Card numbers in a chat window are a containment problem, not a policy problem. Detect the pattern at ingest, replace it with a masked token before the message is persisted, notify the agent that redaction occurred, and have the agent respond with a secure payment link rather than acknowledging the digits. If a raw number ever reaches your transcript store, you have widened your compliance scope to include your helpdesk — which is exactly what you do not want.
Practically, that means three things. Pattern-match and mask on the way in, so the unredacted string never lands in durable storage. Train the AI agent and your human team on the same script: never confirm, never repeat back, never ask for the remaining digits. And route every payment retry to a hosted payment page generated through Stripe, so the card is captured by the processor and not by you. Aftersales holds SOC 2 Type II, supports GDPR obligations and is HIPAA-ready; it is not, and does not claim to be, a card data environment. Keeping cards out of the conversation is what keeps that true.
The same logic applies to bank details, government identifiers and gift card codes. Gift card codes in particular get pasted constantly, and an unredacted code in a transcript is a bearer instrument sitting in plain text. Mask them.
GDPR obligations for EU shoppers
If you sell into the EU or UK, three obligations bite directly on your support stack.
Right to erasure applied to conversation history. When a customer asks to be deleted, most teams remember to delete the account in Shopify and forget the two years of chat transcripts in the helpdesk. Those transcripts contain personal data and are in scope. You need a deletion path that finds every conversation associated with an identifier, removes the message bodies and the identifiers, and leaves behind only non-identifying aggregates. Test this before you need it. Ask the vendor how long deletion takes to propagate to backups and search indexes, and get the answer in writing.
Lawful basis for proactive messaging. Proactive delivery-exception messages are the highest-value automation in ecommerce support, and they are also outbound messages to a person. If the message is strictly about performing the contract — your parcel is delayed, here is the new date — you are generally on solid ground under contractual necessity. The moment the same message carries a discount code, a cross-sell or a review request, it changes character and you need consent or a defensible legitimate interest assessment. Keep the two message types separate in automation and routing rules so nobody accidentally bolts a promo onto a service notification.
Data residency. Ask any vendor four questions and write down the answers. Where is conversation data stored at rest? Where is it processed, including by any model provider? Which sub-processors are involved and is there a published list you will be notified about when it changes? Is a standard data processing agreement available with the current transfer mechanism in place? A vendor who cannot answer those quickly in a sales call will not answer them quickly during an audit either. Aftersales publishes its posture on the security and compliance page.
Consent for marketing-adjacent chat
There is a grey zone between service and marketing that ecommerce lives in constantly. "Your item is back in stock" is arguably service if the customer asked to be notified, and marketing if they did not. "Here is 10% off your next order because of the delay" is a goodwill gesture that also happens to be a promotion.
Draw the line explicitly. Maintain a consent flag per channel — email, SMS and WhatsApp each have their own rules and their own opt-in records — and have the agent check it before any message that is not strictly transactional. WhatsApp in particular enforces template and opt-in rules at the platform level, so a sloppy approach gets your number restricted rather than just generating a complaint. Record the source and timestamp of every opt-in. When someone complains, the only useful answer is a record showing when and where they agreed.
The audit trail when AI issues a refund
The moment you let an AI agent move money, you need a record that would satisfy a sceptical finance lead. Six fields, minimum, on every automated refund:
- The conversation ID and full transcript that led to the decision
- The customer identifier and the order ID being refunded
- The rule or policy that authorised it, named explicitly, not "AI decision"
- The amount, currency, and whether it was full or partial
- The evidence checked — delivery status, return receipt, item value against threshold
- The timestamp, and the human who reviewed it if the amount exceeded the auto-approval limit
Store that as structured data alongside the ticket in the shared helpdesk queue, not buried in free text. Then reconcile weekly: pull every automated refund, match it against Stripe, and look for gaps in both directions. Refunds in the helpdesk with no matching processor record mean a failed write. Refunds in Stripe with no helpdesk record mean somebody bypassed the workflow.
Set the auto-approval ceiling low at launch — low enough that the worst plausible week of errors is an amount you would shrug at — and raise it only after you have a month of clean reconciliation. Every refund above the ceiling goes to a human with the AI's recommendation and evidence attached. That is a review step that takes fifteen seconds, not a queue.
Frequently asked questions
Is Aftersales good for ecommerce?
Aftersales is built for high-volume, repetitive support queues, which describes ecommerce precisely. Order status, delivery exceptions, returns, exchanges and refund requests all follow predictable patterns that an AI agent can resolve end to end using live order data. Native Shopify and Stripe connections mean the agent reads real order records rather than guessing from a knowledge base article.
Does Aftersales integrate with Shopify?
Yes. Shopify is a native integration. Connecting a Shopify store gives the AI agent read access to orders, fulfilment status, tracking numbers, line items and customer records, so it can answer order questions with live data instead of generic help text. The connection also supports actions such as looking up return eligibility against the original purchase date.
Does Aftersales work with Stripe?
Yes. Stripe is a native integration. It lets the AI agent see payment status, confirm whether a charge succeeded or failed, check refund state, and issue refunds within limits set by the support team. Combined with order data, a payment-related question can usually be resolved in a single conversation without a human reading two dashboards.
Can the AI answer where-is-my-order questions?
Yes, and this is the highest-volume use case in ecommerce support. The AI agent identifies the order, reads live fulfilment and tracking status, and gives a specific answer with a date and a carrier rather than a link to a tracking page. When tracking shows an exception such as a failed delivery attempt, the agent explains what happened and offers the next step.
Can the AI issue refunds?
Yes, within limits the support team defines. Typical guardrails include a maximum automatic refund value, a list of eligible reasons, and a requirement that delivery or return status is confirmed first. Anything outside those limits goes to a human with the recommendation and supporting evidence attached. Every automated refund is logged with the rule that authorised it.
How does Aftersales handle returns?
The AI agent checks return eligibility against purchase date, item condition rules and the store's return window, then walks the customer through the process and generates the next step. Ineligible requests get a clear explanation rather than silence. Complex cases such as damaged high-value items or partial returns from split shipments can be routed to a human automatically.
Does Aftersales support WhatsApp?
Yes. WhatsApp is a supported channel alongside email, SMS, live chat and Slack. Conversations from every channel land in one shared queue, so a customer who starts on WhatsApp and follows up by email is handled as a single thread rather than two disconnected tickets. Channel choice does not change which automations or policies apply.
Is there live chat for a Shopify store?
Yes. Aftersales provides a website chat widget that can be installed on a Shopify storefront, backed by the same AI agent that handles email and messaging. Because the agent reads live Shopify order data, a logged-in shopper asking about an order gets a specific answer in chat rather than a form to fill in.
How is Aftersales different from Gorgias?
Gorgias is a helpdesk built for ecommerce with automation layered on top of a ticketing core. Aftersales is AI-native: the AI agent is designed to resolve conversations end to end rather than deflect them into macros, and the human helpdesk exists to catch what the agent hands off. Pricing is per seat rather than per automated interaction.
How is Aftersales different from Intercom?
The clearest difference is commercial. Intercom's Fin is priced per resolution at $0.99, so support costs scale directly with conversation volume — painful during peak trading. Aftersales starts at $24 per seat per month with AI resolution included, so a Black Friday spike does not produce a proportional invoice. A fuller comparison is available in the Intercom alternatives guide.
What does Aftersales cost?
Pricing starts at $24 per seat per month. AI resolution is included rather than billed separately per conversation, which matters for ecommerce because volume is seasonal and spiky. Current plans, included features and any volume considerations are listed on the pricing page. There are no per-resolution charges layered on top of the seat price.
Is there a free trial?
Yes, a 14-day free trial is available without a sales call. Two weeks is enough to connect a Shopify store, import an existing help centre, run the AI agent in shadow mode against real conversations, and compare its proposed answers against what the human team actually sent. Starting a trial takes a few minutes from the sign-up page.
How long does migration take?
For a typical ecommerce store with an existing help centre and a Shopify connection, a working setup takes days rather than months. Importing knowledge content and connecting order data is the fast part. The realistic timeline is one to two weeks of parallel running in shadow mode before the AI agent handles live conversations, plus a tuning period afterwards.
Can Aftersales handle Black Friday volume?
Yes, and peak trading is where AI resolution matters most. Because the AI agent handles unlimited concurrent conversations, a five-times volume spike does not create a five-times backlog or require seasonal hiring. Per-seat pricing means the cost of that spike is not billed per conversation, unlike per-resolution pricing models that invoice proportionally to peak demand.
Does Aftersales support multiple languages?
Yes. The AI agent responds in the customer's language, and Copilot translates in both directions so a human agent writing in English can answer a message received in German and have the reply delivered translated. For cross-border ecommerce this removes the need to staff native speakers for every market served, particularly on lower-volume languages.
Does Aftersales work for multi-store or multi-brand setups?
Yes. Multiple Shopify stores can be connected, with separate knowledge sources, policies, tone and routing per brand, while the support team works from one shared queue. That keeps a returns policy that differs between brands from leaking across, while avoiding the overhead of running separate helpdesk instances for each storefront.
Can Aftersales handle subscription orders?
Subscription questions such as next charge date, payment failures and renewal status can be answered using the Stripe connection, which exposes payment and subscription state. Cancellation and pause requests can be routed according to policy — resolved automatically where permitted, or handed to a retention specialist with full context where a save attempt is worthwhile.
What happens when the AI does not know the answer?
The AI agent hands off to a human rather than guessing or looping. The handoff carries the full conversation, the order data already retrieved and a summary of what was attempted, so the human agent does not restart the conversation. Frequent handoffs on the same topic signal a knowledge gap worth filling rather than a model failure.
How do I measure whether Aftersales works?
Track resolution rate, not deflection rate — the share of conversations closed without a human touching them and without the customer returning within a week. Pair it with CSAT split by AI-handled and human-handled conversations, and with first response time. The resolution rate guide explains how to define and measure the metric honestly.
Does Aftersales need a developer to install?
No. Connecting Shopify, Stripe and email channels is an authorisation flow, and the chat widget installs as a snippet on a storefront theme. A developer is only needed for custom work such as connecting an in-house order management system through the API. Most stores get a working configuration without engineering involvement.
Can I keep my existing help centre?
Yes. Existing help centre content can be imported so the AI agent answers from articles already written and approved, rather than requiring a rewrite. The customer-facing help centre can stay where it is. Keeping a single source of truth matters more than where it lives, since drift between published policy and AI answers causes most quality problems.
Is my customer data used to train models?
Customer conversation data is not used to train shared foundation models. Data handling, sub-processors and retention are documented on the security and compliance page, and a data processing agreement is available. Support teams handling EU shoppers should confirm data residency and deletion propagation timelines in writing during evaluation rather than after going live.
Is Aftersales GDPR compliant?
Aftersales supports GDPR obligations and holds SOC 2 Type II certification. That includes deletion of conversation history in response to erasure requests, a data processing agreement, and documented sub-processors. Merchants remain the data controller and are responsible for lawful basis decisions such as consent for marketing-adjacent messaging on WhatsApp, SMS and email channels.
Does Aftersales replace my support team?
No. The AI agent absorbs repetitive, high-volume work such as order status and return eligibility, which frees human agents for judgement calls, angry escalations, high-value customers and retention conversations. Most teams redeploy rather than shrink, moving people from queue-clearing into work that affects revenue. Headcount planning is a business decision, not a product requirement.
What if I only have one support person?
A single-person support function is arguably where AI resolution has the most impact, because there is no capacity buffer at all. The AI agent covers nights, weekends and volume spikes that one person cannot, while the shared inbox keeps everything in one queue. Per-seat pricing from $24 per month keeps the cost proportional to a one-person team.
Glossary of ecommerce support terms
WISMO
Short for "where is my order". The single highest-volume support enquiry in ecommerce, covering any question about whether a parcel has shipped, where it currently is, and when it will arrive. WISMO volume rises with delivery promise aggressiveness and falls with proactive shipping notifications that give customers specific dates.
RMA
Return merchandise authorisation. The reference number issued when a merchant approves a return, used to match the inbound parcel to the original order when it arrives at the warehouse. Without an RMA, returned goods arrive unidentified and sit in a reconciliation pile while the customer waits for a refund.
Chargeback
A forced reversal of a card payment initiated by the cardholder through their issuing bank rather than through the merchant. The bank pulls the funds back, usually with a fee attached, and the merchant must prove the transaction was legitimate to recover them. Chargebacks damage processor standing beyond the immediate loss.
Dispute
The broader term for a customer contesting a charge, of which a chargeback is the card-network form. Payment processors generally use "dispute" to describe the whole lifecycle from the customer's initial claim through evidence submission to final ruling. Disputes have hard response deadlines, and missing one forfeits the funds automatically.
Representment
The merchant's formal response to a chargeback, submitting evidence that the transaction was valid. Typical evidence includes proof of delivery, the customer's order confirmation, communication history and the refund policy shown at checkout. Well-organised support transcripts frequently decide representment outcomes, which is why conversation retention policy matters commercially.
Contact rate
The number of support conversations divided by the number of orders over the same period, usually expressed as a percentage. A store at five percent contact rate receives five conversations per hundred orders. Contact rate is the cleanest measure of whether operational problems upstream are generating avoidable support demand.
Deflection
Preventing a conversation from reaching a human, typically by showing a help article or a self-service flow. Deflection counts avoided contacts, not solved problems, which makes it a flattering and frequently misleading metric. A customer who gives up after reading an unhelpful article counts as deflected and as dissatisfied.
Resolution rate
The share of conversations fully closed without human involvement and without the customer coming back on the same issue. Unlike deflection, resolution rate only counts outcomes the customer accepted. It is the honest measure of AI performance in support, and the one worth tying targets and budgets to.
First contact resolution
The proportion of issues solved in the customer's first interaction, with no follow-up required. High first contact resolution correlates strongly with satisfaction because it eliminates waiting and repetition. Measuring it requires linking follow-up conversations back to the original issue rather than treating each inbound message as new.
CSAT
Customer satisfaction score, collected by asking the customer to rate a specific interaction, usually immediately after it closes. Reported as the percentage of responses that are positive. CSAT measures the interaction, not the product, and should be segmented by AI-handled versus human-handled conversations to be interpretable.
NPS
Net promoter score, derived from asking how likely a customer is to recommend a brand on a zero to ten scale. Promoters minus detractors gives the score. NPS measures brand relationship rather than a single support interaction, so it moves slowly and is a poor tool for week-to-week support tuning.
AOV
Average order value, calculated as total revenue divided by order count over a period. Used in support to decide how much service investment an order justifies — a store with a high average order value can afford richer handling and more generous goodwill than one selling low-value consumables at thin margins.
LTV
Lifetime value, the total margin a customer is expected to generate across their entire relationship with a brand. Support decisions look different when judged against lifetime value rather than the value of a single order, because retaining a repeat buyer usually justifies absorbing a cost on one transaction.
Repeat purchase rate
The share of customers who place more than one order within a defined window. It is the most direct measure of whether a store is building a customer base or renting one. Support quality influences it heavily, since a badly handled delivery problem reliably ends an otherwise healthy relationship.
Backorder
An order accepted for an item that is currently out of stock, to be fulfilled when inventory arrives. Backorders generate sustained support volume because the promised date is uncertain and often moves. Proactive updates at each date change dramatically reduce the resulting enquiries and cancellation requests.
Pre-order
An order placed before a product is available for sale at all, typically ahead of a launch date. Pre-orders create concentrated support spikes around the launch window, particularly if the date slips or allocation runs short. Clear expectations set at purchase time are the main lever on pre-order support volume.
Split shipment
A single order dispatched in multiple parcels, usually because items ship from different locations or become available at different times. Split shipments are a major source of confusion, since customers receive a partial delivery and assume items are missing. Clear per-parcel notifications prevent most of the resulting contacts.
Delivery exception
Any event that interrupts normal carrier progress: a failed delivery attempt, an incorrect address, a customs hold, weather delay, or a parcel damaged in transit. Exceptions generate the most emotionally charged support conversations, and the fastest way to reduce them is to message the customer before the customer messages you.
Porch piracy
Theft of a parcel after the carrier has marked it delivered, typically from a doorstep or lobby. It leaves proof of delivery showing success while the customer genuinely has nothing. Merchants need an explicit policy on how these claims are handled, since the evidence will always look contradictory.
Proof of delivery
Carrier-supplied evidence that a parcel reached its destination, which may include a timestamp, GPS coordinates, a signature, or a photograph of the delivery location. Proof of delivery is the primary evidence in both goods-not-received claims and chargeback representment, so retrieving and storing it promptly has direct financial value.
SKU
Stock keeping unit. The unique identifier for a specific sellable item in inventory, distinct from the general product. Accurate SKU-level reasoning matters in support because return eligibility, stock availability and pricing all attach to the SKU rather than to the product name the customer uses.
Variant
A specific purchasable configuration of a product, such as a particular size and colour combination. Variants are where most ecommerce support confusion originates, because customers describe the product while the order record, inventory and fulfilment systems all operate on the variant identifier underneath.
Fulfilment centre
The warehouse where inventory is stored, picked, packed and dispatched. Support quality depends heavily on visibility into fulfilment centre status, because the gap between order placed and carrier scan is invisible to the customer and is where a large share of anxious enquiries are generated.
3PL
Third-party logistics provider. An external company that stores inventory and handles fulfilment on a merchant's behalf. Working with a 3PL introduces a data boundary: the merchant's own systems may not know a parcel's true status until the provider reports it, which creates lag that support has to explain.
Dropship
A model where the merchant never holds inventory and orders are shipped directly from a supplier to the customer. Dropshipping makes support harder because delivery timelines, packaging quality and stock accuracy sit with a third party the support team cannot see into or influence directly.
Subscription churn
The rate at which recurring customers cancel their subscriptions over a period. Churn splits into voluntary cancellation, where the customer actively chooses to leave, and involuntary churn caused by payment failures. The two require completely different interventions, and conflating them hides where the losses are actually coming from.
Dunning
The process of recovering failed subscription payments through a sequence of retries and customer notifications. Good dunning retries at times when a card is more likely to succeed and messages the customer before the subscription lapses. Poor dunning silently cancels, converting a recoverable payment failure into permanent churn.
Macro
A saved action in a helpdesk that applies a pre-written reply along with tags, status changes or assignment in one click. Macros speed up human agents on repetitive work, but large macro libraries become unmaintainable and drift out of line with current policy faster than anyone notices.
Canned response
A stored block of reply text an agent inserts into a conversation. Simpler than a macro, since it changes the message only. Canned responses create a recognisable robotic tone when overused, and customers who receive two identical replies to different questions lose trust in the entire support channel.
SLA
Service level agreement. A committed target for support responsiveness, usually expressed as a maximum first response time and a maximum resolution time, sometimes varying by channel or customer tier. SLAs are only useful when breaches trigger visible alerts and reassignment rather than appearing in a monthly report.
Queue
The ordered set of conversations waiting for attention. Queue design — how work is prioritised, filtered and assigned — determines whether urgent issues surface or drown. A well-configured queue makes the next conversation an agent should handle obvious without the agent having to search for it.
Concurrency
The number of conversations being handled simultaneously. A skilled human agent manages perhaps three or four live chats at once before quality degrades measurably. An AI agent has no practical concurrency ceiling, which is why peak trading periods are where automated resolution changes the operating model rather than just trimming cost.
Escalation
Moving a conversation to someone with more authority or specialist knowledge, such as a supervisor approving an out-of-policy refund. Escalation paths should be explicit and time-bounded. Undefined escalation is how conversations sit for three days while each person assumes someone else has picked it up.
Handoff
The transfer of a conversation from the AI agent to a human, or between humans. A good handoff carries the full transcript, retrieved order data and a summary of what has already been attempted, so the customer never repeats themselves. A bad handoff resets the conversation and undoes any goodwill earned.
Shadow mode
Running an AI agent against live conversations while it drafts answers that are never sent, so its output can be compared against what human agents actually replied. Shadow mode is the safest way to measure readiness before launch, and the cheapest way to find knowledge gaps without exposing customers to them.
The bottom line on ecommerce customer service software
Here is the honest version.
Aftersales is a strong fit if your support queue is dominated by repetitive, data-driven questions — order status, delivery exceptions, returns, exchanges, refund requests — and if that data lives in Shopify and Stripe. That is the configuration the product handles best, because the AI agent is reading live order records rather than paraphrasing help articles. If you run a Shopify store doing meaningful volume with a support team between one and fifty people, the maths works and the setup is measured in days.
It is also a strong fit if your volume is seasonal. Per-resolution pricing punishes exactly the weeks when you are making your year. Per-seat pricing from $24 per month does not. If you dread the November invoice more than the November queue, that difference is the whole argument.
It is a good fit if you sell internationally and cannot justify native-speaking staff for every market, because the agent answers in the customer's language and Copilot's two-way translation lets your existing team handle the rest.
Now the other side.
Aftersales is not the right choice if your support volume is genuinely low — a handful of conversations a week does not justify any platform, and a shared mailbox will serve you fine until it does not. It is not right if your order data lives in a bespoke system with no API, because the agent's advantage collapses the moment it cannot see the order. It is not right if the bulk of your queue is complex, consultative pre-sales conversation rather than post-purchase service; that work rewards human judgement and product expertise, and automating it produces worse outcomes and worse revenue.
It is not right if you want a deflection bot that suppresses ticket counts while customers quietly give up. The product is built to resolve conversations, which means it will sometimes hand off, and the metric that matters is resolution, not avoidance. Teams looking for a cheaper way to ignore customers should buy something else.
And it is not the right choice if you need a vendor claiming certifications it does not hold. Aftersales is SOC 2 Type II, supports GDPR, and is HIPAA-ready. It is not PCI certified and does not pretend to be, which is precisely why card data belongs in a hosted payment page and never in a chat window.
What to do next depends on where you are. If you already know the shape of your problem, start the 14-day trial, connect your store, and run the agent in shadow mode for a week against real conversations. You will learn more from that than from any demo, because you will be comparing its drafts against what your own team actually sent. Check the current plans and per-seat costs first so you know what a rollout looks like at your headcount.
If your setup is more complicated — multiple brands, a 3PL with patchy data, a migration off an incumbent helpdesk with years of history — establish the messy parts first. Whether your order data is reachable is worth settling before two weeks of trialling something that was never going to fit. Either way, the question to answer is the same: how many of last month's conversations could have been resolved from data you already had?