The draft looked perfect. It thanked the customer for their patience, laid out the return steps in order, and promised the refund would land within five business days. The five business days were the problem: the policy said ten, the customer kept the email, and the argument about which number counted took longer than the refund did.
Nothing had malfunctioned. The model was asked to write a refund email, so it wrote one, and where the company's actual policy should have been it put the most ordinary number it knew. It had no way to find out what anyone had promised. Nobody had told it.
That's the shape of nearly every problem teams run into when they point ChatGPT at a support queue, and it's the part the "8 ways to use ChatGPT for customer service" posts skip. The tool is useful. What it's useful for is narrower and more specific than the pitch, and the reasons are worth knowing, because they tell you which limits a better prompt will fix. Almost none of them.

The short version
- ChatGPT is good behind an agent and unreliable in front of a customer. Drafting, summarizing, translating and rewriting are real wins. Answering your customers unsupervised isn't one yet.
- It doesn't know your business unless somebody tells it, every time. Not your refund window, not this customer's order, not what a colleague promised on Tuesday. Every confident wrong answer traces back to that.
- Its training stops at a published date. The current flagship models list a cutoff of February 16, 2026 (checked September 2026), so anything later it has to be told or has to look up.
- The reliability numbers are worse than "it makes mistakes sometimes." On a published customer-service benchmark, the best agents of their generation finished under half the tasks overall, and in the retail half they got the same task right on all eight attempts less than a quarter of the time.
- Where you paste matters more than what you paste. The consumer chat window, the API and the business plans are three different doors, and most of the reassuring quotes online are about the API.
- If a system talks directly to your customers in the EU, it has to tell them it's an AI. That obligation applied from 2 August 2026, so it's already live.
- If you're trying to decide which category of AI tool to buy rather than what this one does, start with the three places AI can sit in a support workflow instead.
What's actually happening when you paste a support email into it
Using ChatGPT for customer support means putting a general-purpose text predictor between a customer's message and your reply. You give it words, it produces the words that most plausibly follow, shaped by an enormous amount of training text and by instructions about how to behave. It isn't reasoning from your records. It's writing the most likely version of the email you asked for.
That single sentence explains most of what follows. When the shape of the answer is the whole job (a polite apology, a clear explanation of a process, the same message in German), it's excellent, because the shape is exactly what it learned. When the answer depends on a fact only your company holds, it will still produce something confident and well-formed, and the fact will be invented.
What ChatGPT is good at in a support queue today
Four jobs, each with its boundary in the same breath.
Drafting a first reply. Give it the customer's message plus the facts (the policy, the order state, what you're offering) and it will write a competent, appropriately warm version faster than you would. The boundary: it drafts, you send. The moment nobody reads before sending, you've moved it into the category it's bad at.
Catching up on a long thread. Forty-message escalations where you need to know what was promised are where it earns its keep. Ask for the commitments made and the open questions rather than "summarize this," and you get something you can act on. The boundary: it summarizes what's in the text you gave it, so anything agreed on a call is invisible to it.
Translating in both directions. Reading an inbound message in a language nobody on the team speaks, and answering in it, is a job it does well enough to unblock a queue. The boundary: you can't check the tone of the output, so keep it to factual replies and away from apologies and refusals, where register matters.
Rewriting a reply that's correct and reads badly. The engineer's answer that's accurate and sounds hostile, turned into something a customer can read, with the facts kept. The boundary: "make this friendlier" will quietly soften a refusal, so a hard no comes back as a maybe. Read the rewrite for what it now promises, not just for how it sounds.
Sorting and routing is a different job with a different answer, handled by a model that classifies instead of writes. This page is about the drafting window.
What ChatGPT can't do, and why the reason matters more than the limit
| What you want | Can ChatGPT do it? | Why |
|---|---|---|
| Look up this customer's order | No | Nothing connects it to your order system |
| Quote your refund window | No | It never saw your policy |
| Know what a colleague promised | No | Nothing in a chat window tracks a thread or its owner |
| Tell you when it isn't sure | Unreliably | The training rewards a confident answer |
| Know what changed last month | Only if told or looked up | Fixed training cutoff |
| Draft, summarize, translate, rewrite | Yes | The shape of the text is the whole job |
The first three rows are one limitation in three costumes: the model works from what's in front of it. That's less absolute than it used to be, because paid plans can connect it to a mailbox or a file store, and what's on offer varies by plan and by region, so check what applies to yours. It changes less than you'd hope. A connector reads messages and documents. It still can't tell which of four policy drafts is the one in force, who owns this thread, or what got agreed on a call, and it starts the next conversation with none of it.
The training cutoff is published and worth knowing rather than guessing at. OpenAI's model documentation lists the current flagship models with a knowledge cutoff of "Feb 16, 2026" (checked September 2026). Anything after that isn't in the model. It can sometimes retrieve a public page, so your published prices may be reachable that way, but your internal escalation policy and this customer's history are not.
The fourth row is the interesting one, because it's the limit people assume they can prompt their way around. They can't, and there's a good explanation of why. In Why Language Models Hallucinate (Kalai, Nachum, Vempala and Zhang, 2025), the authors argue that these systems "are optimized to be good test-takers, and guessing when uncertain improves test performance," so they "hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty." A model that says "I don't know" scores worse on the benchmarks it's graded against than one that guesses. Confident wrongness is what the scoring encouraged.
OpenAI's own Model Spec carries guidelines telling the model to "express uncertainty" and to "consider uncertainty, state assumptions, and ask clarifying questions when appropriate." That's the right instruction, and it's an instruction about behavior. Telling a system to flag its uncertainty is a different thing from that system reliably detecting its own uncertainty.
Practically: ask it for things where being confidently wrong is visible to you immediately. You'll catch a badly worded apology. An invented refund window looks exactly like a real one.
Five prompts that hold up, and the one that keeps failing
The pattern that works is the same every time: supply the facts, name the constraints, say what the output should look like. These are instructions to a model, not text to send a customer. If what you want is reusable reply text, that's a different tool and we keep a library of saved reply blocks for it.
- Draft with the facts supplied. "Here's a customer email. Our policy: returns accepted within 30 days, refund processed in 10 business days, customer pays return shipping. This order shipped 12 days ago. Draft a reply that approves the return, states the timeline, and doesn't apologize more than once."
- Extract the commitments. "Here's a support thread. List every commitment our side made, with who made it and when. Then list every question the customer asked that nobody answered. Don't summarize anything else."
- Rewrite without changing the facts. "Rewrite this reply so a non-technical customer can follow it. Keep every number and every step exactly as written. Flag anything you think is factually unclear instead of fixing it."
- Find the actual question. "This message is long and upset. In one sentence, what does this person want to happen? Then list what we'd need to check before we can answer."
- Pressure-test a reply before it goes. "Here's a draft reply and the policy it's based on. What could a customer reasonably misread here, and which sentence would you argue about if you were them?"
The anti-pattern is the prompt that circulates most: "Write a reply to this angry customer." No policy, no facts, no constraints. It produces something fluent, generically apologetic, and full of invented specifics, because you left every gap for it to fill and filling gaps is precisely what it does. The first refund email in this article came out of that prompt.
What the benchmarks say about letting ChatGPT work alone
"It makes mistakes sometimes" is the standard limitation paragraph, and it's too soft to make a decision with. There's a measured version.
τ-bench (Yao, Shinn, Razavi and Narasimhan, 2024) tests AI agents on realistic tool-using conversations in retail and airline domains, with the customer side played by a simulated user, which is about as close to a support queue as published benchmarks get. Two findings. The state of the art function-calling agents of that generation "succeed on <50% of the tasks" across both domains. And "pass^8 <25% in retail": run the same retail task eight independent times, and fewer than a quarter of tasks come out right all eight times.
That second number is the one to bring to a meeting where somebody wants a customer-facing bot. It measures whether the agent does the same thing twice. A support process that answers the same question differently depending on the attempt isn't a process, and you can't write a policy around it or audit it afterwards.
Two honest caveats. This benchmarked models from 2024, and the field has moved. And it tested autonomous agents completing tasks, not a person using a chat window as a drafting aid, which is why the recommendation on this page splits the way it does. Behind a human who reads before sending, inconsistency costs you a re-roll. In front of a customer, it costs you the answer.
Where your customer data goes: three doors
This is where the guides get vaguest, usually landing on "be mindful of compliance." There are three separate doors and they don't have the same answer.
The consumer chat window is the one your team is actually using, in a browser tab, often on a personal account. The UK's National Cyber Security Centre is direct about this class of tool: don't "include sensitive information in queries to public LLMs," and don't "submit queries to public LLMs that would lead to issues were they made public." Its assessment points out that a query "will be visible to the organisation providing the LLM (so in the case of ChatGPT, to OpenAI)," that stored queries "may be hacked, leaked, or more likely accidentally made publicly accessible," and that a future owner of the service could inherit them. None of that is exotic. It's the normal risk of putting text on someone else's computer. Whether your conversations feed model improvement, and whether you can switch that off, is answered in the app's own data controls and terms, and that's where to read it rather than trusting any third-party summary, this page included.
The API is the one every reassuring quote comes from. OpenAI's data controls documentation states that "as of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)," that abuse monitoring logs are "retained for up to 30 days, unless longer retention is required by law," and that Zero Data Retention excludes customer content from those logs. Read the scope line: that document covers the API platform. It isn't a statement about the chat window your agent has open, and quoting it as if it were is the most common mistake in this whole conversation.
Business plans are the door to walk through if this is going to be a work habit rather than an experiment, and they're also the door this article can't answer for you. The guarantees sit in the terms for the specific plan, they're the layer that changes most often, and last year's answer may not be this year's. Read them for the plan you're on before you rely on anything, including anything you read here.
Whichever door you're on, the practical rule costs nothing: strip what the model doesn't need. It doesn't need the customer's name, their full address, the order number or the account ID to rewrite a paragraph. Replace them with placeholders and put the real values back yourself. The model gets the shape, which is all it was ever good for.
The rule that already applies to you
If an AI system talks to your customers directly, it has to say so. Under the EU AI Act's Article 50, providers must design such systems so that "the individuals concerned are informed that they are interacting with an AI system, unless this is obvious." The European Commission's guidance on Article 50 confirms it "applies as from 2 August 2026," which means it's in force now, and that the obviousness exception is "interpreted in a restrictive manner." Assume it doesn't cover you.
Two details worth getting right. The obligation in Article 50(1) sits on the provider of the system rather than on the support team deploying it, so if you're switching on a bot somebody else built, the design duty is theirs. In practice you're the one the customer is talking to, and the disclosure either reaches them or it doesn't, so treat it as a question to settle with the vendor in writing rather than one that isn't yours. And note the trigger: this is about systems that interact with people directly. A draft that a human reads, edits and sends is not the same case as a bot answering on its own.
Making it work in a real inbox without the copy-paste tax
Here's the workflow nobody costs out. A message arrives in the shared inbox. Somebody reads it, copies it, switches to a browser tab, decides on the fly which parts are safe to paste, types the policy from memory because the model can't see it, gets a draft, copies it back, fixes what the model got wrong about the order, and sends. Then the next message, from the top.
That round trip is where the time saving goes. It's also where the data risk lives, because "which parts are safe to paste" is a judgment call made forty times a day by a tired person, and it only has to go wrong once. And the model is doing its work with the least context of anyone involved: it doesn't know this customer wrote in twice last week, or that a colleague already replied.
Prompting won't fix that. The drafting has to happen where the context already lives. A shared inbox tool like TriageFlow keeps the thread, the customer's history and the record of who owns it in one place, which is the thing that removes both the clipboard trip and the daily decision about what's safe to paste. If you're at the point of comparing that class of tool properly, what a small team actually needs from a help desk is the better starting page.
If you're not there yet, the cheap version works: keep a document with your actual policies, paste the relevant one into the prompt along with the customer's message, and never let a draft go out unread.
Frequently asked questions
Can ChatGPT be used for customer service?
It works well as a drafting aid and badly as an autonomous responder. Writing first replies, summarizing threads, translating and rewriting all save real time when a person reads the output before it goes. What it isn't ready for is answering your customers without that person, and the reason is structural rather than a question of the current version.
Is it safe to put customer emails into ChatGPT?
That depends which door you go through, and the safe default is to assume the answer is no for the consumer chat window. The UK's NCSC advises against putting sensitive information into public LLMs at all. If you're going to do it routinely, use a business plan, read its current terms yourself, and strip names, addresses, order numbers and account IDs first. The model doesn't need them.
Can ChatGPT replace customer support agents?
The published benchmarks say no, and the interesting part is which number says it. Consistency, not accuracy, is what breaks a support process: on the closest available benchmark, agents got the same retail task right on all eight attempts less than a quarter of the time. A queue that answers the same question differently depending on the day can't be audited or improved.
What are the limitations of ChatGPT for customer service?
They're architectural, which is why prompting doesn't move them. Nothing connects the chat window to your records by default. The training has a fixed cutoff, published as February 16, 2026 on the current flagship models. And the training objective quietly prefers a confident answer to an honest "I'm not sure."
Do we have to tell customers when a reply was written by AI?
In the EU, a system that interacts with them directly has to disclose that it's an AI, and that has applied since 2 August 2026. The exception for cases where it's obvious is read narrowly. The design obligation formally sits with whoever provides the system rather than with you as the deployer, so if you're buying a bot, get the disclosure confirmed in writing. A human editing and sending a drafted reply is a different situation.
Is ChatGPT free to use for business?
Cost is the wrong axis to decide this on. The question is what each plan commits to doing with what your team types into it, and that's set out in the terms rather than in the price. Read those for the tier you're on, then decide whether it belongs anywhere near customer data. For rewriting text you'd be comfortable seeing in public, any of them is fine.