How to Build a Secure AI-Powered Customer Support System
A support chatbot speaks for your company and touches customer data. Here is the architecture, the threats from real incidents, and the guardrails we use to build a secure AI-powered customer support system.
On this page
To build a secure AI-powered customer support system, treat the AI as a helpful but untrusted assistant inside a controlled system. Ground it in approved content, give it least-privilege access to customer data enforced in code, screen what goes in and what comes out, keep a fast route to a human, and log everything. The model is the easy part. The controls around it decide whether it is safe.
This guide is for founders and support leads who want AI to answer customers without leaking data or inventing promises. We reviewed three well-documented incidents, the OWASP risk list for LLM applications and the EU's chatbot disclosure rule, then combined them with how we design these systems at Bracket Coder. Where a fact comes from a source, we name it. Where it is our own judgement or estimate, we say so.
- A support bot speaks for your company: a tribunal held Air Canada responsible for its chatbot's wrong answer.
- Customer-supplied text (tickets, emails, form fields) can carry hidden instructions, as the 2025 ForcedLeak flaw in Salesforce Agentforce showed.
- Enforce identity and permissions in code, at retrieval and at every tool call, never inside the prompt.
- Let the AI answer from approved content, and require human approval for money, policy and account changes.
- Test after every change, keep a kill switch, and tell customers they are talking to AI.
Why does AI customer support need its own security design?
A support assistant sits in an unusual position. It reads text written by strangers, it can see private customer records, and it replies in your voice. That combination has already gone wrong publicly. Three cases show three different failure types.
| Incident | What went wrong | Control that would have helped |
|---|---|---|
| Air Canada chatbot (ruling of 19 February 2024) | The bot gave wrong information about bereavement fares; the tribunal held the airline responsible | Ground answers in current policy, cite the source, human check for money and policy |
| DPD chatbot (18 January 2024) | After a system update, the bot swore at a customer and criticised the company; DPD disabled the AI element | Output guardrails, regression tests after every update, a kill switch |
| Salesforce Agentforce "ForcedLeak" (reported September 2025) | Hidden instructions in a web form field made the agent leak CRM data | Treat customer text as untrusted, allow-list outbound URLs, least privilege |
You own what the bot says
In Moffatt v. Air Canada (2024 BCCRT 149), the airline argued its chatbot was a separate entity. The tribunal rejected that. As McCarthy Tétrault reports, it held that it makes no difference whether information comes from a static page or a chatbot. Every support answer is a statement your business made.
Bots can be talked out of their rules
According to TIME, a customer frustrated that the DPD bot could not connect him to a person began experimenting, and got it to use profanity and call DPD the worst delivery firm in the world. DPD said an error occurred after a system update on 18 January 2024 and that the AI element was immediately disabled. The lesson is less about swearing and more about change: any update can alter behaviour, so every update needs testing.
Your customers' text is an attack surface
Security firm Noma reported ForcedLeak, rated CVSS 9.4 (critical). The Register describes how attackers could hide instructions in the roughly 42,000-character description field of a Web-to-Lead form. When an employee later asked the agent to process the lead, it followed the hidden instructions and sent CRM data out. The content security policy still trusted an expired domain, which researchers bought for USD 5 to receive the stolen data. Salesforce enforced trusted-URL allow-lists by 8 September 2025 and declined to confirm whether the flaw was exploited before the fix.
What does a secure AI support architecture look like?
Think of the assistant as one component in a pipeline with checkpoints, not as a chat box wired to a model. Each checkpoint is code you control.

The request path, step by step
- Identify the customer. Take identity from your authenticated session, never from what the customer types. "I am the account owner" is a claim, not proof.
- Screen the input. Apply rate limits, size limits and basic abuse checks. Strip or flag content that looks like instructions aimed at the model.
- Retrieve with permissions. Search your knowledge base and the customer's own records, filtered by what this customer may see. The filter runs in your code.
- Call the model with minimum context. Send only what the answer needs, with personal data masked where possible.
- Check the output. Validate the reply before showing it: no links or images to unknown domains, no leaked internal text, no forbidden topics.
- Gate any action. Read-only lookups may run automatically. Changes need verification and often human approval.
- Log the exchange. Record the question, sources used, action taken and who approved it, in a store with restricted access.

Where your data actually goes
Most builds send text to an external model provider. Before you choose one, get written answers on four things: whether your data is used for training, how long it is retained, where it is processed, and what audit options exist. These belong in the vendor's data-processing terms, not in a sales call. We cover the wider web-app side in how to protect your business web app from data breaches.
What are the main security threats to an AI support system?
OWASP publishes a Top 10 for LLM Applications (2025 edition). The table maps the entries that matter most for support. The middle column is our plain-English reading, not OWASP's wording.
| OWASP LLM risk | How it shows up in support | Defence |
|---|---|---|
| LLM01 Prompt Injection | A ticket, email, attachment or form field tells the bot to ignore its rules or send data out | Treat all customer text as untrusted; no auto-actions from it; least privilege |
| LLM02 Sensitive Information Disclosure | The bot reveals another customer's details or internal notes | Per-customer retrieval filters; mask personal data; minimal context |
| LLM05 Improper Output Handling | Bot output is rendered or executed without checks, for example as a link or image that leaks data | Sanitise output; allow-list outbound domains |
| LLM06 Excessive Agency | The bot can issue refunds or edit accounts on its own | Narrow tools, verification, human approval for irreversible steps |
| LLM07 System Prompt Leakage | Customers extract your hidden instructions | Keep secrets and keys out of prompts |
| LLM08 Vector and Embedding Weaknesses | Search over your knowledge base returns another tenant's or an internal document | Permission filters inside the search query, separate indexes for private data |
| LLM09 Misinformation | Confident wrong answers about policy or pricing | Grounding, citations, an "I don't know" path |
| LLM10 Unbounded Consumption | Abuse or loops run up your model bill or slow the service | Rate limits, per-user caps, spend alerts |
Indirect prompt injection is the support desk's special problem
Direct prompt injection is a customer typing "ignore your rules". Indirect injection is harder: the instructions hide in content the bot is asked to read, such as a pasted email chain, a PDF attachment or a lead form. ForcedLeak was exactly this. A model cannot reliably tell your instructions apart from text a customer supplies, so the defence lives in the system around it.
- Separate roles. Keep the assistant that reads untrusted text away from tools that can change things or send data.
- Restrict outputs. Only allow links and images from domains you own. ForcedLeak leaked data through an image request to a trusted-but-expired domain.
- Audit your allow-lists. Expired or unused domains in a security policy are a hazard. Review them regularly.
- Never act on untrusted text alone. A human confirms any action triggered by a customer-supplied message.
An instruction such as "never reveal customer data" in the prompt is a request, not a control. If a customer's record is in the model's context, assume it can be extracted. Keep it out of context unless that customer is authenticated.
How do you protect customer data and privacy?
Collect and send less
Mask names, emails, phone numbers and account numbers before text reaches the model, and substitute real values only when a verified action needs them. Do not let the chat collect card numbers. Send customers to your payment provider's hosted page instead.
Enforce access at retrieval and at every tool
The most damaging bug in support AI is a cross-customer leak. Filter every search and every lookup by the authenticated customer's ID inside your backend query. Test it with the two-account check: sign in as customer A, ask for something that belongs to customer B, and confirm you get a refusal. Automate that test so it runs on every release.
Treat transcripts as sensitive data
Chat logs contain personal information. Restrict who can read them, set a retention period, and delete on request where the law requires it. If you serve people in the EU or the UK, data protection law applies to these transcripts, so take advice on lawful basis and retention.
How do you stop the AI giving wrong answers?
Wrong answers are a security issue when they promise refunds, invent policies or misstate terms. Reduce them at the source.
- One source of truth. Keep policies in one reviewed place. If two pages disagree, the assistant will quote either.
- Answer from retrieved content only, and show the source page so customers and agents can check it.
- Define the scope. List topics the bot handles and topics it hands off. Anything else gets a polite refusal.
- Use a handoff trigger when retrieval finds nothing relevant, rather than letting the model improvise.
A support bot should be allowed to say "I don't know". It should never be allowed to guess about money, policy or someone else's account.
When should a human take over?
Design the handoff before launch. Trigger it when the customer asks for a person, after repeated failed answers, on high-stakes topics, and on signs of distress or anger. Pass the human a summary and the source pages used, so the customer does not repeat themselves.

A simple autonomy matrix
| Action | Autonomy level | Control |
|---|---|---|
| Answer from published policy or docs | AI alone | Grounded in sources, with citation |
| Look up order or ticket status for a signed-in customer | AI alone, read-only | Identity from session, per-customer filter |
| Change a delivery address or contact detail | AI proposes, verification required | Re-authentication or a confirmed code |
| Issue a refund, credit or discount | Human approval | Approval recorded in code before the action runs |
| Cancel an account or handle a complaint | Human approval | Escalate with a summary |
| Share another person's data, or give legal or medical advice | Never | Hard block |
We go deeper on this split in AI agent development: what to automate vs keep human, and on what happens when approval lives only in a prompt in the Meta Muse address incident.
What do the rules require?
The European Commission says Article 50 of the AI Act applies from 2 August 2026. For chatbots, people must be informed they are interacting with an AI system unless that is obvious, clearly and from the start of the first interaction. The Commission lists maximum fines of 15 million euros or 3% of worldwide annual turnover, with proportionality for small businesses. Whether you count as provider or deployer is a question for a lawyer, and other regions have their own rules. This is not legal advice.
Open every chat with one plain line: "You are chatting with an AI assistant. A person can take over anytime." It meets the spirit of the rule and sets honest expectations.
How do you build it, and how long does it take?
A support assistant that only answers from published help content is a smaller job than one that looks up accounts and triggers actions. We would plan it in stages, so each stage adds capability only after the previous one is safe. The durations are our planning estimates for a typical mid-sized support operation, not a quote.
| Stage | Focus | Deliverables |
|---|---|---|
| Weeks 1 to 2 | Scope and content | Topic list, cleaned policies, handoff rules, vendor data terms reviewed |
| Weeks 2 to 4 | Core pipeline | Identity, retrieval with permissions, grounded answers, output checks, logging |
| Weeks 4 to 6 | Actions and handoff | Read-only lookups, approval gates, human handoff with summaries |
| Weeks 6 to 8 | Testing and pilot | Real-question test set, hostile tests, pilot on a slice of traffic |
| Ongoing | Operate | Transcript review, regression tests, allow-list and permission audits |
- 2 Aug 2026EU AI Act Article 50 disclosure duties apply (European Commission)
- CVSS 9.4severity rating reported for ForcedLeak (The Register)
- USD 5cost of the lapsed domain used in the ForcedLeak proof of concept (The Register)
- 6 to 8 weekstypical staged build to pilot (our estimate)
What drives the cost
- Integration depth: answers only, read-only lookups, or actions.
- Content readiness: tidy policies are fast; scattered documents are slow.
- Testing effort: the security and quality test set is real work and worth budgeting.
- Model usage and hosting, which scale with traffic. Check your provider's current pricing rather than a blog figure.
- Operations: monitoring and review, which our application maintenance and support covers.
For a scoped figure, see our pricing page or ask for a quote.
How do you test and monitor it after launch?
Test before every release
Keep a fixed set of real customer questions with expected outcomes, plus hostile ones: attempts to override rules, requests for other customers' data, and injected instructions inside pasted text. Run the set after every prompt, model, content or vendor change. DPD's problem followed a system update, which is exactly when a regression suite earns its keep.
Watch these signals
- Handoff rate and reasons, which show content gaps.
- Weekly sample reviews, marking answers right, wrong or unsafe.
- Refusals and blocked outputs, which show attack attempts.
- Cost per conversation, with alerts for spikes.
Have a kill switch and a plan
Make it possible to switch the assistant off, or fall back to a static contact form, in minutes. Decide who can press that button. If something does go wrong, follow the same containment and evidence steps we describe for a suspected breach in our breach protection guide.
Should you build custom or use a helpdesk vendor's AI?
Many helpdesk and CRM platforms now ship AI agents. That can be the fastest start, but the security questions do not disappear. You still control what data it can see, what actions it can take and how it is tested. Whichever route you take, ask the vendor:
- Is our data used to train models, and can we opt out in writing?
- Where is it processed, and how long is it retained?
- Can we restrict which records and actions the agent can reach?
- Can we allow-list outbound domains, and how is output sanitised?
- Do we get audit logs, and is there a kill switch?
Build custom when you need tighter data control, deep integration with your own product, or behaviour the vendor will not let you configure. We do this as part of SaaS development and web application development. For a lighter first step, see how to add AI to an existing business website, and browse our AI Applications hub.
Frequently asked questions
Is it safe to let an AI chatbot access customer data?
Only with strict limits. Take identity from your authenticated session, filter every search and lookup by that customer, send the minimum data to the model, and keep logs protected. Never rely on the prompt alone to keep records private.
What is prompt injection in customer support?
It is when text the bot reads, such as a message, email, attachment or form field, contains instructions that the model follows. Indirect injection, where the instructions hide in content the bot is asked to process, was behind the ForcedLeak flaw reported in Salesforce Agentforce in 2025.
Can a company be held responsible for what its chatbot says?
Yes. In Moffatt v. Air Canada (2024), a Canadian tribunal held the airline responsible for wrong information its chatbot gave, and rejected the argument that the bot was a separate entity. Treat every answer as your own statement.
Do I have to tell customers they are talking to an AI?
Under the EU AI Act, Article 50 requires that people are informed they are interacting with an AI system, unless it is obvious, from the first interaction. It applies from 2 August 2026. Other regions have their own rules, so check with a lawyer and disclose clearly either way.
Should the AI be allowed to issue refunds?
Not on its own. Let it gather details and propose the action, but require human approval that is enforced in code before money moves. Start with read-only tools and add actions one at a time.
How do I know my support AI is secure enough to launch?
You should be able to show that per-customer access is tested, hostile inputs are blocked, outputs are checked, actions are gated, logs exist, a kill switch works and customers are told it is AI. If any of those is missing, keep the pilot small.
Planning an AI support system? Get a free scope and quote and we will map what the assistant can safely do on day one and what should stay with your team. You can also see all our services.
Sources
- OWASP GenAI Security Project: Top 10 for LLM Applications 2025
- The Register: Salesforce Agentforce tricked into leaking sales leads
- McCarthy Tétrault: Moffatt v. Air Canada, a misrepresentation by an AI chatbot
- TIME: AI chatbot curses at customer and criticizes company (DPD)
- European Commission: Transparency obligations under Article 50 of the AI Act
Have a project like this in mind?
Tell us what you're building and we'll map out the scope, timeline and a fixed starting quote — no obligation.
Start your project


