I shipped a customer support chatbot for a client last year that passed every demo we ran. Two weeks after launch, a user pasted text that looked like a shipping question but ended with instructions to ignore previous rules and quote internal pricing. The bot complied. Nothing in our carefully written system prompt stopped it, because system prompts are suggestions written in the same language the attacker speaks. That incident rewired how I build LLM features. Prompts set intent. Guardrails enforce it.
This post covers the guardrail stack I now install on every production LLM feature: input checks, output checks, deterministic rules, the tools worth using, and what each layer costs in latency and money. If you are still picking models and wiring your first feature, start with our guide to building LLM applications and come back before launch day.
Why System Prompts Are Not a Security Boundary
A system prompt is text. The model reads it alongside user text and decides what to do. A determined user can argue with it, contradict it, translate around it, or bury an instruction inside a document your pipeline retrieves. OWASP ranks prompt injection as the top risk for LLM applications for exactly this reason. Our prompt injection protection guide covers the attack patterns, so I will not repeat them here.
The practical takeaway: treat the model as an untrusted component that usually behaves. Everything you truly care about, like never leaking PII or never emitting invalid JSON, must be enforced by code. Code cannot be talked out of its rules.

Input Checks: Stop Bad Requests Before They Cost You
Input guardrails run before or alongside the model call. Four categories cover most of what matters:
- Jailbreak and injection screening. A classifier scores the incoming message for attack patterns. This can be a dedicated safety model such as Llama Guard, a smaller fine-tuned classifier, or a service like the Azure AI Content Safety prompt shield. Expect 50 to 200 milliseconds and near-zero cost if you self-host a small model. It will not catch everything, and that is fine. Its job is to raise the attacker’s effort, not to be perfect.
- PII detection. Users paste phone numbers, emails, and account IDs into chat boxes constantly. If your logs or your model provider retain prompts, that is a compliance problem. Run a recognizer such as Microsoft Presidio, and either redact before the model call or reject with a helpful message. Redaction adds under 20 milliseconds for typical messages and costs nothing beyond compute.
- Topic scoping. If your bot exists to answer billing questions, it should not write poetry or debug Python. A cheap zero-shot classifier or an embedding similarity check against your allowed topics handles this. Reject politely and redirect. This also cuts your token bill, since off-topic conversations tend to run longest.
- Rate and length limits. Cap message length, requests per session, and daily spend per user. Injection payloads are often long. A 4,000-character ceiling blocks a surprising share of attacks and runaway costs at once.
Output Checks: Trust Nothing the Model Says
Output guardrails run after generation, before the user sees anything. This is where most teams underinvest, because the demo always looks fine.

- Schema validation. If you asked for JSON, parse it. If parsing fails, retry once with the error fed back, then fail closed. Pydantic makes this trivial in Python. This check is nearly free and catches the most common production breakage: a model that returns prose where your code expected structure. For structured pipelines, pair this with the quality gates in our post on evaluating LLM outputs with automated checks.
- Groundedness checks. For retrieval-augmented systems, verify that the answer is supported by the retrieved sources. The cheap version checks citation overlap or string similarity against source chunks. The expensive version asks a second model to grade entailment, adding 300 to 800 milliseconds and one extra API call. Use the cheap version on every response and the expensive one on a sampled 5 to 10 percent, plus anything flagged low-confidence.
- Refusal and topic drift detection. Check that the answer addresses the allowed topic and did not slip into an unwarranted refusal, or worse, a compliance where refusal was required. Simple classifiers or keyword heuristics catch the obvious cases.
- Toxicity and brand screening. Even a well-behaved model can produce something your client cannot show a customer. A moderation endpoint or a local classifier screens output for toxicity, profanity, and competitor mentions. Budget 50 to 150 milliseconds. Flagged output can route to a canned apology instead of a retry, which is cheaper and safer.
Deterministic Rules: The Underrated Layer
Regex and blocklists are unfashionable in LLM circles and I do not care. They are fast, free, and unambiguous. If a response must never contain a competitor’s name, a profanity, an internal hostname, or a credit card pattern, a regex enforces that with zero false-confidence. Every non-LLM check you add is one less thing you ask a probabilistic system to promise.
A minimal output gate in Python looks like this:
import re
BLOCKED = [
re.compile(r"\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b"), # card numbers
re.compile(r"\b(internal\.acme\.corp|10\.\d+\.\d+\.\d+)\b"), # infra
re.compile(r"\b(competitorone|competitortwo)\b", re.I),
]
def output_gate(text: str) -> str:
for pattern in BLOCKED:
if pattern.search(text):
return "I can't share that. Let me connect you with our team."
return text
Fifteen lines, runs in microseconds, and it has never been talked out of its job by a clever prompt. Layer this under the model-based checks, not instead of them.

The Tooling Landscape
You can hand-roll all of this, and for simple features you should. When the check count grows, frameworks earn their keep.
| Tool | Backing | Strengths | Watch out for |
|---|---|---|---|
| Guardrails AI | Open source (Python) | Huge validator hub, RAIL specs, structured output repair | Validator quality varies; test each one |
| NeMo Guardrails | NVIDIA, open source | Dialog flows (Colang), topical rails, integrates LangChain | Heavier learning curve, opinionated architecture |
| Llama Guard 3 | Meta, open weights | Strong safety classifier for input and output screening | GPU or inference cost to self-host |
Guardrails AI gives you a library of validators you can chain on any output: PII detection, profanity, schema checks, regex, citation checks. Its validator hub is the fastest way to prototype. The Guardrails AI documentation covers the hub and the RAIL spec format.
NeMo Guardrails from NVIDIA takes a different angle. You define conversational rails in a language called Colang: allowed topics, canned responses for sensitive paths, and execution flows that route around the model entirely. It is the best fit when your bot has a defined job with hard boundaries. The NeMo Guardrails repository has working examples.
Llama Guard is not a framework; it is a safety classifier you point at inputs, outputs, or both. Self-hosting the smaller variants keeps per-call cost near zero, and one mid-range GPU instance covers a small business chatbot doing a few thousand requests a day.
My default stack for a client project: deterministic rules I write myself, Guardrails AI validators for PII and schema, and Llama Guard or a hosted moderation endpoint for screening. I reach for NeMo when the client needs strict dialog flows.
The Latency and Cost Budget
Every layer adds milliseconds and, sometimes, dollars. Here is the honest math for one user message with a cloud LLM for the main call:
| Layer | Added latency | Added cost per 1K requests |
|---|---|---|
| Length and rate limits | < 1 ms | $0 |
| Regex and blocklists | < 1 ms | $0 |
| PII detection (Presidio, self-hosted) | 10-30 ms | ~$0 |
| Injection classifier (Llama Guard, self-hosted) | 50-200 ms | GPU share, roughly $0.10-$0.50 |
| Output schema validation | < 5 ms | $0 |
| Moderation endpoint (in + out) | 50-150 ms | often free tiers; cents at scale |
| Groundedness judge (second model) | 300-800 ms | $1-$5 depending on model |
The full stack adds roughly half a second and a few dollars per thousand requests if you run everything on every call. That works for chat and is too slow for autocomplete-style features. My rule: deterministic and classifier checks on everything, and the expensive judge model only on sampled traffic and flagged responses. That lands around 150 to 250 milliseconds of overhead and keeps marginal cost under a dollar per thousand requests.

Compare that to the cost of one incident. A leaked system prompt is embarrassing. A bot that quotes a 90 percent discount or leaks a customer’s email address is a refund, a support storm, and possibly a legal letter. The latency is the cheap part.
When Guardrails Are Overkill
Be honest about your risk surface. An internal tool that summarizes meeting notes for five trusted employees does not need an injection classifier; it needs logging and access control. A marketing copy generator that never touches customer data and is human-reviewed before publishing needs a profanity check and little else.
Guardrails also fight you when over-tightened. Aggressive toxicity filters mangle legitimate medical, legal, and security discussions. Strict topic scoping frustrates users whose questions sit near the boundary. Every layer adds false positives, and false positives are user-facing bugs. Start with the deterministic and schema layers, add classifiers where the threat is real, and measure your false-block rate as rigorously as your block rate.
The Defense-in-Depth Checklist
Before any LLM feature ships, I run down this list:
- System prompt states scope and refusal policy, treated as intent only.
- Input length and rate limits enforced at the application layer.
- Injection and jailbreak classifier on every user message.
- PII detection with redaction before the model call and before logging.
- Topic scoping with a polite redirect for out-of-scope requests.
- Output schema validation with one retry, then fail closed.
- Deterministic regex and blocklist gate on every response.
- Toxicity and brand screening on output, with a canned fallback.
- Groundedness check on retrieved answers, sampled plus flagged.
- Full request and response logging with redaction, so incidents are auditable.
- A red-team session before launch, and one every quarter after.
- A kill switch that disables the feature without a deploy.
If you run agents or tool-calling pipelines, add one more: every tool the model can invoke has its own allowlist and parameter validation, because a guardrail on text means nothing if the model can call send_email directly. Our guide to production LLM applications covers that side of the fence.
Conclusion
Prompts are policy documents. Guardrails are enforcement. The stack that works in production is boring on purpose: cheap deterministic checks everywhere, one or two classifiers on inputs and outputs, schema validation on anything structured, and an expensive judge model only where the math justifies it. Budget about 250 milliseconds and pennies per request for the whole thing.
Build the layers before you need them. Retrofitting after an incident costs more in trust than the stack ever will in latency, and the teams that ship safely in 2026 are the ones that assumed the model would eventually misbehave and planned for it.



