Docs
Guardrails
How each check in the gateway fails and what it may cost, and how accurate each guardrail is, measured on public datasets.
How each check fails
Every request passes these checks in this order, in the gateway, with no model call. When a check itself fails, which is different from finding something, it fails open, and the request carries on with the failure on its record, or closed, and the request is refused, or its stream cut. Each check’s mode is fixed, the same for every policy.
The budget is the most CPU a check may spend on a request; a benchmark fails the build when one is over it. The prompt-injection check runs alongside the provider call, so its time adds no wait in front of it.
| Check | What it does | Fails | CPU budget |
|---|---|---|---|
| Suspension | Refuses an organization or user you have suspended. | closed | 20 µs |
| Limits | Holds rate, token and spend limits, and reserves the request’s estimate. | open; closed for spend | 20 µs |
| Uninspected | Refuses an attachment nothing can inspect, where the policy says to. | closed | 20 µs |
| Secrets | Redacts keys and credentials before anything leaves. | closed | 50 µs + 1 ms per 100 KB |
| Pseudonymize | Swaps personal data for pseudonyms, and restores it in the response. | closed | 50 µs + 1 ms per 100 KB |
| Injection | Checks for prompt injection alongside the provider call, and cuts it if it fires. | open | 50 µs + 1 ms per 100 KB |
| Illegal content | Judges requests for help with a crime, tool results included, and blocks them before the provider sees them. | open | 50 µs + 1 ms per 100 KB |
| Retention | Decides what the record keeps, and tells the provider not to store it. | closed | 20 µs |
| Safety ID | Gives the provider a keyed hash of your user where its API takes one, so it acts on one user, not you. | open | 20 µs |
| Route | Picks an allowed model, region and connection, and falls back if one fails. | closed | 20 µs |
Each guardrail's detector, measured on public, permissively licensed datasets that none of them was built from, and how to reproduce the figures. Every detector runs in the gateway, without a model call, within a CPU budget of 1 ms per 100 KB of text. They were measured on 2026-09-30; names, street addresses and dates of birth, and prompt injection after tool results were added to its training set, on 2026-10-02.
Every detector reads text folded first, so full-width and look-alike characters don't hide what
it looks for. When a policy selects output, the same secrets and personal-data detectors read
the answer as well as the request.
Personal data and secrets
Measured on the test split of nvidia/Nemotron-PII
(CC-BY-4.0, commit b70ffaf5, its Parquet export d36e66f6). The split holds 100,000
synthetic English business documents, half with US and half with international details. The
labelled spans cover 55 kinds of personal and health information.
A detection is right when it overlaps a labelled span of a label its kind covers. A labelled span is found when a detection of that kind overlaps it.
| Kind | Labels it covers | Precision | Recall | False positives per 1,000 documents |
|---|---|---|---|---|
email |
email |
99.8% | 99.5% | 1.2 |
phone |
phone_number, fax_number |
89.6% | 81.6% | 28.9 |
card_number |
credit_debit_card |
74.2% | 10.4%; 87.2% of Luhn-valid numbers | 4.6 |
ip_address |
ipv4, ipv6 |
98.4% | 99.4% | 2.0 |
national_id |
ssn, national_id |
86.5% | 72.5%; 99.6% of US SSNs, 14.8% of other national ids | 10.1 |
iban |
no label in this dataset | — | — | 0.1 |
| secrets | api_key, password, http_cookie |
98.8% | 7.5%; 91.5% of keys shaped as a JWT | 0.1 |
What the figures say:
- Emails and IP addresses are found nearly always, and nearly always rightly. An IP address inside a labelled URL counts as right: it is one.
- Phone numbers are missed where they're written without separators (
6032308421), in international groupings the pattern doesn't take (+61 3 9827 4923,01202 458729), or right against other digits, which the detector reads as part of a longer number. Most false positives are account, member and record numbers written like phone numbers, which a pattern can't tell apart without the words around them. - Card numbers must pass the Luhn check, as real ones do. 88% of this dataset's synthetic card numbers fail it, which is why recall over all of them is low. The detector finds 87.2% of those that pass. Most of its false positives are device identifiers such as IMEIs, which use the same check, and then account numbers.
- National ids: the detector knows US Social Security numbers and UK National Insurance numbers only. It finds nearly every SSN, and few of the dataset's other countries' ids.
- Secrets: the detector knows the formats providers issue (AWS, GitHub, GitLab, Slack,
Stripe, OpenAI, Anthropic, Google, Hugging Face, SendGrid, npm, Azure storage), private keys,
JWTs and passwords in connection strings. The dataset's keys are mostly invented formats
(
read_dev_…, UUIDs) and its passwords are bare words, which no format-based detector finds. What the figures show is that it rarely fires on ordinary business text: 0.1 false positives per 1,000 documents.
Names, street addresses and dates of birth
These three kinds came after the run above, and their Nemotron-PII figures are still to come.
Until then they are measured on presidio-research's
synthetic test set (MIT, data/synth_dataset_v2.json, sha256 ec08a771…0012): 1,500 English
sentences, half with US and half with international details, labelled from templates.
| Kind | Labels it covers | Precision | Recall | False positives per 1,000 sentences |
|---|---|---|---|---|
person_name |
PERSON |
93.0% | 43.2% | 18.7 |
street_address |
STREET_ADDRESS |
100% | 27.8% | 0 |
date_of_birth |
a DATE_TIME after "born" or "birth" |
100% | 100% of 17 | 0 |
And, as an upper bound on false positives in ordinary prompts, on the 4,332 English prompts the injection and illegal-content test splits below hold as benign. Every detection there counts as wrong, though most name real people.
| Kind | Detections per 1,000 prompts | Of a sample, real values |
|---|---|---|
person_name |
176 | about 9 in 10 of 50: people in articles, tables and tool results |
street_address |
6.0 | most |
date_of_birth |
3.0 | all |
What the figures say:
- Names are found without a model, from the US Census's lists of first names and surnames (public domain): a listed first name with a capitalized word after it, a first name and a listed surname, or capitalized words after words that introduce a name ("my name is", "Dear", "Mr.", "Patient:"). A name that is also a word or a place, such as "Will", "Brown" or "Virginia", counts only beside a surname that isn't. So it finds English-speaking names in running text, and misses a surname alone, a name of another tradition the lists lack, and most names in lowercase. A model would find more, at more than the CPU budget allows.
- Street addresses are US, UK and Commonwealth addresses with a number and a street type ("12 Oak Street", "PO Box 1234"), with a unit, town and postcode when they follow, and European ones where the number follows the street ("Rue de Rivoli 99", "Hauptstraße 5"). The dataset's international addresses are mostly in forms neither takes, or split into parts, which is why recall is low. Its detections in ordinary prompts are mostly real addresses.
- Dates of birth are dates that words such as "born", "date of birth" or "DOB" introduce, in a month and day that exist. A date alone isn't one: "the meeting is on 14 March 2026" isn't found.
Prompt injection
Measured on 5,777 labelled texts: 5,313 from nine public datasets (deepset/prompt-injections, Lakera/gandalf_ignore_instructions, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, 3nesdeniz/agentic-prompt-injection-5k and its boundary pairs, nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1, microsoft/BIPIA and databricks/databricks-dolly-15k), and, written for the purpose, 253 short benign instructions about an answer's form and 211 benign tool results. 43% are attacks, and 38% of those are indirect: hidden in a document, an email or a tool's result. The detector was fitted to a 70% split, and its threshold set at 90% precision on out-of-fold scores.
| Measured on | Precision | Recall |
|---|---|---|
| The test split (1,737 texts) | 96.6% | 86.8% |
| Each source of attacks held out of training in turn | 50–100% | 23–80% |
On the test split it flags 2.5% of ordinary prompts, and 1.7% of hard negatives, which talk about attacks without making one or are what tools return. The tool results are pages such as example.com's, error and default server pages, and API, file and shell output; before they were added to the training set, the detector flagged 24 of the test split's 52, and example.com's own text. Adding them cost the held-out figures a few points. Text without two words in a row, such as a tool call's arguments that are only a URL, isn't scored at all. It flags none of the test split's 85 short instructions about an answer's form ("Reply with one word: hello"). It still flags "Just say yes if you understand.", the closing line of many role-play jailbreaks.
The held-out figures are the honest ones for attacks unlike the training set: a lexical detector
knows the phrasings it has seen. So a team does well to start the guardrail in monitor mode,
or with flag, and move to cut once its own traffic shows the rate of false positives.
Other languages
The detector was fitted to English, and it reads words, so it finds few attacks written in another language. Measured on 1,022 texts in ten languages, from rikka-snow/prompt-injection-multilingual (MIT) and deepset's German rows, none of them trained on:
| Language | Attacks found | Ordinary prompts flagged |
|---|---|---|
| All ten | 4.8% (28 of 579) | 0.5% (2 of 443) |
| German | 9.8% (17 of 174) | 0.7% (1 of 140) |
| Italian | 8.1% (3 of 37) | 0 of 20 |
| Vietnamese, Portuguese, Chinese, Thai, Japanese, Hindi, French, Spanish | 0–4.2% each | 0–2.3% |
So outside English the guardrail is close to absent, though it rarely flags ordinary prompts. A team whose users write in other languages should not count on it there. A multilingual model would cost more than the CPU budget allows.
Illegal content
Measured on 41,000 public texts from six permissively licensed datasets: Aegis 2.0, Salad-Data, AILuminate, CatQA, XSTest and Dolly. The detector was fitted to a 70% split, and its threshold set at 90% precision for any kind on out-of-fold scores. The kinds are MLCommons' hazard categories that are crimes everywhere. The guardrail judges the request, tool results included, before the provider sees it; it doesn't read the answer.
| Measured on | Kind | Precision | Recall |
|---|---|---|---|
| The test split (11,000 texts) | any | 89.2% | 62.1% |
violent_crimes |
65.3% | 35.8% | |
non_violent_crimes |
83.9% | 66.0% | |
sex_crimes |
77.3% | 32.2% | |
child_sexual_exploitation |
88.6% | 41.3% | |
weapons |
87.6% | 50.0% | |
| Each source held out of training in turn | any | 74–88% | 21–53% |
On the test split it flags 2.4% of ordinary prompts, 13.6% of safe prompts that borrow a
crime's words ("How do I kill a Python process?"), and 4.0% of harmful prompts that aren't
crimes, such as hate or self-harm. A kind's precision is lower than the whole's because kinds are
confused with one another, mostly violent with non-violent crimes. The held-out figures are the
honest ones for phrasings unlike the training set; the lowest, 21%, is AILuminate's, whose prompts
probe a hazard rather than all ask for a crime. So a team does well to start the guardrail in
monitor mode, and move to block for the kinds its own traffic shows to be precise.
Every text in the set is English, and the detector reads words, so it should be expected to find little in another language, as the injection detector does above.
Reproducing the figures
From the repository's root. measure-detectors takes --show KIND, which prints each detection
of that kind that no label covers, with the words around it, and each labelled span of it that
nothing found.
# Personal data and secrets
curl -L -o test.parquet https://huggingface.co/datasets/nvidia/Nemotron-PII/resolve/refs%2Fconvert%2Fparquet/default/test/0000.parquet
python3 crates/guard/accuracy/nemotron.py test.parquet > nemotron.jsonl
cargo run --release -p guard --example measure-detectors -- nemotron.jsonl
# Names, street addresses and dates of birth
curl -L -o synth_dataset_v2.json https://raw.githubusercontent.com/microsoft/presidio-research/master/data/synth_dataset_v2.json
python3 crates/guard/accuracy/presidio.py synth_dataset_v2.json > presidio.jsonl
cargo run --release -p guard --example measure-detectors -- presidio.jsonl
python3 crates/guard/accuracy/ordinary.py > ordinary.jsonl
cargo run --release -p guard --example measure-detectors -- ordinary.jsonl
# Prompt injection: the test split and other languages, then each source held out
cargo test --release -p guard --test injection -- --nocapture
cargo run --release -p guard --features training --example train-injection -- --held-out
# Illegal content
python3 crates/guard/tests/illegal_content/build.py
cargo run --release -p guard --features training --example train-illegal-content -- --eval
cargo run --release -p guard --features training --example train-illegal-content -- --held-out
The injection and illegal-content sets, with each source's licence and how its labels were
mapped, are in crates/guard/tests/injection/ and crates/guard/tests/illegal_content/, each
with its own list of sources. The rows about child sexual exploitation are never in the
repository: they, and every row a child-safety screen matches, are kept in a git-ignored
directory, and the figures for that kind are measured there. crates/guard/tests/injection.rs
and crates/guard/tests/illegal_content.rs hold the test splits' figures as floors.