Labels and trust
Quard labels every piece of content an agent reads: the user’s message, a web page, an email, a tool result or a message from another agent. The label says where the content came from, and that decides how far Quard trusts it.
The three parts of a label
| Part | What it says | Example |
|---|---|---|
| Origin | The tool, domain, sender, server or agent that let the content in | web:supplier-portal.example |
| Trust | Trusted or untrusted | untrusted |
| Sensitivity | Internal or public | public |
Where the origin comes from
The origin comes from the guarded tool that let the content in, such as the web fetch tool that loaded a page or the inbox tool that read an email. It never comes from the content itself or from an AI. A page cannot claim to be trusted, and no model is asked to judge where text came from.
Trust and sensitivity follow from the origin.
Default trust
| Origin | Trust | Sensitivity |
|---|---|---|
| The user | trusted | internal |
| Your own tools, with any guard except a source guard | trusted | internal |
| Web pages and hosted web search | untrusted | public |
| Email read by inbox tools, colleagues’ mail included | untrusted | public |
| MCP servers | untrusted | public |
| Files read by source tools | untrusted | internal |
| Another agent | the labels of what it sent | the labels of what it sent |
| Unknown: tools without a guard, unlabeled memory and unwatched channels | untrusted | internal |
Changing the defaults
Teams change the trust or sensitivity of one origin at a time, in code. For example, a team can trust its own MCP server and mark it internal, because it only serves the team’s own CRM data. Every change is recorded in the run, and the dashboard lists them all under Origin overrides.
Apps open to the public can mark the user as untrusted.
The context label
Each call also gets a context label: the least trusted and most sensitive label among everything the model read before it asked for the call. When that label is untrusted, the call is influenced. An agent that reads a web page and then sends an email makes an influenced call.
An agent that browses the web is influenced almost all the time. So rules for sensitive actions also look at where each argument came from, not only at the context. See Value tracing.
In the dashboard, the run timeline colors every step by its context label.
Labels across agents
Content keeps its labels when it moves between agents. A brief from another agent that arrives as a user message still carries the web label of the page it came from, because Quard labels content by matching it against what the run has already read, not by the role of the message.
- Across processes, a message carries only the run id, the step that sent it and a reference to its labels. The receiving side looks the labels up. A missing or broken reference counts as untrusted. See Multi-agent.
- Shared memory keeps its labels too. They are stored with a hash of the content and come back on read, even in a later run. Content changed outside Quard fails the hash check and reads back as untrusted. See Shared memory.
Labels and AI
AI never decides where content came from, or how far Quard trusts it. Content that a model judges could argue its way to trusted, so the origin stays a recorded fact.
An AI detector adds a label that says what the content is, such as invoice, article or payment_fraud. Quard’s detector, Jev, labels public content only. Web pages, email and MCP results count as public, intranet hosts included, so mark your own servers internal to keep their data away from it. When its risky labels are likely enough, the content gets a flag such as detector:payment_fraud. A value from flagged content no longer counts as coming from its origin, so rules for sensitive actions treat it with suspicion.
AI detectors can only make things stricter. A risky label can flag content, or strip the chunk, up to 4,000 characters, that holds a likely prompt injection. It can never allow what a rule blocked, or raise trust. Content the detector could not check, because it failed, answered late or could not reach its API, is flagged detector:unchecked. The detector acts by default. A team can switch it to observe mode, where it only saves its labels. People check the labels in the review queue.
Turn on the detector
The detector is off until you set it up in code, and until then nothing is labeled. Get an API key from TypeSafe, the company that runs Jev, and pass it to jevDetector():
import { jevDetector, quard } from "quard";
quard.configure({
detector: jevDetector({ apiKey: process.env.TYPESAFE_API_KEY ?? "" }),
});jevDetector() throws when the key is empty, so a missing key shows up at startup.
Once set up, the detector runs in enforce mode: it waits up to 5 seconds for its answers, then flags or strips. To see its labels before it changes anything, switch it to observe mode. It saves the labels without waiting, and changes nothing:
quard.configure({ detectorRules: { mode: "observe" } });The policy file can switch it as well, with "detector": { "mode": "observe" }.
Email and the detector
Every email origin counts as public, so mail from colleagues read by an inbox tool goes to the detector too. The origin comes from the inbox tool and its arguments, such as email:readInbox, never from the sender. To keep that mail away from the detector, mark the origin internal:
quard.configure({
origins: { "email:readInbox": { sensitivity: "internal" } },
});This keeps all the mail that tool reads away from the detector, outside mail included, so none of it is labeled.