AI security needs a chain of provenance from context to action
Securing production AI requires more than a safe model: identity, retrieved context, permissions, policy decisions, and tool execution must remain linked in an auditable chain.
AI security discussions often centre on the model: whether it can be jailbroken, induced to leak sensitive data, or manipulated through its prompt. Those questions matter, but production systems increasingly connect retrieval, model inference, tool calls, APIs, and actions in other systems.
An analysis published by AI News argues that the critical property is not merely whether each component is secure in isolation. Organisations must preserve provenance across every transformation: where information came from, which identity introduced it, what the model was permitted to do, which policy governed the next step, and what action ultimately occurred.
Risk often appears at the seams
A model can behave correctly while the application around it creates risk. Retrieval may place untrusted instructions in context. A tool may accept arguments that the model should never generate. An application may authorise an action using broad account permissions even though the specific task did not require them.
Meaning changes at these boundaries. A retrieved document is data to an application but language—and potentially an instruction—to a model. Model output is a suggestion until an execution layer treats it as a command. Individually legitimate components can therefore combine into an unsafe chain.
Provenance begins before the prompt
The first question is not only what the user typed, but who or what initiated the task, under which identity, and with what authority. An enterprise workflow may begin with a person, a scheduled process, an application responding to a business event, or another agent delegating a subtask.
Security metadata should travel with the request, including initiating identity, device or workload context, data classification, approved purpose, and the maximum authority available. If that context disappears after the first model call, downstream controls have to reconstruct intent from incomplete evidence.
Retrieved context needs trust labels
RAG systems ground responses in enterprise data, but retrieved material may contain malicious instructions, stale permissions, manipulated text, or information the user is not entitled to combine with other sources. The retrieval layer should preserve source identity, access rights, freshness, and trust level instead of flattening everything into an anonymous block of text.
This record helps determine whether a conclusion relied on authoritative data, whether sensitive information crossed a policy boundary, and whether a source was still valid when the decision was made. The article also points to NIST’s 2026 Cyber AI Profile workshop report as an example of connecting AI-specific attack surfaces with broader cybersecurity governance.
Structured output is not trusted output
Valid JSON, a correct function name, or schema-compliant arguments prove structure, not authorisation. Model output should be treated as untrusted input to a separate decision layer. Before a tool executes, trusted code should verify the initiating identity, the target resource, whether the data use is permitted, and whether the risk requires human confirmation.
The model can recommend an action, but it should not define its own authority. This separation prevents a well-formed response from becoming an unauthorised real-world operation.
One identifier for the full agent run
If an agent reads a support ticket, retrieves customer records, modifies a CRM entry, drafts an email, and sends it, the security record should not be five unrelated logs. A single task identifier should connect the initiating request, data sources, authorisation decisions, tool invocations, user confirmations, and resulting changes.
That chain helps investigators distinguish legitimate automation from manipulated automation. It also enables controls to tighten when the task changes from reading to modifying, or from internal processing to external communication.
Traceability supports revocation and repair
Documents are corrected, permissions change, models are updated, and tools are replaced. Strong provenance lets an organisation identify which outputs or actions depended on a compromised source, vulnerable component, or revoked permission. Without those links, estimating the blast radius becomes difficult.
The objective is not to record hidden model reasoning. Security needs observable causality around the system: what entered, which sources were used, which identity acted, which tools were available, which policy decision occurred, and what changed. A practical chain is: request identity → retrieved context → model interaction → proposed action → authorisation decision → tool execution → resulting state.
As AI moves from generating content to changing business systems, provenance becomes a security primitive. The realistic goal is not to assume every component can be made perfectly trustworthy, but to trace how trust changes from context to action and stop execution when one link no longer deserves it.

Source: AI News