A short, practical primer across the AI stack: what makes each risk dangerous, and the exact AMARA scan that catches it.
Before an LLM, an agent, or an MCP server ever enters the picture, there's a simpler question: is the checkpoint you just pulled from a hub actually safe to load? This is AMARA's original, always-on capability, the malware/backdoor and dataset scans run by default, with no opt-in flag required.
A pretrained checkpoint is not inert data, pickle-family formats
(.pt, .pth, .bin, .ckpt) execute
arbitrary Python the moment they're deserialized, and a Keras Lambda layer or
a TensorFlow PyFunc op runs code the moment the model loads. There is nothing
in a hub's UI that visually distinguishes a clean checkpoint from a backdoored one.
torch.load/pickle.load an unreviewed file.safetensors for untrusted sources, it carries no code-execution surface at all.pickletools, never deserializing, and treats every global reference as suspicious until affirmatively matched as a known-safe helper. Runs by default, no flag needed.Fine-tuning on a public dataset trains in whatever it contains, a rare token engineered to correlate almost perfectly with one label plants a backdoor trigger invisible to a human skimming the rows; a malicious pickle smuggled into a Parquet column executes the moment the shard is read.
--dataset-scan for PII, poisoned rows, and backdoor-trigger statistics.A model fetched without a pinned revision isn't reproducible, the "same" repo id can serve different bytes tomorrow, silently re-poisonable after you've already reviewed it once. A digest mismatch means the bytes you received don't match what the source advertised: tampering in transit, or at the source itself.
--revision, never trust an unpinned "latest."--manifest.--data-bom on every scan and diff it between runs in CI to catch upstream drift.Anywhere your application calls an LLM and does something with the answer, two new trust boundaries appear: what goes into the prompt, and what happens to what comes out. Most LLM incidents start at one of these two boundaries.
Anything the LLM reads, a retrieved document, a web page, a user message, another tool's output, can contain text written to look like an instruction. The LLM can't reliably tell "content to summarize" from "commands to follow," so if that text reaches the instruction stream, the attacker is now talking to your LLM directly.
<reference>...</reference>) and tell the LLM in the static part that it's data, not instructions.LLM.UNTRUSTED_INPUT_IN_PROMPT, LLM.DYNAMIC_PROMPT_TEMPLATE.An LLM's answer is not proof of intent. It's text, produced by a system a prompt injected attacker can influence. Handing that text to a shell, a SQL cursor, or an interpreter means the LLM's word is treated as a trusted command, and by extension so is whoever managed to steer it.
LLM.OUTPUT_TO_COMMAND_EXEC, LLM.OUTPUT_TO_CODE_EXEC, LLM.OUTPUT_TO_SQL.A committed API key is a standing credential anyone with repo access, or a leaked repo, now owns. Quieter but just as real: disabling certificate verification on LLM traffic turns every prompt and completion into something a network position can read or rewrite.
os.environ["API_KEY"], never a literal.LLM.HARDCODED_API_KEY, LLM.TLS_VERIFY_DISABLED.An agent is an LLM with hands: tools it can call, permissions it can exercise, and often a loop that lets it keep going without a human in between. Every one of those capabilities is also something a prompt-injected LLM can misuse, the risk isn't the LLM being wrong, it's the LLM being steered.
A shell tool, a Python REPL, or an unbounded loop is a general-purpose capability handed to a system that a document it reads can influence. Granting more capability than a task needs doesn't make the agent more useful on the happy path, it just widens what a successful injection can do.
restart_service(name) function, not a shell.max_iterations / time budget on every agent loop.allow_dangerous_code=True as a reviewed, logged exception, not a default.AGENT.CODE_EXECUTION_TOOL, AGENT.DANGEROUS_CODE_ENABLED, AGENT.UNBOUNDED_ITERATIONS.A tool's name and description aren't just documentation for humans, the framework hands them to the LLM verbatim so it knows when and how to call the tool. Text planted there is read by the LLM as instructions, even though nobody reviewing the code thinks of a docstring as an attack surface.
AGENT.TOOL_POISONING, AGENT.TOOL_HIDDEN_INSTRUCTIONS.The LLM, not your code, chooses the arguments it passes to a tool. If that argument reaches a shell or a query unchanged, a prompt-injected LLM is effectively calling that shell or query itself, your own tool becomes the attacker's execution primitive.
{"restart-api": [...]}, the LLM can only pick a name.AGENT.TOOL_COMMAND_INJECTION, AGENT.TOOL_SQL_INJECTION, AGENT.TOOL_SSRF.OWASP's GenAI Security Project also publishes agentic-AI-specific threat guidance alongside the numbered LLM Top 10, and the wider security community has converged on a similar set of categories for autonomous, tool-using systems. These aren't a fixed numbered list the way the LLM Top 10 is, so treat the rows below as risk categories, not citations.
| Category | What it means | AMARA detects via | Flag |
|---|---|---|---|
| Tool / resource poisoning | A tool's description or parameter text carries instructions read by the LLM, not the developer. | AGENT.TOOL_POISONING, AGENT.TOOL_HIDDEN_INSTRUCTIONS, including base64-decoded and invisible-Unicode variants | --agent-scan |
| Excessive agency / critical-systems interaction | An agent wired directly to a shell, interpreter, or live database with no gate. | AGENT.CODE_EXECUTION_TOOL, AGENT.SQL_AGENT, AGENT.DANGEROUS_CODE_ENABLED, AGENT.DANGEROUS_DESERIALIZATION_ENABLED | --agent-scan |
| Authorization / control-flow hijack | An LLM-controlled value reaching a sink it should never influence, the LLM becomes the attacker's proxy. | AGENT.TOOL_COMMAND_INJECTION, AGENT.TOOL_SQL_INJECTION, AGENT.TOOL_SSRF, AGENT.TOOL_FILE_WRITE | --agent-scan |
| Credential / secret exposure to the agent | A tool that reads private keys, cloud credentials, or config secrets on the LLM's behalf. | AGENT.TOOL_CREDENTIAL_ACCESS, LLM.HARDCODED_API_KEY | --agent-scan / --llm-scan |
| Unbounded blast radius | An agent loop with no iteration or time budget, able to compound a mistake indefinitely. | AGENT.UNBOUNDED_ITERATIONS | --agent-scan |
| Knowledge-base / memory poisoning | A "long-term memory" or retrieval store used as a second, less-reviewed injection surface. | The same taint + poisoning engine applied to memory-store tools; DATASET.BACKDOOR_TRIGGER on the training side | --agent-scan / --dataset-scan |
| Agent supply chain | A chain, executor, or plugin loaded from a remote or unpinned source. | AGENT.REMOTE_CHAIN_LOAD, AGENT.UNSAFE_CHAIN, MCP.SERVER_UNPINNED_PACKAGE | --agent-scan / --mcp-scan |
An MCP manifest is a list of programs your AI client launches with your privileges, the moment it starts. Nobody reads a JSON config the way they'd review a pull request, which is exactly why it's become a favorite place to hide something.
A server entry's command is just a program the client will run.
There's nothing stopping it from being a shell one-liner that pulls a script from the
internet and runs it, and nothing in the UI that would show you before it happens.
curl | sh or any command that fetches and runs code inline.npx/uvx packages to an exact version, an unpinned name resolves to whatever the registry serves today.command/args the same way you'd review a CI pipeline step.MCP.SERVER_SHELL_PAYLOAD, MCP.SERVER_UNPINNED_PACKAGE.Approval prompts exist so a human sees a sensitive tool call before it runs.
autoApprove exists to skip that for trusted, low-risk tools, but set broadly,
it turns every tool on that server into one the LLM can invoke with nobody watching,
including tools that were never meant to run unattended.
autoApprove, never a wildcard./.MCP.SERVER_AUTO_APPROVE, MCP.SERVER_BROAD_FILESYSTEM_ROOT.An env block in a manifest is plaintext on disk, inherited by
whatever the server process does next, and a remote server reached over plain
http:// puts every prompt, tool result, and bearer token on the wire for
anyone on the network path.
https:// for every remote MCP server.env blocks and flags plaintext remote transports, MCP.SERVER_HARDCODED_SECRET, MCP.SERVER_INSECURE_TRANSPORT.MCP is new enough that there's no single settled "Top 10" the way there is for LLM
applications, guidance is still forming across the security community, OWASP's GenAI
Security Project included. The categories below are the risk classes that keep coming up
in MCP security write-ups, mapped to what --mcp-scan actually checks.
| Category | What it means | AMARA detects via |
|---|---|---|
| Launch-time command injection | A server manifest that boots via a shell one-liner, curl | sh, an inline script, a dangerous CLI flag. | MCP.SERVER_SHELL_PAYLOAD, MCP.SERVER_DANGEROUS_FLAG |
| Supply chain / unpinned execution | An npx/uvx package with no pinned version, or code fetched from a URL at launch. | MCP.SERVER_UNPINNED_PACKAGE, MCP.SERVER_REMOTE_CODE_FETCH |
| Credential / secret exposure | A provider API key or credential pasted directly into a manifest's env block, inherited by the server process, sitting in plaintext on disk. | MCP.SERVER_HARDCODED_SECRET |
| Excessive permission / broad root | A filesystem server rooted at / or a home directory instead of a scoped folder. | MCP.SERVER_BROAD_FILESYSTEM_ROOT, MCP.SERVER_SENSITIVE_PATH |
| Confused deputy / approval bypass | autoApprove / alwaysAllow entries that skip the human-in-the-loop check the client was designed around. | MCP.SERVER_AUTO_APPROVE |
| Insecure transport | A remote server reached over plain http://, exposing every prompt, tool result, and bearer token on the wire. | MCP.SERVER_INSECURE_TRANSPORT |
| Tool poisoning | A tool's name, description, or parameter text carrying directives the LLM reads as instructions, including text hidden in invisible Unicode. | MCP.TOOL_POISONING, MCP.TOOL_HIDDEN_INSTRUCTIONS, MCP.CONFIG_PROMPT_INJECTION, MCP.CONFIG_HIDDEN_INSTRUCTIONS |
| Server-side tool injection | The server's own tool implementation handing an LLM-supplied parameter to a shell, query, or filesystem call. | MCP.TOOL_COMMAND_INJECTION, MCP.TOOL_SSRF, MCP.TOOL_PATH_TRAVERSAL, MCP.TOOL_SQL_INJECTION |
The OWASP GenAI Security Project's Top 10 for LLM Applications is the closest thing the field has to a common vocabulary, and most of it is decided by code, before the LLM is ever called. LLM03 · Supply Chain is where AMARA's original, always-on scan lives, see the supply chain risk walkthrough above for the full detail.
| Category | What it means | AMARA detects via | Flag |
|---|---|---|---|
| LLM01 · Prompt Injection | Instructions smuggled into content the LLM reads, overriding the developer's intent. | PROMPT.INJECTED_INSTRUCTION, PROMPT.HIDDEN_INSTRUCTION, AGENT.TOOL_POISONING, MCP.CONFIG_PROMPT_INJECTION, AMARA.HIDDEN_SOURCE_CHARACTERS | --agent-scan / --mcp-scan / --llm-scan |
| LLM02 · Sensitive Information Disclosure | Secrets or PII leaking out through the LLM, its config, or its training data. | LLM.HARDCODED_API_KEY, MCP.SERVER_HARDCODED_SECRET, AGENT.TOOL_CREDENTIAL_ACCESS, DATASET.PII | --llm-scan / --mcp-scan / --agent-scan / --dataset-scan |
| LLM03 · Supply Chain | A compromised model, dataset, or dependency pulled into the application, backdoored weights, poisoned data, tampered or unpinned artifacts. This is the original reason AMARA exists, and the only category here that runs by default with no flag at all. | PICKLE.DANGEROUS_GLOBAL, KERAS.LAMBDA_LAYER, ARCHIVE.PATH_TRAVERSAL, DATASET.EMBEDDED_PICKLE, DATASET.PROVENANCE_UNPINNED, INTEGRITY.SOURCE_DIGEST_MISMATCH, FORMAT.LFS_POINTER_UNRESOLVED, plus MCP.SERVER_UNPINNED_PACKAGE | default scan / --mcp-scan / --dataset-scan |
| LLM04 · Data and Model Poisoning | Training or fine-tuning data engineered to plant a backdoor or bias. | DATASET.BACKDOOR_TRIGGER, DATASET.LABEL_ANOMALY, DATASET.UNSAFE_COMPLETION, DATASET.SPECTRAL_OUTLIER | --dataset-scan |
| LLM05 · Improper Output Handling | LLM output trusted and passed straight to a shell, query, or deserializer. | LLM.OUTPUT_TO_CODE_EXEC, LLM.OUTPUT_TO_COMMAND_EXEC, LLM.OUTPUT_TO_SQL, LLM.OUTPUT_TO_UNSAFE_DESERIALIZE, AGENT.TOOL_COMMAND_INJECTION, MCP.TOOL_SSRF | --llm-scan / --agent-scan / --mcp-scan |
| LLM06 · Excessive Agency | An agent granted more capability, tools, or autonomy than its task needs. | AGENT.CODE_EXECUTION_TOOL, AGENT.DANGEROUS_CODE_ENABLED, AGENT.SQL_AGENT, AGENT.UNBOUNDED_ITERATIONS, AGENT.REMOTE_CHAIN_LOAD | --agent-scan / --llm-scan |
| LLM08 · Vector and Embedding Weaknesses | Poisoned or manipulated vectors in a RAG/embedding pipeline. | DATASET.SPECTRAL_OUTLIER, DATASET.FEATURE_OUTLIER, heuristics scored over an embedding matrix | --dataset-scan |
| LLM10 · Unbounded Consumption | Resource exhaustion via unbounded loops, oversized inputs, or amplification. | AGENT.UNBOUNDED_ITERATIONS, ARCHIVE.HIGH_COMPRESSION_RATIO, ARCHIVE.OVERSIZED_MEMBER, plus --max-size as an operator control | --agent-scan / default scan |
Every finding above is real, running code, not a mockup. One command each, or one flag to run all three at once.
The same detection engine from this page, run against a live console session.
See the console output → Every Flag, Explained--llm-scan, --mcp-scan, --agent-scan, --full-scan, what each widens and why.
Every finding ID on this page, and every other one AMARA emits, in one catalog.
Read the catalog →