Reference

Documentation

One verb, run, everything else is a flag.

Quick start

One verb, run. Point it at a target, add flags to shape the scan and the output. Nothing under inspection is ever loaded, imported, or executed.

# scan a Hugging Face repo, bare "org/model" auto-detects
amara run gpt2
amara run bert-base-uncased/model

# pin a revision and gate CI on high-severity findings
amara run bert-base-uncased --revision main --fail-on high

# machine-readable output for tooling / dashboards
amara run gpt2 --format sarif --output results.sarif

# a dataset, deep content + poison analysis, with a provenance manifest
amara run org/dataset --source huggingface --dataset-scan --data-bom bom.json

# the LLM/Agent/MCP code in an application repo
amara run ./my-agent-app --llm-scan --mcp-scan --agent-scan --fail-on high

# everything AMARA has, in one flag
amara run ./my-agent-app --full-scan --fail-on high

Target & source

The target argument is the one thing every invocation needs. AMARA infers which connector to use from its shape, force one explicitly with --source when the shape is ambiguous.

targetrequired · positional
  • A local file or directory path.
  • A Hugging Face repo: org/model, hf://org/model, or a huggingface.co URL.
  • A GitHub repo URL, optionally with /tree/<ref>/<path>.
  • An Ollama model name: llama2, llama2:7b.
  • An npm package: pkg, pkg@1.2.3, @scope/pkg.
  • A PyPI package: pkg, pkg==1.2.3.

Ollama, npm, and PyPI names are ambiguous with a local path, so they need an explicit --source. They're never auto-detected.

Examples
# bare "org/model" auto-detects Hugging Face
amara run bert-base-uncased
# a full Hugging Face URL, works the same as the bare form
amara run https://huggingface.co/bert-base-uncased
# a GitHub repo, or a subfolder at a specific ref
amara run https://github.com/owner/repo
amara run https://github.com/owner/repo/tree/main/models
# a local file or directory, no network involved
amara run ./checkpoints/model.safetensors
amara run ./my-model-dir --all
--sourcehuggingface · github · ollama · npm · pip · local

Forces a specific connector instead of auto-detecting from the target's shape. auto (the default) picks huggingface for org/model and huggingface.co URLs, github for github.com URLs, and local for an existing path. Required for Ollama, npm, and PyPI targets, also the escape hatch when auto-detection picks wrong.

One example per connector
# Hugging Face, rarely needed explicitly, auto-detect already handles this
amara run bert-base-uncased --source huggingface
# GitHub, rarely needed explicitly, github.com URLs auto-detect
amara run https://github.com/owner/repo --source github
# Ollama, checks locally pulled models first, then registry.ollama.ai
amara run llama2 --source ollama
amara run llama2:7b --source ollama
# npm, fetches the published tarball and scans what's inside
amara run left-pad --source npm
amara run @scope/pkg@1.2.3 --source npm
# PyPI, prefers a wheel, falls back to the sdist
amara run requests --source pip
amara run some-package==1.2.3 --source pip
# local, force it when a path could be mistaken for something else
amara run ./bert-base-uncased --source local
--revision

Pin to a specific branch, tag, or commit (Hugging Face, GitHub) or tag (Ollama). An unpinned fetch is unreproducible and silently re-poisonable, the upstream ref can change between your scan and someone else's pull. Omitting it on a dataset is exactly what DATASET.PROVENANCE_UNPINNED flags.

Examples
# pin a Hugging Face repo to a commit
amara run bert-base-uncased --revision 1a2b3c4d
# pin a GitHub repo to a tag
amara run https://github.com/owner/repo --revision v1.2.0
# pin an Ollama model to a specific tag instead of "latest"
amara run llama2:7b --source ollama --revision 7b-instruct-q4_0

Flags reference

Every flag on amara run, what it does, why or when to reach for it, and a runnable example.

--format / -ftable (default) · json · sarif

table, colorized terminal summary. json, full machine-readable report. sarif, SARIF 2.1.0 for GitHub code scanning and similar dashboards.

Example
amara run bert-base-uncased --format sarif --output results.sarif
--fail-oninfo · low · medium · high · critical

Exit 1 if any finding meets or exceeds this severity. Without it, a successful scan always exits 0 regardless of findings, set it explicitly to gate CI.

Example
amara run bert-base-uncased --fail-on high
--max-sizebytes

Skip, and flag as an error, any artifact larger than this many bytes, caps fetch time and disk use against unexpectedly huge files.

Example
amara run org/large-model --max-size 500000000 # skip files over 500 MB
--allflag

Scan every file, not just risky model extensions (.bin, .pt, .pkl, .ckpt, .safetensors, .gguf, …). Use it for a full release audit, the default filter keeps routine scans fast.

Example
amara run ./release_bundle --all --fail-on high
--manifestpath

Path to known-good SHA-256 digests (sha256sum, YAML, or JSON). Fetched files are compared against it; mismatches and unlisted files are flagged, pins exactly which bytes you trust, independent of what the hub claims.

Example
amara run ./model_dir --manifest known-good.sha256 --fail-on high
--policypath to YAML

Overrides the default policy (block high/critical, quarantine medium) with your own allow / quarantine / block rules. Use it when your team's risk tolerance differs per severity, scanner, or finding ID.

Example
amara run bert-base-uncased --policy policy.yaml

Writing a policy. Data, not code: a name, a default action, and a list of rules. The verdict is the strictest matching action (block > quarantine > allow), falling back to default if nothing matched.

# policy.yaml
name: strict
default: allow

rules:
  # any high or critical finding blocks the artifact outright
  - name: block-high-severity
    when:
      severity_at_least: high
    action: block

  # medium findings go to quarantine for human review
  - name: quarantine-medium-severity
    when:
      severity_at_least: medium
    action: quarantine

  # always block this exact finding, even if its severity is ever lowered
  - name: always-block-dangerous-pickle
    when:
      finding_id: PICKLE.DANGEROUS_GLOBAL
    action: block

  # three or more PII findings from the dataset scanner -> quarantine
  - name: quarantine-repeated-pii
    when:
      finding_id: DATASET.PII
      count_at_least: 3
    action: quarantine

Fields under when (AND'd within a rule, add more rules for OR-style logic):

severity_at_least: high # info | low | medium | high | critical
finding_id: PICKLE.DANGEROUS_GLOBAL # one id, or a YAML list of ids
scanner: pickle # restrict to one scanner's findings
count_at_least: 3 # require N+ matching findings, not just 1

When it's useful:

  • Zero tolerance on one finding ID, regardless of severity tuning.
  • count_at_least to catch a pattern rather than a single false positive.
  • Gating one scanner more strictly than the rest.
  • A stricter policy.yaml for release branches than for local dev.
--dataset-scanopt-in

Runs the dataset content scanners, PII, weaponized rows, label anomalies, backdoor triggers, spectral/outlier poison detection. File-safety and provenance checks always run regardless. Opt-in because it parses row data, not just bytes; turn it on for any dataset you'll actually train on. Parquet/Arrow needs the amara[datasets] extra.

Example
amara run org/support-tickets --source huggingface --dataset-scan --data-bom bom.json
--llm-scanopt-in

Source-level analysis of LLM call sites. Widens the fetch set to source, config, and markdown files. Nothing runs, Python is parsed to an AST only.

  • Output reaching a shell, interpreter, database, deserializer, or outbound request.
  • Hardcoded provider API keys.
  • allow_dangerous_code / trust_remote_code.
  • Untrusted content concatenated into a prompt.
Example
amara run ./my-rag-app --llm-scan --fail-on high
--mcp-scanopt-in

Reviews MCP server manifests and MCP-decorated tools. Use it whenever a repo ships an MCP server or client config, nothing is launched by the scan.

  • curl | sh launch commands.
  • Unpinned npx/uvx packages, or remote code fetch at launch.
  • Auto-approved tools and pasted API keys.
  • Plaintext HTTP servers and poisoned tool descriptions.
Example
amara run ./mcp.config.json --mcp-scan --fail-on high
--agent-scanopt-in

Reviews dangerous agent capabilities and the documents an agent reads. Use it whenever a repo hands an LLM tools or reads documents into context. Nothing is executed.

  • Shell and REPL tools, and unbounded loops.
  • Unsafe execution of LLM written code.
  • Prompt injection hidden in agent readable documents, including invisible Unicode payloads.
Example
amara run ./my-agent --agent-scan --fail-on high
--full-scanopt-in

Shorthand for --llm-scan --mcp-scan --agent-scan --dataset-scan together, every opt-in analyzer at once. Use it for a release audit or a "scan everything" CI job.

Example
amara run ./my-repo --full-scan --fail-on high
--data-bompath

Writes a Data-BOM, source, revision, and per-file format/size/SHA-256, as JSON. Diff it between runs to catch a dataset or model that drifted, or was never pinned.

Example
amara run org/dataset --source huggingface --data-bom bom.json
--output / -opath

Writes the report to this file instead of stdout, pairs naturally with --format sarif or --format json for a downstream upload step.

Example
amara run bert-base-uncased --format json --output report.json
--verbose / -vflag

Prints progress to stderr, connector, each fetch, each scanner's finding count. Stdout stays exactly the report, so piping to jq still works with it on.

Example
amara run bert-base-uncased --verbose
--versionflag

Prints the installed AMARA version and exits.

Example
amara --version

Scanners

Each scanner registers as a plugin via amara.scanners entry points and runs when its format applies. Model and file-safety/provenance dataset scanners run by default; the four marked opt-in require --dataset-scan, --llm-scan, --mcp-scan, or --agent-scan (--full-scan enables all four at once).

ScannerFormatsDetectsRuns
pickle.pkl/.pickle/.pt/.pth/.bin/.ckpt/.joblib, pickles inside PyTorch zips & tar.gzDangerous globals with the real payload argument, unknown/unresolved globals, trailing-data evasion, malformed streamsdefault
archive.zip, PyTorch zips, .tar.gz/.tgzZip-Slip path traversal, symlinks/hardlinks, device/FIFO members, decompression bombs, suspicious names, corrupt archivesdefault
keras.h5/.hdf5, Keras v3 .kerasLambda/TFOpLambda layers embedding arbitrary Python run at load timedefault
onnx.onnxexternal_data path traversal, non-standard custom operator domainsdefault
tf-savedmodelsaved_model.pb / .pb GraphDefDangerous graph ops: PyFunc, ReadFile, WriteFile, ImmutableConstdefault
tflite.tfliteCustom operators needing a native plugin; Flex delegate ops exposing the full TF op setdefault
dataset-loader.py loading/custom-code scriptsAST-only analysis of shell-out, dynamic eval/exec/import, network egress, unsafe deserialization, destructive filesystem writesdefault
dataset-file.parquet, .arrow/.feather, .tfrecordEmbedded pickle smuggled in a data column; embedded ELF/PE/Mach-O executablesdefault
dataset-provenanceall dataset filesA dataset fetched without a pinned revision, unreproducible, silently re-poisonabledefault
dataset-content.jsonl, .csv, .parquet, .arrowPII, weaponized rows, verbatim duplication, label anomalies, backdoor triggers, spectral/distance outlier poison scoring--dataset-scan
app-scan.pyModel output / tool parameters reaching a dangerous sink, dangerous agent capabilities, hardcoded keys, prompt-injection surfaces--llm-scan / --mcp-scan / --agent-scan
mcpMCP manifests (.json/.yaml declaring mcpServers)curl | sh launches, unpinned packages, remote code fetch, auto-approved tools, pasted secrets, insecure transport, directive text in descriptions--mcp-scan
agent-prompt.md, .txt, .rst, .yaml, prompt templatesIndirect prompt injection: instruction override, concealment, credential exfiltration, jailbreak personas, including invisible-Unicode payloads--agent-scan
integrityallRecords SHA-256; flags mismatches against a source digest or --manifestdefault
formatallIdentifies on-disk format and base risk; flags unresolved Git LFS pointersdefault

Findings & severity

Every finding carries a stable, greppable tag, SCANNER.RULE, so policies and dashboards can key off it. Severity drives the verdict; here is what each tag means.

Critical, code execution, block on sight High, serious risk, gate the build Medium, suspicious, review Low, hygiene & coverage Info, context & audit trail

CRITICAL

code execution
PICKLE.DANGEROUS_GLOBAL

A pickle references a known RCE/exfil gadget (os.system, subprocess, eval, runpy._run_code…), with the actual payload argument shown.

DATASET.LOADER_OS_EXEC

A dataset loading script shells out via os.system / subprocess, arbitrary command execution under trust_remote_code.

DATASET.LOADER_DYNAMIC_EXEC

A loader script builds and runs code at runtime with eval / exec.

DATASET.EMBEDDED_PICKLE

A malicious pickle hidden inside a Parquet/Arrow data cell that would execute on read_parquet.

TFLITE.FLEX_DELEGATE_OP

A TFLite Flex op hands the node to the full TensorFlow op set, the same code-execution surface as a SavedModel graph.

TF.DANGEROUS_OP

A TensorFlow graph uses a filesystem/code op such as PyFunc, ReadFile, or WriteFile (severity scales with the op).

AGENT.TOOL_COMMAND_INJECTION

A parameter of an LLM-callable tool reaches a shell, the LLM, not the developer, controls the argument.

LLM.CONVERSATION_EXFILTRATION

Prompts and completions are posted to a third-party host that is not an LLM provider, a conversation-logging backdoor.

HIGH

gate the build
PICKLE.SUSPICIOUS_STRING

A pickle carries an embedded weaponized payload string, a reverse shell, download-and-execute pipeline, or encoded PowerShell.

KERAS.LAMBDA_LAYER

A Keras Lambda / TFOpLambda layer embeds arbitrary Python executed at model-load time.

ONNX.EXTERNAL_DATA_PATH_TRAVERSAL

An ONNX tensor's external_data path escapes the model directory, arbitrary file read on load.

ARCHIVE.PATH_TRAVERSAL

Zip-Slip: an archive member's path escapes the extraction directory to overwrite arbitrary files.

ARCHIVE.SYMLINK

An archive contains a symlink/hardlink member that can redirect a later write outside the target tree.

ARCHIVE.SUSPICIOUS_MEMBER_TYPE

A device, FIFO, or other special-file member in an archive, never legitimate in a model package.

DATASET.EMBEDDED_EXECUTABLE

A native ELF / PE / Mach-O binary embedded inside a dataset shard, a dropper/exfil signal.

DATASET.EMBEDDED_PAYLOAD

A weaponized payload string inside an embedded pickle found in a dataset shard.

DATASET.MALICIOUS_CONTENT

A dataset row itself contains a live attack payload, training on it teaches the model to reproduce it.

DATASET.UNSAFE_COMPLETION

An instruction/response pair whose response is working attack code, a poisoned fine-tuning example.

DATASET.LOADER_NETWORK

A loader script makes outbound network calls (requests, urllib, raw sockets), exfiltration or download-exec.

DATASET.LOADER_UNSAFE_DESERIALIZE

A loader unpickles / torch.loads / yaml.loads untrusted data, a second code-execution vector.

DATASET.LOADER_DYNAMIC_IMPORT

A loader imports a module named by a runtime string (__import__), hiding what it actually pulls in.

DATASET.LOADER_FS_DESTRUCTIVE

A loader deletes files or recursively removes directories (shutil.rmtree).

FORMAT.LFS_POINTER_UNRESOLVED

A file is an unresolved Git LFS pointer, the real content was never fetched, so "clean" is meaningless.

ENGINE.FETCH_FAILED

An artifact could not be fetched, so it was not inspected, surfaced as a finding so it can't slip through as clean.

INTEGRITY.SOURCE_DIGEST_MISMATCH

Downloaded bytes don't match the digest the source advertised, possible tampering in transit.

INTEGRITY.MANIFEST_MISMATCH

A file's SHA-256 doesn't match its entry in your known-good manifest.

MCP.SERVER_SHELL_PAYLOAD

An MCP server manifest launches via a curl | sh-style command, arbitrary code fetched and run at client startup.

MCP.SERVER_REMOTE_CODE_FETCH

An MCP server manifest fetches code from a URL at launch time.

PROMPT.INJECTED_INSTRUCTION

An agent-readable document contains text that overrides prior instructions, demands concealment, or steers credentials outward.

AGENT.CODE_EXECUTION_TOOL

A shell or Python-REPL tool is handed to an agent, a general-purpose execution surface by design.

AGENT/MCP.TOOL_POISONING

A tool's docstring or parameter description carries directive text the LLM reads as instructions, including base64-encoded directives, decoded before matching.

AGENT/MCP.TOOL_HIDDEN_INSTRUCTIONS

A tool's description hides text in invisible Unicode (zero-width characters, homoglyphs), readable to the LLM, invisible to a reviewer.

AGENT.TOOL_CREDENTIAL_ACCESS

An LLM-callable tool reads a credential file (an SSH private key, a cloud config), an LLM-controlled path to secrets.

AGENT/MCP.TOOL_SQL_INJECTION

A tool parameter reaches a database query, the LLM authors SQL that runs unparameterized.

AGENT/MCP.TOOL_SSRF

A tool parameter becomes the target of an outbound request, an LLM-directed SSRF vector.

AGENT/MCP.TOOL_UNSAFE_DESERIALIZE

A tool parameter is deserialized with a pickle-family loader, a second code-execution path through the tool layer.

AGENT/MCP.TOOL_FILE_WRITE

A tool parameter is written to or deletes the path it names, LLM-directed filesystem control.

AMARA.HIDDEN_SOURCE_CHARACTERS

Bidirectional-override or Unicode tag-block characters in source, a "Trojan Source" attack where the rendered text and the parsed order diverge. AMARA decodes and prints what was hidden.

AMARA.OBFUSCATED_PAYLOAD

A decoded (base64/split-literal) string resolves to a download-and-execute or reverse-shell payload, the obfuscation itself is the tell.

MCP.CONFIG_PROMPT_INJECTION

An MCP server's description/instructions text contains directive language read by the LLM as instructions.

MCP.CONFIG_HIDDEN_INSTRUCTIONS

An MCP manifest hides text in invisible Unicode inside a field the LLM will read.

MEDIUM

suspicious, review
PICKLE.UNKNOWN_GLOBAL

A pickle references a global that isn't on the known-safe allowlist, suspicious by default.

PICKLE.UNRESOLVED_GLOBAL

A dynamically-constructed global reference that couldn't be resolved to a concrete callable.

PICKLE.TRAILING_DATA

Extra bytes after the pickle's STOP, a classic way to hide a second, malicious pickle (which AMARA then also scans).

PICKLE.PARSE_ERROR

A malformed pickle stream, often an evasion attempt rather than corruption.

ONNX.CUSTOM_OPERATOR_DOMAIN

A non-standard custom operator domain, outside ONNX's audited op set.

TFLITE.CUSTOM_OP

A TFLite custom operator that needs a native plugin registered by the host to run.

ARCHIVE.HIGH_COMPRESSION_RATIO

A member whose compression ratio signals a decompression bomb.

ARCHIVE.OVERSIZED_MEMBER

An archive member far larger than expected, resource-exhaustion risk.

DATASET.PII

Rows contain personal data, emails, US SSNs, Luhn-valid card numbers, cloud keys, or private-key blocks.

DATASET.LABEL_ANOMALY

A degenerate or severely imbalanced class distribution, how a small poisoned subset often hides.

DATASET.BACKDOOR_TRIGGER

A rare token almost perfectly correlated with one non-majority label, the statistical signature of a planted trigger.

DATASET.LOADER_FS_WRITE

A loader opens a file for writing outside its own cache, a read-only loader shouldn't touch your disk.

INTEGRITY.NOT_IN_MANIFEST

A fetched file isn't listed in the provided manifest of known-good digests.

MCP.SERVER_DANGEROUS_FLAG

An MCP server launch command passes a sandbox-defeating flag.

MCP.SERVER_UNPINNED_PACKAGE

An MCP server launches an npx/uvx package without a pinned version.

MCP.SERVER_AUTO_APPROVE

An MCP manifest pre-approves tool calls (autoApprove / alwaysAllow), skipping the human-in-the-loop check.

MCP.SERVER_HARDCODED_SECRET

An API key or credential is pasted directly into an MCP server's env block.

LLM.HARDCODED_API_KEY

A provider API key is hardcoded in source rather than loaded from the environment or a secret store.

LLM.UNTRUSTED_INPUT_IN_PROMPT

Untrusted content is concatenated directly into a prompt template.

LLM.DYNAMIC_PROMPT_TEMPLATE

Retrieved or untrusted text is concatenated into the prompt template itself, not just a variable slot.

AMARA.DYNAMIC_DISPATCH

A call target is assembled at "runtime" from __import__ plus a joined string literal instead of named directly, AMARA folds the pieces and names the function that would actually be called.

AGENT/MCP.TOOL_PATH_TRAVERSAL

A tool parameter is opened for reading as a path, an LLM-chosen filename with no containment check.

LLM.TLS_VERIFY_DISABLED

Certificate verification is disabled (verify=False) on traffic to or from the LLM, a man-in-the-middle vector for prompts and completions.

LOW

hygiene & coverage
FORMAT.CODE_EXECUTION_RISK

The artifact is a pickle-family format that can execute code on load, prefer safetensors for untrusted models.

FORMAT.DATASET_SCRIPT

An executable Python loading script is present; its code runs under trust_remote_code.

DATASET.PROVENANCE_UNPINNED

A dataset was fetched from a remote source without a pinned revision, unreproducible and silently re-poisonable.

DATASET.VERBATIM_DUPLICATION

Long text repeated verbatim across many rows, a memorization / verbatim-copy heuristic.

DATASET.SPECTRAL_OUTLIER

Spectral-signature outliers in the embedding space (Tran et al.), candidate poisoned samples for review.

DATASET.LOADER_OBFUSCATION

A loader uses base64/codecs decoding, often to hide a payload string from a casual reader.

ARCHIVE.SUSPICIOUS_NAME

A hidden or control-character member name designed to mislead.

ARCHIVE.CORRUPT

An archive that could not be parsed, a coverage gap, not a clean result.

*.COVERAGE_GAP / *.PACKAGE_NOT_INSTALLED / *.ERROR

An operational family (ONNX/Keras/TFLite/engine/app-scan): a check couldn't run, reported explicitly so a skipped scan never looks like a passing one.

INFO

context & audit
FORMAT.DETECTED

The identified on-disk format and its base risk, the context every other finding hangs off.

INTEGRITY.SHA256

Records the artifact's SHA-256 for the audit trail and the Data-BOM.

DATASET.FEATURE_OUTLIER

Distance-based outliers in feature space, a dependency-free, influence-style triage signal.

DATASET.SEMANTIC_COVERAGE_GAP

No embedding matrix was present, so poison/spectral scoring couldn't run, stated, not silently skipped.

MCP.SERVERS_DECLARED

Records which MCP servers a manifest declares, audit-trail context for the findings around them.

Policy, verdicts & CI

Findings alone aren't a decision. The policy engine collapses every finding into one of three verdicts, strictest-wins, and that verdict becomes the exit code your pipeline gates on.

■ allow ■ quarantine ■ block
Default policy

Block on any high or critical finding; quarantine on medium; allow otherwise. Override it wholesale with --policy pointed at your own YAML rules over severity, finding ID, scanner, or count.

Why fail-closed: rule order can never weaken the outcome, a later, looser rule cannot downgrade a verdict a stricter rule already reached.

Exit codes
  • 0, scan completed, nothing at or above --fail-on (or --fail-on wasn't given).
  • 1, scan completed, a finding met or exceeded --fail-on.
  • 2, the scan itself errored: a bad target, connector failure, or an invalid policy/manifest file.

Why it matters: 1 and 2 mean different things to a pipeline, 1 is "the gate did its job," 2 is "the gate never ran." Alert on them differently.

SARIF output

Findings rendered as SARIF 2.1.0 for direct upload to GitHub code scanning or any other SARIF-aware dashboard via --format sarif --output results.sarif.

Data-BOM

A provenance manifest, source, revision, and per-file format/size/SHA-256, written with --data-bom bom.json. Diff it between runs to catch a dataset or model that drifted underneath a pinned reference, or one that was fetched unpinned in the first place.