Trust Evidence-based security grades for MCP servers and agent tools.

How grades work

Each audit reads one repository at one commit, extracts its tool definitions and runs deterministic rules. A Claude reviewer may confirm or dismiss rule findings with a reason code, and may add findings only when it quotes the source verbatim. Code, not the model, computes the score.

Score

A finding costs its category severity points × confidence factor. One rule can cost a category at most 2× its severity points. Categories start at 100; the total is their weighted mean.

severitypoints
critical40
high20
medium8
low3
confidencefactor
high1.0
medium0.6
low0.3
categoryweight
injection0.3
capabilities0.2
exfiltration0.25
supply chain0.15
transparency0.1

Grades: A ≥ 95, B ≥ 85, C ≥ 70, D ≥ 55, otherwise F. Any critical finding caps the grade at D.

Rules

rulecategoryseveritydismissablechecks
hidden-instructionsinjectioncriticalnoTool descriptions with hidden instruction blocks, instruction overrides, requests to hide behaviour from the user, or references to local credential files.
invisible-unicodeinjectioncriticalnoZero-width and other invisible characters in tool descriptions.
tool-shadowinginjectioncriticalnoDescriptions that instruct the agent about other tools.
auditor-addressedinjectionhighnoText addressed to security reviewers or scanners.
encoded-blobinjectionhighyesLong encoded blobs inside tool descriptions.
name-collisioninjectionmediumyesTool names that collide with common tools, so agents may call the wrong one.
opaque-descriptiontransparencymediumyesDescriptions built at runtime that could only be partly checked.
shell-execcapabilitieshighyesRuns shell commands or external programs.
dynamic-evalcapabilitieshighyesEvaluates dynamically built code.
sensitive-pathexfiltrationhighyesReferences local credential files (SSH, cloud, agent client configs).
hardcoded-url-callexfiltrationmediumyesNetwork calls to hard-coded external URLs.
hardcoded-secretsupply_chainhighyesCredentials committed in source (the published quote stops before the secret).
unpinned-depssupply_chainlowyesDependencies without version pins.
archivedsupply_chainmediumyesArchived repository.
stalesupply_chainlowyesNo commits for over a year.
no-licensetransparencylowyesNo license declared.
no-readmetransparencylowyesNo README.
llm-reviewanyanyyesAdded by the Claude reviewer; kept only when its quote matches the source verbatim.

Accuracy

On 10 labelled cases the critical/high rules score precision 1.0 and recall 1.0. CI fails if either drops.

Disclosure

Harsh results are approved by a person before publishing. Apparently unintentional vulnerabilities are held for 30 days while the maintainer is told; the listing shows only counts until then. Audits older than 30 days, or whose tool definitions changed since, are marked stale.

rules v1.0.0 · scoring v1.0.0 · trust.review@1