How grades work
Each audit reads one repository at one commit, extracts its tool definitions and runs deterministic rules. A Claude reviewer may confirm or dismiss rule findings with a reason code, and may add findings only when it quotes the source verbatim. Code, not the model, computes the score.
Score
A finding costs its category severity points × confidence factor. One rule can cost a category at most 2× its severity points. Categories start at 100; the total is their weighted mean.
| severity | points |
|---|---|
| critical | 40 |
| high | 20 |
| medium | 8 |
| low | 3 |
| confidence | factor |
|---|---|
| high | 1.0 |
| medium | 0.6 |
| low | 0.3 |
| category | weight |
|---|---|
| injection | 0.3 |
| capabilities | 0.2 |
| exfiltration | 0.25 |
| supply chain | 0.15 |
| transparency | 0.1 |
Grades: A ≥ 95, B ≥ 85, C ≥ 70, D ≥ 55, otherwise F. Any critical finding caps the grade at D.
Rules
| rule | category | severity | dismissable | checks |
|---|---|---|---|---|
| hidden-instructions | injection | critical | no | Tool descriptions with hidden instruction blocks, instruction overrides, requests to hide behaviour from the user, or references to local credential files. |
| invisible-unicode | injection | critical | no | Zero-width and other invisible characters in tool descriptions. |
| tool-shadowing | injection | critical | no | Descriptions that instruct the agent about other tools. |
| auditor-addressed | injection | high | no | Text addressed to security reviewers or scanners. |
| encoded-blob | injection | high | yes | Long encoded blobs inside tool descriptions. |
| name-collision | injection | medium | yes | Tool names that collide with common tools, so agents may call the wrong one. |
| opaque-description | transparency | medium | yes | Descriptions built at runtime that could only be partly checked. |
| shell-exec | capabilities | high | yes | Runs shell commands or external programs. |
| dynamic-eval | capabilities | high | yes | Evaluates dynamically built code. |
| sensitive-path | exfiltration | high | yes | References local credential files (SSH, cloud, agent client configs). |
| hardcoded-url-call | exfiltration | medium | yes | Network calls to hard-coded external URLs. |
| hardcoded-secret | supply_chain | high | yes | Credentials committed in source (the published quote stops before the secret). |
| unpinned-deps | supply_chain | low | yes | Dependencies without version pins. |
| archived | supply_chain | medium | yes | Archived repository. |
| stale | supply_chain | low | yes | No commits for over a year. |
| no-license | transparency | low | yes | No license declared. |
| no-readme | transparency | low | yes | No README. |
| llm-review | any | any | yes | Added by the Claude reviewer; kept only when its quote matches the source verbatim. |
Accuracy
On 10 labelled cases the critical/high rules score precision 1.0 and recall 1.0. CI fails if either drops.
Disclosure
Harsh results are approved by a person before publishing. Apparently unintentional vulnerabilities are held for 30 days while the maintainer is told; the listing shows only counts until then. Audits older than 30 days, or whose tool definitions changed since, are marked stale.
rules v1.0.0 · scoring v1.0.0 · trust.review@1