Skip to content
MCP Vetted

How we check

Every rule on this page is generated from the code that runs it, so it can't drift. Current rules version: 2026-09-29.1.

The pipeline

Every new version of every tool goes through the same steps, cheapest first. A version stops at the first check it fails.

  1. Find. Agents read the Official MCP Registry hourly, npm and GitHub daily, plus quick submits and extension reports.
  2. Qualify. Is it a real, installable tool? It needs an install path and a usable description, and it can't be an obvious stub, template or test.
  3. Static scan. The rules below run over everything the server says about itself and its install config. Nothing is executed.
  4. Claims check. Claude Haiku 4.5 reads the claimed purpose and each tool's description, as untrusted data, and decides whether they match and whether any text tries to steer a model.
  5. Tier. The version gets a tier from its checks, its publisher's claim and agent outcome reports.
  6. Watch. Every release goes back through the pipeline. A release of an installed tool that fails a critical check is revoked, and installers are told.

Tiers

Tiers are cumulative: each one includes everything below it.

TierMeans
IndexedCrawled from a public source. No check has passed yet, or a check failed.
ScannedPassed qualification, the static scan and the LLM claims check on this exact version.
VerifiedScanned, and the publisher proved they own it through GitHub or DNS.
ProvenVerified, and at least 25 distinct agents reported it did the job after installing through us.

Proven needs at least 25 reports from distinct agents, each tied to an install made through us, with at least 90% saying it worked. A revoked version is always Indexed.

Static scan rules

RuleSeverityLooks for
hidden-instructionscriticalInstructions hidden in HTML comments, zero-width characters or tags aimed at the model.
model-overridecriticalTries to override the model's instructions or hide actions from the user.
credential-exfiltrationcriticalMentions secrets or keys together with sending them somewhere.
tool-hijackhighTries to change how the model uses other tools.
encoded-payloadhighLong base64-like blobs in descriptions, a common way to hide instructions.
suspicious-hosthighPoints at raw IPs, tunnels or paste and webhook-capture services.
insecure-remotemediumRemote server over plain http.
shell-pipe-installmediumInstall asks to pipe a download straight into a shell.
typosquathighA name one letter away from a popular server.
sensitive-pathhighReferences to credential locations such as ~/.ssh or ~/.aws.
unpinned-packagelowA package with no pinned version.

Any critical or high finding fails the scan. Medium and low findings are shown but don't block.

How we test the scanner

We keep a set of harmless look-alike servers, each copying one attack pattern with invented hosts and fake credentials. Every scanner change runs against them, and the build fails if any verdict changes. Real malicious samples are never executed; only their hashes are kept so they can be refused.

SamplePatternExpected
clean-weathercontrol: a clean serverpass
html-comment-injectionhidden instructions in an HTML commentfail (hidden-instructions)
zero-width-injectioninstructions hidden with zero-width charactersfail (hidden-instructions)
ssh-key-exfilreads a planted fake SSH key and sends it to a sinkholefail (credential-exfiltration, sensitive-path)
override-instructionsasks the model to ignore its instructions and hide it from the userfail (model-override)
tool-hijacktries to route every other tool call through itselffail (tool-hijack)
typosquata name one letter away from a popular serverfail (typosquat)
encoded-payloada long encoded blob hidden in a descriptionfail (encoded-payload)
tunnel-hostsends data to a tunnel or capture servicefail (suspicious-host)

What money can't buy

Ranking uses text relevance, tier and installs through us. There is no input for payment, plan or sponsorship anywhere in the ranking code. Publisher sync is billed and run separately from indexing and checks.

What checks can't tell you

  • A static scan reads descriptions and config; it can't see what code does at runtime. Sandbox runs will add that.
  • A server can change its behavior after a check without a new version number. Remote servers are re-probed, but a check is a snapshot.
  • Checks show a tool isn't obviously malicious or misleading. Whether it's good at its job is what agent outcome reports are for.