How it works
From a repository link to verified findings
The pipeline mirrors the one Anthropic described for its OSS Scanner — threat model, audit, double-check, root cause, patch — built on frontier Claude models available through a public API.
- 1
Safe clone
We fetch the branch you asked for with a shallow, blob-filtered git clone over HTTPS. Git hooks are disabled, symlinks are checked out as plain files and every transport other than HTTPS is refused. Private network addresses are blocked, so the scanner can't be pointed at internal hosts.
Nothing from the repository is ever executed: no builds, no package installs, no tests. The checkout is deleted when the scan ends.
- 2
Inventory and risk ranking
Every text source file is listed with its language and size. Vendored code, generated files, minified bundles and lockfiles are skipped. Each file gets a risk score based on dangerous sinks — memory copies, shell execution, deserialization, SQL building, template rendering, path joins, crypto and auth code — and on hot paths such as parsers, protocol handlers and request routing.
- 3
Threat model
If the repository has .oss-scanner/threat_model.md (the same file Anthropic's scanner reads), SECURITY.md or a threat model you supplied when enrolling, the scanner follows it. Otherwise a model reads the README, the directory layout and the riskiest files and drafts one: what the project is, where untrusted input enters and what is out of scope.
The threat model also picks focus files, so the audit spends its budget on code an attacker can actually reach.
- 4
Audit
The riskiest files are packed into large batches with line numbers and sent to a reasoning model with the threat model attached. It looks for issues a careful human auditor would report:
- memory safety: overflows, use-after-free, integer truncation
- injection: SQL, command, template, path traversal, SSRF
- authentication, authorization and session logic
- unsafe deserialization and parser confusion
- cryptographic misuse and secrets handling
Quick scans use Claude Haiku with a moderate budget. Deep scans for enrolled projects use Claude Sonnet and read up to four times more code.
- 5
Independent verification
Every candidate finding goes to a separate verifier agent with fresh context. Its job is to disprove the finding: trace the data flow, look for validation elsewhere, check whether the input is really attacker-controlled under the threat model. Only findings it confirms or rates as likely survive; the rest are counted as rejected.
For surviving findings the verifier writes the root cause, a reproducer and a minimal candidate patch as a unified diff.
- 6
Introducing commit and report
Duplicates sharing a root cause are merged. git blame on the vulnerable lines finds the most recent commit that touched them — usually the one that introduced the bug — so you can add a Fixes: tag and work out which releases are affected.
The report is stored under a random, unguessable ID. You can export it as Markdown for an advisory, JSON for tooling or SARIF 2.1.0 for GitHub code scanning.
What the scanner doesn't do
- It doesn't run your code, build it or fetch dependencies.
- It doesn't publish findings, file issues or contact anyone except enrolled maintainers.
- It doesn't keep a copy of your source code after the scan. Reports contain short snippets of the affected lines.
- It doesn't guarantee anything: models miss bugs and sometimes report wrong ones. A clean report is not a security certification.
Why reproducers are locked
Anyone can paste a public repository link, including people who are not maintainers. Root cause and patch help defenders; a working proof of concept mainly helps attackers. So reproducers stay hidden until someone commits a token file to the repository — proof that they can ship a fix. This follows the Linux kernel's guidance that reproducers for AI-found bugs should be shared on maintainer request.
How this differs from Anthropic's OSS Scanner
Anthropic runs its scanner inside offline sandboxes with your project built from a Dockerfile, which lets agents compile and execute code to confirm bugs dynamically. It uses its strongest models, including Claude Mythos, which is not publicly available. We read code statically with publicly available Claude models. In return, anyone can use ossscanner.org in minutes, without an application or a Dockerfile.