Vulnerability validation: why a failed attempt is no proof of security
A commissioned security test (penetration test) checks whether known vulnerabilities in IT systems are actually exploitable — not merely present in theory. The laborious, error-prone part is working through each individual candidate by hand. And an uncomfortable truth applies: when an exploitation attempt fails, that does not mean the system is secure — the attempt can fail for reasons that have nothing to do with the vulnerability. A robust test result is therefore not a yes/no, but a graded verdict backed by evidence. This note describes an internal tool that enforces that grading.
Context
The tool — internally EasyMS — drives the established
Metasploit framework through its RPC interface
(pymetasploit3 → msfrpcd) and works through a
defined test assignment: given a target and a list of candidates
(a CVE identifier or an explicitly named module), it runs a low-risk
check per candidate, on request a controlled exploitation attempt,
and assigns the result to an evidence-backed class. Operation is restricted to
authorised targets: check-only is the default, an exploit runs only
with an explicit --fire, and a scope allowlist hard-rejects hosts
outside the assignment.
The crux: a failed attempt is no proof of security
An exploit that opens no session only shows that this attempt failed — not that the vulnerability is absent. In practice the cause is regularly a boundary condition rather than the absence of the vulnerability: a payload incompatible with the target, a firewall on the return channel, timing. A binary result collapses exactly this distinction and thereby produces false all-clears — the most expensive class of error in a security test.
The way out is to keep two signals separate that are usually conflated:
- the module's CheckCode — Metasploit's own, low-risk assessment “vulnerable / not vulnerable / unclear”, without exploitation;
- the actual session — the diff over
sessions.listbefore and after the attempt, the reliable truth about “did I gain access”.
Combining both signals yields four evidence-backed verdict classes
(plus ERROR for operational failures such as module or
connection problems):
- CONFIRMED — an interactive session was opened; the access itself is the proof.
- LIKELY — the
checkreports vulnerable, but no session was established. This is explicitly not “secure”, but the candidate for targeted manual follow-up. - NOT_VULNERABLE — the
checkrobustly reports “not vulnerable”. - INCONCLUSIVE — the
checkcannot decide (unsupported by the module or ambiguous).
The practical value lies in LIKELY: this class surfaces the candidates that a yes/no tool would silently have booked as “secure”, and steers scarce manual review attention exactly where it makes a difference.
Why default payloads fail — two examples reproduced in the lab
The LIKELY case is not a fringe phenomenon but follows from the
defaults. Two cases, reproduced against an isolated test lab:
- UnrealIRCd: the module chooses no payload of its own — without an explicit choice, simply no session is created.
- distcc: the default is
reverse_bash, which requires/dev/tcp; the target's aged Bash does not provide it — the attempt fails on the shell, not on the vulnerability.
The answer to this is a short, compatible payload fallback list that is tried in
turn until a real session is created. This moves cases that would have failed on
the payload choice alone from LIKELY to CONFIRMED — the
difference between “presumed” and “proven”.
Traceability
Every run leaves three artefacts: a run.jsonl as a
SHA-256 hash chain (verifiable via --verify; any later change,
deletion or reordering becomes detectable — the same mechanism as in our
audit-log format,
here as a standalone evidence log), the console transcript as proof,
and a replayable .rc with which any finding can be reproduced. A test
result is thereby not merely an assertion, but one that a third party can
recompute.
Status and limits
The verifiable state is a prototype, verified in July 2026 against an isolated Metasploitable2 lab (Docker, its own network, never exposed to the LAN). On throughput or use in real engagements this deliberately says nothing — that could not currently be substantiated. And the tool does not replace a tester's judgement; it structures the preparatory work, makes the sources of error visible and the results traceable. That is precisely where the value lies: a graded, evidenced verdict is more honest than a tidy yes/no — and steers review time to where it is needed.