PROJECTSKILLSCONNECTING…

SYS.33 / AUDIT

Auditing the Skills

A Skill is prose an agent executes with your privileges. The material it was written from is untrusted by construction, so the finished Skill is read back.

What comes out6 SECTIONS

Why this exists

The generator reads your repository, and your repository contains text other people wrote.

A README, an issue template, a comment, a request in somebody’s own words — all of it is input to the generator, which then writes a document an agent will follow. That makes the finished Skill worth reading back with rules, looking for text that would make an agent exfiltrate data, reach a credential, override its own safety rules, or act unasked.

No model is involved in the audit, so it runs on every plan and on every analysis.

Polarity, which is the whole of it

A sentence forbidding X does not instruct X — and getting that wrong hides findings.

The rule is one line: a construction may be read as non-instructing only when an agent reading the Skill would also not act on it. That is not a suppression list; it is the test for whether the rule’s own claim is true.

The criterion decides the hard cases without anybody picking them, and not always in the flattering direction. A fenced code block stays instructing — an agent may well run what is in a fence, which is what fences in a Skill are mostly for.

The bar a rule has to clear

Zero false positives on a measured corpus, or the rule does not ship.

At a high false-positive rate a reader learns to dismiss the seal, and the feature fails silently while every screen still shows a verdict — which is worse than not shipping it, because it looks like a control.

A rule that cannot clear the corpus is removed from the shipped set rather than tuned. Its risk category then appears in the list of categories and not in what the audit covered, and the interface reports that risk as not checked.

What it refuses to do

Three things, each of which would be easy and wrong.

No score
The data model makes a score unrepresentable rather than merely absent, for the same reason findings have none.
No blocking
The verdict is advisory. pskl pull warns and continues. Given the measured false-positive rate of this class of analysis, anything else would be stopping your work on a guess.
No claim of safety
Polarity is decided from syntax, and syntax is sometimes ambiguous in ways a reader resolves instantly and a parser does not. Where they could disagree, the audit reports.

Suppressing a finding

A judgement binds to the exact text it was made about.

A suppression carries to a new version only while that text appears byte-identically there. Where it changed, the judgement lapses, the finding returns, and somebody looks again. A lapsed judgement keeps its reason and its date, and is never deleted.

Reading the seal

A registration mark, not a badge.

The corner ticks are the attestation: two means measured, none means nobody looked, one means the audit did not finish. Without them there is no way to draw the difference between a Skill with no findings and a Skill nobody examined — and the only remaining options are a green tick or a score, both of which are claims this does not make.

In the lattice beside it, density carries how many and the status square carries how bad. From the terminal, pskl skills audit prints the same material and leads with how many Skills nobody has audited.

DOCS23 CHAPTERS