# PS-01 — unsafe-code tag contract

Status: candidate specification with a working development integration; not release-verified.

## Baseline job and outcome

Identify supported network-to-shell and encoded-payload execution patterns in an explicitly identified AI response. Show the person the matched instruction, a plain-language explanation, and the limits of the finding before they act.

The tag does not establish a scam, malicious intent, human/AI authorship, or that untagged code is safe.

## Architecture before implementation

The phone/paste route runs the pinned method in a self-hosted browser worker, without uploading or redacting answer text; its result is a device-preview record. The optional local-service route is distinct: one selected site adapter binds captured text to a response ID and content revision. The extension sends bounded input to an authenticated loopback gateway. Local redaction precedes an isolated evaluator. The gateway validates an assertion; the extension displays it only on the matching response revision. Unsupported detectors are excluded through explicit production capability selection, not invalid confidence floors.

Manual selection is preserved scope; it is not implemented by the current response-check controls. No new endpoint is authorized by this specification.

## In scope

- Download piped into a shell; bounded command, process and backtick substitution with stdout and evaluation-layer checks; recognized encoded-payload execution patterns.
- Use/mention context: recommendations, warnings, quotations, comments, and fenced code.
- Initial extension route: Chrome and ChatGPT; additional adapters remain separate work. The phone/paste route works independently of that adapter and keeps its own device/offline/accessibility gate.
- Personal local workflow, accessible evidence, identifiable/versioned record.
- Distinct pending, no-supported-finding, uncertain, unsupported, and unavailable states.

## Required acceptance scenarios

1. Supported recommendations produce the expected finding and exact redacted-text offsets.
2. Explicit warnings and explanatory mentions do not become execution recommendations.
3. Fenced commands are not blanket-exempt; nearby unrelated negations do not suppress hazards.
4. Delayed results, regeneration, navigation, and changed content cannot receive stale tags.
5. Repeated capture produces one current presentation per response revision.
6. Invalid authentication, dependency outage, malformed upstream output, and oversized input have defined outcomes and never masquerade as evaluated uncertainty.
7. Only redactor output is forwarded to the evaluator; sensitive content can remain when redaction misses it. Application logs and persistence contain no content in tested paths. Redaction coverage limitations are disclosed.
8. Keyboard, screen-reader, zoom, forced-colors, and reduced-motion paths meet prewritten criteria.
9. A clean installation and removal work for an outside pilot participant.
10. Tag explanations expose method/version, subject binding, evidence, limitations, and a challenge path.

## Evaluation gate proposal

Keep the existing precision 0.93 and recall 0.75 goals as candidate thresholds. Freeze independent evaluation data and agree the interval method and category coverage before using them as release gates. Publish TP/FP/FN/TN, denominators, uncertainty intervals, and category failures. The existing 24 examples are regression material, not a representative held-out corpus. No mandatory MT agreement gate applies to this tag. Heuristic scores are not probabilities without calibration evidence.

Latency must be measured on declared hardware and input lengths before accepting a budget. Do not force a minimum abstention rate on a deterministic detector without a defensible rationale.

## Failure containment and maintenance

A required redaction failure stops affected evaluation. Optional integrations cannot stop the personal workflow. New detector versions run the existing regression suite plus independent evaluation. Falling below an accepted gate narrows or disables affected issuance pending correction. Issued records remain identifiable; corrections supersede rather than silently rewrite them.

## Ownership

Product owner: accepts the product outcome and release scope. Claim and evidence reviewers: claim wording, labeling guidance, independent review. Engineering lead: architecture, implementation, executed verification. Independent human reviewers: ground truth and accessibility/usability judgment.

## Remaining decisions

The installed Chrome extension has been verified on one live ChatGPT Homebrew answer, including evidence and source-hash binding. Broader live-site coverage remains unverified. Confirm statistical acceptance policy before independent evaluation. Team, enterprise, and proprietary applications stay preserved; they do not block this personal tag workflow. Signing and lifecycle authentication require an explicit threat-model decision before release.

## Detailed acceptance packet

See [PS-01 acceptance](PS-01-acceptance.md). RFC-0004 proposes UC as a successor; no rename or registration is recorded as accepted.


## Current development evidence

The context-v5 routing correction passes 86 examples across five active datasets:
TP 31, FP 0, FN 0, TN 55. This is development regression evidence. The original
24 examples and earlier runtime reports remain historical references. This pass
also executed the evaluator HTTP endpoint and full-schema gateway tests with
controlled upstream responses, plus a harmless authored file-only curl routing
experiment. It did not restart the local container daemon or establish independent
accuracy. The downloadable CLI refuses a mismatched old evaluator source hash.

Use instructions are included in the site's PS modal and crawlable reference.
The same installation provides PII redaction before PS. The technical local-service route requires Linux/macOS setup and optional unpacked Chrome installation. The separate phone/paste preview needs no Docker, account or extension; browser-engine checks have executed, while actual-phone and human accessibility evidence remain incomplete.

## Execution work package

[TAG-PS](../scope-delivery.md#tag-ps) preserves the broader instruction-risk scope and separately validates each enhancement. Credential-exfiltration and typosquat rows are retained work, not coverage inferred from the current 86 command regressions. No unrelated tag depends on PS completion.
