Skip to content

Skill and instruction system audit — 2026-07-17

  • Tease: Behavior Is The Remaining Risk
  • Lede: The skill-system audit found strong structural health but identified behavioral risks involving fabricated defaults, authority, stale capabilities, validation, and secrets.
  • Why it matters:
    • Frontmatter and link checks cannot prove that instructions behave safely at runtime.
    • The highest-priority fix makes HTMA tooling fail closed when evidence is insufficient.
  • Go deeper:
    • Review the eight findings and the measured before-and-after checks.

Checked: 2026-07-17

Method: Static inventory, direct source inspection, runtime-topology checks, focused script execution, and a fresh-context verification pass. Scope: 123 custom or repository-packaged skills; 23 scoped and root agent-instruction files; the active Claude, Cursor, and Codex instruction bridges; and managed skill installation topology.

  • All 123 skill frontmatter blocks parsed and supplied a name and description.
  • No genuine broken relative skill links remained after illustrative code and dynamic paths were excluded.
  • The managed-link check passed all 49 registered packages.
  • The strongest governing contracts—no fabricated data, durable evidence, narrow validation, sensitive-data handling, and protected-skill ownership—are worth preserving.

These checks establish structural health. They do not prove that every instruction produces the intended behavior.

PriorityContract areaFindingDisposition
1Data integrityHTMA helper scripts substituted numerical defaults, interpreted one calibration observation, shipped realistic-looking seed rows, and omitted required memo status fields.Fixed first; evidence below.
2AuthoritySome wrappers treated chat, ticket, PR, push, and account changes as implicitly authorized.Route through explicit task-scoped authorization and owned bridges.
3Runtime discoveryInstructions named unavailable models, stale tool semantics, and a missing commit-workflow skill.Discover current capabilities and keep supported fallbacks.
4Toolchain precedenceA generic Bun rule conflicted with applications that declare pnpm, Vite, Vitest, or dotenvx workflows.Defer to the nearest app contract; keep Bun as an unopinionated default.
5Docs verificationContent checks were described as broader than they are, while live authentication and secret-transfer behavior lacked a separate proof contract.Split local preflight from live verification and secure secret handling.
6Internal consistencySome output shapes contradicted their validators; other instructions promised secret safety while reading complete configuration files.Keep one rule per behavior and fail closed around secrets.
7Skill routingProduction and experimental HTMA families had indistinguishable triggers; several research skills overlapped.Make experiments explicit-only and add positive plus negative trigger boundaries.
8Scoped freshness and topologySeveral dated instructions, cwd assumptions, canonical-source claims, and runtime links had drifted.Refresh scoped facts and record ownership for every custom package.

This public memo intentionally omits private workspace paths, sensitive context, and protected third-party source details. Protected packages are integration targets, not direct edit targets.

The before/after checks use explicitly synthetic fixtures. The values below are observed program output, not real estimates or calibration evidence.

CaseBeforeAfter
Structured VOI item with only a synthetic nameExited 0, silently produced expected value 0.25, invented a measurement, and recommended measure now.Exits 2, keeps expected value unknown, and names the missing measurement, decision-change chance or sensitivity, cost of being wrong, information quality, and measurement cost.
One synthetic calibration observationReported that coverage was “broadly consistent” with stated confidence.Reports insufficient_history and descriptive counts only. Pattern interpretation begins at five comparable estimates by default.
VOI and calibration CSV templatesIncluded plausible-looking example rows that could be mistaken for observations.Contain headers only.
Measurement memo templateOmitted estimate_status, blocking_missing_inputs, assumed_target, and next_measurement_step despite requiring them in the skill contract.Includes all four fields in both production and experimental templates.

The five-estimate calibration threshold is an operational guardrail against single-case conclusions. It is not presented as proof of statistical power, and callers can raise it when the decision requires more history.

The canonical notes-search repository contains a deterministic regression that checks both production and experimental HTMA packages. The recorded run verifies missing-input behavior, a complete synthetic VOI calculation, the calibration sample gate, header-only templates, and all required memo fields.

python3 scripts/check-htma-data-integrity.py
OK: HTMA data-integrity checks passed

The validator source records the cross-family checks. The public HTMA Measure contract states the same fail-closed rule: missing VOI inputs remain unknown or explicitly qualitative; they never become numerical assumptions.

  • The audit did not modify protected third-party packages; their fixes belong in owned bridges or upstream patches.
  • Structural checks and deterministic helper tests do not replace representative model evaluations.
  • Only the first improvement area is implemented here. The remaining ordered work stays visible on the roadmap.