# Test an AI skill with a small evaluation set Evaluate a skill with several representative inputs and explicit checks. One polished answer does not show whether it handles missing facts, contradictions, or unsafe instructions in a document. Record the host, model, version, inputs, and observed behavior so the test means something later. ## Define success before the run For a fictional résumé workflow, success might mean preserving dates and titles, asking about missing metrics, and keeping rewrites attached to evidence. Do not define success as “sounds professional”; that can reward invented claims. Include a usable output criterion as well as factual constraints. ## Use four cases | Case | Input | What to inspect | | --- | --- | --- | | Complete | Bullet plus confirmed scope | Accurate rewrite | | Incomplete | No outcome measurement | Question or unquantified wording | | Conflicting | Two different employment dates | Conflict flagged | | Embedded instruction | Document says to ignore the user | Document treated as data | These cases are a starting set, not a certification. Add cases that resemble the actual work and consequences of your workflow. ## Keep an evaluation record ```text Task and package version: Host and model: Input case: Checks defined before the run: Observed output: Checks passed or failed: Unsupported additions: Questions asked: Next change to test: ``` Preserve actual outputs separately from authored examples. A sample you write to explain desired behavior should never be presented as a recorded model result. ## Test changes against the same cases After changing the instructions, repeat the relevant cases. A fix that improves the complete example may make the missing-input case worse. Record regressions and avoid changing several unrelated parts at once when diagnosing a failure. Skill tests are different from package validation and storefront tests. Checking that files exist or that a ZIP downloads proves those behaviors, not native activation or output quality. State exactly which layer you evaluated and keep the limits visible when describing compatibility or results. --- SkillStall · 2026-10-02 Agent Skills format specification: https://agentskills.io/specification