Method

How we test skills against Claude.

Some product pages carry a line that says a skill was tested against a named model on a named date. This page says exactly what that means, and where it stops meaning anything.

What a test is

We load the skill file as the assistant's instructions, exactly as shipped. We then give the assistant three fixed tasks, one at a time, and record what it does. A second, larger model reads each response against a written rubric and marks each check pass or fail. The result is stored with the model name, the date and the rubric version.

The line on a product page is that record, nothing more: which model, which day, and how many of the checks it passed.

The three probes

  • A task the skill is for. A realistic request that sits squarely inside what the skill says it does.
  • An adjacent task. A nearby request the skill could inform but was not built for. We want it to help where it fits and not force its whole procedure onto the task.
  • A task it should refuse to take over. A request that the skill's own “Do NOT use” clause excludes. We want the assistant to answer plainly, not to apply the skill's voice or ritual anyway.

Probes are written once, from each skill's own description, and then frozen. The same tasks are used at every re-test, so results are comparable over time.

The eight checks

  • Triggering (3 checks, one per probe). Did the response adopt the skill where it should, and hold back where it should? Generic advice that any assistant would give does not count as adopting it.
  • Adherence (2 checks, task and adjacent probes). Does the response use the skill's own concepts, steps or questions, and avoid what the skill says it would not do?
  • Format (3 checks, one per probe). Is the response complete and intact: not cut off, no unfilled placeholders, no leaked instructions?

The grader sees the skill, the task and the response, and returns a structured verdict. The rubric is versioned. Records currently use rubric version 1.0.

Where the results stand

72 items have a test record so far: 24 Tools and 48 featured frameworks, one per category. Across them, 569 of 576 checks passed. Responses were generated by Claude Sonnet 5.5 and graded by Claude Opus 5.5. The most recent run was Oct 11, 2026.

An item with no record on its page has not been tested yet. That is not a failure and not a warning. It means it is waiting its turn. We would rather show nothing than show a placeholder.

When we re-test

Every time a new major Claude model is released, and at least once a quarter. Each re-test writes a new dated record over the old one, so the date on a product page is always the date of the most recent run.

What this does not tell you

  • A pass means the assistant behaved as the rubric describes on three tasks. It does not mean the framework gives good advice for your situation.
  • The grader is a model. It can be wrong, and model output varies between runs. A record is one run, not an average.
  • We load the skill as the assistant's instructions. We do not test whether Claude decides to load the skill by itself in a long conversation, which depends on the product you use it in.
  • Three probes cannot cover everything a skill might be asked. They catch gross failures, not subtle ones.
  • This is our own testing, not an independent audit. Results are published as we record them, including low scores.

Related

For what is inside a purchase and how to verify the files, see Security. For the catalog itself, see Browse and Tools.