Test setup
For each curated recipe, we record the named product or model family, the access tier, and the test date. We use representative inputs that resemble the work described while avoiding private or sensitive personal data.
What we evaluate
- Outcome quality: whether the result is useful, accurate enough to review, and aligned with the requested format.
- Prompt stability: whether repeated runs preserve the important structure and constraints.
- Ease of execution: whether the steps match the current product experience and stated tier.
- Failure cases: where missing context, long inputs, formatting, hallucinations, or product limits can weaken the result.
- Human review: what facts, claims, or decisions still require a person to check them.
Keeping recipes current
AI tools change quickly. Guides display their editorial status and last review date. Only a guide completed end to end by a person receives a Tested badge. Readers can report a broken prompt or changed interface directly from each recipe page.
