Measure whether the feature helps before expanding it.
1. Evaluation cases
Collect representative inputs with expected properties, not just one favorite demonstration. Separate development examples from a held-out set. Track omissions, invented facts and unacceptable actions individually.
2. Data minimization
Send only the information needed for the task. Remove unnecessary identifiers and define retention for prompts and responses. Logs should help diagnose failures without becoming a second copy of private documents.
3. Operational budgets
Measure time to first useful output, end-to-end latency, errors and estimated cost per successful task. Retry budgets and quotas prevent one broken interaction from producing uncontrolled work.
Worked scenario
Two prompts produce equally fluent summaries, but one invents dates in three of ten notes. Fluency alone would hide the regression.
Apply it
Create an evaluation sheet with correctness, missing facts, latency and fallback outcome. Compare a deterministic baseline and a model response.
Check your understanding
Your evaluation can identify a worse prompt even when its prose sounds better. Explain the decision and show evidence from your implementation or design. If you cannot demonstrate it yet, revisit the relevant section before continuing.