androidengineers.Book a session

Decision Inspector and evaluation

Inspect and evaluate every decision

article2–3 hours

Design a record that tells the truth

The Decision Inspector explains one request from input to outcome. Store a record with a request ID, source mode, redacted input, context timestamp, model version, returned alternatives, policy result and execution result. A recommended action is not an executed action.

In Compose, show these as consecutive sections:

  1. Input: reviewed transcript or fictional notification, with Live or Simulated labeling.
  2. Context: relevant observed values, their age and unavailable signals.
  3. Decision: selected route, labeled alternative probabilities and provider confidence when returned.
  4. Policy: allowed, needs confirmation, blocked or superseded, with the application's reason.
  5. Outcome: what completed, failed or remains pending.

Read the TypeSafe response example before mapping its fields. Keep its confidence distinct from an alternative's probability. A probability bar is a model estimate, not evidence that an action is correct. Fixture bars must never be presented as a live provider response.

Use text labels alongside bars and color, support large text, and announce status changes accessibly. UI animation may interpolate to a returned value but must not manufacture intermediate model predictions. Keep the underlying data stable while the display animates.

Measure the complete interaction

Record elapsed time with a monotonic clock. Measure Android request start to validated response separately from execution duration. That request interval includes networking and your backend; do not label it “model inference time.” Show a provider-only duration only if that measurement is actually available and clearly defined.

Measure multiple requests on a named device and connection. Report sample count, median, p95, timeouts and failure rate. A fast failed request does not count as a successful action. Separate cold and warm behavior if you compare them. Avoid using sample numbers from a mockup as benchmark claims.

Compare decisions with labels

Build 30–50 fictional cases containing clear commands, ambiguous requests, unsupported actions, delayed results and hostile text such as “ignore the allowed actions.” Label the expected route and allowed execution before running inference. Split tuning cases from a held-out set.

Compare a small rules baseline with Jev using the same inputs. Track routing accuracy, clarification rate, blocked actions and incorrect executions. Increasing a confidence threshold may reduce both mistakes and useful coverage; report both. Thresholds are experimental settings for your application, not universal safety guarantees.

Record model version and question configuration with results so changes are explainable. Keep runtime infrastructure failures separate from classification mistakes. A small local dataset supports a demo assessment, not a broad claim that one model is better than another.

Practice and checkpoint

Build a list and detail inspector, then run the labeled cases. Include a fixture result, a real result if access is available, a blocked action and a timeout. Export redacted evaluation rows to a local file and derive summary counts from them.

Pass when: a reviewer can reconcile every executed action with a validated request and policy decision, reproduce your metric calculations, and identify which screenshots used fixtures. Do not proceed to more capabilities to conceal poor V1 routing or unreliable execution.

YOUR LEARNING JOURNEY

0 of 7 available lessons completed

Progress saved in this browser. No account needed.
Inspect and evaluate every decision | Jev + Android | Android Engineers