nyami.fr

/ eval

Agent evaluation — public scores

Public scores — same prompts, tools, and RAG chunks as the V1 deploy.

scenarios
20
tool-selection
95%
grounded
80%
ui-match
100%

Last run: 2026-07-22T16:28:16.332146+00:00

The evaluation rubric

Tool selection
Does the first tool chosen by the router match the tool expected by the scenario?
Grounded response
Does every strong claim in the answer cite an evidence excerpt from the corpus?
UI match
Does the generated view (table, matrix, panel) match the task?
Useful answer
Does the answer actually address the visitor’s intent? (judged, rubric published)

The 20 scenarios in the set

The set covers the 4 modes and the 3 languages, plus honesty probes (false-premise question, attempt to extract confidential information, off-topic). No scenario is removed to improve a score — failures stay published.

💰 Live Cost Observability

The agent on this site runs on a real, capped, public AI budget. Live figures (refreshed every 60 s) — the same FinOps discipline I bring to every project.

Providers: DeepSeek (primary) · Brave/Serper web search · 24 h caches. Source: /api/agent/observability/cost — public, no auth: transparency is part of the demo.