Riferimento API / Evaluation

Evaluation

Gli endpoint del gruppo Evaluation dell'API Voiceland AI, con parametri, schemi ed esempi curl.

Ultimo aggiornamento:

Le descrizioni degli endpoint e dei campi restano in inglese, esattamente come le pubblica l'API. È una scelta voluta: qui legge ciò che vedrà anche nelle risposte.

GET /v1/agents/{name}/evals#

List eval runs. One page of the agent's eval runs, **newest first**, without the per-case matrix (fetch a run for the full detail). ?version_id= narrows to the runs of one agent version, it is applied after the read, so short pages are normal. **Paging.** Pass the next_cursor back as ?cursor=; **an empty next_cursor is the only signal that you have seen everything.**

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
version_id query string Only runs of this agent version. Applied after the read, short pages are normal.
limit query integer Rows per page, 1 to 100 (default 50).
cursor query string The next_cursor from the previous page. Omit for the first page.

Risposte

Codice Descrizione
200 Success.
400 limit is not a positive whole number (bad_request), or the cursor was not issued by this listing (invalid_cursor).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl "https://api.voiceland.ai/v1/agents/{name}/evals" \
  -H "Authorization: Bearer VL_API_KEY"

Una risposta riuscita restituisce un EvalRunsResponse.

Esempio di risposta

{
  "items": [
    {
      "agent_name": "front-desk",
      "agent_version_id": "00001754540000000000000",
      "failed": 1,
      "id": "00001754550000000000000",
      "pass_rate": 0.5,
      "passed": 1,
      "status": "failed",
      "total": 2
    }
  ],
  "next_cursor": ""
}

POST /v1/agents/{name}/evals#

Run the agent's golden set (pre-deploy eval). Executes every golden case against the agent **draft** over the text path, the question is answered with the agent's own system prompt, model and knowledge retrieval, no live call involved, and grades each answer with the platform's LLM judge. The graded run is stored and returned: the per-case pass/fail matrix, the aggregate score, and the diff against the agent's previous run. version_id pins a specific version snapshot; omitted, the latest snapshot (your current draft) runs. The run is what the eval deploy gate reads, see PUT /eval-policy. The call is synchronous and bounded: a golden set larger than the per-run case limit is refused rather than partially graded.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.

Corpo della richiesta

application/json Schema: EvalRunRequest

Campo Tipo Obbligatorio Descrizione
triggered_by string
version_id string

Risposte

Codice Descrizione
201 The stored run, matrix included.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found), or version_id names no retained snapshot (version_not_found).
409 The agent has no golden cases (golden_set_empty), or the set exceeds the per-run case limit (golden_set_too_large).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
503 No eval judge is configured on this deployment (eval_judge_unavailable).
5XX Server or upstream error.

Esempio di richiesta

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/evals" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "triggered_by": "maria@acme.example",
  "version_id": ""
}'

Esempio di risposta

{
  "agent_name": "front-desk",
  "agent_version_id": "00001754540000000000000",
  "cases": [
    {
      "answer": "The premium plan costs 10 euros.",
      "case_id": "gc_9d8c7b6a",
      "fail_reasons": [
        "does_not_match_expected_answer"
      ],
      "pass": false,
      "question": "How much is the premium plan?"
    }
  ],
  "diff": {
    "newly_failing": [
      "gc_9d8c7b6a"
    ],
    "previous_run_id": "00001754530000000000000"
  },
  "failed": 1,
  "id": "00001754550000000000000",
  "pass_rate": 0.5,
  "passed": 1,
  "status": "failed",
  "tenant_slug": "acme",
  "total": 2
}

GET /v1/agents/{name}/evals/{run_id}#

Get an eval run. One run in full: the per-case pass/fail matrix (each case's question, the draft answer, its sources, the judge's verdict and the fail reasons), the aggregate score, and the diff against the previous run.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
run_id path string Run id.

Risposte

Codice Descrizione
200 Success.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no run with that id (run_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl "https://api.voiceland.ai/v1/agents/{name}/evals/{run_id}" \
  -H "Authorization: Bearer VL_API_KEY"

Esempio di risposta

{
  "cases": [
    {
      "case_id": "gc_9d8c7b6a",
      "fail_reasons": [
        "not_grounded"
      ],
      "pass": false
    }
  ],
  "failed": 1,
  "id": "00001754550000000000000",
  "passed": 1,
  "status": "failed",
  "total": 2
}

GET /v1/agents/{name}/golden-cases#

List an agent's golden cases. One page of the agent's golden set, the curated question → expectation cases the pre-deploy evaluation runs, in case-id order. **Paging.** Pass the next_cursor from the previous response back as ?cursor=. **An empty next_cursor is the only signal that you have seen everything.**

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
limit query integer Rows per page, 1 to 100 (default 50). Larger values are capped rather than refused.
cursor query string The next_cursor from the previous page. Omit for the first page.

Risposte

Codice Descrizione
200 Success.
400 limit is not a positive whole number (bad_request), or the cursor was not issued by this listing (invalid_cursor).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl "https://api.voiceland.ai/v1/agents/{name}/golden-cases" \
  -H "Authorization: Bearer VL_API_KEY"

Una risposta riuscita restituisce un GoldenCasesResponse.

Esempio di risposta

{
  "items": [
    {
      "agent_name": "front-desk",
      "authored_by": "maria@acme.example",
      "expected_answer": "The premium plan costs 12 euros per month.",
      "id": "gc_9d8c7b6a",
      "must_cite": true,
      "question": "How much is the premium plan?",
      "source": "operator",
      "tenant_slug": "acme"
    }
  ],
  "next_cursor": ""
}

POST /v1/agents/{name}/golden-cases#

Author a golden case. Adds one case to the agent's golden set. A case is a question plus **at least one expectation**, expected_answer (graded for meaning, not string equality), wrong_answer (an answer the agent must never repeat), or the behaviour flags must_cite, must_escalate, must_not_answer. A case with no expectation cannot fail and is refused. must_not_answer contradicts expected_answer and must_cite (a refusal answers nothing and cites nothing); the contradiction is refused rather than silently resolved.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.

Corpo della richiesta

application/json Schema: GoldenCaseRequest

Campo Tipo Obbligatorio Descrizione
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string
wrong_answer string

Risposte

Codice Descrizione
201 The stored case, with its server-minted id.
400 The case has no question, no expectation that could fail, or a contradictory expectation pair (invalid_golden_case).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/golden-cases" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "expected_answer": "The premium plan costs 12 euros per month.",
  "must_cite": true,
  "question": "How much is the premium plan?"
}'

Esempio di risposta

{
  "agent_name": "front-desk",
  "expected_answer": "The premium plan costs 12 euros per month.",
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?",
  "source": "operator",
  "tenant_slug": "acme"
}

POST /v1/agents/{name}/golden-cases/import-regressions#

Seed golden cases from confirmed-wrong reports. Imports the agent's confirmed-wrong inaccuracy reports (see GET /feedback/regressions) into its golden set. Each imported case carries the reported answer as wrong_answer, the expectation "never assert this again", plus the caller's question recovered from the call transcript, and must_cite when the wrong answer had been grounded. **Idempotent per report:** the case id is the report id, and a report whose case already exists is skipped, re-importing never duplicates a case or overwrites your edits. Reports for other agents, erased reports, and reports whose transcript no longer yields a question are counted in skipped.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.

Risposte

Codice Descrizione
200 What the import did.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/golden-cases/import-regressions" \
  -H "Authorization: Bearer VL_API_KEY"

Una risposta riuscita restituisce un GoldenImportResponse.

Esempio di risposta

{
  "imported": 2,
  "imported_ids": [
    "fbr_9d8c7b6a",
    "fbr_1a2b3c4d"
  ],
  "scanned": 3,
  "skipped": 1,
  "truncated": false
}

DELETE /v1/agents/{name}/golden-cases/{id}#

Delete a golden case. Removes one case from the set. Deleting is idempotent, an absent id answers 204 too. Past eval runs keep their own snapshot of every case they executed, so history is unaffected.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
id path string Case id.

Risposte

Codice Descrizione
204 Removed (or was never there).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl -X DELETE "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY"

GET /v1/agents/{name}/golden-cases/{id}#

Get a golden case. One case of the agent's golden set.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
id path string Case id.

Risposte

Codice Descrizione
200 Success.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no case with that id (golden_case_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY"

Esempio di risposta

{
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?",
  "source": "operator"
}

PUT /v1/agents/{name}/golden-cases/{id}#

Edit a golden case. Replaces the case's question and expectations. Provenance is preserved: a case imported from a confirmed-wrong report keeps its source and source_report_id whatever the edit says, so its origin cannot be laundered. The same at-least-one-expectation rule as create applies.

Parametri

Nome Posizione Tipo Obbligatorio Descrizione
name path string Agent name.
id path string Case id.

Corpo della richiesta

application/json Schema: GoldenCaseRequest

Campo Tipo Obbligatorio Descrizione
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string
wrong_answer string

Risposte

Codice Descrizione
200 The updated case.
400 The edit leaves the case unable to fail, or contradictory (invalid_golden_case).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no case with that id (golden_case_not_found).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl -X PUT "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "expected_answer": "12 euros per month.",
  "must_cite": true,
  "question": "How much is the premium plan?"
}'

Esempio di risposta

{
  "expected_answer": "12 euros per month.",
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?"
}

GET /v1/eval-policy#

Get the eval deploy gate. Your project's eval policy. The **deploy gate** is off (evals never affect deploys, the default), warn (a deploy over a failing latest eval succeeds but carries eval_warning), or block (a deploy over a failing latest eval is refused with 409 eval_failed naming the run). **Drift detection** is the second half: with drift_enabled, your golden set is re-run on a schedule AND whenever the models behind your calls change, and you are alerted when the score falls further than drift_pass_rate_drop_max below the previous run.

Risposte

Codice Descrizione
200 Success.
401 Missing or invalid API key.
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl "https://api.voiceland.ai/v1/eval-policy" \
  -H "Authorization: Bearer VL_API_KEY"

Una risposta riuscita restituisce un EvalPolicyResponse.

Esempio di risposta

{
  "deploy_gate": "off",
  "drift_enabled": false
}

PUT /v1/eval-policy#

Set the eval deploy gate and drift detection. Sets the deploy gate to off, warn or block. The gate acts only on a latest run whose status is failed: a run the harness could not fully execute (status error) never blocks a deploy, and an agent with no runs at all is not gated. drift_enabled turns on scheduled re-runs of every agent's golden set at drift_interval_hours (1..720, default 24), plus an extra re-run whenever the models behind your calls change. Each re-run is compared with the previous one for the same agent, and a fall of more than drift_pass_rate_drop_max (0..1, default 0.05, five points of the pass rate) alerts through your notification channels. It is a **drop**, not a floor: an assistant that has always scored 60% is not drifting. The alert is its OWN event type, eval.drift_regression, and not one of the kpi.* ones. If a subscription lists its event types explicitly, add eval.drift_regression to that list (or subscribe to "*"), or the re-runs will happen and nobody will be told. The first re-run of an agent is a baseline and alerts on nothing; an agent with no golden set is skipped, not failed. Drift is off by default because each re-run costs one assistant answer and one grading pass per golden case.

Corpo della richiesta

application/json Schema: EvalPolicyRequest

Campo Tipo Obbligatorio Descrizione
deploy_gate string
drift_enabled boolean
drift_interval_hours integer
drift_pass_rate_drop_max number

Risposte

Codice Descrizione
200 The stored policy.
401 Missing or invalid API key.
422 deploy_gate is not off, warn or block, drift_interval_hours is outside 1..720, or drift_pass_rate_drop_max is outside 0..1 (invalid_policy).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Esempio di richiesta

curl -X PUT "https://api.voiceland.ai/v1/eval-policy" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "deploy_gate": "block",
  "drift_enabled": true,
  "drift_interval_hours": 24,
  "drift_pass_rate_drop_max": 0.05
}'

Esempio di risposta

{
  "deploy_gate": "block",
  "drift_enabled": true,
  "drift_interval_hours": 24,
  "drift_pass_rate_drop_max": 0.05
}

Schemi#

Error#

Error envelope returned for non-2xx responses.

Campo Tipo Obbligatorio Descrizione
error object

EvalPolicyRequest#

Campo Tipo Obbligatorio Descrizione
deploy_gate string
drift_enabled boolean
drift_interval_hours integer
drift_pass_rate_drop_max number

EvalPolicyResponse#

Campo Tipo Obbligatorio Descrizione
deploy_gate string
drift_enabled boolean
drift_interval_hours integer
drift_pass_rate_drop_max number

EvalRunRequest#

Campo Tipo Obbligatorio Descrizione
triggered_by string
version_id string

EvalRunSummary#

Campo Tipo Obbligatorio Descrizione
agent_name string
agent_version_id string
created_at string (date-time)
errored integer
failed integer
id string
judge_model string
pass_rate number
passed integer
status string
total integer
triggered_by string

EvalRunsResponse#

Campo Tipo Obbligatorio Descrizione
items array of EvalRunSummary
next_cursor string

GoldenCase#

Campo Tipo Obbligatorio Descrizione
agent_name string
authored_by string
call_id_hash string
created_at string (date-time)
erased boolean
expected_answer string
id string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string
source string
source_report_id string
tenant_slug string
updated_at string (date-time)
wrong_answer string

GoldenCaseRequest#

Campo Tipo Obbligatorio Descrizione
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string
wrong_answer string

GoldenCasesResponse#

Campo Tipo Obbligatorio Descrizione
items array of GoldenCase
next_cursor string

GoldenImportResponse#

Campo Tipo Obbligatorio Descrizione
imported integer
imported_ids array of string
scanned integer
skipped integer
truncated boolean

La console

Queste pagine sono di sola lettura. La chiamata di prova, le chiavi API e il riferimento API aggiornato si trovano nella console, dove il suo account è connesso.

Apri la console