API reference / Evaluation

Evaluation

The Evaluation endpoints of the Voiceland AI API, with parameters, schemas and curl examples.

Last updated:

GET /v1/agents/{name}/evals#

List eval runs. One page of the agent's eval runs, **newest first**, without the per-case matrix (fetch a run for the full detail). ?version_id= narrows to the runs of one agent version, it is applied after the read, so short pages are normal. **Paging.** Pass the next_cursor back as ?cursor=; **an empty next_cursor is the only signal that you have seen everything.**

Parameters

Name In Type Required Description
name path string Yes Agent name.
version_id query string Only runs of this agent version. Applied after the read, short pages are normal.
limit query integer Rows per page, 1 to 100 (default 50).
cursor query string The next_cursor from the previous page. Omit for the first page.

Responses

Code Description
200 Success.
400 limit is not a positive whole number (bad_request), or the cursor was not issued by this listing (invalid_cursor).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl "https://api.voiceland.ai/v1/agents/{name}/evals" \
  -H "Authorization: Bearer VL_API_KEY"

A successful response returns a EvalRunsResponse.

Example response

{
  "items": [
    {
      "agent_name": "front-desk",
      "agent_version_id": "00001754540000000000000",
      "failed": 1,
      "id": "00001754550000000000000",
      "pass_rate": 0.5,
      "passed": 1,
      "status": "failed",
      "total": 2
    }
  ],
  "next_cursor": ""
}

POST /v1/agents/{name}/evals#

Run the agent's golden set (pre-deploy eval). Executes every golden case against the agent **draft** over the text path, the question is answered with the agent's own system prompt, model and knowledge retrieval, no live call involved, and grades each answer with the platform's LLM judge. The graded run is stored and returned: the per-case pass/fail matrix, the aggregate score, and the diff against the agent's previous run. version_id pins a specific version snapshot; omitted, the latest snapshot (your current draft) runs. The run is what the eval deploy gate reads, see PUT /eval-policy. The call is synchronous and bounded: a golden set larger than the per-run case limit is refused rather than partially graded.

Parameters

Name In Type Required Description
name path string Yes Agent name.

Request body

application/json Schema: EvalRunRequest

Field Type Required Description
triggered_by string
version_id string

Responses

Code Description
201 The stored run, matrix included.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found), or version_id names no retained snapshot (version_not_found).
409 The agent has no golden cases (golden_set_empty), or the set exceeds the per-run case limit (golden_set_too_large).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
503 No eval judge is configured on this deployment (eval_judge_unavailable).
5XX Server or upstream error.

Example request

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/evals" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "triggered_by": "maria@acme.example",
  "version_id": ""
}'

Example response

{
  "agent_name": "front-desk",
  "agent_version_id": "00001754540000000000000",
  "cases": [
    {
      "answer": "The premium plan costs 10 euros.",
      "case_id": "gc_9d8c7b6a",
      "fail_reasons": [
        "does_not_match_expected_answer"
      ],
      "pass": false,
      "question": "How much is the premium plan?"
    }
  ],
  "diff": {
    "newly_failing": [
      "gc_9d8c7b6a"
    ],
    "previous_run_id": "00001754530000000000000"
  },
  "failed": 1,
  "id": "00001754550000000000000",
  "pass_rate": 0.5,
  "passed": 1,
  "status": "failed",
  "tenant_slug": "acme",
  "total": 2
}

GET /v1/agents/{name}/evals/{run_id}#

Get an eval run. One run in full: the per-case pass/fail matrix (each case's question, the draft answer, its sources, the judge's verdict and the fail reasons), the aggregate score, and the diff against the previous run.

Parameters

Name In Type Required Description
name path string Yes Agent name.
run_id path string Yes Run id.

Responses

Code Description
200 Success.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no run with that id (run_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl "https://api.voiceland.ai/v1/agents/{name}/evals/{run_id}" \
  -H "Authorization: Bearer VL_API_KEY"

Example response

{
  "cases": [
    {
      "case_id": "gc_9d8c7b6a",
      "fail_reasons": [
        "not_grounded"
      ],
      "pass": false
    }
  ],
  "failed": 1,
  "id": "00001754550000000000000",
  "passed": 1,
  "status": "failed",
  "total": 2
}

GET /v1/agents/{name}/golden-cases#

List an agent's golden cases. One page of the agent's golden set, the curated question → expectation cases the pre-deploy evaluation runs, in case-id order. **Paging.** Pass the next_cursor from the previous response back as ?cursor=. **An empty next_cursor is the only signal that you have seen everything.**

Parameters

Name In Type Required Description
name path string Yes Agent name.
limit query integer Rows per page, 1 to 100 (default 50). Larger values are capped rather than refused.
cursor query string The next_cursor from the previous page. Omit for the first page.

Responses

Code Description
200 Success.
400 limit is not a positive whole number (bad_request), or the cursor was not issued by this listing (invalid_cursor).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl "https://api.voiceland.ai/v1/agents/{name}/golden-cases" \
  -H "Authorization: Bearer VL_API_KEY"

A successful response returns a GoldenCasesResponse.

Example response

{
  "items": [
    {
      "agent_name": "front-desk",
      "authored_by": "maria@acme.example",
      "expected_answer": "The premium plan costs 12 euros per month.",
      "id": "gc_9d8c7b6a",
      "must_cite": true,
      "question": "How much is the premium plan?",
      "source": "operator",
      "tenant_slug": "acme"
    }
  ],
  "next_cursor": ""
}

POST /v1/agents/{name}/golden-cases#

Author a golden case. Adds one case to the agent's golden set. A case is a question plus **at least one expectation**, expected_answer (graded for meaning, not string equality), wrong_answer (an answer the agent must never repeat), or the behaviour flags must_cite, must_escalate, must_not_answer. A case with no expectation cannot fail and is refused. must_not_answer contradicts expected_answer and must_cite (a refusal answers nothing and cites nothing); the contradiction is refused rather than silently resolved.

Parameters

Name In Type Required Description
name path string Yes Agent name.

Request body

application/json Schema: GoldenCaseRequest

Field Type Required Description
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string Yes
wrong_answer string

Responses

Code Description
201 The stored case, with its server-minted id.
400 The case has no question, no expectation that could fail, or a contradictory expectation pair (invalid_golden_case).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/golden-cases" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "expected_answer": "The premium plan costs 12 euros per month.",
  "must_cite": true,
  "question": "How much is the premium plan?"
}'

Example response

{
  "agent_name": "front-desk",
  "expected_answer": "The premium plan costs 12 euros per month.",
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?",
  "source": "operator",
  "tenant_slug": "acme"
}

POST /v1/agents/{name}/golden-cases/import-regressions#

Seed golden cases from confirmed-wrong reports. Imports the agent's confirmed-wrong inaccuracy reports (see GET /feedback/regressions) into its golden set. Each imported case carries the reported answer as wrong_answer, the expectation "never assert this again", plus the caller's question recovered from the call transcript, and must_cite when the wrong answer had been grounded. **Idempotent per report:** the case id is the report id, and a report whose case already exists is skipped, re-importing never duplicates a case or overwrites your edits. Reports for other agents, erased reports, and reports whose transcript no longer yields a question are counted in skipped.

Parameters

Name In Type Required Description
name path string Yes Agent name.

Responses

Code Description
200 What the import did.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl -X POST "https://api.voiceland.ai/v1/agents/{name}/golden-cases/import-regressions" \
  -H "Authorization: Bearer VL_API_KEY"

A successful response returns a GoldenImportResponse.

Example response

{
  "imported": 2,
  "imported_ids": [
    "fbr_9d8c7b6a",
    "fbr_1a2b3c4d"
  ],
  "scanned": 3,
  "skipped": 1,
  "truncated": false
}

DELETE /v1/agents/{name}/golden-cases/{id}#

Delete a golden case. Removes one case from the set. Deleting is idempotent, an absent id answers 204 too. Past eval runs keep their own snapshot of every case they executed, so history is unaffected.

Parameters

Name In Type Required Description
name path string Yes Agent name.
id path string Yes Case id.

Responses

Code Description
204 Removed (or was never there).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl -X DELETE "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY"

GET /v1/agents/{name}/golden-cases/{id}#

Get a golden case. One case of the agent's golden set.

Parameters

Name In Type Required Description
name path string Yes Agent name.
id path string Yes Case id.

Responses

Code Description
200 Success.
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no case with that id (golden_case_not_found).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY"

Example response

{
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?",
  "source": "operator"
}

PUT /v1/agents/{name}/golden-cases/{id}#

Edit a golden case. Replaces the case's question and expectations. Provenance is preserved: a case imported from a confirmed-wrong report keeps its source and source_report_id whatever the edit says, so its origin cannot be laundered. The same at-least-one-expectation rule as create applies.

Parameters

Name In Type Required Description
name path string Yes Agent name.
id path string Yes Case id.

Request body

application/json Schema: GoldenCaseRequest

Field Type Required Description
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string Yes
wrong_answer string

Responses

Code Description
200 The updated case.
400 The edit leaves the case unable to fail, or contradictory (invalid_golden_case).
401 Missing or invalid API key.
404 No agent with that name (agent_not_found) or no case with that id (golden_case_not_found).
422 The body is not valid JSON for this shape (invalid_request).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl -X PUT "https://api.voiceland.ai/v1/agents/{name}/golden-cases/{id}" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "expected_answer": "12 euros per month.",
  "must_cite": true,
  "question": "How much is the premium plan?"
}'

Example response

{
  "expected_answer": "12 euros per month.",
  "id": "gc_9d8c7b6a",
  "must_cite": true,
  "question": "How much is the premium plan?"
}

GET /v1/eval-policy#

Get the eval deploy gate. Your project's eval policy. The **deploy gate** is off (evals never affect deploys, the default), warn (a deploy over a failing latest eval succeeds but carries eval_warning), or block (a deploy over a failing latest eval is refused with 409 eval_failed naming the run). **Drift detection** is the second half: with drift_enabled, your golden set is re-run on a schedule AND whenever the models behind your calls change, and you are alerted when the score falls further than drift_pass_rate_drop_max below the previous run.

Responses

Code Description
200 Success.
401 Missing or invalid API key.
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl "https://api.voiceland.ai/v1/eval-policy" \
  -H "Authorization: Bearer VL_API_KEY"

A successful response returns a EvalPolicyResponse.

Example response

{
  "deploy_gate": "off",
  "drift_enabled": false
}

PUT /v1/eval-policy#

Set the eval deploy gate and drift detection. Sets the deploy gate to off, warn or block. The gate acts only on a latest run whose status is failed: a run the harness could not fully execute (status error) never blocks a deploy, and an agent with no runs at all is not gated. drift_enabled turns on scheduled re-runs of every agent's golden set at drift_interval_hours (1..720, default 24), plus an extra re-run whenever the models behind your calls change. Each re-run is compared with the previous one for the same agent, and a fall of more than drift_pass_rate_drop_max (0..1, default 0.05, five points of the pass rate) alerts through your notification channels. It is a **drop**, not a floor: an assistant that has always scored 60% is not drifting. The alert is its OWN event type, eval.drift_regression, and not one of the kpi.* ones. If a subscription lists its event types explicitly, add eval.drift_regression to that list (or subscribe to "*"), or the re-runs will happen and nobody will be told. The first re-run of an agent is a baseline and alerts on nothing; an agent with no golden set is skipped, not failed. Drift is off by default because each re-run costs one assistant answer and one grading pass per golden case.

Request body

application/json Schema: EvalPolicyRequest

Field Type Required Description
deploy_gate string Yes
drift_enabled boolean
drift_interval_hours integer
drift_pass_rate_drop_max number

Responses

Code Description
200 The stored policy.
401 Missing or invalid API key.
422 deploy_gate is not off, warn or block, drift_interval_hours is outside 1..720, or drift_pass_rate_drop_max is outside 0..1 (invalid_policy).
4XX Request error (validation, not-found, etc.).
5XX Server or upstream error.

Example request

curl -X PUT "https://api.voiceland.ai/v1/eval-policy" \
  -H "Authorization: Bearer VL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "deploy_gate": "block",
  "drift_enabled": true,
  "drift_interval_hours": 24,
  "drift_pass_rate_drop_max": 0.05
}'

Example response

{
  "deploy_gate": "block",
  "drift_enabled": true,
  "drift_interval_hours": 24,
  "drift_pass_rate_drop_max": 0.05
}

Schemas#

Error#

Error envelope returned for non-2xx responses.

Field Type Required Description
error object Yes

EvalPolicyRequest#

Field Type Required Description
deploy_gate string Yes
drift_enabled boolean
drift_interval_hours integer
drift_pass_rate_drop_max number

EvalPolicyResponse#

Field Type Required Description
deploy_gate string Yes
drift_enabled boolean Yes
drift_interval_hours integer
drift_pass_rate_drop_max number

EvalRunRequest#

Field Type Required Description
triggered_by string
version_id string

EvalRunSummary#

Field Type Required Description
agent_name string Yes
agent_version_id string
created_at string (date-time) Yes
errored integer
failed integer Yes
id string Yes
judge_model string
pass_rate number Yes
passed integer Yes
status string Yes
total integer Yes
triggered_by string

EvalRunsResponse#

Field Type Required Description
items array of EvalRunSummary Yes
next_cursor string Yes

GoldenCase#

Field Type Required Description
agent_name string Yes
authored_by string
call_id_hash string
created_at string (date-time)
erased boolean
expected_answer string
id string Yes
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string Yes
source string Yes
source_report_id string
tenant_slug string Yes
updated_at string (date-time)
wrong_answer string

GoldenCaseRequest#

Field Type Required Description
authored_by string
expected_answer string
must_cite boolean
must_escalate boolean
must_not_answer boolean
question string Yes
wrong_answer string

GoldenCasesResponse#

Field Type Required Description
items array of GoldenCase Yes
next_cursor string Yes

GoldenImportResponse#

Field Type Required Description
imported integer Yes
imported_ids array of string
scanned integer Yes
skipped integer Yes
truncated boolean Yes

The console

These pages are read only. The test call, the API keys and the live API reference are in the console, where your account is signed in.

Open the console