> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sovara-labs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluations

> Run a set of agent samples, record verdicts, and compare evaluation runs.

An evaluation run groups the executions of one test set. Each sample is a
separate agent run, so you can inspect its trace and compare the results of
successive evaluations. The Python SDK records the samples and their fields;
your code supplies the test cases and verdicts.

## Record an evaluation

[Connect to Sovara](/get-started/installation) and install the Python SDK with
`uv add sovara` (or your Python package manager). Save this example as
`evaluate.py` and run it with `uv run python evaluate.py`:

```python theme={"theme":{"light":"github-light","dark":"github-dark"}}
from sovara import SovaraClient, trace

client = SovaraClient(project_name="capital-answers")
cases = [
    ("What is the capital of France?", "Paris"),
    ("What is the capital of Italy?", "Rome"),
]


@trace
def answer(question: str) -> str:
    # Replace this example with your agent. Supported model calls are traced too.
    return {
        "What is the capital of France?": "Paris",
        "What is the capital of Italy?": "Rome",
    }[question]


eval_run_id = client.create_eval_run()
for question, expected in cases:
    with client.run("capital-answer", eval_run_id=eval_run_id) as run_key:
        if run_key is None:
            raise RuntimeError("Sovara could not record this sample")

        client.log(run_key=run_key, run_input=question, groundtruth=expected)
        output = answer(question)
        correct = output == expected
        client.log(
            run_key=run_key,
            run_output=output,
            llm_judge_is_correct=correct,
            llm_judge_output="Exact match" if correct else "Does not match",
            test_set="capital-questions",
        )

client.log(eval_run_id=eval_run_id, test_set="capital-questions", dataset_version="v1")
print(f"Recorded evaluation {eval_run_id}")
```

The exact-match check is only an example evaluator. Sovara does not run a judge
when you log `llm_judge_is_correct`; use your own rule or model to produce the
verdict. Log `groundtruth` while the sample run is active so its run analysis can
use it. The `test_set` value is a custom field on each sample and the evaluation
run; `dataset_version` is another evaluation-level field. You can log other
string, boolean, or numeric fields the same way. Sovara derives sample counts
from the recorded runs, so you do not need to log them.

Open the project in Sovara and choose **Evaluations** to see the new evaluation
run. Select it to inspect each sample and open its trace. The displayed accuracy
is calculated from samples with a boolean verdict; samples without one are
excluded. You can log an `accuracy` fraction on the evaluation run to override
that calculated value, and log other aggregate fields with
`client.log(eval_run_id=eval_run_id, **fields)`.

Run the script again after changing your agent to create a separate evaluation
run for comparison. See the [Python SDK API reference](/sdks/python/api-reference)
for the exact `run()`, `create_eval_run()`, and `log()` contracts.
