Skip to main content
An evaluation run groups the executions of one test set. Each sample is a separate agent run, so you can inspect its trace and compare the results of successive evaluations. The Python SDK records the samples and their fields; your code supplies the test cases and verdicts.

Record an evaluation

Connect to Sovara and install the Python SDK with uv add sovara (or your Python package manager). Save this example as evaluate.py and run it with uv run python evaluate.py:
The exact-match check is only an example evaluator. Sovara does not run a judge when you log llm_judge_is_correct; use your own rule or model to produce the verdict. Log groundtruth while the sample run is active so its run analysis can use it. The test_set value is a custom field on each sample and the evaluation run; dataset_version is another evaluation-level field. You can log other string, boolean, or numeric fields the same way. Sovara derives sample counts from the recorded runs, so you do not need to log them. Open the project in Sovara and choose Evaluations to see the new evaluation run. Select it to inspect each sample and open its trace. The displayed accuracy is calculated from samples with a boolean verdict; samples without one are excluded. You can log an accuracy fraction on the evaluation run to override that calculated value, and log other aggregate fields with client.log(eval_run_id=eval_run_id, **fields). Run the script again after changing your agent to create a separate evaluation run for comparison. See the Python SDK API reference for the exact run(), create_eval_run(), and log() contracts.