> For the complete documentation index, see [llms.txt](/llms.txt).
> A full single-fetch corpus is available at [llms-full.txt](/llms-full.txt).
---
title: Batch eval
description: Run a component against multiple inputs in parallel, score every output, and inspect aggregate results using the Python, TypeScript, or Go SDK.
last_verified: 2026-09-29
---



**Batch eval** lets you evaluate a component against a list of inputs in a single call. The SDK runs each input in parallel, scores every output, and returns a single result object with per-item scores and aggregate statistics. Use batch eval for local regression testing, CI pipelines that don't need the full [Experiments](/docs/improve/experiments.md) workflow, or quick quality checks during development.

You need a running AGNT5 worker and at least one scorer. Point the client at your worker by setting `AGNT5_GATEWAY_URL` in your environment — `Client()`, `AsyncClient()`, and Go's `agnt5.NewClient("")` pick it up automatically and fall back to `https://gw.agnt5.com` when the variable is not set.

```bash
# local dev
export AGNT5_GATEWAY_URL=http://localhost:34181
```

## Run a single eval



**Python:**

`client.eval()` evaluates one input and returns a scored result:

```python
from agnt5 import Client
from agnt5.eval import Correctness

client = Client()

result = client.eval(
    component="support_agent",
    component_type="agent",
    input_data={"message": "Where is my order #1234?"},
    expected="Your order #1234 is in transit",
    scorers=[Correctness()],
)

print(result.passed)           # True / False
print(result.output)           # Component output
for score in result.scores:
    print(score.scorer, score.score, score.passed, score.explanation)
```





**TypeScript:**

`client.eval()` evaluates one input and returns a scored result:

```typescript
import { Client, Correctness } from '@agnt5/sdk';

const client = new Client();

const result = await client.eval(
  'support_agent',
  { message: 'Where is my order #1234?' },
  {
    componentType: 'agent',
    expected: 'Your order #1234 is in transit',
    scorers: [new Correctness()],
  },
);

console.log(result.passed);   // true / false
console.log(result.output);   // Component output
for (const score of result.scores) {
  console.log(score.scorer, score.score, score.passed, score.explanation);
}
```





**Go:**

`client.Eval(ctx, request)` evaluates one input and returns a scored result:

```go
import "github.com/agnt5dev/sdk-go/agnt5"

client, err := agnt5.NewClient("") // "" falls back to AGNT5_GATEWAY_URL

response, err := client.Eval(ctx, agnt5.EvalRequest{
    Component:     "support_agent",
    ComponentType: agnt5.ComponentTypeAgent,
    Input:         map[string]any{"message": "Where is my order #1234?"},
    Expected:      "Your order #1234 is in transit",
    Scorers:       []agnt5.EvalScorerSpec{agnt5.Correctness{}.ToEvalScorerSpec()},
})

fmt.Println(response.Passed) // true / false
fmt.Println(string(response.Output))
for _, score := range response.Scores {
    fmt.Println(score.Scorer, score.Score, score.Passed, score.Explanation)
}
```

`Scorers` takes `[]agnt5.EvalScorerSpec` — `{Name: "exact_match"}` for a built-in by name (the same shape as the CLI's `--builtin-scorer` flag), or the `ToEvalScorerSpec()` output of a typed preset such as `agnt5.Correctness{}` (see [Scorers](#scorers) below).



## Run a batch



**Python:**

`client.batch_eval()` accepts a list of inputs and runs them in parallel:

```python
from agnt5 import Client, BatchEvalItem

client = Client()

result = client.batch_eval(
    component="support_agent",
    component_type="agent",
    items=[
        BatchEvalItem(
            input={"message": "Where is my order #1234?"},
            expected="Your order #1234 is in transit",
            item_id="order-status",
        ),
        BatchEvalItem(
            input={"message": "I want to cancel order #5678"},
            expected="Order #5678 has been cancelled",
            item_id="order-cancel",
        ),
    ],
    scorers=["exact_match"],
)

print(f"Pass rate: {result.pass_rate:.0%}")
for item in result.results:
    status = "PASS" if item.passed else "FAIL"
    print(f"{item.item_id}: {status} ({item.duration_ms}ms)")
```





**TypeScript:**

`client.batchEval()` accepts a list of inputs and runs them in parallel. `BatchEvalItem` is a plain object shape, not a class — `{ input, expected?, itemId? }`.

```typescript
import { Client, BatchEvalItem } from '@agnt5/sdk';

const client = new Client();

const items: BatchEvalItem[] = [
  { input: { message: 'Where is my order #1234?' }, expected: 'Your order #1234 is in transit', itemId: 'order-status' },
  { input: { message: 'I want to cancel order #5678' }, expected: 'Order #5678 has been cancelled', itemId: 'order-cancel' },
];

const result = await client.batchEval('support_agent', items, {
  componentType: 'agent',
  scorers: ['exact_match'],
});

console.log(`Pass rate: ${(result.passRate * 100).toFixed(0)}%`);
for (const item of result.results) {
  const status = item.passed ? 'PASS' : 'FAIL';
  console.log(`${item.itemId}: ${status} (${item.durationMs}ms)`);
}
```





**Go:**

`client.BatchEval(ctx, component, items, options)` accepts a slice of `agnt5.BatchEvalItem` and runs them in parallel. It returns a `*agnt5.BatchEvalResult` and no error: an item that fails to evaluate is reported in its own `Error` field (see [Handling partial failures](#handling-partial-failures)).

```go
import "github.com/agnt5dev/sdk-go/agnt5"

client, err := agnt5.NewClient("")

result := client.BatchEval(ctx, "support_agent", []agnt5.BatchEvalItem{
    {
        Input:    map[string]any{"message": "Where is my order #1234?"},
        Expected: "Your order #1234 is in transit",
        ItemID:   "order-status",
    },
    {
        Input:    map[string]any{"message": "I want to cancel order #5678"},
        Expected: "Order #5678 has been cancelled",
        ItemID:   "order-cancel",
    },
}, agnt5.BatchEvalOptions{
    ComponentType: agnt5.ComponentTypeAgent,
    Scorers:       []agnt5.EvalScorerSpec{{Name: "exact_match"}},
})

fmt.Printf("Pass rate: %.0f%%\n", result.PassRate()*100)
for _, item := range result.Results {
    status := "FAIL"
    if item.Passed {
        status = "PASS"
    }
    fmt.Printf("%s: %s (%dms)\n", item.ItemID, status, item.DurationMS)
}
```



## Input formats



**Python:**

`batch_eval()` accepts items in several forms — mix them freely:

**Plain dicts with a separate `expected` list:**

```python
result = client.batch_eval(
    component="greet",
    items=[{"name": "Alice"}, {"name": "Bob"}],
    expected=["Hello, Alice!", "Hello, Bob!"],
    scorers=["exact_match"],
)
```

**Dicts with `input` and `expected` keys:**

```python
result = client.batch_eval(
    component="add",
    items=[
        {"input": {"a": 1, "b": 2}, "expected": 3},
        {"input": {"a": 3, "b": 4}, "expected": 7, "item_id": "add-2"},
    ],
    scorers=["exact_match"],
)
```

**`BatchEvalItem` objects for full control:**

```python
from agnt5 import BatchEvalItem

result = client.batch_eval(
    component="add",
    items=[
        BatchEvalItem(input={"a": 1, "b": 2}, expected=3, item_id="add-1"),
        BatchEvalItem(input={"a": 3, "b": 4}, expected=7, item_id="add-2"),
    ],
    scorers=["exact_match"],
)
```





**TypeScript:**

`batchEval()` items are always `{ input, expected?, itemId? }` object literals — there's no separate positional-`expected`-list form like Python's plain-dict shorthand.

```typescript
const result = await client.batchEval(
  'add',
  [
    { input: { a: 1, b: 2 }, expected: 3, itemId: 'add-1' },
    { input: { a: 3, b: 4 }, expected: 7, itemId: 'add-2' },
  ],
  { scorers: ['exact_match'] },
);
```





**Go:**

`BatchEval` always takes `[]agnt5.BatchEvalItem` — `Input` is a `map[string]any`; `Expected` and `ItemID` are optional. For the positional-`expected` form, build the slice with `agnt5.NormalizeBatchEvalItems(inputs, expected)`, or leave `Expected` nil on the items and pass the list as `BatchEvalOptions.Expected`, which is matched by index.

```go
// Plain input maps with a separate expected list
items := agnt5.NormalizeBatchEvalItems(
    []map[string]any{{"name": "Alice"}, {"name": "Bob"}},
    []any{"Hello, Alice!", "Hello, Bob!"},
)
result := client.BatchEval(ctx, "greet", items, agnt5.BatchEvalOptions{
    Scorers: []agnt5.EvalScorerSpec{{Name: "exact_match"}},
})

// BatchEvalItem values for full control
result = client.BatchEval(ctx, "add", []agnt5.BatchEvalItem{
    {Input: map[string]any{"a": 1, "b": 2}, Expected: 3, ItemID: "add-1"},
    {Input: map[string]any{"a": 3, "b": 4}, Expected: 7, ItemID: "add-2"},
}, agnt5.BatchEvalOptions{
    Scorers: []agnt5.EvalScorerSpec{{Name: "exact_match"}},
})
```



## Scorers



**Python:**

Pass scorer names, SDK preset classes, or `LLMJudge` instances. Combine them freely:

```python
from agnt5.eval import Correctness, Helpfulness, LLMJudge

scorers = [
    "json_valid",          # Fast structure check
    "contains",            # Required substring
    Correctness(),         # Managed correctness preset
    Helpfulness(model="openai/gpt-4o"),
    LLMJudge(
        criteria="Is the response under 50 words?",
        model="openai/gpt-4o-mini",
    ),
]
```

All [built-in deterministic scorers](/docs/improve/scorers.md#built-in-deterministic-scorers) work as strings. All [SDK evaluator presets](/docs/improve/scorers.md#sdk-evaluator-presets) work as class instances.





**TypeScript:**

Pass scorer names, SDK preset classes, or `LLMJudge` instances. Combine them freely:

```typescript
import { Correctness, Helpfulness, LLMJudge } from '@agnt5/sdk';

const scorers = [
  'json_valid',          // Fast structure check
  'contains',            // Required substring
  new Correctness(),     // Managed correctness preset
  new Helpfulness({ model: 'openai/gpt-4o' }),
  new LLMJudge({
    criteria: 'Is the response under 50 words?',
    model: 'openai/gpt-4o-mini',
  }),
];
```

All [built-in deterministic scorers](/docs/improve/scorers.md#built-in-deterministic-scorers) work as strings. All [SDK evaluator presets](/docs/improve/scorers.md#sdk-evaluator-presets) work as class instances.





**Go:**

`BatchEvalOptions.Scorers` and `EvalRequest.Scorers` take `[]agnt5.EvalScorerSpec`. Build the specs from scorer names, typed presets, or `agnt5.NewLLMJudge` — every typed scorer has a `ToEvalScorerSpec()` method — or hand a mix of all three to `agnt5.NormalizeEvalScorers`:

```go
helpfulness := agnt5.Helpfulness{}
helpfulness.Model = "openai/gpt-4o"

scorers := agnt5.NormalizeEvalScorers(
    "json_valid",        // Fast structure check
    "contains",          // Required substring
    agnt5.Correctness{}, // Managed correctness preset
    helpfulness,
    agnt5.NewLLMJudge(agnt5.LLMJudgeConfig{
        Criteria: "Is the response under 50 words?",
        Model:    "openai/gpt-4o-mini",
    }),
)
```

All [built-in deterministic scorers](/docs/improve/scorers.md#built-in-deterministic-scorers) work as strings. All [SDK evaluator presets](/docs/improve/scorers.md#sdk-evaluator-presets) work as struct values (`agnt5.Correctness{}`, `agnt5.Helpfulness{}`, ...).

Use `NewLLMJudge` or a preset rather than a hand-written `{Name: "llm_judge", Config: map[string]any{"model": "openai/gpt-4o-mini"}}` spec: the typed helpers split `provider/model` into the separate `provider` and `model` config keys the gateway expects, while a raw config sends the string verbatim and the gateway rejects it with a 400. If you do write a raw config, pass a bare model name (`"gpt-4o-mini"`) and a separate `"provider"` key.



## Concurrency and timeouts



**Python:**

```python
result = client.batch_eval(
    component="slow_agent",
    component_type="agent",
    items=test_items,
    scorers=[Correctness()],
    max_concurrency=5,   # Parallel evaluations (default 10)
    timeout=60.0,        # Per-item timeout in seconds
)
```

For large batches, start with `max_concurrency=3` and `timeout=30.0` during development, then increase for production runs.





**TypeScript:**

```typescript
const result = await client.batchEval('slow_agent', testItems, {
  componentType: 'agent',
  scorers: [new Correctness()],
  maxConcurrency: 5, // Parallel evaluations (default 10)
  timeout: 60_000,    // Per-item timeout in milliseconds
});
```

For large batches, start with `maxConcurrency: 3` and `timeout: 30_000` during development, then increase for production runs.





**Go:**

```go
result := client.BatchEval(ctx, "slow_agent", testItems, agnt5.BatchEvalOptions{
    ComponentType:  agnt5.ComponentTypeAgent,
    Scorers:        []agnt5.EvalScorerSpec{agnt5.Correctness{}.ToEvalScorerSpec()},
    MaxConcurrency: 5,                // Parallel evaluations (default 10)
    Timeout:        60 * time.Second, // Per-item timeout
})
```

For large batches, start with `MaxConcurrency: 3` and `Timeout: 30 * time.Second` during development, then increase for production runs. The `ctx` you pass is used for every item's request, so its deadline or cancellation bounds the whole batch.



## Async client



**Python:**

```python
import asyncio
from agnt5 import AsyncClient, BatchEvalItem

async def run_evals():
    async with AsyncClient() as client:
        return await client.batch_eval(
            component="analyze",
            items=[
                BatchEvalItem(input={"text": "Hello"}, expected="greeting"),
                BatchEvalItem(input={"text": "Goodbye"}, expected="farewell"),
            ],
            scorers=["exact_match"],
            max_concurrency=10,
        )

result = asyncio.run(run_evals())
```





**TypeScript:**

The TypeScript `Client` is async by default — there's no separate `AsyncClient`.

```typescript
import { Client, BatchEvalItem } from '@agnt5/sdk';

async function runEvals() {
  const client = new Client();
  const items: BatchEvalItem[] = [
    { input: { text: 'Hello' }, expected: 'greeting' },
    { input: { text: 'Goodbye' }, expected: 'farewell' },
  ];
  return client.batchEval('analyze', items, {
    scorers: ['exact_match'],
    maxConcurrency: 10,
  });
}

const result = await runEvals();
```





**Go:**

There's no separate async client in Go. `BatchEval` takes a `context.Context`, blocks until every item has finished, and returns the complete `*BatchEvalResult` — run it in a goroutine if you have other work to do meanwhile.

```go
func runEvals(ctx context.Context) (*agnt5.BatchEvalResult, error) {
    client, err := agnt5.NewClient("")
    if err != nil {
        return nil, err
    }
    return client.BatchEval(ctx, "analyze", []agnt5.BatchEvalItem{
        {Input: map[string]any{"text": "Hello"}, Expected: "greeting"},
        {Input: map[string]any{"text": "Goodbye"}, Expected: "farewell"},
    }, agnt5.BatchEvalOptions{
        Scorers:        []agnt5.EvalScorerSpec{{Name: "exact_match"}},
        MaxConcurrency: 10,
    }), nil
}
```



## Reading results

### Aggregate result



**Python:**

```python
result.batch_id        # Unique batch identifier
result.status          # "completed", "partial_failure", or "failed"
result.pass_rate       # Float 0.0–1.0 (passed_items / total_items)
result.stats           # BatchEvalStats object

result.passing_items() # Items where passed=True
result.failing_items() # Items where passed=False (not errors)
result.failed_items()  # Items with evaluation errors
```





**TypeScript:**

```typescript
result.batchId;        // Unique batch identifier
result.status;         // "completed", "partial_failure", or "failed"
result.passRate;       // Number 0.0–1.0 (passed items / total items)
result.stats;          // BatchEvalStats object

result.passingItems(); // Items where passed=true
result.failingItems(); // Items where passed=false (not errors)
result.failedItems();  // Items with evaluation errors
```





**Go:**

```go
fmt.Println(result.BatchID)        // Unique batch identifier
fmt.Println(result.Status)         // "completed", "partial_failure", or "failed"
fmt.Println(result.PassRate())     // float64 0.0–1.0 (passed items / total items)
fmt.Println(result.Stats)          // BatchEvalStats value

fmt.Println(result.PassingItems()) // Items where Passed is true
fmt.Println(result.FailingItems()) // Items where Passed is false (not errors)
fmt.Println(result.FailedItems())  // Items with evaluation errors
```

`result.IsSuccess()` and `result.IsPartialFailure()` test `Status` for you.



### Per-item results



**Python:**

```python
for item in result.results:
    item.index          # Position in batch
    item.item_id        # Custom identifier (if provided)
    item.run_id         # Platform run ID
    item.output         # Component output
    item.passed         # True if all scorers passed
    item.scores         # List of ScorerResultSummary
    item.duration_ms    # Execution time
    item.trace_id       # OpenTelemetry trace ID
    item.error          # Error message (if evaluation failed)

    score = item.get_score("json_valid")  # Get a specific scorer result
```





**TypeScript:**

```typescript
for (const item of result.results) {
  item.index;       // Position in batch
  item.itemId;      // Custom identifier (if provided)
  item.runId;       // Platform run ID
  item.output;      // Component output
  item.passed;      // true if all scorers passed
  item.scores;      // Array of ScorerResultSummary
  item.durationMs;  // Execution time
  item.traceId;     // OpenTelemetry trace ID
  item.error;       // Error message (if evaluation failed)

  const score = item.getScore('json_valid'); // Get a specific scorer result
}
```





**Go:**

```go
for _, item := range result.Results {
    fmt.Println(item.Index)      // Position in batch
    fmt.Println(item.ItemID)     // Custom identifier (if provided)
    fmt.Println(item.RunID)      // Platform run ID
    fmt.Println(item.Output)     // Component output (json.RawMessage)
    fmt.Println(item.Passed)     // true if all scorers passed
    fmt.Println(item.Scores)     // []EvalScore
    fmt.Println(item.DurationMS) // Execution time
    fmt.Println(item.TraceID)    // OpenTelemetry trace ID
    fmt.Println(item.Error)      // Error message (empty unless evaluation failed)

    if score, ok := item.GetScore("json_valid"); ok { // Get a specific scorer result
        fmt.Println(score.Score, score.Passed)
    }
}
```



### Aggregate statistics



**Python:**

```python
stats = result.stats

stats.total_items      # Total items submitted
stats.completed_items  # Items evaluated without errors
stats.failed_items     # Items with evaluation errors
stats.passed_items     # Items where all scorers passed
stats.avg_duration_ms  # Average duration per item
stats.duration_ms      # Total batch wall-clock time
```





**TypeScript:**

```typescript
const stats = result.stats;

stats.totalItems;      // Total items submitted
stats.completedItems;  // Items evaluated without errors
stats.failedItems;     // Items with evaluation errors
stats.passedItems;     // Items where all scorers passed
stats.avgDurationMs;   // Average duration per item
stats.durationMs;      // Total batch wall-clock time
```





**Go:**

```go
stats := result.Stats

fmt.Println(stats.TotalItems)     // Total items submitted
fmt.Println(stats.CompletedItems) // Items evaluated without errors
fmt.Println(stats.FailedItems)    // Items with evaluation errors
fmt.Println(stats.PassedItems)    // Items where all scorers passed
fmt.Println(stats.AvgDurationMS)  // Average duration per item
fmt.Println(stats.DurationMS)     // Total batch wall-clock time
```



## Handling partial failures

An item can fail at evaluation time (component error) separately from failing scoring. The example below checks both.



**Python:**

```python
if result.status == "partial_failure":
    for item in result.failed_items():
        print(f"Evaluation error: {item.item_id} — {item.error}")

failing_scores = result.failing_items()
for item in failing_scores:
    print(f"Score fail: {item.item_id} (pass rate: {sum(s.score for s in item.scores) / len(item.scores):.0%})")
```





**TypeScript:**

```typescript
if (result.status === 'partial_failure') {
  for (const item of result.failedItems()) {
    console.log(`Evaluation error: ${item.itemId} — ${item.error}`);
  }
}

const failingScores = result.failingItems();
for (const item of failingScores) {
  const passRate = item.scores.reduce((sum, s) => sum + s.score, 0) / item.scores.length;
  console.log(`Score fail: ${item.itemId} (pass rate: ${(passRate * 100).toFixed(0)}%)`);
}
```





**Go:**

```go
if result.IsPartialFailure() {
    for _, item := range result.FailedItems() {
        fmt.Printf("Evaluation error: %s — %s\n", item.ItemID, item.Error)
    }
}

for _, item := range result.FailingItems() {
    if item.IsFailed() {
        continue // evaluation error, reported above
    }
    total := 0.0
    for _, score := range item.Scores {
        total += score.Score
    }
    fmt.Printf("Score fail: %s (pass rate: %.0f%%)\n", item.ItemID, total/float64(len(item.Scores))*100)
}
```



`Client` reads `AGNT5_GATEWAY_URL` from the environment automatically in all three SDKs; in Go, pass `""` as the gateway URL to `agnt5.NewClient` to use it.


**Python API**: `client.eval(component, input, expected?, scorers?, component_type?)` -> `EvalResponse`; `client.batch_eval(component, items, scorers?, expected?, component_type?, max_concurrency=10, timeout?)` -> `BatchEvalResult`.
**Go API**: `client.Eval(ctx, agnt5.EvalRequest{Component, ComponentType, Input, Expected, Scorers}, ...RunOption)` -> `(*EvalResponse, error)`; `client.BatchEval(ctx, component, []BatchEvalItem{Input, Expected?, ItemID?, Index?}, BatchEvalOptions{Scorers, Expected, ComponentType, DeploymentID, MaxConcurrency, Timeout}, ...RunOption)` -> `*BatchEvalResult` (never an error; per-item `Error`). Scorers are `[]EvalScorerSpec`: `{Name}`, a typed preset (`Correctness{}` ... with promoted `EvaluatorPresetConfig` fields) or `NewLLMJudge(LLMJudgeConfig{...})` via `ToEvalScorerSpec()`, or `NormalizeEvalScorers(...)` over a mix. `BatchEvalResult{BatchID, Status, Results, Stats}` with `PassRate()`, `IsSuccess()`, `IsPartialFailure()`, `PassingItems()`, `FailingItems()`, `FailedItems()`, `Outputs()`; `BatchEvalItemResult` adds `IsSuccess()`, `IsFailed()`, `GetScore(name) (EvalScore, bool)`.
**Input normalization**: plain dicts use positional `expected` list; dicts with `input` key use embedded `expected`; `BatchEvalItem(input, expected?, item_id?, index?)` for full control.
**Scorers**: strings ("exact_match"), SDK preset instances (`Correctness()`, `Helpfulness(model=...)`), or `LLMJudge(criteria, model, include_input, temperature)`.
**BatchEvalResult**: `batch_id`, `status` ("completed"/"partial_failure"/"failed"), `results: BatchEvalItemResult[]`, `stats: BatchEvalStats`, `pass_rate`, `passing_items()`, `failing_items()`, `failed_items()`.
**BatchEvalItemResult**: `index`, `run_id`, `output`, `scores: ScorerResultSummary[]`, `passed`, `duration_ms`, `item_id?`, `trace_id?`, `error?`, `get_score(name)`.


## Next steps

* [Scorers](/docs/improve/scorers.md): full reference for built-in scorers, SDK presets, and custom scorer code.
* [Experiments](/docs/improve/experiments.md): run a component against a versioned dataset, compare runs, and gate CI.
* [Datasets](/docs/improve/datasets.md): curate test cases from production runs and publish immutable versions.
* [Agents](/docs/build/agents.md): structure agent tool use so scorers have predictable events to check.
