Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Budget-Aware Testing

Soroban contracts execute within strict CPU and memory budgets. Tests that pass locally can fail on-chain once resource consumption exceeds the transaction limits — Testkit makes those costs visible in ordinary assertions.

Cratesoroban-testkit-core
APIBudgetSnapshot · BudgetGuard · BudgetBaseline · budget_guard!
Runs incargo test / CI

Why budget matters

The Soroban runtime enforces resource budgets per transaction. The SDK’s test environment does not enforce those limits by default, so a contract can pass every unit test and still fail during deployment or execution on testnet and mainnet.

Symptom Green test suite, HOST_VALUE_SIZE / budget exhausted error on-chain. Fix: assert on the CPU and memory cost of each call, not just its result.

Basic budget tracking

The SDK meters one call at a time, and the metering of the call that just finished is what a test should read:

#![allow(unused)]
fn main() {
use soroban_testkit_core::budget::BudgetSnapshot;

#[test]
fn test_budget_regression() {
    let env = Env::default();
    env.mock_all_auths();

    // ... invoke function
    client.increment(&caller, &1);

    let cost = BudgetSnapshot::last_invocation(&env);
    assert!(cost.cpu_insns < 1_000_000, "CPU: {}", cost.cpu_insns);
    assert!(cost.mem_bytes < 100_000, "Memory: {}", cost.mem_bytes);
}
}
Where the numbers come from

SDK v28 meters each top-level invocation separately: env.cost_estimate().resources() returns what the call that just ran consumed, and env.cost_estimate().budget() returns the same metering as a running total for that call. The older env.budget() accessor is deprecated and should not appear in new tests.

What a native test contract hides

A contract registered with env.register(...) runs as host code, so VM instantiation, wasm execution and rent reads are never metered — the reading is the storage and host-work half of the real cost, not all of it. Treat these numbers as a comparison between builds of the same contract rather than as a fee quote, and use a deployed wasm contract when you need the full figure.

Reading a snapshot

FieldMeaningSuggested use
cpu_insnsInstructions metered for one invocationRegression guard in CI
mem_bytesMemory metered for one invocationKeeps entries under ledger limits
last_invocation()Snapshot of the call that just ranThe reading a test asserts on
diff()Saturating delta of two snapshotsComparing two readings of one running budget

Asserting a limit instead of a number

BudgetGuard carries the two limits a reviewer can argue about: a ceiling this call may not cross, and a growth allowance against the cost recorded the last time the case was measured.

  1. Name the case

    BudgetGuard::new("transfer"). The name is what every failure line carries, so a CI log points at an operation rather than at a test file.

  2. Set the ceilings

    .cpu_ceiling(2_000_000).mem_ceiling(500_000) — absolute limits, independent of history.

  3. Add the baseline

    .baseline(Some(recorded)).tolerance_percent(10) rejects a call that grew more than 10% past the recorded cost. A tolerance of 0 means not one instruction more.

  4. Measure the call

    .run(&env, || client.transfer(&from, &to, &1000)) makes the call, reads the metering that call left behind, asserts against it, and returns the invocation's value.

#![allow(unused)]
fn main() {
use soroban_testkit_core::budget::BudgetGuard;

#[test]
fn increment_stays_inside_its_budget() {
    let ctx = TestContextBuilder::new().with_users(1).build();
    let (id, client) = register(&ctx);
    let caller = ctx.users[0].clone();

    BudgetGuard::new("increment")
        .cpu_ceiling(20_000_000)
        .mem_ceiling(20_000_000)
        .run(&ctx.env, || client.increment(&caller, &1));
}
}

The budget_guard! macro is the same assertion written around the call:

#![allow(unused)]
fn main() {
use soroban_testkit_core::budget_guard;

budget_guard!(&env, "transfer", {
    cpu_max: 2_000_000,
    mem_max: 500_000,
    baseline: recorded,        // Option<BudgetSnapshot>
    tolerance: 10,             // percent
}, || client.transfer(&sender, &receiver, &1000));
}
Both readings, once

A guard reports every limit a cost breaks — ceilings first, then growth — instead of stopping at the first. One run tells you whether a change moved CPU, memory, or both.

Committing the numbers: baseline files

A baseline is a JSON file mapping case names to the cost recorded for them, so the limits live in the repository and a pull request shows a cost change as a diff.

{
  "version": 1,
  "cases": {
    "get": { "cpu_insns": 7861, "mem_bytes": 1510 },
    "increment": { "cpu_insns": 32669, "mem_bytes": 5252 }
  }
}
#![allow(unused)]
fn main() {
use soroban_testkit_core::budget::BudgetBaseline;

let mut baseline = BudgetBaseline::load(std::path::Path::new("tests/budget.json")).unwrap();

baseline
    .guard("increment")            // carries the recorded cost for the case
    .tolerance_percent(5)
    .run(&env, || client.increment(&caller, &1));

// Re-record and commit the file when a cost change is intended:
baseline.record("increment", measured_cost);
baseline.save(std::path::Path::new("tests/budget.json")).unwrap();
}

guard() on a case the file does not know returns a guard with no baseline, so a new operation is only bounded by whatever ceilings the test adds — the missing case never fails a suite by itself. A file written by a newer Testkit is refused rather than half-read.

Symptom A baseline that regenerates on every run accepts whatever the last run cost. Fix: commit the file, and rewrite it only in a deliberate "record budgets" change.

Machine-readable output

Every breach renders as one line of key=value pairs, which is what lets a CI step grep, annotate, or trend them:

BUDGET kind=ceiling case=transfer metric=cpu_insns actual=2500000 limit=2000000
BUDGET kind=growth case=transfer metric=mem_bytes actual=610000 limit=550000 baseline=500000 tolerance_percent=10
FieldMeaning
kindceiling (absolute limit) or growth (past a baseline)
caseThe name the guard was built with
metriccpu_insns or mem_bytes
actual / limitMeasured cost and the largest value still accepted
baseline / tolerance_percentOnly on growth, so the ratio is recoverable from the line

Running the check in CI

The counter example ships the whole loop: examples/counter/budget.json records the cost of its two hot paths, and two #[ignore]d tests in examples/counter/src/test.rs drive it.

# record the costs where they will be enforced
cargo test -p soroban-testkit-example-counter -- --ignored --nocapture \
    budget_baseline_records_current_costs

# compare a run against the committed file
TESTKIT_BUDGET_TOLERANCE=5 cargo test -p soroban-testkit-example-counter -- --ignored --nocapture \
    budget_baseline_rejects_drifted_costs

Both are ignored on purpose. Soroban metering is deterministic, so the reading is a property of the contract and the pinned SDK rather than of the machine that ran it — the file recorded on Windows measured identical on ubuntu-latest, drift=+0.00% on every metric. What that means is that an SDK or env-host bump moves every number at once, and the recording test rewrites a file in the repository, so neither belongs in cargo test --workspace. The Budget baseline workflow runs them where Cargo.lock and the runner are fixed: one pinned ubuntu-latest, the numbers printed in the job summary, a comment on the pull request, and a failure when a case grows past the tolerance — 10% by default, tolerance on a manual run, fail_on_drift off when you want the report without the verdict.

BUDGET_SUMMARY case=get metric=cpu_insns actual=7861 baseline=7861 drift=+0.00%
BUDGET_SUMMARY case=increment metric=cpu_insns actual=32669 baseline=32669 drift=+0.00%
FieldMeaning
caseThe key in the committed budget.json
actualWhat the call cost on this runner
baselineWhat the committed file says it cost
driftactual against baseline, signed percent

A breach then follows as the usual BUDGET kind=growth … line, so the same grep covers both the check and an in-test guard.

  1. An unintended rise

    The job fails and the comment says which case, which metric and by how much. Nothing to record — the code is what has to change.

  2. An intended rise

    Run the workflow with Record checked: it rewrites budget.json on the same runner that enforces it and uploads the file as an artifact. Commit that file on the branch, and the pull request diff reads as a reviewable statement — increment: 32669 → 41000 — rather than a red check someone retried.

One guard, one call

The metering the SDK reports belongs to the top-level invocation that has just finished, so a closure that makes two calls is judged on the second one. Give each call its own guard, and name the case after the call you care about.


Related: Property testing for inputs, soroban-testkit-core for the full budget API.