Guest loop and closure allocation benchmark

Run ./gradlew-sandbox-dev-parallel-summary benchmarkTransientAllocations. The opt-in generator compiles eight programs with the same 256-element List<Int> and 300 summation rounds. Each program checks the same checksum, 86,957,440, and retains the original list for a final scan. The harness verifies the artifacts with the Rust VM and separates construction from the repeated operation at println("ready"). It takes three interleaved timing samples and a separate pressure run at a 16 KiB Guest heap. This is managed Guest heap evidence, not total host process memory per VM.

Sources are retained under modules/common/compiler-k2/build/generated/benchmarks/transient-allocations. Reports are under build/reports/benchmarks/transient-allocations. The measurement artifact removes the compiler’s normal 64 KiB admission floor; it does not bypass the Rust verifier. Fixed-budget runs leave minimum-heap columns empty. timing_heap_bytes and pressure_heap_bytes distinguish timing fallback from the requested pressure budget. No fallback was needed in the recorded baseline.

Baseline captured on 2026-09-27, before changing capture-cell selection: measurements, raw samples. All eight checksums passed. fold-read-var alone incurred GC work (737 maintenance units); its otherwise identical fold-val control incurred none. The compiler allocated a separate typed cell for the read-only captured var. The mutable control reassigns its captured variable after constructing the lambda and checks that the lambda sees that new value. This checks aliasing rather than just a numerically equivalent sum.

Instruction and resource units are deterministic VM work counters, not allocation counts or allocated bytes. Three timing samples are exploratory and insufficient to establish a small CPU speedup. The cases compare ordinary indexed loops, iterator loops, uncaptured fold, scalar captures, shared mutable captures, reference captures, and reuse of a prebuilt function value. They do not establish that all loops or all collection operations have the same cost.

Read-only captured variables

The compiler now retains a shared cell only when it finds an assignment to that captured variable anywhere in the relevant IR roots, including nested closures. This is conservative: even a dead or pre-capture assignment retains its cell. No escape analysis, closure lifetime assumption, or change to Guest APIs is required.

After the change: measurements, raw samples. All eight programs and all 24 timed samples passed at the requested 16 KiB heap with no fallback.

fold-read-var, 300 rounds Before After
Extra captured-variable cells 300 0
Cumulative bytes allocated for those cells 4,800 0
Hot instructions 4,707,575 4,630,175
Hot fixed units 10,030,866 9,875,466
Hot dynamic units 911 611
Hot GC maintenance units 737 0
Artifact bytes 57,944 57,688

The byte saving follows from one 16-byte Int capture cell per round and the emitted allocation path; it is cumulative allocation traffic, not a 4,800-byte reduction in live memory. The closure and iterator still allocate. The optimized case now matches the fold-val control in artifact size and deterministic hot work counters. All seven other cases retained their artifact size and deterministic hot counters, including the mutable control. Fixed work decreased by approximately 1.55% in this program. This is not a wall-clock CPU speedup claim: the after run overlapped focused compiler/conformance checks, and wall-clock times also changed for unchanged controls.

Verification: compiler allocation-contract test, testKotlinFunctionValuesVmConformance, testKotlinDestinationVmConformance, testKotlinGenericFunctionsVmConformance, compiler/engine Kotlin lint, build-script tests, and the benchmark. These are focused checks, not complete checkout or release verification.