Skip to main content
BM-STAT CERTIFIED · 246-TURN REPLAY

Same answers.62.75% less context.

Every figure below is replayable from raw logs. Nothing is estimated.

Mean lifecycle reductionBM-STAT
-62.75%
-85.7%
peak
5.8ms
p50 CPU
0.00pp
parity delta
Compression ramp · 10-turn codebase replay
≥70% a regime dal turno 8
turni 1–4 warm-up · picco 85.7%
—
—
—
—
70.5%
62%
74.6%
70.7%
84.3%
85.7%
70%
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10

Condenses context in-memory by -62.75% (up to -85.7% peak) with pure CPU processing (5.8ms p50) and zero token inflation.

Codebase review cut
-61.1%85.7% peak
Needle survival 128k
100.0%0% expansion
Proxy overhead p50
17.5msisolated
Pass-rate parity
0.00ppzero regression
TRANSPARENT PROXY TRANSIT TOPOLOGY

Neither an LLM nor an Agent: The Zero-Regression Gateway

Click a node to inspect payload transit
Transit Inspection: Velora Policy Gateway
-62.75% Mean Lifecycle Cut·0.0pp Regression

Condenses context in-memory by -62.75% (up to -85.7% peak) with pure CPU processing (5.8ms p50) and zero token inflation.

SMART ROUTING · MULTI-MODEL DISPATCH

How velora-auto routes across the model pool

Set model: "velora-auto" and each turn is classified into one of four complexity classes.

CLASS 1: TRIVIAL
Linting, Diff & Syntax

Cost-efficient sub-agent models from the pool.

CLASS 2: STANDARD
Single-Function Edits

High-throughput instruct models, exact tool-schema retention.

CLASS 3: COMPLEX
Multi-File Refactoring

Senior coding models with cross-file context.

CLASS 4: FRONTIER
System Architecture

Frontier reasoning models for deep design work.

Strict tool-call ID & signature preservationMeasured ~89.1% session saving (BM6 · lower bound)
AUTONOMOUS AGENT BENCHMARK · 246 CONTINUOUS TURNS

Why Context Explodes in Agentic Workflows

Coding agents in Docker testbeds repeatedly execute pytest, inspect files, and accumulate verbose compiler logs. Replay the turns below to see how context compounds and how Velora condenses it.

Mean Session Cut
-62.75%
246 continuous turns
Pass-Rate Parity
Delta 0.0pp
14/14 patch match
Trajectory Replay Turn ScrubberOptimization active
Cumulative Buffer at Turn #5
Pytest runner trace & refactoring coupon logic diff
Savings:-70.5%
Direct to Provider (Uncompressed Context Buffer)6,518 tokens
Through Velora Gateway (optimized context)2,155 tokens (-70.5%)
Turn Dynamics: Redundant stack traces and recurring compiler error logs are stripped; only essential failure lines and code changes are kept.
Parity Delta 0.0pp
Zero output regressions across 14/14 evaluated real SWE-agent trajectories.
246/246 200 OK
100% tool-call schema integrity: exact IDs and JSON parameters preserved.
2.00M Tokens Cut
Reduced 3,188,351 raw tokens to 1,187,672 billed tokens across replay.
Up to -75.3% Peak
Steady-state compression on mature turns with heavy terminal logs.
Sample of Evaluated SWE-agent Trajectories (Click to Inspect)Selected: astropy-13453 (-75.33% cut)

Global Benchmark Suite Evaluations

Rigorous testing on standardized datasets proving compression ratio, factual preservation, and context fidelity.

SWE-agent Trajectory Replay

246 Continuous Agent Turns
-62.75%Mean Lifecycle Cut
Peak Turn Reduction: Up to -75.33%

Evaluates authentic tool outputs, diffs, and shell traces. 3.19M tokens reduced to 1.19M billed tokens. Verified lossless string roundtrip and 0 protocol errors.

10 Multi-Turn Sessions100% 200 OK

Stack Reduction (BM1)

150-Turn End-to-End Pipeline Replay
-84.61%Token Stack Cut
p95 Latency Overhead: +3.44 ms

Full optimization pipeline. Condenses 3.23M raw tokens down to 497k forwarded tokens with 100% key-detail retention.

150 Turns Replay100% Needle Survival

NVIDIA RULER 128k

Needle-in-a-Haystack Factual Recall
100.0%Needle Survival

1.23M tokens tested across 128k contexts. 10/10 needles preserved. Intelligent Fast-Path automatically bypasses incompressible text with 0.00% token inflation.

128k Token Contexts0.00% Expansion
ISOLATED GATEWAY PERFORMANCE PROFILING

Gateway Transit Overhead Isolated From GPU Latency

A common objection from engineering buyers is whether adding an intelligent proxy introduces noticeable latency. We profiled Velora with a zero-delay mock upstream across 150 realistic requests to measure the exact internal proxy transformation and routing duration, isolated from the 3,000ms–5,000ms required for LLM GPU generation.

Workload Classp50 Latencyp90 Latencyp95 LatencyContext Characteristics
Small Payloads (~500 tok)6.31 ms7.40 ms8.43 msInteractive developer turns, single shell tool outputs, git status
Medium Payloads (~4k tok)17.47 ms18.61 ms19.42 msMulti-file inspections, unit test runners, structured diffs
Saturated Payloads (~20k tok)69.78 ms72.85 ms74.39 msDeep multi-turn logs, complete stack traces, cumulative agent buffers
Overall Gateway Overhead p50: 17.47 ms · Pure CPU Transform p50: 5.80 ms · Bootstrap p95 CI: [70.57, 73.63] ms on saturated buffers (<1% of end-to-end LLM turn duration).
Empirical Dynamics

The 3-Phase Context Lifecycle in Software Engineering

Understanding why full-session averages land at 61%–63% while mature turns achieve 75% to 85%+ steady-state compression.

Phase 1: Warmup (Turns 1–4)
0% (Fast-Path)

Initial task description, user prompt, and preliminary directory exploration. Fast-Path automatically passes short prompts verbatim with 0ms overhead.

Phase 2: Lifecycle Average
-62.75% Mean Cut

Certified statistical average across the complete lifespan of 246 continuous agent turns (and 61.1% on codebase review), factoring in early cold turns.

Phase 3: Saturated Peak (Turns 10+)
-75.3% to -85.7%

Where developer invoices explode: re-executed test suites, build traces, and repeated file inspections are compressed at peak efficiency.

DYNAMIC BUFFER BEHAVIOR ACROSS MULTI-TURN SESSIONS

How Compression Compounds Across Conversation Turns (Up to -85.7% Peak)

Notice the fundamental dynamic of autonomous software engineering: compression starts at 0% during initial cold reads via Fast-Path, averages 61.1% across full task lifecycles (72.4% across active compression turns), and steadily compounds to an 85.7% steady-state peak as repeated file reads, large stack traces, and verbose build logs saturate the context buffer.

← Swipe table horizontally →
TurnAgent Command / ActionPipeline PhaseRaw TokensForwarded TokensToken CutGateway Overhead
#1package.json stack inspectWarmup / Fast-Path4054140.0%2.5s
#2src/lib/billing.ts reviewWarmup / Fast-Path2,2202,3400.0%3.7s
#3src/lib/types.ts verificationWarmup / Fast-Path3,5973,7430.0%3.6s
#4src/app/api/checkout/route.tsWarmup / Fast-Path5,9306,2170.0%5.0s
#5Refactoring checkout coupon logicOptimization active6,5182,155-70.5%2.8s
#6src/components/api-key-modal.tsxOptimization active8,6333,678-62.0%3.1s
#7src/components/quota-banner.tsxOptimization active10,2602,921-74.6%2.9s
#8status/route.ts + re-read billing.tsOptimization active12,6664,148-70.7%3.4s
#9Build diagnostics tsc --noEmitOptimization active13,4592,360-84.3%3.2s
#10Architecture & security final auditSteady-State Peak13,9892,241-85.7%3.0s
COMPOUND COST EFFICIENCY ARCHITECTURE

The Dual Lever: Context Compression × Autonomous Routing

Inference spend is the product of two variables: token volume and model unit cost. Velora compresses token volume by -62.75% up to -85.7% (Lever 1). Enabling velora-auto adds semantic routing across the model pool, sending each turn to the optimal model for its complexity (measured ~89.1% session saving, BM6, lower bound).

Path A: Raw Baseline

Direct Upstream

100% Cost

All file history, bloated logs, and tool outputs sent verbatim to a single expensive frontier model ($3–15/M tokens).

Relative Cost: 1.00×
Path B: Lever 1

Compression Only (BYOK)

−63% to −85% Cost

Compresses prompt context across turns. Preserves your existing API keys and models while cutting billing by more than half.

Relative Cost: 0.37× to 0.15×
Max Capital Efficiency
Path C: Lever 1 × Lever 2

velora-auto Compound

~89% Measured Cut

Combines -62.75% token compression with semantic routing across the model pool: routine steps go to instruct models, architecture stays on frontier models.

Measured: $5.00 → $0.55 per session (BM6)

Inspect the raw data and replication artifacts.

We publish complete trajectory logs, PostgreSQL ledger dumps, and test harness code so engineering teams can independently audit every figure.