Same answers.62.75% less context.
Every figure below is replayable from raw logs. Nothing is estimated.
Condenses context in-memory by -62.75% (up to -85.7% peak) with pure CPU processing (5.8ms p50) and zero token inflation.
Neither an LLM nor an Agent: The Zero-Regression Gateway
Condenses context in-memory by -62.75% (up to -85.7% peak) with pure CPU processing (5.8ms p50) and zero token inflation.
How velora-auto routes across the model pool
Set model: "velora-auto" and each turn is classified into one of four complexity classes.
Cost-efficient sub-agent models from the pool.
High-throughput instruct models, exact tool-schema retention.
Senior coding models with cross-file context.
Frontier reasoning models for deep design work.
Why Context Explodes in Agentic Workflows
Coding agents in Docker testbeds repeatedly execute pytest, inspect files, and accumulate verbose compiler logs. Replay the turns below to see how context compounds and how Velora condenses it.
Global Benchmark Suite Evaluations
Rigorous testing on standardized datasets proving compression ratio, factual preservation, and context fidelity.
SWE-agent Trajectory Replay
Evaluates authentic tool outputs, diffs, and shell traces. 3.19M tokens reduced to 1.19M billed tokens. Verified lossless string roundtrip and 0 protocol errors.
Stack Reduction (BM1)
Full optimization pipeline. Condenses 3.23M raw tokens down to 497k forwarded tokens with 100% key-detail retention.
NVIDIA RULER 128k
1.23M tokens tested across 128k contexts. 10/10 needles preserved. Intelligent Fast-Path automatically bypasses incompressible text with 0.00% token inflation.
Gateway Transit Overhead Isolated From GPU Latency
A common objection from engineering buyers is whether adding an intelligent proxy introduces noticeable latency. We profiled Velora with a zero-delay mock upstream across 150 realistic requests to measure the exact internal proxy transformation and routing duration, isolated from the 3,000ms–5,000ms required for LLM GPU generation.
| Workload Class | p50 Latency | p90 Latency | p95 Latency | Context Characteristics |
|---|---|---|---|---|
| Small Payloads (~500 tok) | 6.31 ms | 7.40 ms | 8.43 ms | Interactive developer turns, single shell tool outputs, git status |
| Medium Payloads (~4k tok) | 17.47 ms | 18.61 ms | 19.42 ms | Multi-file inspections, unit test runners, structured diffs |
| Saturated Payloads (~20k tok) | 69.78 ms | 72.85 ms | 74.39 ms | Deep multi-turn logs, complete stack traces, cumulative agent buffers |
The 3-Phase Context Lifecycle in Software Engineering
Understanding why full-session averages land at 61%–63% while mature turns achieve 75% to 85%+ steady-state compression.
Initial task description, user prompt, and preliminary directory exploration. Fast-Path automatically passes short prompts verbatim with 0ms overhead.
Certified statistical average across the complete lifespan of 246 continuous agent turns (and 61.1% on codebase review), factoring in early cold turns.
Where developer invoices explode: re-executed test suites, build traces, and repeated file inspections are compressed at peak efficiency.
How Compression Compounds Across Conversation Turns (Up to -85.7% Peak)
Notice the fundamental dynamic of autonomous software engineering: compression starts at 0% during initial cold reads via Fast-Path, averages 61.1% across full task lifecycles (72.4% across active compression turns), and steadily compounds to an 85.7% steady-state peak as repeated file reads, large stack traces, and verbose build logs saturate the context buffer.
| Turn | Agent Command / Action | Pipeline Phase | Raw Tokens | Forwarded Tokens | Token Cut | Gateway Overhead |
|---|---|---|---|---|---|---|
| #1 | package.json stack inspect | Warmup / Fast-Path | 405 | 414 | 0.0% | 2.5s |
| #2 | src/lib/billing.ts review | Warmup / Fast-Path | 2,220 | 2,340 | 0.0% | 3.7s |
| #3 | src/lib/types.ts verification | Warmup / Fast-Path | 3,597 | 3,743 | 0.0% | 3.6s |
| #4 | src/app/api/checkout/route.ts | Warmup / Fast-Path | 5,930 | 6,217 | 0.0% | 5.0s |
| #5 | Refactoring checkout coupon logic | Optimization active | 6,518 | 2,155 | -70.5% | 2.8s |
| #6 | src/components/api-key-modal.tsx | Optimization active | 8,633 | 3,678 | -62.0% | 3.1s |
| #7 | src/components/quota-banner.tsx | Optimization active | 10,260 | 2,921 | -74.6% | 2.9s |
| #8 | status/route.ts + re-read billing.ts | Optimization active | 12,666 | 4,148 | -70.7% | 3.4s |
| #9 | Build diagnostics tsc --noEmit | Optimization active | 13,459 | 2,360 | -84.3% | 3.2s |
| #10 | Architecture & security final audit | Steady-State Peak | 13,989 | 2,241 | -85.7% | 3.0s |
The Dual Lever: Context Compression × Autonomous Routing
Inference spend is the product of two variables: token volume and model unit cost. Velora compresses token volume by -62.75% up to -85.7% (Lever 1). Enabling velora-auto adds semantic routing across the model pool, sending each turn to the optimal model for its complexity (measured ~89.1% session saving, BM6, lower bound).
Direct Upstream
All file history, bloated logs, and tool outputs sent verbatim to a single expensive frontier model ($3–15/M tokens).
Compression Only (BYOK)
Compresses prompt context across turns. Preserves your existing API keys and models while cutting billing by more than half.
velora-auto Compound
Combines -62.75% token compression with semantic routing across the model pool: routine steps go to instruct models, architecture stays on frontier models.
Inspect the raw data and replication artifacts.
We publish complete trajectory logs, PostgreSQL ledger dumps, and test harness code so engineering teams can independently audit every figure.