benchmarks
what we measured, how, and on what hardware. no projections: every number below comes from a recorded run of our benchmark suite
Queries on real S3
| Setup | Phase | p50 ms | p90 ms | p99 ms | S3 GETs / query | S3 $ / M queries |
|---|---|---|---|---|---|---|
| 1 server | cold | 1.0 | 33.0 | 130.5 | 1.44 | $0.58 |
| warm | 0.8 | 25.0 | 64.3 | 0.43 | $0.17 | |
| after restart | 0.9 | 24.3 | 49.7 | 0.35 | $0.14 | |
| 3 servers | cold | 2.0 | 40.3 | 413.9 | 1.57 | $0.63 |
| warm | 1.4 | 26.4 | 71.6 | 0.44 | $0.18 | |
| after restart | 1.5 | 25.1 | 51.1 | 0.28 | $0.11 |
coldbench: one store with 12,000 real web pages (1.77 GB of HTML), 100,000 arxiv vectors (768-d) and 2,000 small tenants. A Zipf stream from 8 clients: 50% text, 20% vector, 30% small tenants. Cold = empty disks; restart = every server restarted on the same disks. c7gd.2xlarge, S3 Standard in us-east-1, 400 MB cache per server.
Vector search
| Measure | arxiv-nomic-768 (1.34M, cosine) | GIST-960 (1M, euclidean) |
|---|---|---|
| Recall@10, defaults (p50 / p99) | 0.970 (0.80 / 1.72 ms) | 0.913 (0.86 / 1.17 ms) |
| Recall@10, tuned | 0.982 at 1.27 ms | 0.968 at 1.65 ms |
| Queries / s, 8 at a time | 5,600 | 5,560 |
| With a filter passing 10% | 0.989 at 4.38 ms | 0.989 at 4.49 ms |
| Hybrid: text + filter (1.5% pass) | 0.982 at 50 ms | 0.976 at 58 ms |
| Ingest, acknowledged | 39,100 vectors / s | 28,500 vectors / s |
| Write to searchable (p50 / p99) | 22 / 28 ms | 23 / 28 ms |
| Cold query (p50 / p90), 2 parallel rounds | 63 / 70 ms | 60 / 66 ms |
| Cold query, a 10,000-vector namespace | 46 / 50 ms | 45 / 46 ms |
vbench, in-process, 8 threads on a 14-core laptop, against the datasets’ brute-force ground truth. Defaults: nprobe 48, rerank 64 (int8). Cold rows add 20 ms to every object-store request.
Writes and text search
| Measure | store +0 ms | store +20 ms |
|---|---|---|
| Push, documents / s | 152,580 | 50,470 |
| Time to searchable, occasional writer | 9.3 ms | 32.8 ms |
| Write with wait: true | 30.5 ms | 103.4 ms |
| Small namespace, cold / warm query | 2.01 / 0.07 ms | 25.6 / 0.14 ms |
| Big namespace, cold / warm query | 13.1 / 2.10 ms | 34.8 / 2.09 ms |
| Exact count with a filter | 12.6 ms | 14.0 ms |
| Vector query, warm (p50) | 0.53 ms | 0.53 ms |
The product benchmark (bench/run.sh) over HTTP: 50,000 ~1 KB documents, 300 small namespaces, 20,000 128-d vectors. Server 6 threads; local object store with +0 or +20 ms per request (+20 models S3). Medians of the runs’ p50s.
Noisy neighbours
| Tenant B | p50 | p99 |
|---|---|---|
| Search, server quiet | 0.10 ms | 0.63 ms |
| Search, A flooding (held to its plan) | 0.39 ms | 1.56 ms |
| Write with wait, A flooding (held to its plan) | 39.6 ms |
tenancy_test: tenant A floods one server from 16 processes (2 MB pushes and regex scans, back to back) while tenant B searches and writes one at a time. Per-tenant admission keeps A to its plan’s concurrency.
What these numbers mean
- Warm means the namespace's data is in a server's RAM or NVMe cache. Cold means it has to be read from object storage, which costs a round trip or two to S3 (20 to 50 ms each).
- The p90 on S3 is mostly first-time queries: data no cache has seen yet. Repeated queries stay near the p50.
- S3 cost per query is what the reads themselves cost at S3's list price ($0.40 per million GETs). Warm queries make none.
- Runs on a laptop share it with other work; the S3 runs used dedicated spot instances. We publish medians, not best runs.