benchmarks

what we measured, how, and on what hardware. no projections: every number below comes from a recorded run of our benchmark suite

0.8 mswarm query p50 on S3
0.97recall@10 on 1.34M vectors
2parallel round trips for a cold vector query
$0.17S3 reads per million warm queries

Queries on real S3

SetupPhasep50 msp90 msp99 msS3 GETs / queryS3 $ / M queries
1 servercold1.033.0130.51.44$0.58
warm0.825.064.30.43$0.17
after restart0.924.349.70.35$0.14
3 serverscold2.040.3413.91.57$0.63
warm1.426.471.60.44$0.18
after restart1.525.151.10.28$0.11

coldbench: one store with 12,000 real web pages (1.77 GB of HTML), 100,000 arxiv vectors (768-d) and 2,000 small tenants. A Zipf stream from 8 clients: 50% text, 20% vector, 30% small tenants. Cold = empty disks; restart = every server restarted on the same disks. c7gd.2xlarge, S3 Standard in us-east-1, 400 MB cache per server.

Vector search

Measurearxiv-nomic-768 (1.34M, cosine)GIST-960 (1M, euclidean)
Recall@10, defaults (p50 / p99)0.970 (0.80 / 1.72 ms)0.913 (0.86 / 1.17 ms)
Recall@10, tuned0.982 at 1.27 ms0.968 at 1.65 ms
Queries / s, 8 at a time5,6005,560
With a filter passing 10%0.989 at 4.38 ms0.989 at 4.49 ms
Hybrid: text + filter (1.5% pass)0.982 at 50 ms0.976 at 58 ms
Ingest, acknowledged39,100 vectors / s28,500 vectors / s
Write to searchable (p50 / p99)22 / 28 ms23 / 28 ms
Cold query (p50 / p90), 2 parallel rounds63 / 70 ms60 / 66 ms
Cold query, a 10,000-vector namespace46 / 50 ms45 / 46 ms

vbench, in-process, 8 threads on a 14-core laptop, against the datasets’ brute-force ground truth. Defaults: nprobe 48, rerank 64 (int8). Cold rows add 20 ms to every object-store request.

Writes and text search

Measurestore +0 msstore +20 ms
Push, documents / s152,58050,470
Time to searchable, occasional writer9.3 ms32.8 ms
Write with wait: true30.5 ms103.4 ms
Small namespace, cold / warm query2.01 / 0.07 ms25.6 / 0.14 ms
Big namespace, cold / warm query13.1 / 2.10 ms34.8 / 2.09 ms
Exact count with a filter12.6 ms14.0 ms
Vector query, warm (p50)0.53 ms0.53 ms

The product benchmark (bench/run.sh) over HTTP: 50,000 ~1 KB documents, 300 small namespaces, 20,000 128-d vectors. Server 6 threads; local object store with +0 or +20 ms per request (+20 models S3). Medians of the runs’ p50s.

Noisy neighbours

Tenant Bp50p99
Search, server quiet0.10 ms0.63 ms
Search, A flooding (held to its plan)0.39 ms1.56 ms
Write with wait, A flooding (held to its plan)39.6 ms

tenancy_test: tenant A floods one server from 16 processes (2 MB pushes and regex scans, back to back) while tenant B searches and writes one at a time. Per-tenant admission keeps A to its plan’s concurrency.

What these numbers mean

  • Warm means the namespace's data is in a server's RAM or NVMe cache. Cold means it has to be read from object storage, which costs a round trip or two to S3 (20 to 50 ms each).
  • The p90 on S3 is mostly first-time queries: data no cache has seen yet. Repeated queries stay near the p50.
  • S3 cost per query is what the reads themselves cost at S3's list price ($0.40 per million GETs). Warm queries make none.
  • Runs on a laptop share it with other work; the S3 runs used dedicated spot instances. We publish medians, not best runs.