Skip to content

Research

What we measure, and what we can't yet show.

Two write-ups. One is a reliability result we are pleased with. The other is an evaluation result that does not support the thing we sell, and we publish it anyway.

The result we would rather not have.

Graph-derived retrieval participates in rank fusion on 201 of 231 question runs. On the question sets we can score, it has not answered a single question that vector-only retrieval missed — and one question went the other way.

We publish it because a category where everyone posts a favourable chart is a category where charts carry no information, and because someone would eventually run the comparison themselves.

Read the full methodology

What we can show

Reliability, measured on our own pipeline.

Before-and-after on the same system, about throughput and loss rather than answer quality.

Documents lost
9 in 205improved tozero

Across the same corpus, after moving orchestration onto durable steps with per-step retries.

Extraction p90
~103 simproved to5.2 s

Per document, sharded and run concurrently under a per-organisation ceiling.

Throughput
3.0 / minimproved to17.1 / min

Documents fully processed per minute, same corpus, same hardware shape.

Worst pipeline step
348% of budgetimproved to4%

As a share of the load balancer timeout. Above 100% the request dies before it finishes.

Run your own evaluation.

The free tier is enough to build a question set over your own content and measure it — including the refusal controls, which are the part that matters.

  • no card required
  • first workspace is free
  • one POST to ingest