Your samples
Separate numbers with spaces, commas, semicolons, tabs or newlines. Negative and zero values are legal. You can also drop a plain text file onto a shard box.
Result
How it computes
Quantiles use the type 7 definition, the default in NumPy's percentile and the same one Excel's PERCENTILE.INC uses. For a sorted sample of n values and p in [0, 1], let h = (n - 1)p, take lo = floor(h), and return xs[lo] + (h - lo) x (xs[lo+1] - xs[lo]), clamping lo+1 to the last index. Nothing is rounded and no other quantile type is offered.
The pooled value is that same function applied once to every sample from every shard concatenated. The unweighted estimate is the plain mean of the per-shard quantiles. The traffic-weighted estimate weights each shard's quantile by that shard's sample count. Signed error is 100 x (estimate - pooled) / |pooled|, divided by the size of the pooled value so that a negative pooled value cannot flip the sign, and the word "overstates" or "understates" is chosen from the sign of that number at run time.
When the naive answer is exactly right, and the limit of that claim. For population quantiles the argument is clean: if every component CDF equals p at x, then every convex mixture of them equals p at x, so if every shard's population quantile is x the pooled population quantile is x too and the mean of the shard quantiles is exact, for any shard shapes and any mixing weights. That argument does not carry over to the finite-sample type 7 quantile this page actually computes, because an interpolated order statistic is not a CDF inverse. Two shards can return a bit-identical p95 and still pool to a different one. Hand-checkable counterexample at p = 0.9: shard A is 0 1 3 3 4 and shard B is 2 2 3 3 4, both p90 are 3.60 so the spread is exactly zero, and the pooled p90 over all ten samples is 4.00, so the mean of the shard quantiles is off by -10.00 percent. The self test asserts that case. So read the spread as a hint about how close the shards are, and read the signed error above it as the answer.
A KS test is deliberately not used here; at n=10000 it rejects homogeneity on shards whose mean-of-p95 error is 0.02 percent.
Self test
Runs the shipped math against fixtures you can check by hand, including a positive control that is designed to fail. If the control ever reports a pass, the test harness itself is broken and nothing below it should be trusted.