Before You Call an LLM Endpoint "Nerfed": A Small Statistics Checklist in Python

Every few weeks someone posts "provider X is serving a watered-down model" with a handful of screenshots, and every few weeks the replies split into "same here" and "works fine for me." Both camps are usually arguing from data that can't settle the question.

My last post argued that a single output tells you almost nothing. A commenter on it made a sharper point that I want to build on: even with repeated runs, the way you pick what to compare can manufacture a difference. Comparing two pre-chosen providers at 16/20 vs 8/20 is meaningful (Fisher p ≈ 0.02). Scanning a table of ten providers and quoting the best and worst is not, because with ten identical providers a gap that large shows up by luck about 30% of the time. Their suggested fix — compare against a pooled rate or a reference endpoint chosen in advance — is the backbone of this post.

So here's the checklist I now run before believing that an endpoint regressed. Everything assumes the simplest useful setup: one fixed probe (for example, "draw a pelican riding a bicycle as SVG"), a fixed pass/fail rubric, and a count of passes out of N attempts. All numbers below are illustrative — they come from the code shown, not from any real provider. Every snippet runs as-is with scipy and numpy.

1. Put an interval on every rate

A pass rate without an interval is a vibe. For binomial counts, the Wilson interval behaves well at small N:

from scipy.stats import binomtest

# Illustrative numbers, not real measurements
for label, k, n in [("last week", 34, 40), ("this week", 25, 40)]:
    ci = binomtest(k, n).proportion_ci(confidence_level=0.95, method="wilson")
    print(f"{label}: {k}/{n} = {k/n:.1%}   95% Wilson CI {ci.low:.1%} to {ci.high:.1%}")
last week: 34/40 = 85.0%   95% Wilson CI 70.9% to 92.9%
this week: 25/40 = 62.5%   95% Wilson CI 47.0% to 75.8%

The intervals barely overlap. That alone doesn't settle it (overlapping intervals are not a significance test, and non-overlap isn't required for one), but it tells you how loosely 40 runs pin down each rate. If you only had 10 runs, both intervals would span roughly half the scale.

2. One pre-chosen comparison: Fisher's exact test

If you decided before running anything that you would compare "last week" against "this week" on this probe, a 2×2 exact test is the honest tool:

from scipy.stats import fisher_exact

#          pass  fail
table = [[34,    6],   # last week (illustrative)
         [25,   15]]   # this week (illustrative)

two_sided = fisher_exact(table, alternative="two-sided").pvalue
one_sided = fisher_exact(table, alternative="greater").pvalue  # only if "drop" was pre-registered
print(f"two-sided p = {two_sided:.4f}")
print(f"one-sided p = {one_sided:.4f}")
two-sided p = 0.0406
one-sided p = 0.0203

Two caveats. The one-sided p-value is only legitimate if "this week is worse" was the hypothesis you wrote down in advance; choosing the direction after seeing the data is just halving your p-value for free. And p = 0.04 means "data this lopsided would be unusual if nothing changed," not "96% chance the provider swapped models."

3. Count how many comparisons you could have made

This is the commenter's point, and it deserves a simulation. Give every provider the exact same true pass rate (60%), run each 20 times, and look at the gap between the best and worst:

import numpy as np

rng = np.random.default_rng(42)
n_runs, p_true, sims = 20, 0.6, 200_000   # every provider is identical: 60% pass rate

for k in (2, 5, 10, 20):
    passes = rng.binomial(n_runs, p_true, size=(sims, k))
    gap = passes.max(axis=1) - passes.min(axis=1)
    print(f"{k:>2} providers: P(best-worst gap >= 8) = {(gap >= 8).mean():5.1%}   "
          f"95th percentile of gap = {np.percentile(gap, 95):.0f}")
 2 providers: P(best-worst gap >= 8) =  1.4%   95th percentile of gap = 6
 5 providers: P(best-worst gap >= 8) = 10.4%   95th percentile of gap = 8
10 providers: P(best-worst gap >= 8) = 29.9%   95th percentile of gap = 10
20 providers: P(best-worst gap >= 8) = 62.4%   95th percentile of gap = 11

With two pre-chosen providers, an 8-pass gap is rare. With ten, it happens about 30% of the time with nothing wrong — matching the commenter's estimate. With twenty, it's the norm. The 95th-percentile column is the gap you'd need just to clear the noise floor of "best vs worst," and it keeps rising as the table grows.

If your claim comes from eyeballing a leaderboard, the number of comparisons is the number of rows, not two.

4. Compare each endpoint to the pool, then correct for multiplicity

Instead of best-vs-worst, ask a question that doesn't depend on which rows happen to land at the extremes: is any endpoint off from the others by more than chance allows for this many endpoints?

import numpy as np
from scipy.stats import chi2_contingency, fisher_exact

# Illustrative pass counts for 8 endpoints serving the "same" model, 30 runs each
passes = {"A": 22, "B": 20, "C": 23, "D": 19, "E": 21, "F": 12, "G": 22, "H": 20}
n = 30

table = np.array([[k, n - k] for k in passes.values()])
chi2, p_omni, dof, _ = chi2_contingency(table)
print(f"omnibus chi2 = {chi2:.2f}, dof = {dof}, p = {p_omni:.4f}")

# Each endpoint vs the pooled pass rate of all the OTHER endpoints
raw = {}
for name, k in passes.items():
    others_pass = sum(passes.values()) - k
    others_fail = n * (len(passes) - 1) - others_pass
    raw[name] = fisher_exact([[k, n - k], [others_pass, others_fail]]).pvalue

# Holm step-down correction for 8 tests
m = len(raw)
running_max = 0.0
for i, (name, p) in enumerate(sorted(raw.items(), key=lambda kv: kv[1])):
    running_max = max(running_max, min(1.0, (m - i) * p))
    print(f"{name}: {passes[name]}/{n}  raw p = {p:.4f}  Holm p = {running_max:.4f}")
omnibus chi2 = 12.35, dof = 7, p = 0.0895
F: 12/30  raw p = 0.0018  Holm p = 0.0144
C: 23/30  raw p = 0.2221  Holm p = 1.0000
A: 22/30  raw p = 0.4178  Holm p = 1.0000
G: 22/30  raw p = 0.4178  Holm p = 1.0000
E: 21/30  raw p = 0.6862  Holm p = 1.0000
D: 19/30  raw p = 0.8367  Holm p = 1.0000
B: 20/30  raw p = 1.0000  Holm p = 1.0000
H: 20/30  raw p = 1.0000  Holm p = 1.0000

Two things worth noticing:

  • The omnibus chi-square doesn't reach 0.05. It spreads its attention over all 7 degrees of freedom, so one clear outlier among mostly similar endpoints can slip under it.
  • The per-endpoint test against the pooled rate of the others flags F, and it survives Holm correction across all 8 tests (adjusted p ≈ 0.014). Everyone else is unremarkable — including C at 23/30, which a "top of the table" screenshot would happily crown.

Holm is a reasonable default: it controls the chance of any false flag across the family and is never less powerful than plain Bonferroni. The leave-one-out pool matters too; if F were included in its own baseline, it would drag the baseline toward itself.

5. Pre-register the plan (it takes 10 lines)

"Pre-registration" sounds academic. In practice it means writing down the comparison before you collect data and making it hard to quietly edit later:

import hashlib, json

plan = {
    "question": "Has endpoint X's first-attempt pass rate dropped vs reference R?",
    "endpoints": {"test": "provider-x/model-y", "reference": "official-api/model-y"},
    "probe": "pelican-static-v1",
    "rubric": "pass = all required checks pass; timeouts count as fail",
    "runs_per_endpoint": 80,
    "schedule": "4 blocks, interleaved, random order within block",
    "primary_test": "Fisher exact, one-sided (X < R), alpha = 0.05",
    "stopping_rule": "no interim looks",
}
blob = json.dumps(plan, sort_keys=True).encode()
print(hashlib.sha256(blob).hexdigest())
1e240df48b2c46c2cbd1782eecd2fe87c69f1ef6b469b3a97f3625e803ab884e

Post the hash (in a commit, an issue, a chat message) before the first run, and publish the plan alongside the results. Note what goes in it: a reference endpoint fixed in advance, the sample size, the schedule, how timeouts are counted, the primary test, and the fact that there are no interim looks. If you later add a second probe or a third provider, that's a new plan, not an edit.

The reference endpoint matters most for "did it regress over time?" questions. Re-run the reference in the same window as the endpoint you suspect. If both dropped together, the culprit is more likely your probe, your judge, or something shared upstream — not a secret model swap at one provider.

6. Check that you can actually detect the drop you care about

Twenty runs per side feels like a lot when you're doing it by hand. It isn't, if the drop you're worried about is moderate. Here's a simulated power check for a real 80% → 60% drop:

import numpy as np
from scipy.stats import fisher_exact

rng = np.random.default_rng(7)
p_before, p_after, sims = 0.80, 0.60, 2000   # a real 20-point drop

for n in (20, 40, 80, 120):
    a = rng.binomial(n, p_before, sims)
    b = rng.binomial(n, p_after, sims)
    hits = sum(
        fisher_exact([[x, n - x], [y, n - y]], alternative="greater").pvalue < 0.05
        for x, y in zip(a, b)
    )
    print(f"n = {n:>3} per arm: power ~ {hits / sims:.0%}")
n =  20 per arm: power ~ 30%
n =  40 per arm: power ~ 56%
n =  80 per arm: power ~ 82%
n = 120 per arm: power ~ 95%

At 20 runs per arm you'd miss a genuine 20-point drop about 70% of the time. Around 80 per arm gets you to the conventional 80% power. Two consequences:

  • "I ran it 20 times and saw no difference" is weak evidence that nothing changed. Absence of significance is not equivalence.
  • Decide the drop size you care about first, then the N. Not the other way round.

7. Don't peek and stop when it looks good

The most natural thing in the world is to check the numbers after every batch and post as soon as the gap looks significant. It quietly breaks the test:

import numpy as np
from scipy.stats import fisher_exact

rng = np.random.default_rng(1)
p_true, sims, looks = 0.7, 2000, range(10, 61, 5)   # no real difference at all

false_alarms = 0
for _ in range(sims):
    a = rng.random(60) < p_true
    b = rng.random(60) < p_true
    for n in looks:
        x, y = a[:n].sum(), b[:n].sum()
        if fisher_exact([[x, n - x], [y, n - y]]).pvalue < 0.05:
            false_alarms += 1
            break

fixed = 0
for _ in range(sims):
    x, y = rng.binomial(60, p_true, 2)
    fixed += fisher_exact([[x, 60 - x], [y, 60 - y]]).pvalue < 0.05

print(f"peek every 5 runs, stop at p<0.05: false alarm rate ~ {false_alarms / sims:.1%}")
print(f"single test at n=60:              false alarm rate ~ {fixed / sims:.1%}")
peek every 5 runs, stop at p<0.05: false alarm rate ~ 8.7%
single test at n=60:              false alarm rate ~ 3.4%

Here both arms are identical. Peeking every 5 runs and stopping at the first p < 0.05 more than doubles the false-alarm rate compared with one test at the planned N. (Fisher's test is conservative at these sizes, which is why the fixed-N rate sits below 5%.) If you genuinely need to monitor continuously, use a method designed for it, such as sequential or always-valid tests. Otherwise, fix N and look once.

The checklist, condensed

  1. Report pass rates with intervals, not bare percentages.
  2. For one pre-chosen pair, use an exact test; pick one- vs two-sided in advance.
  3. Count every comparison you could have made, not the one you ended up quoting.
  4. Compare each endpoint to a pooled or pre-registered reference, and correct for multiplicity.
  5. Write the plan down (and hash it) before the first run.
  6. Size N for the smallest drop you care about.
  7. Fix N and look once — or use a sequential method on purpose.

None of this tells you why an endpoint got worse. Timeouts, truncated outputs, a different default reasoning effort, and an actual model change can all lower a pass rate. Statistics only tells you whether there's a difference worth explaining. Keep timeouts and wrong answers as separate categories in your logs, so that when something does show up you can tell which kind of failure moved.

Where the tooling fits

The statistics above fit in a page. What actually costs time is the bookkeeping around it: same prompt, same parameters, timestamps on every run, a reference endpoint re-tested in the same window, and failures kept instead of dropped.

That bookkeeping is the problem I work on at Folkbench. It's a public leaderboard that compares different services (official APIs and third-party relays) serving the same model, using one published methodology, with the test time shown on each result and missing data left blank rather than guessed. There's also a pelican-on-a-bicycle gallery, which the site itself labels as a fun side-by-side that does not feed into the rankings and should not be read as a capability score (given everything above, I think that is the right call). If you'd rather start from existing results than build the harness yourself, it's at folkbench.com. Apply the same checklist to it: ask what was tested, when, and how many runs are behind each number.

If you run this kind of comparison yourself, I'd like to hear how you choose the reference endpoint. That's the decision I'm least sure about.

Disclosure: I work on Folkbench. Drafted with AI assistance; code and numbers checked by me. All numbers in this post are illustrative outputs of the code shown, not measurements of any real provider.

Story originally reported by Dev.to. View at Dev.to →
← Back to all news