Antibody developability · rank correlation
A model's place on this board is mostly a property of which assay you picked.
Eight benchmarks ask general-purpose models to rank real antibodies by a measured developability property — how long they stick to a column, how much they aggregate, what temperature they melt at, how much protein you get out. Each is scored by Spearman's rank correlation against the experiment. Higher is better, zero means no better than guessing the order, and negative means backwards.
The results do not support a single ranking of models. The benchmarks barely agree with each other, half the results sit at or below the point where the metric stops meaning anything, and no model is yet a usable predictor of any of these properties. All of that is on this page, before the leaderboard, because it is the more important finding.
What is measured
These are developability assays: not whether an antibody binds its target, but whether it can be manufactured, stored and injected. An antibody that binds beautifully and aggregates in the vial is not a drug.
Each benchmark holds one assay. The model is given antibodies and asked to put them in order by the measured property; the score is how well that order matches the laboratory's. Rank correlation is the right metric for this because the useful question is which of these candidates is worse, not what the retention time is in minutes.
| Benchmark | Every model, ρ from −0.30 to +0.70 | Best ρ | Best by | Median ρ | ρ<0.10 |
|---|
One dot per model. The vertical line is ρ = 0, where the
predicted order is no better than random; dots left of it are
ranked backwards. Click a dot to pin its model, or a benchmark
name for its full ranking.
Pinning joins that model's dots down the column with a
solid line, and hovering another joins it with a
dashed one — two at a time, to compare. A line
breaks at a benchmark the model has no result for, rather than
crossing a value that does not exist.
The leaderboard, and how far to trust it
Ranked by mean ρ across the benchmarks in the current slice. The rank range column is the part worth reading: it is the best and worst position the same model takes on individual benchmarks. Almost every model in this table has been near the top of something and near the bottom of something else.
| # |
|---|
Positions are always by mean ρ, whichever column you sort by, so re-sorting never renumbers the standings. A model with results for fewer than 90% of the benchmarks in the slice is flagged with its count rather than dropped — and its mean is taken over the benchmarks it has, never over a zero it did not score. Mean rank is computed within each benchmark, so a model that skipped one is absent from that ranking rather than placed last.
The bar spans the best and worst rank that model reaches on any single benchmark; the notch is its mean rank. Rows are ordered by mean ρ, so the leaderboard's order runs top to bottom.
The benchmarks disagree with each other
If these eight assays were all measuring "understands antibodies", a model that ranked well on one would rank well on the rest, and every cell below would be strongly positive. Take each pair of benchmarks, compare how they order the models they share, and this is what comes out.
Computed in the browser over the models the two benchmarks have in common; pairs sharing fewer than three models are left empty rather than estimated. This is agreement between benchmarks, so it is on a deliberately different colour scale from every performance figure on this page — blue for alike, red for opposed, and the value printed in every cell.
Every result
One row per model, one column per benchmark
The whole table, so nothing above has to be taken on trust. Colour marks the five interpretation bands; the two lowest bands share one recessive treatment because both mean the same thing — no usable signal — and are told apart by texture rather than by hue.
▲ marks the best model for that benchmark. A · means that model has no result for that benchmark — not a failed one — so it is left out of the model's mean instead of counted as zero. Click a model's name to pin it; it stays marked in every figure on the page.
Does more thinking help?
Three of the models in the feed were run twice, at medium and high reasoning effort, on the same benchmarks. That is a controlled comparison — same model, same task, one variable — and it is the only one this data set contains.
Each cell is high minus medium on that benchmark: positive means the extra effort helped. These are single runs with no interval attached, so a cell near zero should be read as "no measurable change", not as a small real one.
Method, and what this cannot tell you
The metric
Spearman's ρ compares two orderings. It runs from −1 to +1: +1 is the same order as the laboratory, 0 is an unrelated order, −1 is exactly reversed. It ignores how far apart the values are, which is the point — a model does not need to predict a retention time in minutes to be useful for triage, it needs to put the risky candidates last.
The bands
The five bands used for colour throughout — inverted, negligible, weak, moderate, strong, split at 0, 0.10, 0.30 and 0.50 — are the conventional effect-size labels for a rank correlation, and they live in the data file rather than in the code, so moving a threshold is a data change.
Averaging
Mean ρ is the plain arithmetic mean over the benchmarks a model has a result for. Averaging correlations via Fisher's z is arguably more correct; here it changes no value by more than 0.015 but does swap three adjacent pairs of models. That it swaps anything at all is the clearest evidence on this page that positions this close are not real.
-
No confidence intervals, no sample sizes
The feed carries a point estimate per model per benchmark and nothing else — not how many antibodies each ρ was computed over. Without that, no statement of the form "model A beats model B" is supportable at these margins, and this site does not make one. Slots exist in the data contract for intervals and counts; when they arrive, the figures will show them.
-
Coverage is partial by design
— Nothing here is a placeholder: the benchmarks that are present are complete and real. The ones that are missing simply have not been run yet.
-
Ragged model rosters
— A missing result is never treated as a zero, and models with gaps are flagged in the leaderboard, but a mean over six benchmarks and a mean over eight are still not quite the same measurement.
-
The assays are not equally hard
Mean ρ gives every benchmark equal weight, which quietly rewards the ones where scores run high. Mean rank does not, which is why both columns are in the table and why they do not produce the same order.