Generative Biologics Benchmarks · developability

Antibody developability · rank correlation

A model's place on this board is mostly a property of which assay you picked.

Eight benchmarks ask general-purpose models to rank real antibodies by a measured developability property — how long they stick to a column, how much they aggregate, what temperature they melt at, how much protein you get out. Each is scored by Spearman's rank correlation against the experiment. Higher is better, zero means no better than guessing the order, and negative means backwards.

The results do not support a single ranking of models. The benchmarks barely agree with each other, half the results sit at or below the point where the metric stops meaning anything, and no model is yet a usable predictor of any of these properties. All of that is on this page, before the leaderboard, because it is the more important finding.

Metric Spearman's ρ Updated Coverage Intervals

01

What is measured

These are developability assays: not whether an antibody binds its target, but whether it can be manufactured, stored and injected. An antibody that binds beautifully and aggregates in the vial is not a drug.

Each benchmark holds one assay. The model is given antibodies and asked to put them in order by the measured property; the score is how well that order matches the laboratory's. Rank correlation is the right metric for this because the useful question is which of these candidates is worse, not what the retention time is in minutes.

every model's score, on one axis

Each benchmark's distribution of Spearman correlations across models
Benchmark Every model, ρ from −0.30 to +0.70 Best ρ Best by Median ρ ρ<0.10

One dot per model. The vertical line is ρ = 0, where the predicted order is no better than random; dots left of it are ranked backwards. Click a dot to pin its model, or a benchmark name for its full ranking.
Pinning joins that model's dots down the column with a solid line, and hovering another joins it with a dashed one — two at a time, to compare. A line breaks at a benchmark the model has no result for, rather than crossing a value that does not exist.

Read the spread, not the leader. On most of these benchmarks the whole field lands in a band a tenth of a unit wide, close to zero — which is the shape of a task nobody has solved rather than a contest with a winner.
02

The leaderboard, and how far to trust it

Suite
Developer
Coverage

Ranked by mean ρ across the benchmarks in the current slice. The rank range column is the part worth reading: it is the best and worst position the same model takes on individual benchmarks. Almost every model in this table has been near the top of something and near the bottom of something else.

Models ranked by mean Spearman correlation
#

Positions are always by mean ρ, whichever column you sort by, so re-sorting never renumbers the standings. A model with results for fewer than 90% of the benchmarks in the slice is flagged with its count rather than dropped — and its mean is taken over the benchmarks it has, never over a zero it did not score. Mean rank is computed within each benchmark, so a model that skipped one is absent from that ranking rather than placed last.

Where each model actually lands best to worst rank, one row per model

The bar spans the best and worst rank that model reaches on any single benchmark; the notch is its mean rank. Rows are ordered by mean ρ, so the leaderboard's order runs top to bottom.

The bars overlap almost completely. If ability at this task were a stable property of a model, these bars would be short and would step neatly down the page. Instead the top model's range covers most of the field, which is what it looks like when the benchmark, not the model, decides the result.
03

The benchmarks disagree with each other

If these eight assays were all measuring "understands antibodies", a model that ranked well on one would rank well on the rest, and every cell below would be strongly positive. Take each pair of benchmarks, compare how they order the models they share, and this is what comes out.

Which of them contradict each other? the strongest links only — above the line they agree, below it they rank models in opposite orders

The assays that agree are not the ones grouped together. The links above the line join the surface-interaction assays — HIC, HAC, SMAC and both polyreactivity tests — into one cluster. In the full set, every link below the line involves SEC %Monomer or Titer: a model that ranks antibodies well for surface stickiness tends to rank them badly for aggregation and yield. Note that SEC and SMAC share a suite and sit side by side, and the arc between them still hangs below the line.

Do two benchmarks rank models alike? Spearman ρ between each pair's model rankings

Rank agreement between every pair of benchmarks

Computed in the browser over the models the two benchmarks have in common; pairs sharing fewer than three models are left empty rather than estimated. This is agreement between benchmarks, so it is on a deliberately different colour scale from every performance figure on this page — blue for alike, red for opposed, and the value printed in every cell.

A caution about what this does not show. Two benchmarks can disagree because they measure genuinely different chemistry — thermal stability and expression titer need not go together — or because neither has enough signal for its ordering to be stable. With no confidence intervals in the feed, these two explanations cannot be told apart here.
04

Every result

One row per model, one column per benchmark

The whole table, so nothing above has to be taken on trust. Colour marks the five interpretation bands; the two lowest bands share one recessive treatment because both mean the same thing — no usable signal — and are told apart by texture rather than by hue.

Spearman correlation for every model and benchmark

marks the best model for that benchmark. A · means that model has no result for that benchmark — not a failed one — so it is left out of the model's mean instead of counted as zero. Click a model's name to pin it; it stays marked in every figure on the page.

05

Does more thinking help?

Three of the models in the feed were run twice, at medium and high reasoning effort, on the same benchmarks. That is a controlled comparison — same model, same task, one variable — and it is the only one this data set contains.

Change in Spearman correlation from medium to high reasoning effort

Each cell is high minus medium on that benchmark: positive means the extra effort helped. These are single runs with no interval attached, so a cell near zero should be read as "no measurable change", not as a small real one.

06

Method, and what this cannot tell you

The metric

Spearman's ρ compares two orderings. It runs from −1 to +1: +1 is the same order as the laboratory, 0 is an unrelated order, −1 is exactly reversed. It ignores how far apart the values are, which is the point — a model does not need to predict a retention time in minutes to be useful for triage, it needs to put the risky candidates last.

The bands

The five bands used for colour throughout — inverted, negligible, weak, moderate, strong, split at 0, 0.10, 0.30 and 0.50 — are the conventional effect-size labels for a rank correlation, and they live in the data file rather than in the code, so moving a threshold is a data change.

Averaging

Mean ρ is the plain arithmetic mean over the benchmarks a model has a result for. Averaging correlations via Fisher's z is arguably more correct; here it changes no value by more than 0.015 but does swap three adjacent pairs of models. That it swaps anything at all is the clearest evidence on this page that positions this close are not real.

  • No confidence intervals, no sample sizes

    The feed carries a point estimate per model per benchmark and nothing else — not how many antibodies each ρ was computed over. Without that, no statement of the form "model A beats model B" is supportable at these margins, and this site does not make one. Slots exist in the data contract for intervals and counts; when they arrive, the figures will show them.

  • Coverage is partial by design

    Nothing here is a placeholder: the benchmarks that are present are complete and real. The ones that are missing simply have not been run yet.

  • Ragged model rosters

    A missing result is never treated as a zero, and models with gaps are flagged in the leaderboard, but a mean over six benchmarks and a mean over eight are still not quite the same measurement.

  • The assays are not equally hard

    Mean ρ gives every benchmark equal weight, which quietly rewards the ones where scores run high. Mean rank does not, which is why both columns are in the table and why they do not produce the same order.