Built with Claude — Life Sciences

Causal and Interpretable Modeling of Influenza Antigenic Distance

Separating the hemagglutinin positions that causally drive antigenic escape from the linked positions that merely ride along with them — across three hemagglutination-inhibition datasets, with an interpretable B-spline Kolmogorov–Arnold Network for prediction.

←  Back to the article
01 — The problem

ML predicts escape well. It can't say why.

Models read antigenic distance off HA sequence accurately — but they flag positions that predict titer, not ones that causally drive escape. The obstacle is linkage: in H3N2's dense phylogeny a passenger mutation looks just as predictive as the true driver.

Driver

Physically disrupts antibody binding.

Change it and the virus escapes. This is what we want to isolate.

Hitchhiker

Just travels alongside a driver.

Predicts escape by association, but does nothing at the interface.

02 — Why it matters

From measuring escape to anticipating it.

Stakes

Vaccine effectiveness

Antigenic distance sets how well the vaccine works — and drives the twice-yearly vaccine-strain decision.

Bottleneck

The HI assay is slow

The gold-standard measurement is costly and hard to reproduce across labs — motivating sequence-based models.

The step

Cause, not correlation

Separating drivers from hitchhikers is what turns measuring escape into anticipating it.

03 — The data
Two independent H3N2 panels — replication by design.
Distribution of log2 HI titer for both datasets
Read it: the HI cross-titer is the target — lower titer = more escape. Two independent H3N2 datasets: VHID (n = 2,751) and Bedford 2014 (n = 7,808), the latter carrying collection-year metadata for temporal tests. Both rebuilt from raw pair tables into per-position HA1 feature matrices, verified byte-for-byte (SHA-256). Agreeing across both panels is itself the evidence.
04 — From sequences to features
Turning raw sequences into two feature encodings.
How the feature matrices are built from raw influenza HI data
Read it: raw HI titer tables and GenBank HA sequences are cleaned and HA1-aligned, then every position is encoded two ways — a binary mismatch flag for causal discovery, and a gap-aware Grantham chemical distance for prediction (the KAN). One row per virus–reference pair, with the HI titer as the target.
05 — Finding the causes

Break the convoys, then ask what causes the signal.

Collapse

Residual linkage → 0

Merge co-evolving positions (|φ|≥0.8) into blocks. Strong residual pairs (|φ|≥0.9) drop to zero — restoring the conditions causal search needs.

Discover

Titer as a sink

PC, GES, FCI with HI titer pinned as an effect, never a cause; each position ranked by 200× bootstrap stability.

Calibrate

A well-behaved test

A permutation null shows the Fisher-z test holds near-nominal type-I error at α = 0.01 on binary, left-censored data.

06 — Interpretable vs black box
The glass box nearly matches the black box.
Cross-validated R-squared with 95% CIs for LASSO, Ridge, XGBoost, KAN
Read it: 5×4 repeated cross-validation, matched folds. The interpretable B-spline KAN trails black-box XGBoost by only 0.025 R² (VHID) / 0.028 (Bedford) — while exposing per-position response curves XGBoost can't. It recovers R² ≈ 0.99 on synthetic data, and its learned curves are monotone: bigger substitution → more escape.
07 — The drivers
It lands on the sites biology already knew.
Discovered titer-sink causal graph for both datasets
Read it: each arrow is a direct-cause candidate into HI titer; thickness = stability. From titers alone, no structural priors, the high-stability survivors — VHID 156, 189 and Bedford 133, 158, 189 — land in classical head antigenic sites A and B, ringing the receptor-binding site. Reproducibly, across both datasets.
08 — Four methods, one answer
Position 189 surfaces no matter how you look.
Cross-method convergence across causal, KAN, XGBoost, univariate screens
Read it: a position is red when causal search flags it and it clears every predictive screen (KAN, XGBoost, univariate). Position 189 is the single most robust signal — it survives i.i.d., by-virus, and by-serum resampling. As a textbook cluster-transition determinant (Koel 2013), recovering it structure-free confirms the features track real antigenic biology.
09 — The payoff plot
Drivers and hitchhikers, cleanly separated.
Selection stability versus adjusted effect quadrant plot
Read it: selection stability (y) vs. backdoor-adjusted effect (x). Red = stable, large-effect drivers (VHID 158, 189, 289; Bedford 2, 133, 189). Blue = a stable hitchhiker (VHID 156) — near-fixed, selected every time but with a small non-robust effect. The quadrant does the separating the linkage problem demanded.
10 — How big, and which way
Adjusted effects shrink — but never flip sign.
Backdoor-adjusted per-position causal effects
Read it: dots = backdoor-adjusted effect on titer, × = raw. Adjustment shrinks every one of the 13 parents toward zero yet never flips its sign — consistent with removing phylogenetic confounding rather than manufacturing the signal. Almost all are negative: larger substitution → more escape.
11 — Positions that escape together
Epistasis, captured and tested.
Second-order KAN pairwise interaction surfaces
Read it: the second-order KAN plots interacting position pairs — blue = a pair jointly lowers titer beyond their separate effects (synergistic escape), red = raises it. 10 of 16 nominated pairs show a significant interaction (4/8 VHID, 6/8 Bedford) — structure a black box hides.
12 — What we won't overclaim
Reproducible selection is not proven cause.
Forward-in-time prediction of future antigenic clusters
Read it: train on the past, test on the next 5 years. Accuracy is unstable and, in the worst window, negative (XGBoost 0.60 → −0.37). As drift crosses into clusters the model never saw, prediction fails — so we report selection and prediction separately, and treat the recovered drivers as strong hypotheses still to be tested.
13 — What's next

From recovered drivers to a definitive test.

Prove it

Reverse genetics

Single-substitution mutagenesis at predicted drivers — start with 193 and 158, the richest reservoirs of clean contrasts.

Resolve blocks

Break the linkage

Denser, linkage-breaking serum panels and deep-mutational-scanning escape maps to split within-block ambiguity.

Better target

Cartographic distance

Model antigenic-map distance instead of raw HI titer — separating avidity from true antigenic change.

Generalize

Beyond H3N2

Extend to other subtypes and lineages, and test whether the recovered structure transports to prospective vaccine-strain selection.

Takeaway

A trustworthy lens on how flu escapes.

From antibody titers alone, no structural priors, the pipeline rediscovers the exact residues structural biology already implicates — reproducibly, interpretably, and without overclaiming. Prediction says a strain will escape; this says where to look.

Every figure here regenerates from raw data — one executable notebook.

Read the full article  →
The full story

Read the full article.

Every figure, table, and number in this deck — with the full method, the causal derivation, and the interactive two-column reader.

1 / 16
↑ ↓  or  scroll