9  Sparse designs

Two to three observations per subject.

This is the regime of the motivating application and the one in which method comparison is most frequently misdesigned.

9.1 What is available

At \(T = 2\): a single rate per subject and region. Population-level estimation of one- or two-parameter rate laws, network diffusion, and latent time models, all through pooling. Supervised learning on change, with the cross-section supplying the information.

At \(T = 3\): additionally, one curvature estimate per subject and region, the observed first-interval rate as a feature for supervised models, and — most importantly — the possibility of training on the first interval and evaluating on the second.

9.3 Evaluation

At \(T = 3\), the primary evaluation is to train on the first interval and predict the third visit. This is the only design in the sparse regime that separates the arms on something other than in-sample fit, because it tests extrapolation, where the mechanistic and supervised arms differ by construction (Chapter 3).

At \(T = 2\) this design is unavailable within a single cohort. The alternatives are:

  • An external cohort with longer follow-up, used only for evaluation.
  • A held-out subset of subjects with a third visit, if the study is unbalanced.
  • Restriction of the claim to description rather than prediction, with the comparison reported on population-level fit and calibration rather than forecast accuracy.

Reporting a three-arm forecast comparison from a strictly two-visit design, with evaluation by cross-validation within the same interval, does not support conclusions about the relative merits of the arms.

9.4 What can and cannot be claimed

Supportable in this regime:

  • Population-average rate of change per region, with between-subject variation.
  • Ordering of regions by rate, subject to the measurement floor.
  • Association of baseline features with subsequent rate.
  • Relative forecast accuracy over one interval, against stated baselines.
  • Calibration of prediction intervals.

Not supportable:

  • That a particular nonlinear functional form governs individual trajectories.
  • Subject-specific dynamical parameters, unless the shrinkage diagnostic shows they are data-determined.
  • Directional coupling between regions, in the sense of one region’s change driving another’s. Cross-sectional covariance is not evidence of dynamical coupling at \(T = 2\).
  • Discovery of governing equations.

9.5 Expected outcome

The result that should be anticipated, and is worth reporting whether or not it is the result hoped for: the supervised arm performs comparably or better on short-horizon interpolation; the mechanistic arm performs better on extrapolation and on interval calibration; the learned-parameterization hybrid performs best overall; and in a substantial fraction of regions no method outperforms the no-change baseline because the measurement error exceeds the effect.