11 Sparse designs
Two to three observations per subject.
This is the regime of the motivating application and the one in which method comparison is most frequently misdesigned.
11.1 What is available
At \(T = 2\): a single rate per subject and region. Population-level estimation of one- or two-parameter rate laws, network diffusion, and latent time models, all through pooling. Supervised learning on change, with the cross-section supplying the information.
At \(T = 3\): additionally, one curvature estimate per subject and region, the observed first-interval rate as a feature for supervised models, and — most importantly — the possibility of training on the first interval and evaluating on the second.
11.2 Recommended analysis structure
- Establish the measurement floor (Chapter 4) and stratify all subsequent results by per-region signal-to-noise ratio.
- Fit the exponential mixed model as the reference mechanistic specification. Report it on the log scale.
- Add latent time. This is the step that permits any statement about trajectory shape in this regime.
- Add the structural prior. Network diffusion contributes a coupled model at a parameter cost that the design can support.
- Fit the supervised arm on change, pooled across regions, with explicit handling of regression toward the mean.
- Fit H2, the learned parameterization, as the primary hybrid.
- Fit H1, the learned residual, as the interpretable hybrid decomposition.
- Ensemble the three arms and report it as a fourth condition.
H3 may be included at \(T = 3\) with strong regularization, reported with its ablation. It should not be the headline result in this regime.
11.3 Evaluation
At \(T = 3\), the primary evaluation is to train on the first interval and predict the third visit. This is the only design in the sparse regime that separates the arms on something other than in-sample fit, because it tests extrapolation, where the mechanistic and supervised arms differ by construction (Chapter 3).
At \(T = 2\) this design is unavailable within a single cohort. The alternatives are:
- An external cohort with longer follow-up, used only for evaluation.
- A held-out subset of subjects with a third visit, if the study is unbalanced.
- Restriction of the claim to description rather than prediction, with the comparison reported on population-level fit and calibration rather than forecast accuracy.
Reporting a three-arm forecast comparison from a strictly two-visit design, with evaluation by cross-validation within the same interval, does not support conclusions about the relative merits of the arms.
11.4 A worked example
Brown et al. (2019) forecast region-wise gray matter loss in individual patients from roughly three annual scans, using patient-specific epicenters and two network-derived features — shortest path length to the epicenter and cumulative atrophy in connected neighbours — in a generalized additive model. It is an instance of the recommended pattern at this density: the mechanism supplies the features, the flexible model fits them. Chapter 6 discusses it further.
11.5 What can and cannot be claimed
Supportable in this regime:
- Population-average rate of change per region, with between-subject variation.
- Ordering of regions by rate, subject to the measurement floor.
- Association of baseline features with subsequent rate.
- Relative forecast accuracy over one interval, against stated baselines.
- Calibration of prediction intervals.
Not supportable:
- That a particular nonlinear functional form governs individual trajectories.
- Subject-specific dynamical parameters, unless the shrinkage diagnostic shows they are data-determined.
- Directional coupling between regions, in the sense of one region’s change driving another’s. Cross-sectional covariance is not evidence of dynamical coupling at \(T = 2\).
- Discovery of governing equations.
11.6 Expected outcome
The result that should be anticipated, and is worth reporting whether or not it is the result hoped for: the supervised arm performs comparably or better on short-horizon interpolation; the mechanistic arm performs better on extrapolation and on interval calibration; the learned-parameterization hybrid performs best overall; and in a substantial fraction of regions no method outperforms the no-change baseline because the measurement error exceeds the effect.