6  Spatial structure

Why regions are not exchangeable variables.

Chapter 5 treats the variable count \(R\) and the driver count \(d\) in general terms. In brain-wide parcellations the variables carry additional structure: they have positions, neighbours, and connections, and their change is spatially correlated for both biological and methodological reasons. This chapter treats that structure — what it invalidates, what it enables, and how to avoid assuming it.

6.1 Sources of spatial correlation

Change in neighbouring regions is correlated through at least four mechanisms, which must be separated because they have opposite implications:

  1. Shared biological process. Regions belonging to a common functional or structural system change together. This is the signal.
  2. Connectivity-mediated propagation. Change in one region is followed by change in regions connected to it, which need not be spatially adjacent.
  3. Partial volume and boundary effects. Adjacent regions share tissue at their border; segmentation error in one appears as opposite-signed error in the other.
  4. Registration error and smoothing. Misregistration displaces signal coherently across neighbouring regions, and any smoothing kernel imposes correlation by construction.

Mechanisms 3 and 4 are methodological and produce correlation that is local and decays monotonically with distance. Mechanism 2 produces correlation that follows connectivity and can be long-range. A useful diagnostic follows directly: compare the distance-decay profile of the change covariance against that of the noise field estimated from scan–rescan data (Chapter 4). Correlation that persists at distances where the noise field has decorrelated is candidate biological structure; correlation confined to the noise field’s range is not.

6.2 Consequences

The effective number of independent variables is far below \(R\). Standard false-discovery procedures assume a dependence structure that spatially correlated data violate, so nominal error rates are wrong in a direction that depends on the parcellation.

Apparent low-dimensional structure is partly induced. Smoothing and registration impose correlation, which appears in a decomposition as leading components. The whitening step of Chapter 5 is what separates this from real structure; it is not optional when the variables are spatially arranged.

Effective sample size for spatial contrasts is smaller than \(R\) suggests. A contrast between two large systems is supported by far fewer independent units than the number of regions in them.

6.3 Parcellation is a modeling assumption

\(R\) is chosen, not given. A parcellation asserts that its boundaries align with the units over which the process of interest is homogeneous. Where they do not, a region spanning two processes dilutes both, and adjacent regions sharing one process appear coupled when they are one unit split in half.

This should be treated as an assumption to be probed, not a preprocessing detail. Repeating the primary analysis across parcellations of different resolution and construction, and reporting the stability of conclusions, is the minimum. Different morphometric measures also behave differently under the same parcellation: Gennatas et al. (2017) reported divergent age trajectories for gray matter density, volume, mass, and cortical thickness over the same regions, so the choice of measure is not separable from the choice of parcellation.

6.4 Candidate spatial decompositions

Three sources of prior spatial structure are available in brain data, each with supporting evidence.

Intrinsic functional connectivity networks. Seeley et al. (2009) showed that distinct neurodegenerative syndromes produce atrophy circumscribed within distinct healthy intrinsic connectivity networks, and additionally that regions functionally connected within those networks show correlated gray matter volumes across healthy individuals.

Structural covariance networks. Zielinski et al. (2010) showed that networks derived from covariance of regional gray matter volume recapitulate intrinsic connectivity architecture and are detectable during development.

The structural connectome. Diffusion-derived connectivity, used as the Laplacian \(\mathbf{L}\) in the network diffusion model of Chapter 7.

These three are related but not identical, and the correspondence between them is an empirical finding with limits rather than a definition. Choosing among them is a modeling decision with consequences for every downstream result, and it should be stated and, where the design permits, tested rather than defaulted to.

6.5 Testing a spatial prior

The general argument of Chapter 5 applies: a model constrained by a spatial prior will fit plausibly whether or not that structure governs the data, because the structure is supplied by the constraint. The controls have a spatial-specific form.

The null must preserve spatial autocorrelation. This is the point most often missed. Permuting region labels destroys spatial smoothness, so any smooth empirical map beats such a null trivially and the test has no power to distinguish a specific structure from generic smoothness. Valid nulls preserve the nuisance property while destroying the feature of interest: degree-preserving rewiring for a connectome, position-shuffling that retains connection profiles, and rotation-based procedures that preserve spatial autocorrelation. Zheng et al. (2019) used rewired and spatial nulls together and reported the empirical network outperforming both across all densities; Váša and Mišić (2022) catalogue the available nulls and the property each preserves.

Report which null was used and what it preserves. A result stated against an unspecified null is uninterpretable, because the strength of the claim is entirely determined by how much structure the null retained.

6.6 Discriminating among spatial mechanisms

Where several mechanisms predict spatially organized change, they are separated by deriving distinct predictions and comparing them on the same data, rather than by fitting one and reporting that it fits.

Zhou et al. (2012) is the template. Four candidate mechanisms — nodal stress, transneuronal spread, trophic failure, and shared vulnerability — each imply a different relationship between a region’s position relative to a disease epicenter and its degree of atrophy. Evaluated against five neurodegenerative syndromes, regions with shorter functional paths to the epicenter showed greater atrophy, a pattern best fitting transneuronal spread.

The methodological point generalizes beyond neurodegeneration: a mechanism is supported by out-predicting its rivals on a shared target, not by achieving an acceptable fit in isolation.

6.7 Spatially informed estimation

Spatial structure is not only a nuisance to be controlled; it is information that mitigates the loss of per-region signal-to-noise as \(R\) grows (Chapter 5).

  • Neighbourhood priors. Conditional autoregressive or Gaussian Markov random field priors on region-level parameters borrow strength from neighbours, reducing variance most in exactly the small, low-reliability regions that a fine parcellation produces.
  • Connectivity-based neighbourhoods. Where propagation is connectivity-mediated, the relevant neighbourhood is the connectome’s, not physical adjacency. The prior should match the hypothesized mechanism.
  • Distance-decay and spatially varying coefficients. Allow effects to vary smoothly over the cortex rather than independently per region, at a parameter cost far below \(R\).

These are the spatial analogue of pooling across regions in Chapter 8, and they are frequently more effective than increasing \(N\).

6.8 Worked example at small \(T\)

Brown et al. (2019) forecast region-wise future gray matter loss in individual patients with behavioural variant frontotemporal dementia and semantic variant primary progressive aphasia. Patients contributed roughly three scans at approximately annual intervals. Patient-specific epicenters were identified, and two network-derived quantities — shortest path length to the epicenter, and cumulative atrophy in connected neighbours — entered a generalized additive model predicting subsequent regional atrophy.

This is a concrete instance of the pattern this book recommends for sparse designs: the mechanism supplies the features and the flexible model fits them, rather than either component operating alone. In the taxonomy of Chapter 9 it sits between H2 and H4. It is also a realistic calibration of achievable accuracy: the reported within-scan variance explained was moderate, which is consistent with the measurement floor argument of Chapter 4 rather than a shortcoming of the approach.

6.9 Interaction with \(T\)

Spatial priors supply structure that temporal sampling cannot. At \(T = 2\)\(3\) the data contain no within-subject evidence about coupling (Chapter 3), so any coupled model must obtain its structure externally. This is precisely why connectivity-constrained specifications are the strongest available option in sparse designs, and equally why their apparent success is uninformative without the null controls above.