MSc dissertation, University of Bristol, 2025 to 2026
Why kinisi’s covariance matrices fail without warning
kinisi extracts diffusion coefficients from molecular dynamics trajectories. It does something unusually careful: it carries the full covariance matrix through its fit instead of treating each lag time as independent. That care is also the source of a fragility. In a minority of runs the matrix stops being positive definite, which quietly degrades the fitted coefficient and its uncertainty without raising any error.
Rather than patch the symptom, I built a random-walk testbed with a known analytical answer, so that a numerical artefact could be separated from a real physical signal. The work is supervised by Dr Andrew McCluskey, one of kinisi’s developers.
- Matrix failure rate fell from 90% to 2% as sampling increased, showing the matrices were under-determined rather than wrong.
- Root cause isolated to sample noise in long lag-time variance estimates, through a systematic convergence study.
- Minimum-eigenvalue, ridge, linear and non-linear shrinkage reconditioning benchmarked against the analytical matrix.
- A data-adaptive eigenvalue floor recovered parameters better than the library’s fixed default on the testbed.
- Validation across noise levels is ongoing, and the adaptive rule has only been tested offline.
- A consistency check is open upstream as pull request #222, awaiting maintainer review.
Python, NumPy, MDAnalysis, eigenvalue analysis, matrix conditioning
Generalised least squares needs the inverse covariance
$$\hat{\beta} = \left(X^{\top}\Sigma^{-1}X\right)^{-1}X^{\top}\Sigma^{-1}y$$Σ is the MSD covariance. Inversion fails once the smallest eigenvalue reaches zero.