A fifteen-part interactive tour of the geometry of probability. Visualize distributions as points, understand Fisher information as the ruler, and measure distance with KL divergence.
“The space of probability distributions is a landscape, and everything follows.”
Information geometry treats probability distributions as physical locations on a map. Instead of dealing with abstract numbers, we can visualize the shape of data, measure distances between distributions, and see how statistical algorithms naturally navigate this space.
These interactive explainers explore the Fisher metric, the hidden structure of exponential families, and how gradient descent moves across the terrain. Venture into the frontier, connecting these statistical shapes to string theory and Monstrous Moonshine.
A parameterised family of distributions is a curved surface, not a table. The Gaussian family is a 2D manifold whose points are the densities themselves; coordinate distance is the wrong way to measure how far apart they are.
The score function, its covariance, and the metric tensor it produces. For Gaussians, the Fisher metric is the Poincaré half-plane: geodesics are semicircles, not straight lines.
Fisher is the only Riemannian metric invariant under Markov morphisms (information-preserving data transformations). Euclidean distance on the simplex fails the test immediately.
The global companion to the Fisher metric. Forward KL gives mode-covering fits; reverse KL gives mode-seeking fits. The asymmetry is part of why VAEs blur and GANs mode-collapse.
Taylor-expand KL around coincident arguments and the second-order term is half the Fisher quadratic form. KL's rough edges (asymmetry, triangle failures) live at third order and above.
Amari's one-parameter family of connections ∇^(α). Two of them (e and m) are dual under the Fisher metric. Dually flat manifolds are exactly the exponential families.
Natural parameters, expectation parameters, and the Legendre duality between them. Gaussians, Bernoullis, Dirichlets — the workhorses of statistics are dually flat, which is what makes them tractable.
When an e-geodesic and an m-geodesic meet at a right angle, KL decomposes additively like squared Euclidean distance. The projection theorem behind maximum likelihood, EM, variational inference, and belief propagation.
Every strictly convex potential induces its own divergence and geometry. KL comes from negative Shannon entropy, squared Euclidean from half the norm squared, Itakura-Saito from negative log. All share the local-Fisher structure of Part 5.
Vanilla SGD is Euclidean descent on a non-Euclidean space. Pre-multiplying by the inverse Fisher matrix makes it parameterisation-invariant and second-order-like. Now standard in variational quantum circuits as Quantum Natural Gradient.
Voronoi diagrams on statistical manifolds use the divergence, not squared Euclidean. Gaussian cells warp into hyperbolic geometry; categorical cells become spherical; autoregressive data lives in Siegel space.
EM is alternating projection between two flat submanifolds: m-projection for the E-step, e-projection for the M-step. The Pythagorean theorem forces monotone convergence.
Moonshine's partition functions, moduli spaces, relative entropies, and modular invariants are fundamentally statistical objects. Information geometry has barely touched this territory.
The Weil-Petersson metric on a Calabi-Yau or K3 moduli space (from complex algebraic geometry) is the Fisher metric on the same moduli space read as a statistical manifold of normalised characters. For CP¹ instantons, both are AdS₃.
The j-function appears as an entropic potential in a Koszul-Souriau dually flat geometry on the Lie algebra of the Monster. Moonshine's coincidences reframed as moments of a thermal state. Speculative; rigor runs out.