← All series

Information Geometry

A fifteen-part interactive tour of the geometry of probability. Visualize distributions as points, understand Fisher information as the ruler, and measure distance with KL divergence.

“The space of probability distributions is a landscape, and everything follows.”

Information geometry treats probability distributions as physical locations on a map. Instead of dealing with abstract numbers, we can visualize the shape of data, measure distances between distributions, and see how statistical algorithms naturally navigate this space.

These interactive explainers explore the Fisher metric, the hidden structure of exponential families, and how gradient descent moves across the terrain. Venture into the frontier, connecting these statistical shapes to string theory and Monstrous Moonshine.

The Map distributions as points, Fisher as ruler
01

Parameterised Distributions as a Manifold

A parameterised family of distributions is a curved surface, not a table. The Gaussian family is a 2D manifold whose points are the densities themselves; coordinate distance is the wrong way to measure how far apart they are.

Warm-up · Distinguishability
02

The Fisher Information Metric

The score function, its covariance, and the metric tensor it produces. For Gaussians, the Fisher metric is the Poincaré half-plane: geodesics are semicircles, not straight lines.

Fisher · Poincaré half-plane
03

Chentsov's Uniqueness Theorem

Fisher is the only Riemannian metric invariant under Markov morphisms (information-preserving data transformations). Euclidean distance on the simplex fails the test immediately.

Uniqueness · Markov invariance
04

Kullback-Leibler Divergence

The global companion to the Fisher metric. Forward KL gives mode-covering fits; reverse KL gives mode-seeking fits. The asymmetry is part of why VAEs blur and GANs mode-collapse.

KL · Forward vs reverse
05

Local KL Equals Fisher

Taylor-expand KL around coincident arguments and the second-order term is half the Fisher quadratic form. KL's rough edges (asymmetry, triangle failures) live at third order and above.

Taylor expansion · Bregman
The Hidden Structure exponential families, projections, natural gradients
06

α-Connections and Duality

Amari's one-parameter family of connections ∇^(α). Two of them (e and m) are dual under the Fisher metric. Dually flat manifolds are exactly the exponential families.

Connections · Duality
07

Exponential Families and Dual Flatness

Natural parameters, expectation parameters, and the Legendre duality between them. Gaussians, Bernoullis, Dirichlets — the workhorses of statistics are dually flat, which is what makes them tractable.

Legendre · Log-partition
08

The Generalised Pythagorean Theorem

When an e-geodesic and an m-geodesic meet at a right angle, KL decomposes additively like squared Euclidean distance. The projection theorem behind maximum likelihood, EM, variational inference, and belief propagation.

Pythagoras · MLE as projection
09

Bregman Divergences

Every strictly convex potential induces its own divergence and geometry. KL comes from negative Shannon entropy, squared Euclidean from half the norm squared, Itakura-Saito from negative log. All share the local-Fisher structure of Part 5.

Bregman · Mirror descent
10

Natural Gradient Descent

Vanilla SGD is Euclidean descent on a non-Euclidean space. Pre-multiplying by the inverse Fisher matrix makes it parameterisation-invariant and second-order-like. Now standard in variational quantum circuits as Quantum Natural Gradient.

NGD · Mirror descent
Exploring the Frontier hyperbolic Voronoi, EM, and the unclaimed territory
11

Hyperbolic and Bregman Voronoi Diagrams

Voronoi diagrams on statistical manifolds use the divergence, not squared Euclidean. Gaussian cells warp into hyperbolic geometry; categorical cells become spherical; autoregressive data lives in Siegel space.

Voronoi · Bregman clustering
12

The EM Algorithm as Alternating Projection

EM is alternating projection between two flat submanifolds: m-projection for the E-step, e-projection for the M-step. The Pythagorean theorem forces monotone convergence.

EM · e- and m-projections
13

Information Geometry Meets Moonshine

Moonshine's partition functions, moduli spaces, relative entropies, and modular invariants are fundamentally statistical objects. Information geometry has barely touched this territory.

Frontier · Pivot
14

Fisher as Weil-Petersson on Moduli Spaces

The Weil-Petersson metric on a Calabi-Yau or K3 moduli space (from complex algebraic geometry) is the Fisher metric on the same moduli space read as a statistical manifold of normalised characters. For CP¹ instantons, both are AdS₃.

Frontier · Weil-Petersson
15

Koszul-Souriau and the j-Function as Thermodynamic Entropy

The j-function appears as an entropic potential in a Koszul-Souriau dually flat geometry on the Lie algebra of the Monster. Moonshine's coincidences reframed as moments of a thermal state. Speculative; rigor runs out.

Frontier · Speculative closer