Skip to main content

Subsection 1.3.1 Compositional Data and the CLR Transform

Note Id: 202608290001 | Tags: compositional data spurious correlations centered log-ratio (CLR)

Compositional Data.

Many scientific datasets naturally consist of compositions. Examples include molecular abundances, microbiome compositions, mineral fractions, metabolite concentrations, and probability distributions over discrete states.

Remark 1.3.2 Constant-Sum Constraint.

Compositional vectors satisfy a constant-sum constraint. Consequently, increasing one component necessarily decreases one or more of the other components. This creates spurious correlations when analyzed using conventional Euclidean methods.
Due to this geometric constraint, methods such as principal component analysis (PCA), singular value decomposition (SVD), and many clustering algorithms are not appropriate for analyzing compositional data.

The Centered Log-Ratio (CLR) Transform.

Definition 1.3.3 The CLR Transform.

The centered log-ratio (CLR) transform of the \(i^\text{th}\) element \(x_i\) in a compositional data vector
\begin{equation*} \mathbf{x} = (x_1, \dots , x_n) : \sum_i^n{x_i}=1 \end{equation*}
is defined as
\begin{equation*} \operatorname{CLR}(x_i) = \ln {\left(\frac{x_i}{g(\mathbf{x})} \right)} \end{equation*}
where \(g(\mathbf{x})\) denotes the geometric mean of the elements in \(\mathbf{x}\text{.}\)
After applying the CLR transform, conventional Euclidean techniques such as PCA, POD, SVD, clustering, and manifold learning become mathematically appropriate. This is because distances under the CLR transformation between two elements correspond with relative rather than absolute differences.