Many scientific datasets naturally consist of compositions. Examples include molecular abundances, microbiome compositions, mineral fractions, metabolite concentrations, and probability distributions over discrete states.
Subsection 1.3.1 Compositional Data and the CLR Transform
Note Id:
202608290001 | Tags: compositional data spurious correlations centered log-ratio (CLR)
Compositional Data.
Remark 1.3.2 Constant-Sum Constraint.
Compositional vectors satisfy a constant-sum constraint. Consequently, increasing one component necessarily decreases one or more of the other components. This creates spurious correlations when analyzed using conventional Euclidean methods.
Due to this geometric constraint, methods such as principal component analysis (PCA), singular value decomposition (SVD), and many clustering algorithms are not appropriate for analyzing compositional data.
The Centered Log-Ratio (CLR) Transform.
The centered log-ratio (CLR) transform can be used to overcome the constant-sum constraint. It acts as a map from the original simplex into an unconstrained Euclidean space as it effectively compares each element of the composition to the geometric mean of the entire composition.
Definition 1.3.3 The CLR Transform.
The centered log-ratio (CLR) transform of the \(i^\text{th}\) element \(x_i\) in a compositional data vector
\begin{equation*}
\mathbf{x} = (x_1, \dots , x_n) : \sum_i^n{x_i}=1
\end{equation*}
is defined as
\begin{equation*}
\operatorname{CLR}(x_i) = \ln {\left(\frac{x_i}{g(\mathbf{x})} \right)}
\end{equation*}
After applying the CLR transform, conventional Euclidean techniques such as PCA, POD, SVD, clustering, and manifold learning become mathematically appropriate. This is because distances under the CLR transformation between two elements correspond with relative rather than absolute differences.