Information-geometric hypothesis testing on the categorical statistical manifold: A geodesic test statistic for anomaly detection in decision support systems
International Journal of Approximate Reasoning, cilt.199, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 199
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.ijar.2026.109830
- Dergi Adı: International Journal of Approximate Reasoning
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, MathSciNet, zbMATH, DIALNET, Academic Search Ultimate (EBSCO)
- Anahtar Kelimeler: Anomaly detection, Decision support systems, Decision under uncertainty, Fisher–Rao distance, Hypothesis testing, Information geometry, Statistical manifold, Uncertainty quantification
- Hatay Mustafa Kemal Üniversitesi Adresli: Evet
Özet
We study hypothesis testing for categorical monitoring data using the squared Fisher–Rao distance Λn=ndFR(p^n,p0)2. The statistic is an exact strictly monotone transformation of the Freeman–Tukey/Hellinger statistic, the λ=−1/2 member of the Cressie–Read power-divergence family. Consequently, with correspondingly transformed critical values the two statistics define the same test, power function, and ROC ranking. The geometric formulation instead provides an intrinsic representation and extends naturally from a point null to a smooth null submanifold. We give a self-contained derivation of the χK−12 point-null limit and obtain explicit O(n−1) coefficients for both the mean and variance on the geodesic scale. The resulting mean-matching Bartlett-type scaling improves small-sample calibration. More importantly, writing P1=∑ip0,i−1, we show analytically that the residual second-cumulant coefficient is v−4b=(9P1−K2−12K+4)/6>0 for every interior null distribution and K ≥ 2. Thus no single scalar can match both first two chi-square cumulants to this order. For a smooth composite null M 0, the distance-to-submanifold statistic has a χr2 limit with r=codimM0; the independence submanifold illustrates detection of marginal-preserving dependence changes. Simulations assess calibration and compare the intrinsic statistic with a coordinate-dependent Euclidean benchmark. A Hurwicz loss illustrates cost-sensitive threshold selection, while an NSL-KDD application shows why an empirical null is needed under real-data overdispersion.