<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Multiple Testing | Sieun Kim Portfolio</title><link>https://siuunni.github.io/tags/multiple-testing/</link><atom:link href="https://siuunni.github.io/tags/multiple-testing/index.xml" rel="self" type="application/rss+xml"/><description>Multiple Testing</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>ko-kr</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://siuunni.github.io/media/icon_hu_1c0e9cb08cfb822a.png</url><title>Multiple Testing</title><link>https://siuunni.github.io/tags/multiple-testing/</link></image><item><title>Unsupervised Conformal Novelty Detection for Hierarchical Data</title><link>https://siuunni.github.io/publications/unsupervised-conformal-novelty-detection/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://siuunni.github.io/publications/unsupervised-conformal-novelty-detection/</guid><description>&lt;hr&gt;
&lt;div style="display:grid;grid-template-columns:160px 1fr;gap:1.5rem 2rem;margin-bottom:2rem;"&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Abstract&lt;/div&gt;
&lt;div&gt;
&lt;p&gt;Electric vehicle (EV) battery packs exhibit a natural pack–module–cell hierarchy, which induces dependence among measurements within the same module. Such hierarchical dependence poses challenges for the direct application of conventional novelty detection methods. To address these challenges, we develop conformal e-value procedures for hierarchical novelty detection, with the goal of controlling the false discovery rate (FDR) at both the group and unit levels. We combine hierarchical conformal score construction with eBH and U-eBH multiple testing procedures, and consider split conformal, full conformal, and group-wise conformal regimes. The proposed methods are evaluated through simulation studies and an analysis of EV battery pack data, where they provide empirical FDR control and detect localized module- and cell-level irregularities.&lt;/p&gt;
&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Type&lt;/div&gt;
&lt;div&gt;Preprint&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Publication&lt;/div&gt;
&lt;div&gt;To be submitted to &lt;em&gt;Journal of the Korean Statistical Society&lt;/em&gt; (JKSS, SCIE)&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Keywords&lt;/div&gt;
&lt;div&gt;novelty detection · hierarchical structure · conformal inference · false discovery rate · multiple testing&lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id="motivation"&gt;Motivation&lt;/h2&gt;
&lt;p&gt;Real-world industrial data — such as EV battery manufacturing — presents three compounding challenges that standard anomaly detection cannot handle jointly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hierarchical dependence:&lt;/strong&gt; Cells within the same module share a common group effect, violating the IID assumption of most methods.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No anomaly labels:&lt;/strong&gt; Ground-truth defect labels are expensive or infeasible to obtain in manufacturing, ruling out supervised approaches.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multiple comparisons:&lt;/strong&gt; Simultaneously testing hundreds of units inflates false discoveries without a principled error-control mechanism.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="method-overview"&gt;Method Overview&lt;/h2&gt;
&lt;p&gt;We frame hierarchical novelty detection as a &lt;strong&gt;multiple hypothesis testing problem&lt;/strong&gt; at two levels simultaneously:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Group-level (HC-GND)&lt;/strong&gt; — Does a test module contain any anomalous cells?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unit-level (HC-UND)&lt;/strong&gt; — Which specific cells within each module are anomalous?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Feature extraction:&lt;/strong&gt; Charging voltage time series are treated as functional observations. Functional PCA (FPCA) projects each cell&amp;rsquo;s charging curve onto a low-dimensional score vector, placing cells from all packs on a common feature space.&lt;/p&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_charge_curves.png"
alt="Preprocessed charging voltage curves for all cells in the six test battery packs"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
Preprocessed charging voltage curves for all cells in the six test packs. Normal packs (blue) show regular charging behavior; Abnormal packs 4 and 5 (red) exhibit irregular patterns that are difficult to distinguish visually at the pack level.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_fpca_scatter.png"
alt="FPCA scatter plot separating normal and abnormal battery packs"
style="width:80%;display:block;margin:0 auto;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
Two-dimensional FPCA score vectors under the split conformal regime. Normal reference packs (black) form a tight cluster, while abnormal test packs (red) appear in a distinct region, validating FPCA as a discriminative feature extraction step.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Nonconformity score:&lt;/strong&gt; For each cell, we compute a &lt;strong&gt;Mahalanobis-distance-based nonconformity score&lt;/strong&gt; relative to a robust trimmed-mean group center. This naturally respects within-group dependence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multiple testing:&lt;/strong&gt; Scores are converted to &lt;strong&gt;conformal e-values&lt;/strong&gt; — non-negative statistics satisfying E[e] ≤ 1 under the null — and tested jointly via the &lt;strong&gt;eBH / U-eBH procedure&lt;/strong&gt;, guaranteeing FDR control under minimal distributional assumptions.&lt;/p&gt;
&lt;p&gt;We implement three conformal regimes with different data-efficiency / robustness trade-offs:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: left"&gt;Regime&lt;/th&gt;
&lt;th style="text-align: left"&gt;Training data&lt;/th&gt;
&lt;th style="text-align: left"&gt;Contamination risk&lt;/th&gt;
&lt;th style="text-align: left"&gt;Theoretical FDR guarantee&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HSC&lt;/strong&gt; (Split)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference only&lt;/td&gt;
&lt;td style="text-align: left"&gt;None&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Both levels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HFC&lt;/strong&gt; (Full)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference + all test&lt;/td&gt;
&lt;td style="text-align: left"&gt;Higher&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Group level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HGC&lt;/strong&gt; (Group-wise)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference + one test group&lt;/td&gt;
&lt;td style="text-align: left"&gt;Moderate&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Group level&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="results"&gt;Results&lt;/h2&gt;
&lt;h3 id="simulation-study"&gt;Simulation Study&lt;/h3&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_h_gnd.png"
alt="HC-GND simulation results: FDR and power under varying outlier proportion and signal strength"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;HC-GND simulation results.&lt;/strong&gt; Empirical FDR (top) and power (bottom) under varying outlier proportions (left) and signal strengths (right). All three regimes maintain FDR below the target α = 0.1. HFC and HGC achieve higher power than HSC, especially at low outlier proportions, by exploiting more calibration data.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_comparison.png"
alt="HC-UND vs non-hierarchical baselines: FDR and power comparison"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;HC-UND vs. non-hierarchical baselines.&lt;/strong&gt; The proposed hierarchical methods (HSC, HFC, HGC) consistently outperform non-hierarchical counterparts (SC, FC, AdaDetect) in power while maintaining empirical FDR control. Unit-level anomalies that are subtle in the pooled population become detectable within their own group context.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Key findings across 1,000 simulation replicates (α = 0.1):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Empirical FDR remains &lt;strong&gt;at or below the target level&lt;/strong&gt; across all outlier proportions and signal strengths.&lt;/li&gt;
&lt;li&gt;At weak signal, power reaches ~&lt;strong&gt;40%&lt;/strong&gt;; as signal increases, power &lt;strong&gt;approaches 100%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Hierarchical methods &lt;strong&gt;outperform non-hierarchical baselines&lt;/strong&gt; by leveraging within-group structure.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="ev-battery-pack-application"&gt;EV Battery Pack Application&lt;/h3&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_pack_wise.png"
alt="Pack-wise comparison of hierarchical and non-hierarchical detection methods"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;Pack-wise detection results (α = 0.1).&lt;/strong&gt; Left three columns: hierarchical HC-GND/HC-UND results (yellow = module rejection, red = cell rejection). Right three columns: non-hierarchical baselines (SVM, Isolation Forest, AdaDetect). The proposed methods identify localized irregularities in Abnormal Pack 5 across multiple modules and cells that non-hierarchical methods largely miss.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul&gt;
&lt;li&gt;The proposed procedures detect &lt;strong&gt;localized module- and cell-level irregularities&lt;/strong&gt; not captured by pack-level labels.&lt;/li&gt;
&lt;li&gt;Non-hierarchical baselines concentrate false detections on Normal Pack 0 and miss structured signals in Abnormal Pack 5.&lt;/li&gt;
&lt;li&gt;Results provide an &lt;strong&gt;additional diagnostic layer&lt;/strong&gt; when only coarse pack-level labels are available.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="presentation-slides"&gt;Presentation Slides&lt;/h2&gt;
&lt;div style="margin-top:1rem;"&gt;
&lt;img
src="https://siuunni.github.io/uploads/papers/thesis-page-4.png"
alt="Presentation slide 1 — Overview"
style="width:100%;border:1px solid rgba(148,163,184,.3);border-radius:8px;margin-bottom:1.25rem;"
&gt;
&lt;img
src="https://siuunni.github.io/uploads/papers/thesis-page-5.png"
alt="Presentation slide 2 — Real data application"
style="width:100%;border:1px solid rgba(148,163,184,.3);border-radius:8px;"
&gt;
&lt;/div&gt;</description></item></channel></rss>