<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Conformal Prediction | Sieun Kim Portfolio</title><link>https://siuunni.github.io/tags/conformal-prediction/</link><atom:link href="https://siuunni.github.io/tags/conformal-prediction/index.xml" rel="self" type="application/rss+xml"/><description>Conformal Prediction</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>ko-kr</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://siuunni.github.io/media/icon_hu_1c0e9cb08cfb822a.png</url><title>Conformal Prediction</title><link>https://siuunni.github.io/tags/conformal-prediction/</link></image><item><title>Unsupervised Conformal Novelty Detection for Hierarchical Data</title><link>https://siuunni.github.io/publications/unsupervised-conformal-novelty-detection/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://siuunni.github.io/publications/unsupervised-conformal-novelty-detection/</guid><description>&lt;hr&gt;
&lt;div style="display:grid;grid-template-columns:160px 1fr;gap:1.5rem 2rem;margin-bottom:2rem;"&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Abstract&lt;/div&gt;
&lt;div&gt;
&lt;p&gt;Electric vehicle (EV) battery packs exhibit a natural pack–module–cell hierarchy, which induces dependence among measurements within the same module. Such hierarchical dependence poses challenges for the direct application of conventional novelty detection methods. To address these challenges, we develop conformal e-value procedures for hierarchical novelty detection, with the goal of controlling the false discovery rate (FDR) at both the group and unit levels. We combine hierarchical conformal score construction with eBH and U-eBH multiple testing procedures, and consider split conformal, full conformal, and group-wise conformal regimes. The proposed methods are evaluated through simulation studies and an analysis of EV battery pack data, where they provide empirical FDR control and detect localized module- and cell-level irregularities.&lt;/p&gt;
&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Type&lt;/div&gt;
&lt;div&gt;Preprint&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Publication&lt;/div&gt;
&lt;div&gt;To be submitted to &lt;em&gt;Journal of the Korean Statistical Society&lt;/em&gt; (JKSS, SCIE)&lt;/div&gt;
&lt;div style="font-weight:700;padding-top:.1rem;"&gt;Keywords&lt;/div&gt;
&lt;div&gt;novelty detection · hierarchical structure · conformal inference · false discovery rate · multiple testing&lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id="motivation"&gt;Motivation&lt;/h2&gt;
&lt;p&gt;Real-world industrial data — such as EV battery manufacturing — presents three compounding challenges that standard anomaly detection cannot handle jointly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hierarchical dependence:&lt;/strong&gt; Cells within the same module share a common group effect, violating the IID assumption of most methods.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No anomaly labels:&lt;/strong&gt; Ground-truth defect labels are expensive or infeasible to obtain in manufacturing, ruling out supervised approaches.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multiple comparisons:&lt;/strong&gt; Simultaneously testing hundreds of units inflates false discoveries without a principled error-control mechanism.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="method-overview"&gt;Method Overview&lt;/h2&gt;
&lt;p&gt;We frame hierarchical novelty detection as a &lt;strong&gt;multiple hypothesis testing problem&lt;/strong&gt; at two levels simultaneously:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Group-level (HC-GND)&lt;/strong&gt; — Does a test module contain any anomalous cells?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unit-level (HC-UND)&lt;/strong&gt; — Which specific cells within each module are anomalous?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Feature extraction:&lt;/strong&gt; Charging voltage time series are treated as functional observations. Functional PCA (FPCA) projects each cell&amp;rsquo;s charging curve onto a low-dimensional score vector, placing cells from all packs on a common feature space.&lt;/p&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_charge_curves.png"
alt="Preprocessed charging voltage curves for all cells in the six test battery packs"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
Preprocessed charging voltage curves for all cells in the six test packs. Normal packs (blue) show regular charging behavior; Abnormal packs 4 and 5 (red) exhibit irregular patterns that are difficult to distinguish visually at the pack level.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_fpca_scatter.png"
alt="FPCA scatter plot separating normal and abnormal battery packs"
style="width:80%;display:block;margin:0 auto;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
Two-dimensional FPCA score vectors under the split conformal regime. Normal reference packs (black) form a tight cluster, while abnormal test packs (red) appear in a distinct region, validating FPCA as a discriminative feature extraction step.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Nonconformity score:&lt;/strong&gt; For each cell, we compute a &lt;strong&gt;Mahalanobis-distance-based nonconformity score&lt;/strong&gt; relative to a robust trimmed-mean group center. This naturally respects within-group dependence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multiple testing:&lt;/strong&gt; Scores are converted to &lt;strong&gt;conformal e-values&lt;/strong&gt; — non-negative statistics satisfying E[e] ≤ 1 under the null — and tested jointly via the &lt;strong&gt;eBH / U-eBH procedure&lt;/strong&gt;, guaranteeing FDR control under minimal distributional assumptions.&lt;/p&gt;
&lt;p&gt;We implement three conformal regimes with different data-efficiency / robustness trade-offs:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: left"&gt;Regime&lt;/th&gt;
&lt;th style="text-align: left"&gt;Training data&lt;/th&gt;
&lt;th style="text-align: left"&gt;Contamination risk&lt;/th&gt;
&lt;th style="text-align: left"&gt;Theoretical FDR guarantee&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HSC&lt;/strong&gt; (Split)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference only&lt;/td&gt;
&lt;td style="text-align: left"&gt;None&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Both levels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HFC&lt;/strong&gt; (Full)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference + all test&lt;/td&gt;
&lt;td style="text-align: left"&gt;Higher&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Group level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;strong&gt;HGC&lt;/strong&gt; (Group-wise)&lt;/td&gt;
&lt;td style="text-align: left"&gt;Reference + one test group&lt;/td&gt;
&lt;td style="text-align: left"&gt;Moderate&lt;/td&gt;
&lt;td style="text-align: left"&gt;✓ Group level&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="results"&gt;Results&lt;/h2&gt;
&lt;h3 id="simulation-study"&gt;Simulation Study&lt;/h3&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_h_gnd.png"
alt="HC-GND simulation results: FDR and power under varying outlier proportion and signal strength"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;HC-GND simulation results.&lt;/strong&gt; Empirical FDR (top) and power (bottom) under varying outlier proportions (left) and signal strengths (right). All three regimes maintain FDR below the target α = 0.1. HFC and HGC achieve higher power than HSC, especially at low outlier proportions, by exploiting more calibration data.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_comparison.png"
alt="HC-UND vs non-hierarchical baselines: FDR and power comparison"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;HC-UND vs. non-hierarchical baselines.&lt;/strong&gt; The proposed hierarchical methods (HSC, HFC, HGC) consistently outperform non-hierarchical counterparts (SC, FC, AdaDetect) in power while maintaining empirical FDR control. Unit-level anomalies that are subtle in the pooled population become detectable within their own group context.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Key findings across 1,000 simulation replicates (α = 0.1):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Empirical FDR remains &lt;strong&gt;at or below the target level&lt;/strong&gt; across all outlier proportions and signal strengths.&lt;/li&gt;
&lt;li&gt;At weak signal, power reaches ~&lt;strong&gt;40%&lt;/strong&gt;; as signal increases, power &lt;strong&gt;approaches 100%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Hierarchical methods &lt;strong&gt;outperform non-hierarchical baselines&lt;/strong&gt; by leveraging within-group structure.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="ev-battery-pack-application"&gt;EV Battery Pack Application&lt;/h3&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/fig_pack_wise.png"
alt="Pack-wise comparison of hierarchical and non-hierarchical detection methods"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
&lt;strong&gt;Pack-wise detection results (α = 0.1).&lt;/strong&gt; Left three columns: hierarchical HC-GND/HC-UND results (yellow = module rejection, red = cell rejection). Right three columns: non-hierarchical baselines (SVM, Isolation Forest, AdaDetect). The proposed methods identify localized irregularities in Abnormal Pack 5 across multiple modules and cells that non-hierarchical methods largely miss.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul&gt;
&lt;li&gt;The proposed procedures detect &lt;strong&gt;localized module- and cell-level irregularities&lt;/strong&gt; not captured by pack-level labels.&lt;/li&gt;
&lt;li&gt;Non-hierarchical baselines concentrate false detections on Normal Pack 0 and miss structured signals in Abnormal Pack 5.&lt;/li&gt;
&lt;li&gt;Results provide an &lt;strong&gt;additional diagnostic layer&lt;/strong&gt; when only coarse pack-level labels are available.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="presentation-slides"&gt;Presentation Slides&lt;/h2&gt;
&lt;div style="margin-top:1rem;"&gt;
&lt;img
src="https://siuunni.github.io/uploads/papers/thesis-page-4.png"
alt="Presentation slide 1 — Overview"
style="width:100%;border:1px solid rgba(148,163,184,.3);border-radius:8px;margin-bottom:1.25rem;"
&gt;
&lt;img
src="https://siuunni.github.io/uploads/papers/thesis-page-5.png"
alt="Presentation slide 2 — Real data application"
style="width:100%;border:1px solid rgba(148,163,184,.3);border-radius:8px;"
&gt;
&lt;/div&gt;</description></item><item><title>배터리 전압·온도 시계열 기반 이상 셀 탐지</title><link>https://siuunni.github.io/projects/battery-anomaly-cell/</link><pubDate>Fri, 12 Dec 2025 00:00:00 +0000</pubDate><guid>https://siuunni.github.io/projects/battery-anomaly-cell/</guid><description>&lt;div style="display:grid;grid-template-columns:130px 1fr;gap:1rem 1.5rem;margin-bottom:2rem;"&gt;
&lt;div style="font-weight:700;"&gt;개요&lt;/div&gt;
&lt;div&gt;배터리팩 충전 과정의 &lt;strong&gt;전압 및 온도&lt;/strong&gt; 시계열을 함수형 데이터로 변환하고, 데이터 특성에 맞춰 FPCA와vd-FPCA로 차원을 줄인 뒤, GMM 기반 등각 예측(Conformal Prediction)으로 통계적으로 보장된 이상 탐지 기준을 설계한 프로젝트입니다.&lt;/div&gt;
&lt;div style="font-weight:700;"&gt;연구 기간&lt;/div&gt;
&lt;div&gt;2024.05~2025.03&lt;/div&gt;
&lt;div style="font-weight:700;"&gt;데이터&lt;/div&gt;
&lt;div&gt;KAMP 「전기차 배터리 충전 실험 데이터(품질보증)」 — 셀 176개(16 모듈 × 11 셀) 전압, 온도 측정점 32개(모듈당 2개). 학습은 정상만, 테스트는 정상,이상 혼합&lt;/div&gt;
&lt;div style="font-weight:700;"&gt;기술 스택&lt;/div&gt;
&lt;div&gt;Python,R,scikit-fda, scikit-learn,&lt;/div&gt;
&lt;div style="font-weight:700;"&gt;성과&lt;/div&gt;
&lt;div&gt;전압·온도 이상을 &lt;strong&gt;위양성 없이 전수 식별&lt;/strong&gt;(실제,시뮬레이션 데이터 모두 F1 ≈ 1), 유의수준 $\alpha=0.01$ 에서 통계적으로 보장된 커버리지 달성&lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id="문제-정의"&gt;문제 정의&lt;/h2&gt;
&lt;p&gt;제조 현장의 이상 탐지는 두 가지 제약이 있습니다. &lt;strong&gt;① 이상 데이터를 얻기 어렵고, ② 라벨을 얻기는 더 어렵습니다.&lt;/strong&gt; 이 배터리팩 데이터도 같은 상황이었습니다.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;학습 데이터:&lt;/strong&gt; 전부 &lt;strong&gt;정상&lt;/strong&gt; 상태의 충전 과정만 수집 — 176개 셀 전압 + 32개 온도&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;테스트 데이터:&lt;/strong&gt; 정상·이상이 섞여 있지만, 이상 샘플엔 &lt;strong&gt;이상 발생 시점과 &amp;ldquo;이상이다&amp;quot;라는 이진 라벨만&lt;/strong&gt; 제공. 이상의 &lt;strong&gt;구체적 유형(온도/전압)은 알려주지 않음&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;판정 규칙:&lt;/strong&gt; 전압이든 온도든 한쪽이라도 이상이면 그 배터리팩 전체를 이상으로 분류&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;즉 정상 데이터만으로 정상의 분포를 학습하고, 새 샘플이 그 분포를 벗어나는지를 판단하는 &lt;strong&gt;단일 클래스(novelty) 탐지&lt;/strong&gt; 문제로 접근했습니다. 그리고 &amp;ldquo;몇 % 신뢰수준에서 이상이다&amp;quot;라고 말할 수 있도록, 통계적 보장이 있는 판정 기준이 필요했습니다.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="배터리팩-계층-구조"&gt;배터리팩 계층 구조&lt;/h2&gt;
&lt;p&gt;배터리팩은 계층적으로 구성됩니다. 팩 1개 = &lt;strong&gt;모듈 16개&lt;/strong&gt;, 모듈 1개 = &lt;strong&gt;셀 11개&lt;/strong&gt;, 따라서 팩 전체는 &lt;strong&gt;176개 셀&lt;/strong&gt;입니다. 온도는 모듈마다 센서 2개가 달려 총 32개 측정점을 갖습니다. 이 &lt;strong&gt;셀–모듈–팩 계층&lt;/strong&gt;을 그대로 분석 단위에 반영했습니다(온도는 모듈당 2개 센서를 평균내 모듈 대표 온도로 사용).&lt;/p&gt;
&lt;figure style="margin:1.5rem auto;max-width:560px;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/battery_pack.png"
alt="전기 자동차용 배터리 팩-모듈-셀"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);background:#fff;padding:.5rem;"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
배터리 팩은 여러 개의 모듈로 구성되며, 각 모듈은 다수의 셀로 이루어진 계층적 구조를 가졌습니다.[사진 출처: 삼성 SDI]
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;hr&gt;
&lt;h2 id="데이터-전처리"&gt;데이터 전처리&lt;/h2&gt;
&lt;p&gt;함수형 데이터 분석은 모든 곡선이 같은 충전 사이클을 담고 있어야 합니다. 그래서 전처리의 핵심은 실제 충전 구간만 정확히 절단하는 것이었습니다.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;① 기본 절단.&lt;/strong&gt; 각 곡선에서 전압 및 온도의 &lt;strong&gt;최솟값 인덱스를 충전 시작점, 최댓값 인덱스를 충전 완료점&lt;/strong&gt;으로 잡아 완전한 충전 사이클을 추출했습니다.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;② 변화점 탐지로 정밀 절단.&lt;/strong&gt; 그런데 일부 데이터는 충전 초반에 &lt;strong&gt;최저 전압이 한동안 평평하게 유지&lt;/strong&gt;되는 구간이 있었습니다. 단순히 최솟값의 첫 인덱스를 시작점으로 잡으면, 충전 전 평형 상태까지 끌려 들어가는 문제가 생겼습니다. 그래서 이런 경우에만 선택적으로 &lt;strong&gt;변화점 탐지(ruptures 라이브러리)&lt;/strong&gt; 를 적용해, 감지된 변화점 중 &lt;strong&gt;가장 이른 시점&lt;/strong&gt;을 실제 충전 시작점으로 삼았습니다. 모든 데이터에 적용하면 계산 비용이 크기 때문에, &lt;strong&gt;문제가 되는 곡선에만 선택 적용&lt;/strong&gt;해 효율적으로 처리했습니다.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;③ 길이 통일.&lt;/strong&gt; 곡선마다 측정 길이가 달라, 모두 &lt;strong&gt;1초 단위 격자로 선형보간&lt;/strong&gt;하고, 가장 긴 곡선을 기준으로 짧은 곡선은 마지막 값을 연장해 길이를 맞췄습니다. 다만 &lt;strong&gt;온도는 센서별 변화 양상이 너무 달라&lt;/strong&gt; 같은 방식으로 늘리면 본래 패턴이 손상되고 탐지 성능이 떨어졌습니다. 그래서 온도에는 &lt;strong&gt;가변 도메인 함수형 주성분 분석(vd-FPCA)&lt;/strong&gt; 을 적용했습니다. vd-FPCA는 길이가 다른 곡선을 공통 격자로 억지로 맞추지 않고, &lt;strong&gt;도메인 길이를 변수로 한 삼변량 평활(penalized thin-plate spline)로 공분산을 추정하고 길이에 조건부로 고유분해&lt;/strong&gt;하는 방법입니다(
). 덕분에 측정 길이가 제각각인 온도 곡선의 고유한 패턴을 보존하면서 차원을 줄일 수 있었습니다.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="함수형-분석"&gt;함수형 분석&lt;/h2&gt;
&lt;p&gt;전압과 온도는 데이터 성격이 달랐습니다. 전압은 곡선 길이가 비교적 일정해 일반 FPCA를 적용했고, 온도는 센서마다 측정 길이가 제각각이라 길이 차이를 다룰 수 있는 vd-FPCA를 적용했습니다.&lt;/p&gt;
&lt;h3 id="전압--fpca--mahalanobis-거리"&gt;전압 — FPCA + Mahalanobis 거리&lt;/h3&gt;
&lt;p&gt;먼저 데이터를 &lt;strong&gt;팩 단위 5:5로 Train / Calibration 으로 분할&lt;/strong&gt;한 뒤, 모든 전압 곡선을 함수형 데이터로 변환했습니다. Train 세트에서 &lt;strong&gt;FPCA&lt;/strong&gt;를 학습해 고유함수(eigenfunction) 집합을 추정하고, Calibration,Test 곡선은 같은 함수공간으로 &lt;strong&gt;정사영&lt;/strong&gt;해 2차원 FPCA 점수 공간에 표현했습니다.&lt;/p&gt;
&lt;p&gt;이어 Train 세트에서 모듈 안 셀들의 전압 벡터로 평균 벡터와 공분산 행렬을 추정했습니다. 팩 $i$, 모듈 $j$, 셀 $k$의 전압 벡터를 $y_{ijk}$ 라 하면 모듈 단위 평균은&lt;/p&gt;
$$\bar{y}_{ij} = \frac{1}{K} \sum_{k=1}^{K} y_{ijk},$$&lt;p&gt;Train 세트 공분산 행렬 $\Sigma$ 는&lt;/p&gt;
$$\Sigma = \frac{1}{|D_{\mathrm{tr}}|}\sum_{i \in D_{\mathrm{tr}}} \sum_{j=1}^{M}\frac{1}{K} \sum_{k=1}^{K}(y_{ijk} - \bar{y}_{ij})(y_{ijk} - \bar{y}_{ij})^{\top}.$$&lt;p&gt;이 평균과 공분산 통해 각 셀이 모듈 분포에서 얼마나 벗어났는지를 &lt;strong&gt;Mahalanobis 거리&lt;/strong&gt;로 측정해, 셀 단위 이상 점수를 정의했습니다. 셀–모듈 계층을 반영하므로, 팩이 이상으로 판정됐을 때 &lt;strong&gt;어느 셀이 원인인지까지 해석&lt;/strong&gt;할 수 있습니다.&lt;/p&gt;
&lt;h3 id="온도--vd-fpca"&gt;온도 — vd-FPCA&lt;/h3&gt;
&lt;p&gt;온도 곡선엔 vd-FPCA를 적용한 후 GMM을 통해 정상 집단과 이상 집단을 군집화했습니다. 주성분 공간에서 온도 이상 데이터는 정상 군집과 멀리 떨어진 $(-25,\,26)$ 부근에 &lt;strong&gt;고립되어 나타났습니다.&lt;/strong&gt; 이런 뚜렷한 분리는 &lt;strong&gt;vd-FPCA가 길이가 다른 온도 곡선의 정상·이상 패턴 차이를 효과적으로 포착&lt;/strong&gt;했음을 보여줍니다.&lt;/p&gt;
&lt;figure style="margin:1.5rem auto;max-width:560px;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/battery_temp.png"
alt="온도 vd-FPCA 주성분 점수 공간"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);background:#fff;padding:.5rem;"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
온도 vd-FPCA 점수 공간. 정상(train·cal·test_temp_ok)은 한 군집을 이루고, 온도 이상(test_temp_ng)은 $(-25,26)$ 부근에 고립됩니다.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;hr&gt;
&lt;h2 id="등각-예측"&gt;등각 예측&lt;/h2&gt;
&lt;p&gt;차원을 줄인 점수 공간 위에서, &lt;strong&gt;등각 예측(Conformal Prediction)&lt;/strong&gt; 으로 이상을 판정했습니다.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;왜 등각 예측인가.&lt;/strong&gt; 이 데이터는 이상 샘플이 적고 라벨도 부족합니다. 등각 예측은 점 예측 대신 예측 집합(prediction set)을 내놓는데, 교환가능성(exchangeability)이라는 최소한의 가정만으로 불확실성을 정량화합니다. 핵심은 &lt;strong&gt;모델의 유효성과 무관하게 예측 집합이 참값을 포함할 확률이 설정한 수준 이상으로 유지된다&lt;/strong&gt;는 점입니다. 덕분에 분포 가정을 강하게 두기 어렵고 데이터 수가 적은 이번 같은 상황에서도 유효하게 작동하고, &amp;ldquo;$\alpha$ 수준에서 정상을 이상으로 잘못 볼 확률&amp;quot;을 통제할 수 있습니다.&lt;/p&gt;
&lt;p&gt;적합성 점수와 예측 집합 구성 방식이 다른 &lt;strong&gt;세 가지 임계값 설정 방법&lt;/strong&gt;을 모두 적용해 강건성을 검증했습니다.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;정상 테스트 데이터&lt;/strong&gt; — 세 방법 모두에서 예측 영역 &lt;strong&gt;내부&lt;/strong&gt;에 위치해 정상으로 올바르게 분류&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;이상 데이터&lt;/strong&gt; — 세 방법 모두에서 예측 영역을 &lt;strong&gt;명확히 벗어나&lt;/strong&gt; 이상으로 정확히 식별. 특히 한 이상 샘플은 정상 군집에서 상당히 멀리 떨어져, 높은 신뢰도로 탐지됨&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;위양성(false positive) 0건&lt;/strong&gt; — 세 방법 모두에서 완벽한 분류&lt;/li&gt;
&lt;/ul&gt;
&lt;figure style="margin:1.5rem auto;max-width:600px;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/battery_voltage.png"
alt="전압 적합성 점수 분포와 1% 임계값"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);background:#fff;padding:.5rem;"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
전압 적합성 점수 분포. 정상(ok)은 낮은 값에 모이고 이상(test_ng)은 큰 값으로 퍼지며, $\alpha=0.01$ 임계값(26.65)을 기준으로 이상이 분리됩니다.
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;→ 전압·온도 양쪽 모두에서 $\alpha=0.01$ 수준의 &lt;strong&gt;통계적으로 보장된 커버리지를 유지하면서&lt;/strong&gt;, 실제·시뮬레이션 데이터 모두 &lt;strong&gt;F1 ≈ 1&lt;/strong&gt; 의 높은 정확도를 달성했습니다.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="시뮬레이션-검증"&gt;시뮬레이션 검증&lt;/h2&gt;
&lt;p&gt;실제 이상 데이터가 적어, 평균 이동·최댓값/최솟값 이동·시간가변 이동 등 다양한 이상 시나리오를 시뮬레이션으로 생성해 방법론의 강건성을 검증했습니다. 대부분의 시나리오에서 &lt;strong&gt;Accuracy·F1·MCC가 0.94~0.99 수준&lt;/strong&gt;으로 안정적이었습니다.&lt;/p&gt;
&lt;figure style="margin:1.5rem 0;"&gt;
&lt;img src="https://siuunni.github.io/uploads/papers/battery_sim.png"
alt="시뮬레이션 데이터 성능 요약 표"
style="width:100%;border-radius:8px;border:1px solid rgba(148,163,184,.25);background:#fff;padding:.5rem;"&gt;
&lt;figcaption style="font-size:.8rem;color:#94a3b8;margin-top:.5rem;text-align:center;"&gt;
시뮬레이션 이상 시나리오별 성능 요약 (Accuracy·Precision·Recall·F1·MCC).
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;hr&gt;
&lt;h2 id="핵심-요약"&gt;핵심 요약&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;라벨 없는 제조 데이터에 맞춘 설계:&lt;/strong&gt; 정상만으로 분포를 학습하는 novelty 탐지로 접근하고, 셀–모듈–팩 계층 구조를 분석 단위에 반영했습니다.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;전압·온도를 특성에 맞게:&lt;/strong&gt; 길이가 일정한 전압엔 FPCA + Mahalanobis 거리를, 길이가 제각각인 온도엔 vd-FPCA를 적용해 각 신호의 패턴을 살렸습니다.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;통계적으로 보장된 판정:&lt;/strong&gt; GMM 기반 등각 예측으로 세 가지 임계값 방법 모두에서 &lt;strong&gt;위양성 없이&lt;/strong&gt; 이상을 전수 식별(F1 ≈ 1)했고, $\alpha=0.01$ 의 커버리지 보장을 확보했습니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="참고문헌"&gt;참고문헌&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Johns, J. T., Crainiceanu, C., Zipunnikov, V., &amp;amp; Gellar, J. (2019).
. &lt;em&gt;Journal of Computational and Graphical Statistics&lt;/em&gt;, 28(4), 993–1006.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="출처"&gt;출처&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;배터리 팩 사진
&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>