August 23, 2026 · Fault Diagnosis

The accuracy question

Machine learning (ML) is one of the most discussed directions in dissolved gas analysis (DGA). The proposition is simple: algorithms can interpret fault gases more consistently than rule-based methods. The reality is more nuanced — and the nuance matters to asset managers deciding whether to act on algorithm outputs for critical transformers.

This article examines what the published accuracy numbers actually say, where they come from, and how to deploy ML-assisted DGA diagnosis without giving up the safeguards that standards such as IEC 60599, IEEE C57.104, and GB/T 7252 provide.

Where the accuracy numbers come from

The accuracy figures that circulate for DGA interpretation methods originate from laboratory comparative studies. They are research comparisons, not vendor claims or product specifications — they measure how often each method matched a known fault type across a study dataset.

The most-quoted figures come from method comparisons run on the IEC TC 10 fault database. The key gas method is consistently the weakest, at roughly 42% correct. The Duval triangle is among the strongest: Duval’s own comparison on 122 database cases puts it at about 95%, and independent academic benchmarks generally land between 65% and 87%. Multi-method voting and machine-learning fusion are reported anywhere from about 76% to 98%, depending on sample size and balance.

Diagnostic method Approach Consistency (research comparison)
Key gas method Single-gas assessment; simple to implement ~42%
Rogers ratio method Gas-ratio coding ~62% (poorer for some combinations)
IEC ratio method Standardized ratio coding Only 11 of 27 code combinations defined
Duval triangle Intuitive; simple to implement ~95%
Duval pentagon Seven zones; finer fault coverage Best consistency (single method)
Multi-method voting / ML fusion Combines multiple sources and historical samples 76–98% (varies by study)

Figure sources: the Duval triangle figure is from M. Duval’s comparison on 122 cases from the IEC TC 10 fault database (95–96% for Duval triangles I and II); the key-gas figure comes from the same IEC TC 10 database comparisons. The ratio-method limitation is structural: the IEC ratio scheme defines only 11 of its 27 possible code combinations. Independent benchmarks report a wide spread — roughly 65–87% for the Duval triangle. These are fault-classification hit rates on research datasets, not measurement accuracy or product specifications.

Treat these as relative strengths, not absolute guarantees. And read the ML number with particular care: a model reporting 76% looks impressive beside a 42% key-gas baseline, but far less so beside a well-applied Duval triangle at 65–87%. Published ML results are also often obtained on small or unbalanced datasets, where the apparent gain can shrink or vanish on a different fleet. What ML reliably adds is consistency and triage — a research-comparison figure is not a promise that a specific product will reach that accuracy on your fleet.

What machine learning actually adds

ML contributes to DGA interpretation in three roles, and distinguishing them sets realistic expectations.

A fusion engine

ML takes the outputs of multiple traditional methods — key gas, ratios, Duval — plus concentration vectors, gas generation rates, and operating conditions (load, temperature, oil temperature), and feeds them into models such as random forest, XGBoost, fuzzy logic, or neural networks. The result is better class-discrimination than any single rule-based method.

A memory engine

Supervised learning on a utility’s own historical fault samples lets the model align with that fleet’s oil quality and operating characteristics. This is where ML is most valuable: it remembers how this utility’s transformers behave as they approach failure.

A filter

Before an alarm is raised, an ML model can perform multi-factor cross-checking, reducing false alarms caused by the transient fluctuation of a single gas.

The limits: labeled data and explainability

Engineering practice should respect three boundaries. First, model training depends on annotation quality and sample balance; transfer-generalization across manufacturers, oil types, and voltage classes is limited. A model trained on one fleet may not transfer cleanly to another.

Second, algorithm outputs can be a “black box.” Before procurement, require that any health-index or ML output be interpretable and its parameters configurable, so results can be verified rather than accepted on faith. Third, well-labeled training data remain scarce across the industry — a genuine constraint on accuracy claims.

Standards alongside, and expert sign-off

The safe deployment pattern is not “ML replaces standards.” It is ML presented side by side with, and corroborated by, IEC 60599, IEEE C57.104, and GB/T 7252. A fault conclusion should ultimately be reviewed by diagnostic personnel and confirmed with electrical tests and off-line laboratory analysis. The output is best framed as a hypothesis ranking with an evidence chain rather than a single verdict.

This fits the way online DGA data already flows through the decision chain — concentration, trend, gas generation rate, health index — where diagnostic methods are cross-validated against each other. For a practical look at combining methods, see our multi-method DGA diagnosis workflow and the IEC 60599 vs IEEE C57.104 comparison.

What this means for an online DGA program

ML fusion is only as good as the data it consumes. Continuous online DGA data — with complete timestamps, calibration traceability, and consistent measurement across the fleet — is the foundation that makes fusion, trend, and health-index outputs credible. A monitor that covers all nine gases plus moisture gives the model the full feature set that the IEC 60599 fault taxonomy requires, without gaps or reconstruction.

PAS DGA for AI-ready monitoring

The PAS DGA-900 online monitor measures 9 gases plus moisture with laser photoacoustic spectroscopy (L-PAS), needs no carrier gas and no consumables, and delivers continuous, standards-consistent data. That is exactly the input ML-assisted diagnosis needs — complete gas coverage feeding the interpretation chain you already trust.

Questions about how machine learning fits into your DGA program? Contact PAS DGA to discuss.