DC converter transformer fault prediction system based on intelligent power grid

By utilizing a fault prediction system for DC converter transformers in smart grids, multi-parameter sensors and deep forest models are employed to achieve accurate fault diagnosis and proactive early warning for DC converter transformers. This solves technical problems that traditional technologies cannot address, enabling a leap from rough judgment to precise decision-making and enhancing the intelligence level and safety and reliability of power grid asset management.

CN122046006APending Publication Date: 2026-05-15STATE GRID SHANDONG ELECTRIC POWER CO YUNCHENG POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANDONG ELECTRIC POWER CO YUNCHENG POWER SUPPLY CO
Filing Date
2026-01-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional DC converter transformer maintenance strategies cannot achieve proactive early warning, resulting in unplanned operation, high repair costs, and problems of "over-maintenance" or "under-maintenance," which affect power grid safety and economy.

Method used

The fault prediction system for DC converter transformers based on smart grids adopts a decomposition-prediction-reconstruction strategy. By fusing multi-parameter sensors, SCADA data, historical data, and environmental data, and combining deep forest models and expert rules, it achieves accurate fault diagnosis and forward-looking prediction.

Benefits of technology

It enables the extraction of potential trends and periodic patterns from complex data, improving the accuracy and reliability of fault diagnosis, providing a clear decision window, transforming into pre-intervention, reducing the risk of misjudgment, and improving the initiative and security of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122046006A_ABST
    Figure CN122046006A_ABST
Patent Text Reader

Abstract

The invention discloses a DC converter transformer fault prediction system based on an intelligent power grid, and relates to the technical field of DC converter transformers. Comprising a multi-parameter sensor unit, an SCADA data interface unit, an equipment ledger and historical data management unit and an environment data monitoring unit. The feature extraction module is used for processing and converting original data into useful information and comprises a data preprocessing and cleaning unit, a multi-source data fusion unit and a feature engineering unit; and the fault diagnosis and prediction analysis module is responsible for carrying out deep analysis and prospective prediction on the state of the direct current converter transformer and comprises a fault diagnosis model unit, a fault prediction model unit and a health state evaluation unit. The traditional'post remedy 'or'in-event emergency' is converted into'pre-intervention ', so that the initiative and safety of operation and maintenance are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of DC converter transformer technology, specifically a fault prediction system for DC converter transformers based on smart grids. Background Technology

[0002] As the "heart" of the high-voltage direct current (HVDC) transmission system, the operational reliability of the DC converter transformer directly affects the safety and stability of the entire inter-regional power grid. This equipment operates under complex electromagnetic and mechanical stress environments of high voltage and high current for extended periods, making its internal insulation system prone to progressive degradation, potentially leading to catastrophic failure and causing significant economic losses and social impact. Traditional equipment maintenance strategies are mainly divided into two categories: reactive maintenance, which involves repairing equipment after a failure occurs. This approach is highly unplanned, has extremely high repair costs, and incurs substantial losses due to power outages. Regular preventative maintenance involves inspecting, testing, and maintaining equipment at fixed time intervals. While this approach offers some preventative benefits, it lacks specificity and may result in either "over-maintenance" (equipment in good condition being shut down for maintenance) or "under-maintenance" (sudden failures during the maintenance cycle), leading to unsatisfactory economic efficiency and safety.

[0003] To this end, the present invention provides a fault prediction system for DC converter transformers based on smart grids. Summary of the Invention

[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a fault prediction system for DC converter transformers based on smart grids. This invention achieves truly proactive early warning through a "decomposition-prediction-reconstruction" strategy. It not only accurately extracts potential trends and periodic patterns from complex and noisy gas data, but more importantly, it can quantify and predict the trajectory of fault development. When the system issues a warning that "the threshold is expected to be exceeded in 5 days," maintenance personnel no longer receive a simple status alarm, but rather a decision window with a clear timeframe. This transforms traditional "post-event remediation" or "in-event emergency response" into "pre-event intervention," significantly improving the initiative and safety of maintenance, thereby solving the technical problems described in the background section.

[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a fault prediction system for DC converter transformers based on smart grids, comprising: The data acquisition module is used to comprehensively collect various status information of the converter transformer, including a multi-parameter sensor unit, a SCADA data interface unit, an equipment ledger and historical data management unit, and an environmental data monitoring unit. The feature extraction module is used to transform raw data into useful information, including a data preprocessing and cleaning unit, a multi-source data fusion unit, and a feature engineering unit. The fault diagnosis and prediction analysis module is responsible for in-depth analysis and forward-looking prediction of the condition of DC converter transformers, including fault diagnosis model unit, fault prediction model unit and health status assessment unit.

[0006] Furthermore, the multi-parameter sensor unit directly monitors the physical and chemical state of the converter transformer by deploying high-frequency or ultrasonic partial discharge sensors, dissolved gas analysis sensors in oil, temperature sensors, vibration sensors, and core grounding current sensors on the converter transformer. The main insulating materials inside a transformer are insulating oil and solid insulation. During normal operation, they will slowly age due to heat and electrical stress, producing a small amount of gas. However, when an internal fault occurs, the fault point will generate local high temperature or high energy discharge, which will rapidly accelerate the chemical decomposition of these insulating materials and produce a large amount of characteristic gas.

[0007] Under high temperature or discharge energy, the carbon-hydrogen bond (CH) and carbon-carbon bond (CC) of insulating oil will break, generating free hydrogen atoms and hydrocarbon radicals. These radicals recombine to form hydrogen gas (H2) and a series of low-molecular-weight hydrocarbon gases, such as methane (CH4), ethane (C2H6), ethylene (C2H4), and acetylene (C2H2).

[0008] Solid insulation is mainly composed of carbon, hydrogen, and oxygen. When overheated, it undergoes pyrolysis, primarily producing carbon monoxide (CO) and carbon dioxide (CO2).

[0009] Different types and severity of faults will produce different kinds and proportions of gas. Each gas acts like a specific "warning light," indicating a certain type of fault. Hydrogen (H2) is produced by the decomposition of water molecules in insulating oil during low-energy discharges, such as partial discharges and low-temperature corona discharges. It is a sensitive indicator of low-energy discharges and arc discharges, but it can also be produced by a variety of other factors, so it needs to be considered in conjunction with other gases for accurate judgment.

[0010] Methane (CH4) and ethane (C2H6) are produced by the decomposition of insulating oil at low to medium temperatures (~150℃~500℃). They usually indicate general overheating faults, such as poor contact of tap changers or eddy current overheating caused by multiple grounding points of the iron core.

[0011] Ethylene (C₂H₄) is the main characteristic gas produced by insulating oil when it is overheated to above 700°C. Its production rate increases sharply with increasing temperature, a clear indication of severe overheating. When ethylene becomes the dominant hydrocarbon gas, it indicates that the temperature of the overheating point is extremely high.

[0012] Acetylene (C2H2) is only produced in large quantities by insulating oil during high-energy arc discharges, where temperatures can reach over 3000℃. It is almost never generated during normal overheating and low-energy discharges, making it the sole and decisive indicator of arc discharge. The presence of a significant amount of acetylene in the oil indicates a potential serious internal discharge fault, such as short circuits between coil turns or layers, or flashover of leads to ground—a very dangerous signal.

[0013] Carbon monoxide (CO) and carbon dioxide (CO2) are the byproducts of the pyrolysis and aging of solid insulating materials, reflecting the deterioration of the solid insulation. When the content of hydrocarbon gases increases significantly at the same time, it indicates that the overheating fault not only involves the oil but has also endangered the solid insulation of the coil, making the fault more severe.

[0014] The SCADA data interface unit obtains real-time operating parameters of load current, voltage, and power from the power grid SCADA system; the equipment ledger and historical data management unit centrally manages the equipment model, design parameters, maintenance records, test reports, and family defect information; and the environmental data monitoring unit collects data on ambient temperature, humidity, and pollution level.

[0015] Furthermore, the multi-source data fusion unit performs time alignment and integration of heterogeneous data from different sensors and systems to form a unified and continuous panoramic view of the device status; the feature engineering unit extracts key features from the processed data.

[0016] Furthermore, the fault diagnosis model unit uses intelligent algorithms to identify and classify the current fault; it adopts a Deep Forest model to perform a preliminary diagnosis of the fault nature of the transformer based on oil chromatography data; the feature vector is input into the pre-trained Deep Forest model, which first generates feature subsets with multiple random dimensions to capture patterns under different feature combinations. The data passes through multiple forest layers in sequence, with each layer receiving the feature information processed by the previous layer and further optimizing the feature representation. The last forest layer outputs a probability distribution, representing the probability of belonging to each fault type. The system selects the category with the highest probability as the preliminary diagnosis result.

[0017] Furthermore, upon receiving the initial diagnostic results, the inference engine immediately searches the association rule base, retrieving rules related to all conclusions. These rules are then triggered and placed into the rule pool to be evaluated. The system begins to examine each rule in the pool, checking whether the conditions of its IF part are satisfied by the current data. It selects the result with the highest confidence. After completing comprehensive inference, the inference engine generates the final precise location diagnostic report. The report includes the final diagnostic conclusion, the most likely fault location, and the overall confidence level.

[0018] Furthermore, the fault prediction model unit adopts two complementary paradigms: data-driven and physical mechanism-driven. The data-driven prediction adopts the VMD-GRU model, which inputs the historical gas data sequence into the VMD algorithm. VMD adaptively decomposes this complex signal into K relatively stable subsequences with different frequencies, which are called intrinsic mode functions (IMFs).

[0019] Furthermore, a separate GRU model is trained for each IMF component. The system uses historical data to teach each GRU model to learn the changing patterns of its corresponding component. The GRU responsible for the IMF1 trend component learns to identify and extrapolate long-term, slow-changing patterns, while the GRU responsible for the IMF2 periodic component learns to identify and repeat periodic fluctuation patterns. After training, the most recent K IMF sequences are input into their respective GRU models. Each GRU model outputs its component's predicted value for the next N steps. GRU1 outputs the trend for the next 7 days, and GRU2 outputs the periodic fluctuations for the next 7 days. The system sums the future prediction sequences of all K IMF components according to their corresponding time points to output the final gas concentration prediction curve for the next N days. If the prediction curve exceeds the warning value at some future time point, the system will immediately trigger an early warning.

[0020] Furthermore, the physical mechanism model is based on the physical and chemical laws inside the transformer to construct a mathematical model to simulate the operation and aging process of the equipment, including a thermal balance model, an insulation aging model, and an overload model.

[0021] Furthermore, the health status assessment unit integrates all information to provide an intuitive and quantitative overview of equipment health. It uses a weighted comprehensive evaluation method or machine learning regression model to determine key evaluation indicators, outputs a health score of 0-100, and integrates the uncertainty of multi-source evidence based on DS evidence theory or Bayesian inference.

[0022] (III) Beneficial Effects This invention provides a fault prediction system for DC converter transformers based on smart grids, which has the following beneficial effects: 1. A fault diagnosis system based on the fusion of deep forest models and expert rules has brought revolutionary improvements to the operation and maintenance of DC converter transformers. Its core value lies in achieving a leap from rough judgment to precise decision-making. This system greatly improves the accuracy and reliability of fault diagnosis, significantly reducing the risk of misjudgment and missed judgment. Simultaneously, its multi-layered cascaded structure can automatically optimize feature representation, resulting in stronger generalization capabilities. It enables precise fault location and root cause analysis.

[0023] 2. Through the "decomposition-prediction-reconstruction" strategy, true forward-looking early warning is achieved. It not only accurately extracts potential trends and periodic patterns from complex and noisy gas data, but more importantly, it can quantitatively predict the trajectory of fault development. When the system issues a warning that "the threshold is expected to be exceeded in 5 days," maintenance personnel no longer receive a simple status alarm, but a decision window with a clear timeframe. This transforms traditional "post-event remediation" or "in-event emergency response" into "pre-event intervention," significantly improving the initiative and safety of maintenance.

[0024] 3. By integrating quantitative scoring with confidence levels, the system achieves digitalization and visualization of equipment status, enabling managers to quickly grasp the overall condition of equipment without delving into technical details. This effectively solves the uncertainty problem in complex diagnostics, equipping each diagnostic conclusion with a "confidence scale" to help maintenance personnel determine whether a diagnosis is "basically certain" or "requires further observation," significantly reducing the risk of misjudgment. This dual output of "status + confidence" makes maintenance strategy formulation both comprehensive and robust, successfully transforming traditional experience-based qualitative judgments into data-driven quantitative decision-making, and significantly improving the intelligence level and reliability of power grid asset management. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the DC converter transformer fault prediction system based on smart grid of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Please see Figure 1 This invention provides a fault prediction system for DC converter transformers based on smart grids, comprising: The data acquisition module is used to comprehensively collect various status information of the converter transformer, including a multi-parameter sensor unit, a SCADA data interface unit, an equipment ledger and historical data management unit, and an environmental data monitoring unit.

[0028] Multi-parameter sensor units directly monitor the physical and chemical states of converter transformers by deploying high-frequency or ultrasonic partial discharge sensors, dissolved gas analysis sensors in oil, temperature sensors, vibration sensors, and core grounding current sensors on the transformer. For example, DGA sensors can monitor the content of characteristic gases such as hydrogen (H2), methane (CH4), ethane (C2H6), ethylene (C2H4), and acetylene (C2H2).

[0029] The main insulating materials inside a transformer are insulating oil and solid insulation. During normal operation, they will slowly age due to heat and electrical stress, producing a small amount of gas. However, when an internal fault occurs, the fault point will generate local high temperature or high energy discharge, which will rapidly accelerate the chemical decomposition of these insulating materials and produce a large amount of characteristic gas.

[0030] Under high temperature or discharge energy, the carbon-hydrogen bond (CH) and carbon-carbon bond (CC) of insulating oil will break, generating free hydrogen atoms and hydrocarbon radicals. These radicals recombine to form hydrogen gas (H2) and a series of low-molecular-weight hydrocarbon gases, such as methane (CH4), ethane (C2H6), ethylene (C2H4), and acetylene (C2H2).

[0031] Solid insulation is mainly composed of carbon, hydrogen, and oxygen. When overheated, it undergoes pyrolysis, primarily producing carbon monoxide (CO) and carbon dioxide (CO2).

[0032] Different types and severity of faults will produce different kinds and proportions of gas. Each gas acts like a specific "warning light," indicating a certain type of fault. Hydrogen (H2) is produced by the decomposition of water molecules in insulating oil during low-energy discharges, such as partial discharges and low-temperature corona discharges. It is a sensitive indicator of low-energy discharges and arc discharges, but it can also be produced by a variety of other factors, so it needs to be considered in conjunction with other gases for accurate judgment.

[0033] Methane (CH4) and ethane (C2H6) are produced by the decomposition of insulating oil at low to medium temperatures (~150℃~500℃). They usually indicate general overheating faults, such as poor contact of tap changers or eddy current overheating caused by multiple grounding points of the iron core.

[0034] Ethylene (C₂H₄) is the main characteristic gas produced by insulating oil when it is overheated to above 700°C. Its production rate increases sharply with increasing temperature, a clear indication of severe overheating. When ethylene becomes the dominant hydrocarbon gas, it indicates that the temperature of the overheating point is extremely high.

[0035] Acetylene (C2H2) is only produced in large quantities by insulating oil during high-energy arc discharges, where temperatures can reach over 3000℃. It is almost never generated during normal overheating and low-energy discharges, making it the sole and decisive indicator of arc discharge. The presence of a significant amount of acetylene in the oil indicates a potential serious internal discharge fault, such as short circuits between coil turns or layers, or flashover of leads to ground—a very dangerous signal.

[0036] Carbon monoxide (CO) and carbon dioxide (CO2) are the byproducts of the pyrolysis and aging of solid insulating materials, reflecting the deterioration of the solid insulation. When the content of hydrocarbon gases increases significantly at the same time, it indicates that the overheating fault not only involves the oil but has also endangered the solid insulation of the coil, making the fault more severe.

[0037] The SCADA data interface unit obtains real-time operating parameters such as load current, voltage, and power from the power grid SCADA system.

[0038] The equipment ledger and historical data management unit centrally manages equipment models, design parameters, maintenance records, test reports, and family defect information.

[0039] The environmental data monitoring unit collects data such as ambient temperature, humidity, and pollution level. These external factors can affect the operating status and insulation performance of the equipment.

[0040] The feature extraction module is used to transform raw data into useful information, and includes a data preprocessing and cleaning unit, a multi-source data fusion unit, and a feature engineering unit.

[0041] The data preprocessing and cleaning unit is responsible for handling noise, outliers, and missing values ​​in the raw data. For example, for vibration and electrical signals, filtering algorithms such as Sensitive Singular Value Decomposition (SSVD) or Fast Fourier Transform (FFT) are used to extract the effective components.

[0042] The multi-source data fusion unit aligns and integrates heterogeneous data from different sensors and systems over time to form a unified and continuous panoramic view of the device status.

[0043] The feature engineering unit extracts key features from the processed data, such as calculating the gas content and gas production rate of DGA data, or extracting the amplitude of specific frequency components from vibration signals. These features can more directly reflect the health status of the equipment.

[0044] The fault diagnosis and prediction analysis module is responsible for in-depth analysis and forward-looking prediction of the condition of DC converter transformers, including fault diagnosis model unit, fault prediction model unit and health status assessment unit.

[0045] The fault diagnosis model unit uses intelligent algorithms to identify and classify the current fault.

[0046] A Deep Forest model is used to perform a preliminary diagnosis of transformer fault characteristics based on oil chromatography data, specifically: A large amount of DGA data records were collected from the transformer's historical operation and maintenance records. Each record should include the content of seven key gases: H2, CH4, C2H2, C2H4, C2H6, CO, and CO2. For each DGA record, a clear "fault type" label was assigned based on subsequent core inspections, electrical tests, or expert consultation conclusions. Common labels include: normal, low-temperature overheating, medium-temperature overheating, high-temperature overheating, partial discharge, low-energy discharge, and high-energy discharge.

[0047] Check the dataset for obvious errors or missing records. For missing values, interpolation between adjacent records or using the average can be used to fill in the gaps. Outliers that are significantly outside the reasonable range should be removed or corrected. Since the concentration ranges of different gases vary greatly—for example, H2 might be tens of ppm while CO2 might be thousands of ppm—it is necessary to standardize the values ​​of all gases to a uniform scale. The most common method is Z-score standardization, which involves subtracting the average value of all samples for that gas from its value for each gas, and then dividing by the standard deviation.

[0048] The standardized values ​​of the seven gases are retained as basic features. The gas ratio features are calculated based on the three-ratio method. The sum of CH4 + C2H2 + C2H4 + C2H6 is calculated as the total hydrocarbon content, and the proportion of key gases is calculated.

[0049] The gas ratio characteristics calculated based on the three-ratio method are as follows: calculating the C2H2 / C2H4 ratio is the key to distinguishing between discharge faults and overheating faults; calculating the CH4 / H2 ratio helps to distinguish between low-temperature overheating and discharge faults; and calculating the C2H4 / C2H6 ratio reflects the temperature range of overheating. The larger the ratio, the higher the temperature is usually.

[0050] The proportion of key gases, such as the C2H2 / total hydrocarbon ratio, can highlight the importance of characteristic gases in high-energy discharges.

[0051] Each DGA record is transformed into a feature vector, for example: [normalized H2, normalized CH4, ..., log(C2H2 / C2H4), log(CH4 / H2), log(C2H4 / C2H6), total hydrocarbons].

[0052] The prepared labeled dataset is randomly divided into three parts: a training set (approximately 70%), used to "teach" the model the relationship between features and fault types; a validation set (approximately 15%), used to tune model hyperparameters during training and monitor for overfitting; and a test set (approximately 15%), used to finally and independently evaluate the model's generalization performance, simulating real-world scenarios.

[0053] The original feature vector is input into the first layer. This layer typically consists of multiple base learners, such as two random forests and two extreme random trees. Each base learner independently classifies the input sample and outputs a class probability vector. For example, for a given sample, a random forest might output [0.02, 0.1, 0.05, 0.8, 0.01, 0.01, 0.01], indicating that it considers the sample to have an 80% probability of being "overheated".

[0054] The probability vectors output by all base learners in the first layer are concatenated, and then concatenated with the original input feature vector to form a completely new, greatly expanded enhanced feature vector. This new feature vector not only contains the original information but also incorporates the first-layer model's "understanding and judgment" of the data, resulting in a richer information content.

[0055] The enhanced feature vectors are then fed into the second layer of the forest. The second layer learns at a higher level of feature abstraction. This process is repeated layer by layer. The model continuously monitors its performance on the validation set. Training automatically stops when adding new layers no longer significantly improves validation accuracy. This is called the "early stopping" mechanism, which effectively prevents overfitting and ensures the model's generalization ability.

[0056] The final model is evaluated using a test set that has never been used for training. Once the model passes the evaluation, it can be integrated into the fault prediction system.

[0057] The metrics used to evaluate the final model include overall accuracy, confusion matrix, precision, recall and F1 score, and recall rate.

[0058] Overall accuracy represents the proportion of samples correctly classified by the model. The confusion matrix is ​​a crucial table that clearly shows which fault types the model is prone to confusion with. For example, does the model consistently misclassify "medium-temperature overheating" as "high-temperature overheating"? Precision, recall, and F1 score are analyzed in detail for each fault type. Recall measures the model's "complete coverage" of a particular fault. For severe faults like "high-energy discharge," we aim for extremely high recall to ensure no missed cases.

[0059] The classification model based on Deep Forest takes a structured feature vector from the feature engineering unit as input. The core features include the three-ratio encoding of DGA data, gas production rate, and percentage of each gas. Auxiliary features include load rate, operating temperature, and partial discharge level.

[0060] The aforementioned feature vectors are input into a pre-trained deep forest model. The model first generates feature subsets with multiple random dimensions to capture patterns under different feature combinations. The data is then passed through multiple forest layers. Each layer receives the feature information processed by the previous layer and further optimizes the feature representation. The final forest layer outputs a probability distribution representing the likelihood of belonging to each fault type (e.g., normal: 5%, low-energy discharge: 10%, high-temperature overheating: 80%, arc discharge: 5%). The system selects the category with the highest probability as the initial diagnostic result, such as "high-temperature overheating fault".

[0061] The rules are written by domain experts based on their theoretical knowledge and field experience. These rules are logical statements in the form of "IF-THEN". The rules specify what kind of combination of evidence will lead to what kind of conclusion, and are accompanied by a confidence level.

[0062] Example: Rule 1: IF "C2H4 is the main hydrocarbon gas" AND "The energy of the 100Hz and its harmonic components in the vibration spectrum is significantly increased" THEN "Fault location: Multiple grounding points in the iron core" (Confidence level: 85%) Rule 2: IF "C2H4 is the main hydrocarbon gas" AND "CO / CO2 ratio is normal" AND "Infrared thermal imaging shows abnormal temperature at the tap changer connection point, higher than other phases" THEN "Fault location: poor contact in the tap changer" (Confidence level: 92%) Rule 3: IF "H2 and CH4 levels significantly increased" AND "Trace amounts of C2H2 present in oil" AND "Pulse signals detected in high-frequency grounding current" THEN "Fault involving partial discharge in solid insulation" (Confidence level: 88%) The family-based defect database records known design, material, or workmanship defects in specific models and batches of transformers, sourced from technical bulletins issued by equipment manufacturers, industry-shared failure case databases, and the organization's historical experience.

[0063] Example: Model: ZZZ-1000; Production batch: Second half of 2020; Family defect: Poor welding process of the B phase winding lead wire, which is prone to overheating.

[0064] The equipment ledger and historical maintenance database include equipment structure diagrams, the last maintenance time, replaced parts, and past fault records.

[0065] When a diagnostic request arrives, the inference engine operates according to the following steps: Step 1: The system receives a preliminary diagnostic result, such as "high temperature overheating". The inference engine immediately searches the association rule base and "fetches" all rules whose conclusions are related to "overheating". Rule 1 and Rule 2 are then triggered and placed into the "rule pool to be evaluated".

[0066] Step 2: The system begins to check each rule in the "Rule Pool to be Evaluated" to see if the conditions in its "IF" part are met by the current data.

[0067] Regarding rule 1: Condition 1: "C2H4 is the main hydrocarbon gas" → The system checks the DGA data to confirm whether the C2H4 content is significantly higher than that of other hydrocarbon gases. (Satisfied) Condition 2: "100Hz component is significant in the vibration spectrum" → The system calls the output of the vibration analysis unit to check whether the amplitude of the 100Hz frequency exceeds the normal threshold. (Satisfied) Since both conditions are met, rule 1 is activated, and the system draws a preliminary conclusion: "It may be that the iron core is grounded at multiple points," with a confidence level of 85%.

[0068] Regarding rule 2: Condition 1: "C2H4 is the main hydrocarbon gas" → (Satisfied) Condition 2: "CO / CO2 ratio is normal" → The system calculates the ratio and confirms that it is within the range indicating that the solid insulation has not severely decomposed. (Satisfied) Condition 3: "Infrared indicator shows high tap changer temperature" → The system retrieves the latest infrared inspection report and confirms that the tap changer has an overheating point. (Saved) Rule 2 was also activated, leading to the conclusion: "It may be a problem with the tap changer contact," with a confidence level of 92%.

[0069] Step 3: Now the system has two possible results, both with a certain degree of confidence. It needs to decide which is more credible, or how to combine this information. The simplest method is to choose the result with the highest confidence level. Here, 92% of Rule 2 is greater than 85% of Rule 1, so the system will initially lean towards "tap switch poor contact".

[0070] A more sophisticated approach involves uncertainty fusion. For example, using DS evidence theory or Bayesian networks. The system treats each activated rule and its confidence level as an independent "evidence body," mathematically calculating the final probability supported by these pieces of evidence. This considers not only the level of confidence but also whether the evidence supports or contradicts each other.

[0071] During the integration process, the system queries the family defect database in parallel. If it finds that the current transformer matches the record for "poor welding process of B-phase winding lead wire", it generates strong corroborating evidence and significantly increases the confidence level of conclusions related to "winding lead wire".

[0072] After the inference engine completes comprehensive reasoning, it generates a final, precise diagnostic report. The report includes the final diagnostic conclusion, the most likely location of the fault, and the overall confidence level.

[0073] Key evidence supporting this conclusion includes: “DGA shows C2H4 dominance, infrared detection of tap changer overheating, and consistency with familial defect records.”

[0074] This report is no longer a simple fault classification, but a precise decision support information that includes the nature of the fault, its specific location, its reliability, and the reasoning behind it, enabling maintenance personnel to take direct action. This is the true value of a fault prediction system.

[0075] A fault diagnosis system based on the fusion of deep forest models and expert rules has brought revolutionary improvements to the operation and maintenance of DC converter transformers. Its core value lies in achieving a leap from rough judgment to precise decision-making. This system greatly improves the accuracy and reliability of fault diagnosis, significantly reducing the risk of misdiagnosis and missed diagnosis. Simultaneously, its multi-layered cascaded structure can automatically optimize feature representation, resulting in stronger generalization capabilities. It enables precise fault location and root cause analysis.

[0076] The fault prediction model unit adopts two complementary paradigms: data-driven and physical mechanism-driven.

[0077] Data-driven prediction employs the VMD-GRU model, disregarding physical principles and relying solely on hidden patterns within the data. Its core idea is "decomposition-prediction-reconstruction," and its complete workflow can be clearly illustrated as follows: The system inputs historical gas data sequences into the VMD algorithm, which adaptively decomposes this complex signal into K relatively stable subsequences with different frequencies, called intrinsic mode functions (IMFs). Through decomposition, a complex problem is transformed into a set of multiple simple problems.

[0078] For example, the C2H2 sequence may be decomposed into: IMF1: A smooth, slowly rising curve, representing a long-term growth trend due to the continuous deterioration of insulation materials.

[0079] IMF2: A periodically fluctuating curve that may be related to seasonal or cyclical changes in load.

[0080] IMF3 and later: curves with high-frequency fluctuations, mainly representing random noise and measurement errors.

[0081] By decomposing, we have transformed a complex problem into a set of multiple simple problems.

[0082] A separate GRU model is trained for each IMF component. The system uses historical data (such as data from the previous 30 days) to teach each GRU model to learn the patterns of change in its corresponding component. The GRU responsible for IMF1 (the trend component) learns to identify and extrapolate long-term, slow-changing patterns. The GRU responsible for IMF2 (the cycle component) learns to identify and repeat patterns of cyclical fluctuations.

[0083] After training, the K most recent IMF sequences are input into their respective GRU models. Each GRU model outputs the predicted value of its components for the next N steps (e.g., 7 days). GRU1 outputs the trend for the next 7 days, GRU2 outputs the cyclical fluctuations for the next 7 days, and so on.

[0084] The system sums the future prediction sequences of all K IMF components from the second step according to their corresponding time points, outputting the final gas concentration prediction curve for the next N days. This curve integrates predictions of trends, periods, and noise. The system compares this prediction curve with preset safety thresholds, attention values, and warning values.

[0085] If the predicted curve exceeds the warning value at some point in the future (e.g., day 5), the system will not wait until day 5 to issue an alarm, but will immediately trigger an early warning, informing that "based on the current trend, it is expected to enter a high-risk state in 5 days", giving maintenance personnel sufficient time to intervene.

[0086] The physical mechanism model is based on the physical and chemical laws governing the transformer to construct a mathematical model that simulates the operation and aging process of the equipment. It includes a thermal balance model, an insulation aging model, and an overload model.

[0087] Thermal balance models predict winding hot spot temperatures, which is the primary driver of insulation aging.

[0088] Input real-time data, equipment parameters, and core equations. Real-time data includes load current (from SCADA), ambient temperature, and cooler operating status (fan / oil pump on / off). Equipment parameters include design heat dissipation coefficient, thermal resistance, heat capacity, etc. (from equipment ledger). Core equations include differential equations or equivalent thermal circuit models based on the first law of thermodynamics (energy conservation). Heat generated = Heat dissipated to the environment + Heat required for equipment to heat up. The model solves the above equations in real time to calculate the temperature of the hottest spot in the winding under the current and predicted future loads. It predicts whether the hot spot temperature will exceed the safety limit (e.g., 120°C) during the upcoming peak electricity consumption period. If the actual oil temperature is significantly higher than the model's predicted value, it indicates that the cooling system may be inefficient or blocked.

[0089] The insulation aging model predicts the remaining life of the transformer. The input is the winding hot spot temperature calculated by the thermal balance model. The model converts the hot spot temperature at each moment into a relative aging rate based on the Arrhenius equation. For example, the aging rate might be 1 (normal) at 98°C, but could surge to 8 (8 times the normal aging rate) at 110°C. The system integrates the aging rate over time to obtain the cumulative life loss. Remaining life percentage = 1 - (lost life / total expected life), outputting the real-time aging rate and the percentage of life lost (e.g., "40% lost").

[0090] The Arrhenius rate equation describes how the rate of a chemical reaction (in this case, the rate at which the cellulose chains of the insulating paper break) increases exponentially with temperature.

[0091] The overload model dynamically assesses the maximum load a transformer can withstand while ensuring safety; essentially, it's the inverse application of the thermal balance model. Inputting real-time load, ambient temperature, and current health status (e.g., insulation aging level, cooling system status), it solves the problem: "Given the current ambient temperature and health status, what is the maximum permissible load current to ensure the hot spot temperature does not exceed the absolute safety limit?" This comprehensively considers both short-term thermal tolerance and long-term lifespan degradation. The output includes a safe operating range and decision-making recommendations.

[0092] For example, "Under the current ambient temperature of 35°C, overload to 105% of rated capacity is permissible within 2 hours." Decision recommendations provide the power grid dispatch center with specific, quantifiable load adjustment suggestions, thereby achieving seamless coordination between power grid dispatch and equipment status, maximizing transmission potential while ensuring equipment safety.

[0093] By employing a "decomposition-prediction-reconstruction" strategy, true proactive early warning is achieved. It not only accurately extracts potential trends and periodic patterns from complex and noisy gas data, but more importantly, it can quantify and predict the trajectory of fault development. When the system issues a warning that "the threshold is expected to be exceeded in 5 days," maintenance personnel no longer receive a simple status alarm, but rather a decision window with a clear timeframe. This transforms traditional "post-event remediation" or "in-event emergency response" into "pre-event intervention," significantly improving the proactiveness and safety of maintenance.

[0094] The health status assessment unit integrates all information to provide an intuitive and quantitative overview of equipment health. It typically employs a weighted comprehensive evaluation method or a machine learning regression model to determine key evaluation indicators, such as DGA data score, partial discharge level, vibration level, load capacity, and insulation aging rate. The actual value of each indicator is mapped to a score between 0 and 100 (100 being optimal), and the weights are dynamically adjusted based on diagnostic and predictive results. For example, when a rapid increase in acetylene concentration is predicted, the weight of the "DGA score" will automatically increase. Using the formula: Health Score = Σ(weight_i * indicator score_i), a health score of 0-100 is output, similar to the health status of a mobile phone battery, making it intuitive and easy to read.

[0095] Based on DS evidence theory or Bayesian inference, this method fuses the uncertainties of multi-source evidence. Evidence is obtained from different models. For example, a deep forest model outputs an 80% probability of "overheating" (with a 20% uncertainty). Association rules match a rule with 90% confidence. Vibration monitoring data shows a slight anomaly, but the uncertainty is associated with a 60% fault. Mathematical methods (such as DS synthesis rules) are used to fuse this evidence and its uncertainties. The output calculates a final composite confidence level (e.g., 85%), representing the reliability of the current diagnosis or prediction.

[0096] By integrating quantitative scoring with confidence levels, the system achieves the digitization and visualization of equipment status, enabling managers to quickly grasp the overall condition of equipment without delving into technical details. This effectively solves the uncertainty problem in complex diagnostics, equipping each diagnostic conclusion with a "confidence scale" to help maintenance personnel determine whether a diagnosis is "basically certain" or "requires further observation," significantly reducing the risk of misjudgment. This dual output of "status + confidence" makes maintenance strategy formulation both comprehensive and robust, successfully transforming traditional experience-based qualitative judgments into data-driven quantitative decision-making, and significantly improving the intelligence level and reliability of power grid asset management.

[0097] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0098] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0099] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A fault prediction system for DC converter transformers based on smart grids, characterized in that: include: The data acquisition module is used to comprehensively collect various status information of the converter transformer, including a multi-parameter sensor unit, a SCADA data interface unit, an equipment ledger and historical data management unit, and an environmental data monitoring unit. The feature extraction module is used to transform raw data into useful information, including a data preprocessing and cleaning unit, a multi-source data fusion unit, and a feature engineering unit. The fault diagnosis and prediction analysis module is responsible for in-depth analysis and forward-looking prediction of the condition of DC converter transformers, including fault diagnosis model unit, fault prediction model unit and health status assessment unit.

2. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The multi-parameter sensor unit directly monitors the physical and chemical state of the converter transformer by deploying high-frequency or ultrasonic partial discharge sensors, dissolved gas analysis sensors in oil, temperature sensors, vibration sensors, and core grounding current sensors on the converter transformer; the SCADA data interface unit obtains real-time operating parameters of load current, voltage, and power from the power grid SCADA system; the equipment ledger and historical data management unit centrally manages the equipment model, design parameters, maintenance records, test reports, and family defect information; and the environmental data monitoring unit collects data on ambient temperature, humidity, and pollution level.

3. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The multi-source data fusion unit aligns and integrates heterogeneous data from different sensors and systems over time to form a unified and continuous panoramic view of the device status. The feature engineering unit extracts key features from the processed data.

4. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The fault diagnosis model unit uses intelligent algorithms to identify and classify current faults; it employs the Deep Forest model to perform a preliminary diagnosis of the fault characteristics of the transformer based on oil chromatography data. The feature vector is input into the pre-trained deep forest model. The model first generates feature subsets of multiple random dimensions to capture patterns under different feature combinations. The data passes through multiple forest layers in sequence. Each layer receives the feature information processed by the previous layer and further optimizes the feature representation. The last forest layer outputs a probability distribution, representing the probability of belonging to each fault type. The system selects the category with the highest probability as the primary diagnostic result.

5. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: Once the system receives the initial diagnostic results, the inference engine immediately searches the association rule base, retrieving rules related to all conclusions. These rules are then triggered and added to the rule pool for evaluation. The system then checks each rule in the pool to see if its IF condition is met by the current data, selecting the result with the highest confidence. After completing comprehensive inference, the inference engine generates the final precise location diagnostic report. The report includes the final diagnostic conclusion, the most likely fault location, and the overall confidence level.

6. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The fault prediction model unit adopts two complementary paradigms: data-driven and physical mechanism-driven. The data-driven prediction adopts the VMD-GRU model, which inputs the historical gas data sequence into the VMD algorithm. VMD adaptively decomposes this complex signal into K relatively stable subsequences with different frequencies, which are called intrinsic mode functions (IMFs).

7. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: A separate GRU model is trained for each IMF component. The system uses historical data to teach each GRU model to learn the changing patterns of its corresponding component. The GRU responsible for the IMF1 trend component learns to identify and extrapolate long-term, slow-changing patterns, while the GRU responsible for the IMF2 periodic component learns to identify and repeat patterns of periodic fluctuations. After training, the K most recent IMF sequences are input into their respective GRU models. Each GRU model outputs the predicted value of its components for the next N steps. GRU1 outputs the trend for the next 7 days, and GRU2 outputs the periodic fluctuations for the next 7 days. The system sums the future prediction sequences of all K IMF components according to their corresponding time points to output the final gas concentration prediction curve for the next N days. If the prediction curve exceeds the warning value at some future time point, the system will immediately trigger an early warning.

8. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The physical mechanism model is based on the physical and chemical laws inside the transformer to construct a mathematical model to simulate the operation and aging process of the equipment, including the thermal balance model, insulation aging model and overload model.

9. The fault prediction system for DC converter transformers based on smart grids according to claim 1, characterized in that: The health status assessment unit integrates all information to provide an intuitive and quantitative overview of equipment health. It uses a weighted comprehensive evaluation method or machine learning regression model to determine key evaluation indicators, outputs a health score of 0-100, and integrates the uncertainty of multi-source evidence based on DS evidence theory or Bayesian inference.