Dissolved gas regression prediction and abnormal state detection classification system and method based on multi-expert fusion learning and medium
Patent Information
- Application Number
- CN202511530679.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-10-24
AI Technical Summary
[0005]本发明的目的是为了解决现有变压器油中溶解气体预测方法中,尤其在面对复杂工况或突发性异常事件时,现有模型的鲁棒性不足,气体浓度预测性能大幅下降等问题
1、构建了基于注意力机制的多专家学习框架,通过门控网络实现不同运行状态下预测策略的智能切换;
Smart Images

Figure CN121479727B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power equipment condition monitoring and fault diagnosis technology, specifically to a system, method and medium for regression prediction of dissolved gases in oil and abnormal condition detection classification based on multi-expert fusion learning. Background Technology
[0002] During long-term operation, oil-immersed power equipment (such as transformers and reactors) experiences gradual decomposition of insulating oil due to electrical, thermal, and mechanical stresses, producing characteristic gases such as H2, CH4, C2H6, C2H4, C2H2, CO, and CO2. The concentration and growth trend of these gases can reflect the degree of insulation degradation and potential fault types (such as partial discharge, overheating, and arc discharge) within the equipment. Therefore, dissolved gas analysis in oil has become one of the important methods for condition monitoring of power equipment.
[0003] Against the strategic backdrop of "digital transformation" and the construction of "new power systems," intelligent operation and maintenance technology for power equipment is ushering in significant development opportunities. As a core piece of equipment in the power grid, the operating status of oil-immersed transformers directly affects the safety and stability of the entire power system.
[0004] Currently, the prediction of dissolved gases in transformer oil faces three major technical challenges: First, traditional ratio methods and single machine learning models have poor versatility and cannot adapt to the differences in insulation characteristics of equipment of different voltage levels and manufacturers, resulting in significant errors in gas concentration prediction across different equipment scenarios; second, existing methods are insufficient in modeling the complex coupling relationships between multi-component gases (H2, CH4, C2H6, etc.), making it difficult to capture the intrinsic correlation of gas concentration changes; third, when faced with complex operating conditions such as equipment aging, load fluctuations, and changes in ambient temperature and humidity, existing models lack robustness, resulting in a significant decrease in gas concentration prediction performance and an inability to stably output reliable trend prediction results. Summary of the Invention
[0005] The purpose of this invention is to address the problems of insufficient robustness and significant degradation in gas concentration prediction performance in existing methods for predicting dissolved gases in transformer oil, especially under complex operating conditions or sudden abnormal events. This invention provides a system, method, and medium for regression prediction and abnormal state detection classification of dissolved gases in oil based on multi-expert fusion learning. This technology integrates sub-models adapted to different equipment characteristics through multi-expert fusion learning, overcoming the limitations of single models and improving versatility to reduce cross-equipment prediction errors. It uses a context-aware expert model to obtain the complex coupling relationships between multiple gas components; and it uses an anomaly-aware expert model to capture real-time interference signals such as equipment aging, load fluctuations, and changes in ambient temperature and humidity, dynamically optimizing the model output strategy to significantly increase the model's robustness under complex operating conditions and ensure stable prediction performance.
[0006] To achieve the above objectives, the present invention provides, in one aspect, a system for regression prediction and abnormal state detection classification of dissolved gases in oil based on multi-expert fusion learning, comprising: The data preprocessing module is used to perform noise filtering on the historical concentration time series datasets of each characteristic gas using the interquartile range method, to obtain the filtered historical concentration time series datasets of each characteristic gas, to form the overall dataset of all the filtered historical concentration time series datasets of all characteristic gases, and to divide the overall dataset into training set and validation set according to time. The filtering module is used to correct the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on the corrected predicted value sequence, the prediction values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, these prediction values are weighted and fused using a gating network model to obtain the final prediction result of the future concentration of each characteristic gas in the training set. Based on the final prediction result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, and the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The prediction module is used to input the overall dataset into the anomaly perception expert model, time series modeling expert model, context-aware expert model and gated network model with optimal neural network model parameters to obtain the predicted concentration data of each characteristic gas.
[0007] Preferably, the historical concentration data of each characteristic gas in the insulating oil are organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
[0008] Preferably, the method of using the interquartile range (INR) to filter the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets of each characteristic gas includes: removing noise from the historical concentration time series datasets of each characteristic gas within the specified intervals. Data points outside the range are used to obtain the range for each characteristic gas. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration dataset within the range is processed by removing data points at these time points to obtain a filtered historical concentration time series dataset for each feature gas. Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas; Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas; k represents the interval parameter of the interquartile range method; and IQR represents the interquartile range.
[0009] Preferably, the method for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set includes: The concentration features of the gas sequences in the training set are constructed using an anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; This represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional. An encoder-decoder structure is used to reconstruct the gas sequences from the obtained training set based on their concentration features, thus obtaining the anomaly perception expert model's prediction of the reconstructed gas sequences from the training set. in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer. The residual values between the concentrations of characteristic gases in the training set and the predicted values reconstructed from the gas sequences in the training set were calculated using an anomaly perception expert model. The calculation method is as follows: Continue to use the anomaly perception expert model to calculate the residual values. Intelligent scoring is performed to identify abnormal concentration points in the training set. The intelligent scoring mechanism is expressed as follows: in, Represents the residual value at time j. Abnormal scores; Represents the sigmoid function; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j; The predicted values of the gas sequence reconstruction at abnormal concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstruction at those abnormal concentration points, resulting in the corrected predicted sequence for the training set. The corrected predicted gas sequence reconstruction value at time j in the training set is then used. The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
[0010] Preferably, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model based on the corrected predicted value sequence of the training set. Then, these predicted values are weighted and fused using a gated network model to obtain the final predicted result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final predicted result of the future concentration of each characteristic gas in the training set includes: The corrected predicted sequence from the training set is input into the encoder-decoder structure to obtain the anomaly perception expert model's prediction of the future concentration of each characteristic gas in the training set. , is represented as: Where f(...) represents a temporal neural network; This represents the corrected predicted sequence from the training set; This represents the predicted output at time T. Input the corrected prediction sequence from the training set. In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: The corrected predicted sequences from the training set are input into the Bi-LSTM network to obtain the context-aware expert model's predictions of the future concentrations of each characteristic gas in the training set. , is represented as: in, This represents the final hidden state of the Bi-LSTM network; express dimensional space; This represents the weight matrix of the output layer; Indicates the bias term of the output layer; A gated network model is used to weight and fuse the predictions of the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model for the future concentrations of each characteristic gas in the training set, resulting in the final prediction of the future concentrations of each characteristic gas in the training set. , is represented as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides; Calculate the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each characteristic gas in the training set The loss value L between the two is calculated using the following formula: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
[0011] Preferably, the method of optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model using the Adam optimizer based on the loss value L, and iteratively updating the weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; This represents the decay coefficient of first-order momentum; It is the decay coefficient of second momentum; Represents the gradient at time t-1 The first-order momentum; Represents the gradient at time. The second momentum; This represents the neural network model parameters of a certain model at time t; This represents the neural network model parameters of a certain model at time t-1; Indicates the learning rate; It represents a very small positive number.
[0012] A second aspect of this invention provides a method for regression prediction and abnormal state detection and classification of dissolved gases in oil based on multi-expert fusion learning, comprising: The quartile range method was used to filter the noise in the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets of each characteristic gas. All the filtered historical concentration time series datasets of the characteristic gases were combined into a whole dataset, and the whole dataset was divided into training set and validation set according to time. The predicted values of gas sequences reconstructed from abnormal concentration points in the training set are corrected using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on this corrected predicted value sequence, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, a gating network model is used to weight and fuse these predicted values to obtain the final predicted result of the future concentration of each characteristic gas in the training set. Based on the final predicted result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training session, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, at which point the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The entire dataset is input into the anomaly perception expert model, time series modeling expert model, context-aware expert model, and gated network model with optimal neural network model parameters to obtain the predicted concentration data for each characteristic gas.
[0013] Preferably, the historical concentration data of each characteristic gas in the insulating oil are organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
[0014] Preferably, the method of using the interquartile range (INR) to filter the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets of each characteristic gas includes: removing noise from the historical concentration time series datasets of each characteristic gas within the specified intervals. Data points outside the range are used to obtain the range for each characteristic gas. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration dataset within the range is processed by removing data points at these time points to obtain a filtered historical concentration time series dataset for each feature gas. Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas; Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas; k represents the interval parameter of the interquartile range method; and IQR represents the interquartile range.
[0015] Preferably, the method for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set includes: The concentration features of the gas sequences in the training set are constructed using an anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; This represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional. An encoder-decoder structure is used to reconstruct the gas sequences from the obtained training set based on their concentration features, thus obtaining the anomaly perception expert model's prediction of the reconstructed gas sequences from the training set. in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer. The residual values between the concentrations of characteristic gases in the training set and the predicted values reconstructed from the gas sequences in the training set were calculated using an anomaly perception expert model. The calculation method is as follows: Continue to use the anomaly perception expert model to calculate the residual values. Intelligent scoring is performed to identify abnormal concentration points in the training set. The intelligent scoring mechanism is expressed as follows: in, Represents the residual value at time j. Abnormal scores; Represents the sigmoid function; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j; The predicted values of the gas sequence reconstruction at abnormal concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstruction at those abnormal concentration points, resulting in the corrected predicted sequence for the training set. The corrected predicted gas sequence reconstruction value at time j in the training set is then used. The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
[0016] Preferably, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model based on the corrected predicted value sequence of the training set. Then, these predicted values are weighted and fused using a gated network model to obtain the final predicted result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final predicted result of the future concentration of each characteristic gas in the training set includes: The corrected predicted sequence from the training set is input into the encoder-decoder structure to obtain the anomaly perception expert model's prediction of the future concentration of each characteristic gas in the training set. , is represented as: Where f(...) represents a temporal neural network; This represents the corrected predicted sequence from the training set; This represents the predicted output at time T. Input the corrected prediction sequence from the training set. In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: The corrected predicted sequences from the training set are input into the Bi-LSTM network to obtain the context-aware expert model's predictions of the future concentrations of each characteristic gas in the training set. , is represented as: in, This represents the final hidden state of the Bi-LSTM network; express dimensional space; This represents the weight matrix of the output layer; Indicates the bias term of the output layer; A gated network model is used to weight and fuse the predictions of the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model for the future concentrations of each characteristic gas in the training set, resulting in the final prediction of the future concentrations of each characteristic gas in the training set. , is represented as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides; Calculate the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each characteristic gas in the training set The loss value L between the two is calculated using the following formula: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
[0017] Preferably, the method of optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model using the Adam optimizer based on the loss value L, and iteratively updating the weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; This represents the decay coefficient of first-order momentum; It is the decay coefficient of second momentum; Represents the gradient at time t-1 The first-order momentum; Represents the gradient at time. The second momentum; This represents the neural network model parameters of a certain model at time t; This represents the neural network model parameters of a certain model at time t-1; Indicates the learning rate; It represents a very small positive number.
[0018] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above.
[0019] Compared with the prior art, the present invention has the following advantages: 1. A multi-expert learning framework based on attention mechanism was constructed, and intelligent switching of prediction strategies under different operating states was realized through gating network; 2. Innovatively combining spatiotemporal feature extraction with multimodal data fusion, and employing an improved Transformer architecture to capture the long-term dependence of gas concentration; 3. An anomaly detection algorithm based on dynamic thresholds was designed to achieve early and accurate warning of faults; 4. A lightweight model deployment solution was developed to ensure the real-time operation of the algorithm on edge devices.
[0020] 5. Integrating multi-expert modeling capabilities to enhance system robustness: This invention constructs three expert models with significant structural differences and complementary functions, which can simultaneously model abnormal modes, time series trends, and gas relationships, significantly improving the system's adaptability and robustness in complex scenarios.
[0021] 6. An anomaly perception mechanism is introduced to achieve dual functions of anomaly detection and correction: Anomaly perception experts can identify sudden anomalies in real time and make dynamic corrections, effectively mitigating the interference of outliers on regression prediction results and improving the stability of the model in non-stationary environments.
[0022] 7. Supports multi-task fusion of regression prediction and classification judgment: The system can not only output future gas concentration trends, but also simultaneously judge the current gas state level or warning category, realizing integrated modeling of "trend + diagnosis" to meet the intelligent monitoring needs of power equipment.
[0023] 8. Possesses excellent engineering adaptability and real-time response capability: The model structure is lightweight and can be deployed on power equipment condition monitoring platforms to meet online inference requirements, providing data support and decision-making basis for multi-level early warning of gas in oil. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the structure of the oil dissolved gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning of the present invention; Figure 2 This is a schematic diagram of the hardware device structure of the oil dissolved gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning according to the present invention; Figure 3 This is a flowchart of the method for regression prediction and abnormal state detection classification of dissolved gases in oil based on multi-expert fusion learning, which is based on the present invention. Detailed Implementation
[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0026] The endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.
[0027] Furthermore, the technical solutions provided in the various embodiments of the present invention can be combined with each other, but only if they are feasible to those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0028] Example 1 like Figure 1 The system shown is a multi-expert fusion learning-based regression prediction and anomaly detection classification system for dissolved gases in oil, including: The data preprocessing module is used to perform noise filtering on the historical concentration time series datasets of each characteristic gas using the interquartile range method, to obtain the filtered historical concentration time series datasets of each characteristic gas, to form the overall dataset of all the filtered historical concentration time series datasets of all characteristic gases, and to divide the overall dataset into training set and validation set according to time. The filtering module is used to correct the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on the corrected predicted value sequence, the prediction values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, these prediction values are weighted and fused using a gating network model to obtain the final prediction result of the future concentration of each characteristic gas in the training set. Based on the final prediction result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, and the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The prediction module is used to input the overall dataset into the anomaly perception expert model, time series modeling expert model, context-aware expert model and gated network model with optimal neural network model parameters to obtain the predicted concentration data of each characteristic gas.
[0029] In this invention, the characteristic gas refers to the characteristic gas in insulating oil. The characteristic gases in insulating oil typically include seven types: H2, CH4, C2H6, C2H4, C2H2, CO, and CO2. In this embodiment, the concentration data of these seven characteristic gases are detected and predicted.
[0030] In this invention, the historical concentration data of each characteristic gas in insulating oil are organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
[0031] Specifically, this invention first requires extracting historical concentration data (unit: ppm) of seven characteristic gases (H2, CH4, C2H6, C2H4, C2H2, CO, CO2) from insulating oil. The gas concentration detection strictly follows the national standard "Gas Chromatography Determination of Dissolved Gas Components in Insulating Oil," and its detection process includes three main steps: first, standardized collection of transformer oil samples; second, degassing of dissolved gases in the oil; and finally, separation and quantitative detection of gas components using a gas chromatograph, i.e., detecting the concentration of each characteristic gas in the oil sample. The original data comes from long-term monitoring records of ultra-high voltage transformers across the entire network. The historical oil and gas concentration readings of each independent substation are organized chronologically and constructed into a time-series dataset of historical concentrations for each characteristic gas.
[0032] Preferably, the present invention employs the interquartile range method to perform noise filtering on the historical concentration time-series datasets of each characteristic gas, and the method for obtaining the filtered historical concentration time-series datasets of each characteristic gas includes: calculating the quartiles of the historical concentration time-series datasets of each characteristic gas, and removing the historical concentration time-series datasets of each characteristic gas that fall within the specified intervals. The data points outside of these (each characteristic gas's historical concentration time series dataset corresponds to a different interval) Any interval in the historical concentration time series dataset of the characteristic gas that falls within the range of the characteristic gas. Data points outside the range will be considered outliers and removed. This will allow you to obtain the values for each characteristic gas within its respective range. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration datasets within these time points are all removed to improve data quality, resulting in a filtered historical concentration time series dataset for each characteristic gas (the time series in the filtered historical concentration time series datasets for each characteristic gas are the same). Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas (i.e., the value located at the 25th percentile of the sorted data); Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas (i.e., the value located at the 75th percentile of the sorted data); k represents the interval parameter of the interquartile range method, which is used to adjust the normal concentration range of the characteristic gas, and is generally taken as 1.5; IQR represents the interquartile range, which is a statistic describing the degree of data dispersion, and its core function is to reflect the distribution range of the middle 50% of the data in the dataset.
[0033] In a specific implementation, the overall dataset (which still consists of historical concentrations of characteristic gases arranged chronologically) is divided into training and validation sets according to time. Simultaneously, a regression target is automatically generated for each sample to support the system's learning. Specifically, the training and validation sets can be divided in a time ratio of 8:2 or 9:1. For example, all historical concentration data corresponding to the first 80% or 90% of the time period in the overall dataset can be assigned to the training set, and all historical concentration data corresponding to the last 20% or 10% of the time period can be assigned to the validation set. This division method strictly avoids future information leakage, ensuring the reliability and generalization ability of the model evaluation results.
[0034] Furthermore, this invention constructs an anomaly perception expert model, a time-series modeling expert model, and a context-aware expert model. The anomaly perception expert scores the input sequence for anomalies, locates potential mutation points, and performs soft adjustments on detected abnormal regions. This expert simultaneously outputs corrected concentration predictions and anomaly classification results. The time-series modeling expert models gas volatility, captures stable concentration fluctuations, focuses on the trend of gas concentration changes over time, and outputs regression predictions. The context-aware expert focuses on the interrelationships between different gases, models the dependencies between different gases, improves the ability to model multivariate time-series features, and outputs regression predictions based on the influence between gases. The design goal of the anomaly perception expert is to accurately identify outliers in the characteristic gas concentration sequence and correct the sequence accordingly, thereby improving the accuracy and stability of subsequent regression predictions. The main process of the anomaly perception expert model can be divided into three parts: sequence encoding and reconstruction, residual calculation and anomaly scoring, and sequence correction.
[0035] Specifically, methods for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set include: The concentration features of the gas sequences in the training set are constructed using the sequence encoding and reconstruction part of the anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; The value represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional.
[0036] To further effectively identify the location of abnormal gas values in the sequence, the anomaly perception expert model uses an encoder-decoder structure to reconstruct the gas sequence from the concentration features of the obtained training set, thus obtaining the predicted value of the anomaly perception expert model for the reconstructed gas sequence from the training set: in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer.
[0037] This invention's anomaly perception expert model is built upon an encoder-decoder architecture, with its core employing standard Transformer modules. The encoder consists of N identical layers connected in series, each layer containing a multi-head self-attention sublayer and a feedforward neural network sublayer. The self-attention mechanism captures the global dependencies between all time steps in the input gas sequence, and its multi-head design allows the model to focus on different feature representations of the sequence in parallel. The feedforward network further processes the features through nonlinear transformations. Each sublayer is supplemented with residual connections and layer normalization to ensure stable training of the deep network. The encoder's ultimate task is to encode the original input sequence into a high-dimensional feature vector containing rich contextual information. The decoder adopts a design symmetrical to the encoder and additionally introduces a masked attention mechanism to progressively decode based on the aforementioned feature representations in an autoregressive manner, generating an accurate gas concentration prediction sequence.
[0038] After sequence reconstruction is completed, residual calculation and anomaly scoring are performed. Anomaly detection analysis is conducted on the feature vectors at each time step. Specifically, the anomaly-aware expert model calculates the residual (i.e., the deviation) between the reconstructed feature vector and the original input features at that time step. This residual calculation process reflects the model's ability to "understand" the current time-series features—the larger the residual value, the more difficult it is for the model to accurately reconstruct the feature information at that position, often indicating a significant difference between the data pattern at that time point and the normal state, and a higher probability of an anomaly.
[0039] Therefore, this invention utilizes an anomaly perception expert model to calculate the residual value between the concentration of the characteristic gas in the training set and the estimated value reconstructed from the gas sequence in the training set. The magnitude of the residuals represents the quality of the model's predictions, and the calculation method is as follows: Obtaining residual values Furthermore, to quantify the probability of this anomaly, the anomaly perception expert model also designed an intelligent scoring mechanism: through a learnable linear transformation layer, the residual vector at each time step is mapped to an anomaly probability of gas concentration between 0 and 1. The weight matrix of this linear transformation layer... and bias terms These scores are automatically optimized during model training, adaptively learning the contribution of different feature dimensions to anomaly detection, ultimately outputting anomaly scores. It intuitively reflects the probability of an anomaly occurring at the current moment; the higher the value, the greater the likelihood of an anomaly.
[0040] Continue to use the anomaly perception expert model to calculate the residual values. Intelligent scoring is performed to identify abnormal concentration points in the training set. The intelligent scoring mechanism is expressed as follows: in, Represents the residual value at time j. Abnormal scores; This represents the sigmoid function, which projects a fraction onto a real number between 0 and 1. This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j.
[0041] In practical applications, a dynamic threshold (such as 0.8) can be set according to specific scenario requirements. When the abnormal score at a certain point in time exceeds this threshold, it will be determined that an abnormal event has occurred at that location, and an investigation will be conducted to confirm it.
[0042] To effectively reduce the impact of outliers on model prediction accuracy, the system employs a sequence correction strategy based on outlier scores. This strategy achieves intelligent processing of outlier data by dynamically fusing original observations and model reconstructed values. Specifically, the predicted values of gas sequence reconstructions at outlier concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstructions at those outlier concentration points (high outlier scores will prompt the use of reconstructed values instead of the original input at that time step, thus smoothing outliers). This yields the corrected predicted value sequence for the training set. For each data point at time step j, the corrected predicted value of the gas sequence reconstruction at time j in the training set is... The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
[0043] Furthermore, based on the corrected prediction sequence of the training set, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model, respectively. Then, these predicted values are weighted and fused using a gated network model to obtain the final prediction result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final prediction result of the future concentration of each characteristic gas in the training set includes: The corrected prediction sequence from the training set is input into the encoder-decoder structure, which outputs the gas concentration prediction for future time steps.
[0044] Specifically, the corrected predicted sequence from the training set is input into the encoder-decoder structure to obtain the anomaly perception expert model's predicted future concentrations of each characteristic gas in the training set. , is represented as: Where f(...) represents a temporal neural network, which can be a common temporal neural network, such as an LSTM model or a Transformer model. The choice of the two depends on the length of the time series. When the time series is long, the Transformer architecture can be used, while when it is short, the LSTM module can capture its features better. This represents the corrected predicted sequence from the training set; This represents the predicted output at time T.
[0045] In this invention, a time-series modeling expert model is used to accurately capture the dynamic characteristics of dissolved gas concentration evolution over time, focusing on modeling its inherent deterministic trends and complex periodic variation patterns. This model is constructed using a gated recurrent unit (GRU) network, taking the hidden state from the previous time step and the observed gas concentration at the current time step as input. Through a gating mechanism, it adaptively fuses old and new information, outputting an updated representation of the current state, and thereby achieving multi-step prediction of a single gas concentration sequence.
[0046] At each time step, the hidden state is dynamically updated based on the current input and the hidden state of the previous time step: the update gate determines the proportion of historical information to be retained by weighted combination of the current input and the previous hidden state; the reset gate uses a similar mechanism to determine the information to be forgotten. This gating mechanism enables the model to explicitly control the information flow, effectively balancing short-term fluctuations and long-term dependencies, thereby enhancing its ability to express complex temporal patterns.
[0047] Specifically, in this invention, the predicted sequence after the training set is corrected is input into... In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: GRU stands for Gated Recurrent Unit, which uses a gating mechanism to control the flow of information and can effectively capture long-term dependencies in time series modeling.
[0048] In the design of the context-aware expert model, the system adopts a deep temporal feature extraction architecture based on a bidirectional long short-term memory network (Bi-LSTM) to enhance the collaborative perception and joint modeling capabilities among multiple gas concentration sequences.
[0049] Specifically, a Bi-LSTM network is used to process the input matrix. This allows for bidirectional capture of the dependencies between gases. Each hidden unit of the Bi-LSTM processes all gas features at the current time step and the hidden state from the previous time step. This process integrates and generates the representation for the current time step. The hidden state vector from the last time step is then used as the global context representation.
[0050] Specifically, the corrected predicted sequences from the training set are input into the Bi-LSTM network to obtain the context-aware expert model's predictions of the future concentrations of each characteristic gas in the training set. , is represented as: Bi-LSTM stands for Bidirectional Long Short-Term Memory, which can simultaneously capture the contextual information of a sequence (features from the past and future). This represents the final hidden state of the Bi-LSTM network, which includes contextual information about the input sequence; express dimensional space; This represents the weight matrix of the output layer; This represents the bias term of the output layer.
[0051] Next, this invention requires weighted fusion of the outputs of three expert models (i.e., anomaly perception expert, time series modeling expert, and context-aware expert) to obtain the final gas concentration prediction result. To enable dynamic adjustment, a learnable gating network is introduced, which adaptively allocates the weights of each expert model based on the current input gas concentration characteristics. This achieves refined modeling and fusion prediction of multi-gas time series signals. The gating network dynamically selects models based on concentration fluctuation characteristics within the current time window and historical residual statistics, enhancing the model's adaptability to different types of gas signals, achieving collaborative reasoning and dynamic fusion. The fused prediction result is a weighted sum of the regression outputs of each expert, thereby improving the overall prediction accuracy and robustness.
[0052] Specifically, a gated network model is used to weight and fuse the predictions of the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model for the future concentrations of each characteristic gas in the training set, to obtain the final prediction result for the future concentrations of each characteristic gas in the training set. , is represented as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation, the purpose of which is to make all fractions between 0 and 1, and their sum equal to 1; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides.
[0053] Further calculations were performed on the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each characteristic gas in the training set The loss value L (to optimize the performance of the multi-expert model in the task of predicting dissolved gas concentration in transformer oil, an end-to-end joint training strategy is adopted. This training process coordinates the learning objectives of each expert model through a unified loss function, using mean squared error (MSE) as the loss value), is calculated as follows: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
[0054] Preferably, the method of optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model using the Adam optimizer based on the loss value L, and iteratively updating the weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; The decay coefficient of first-order momentum is used to control the first-order momentum. The degree of decay of the historical gradient is usually taken as a value such as 0.9; It is the decay coefficient of second momentum, used to control the second momentum. The degree of decay of the squared historical gradient is usually taken as 0.999; Represents the gradient at time t-1 The first-order momentum is used to transmit the first-order statistical information of the gradient at the previous time step; Represents the gradient at time. The second-order momentum is used to transmit the second-order statistical information of the gradient from the previous time step; Represents the neural network model parameters of a certain model (anomaly perception expert model, time series modeling expert model, context-aware expert model, and gated network model) at time t; This represents the neural network model parameters at time t-1, which is the state of the neural network model before the parameters are updated. This represents the learning rate, used to control the step size for parameter updates; Represents extremely small positive numbers (such as 10). -8 ), used to prevent the denominator from being 0.
[0055] In this invention, after each round of training (i.e., the process of optimizing the neural network model parameters) using the training set, it is also necessary to use the validation set to validate the anomaly perception expert model, time series modeling expert model, context perception expert model, and gating network model after this training. The validation set is used to perform hyperparameter tuning and early stopping, calculate the model's MSE, and select the optimal parameters of the model based on the mean squared error calculated from the validation set.
[0056] In this invention, the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gating network model are trained multiple times. The number of training rounds is generally 100, and the number of validation rounds is the same as the number of training rounds. After multiple training rounds, the entire hybrid expert structure is jointly optimized. The gating strategy and expert parameters are driven to converge collaboratively through joint regression loss. The minimum index value V is selected from the multiple validation operations (the index value V is calculated in the same way as the loss value L during training). This minimum index value V corresponds to the optimal neural network model parameters of each model at this time.
[0057] In this invention, after obtaining the predicted concentration data for each characteristic gas, a safety threshold is given for the predicted value to determine whether the predicted concentration data will exceed the safety threshold of the gas. If it exceeds a certain danger threshold, an alarm will be triggered, which can be used for early warning.
[0058] Furthermore, if it is necessary to test and compare the anomaly perception expert model, time series modeling expert model, context perception expert model, and gated network model with the optimal parameters of the present invention, and to measure their performance, the overall dataset can be further divided into training set, validation set, and test set according to a time ratio of 8:1:1. For example, all historical concentration data corresponding to the first 80% of the time in the overall dataset can be divided into training set, all historical concentration data corresponding to the middle 10% of the time can be divided into validation set, and all historical concentration data corresponding to the last 10% of the time can be divided into test set. Then, the test set can be used to perform the test in the same way as the validation operation.
[0059] During the test set evaluation process, the gated network model integrates the outputs of three expert models from the validation set and receives oil and gas concentration data from the test set to predict the dissolved gas content in transformer oil. The mean squared error is calculated by comparing the model's predictions for the test set with the characteristic gas concentration values of the corrected prediction sequence, serving as the core indicator for quantifying prediction accuracy. This method effectively measures the model's generalization performance in time series prediction tasks, thus providing validation for its reliability in practical applications.
[0060] In practical applications, this invention can deploy a trained and validated multi-expert learning model in a real-world power equipment monitoring system. By collecting dissolved gas concentration parameters in oil from AC equipment status sensors in real time, and leveraging the collaborative predictive capabilities of anomaly perception experts, time-series modeling experts, and context-aware experts, it achieves high-precision prediction of gas concentrations at future moments. Based on the consistency index output by the experts, it detects potential anomalies or sensor drift. The system employs a dynamic threshold determination mechanism. When the deviation between the measured concentration value and the predicted value exceeds a threshold set based on the historical error standard deviation, a multi-level early warning mechanism is immediately triggered: from local alarms via LED indicators and buzzers to remote notifications automatically pushed to the monitoring center, forming a complete anomaly response closed loop. Simultaneously, the system automatically initiates sensor self-test procedures, generating a diagnostic report containing abnormal gas composition analysis, potential fault type inference, and maintenance recommendations. Through an industrial IoT platform, it enables functions such as predictive data visualization, historical anomaly tracing, and automatic maintenance work order generation, providing intelligent decision support for early warning and trend analysis of latent transformer faults and for equipment condition-based maintenance.
[0061] like Figure 2 As shown, the deployment device can consist of the following parts, which are connected through a system bus to work together to complete the real-time inference and early warning functions of the multi-expert model. The functions of each module are as follows: The Central Processing Unit (CPU) is responsible for the control and scheduling of the entire system, loading and executing the multi-expert model programs resident in memory. These include: anomaly detection and sequence correction by the anomaly perception expert module; parallel prediction by the time-series modeling expert and context-aware expert; and dynamic weight calculation and output fusion by the gating fusion module. Non-transitory memory is used to persistently store: weight files and network structure descriptions for each expert model and gating network; the system bootloader and inference engine; historical concentration data cache; and intermediate results. The memory and CPU exchange data via a bus. The data acquisition interface serves as the input channel, acquiring real-time concentration sequences from the dissolved gas sensor in transformer oil or the upper-level monitoring system and forwarding them to the processing unit. Visualization and alarm interfaces are also included. Output modules include: an industrial touchscreen, LED indicators, buzzers, or pushing prediction results and anomaly alarms to the monitoring center via Ethernet / wireless network, visually displaying concentration curves and warning levels.
[0062] During runtime, the CPU first reads the latest sequence from the data acquisition interface and then sequentially calls each expert module for calculation, while the GPU assists in completing the deep network's forward inference. All intermediate data and model execution files are loaded or written from memory. The final fused concentration prediction and anomaly level are fed back in real time through a visualization interface, enabling online monitoring and early warning of the transformer oil gas status.
[0063] The method provided by this invention has good engineering feasibility and adaptability to time series modeling. Combining machine learning, time series data analysis, and multi-model fusion technology, it constructs an expert system with differentiated modeling capabilities and achieves dynamic fusion through a gating mechanism. It has strong robustness and scalability, and can simultaneously complete trend prediction, anomaly detection, and state classification of multiple characteristic gas concentrations. It can efficiently complete regression prediction and anomaly detection tasks for multiple dissolved gas concentrations, and has good generalization and real-time response capabilities. It provides reliable data support and decision-making basis for multi-level early warning of dissolved gases in oil, and provides important basis for early fault warning and health status assessment of oil-immersed power equipment such as power transformers, thereby improving the intelligent monitoring capabilities and prediction accuracy of the system.
[0064] This invention addresses the industry pain point of complex operating environments in ultra-high voltage (UHV) transformers, where single regression models struggle to accurately capture the temporal characteristics of dissolved gases. It proposes an integrated solution for dissolved gas prediction and anomaly detection in oil based on multi-expert learning. By integrating a regression model driven by multi-source big data and an adaptive anomaly detection algorithm, an intelligent expert system with continuous optimization capabilities is constructed, significantly improving prediction accuracy and environmental robustness, thus meeting the stringent requirements of UHV equipment condition monitoring.
[0065] Technical advantages of this invention: 1. Accurate Prediction: Integrating the advantages of multiple models to achieve high-precision time series regression. 2. Intelligent Early Warning: Prediction results are directly integrated with multi-level early warning modules to support preventative decision-making. 3. Engineering-friendly: Flexible deployment, efficient computation, and adaptable to complex substation operating conditions. This invention provides reliable technical support for transformer condition monitoring, and its prediction results can directly serve the insulating oil fault early warning system, contributing to the safe and stable operation of the power grid.
[0066] Example 2 like Figure 3 The method for regression prediction and anomaly detection classification of dissolved gases in oil based on multi-expert fusion learning, as shown, includes: The quartile range method was used to filter the noise in the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets of each characteristic gas. All the filtered historical concentration time series datasets of the characteristic gases were combined into a whole dataset, and the whole dataset was divided into training set and validation set according to time. The predicted values of gas sequences reconstructed from abnormal concentration points in the training set are corrected using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on this corrected predicted value sequence, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, a gating network model is used to weight and fuse these predicted values to obtain the final predicted result of the future concentration of each characteristic gas in the training set. Based on the final predicted result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training session, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, at which point the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The entire dataset is input into the anomaly perception expert model, time series modeling expert model, context-aware expert model, and gated network model with optimal neural network model parameters to obtain the predicted concentration data for each characteristic gas.
[0067] In this invention, the characteristic gas refers to the characteristic gas in insulating oil. The characteristic gases in insulating oil typically include seven types: H2, CH4, C2H6, C2H4, C2H2, CO, and CO2. In this embodiment, the concentration data of these seven characteristic gases are detected and predicted.
[0068] In this invention, the historical concentration data of each characteristic gas in insulating oil are organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
[0069] Specifically, this invention first requires extracting historical concentration data (unit: ppm) of seven characteristic gases (H2, CH4, C2H6, C2H4, C2H2, CO, CO2) from insulating oil. The gas concentration detection strictly follows the national standard "Gas Chromatography Determination of Dissolved Gas Components in Insulating Oil," and its detection process includes three main steps: first, standardized collection of transformer oil samples; second, degassing of dissolved gases in the oil; and finally, separation and quantitative detection of gas components using a gas chromatograph, i.e., detecting the concentration of each characteristic gas in the oil sample. The original data comes from long-term monitoring records of ultra-high voltage transformers across the entire network. The historical oil and gas concentration readings of each independent substation are organized chronologically and constructed into a time-series dataset of historical concentrations for each characteristic gas.
[0070] Preferably, the present invention employs the interquartile range method to perform noise filtering on the historical concentration time-series datasets of each characteristic gas, and the method for obtaining the filtered historical concentration time-series datasets of each characteristic gas includes: calculating the quartiles of the historical concentration time-series datasets of each characteristic gas, and removing the historical concentration time-series datasets of each characteristic gas that fall within the specified intervals. The data points outside of these (each characteristic gas's historical concentration time series dataset corresponds to a different interval) Any interval in the historical concentration time series dataset of the characteristic gas that falls within the range of the characteristic gas. Data points outside the range will be considered outliers and removed. This will allow you to obtain the values for each characteristic gas within its respective range. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration datasets within these time points are all removed to improve data quality, resulting in a filtered historical concentration time series dataset for each characteristic gas (the time series in the filtered historical concentration time series datasets for each characteristic gas are the same). Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas (i.e., the value located at the 25th percentile of the sorted data); Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas (i.e., the value located at the 75th percentile of the sorted data); k represents the interval parameter of the interquartile range method, which is used to adjust the normal concentration range of the characteristic gas, and is generally taken as 1.5; IQR represents the interquartile range, which is a statistic describing the degree of data dispersion, and its core function is to reflect the distribution range of the middle 50% of the data in the dataset.
[0071] In a specific implementation, the overall dataset (which still consists of historical concentrations of characteristic gases arranged chronologically) is divided into training and validation sets according to time. Simultaneously, a regression target is automatically generated for each sample to support the system's learning. Specifically, the training and validation sets can be divided in a time ratio of 8:2 or 9:1. For example, all historical concentration data corresponding to the first 80% or 90% of the time period in the overall dataset can be assigned to the training set, and all historical concentration data corresponding to the last 20% or 10% of the time period can be assigned to the validation set. This division method strictly avoids future information leakage, ensuring the reliability and generalization ability of the model evaluation results.
[0072] Furthermore, this invention constructs an anomaly perception expert model, a time-series modeling expert model, and a context-aware expert model. The anomaly perception expert scores the input sequence for anomalies, locates potential mutation points, and performs soft adjustments on detected abnormal regions. This expert simultaneously outputs corrected concentration predictions and anomaly classification results. The time-series modeling expert models gas volatility, captures stable concentration fluctuations, focuses on the trend of gas concentration changes over time, and outputs regression predictions. The context-aware expert focuses on the interrelationships between different gases, models the dependencies between different gases, improves the ability to model multivariate time-series features, and outputs regression predictions based on the influence between gases. The design goal of the anomaly perception expert is to accurately identify outliers in the characteristic gas concentration sequence and correct the sequence accordingly, thereby improving the accuracy and stability of subsequent regression predictions. The main process of the anomaly perception expert model can be divided into three parts: sequence encoding and reconstruction, residual calculation and anomaly scoring, and sequence correction.
[0073] Specifically, methods for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set include: The concentration features of the gas sequences in the training set are constructed using the sequence encoding and reconstruction part of the anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; The value represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional.
[0074] To further effectively identify the location of abnormal gas values in the sequence, the anomaly perception expert model uses an encoder-decoder structure to reconstruct the gas sequence from the concentration features of the obtained training set, thus obtaining the predicted value of the anomaly perception expert model for the reconstructed gas sequence from the training set: in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer.
[0075] This invention's anomaly perception expert model is built upon an encoder-decoder architecture, with its core employing standard Transformer modules. The encoder consists of N identical layers connected in series, each layer containing a multi-head self-attention sublayer and a feedforward neural network sublayer. The self-attention mechanism captures the global dependencies between all time steps in the input gas sequence, and its multi-head design allows the model to focus on different feature representations of the sequence in parallel. The feedforward network further processes the features through nonlinear transformations. Each sublayer is supplemented with residual connections and layer normalization to ensure stable training of the deep network. The encoder's ultimate task is to encode the original input sequence into a high-dimensional feature vector containing rich contextual information. The decoder adopts a design symmetrical to the encoder and additionally introduces a masked attention mechanism to progressively decode based on the aforementioned feature representations in an autoregressive manner, generating an accurate gas concentration prediction sequence.
[0076] After sequence reconstruction is completed, residual calculation and anomaly scoring are performed. Anomaly detection analysis is conducted on the feature vectors at each time step. Specifically, the anomaly-aware expert model calculates the residual (i.e., the deviation) between the reconstructed feature vector and the original input features at that time step. This residual calculation process reflects the model's ability to "understand" the current time-series features—the larger the residual value, the more difficult it is for the model to accurately reconstruct the feature information at that position, often indicating a significant difference between the data pattern at that time point and the normal state, and a higher probability of an anomaly.
[0077] Therefore, this invention utilizes an anomaly perception expert model to calculate the residual value between the concentration of the characteristic gas in the training set and the estimated value reconstructed from the gas sequence in the training set. The magnitude of the residuals represents the quality of the model's predictions, and the calculation method is as follows: Obtaining residual values Furthermore, to quantify the probability of this anomaly, the anomaly perception expert model also designed an intelligent scoring mechanism: through a learnable linear transformation layer, the residual vector at each time step is mapped to an anomaly probability of gas concentration between 0 and 1. The weight matrix of this linear transformation layer... and bias terms These scores are automatically optimized during model training, adaptively learning the contribution of different feature dimensions to anomaly detection, ultimately outputting anomaly scores. It intuitively reflects the probability of an anomaly occurring at the current moment; the higher the value, the greater the likelihood of an anomaly.
[0078] Continue to use the anomaly perception expert model to calculate the residual values. Intelligent scoring is performed to identify abnormal concentration points in the training set. The intelligent scoring mechanism is expressed as follows: in, Represents the residual value at time j. Abnormal scores; This represents the sigmoid function, which projects a fraction onto a real number between 0 and 1. This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j.
[0079] In practical applications, a dynamic threshold (such as 0.8) can be set according to specific scenario requirements. When the abnormal score at a certain point in time exceeds this threshold, it will be determined that an abnormal event has occurred at that location, and an investigation will be conducted to confirm it.
[0080] To effectively reduce the impact of outliers on model prediction accuracy, the system employs a sequence correction strategy based on outlier scores. This strategy achieves intelligent processing of outlier data by dynamically fusing original observations and model reconstructed values. Specifically, the predicted values of gas sequence reconstructions at outlier concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstructions at those outlier concentration points (high outlier scores will prompt the use of reconstructed values instead of the original input at that time step, thus smoothing outliers). This yields the corrected predicted value sequence for the training set. For each data point at time step j, the corrected predicted value of the gas sequence reconstruction at time j in the training set is... The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
[0081] Furthermore, based on the corrected prediction sequence of the training set, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model, respectively. Then, these predicted values are weighted and fused using a gated network model to obtain the final prediction result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final prediction result of the future concentration of each characteristic gas in the training set includes: The corrected prediction sequence from the training set is input into the encoder-decoder structure, which outputs the gas concentration prediction for future time steps.
[0082] Specifically, the corrected predicted sequence from the training set is input into the encoder-decoder structure to obtain the anomaly perception expert model's predicted future concentrations of each characteristic gas in the training set. , is represented as: Where f(...) represents a temporal neural network, which can be a common temporal neural network, such as an LSTM model or a Transformer model. The choice of the two depends on the length of the time series. When the time series is long, the Transformer architecture can be used, while when it is short, the LSTM module can capture its features better. This represents the corrected predicted sequence from the training set; This represents the predicted output at time T.
[0083] In this invention, a time-series modeling expert model is used to accurately capture the dynamic characteristics of dissolved gas concentration evolution over time, focusing on modeling its inherent deterministic trends and complex periodic variation patterns. This model is constructed using a gated recurrent unit (GRU) network, taking the hidden state from the previous time step and the observed gas concentration at the current time step as input. Through a gating mechanism, it adaptively fuses old and new information, outputting an updated representation of the current state, and thereby achieving multi-step prediction of a single gas concentration sequence.
[0084] At each time step, the hidden state is dynamically updated based on the current input and the hidden state of the previous time step: the update gate determines the proportion of historical information to be retained by weighted combination of the current input and the previous hidden state; the reset gate uses a similar mechanism to determine the information to be forgotten. This gating mechanism enables the model to explicitly control the information flow, effectively balancing short-term fluctuations and long-term dependencies, thereby enhancing its ability to express complex temporal patterns.
[0085] Specifically, in this invention, the predicted sequence after the training set is corrected is input into... In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: GRU stands for Gated Recurrent Unit, which uses a gating mechanism to control the flow of information and can effectively capture long-term dependencies in time series modeling.
[0086] In the design of the context-aware expert model, the system adopts a deep temporal feature extraction architecture based on a bidirectional long short-term memory network (Bi-LSTM) to enhance the collaborative perception and joint modeling capabilities among multiple gas concentration sequences.
[0087] Specifically, a Bi-LSTM network is used to process the input matrix. This allows for bidirectional capture of the dependencies between gases. Each hidden unit of the Bi-LSTM processes all gas features at the current time step and the hidden state from the previous time step. This process integrates and generates the representation for the current time step. The hidden state vector from the last time step is then used as the global context representation.
[0088] Specifically, the corrected predicted sequences from the training set are input into the Bi-LSTM network to obtain the context-aware expert model's predictions of the future concentrations of each characteristic gas in the training set. , is represented as: Bi-LSTM stands for Bidirectional Long Short-Term Memory, which can simultaneously capture the contextual information of a sequence (features from the past and future). This represents the final hidden state of the Bi-LSTM network, which includes contextual information about the input sequence; express dimensional space; This represents the weight matrix of the output layer; This represents the bias term of the output layer.
[0089] Next, this invention requires weighted fusion of the outputs of three expert models (i.e., anomaly perception expert, time series modeling expert, and context-aware expert) to obtain the final gas concentration prediction result. To enable dynamic adjustment, a learnable gating network is introduced, which adaptively allocates the weights of each expert model based on the current input gas concentration characteristics. This achieves refined modeling and fusion prediction of multi-gas time series signals. The gating network dynamically selects models based on concentration fluctuation characteristics within the current time window and historical residual statistics, enhancing the model's adaptability to different types of gas signals, achieving collaborative reasoning and dynamic fusion. The fused prediction result is a weighted sum of the regression outputs of each expert, thereby improving the overall prediction accuracy and robustness.
[0090] Specifically, a gated network model is used to weight and fuse the predictions of the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model for the future concentrations of each characteristic gas in the training set, to obtain the final prediction result for the future concentrations of each characteristic gas in the training set. , is represented as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation, the purpose of which is to make all fractions between 0 and 1, and their sum equal to 1; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides.
[0091] Further calculations were performed on the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each characteristic gas in the training set The loss value L (to optimize the performance of the multi-expert model in the task of predicting dissolved gas concentration in transformer oil, an end-to-end joint training strategy is adopted. This training process coordinates the learning objectives of each expert model through a unified loss function, using mean squared error (MSE) as the loss value), is calculated as follows: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
[0092] Preferably, the method of optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model using the Adam optimizer based on the loss value L, and iteratively updating the weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; The decay coefficient of first-order momentum is used to control the first-order momentum. The degree of decay of the historical gradient is usually taken as a value such as 0.9; It is the decay coefficient of second momentum, used to control the second momentum. The degree of decay of the squared historical gradient is usually taken as 0.999; Represents the gradient at time t-1 The first-order momentum is used to transmit the first-order statistical information of the gradient at the previous time step; Represents the gradient at time. The second-order momentum is used to transmit the second-order statistical information of the gradient from the previous time step; Represents the neural network model parameters of a certain model (anomaly perception expert model, time series modeling expert model, context-aware expert model, and gated network model) at time t; This represents the neural network model parameters at time t-1, which is the state of the neural network model before the parameters are updated. This represents the learning rate, used to control the step size for parameter updates; Represents extremely small positive numbers (such as 10). -8 ), used to prevent the denominator from being 0.
[0093] In this invention, after each round of training (i.e., the process of optimizing the neural network model parameters) using the training set, it is also necessary to use the validation set to validate the anomaly perception expert model, time series modeling expert model, context perception expert model, and gating network model after this training. The validation set is used to perform hyperparameter tuning and early stopping, calculate the model's MSE, and select the optimal parameters of the model based on the mean squared error calculated from the validation set.
[0094] In this invention, the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gating network model are trained multiple times. The number of training rounds is generally 100, and the number of validation rounds is the same as the number of training rounds. After multiple training rounds, the entire hybrid expert structure is jointly optimized. The gating strategy and expert parameters are driven to converge collaboratively through joint regression loss. The minimum index value V is selected from the multiple validation operations (the index value V is calculated in the same way as the loss value L during training). This minimum index value V corresponds to the optimal neural network model parameters of each model at this time.
[0095] In this invention, after obtaining the predicted concentration data for each characteristic gas, a safety threshold is given for the predicted value to determine whether the predicted concentration data will exceed the safety threshold of the gas. If it exceeds a certain danger threshold, an alarm will be triggered, which can be used for early warning.
[0096] Furthermore, if it is necessary to test and compare the anomaly perception expert model, time series modeling expert model, context perception expert model, and gated network model with the optimal parameters of the present invention, and to measure their performance, the overall dataset can be further divided into training set, validation set, and test set according to a time ratio of 8:1:1. For example, all historical concentration data corresponding to the first 80% of the time in the overall dataset can be divided into training set, all historical concentration data corresponding to the middle 10% of the time can be divided into validation set, and all historical concentration data corresponding to the last 10% of the time can be divided into test set. Then, the test set can be used to perform the test in the same way as the validation operation.
[0097] During the test set evaluation process, the gated network model integrates the outputs of three expert models from the validation set and receives oil and gas concentration data from the test set to predict the dissolved gas content in transformer oil. The mean squared error is calculated by comparing the model's predictions for the test set with the characteristic gas concentration values of the corrected prediction sequence, serving as the core indicator for quantifying prediction accuracy. This method effectively measures the model's generalization performance in time series prediction tasks, thus providing validation for its reliability in practical applications.
[0098] In practical applications, this invention can deploy a trained and validated multi-expert learning model in a real-world power equipment monitoring system. By collecting dissolved gas concentration parameters in oil from AC equipment status sensors in real time, and leveraging the collaborative predictive capabilities of anomaly perception experts, time-series modeling experts, and context-aware experts, it achieves high-precision prediction of gas concentrations at future moments. Based on the consistency index output by the experts, it detects potential anomalies or sensor drift. The system employs a dynamic threshold determination mechanism. When the deviation between the measured concentration value and the predicted value exceeds a threshold set based on the historical error standard deviation, a multi-level early warning mechanism is immediately triggered: from local alarms via LED indicators and buzzers to remote notifications automatically pushed to the monitoring center, forming a complete anomaly response closed loop. Simultaneously, the system automatically initiates sensor self-test procedures, generating a diagnostic report containing abnormal gas composition analysis, potential fault type inference, and maintenance recommendations. Through an industrial IoT platform, it enables functions such as predictive data visualization, historical anomaly tracing, and automatic maintenance work order generation, providing intelligent decision support for early warning and trend analysis of latent transformer faults and for equipment condition-based maintenance.
[0099] like Figure 2As shown, the deployment device can consist of the following parts, which are connected through a system bus to work together to complete the real-time inference and early warning functions of the multi-expert model. The functions of each module are as follows: The Central Processing Unit (CPU) is responsible for the control and scheduling of the entire system, loading and executing the multi-expert model programs resident in memory. These include: anomaly detection and sequence correction by the anomaly perception expert module; parallel prediction by the time-series modeling expert and context-aware expert; and dynamic weight calculation and output fusion by the gating fusion module. Non-transitory memory is used to persistently store: weight files and network structure descriptions for each expert model and gating network; the system bootloader and inference engine; historical concentration data cache; and intermediate results. The memory and CPU exchange data via a bus. The data acquisition interface serves as the input channel, acquiring real-time concentration sequences from the dissolved gas sensor in transformer oil or the upper-level monitoring system and forwarding them to the processing unit. Visualization and alarm interfaces are also included. Output modules include: an industrial touchscreen, LED indicators, buzzers, or pushing prediction results and anomaly alarms to the monitoring center via Ethernet / wireless network, visually displaying concentration curves and warning levels.
[0100] During runtime, the CPU first reads the latest sequence from the data acquisition interface and then sequentially calls each expert module for calculation, while the GPU assists in completing the deep network's forward inference. All intermediate data and model execution files are loaded or written from memory. The final fused concentration prediction and anomaly level are fed back in real time through a visualization interface, enabling online monitoring and early warning of the transformer oil gas status.
[0101] The method provided by this invention has good engineering feasibility and adaptability to time series modeling. Combining machine learning, time series data analysis, and multi-model fusion technology, it constructs an expert system with differentiated modeling capabilities and achieves dynamic fusion through a gating mechanism. It has strong robustness and scalability, and can simultaneously complete trend prediction, anomaly detection, and state classification of multiple characteristic gas concentrations. It can efficiently complete regression prediction and anomaly detection tasks for multiple dissolved gas concentrations, and has good generalization and real-time response capabilities. It provides reliable data support and decision-making basis for multi-level early warning of dissolved gases in oil, and provides important basis for early fault warning and health status assessment of oil-immersed power equipment such as power transformers, thereby improving the intelligent monitoring capabilities and prediction accuracy of the system.
[0102] This invention addresses the industry pain point of complex operating environments in ultra-high voltage (UHV) transformers, where single regression models struggle to accurately capture the temporal characteristics of dissolved gases. It proposes an integrated solution for dissolved gas prediction and anomaly detection in oil based on multi-expert learning. By integrating a regression model driven by multi-source big data and an adaptive anomaly detection algorithm, an intelligent expert system with continuous optimization capabilities is constructed, significantly improving prediction accuracy and environmental robustness, thus meeting the stringent requirements of UHV equipment condition monitoring.
[0103] Technical advantages of this invention: 1. Accurate Prediction: Integrating the advantages of multiple models to achieve high-precision time series regression. 2. Intelligent Early Warning: Prediction results are directly integrated with multi-level early warning modules to support preventative decision-making. 3. Engineering-friendly: Flexible deployment, efficient computation, and adaptable to complex substation operating conditions. This invention provides reliable technical support for transformer condition monitoring, and its prediction results can directly serve the insulating oil fault early warning system, contributing to the safe and stable operation of the power grid.
[0104] Example 3 A computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of Embodiment 2.
[0105] It should be understood that any parts not described in detail in this specification belong to the prior art.
[0106] The preferred embodiments of the present invention have been described in detail above; however, the present invention is not limited thereto. Within the scope of the inventive concept, various simple modifications can be made to the technical solutions of the present invention, including combinations of various technical features in any other suitable manner. These simple modifications and combinations should also be considered as the content disclosed in the present invention and are all within the protection scope of the present invention.
Claims
1. A system for regression prediction and abnormal state detection classification of dissolved gases in oil based on multi-expert fusion learning, characterized in that, include: The data preprocessing module is used to perform noise filtering on the historical concentration time series datasets of each characteristic gas using the interquartile range method, to obtain the filtered historical concentration time series datasets of each characteristic gas, to form the overall dataset of all the filtered historical concentration time series datasets of all characteristic gases, and to divide the overall dataset into training set and validation set according to time. The filtering module is used to correct the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on the corrected predicted value sequence, the prediction values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, these prediction values are weighted and fused using a gating network model to obtain the final prediction result of the future concentration of each characteristic gas in the training set. Based on the final prediction result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, and the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The prediction module is used to input the overall dataset into the anomaly perception expert model, time series modeling expert model, context perception expert model and gated network model with optimal neural network model parameters to obtain the predicted concentration data of each characteristic gas. The method for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set includes: The concentration features of the gas sequences in the training set are constructed using an anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; This represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional. An encoder-decoder structure is used to reconstruct the gas sequences from the obtained training set based on their concentration features, thus obtaining the anomaly perception expert model's prediction of the reconstructed gas sequences from the training set. in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer. The residual values between the concentrations of characteristic gases in the training set and the predicted values reconstructed from the gas sequences in the training set were calculated using an anomaly perception expert model. The calculation method is as follows: Continue to use the anomaly perception expert model to calculate the residual values. Intelligent scoring is performed to identify abnormal concentration points in the training set. The intelligent scoring mechanism is expressed as follows: in, Represents the residual value at time j. Abnormal scores; Represents the sigmoid function; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j; The predicted values of the gas sequence reconstruction at abnormal concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstruction at those abnormal concentration points, resulting in the corrected predicted sequence for the training set. The corrected predicted gas sequence reconstruction value at time j in the training set is then used. The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
2. The oil-based dissolved gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning as described in claim 1, characterized in that, The historical concentration data of each characteristic gas in the insulating oil were organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
3. The oil-based dissolved gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning according to claim 1 or 2, characterized in that, The method of using the interquartile range (IMR) to filter noise from the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets for each characteristic gas includes: removing noise from the historical concentration time series datasets of each characteristic gas within the specified intervals. Data points outside the range are used to obtain the range for each characteristic gas. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration dataset within the range is processed by removing data points at these time points to obtain a filtered historical concentration time series dataset for each feature gas. Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas; Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas; k represents the interval parameter of the interquartile range method; and IQR represents the interquartile range.
4. The oil-soluble gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning according to claim 1, characterized in that, Based on the corrected predicted value sequences in the training set, the predicted values of the future concentrations of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model, respectively. Then, these predicted values are weighted and fused using a gated network model to obtain the final predicted result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final predicted result of the future concentration of each characteristic gas in the training set includes: The corrected predicted sequence from the training set is input into the encoder-decoder structure to obtain the anomaly perception expert model's prediction of the future concentration of each characteristic gas in the training set. , is represented as: Where f(...) represents a temporal neural network; This represents the corrected predicted sequence from the training set; This represents the predicted output at time T. Input the corrected prediction sequence from the training set. In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: The corrected predicted sequences from the training set are input into the Bi-LSTM network to obtain the context-aware expert model's predictions of the future concentrations of each characteristic gas in the training set. , is represented as: in, This represents the final hidden state of the Bi-LSTM network; express dimensional space; This represents the weight matrix of the output layer; Indicates the bias term of the output layer; A gated network model is used to weight and fuse the predictions of the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model for the future concentrations of each characteristic gas in the training set, resulting in the final prediction of the future concentrations of each characteristic gas in the training set. , is represented as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides; Calculate the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each feature gas in the training set The loss value L between the two is calculated using the following formula: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
5. The oil-based dissolved gas regression prediction and abnormal state detection classification system based on multi-expert fusion learning according to claim 4, characterized in that, Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gated network model. The method of iteratively updating the weights of the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; This represents the decay coefficient of first-order momentum; It is the decay coefficient of second momentum; Represents the gradient at time t-1 The first-order momentum; Represents the gradient at time. The second momentum; This represents the neural network model parameters of a certain model at time t; This represents the neural network model parameters of a certain model at time t-1; Indicates the learning rate; It represents a very small positive number.
6. A method for regression prediction and abnormal state detection and classification of dissolved gases in oil based on multi-expert fusion learning, characterized in that, include: The quartile range method was used to filter the noise in the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets of each characteristic gas. All the filtered historical concentration time series datasets of the characteristic gases were combined into a whole dataset, and the whole dataset was divided into training set and validation set according to time. The predicted values of gas sequences reconstructed from abnormal concentration points in the training set are corrected using an anomaly perception expert model, resulting in a corrected predicted value sequence for the training set. Based on this corrected predicted value sequence, the predicted values of the future concentration of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context awareness expert model, respectively. Then, a gating network model is used to weight and fuse these predicted values to obtain the final predicted result of the future concentration of each characteristic gas in the training set. Based on the final predicted result of the future concentration of each characteristic gas in the training set, the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are optimized to obtain the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model after this training. After each training session, the trained anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model are validated using a validation set until the training is completed, at which point the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gating network model with the optimal neural network model parameters are selected. The entire dataset is input into the anomaly perception expert model, time series modeling expert model, context perception expert model and gated network model with optimal neural network model parameters to obtain the predicted concentration data of each characteristic gas. The method for correcting the predicted values of gas sequences reconstructed from abnormal concentration points in the training set using an anomaly perception expert model to obtain the corrected predicted value sequences of the training set includes: The concentration features of the gas sequences in the training set are constructed using an anomaly perception expert model, and are represented as follows: in, The value represents the concentration of the characteristic gas within the training window of the training set; L represents the window length of the gas sequence. This represents the concentration feature vector at time step t; This indicates the concentration of the characteristic gas before time L; This represents the characteristic gas concentration at all times within the window; m indicates that the characteristic gas concentration eigenvector is m-dimensional. An encoder-decoder structure is used to reconstruct the gas sequences from the obtained training set based on their concentration features, thus obtaining the anomaly perception expert model's prediction of the reconstructed gas sequences from the training set. in, This represents the concentration characteristics of the feature gas in the training set within the hidden layer; This is the predicted value of the anomaly perception expert model for reconstructing gas sequences in the training set; d represents the value of the characteristic gas concentration in the hidden layer; d represents the dimension of the concentration feature in the hidden layer. The residual value between the concentration of the characteristic gas in the training set and the predicted value of the gas sequence reconstruction is calculated by using the anomaly perception expert model The calculation method is as follows: continue to utilize the anomaly perception expert model on the calculated residual values intelligent scoring is performed, and the abnormal concentration points in the training set are determined according to the intelligent scoring, wherein the intelligent scoring mechanism is represented as: in, Represents the residual value at time j. Abnormal scores; Represents the sigmoid function; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; This represents the residual value at time j; The predicted values of the gas sequence reconstruction at abnormal concentration points in the training set are corrected, and the corrected predicted values are used to replace the predicted values of the gas sequence reconstruction at those abnormal concentration points, resulting in the corrected predicted sequence for the training set. The corrected predicted gas sequence reconstruction value at time j in the training set is then used. The calculation method is as follows: in, This represents the concentration of the characteristic gas at time j within the training window of the training set. This represents the predicted value of the anomaly perception expert model for reconstructing the gas sequence in the training set at time j.
7. The method according to claim 6, wherein the method is characterized by, The historical concentration data of each characteristic gas in the insulating oil were organized in chronological order to obtain a time-series dataset of the historical concentration of each characteristic gas.
8. The method for regression prediction and abnormal state detection classification of dissolved gases in oil based on multi-expert fusion learning according to claim 6 or 7, characterized in that, The method of using the interquartile range (IMR) to filter noise from the historical concentration time series datasets of each characteristic gas to obtain the filtered historical concentration time series datasets for each characteristic gas includes: removing noise from the historical concentration time series datasets of each characteristic gas within the specified intervals. Data points outside the range are used to obtain the range for each characteristic gas. The historical concentration dataset within the range is used to determine the time points corresponding to the data points where each characteristic gas was removed, and then each characteristic gas is placed within the interval. The historical concentration dataset within the range is processed by removing data points at these time points to obtain a filtered historical concentration time series dataset for each feature gas. Where Q1 represents the first quartile of the historical concentration time series dataset of the characteristic gas; Q3 represents the third quartile of the historical concentration time series dataset of the characteristic gas; k represents the interval parameter of the interquartile range method; and IQR represents the interquartile range.
9. The method for regression prediction and abnormal state detection and classification of dissolved gases in oil based on multi-expert fusion learning according to claim 6, characterized in that, Based on the corrected predicted value sequences in the training set, the predicted values of the future concentrations of each characteristic gas in the training set are calculated by the anomaly perception expert model, the time series modeling expert model, and the context-aware expert model, respectively. Then, these predicted values are weighted and fused using a gated network model to obtain the final predicted result of the future concentration of each characteristic gas in the training set. The method for optimizing the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context-aware expert model, and the gated network model based on the final predicted result of the future concentration of each characteristic gas in the training set includes: The revised estimated value sequence of the training set is input into the encoder-decoder structure to obtain the prediction value of the anomaly perception expert model for the future concentration of each characteristic gas of the training set , which is expressed as: Where f(...) represents a temporal neural network; This represents the corrected predicted sequence from the training set; This represents the predicted output at time T. Input the corrected prediction sequence from the training set. In the network, the predicted future concentrations of each characteristic gas in the training set are obtained by the time-series modeling expert model. , is represented as: The revised estimated value sequence of the training set is input into a Bi-LSTM network to obtain a prediction value of the context-aware expert model for the future concentration of each characteristic gas in the training set is represented as: wherein, denotes the final hidden state of the Bi-LSTM network; denotes dimensional space; denotes the weight matrix of the output layer; denotes the bias term of the output layer; The prediction value of each characteristic gas in the training set is obtained by using the gating network model to fuse the prediction value of each characteristic gas in the training set by the anomaly perception expert model, the prediction value of each characteristic gas in the training set by the time series modeling expert model and the prediction value of each characteristic gas in the training set by the context perception expert model , is expressed as: in, This indicates the weight that each expert model should be assigned; It is a normalization operation; This represents the weight matrix of the linear transformation layer; This represents the bias term of the linear transformation layer; Represents the weights of the anomaly detection expert model; Represents the weights of the time series modeling expert model; Represents the weights of the context-aware expert model; express The real number space in which it resides; Calculate the concentration of each characteristic gas in the corrected predicted sequence from the training set. The final prediction results of the future concentration of each characteristic gas in the training set The loss value L between the two is calculated using the following formula: Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network model. The weights of the anomaly perception expert model, the time series modeling expert model, the context awareness expert model, and the gated network are iteratively updated using the gradient descent algorithm.
10. The method of claim 9, wherein the method is based on multi-expert fusion learning for dissolved gas regression prediction and abnormal condition detection classification. Based on the loss value L, the Adam optimizer is used to optimize the neural network model parameters of the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gated network model. The method of iteratively updating the weights of the anomaly perception expert model, the temporal modeling expert model, the context-aware expert model, and the gated network using the gradient descent algorithm includes: in, This represents the gradient of the model during backpropagation at time t. The first-order momentum; This represents the gradient of the model during backpropagation at time t. The second momentum; This represents the decay coefficient of first-order momentum; It is the decay coefficient of second momentum; Represents the gradient at time t-1 The first-order momentum; Represents the gradient at time. The second momentum; This represents the neural network model parameters of a certain model at time t; This represents the neural network model parameters of a certain model at time t-1; Indicates the learning rate; It represents a very small positive number.
11. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 6-10.