Prediction method of dissolved gas in transformer oil based on improved BiGRU network

Through the improved BiGRU network and a variety of optimization technologies, the problem of insufficient local anomaly sensitivity and low prediction accuracy for the identification of dissolved gas abnormal data in transformer oil is solved, and higher prediction accuracy and model generalization capabilities are achieved.

CN119760324BActive Publication Date: 2025-05-09SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510255970.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-09
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The prior art lacks sensitivity to local anomalies in the identification of abnormal data of dissolved gases in transformer oil, and the empirical modal decomposition method is susceptible to modal aliasing and endpoint effects, resulting in a decrease in prediction accuracy.

Method used

The improved BiGRU network is adopted, combined with the SDROF anomaly detection algorithm, weighted fusion filling strategy, variational modal decomposition VMD and Black-winged Kitchen optimization algorithm BKA, and the prediction window length and transition points are optimized, and the initial parameters of the BiGRU network are optimized through the meta-learning MAML algorithm.

Benefits of technology

It improves the ability to identify local anomalies, enhances prediction accuracy and generalization performance of the model, and overcomes the limitations of traditional methods in nonlinear and non-stationary signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760324B_ABST
    Figure CN119760324B_ABST
Patent Text Reader

Abstract

A prediction method for dissolved gas in transformer oil based on an improved BiGRU network is used to solve the problem of low prediction accuracy in existing abnormal data identification. The present invention first uses SDROF to identify abnormal values ​​of the original data of gas content in transformer oil. Subsequently, the VMD algorithm is applied to the pre-processed gas content data, and the adjustment parameter α and the number of modes M are precisely tuned in combination with BKA to obtain a set of intrinsic mode function IMF components. At the same time, the prediction window length and transition point in the prediction model are further optimized based on BKA, the processed data set is divided into training set, test set and validation set, and a prediction model is constructed based on MAML and BiGRU, and each decomposed IMF component is predicted, and finally a complete prediction time series is obtained by reconstruction. The present invention can realize the accurate prediction of abnormal gas data in transformer oil.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of transformer anomaly detection, in particular to a method for predicting dissolved gas in transformer oil based on an improved BiGRU network. Background Art

[0002] Oil-immersed power transformers play a core role in the power system, and the reliability of their operation is directly related to the stability of the power grid and the safety of electricity use. Dissolved gas analysis (DGA), as a widely used diagnostic technology, can effectively evaluate the operating status of the transformer by detecting the concentration changes of characteristic gases in the insulating oil. When the transformer fails such as discharge or overheating, the insulating oil and oil-paper materials will decompose and release a variety of characteristic gases. As the operating time increases, these gases gradually accumulate, resulting in a decrease in insulation performance and may cause serious faults. The application of DGA technology provides a scientific basis for transformer status assessment and fault diagnosis.

[0003] In recent years, time series prediction and analysis of dissolved gases in transformer oil has gradually become a key area of ​​academic research. The concentration of dissolved gases usually shows a trend of gradual accumulation during normal operation of the equipment, but may jump sharply in the event of sudden failure or abnormal conditions, and the concentration changes of these gases are regarded as key indicators for evaluating fault characteristics. The prediction model based on time series can accurately capture the dynamic evolution of dissolved gas concentration, providing reliable data support for early identification of potential risks and preventive maintenance measures. Among artificial intelligence methods, support vector regression SVR and multi-layer perceptron MLP are commonly used prediction models, but they often ignore the intrinsic dependencies in time series data. The long short-term memory network LSTM has a significant long-term memory function and can effectively solve the problems of gradient disappearance and gradient explosion in long sequence training. However, it still has certain limitations in parallel task processing.

[0004] Affected by the synergistic effects of operating conditions, environmental influences and complex electromagnetic environments, the time series data of dissolved gas in transformer oil exhibits significant nonlinearity and instability. Traditional statistical methods have limited ability to fit nonlinear time series and are difficult to reveal the intrinsic characteristics of time series data of dissolved gas in transformer oil. The combination of sequence decomposition technology and deep learning models has become a hot topic in current research. In the existing technology, empirical mode decomposition (EMD) is used to decompose non-stationary signals into multiple relatively stable sub-signals. However, the EMD method is susceptible to modal aliasing and endpoint effects, resulting in a decrease in prediction accuracy.

[0005] Abnormal data identification is a key step in improving data quality. Its core lies in accurately locating the abnormal values ​​in the data sequence that significantly deviate from the normal mode, and correcting them with corresponding correction strategies, so as to effectively reflect the actual operating status and change trend of the transformer and improve the credibility of the data. In the fields of power equipment monitoring and load sensing, abnormal data identification has become a research hotspot. Existing studies have successfully identified various types of anomalies using local outlier factors and isolation forest methods, but these methods usually identify anomalies based on global metrics, resulting in insufficient sensitivity to local anomalies. Summary of the invention

[0006] The purpose of the present invention is to provide a method for predicting dissolved gas in transformer oil based on an improved BiGRU network, so as to solve the problems that the existing abnormal data recognition is insufficiently sensitive to local anomalies, and the EMD method is susceptible to modal aliasing and endpoint effects, resulting in a decrease in prediction accuracy.

[0007] The technical solution adopted by the present invention to solve the technical problem is: a method for predicting dissolved gas in transformer oil based on an improved BiGRU network, comprising the following steps:

[0008] S1: Collect the CO, H2, CH4, C2H4, C2H2, and CH3CH3 gases dissolved in the transformer oil and input them into the SDROF anomaly detection algorithm model to detect and identify outliers in the original data and remove them;

[0009] S2: Use multiple filling methods for weighted fusion, calculate the error performance of different filling methods in data filling, determine the weight coefficient of each filling method, and obtain the time series after filling preprocessing;

[0010] S3: Perform variational mode decomposition (VMD) on the time series after filling preprocessing, and use the Black Kite Optimization Algorithm (BKA) to optimize the adjustment parameters. and the number of decomposition modes M;

[0011] S4: Divide the training set and the test set, and use the Black Kite Optimization Algorithm (BKA) to optimize the prediction window length L and transition point P of the model;

[0012] S5: Using the optimized prediction window length L and transition point P, the training set is input into the meta-learning MAML algorithm framework, and the initial parameters of the bidirectional gated recurrent unit BiGRU network are updated by minimizing the training error to enhance the model's rapid adaptation ability and generalization performance;

[0013] S6: Based on the BiGRU network optimized by MAML, each intrinsic mode function IMF after VMD decomposition is predicted to obtain the prediction results of each component;

[0014] S7: reconstruct the prediction results of all IMF components to generate a complete prediction time series of dissolved gas concentration in oil;

[0015] S8: Evaluate the prediction results, verify the prediction accuracy and reliability of the model by comparing with the real data, and analyze the prediction results of the test set data based on the model's evaluation indicators.

[0016] Furthermore, the steps of the SDROF anomaly detection algorithm of step S1 include: S1.1 calculating the relative distance , (1); among which, Represents a data object , Represents a data object ; S1.2 Calculate relative skewness , (2) Among them, Represents the skewness of the data object; k represents the number of data in the natural neighbor; Represented as the natural neighbor set of the data object; x ij represents the jth eigenvalue of the i-th nearest neighbor of the data object; x pj Represents the data object x p The eigenvalue in the jth dimension; j represents the dimension index of the data, and the value of j ranges from 1 to m, where m is the number of dimensions of the data set; p represents the reference data object, which is used to calculate the mean or center point when calculating the skewness; S1.3 Calculate the local density ratio , (3); S1.4 Calculate the data object x i The relative skewness density ratio outlier factor SDROF i , (4) Calculate the factor difference , (5); SDROF j is the data object x j The relative skewness density ratio outlier factor; according to the number of marked outliers o in the data set, the first o data objects with the largest factor difference are selected as outliers.

[0017] Furthermore, the filling method in step S2 includes nearest neighbor interpolation method, Lagrange interpolation method, local weighted regression interpolation method, spline interpolation method, mean filling method, least squares method, forward filling method and backward filling method.

[0018] Furthermore, in step S3, the formula for VMD decomposition is: (6); (7); where M is the number of modes, is the mth mode function, is its center frequency, is the pulse function, x(t) is the original signal, j is the imaginary unit, t is the time, and the Lagrange multiplier is introduced. and penalty parameters >0, construct a Lagrangian function to eliminate the constraints: (8) Using the alternating direction multiplier method ADMM, the modal components are updated alternately and center frequency , to achieve the optimal solution of the Lagrangian function; in the n+1th iteration, the update formula is: (9); (10); among which, is the Fourier transform of the signal x(t); is the modal component u m Fourier transform of (t); is the Lagrange multiplier Fourier transform of ; n represents the number of iterations; Represents the signal frequency; represents the frequency domain representation of the kth mode in the nth iteration; by optimizing the frequency domain characteristics of the VMD separation signal, accurate modal decomposition of the complex signal is achieved.

[0019] Furthermore, the optimization steps of BKA include: 1) randomly assigning the initial position of each black-winged kite (11); where i is an integer between 1 and N, and N is the number of potential solutions, i.e., the population size of black-winged kites; is the position of the i-th black-winged kite; rand is a random number between [0,1]; represents the upper bound of the position of the i-th black-winged kite, Represents the next position of the i-th black-winged kite; during initialization, the individual with the best fitness is taken as the optimal position of the black-winged kite in the initial group , the calculation formula is: (12); (13); among them, is the best fitness value in the population; f(X i ) is represented by the population X i 2) The position of the black kite attack behavior pattern is updated as follows: (14); represents the position of the i-th black-winged kite in the t-th iteration on the j-th dimension; represents the position of the ith black-winged kite in the t+1th iteration on the jth dimension; r is a random number between 0 and 1; p is a constant with a value of 0.9; n is a nonlinear factor, (15), t is the number of iterations completed so far, and T is the total number of iterations; 3) The formula for calculating the position of the population during migration behavior is as follows: (16); (17); represents the optimal position of the black kite in the jth dimension at the tth iteration; F i is the fitness value of the i-th black-winged kite in the t-th iteration, is the fitness value of the randomly selected position in the tth iteration; C(0,1) represents the Cauchy mutation operation.

[0020] Furthermore, in step S5, the BiGRU network is composed of two unidirectional gated recurrent neural networks GRU stacked up and down.

[0021] Furthermore, in step S5, the calculation formula for GRU parameter update and state output is: (18); (19); where x t is the input at the current moment; h t-1 Indicates the state output at the previous moment; h t Indicates the status output at the current moment; is the candidate state at the current moment; r t is the output of the reset gate; z t is the output of the update gate; W r is to reset the weight of the gate; b r is the bias of the reset gate; W z is the weight of the update gate; b z is the bias of the update gate; W d is the weight of the gate control unit; b d is the bias of the gated unit; σ( ) is the sigmoid function, and tanh( ) is the hyperbolic tangent activation function.

[0022] Furthermore, in step S5, the specific steps of the MAML algorithm are as follows: in each source task T, first, in the support set The training error is calculated and the parameters are updated by the following formula: (20); where the training error Defined as: (21); where represents the initial model parameters, The specific parameters after the task is updated. is the learning rate, Expressed as parameter gradient; f θ (x) represents the output of the model on input x, that is, the predicted value; y represents the true label or target value corresponding to input x. In the query set Calculate the test error , the formula is: (twenty two); Represents the model output value after updating the parameters; the initial model parameter θ is optimized through the query set error of all tasks, and the update formula of θ is: (23); among them, is the global learning rate, is the sum of the test errors of all tasks.

[0023] Furthermore, the optimization method of the prediction window length L is as follows: Assuming that the time series is X={X1,X2,X3,…,X N}, define X t ={(x t ,z t )} is the prediction task dataset, where x t is the input of the prediction model, z t is the output of the prediction model; (twenty four); (25); wherein the prediction window length L satisfies: 1≤L≤N.

[0024] Furthermore, the optimization method of the transition point P is as follows: Assuming that the time series is X={X1,X2,X3,…,X N}, the prediction target is X at the next T time points, and the prediction task dataset is defined as ,in is the input of the prediction model, is the output of the prediction model, (26); (27); where the transition point P satisfies: L≤P≤N.

[0025] The beneficial effects of the present invention are as follows: (1) The SDROF outlier detection algorithm is used, and the weighted fusion filling strategy is combined for data preprocessing to enhance the ability to identify local anomalies. (2) The VMD parameters are optimized by BKA to achieve decomposition prediction of non-stationary nonlinear signals, thereby improving the prediction accuracy. (3) The front-to-back dependency relationship of time series data is fully explored, and MAML is used to optimize the key parameters that affect the prediction accuracy of the neural network model, effectively overcoming the low prediction accuracy of the model caused by the traditional selection of parameters based on experience, thereby improving the prediction performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a network structure diagram of the BiGRU model of the present invention;

[0027] Figure 2 It is the structure diagram of weighted coefficient preprocessing;

[0028] Figure 3 To predict the window length and transition point:

[0029] Figure 4 The present invention is a flow chart for predicting dissolved gas in transformer oil;

[0030] Figure 5 A graph with anomaly points labeled for the original data;

[0031] Figure 6 is the restoration image after weighted fusion;

[0032] Figure 7 VMD decomposition diagram of H2;

[0033] Figure 8 This is a comparison chart of the prediction results of each prediction model for C2H6 gas in a certain test period;

[0034] Fig. 9 This is a comparison chart of the prediction results of each prediction model for C2H4 gas in a certain test period;

[0035] Fig.10 This is a comparison chart of the prediction results of CH4 gas by various prediction models during a certain test period;

[0036] Fig.11 This is a comparison chart of the prediction results of H2 gas by various prediction models during a certain test period;

[0037] Fig.12 This is a comparison chart of the prediction results of each prediction model for C0 gas in a certain test period. DETAILED DESCRIPTION

[0038] The present invention first uses the SDROF outlier detection algorithm to identify outliers in the raw data of gas content in transformer oil, and performs weighted fusion preprocessing on the detected outliers to improve the reliability and accuracy of subsequent analysis. Subsequently, the variational mode decomposition (VMD) method is applied to the preprocessed gas content data, and the parameters are precisely tuned in combination with the black kite optimization algorithm (BKA). and the number of modes M, to obtain a set of intrinsic mode function IMF components. At the same time, the prediction window length and transition point in the prediction model are further optimized based on BKA, so as to systematically optimize the input and output design and training mechanism of the prediction model. After completing the above steps, the processed data set is scientifically divided into training set, test set and validation set, and a prediction model is constructed based on meta-learning MAML and bidirectional gated recurrent unit network BiGRU, and each decomposed IMF component is predicted separately, and finally a complete prediction time series is obtained by reconstruction. The algorithm of the present invention is described below:

[0039] 1. SDROF algorithm.

[0040] The relative skewness density ratio anomaly factor (SDROF) is an anomaly detection algorithm that combines global and local perspectives. It achieves an accurate assessment of the abnormality of data objects by comprehensively considering the characteristics of relative skewness and local density ratio. Relative skewness introduces the concept of relative distance to characterize the distribution relationship between data objects and their natural neighbors, and evaluates data abnormality from a global perspective. The local density ratio is used to describe the relative density difference between data objects and their neighbors, revealing abnormal characteristics from a local perspective. By combining relative skewness and local density ratio, SDROF can comprehensively characterize the abnormal characteristics of data objects. At the same time, the algorithm further proposes a factor difference indicator to quantify the local distribution difference between data objects and their neighbors, so as to more clearly identify potential anomalies.

[0041] Proximity algorithm KNN relative distance Data object The shortest distance to the data object with a higher density than its k nearest neighbor is calculated as follows: (1). Among them, Represents a data object , Represents a data object .

[0042] Relative skewness It is calculated by multiplying the relative distance of natural neighbors by the skewness to measure the degree of deviation of the data object relative to the neighborhood. The calculation formula is: (2). Among them, Represents the skewness of the data object; k represents the number of data in the natural neighbor; Represented as the natural neighbor set of the data object; x ij represents the jth eigenvalue of the i-th nearest neighbor of the data object; x pj Represents the data object x p The eigenvalue at the jth dimension; j represents the dimension index of the data, and the value of j ranges from 1 to m, where m is the number of dimensions of the data set. p represents the reference data object, which is used to calculate the mean or center point when calculating skewness.

[0043] Local density ratio It is expressed by the ratio of the density of the target object to the density of its natural neighbors, which is used to describe the density distribution difference within the neighborhood. The calculation formula is: (3).

[0044] Further combined with relative skewness and local density ratio, the SDROF index reflecting the comprehensive abnormality of the data object is constructed to obtain the data object x i The relative skewness density ratio outlier factor SDROF i , the calculation formula is: (4).

[0045] In order to further characterize the local distribution difference between the data object and its neighborhood, the algorithm introduces the factor difference Concept: (5) SDROF j is the data object x j The relative skewness density ratio outlier factor.

[0046] The factor difference is defined by calculating the deviation of the outlier factor between each data object and its neighbors As the final abnormality metric. Factor difference Based on the relative skewness density ratio outlier factor, local difference information is further extracted to more accurately characterize the outlier characteristics of the data object. The SDROF algorithm selects the first o data objects with the largest factor difference as outliers based on the number of marked outliers in the data set.

[0047] 2. VMD algorithm.

[0048] Variational mode decomposition (VMD) is a non-recursive signal processing algorithm that aims to decompose complex signals into intrinsic mode function (IMF) components with specific bandwidths. The algorithm achieves frequency domain decomposition of the signal by constructing and solving a variational problem. Specifically, VMD ensures that the sum of all modes is equal to the original signal through constraints, and minimizes the sum of the bandwidths of each mode. Its constrained variational model is expressed as: (6); (7). Where M is the number of modes, is the mth mode function, is its center frequency, is the impulse function, x(t) is the original signal, j is the imaginary unit, and t is the time. To solve the above constrained variation problem, the Lagrange multiplier is introduced and penalty parameters >0, construct the following Lagrangian function to eliminate the constraints: (8).

[0049] Using the alternating direction multiplier method ADMM, the modal components are updated alternately and center frequency , to achieve the optimal solution of the Lagrangian function. In the n+1th iteration, the update formula is: (9); (10). Among them, is the Fourier transform of the signal x(t); is the modal component u m Fourier transform of (t); is the Lagrange multiplier Fourier transform of ; n represents the number of iterations; Represents the signal frequency; represents the frequency domain representation of the kth mode in the nth iteration; for The Fourier transform of .

[0050] Through the above optimization process, VMD can effectively separate the frequency domain characteristics of the signal and achieve accurate modal decomposition of complex signals. The number of modes M can further improve the adaptability and decomposition accuracy of the algorithm to different signals.

[0051] 3. Black Kite Optimization Algorithm.

[0052] Compared with the existing bionic optimization algorithms, the black kite optimization algorithm BKA has the characteristics of fast convergence speed and more fitness function calls. This algorithm is a simulation of the behavior of black kites in attack and migration. Due to its small number of parameters and simple operation, as well as the combination of the algorithm's unique dual-mode attack behavior and dynamic leadership mechanism, it has excellent global optimization ability and adaptability. The black kite optimization algorithm is an optimization algorithm that simulates the hunting skills and migration behavior of black kites. The algorithm can be divided into three stages: population initialization stage, attack stage, and migration stage.

[0053] (1) In the population initialization stage, a random initialization strategy is adopted to randomly assign the initial position of each black kite using the following formula: (11). Where i is an integer between 1 and N, and N is the number of potential solutions, i.e., the population size of black-winged kites; is the position of the i-th black-winged kite; rand is a random number between [0,1]; represents the upper bound of the position of the i-th black-winged kite, Represents the next position of the i-th black-winged kite. When initialized, the individual with the best fitness is regarded as the optimal position of the black-winged kite in the initial group. , the calculation formula is as follows: (12); (13). Among them, is the best fitness value in the population; f(X i ) is represented by the population X i The fitness of .

[0054] (2) As a predator of small mammals and insects in grasslands, black-winged kites adjust the angle of their wings and tail according to wind speed during flight, hovering to observe prey, and then swooping down quickly to hunt. Their attack behavior mainly includes hovering to observe and striking prey. The position update is as follows: (14). represents the position of the i-th black-winged kite in the t-th iteration on the j-th dimension; represents the position of the ith black-winged kite in the t+1th iteration on the jth dimension; r is a random number between 0 and 1; p is a constant with a value of 0.9; n is a nonlinear factor, (15), t is the number of iterations completed so far, and T is the total number of iterations.

[0055] (3) Migration behavior is led by the leader. When the fitness value of the current population is less than that of the randomly selected population, the leader will give up the leadership role and join the migrating population; conversely, when the fitness value of the current population is greater than that of the random population, the leader will continue to guide the population towards the target location. The specific formula is as follows: (16); (17). represents the optimal position of the black kite in the jth dimension at the tth iteration; F i is the fitness value of the i-th black-winged kite in the t-th iteration, is the fitness value of the randomly selected position in the tth iteration; C(0,1) represents the Cauchy mutation operation.

[0056] 4. Bidirectional gated recurrent unit neural network.

[0057] GRU is a widely used gated recurrent neural network, known for its strong learning ability for long-term dependent information. The GRU model combines the advantages of RNN in time series calculation and the advantages of LSTM in data correlation processing. By merging the forget gate and input gate of LSTM into an update gate, the model structure is significantly simplified, the training parameters are reduced, and the computational efficiency is improved and the gradient explosion problem is effectively avoided. The calculation formula for GRU model parameter update and state output is shown in the following formula: (18); (19). In the formula, x t is the input at the current moment; h t-1 Indicates the state output at the previous moment; h t Indicates the status output at the current moment; is the candidate state at the current moment; r t is the output of the reset gate; z t is the output of the update gate; W r is to reset the weight of the gate; b r is the bias of the reset gate; W z is the weight of the update gate; b z is the bias of the update gate; W d is the weight of the gate control unit; b dis the bias of the gated unit; σ( ) is the sigmoid function, and tanh( ) is the hyperbolic tangent activation function.

[0058] However, the traditional GRU network can only effectively capture the forward feature information of the sequence data, and has insufficient modeling capabilities for the backward features. To this end, the present invention introduces a bidirectional GRU network BiGRU to achieve bidirectional feature extraction of the time series of gas content in transformer oil, thereby improving the prediction accuracy. Figure 1 As shown, BiGRU is composed of two unidirectional GRUs stacked up and down, and the output is determined by the states of these two GRUs. The model structure is as follows Figure 1 As shown. Among them, X i Indicates the dissolved gas content in oil and external conditions input data, e i Represents the input data X i The vector of represents the positive hidden layer state, Represents the reverse hidden layer state. BiGRU overcomes the defect that the unidirectional GRU can only process unidirectional state information and cannot combine historical and future bidirectional information. By performing forward and reverse calculations on the time series at the same time, it enhances the extraction effect of the original sequence information and improves the accuracy of the model output results.

[0059] 5. Meta-learning network.

[0060] As an important frontier direction in the field of machine learning, meta-learning aims to solve the challenge of models adapting to new tasks quickly and efficiently under limited data and computing resources. In the present invention, the meta-learning MAML algorithm is introduced as the core meta-learning layer to significantly enhance the generalization performance of the model. The MAML algorithm is a general and efficient meta-learning framework that optimizes the initial parameters of the model in a multi-task environment so that it has the ability to quickly adapt to the distribution characteristics of new tasks. This method mines shared features between tasks and learns parameter initialization with good migration capabilities, thereby achieving rapid convergence of the model on new tasks. Even when data is scarce and the number of update iterations is limited, it can still show excellent learning performance, providing theoretical and practical support for multi-task learning in complex scenarios.

[0061] When using BiGRU to predict the dissolved gas content in transformer oil, the data is divided into a training set and a test set. In the MAML-BiGRU framework, the support set Corresponding to the training set, query set Corresponding to the test set. The specific steps of the MAML algorithm are as follows.

[0062] In each source task T, the MAML model first runs The training error is calculated and the parameters are updated by the following formula: (20). Among them, the training error Defined as: (21). Where, represents the initial model parameters; Updated specific parameters for the task; is the learning rate; Expressed as parameter gradient; f θ (x) represents the output of the model on input x, that is, the predicted value; y represents the true label or target value corresponding to input x. On this basis, using the updated parameters In the query set Calculate the test error , the formula is: (twenty two). Represents the model output value after updating the parameters; optimize the initial model parameters through the query set error of all tasks , The update formula is: (23). Among them, is the global learning rate, is the sum of the test errors of all tasks. This optimization process can ensure that the initial parameters of the model have strong generalization ability and can quickly adapt to new tasks.

[0063] In summary, the MAML algorithm uses the support set of the target task to The initial parameters are optimized on the query set, which significantly enhances the model's adaptability to the target data distribution and The effective evaluation of model performance is achieved in this way. Its core advantage is that it can quickly adapt to new task distributions through a small number of updates, especially in scenarios where data distribution changes little, which can significantly improve the prediction accuracy of the model. For the BiGRU model, the model optimized by introducing the MAML algorithm not only improves the prediction accuracy of gas concentration changes, but also enhances the sensitivity to extreme values ​​and mutation phenomena, showing stronger robustness. Especially in complex multivariate scenarios, this method effectively mines the correlation characteristics and potential patterns between gases, greatly improving the prediction effect of traditional models in cases of scarce data and heterogeneous distribution, and providing a more forward-looking solution for gas analysis in oil.

[0064] Based on the above algorithm, the present invention constructs an outlier detection and BKA-VMD-MAML-BiGRU prediction model based on the SDROF algorithm. The prediction model is described below:

[0065] 1. Data preprocessing based on weighted fusion strategy.

[0066] In view of the large differences in the trends and magnitudes of dissolved gases in transformer oil, the traditional single missing data filling method is difficult to apply to complex multivariate historical sequences. In order to improve data integrity and filling accuracy, this paper proposes a multi-method data preprocessing strategy based on weighted fusion.

[0067] like Figure 2 As shown, the SDROF outlier detection algorithm is used to screen out the outliers in the original data, and after the outliers are removed, a variety of filling methods are used for filling. The present invention selects eight common and effective filling methods, including nearest neighbor interpolation, Lagrange interpolation, local weighted regression interpolation, spline interpolation, mean filling method, least squares method, forward filling method and backward filling method, and systematically evaluates the filling performance of each method. By calculating the error performance of different filling methods in data filling, a weight-based comprehensive scoring mechanism is constructed to determine the weight coefficient of each filling method. Finally, combining the advantages of all filling methods, a weighted fusion high-precision data filling method is proposed. This method can adaptively select the best filling scheme in a multivariate time series, significantly improving the robustness and accuracy of subsequent data analysis and modeling. The architecture of weighted coefficient preprocessing is shown in the figure. Figure 2 shown.

[0068] 2. Prediction window length L and transition point P optimized based on BKA algorithm.

[0069] Determining the optimal prediction window length L is a key issue in medium- and long-term prediction tasks. Under the same prediction target and prediction model, choosing different prediction window times will lead to significant differences in prediction accuracy. Figure 3 As shown, assuming that the time series is X={X1,X2,X3,…,X N}, define X t ={(x t ,z t )} is the prediction task dataset, where x t is the input of the prediction model, z t is the output of the prediction model and is expressed as follows: (twenty four); (25). The prediction window length L is a parameter that needs to be determined in the prediction task, satisfying the condition 1≤L≤N. If the L value is too small, the historical information content is insufficient, resulting in the inability of the sequence information extracted by the prediction model to support the prediction of the next data point, thereby affecting the accuracy of the prediction result. If the L value is too large, the time series with a length of L may contain multiple subsequences with different trends, resulting in information redundancy, thereby affecting the accuracy of the prediction. In addition, an excessively large L may also increase the model's demand for historical data and cause training problems such as gradient explosion.

[0070] like Figure 4As shown in the figure, when performing medium- and long-term forecasting tasks, it is usually necessary to use the previous forecast values ​​to predict the subsequent results. In order to improve the accuracy of the subsequent forecast values, it is necessary to select a suitable transition point P. This is because during the model training process, some early forecast values ​​will be used as input together with the true value to help the model learn how to predict the target value based on the previous forecast value, and continuously adjust and optimize the performance of the model through training. Through the positive and negative feedback mechanism, the model can be gradually corrected to achieve the best forecasting effect.

[0071] Assume that the time series is X={X1,X2,X3,…,X N}, the prediction target is X at the next T time points, and the prediction task dataset is defined as ,in is the input of the prediction model, is the output of the prediction model, as shown below: (26); (27). The transition point P is a key parameter to be determined in the prediction task, satisfying the condition L≤P≤N. In the early stage of training, due to the limited training data, it is difficult for the prediction model to fully learn the sequence information, resulting in large errors in the early prediction values. When the value of P is small, the prediction data is used for model training too early, which may cause error accumulation, making it difficult for the model to converge or even training failure. In medium- and long-term predictions, the prediction value of the previous moment is usually used as input for subsequent predictions. When the value of P is too large, the prediction information of the previous training cannot be used in the subsequent model training in time, which may lead to insufficient information in subsequent tasks or failure of the model to converge effectively, and ultimately introduce large prediction errors.

[0072] The Black Kite Optimization Algorithm (BKA) is used to optimize the prediction window length L and transition point P. During the optimization process, the fitness function is defined as the error performance of the MAML-BiGRU prediction model in the training phase to measure the pros and cons of different L and P parameter combinations. BKA simulates predation and migration behaviors to achieve an effective combination of local search and global exploration, so that the optimization process can take into account the comprehensiveness and sophistication of the solution space. After the algorithm converges, the optimal prediction window and transition point obtained can not only effectively balance historical information and future needs, but also significantly improve the prediction accuracy and generalization ability of the model, providing reliable parameter support for complex time series prediction tasks.

[0073] 3. Prediction of dissolved gas concentration in oil based on BKA-VMD-MAML-BiGRU model.

[0074] The concentration of dissolved gas in transformer oil plays a key role in evaluating the operating status of the transformer. When electrical or thermal failure or insulation cracking occurs, a small amount of CO, H2, CH4, C2H4, C2H2, and CH3CH3 gas will be produced in the oil, and when the fault is serious, the gas concentration will increase significantly. The content of dissolved gas is closely related to the operating status of the transformer. By predicting the gas concentration, potential faults can be identified in advance, which facilitates timely maintenance and ensures the stability and safety of the power system. The historical data of dissolved gas in transformer oil usually shows significant nonlinear and non-stationary characteristics, which makes the modeling and prediction of the original data face great challenges. In order to more effectively extract the essential characteristics of the data and improve the prediction accuracy, the present invention proposes to use VMD to decompose the preprocessed data. The VMD decomposition effect is affected by the adjustment parameters and the number of modes M. Improper parameter settings may lead to over-decomposition or under-decomposition, which in turn affects the frequency distribution and stability of the decomposition results, thereby affecting the accuracy of the subsequent prediction model. To this end, the present invention introduces the parameter optimization mechanism of BKA, explores the The optimal combination of and M is used to ensure the rationality and effectiveness of the VMD decomposition results.

[0075] In the BiGRU network, the key hyperparameters to be optimized include the number of neurons in the hidden layer, the learning rate, the Dropout rate, and the Batchsize, where the learning rate includes the learning rate of BiGRU and the learning rate of Meta-Learner. These hyperparameters play a vital role in the model's expressiveness, training dynamics, and generalization performance. Specifically, the number of neurons in the hidden layer determines the model's feature extraction ability and expression capacity, affecting the network's ability to capture deep temporal dependencies in sequence data; the learning rate controls the step size of each gradient update, directly affecting the convergence speed and stability of the model; the Dropout rate, as a regularization method, effectively prevents the model from overfitting on the training data by randomly discarding some neurons during the training process, and improves its generalization ability; the Batchsize plays a balancing role between sample parallel processing and optimization speed during the training process. Based on this, the MAML algorithm optimizes the initial parameters and hyperparameters of the BiGRU network by minimizing the training error on the support set using a meta-learning strategy. This process not only accelerates the convergence of the model on new tasks, but also significantly improves its performance in complex time series prediction tasks. In the optimized model, the tuning of hyperparameters enables the BiGRU network to extract time series features more efficiently, thereby showing significant improvement in accuracy and robustness in the prediction of dissolved gas concentration in oil.

[0076] The implementation process of the method for predicting dissolved gas in transformer oil proposed by the present invention is as follows: Figure 4As shown, it mainly includes feature sequence data input, SDROF data preprocessing, BKA optimization, MAML optimization, and BiGRU prediction output. The specific steps are as follows:

[0077] S1: The CO, H2, CH4, C2H4, C2H2, and CH3CH3 gases dissolved in the transformer oil collected by the monitoring system are input into the SDROF anomaly detection algorithm model to detect and identify outliers in the original data and remove them;

[0078] S2: Use multiple filling methods for weighted fusion, calculate the error performance of different filling methods in data filling, determine the weight coefficient of each filling method, and obtain the time series after filling preprocessing;

[0079] S3: Perform variational mode decomposition (VMD) on the time series after filling preprocessing, and use the black kite optimization algorithm to optimize the adjustment parameters. and the number of decomposition modes M;

[0080] S4: Divide the training set and the test set, and use the black kite optimization algorithm to optimize the prediction window length L and transition point P of the model;

[0081] S5: Using the optimized prediction window length L and transition point P, the training set is input into the meta-learning MAML algorithm framework, and the initial parameters of the bidirectional gated recurrent unit BiGRU network are updated by minimizing the training error to enhance the model's rapid adaptation ability and generalization performance;

[0082] S6: Based on the BiGRU network optimized by MAML, each intrinsic mode function IMF after VMD decomposition is predicted to obtain the prediction results of each component;

[0083] S7: reconstruct the prediction results of all IMF components to generate a complete prediction time series of dissolved gas concentration in oil;

[0084] S8: Evaluate the prediction results, verify the prediction accuracy and reliability of the model by comparing with the real data, and analyze the prediction results of the test set data based on the model's evaluation indicators.

[0085] In order to verify the performance of the prediction model proposed in the present invention, the mean absolute percentage error MAPE, root mean square error RMSE, prediction accuracy FA and determination coefficient R are selected. 2 As evaluation indicators. Among these evaluation criteria, the smaller the values ​​of mean absolute percentage error and root mean square error, the higher the prediction accuracy, and the closer the coefficient of determination is to 1 (0≤R 2 ≤10), it means that the prediction effect of the model is better and closer to the change trend of the actual data. Among them, the mean absolute percentage error (28); Root mean square error (29); Prediction accuracy (30); Determination coefficient (31). In the above formula, n represents the number of samples in the test set; represents the true value of the gas concentration at the i-th moment, represents the predicted value of gas concentration at the i-th moment, is the mean of the true values.

[0086] Based on the analysis of the above theoretical part, the present invention adopts the online monitoring data of dissolved gas in the power transformer oil of a 500kV substation, with a collection period of 8h and a gas volume fraction unit of 10 -6 . The first 530 sets of data collected continuously are used as training samples, and the last 100 sets of data collected continuously are used as test samples. The present invention verifies the effectiveness of the proposed method from the following aspects: first, the reliability of the data preprocessing method is evaluated by performing outlier detection and preprocessing based on weighted coefficients on the original data; secondly, the prediction results based on the BKA-VMD-MAML-BiGRU model are compared with several prediction methods, and the optimal prediction window and transition point of each gas are analyzed in combination with the preprocessed data to comprehensively verify the effectiveness and superiority of the combined prediction method.

[0087] 1. Data preprocessing.

[0088] Based on the SDROF algorithm, the present invention performs outlier identification on the first 530 groups of data to obtain the original data distribution map containing outlier annotations, such as Figure 5 As shown. For the detected abnormal points, the high-precision abnormal data filling method proposed in the present invention is used to repair them. Table 1 lists in detail the filling errors and corresponding weight coefficient results of each filling method in different gas data processing. Figure 6 It can be seen that the weighted fusion filling method proposed in the present invention has a filling error of 0.84% ​​for CO, a filling error of 0.33% for H2, a filling error of 0.68% for CH4, a filling error of 0.26% for C2H4, and a filling error of 0.47% for CH3CH3. Among them, since the C2H2 gas content is stable at 0μL / L, its filling accuracy is not included in the analysis. It can be seen from the error results that the filling error of this method shows good stability and does not fluctuate significantly due to differences in data point characteristics. Although the errors of the nearest neighbor interpolation method, the forward filling method, and the backward filling method may be slightly lower than the weighted fusion method in certain gas contents, this is mainly limited by the characteristics of the original signal. However, this type of method shows low applicability and versatility in data samples with significant differences in characteristics, and the filling error fluctuates greatly between different abnormal points.

[0089] Table 1 Weighted filling results

[0090]

[0091] 2. Prediction results.

[0092] The variational mode decomposition (VMD) algorithm is used to decompose the contents of dissolved gases CO, H2, CH4, C2H4 and CH3CH3 in transformer oil. Taking H2 gas as the research object, the adjustment parameters in the VMD decomposition model are optimized with the help of the BKA optimization algorithm. The optimal parameters are α=1843 and M=7. Figure 7 As shown in the figure, the fluctuation amplitude of the decomposed IMF component is significantly reduced and the stability is significantly enhanced compared with the undecomposed gas concentration time series, which more clearly reflects the time series characteristics from high frequency to low frequency. This decomposition method helps to deeply explore the changing laws of dissolved gas data in oil and provides a more stable foundation for subsequent feature extraction and predictive modeling.

[0093] In order to optimize the prediction window length L and transition point P, the present invention adopts the black kite optimization algorithm BKA, and uses the mean absolute percentage error MAPE predicted by the MAML-BiGRU model as the fitness function to systematically optimize the optimal values ​​of L and P. During the optimization process, the population size of BKA is set to 50, and the maximum number of iterations is set to 80 to fully balance the global search capability and local development efficiency. The optimal L and P values ​​of each characteristic gas are shown in Table 2.

[0094] Table 2 Optimal prediction window length L and transition point P optimization

[0095]

[0096] In the optimization process of the MAML algorithm, the meta-learning strategy is used to minimize the training error on multiple tasks. When optimizing the hyperparameters of the BiGRU network, MAML performs backpropagation based on the support set and updates the initial parameters and hyperparameters in each iteration. Through iterative optimization, the learning rate, number of hidden layer neurons, and Dropout rate are gradually adjusted to adapt to the specific characteristics of different tasks. Finally, the MAML algorithm obtains the hyperparameter combination with the best generalization ability through optimization on the support set, which includes 128 hidden layer neurons, 0.008 BiGRU learning rate, 0.0009 Meta-Learner learning rate, and 0.3 Dropout rate.

[0097] Based on the collected data of the last 100 test samples, the present invention compares the prediction performance of the proposed BKA-VMD-MAML-BiGRU model with other comparison models, such as VMD-BKA-BiGRU, MAML-BiGRU and BiGRU. The relevant results are as follows: Figures 8 to 12 As shown in Table 3. Figures 8 to 12 It can be clearly seen from the prediction results that the BKA-VMD-MAML-BiGRU model is significantly better than other comparison models in terms of prediction effect. Specifically, compared with each comparison model, the prediction curve of the model proposed in the present invention has a significantly higher degree of fit with the actual curve, which shows that the model of the present invention performs better in feature learning and regularity capture of dissolved gas data in transformer oil. The model shows a stronger sensitivity to the dynamic changes of gas data, can more accurately reflect the peaks and fluctuations of the data, and can still maintain a high prediction accuracy when facing complex changing situations, which fully reflects the advantages of the model of the present invention in prediction performance.

[0098] The comparison of the quantitative evaluation indicators in Table 3 shows that, especially Fig.11 In the H2 sample shown in the figure, the four indicators of the BKA-VMD-MAML-BiGRU model are 26.2%, 0.5%, 98.6%, and 97.6%, respectively. Compared with the BiGRU, MAML-BiGRU, and VMD-BKA-BiGRU models, the y MAPE The decreases were 73.5%, 63.7% and 61.7%, respectively. RMSE They decreased by 2.3%, 1.1% and 1.1% respectively. FA The increases were 11.9%, 11.0% and 10.8%, respectively. R2 They increased by 54.4%, 40.4% and 41.3% respectively.

[0099] According to the analysis of experimental data, the generation process of CH4 and CH3CH3 gases is significantly affected by external variables, such as load, temperature and load, and shows strong nonlinear and non-stationary characteristics. In contrast, the generation trend of H2 and CO gases is relatively stable and less affected by changes in the external environment. For CH4 and CH3CH3 gases, the BKA-VMD-MAML-BiGRU model has a better performance than the MAML-BiGRU model. MAPE decreased by 32.9% and 36.4% respectively. RMSE They decreased by 1.7% and 10% respectively. FA They increased by 4.7% and 5.3% respectively. R2The performance of the MAML optimization module was verified by Fig.12 In the CO sample shown, the y of the BKA-VMD-MAML-BiGRU model is lower than that of the VMD-BKA-BiGRU model. MAPE Reduced by 32.4%, RMSE Down 2%, FA Increase by 4%, R2 The above results fully demonstrate that the proposed MAML optimization module has significant advantages in improving feature learning and model generalization capabilities, and effectively enhances the model's adaptability and prediction accuracy to complex gas generation laws.

[0100] Table 3 Characteristic gas prediction

[0101]

[0102] The present invention aims to improve the prediction accuracy of dissolved gas concentration in transformer oil. Aiming at the limitations of traditional prediction methods in dealing with dynamic changes of dissolved gas in oil, a new method based on a joint prediction model is proposed. The model combines the SDROF outlier detection algorithm, the weighted fusion filling method, the variational mode decomposition VMD, the black kite optimization algorithm BKA, the meta-learning method MAML and the bidirectional gated recurrent unit network BiGRU, which can effectively capture and analyze the time series characteristics of dissolved gas concentration in oil. (1) By introducing the SDROF outlier detection algorithm and the weighted fusion strategy, the problem of the gas concentration in oil being affected by outliers due to the interference of existing dissolved gas analysis equipment and multiple factors, which has an adverse effect on the prediction accuracy, is effectively solved. (2) The BKA optimization algorithm is applied to adaptively optimize the prediction window length L and the position of the transition point P, which effectively overcomes the problem of insufficient model fitting ability and low prediction accuracy caused by the traditional reliance on experience to select the prediction window length L and the transition point P. At the same time, by optimizing the VMD algorithm with BKA and accurately adjusting the parameter α and the number of modes M, the algorithm's adaptability to nonlinear and non-stationary signals and decomposition accuracy are significantly improved, laying the foundation for further optimization of the model performance. (3) A prediction model for dissolved gas in transformer oil that integrates the MAML meta-learning framework and the bidirectional gated recurrent unit BiGRU is proposed. By making full use of the rapid adaptation characteristics of the MAML meta-learning algorithm and the advantages of BiGRU in deep feature modeling of time series data, the model parameters are efficiently optimized, significantly improving the prediction accuracy and generalization ability. The practicality and effectiveness of the proposed model are verified by analyzing transformer fault cases. The experimental results show that the optimized dissolved gas prediction model is significantly superior to traditional deep learning methods in terms of accuracy and stability, further proving the effectiveness and potential application value of the model in predicting complex time series data.

Claims

1. A method for predicting dissolved gas in transformer oil based on an improved BiGRU network, characterized in that: The following steps are involved: S1: Collect the CO, H2, CH4, C2H4, C2H2, and CH3CH3 gases dissolved in the transformer oil and input them into the SDROF anomaly detection algorithm model to detect and identify outliers in the original data and remove them; S2: Use multiple filling methods for weighted fusion, calculate the error performance of different filling methods in data filling, determine the weight coefficient of each filling method, and obtain the time series after filling preprocessing; S3: Perform variational mode decomposition (VMD) on the time series after filling preprocessing, and use the Black Kite Optimization Algorithm (BKA) to optimize the adjustment parameters. and the number of decomposition modes M; S4: Divide the training set and the test set, and use the Black Kite Optimization Algorithm (BKA) to optimize the prediction window length L and transition point P of the model; S5: Using the optimized prediction window length L and transition point P, the training set is input into the meta-learning MAML algorithm framework, and the initial parameters of the bidirectional gated recurrent unit BiGRU network are updated by minimizing the training error; S6: Based on the BiGRU network optimized by MAML, each intrinsic mode function IMF after VMD decomposition is predicted to obtain the prediction results of each component; S7: reconstruct the prediction results of all IMF components to generate a complete prediction time series of dissolved gas concentration in oil; S8: Evaluate the prediction results, verify the prediction accuracy and reliability of the model by comparing with the real data, and analyze the prediction results of the test set data based on the model's evaluation indicators.

2. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 1, characterized in that: The steps of the SDROF anomaly detection algorithm in step S1 include: S1.1 Calculate the relative distance , (1); among which, Represents a data object , Represents a data object ; S1.2 Calculate relative skewness , (2); among which, Represents the skewness of the data object; k represents the number of data in the natural neighbor; Represented as the natural neighbor set of the data object; x ij represents the jth eigenvalue of the i-th nearest neighbor of the data object; x pj Represents the data object x p The eigenvalue in the jth dimension; j represents the dimension index of the data, and the value of j ranges from 1 to m, where m is the number of dimensions of the data set; p represents the reference data object, which is used to calculate the mean or center point when calculating the skewness; S1.3 Calculate the local density ratio , (3); S1.4 Calculate the data object x i The relative skewness density ratio outlier factor SDROF i , (4) Calculate the factor difference , (5); SDROF j is the data object x j The relative skewness density ratio outlier factor; according to the number of marked outliers o in the data set, the first o data objects with the largest factor difference are selected as outliers.

3. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 2, characterized in that: The filling methods in step S2 include nearest neighbor interpolation, Lagrange interpolation, local weighted regression interpolation, spline interpolation, mean filling, least squares method, forward filling and backward filling.

4. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 3, characterized in that: In step S3, the formula for VMD decomposition is: (6); (7); where M is the number of modes, is the mth mode function, is its center frequency, is the pulse function, x(t) is the original signal, j is the imaginary unit, t is the time, and the Lagrange multiplier is introduced. and penalty parameters >0, construct a Lagrangian function to eliminate the constraints: (8) Using the alternating direction multiplier method ADMM, the modal components are updated alternately and center frequency , to achieve the optimal solution of the Lagrangian function; in the n+1th iteration, the update formula is: (9); (10); among which, is the Fourier transform of the signal x(t); is the modal component u m Fourier transform of (t); is the Lagrange multiplier Fourier transform of ; n represents the number of iterations; Represents the signal frequency; represents the frequency domain representation of the kth mode in the nth iteration; by optimizing the frequency domain characteristics of the VMD separation signal, accurate modal decomposition of the complex signal is achieved.

5. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 4, characterized in that: The optimization steps of BKA include: 1) Randomly assigning the initial position of each black-winged kite (11); where i is an integer between 1 and N, and N is the number of potential solutions, i.e., the population size of black-winged kites; is the position of the i-th black-winged kite; rand is a random number between [0,1]; represents the upper bound of the position of the i-th black-winged kite, Represents the next position of the i-th black-winged kite; during initialization, the individual with the best fitness is taken as the optimal position of the black-winged kite in the initial group , the calculation formula is: (12); (13); among them, is the best fitness value in the population; f(X i ) is represented by the population X i 2) The position of the black kite attack behavior pattern is updated as follows: (14); represents the position of the i-th black-winged kite in the t-th iteration on the j-th dimension; represents the position of the ith black-winged kite in the t+1th iteration on the jth dimension; r is a random number between 0 and 1; p is a constant with a value of 0.9; n is a nonlinear factor, (15), t is the number of iterations completed so far, and T is the total number of iterations; 3) The formula for calculating the position of the population during migration behavior is as follows: (16); (17); represents the optimal position of the black kite in the jth dimension at the tth iteration; F i is the fitness value of the i-th black-winged kite in the t-th iteration, is the fitness value of the randomly selected position in the tth iteration; C(0,1) represents the Cauchy mutation operation.

6. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 5, characterized in that: In step S5, the BiGRU network is composed of two unidirectional gated recurrent neural networks GRU stacked up and down.

7. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 6, characterized in that: In step S5, the calculation formula for GRU parameter update and state output is: (18); (19); where x t is the input at the current moment; h t-1 Indicates the state output at the previous moment; h t Indicates the status output at the current moment; is the candidate state at the current moment; r t is the output of the reset gate; z t is the output of the update gate; W r is to reset the weight of the gate; b r is the bias of the reset gate; W z is the weight of the update gate; b z is the bias of the update gate; W d is the weight of the gate control unit; b d is the bias of the gated unit; σ( ) is the sigmoid function, and tanh( ) is the hyperbolic tangent activation function.

8. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 7, characterized in that: In step S5, the specific steps of the MAML algorithm are as follows: in each source task T, first, in the support set The training error is calculated and the parameters are updated by the following formula: (20); where the training error Defined as: (21); where represents the initial model parameters, The specific parameters after the task is updated. is the learning rate, Expressed as parameter gradient; f θ (x) represents the output of the model on input x, that is, the predicted value; y represents the true label or target value corresponding to input x; secondly, using the updated parameters In the query set Calculate the test error , the formula is: (twenty two); Represents the model output value after updating the parameters; the initial model parameter θ is optimized through the query set error of all tasks, and the update formula of θ is: (23); among them, is the global learning rate, is the sum of the test errors of all tasks.

9. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 8, characterized in that: The optimization method of the prediction window length L is as follows: Assume that the time series is X={X1,X2,X3,…,X N }, define X t ={(x t ,z t )} is the prediction task dataset, where x t is the input of the prediction model, z t is the output of the prediction model; (twenty four); (25); wherein the prediction window length L satisfies: 1≤L≤N.

10. The method for predicting dissolved gas in transformer oil based on improved BiGRU network according to claim 9, characterized in that: The optimization method of transition point P is: Assume that the time series is X={X1,X2,X3,…,X N }, the prediction target is X at the next T time points, and the prediction task dataset is defined as ,in is the input of the prediction model, is the output of the prediction model, (26); (27); where the transition point P satisfies: L≤P≤N.

Citation Information

Patent Citations

  • Improved variational mode decomposition and BiGRU fused water quality prediction method

    CN116451553A

  • Power transmission line icing thickness prediction method based on VMD-SSA-LSTM

    CN117592592A