Current transformer error prediction method based on improved adaptive composite mode decomposition

By improving the adaptive composite modal decomposition and deep learning methods, the current transformer error prediction model is built, which solves the problems of time-consuming and online calibration risks of traditional calibration methods, and achieves efficient and accurate error prediction.

CN120336779APending Publication Date: 2025-07-18CHINA THREE GORGES UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510335421.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when the measurement accuracy of the current transformer is reduced, traditional calibration methods take a long time and there is a risk of high-voltage environmental protection in the online calibration, and it is difficult to obtain model parameters, resulting in inaccurate error evaluation.

Method used

The improved adaptive composite modal decomposition method is adopted, including ICEEMDAN, multi-scale arrangement entropy and probability density function, variational modal decomposition, and CNN-BiGRU-MHA model, and the current transformer error prediction model is constructed and the prediction is carried out through target optimization matching parameters.

Benefits of technology

It improves the accuracy of current transformer error prediction and model generalization ability, reduces modal aliasing and noise interference, and enhances the accuracy of signal processing and model adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336779A_ABST
    Figure CN120336779A_ABST
Patent Text Reader

Abstract

The invention discloses a current transformer error prediction method based on improved adaptive composite mode decomposition. The method comprises the following steps: preprocessing current transformer measurement data based on improved fully adaptive noise ensemble empirical mode decomposition (ICEEMDAN); performing feature selection and reconstruction on the complex components decomposed by the ICEEMDAN based on a multi-scale permutation entropy and a probability density function; further decomposing the reconstructed complex component based on variational mode decomposition (VMD); and constructing a depth prediction model based on CNN-BiGRU, optimizing sub-model network parameters by using a target optimization method, and finally superposing sub-results to obtain a final prediction result of the current transformer. According to the method, ICEEMDAN is comprehensively utilized for signal decomposition, MPE and PDF are utilized for feature selection and reconstruction, VMD is utilized for further decomposition of complex components, and CNN-BiGRU-MHA deep learning architecture is utilized for feature learning and prediction, so that nonlinear and non-stationary signals can be effectively processed, deep-level features can be extracted, a complex dynamic relationship in time sequence data can be captured, and the time sequence data can be obtained. And the prediction accuracy and the generalization ability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of current transformer error prediction, and specifically relates to a current transformer error prediction method based on improved adaptive composite mode decomposition. Background Art

[0002] Current transformers (ECTs), as digital metering devices in smart substations, undertake electrical measurement and protection tasks. Their measurement accuracy must reach the 0.2-level measurement standard. However, when working in a complex industrial environment for a long time, various factors such as changes in temperature and humidity, electromagnetic interference, vibration, and equipment aging will cause their performance to decline and measurement accuracy to decrease. To ensure their measurement standards, regular verification is required, usually once every 1 to 4 years. Traditional verification methods rely on offline comparison with standard transformers, which is time-consuming.

[0003] To address these challenges, researchers have proposed an online calibration method for real-time monitoring and correction of the performance of current transformers (ECTs). However, due to the high-voltage environment, this method has many limitations, resulting in a "vacuum" state in monitoring between verification cycles and bringing potential risks of faulty operations. To solve this problem, scholars have proposed methods that do not require a physical standard transformer and infer measurement errors by establishing and solving an equivalent circuit model. However, accurately obtaining model parameters that match the device operating conditions remains a major technical challenge. Some researchers have transformed the evaluation of current transformer measurement errors into anomaly detection through principal component analysis (PCA), separating the inherent error changes from the primary disturbances of the power grid to achieve online status evaluation. However, this method has a large error and requires the sensor signals to follow a Gaussian distribution. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a current transformer error prediction method based on improved adaptive composite mode decomposition. First, the method uses Improved Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (ICEEMDAN) to preprocess the measurement data of the electronic current transformer (ECT) to achieve multi-band amplitude mode decomposition. Secondly, multi-scale permutation entropy (MPE) and probability density function (PDF) are used to select and reconstruct components, and the variational mode decomposition (VMD) is further used to decompose the complex part to minimize the MPE. Then, an ECT ratio error prediction model composed of a convolutional neural network (CNN), a bidirectional gated recurrent unit network (BiGRU), and a multi-head attention mechanism (MHA) is constructed. Finally, the parameters of the target optimization matching sub-network model are obtained, and the final prediction result is obtained. It can improve the current transformer error prediction.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A current transformer error prediction method based on improved adaptive composite mode decomposition includes the following steps:

[0007] Step 1: Preprocess the current transformer measurement data based on ICEEMDAN;

[0008] Step 2: Feature selection and reconstruction of the complex components decomposed by ICEEMDAN based on multi-scale permutation entropy and probability density function;

[0009] Step 3: Further decompose the reconstructed complex components based on variational mode decomposition VMD;

[0010] Step 4: Construct a deep prediction model based on CNN-BiGRU, and use the target optimization method to optimize the sub-model network parameters. Finally, the sub-results are superimposed to obtain the final prediction result of the current transformer.

[0011] The said Step 1 includes the following steps:

[0012] Step 1.1: Add a Gaussian white noise signal that follows a normal distribution to the measured data of the original current transformer to obtain a new signal sequence

[0013]

[0014] In Equation (1): represents the original signal, i represents the number of times of adding white noise, is the first-order modal component after EMD decomposition, α0 is the noise intensity adjustment coefficient; w i is the i-th Gaussian white noise with zero mean and zero variance.

[0015] Step 1.2: Perform a local mean operation on the new signal sequence to calculate the first set of residuals:

[0016]

[0017] In Equation (2): R1 is the first residual component; M(·) represents the local mean calculated by applying the EMD algorithm to the sequence

[0018] Step 1.3: Subtract the first set of residuals from the original signal to obtain the first set of intrinsic mode functions IMF1:

[0019]

[0020] Step 1.4: Iteratively calculate the residuals and modal components. Repeat the above process, add white noise to each round of residual signals, calculate the new residuals and modal components until all decompositions are completed;

[0021]

[0022] In Equation (4): R k is the k-th set of residuals; R k-1 is the (k - 1)-th set of residuals, α k-1 is the noise adjustment coefficient when adding noise for the (k - 1)-th time, is the k-th order modal component after EMD decomposition, is the k-th modal component; the noise intensity adjustment coefficient α k controls the proportion of adding white noise. It is determined according to the characteristics of the signal and can be expressed as:

[0023]

[0024] In Equation (5): ε0 is the reciprocal of the signal-to-noise ratio between the initially added noise and the signal to be analyzed; std(λ) represents the standard deviation of the threshold, E1(w i) represents the first-order modal component after EEMD decomposition, std(E1(w i )) represents the standard deviation of the first-order modal component after EEMD decomposition, γ represents the residual signal in the current decomposition step, represents the local mean calculated by applying the EMD algorithm to the original signal sequence before decomposition starts, k represents the current decomposition times, and K represents the maximum decomposition times.

[0025] Step 2 includes the following steps:

[0026] Step 2.1: Coarsely process the original signal x(t) to obtain multiple coarse-grained sequences, and calculate the average value y(i) of each coarse-grained sequence:

[0027]

[0028] In formula (6): y(i) is the average value of the i-th coarse-grained sequence; s is the scaling factor; N is the length of the original signal; i is the index of the coarse-grained sequence, j is the sampling point index of the original sequence, and x(j) represents the j-th sampling point of the original signal.

[0029] Step 2.2: Reconstruct the phase space of each coarse-grained sequence using the mean value of each coarse-grained sequence calculated in Step 2.1 to obtain the matrix {y(i), i = 1, 2, ···, N / s}:

[0030]

[0031] In Equation (7): m is the embedding dimension; τ is the time delay, k = n - (m - 1)τ; h is the maximum index in the coarse-grained sequence that can form a complete phase space vector; n represents the length of the coarse-grained matrix, which is N / s. Each row of the matrix represents a reconstructed component. y(1) represents the first element in the coarse-grained sequence, y(2) represents the second element in the coarse-grained sequence, y(j) represents the j-th element in the coarse-grained sequence, and y(h) represents the h-th element in the coarse-grained sequence. y(1 + τ) represents the value of the (1 + τ)-th element in the coarse-grained sequence, that is, the element value after one time step interval starting from the first element, y(2 + τ) represents the element value after one time step interval starting from the second element, y(j + τ) represents the element value after one time step interval starting from the j-th element, and y(h + τ) represents the element value after one time step interval starting from the h-th element. y[1 + (m - 1)τ] represents the element value after (m - 1) time step intervals starting from the first element, y[2 + (m - 1)τ] represents the element value after (m - 1) time step intervals starting from the second element, y[j + (m - 1)τ] represents the element value after (m - 1) time step intervals starting from the j-th element, and y[h + (m - 1)τ] represents the element value after (m - 1) time step intervals starting from the h-th element.

[0032] Step 2.3: Arrange the elements of each reconstructed component in ascending order:

[0033] {y(j + (k1 - 1)τ) y(j + (k2 - 1)τ) … y(j + (k m -1)τ)}(8);

[0034] In Equation (8): k1, k2, …, k m are the indices of the ordered elements in the original reconstructed component. If there are equal elements in the reconstructed component, they are sorted according to their indices. Therefore, each reconstructed component corresponds to a set of symbol sequences.

[0035] S(g) = (k1, k2, …, k m ) g = 1, 2, …, v; v ≤ m! (9);

[0036] In Equation (9): S(g) represents the g-th symbol sequence; g represents the index of the symbol sequence, ranging from 1 to v, where v is the number of all possible permutations and combinations, and the maximum value is m!; k1, k2, …, k m represent the indices of the elements in the reconstructed phase space vector, and these indices are the results after sorting the elements from small to large; m represents the embedding dimension;! represents factorial;

[0037] Step 2.4: Denote the probability of each symbol sequence appearing as P(i), and define the permutation entropy H pe as:

[0038]

[0039] In Equation (10): 0 ≤ Hpe ≤ 1, which reflects the randomness of the time series. The degree of randomness increases with the increase of the entropy value; P i represents the probability of occurrence of each symbol sequence, and b represents the normalization factor of the permutation entropy H pe .

[0040] The said step 3 includes the following steps:

[0041] Step 3.1: Select the minimum multi-scale permutation entropy as the objective function and perform secondary decomposition on the reconstructed complex component:

[0042] First, construct a constrained optimization model as described in Equation (11) below. The objective is to minimize the deviation of the bandwidth and center frequency of each mode function while ensuring that the sum of all mode functions is equal to the original signal;

[0043]

[0044] In Equation (11): {u k} and {w k} are the k-th mode component and its corresponding center frequency after inputting the component respectively; K is the number of decomposition times; δ t is the Dirac distribution; G CIMF (t) is the original signal; represents the exponential term of the center frequency; t is the sampling time u k (t) is the k-th mode component corresponding after decomposition; is the distribution of the mode function u k (t) in the frequency domain obtained by taking the derivative with respect to time t;.2 is the Euclidean norm.

[0045] Step 3.2: Introduce the augmented Lagrangian function to solve the constructed constrained optimization model as described in Equation (11);

[0046]

[0047] In Equation (12): L(.) is the Lagrangian operator; {u k}{ω k} is the set of the k-th mode component and the center frequency corresponding after decomposition, α is the quadratic penalty factor; λ is the Lagrange multiplier; λ(t) is the Lagrangian operator at the t-th sampling time;

[0048] Step 3.3: Initialize the parameters and set {u k 1}, {w k 1}, λ 1 , where {u k 1} is the first modal component corresponding after decomposition; {w k 1} is the first modal center frequency corresponding after decomposition; λ 1 is the first Lagrangian operator corresponding after decomposition.

[0049] Then update {u k} and {w k} through the following formula, looping from 1 to K;

[0050]

[0051] In Equation (13): represents the Fourier transform of the k-th modal function in the (n + 1)-th iteration; represents the Fourier transform of the original signal f(t); represents the sum of the modal functions in the frequency domain; represents the frequency domain representation of the Lagrange multiplier; ω represents frequency; ω k represents the center frequency of the k-th modal function;

[0052]

[0053] In Equation (14): represents the updated value of the center frequency of the k-th modal function, and n + 1 represents the current iteration step; represents the k-th modal function.

[0054] The value of λ is also updated through the following formula:

[0055]

[0056] In Equation (15): represents the updated value of the Lagrange multiplier λ in the frequency domain; represents the updated value of the Lagrange multiplier in the frequency domain in the n-th iteration, represents the updated value of the k-th modal function in the frequency domain at the current (n + 1)-th iteration step.

[0057] Given ε > 0, the iteration terminates when the following conditions are met.

[0058]

[0059] In Equation (16): ε is the criterion for terminating the decomposition; represents the updated value of the k-th modal function at the current (n + 1)-th iteration step; Denote the updated value of the k-th mode function at the current n-th iteration step; k represents the order of the mode function.

[0060] Step 4 includes the following steps:

[0061] Step 4.1: Use CNN to extract deep features from the decomposed component data to generate a multi-dimensional feature matrix:

[0062] First, input the mode components IMFs obtained by ICEEMDAN and VMD into CNN, and slide the convolutional kernels in CNN on the IMFs to capture local features. Specifically, as Figure 9 shown.

[0063] After the convolution operation, a ReLU (Rectified Linear Unit) activation function is connected to introduce non-linearity into CNN, enabling CNN to capture more complex features and patterns.

[0064] Each convolutional layer consists of multiple convolutional kernels (or filters), and each convolutional kernel slides on the input data to generate a feature map F k , that is, the k-th feature map. The feature map is the result of the input data after the convolution operation, denoted as

[0065] F q =φ(M q *X + C q )(17);

[0066] In formula (17): φ represents the activation function, introducing non-linearity; M q is the weight matrix of the q-th convolutional kernel; C q is the bias term corresponding to the q-th convolutional kernel; q is the number of convolutional kernels; X is the input data.

[0067] Then, enter the pooling layer. Through max-pooling and average-pooling, reduce the spatial size of the data, thereby reducing the number of parameters and computational complexity, while retaining the most important features and enhancing the generalization ability of the model. The feature output P after max-pooling is expressed as:

[0068] P = max(F q )(18);

[0069] In formula (18): F q is the feature output after CNN feature extraction.

[0070] After multiple layers of convolution, activation, and pooling operations, the CNN converts the original IMFs into a feature vector in a high-dimensional space. The CNN reduces the free parameters in the model by reusing the same convolutional filters at all positions. Weight sharing means that each filter learns a set of weights, ensuring parameter and computational efficiency when extracting features from the network. This allows for effective encoding of local features in the input data and helps reduce overfitting. The pooling layer reduces the dimensionality of the information output by the convolutional layer, reducing the number of parameters and the computational load in the subsequent processing stages. Max pooling and average pooling extract the maximum and average values, respectively, from local regions of the input data to achieve feature compression and feature extraction. The convolutional and pooling operations are as Figure 10 shown.

[0071] Step 4.2: The multi-dimensional feature matrix is transformed into one-dimensional features through a flattening layer and fed into the BiGRU-MHA unit for training. The BiGRU is used to capture the dependencies in the input features. The computational process of the GRU is as follows:

[0072]

[0073] In Equation (19): z t and r t represent the update gate and reset gate activation vectors, respectively, and represent the current candidate hidden state; is the transfer state at time t; h t and h t-1 represent the output states of the GRU at the current and previous time steps, respectively; W z , W r and W h are the weight matrices of the update gate, reset gate, and candidate hidden state, respectively; b z , b r and b h represent the corresponding bias matrices; σ is the sigmoid activation function; tanh is the hyperbolic tangent function; the symbol ⊙ represents the Hadamard product, i.e., the element-wise multiplication of vectors or matrices.

[0074] The present invention uses a bidirectional GRU, i.e., two GRU units with opposite directions, to construct the BiGRU, achieving parallel processing of forward and backward information in sequence data and avoiding time information entanglement. The computational process of the BiGRU is as follows:

[0075] L Bi = G(L for , L rev )(20);

[0076] In Equation (20): L Bi represents the output of the BiGRU; L for and L revThey are the forward and reverse GRU outputs respectively. By processing the forward and reverse time information of the sequence in parallel, BiGRU avoids the parameter entanglement caused by different inputs.

[0077] Step 4.3: Use MHA to calculate the similarity between elements in the BiGRU output feature sequence and generate weight factors, thereby enhancing the relationship between elements and guiding the CNN-BiGRU-MHA model to pay more attention to the key information in the sequence. The multi-head attention mechanism captures information from different subspace perspectives by executing multiple self-attention processes in parallel, enhancing the model's ability to process complex data sequences. The specific steps are as follows:

[0078]

[0079] In Equation (21): Attention(q, k, v) represents the output of the self-attention mechanism, that is, the self-attention weights; q, k, and v represent the query vector, key vector, and value vector respectively. These are obtained by multiplying the input data X by three different weight matrices to improve learnability. k T represents the transpose of the key vector; is a scaling factor, and the weights are obtained by taking the dot product of the query vector and the key vector and scaling appropriately.

[0080] [q i , k i , v i T =[w q , w k , w ν T ·a i

[0081] a 1,i =q i ·k i (22);

[0082] Among them: q i , k i and v i represent the query vector, key vector, and value vector in the i-th attention head respectively; the weight matrix w converts the embedding vector a i of the input sequence into the corresponding query vector, key vector, and value vector; w q represents the weight matrix corresponding to the query vector; w k represents the weight matrix corresponding to the key vector; w ν represents the weight matrix corresponding to the value vector T represents matrix transpose; a i is the embedding representation of the i-th element of the input sequence; a 1,i ​​It represents the output value of the i-th element after attention weighting in the first attention head;

[0083] The outputs a of all heads 1i , a 2i , …, a Hi will be concatenated. a 1i , a 2i , …, a Hi It represents the attention weights of the outputs of each head in the multi-head attention mechanism; the calculated weights are normalized using the SoftMax function to form a probability distribution, reflecting the relative importance of the elements. Specifically as follows:

[0084]

[0085] In Equation (23): SoftMax(a hi ) represents the result of normalizing each attention weight, a hi represents the attention weights of the outputs of each head, represents the exponential operation of the attention weights, h represents the index of the number of attention weights, u represents the index of the number of attention weights, and H represents the number of heads in the multi-head attention mechanism.

[0086] Finally, the normalized weights are weighted and summed to obtain a comprehensive sequence representation bi.

[0087] b i = ∑SoftMax(a hi ) · v i (24)

[0088] In Equation (24): v i represents the value vector in Equation (22).

[0089] Step 4.4: Use SAO (Snow Ablation Optimizer) to optimize the BiGRU parameters, including adjusting the learning rate, regularization, convolution kernel size, and number of hidden layers. Specifically, introduce the SAO module in the pytorch framework to dynamically adjust the BiGRU parameters, including the learning rate (learning), regularization (Regularization), convolution kernel size (kernel_size), and number of hidden layers (Hidden Layers). Just introduce the module without manual adjustment to ensure the adaptation of the BiGRU model to the data features.

[0090] Use the optimized CNN - BiGRU - MHA model to predict the test dataset, and synthesize the prediction results of different components to generate the final ratio error prediction value.

[0091] In step 4, the depth prediction model is the CNN-BiGRU-MHA overall prediction model, which includes equations (17) to (24).

[0092] The sub-model network is the BiGRU module, which includes equations (19) to (20).

[0093] For a current transformer error prediction method based on improved adaptive composite mode decomposition according to the present invention, the technical effects are as follows: 1) The advantages of step 1 of the present invention: The measurement data of the current transformer is preprocessed based on the improved complete ensemble empirical mode decomposition with adaptive noise (ICEEMDAN), which can effectively reduce mode mixing and noise interference, decompose the measurement data of the current transformer into multiple clear intrinsic mode functions, capture multi-band features, improve the accuracy of signal processing and the generalization ability of the model, and provide high-quality feature inputs for subsequent deep learning prediction models.

[0094] 2) The advantages of step 2 of the present invention: Feature selection and reconstruction are performed on the complex components decomposed by ICEEMDAN based on MPE and PDF, which can quantify the complexity of each mode component, identify the main trends and key features in the signal, and then screen out the most representative mode components through the probability density function, thereby reducing the complexity of model learning and improving the accuracy and efficiency of prediction.

[0095] 3) The advantages of step 3 of the present invention: The reconstructed complex components are further decomposed based on variational mode decomposition (VMD), effectively decomposing the signal into stable and reliable mode functions with different center frequencies, thereby improving the accuracy of signal processing and enhancing the model's ability to identify different frequency components in the signal.

[0096] 4) The advantages of step 4 of the present invention: A depth prediction model is constructed based on CNN-BiGRU, and the target optimization method is used to match the parameters of the sub-model network. The superposition of the sub-results generates the final prediction result, which can take into account the efficient extraction ability of the convolutional neural network for local features of the signal and the capture ability of the bidirectional gated recurrent unit for long-term dependence relationships in time series data. And using the target optimization method to finely match the parameters of the sub-model network can significantly improve the adaptability of the model to data features, make the final prediction result generated by the superposition of sub-results more accurate, and at the same time enhance the generalization ability and prediction performance of the model. Description of the Drawings

[0097] The present invention will be further described below in conjunction with the drawings and embodiments:

[0098] Figure 1 It is the overall framework diagram of the current transformer ratio error prediction model based on improved adaptive composite mode decomposition.

[0099] Figure 2It is the decomposition effect diagram of ICEEMDAN.

[0100] Figure 3 It is the multi-scale permutation entropy effect diagram.

[0101] Figure 4 It is the probability distribution estimation diagram.

[0102] Figure 5 It is the complex component reconstruction diagram.

[0103] Figure 6 It is the VMD decomposition effect diagram.

[0104] Figure 7 It is the schematic diagram of the CNN-BiGRU-MHA overall prediction model.

[0105] Figure 8 It is the prediction effect comparison diagram.

[0106] Figure 9 It is the schematic diagram of the convolutional kernel in CNN sliding on IMFs to capture local features.

[0107] Figure 10 It is the schematic diagram of convolution and pooling operations. Detailed implementation method

[0108] In the specific embodiments of the present invention, the proposed method is used to predict the ratio error data of current transformers.

[0109] Figure 1 It is the overall framework diagram of the model. It includes data collection and preprocessing: collecting the historical ratio error sequence data of the transformer and removing zero values. ICEEMDAN decomposition: processing the original data to obtain intrinsic mode functions (IMFs). Feature selection and reconstruction: using PE to calculate the complexity of components, retaining the main distribution through PDF, and reconstructing complex components. VMD decomposition: by selecting the minimum multi-scale permutation entropy as the objective function, performing secondary decomposition on the reconstructed complex components, and using an optimization algorithm to iteratively solve the optimal decomposition parameters of the variational problem. CNN-BiGRU-MHA model training: using CNN to extract deep features from the decomposed component data to generate a multi-dimensional feature matrix, and then sending it to the BiGRU-MHA unit through a flattening layer for training. Model parameter optimization: using SAO to optimize the BiGRU parameters, including adjusting the learning rate, regularization, convolutional kernel size, and number of hidden layers, to ensure the model's adaptation to data features. Ratio error prediction: using the optimized model to predict the test data set, integrating the prediction results of different components, and generating the final ratio error prediction value.

[0110] Figure 2 It is the decomposition effect diagram of ICEEMDAN. From Figure 2It can be seen that ICEEMDAN can decompose complex non-stationary signals into multiple mode functions with different frequency characteristics. The decomposition result reduces the mode mixing phenomenon and improves the accuracy and reliability of the decomposition. The low-frequency IMFs reflect the main trend of the signal, while the high-frequency IMFs contain the detailed information of the signal. All the information of the original signal is retained during the decomposition process, making subsequent analysis easier.

[0111] Figure 3 It is the multi-scale permutation entropy effect diagram. From Figure 3 it can be seen that the permutation entropy value reflects the complexity and randomness of the signal at different scales. The mode functions with lower permutation entropy values usually contain the main trend or low-frequency components of the signal, while the mode functions with higher permutation entropy values may contain noise or high-frequency details. The distribution of the permutation entropy values can be used as a basis for feature selection, and the mode functions with lower permutation entropy values are selected for further analysis.

[0112] Figure 4 It is the probability distribution estimation diagram. The analysis shows that in the range of 0 to 0.2, the PE values of the sub-components are lower and the probability density is higher. This part constitutes the main part of the overall distribution of the decomposition sequence. Therefore, this main part is retained, and the components with higher PE values and the non-main part (range 0.2 to 1) are selected for further reconstruction and decomposition.

[0113] Figure 5 It is the complex component reconstruction diagram. IMF1 to IMF6 are the complex components of the secondary decomposition. In addition, in order to ensure the rigor of the decomposition process, the residual is also considered. Therefore, the reconstructed signal components mainly consist of IMF1, IMF2, IMF3, IMF4, IMF5, IMF6 and the decomposition residual.

[0114] Figure 6 It is the VMD decomposition effect diagram. From Figure 6 it can be seen that VMD can decompose complex non-stationary signals into multiple mode functions with different frequency characteristics, and the decomposition result retains all the information of the original signal, including the main trend and the detailed information.

[0115] Figure 8 It is the prediction effect comparison diagram. Four model configurations are used as control groups for comparative experiments: CNN-BiGRU, CNN-BiGRU-MHA, single-mode decomposition CNN-BiGRU-MHA, and combined-mode decomposition CNN-BiGRU-MHA to confirm the superiority of the method proposed in the present invention.

[0116] From the perspective of data prediction, comprehensively using ICEEMDAN for signal decomposition, MPE and PDF for feature selection and reconstruction, VMD for further decomposition of complex components, and the CNN-BiGRU-MHA deep learning architecture for feature learning and prediction can effectively process non-linear and non-stationary signals, extract deep features, capture complex dynamic relationships in time series data, and precisely adjust model parameters through objective optimization methods, improving the accuracy of prediction and the generalization ability of the model.

Claims

1. A current transformer error prediction method based on improved adaptive composite mode decomposition, characterized in that It includes the following steps: Step 1: Preprocess the measurement data of the current transformer based on ICEEMDAN; Step 2: Perform feature selection and reconstruction on the complex components decomposed by ICEEMDAN based on multi-scale permutation entropy and probability density function; Step 3: Further decompose the reconstructed complex components based on variational mode decomposition VMD; Step 4: Build a deep prediction model based on CNN-BiGRU, optimize the network parameters of the sub-model using the target optimization method, and finally superimpose the sub-results to obtain the final prediction result of the current transformer.

2. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 1, wherein: The said Step 1 includes the following steps: Step 1.1: Add a Gaussian white noise signal that follows a normal distribution to the measured data of the original current transformer to obtain a new signal sequence In formula (1): represents the original signal, i represents the number of times of adding white noise, is the first-order modal component after EMD decomposition, α0 is the noise intensity adjustment coefficient; w i is the i-th Gaussian white noise with zero mean and zero variance; Step 1.2: Perform a local mean operation on the new signal sequence to calculate the first set of residuals: In Equation (2): R1 is the first residual component; M(·) represents the local mean calculated by applying the EMD algorithm to the sequence. Step 1.3: Subtract the first set of residuals from the original signal to obtain the first set of intrinsic mode functions IMF1: Step 1.4: Iteratively calculate the residual and modal components; repeat the above process, add white noise to the residual signal in each round, and calculate the new residual and modal components until all decompositions are completed; In formula (4): R k is the k-th set of residuals; R k-1 is the (k-1)-th set of residuals, and α k-1 is the noise adjustment coefficient when adding noise for the (k-1)-th time. is the k-th order modal component after EMD decomposition, is the k-th modal component; the noise intensity adjustment coefficient α k in the decomposition process controls the proportion of added white noise; it is determined according to the characteristics of the signal and can be expressed as: In formula (5): ε0 is the reciprocal of the signal-to-noise ratio between the initially added noise and the signal to be analyzed; std(λ) represents the standard deviation of the threshold, and E1(w i ) represents the first-order mode component after EEMD decomposition, and std(E1(w i )) represents the standard deviation of the first-order mode component after EEMD decomposition. γ represents the residual signal in the current decomposition step, represents the local mean calculated by applying the EMD algorithm to the original signal sequence before the decomposition starts. k represents the current decomposition times, and K represents the maximum decomposition times.

3. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 1, wherein: The said Step 2 includes the following steps: Step 2.1: Coarsely grain the original signal x(t) to obtain multiple coarsely grained sequences, and calculate the average value y(i) of each coarsely grained sequence: In Equation (6): y(i) is the average value of the i-th coarsely grained sequence; s is the scale factor; N is the length of the original signal; i is the index of the coarsely grained sequence, j is the sampling point index of the original sequence, and x(j) represents the j-th sampling point of the original signal; Step 2.2: Reconstruct the phase space of each coarsely grained sequence using the mean value of each coarsely grained sequence calculated in Step 2.1 to obtain the matrix {y(i), i = 1, 2, ···, N / s}: In Equation (7): m is the embedding dimension; τ is the delay time, k = n - (m - 1)τ; h is the maximum index in the coarsely grained sequence that can form a complete phase space vector; n represents the length of the coarsely grained matrix, which is N / s; y(1) represents the first element in the coarsely grained sequence, y(2) represents the second element in the coarsely grained sequence, y(j) represents the j-th element in the coarsely grained sequence, and y(h) represents the h-th element in the coarsely grained sequence; y(1 + τ) represents the value of the (1 + τ)-th element in the coarsely grained sequence, that is, starting from the first element, the element value after an interval of 1 time step, y(2 + τ) represents the element value after an interval of 1 time step starting from the second element, y(j + τ) represents the element value after an interval of 1 time step starting from the j-th element, and y(h + τ) represents the element value after an interval of 1 time step starting from the h-th element; y[1 + (m - 1)τ] represents the element value after an interval of m - 1 time steps starting from the first element, y[2 + (m - 1)τ] represents the element value after an interval of m - 1 time steps starting from the second element, y[j + (m - 1)τ] represents the element value after an interval of m - 1 time steps starting from the j-th element, and y[h + (m - 1)τ] represents the element value after an interval of m - 1 time steps starting from the h-th element; Step 2.3: Arrange the elements of each reconstructed component in ascending order: {y(j+(k1-1)τ)y(j+(k2-1)τ)…y(j+(k m -1)τ)}(8); In formula (8): k1, k2, …, k m are the indices of the ordered elements in the original reconstructed component; if there are equal elements in the reconstructed component, they are sorted according to their indices; thus, each reconstructed component corresponds to a set of symbol sequences; S(g) = (k1, k2, …, k m ) where g = 1, 2, …, v; v ≤ m!(9); In formula (9): S(g) represents the g-th symbol sequence; g represents the index of the symbol sequence, ranging from 1 to v, where v is the number of all possible permutations and combinations, and the maximum value is m!; k1, k2, …, k m represent the indices of the elements in the reconstructed phase space vector, and these indices are the results after sorting the elements from small to large; m represents the embedding dimension;! represents the factorial; Step 2.4: Denote the probability of each symbol sequence occurrence as P(i), and define the permutation entropy H pe as: In formula (10): 0 ≤ Hpe ≤ 1, which reflects the randomness of the time series; the degree of randomness increases with the increase of the entropy value; P i represents the probability of the occurrence of each symbol sequence, and b represents the normalization factor of the permutation entropy H pe of.

4. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 1, wherein: The said Step 3 includes the following steps: Step 3.1: Select the minimum multi-scale permutation entropy as the objective function and perform secondary decomposition on the reconstructed complex component: First, a constrained optimization model is constructed as described by Equation (11) below. The goal is to minimize the deviation of the bandwidth and center frequency of each modal function while ensuring that the sum of all modal functions is equal to the original signal; In Equation (11): {u k} and {w k} are the k-th modal component after the input components and their corresponding center frequencies respectively; K is the number of decomposition times; δ t is the Dirac distribution; G CIMF (t) is the original signal; is the exponential term representing the center frequency; t is the sampling time u k (t) is the k-th modal component corresponding after decomposition; is the distribution of the modal function u k (t) in the frequency domain obtained by differentiating with respect to time t;.2 is the Euclidean norm; Step 3.2: Introduce the augmented Lagrangian function to solve the constructed constrained optimization model as described by Equation (11); In Equation (12): L(.) is the Lagrangian operator; {u k}{ω k} is the set of the k-th modal component and the center frequency corresponding to the decomposition. α is the quadratic penalty factor; λ is the Lagrange multiplier; λ(t) is the Lagrangian operator at the t-th sampling moment; Step 3.3: Initialize the parameters, and set {u k 1}, {w k 1}, and λ 1 , where {u k 1} is the first modal component after decomposition; {w k 1} is the first modal center frequency after decomposition; and λ 1 is the first Lagrange operator after decomposition. Then update {u k} and {w k} according to the following formula, looping from 1 to K; In Equation (13): represents the Fourier transform of the k-th mode function in the (n + 1)-th iteration; represents the Fourier transform of the original signal f(t); represents the sum of the mode functions in the frequency domain; represents the frequency domain representation of the Lagrange multiplier; ω represents the frequency; ω k represents the central frequency of the k-th mode function; In formula (14): represents the updated value of the central frequency of the k-th mode function, and n + 1 represents the current iteration step; represents the k-th mode function; The value of λ is also updated by the following formula: In Equation (15): represents the updated value of the Lagrange multiplier λ in the frequency domain; represents the updated value of the Lagrange multiplier in the frequency domain at the n-th iteration, represents the updated value of the k-th mode function in the frequency domain at the current (n + 1)-th iteration step; Given ε > 0, the iteration terminates when the following conditions are met; In formula (16): ε is the criterion for terminating the decomposition; represents the updated value of the k-th mode function in the current (n + 1)-th iteration step; represents the updated value of the k-th mode function in the current n-th iteration step; k represents the order of the mode function.

5. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 1, wherein: The said Step 4 includes the following steps: Step 4.1: Use CNN to extract deep features from the decomposed component data to generate a multi-dimensional feature matrix: Step 4.2: The multi-dimensional feature matrix is transformed into one-dimensional features through a flattening layer and fed into the BiGRU-MHA unit for training. The BiGRU is used to capture the dependencies in the input features. Step 4.3: Use MHA to calculate the similarity between elements in the BiGRU output feature sequence and generate a weight factor to enhance the relationship between elements; Step 4.4: Use SAO to optimize the BiGRU parameters, including adjusting the learning rate, regularization, convolutional kernel size, and number of hidden layers; Use the optimized CNN-BiGRU-MHA model to predict the test dataset, and synthesize the prediction results of different components to generate the final ratio error prediction value.

6. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 5, wherein: Step 4.1 includes: First, input the modal components IMFs obtained by ICEEMDAN and VMD into CNN, and slide the convolutional kernels in CNN on the IMFs to capture local features; After the convolution operation, a ReLU (Rectified Linear Unit) activation function is connected to introduce non-linearity into CNN, enabling CNN to capture more complex features and patterns; Each convolutional layer consists of multiple convolutional kernels. Each convolutional kernel slides over the input data to generate a feature map F k , which is the k-th feature map; the feature map is the result of the input data after the convolution operation and is denoted as F q = φ(M q *X + C q )(17); In formula (17): φ represents an activation function, introducing non-linearity; M q is the weight matrix of the q-th convolutional kernel; C q is the bias term corresponding to the q-th convolutional kernel; q is the number of convolutional kernels; X is the input data; Next, enter the pooling layer. Through max-pooling and average-pooling, the feature output P after max-pooling is expressed as: P = max(F q )(18); In formula (18): F q The feature output after CNN feature extraction; Max-pooling and average-pooling respectively extract the maximum value and the average value from the local area of the input data to achieve feature compression and feature extraction.

7. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 5, wherein: In Step 4.2, the calculation process of GRU is expressed as follows: In Equation (19): z t and r t represent the activation vectors of the update gate and the reset gate respectively, and represent the current candidate hidden state respectively; is the transfer state at time t; h t and h t-1 represent the output states of the GRU at the current and the previous time steps respectively; W z , W r and W h are the weight matrices of the update gate, the reset gate and the candidate hidden state respectively; b z , b r and b h represent the corresponding bias matrices; σ is the sigmoid activation function; tanh is the hyperbolic tangent function; the symbol ⊙ represents the Hadamard product, i.e., the element-wise multiplication of vectors or matrices; Use bidirectional GRU to construct BiGRU, which realizes the parallel processing of forward and backward information in sequence data. The calculation process of BiGRU is as follows: L Bi = G(L for , L rev )(20); In formula (20): L Bi represents the output of BiGRU; L for and L rev are the forward and backward GRU outputs respectively.

8. The method for predicting the error of a current transformer based on improved adaptive composite mode decomposition according to claim 5, wherein: Step 4.3 specifically includes: In formula (21): Attention(q, k, v) represents the output of the self-attention mechanism, that is, the self-attention weight; q, k, and v represent the query vector, key vector, and value vector respectively; k T represents the transpose of the key vector; is a scaling factor, and the weight is obtained by taking the dot product of the query vector and the key vector and scaling appropriately; a 1,i = q i · k i (22); where: q i , k i and v i represent the query vector, key vector, and value vector in the i-th attention head respectively; the weight matrix w transforms the embedding vector a i of the input sequence into the corresponding query vector, key vector, and value vector; w q represents the weight matrix corresponding to the query vector; w k represents the weight matrix corresponding to the key vector; w ν represents the weight matrix corresponding to the value vector, T represents matrix transpose; a i is the embedding representation of the i-th element of the input sequence; a 1,i represents the output value after attention weighting of the i-th element in the first attention head; Output a of all heads 1i , a 2i , …, a Hi will be concatenated, a 1i , a 2i , …, a Hi The attention weights of the outputs of each head of the multi-head attention mechanism represented; the calculated weights are normalized using the SoftMax function to form a probability distribution that reflects the relative importance of the elements; specifically as follows: In formula (23): SoftMax(a hi ) represents the result of normalizing each attention weight, and a hi represents the attention weights output by each head. represents the exponential operation of the attention weights. h represents the index of the number of attention weights, u represents the index of the number of attention weights, and H represents the number of heads in the multi-head attention mechanism. Finally, perform weighted summation on the normalized weights to obtain a comprehensive sequence representation bi; b i = ∑ SoftMax(a hi )·v i (24) In Equation (24): v i represents the value vector in Equation (22).

Citation Information

Cited By

  • New energy station voltage transformer signal processing method and system

    CN121524609A