Fault detection method based on combination of Transform model of VMD and BiGRU
By combining VMD, Transformer model and BiGRU network, extracting and processing the characteristics of voltage and current signals, the problems of information loss and modal aliasing in the existing fault detection methods are solved, and high-precision and efficient fault detection are achieved.
Patent Information
- Application Number
- CN202510226674.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing fault detection methods have problems of information loss and modal aliasing when processing voltage and current signals, making it difficult to fully utilize the timing characteristics of the signal, and the accuracy and efficiency are insufficient.
The fault detection method combined with VMD-based Transformer model and BiGRU is adopted to obtain multiple modal components through variational modal decomposition, extract and process features, and use the Transformer model and BiGRU network to capture global and timing features to diagnose and classify faults.
It improves the accuracy and efficiency of power system fault diagnosis, realizes high-precision and high-rootability detection of voltage and current faults, and has strong adaptability and good real-time performance.
Smart Images

Figure CN120162669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system fault detection, and in particular to a fault detection method combining a Transformer model based on VMD and a BiGRU. Background Art
[0002] In modern power systems, the abnormal detection of voltage and current signals is crucial for ensuring the safety of the power grid. Traditional fault detection methods mainly rely on signal processing technologies such as Fourier transform, wavelet transform, and empirical mode decomposition (EMD). However, these methods have certain limitations. For example, the Fourier transform is difficult to process non-stationary signals, the selection of wavelet transform basis functions has a great impact on the analysis results, and the EMD method is prone to mode mixing.
[0003] In recent years, variational mode decomposition (VMD) technology has been widely applied in the field of signal processing due to its good time-frequency decomposition performance. VMD can effectively suppress noise, improve the decomposition accuracy, and avoid the mode mixing problem in the traditional EMD method. However, there is still a problem of information loss when using VMD alone for feature extraction, and it is difficult to fully utilize the temporal characteristics of signals.
[0004] To further improve the accuracy of fault detection, deep learning methods have gradually been introduced into the analysis of power system signals. The Transformer model, due to its powerful feature extraction ability, can effectively capture long-range dependencies and is suitable for modeling complex signals. However, the Transformer model has certain limitations when processing time series data. For example, its ability to capture local temporal features is weak. Therefore, combining the Transformer with a bidirectional gated recurrent unit (BiGRU) can fully utilize the global feature extraction ability of the Transformer and the temporal information capture ability of the BiGRU to improve the accuracy of fault classification.
[0005] In summary, the existing fault detection methods still pose certain challenges when dealing with voltage and current signals. To address this issue. Summary of the Invention
[0006] To solve the current technical problems, the main objective of the present invention is to provide a fault detection method combining a Transformer model based on VMD and a BiGRU. By using a method that combines variational mode decomposition (VMD), a Transformer model, and a bidirectional gated recurrent unit (BiGRU), the original electrical signal is decomposed to obtain multiple modal components, features are extracted and processed, and fault diagnosis and classification are carried out, thereby improving the accuracy and efficiency of fault diagnosis and achieving high-precision and high-robustness detection of voltage and current faults.
[0007] To achieve the above technical features, the object of the present invention is achieved as follows: A fault detection method combining a Transformer model based on VMD and BiGRU includes the following steps:
[0008] Step 1, data acquisition:
[0009] Collect voltage and current signal data of the power grid to be detected;
[0010] Step 2, data preprocessing:
[0011] Preprocess the collected voltage and current signal data to obtain the processed signal data;
[0012] Step 3, perform variational mode decomposition (VMD):
[0013] Perform variational mode decomposition on the signal data after preprocessing in Step 2 to obtain modal components;
[0014] Step 4, feature extraction and encoding:
[0015] Extract features from each modal component obtained by VMD decomposition in Step 3, perform encoding processing, and obtain feature vectors;
[0016] Step 5, Transformer model processing:
[0017] Process the feature vectors in Step 4 using the Transformer model to obtain high-dimensional features;
[0018] Step 6, bidirectional gated recurrent unit (BiGRU) processing:
[0019] Process the high-dimensional features in Step 5 using BiGRU to obtain the feature vectors output by bidirectional GRU;
[0020] Step 7, fault classification:
[0021] Classify the feature vectors output in Step 6 using a softmax classifier to obtain classification results;
[0022] Step 8, fault detection and warning:
[0023] Based on the classification results output in Step 7, determine whether there is a fault in the current voltage and current, and issue an alarm when a fault occurs.
[0024] Preferably, the specific steps of Step 1 include: selecting appropriate capacitive voltage transformers (CVTs) and protective current transformers (CTs), installing and connecting these transformers to the detection module, transmitting signals to the information processing module through the wireless transmission module, and recording the voltage and current signal data.
[0025] Preferably, the specific steps of Step 2 include: amplifying and denoising the collected voltage and current signals, then normalizing the amplitude of the signals, scaling the amplitude of the signals to a specific range, and finally cleaning the data, processing missing values and outliers to obtain the signal data x(t).
[0026] Preferably, the specific steps of Step 3 include: inputting the preprocessed voltage and current signal x(t) into the VMD algorithm to solve the variational model;
[0027] Specific solution process:
[0028] First, initialize the modal components and the linear frequency , and use the original signal data x(t) as the initial modal component; then, for each modal component u k (t), given other modal components and the center frequency, update it by solving the following optimization problem:
[0029]
[0030] where represents the inverse Fourier transform, represent the Fourier transforms of x(t) and respectively, α is the bandwidth parameter of the modal component, and for the center frequency ω k of each modal component, given the modal component, update it by solving the following optimization problem:
[0031]
[0032] Repeat the above update steps until the convergence condition is met, and finally decompose x(t) into K modal components u k (t).
[0033] Preferably, the specific steps of Step 4 include: extracting features from each modal component obtained by VMD decomposition. The extracted features include the energy E k of the modal component, the mean M k , the variance V k , and the time-frequency feature F k (τ, ω) obtained by the wavelet transform method. Normalize and encode the statistical features and time-frequency features so that their ranges are between [0, 1]. The normalization formulas are as follows:
[0034]
[0035] Among them, E max and E min and M max and M min and V max and V min and F k,max and F k,min are respectively the maximum and minimum values of each modulus feature. The normalized statistical and time-frequency features are combined into a feature vector N with a sequence length of n and a feature dimension of d.
[0036] Preferably, the specific steps of step five include: inputting the feature vector N into a Transformer model. Based on the self-attention mechanism, the Transformer model performs weighted summation on the feature vector to obtain a weighted feature representation T. Then, through a feed-forward neural network, non-linear transformation is performed on the weighted feature to obtain a feature S. Calculate the output T′ and S′ of the residual connection and layer normalization, and use the output S′ of the feed-forward neural network as the high-dimensional feature representation.
[0037] Preferably, the specific steps of step six include: inputting the feature representation S′ output by the Transformer model into a BiGRU network. According to the update formula of the forward GRU, calculate from the start position to the end position of the sequence. According to the update formula of the reverse GRU, calculate from the end position to the start position of the sequence. Concatenate the hidden states of the forward and reverse GRUs at each time step to obtain the output G of the bidirectional GRU t .
[0038] Preferably, the specific steps of step seven include: inputting the feature vector G t output by the BiGRU network into a softmax classifier. The classifier performs a linear transformation on the input feature vector to obtain Z, and converts Z into a probability distribution P through the softmax function. According to the probability distribution P, select the fault type with the highest probability as the final classification result.
[0039] Preferably, the specific steps of step eight include: judging whether there is a fault in the current voltage and current signals according to the output result of the fault classifier;
[0040] Among them, one of the fault types is "normal state", and its index is i0. If then it is determined that there is a fault. Once a fault is detected, the system immediately generates a warning message and emits an alarm sound.
[0041] Preferably, in the third step, the variational model of VMD is expressed as:
[0042]
[0043] where u k (t) = A k (t)cos(φ k (t));
[0044] A k (t) is the instantaneous amplitude of the modal component, φ k (t) is the instantaneous phase, δ(t) is the Dirac function, * represents the convolution operation, ω k is the central frequency of the modal component u k (t), and λ is the Lagrange multiplier for balancing the reconstruction error and the bandwidth penalty term;
[0045] The convergence of VMD decomposition is judged by monitoring the change amount of the modal component; when the change amount of the modal component is less than a preset threshold, it is considered that the VMD decomposition has converged. Specifically, by calculating the Euclidean distance between the modal components of two adjacent iterations:
[0046]
[0047] When is less than the threshold ε, stop the iteration;
[0048] In the fourth step, the energy, mean, variance, and time-frequency characteristics are respectively:
[0049]
[0050] where T is the total duration of the signal, ψ(τ, ω, t) is the wavelet function, τ is the time shift parameter, ω is the frequency parameter, and * represents the complex conjugate;
[0051] In the fifth step, the multi-head self-attention decomposes the input feature vector N into h heads, each head corresponding to a self-attention mechanism, and the output feature of the weighted multi-head self-attention:
[0052] T = Concat(head1, head2,..., head h )W O ; (13)
[0053] In the formula, head i is the self-attention output of the i-th head, and W O is the output weight matrix;
[0054] The feature obtained through the non-linear transformation:
[0055] S = FFN(T) = max(0, TW 1 + b 1 )W 2 + b 2 ; (14)
[0056] Wherein, W 1 , W 2 , b 1 , b 2 are the weight and bias parameters of the feed - forward neural network;
[0057] Output features of residual connection and layer normalization:
[0058] T' = LayerNorm(N + T); (15)
[0059] S' = LayerNorm(T' + S); (16)
[0060] Wherein, LayerNorm is the layer normalization function;
[0061] In the above step seven, the linear transformation output representation:
[0062] Z = GW + b; (17)
[0063] Wherein, W is the weight matrix and b is the bias vector;
[0064] Probability distribution:
[0065]
[0066] The fault type with the maximum predicted probability:
[0067]
[0068] Wherein, P i is the probability of the i - th fault type, and A is the number of fault types.
[0069] The present invention has the following beneficial effects:
[0070] The present invention provides an electrical signal fault detection method based on VMD + Transformer + BiGRU, aiming to improve the accuracy and efficiency of power system fault diagnosis. This method first collects voltage and current signals through a capacitive voltage transformer (CVT) and a protective current transformer (CT), and performs preprocessing, including amplification, denoising, normalization, and data cleaning. Then, VMD is used to decompose the preprocessed signal into multiple modal components, and the energy, mean, variance, and time-frequency characteristics of each modal component are extracted, and normalized encoding is performed to form feature vectors. These feature vectors are then input into the Transformer model, and deep features are extracted through the multi-head self-attention mechanism and the feed-forward neural network, and residual connection and layer normalization processing are performed. The output of the Transformer model is then input into the BiGRU network to capture the time-series change characteristics, and a feature vector of a fixed length is output. Finally, the output of the BiGRU network is input into the softmax classifier, and the probability distribution is output through linear transformation and the softmax function, and the fault type with the highest probability is selected as the final result. If the fault type is not "normal state", a warning message is generated and an alarm is issued. The above method has the advantages of high-precision fault detection, strong robustness, strong adaptability, and good real-time performance, and can effectively improve the accuracy and efficiency of power system fault diagnosis. Description of the Drawings
[0071] The present invention will be further described below in conjunction with the drawings and embodiments.
[0072] Figure 1 It is a schematic flow chart of the method of the present invention.
[0073] Figure 2 It is a schematic flow chart of the Transformer model.
[0074] Figure 3 It is a visualization diagram of the t-SNE features of the original data.
[0075] Figure 4 It is a visualization diagram of the t-SNE features after model training. Detailed Embodiments
[0076] The embodiments of the present invention will be further described below in conjunction with the drawings.
[0077] Embodiment 1:
[0078] Refer to Figure 1 , a fault detection method combining a Transformer model based on VMD and BiGRU, specifically including the following steps:
[0079] Step 1. Data acquisition: Collect voltage and current signals, select appropriate capacitive voltage transformers (CVTs) and protective current transformers (CTs), install and connect these transformers to the detection module, transmit the signals to the information processing module through the wireless transmission module, and record the voltage and current signal data;
[0080] Step 2. Data preprocessing: Amplify and denoise the collected voltage and current signals, then normalize the amplitude of the signals, scale the amplitude of the signals to a specific range, and finally clean the data, handle missing values and outliers to obtain the signal data x(t);
[0081] Step 3. Perform variational mode decomposition (VMD): Input the preprocessed voltage and current signal x(t) into the VMD algorithm to solve the variational model; First, initialize the modal components and the linear frequency , and use the original signal x(t) as the initial modal component; Then, for each modal component u k (t), given other modal components and the center frequency, update it by solving the following optimization problem:
[0082]
[0083] where, represents the inverse Fourier transform, represent the Fourier transforms of x(t) and respectively, α is the bandwidth parameter of the modal component, and for the center frequency ω k of each modal component, given the modal component, update it by solving the following optimization problem:
[0084]
[0085] Repeat the above update steps until the convergence condition is met, and finally decompose x(t) into K modal components u k (t).
[0086] Step 4. Feature extraction and encoding: Extract features from each modal component obtained by VMD decomposition. The extracted features include the energy E k , mean M k , variance V k , and the time-frequency feature F k (τ, ω) obtained by the wavelet transform method. Normalize and encode the statistical features and time-frequency features so that their ranges are between [0, 1]. The normalization formulas are respectively:
[0087]
[0088] Among them, E max , E min , M max , M min , V max , V min , F k,max and F k,min are respectively the maximum and minimum values of each modulus feature. The normalized statistical and time-frequency features are combined into a feature vector N with a sequence length of n and a feature dimension of d.
[0089] Step Five: Transformer Model Processing: Input the feature vector N into the Transformer model. The Transformer model, based on the self-attention mechanism, performs weighted summation on the feature vector to obtain the weighted feature representation T, and then performs a non-linear transformation on the weighted feature through a feed-forward neural network to obtain the feature S. Calculate the output T′ and S′ of the residual connection and layer normalization, and use the output S′ of the feed-forward neural network as the high-dimensional feature representation;
[0090] Step Six: Bidirectional Gated Recurrent Unit Processing: Input the feature representation S′ output by the Transformer model into the BiGRU network. Calculate according to the update formula of the forward GRU from the start position to the end position of the sequence, and calculate according to the update formula of the reverse GRU from the end position to the start position of the sequence. Concatenate the hidden states of the forward and reverse GRUs at each time step to obtain the output G of the bidirectional GRU t ;
[0091] Step Seven: Fault Classification: Input the feature vector G t output by the BiGRU network into the softmax classifier. The classifier performs a linear transformation on the input feature vector to obtain Z, converts Z into a probability distribution P through the softmax function, and selects the fault type with the highest probability as the final classification result;
[0092] Step Eight: Fault Detection and Warning: Determine whether there is a fault in the current voltage and current signals according to the output result of the fault classifier . There is a category "normal state" in the fault types, and its index is i0. If then it is determined that there is a fault. Once a fault is detected, the system immediately generates a warning message and emits an alarm sound.
[0093] Furthermore, in Step Two, the data is normalized using min-max normalization, and the formula is as follows:
[0094]
[0095] Further, the goal of the variational mode decomposition (VMD) in step three is to decompose the signal x(t) into K modal components u k (t), each modal component having a finite bandwidth and being orthogonal to each other. Minimize the sum of each modal component while satisfying the constraint that the sum of all modal components is equal to the original signal. It suppresses noise through bandwidth constraints and effectively avoids the occurrence of modal aliasing problems. The variational model of VMD can be expressed as:
[0096]
[0097] where, u k (t) = A k (t)cos(φ k (t));
[0098] A k (t) is the instantaneous amplitude of the modal component, φ k (t) is the instantaneous phase, δ(t) is the Dirac function, * represents the convolution operation, ω k is the center frequency of the modal component u k (t), and λ is the Lagrange multiplier that balances the reconstruction error and the bandwidth penalty term. By transforming the signal decomposition problem through a variational optimization framework into a minimization problem under constraints, the objective function contains two terms. The first term is the bandwidth constraint term, which minimizes the bandwidth of each mode to ensure the compactness of the modal component in the frequency domain. The second term is the reconstruction error term, which balances the accuracy of signal reconstruction through the Lagrange multiplier. The constraint condition is that the sum of all modal components must be equal to the original signal x(t).
[0099] The specific steps of VMD decomposition are as follows: First, initialize the modal components and the center frequency. Take the original signal x(t) as the initial estimate of all modal components, that is At the same time, preset according to spectral analysis Then iteratively update the modal components and the center frequency. Alternately optimize the modal components and the center frequency through the alternating direction method of multipliers (ADMM). The update formula for the modal components is Equation (0.1). Update each modal component one by one in the Fourier domain, considering the remaining modes as known. Control the bandwidth of the mode through the bandwidth parameter α. The larger α is, the narrower the mode. The Lagrange multiplier is indirectly reflected through α and the iterative process. The update formula for the center frequency is Equation (2). Determine the new center frequency by calculating the energy centroid of the modal component in the frequency domain, that is, the central position of the frequency domain energy distribution;
[0100] Finally, perform convergence judgment. When the change amount of the modal component is less than a preset threshold, that is, when the following formula is satisfied:
[0101]
[0102] When When it is less than the threshold ε, it can be considered that the VMD decomposition has converged and the iteration terminates; otherwise, continue to update.
[0103] Furthermore, in step 4, feature extraction is performed on each modal component obtained by VMD decomposition. The extracted features include the energy E k of the modal component, the mean value M k and the variance V k , as well as the time-frequency feature F k (τ, ω) obtained by the wavelet transform method. First, the energy E k is a basic feature of the signal in the time domain or frequency domain, which can reflect the intensity and activity of the signal. There are usually significant differences in the energy between normal signals and fault signals. In fault detection, changes in energy may indicate abnormalities or faults in the signal. Its calculation formula is as follows:
[0104]
[0105] Secondly, the mean value M k is the average value of the signal, which can reflect the baseline level of the signal. In some cases, faults may cause a shift in the baseline of the signal. By monitoring changes in the mean value, the baseline shift of the signal can be detected, which can be used as an important indicator for fault diagnosis. Its calculation formula is as follows, where T is the total duration of the signal.
[0106]
[0107] The variance V k is a measure of the degree of fluctuation of the signal, reflecting the degree of dispersion of the signal. Faults may cause an increase or decrease in the degree of fluctuation of the signal. By monitoring changes in the variance, these abnormalities can be detected in a timely manner. Its calculation formula is as follows:
[0108]
[0109] The time-frequency feature F k (τ, ω) can reflect the time and frequency characteristics of the signal at the same time. By analyzing the time-frequency feature, the change of the frequency component of the signal over time can be detected, so as to effectively identify faults. Its calculation formula is as follows:
[0110]
[0111] where T is the total duration of the signal, ψ(τ, ω, t) is the wavelet function, τ is the time shift parameter, ω is the frequency parameter, and * represents complex conjugate;
[0112] Finally, these features are normalized to form the feature vector N.
[0113] Further, in the fifth step, the multi-head self-attention decomposes the input feature vector N into h heads, each head corresponding to a self-attention mechanism. The output feature of the weighted multi-head self-attention is:
[0114] T = Concat(head1, head2, …, head h )W O ; (13)
[0115] In the formula, head i is the self-attention output of the i-th head, and W O is the output weight matrix;
[0116] The feature obtained through non-linear transformation is:
[0117] S = FFN(T) = max(0, TW 1 +b 1 )W 2 +b 2 ; (14)
[0118] In the formula, W 1 , W 2 , b 1 , b 2 are the weight and bias parameters of the feed-forward neural network;
[0119] The output features of residual connection and layer normalization are:
[0120] T′ = LayerNorm(N + T); (15)
[0121] S′ = LayerNorm(T′ + S); (16)
[0122] In the formula, LayerNorm is the layer normalization function;
[0123] Further, in the fifth step, the high-dimensional feature generation system based on the Transformer network is composed of a feature input interface, a self-attention module, a feed-forward neural network module, and a feature output interface. The input feature sequence is weighted by the self-attention module and then outputs high-dimensional features after expanding the dimensions through the feed-forward network, as Figure 2 shown.
[0124] Further, in the sixth step, the update formula of the forward GRU is:
[0125]
[0126] Among them, are the update gate, reset gate, and hidden state of the forward GRU at time step t, respectively, are the weight and bias parameters of the forward GRU, σ is the sigmoid function, and ⊙ is the element-wise multiplication.
[0127] The update formula for the reverse GRU is:
[0128]
[0129] where are the update gate, reset gate, and hidden state of the reverse GRU at time step t, respectively, are the weight and bias parameters of the reverse GRU.
[0130] Furthermore, in step seven, the linear transformation output represents:
[0131] Z = GW + b; (17)
[0132] In the formula, W is the weight matrix and b is the bias vector;
[0133] Probability distribution:
[0134]
[0135] The fault type with the maximum predicted probability:
[0136]
[0137] In the formula, P i is the probability of the i-th fault type, and A is the number of fault types.
[0138] In step seven, the fault diagnosis system includes a feature input module, a linear transformation unit, a probability conversion module, and a classification decision module. The probability conversion module is a tool that converts input data into a probability distribution. Its core goal is to map the original data feature values to probability values, making them satisfy the properties of the probability distribution, non-negative and summing to 1. In this embodiment, the Sigmiod function is used for binary classification, and the output value is the probability, and the range of the output value is 0 - 1. The Sigmiod function formula is:
[0139]
[0140] Perform threshold filtering on the probability distribution P (for example, the category with a probability value < 0.1 is regarded as invalid), and combine the temporal context information to smooth the classification result with a sliding window. Sliding window smoothing is a commonly used signal processing technique for smoothing temporal data or sequence data to remove noise, suppress outliers, or extract trend information. Its core idea is to slide a fixed-size window over the data and perform mean smoothing on the data within the window to generate a smoothed output sequence.
Claims
1. A fault detection method based on VMD-based Transformer model and BiGRU, characterized in that: The following steps are involved: Step 1: Data collection: Collect voltage and current signal data of the power grid to be tested; Step 2: Data preprocessing: Preprocessing the collected voltage and current signal data to obtain processed signal data; Step 3: Perform variational mode decomposition (VMD): Perform variational mode decomposition on the signal data after preprocessing in step 2, and obtain modal components; Step 4: Feature extraction and encoding: Perform feature extraction on each modal component obtained by VMD decomposition in step 3, perform encoding processing, and obtain a feature vector; Step 5: Transformer model processing: The feature vector in step 4 is processed using the Transformer model to obtain high-dimensional features; Step 6: Bidirectional Gated Recurrent Unit (BiGRU) processing: Process the high-dimensional features in step 5 with BiGRU and obtain the feature vector output by bidirectional GRU; Step 7, fault classification: The feature vector output in step 6 is classified using a softmax classifier and the classification result is obtained; Step 8: Fault detection and early warning: According to the classification result outputted in step seven, it is determined whether there is a fault in the current voltage and current, and an alarm is issued when a fault occurs.
2. According to claim 1, a fault detection method combining a VMD-based Transformer model and BiGRU, characterized in that: The step 1 specifically includes: selecting a suitable capacitive voltage transformer (CVT) and a protective current transformer (CT), installing and connecting these transformers to a detection module, transmitting signals to an information processing module via a wireless transmission module, and recording voltage and current signal data.
3. According to claim 2, a fault detection method combining a VMD-based Transformer model and BiGRU, characterized in that: The step 2 specifically includes: amplifying and denoising the collected voltage and current signals, then normalizing the amplitude of the signals, scaling the amplitude of the signals to a specific range, and finally cleaning the data, processing missing values and outliers, and obtaining signal data x(t).
4. According to claim 3, a fault detection method combining a VMD-based Transformer model and BiGRU is characterized in that: The step three specifically includes: inputting the preprocessed voltage and current signal x(t) into the VMD algorithm to solve the variational model; Specific solution process: First, for the modal components and linear frequency Perform initialization processing and use the original signal data x(t) as the initial modal component; then, for each modal component u k (t), given the other modal components and the center frequency, it is updated by solving the following optimization problem: in, represents the inverse Fourier transform, and They represent x(t) and Fourier transform, α is the bandwidth parameter of the modal component, for the center frequency ω of each modal component k , given the modal components, is updated by solving the following optimization problem: Repeat the above update steps until the convergence condition is met, and finally decompose x(t) into K modal components u k (t).
5. According to claim 4, a fault detection method combining a VMD-based Transformer model and BiGRU, characterized in that: The step 4 specifically includes: extracting features from each modal component obtained by VMD decomposition, and the extracted features include the energy E of the modal component. k , mean M k , Variance V k , and the time-frequency features F obtained by wavelet transform method k (τ, ω), the statistical features and time-frequency features are normalized and encoded so that their range is between [0, 1]. The normalization formulas are: Among them, E max , E min , M max , M min , V max , V min , F k,max and F k,min are the maximum and minimum values of each modulus feature respectively. The normalized statistical and time-frequency features are combined into a feature vector N with a sequence length of n and a feature dimension of d.
6. According to claim 5, a fault detection method combining a VMD-based Transformer model and BiGRU, characterized in that: The step five specifically includes: inputting the feature vector N into the Transformer model, the Transformer model performs weighted summation on the feature vector based on the self-attention mechanism to obtain a weighted feature representation T, then performing a nonlinear transformation on the weighted feature through a feedforward neural network to obtain a feature S, calculating the outputs T′ and S′ of the residual connection and layer normalization, and using the output S′ of the feedforward neural network as a high-dimensional feature representation.
7. The fault detection method combining a VMD-based Transformer model and BiGRU according to claim 6, characterized in that: The step six specifically includes: inputting the feature representation S′ output by the Transformer model into the BiGRU network, calculating from the start position to the end position of the sequence according to the update formula of the forward GRU, calculating from the end position to the start position of the sequence according to the update formula of the reverse GRU, concatenating the hidden states of the forward and reverse GRUs at each time step, and obtaining the output G of the bidirectional GRU t .
8. The fault detection method combining a VMD-based Transformer model and BiGRU according to claim 7, characterized in that: The step 7 specifically includes: converting the feature vector G output by the BiGRU network t Input to the softmax classifier, the classifier performs linear transformation based on the input feature vector to obtain Z, and converts Z into probability distribution P through the softmax function. According to the probability distribution P, the fault type with the highest probability is selected. as the final classification result.
9. The fault detection method combining a VMD-based Transformer model and BiGRU according to claim 8, characterized in that: The step eight specifically includes: according to the output result of the fault classifier Determine whether there is a fault in the current voltage and current signal; Among the fault types, there is a category called "normal state" with index i0. If Once a fault is detected, the system immediately generates a warning message and sounds an alarm.
10. The fault detection method combining a VMD-based Transformer model and BiGRU according to claim 9, characterized in that: In step 3, the variational model of VMD is expressed as: wherein, u k (t)=A k (t)cos(φ k (t)); A k (t) is the instantaneous amplitude of the modal component, φ k (t) is the instantaneous phase, δ(t) is the Dirac function, * represents the convolution operation, ω k is the modal component u k (t) is the center frequency, λ is the Lagrange multiplier that balances the reconstruction error and bandwidth penalty; The convergence of VMD decomposition is determined by monitoring the change in the modal components. When the change in the modal components is less than a preset threshold, the VMD decomposition is considered to have converged. Specifically, the Euclidean distance between the modal components of two adjacent iterations is calculated: when When it is less than the threshold ε, the iteration stops; In step 4, the energy, mean, variance and time-frequency features are: Where T is the total duration of the signal, ψ(τ, ω, t) is the wavelet function, τ is the time shift parameter, ω is the frequency parameter, and * represents the complex conjugate; In step 5, the multi-head self-attention decomposes the input feature vector N into h heads, each head corresponds to a self-attention mechanism, and the output features of the weighted multi-head self-attention are: T=Concat(head1,head2,…,head h )W O ;(13) In the formula, head i is the self-attention output of the i-th head, W o is the output weight matrix; Features obtained by nonlinear transformation: S=FFN(T)=max(0,TW 1 +b 1 )W 2 +b 2 ;(14) Where W 1 , W 2 , b 2 , b 2 are the weight and bias parameters of the feedforward neural network; Output features of residual connection and layer normalization: T′=LayerNorm(N+T); (15) S′=LayerNorm(T′+S); (16) Where LayerNorm is the layer normalization function; In step 7 above, the linear transformation output is represented as: Z=GW+b;(17) Where W is the weight matrix and b is the bias vector; Probability distribution: The fault type with the highest predicted probability: Where P i is the probability of the i-th fault type, and A is the number of fault types.
Citation Information
Cited By
Self-adaptive frequency spectrum monitoring and interference suppression method for railway power transformer
CN120832573A
Adaptive spectrum monitoring and interference suppression method for railway power transformer
CN120832573B
Marine autonomous surface ship power system reliability real-time evaluation method based on diagnosis input
CN121389313A
Model training method, fault diagnosis method, device, equipment, medium and product
CN121614863A