Distribution network single-phase earth fault prediction method based on LSTM-Transform neural network

By applying the LSTM-Transformer neural network in the distribution network, combining time-frequency characteristics and frequency domain characteristics, the problems of low prediction accuracy and poor real-time performance of single-phase ground faults are solved, and higher prediction accuracy and robustness are achieved.

CN120178095APending Publication Date: 2025-06-20STATE GRID HENAN ELECTRIC POWER COMPANY ANYANG POWER SUPPLY +1

Patent Information

Application Number
CN202510261405.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has low prediction accuracy and poor real-time performance in the distribution network, so it is impossible to effectively handle timing data.

Method used

The single-phase grounding fault prediction method of distribution network based on LSTM-Transformer neural network is adopted. By combining the data in the historical fault database, local time-frequency characteristics and global frequency domain characteristics are extracted, and the LSTM network captures timing dependencies, and the Transformer network captures global dependencies to perform fault prediction.

Benefits of technology

It improves the accuracy and real-timeness of fault prediction, enhances the robustness and noise resistance of the model, and can more accurately identify single-phase grounding fault characteristics, reducing the delay in fault processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178095A_ABST
    Figure CN120178095A_ABST
Patent Text Reader

Abstract

The invention provides a distribution network single-phase earth fault prediction method based on an LSTM-Transformer neural network. The method comprises the steps that a big data set is established according to distribution network single-phase earth fault information in an existing historical fault database; preprocessing the big data set, and dividing the big data set into a training set, a verification set and a test set according to a time sequence; extracting a local time-frequency feature and a global frequency-domain feature in the sample data, fusing the local time-frequency feature and the global frequency-domain feature, inputting the fused local time-frequency feature and global frequency-domain feature into an LSTM-Transform neural network, and outputting a predicted fault result; defining a loss function, and optimizing the loss function by using a composite optimizer; inputting the training set into the established LSTM-Transform neural network for training, verifying the trained LSTM-Transform neural network by using the verification set, and storing model parameters to obtain a fault prediction model; and inputting sample data in the test set into the fault prediction model to obtain a prediction result of the single-phase earth fault of the power distribution network. According to the invention, the real-time performance and the accuracy of single-phase earth fault prediction can be improved, so that the reliability and the stability of a power distribution network are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network fault monitoring, and particularly to a method for predicting single-phase grounding faults in a distribution network based on an LSTM-Transformer neural network. Background Art

[0002] With the continuous development of modern power systems, the degree of automation and intelligence of distribution networks has gradually increased. In power distribution networks, fault detection and diagnosis are important links to ensure the stability and reliability of the power grid. Especially for single-phase grounding faults, such faults usually occur in low-voltage distribution networks and have a greater impact on the power system. Due to its unclear fault characteristics, traditional fault detection methods often rely on manual experience or simple threshold judgment, and these methods have certain deficiencies in terms of real-time performance and accuracy. With the rapid development of deep learning technology, data-driven intelligent fault prediction methods have gradually become a research hotspot, especially in the early prediction of single-phase grounding faults.

[0003] A single-phase grounding fault in a distribution network refers to the contact of a phase wire in the power system with the ground, resulting in current flowing through the grounding path. Compared with other types of faults, single-phase grounding faults have the following characteristics:

[0004] 1. Small fault current; The fault current of a single-phase grounding fault is usually small and difficult to be accurately detected by traditional protection devices. This makes it difficult to timely identify single-phase grounding faults, thus delaying fault handling and affecting the reliability of power supply.

[0005] 2. Difficult to detect in the initial stage; The current of a single-phase grounding fault usually shows asymmetric changes, and after the fault occurs, there may be long-term continuous current fluctuations. Due to the insufficient ability of the system to identify these small changes, it is often difficult to accurately predict in the initial stage.

[0006] 3. Rapid fault development; If not repaired in time, a single-phase grounding fault may develop into other more serious fault types, such as short-circuit faults, which may lead to large-scale power outages and even damage to power equipment.

[0007] Traditional fault detection methods are mainly based on the threshold detection of electrical parameters such as current and voltage. Generally, when the value of current or voltage exceeds the set threshold, it is considered that a fault has occurred. This method has many disadvantages, poor adaptability to complex fault patterns, difficult to achieve precise early warning, and sensitive to noise.

[0008] As a neural network model specialized in processing time series data, the Long Short-Term Memory network (LSTM) has become an important technical means for distribution network fault prediction due to its advantages in dealing with long-term dependencies and time series characteristics. LSTM (Long Short-Term Memory) is an improvement of the Recurrent Neural Network (RNN) proposed by Hochreiter and Schmidhuber in 1997. Compared with traditional RNNs, LSTM can effectively avoid the problems of gradient vanishing and gradient explosion and can maintain long-term dependencies. LSTM is particularly suitable for learning tasks of time series data in power systems, such as time series characteristics like current and voltage fluctuations. In the prediction of single-phase grounding faults in distribution networks, LSTM can identify potential fault characteristics by learning the fluctuation patterns of historical current and voltage and issue early warnings. Especially in the initial stage of a fault, LSTM can make predictions based on past input data and detect weak current fluctuations in a timely manner, thereby improving the accuracy of fault prediction.

[0009] Transformer is a deep learning model based on parallel computing with self-attention mechanism (Self-Attention), widely used in natural language processing and time series data analysis. Compared with traditional RNNs and LSTMs, Transformer can be more efficient in processing long time series data because it does not rely on a recursive structure but captures dependencies in the sequence by calculating global information. This parallel computing ability makes Transformer have higher computing efficiency and accuracy when dealing with large-scale data. In the prediction of distribution network faults, Transformer can capture global dependencies in the input data through the self-attention mechanism and model complex fault patterns. In the processing of multi-channel data (such as current, voltage, frequency, etc.), Transformer can make full use of the correlation between each channel to more accurately identify fault characteristics.

[0010] Single-phase grounding faults in distribution networks are difficult to effectively handle with traditional fault prediction methods due to their small current and weak changes.

[0011] The invention patent with the application number 202410686058.X discloses a transformer fault detection method based on the fusion of mechanism model and data model, including the following steps: collecting the actual vibration data D of key points of on-site transformers and dividing D proportionally train and D testDataset; Obtain the simulated vibration data M output by the simulation mechanism model; Perform data-mechanism fusion to obtain the fusion data F and format processing; Based on the multi-scale convolutional neural network MSCNN, long short-term memory network LSTM, and attention mechanism ATTENTION, establish a transformer fault detection network, with F and Dtrain as inputs and the output being the transformer detection status; Train the transformer fault detection network and obtain the optimized network structure; Use the trained transformer fault detection network to diagnose the actual vibration data Dtest and output the working status. The above invention adopts a diagnostic model of a multi-scale convolutional neural network, a long short-term memory network, and an attention mechanism, which can accurately extract transformer fault features and perform fault identification, and has good practical value. However, in the above invention, although MSCNN can extract multi-scale local features, it relies on the series integration of LSTM and the attention mechanism to integrate global temporal information, with high model complexity and possible information loss; This method has a certain dependence on the acquisition and accuracy of the mechanism model. If the mechanism model has deviations or is incomplete, it may affect the performance of fault detection, and actual vibration data needs to be collected. If the data quality is poor or there is a lack of sufficient sample data, it may also affect the effect of the model. Summary of the Invention

[0012] Aiming at the technical problems of low accuracy, poor real-time performance, and inability to effectively process time series data in the existing fault prediction methods for distribution networks, the present invention proposes a method for predicting single-phase grounding faults in distribution networks based on the LSTM-Transformer neural network, which can not only effectively capture long-term and short-term time series dependencies, improve prediction accuracy, but also process multi-channel input data, enhancing the robustness and anti-noise ability of the entire model.

[0013] To achieve the above object, the technical solution of the present invention is realized as follows: A method for predicting single-phase grounding faults in distribution networks based on the LSTM-Transformer neural network, the steps are as follows:

[0014] S1: Establish a large dataset according to the single-phase grounding fault information of the distribution network in the existing historical fault database;

[0015] S2: Preprocess the sample data in the large dataset, and divide the preprocessed large dataset into a training set, a validation set, and a test set according to the time sequence;

[0016] S3: Extract the local time-frequency features and global frequency domain features in the sample data of the preprocessed large dataset, build an LSTM-Transformer neural network, fuse the local time-frequency features and global frequency domain features and input them into the LSTM-Transformer neural network, and output the predicted fault results;

[0017] S4: Define the loss function and optimize the LSTM-Transformer neural network using a composite optimizer;

[0018] S5: Input the sample data of the training set into the constructed LSTM-Transformer neural network for training, use the sample data in the validation set to verify the trained LSTM-Transformer neural network, and save the model parameters to obtain the fault prediction model;

[0019] S6: Input the sample data in the test set into the trained fault prediction model to obtain the prediction result of the single-phase grounding fault in the distribution network.

[0020] Preferably, the large dataset includes the multi-modal data fusion of electrical parameters, environmental data, equipment status data, load fluctuation data, and historical fault data of single-phase grounding in the distribution network; the large dataset comes from the historical fault dataset of the power company, the distribution network monitoring system data, the power equipment monitoring and status data, the meteorological data, the load fluctuation data, and the publicly available power fault dataset;

[0021] The preprocessing includes: checking whether there are missing values, outliers, and duplicate data in the data, filling the missing values using the linear interpolation method, correcting the outliers, and deleting the duplicate data; performing Z-score standardization and Min-Max normalization on the corrected sample data;

[0022] Use isnull() in the Pandas library in Python to detect missing values, use Z-score to detect outliers, and use the drop_duplicates() method in the Pandas library in Python to directly delete duplicate data.

[0023] Preferably, the method for extracting local time-frequency features is: use adaptive wavelet transform to extract the instantaneous changes and time-frequency local features of the signal, and decompose the time-series data in the large dataset into the low-frequency part and high-frequency part of multiple frequency bands;

[0024] The method for extracting global frequency-domain features is: perform FFT on the low-frequency part and high-frequency part after adaptive wavelet transform respectively to obtain global frequency-domain features;

[0025] The method for fusing local time-frequency features and global frequency-domain features is: fuse local time-frequency features and global frequency-domain features into a single feature through statistical operations;

[0026] Using an LSTM network as a feature extractor to capture local temporal dependencies and short-term memory in the fused features, extract features of local temporal dependencies, short-term memory, and long-term memory, and obtain a hidden state sequence; using the multi-head self-attention mechanism of the Transformer network to capture the long-range dependency relationships between time steps in the hidden state sequence, obtain output features, and the output layer maps the output features to the prediction space through a linear transformation, and uses the Softmax activation function for the classification task to obtain the predicted fault results.

[0027] Preferably, the adaptive wavelet transform is as follows:

[0028]

[0029] where x(t) is the input signal, a is the scale factor, and b is the translation factor; x low (n), x high (n) represent the low-frequency signal and the high-frequency signal of the time-domain signal respectively, ψ opt (t) is the optimal mother wavelet adaptively selected according to the signal characteristics; ψ low (t) is the low-frequency wavelet basis function, ψ high (t) is the high-frequency wavelet basis function, W x (a, b, ψ) represents the transformation response of the input signal x(t) and the wavelet basis function ψ opt (t) at the scale a and the translation factor b;

[0030] The method for performing FFT is as follows:

[0031]

[0032] where X low (k) is the k-th frequency component of the frequency-domain representation of the low-frequency signal, X high (k) is the k-th frequency component of the frequency-domain representation of the high-frequency signal, and N is the total length of the signal; is the complex exponential term;

[0033] The fused features are:

[0034] X t = [Max(X low ), Min(X high ), μ low , σ low , μ high , σ high

[0035] where Max(X low ) is the maximum value of the low-frequency part X low ), Min(X​high ) is the minimum value of the high-frequency part X high , μ low , σ low are the mean and standard deviation of the low-frequency part respectively, μ high , σ high are the mean and standard deviation of the high-frequency part respectively;

[0036] In the unit structure of the LSTM network, there are three gating mechanisms: the forget gate, the input gate, and the output gate. The forget gate calculates the output of the forget gate based on the currently fused feature X t input and the hidden state h at the previous moment t-1 t-1 through the forget weight matrices W f ' and W f to control which long-term memories in the LSTM network can be forgotten; the output of the input gate is obtained by processing the fused feature Xt and the hidden state ht-1 at the previous moment t-1 with the weight matrix together with the bias term. The input gate and the candidate cell state together determine the update of the memory state C t at the current moment. The input gate controls the writing ratio of the candidate cell state into the cell state C t at the current moment; the update of the memory state C t at the current moment is controlled by the outputs of the input gate and the forget gate to update the long-term memory cells in the cell state; the output state o t of the output gate is obtained by adding the fused feature Xt and the hidden state ht-1 at the previous moment t-1 processed by the output gate weight matrix with the bias term. The hidden state h t at time step t is: h t = o t *tanh(C t ), where tanh is the hyperbolic tangent function.

[0037] Preferably, at each time step t, the LSTM network outputs a hidden state h t ; through the forward propagation of all time steps, the LSTM network generates a hidden state sequence H = [h1, h2,..., h t to record the hidden state at each time step; the hidden state sequence H = [h1, h2,..., ht] output by the LSTM network is position-encoded to obtain the position encodings of the even and odd positions;

[0038] The hidden state h t output by the LSTM network and the corresponding position encoding PE(t) are added together to obtain the information with position encoding: H encoded= H + P; where P = [PE(1), PE(2),..., PE(t)] is the position encoding corresponding to the time step.

[0039] Preferably, the Transformer network uses the multi-head self-attention mechanism to capture the information H with position encoding encoded the relationships between time steps in, and through the processing of the feed-forward neural network layer, further strengthen the feature expression ability, calculate the dot product of queries, keys, and values and normalize it to obtain the attention weights, calculate the output of each attention head through the attention weights, concatenate the outputs of all attention heads and multiply by the linear transformation matrix of the learned output to obtain the total output of the multi-head self-attention; the feed-forward neural network uses two fully connected layers to process the total output of the multi-head self-attention to obtain the output of the feed-forward neural network, and the total output of the multi-head self-attention of the current layer and the output of the feed-forward neural network are connected by a residual connection to obtain the final output features.

[0040] Preferably, the structure of the Transformer network has four parts: input, encoder, decoder, and output. Among them, the input includes a source text embedding layer, a position encoder, and a position encoder. The source text embedding layer converts the vocabulary digital representation in the hidden state sequence H = [h1, h2,..., h t into a vector representation to capture the relationships between vocabulary; the position encoder generates position vectors for each position of the hidden state sequence H = [h1, h2,..., h t , the encoder is stacked by multiple encoder layers to capture more complex and abstract features; each encoder consists of two sub-layer connection structures. The first sub-layer is a multi-head self-attention sub-layer, and the second sub-layer is a feed-forward fully connected sub-layer. After each sub-layer, there is a normalization layer and a residual connection, calculate the dot product of queries, keys, and values and normalize it to obtain the attention weights;

[0041] The decoder is stacked by multiple decoder layers. Each decoder layer consists of three sub-layer connection structures: the first sub-layer is a masked multi-head self-attention sub-layer, the second sub-layer is a multi-head self-attention sub-layer, and the third sub-layer is a feed-forward fully connected sub-layer. The output of each self-attention layer in the multi-head self-attention sub-layer will be sent to the next self-attention layer. After each sub-layer, there is a normalization layer and a residual connection; the output consists of a linear layer that converts the vector output by the decoder into the final output dimension and a Softmax layer that converts the output of the linear layer into a probability to obtain the prediction distribution.

[0042] Preferably, the total output of the multi-head self-attention: A Multi-Head = Concat(A1, A2,…, A h )W O; where, W O is a linear transformation matrix of the learned output, A1, A2, …, A h are the outputs of h attention heads, and Concat means concatenating the outputs of all attention heads together;

[0043] The output of the feed-forward neural network is z FFN = max(0, A Multi-Head W1 + b1)W2 + b2, where, W1 and W2 are the weight matrices of the fully connected layers, and b1 and b2 are the bias terms;

[0044] The Transformer network uses residual connections to enhance the gradient flow during training. For the total output A Multi-Head of the multi-head self-attention of the current layer and the output z FFN of the feed-forward neural network, a residual connection is performed to obtain the feature z res = A Multi-Head + z FFN ; The feature z res is subjected to layer normalization to obtain the final output feature: z norm = LayerNorm(z res );

[0045] The method for the output layer to use the Softmax activation function for the classification task is: p = Softmax(W out ·z norm + b out ); where, W out and b out are the weights and biases of the output layer, and p is the final fault prediction result.

[0046] Preferably, the method for optimizing the LSTM-Transformer neural network using a composite optimizer in step S4 is: optimizing the LSTM-Transformer neural network through the Adam optimizer to make it converge quickly; when the loss function is stable, switch to the SGD optimizer for fine-tuning and convergence of the LSTM-Transformer neural network.

[0047] Preferably, the loss function in step S4 is:

[0048] E L = -[ylog(p) + (1 - y)log(1 - p)]

[0049] Among them, y is the true label, y ∈ {0, 1}, y = 1 indicates a fault, and y = 0 indicates no fault. p is the predicted probability output by the LSTM-Transformer neural network, representing the probability that the sample data belongs to a fault, and p ∈ [0, 1].

[0050] The method for the SGD optimizer to perform fine-tuning and convergence is as follows: Among them, η is the learning rate. is the gradient of the loss function E(θ t ) with respect to the current parameter θ t corresponding to.

[0051] Compared with the existing methods, the beneficial effects of the present invention are as follows: The LSTM-Transformer hybrid model of the present invention combines the advantages of the LSTM and Transformer models. The LSTM part can effectively capture the long-term dependence of time series, and the Transformer part can enhance the global information processing ability of the model through the self-attention mechanism. The combination of the two can improve the prediction accuracy, enhance the robustness of the model, and optimize the calculation efficiency, becoming an effective method for dealing with the problem of single-phase grounding fault prediction in the distribution network. The LSTM-Transformer neural network model is significantly superior to the mechanism model fusion method in terms of long-time series modeling, time-frequency feature decoupling, real-time performance, etc., and is especially suitable for the distribution network fault prediction scenario with high noise and multi-modal.

[0052] The present invention is based on a hybrid model of LSTM and Transformer, which combines the advantages of the two neural networks in processing time series data and can perform more accurate fault prediction and early warning; the present invention uses the LSTM-Transformer hybrid model to process the historical data of distribution network faults, and has a significant improvement in the ability to predict distribution network faults in complex and diverse power system fault detection tasks; it can realize early warning of single-phase grounding faults in the distribution network and further avoid the occurrence of major distribution network fault accidents. The present invention can improve the real-time performance and accuracy of single-phase grounding fault prediction, thereby enhancing the reliability and stability of the distribution network and promoting the further development of the smart grid. The parallel computing ability and end-to-end learning characteristics of the entire hybrid model of the present invention give it significant advantages in the tasks of distribution network intelligent monitoring and fault prediction, making up for some limitations of traditional methods and single deep learning models in fault prediction. The present invention significantly improves the accuracy, response speed and system reliability of distribution network fault prediction, ensuring that the power system operates more intelligently and efficiently. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0054] Figure 1 This is the flowchart of the present invention.

[0055] Figure 2 This is a schematic diagram of the calculation of the memory unit of LSTM in the distribution network fault prediction of the present invention.

[0056] Figure 3 This is the principle diagram of the multi-head self-attention mechanism of Transformer of the present invention

[0057] Figure 4 This is the architecture diagram of the Transformer model of the present invention. Detailed implementation manners

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0059] As Figure 1 shown, the present invention proposes a method for predicting single-phase grounding faults in a distribution network based on an LSTM-Transformer neural network, including the following steps:

[0060] S1: Establish a large dataset according to the distribution network fault information in the existing historical fault database.

[0061] The established large dataset includes the multimodal data fusion of electrical parameters (voltage, current, power, frequency) of single-phase grounding in the distribution network, environmental factors (temperature, humidity, meteorological data), equipment status data (equipment temperature, insulation status, noise monitoring), load fluctuations, and historical fault data, providing rich training samples for the hybrid model, improving the accuracy, robustness, and real-time performance of fault prediction. With the support of a large amount of historical data, the hybrid model can continuously learn and adapt to the complex operating environment of the actual power grid, improving the fault detection and early warning capabilities of the distribution network, thereby enhancing the stability and reliability of the power system.

[0062] The large dataset comes from the historical fault dataset of power companies, the data collected in real time by the distribution network monitoring system (such as the SCADA system and EMS system of the distribution network), the monitoring and status data of power equipment (equipment sensors or intelligent power equipment management systems), meteorological data (historical meteorological data obtained from meteorological bureaus or meteorological data platforms), load fluctuation data (intelligent meters or power dispatching systems of the distribution network), and publicly available power fault datasets (universities, research institutions, and open-source communities for fault prediction research will provide publicly available datasets).

[0063] S2: Preprocess the sample data in the large dataset, and divide the preprocessed large dataset into a training set, a validation set, and a test set in chronological order.

[0064] The preprocessing in step S2 includes the following sub-steps:

[0065] S21: Check whether the data has missing values, outliers, and duplicate data, use the linear interpolation method to fill in the missing values and correct the outliers; delete the duplicate data. Use isnull() in the Pandas library in Python to detect missing values. Use Z-score to detect outliers to ensure data quality. Z-score is a standardized index that measures how much a data point deviates from the mean. If the Z-score of a data point is higher than a certain threshold, it is considered an outlier. Use the drop_duplicates() method in the Pandas library in Python to directly delete duplicate data.

[0066] S22: To eliminate the differences in data feature scales and improve the efficiency and convergence speed of model training, perform Z-score standardization and Min-Max normalization on the sample data corrected in step S21. When performing distribution network fault prediction, electrical parameters such as current, voltage, and load are usually used to determine whether a system failure has occurred. Z-score standardization processes the offset and scale differences of the data, avoiding the deviation of model training caused by the dimensional differences of different features, and ensuring that these parameters have similar scales. Min-Max normalization ensures that the values of each feature are within the same interval, usually [0,1], which is suitable for the training process of neural networks. For neural networks using the Sigmoid activation function, it avoids the problem of gradient disappearance or explosion. Thus, the LSTM network and the Transformer network can simultaneously focus on all features, rather than the dominant role of a certain feature in model training.

[0067] S23: Divide the data into a training set (the first 70%), a validation set (the next 15%), and a test set (the last 15%) according to time. This ensures the scientific nature of the training process, avoids data leakage, simulates the time-series data stream in actual applications, helps improve the generalization ability, stability, and reliability of the model. This data division method can help the model adapt to the long-term changes in the power system, maintain efficient and accurate fault prediction capabilities, and lay a solid foundation for the application of the model in actual scenarios. Divide the data into a training set, a validation set, and a test set following the chronological order, where the three sets represent the past, the future, and the more distant future respectively, which conforms to the definition of a time-series data stream: a data stream that is generated, updated, and gradually enters the system in chronological order.

[0068] S3: Extract the local time-frequency features and global frequency-domain features from the sample data, build an LSTM-Transformer neural network, fuse the local time-frequency features and global frequency-domain features and then input them into the LSTM-Transformer neural network, and the LSTM-Transformer neural network outputs the fault result.

[0069] Use adaptive wavelets to extract the local time-frequency features of the time series in the input sample data, use FFT to analyze the local time-frequency features to extract the global frequency-domain features, merge the extracted local time-frequency features and global time-frequency features as the input data of the LSTM network, the output of the LSTM network is used as the input of the Transformer network, and the output of the Transformer network determines whether there is a fault through the Sigmoid function. The specific implementation method includes the following sub-steps:

[0070] S31: First, use the Adaptive-Wavelet-Transform to extract the instantaneous changes and local time-frequency features of the signal, decompose the time-series data in the preprocessed large dataset into the low-frequency part and high-frequency part of multiple frequency bands, and reveal the local changes of the signal (such as sudden faults, short circuits, etc.).

[0071] The traditional wavelet is defined as:

[0072]

[0073] where x(t) is the input signal; ψ is the mother wavelet; a is the scale factor, which controls the compression degree (corresponding frequency) of the wavelet, that is, the scaling factor; b is the translation factor, which controls the time positioning of the wavelet, that is, the translation factor.

[0074] The Adaptive Wavelet Transform is an extension of the traditional Wavelet Transform. The selection of the mother wavelet ψ and the scaling factor a is dynamically changed and adjusted adaptively according to the characteristics of the input signal (such as frequency, instantaneous frequency, etc.).

[0075] The adaptive wavelet is defined as:

[0076]

[0077] where x low (n), x high (n) represent the low-frequency signal and the high-frequency signal of the time-domain signal respectively. ψ opt (t) is the optimal mother wavelet adaptively selected according to the signal characteristics; ψ low (t) is the low-frequency wavelet basis function, ψ high (t) is the high-frequency wavelet basis function. W x (a, b, ψ) represents the transformation response of the input signal x(t) and the wavelet basis function ψ(t) at scale a and position b. The low-frequency and high-frequency wavelet basis functions complement each other and represent the local characteristics of the signal at different scales. a is the scaling factor; large scales such as a ≈ 10 to a ≈ 1000 are applicable to the low-frequency components in the signal; small scales such as a ≈ 0.1 to a ≈ 10 are applicable to the high-frequency components in the signal. b is the translation factor. A larger step size such as every 5 sampling points is selected for the low frequency, and a smaller step size such as every 1 sampling point is selected for the high frequency.

[0078] Using the Adaptive Wavelet Transform (AWT) to extract the instantaneous changes of the signal and obtain local time-frequency characteristics can effectively reveal the local changes in the signal, especially sudden faults and short-circuit phenomena. By performing time-frequency decomposition on the signal, AWT provides richer and more accurate feature information for the subsequent LSTM-Transformer neural network hybrid model, enabling the fault prediction model to better identify fault patterns and improve prediction accuracy and robustness.

[0079] S32: Then, perform FFT on the low-frequency part and the high-frequency part after the adaptive wavelet transform respectively to extract global frequency-domain characteristics.

[0080] The FFT is defined as:

[0081]

[0082] where X low (k) is the k-th frequency component of the frequency-domain representation of the low-frequency signal, X high(k) is the k-th frequency component of the frequency-domain representation of the high-frequency signal, N is the total length (number of sampling points) of the signal; k is the frequency index, representing the frequency corresponding to each point in the frequency domain; is a complex exponential term, representing the oscillation of each frequency component.

[0083] Perform fast Fourier transform (FFT) on the low-frequency part and high-frequency part after adaptive wavelet transform (AWT) to extract global frequency-domain features, which helps to understand the behavior of the signal (referring to the previous electrical signal including voltage, current, frequency, and power) at different frequency levels. These frequency-domain features can help the model capture long-range changes (such as load changes, trend faults) and short-range changes (such as instantaneous faults, emergencies) in the signal, providing richer and more accurate input features for the subsequent LSTM-Transformer neural network model.

[0084] S33: Fuse the local time-frequency features and global frequency-domain features into a single feature through statistical operations. Defined as:

[0085] X t =[Max(X low ),Min(X high ),μ low ,σ low ,μ high ,σ high

[0086] where, Max(X low ) is the maximum value of the low-frequency part, Min(X high ) is the minimum value of the high-frequency part, μ low ,σ low are the mean and standard deviation of the low-frequency part respectively, μ high ,σ high are the mean and standard deviation of the high-frequency part. The mean and standard deviation are mainly selected for the low-frequency part to capture the long-term trend and stability of the signal, helping the model understand the global behavior of the signal, such as the long-term changes or periodic fluctuations of the power grid load. The maximum and minimum values are selected for the high-frequency part to capture the instantaneous changes or short-term fluctuations in the signal. These features help the hybrid model identify fault signals or sudden fluctuations, especially single-phase ground faults. The high-frequency mean provides the central tendency of the signal in the high-frequency part, which can help identify continuous fluctuations or slight changes in the equipment. The high-frequency standard deviation can quantify the fluctuation amplitude of the signal, helping to reveal the instability and sudden changes in the signal, which is very helpful for fault detection. Especially when the equipment has a sudden fault or the system is abnormal, the standard deviation can reflect the imbalance of current and voltage.

[0087] ​Fusing multiple features into a single feature through statistical operations is a key step in improving the performance of the LSTM-Transformer neural network model. The main functions of this operation are to simplify the input data, reduce dimensions, improve the model training efficiency, enhance the representativeness of features, extract more representative comprehensive information; improve the generalization ability of the model and avoid overfitting; enhance the stability of features and reduce the sensitivity to noise and local fluctuations; improve the prediction accuracy, especially being able to capture key changes when dealing with fault signals. Through this feature fusion operation, the LSTM-Transformer neural network model can predict single-phase grounding faults in the distribution network more efficiently and accurately, and has good adaptability to different types of faults.

[0088] S34: Use LSTM as a feature extractor to capture feature X t Extract features of local temporal dependencies, short-term memory, and long-term memory by capturing local temporal dependencies and short-term memory in it.

[0089] Refer to Figure 2 As shown, in the LSTM cell structure, there are three gating mechanisms: the forget gate, the input gate, and the output gate.

[0090] Forget gate f t Controls which long-term memories in the LSTM network can be forgotten, defined as:

[0091] f t = σ(W f h t-1 + W f 'X t + b f )

[0092] Where σ is the Sigmoid activation function, W f and W f ' are forget weight matrices, acting on the current feature input X t and the hidden state h t-1 at time t-1 respectively. The current input feature X t and the previous hidden state h t-1 calculate the output f f of the forget gate through the weight matrices W f ' and W t , thereby controlling the retention degree of the previous moment's memory at the current moment. b f represents the bias vector of the forget gate. The bias term is usually initialized to zero or a small constant (such as 0.1), here it is taken as 0. During the training process, the bias term b f will be updated through gradient descent to learn how to optimize the decision of the forget gate. f tRepresents the output of the forget gate, which determines how much of the previous moment's memory should be retained, and the output value ranges from 0 to 1.

[0093] Input gate i t and the candidate cell state together determine the update of the cell state C t at the current moment, defined as:

[0094] i t = σ(W i h t-1 + W i 'X t + b i )

[0095]

[0096] where i t represents the output of the input gate, is the candidate cell state, representing the new memory at the current moment, which is the value obtained after processing the current input information. W i ' and W c ' are the weight matrices related to the current input feature X t , W i and W C are the weight matrices related to the hidden state h t-1 at the previous moment. b i is the bias term, used to adjust the calculation result of the input gate. The initial value is taken as 0.1. Tanh is the hyperbolic tangent activation function, which compresses the value of the candidate cell state to the range [-1, 1], helping to control the value range of the cell state and preventing gradient explosion. b c is the bias term used to adjust the calculation result of the candidate cell state.

[0097] where is the vector to be judged processed by the tanh activation function of the cell state C t at the current moment. It is determined whether to rely on the vector to be judged t for update by the input gate i : The input gate i t and the candidate cell state together determine the update of the cell state C t at the current moment. The input gate i t controls the writing ratio of the candidate cell state in the cell state C t at the current moment. W i and W i ' are the judgment matrices for the current time step, and W C and W c ' are the candidate value vector matrices for the current time step, ht-1 is the hidden state of the hidden layer, b i and b c are bias vectors. The initial values of the weight matrices are usually randomly initialized. During training, they are automatically adjusted according to gradient descent and backpropagation, and gradually updated to minimize the loss function and learn the optimal input gate behavior.

[0098] The update of the memory state C at the current time step is controlled by the input gate and the forget gate: t , updating the long-term memory cells in the cell state:

[0099]

[0100] where C t is the cell state at the current moment, C t-1 is the cell state at the previous moment. f t is the output of the forget gate, which determines the forgetting ratio of the cell state C t-1 at the previous moment, i t represents the output of the input gate, which determines the writing ratio of the current input information into the cell state, is the candidate cell state at the current moment, representing new memory.

[0101] Define the output gate:

[0102] o t = σ(W o h t-1 + W o 'X t + b o )

[0103] h t = o t * tanh(C t )

[0104] where o t is the output state, W o and W o ' are both output gate weight matrices, b o is the bias vector. h t is the hidden state at time t.

[0105] The tanh function is defined as:

[0106]

[0107] The Sigmoid activation function is defined as:

[0108]

[0109] Among them, x represents the independent variable.

[0110] S35: Use the Transformer network to effectively capture the long-range dependencies between each time step in the sequence.

[0111] At each time step t, the LSTM network outputs a hidden state h t ; Finally, after the forward propagation through all time steps, the LSTM network generates a hidden state sequence H = [h1, h2,..., h t , which records the hidden state at each time step. Since the Transformer network itself does not have the ability to perceive the sequence order, positional encoding is needed to provide the position of each time step.

[0112] Perform positional encoding on the hidden state sequence H = [h1, h2,..., ht] output by the LSTM network.

[0113]

[0114] Among them, t is the current position, i is the dimension index of the positional encoding, that is, the different dimensions of each positional encoding vector. d is the dimension of the positional encoding vector, which is equal to the input dimension. PE(t, 2i) and PE(t, 2i + 1) represent the encodings of the even and odd positions respectively.

[0115] Add the hidden state h t output by the LSTM network and the corresponding positional encoding PE(t) to obtain the information with positional encoding input to the Transformer network:

[0116] H encoded = H + P

[0117] Among them, H = [h1, h2,..., h t is the hidden state sequence output by the LSTM network, and P = [PE(1), PE(2),..., PE(t)] is the positional encoding corresponding to the time steps.

[0118] Use the multi-head self-attention mechanism to capture the relationships between each time step in the information H with positional encoding encoded ; After being processed by the feed-forward neural network layer, the feature expression ability is further enhanced, the dot products of the query, key, and value are calculated and normalized to obtain the attention weights.

[0119] Refer to Figure 3 , the self-attention mechanism of the Transformer network enables the features at each time step (i.e., the information H encodedEach position (in each time step) can be associated with the features of other time steps simultaneously and is independent of the order of time steps. This is a fundamental difference from recurrent networks such as LSTM or GRU, which process sequential data step by step. The self-attention mechanism enables the Transformer network to compute the dependencies between each time step and other time steps in parallel when processing sequences, obtaining a global representation for each position. Specifically, the Transformer network calculates queries Q, keys K, and values V for each input position, and then assigns a weight to each position by computing the similarity between the query and the key, so that each position can obtain information from other positions, realizing the interaction of global information.

[0120] Q = H encoded W Q , K = H encoded W K , V = H encoded W V

[0121] Among them, H encoded represents the information with positional encoding input to the encoder, and W Q , W K , W V are the weight matrices for queries, keys, and values (initialized randomly). According to the vector values of queries Q, keys K, and values V, calculate the dot product of the query, key, and value and normalize it to calculate the result of the dot product attention mechanism, as shown in the following formula:

[0122]

[0123] The multi-head self-attention mechanism divides queries Q, keys K, and values V into h subspaces, and each subspace calculates an independent attention weight through different transformations. Assuming there are h heads, and concatenating the outputs of all heads, the output A h of each head is:

[0124]

[0125] Among them, d k is the scaling factor (used to control the size of the dot product), T represents matrix transpose, K T represents the transpose of key K, and softmax is used to calculate the probability distribution of each row, representing the attention weights of different time steps to the current time step.

[0126] Obtain the total output of the multi-head self-attention:

[0127] A Multi-Head = Concat(A1, A2, …, A h )WO

[0128] Among them, W O is a linear transformation matrix of the learned output (initialized randomly), and Concat means concatenating the outputs of all heads. The total output A of the multi-head self-attention Multi-Head is input into the feed-forward neural network (FFN). The feed-forward neural network is a standard two-layer fully connected network, in the form as follows:

[0129] z FFN = max(0, zW1 + b1)W2 + b2

[0130] Among them, z is the output from the multi-head self-attention mechanism, that is, A Multi-Head , and W1, W2 are the weight matrices of the fully connected layers (initialized randomly), and b1, b2 are the bias terms.

[0131] The Transformer network uses residual connections to enhance the gradient flow during training. For the total output A Multi-Head of the current layer and the output z FFN of the feed-forward network, they perform a residual connection to obtain:

[0132] z res = A Multi-Head + z FFN

[0133] Then, this result is subjected to layer normalization to obtain the final layer output:

[0134] z norm = LayerNorm(z res )

[0135] The output layer maps the sequence output by the Transformer network to the prediction space through a linear transformation, and uses the Softmax activation function for the classification task (fault and non-fault).

[0136] p = Softmax(W out · z norm + b out )

[0137] Among them, W out and b out are the weights and biases of the output layer (with the same values as above), and p is the final fault prediction result.

[0138] In the encoder of the Transformer network, multiple self-attention layers are stacked together. The output of each self-attention layer is fed into the next self-attention layer. The self-attention mechanism of each self-attention layer can capture local and global dependencies at different levels.

[0139] Layer Normalization and Feed-Forward Network are common components in Transformer. Usually, a Feed-Forward Network follows each self-attention layer, and a Residual Connection is used to avoid the vanishing gradient problem.

[0140] See Figure 4 As shown, in the Transformer network structure, there are four parts: input, encoder, decoder, and output. The input part includes a source text embedding layer, a position encoder, and a positional encoder. The source text embedding layer converts the vocabulary numbers in the source text, which is the output hidden state sequence H = [h1, h2,..., h t into vector representations to capture the relationships between words. The position encoder generates position vectors for each position in the hidden state sequence H = [h1, h2,..., h t so that the model can understand the position information in the sequence. The encoder part is stacked by N (usually 6) encoder layers to capture more complex and abstract features, thereby improving the model's expressive ability, learning longer-range dependencies, enhancing the ability to model global information, improving the parallel computing ability, accelerating the training and inference processes, enhancing the learning ability of the deep network, and better handling complex tasks. The model can be flexibly extended by increasing or decreasing the number of layers to adapt to tasks of different scales. Each encoder layer consists of two sub-layer connection structures. The first sub-layer is a multi-head self-attention sub-layer, and the second sub-layer is a feed-forward fully connected sub-layer. A normalization layer and a residual connection are connected after each sub-layer. The multi-head self-attention mechanism is used to capture the relationships between time steps in the input sequence. After passing through the feed-forward neural network layer, the feature expression ability is further enhanced, and the dot products of queries, keys, and values are calculated and normalized to obtain attention weights.

[0141] The decoder part is stacked by N decoder layers. The stacking of the N layers in the decoder part can significantly improve the model's long-term sequence dependence modeling ability, multi-modal feature fusion ability, global and local pattern recognition ability, and enhance the robustness of the model under noise interference and environmental changes. Through the stacking of decoder layers, the model can more accurately capture fault features and time sequence patterns, improving the accuracy and practicality of fault prediction. Each decoder layer consists of three sub-layer connection structures: the first sub-layer is a masked multi-head self-attention sub-layer, which maintains causality, ensuring that the model only uses current and historical data to predict the future and avoiding future information leakage. It can perform parallel computing, improving the efficiency of the model in capturing long-range dependence relationships, especially suitable for processing long-time sequence data. It has diverse attention focuses, allowing the model to understand input data from multiple perspectives and enhancing the ability to capture complex fault patterns. It is more adaptable and can handle the dynamically changing environment in the power system, improving the ability to predict faults. Temporal consistency ensures that the time sequence is not violated during the prediction process, maintaining the rationality of the prediction. Through these mechanisms, the masked multi-head self-attention sub-layer can effectively improve the accuracy, robustness, and temporal stability of the single-phase grounding fault prediction in the distribution network. The second sub-layer is a multi-head self-attention sub-layer (from encoder to decoder), and this mechanism enables the model to more accurately predict single-phase grounding faults, ensuring that the precursors or patterns of fault occurrence can be efficiently and accurately captured in a complex power system. The third sub-layer is a feed-forward fully connected sub-layer, which plays multiple functions such as converting information, extracting features, and enhancing the expression ability in the decoder and is an indispensable key component in the Transformer architecture. A normalization layer and a residual connection are connected after each sub-layer. The output part consists of a linear layer (converting the vector output by the decoder into the final output dimension) and a Softmax layer (converting the output of the linear layer into a probability distribution for final prediction).

[0142] S4: Define the loss function and optimize the LSTM-Transformer neural network using a composite optimizer.

[0143] It includes the following sub-steps:

[0144] S41: Define the loss function, and its formula is:

[0145] E L = -[ylog(p)+(1 - y)log(1 - p)]

[0146] where E L represents the loss function, y is the true label, y ∈ {0, 1}, y = 1 represents a fault, y = 0 represents no fault, p is the predicted probability output by the model, representing the probability that the sample belongs to a fault, and p ∈ [0, 1].

[0147] S42: Optimize the LSTM-Transformer neural network using the Adam optimizer to enable it to converge quickly.

[0148] First, calculate the cross-entropy loss function E L for the gradient of the predicted probability p output by the LSTM-Transformer neural network:

[0149]

[0150] Then, calculate the gradients of the loss function with respect to W out and b out using the chain rule:

[0151]

[0152] First moment:

[0153] m t = β1m t-1 + (1 - β1)g t

[0154] Second moment:

[0155]

[0156] Bias correction:

[0157]

[0158] Update parameters:

[0159]

[0160] where m t is the first moment estimate at the current time step t, m t-1 represents the first moment estimate at the previous time step, β1 is the decay rate of the first moment estimate, taken as 0.9 here, v t is the second moment estimate at the current time step t, v t-1 represents the first moment estimate at the previous time step, β2 is the decay rate of the second moment, taken as 0.999 here, are the bias corrections for the first moment estimate m t and the second moment estimate v t respectively, t is the number of iteration steps, α is the learning rate, set to 0.001, and ε is a very small constant (set to 10 -8 ).

[0161] The update rule of the Adam optimizer combines first - moment estimation (momentum) and second - moment estimation (gradient squared), enabling each parameter to have an adaptive learning rate, thus optimizing quickly and effectively. Through bias correction, Adam can maintain stable updates in the early stage of training, enabling the neural network to converge in a shorter time.

[0162] S43: When the loss function E L of the model is stable, switch to the SGD optimizer for fine - tuning and convergence. The formula is:

[0163]

[0164] where η is the learning rate, taking 0.001, is the gradient of the loss function E(θ t ) with respect to the current parameter θ t .

[0165] The role of the loss function is to guide the model on how to adjust its parameters to improve prediction accuracy. The optimizer helps to accelerate convergence, avoid overfitting, improve generalization ability, and ensure the stability of the training process through appropriate strategies. The use of a composite optimizer can improve the performance of the LSTM - Transformer neural network in processing complex time - series data and fault prediction tasks, making the model more accurate, stable, and reliable in practical applications.

[0166] S5: Input the sample data of the training set and validation set into the constructed LSTM - Transformer neural network for training to obtain a fault prediction model.

[0167] In step S5, the training set contains a large amount of historical data for the neural network model to learn the relationship between its input features and target outputs. The validation set is used to adjust the hyperparameters of the neural network model and prevent overfitting or underfitting. Set the neural network association parameters Epoch, Interval, and Batch Size. Among them, Epoch refers to the number of iterations of the data set during the entire training process, Interval refers to the frequency of controlling the execution of operations during the training process, and Batch Size refers to the number of samples processed by the neural network in each iteration. Select Epoch as 100, Interval as every 5 Epochs, and Batch Size as 64. Inputting the data samples of the training set and the validation set into the LSTM-Transformer neural network model for training can effectively improve the learning ability and generalization ability of the model. The training set helps the model learn the characteristics and rules of the single-phase grounding fault in the distribution network, while the validation set is used to monitor and optimize the performance of the model to ensure that it does not overfit and has good generalization ability. Through this training process, the finally obtained fault prediction model can more accurately predict the possible faults in actual applications and improve the fault diagnosis and response capabilities of the distribution network.

[0168] S6: Input the sample data in the test set into the trained fault prediction model to obtain the prediction results of the single-phase grounding fault in the distribution network.

[0169] Using the test set to evaluate the prediction ability of the model for the single-phase grounding fault in the distribution network is a key step in model development. Through the evaluation of the test set, the generalization ability, accuracy, robustness, stability and other aspects of the model can be comprehensively measured. The test set helps to verify the feasibility of the fault prediction model in actual applications, ensure that it can accurately and reliably perform fault prediction, and provide feedback and basis for further model optimization. Finally, this evaluation process will help ensure that the model can effectively improve the fault detection ability and operation safety of the distribution network.

[0170] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network, characterized in that: The steps are as follows: S1: Establish a large data set based on the distribution network single-phase grounding fault information in the existing historical fault database; S2: Preprocess the sample data in the large data set, and divide the preprocessed large data set into a training set, a validation set, and a test set in chronological order; S3: Extract local time-frequency features and global frequency domain features from the sample data of the preprocessed large data set, build an LSTM-Transformer neural network, fuse the local time-frequency features and the global frequency domain features, and input them into the LSTM-Transformer neural network to output the predicted fault results; S4: Define the loss function and use the composite optimizer to optimize the LSTM-Transformer neural network; S5: Input the sample data of the training set into the constructed LSTM-Transformer neural network for training, use the sample data in the validation set to verify the trained LSTM-Transformer neural network, save the model parameters to obtain the fault prediction model; S6: Use the sample data in the test set to input the trained fault prediction model to obtain the prediction result of the single-phase grounding fault in the distribution network.

2. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 1 is characterized in that: The big data set includes multimodal data fusion of electrical parameters, environmental data, equipment status data, load fluctuation data, and historical fault data of single-phase grounding of the distribution network; the big data set comes from the power company's historical fault data set, distribution network monitoring system data, power equipment monitoring and status data, meteorological data, load fluctuation data, and public power fault data set; The preprocessing includes: checking whether the data has missing values, outliers and duplicate data, using linear interpolation method to fill the missing values, correct the outliers, and delete the duplicate data; performing Z-score standardization and Min-Max normalization on the corrected sample data; Use isnull() in the Pandas library in Python to detect missing values, use Z-score to detect outliers, and use the drop_duplicates() method of the Pandas library in Python to directly delete duplicate data.

3. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 1 or 2 is characterized in that: The method for extracting local time-frequency features is: using adaptive wavelet transform to extract instantaneous changes and local time-frequency features of the signal, and decomposing the time series data in the large data set into low-frequency parts and high-frequency parts of multiple frequency bands; The method for extracting the global frequency domain feature is: performing FFT on the low frequency part and the high frequency part after the adaptive wavelet transform respectively to obtain the global frequency domain feature; The method of fusing local time-frequency features and global frequency domain features is as follows: fusing local time-frequency features and global frequency domain features into a single feature through statistical operations; The LSTM network is used as a feature extractor to capture the local temporal dependency and short-term memory in the fused features, extract the features of local temporal dependency, short-term memory, and long-term memory, and obtain the hidden state sequence. The multi-head self-attention mechanism of the Transformer network is used to capture the long-range dependency between each time step in the hidden state sequence to obtain the output features. The output layer maps the output features to the prediction space through linear transformation, and uses the Softmax activation function for classification tasks to obtain the predicted fault results.

4. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 3 is characterized in that: The adaptive wavelet transform is: Where x(t) is the input signal, a is the scale factor, and x low (n),x high (n) represent the low-frequency signal and high-frequency signal of the time domain signal, respectively, ψ low (t) is the low-frequency wavelet basis function, ψ high (t) is the high-frequency wavelet basis function; The method for performing FFT is: Among them, X low (k) is the kth frequency component of the low-frequency signal in the frequency domain, X high (k) is the kth frequency component of the frequency domain representation of the high-frequency signal, and N is the total length of the signal; is a complex exponential term; The fused features are: X t =[Max(X low ),Min(X high ),m low ,s low ,m high ,s high ] Among them, Max(X low ) is the low frequency part X low The maximum value, Min(X high ) is the high frequency part X high The minimum value of μ low , σ low are the mean and standard deviation of the low-frequency part, μ high , σ high are the mean and standard deviation of the high frequency part respectively; In the unit structure of the LSTM network, there are three gating mechanisms: forget gate, input gate and output gate. The forget gate is based on the current fused feature X of the input. t and the hidden state h at the previous time t-1 t-1 By forgetting the weight matrix W′ f and W f Calculate the output of the forget gate to control which long-term memories in the LSTM network can be forgotten; use the weight matrix to process the fused feature Xt and the hidden state ht-1 at the previous moment t-1 together with the bias term to obtain the output of the input gate, the input gate and the candidate cell state Together they determine the memory state C at the current moment t The input gate controls the candidate cell state At the current moment, the cell state is C t The write ratio in; the output of the input gate and the forget gate controls the update of the current memory state C t , update the long-term memory cells in the unit state; use the output gate weight matrix to process the fused feature Xt and the hidden state ht-1 at the previous moment t-1 and add the bias term to obtain the output state o of the output gate t , the hidden state h at time t t For: h t =o t *tanh(C t ), tanh is the hyperbolic tangent function.

5. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 3 is characterized in that: At each time step t, the LSTM network outputs a hidden state h t ; After forward propagation of all time steps, the LSTM network generates a hidden state sequence H = [h1,h2,...,h t ], record the hidden state of each time step; perform position encoding on the hidden state sequence H = [h1,h2,...,ht] output by the LSTM network to obtain the position encoding of the even position and the odd position; Output the hidden state h of the LSTM network t Add the corresponding position code PE(t) to get the information with position code: H encoded =H+P; where P=[PE(1), PE(2), ..., PE(t)] is the position code corresponding to the time step.

6. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 4 or 5 is characterized in that: The Transformer network uses a multi-head self-attention mechanism to capture information H with position encoding encoded The relationship between each time step in the multi-head self-attention is further enhanced after being processed by the feedforward neural network layer. The dot product of the query, key and value is calculated and normalized to obtain the attention weight. The output of each attention head is calculated by the attention weight, and the output of all attention heads is concatenated and multiplied with the linear transformation matrix of the learned output to obtain the total output of the multi-head self-attention. The feedforward neural network uses two layers of fully connected layers to process the total output of the multi-head self-attention to obtain the output of the feedforward neural network. The total output of the multi-head self-attention of the current layer is residually connected with the output of the feedforward neural network to obtain the final output feature.

7. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 6 is characterized in that: The structure of the Transformer network has four parts: input, encoder, decoder, and output. The input includes a source text embedding layer, a position encoder, and a position encoder. The source text embedding layer converts the output hidden state sequence H of the LSTM network into [h1, h2, ..., h t ] is converted into a vector representation to capture the relationship between words; the position encoder is a hidden state sequence H = [h1,h2,...,h t ] generates a position vector for each position. The encoder is composed of multiple encoder layers stacked together to capture more complex and abstract features. Each encoder consists of a two-sublayer connection structure. The first sublayer is a multi-head self-attention sublayer, and the second sublayer is a feed-forward fully connected sublayer. Each sublayer is followed by a normalization layer and a residual connection to calculate the dot product of the query, key, and value and normalize them to obtain the attention weight. The decoder is composed of multiple decoder layers stacked together, each decoder layer consists of a three-sublayer connection structure: the first sublayer is a masked multi-head self-attention sublayer, the second sublayer is a multi-head self-attention sublayer, and the third sublayer is a feedforward fully connected sublayer. The output of each self-attention layer in the multi-head self-attention sublayer will be sent to the next self-attention layer. Each sublayer is followed by a normalization layer and a residual connection; the output consists of a linear layer that converts the decoder output vector into the final output dimension and a Softmax layer that converts the output of the linear layer into a probability to obtain a predicted distribution.

8. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 7 is characterized in that: The total output of the multi-head self-attention: A Multi-Head =Concat(A1,A2,…,A h )W O ; Among them, W O is the linear transformation matrix of the learned output, A1, A2, …, A h is the output of h attention heads, and Concat means concatenating the outputs of all attention heads together; The output of the feedforward neural network is z FFN =max(0,A Multi-Head W1+b1)W2+b2, where W1, W2 are the weight matrices of the fully connected layer, and b1, b2 are bias terms; The Transformer network uses residual connections to enhance the gradient flow during training. For the total output A of the multi-head self-attention of the current layer Multi-Head and the output z of the feedforward neural network FFN Perform residual connection to obtain feature z res =A Multi-Head +z FFN ; for feature z res Perform layer normalization to obtain the final output feature: z norm =LayerNorm(z res ); The method of using the Softmax activation function in the output layer to perform the classification task is: p = Softmax (W out ·z norm +b out ), where W out and b out are the weights and biases of the output layer, and p is the final fault prediction result.

9. The method for predicting single-phase grounding fault in distribution network based on LSTM-Transformer neural network according to any one of claims 4, 5, 7 and 8, characterized in that: The method for optimizing the LSTM-Transformer neural network using the composite optimizer in step S4 is: optimizing the LSTM-Transformer neural network through the Adam optimizer to make it converge quickly; when the loss function is stable, switching to the SGD optimizer to fine-tune and converge the LSTM-Transformer neural network.

10. The distribution network single-phase grounding fault prediction method based on LSTM-Transformer neural network according to claim 9 is characterized in that: The loss function in step S4 is: E L n-[ylog(p)+(1-y)log(1-p)] Where y is the true label, y∈{0,1}, y=1 indicates fault, y=0 indicates non-fault, p is the predicted probability output by the LSTM-Transformer neural network, indicating the probability that the sample data belongs to fault, p∈[0,1]; The method for fine-tuning and converging the SGD optimizer is: Where η is the learning rate, is the loss function E(θ t ) for the current parameter θ t The corresponding gradient.

Citation Information

Patent Citations

  • Transformer fault detection method based on fusion of mechanism model and data model

    CN118673399A

Cited By

  • EEG signal ratchet wave detection method and system based on KAN feature fusion

    CN120408385A

  • Transform-based physical information neural network tool wear prediction method

    CN120524154A

  • A tool wear prediction method based on physical information neural network of Transformer

    CN120524154B

  • Multi-feature fusion-based multi-element ocean observation data prediction method

    CN120653940A

  • Machine room precision air conditioner fault early warning method and system based on data fusion and medium

    CN121350922A