Partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer
By using the Legendre multi-wavelet Patch Embedding Transformer method in local discharge detection, combined with the Legendre multi-wavelet transformation and attention mechanism, the diagnostic problems under long-time series and sparse PD mode distribution are solved, and more efficient feature extraction and diagnostic accuracy are achieved.
Patent Information
- Application Number
- CN202510186058.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively diagnose and handle long-time series and sparsely distributed local discharge (PD) patterns, resulting in challenges in learning and parallel capabilities of deep learning models.
The local discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer is adopted, and the Patch Embedding module is improved by combining Legendre multi-wavelet transformation and attention mechanism, which enhances feature extraction capabilities, solves information loss problems, and improves the parallelism and information extraction capabilities of the model.
It effectively solves the PD diagnosis problem under long-time series and sparse PD mode distribution conditions, and improves the model's feature extraction ability and diagnostic accuracy.
Smart Images

Figure CN120196981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of partial discharge diagnosis of insulated wires and artificial intelligence technology, and specifically to a partial discharge detection method based on Legendre multi - wavelet Patch Embedding Transformer. Background Technique
[0002] The health status of the covered wire directly affects the safety and stability of the power system and the safety of maintenance personnel. As a state quantity, PD can effectively reflect the state of the covered wire. Therefore, if the PD pattern in the signal can be accurately diagnosed, maintenance can be effectively carried out to avoid huge losses caused by faults. However, in the actual environment, since PD is a rare event and its pulse duration is usually less than 1 to 4 microseconds, the distribution of PD patterns in the entire dataset is sparse. At the same time, the short duration requires a higher high - sampling rate, and the signals collected in this case belong to long - time series, which poses a huge challenge to the learning and parallel capabilities of deep - learning models.
[0003] PD detection techniques based on data - driven feature engineering can be divided into two categories: methods that rely on traditional machine - learning feature extraction engineering and end - to - end feature extraction methods based on deep - learning models. Machine - learning - based feature extraction methods focus on extracting complex features from the original signal and usually combine traditional machine - learning models for PD detection. These methods usually have strong robustness because they rely on expert - knowledge features to ensure that key information is captured. In contrast, end - to - end feature extraction methods based on deep - learning models directly input more original signal information into the deep - learning model after pre - processing, reducing manual intervention and improving scalability.
[0004] Patch Embedding (PE) technology is widely used to reduce the input burden of deep - learning models to improve the throughput of diagnostic models. However, the traditional Patch Embedding module is only composed of simple multi - layer perceptrons or multi - layer convolutional structures, resulting in relatively weak feature - extraction capabilities in this part.
[0005] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0006] The objective of the present invention is to provide a partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer. The present invention combines the Legendre multi-wavelet transform (LW) and the attention mechanism to improve the feature extraction ability of the Patch Embedding (PE) module. This improvement enhances the PE module's ability to extract context and detail features, while alleviating the problem of information loss caused by a large block size, effectively solving the PD diagnosis problem under long time series and sparse PD pattern distribution conditions.
[0007] To achieve the above objective, the present invention provides the following technical solution: A partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer, at least including the following steps:
[0008] S1: Improve the traditional PE module to establish an LW Patch Embedding model, namely the LWPE model. The LWPE model includes an LWCA module and an LWSA module. The LWPE model combines the channel attention mechanism, the spatial attention mechanism, and the Legendre multi-wavelet transform, and uses the LWCA module and the LWSA module to efficiently complete the feature extraction of PD signals under complex environmental conditions, improving the framework parallelism and information extraction ability;
[0009] S2: Propose an LWPET framework based on the LWPE model, the LW coefficients, and the Transformer Encoder classifier. The LWPET framework performs feature extraction at different granularities on the LW coefficients of the signal and retains important information for PD diagnosis. The LWPET framework is used to improve data parallelism and learning ability;
[0010] S3: Use the ENET partial discharge dataset to evaluate the LWPET framework. During the evaluation process, in order to meet the actual diagnosis requirements and the repeatable characteristics, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2;
[0011] S4: Complete partial discharge detection using the evaluated LWPET framework.
[0012] Preferably, the LWCA module is used to calculate the attention weights of each sub-band in the multi-scale time-frequency domain of the LW coefficients, with the aim of discovering sub-bands helpful for classification;
[0013] For the input feature map The LWCA module first slices along the dimension W represented by the number of multi-wavelets to generate tensors where \(j\in\{1,2,\cdots,W\}\), \(N\) represents the number of patches, and \(P\) represents the patch length. indicates that the data is of real number type;
[0014] Each slice corresponds to the coefficient distribution of a single wavelet over the patch length \(P\) and consists of \(M\) subbands.
[0015] Preferably, for each slice, calculating the attention weights of the subbands includes at least the following steps:
[0016] S1.1.1: Apply sum pooling and max pooling to each subband of along the patch length \(P\) to generate two matrices and is the sum pooling result, is the max pooling result;
[0017] S1.1.2: Then, input the pooled matrices and into a shared multi-layer perceptron respectively for processing to achieve interaction between multi-wavelet channels, as shown in Equation (1);
[0018]
[0019] where that is, and are both tensors of shape ; MLP is the multi-layer perceptron;
[0020] S1.1.3: Add the outputs of the shared multi-layer perceptron and apply the Sigmoid activation function to map the values to \([0, 1]\), as shown in Equation (2);
[0021]
[0022] where, represents the attention weight of the subbands in slice ;
[0023] S1.1.4: Repeat the steps of S1.1.1 to S1.1.3, and splice the attention weights of each slice along the wavelet dimension to generate an attention tensor, as shown in Equation (3);
[0024]
[0025] where, \(N\) represents the number of patches, and \(W\) represents the dimension represented by the number of multi-wavelets.
[0026] Preferably, the LWSA module is designed to calculate attention weights along the patch length, enabling the model to capture the spatial correlation features in the wavelet coefficients. The input feature map of the LWSA module comes from the LWCA module, and the following steps are executed starting from this:
[0027] S1.2.1: Perform max-pooling and average-pooling with a kernel size of k and a stride of s along the patch length P to generate two reduced tensors and
[0028] is the max-pooling result, is the average-pooling structure;
[0029] where N represents the number of patches, W represents the wavelet dimension, and P′ represents the length of the patch after pooling;
[0030] Through the pooling operation, the resolution of the feature map on the patch length P′ is reduced, reducing the computational amount while retaining the key information;
[0031] S1.2.2: Pass the two pooled tensors through two independent fully connected layers respectively to achieve interaction in the patch dimension. And, to increase non-linearity, use the ReLU activation function, as shown in formula (4);
[0032]
[0033] where FC represents the fully connected layer, and are the tensors obtained after being processed by the fully connected layer and have undergone non-linear transformation by the ReLU activation function;
[0034] S1.2.3: Unfold the and two tensors along the patch dimension and concatenate them into a new matrix
[0035]
[0036] The concatenated matrix contains the information from the max-pooling and average-pooling operations and becomes N×2WP′ in dimension;
[0037] S1.2.4: Pass through an MLP. During the MLP processing, first map the second dimension to 2WP′, then restore it to the WP dimension to calculate the spatial attention weights, and finally pass it to the sigmoid function to generate the spatial attention tensor Refer to formula (5);
[0038]
[0039] Among them,
[0040] Preferably, the LWPET framework for completing partial discharge detection at least includes the following steps:
[0041] Instance normalization and slicing. Let the i-th signal be where L represents the signal length. First, instance normalization is used for the input signal, and then the signal is sliced into a patch sequence where N represents the number of patches and P represents the patch length;
[0042] Perform Legendre multi-wavelet transform on the patch sequence. The sliced patch sequence is decomposed into low-frequency and high-frequency multi-scale and multi-resolution feature representation coefficients through Legendre multi-wavelet, and then stitched along the multi-wavelet dimension to generate a multi-wavelet coefficient tensor where W represents the dimension represented by the number of multi-wavelets;
[0043] Process the multi-wavelet coefficient tensor with the LWPE model. The LWPE model consists of the LWCA module and the LWSA module. Its core is to enhance the classification ability of the framework through the attention mechanism, which is used to enhance the representation of each wavelet's sub-band and patch dimension respectively, and finally generate a feature tensor where D represents the specified embedding space dimension;
[0044] Process the embedded features with a Transformer encoder. Use an improved Transformer encoder model with batch normalization as the regularization method to process the feature tensor Generate an output feature map
[0045] Complete PD diagnosis with a multi-layer perceptron. The output feature map is first flattened into a vector, and then a label is output through a multi-layer perceptron Complete PD detection.
[0046] Compared with the prior art, the beneficial effects of the present invention are:
[0047] The present invention constructs an LWPET framework. The internal LW Patch Embedding (LWPE) model combines the designed LWCA and LWSA modules, which are used to effectively mine the PD patterns existing in the LW coefficients of the signal;
[0048] Among them, the LWCA module calculates the attention weights of each frequency sub-band from the perspective of multi-wavelet multi-scale, enabling the model to focus on the sub-bands most relevant to the classification task;
[0049] The LWSA module calculates the attention weights along the patch length, enabling the model to capture spatially correlated features in the wavelet coefficients;
[0050] The concatenation of the LWCA module and the LWSA module is connected in a residual manner with the input LW coefficients, which not only optimizes the gradient backpropagation flow but also improves the representation ability of the PE module, thus effectively solving the PD diagnosis problem under the conditions of long sequences and sparse PD pattern distributions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic diagram of the data flow of the LWPET framework of the present invention;
[0053] Figure 2 It is a schematic diagram of the LWPE model of the present invention;
[0054] Figure 3 It is a schematic diagram of the LWCA module of the present invention;
[0055] Figure 4 It is a schematic diagram of the LWSA module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0057] The present invention uses the Legendre multi-wavelet transform (LW) and the attention mechanism to enhance the feature extraction ability of the PatchEmbedding module in the proposed framework. Through the designed Legendre multi-wavelet channel attention (LWCA) module and Legendre multi-wavelet spatial attention (LWSA) module, learnable feature extraction is performed on the Legendre multi-wavelet coefficients of the signal at the Patch scale, effectively solving the problems of long sequence representation learning, sparse partial discharge patterns, and poor interpretability of the deep learning-based PD detection framework in the partial discharge (PD) signals collected at high sampling rates.
[0058] Please refer toFigures 1-4 , a partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer, at least including the following steps:
[0059] S1: Improve the traditional PE module to establish an LW Patch Embedding model, i.e., the LWPE model. The LWPE model includes an LWCA module and an LWSA module. The LWPE model combines the channel attention mechanism, the spatial attention mechanism, and the Legendre multi-wavelet transform, and uses the LWCA module and the LWSA module to efficiently complete the feature extraction of PD signals under complex environmental conditions, improving the framework parallelism and information extraction ability;
[0060] S2: Based on the LWPE model, the LW coefficients, and the Transformer Encoder classifier, propose the LWPET framework. The LWPET framework performs feature extraction at different granularities on the LW coefficients of the signal and retains important information for PD diagnosis. The LWPET framework is used to improve data parallelism and learning ability;
[0061] S3: Use the ENET partial discharge dataset to evaluate the LWPET framework. During the evaluation process, in order to meet the actual diagnosis needs and the repeatable characteristics, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2;
[0062] S4: Use the evaluated LWPET framework to complete partial discharge detection.
[0063] The LWCA module is used to calculate the attention weights of each sub-band in the multi-scale time-frequency domain of the LW coefficients, aiming to discover sub-bands helpful for classification;
[0064] For the input feature map The LWCA module first slices along the dimension W represented by the number of multi-wavelets and generates tensors where j ∈ {1, 2,..., W}, N represents the number of patches, P represents the patch length, indicates that the data is of real number type;
[0065] Each slice corresponds to the coefficient distribution of a single wavelet on the patch length P and consists of M sub-bands.
[0066] For each slice, calculating the attention weights of the sub-bands at least includes the following steps:
[0067] S1.1.1: For Apply sum pooling and max pooling to each sub-band of along the patch length P to generate two matrices and is the sum pooling result, is the max pooling result;
[0068] S1.1.2: Next, input the pooled matrices and into the shared multi-layer perceptron respectively for processing to achieve the interaction between multi-wavelet channels, as shown in formula (1);
[0069]
[0070] where that is and are both tensors of shape ; MLP is the multi-layer perceptron;
[0071] S1.1.3: Add the outputs of the shared multi-layer perceptron and apply the Sigmoid activation function to map the values to [0, 1], as shown in formula (2);
[0072]
[0073] where, represents the attention weight of the sub-band in the slice ;
[0074] S1.1.4: Repeat the steps of S1.1.1 to S1.1.3, and splice the attention weights of each slice along the wavelet dimension to generate an attention tensor, as shown in formula (3);
[0075]
[0076] where, N represents the number of patches, and W represents the dimension represented by the number of multi-wavelets.
[0077] The LWSA module aims to calculate the attention weights along the patch length, enabling the model to capture the spatial correlation features in the wavelet coefficients. The input feature map of the LWSA module comes from the LWCA module, and the following steps are executed starting from this:
[0078] S1.2.1: Perform max pooling and average pooling with a kernel size of k and a stride of s along the patch length P on to generate two reduced tensors and
[0079] is the max pooling result, is the average pooling structure;
[0080] Where N represents the number of patches, W represents the wavelet dimension, and P′ represents the length of the patch after pooling;
[0081] Through the pooling operation, the resolution of the feature map in the patch length P′ is reduced, reducing the computational amount while retaining the key information;
[0082] S1.2.2: Pass the two pooled tensors through two independent fully connected layers respectively to achieve interaction in the patch dimension. And, in order to increase non-linearity, use the ReLU activation function, as shown in formula (4);
[0083]
[0084] Where FC represents the fully connected layer, and are tensors obtained after being processed by the fully connected layer and have undergone non-linear transformation by the ReLU activation function;
[0085] S1.2.3: Unfold the two tensors and along the patch dimension and splice them into a new matrix
[0086]
[0087] The spliced matrix contains information from the max-pooling and average-pooling operations and becomes N×2WP′ in dimension;
[0088] S1.2.4: Process through an MLP. During the MLP processing, first map the second dimension to 2WP′, then restore it to the WP dimension to calculate the spatial attention weight, and finally pass it to the sigmoid function to generate the spatial attention tensor Refer to formula (5);
[0089]
[0090] Where,
[0091] The LWPET framework for completing partial discharge detection includes at least the following steps:
[0092] Instance normalization and slicing. Let the i-th signal be where L represents the signal length. First, use instance normalization for the input signal, and then slice the signal into a patch sequence where N represents the number of patches and P represents the patch length;
[0093] Perform the Legendre multi-wavelet transform on the patch sequence, and the segmented patch sequence is decomposed into low-frequency and high-frequency multi-scale multi-resolution feature representation coefficients through Legendre multi-wavelets, and then stitched along the multi-wavelet dimension to generate a multi-wavelet coefficient tensor where W represents the dimension represented by the number of multi-wavelets;
[0094] Process the multi-wavelet coefficient tensor with the LWPE model. The LWPE model consists of the LWCA module and the LWSA module. Its core is to enhance the classification ability of the framework through the attention mechanism, which is used to enhance the representation of each wavelet's sub-band and patch dimension respectively, and finally generate a feature tensor where D represents the specified embedding space dimension;
[0095] Process the embedded features with a Transformer encoder, and use an improved Transformer encoder model with batch normalization as the regularization method to process the feature tensor Generate an output feature map
[0096] Complete PD diagnosis with a multi-layer perceptron. The output feature map is first flattened into a vector, and then a label is output through a multi-layer perceptron Complete PD detection.
[0097] Experimentally verify the framework proposed in the present invention on the long-sequence partial discharge signal dataset. The specific process is as follows;
[0098] To verify the effectiveness of the proposed method, we tested it on a large partial discharge signal dataset.
[0099] Effect of Legendre multi-wavelet decomposition on partial discharge detection
[0100] Table 1 Influence of Legendre multi-wavelet decomposition parameters on the results
[0101]
[0102] As shown in Table 1, after introducing the Legendre multi-wavelet transform, the performance is improved compared with the baseline model without decomposition. The specific results are as follows:
[0103] Using three LW bases and three-layer decomposition can achieve the highest MCC (0.757), which is an improvement over the baseline MCC (0.747).
[0104] Increasing the decomposition depth to more than three layers slightly reduces the performance, indicating that over-decomposition may introduce noise or redundant information.
[0105] These findings indicate that LWT can effectively capture the temporal features in PD data, and the three LW bases and three-layer decomposition achieve the best balance between feature representation and computational efficiency.
[0106] Effect of different patch sizes on partial discharge detection
[0107] Table 2 Effect of different patch sizes and batch sizes on the results
[0108]
[0109] As shown in Table 2, the analysis results are as follows:
[0110] When the patch size is 4000, the highest MCC (0.767) is achieved, but its recall rate (0.709) is not sufficient to support online deployment.
[0111] When the patch size is shortened to 2000, the MCC remains stable (0.760), while the recall rate increases to 0.764.
[0112] Considering the importance of recall rate in reducing missed detections in online scenarios, a patch length of 2000 is selected as the best configuration between detection accuracy and recall rate.
[0113] Effect of module combinations in the LWPE model on partial discharge detection
[0114] Table 3 Ablation experiments of the LWPE model
[0115]
[0116]
[0117] The ablation experiment results of the LWPE model are summarized in Table 3, which details the individual and combined contributions of the LWCA and LWSA modules. Specifically, LWCA and LWSA increase the MCC to 0.807 and 0.803 respectively, showing a significant improvement compared to the baseline MCC (0.769). When combined, these modules achieve the highest MCC (0.812), recall rate (0.803), and AUC (0.897), indicating their complementary effect. The integration of these modules enhances feature representation and improves the gradient flow during training, thus achieving robust detection performance under long time series conditions.
[0118] In summary:
[0119] The present invention improves the traditional PE module, establishes the LWPE model, and combines the LWPE model with the LW and Transformer Encoder classifiers to propose the LWPET framework. This framework extracts features at different granularities on the LW coefficients of the signal and retains important information for PD diagnosis.
[0120] The construction process of the LWPE model is completed by integrating the LWCA and LWSA modules through the attention mechanism, which is used to optimize the multi-wavelet time-frequency domain features. The LWCA module focuses on frequency-specific attention to highlight the most relevant sub-bands, while the LWSA module emphasizes the spatial significant features across patches. These two modules work together to reflect the importance of specific features by assigning attention weights, thereby improving the interpretability of the model. Such a combination solves the problem that the classification ability of the classifier is limited due to simple mapping in the traditional PE module and can effectively solve the PD diagnosis problem under long sequences and sparse PD patterns.
[0121] Through a large number of comparison and ablation experiments on the ENET PD detection large dataset, the effectiveness and robustness of the LWEPT framework proposed by the present invention are verified. The Matthews correlation coefficient, precision, recall rate, F1 score, area under the curve, and accuracy of the method proposed by the present invention on this PD detection dataset reach 0.812, 0.845, 0.803, 0.822, 0.897, and 0.979 respectively.
[0122] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. A partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer, characterized by: At least the following steps are included: S1: Improve the traditional PE module and establish the LW Patch Embedding model, namely the LWPE model. The LWPE model includes the LWCA module and the LWSA module. The LWPE model combines the channel attention mechanism, the spatial attention mechanism and the Legendre multi-wavelet transform, and uses the LWCA module and the LWSA module to efficiently complete the feature extraction of the PD signal under complex environmental conditions, thereby improving the framework parallelism and information extraction capabilities; S2: Based on the LWPE model, LW coefficients and Transformer Encoder classifier, an LWPET framework is proposed. The LWPET framework performs feature extraction of different granularities on the LW coefficients of the signal and retains important information for PD diagnosis. The LWPET framework is used to improve data parallelism and learning ability. S3: The ENET partial discharge dataset is used to evaluate the LWPET framework. In order to meet the actual diagnostic needs and repeatable characteristics, the dataset is divided into training set, validation set and test set in a ratio of 6:2:2 during the evaluation process; S4: Partial discharge detection is completed using the evaluated LWPET framework.
2. The method for partial discharge detection based on Legendre multi-wavelet Patch Embedding Transformer according to claim 1, characterized in that: The LWCA module is used to calculate the attention weight of each sub-band in the multi-scale time-frequency domain of the LW coefficient, with the aim of discovering sub-bands that are helpful for classification; For the input feature map The LWCA module first splits along the dimension W represented by the number of multi-wavelets and generates a tensor Where j∈{1,2,...,W}, N represents the number of patches, P represents the patch length, Indicates that the data is of real number type; Each slice corresponds to the coefficient distribution of a single wavelet over a patch length P and consists of M subbands.
3. The partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer according to claim 2 is characterized in that: For each slice, calculating the sub-band attention weights includes at least the following steps: S1.1.1: Yes Sum pooling and max pooling are applied to each subband of along the patch length P, generating two matrices and To sum the pooling results, is the maximum pooling result, M represents the number of sub-bands; S1.1.2: Next, the pooled matrix and They are input into the shared multi-layer perceptron for processing, and the interaction between multiple wavelet channels is realized, as shown in formula (1); in Right now and All shapes Tensor of; MLP is a multi-layer perceptron; S1.1.3: Add the outputs of the shared multilayer perceptron and apply the sigmoid activation function to map the values to [0, 1], as shown in formula (2); in, Representing slices Attention weights of neutron bands; S1.1.4: Repeat steps S1.1.1 to S1.1.3 to set the attention weight of each slice Concatenation along the wavelet dimension generates the attention tensor, as shown in formula (3); in, N represents the number of patches, and W represents the dimension represented by the number of multiwavelets.
4. The partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer according to claim 3 is characterized in that: The LWSA module is designed to calculate the attention weights along the patch length so that the model can capture the spatial correlation characteristics in the wavelet coefficients. The input feature map of the LWSA module comes from the LWCA module, and the following steps are performed based on this: S1.2.1: Yes Perform max pooling and average pooling with kernel size k and stride s along the patch length P, generating two reduced tensors and is the maximum pooling result, It is the average pooling structure; Where N represents the number of patches, W represents the wavelet dimension, and P′ represents the length of the patch after pooling; Through the pooling operation, the resolution of the feature map at the patch length P′ is reduced, reducing the amount of computation while retaining key information; S1.2.2: The two pooled tensors are passed through two independent fully connected layers to achieve interaction in the patch dimension, and, in order to increase nonlinearity, the ReLU activation function is used, as shown in formula (4); Among them, FC represents the fully connected layer, and It is a tensor obtained after processing through the fully connected layer and undergoing a nonlinear transformation using the ReLU activation function; S1.2.3: and The two tensors are expanded along the patch dimension and concatenated into a new matrix The concatenated matrix Contains information from the maximum pooling and average pooling operations, and the dimension becomes N×2WP′; S1.2.4: Processed through an MLP, the second dimension is first mapped to 2WP′, then restored to the WP dimension to calculate the spatial attention weight, and finally passed to the sigmoid function to generate the spatial attention tensor See formula (5); in, 5. The partial discharge detection method based on Legendre multi-wavelet Patch Embedding Transformer according to claim 4 is characterized in that: The LWPET framework for performing partial discharge detection comprises at least the following steps: Instance normalization and slicing, let the i-th signal be Where L represents the signal length. First, the input signal is instance normalized and then the signal is split into a sequence of patches. Where N represents the number of patches and P represents the patch length; Perform Legendre multi-wavelet transform on the patch sequence and transform the segmented patch sequence The multi-wavelet is decomposed into low-frequency and high-frequency multi-scale multi-resolution feature representation coefficients, and then spliced along the multi-wavelet dimension to generate a multi-wavelet coefficient tensor. Where W represents the dimension represented by the number of multi-wavelets; The LWPE model is used to process multi-wavelet coefficient tensors. The LWPE model consists of LWCA module and LWSA module. Its core is to enhance the classification ability of the framework through the attention mechanism, which is used to enhance the representation of the sub-band and patch dimensions of each wavelet, and finally generate a feature tensor. Where D represents the specified embedding space dimension; Use the Transformer encoder to process the embedded features and use the modified Transformer encoder model with batch normalization as the regularization method to process the feature tensor Generate output feature map Use multi-layer perceptron to complete PD diagnosis and output feature map First it is flattened into a vector, then it is passed through a multi-layer perceptron to output the label Complete PD testing.