XLPE cable partial discharge identification model based on multi-channel hierarchical structure capsule attention network
By adopting a multi-channel hierarchical capsule attention network model in the XLPE cable local discharge signal processing, combining ResNet, BiGRU and capsule network, the focus angle interval penalty loss function is used to solve the data imbalance problem, and high-precision local discharge signal recognition is achieved.
Patent Information
- Application Number
- CN202510229783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when processing the local discharge signal of XLPE cable, it is difficult to effectively deal with the problem of local signal acquisition process containing irrelevant information and data imbalance.
Using a model based on a multi-channel hierarchical capsule attention network, local features are extracted using ResNet, BiGRU captures the long-term features of high-frequency time series signals, and signals are recognized through the capsule network and attention mechanism. At the same time, a focus angle interval penalty loss function (F-softmax) is introduced to solve the class imbalance problem of sample data.
It achieved better local discharge signal recognition accuracy, with an identification accuracy of 99.28%. Through ablation experiments and comparative analysis, the superiority of the model in classification results was proved.
Smart Images

Figure CN120197015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of cable partial discharge signal processing, and particularly to an XLPE cable partial discharge recognition model based on a multi-channel hierarchical capsule attention network. Background Art
[0002] With the steady increase in residential and industrial demands brought about by national progress, the power transmission and distribution system is continuously developing towards higher voltage levels and larger capacities. Traditional overhead transmission lines have been replaced by power cables for the transmission of electrical energy. Power cables have now become an indispensable and crucial equipment in the power transmission and distribution system. The cable is very long and is easily damaged by external forces or aging, such as branch impacts or insulator defects. These damages may lead to partial discharge (PD), which may damage the equipment. Engineers can judge whether there is a defect in the conductor by measuring the impulse component of the signal on site. The current detection methods use manual, aerial inspection and lidar technology to routinely inspect overhead power lines. However, these methods are very time-consuming and expensive.
[0003] In recent years, deep learning such as capsule networks and attention mechanisms has received extensive attention in partial discharge signal detection. For example, Yihuang et al. and Chen Lingling made full use of multi-sensor data to construct a decoupled capsule network. Someone introduced an attention mechanism to help the deep network locate information data segmentation, extract discriminant features of the input information, and automatically focus on relevant fault modes under different operating conditions. A robust weight sharing network was used to diagnose faults under different operating conditions, improving the generalization and adaptability of the model. Some literature proposed a new hybrid deep learning model that combines a CNN and a long short-term memory (LSTM) network for PD pattern analysis. The fusion of the spatial feature extraction of the CNN and the sequence learning ability of the LSTM shows that it has good results in capturing the spatio-temporal features of PD signals. Some literature applied decision trees (DTs) and LSTM designs for PD classification, with classification accuracies of 95.3% and 98.5% respectively. Some literature proposed another SVM-assisted PD pattern recognition method. Some literature used different ML model variables for PD recognition, and the ANN classification accuracy using the main test variables was above 80%. Some literature studied the PD pattern recognition and noise separation performance under the action of DC voltage. Some literature used a deep neural network (DNN) and a shape addition explanation method (SHAP) to identify complex PD signals. Some literature used a CNN to detect PD for the Adan optimizer and the Den and Nadam optimizers.
[0004] The above methods all have their own advantages, but they all have certain limitations and cannot effectively solve the problems of irrelevant information and data imbalance in the local signal acquisition process. Summary of the Invention
[0005] The main objective of the present invention is to provide an XLPE cable partial discharge recognition model based on a multi-channel hierarchical capsule attention network. This model uses ResNet to extract local features and pre-filter existing irrelevant information, uses BiGRU to capture the long-term features of high-frequency time series signals, and utilizes capsule network and attention for the recognition of partial discharge signals. Meanwhile, a new loss function (FMCC) is used to pay more attention to the learning of difficult samples during model training to solve the problem of class imbalance in sample data.
[0006] The technical solution adopted by the present invention is as follows: An XLPE cable partial discharge recognition model based on a multi-channel hierarchical capsule attention network, including:
[0007] Using a BiGRU network to extract the feature of discharge signal data;
[0008] Using ResNet to capture the short-term local dependence pattern of different spatial features;
[0009] Using a capsule network to represent multi-dimensional and spatial information, and perform the recognition and classification of fault conditions.
[0010] Furthermore, the use of a BiGRU network to extract the feature of discharge signal data includes:
[0011] Among them, the forward GRU unit of the hidden layer is represented by and the reverse GRU unit of the hidden layer is represented by :
[0012]
[0013] x t represents the input data set at time t, and h t represents the parameter hidden layer, which is also the output parameter and has all the valid representations of the previous time t; the objective of BiGRU is to obtain an effective representation of the feature information of the partial discharge signal.
[0014] Even further, the use of ResNet to capture the short-term local dependence pattern of different spatial features includes:
[0015]
[0016] σ represents the relu activation function, and O t represents the output of the ResNet model;
[0017] For each input feature, using a feature extractor such as a fully connected layer to generate h t at different time steps, and calculating the context vector c t according to the weighted attention mechanism, representing the attention coefficient of different time steps within the feature:
[0018]
[0019] Among them, T represents the total number of time steps of the input sequence, a tj represents the weight calculated for each state h j at different times T; then c is used t to complete the calculation of the new state sequence s, where s t depends on the state s t-1 , and depends on the output of the hidden layer at t - 1; the weight a tj is calculated as follows:
[0020]
[0021] Among them, a represents a learning function, which is mainly determined by h j and the state s at the previous moment t-1 , and its main function is to calculate the importance of h j ; the vector h generated by the hidden layer t is input into the learnable function to obtain the probability vector a, and the context vector c is calculated using the weighted average value of the h t value; the weight is obtained from a.
[0022] Furthermore, the use of the capsule network to represent multi-dimensional and spatial information and perform the identification and classification of fault conditions includes:
[0023] Using the Squash non-linear function as part of the capsule-level activation during training:
[0024] The Squash function realizes the compression of the feature vector:
[0025]
[0026] ||s j ||limits the modulus length to 1, compresses the modulus length to [0, 0.5], and v j can be regarded as the vector output in the capsule network, and s j is used as the overall input of capsule j;
[0027] Dynamic routing algorithm:
[0028] The overall input capsule `s j , except for the first-layer capsule, is a weighted sum of all prediction vectors hatu j|i from the next-layer capsule, and this matrix is generated by the dot product of the output of the next-layer capsule and the weight matrix W ij ;
[0029]
[0030] c ij Called the coupling coefficient, which is continuously iteratively generated by dynamic routing in the capsule network;
[0031] The sum of the coupling coefficients between capsule i and all the above capsules is 1, calculated by routing softmax, and the initial b ik Is the logarithmic prior probability that capsule i is coupled to capsule k;
[0032]
[0033] Focus angle interval penalty loss function
[0034] Define the focus angle interval penalty function F-Softmax:
[0035]
[0036] Among them, S is the scale factor that amplifies the difference in sample distribution; Is the predicted probability value of class i; The loss function can be expressed as follows:
[0037]
[0038] Among them, t is the true label of the current sample, k is a category, a t Is the weight corresponding to the category to which the current sample belongs, G k Is the number of samples in different categories; β = 2 enables the model to learn samples that are not easily classified.
[0039] Advantages of the present invention:
[0040] The present invention first uses ResNet and BiGRU to extract short-term and long-term high-frequency features of partial discharge signals; then uses capsule networks to represent multi-dimensional and spatial information, and performs identification and classification of fault conditions. Considering the class imbalance of the long-tailed distribution of data samples, a focus angle interval penalty loss function (F-softmax) is proposed, which improves the traditional cross-entropy loss function, thereby reducing the weights assigned to well-classified instances using the attention mechanism. Experimental results show that the model has better recognition accuracy compared with existing models, and the recognition accuracy rate reaches 99.28%. Through ablation experiments and comparisons with existing models, it is proved that the MCHSCA model has better classification results.
[0041] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The following will refer to the drawings to further elaborate on the present invention in detail. Description of the Drawings
[0042] The accompanying drawings, which form a part of this application, are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0043] Figure 1 is a schematic diagram of a gated recurrent unit;
[0044] Figure 2 is a diagram of the MCHSCA model architecture;
[0045] Figure 3 is a structural diagram of the BiGRU neural network of the present invention;
[0046] Figure 4 is a diagram of the ResNet model architecture of the present invention;
[0047] Figure 5 is a diagram of the attention mechanism architecture of the present invention;
[0048] Figure 6 is a structural diagram of the capsule network of the present invention;
[0049] Figure 7 is a simulation circuit diagram for identifying cable PD of the present invention;
[0050] Figure 8 is a waveform diagram of the PD signal of the present invention. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Reference Figures 1 to 8 , this application provides an XLPE cable partial discharge recognition model based on a multi-channel hierarchical structure capsule attention network.
[0053] Gated recurrent unit:
[0054] The GRU (gated recurrent unit) neural network is a variant of the recurrent neural network (RNN) for processing sequential data. It solves the problem of gradient vanishing or gradient explosion that traditional RNNs are prone to in long-sequence training by introducing a gating mechanism. Figure 1 is a schematic diagram of the gated recurrent unit (GRU). At time step t, the input is x t , and its forward calculation process is as follows:
[0055] z t =σ(W z [h t-1 ,x t +bz ) (1)
[0056] r t = σ(W r [h t-1 ,x t +b r ) (2)
[0057]
[0058] In equations (1)-(4), z t and r t are the output values of the update gate and the reset gate respectively. h t and g t are the hidden state and the candidate hidden state respectively, W x represents the corresponding weight parameter, b x represents the corresponding bias parameter, denotes the element-wise product operation. r t and z t control the information update process between time steps.
[0059] Convolutional Neural Network:
[0060] CNN is a member of the feedforward neural network with multiple hidden layers and is well-known for its powerful representation learning and automatic feature extraction capabilities. So far, although many variants have evolved, the typical CNN building blocks include an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer.
[0061] The convolutional layer usually consists of a series of learnable kernels (filters) and an additive bias in each feature map. The input of the network or the output of the previous hidden layer is convolved with a predefined filter, where each element corresponds to a weight coefficient w, and then the sum of the local weighted convolution is added to the corresponding bias vector b, and then passed to the activation function to generate the output feature map. Given the input x l-1 at the (l-1)-th layer, the feature map of the next layer x l can be expressed as:
[0062]
[0063] In equation (5), N is the number of filters in the (l-1)-th layer, k and b are the corresponding convolutional filter and bias respectively, f(·) is the activation function, and the most commonly used one is the rectified linear unit (ReLU), which maps the output of the previous layer through relu(v) = max(v, 0).
[0064] The pooling layer, also known as the downsampling layer, is used to reduce the dimensionality of the feature map output by the convolutional layer without updating the weights, which helps to avoid the risk of overfitting and speed up the calculation of the CNN. The pooling operation on the input map is as follows:
[0065]
[0066] where down(·) represents the pooling function, and β and b are the multiplicative bias and additive bias of each feature map respectively. After the pooling operation, the values of different feature maps in the previous layer are usually normalized to ensure that the distribution of the output of each layer does not change too much.
[0067] After multiple alternating operations of the convolutional layer and the pooling layer, a fully connected layer (FC) is used to connect the neurons of the previous layers, perform a non-linear combination of the extracted features, and then flatten them into a vector as the input of the subsequent classifier. The output of the FC layer is as follows:
[0068] y l = f(w l x l-1 + b l ) (7)
[0069] The output layer is controlled by a softmax classifier, which uses a logistic function and a normalized exponential function to provide the posterior probability of the classification label in order to classify the input data into various categories. One stage of CNN training includes a forward step from the input to the output layer and a backward step from the output to the input, where the backpropagation and gradient descent algorithms are used to minimize the loss function.
[0070] This application proposes an XLPE cable partial discharge recognition model (MCHSCA) based on a multi-channel hierarchical structure capsule attention network, as Figure 2 shown.
[0071] BiGRU network:
[0072] This paper uses a BiGRU network to extract the characteristics of discharge signal data, as Figure 3 shown.
[0073] Among them, the forward GRU unit of the hidden layer is represented by , and the backward GRU unit of the hidden layer is represented by .
[0074]
[0075] x t represents the input data set at time t, and h tRepresents the parameter hidden layer. It is also the output parameter and has all the valid representations at the previous time t. The goal of the BiGRU is to obtain the effective representation of the characteristic information of the partial discharge signal. Different characteristic information contributes differently to the partial discharge recognition. In a complete signal, there is a large amount of noise information and useless information, and there is a phenomenon of information redundancy.
[0076] ResNet model:
[0077] CNN is very good at image or video recognition. In addition, it is also used for feature recognition of sequence data. CNN extracts features through convolutional operators with filters. Different local filters extract different scale features for each input sequence. The CNN units in the hidden layer are divided into different feature maps. At the same time, these cells share a weight matrix. In the area with one-dimensional time series, such as sensor signals, a kernel can be regarded as a feature extractor, and different kernels correspond to different spatial features captured in a specific time series. Therefore, this paper uses ResNet to capture the short-term local dependence pattern of different spatial features, as shown in the following formula:
[0078]
[0079] σ represents the relu activation function. O t Represents the output of the ResNet model.
[0080] The architecture of the ResNet model adopted in this paper is as Figure 4 shown.
[0081] Attention mechanism:
[0082] For each input feature, using a feature extractor such as a fully connected layer, generate h t at different time steps, and calculate the "context vector" c t according to the weighted attention mechanism, which represents the attention coefficients of different time steps within the feature.
[0083]
[0084] Among them, T represents the total number of time steps of the input sequence, and a tj represents the weight calculated for each state h j at different times T. Then use c t to complete the calculation of the new state sequence s, where s t depends on the state s t-1 , and depends on the output of the hidden layer at t - 1. The calculation method of the weight a tj is:
[0085]
[0086] Among them, a represents a learning function, which is mainly determined by h j and the state s at the previous moment t-1 and its main function is to calculate the importance of h j This formula is of great significance for the new state sequence s to better access the entire attention-based state sequence h. Figure 5 It is a schematic diagram of the feed-forward attention mechanism. The vector h generated by the hidden layer t is input into the learnable function to obtain the probability vector a. The weighted average value of the h t values is used to calculate the "context vector" c; the weights are obtained from a.
[0087] Capsule Network:
[0088] The classical neural network structure consists of a convolutional layer, a pooling layer, and a fully connected layer. In any case, the presence of the max pooling layer allows some neurons in a layer to be ignored, and only the most active neurons in the local pool of the previous layer are retained. To reduce information loss, a capsule network with a dynamic routing mechanism is proposed, and its main idea is to use the similarity of different capsules to update the weights, and the structure is as Figure 6 shown.
[0089] In the capsule network, capsules can be regarded as different vectors, the length of which represents the probability of the existence of an entity, and the direction represents the attributes of the entity, specifying the characteristics and possibilities of an object. Capsules are a new concept that can contain more information about each "object". The size of each capsule can be described as an attribute, such as pose (position, size, direction), deformation, speed, and rate of change. To more easily classify different capsules, a non-linear function called "Squash" is used as part of the activation at the capsule level during training.
[0090] Squash Non-linear Function:
[0091] The upper-layer features can be obtained through feature combination, but the intensity of the features needs to be compared. The concept of the module length in the capsule network can solve this problem; that is, the length of the capsule output vector represents the probability that the entity represented by the capsule appears in the current input. The modulus length of the feature vector measures the degree of prominence, but the measurement range is preferably within a bounded range. The Squash function can achieve the compression of the feature vector:
[0092]
[0093] ||s j ||limits the modulus length to 1, compresses the modulus length to [0, 0.5], v jCan be regarded as the vector output in the capsule network, s j As the overall input of capsule j.
[0094] Dynamic routing algorithm:
[0095] The overall input capsule s j , except for the first layer of capsules, can be considered as a weighted sum of all prediction vectors hatu j|i From the capsules in the next layer, this matrix is generated by the dot product of the output of the next layer of capsules mathbfW ij Of the weights.
[0096]
[0097] c ij Called the coupling coefficient, which can be continuously iteratively generated through dynamic routing in the capsule network.
[0098] The sum of the coupling coefficients between capsule i and all the above capsules is 1, calculated by routing softmax, and the initial b ik Is the logarithmic prior probability that capsule i is coupled to capsule k.
[0099]
[0100] Focus angle interval penalty loss function:
[0101] The traditional Softmax loss function is defined as follows:
[0102]
[0103] Among them, f i Is the feature vector belonging to class i before the last fully connected layer. Or W i Is the weight corresponding to the feature vector f i Cosθ i Is the cosine value, and θ i Is the weight W i And the interval between the feature f i .
[0104] Although the Softmax function optimizes the task with respect to θ and W, this optimization direction is not strict for the classification task. If the optimization target is concentrated on a specific variable (θ or W), then the optimization direction will become more explicit and ultimately improve the performance. The main purpose of this paper is to obtain a more strict classification boundary, and the realization of this goal depends on the interaction within and between classes. Therefore, this paper defines two functions: the intra-class function ζ(θ i ) and the inter-class function ξ(θ j ), as follows:
[0105] ζ(θ i ) = cos(θ i + m) (17)
[0106] ξ(θ j ) = cos(θ j - m) (18)
[0107] Where m ∈ [0, π] is the focal angle interval.
[0108] The present invention defines a focal angle interval penalty function (F-Softmax) as follows:
[0109]
[0110] Where S is a scaling factor that amplifies the difference in sample distribution. is the predicted probability value of class i. In this paper, F-Softmax is integrated into the loss function, and the imbalance problem, the between-class problem, and the within-class problem can all be solved to a certain extent. That is to say, this study not only focuses on the between-class / within-class embedded feature space but also on the imbalanced dataset. The final loss function can be expressed as follows:
[0111]
[0112] Where t is the true label of the current sample, k is a class, a t is the weight corresponding to the class to which the current sample belongs, G k is the number of samples in different classes. During the training process, to avoid the model tending to a certain class, β = 2 can make the model learning not easy to classify.
[0113] Since the F-Softmax loss function depends on m to produce a classification effect with margins, thus separating the classification boundaries more, and obtaining a better classification effect than the softmax loss function. The loss function of this paper fully combines the advantages of the focal loss and the within-class loss, allowing the model to not only learn difficult samples and mitigate the impact of sample imbalance but also better learn the larger between-class distance.
[0114] Experimental results and comparative analysis:
[0115] In the experiment, the experimental configuration of this paper is: Intel E5-2680 v4 CPU (2.4 GHz), RTX3090 GPU (24 GB), Linux system, and the running environment is PyTorch (Anaconda) which is currently the most widely used.
[0116] Partial discharge pattern recognition aims to determine whether there is a partial discharge phenomenon in the cable and evaluate the corresponding cable fault defect type by analyzing and identifying the characteristics of partial discharge signals. In this paper, PSCAD is used to build the 10kV cable PD identification simulation circuit as shown in Figure 7 , and the simulated PD signal is as shown in Equation (22).
[0117] I t = I0(e -1.3t / τ -e -2.2t / τ )sin(2πf c t) (22)
[0118] Where: I0 is the signal amplitude; τ is the attenuation coefficient; f c is the oscillation frequency. An analog PD signal as shown in Figure 8 is injected into the cable.
[0119] Performance evaluation index:
[0120] The essence of the partial discharge signal pattern recognition task is a classification task. To quantitatively analyze the model's partial discharge signal pattern recognition ability, the accuracy rate and confusion matrix commonly used in classification tasks are adopted in this paper to evaluate the model performance. Among them, the accuracy rate calculation formula can be expressed as:
[0121]
[0122] Where, y i is the true category of sample x i , N represents the total number of samples, and f(·) represents the prediction model.
[0123] Recognition accuracy comparison:
[0124] In the present invention, the MCHSCA model is compared with other commonly used methods. Different models including Transformer, VGG, TCN, ResNet, LSTMFCN and traditional CatBoost are used, and the loss functions used include binary cross-entropy loss function BCE, Combo, DiceBCE and Focal. The experimental results of partial discharge signal recognition of each model under different loss functions are shown in Table 1.
[0125] Table 1 Comparison of recognition results of different models under different loss functions
[0126]
[0127]
[0128] As can be seen from Table 1, on the one hand, generally speaking, the F-softmax loss function proposed in this paper is better than other functions. Since this function pays more attention to the data with small sample characteristics and can prevent parameters from remaining in a poor local minimum by predicting the correlation between labels and actual values, it gradually learns better model parameters. On the other hand, the performance of the MCHSCA model proposed in this paper is better than other commonly used models. This is mainly due to the fact that this model introduces dropout and regularization to reduce the probability of model overfitting. At the same time, the learning rate of the Adam optimizer decreases exponentially, avoiding the situation where the loss does not decrease with the increase of the number of iterations.
[0129] Ablation experiment:
[0130] To verify the necessity of each module in MCHSCA, the present invention observes the changes in model performance through ablation experiments. First, comparing the accuracy of the model after removing the capsule network with that of the complete model, it can be found that there is a certain degree of decline in the recognition accuracy. Second, comparing the accuracy of the model after removing the attention layer with that of the complete model, the results are shown in Table 2. It can be seen from this that removing the attention layer and the capsule network has a greater impact on the model performance. This is mainly because different features have different impacts on the model performance, but if all features are treated equally, it will introduce noise and affect the accuracy of the model. The results prove that using the capsule network can obtain more information about each object, and at the same time, by introducing the attention layer, the importance of different features can be effectively identified and different signals can be located during the weight update process.
[0131] Comparison of ablation experiment results in Table 2
[0132]
[0133] Dynamic routing can be regarded as a parallel attention mechanism, which allows each capsule to focus on some active capsules in the lower layer while ignoring other capsules. The present invention also analyzes the influence of the complexity of the network structure on the results under different dynamic routing numbers, capsule numbers, and dimension parameter values. As shown in Table 3. The experimental results show that as the number of routings increases, the experimental accuracy will not be improved. The more routings there are, the more likely it is to cause overfitting, resulting in a decrease in the experimental accuracy. After many practices, it is proved that choosing 3 or 4 routings in the MCHSCA model has higher accuracy than other cases. Similarly, the number and size of capsules also affect the accuracy. The results show that relatively superior results can be obtained when both the number and size of capsules are 8.
[0134] Comparison of recognition results of different models in Table 3
[0135]
[0136] Conclusion:
[0137] Aiming at the problems of complex distribution and class imbalance existing in the partial discharge signals of XLPE cables, this invention proposes a partial discharge recognition model (MCHSCA) for XLPE cables based on a multi-channel hierarchical structure capsule attention network. This model uses ResNet to extract local features and pre-filter existing irrelevant information, uses BiGRU to capture the long-term features of high-frequency time series signals, and uses capsule network and attention to identify partial discharge signals. At the same time, a new loss function (FMCC) is used to pay more attention to the learning of difficult samples during the model training process, so as to solve the class imbalance problem of sample data. Through ablation experiments and comparisons with existing models, it is proved that the MCHSCA model has better classification results.
[0138] This invention proposes a partial discharge recognition model (MCHSCA) for XLPE cables based on a multi-channel hierarchical structure capsule attention network to solve this problem. In the MCHSCA model, ResNet and BiGRU are used to extract the short-term and long-term high-frequency features of partial discharge signals respectively. Even if the sample data contains a large amount of irrelevant information, it can effectively classify the fault states in the high-frequency signals. To solve the problem of sample data imbalance, the F-softmax loss function is proposed to improve sample imbalance. The results show that this model is significantly better than other existing methods and obtains the best recognition accuracy.
[0139] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. XLPE cable partial discharge recognition model based on multi-channel hierarchical capsule attention network, characterized by: include: The BiGRU network is used to extract the features of discharge signal data; ResNet is used to capture short-term local dependency patterns of different spatial features; Capsule networks are used to represent multidimensional and spatial information and to identify and classify fault conditions.
2. The XLPE cable partial discharge identification model based on multi-channel hierarchical capsule attention network according to claim 1 is characterized in that: The BiGRU network is used to extract the discharge signal data features include: Among them, the forward GRU unit of the hidden layer is used Indicates that the reverse GRU unit of the hidden layer is used express: x t represents the input data set at time t, h t The representation parameter hidden layer, which is also the output parameter, has all valid representations of the previous time t; the goal of BiGRU is to obtain an effective representation of the characteristic information of the partial discharge signal.
3. The XLPE cable partial discharge identification model based on multi-channel hierarchical capsule attention network according to claim 1 is characterized in that: The short-term local dependency patterns of different spatial features captured by ResNet include: The t =σ(Conv1D(σ(Conv1D(σ(Conv1D(x))))))+σ(Conv1D(x)) (10) σ represents the relu activation function, O t Represents the output of the ResNet model; For each input feature, a fully connected feature extractor is used to generate h at different time steps. t , and calculate the context vector c according to the weighted attention mechanism t , characterizing the attention coefficients at different time steps within a feature: Where T represents the total time step of the input sequence, a tj Represents each state h j The weights calculated at different times T; then use c t To complete the calculation of the new state sequence s, where s t Depends on the state t-1 , and depends on the output of the hidden layer at t-1; weight a tj The calculation method is: Among them, a represents a learning function, which is mainly composed of h j and the state s at the previous time t-1 Determine, its main function is to calculate h j The importance of the hidden layer generated vector h t Input the learnable function, get the probability vector a, and use h t The context vector c is calculated by taking the weighted average of the values; the weights are obtained from a.
4. The XLPE cable partial discharge identification model based on multi-channel hierarchical capsule attention network according to claim 1 is characterized in that: The method of using capsule networks to represent multi-dimensional and spatial information and to identify and classify fault conditions includes: Use the Squash nonlinearity as part of the capsule-level activation during training: The Squash function realizes the compression of feature vectors: ||s j ||Limit the modulus length to 1, Compress the modulus length to [0,0.5], v j It can be regarded as the vector output in the capsule network, s j As the overall input of capsule j; Dynamic routing algorithm: Whole input capsule`s j , except for the first layer of capsules, is a weighted sum of all prediction vectors hatu j|i From the next layer of capsules, this matrix is generated by the point output next layer capsule mathbfW ij The weight of c ij It is called the coupling coefficient and is continuously iterated by dynamic routing in the capsule network; The coupling coefficient between capsule i and all the above capsules is added to 1, which is calculated by routing softmax. The initial b ik is the logarithmic prior probability of capsule i coupling to capsule k; Focus Angle Separation Penalty Loss Function Define the focus angle interval penalty function F-Softmax: Among them, S is the proportional factor that amplifies the difference in sample distribution; θ F-softmax is the predicted probability value of class i; the loss function can be expressed as follows: Where t is the true label of the current sample, k is a category, and a t is the weight corresponding to the category to which the current sample belongs, G k is the number of samples in different categories; β = 2 makes it difficult for the model to learn.
Citation Information
Cited By
New energy cable fault early warning method and early warning system based on artificial intelligence
CN121476822A
New energy cable fault early warning method and early warning system based on artificial intelligence
CN121476822B