Method for diagnosing partial discharge of reactor based on multi-source feature fusion and deep network
Through the multi-source feature fusion and deep network methods, the problem of multi-dimensional feature fusion and insufficient cross-domain adaptability in reactor local discharge monitoring is solved, and high-precision and robust local discharge diagnosis is achieved to adapt to reactor insulation state evaluation under complex operating conditions.
Patent Information
- Application Number
- CN202510473525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing reactor local discharge monitoring algorithm lacks multi-dimensional feature fusion capability and cross-domain adaptability, resulting in low monitoring accuracy and robustness, making it difficult to cope with complex nonlinear, non-stationary characteristics and noise interference.
Multi-source feature fusion and deep network methods are adopted, including data preprocessing, SWFK-DPC clustering, deep multi-origin adaptive convolutional neural network EDMSACNN, cross-domain feature compression, dual-channel attention mechanism and adversarial enhancement module, and optimize the model's feature fusion and anti-interference ability.
It significantly improves the adaptive cross-domain capability and robustness of local discharge monitoring of reactors, ensures the reliability and stability of diagnosis, and adapts to the data distribution differences under different operating conditions.
Smart Images

Figure CN120011863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of partial discharge detection, and more specifically, to a method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network. Background Art
[0002] As a key device in the power system, internal insulation defects of the reactor can induce partial discharge (PD), accelerate insulation aging, and threaten the safety of the equipment. Traditional off-line detection methods are limited by long periods and insufficient real-time performance, which has promoted the research focus on partial discharge feature analysis technology based on on-line monitoring. However, the non-linear, non-stationary characteristics of PD signals and the complexity of the high-dimensional feature space pose severe challenges to traditional fault diagnosis methods.
[0003] Currently, the monitoring of partial discharge in reactors mainly relies on a single physical quantity for pattern recognition, which is difficult to comprehensively reflect the essential characteristics of partial discharge, and is relatively sensitive to noise and operating conditions, resulting in limited diagnostic accuracy. In addition, the analysis methods based on statistical features often assume that the signals follow a specific distribution and are difficult to handle complex non-linear and non-stationary characteristics. Therefore, it is necessary to perform deep fusion of multi-dimensional physical quantities to comprehensively characterize the spatio-temporal evolution characteristics of partial discharge signals. Traditional dimensionality reduction techniques mainly rely on linear projection and are difficult to capture the non-linear correlations between different feature modes, resulting in information loss and reduced pattern discrimination. In addition, although deep learning methods have powerful feature extraction capabilities, they often fail to fully exploit the complementary information between physical features and statistical features in multi-modal fusion, which limits the accurate modeling of complex discharge patterns. On the other hand, the distribution of partial discharge signals is affected by environmental factors and equipment aging, and there are significant offsets in the data distribution under different operating conditions. Existing models generally rely on fixed feature extraction methods and lack the adaptive ability to domain offsets, resulting in poor cross-condition generalization and affecting the stability and engineering applicability of diagnosis. Therefore, there is an urgent need to construct an intelligent diagnosis method with both multi-dimensional feature fusion ability and cross-domain adaptability to improve the accuracy and robustness of partial discharge monitoring.
[0004] As a density-based clustering algorithm, density peak clustering (DPC) has been widely applied in many fields because it does not require presetting the number of clusters, has good adaptability to non-spherical clustering structures, and can intuitively identify cluster centers through a decision graph. Especially in partial discharge signal analysis, anomaly detection, and high-dimensional data clustering tasks, DPC relies on the joint measure of local density and relative distance, and can effectively mine the internal structure of data. However, a series of limitations have emerged when the traditional DPC algorithm deals with high-dimensional data. First, standard similarity measures, such as Euclidean distance, do not fully consider the different contributions of different features to the overall distance. Especially when the feature scales vary greatly, it may lead to inaccurate and unstable clustering results. Second, the DPC algorithm is sensitive to noise. Especially in high-dimensional spaces, outliers will interfere with the calculation of local density, thereby affecting the identification of cluster centers. In addition, when the differences between data features are small, DPC is difficult to effectively distinguish similar samples, resulting in a decrease in clustering accuracy. Therefore, in response to these problems, there is an urgent need to improve the traditional DPC algorithm to enhance its adaptability and robustness in high-dimensional data. And existing deep learning models still have significant deficiencies in multi-modal feature processing and robustness improvement. When facing multi-modal features, traditional convolutional neural networks (CNNs) usually adopt simple splicing or early fusion strategies. This method fails to fully mine the high-order interaction information between features, thus limiting the model's representation ability for complex patterns. In addition, current mainstream attention mechanisms, such as channel attention, mainly focus on single-level feature selection, ignoring the collaborative optimization of global and local information, resulting in limited expression ability of the model when dealing with complex signals. More importantly, partial discharge signals are often affected by electromagnetic interference and noise pollution. Traditional classification models are highly sensitive to these adversarial perturbations and are difficult to effectively cope with the challenges brought by noise and distribution shifts. Therefore, how to improve the robustness of the model to noise and environmental changes while ensuring high-precision classification has become a key problem that needs to be solved urgently. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network, which is used to solve the problem that the existing partial discharge monitoring algorithms have low accuracy and robustness in partial discharge monitoring due to the lack of both multi-dimensional feature fusion ability and cross-domain adaptability.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: A method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network includes the following steps:
[0007] S1. Data preprocessing and feature extraction;
[0008] Perform Min-Max normalization processing and missing value filling on the collected partial discharge signals of the reactor;
[0009] S2. Extract multi-dimensional feature information, including phase features, polarity features, amplitude features, time-frequency features, and wavelet packet features;
[0010] S3. Process the high-dimensional feature set through the SWFK-DPC clustering algorithm;
[0011] S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN to achieve the deep fusion of the physical features and statistical features of partial discharge signals;
[0012] S5. Introduce cross-domain feature compression to improve the cross-domain adaptability of the model; introduce a clustering-guided dual-path attention mechanism to optimize the feature weight allocation in space and channels;
[0013] S6. Introduce an adversarial enhancement and classification module to improve the classification accuracy and anti-interference ability of the model;
[0014] S7. Model optimization and iterative training;
[0015] After the model is constructed, train and optimize the partial discharge mode of the reactor under complex working conditions. Through hyperparameter tuning, loss function optimization, and model iterative training, improve the model stability and high recognition rate.
[0016] Furthermore, in step S2, after performing 3-layer wavelet packet decomposition on the signal, extract the energy proportion of 8 frequency bands. The calculation formula is: where E pξ represents the energy proportion of the ξ-th frequency band, x pξ is the wavelet packet coefficient sequence of the corresponding frequency band, and x(i) is the original signal sequence.
[0017] Furthermore, step S3 specifically includes the following steps:
[0018] S3.1. Calculate the similarity between samples using the standard deviation weighted distance metric;
[0019] S3.2. Calculate the local density of each sample, and construct a decision graph based on the local distance to identify the density peak points, and use these density peak points as the clustering centers;
[0020] S3.3. When the maximum neighborhood radius of sample i exceeds the average maximum neighborhood radius of all samples, this point is determined as an outlier, otherwise it is a non-outlier. Assign the non-outliers to the optimal cluster; for the unassigned non-outliers, introduce the fuzzy weighted K-nearest neighbor method FK and propose a DPC clustering assignment strategy.
[0021] Furthermore, in step S3.1, calculate the standard deviation weighted distance between samples i and j where is the weight factor, m is the number of features of each sample, is the value of sample i on the th feature, is the value of sample j on the th feature.
[0022] Furthermore, in step S3.2, the local density of each sample i is defined as where KNN i is the set of K nearest neighbors of sample i, and the local distance δ i is defined as: Based on this, a decision graph is constructed to identify the density peak. The x-axis represents the local density, and the y-axis represents the local distance. By observing the positions of the points in the decision graph, the density peak located in the upper right corner can be visually identified.
[0023] Furthermore, in step S3.3, based on the standard deviation weighted distance d ij a criterion for outlier determination is proposed, specifically: the determination criterion of outliers is defined by the formula where, is the maximum neighborhood radius of sample i within its set of K nearest neighbors, τ is the mean of the maximum neighborhood radii of all sample points in the dataset, n is the total number of samples in the dataset; the membership degree of sample i belonging to cluster c where, κ ij is the weighted similarity measure, γ ij is the normalization coefficient;
[0024] For the dataset X = {x i | i = 1,..., n} ∈ R n×m , calculate the membership degree of each sample i to cluster c and construct the matrix S, where, r is the number of samples to be processed, C is the number of clusters. Select the sample s with the highest membership degree and assign it to the corresponding cluster c, and at the same time set its membership degrees in all clusters to zero. Update the membership degrees of adjacent samples based on their K-nearest neighbor relationships, and iterate this process until all samples are assigned.
[0025] Furthermore, the specific steps of step S4 include:
[0026] S4.1. Through the outer product operation, Hadamard product, and linear weighted fusion method, interactively fuse the phase feature, polarity feature, and amplitude feature;
[0027] S4.2. Jointly encode the time-frequency features using the Short-Time Fourier Transform (STFT) and the Discrete Wavelet Transform (DWT) to characterize the time-frequency domain dynamic features of the partial discharge signal.
[0028] S4.3. For the wavelet packet features, further model them using the Graph Convolutional Network (GCN) in a tree structure to capture the relationships between different energy nodes.
[0029] Furthermore, in step S4.1, the five-dimensional feature vector output by the SWFK-DPC clustering is reconstructed into a three-channel feature tensor. For the phase information φ, polarity p, and amplitude A, use the outer product operation, Hadamard product, and linear weighted fusion. The mathematical expression is: where F phy is the fused physical feature; W p is the learnable parameter matrix corresponding to the polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function to introduce a non-linear mapping to enhance the model's expressive ability.
[0030] Furthermore, in step S4.2, use the Short-Time Fourier Transform (STFT) combined with time-domain convolution for joint encoding. The calculation formula is: where F tfr is the fused time-frequency feature; TFR represents the time-frequency matrix obtained after the Short-Time Fourier Transform; DWT is the Discrete Wavelet Transform; TFR envelope represents the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D represents the realization of time-series feature modeling through a sliding window operation.
[0031] Furthermore, in step S4.3, use the Graph Convolutional Network (GCN) in a tree structure for modeling. The mathematical expression is: where F wpe is the fused wavelet packet feature, WPE l is the wavelet packet feature at the l-th level, α l is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: where Q is the learnable query vector, the softmax function realizes weight normalization, and WPE T l is the wavelet packet feature at the l-th level after adaptive weighting.
[0032] Furthermore, in step S5, for cross-domain feature compression, first align the features and calculate the Sinkhorn distance between the SWFK-DPC clustering center C k and the convolutional feature F i , (16); where M() is the Mahalanobis distance, H(Γ) is the entropy regularization term, is the optimal transport mapping matrix, ε is the regularization coefficient, is calculated from the normalized weights; after feature alignment is completed, a dynamic channel compression strategy based on channel attention is introduced to select a feature subset with key discriminant information in the high-dimensional feature space; according to the normalized weights calculate the channel attention weights where MLP() is the multi-layer perceptron; a tensor contraction operation is used to perform deep fusion between physical features, time-frequency features, and wavelet packet features, and the fused features are calculated as: where w α , w β are the attention weights of different feature modalities respectively.
[0033] Furthermore, the dual-path attention mechanism includes a local attention branch, a global attention branch, and dual-path fusion; the local attention branch uses spatial clustering information to optimize the feature representation, highlighting the significant features of the target region and suppressing background noise; first, based on the SWFK-DPC algorithm, calculate the membership matrix U of the feature points to describe the belonging degree of the input features on different clustering centers. On this basis, by calculating the spatial attention map A local adjust the spatial weights of the feature map, and the calculation method is: A local = softmax(U T W l U)(19); where W l is the learnable parameter matrix; the global attention branch uses the distribution information of the overall clustering to optimize the channel attention, and introduces the clustering compactness index S g as the channel attention adjustment factor, and the global attention weight where C g is the corresponding clustering center; the fusion method in the dual-path attention fusion stage is: F out = (1 - λ)·(A local ☉F) + λ·(A global ☉F) (21); where λ is the adaptive fusion coefficient, calculated by the multi-layer perceptron MLP: λ = σ(MLP([A local , A global )) (22).
[0034] Furthermore, the specific steps of step S6 include:
[0035] S6.1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discriminative ability, and use constrained adversarial perturbations to improve the stability of the model under noise or data distribution shift. First, map the fused features to a low-dimensional space, as expressed below: F cls = ReLU(W d ·F fusion + b d )(23); where W d is the dimensionality reduction matrix; introduce an L2-constrained adversarial perturbation ν on the dimensionality-reduced feature F cls to obtain an adversarially enhanced feature representation: F adv = F cls + ν(24); The L2 constraint is a penalty term that imposes an L2 norm on the model parameters during the optimization process to prevent the model parameters from being too large, thereby improving the generalization ability and stability. b d is the bias term introduced during the feature dimensionality reduction process to adjust the distribution of the dimensionality-reduced features. F fusion is the feature representation after fusion;
[0036] S6.2. Based on the adversarially enhanced features, use the Softmax function to calculate the class probability distribution, and jointly optimize the model through cross-entropy loss and adversarial consistency constraints;
[0037] The class probability distribution P is calculated by the following formula, where W c is the classification weight vector, b c is the bias term of the classifier, b k is the bias term of the clustering center C k W c T is the transpose matrix of the classification weight vector, W k T is the transpose of the feature weight matrix of the clustering center C k ; The cross-entropy loss and the adversarial consistency constraint constitute the loss function: ζ is a hyperparameter, and y c represents the true class label.
[0038] The beneficial effects of the present invention are as follows: The present invention realizes the non-linear interaction of multi-source features through tensor contraction and attention mechanism, breaks through the limitations of traditional linear fusion, and effectively alleviates the data distribution differences under different working conditions based on the feature alignment technology of optimal transport, significantly improving the adaptive cross-domain ability. In addition, combining adversarial training and clustering-guided attention mechanism enhances the robustness of the model in a noisy environment and ensures the diagnostic reliability. This technology provides an intelligent solution for the insulation state assessment of reactors and has important engineering application value for the safe and stable operation of power systems. Description of the Drawings
[0039] Figure 1 This is the framework diagram of the SWFK-DPC clustering algorithm of the present invention;
[0040] Figure 2 This is the algorithm architecture diagram of the deep multi-source adaptive convolutional neural network EDMSACNN of the present invention;
[0041] Figure 3 This is the overall model framework diagram based on SWFK-DPC-EDMSACNN of the present invention. Detailed implementation manners
[0042] The following will describe in detail the method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network of the present invention in conjunction with the figures.
[0043] S1. Data preprocessing and feature extraction.
[0044] Perform Min-Max normalization processing and missing value filling on the collected partial discharge signals of the reactor to ensure the consistency and integrity of the data.
[0045] S2. Extract multi-dimensional feature information, including phase features, polarity features, amplitude features, time-frequency features, and wavelet packet features, to comprehensively capture the dynamic and frequency domain characteristics of partial discharge signals. Among them, the phase feature reflects the change characteristics of partial discharge signals in periodic behavior; the polarity feature describes the polarity change of partial discharge signals within one cycle; the amplitude features include peak value, average value, and root mean square value statistics, which are used to describe the amplitude information of partial discharge signals; the time-frequency features are used to describe the distribution characteristics of partial discharge signals in the time domain and frequency domain, specifically including rise time, signal width, fall time, main frequency, equivalent time length, and equivalent bandwidth; the wavelet packet features extract the energy distribution information of signals in different frequency bands through wavelet packet decomposition. Specifically, after performing 3-layer wavelet packet decomposition on the signal, the energy proportion of 8 frequency bands is extracted. The calculation formula is: Among them, E pξ represents the energy proportion of the ξ-th frequency band, x pξ is the wavelet packet coefficient sequence corresponding to the frequency band, and x(i) is the original signal sequence.
[0046] S3. As Figure 1 shown, process the high-dimensional feature set through the SWFK-DPC clustering algorithm.
[0047] S3.1. Calculate the similarity between samples using the standard deviation weighted distance metric.
[0048] Calculate the standard deviation weighted distance between samples i and j Among them, is the weight factor, m is the number of features of each sample, is the value of sample i on the th feature, is the value of sample j on the th feature. The d ij metric takes into account the differences between different features, making the differences between features have a reasonable impact on the final distance calculation.
[0049] S3.2. Calculate the local density of each sample, and construct a decision graph based on the local distance to identify the density peak points, and use these density peak points as the clustering centers.
[0050] Define the local density of each sample i where KNN i represents the set of K nearest neighbors of sample i. The local distance δ i is defined as: On this basis, construct a decision graph to identify the density peak. Use the x-axis to represent the local density and the y-axis to represent the local distance. By observing the position of the points in the decision graph, visually identify the density peak located in the upper right corner.
[0051] In order to effectively divide the remaining sample points in the dataset into non-outliers and outliers, based on the standard deviation weighted distance d ij metric, an outlier determination criterion is proposed. The outlier determination criterion is: define the determination standard of outliers through the formula where, is the maximum neighborhood radius of sample i within its set of K nearest neighbors, and τ is the mean of the maximum neighborhood radii of all sample points in the dataset.
[0052] n is the total number of samples in the dataset.
[0053] S3.3. For the unassigned non-outliers, introduce the fuzzy weighted K-nearest neighbor method FK to further optimize the clustering effect.
[0054] Based on the above determination criterion, assign the non-outliers to the optimal cluster. When the maximum neighborhood radius of sample i exceeds the average maximum neighborhood radius of all samples, this point is determined as an outlier. For the unassigned non-outliers, introduce the fuzzy weighted K-nearest neighbor method FK and propose an improved DPC clustering assignment strategy. This strategy optimizes the density peak assignment process by calculating the fuzzy membership degree of the samples, so as to ensure that the non-density peak points can be more accurately assigned to the optimal clustering center. Specifically, the membership degree p i c of sample i belonging to cluster c is determined by the membership degrees of the samples in its K nearest neighbors, and its calculation is as follows: where κ ij is the weighted similarity metric, ensuring a reasonable distribution of the membership degrees. γij is a normalization coefficient,
[0055] For the dataset X = {x i | i = 1,..., n} ∈ R n×m , calculate the membership degree of each sample i to the cluster c and construct the matrix S, where r is the number of samples to be processed, and C is the number of clusters. Select the sample s with the highest membership degree and assign it to the corresponding cluster c, and at the same time set its membership degrees in all clusters to zero. Update the membership degrees of adjacent samples based on their K-nearest neighbor relationships, and iterate this process until all samples are assigned.
[0056] S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN.
[0057] For the high-dimensional feature set after SWFK-DPC clustering, an improved deep multi-source adaptive convolutional neural network algorithm Enhanced DMSACNN is proposed. Its core innovation lies in constructing a hybrid coding architecture with multi-feature interaction to achieve the deep fusion of physical features and statistical features of partial discharge signals. The Enhanced DMSACNN framework is as Figure 2 shown. Input the obtained multi-dimensional features into EDMSACNN for further analysis and pattern recognition. EDMSACNN adopts a hybrid coding architecture with multi-feature interaction, including the deep fusion of physical features, time-frequency features, and wavelet packet features.
[0058] S4.1. Through the outer product operation, Hadamard product, and linear weighted fusion method, interactively fuse the phase feature, polarity feature, and amplitude feature.
[0059] Reconstruct the five-dimensional feature vector output by SWFK-DPC clustering into a three-channel feature tensor. For the phase information φ, polarity p, and amplitude A, adopt the outer product operation, Hadamard product, and linear weighted fusion. The mathematical expression is: where F phy is the fused physical feature; W p is the learnable parameter matrix corresponding to the polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function to introduce a non-linear mapping to enhance the model's expressive ability.
[0060] S4.2. Use the short-time Fourier transform STFT and the discrete wavelet transform DWT to jointly encode the time-frequency features, so as to characterize the time-frequency domain dynamic features of partial discharge signals.
[0061] To make full use of time-frequency domain information, the short-time Fourier transform (STFT) is combined with time-domain convolution for joint coding to characterize the time-frequency domain dynamic characteristics of PD signals. The calculation formula is as follows: where F tfr is the fused time-frequency feature; TFR is the time-frequency matrix obtained after the short-time Fourier transform; DWT is the discrete wavelet transform; TFR envelope is the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D is used to model the temporal features through a sliding window operation.
[0062] S4.3. For the wavelet packet features, a graph convolutional network (GCN) is further used for modeling to capture the relationships between different energy nodes.
[0063] Considering the contribution of the wavelet packet energy (WPE) to the PD discharge pattern, a graph convolutional network (GCN) is used for modeling to characterize the relationships between different energy nodes. Its mathematical expression is as follows: where F wpe is the fused wavelet packet feature, WPE l is the wavelet packet feature at the l-th level, and α l is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: where Q is the learnable query vector, and the softmax function is used to normalize the weights to achieve adaptive weighting. WPE T l represents the wavelet packet feature at the l-th level after adaptive weighting.
[0064] S5. Cross-domain feature compression and introduction of the attention mechanism.
[0065] To improve the cross-domain adaptation ability of the model, a cross-domain feature compression method based on the optimal transport theory is introduced. This method measures the similarity of feature distributions by calculating the Sinkhorn distance metric and aligns the features. After feature alignment, a dynamic channel compression strategy based on the channel attention mechanism is adopted to select a subset of features with key discriminative information. In addition, a clustering-guided dual-path attention mechanism is proposed, which includes local and global attention branches to optimize the feature weight allocation in both space and channels.
[0066] To optimize the feature representation and enhance the cross-domain adaptability, a cross-domain feature compression method based on the optimal transport theory is introduced. This method first aligns the features and calculates the Sinkhorn distance between the SWFK-DPC clustering center C k and the convolutional feature F i to measure the similarity of feature distributions. where M() is the Mahalanobis distance and H(Γ) is the entropy regularization term, is the optimal transport mapping matrix, and ε is the regularization coefficient. is calculated from the normalized weights.
[0067] After feature alignment is completed, a dynamic channel compression strategy based on channel attention is introduced to select a feature subset with key discriminant information in the high-dimensional feature space. According to the normalized weights calculate the channel attention weights where MLP() is the multi-layer perceptron.
[0068] To further improve the information interaction ability between different feature modalities, a tensor contraction operation is adopted to achieve deep fusion among physical features, time-frequency features, and wavelet packet features. The fused features are calculated as: where w α , w β are the attention weights of different feature modalities respectively.
[0069] To further enhance the feature fusion effect and strengthen the synergistic effect between global and local information while extracting multi-scale features, a clustering-guided dual-path attention mechanism is proposed. This mechanism includes a local attention branch, a global attention branch, and dual-path fusion, which effectively improves the discriminability and robustness of feature representation through collaborative modeling of space and channels.
[0070] The local attention branch optimizes the feature representation using spatial clustering information to highlight the significant features of the target region and suppress background noise. First, the membership matrix U of feature points is calculated based on the SWFK-DPC algorithm to describe the belonging degree of the input features to different clustering centers. On this basis, by calculating the spatial attention map A local to adjust the spatial weights of the feature map, and its calculation method is: A local = softmax(U T W l U)(19). Where W l is a learnable parameter matrix designed to capture spatial correlations, and the softmax operation ensures the normalization of the attention distribution, making the model pay more attention to the significant regions indicated by the clustering structure. This mechanism enables the network to adaptively adjust the feature weights according to the spatial clustering pattern, thereby enhancing the expression ability of the target region and improving the distinguishability of features.
[0071] The global attention branch focuses on optimizing the channel attention using the overall distribution information of the clustering to enhance the feature selection ability in the channel dimension. For this purpose, the clustering compactness index S g is introduced as a channel attention adjustment factor, and the global attention weight where S greflects the compactness of the k-th cluster, that is, the similarity level within the class; C g represents the corresponding cluster center. In the dual-path attention fusion stage, to fully combine local and global attention information, a gating mechanism is adopted for dynamic fusion, making the final feature representation more abundant and having higher discriminability. The fusion method is: F out =(1 - λ)·(A local ⊙F)+λ·(A global ☉F). Among them, λ is the adaptive fusion coefficient, calculated by the multi-layer perceptron MLP: λ = σ(MLP([A local ,A global ))(22).
[0072] S6. Adversarial enhancement and classification module design.
[0073] To improve the classification accuracy and anti-interference ability of the model, an adversarial enhancement classification module is introduced.
[0074] S6.1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discriminative ability, and use constrained adversarial perturbations to improve the stability of the model under noise or data distribution shift.
[0075] After completing feature fusion and attention weighting, the classification module design is carried out to improve the accuracy and robustness of pattern recognition. First, map the fused features to a low-dimensional space related to the class to enhance the discriminative ability and suppress redundant information. The feature transformation is completed through a fully connected layer, and the formal expression is as follows: F cls =ReLU(W d ·F fusion +b d )(23). Among them, W d is the dimensionality reduction matrix. To improve the anti-interference ability of the model, a constrained L2 adversarial perturbation ν is introduced on the dimensionality-reduced feature F cls , and the adversarial enhanced feature representation is obtained: F adv =F cls +ν(24). The L2 constraint is a penalty term that imposes the L2 norm on the model parameters during the optimization process to prevent the model parameters from being too large, thereby improving the generalization ability and stability. b d is the bias term introduced during the feature dimensionality reduction process to adjust the dimensionality-reduced feature distribution, and F fusion is the feature representation after fusion.
[0076] S6.2. Based on the adversarial enhanced features, use the Softmax function to calculate the class probability distribution, and jointly optimize the model through the cross-entropy loss and the adversarial consistency constraint.
[0077] This perturbation prompts the model to maintain stable discriminative performance when confronting uncertain factors, such as noise or data distribution shift. Based on the adversarial enhanced features, the Softmax function is used to calculate the class probability distribution where, W c is the classification weight vector, which forms an implicit alignment relationship with the SWFK-DPC clustering center C k so that the classifier further combines the pattern structure information based on the deep features, improving the generalization ability; b c is the bias term of the classifier, used to adjust the class probability distribution output by Softmax; b k is the bias term of the clustering center C k used to optimize the matching relationship between the features and the clustering center; W c T is the transpose matrix of the classification weight vector, used to calculate the class probability score; W k T is the transpose of the feature weight matrix of the clustering center C k used to characterize the mapping relationship between the deep features and the pattern structure, so as to improve the classification performance. The loss function consists of the cross-entropy loss and the adversarial consistency constraint, to optimize the classification accuracy while enhancing the robustness of the model: where, the cross-entropy term optimizes the classification decision, and the adversarial consistency constraint term ensures the consistency of the feature distribution before and after perturbation, suppressing the influence of adversarial samples. The hyperparameter ζ balances the classification accuracy and the anti-interference ability, enabling the model to maintain high stability and high recognition in complex environments, and y c represents the true class label.
[0078] S7. Model optimization and iterative training.
[0079] After the model is constructed, it is trained and optimized for the partial discharge patterns of the reactor under complex working conditions. Through hyperparameter tuning, loss function optimization and model iterative training, the performance of the model is further improved. Especially in complex environments, by combining adversarial training and the dual-channel attention mechanism, the high stability and high recognition of the model in practical applications are ensured.
[0080] The present invention realizes the non-linear interaction of multi-source features through tensor contraction and the attention mechanism, breaks through the limitations of traditional linear fusion, and effectively alleviates the data distribution differences under different working conditions based on the feature alignment technology of optimal transport, significantly enhancing the adaptive cross-domain ability. In addition, by combining adversarial training and the clustering-guided attention mechanism, the robustness of the model in a noisy environment is enhanced, ensuring the diagnostic reliability. This technology provides an intelligent solution for the insulation state assessment of reactors, and has important engineering application value for the safe and stable operation of power systems.
Claims
1. A method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network, characterized in that It includes the following steps: S1. Data preprocessing and feature extraction; Perform Min-Max normalization processing and missing value filling on the collected partial discharge signals of the reactor; S2. Extract multi-dimensional feature information, including phase features, polarity features, amplitude features, time-frequency features, and wavelet packet features; S3. Process the high-dimensional feature set through the SWFK-DPC clustering algorithm; S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN to achieve the deep fusion of the physical features and statistical features of partial discharge signals; S5. Introduce a cross-domain feature compression and clustering-guided dual-path attention mechanism; S6. Introduce an adversarial enhancement and classification module to improve the classification accuracy and anti-interference ability of the model; S7. Model optimization and iterative training; After the model is constructed, train and optimize the partial discharge patterns of the reactor under complex working conditions. Through hyperparameter tuning, loss function optimization, and model iterative training, improve the model stability and high recognition; Step S3 specifically includes the following steps: S3.
1. Calculate the similarity between samples using the standard deviation weighted distance metric; S3.
2. Calculate the local density of each sample, and construct a decision graph based on the local distance to identify the density peak points, and use these density peak points as the clustering centers; S3.
3. When the maximum neighborhood radius of sample i exceeds the average maximum neighborhood radius of all samples, this point is determined as an outlier, otherwise it is a non-outlier. Assign the non-outliers to the optimal cluster; for the unassigned non-outliers, introduce the fuzzy weighted K-nearest neighbor method FK and propose a DPC clustering assignment strategy; The specific steps of step S4 include: S4.
1. Through the outer product operation, Hadamard product, and linear weighted fusion method, interactively fuse the phase features, polarity features, and amplitude features; S4.
2. Use the short-time Fourier transform STFT and discrete wavelet transform DWT to jointly encode the time-frequency features to characterize the time-frequency domain dynamic features of partial discharge signals; S4.
3. For the wavelet packet features, adopt the tree diagram convolutional network GCN for further modeling to capture the relationships between different energy nodes.
2. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network according to claim 1, wherein In step S2, after performing three-layer wavelet packet decomposition on the signal, the energy proportion of 8 frequency bands is extracted, and the calculation formula is: Among them, E pξ represents the energy proportion of the ξ-th frequency band, x pξ is the wavelet packet coefficient sequence corresponding to the frequency band, and x(i) is the original signal sequence.
3. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network according to claim 1, characterized in that In step S3.1, calculate the standard deviation weighted distance between samples i and j where is the weight factor, m is the number of features of each sample, is the value of sample i on the th feature, is the value of sample j on the th feature.
4. The partial discharge diagnosis method of the reactor based on multi-source feature fusion and deep network according to claim 3, characterized in that In step S3.2, define the local density of each sample i where KNN i is the set of K nearest neighbors of sample i, and the local distance δ i is defined as: On this basis, construct a decision graph to identify density peaks. Use the x-axis to represent the local density and the y-axis to represent the local distance. By observing the positions of the points in the decision graph, visually identify the density peaks located in the upper right corner.
5. The partial discharge diagnosis method of a reactor based on multi-source feature fusion and a deep network according to claim 4, wherein In step S3.3, based on the standard deviation weighted distance d ij a criterion for outlier determination is proposed based on the metric, specifically: the determination criterion of outliers is defined by the formula , where is the maximum neighborhood radius of sample i within its set of K nearest neighbors, τ is the mean of the maximum neighborhood radii of all sample points in the dataset, n is the total number of samples in the dataset; the membership degree of sample i belonging to cluster c where κ ij is the weighted similarity metric, γ ij is the normalization coefficient; For the dataset X = {x i | i = 1, ..., n} ∈ R n×m , calculate the membership degree of each sample i to the cluster c Select the sample s with the highest membership degree and assign it to the corresponding cluster c. At the same time, set its membership degrees in all clusters to zero, and construct the matrix S where r is the number of samples to be processed, C is the number of clusters, update the membership degrees of adjacent samples based on their K-nearest neighbor relationships, and iterate this process until all samples are assigned 6. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network according to claim 1, wherein In step S4.1, the five-dimensional feature vector output by SWFK-DPC clustering is reconstructed into a three-channel feature tensor. For the phase information φ, polarity p, and amplitude A, outer product operation, Hadamard product, and linear weighted fusion are adopted, and the mathematical expression is as follows: where F phy is the fused physical feature; W p is the learnable parameter matrix corresponding to the polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function to introduce a non-linear mapping to enhance the model's expression ability.
7. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network according to claim 6, wherein In step S4.2, joint encoding is performed by combining the short-time Fourier transform (STFT) with time-domain convolution, and the calculation formula is as follows: Among them, F tfr is the fused time-frequency feature; TFR represents the time-frequency matrix obtained after the short-time Fourier transform; DWT is the discrete wavelet transform; TFR envelope represents the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D represents the realization of time-series feature modeling through a sliding window operation.
8. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network according to claim 7, characterized in that, In step S4.3, a tree-structured graph convolutional network GCN is used for modeling, and the mathematical expression is: where F wpe is the fused wavelet packet feature, WPE l is the wavelet packet feature at the l-th level, and α l is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: where Q is the learnable query vector, the softmax function realizes weight normalization, and WPE T l is the wavelet packet feature at the l-th level after adaptive weighting.
9. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network according to claim 8, wherein In step S5, cross-domain feature compression first aligns the features and calculates the SWFK-DPC clustering center C g and the convolutional features the Sinkhorn distance between them, where M() is the Mahalanobis distance and H(Γ) is the entropy regularization term, is the optimal transport mapping matrix, ε is the regularization coefficient, is calculated from the normalized weights; after feature alignment, a dynamic channel compression strategy based on channel attention is introduced to select a feature subset with key discriminative information in the high-dimensional feature space; according to the normalized weights calculate the channel attention weights where MLP() is the multi-layer perceptron; a tensor contraction operation is used to perform deep fusion between physical features, time-frequency features, and wavelet packet features, and the fused features are calculated as: where w α , w β are the attention weights of different feature modalities respectively.
10. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network according to claim 9, wherein, The dual-path attention mechanism includes a local attention branch, a global attention branch, and dual-path fusion; the local attention branch optimizes the feature representation using spatial clustering information, highlighting the significant features of the target region and suppressing background noise; first, based on the SWFK-DPC algorithm, the membership matrix U of the feature points is calculated to describe the belonging degree of the input features on different clustering centers. On this basis, the spatial attention map A is calculated local to adjust the spatial weight of the feature map. The calculation method is: A local = softmax(U T W l U)(19); Among them, W l is a learnable parameter matrix; the global attention branch optimizes the channel attention by using the distribution information of the overall clustering, and introduces the clustering compactness index S g as the channel attention adjustment factor, the global attention weight Among them, C g is the corresponding clustering center; the fusion method in the dual-path attention fusion stage is: F out =(1 - λ)·(A local ⊙F)+λ·(A global ⊙F)(21); where λ is the adaptive fusion coefficient, calculated by the multi-layer perceptron MLP: λ = σ(MLP([A local , A global ))(22).
11. The method for diagnosing partial discharge of a reactor based on multi-source feature fusion and a deep network according to claim 10, wherein The specific steps of step S6 include: S6.
1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discriminative ability, and use constrained adversarial perturbations to improve the stability of the model under noise or data distribution shift. First, map the fused features to a low-dimensional space, expressed as follows: F cls = ReLU(W d ·F fusion + b d )(23); where W d is the dimensionality reduction matrix; introduce an L2-constrained adversarial perturbation ν on the dimensionality-reduced feature F cls to obtain an adversarial enhanced feature representation: F adv = F cls + ν(24); The L2 constraint is a penalty term that imposes an L2 norm on the model parameters during the optimization process. b d is the bias term introduced during the feature dimensionality reduction process to adjust the dimensionality-reduced feature distribution. F fusion is the feature representation after fusion; S6.
2. Based on the adversarial enhanced features, use the Softmax function to calculate the class probability distribution, and jointly optimize the model through the cross-entropy loss and adversarial consistency constraint; The class probability distribution P is calculated by the following formula: where W c is the classification weight vector, b c is the bias term of the classifier, and b k is the bias term of the clustering center C k ; W c T is the transposed matrix of the classification weight vector, and W k T is the transpose of the feature weight matrix of the clustering center C k ; the cross-entropy loss and the adversarial consistency constraint constitute the loss function: ζ is a hyperparameter, and y c represents the true class label.
Citation Information
Patent Citations
Partial discharge detection network system based on adaptive attention mechanism
CN119646611A
Partial discharge mode recognition method based on residual neural network and attention module
CN119691518A