Reactor partial discharge diagnosis method based on multi-source feature fusion and deep network
By adopting multi-source feature fusion and deep network methods in reactor local discharge monitoring, the problem of low monitoring accuracy and robustness in the prior art is solved, and more efficient and reliable local discharge diagnosis is achieved.
Patent Information
- Application Number
- CN202510473525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art lacks intelligent diagnostic methods that combine multi-dimensional feature fusion capability and cross-domain adaptability in local discharge monitoring of reactors, resulting in low accuracy and robustness of monitoring.
The local discharge diagnosis method of reactor based on multi-source feature fusion and deep network is adopted, including data preprocessing, multi-dimensional feature extraction, SWFK-DPC clustering algorithm, deep multi-origin adaptive convolutional neural network (EDMSACNN), cross-domain feature compression, cluster-guided dual-channel attention mechanism and adversarial enhancement classification module.
It significantly improves the accuracy and robustness of local discharge monitoring of reactors, can effectively deal with noise and environmental changes, and improves the stability of diagnosis and engineering applicability.
Smart Images

Figure CN120011863A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of partial discharge detection, and in particular to a method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network. Background Art
[0002] Reactors are key equipment in power systems. Their internal insulation defects can induce partial discharge (PD), accelerate insulation aging and threaten equipment safety. Traditional offline detection methods are limited by their long cycle and lack of real-time performance, which has promoted the partial discharge feature analysis technology based on online monitoring to become a research hotspot. However, the nonlinear and non-stationary characteristics of PD signals and the complexity of high-dimensional feature space pose severe challenges to traditional fault diagnosis methods.
[0003] At present, the partial discharge monitoring of reactors mainly relies on single physical quantity for pattern recognition, which is difficult to fully reflect the essential characteristics of partial discharge, and is sensitive to noise and changes in working conditions, resulting in limited diagnostic accuracy. In addition, the analysis method based on statistical features often assumes that the signal obeys a specific distribution, which is difficult to deal with complex nonlinear and non-stationary characteristics. Therefore, it is necessary to deeply integrate multidimensional physical quantities to fully characterize the spatiotemporal evolution characteristics of partial discharge signals. Traditional dimensionality reduction technology mainly relies on linear projection, which is difficult to capture the nonlinear correlation between different characteristic modes, resulting in information loss and decreased pattern discrimination. In addition, although deep learning methods have powerful feature extraction capabilities, they often fail to fully explore the complementary information of physical and statistical features in multimodal fusion, which limits the accurate modeling of complex discharge patterns. On the other hand, the distribution of partial discharge signals is affected by environmental factors and equipment aging, and there is a significant shift in the data distribution under different working conditions. The existing models generally rely on fixed feature extraction methods and lack the ability to adapt to domain shifts, resulting in poor generalization across working conditions, which in turn affects the stability of diagnosis and engineering applicability. Therefore, there is an urgent need to build an intelligent diagnosis method that has both multi-dimensional feature fusion capabilities and cross-domain adaptability to improve the accuracy and robustness of partial discharge monitoring.
[0004] Density peak clustering (DPC) is a density-based clustering algorithm. It has good adaptability to non-spherical clustering structures and can intuitively identify cluster centers through decision diagrams because it does not require a preset number of clusters. It has been widely used in many fields. Especially in partial discharge signal analysis, anomaly detection and high-dimensional data clustering tasks, DPC relies on the joint measurement of local density and relative distance, and can effectively mine the intrinsic structure of data. However, the traditional DPC algorithm exposes a series of limitations when processing high-dimensional data. First, standard similarity metrics, such as Euclidean distance, fail to fully consider the different contributions of different features to the overall distance, especially when the feature scales differ greatly, which may lead to inaccurate and unstable clustering results. Second, the DPC algorithm is sensitive to noise, especially in high-dimensional space, where outliers interfere with local density calculations and thus affect the identification of cluster centers. In addition, when the differences between data features are small, DPC is difficult to effectively distinguish similar samples, resulting in a decrease in clustering accuracy. Therefore, in response to these problems, it is urgent to improve the traditional DPC algorithm to enhance its adaptability and robustness in high-dimensional data. In addition, existing deep learning models still have significant deficiencies in multimodal feature processing and robustness improvement. When faced with multimodal features, traditional convolutional neural networks (CNNs) usually adopt simple splicing or early fusion strategies. This method fails to fully tap the high-order interactive information between features, thereby limiting the model's ability to represent complex patterns. In addition, the current mainstream attention mechanisms, such as channel attention, mainly focus on single-level feature selection, ignoring the coordinated optimization of global and local information, resulting in limited expressiveness of the model when processing complex signals. More importantly, local discharge signals are often polluted by electromagnetic interference and noise. Traditional classification models are highly sensitive to these adversarial disturbances and are difficult to effectively cope with the challenges brought by noise and distribution shift. Therefore, how to improve the robustness of the model to noise and environmental changes while ensuring high-precision classification has become a key issue that needs to be solved urgently. Summary of the invention
[0005] The purpose of the present invention is to provide a method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network, so as to solve the problem of low accuracy and robustness of partial discharge monitoring caused by the lack of multi-dimensional feature fusion capability and cross-domain adaptability in existing partial discharge monitoring algorithms.
[0006] The technical solution adopted by the present invention to solve the technical problem is: a method for diagnosing partial discharge of a reactor based on multi-source feature fusion and deep network, comprising the following steps:
[0007] S1, data preprocessing and feature extraction;
[0008] Perform Min-Max normalization processing and missing value filling on the collected reactor partial discharge signal;
[0009] S2, extracting multi-dimensional feature information, including phase feature, polarity feature, amplitude feature, time-frequency feature and wavelet packet feature;
[0010] S3, processing high-dimensional feature sets through SWFK-DPC clustering algorithm;
[0011] S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN to achieve deep fusion of the physical and statistical features of partial discharge signals;
[0012] S5. Introduce cross-domain feature compression to improve the cross-domain adaptability of the model; introduce a clustering-guided dual-path attention mechanism to optimize the feature weight distribution in space and channels;
[0013] S6. Introduce adversarial enhancement and classification modules to improve the classification accuracy and anti-interference ability of the model;
[0014] S7, model optimization and iterative training;
[0015] After the model is built, training and optimization are carried out on the partial discharge mode of the reactor under complex working conditions. The model stability and high recognition are improved through hyperparameter tuning, loss function optimization and model iterative training.
[0016] Furthermore, in step S2, after performing three-layer wavelet packet decomposition on the signal, the energy proportions of the eight frequency bands are extracted, and the calculation formula is: (1); where E pk represents the energy proportion of the kth frequency band, x pk is the wavelet packet coefficient sequence of the corresponding frequency band, and x(i) is the original signal sequence.
[0017] Furthermore, step S3 specifically includes the following steps:
[0018] S3.1, use standard deviation weighted distance metric to calculate the similarity between data points;
[0019] S3.2, calculate the local density of each data point, and build a decision graph based on the local distance, identify the density peak points, and use these density peak points as cluster centers;
[0020] S3.3. When the maximum neighborhood radius of data point i exceeds the average maximum neighborhood radius of all samples, the point is judged as an outlier, otherwise it is a non-outlier and the non-outlier is assigned to the optimal cluster. For unassigned non-outliers, the fuzzy weighted K nearest neighbor method FK is introduced and a DPC clustering allocation strategy is proposed.
[0021] Furthermore, in step S3.1, the standard deviation weighted distance between data points i and j is calculated (2); among which, is the weight factor, m is the number of features for each sample, x ik is the value of sample i on the kth feature, x jk is the value of sample j on the kth feature.
[0022] Furthermore, in step S3.2, the local density of each data point i is defined as (3), where KNN i is the K nearest neighbor set of data point i, and the local distance Defined as: (4) On this basis, a decision graph is constructed to identify the density peak, with the x-axis representing the local density and the y-axis representing the local distance. By observing the position of the points in the decision graph, the density peak located in the upper right corner can be intuitively identified.
[0023] Furthermore, in step S3.3, based on the standard deviation weighted distance d ij The metric proposes the outlier determination criteria, specifically: through the formula (5) Define the criteria for determining outliers, where: is the maximum neighborhood radius of data point i in its K nearest neighbor set, is the mean of the maximum neighborhood radius of all sample points in the data set, (6), (7), n is the total number of samples in the data set; the probability that data point i belongs to cluster c is (8), where is the weighted similarity measure, (9); is the normalization coefficient; (10);
[0024] For the dataset , calculate the membership of each distribution point i to the number of clusters c , and construct the matrix S, (11), where r is the number of points to be assigned and C is the number of clusters. The data point s with the highest membership is selected and assigned to the corresponding cluster k. Its membership in all clusters is reset to zero. The membership of adjacent data points is updated based on their K nearest neighbor relationships. The process is iterated until all data points are assigned.
[0025] Furthermore, the specific steps of step S4 include:
[0026] S4.1, interactively fuse the phase features, polarity features and amplitude features through outer product operation, Hadamard product and linear weighted fusion method;
[0027] S4.2. Use short-time Fourier transform (STFT) and discrete wavelet transform (DWT) to jointly encode the time-frequency features, thereby characterizing the time-frequency domain dynamic characteristics of the partial discharge signal.
[0028] S4.3. Based on the wavelet packet features, the tree graph convolutional network GCN is used to further model and capture the relationship between nodes with different energy.
[0029] Furthermore, in step S4.1, the five-dimensional feature vector output by SWFK-DPC clustering is reconstructed into a three-channel feature tensor. For the phase information φ, polarity p and amplitude A, outer product operation, Hadamard product and linear weighted fusion are used. The mathematical expression is: (12); where F phy is the physical feature after fusion; W p is the learnable parameter matrix corresponding to polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function, and nonlinear mapping is introduced to enhance the model's expressiveness.
[0030] Furthermore, in step S4.2, short-time Fourier transform STFT is combined with time domain convolution to perform joint coding, and the calculation formula is: (13); where F tfr is the fused time-frequency feature; TFR represents the time-frequency matrix obtained after short-time Fourier transform; DWT is discrete wavelet transform; TFR envelope Represents the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D represents the time series feature modeling through sliding window operation.
[0031] Furthermore, in step S4.3, a tree graph convolutional network GCN is used for modeling, and the mathematical expression is: (14); where F wpe is the fused wavelet packet feature, WPE k is the wavelet packet feature of the kth level, is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: (15); where Q is the learnable query vector, the softmax function implements weight normalization, and WPE T k is the k-th level wavelet packet feature after adaptive weighting.
[0032] Furthermore, in step S5, the cross-domain feature compression first aligns the features and calculates the SWFK-DPC cluster center C k With the convolution feature F i The Sinkhorn distance between (16); where M() is the Mahalanobis distance, is the entropy regularization term, is the optimal transmission mapping matrix, is the regularization coefficient, is from The normalized weights are calculated; after feature alignment is completed, a dynamic channel compression strategy based on channel attention is introduced to select feature subsets with key discriminant information in the high-dimensional feature space; according to the normalized weight γ ik Calculate channel attention weights (17); where MLP ( ) is a multi-layer perceptron; the tensor shrinking operation is used to deeply fuse the physical features, time-frequency features and wavelet packet features, and the fused features are calculated as: (18); where w i ,w j are the attention weights of different feature modalities.
[0033] Furthermore, the dual-path attention mechanism includes a local attention branch, a global attention branch, and a dual-path fusion. The local attention branch optimizes the feature representation using spatial clustering information to highlight the salient features of the target area and suppress background noise. First, the membership matrix U of the feature points is calculated based on the SWFK-DPC algorithm to describe the degree of attribution of the input features to different cluster centers. On this basis, the spatial attention map A is calculated. local Adjust the spatial weight of the feature map, calculated as: (19); where W l is a learnable parameter matrix; the global attention branch optimizes channel attention by using the overall distribution information of the cluster and introduces the cluster compactness index S k As a channel attention adjustment factor, the global attention weight (20); where C k is the corresponding cluster center; the fusion method of the dual-path attention fusion stage is: (21); where λ is the adaptive fusion coefficient, which is calculated by the multi-layer perceptron MLP: (twenty two).
[0034] Furthermore, the specific steps of step S6 include:
[0035] S6.1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discrimination ability, and use constraints to combat disturbances to improve the stability of the model under noise or data distribution shift. First, map the fused features to a low-dimensional space, as shown below: (23); where W d is the dimension reduction matrix; in the dimension reduction feature Fcls The adversarial perturbation δ constrained by L2 is introduced to obtain the adversarial enhanced feature representation: (24); L2 constraint is to impose a penalty term of L2 norm on model parameters during optimization to prevent model parameters from being too large, thereby improving generalization ability and stability, b d is the bias term introduced in the feature dimensionality reduction process to adjust the feature distribution after dimensionality reduction. fusion is the feature representation after fusion;
[0036] S6.2. Based on the adversarial enhancement feature, the Softmax function is used to calculate the category probability distribution, and the model is optimized by the cross entropy loss and the adversarial consistency constraint.
[0037] The category probability distribution P is calculated by the following formula, (25), where W c is the classification weight vector, b c is the bias term of the classifier, b k is the cluster center C k The bias term, W c T is the transposed matrix of the classification weight vector, W k T is the cluster center C k The feature weight matrix is transposed; the cross entropy loss and the adversarial consistency constraint constitute the loss function: (26); is a hyperparameter, y c Represents the true category label.
[0038] The beneficial effects of the present invention are as follows: the present invention realizes nonlinear interaction of multi-source features through tensor contraction and attention mechanism, breaks through the limitations of traditional linear fusion, and effectively alleviates the data distribution differences under different working conditions based on the feature alignment technology of optimal transmission, significantly improving the adaptive cross-domain capability. In addition, the adversarial training and cluster-guided attention mechanism are combined to enhance the robustness of the model in a noisy environment and ensure the reliability of diagnosis. This technology provides an intelligent solution for the insulation status assessment of reactors, and has important engineering application value for the safe and stable operation of power systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a framework diagram of the SWFK-DPC clustering algorithm of the present invention;
[0040] Figure 2 This is a diagram of the deep multi-source adaptive convolutional neural network EDMSACNN algorithm architecture of the present invention;
[0041] Figure 3This is the overall model framework diagram based on SWFK-DPC-EDMSACNN of the present invention. DETAILED DESCRIPTION
[0042] The following describes in detail the reactor partial discharge diagnosis method based on multi-source feature fusion and deep network of the present invention in combination with the figure.
[0043] S1. Data preprocessing and feature extraction.
[0044] The collected partial discharge signals of the reactor are subjected to Min-Max normalization processing and missing value filling to ensure the consistency and integrity of the data.
[0045] S2. Extract multi-dimensional feature information, including phase features, polarity features, amplitude features, time-frequency features, and wavelet packet features, to fully capture the dynamic and frequency domain characteristics of the local discharge signal. Among them, the phase feature reflects the changing characteristics of the local discharge signal in periodic behavior; the polarity feature describes the polarity change of the local discharge signal within a cycle; the amplitude feature includes peak value, average value, and root mean square value statistics, which are used to describe the amplitude information of the local discharge signal; the time-frequency feature is used to describe the distribution characteristics of the local discharge signal in the time domain and frequency domain, specifically including rise time, signal width, fall time, main frequency, equivalent time length, and equivalent bandwidth; the wavelet packet feature extracts the energy distribution information of the signal in different frequency bands through wavelet packet decomposition. Specifically, after the signal is decomposed into 3 layers of wavelet packets, the energy proportion of 8 frequency bands is extracted. The calculation formula is: (1). Where, E pk is the energy proportion of the kth frequency band, x pk is the wavelet packet coefficient sequence of the corresponding frequency band, and x(i) is the original signal sequence.
[0046] S3, such as Figure 1 As shown, the high-dimensional feature set is processed by the SWFK-DPC clustering algorithm.
[0047] S3.1. Use the standard deviation weighted distance metric to calculate the similarity between data points.
[0048] Calculate the standard deviation weighted distance between data points i and j (2). Among them, is the weight factor, m is the number of features for each sample, x ik is the value of sample i on the kth feature, x jk is the value of sample j on the kth feature. ij The metric takes into account the differences between different features, so that the differences between features have a reasonable impact on the final distance calculation.
[0049] S3.2. By calculating the local density of each data point and constructing a decision graph based on the local distance, the density peak points are identified and these density peak points are used as cluster centers.
[0050] Define the local density of each data point i (3). Among them, KNN i Represents the K nearest neighbor set of data point i. Local distance Defined as: (4). On this basis, a decision graph is constructed to identify the density peak, with the x-axis representing the local density and the y-axis representing the local distance. By observing the position of the points in the decision graph, the density peak located in the upper right corner can be intuitively identified.
[0051] In order to effectively divide the remaining sample points in the data set into non-outlier points and outlier points, the weighted distance d is calculated based on the standard deviation. ij The outlier determination criterion is: (5) Define the criteria for determining outliers, where: is the maximum neighborhood radius of data point i in its K nearest neighbor set, is the mean of the maximum neighborhood radius of all sample points in the data set. (6), (7), n is the total number of samples in the dataset.
[0052] S3.3. For unassigned non-outlier points, the fuzzy weighted K nearest neighbor method FK is introduced to further optimize the clustering effect.
[0053] Based on the above judgment criteria, non-outlier points are assigned to the optimal cluster. When the maximum neighborhood radius of data point i exceeds the average maximum neighborhood radius of all samples, the point is judged as an outlier. For unassigned non-outlier points, the fuzzy weighted K nearest neighbor method FK is introduced, and an improved DPC clustering allocation strategy is proposed. This strategy optimizes the allocation process of density peaks by calculating the fuzzy membership of data points, thereby ensuring that non-density peaks can be more accurately attributed to the optimal cluster center. Specifically, the probability p of data point i belonging to cluster c is i c It is determined by the membership of the assigned points in its K nearest neighbors, which is calculated as follows: (8). Among them, is the weighted similarity measure, (9) Ensure the reasonable distribution of membership. is the normalization coefficient, (10).
[0054] For the dataset , calculate the membership of each distribution point i to the number of clusters c , and construct the matrix S, (11). Where r is the number of points to be assigned and C is the number of clusters. Select the data point s with the highest membership and assign it to the corresponding cluster k, while setting its membership in all clusters to zero. Update the membership of adjacent data points based on their K nearest neighbor relationships, and iterate the process until all data points are assigned.
[0055] S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN.
[0056] Aiming at the high-dimensional feature set after SWFK-DPC clustering, an improved deep multi-source adaptive convolutional neural network algorithm Enhanced DMSACNN is proposed. Its core innovation lies in building a hybrid coding architecture with multi-feature interaction to achieve deep fusion of the physical and statistical features of partial discharge signals. Figure 2 The obtained multi-dimensional features are input into EDMSACNN for further analysis and pattern recognition. EDMSACNN adopts a hybrid coding architecture with multi-feature interaction, including the deep fusion of physical features, time-frequency features and wavelet packet features.
[0057] S4.1. The phase features, polarity features and amplitude features are interactively fused through outer product operation, Hadamard product and linear weighted fusion method.
[0058] The five-dimensional feature vector output by SWFK-DPC clustering is reconstructed into a three-channel feature tensor. For the phase information φ, polarity p and amplitude A, the outer product operation, Hadamard product and linear weighted fusion are used. The mathematical expression is: (12). Among them, F phy is the physical feature after fusion; W p is the learnable parameter matrix corresponding to polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function, and nonlinear mapping is introduced to enhance the model's expressiveness.
[0059] S4.2. Use short-time Fourier transform (STFT) and discrete wavelet transform (DWT) to jointly encode the time-frequency features, thereby characterizing the time-frequency domain dynamic characteristics of the partial discharge signal.
[0060] In order to make full use of the time-frequency domain information, short-time Fourier transform (STFT) combined with time-domain convolution is used for joint encoding to characterize the time-frequency domain dynamic characteristics of the PD signal. The calculation formula is: (13). Among them, F tfris the fused time-frequency feature; TFR is the time-frequency matrix obtained after short-time Fourier transform; DWT is discrete wavelet transform; TFR envelope It is the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D is used to realize time series feature modeling through sliding window operation.
[0061] S4.3. Based on the wavelet packet features, the tree graph convolutional network GCN is used to further model and capture the relationship between nodes with different energy.
[0062] Considering the contribution of wavelet packet capability WPE to PD discharge mode, a tree graph convolution network GCN is used to model the relationship between different energy nodes, which is mathematically expressed as: (14). Among them, F wpe is the fused wavelet packet feature, WPE k is the wavelet packet feature of the kth level, is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: (15). Where Q is a learnable query vector and the softmax function implements weight normalization to achieve adaptive weighting. T k Represents the k-th level wavelet packet feature after adaptive weighting.
[0063] S5. Cross-domain feature compression and introduction of attention mechanism.
[0064] In order to improve the cross-domain adaptability of the model, a cross-domain feature compression method based on optimal transfer theory is introduced. This method measures the similarity of feature distribution by calculating the Sinkhorn distance metric and aligns the features. After feature alignment, a dynamic channel compression strategy based on the channel attention mechanism is adopted to select feature subsets with key discriminant information. In addition, a cluster-guided two-way attention mechanism is proposed, which includes local and global attention branches to optimize the feature weight distribution in space and channels.
[0065] In order to optimize feature expression and improve cross-domain adaptability, a cross-domain feature compression method based on optimal transmission theory is introduced. This method first aligns the features and calculates the SWFK-DPC cluster center C k With the convolution feature F i The Sinkhorn distance between them is used to measure the similarity of feature distributions. (16). Where M() is the Mahalanobis distance, is the entropy regularization term, is the optimal transmission mapping matrix, is the regularization coefficient, is from The calculated normalized weights.
[0066] After feature alignment is completed, a dynamic channel compression strategy based on channel attention is introduced to select feature subsets with key discriminant information in the high-dimensional feature space. ik Calculate channel attention weights (17). Where MLP( ) is a multi-layer perceptron.
[0067] In order to further enhance the information interaction capability between different feature modes, tensor shrinkage operation is used to achieve deep fusion of physical features, time-frequency features and wavelet packet features. The fusion feature is calculated as: (18). Among them, w i ,w j are the attention weights of different feature modalities.
[0068] In order to further enhance the feature fusion effect and strengthen the synergy of global and local information while extracting multi-scale features, a clustering-guided dual-path attention mechanism is proposed. This mechanism includes a local attention branch, a global attention branch, and dual-path fusion, which effectively improves the discriminability and robustness of feature expression through the collaborative modeling of space and channels.
[0069] The local attention branch optimizes the feature representation using spatial clustering information to highlight the salient features of the target area and suppress background noise. First, the membership matrix U of the feature points is calculated based on the SWFK-DPC algorithm to describe the degree of belonging of the input features to different cluster centers. On this basis, the spatial attention map A is calculated. local To adjust the spatial weight of the feature map, the calculation method is: (19). Among them, W l It is a learnable parameter matrix designed to capture spatial correlation, while the softmax operation ensures the normalization of attention distribution, allowing the model to pay more attention to the salient areas indicated by the clustering structure. This mechanism enables the network to adaptively adjust feature weights according to the spatial clustering pattern, thereby enhancing the expressiveness of the target area and improving the distinguishability of the features.
[0070] The global attention branch focuses on optimizing channel attention by utilizing the overall distribution information of the cluster to enhance the feature selection capability in the channel dimension. To this end, the cluster compactness index S is introduced. k As a channel attention adjustment factor, the global attention weight (20). Among them, S k reflects the closeness of the k-th cluster, that is, the level of similarity within the category; C kIndicates the corresponding cluster center. In the dual-path attention fusion stage, in order to fully combine local and global attention information, a gating mechanism is used for dynamic fusion, making the final feature expression richer and more discriminative. The fusion method is: Among them, λ is the adaptive fusion coefficient, which is calculated by the multi-layer perceptron MLP: (twenty two).
[0071] S6. Adversarial enhancement and classification module design.
[0072] In order to improve the classification accuracy and anti-interference ability of the model, the adversarial enhancement classification module is introduced.
[0073] S6.1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discrimination capability, and use constrained adversarial perturbations to improve the stability of the model under noise or data distribution shift.
[0074] After completing feature fusion and attention weighting, the classification module is designed to improve the accuracy and robustness of pattern recognition. First, the fused features are mapped to a category-related low-dimensional space to enhance the discrimination ability and suppress redundant information. The feature transformation is completed through the fully connected layer, which is formally expressed as follows: (23). Among them, W d is the dimension reduction matrix. In order to improve the anti-interference ability of the model, the dimension reduction feature F cls The adversarial perturbation δ constrained by L2 is introduced to obtain the adversarial enhanced feature representation: (24). The L2 constraint is a penalty term imposed on the model parameters in the optimization process to prevent the model parameters from being too large, thereby improving the generalization ability and stability. d is the bias term introduced in the feature dimensionality reduction process to adjust the feature distribution after dimensionality reduction. fusion It is the feature representation after fusion.
[0075] S6.2. Based on the adversarial enhanced features, the Softmax function is used to calculate the category probability distribution, and the model is optimized jointly by the cross entropy loss and the adversarial consistency constraint.
[0076] This perturbation enables the model to maintain stable discriminative performance when confronting uncertain factors, such as noise or data distribution offset. Based on the adversarial enhancement features, the Softmax function is used to calculate the category probability distribution (25). Among them, W c is the classification weight vector, and the SWFK-DPC cluster center C k Form an implicit alignment relationship, so that the classifier can further combine the pattern structure information based on the deep features to improve the generalization ability; b cis the bias term of the classifier, which is used to adjust the category probability distribution of Softmax output; b k is the cluster center C k The bias term is used to optimize the matching relationship between features and cluster centers; W c T is the transposed matrix of the classification weight vector, which is used to calculate the category probability score; W k T is the cluster center C k The feature weight matrix of is transposed to characterize the mapping relationship between deep features and pattern structures to improve classification performance. The loss function consists of cross entropy loss and adversarial consistency constraints to optimize classification accuracy while improving the robustness of the model: (26). The cross entropy term optimizes the classification decision, and the adversarial consistency constraint term ensures the consistency of feature distribution before and after the perturbation, thereby suppressing the impact of adversarial samples. Hyperparameters Balance classification accuracy and anti-interference ability, so that the model can maintain high stability and high recognition in complex environments, c Represents the true category label.
[0077] S7. Model optimization and iterative training.
[0078] After the model is built, training and optimization are carried out for the partial discharge mode of the reactor under complex working conditions. The performance of the model is further improved through hyperparameter tuning, loss function optimization and model iterative training. Especially in complex environments, the combination of adversarial training and dual-path attention mechanism ensures the high stability and high recognition of the model in practical applications.
[0079] The present invention realizes nonlinear interaction of multi-source features through tensor contraction and attention mechanism, breaks through the limitations of traditional linear fusion, and effectively alleviates the data distribution differences under different working conditions based on the feature alignment technology of optimal transmission, significantly improving the adaptive cross-domain capability. In addition, the attention mechanism guided by adversarial training is combined to enhance the robustness of the model in a noisy environment and ensure the reliability of diagnosis. This technology provides an intelligent solution for the insulation status assessment of reactors, and has important engineering application value for the safe and stable operation of power systems.
Claims
1. A reactor partial discharge diagnosis method based on multi-source feature fusion and deep network, characterized in that: The following steps are involved: S1, data preprocessing and feature extraction; Perform Min-Max normalization processing and missing value filling on the collected reactor partial discharge signal; S2, extracting multi-dimensional feature information, including phase feature, polarity feature, amplitude feature, time-frequency feature and wavelet packet feature; S3, processing high-dimensional feature sets through SWFK-DPC clustering algorithm; S4. Construct a deep multi-source adaptive convolutional neural network EDMSACNN to achieve deep fusion of the physical and statistical features of partial discharge signals; S5, introduces a dual-path attention mechanism guided by cross-domain feature compression and clustering; S6. Introduce adversarial enhancement and classification modules to improve the classification accuracy and anti-interference ability of the model; S7, model optimization and iterative training; After the model is built, training and optimization are carried out on the partial discharge mode of the reactor under complex working conditions. The model stability and high recognition are improved through hyperparameter tuning, loss function optimization and model iterative training.
2. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 1 is characterized in that: In step S2, after performing three-layer wavelet packet decomposition on the signal, the energy proportions of the eight frequency bands are extracted, and the calculation formula is: (1); Among them, E pk represents the energy proportion of the kth frequency band, x pk is the wavelet packet coefficient sequence of the corresponding frequency band, and x(i) is the original signal sequence.
3. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 2 is characterized in that: Step S3 specifically includes the following steps: S3.1, use standard deviation weighted distance metric to calculate the similarity between data points; S3.2, calculate the local density of each data point, and build a decision graph based on the local distance, identify the density peak points, and use these density peak points as cluster centers; S3.
3. When the maximum neighborhood radius of data point i exceeds the average maximum neighborhood radius of all samples, the point is judged as an outlier, otherwise it is a non-outlier and the non-outlier is assigned to the optimal cluster. For unassigned non-outliers, the fuzzy weighted K nearest neighbor method FK is introduced and a DPC clustering allocation strategy is proposed.
4. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 3 is characterized in that: In step S3.1, the standard deviation weighted distance between data points i and j is calculated (2); among which, is the weight factor, m is the number of features for each sample, x ik is the value of sample i on the kth feature, x jk is the value of sample j on the kth feature.
5. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 4 is characterized in that: In step S3.2, the local density of each data point i is defined as (3), where KNN i is the K nearest neighbor set of data point i, and the local distance Defined as: (4) On this basis, a decision graph is constructed to identify the density peak, with the x-axis representing the local density and the y-axis representing the local distance. By observing the position of the points in the decision graph, the density peak located in the upper right corner can be intuitively identified.
6. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 5, characterized in that: In step S3.3, the distance d is weighted based on the standard deviation ij The outlier determination criteria are proposed by the metric: (5) Define the criteria for determining outliers, where: is the maximum neighborhood radius of data point i in its K nearest neighbor set, is the mean of the maximum neighborhood radius of all sample points in the data set, (6), (7), n is the total number of samples in the data set; the probability that data point i belongs to cluster c is (8), where is the weighted similarity measure, (9); is the normalization coefficient; (10); For the dataset , calculate the membership of each distribution point i to the number of clusters c , and construct the matrix S, (11), where r is the number of points to be assigned and C is the number of clusters. The data point s with the highest membership is selected and assigned to the corresponding cluster k. Its membership in all clusters is reset to zero. The membership of adjacent data points is updated based on their K nearest neighbor relationships. The process is iterated until all data points are assigned.
7. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 6, characterized in that: The specific steps of step S4 include: S4.1, interactively fuse the phase features, polarity features and amplitude features through outer product operation, Hadamard product and linear weighted fusion method; S4.2, using short-time Fourier transform STFT and discrete wavelet transform DWT to jointly encode the time-frequency features, so as to characterize the time-frequency domain dynamic characteristics of the partial discharge signal; S4.
3. Based on the wavelet packet features, the tree graph convolutional network GCN is used to further model and capture the relationship between nodes with different energy.
8. The method for diagnosing partial discharge of reactors based on multi-source feature fusion and deep network according to claim 7 is characterized in that: In step S4.1, the five-dimensional feature vector output by SWFK-DPC clustering is reconstructed into a three-channel feature tensor. For the phase information φ, polarity p and amplitude A, the outer product operation, Hadamard product and linear weighted fusion are used. The mathematical expression is: (12); where F phy is the physical feature after fusion; W p is the learnable parameter matrix corresponding to polarity; W φ is the learnable parameter matrix corresponding to the phase; W A is the learnable parameter matrix corresponding to the amplitude; ReLU is used as the activation function, and nonlinear mapping is introduced to enhance the model's expressiveness.
9. The reactor partial discharge diagnosis method based on multi-source feature fusion and deep network according to claim 8 is characterized in that: In step S4.2, short-time Fourier transform (STFT) is combined with time domain convolution to perform joint coding, and the calculation formula is: (13); Among them, F tfr is the fused time-frequency feature; TFR represents the time-frequency matrix obtained after short-time Fourier transform; DWT is discrete wavelet transform; TFR envelope Represents the time-frequency envelope, which is used to extract the envelope information of the signal; Conv1D represents the time series feature modeling through sliding window operation.
10. The reactor partial discharge diagnosis method based on multi-source feature fusion and deep network according to claim 9, characterized in that: In step S4.3, the tree graph convolutional network GCN is used for modeling, and the mathematical expression is: (14); Among them, F wpe is the fused wavelet packet feature, WPE k is the wavelet packet feature of the kth level, is the weight coefficient of the wavelet packet energy, which is adaptively calculated through the softmax mechanism: (15); where Q is the learnable query vector, the softmax function implements weight normalization, and WPE T k is the k-th level wavelet packet feature after adaptive weighting.
11. The reactor partial discharge diagnosis method based on multi-source feature fusion and deep network according to claim 10, characterized in that: In step S5, cross-domain feature compression first aligns the features and calculates the SWFK-DPC cluster center C k With the convolution feature F i The Sinkhorn distance between (16); where M() is the Mahalanobis distance, is the entropy regularization term, is the optimal transmission mapping matrix, is the regularization coefficient, is from The normalized weights are calculated; after feature alignment is completed, a dynamic channel compression strategy based on channel attention is introduced to select feature subsets with key discriminant information in the high-dimensional feature space; according to the normalized weight γ ik Calculate channel attention weights (17); where MLP ( ) is a multi-layer perceptron; the tensor shrinking operation is used to deeply fuse the physical features, time-frequency features and wavelet packet features, and the fused features are calculated as: (18); where w i ,w j are the attention weights of different feature modalities.
12. The reactor partial discharge diagnosis method based on multi-source feature fusion and deep network according to claim 11, characterized in that: The dual-path attention mechanism includes local attention branch, global attention branch and dual-path fusion. The local attention branch optimizes the feature representation by using spatial clustering information to highlight the salient features of the target area and suppress background noise. First, the membership matrix U of the feature points is calculated based on the SWFK-DPC algorithm to describe the degree of belonging of the input features to different cluster centers. On this basis, the spatial attention map A is calculated. local Adjust the spatial weight of the feature map, calculated as: (19); Among them, W l is a learnable parameter matrix; the global attention branch optimizes channel attention by using the overall distribution information of the cluster and introduces the cluster compactness index S k As a channel attention adjustment factor, the global attention weight (20); where C k is the corresponding cluster center; the fusion method of the dual-path attention fusion stage is: (21); where λ is the adaptive fusion coefficient, which is calculated by the multi-layer perceptron MLP: (twenty two).
13. The reactor partial discharge diagnosis method based on multi-source feature fusion and deep network according to claim 12, characterized in that: The specific steps of step S6 include: S6.
1. Map the fused features to a low-dimensional space through a fully connected layer to enhance the discrimination ability, and use constraints to combat disturbances to improve the stability of the model under noise or data distribution shift. First, map the fused features to a low-dimensional space, as shown below: (23); where W d is the dimension reduction matrix; in the dimension reduction feature F cls The adversarial perturbation δ constrained by L2 is introduced to obtain the adversarial enhanced feature representation: (24); L2 constraint is a penalty term that imposes L2 norm on model parameters during optimization, b d is the bias term introduced in the feature dimensionality reduction process to adjust the feature distribution after dimensionality reduction. fusion is the feature representation after fusion; S6.
2. Based on the adversarial enhancement feature, the Softmax function is used to calculate the category probability distribution, and the model is optimized by the cross entropy loss and the adversarial consistency constraint. The category probability distribution P is calculated by the following formula, (25), where W c is the classification weight vector, b c is the bias term of the classifier, b k is the cluster center C k The bias term, W c T is the transposed matrix of the classification weight vector, W k T is the cluster center C k The feature weight matrix is transposed; the cross entropy loss and the adversarial consistency constraint constitute the loss function: (26); is a hyperparameter, y c Represents the true category label.
Citation Information
Patent Citations
Fault adaptive diagnosis method fusing deep convolution and self-attention network
CN117235490A
MMC sub-module open-circuit fault diagnosis method based on multi-source fusion graph and SE-BiGRU-ResNet model
CN118501774A
Self-adaptive data processing and self-diagnosis method and system for partial discharge characteristics of KYN equipment
CN118965207A
Reactor partial discharge fault diagnosis method based on compressed sensing and dynamic Bayesian network
CN119025979A
Hyperspectral image classification method based on multi-scale convolution Fourier and double-branch self-attention
CN119445257A
Cited By
Switch cabinet partial discharge on-line monitoring method and storage medium
CN120408334A
A partial discharge online monitoring method and storage medium for switch cabinet
CN120408334B
Partial discharge intelligent identification method and system based on collaborative reasoning
CN120470462A
Partial discharge intelligent identification method and system based on collaborative reasoning
CN120470462B
Generator partial discharge on-line monitoring method, system, device and medium
CN120577653A