Partial discharge identification method based on feature set simulation evolution optimization

By simulating evolutionary optimization of feature sets and improving convolutional neural networks, the problems of insufficient effectiveness and versatility of feature construction in partial discharge pattern recognition of high-voltage cables are solved, and efficient and accurate partial discharge recognition is achieved.

CN120611236APending Publication Date: 2025-09-09HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510669333.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing methods for recognizing partial discharge patterns of high-voltage cables, the feature construction method has problems such as insufficient effectiveness, complicated data processing and insufficient versatility, resulting in low recognition accuracy and long calculation time.

Method used

A method based on feature set simulation evolutionary optimization is adopted. The feature construction process is transformed through genetic algorithm. Combined with partial discharge phase spectrum and grayscale processing, a high-dimensional feature set is constructed. An improved convolutional neural network is used for pattern recognition, including feature extraction and recognition.

Benefits of technology

The richness and effectiveness of the feature set are improved, the recognition accuracy and versatility are enhanced, the calculation amount and time are reduced, and fast and accurate partial discharge pattern recognition is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611236A_ABST
    Figure CN120611236A_ABST
Patent Text Reader

Abstract

The invention discloses a partial discharge identification method based on feature set simulation evolution optimization, which comprises the following steps: manually constructing defects to obtain partial discharge data, carrying out data preprocessing, graying and gridding to obtain a partial discharge phase spectrogram matrix, and carrying out feature parameter construction on the partial discharge phase spectrogram matrix to obtain an original feature set of the partial discharge phase spectrogram matrix; and constructing high-dimensional features by using a simulated evolution optimization method to obtain a feature set with relatively high importance, and performing pattern recognition by using an improved convolutional neural network to complete recognition of the type of the partial discharge signal. According to the method, pattern recognition is carried out through the feature set obtained through the new feature construction method, the features of the partial discharge phase spectrogram are described better, the pattern recognition algorithm can converge to the optimal solution more quickly, and the accuracy of the result can be effectively improved. The problems that traditional features are insufficient in effectiveness, the data size of multi-dimensional feature construction is large, steps are tedious, and feature combination is single are solved, the accuracy of pattern recognition is improved, and the recognition time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cable detection and testing, and in particular to the field of partial discharge pattern recognition for high-voltage cables. Background Art

[0002] As a key carrier of power transmission, high-voltage cables play an increasingly important role in modern power grids. Compared to traditional overhead transmission lines, high-voltage cables offer significant advantages, such as a smaller footprint, reduced environmental impact, and greater transmission capacity. They are widely used in urban power grids, industrial parks, and renewable energy power generation grids. The integration of a large number of renewable energy sources into the grid has placed higher demands on the flexibility and stability of power transmission, further highlighting the importance of high-voltage cables.

[0003] However, the long-term operation of high-voltage cables under high voltage, high current, and complex environmental conditions inevitably leads to insulation aging, partial discharge, and temperature anomalies, which can lead to cable failures. Once a high-voltage cable is damaged or fails, it will cause a large-scale power outage, which will not only severely impact industrial production and residents' lives, causing huge economic losses, but may also cause safety accidents such as fires. Partial discharge in cable faults is a key indicator of insulation degradation. Long-term undetected partial discharge can lead to insulation breakdown and cause large-scale power outages. Therefore, real-time monitoring of partial discharge signals in high-voltage cables and accurate identification of their types are crucial for preventing failures and extending equipment life.

[0004] Currently, there are two main feature construction methods for partial discharge pattern recognition of high-voltage cables:

[0005] The first method extracts various statistical features and wavelet features, constructs two-dimensional and three-dimensional features, performs feature optimization, and uses features with higher effectiveness for pattern recognition. However, in the actual solution process, since the feature construction method only has product and ratio features, and the multidimensional feature construction uses enumeration method, not only is the data processing cumbersome, but the results are also insufficient in richness and effectiveness.

[0006] Another method is to perform pattern recognition by discovering some new features to provide a new perspective for observing partial discharge. This method is generally more strategic, but the features obtained by this method do not have a good combination with previously recognized effective one-dimensional features such as statistical features and time-frequency features. They can only be used alone, resulting in general adaptability to the environment, inability to identify unknown defect types, and weak versatility.

[0007] Therefore, it is urgent to propose a method for constructing a local discharge discharge feature set that is simple, easy to implement, and has high recognition accuracy. By constructing a new phase spectrum feature construction method, the feature optimization process is optimized, the recognition accuracy is improved, the calculation time is shortened, and it has strong versatility. Summary of the Invention

[0008] The two existing methods have problems such as insufficient feature validity, complicated data processing, and lack of versatility when it comes to partial discharge pattern recognition. To this end, the present invention adopts a new feature construction method, innovatively draws on the evolution mechanism of genetic algorithm and transforms and applies it in feature construction, optimizes the feature optimization process, greatly improves the richness and effectiveness of the feature set, and enhances the versatility of the feature set. A partial discharge identification method based on feature set simulated evolutionary optimization is proposed: partial discharge data is obtained by artificially constructing defects, a partial discharge phase spectrum is obtained by partial discharge signal processing, signal features are constructed according to the partial discharge phase spectrum, the phase spectrum is gray-scaled, and then feature parameters are constructed to obtain signal statistical feature parameters such as average value, root mean square, discharge amount, rise time, signal width, fall time, standard deviation, skewness, kurtosis, crest factor, waveform factor, main frequency, T, W, energy, Ep1~Ep8, time-frequency parameters and wavelet parameters, thereby obtaining its original feature set; then, high-dimensional features are constructed using the simulated evolutionary optimization method, and pattern recognition is performed using artificial intelligence algorithms such as convolutional neural networks to complete the identification of partial discharge signal types.

[0009] The present invention provides a partial discharge identification method based on feature set simulation evolutionary optimization, comprising the following steps:

[0010] A partial discharge identification method based on feature set simulation evolutionary optimization, characterized by comprising the following steps:

[0011] S1. Artificially constructing partial discharge defects to obtain partial discharge signals: Typical partial discharge defects in high-voltage cables are constructed in the laboratory. These defects include ring cutting marks, semiconductor powder defects, stress cone displacement defects, and grinding irregularities. Pressure tests are then performed on these defects to obtain partial discharge signals.

[0012] S2. Partial discharge signal data preprocessing: Perform wavelet decomposition and noise reduction on the original partial discharge signal dataset obtained in S1, identify and filter out abnormal data, remove noise by setting a threshold, and normalize the partial discharge signal data;

[0013] S3. Processing the partial discharge signal to obtain a partial discharge phase spectrum and matrix: Corresponding the amplitude of the partial discharge signal after the normalization process in S2 to the phase angle of the power frequency voltage, traversing a complete voltage cycle to obtain a partial discharge phase spectrum; De-graying the partial discharge phase spectrum, gridding the spectrum, and determining the size of the matrix value corresponding to each grid according to the number of partial discharges in each grid to obtain a partial discharge phase spectrum matrix;

[0014] S4. Extract and construct signal features based on the partial discharge phase spectrum matrix: Input the partial discharge phase spectrum matrix result in S3 into the Inception layer for feature extraction to obtain a feature set, where the series number and convolution kernel size parameters are set in the Inception layer; the feature set includes statistical features, time-frequency features, and wavelet features. The statistical features and wavelet features are obtained by calculation. The statistical features include mean value Mean, variance VA, skewness Sk, kurtosis Ku, and root mean square. The time-frequency features include discharge amount, rise time, signal width, fall time, crest factor, waveform factor, main frequency, equivalent pulse time T, equivalent pulse width W, and energy. The wavelet features are obtained by decomposing the wavelet decomposition energy ratio Ep1 to Ep8 using multiple wavelet bases.

[0015] S5. Simulate evolution and optimize the feature set: There are N generations of evolution in total, and the size of N can be specified or the default value can be used. Before each generation of evolution, the random forest algorithm will rank the importance of all features in the feature set. The ranking is based on the data set of known partial discharge defect types and their corresponding feature values ​​obtained in the laboratory. After that, the simulated evolution process will be performed according to the importance ranking in each generation of evolution. After the evolution is completed, it is first determined whether the recognition accuracy has decreased. If it has decreased, the simulated evolution model parameters are changed. Otherwise, it is determined whether the evolution generation has reached N. If not, the simulated evolution is repeated. If it has reached N, the simulated evolution is terminated and the feature set is output.

[0016] S6. Pattern recognition through improved convolutional neural networks: The optimized feature set is divided into a training set and a test set; a one-dimensional improved convolutional neural network is used for pattern recognition. The feature extraction of the feature set after evolutionary optimization is performed by adding an Inception layer to the neural network. The neural network structure is improved based on GoogLeNet. During training, the numerical labels of the partial discharge defect types are used as training targets, and the training set in the optimized feature set is used as training data to complete the training of the improved convolutional neural network model based on GoogLeNet. When the test set data of the optimized feature set is input, the trained improved convolutional neural network model based on GoogLeNet will give the corresponding numerical labels of the partial discharge defect types for the data set as the result.

[0017] Furthermore, the normalization processing and wavelet decomposition in step 2 are specifically as follows:

[0018] S201: normalize the partial discharge signal using the standard fraction method, the calculation formula is:

[0019]

[0020]

[0021] In the formula, μ is the mean of the characteristic data, σ is the standard deviation; n is the total number of data, i is the data order, x is the i is the i-th data value, Z i is the i-th data after normalization;

[0022] S202: Perform wavelet decomposition on the partial discharge signal. Select the Daubechies (dbN) wavelet basis function and decompose the partial discharge signal into approximate coefficients and detail coefficients layer by layer. The specific formula is as follows:

[0023]

[0024] Among them A J is the approximate coefficient, D J is the detail coefficient, S is the signal function, J is the number of wavelet decomposition layers, and then based on the coefficient of the finest layer D1, the noise standard deviation ε is estimated by the median absolute error to calculate the universal threshold λ. The specific formula is as follows:

[0025]

[0026] Where L is the length of the partial discharge signal. The noise-reduced partial discharge signal is reconstructed by inverse wavelet transform. The specific formula is as follows:

[0027]

[0028] is the partial discharge signal function after processing, To use threshold processing D J The detail coefficient obtained.

[0029] Furthermore, the step 3 is specifically as follows:

[0030] S301: The unit of the partial discharge signal amplitude is pC, the phase of a complete voltage cycle is 0° to 360°, and a scatter plot is drawn with the phase as the independent variable and the product of the amplitude and polarity as the dependent variable to obtain a partial discharge phase spectrum;

[0031] S302: After grayscale processing is performed on the partial discharge phase spectrum, a white image with a black background is displayed. In order to perform feature extraction and pattern recognition and image segmentation, the image is filtered using a filter factor. In actual applications, to reduce the amount of calculation, a pixel value of 255 can be used as the filter function;

[0032] S303: The default grid matrix size is set to 360*100. The horizontal axis takes each degree of phase as one grid, and the vertical axis takes each 0.01 after normalization as one grid, with 50 positive and negative grids. The number of partial discharge signal points in the grid of row a and column b is the size of the matrix of row a and column b, and the partial discharge phase spectrum matrix is ​​obtained. If the classification accuracy of the result is insufficient, the matrix size can be changed to improve the accuracy, but the computational complexity will also increase.

[0033] Furthermore, the step 4 specifically includes:

[0034] S401: The Inception layer contains several convolutions that are performed simultaneously, including 1*1 convolution, 1*1 convolution and 3*3 convolution in series, 1*1 convolution and 5*5 convolution in series, and pooling layer and 1*1 convolution in series. The Inception layer is connected to itself in series multiple times. The convolution size B and the number of series connections C are appropriately adjusted according to the actual situation to improve recognition accuracy.

[0035] S402: The time-frequency parameters are obtained during the partial discharge signal acquisition. The characteristic parameters of the feature set are calculated, including the mean value, variance VA, skewness Sk, kurtosis Ku, and wavelet parameters. The specific formula is as follows:

[0036] Skewness is defined using the third-order moment. The calculation formula for skewness is:

[0037]

[0038] In the formula, S k is the skewness; μ3 is the third-order central moment; σ is the standard deviation;

[0039] When the statistical data is right-skewed, S k >0,S k The larger the value, the higher the degree of right skewness; when the statistical data is left-skewed, S k <0,S k The smaller the value, the higher the degree of left skewness; when the statistical data is symmetrically distributed, there is S k =0;

[0040] In actual calculations, although the definition formula is used for calculation, it is necessary to traverse each sample multiple times to calculate the mean and variance, which is very time-consuming when the sample size is large. Therefore, the 1st to 3rd order origin moments are often used for calculation:

[0041] μ=EX

[0042] σ 2 =E(X-EX) 2 =EX 2 -μ 2

[0043]

[0044] μ is the mean, E is the calculation symbol, EX is the mean of the calculated data X, σ is the standard deviation, S k is the skewness;

[0045] Taking the normal distribution as a reference, the kurtosis describes the steepness of the distribution shape. If b k <3, the distribution is said to have insufficient kurtosis. If b k >3, the distribution is said to have excessive kurtosis. If it is known that the distribution may deviate from the normal distribution in terms of kurtosis, the kurtosis can be used to test the normality of the distribution. The expression of kurtosis is as follows:

[0046]

[0047] b k is the kurtosis, n is the total number of data, m4 is the fourth-order sample central moment, m2 is the second-order central moment, and x is the sample variance. i is the i-th data value, is the sample mean.

[0048] Furthermore, the step 5 specifically includes:

[0049] S501: The following events occur in sequence during the simulated evolution of the feature set:

[0050] (1) Elimination: The container can only hold a maximum of M features, where M ranges from 50 to 100, with a default value of 50. Features with importance greater than M will be eliminated directly. If the total number of features exceeds M after multiplication in the step, this does not conflict with the elimination rule. Elimination only requires that the number of features cannot exceed M at the beginning of evolution, and these features will enter the elimination phase together after the next round of feature optimization.

[0051] (2) Reproduction: The top k features in terms of importance will be combined in pairs to produce their offspring. k is generally 0.2*M. The combination methods include but are not limited to product, division or arithmetic mean, and the probability of occurrence of these methods is specified. A pair of combinations can only produce one offspring. The name of the offspring is determined by the combination method of the parents, and the offspring will not disappear after the parents are combined.

[0052] (3) Mutation: Features that have not been eliminated but whose importance is in the latter t% cannot reproduce. They try to increase their own importance through mutation and will transform themselves into the pth power of the original feature, where p is a random number between 0 and 2. The name of the mutated feature is represented by the original feature name plus the pth power. The mutated feature will replace itself.

[0053] S502: When the features obtained after each round of evolution in the step are subjected to random forest learning, the prediction accuracy of these features for the type of partial discharge will be tested synchronously. If the prediction accuracy is found to have decreased, the parameters of the simulated evolution model will be automatically changed, including the proportion of the number of reproduction and mutation in the total features, the default probability of the three combination methods of reproduction, and the default range of variation of the parameter p during mutation, to ensure that the evolutionary environment will adapt to different feature sets and enhance the versatility of the method.

[0054] Furthermore, the step 6 specifically includes:

[0055] S601: The optimized feature set obtained above is randomly divided into a training set and a test set in a ratio of 5:3. The training target is set as the partial discharge defect type, which is represented by a number. The partial discharge defect type number of the test set is used to calculate the classification error.

[0056] S602: Use a one-dimensional improved convolutional neural network for pattern recognition, and improve the neural network structure based on GoogLeNet. During training, the digital label of the partial discharge defect type is used as the training target. The training set in the optimized feature set is used as the training data. An appropriate number of training times is set to stabilize the recognition rate. The training of the improved convolutional neural network model based on GoogLeNet is completed, and a trained improved convolutional neural network model based on GoogLeNet is obtained. When the test set data of the optimized feature set is input, the trained improved convolutional neural network model based on GoogLeNet will give the corresponding digital label of the partial discharge defect type of the data set as the result. Add several Inception layers to the convolutional neural network, and the number of layers can be changed as needed. After extracting features from the Inception layer, input the activation function layer, and finally input the fully connected layer for pattern recognition. The obtained classification error is fed back to the spectrum matrix feature extraction and construction. Change the convolution size B and the number of series C in the Inception1 layer to improve the feature extraction efficiency, and change the grid accuracy parameters of the gridding of the partial discharge phase spectrum image. To improve the overall accuracy of the improved convolutional neural network model based on GoogLeNet.

[0057] This invention utilizes a novel method for constructing partial discharge signal features for pattern recognition. This method better characterizes the characteristics of the partial discharge phase spectrum, reduces the computational effort involved in feature construction, enables the pattern recognition algorithm to converge to the optimal solution more quickly, and effectively improves the accuracy of the results. This method eliminates redundant feature construction calculations, resulting in rich and effective features. Furthermore, feature construction parameters can be varied to adapt to different circumstances. Therefore, the method offers broad applicability, simplifying practical application while improving algorithm accuracy and computational speed.

[0058] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0059] 1. The present invention constructs high-dimensional features without the need for enumeration and feature optimization, effectively avoiding data redundancy and greatly reducing the amount of calculation. Taking M as 50 as an example, the amount of calculation for new features is reduced from 2500 to 110.

[0060] 2. The present invention uses a variety of methods to construct high-dimensional features. Compared with traditional product features and ratio features, this method can construct multivariate feature combinations such as product, ratio, average, exponent, logarithm, etc., which greatly improves the richness of features.

[0061] 3. The present invention monitors the changes in recognition accuracy during multiple rounds of feature optimization and feature construction. When the accuracy begins to decline, the size of the multi-dimensional feature combination parameters can be automatically changed to have good adaptability to the feature sets obtained in different environments, thereby improving the versatility of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic diagram of the process of the present invention.

[0063] Figure 2 A comparison chart of classification errors between pattern recognition using the original unevolved features and pattern recognition using the evolved features of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0065] The simulated evolution model of the present invention is realized by drawing on the evolution mechanism of genetic algorithms.

[0066] The improved convolutional neural network model based on GoogLeNet of the present invention adds several Inception layers on the basis of GoogLeNet. The number of layers can be determined according to the situation. The present invention adds three layers.

[0067] like Figure 2 As shown, the present invention provides a partial discharge identification method based on feature set simulation evolution optimization, which is characterized by comprising the following steps:

[0068] S1. Artificially constructing partial discharge defects to obtain partial discharge signals: Typical partial discharge defects in high-voltage cables are constructed in the laboratory. These defects include ring cutting marks, semiconductor powder defects, stress cone displacement defects, and grinding irregularities. Pressure tests are then performed on these defects to obtain partial discharge signals.

[0069] S2. Partial discharge signal data preprocessing: performing wavelet decomposition and noise reduction on the original partial discharge signal data set obtained in S1, identifying and filtering out abnormal data and removing noise by setting a threshold, and normalizing the partial discharge signal data; the normalization and wavelet decomposition in step 2 are specifically as follows:

[0070] S201: normalize the partial discharge signal using the standard fraction method, the calculation formula is:

[0071]

[0072] In the formula, μ is the mean of the characteristic data, σ is the standard deviation; n is the total number of data, i is the data order, x is the i is the i-th data value, Z i is the i-th data after normalization;

[0073] S202: Perform wavelet decomposition on the partial discharge signal. Select the Daubechies (dbN) wavelet basis function and decompose the partial discharge signal into approximate coefficients and detail coefficients layer by layer. The specific formula is as follows:

[0074]

[0075] Among them A J is the approximate coefficient, D J is the detail coefficient, S is the signal function, J is the number of wavelet decomposition layers, and then based on the coefficient of the finest layer D1, the noise standard deviation ε is estimated by the median absolute error to calculate the universal threshold λ. The specific formula is as follows:

[0076]

[0077] Where L is the length of the partial discharge signal. The noise-reduced partial discharge signal is reconstructed by inverse wavelet transform. The specific formula is as follows:

[0078]

[0079] is the partial discharge signal function after processing, To use threshold processing D J The detail coefficient obtained later;

[0080] S3. Processing the partial discharge signal to obtain a partial discharge phase spectrum and matrix: Corresponding the amplitude of the partial discharge signal normalized in S2 to the phase angle of the power frequency voltage, traversing a complete voltage cycle to obtain a partial discharge phase spectrum; performing inverse grayscale processing on the partial discharge phase spectrum, and then gridding the spectrum to determine the size of the matrix value corresponding to each grid based on the number of partial discharges in each grid to obtain a partial discharge phase spectrum matrix; specifically:

[0081] S301: The unit of the partial discharge signal amplitude is pG, the phase of a complete voltage cycle is 0° to 360°, and a scatter plot is drawn with the phase as the independent variable and the product of the amplitude and the polarity as the dependent variable to obtain a partial discharge phase spectrum;

[0082] S302: After grayscale processing is performed on the partial discharge phase spectrum, a white image with a black background is displayed. In order to perform feature extraction and pattern recognition and image segmentation, the image is filtered using a filter factor. In actual applications, to reduce the amount of calculation, a pixel value of 255 can be used as the filter function;

[0083] S303: The default grid matrix size is set to 360*100. The horizontal axis is divided into each degree of phase, and the vertical axis is divided into each 0.01 after normalization, with 50 positive and negative grids. The number of partial discharge signal points in the grid at row a and column b is the size of the matrix at row a and column b, and the partial discharge phase spectrum matrix is ​​obtained. If the classification accuracy of the result is insufficient, the matrix size can be changed to improve the accuracy, but the computational complexity will also increase.

[0084] S4. Extract and construct signal features based on the partial discharge phase spectrum matrix: Input the partial discharge phase spectrum matrix result in S3 into the Inception layer for feature extraction to obtain a feature set, wherein the series number and convolution kernel size parameters are set in the Inception layer; the feature set includes statistical features, time-frequency features, and wavelet features. The statistical features and wavelet features are obtained by calculation. The statistical features include mean value Mean, variance VA, skewness Sk, kurtosis Ku, and root mean square. The time-frequency features include discharge amount, rise time, signal width, fall time, crest factor, waveform factor, main frequency, equivalent pulse time T, equivalent pulse width W, and energy. The wavelet features are obtained by decomposing the wavelet decomposition energy ratios Ep1 to Ep8 using multiple wavelet bases. Specifically, it includes:

[0085] S401: The Inception layer contains several convolutions that are performed simultaneously, including 1*1 convolution, 1*1 convolution and 3*3 convolution in series, 1*1 convolution and 5*5 convolution in series, and pooling layer and 1*1 convolution in series. The Inception layer is connected to itself in series multiple times. The convolution size B and the number of series connections C are appropriately adjusted according to the actual situation to improve recognition accuracy.

[0086] S402: The time-frequency parameters are obtained during the partial discharge signal acquisition. The characteristic parameters of the feature set are calculated, including the mean value, variance VA, skewness Sk, kurtosis Ku, and wavelet parameters. The specific formula is as follows:

[0087] Skewness is defined using the third-order moment. The calculation formula for skewness is:

[0088]

[0089] In the formula, S k is the skewness; μ3 is the third-order central moment; σ is the standard deviation;

[0090] When the statistical data is right-skewed, S k >0,S k The larger the value, the higher the degree of right skewness; when the statistical data is left-skewed, S k <0,S k The smaller the value, the higher the degree of left skewness; when the statistical data is symmetrically distributed, there is S k =0;

[0091] In actual calculations, although the definition formula is used for calculation, it is necessary to traverse each sample multiple times to calculate the mean and variance, which is very time-consuming when the sample size is large. Therefore, the 1st to 3rd order origin moments are often used for calculation:

[0092] μ=EX

[0093] σ 2 =E(X-EX) 2 =EX 2 -μ 2

[0094]

[0095] μ is the mean, E is the calculation symbol, EX is the mean of the calculated data X, σ is the standard deviation, S k is the skewness;

[0096] Taking the normal distribution as a reference, the kurtosis describes the steepness of the distribution shape. If b k <3, the distribution is said to have insufficient kurtosis. If b k >3, the distribution is said to have excessive kurtosis. If it is known that the distribution may deviate from the normal distribution in terms of kurtosis, the kurtosis can be used to test the normality of the distribution. The expression of kurtosis is as follows:

[0097]

[0098] bx is the kurtosis, n is the total number of data, m4 is the fourth-order sample central moment, m2 is the second-order central moment, and x is the sample variance.i is the i-th data value, is the sample mean.

[0099] S5. Simulate evolutionary optimization of the feature set: A total of N generations of evolution are set, where N can be specified or the default value is used. Before each generation of evolution, a random forest algorithm will be used to rank the importance of all features in the feature set. The ranking is based on a dataset of known partial discharge defect types and their corresponding feature values ​​obtained in the laboratory. Then, in each generation of evolution, the simulated evolution process is performed according to the importance ranking. After the evolution is completed, it is first determined whether the recognition accuracy has decreased. If so, the simulated evolution model parameters are changed. Otherwise, it is determined whether the evolutionary generation number has reached N. If not, the simulated evolution is repeated. If so, the simulated evolution is terminated and the feature set is output. In this embodiment of the present invention, N is set to 3: Specifically including:

[0100] S501: The following events occur in sequence during the simulated evolution of the feature set:

[0101] (1) Elimination: The container can only hold a maximum of M features. The recommended range of M is 50 to 100, and the default value is 50. Features with importance greater than M will be eliminated directly. The reproduction in the step will cause the total number of features to exceed M, which does not conflict with the elimination rule. Elimination only requires that the number of features cannot exceed M when the evolution starts, and these features will enter the elimination stage together after the next round of feature optimization.

[0102] (2) Reproduction: The top k features in terms of importance will be combined in pairs to produce their offspring. k is generally 0.2*M. The combination methods include but are not limited to product, division or arithmetic mean, and the probability of occurrence of these methods is specified. A pair of combinations can only produce one offspring. The name of the offspring is determined by the combination method of the parents, and the offspring will not disappear after the parents are combined.

[0103] (3) Mutation: Features that have not been eliminated but whose importance is the latter t% cannot be reproduced, where t is generally 20. By trying to increase their own importance through mutation, they will transform themselves into the pth power of the original feature, where p is a random number between 0 and 2. The name of the mutated feature is represented by the original feature name plus the pth power. The mutated feature will replace itself.

[0104] S502: When the features obtained after each round of evolution in the step are subjected to random forest learning, the prediction accuracy of these features for the type of partial discharge will be tested synchronously. If the prediction accuracy is found to have decreased, the parameters of the simulated evolution model will be automatically changed, including the proportion of the number of reproduction and mutation in the total features, the default probability of the three combination methods of reproduction, and the default range of variation of the parameter p during mutation, to ensure that the evolutionary environment will adapt to different feature sets and enhance the versatility of the method.

[0105] S6. Pattern recognition using an improved convolutional neural network: The optimized feature set is divided into a training set and a test set. Pattern recognition is performed using a one-dimensional improved convolutional neural network. The network is enhanced by adding an Inception layer to the neural network to extract features from the evolved optimized feature set. The neural network structure is improved based on GoogLeNet. During training, the numerical labels of the partial discharge defect types are used as training targets, and the training set in the optimized feature set is used as training data to complete the training of the GoogLeNet-based improved convolutional neural network model. When the test set data of the optimized feature set is input, the trained GoogLeNet-based improved convolutional neural network model will output the numerical labels of the partial discharge defect types corresponding to the dataset as the result. This includes:

[0106] S601: The optimized feature set obtained above is randomly divided into a training set and a test set in a ratio of 5:3. The training target is set as the partial discharge defect type, which is represented by a number. The partial discharge defect type number of the test set is used to calculate the classification error.

[0107] S602: Use a one-dimensional improved convolutional neural network for pattern recognition, and improve the neural network structure based on GoogLeNet. During training, the digital label of the partial discharge defect type is used as the training target. The training set in the optimized feature set is used as training data. An appropriate number of training times is set to make the recognition rate tend to be stable. The training number of the present invention is 500 times. The training of the improved convolutional neural network model based on GoogLeNet is completed, and a trained improved convolutional neural network model based on GoogLeNet is obtained. When the test set data of the optimized feature set is input, the trained improved convolutional neural network model based on GoogLeNet will give the corresponding digital label of the partial discharge defect type of the data set as the result. Add several Inception layers to the convolutional neural network, and the number of layers can be changed as needed. After extracting features in the Inception layer, input the activation function layer, and finally input the fully connected layer for pattern recognition. The obtained classification error is fed back to the spectrum matrix feature extraction and construction. Change the convolution size B and the number of series C in the Inception1 layer to improve the feature extraction efficiency, and change the grid accuracy parameters of the gridding of the partial discharge phase spectrum image. To improve the overall accuracy of the convolutional neural network model based on GoogLeNet. For example The default value is 1. Multiplying the original network parameter 360*100 does not affect the grid accuracy. When the classification error is greater than 0.2, The grid parameters will be increased by 0.1 each time, and will become 360*100*1.1. The horizontal coordinate will remain unchanged, and the vertical coordinate will become 110 grids. The subsequent process will be repeated until the classification error is less than 0.2 or the number of vertical coordinate grids reaches 200, taking into account the balance between computational complexity and accuracy.

[0108] The present invention constructs high-dimensional features without the need for enumeration and then feature optimization, effectively avoiding data redundancy and greatly reducing the amount of calculation in actual use. Taking M as 50 as an example, the amount of calculation for new features is reduced from 2500 to 110. When N is 3, only 330 new features need to be calculated to complete the complete simulation evolution process; the overall effectiveness of the features is also high. The comparison of classification errors between pattern recognition using the original unevolved features and pattern recognition using the evolved features is shown in the figure below. Figure 2 As shown in Figure 2, it can be seen that the improved classification error decreases faster with the increase of iteration number, and the final error is also smaller.

[0109] The terms used in the drawings to describe positional relationships are for illustrative purposes only and are not to be construed as limiting the present invention.

[0110] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A partial discharge identification method based on feature set simulation evolution optimization, characterized in that: The following steps are involved: S1. Artificially constructing partial discharge defects to obtain partial discharge signals: Typical partial discharge defects in high-voltage cables are constructed in the laboratory. These defects include ring cutting marks, semiconductor powder defects, stress cone displacement defects, and grinding irregularities. Pressure tests are then performed on these defects to obtain partial discharge signals. S2. Partial discharge signal data preprocessing: Perform wavelet decomposition and noise reduction on the original partial discharge signal dataset obtained in S1, identify and filter out abnormal data, remove noise by setting a threshold, and normalize the partial discharge signal data; S3. Processing the partial discharge signal to obtain a partial discharge phase spectrum and matrix: Corresponding the amplitude of the partial discharge signal after the normalization process in S2 to the phase angle of the power frequency voltage, traversing a complete voltage cycle to obtain a partial discharge phase spectrum; De-graying the partial discharge phase spectrum, gridding the spectrum, and determining the size of the matrix value corresponding to each grid according to the number of partial discharges in each grid to obtain a partial discharge phase spectrum matrix; S4. Extract and construct signal features based on the partial discharge phase spectrum matrix: Input the partial discharge phase spectrum matrix result in S3 into the Inception layer for feature extraction to obtain a feature set, where the series number and convolution kernel size parameters are set in the Inception layer; the feature set includes statistical features, time-frequency features, and wavelet features. The statistical features and wavelet features are obtained by calculation. The statistical features include mean value Mean, variance VA, skewness Sk, kurtosis Ku, and root mean square. The time-frequency features include discharge amount, rise time, signal width, fall time, crest factor, waveform factor, main frequency, equivalent pulse time T, equivalent pulse width W, and energy. The wavelet features are obtained by decomposing the wavelet decomposition energy ratio Ep1 to Ep8 using multiple wavelet bases. S5. Simulate evolutionary optimization of the feature set: There are N generations of evolution. Before each generation, a random forest algorithm will rank the importance of all features in the feature set. The ranking is based on a dataset of known partial discharge defect types and their corresponding feature values ​​obtained in the laboratory. Then, in each generation, the simulated evolution process is performed according to the importance ranking. After the evolution is completed, the recognition accuracy is first determined to see if it has decreased. If so, the simulated evolution model parameters are changed. Otherwise, the number of evolution generations is determined to see if it has reached N. If not, the simulated evolution is repeated. If so, the simulated evolution ends and the feature set is output. S6. Pattern recognition through improved convolutional neural networks: The optimized feature set is divided into a training set and a test set; a one-dimensional improved convolutional neural network is used for pattern recognition. The feature extraction of the feature set after evolutionary optimization is performed by adding an Inception layer to the neural network. The neural network structure is improved based on GoogLeNet. During training, the numerical labels of the partial discharge defect types are used as training targets, and the training set in the optimized feature set is used as training data to complete the training of the improved convolutional neural network model based on GoogLeNet. When the test set data of the optimized feature set is input, the trained improved convolutional neural network model based on GoogLeNet will give the corresponding numerical labels of the partial discharge defect types for the data set as the result.

2. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 1, characterized in that: The normalization processing and wavelet decomposition in step 2 are specifically as follows: S201: normalize the partial discharge signal using the standard fraction method, the calculation formula is: In the formula, μ is the mean of the characteristic data, σ is the standard deviation; n is the total number of data, i is the data order, x is the i is the i-th data value, Z i is the i-th data after normalization; S202: Perform wavelet decomposition on the partial discharge signal. Select the Daubechies (dbN) wavelet basis function and decompose the partial discharge signal into approximate coefficients and detail coefficients layer by layer. The specific formula is as follows: Among them A J is the approximate coefficient, D J is the detail coefficient, S is the signal function, J is the number of wavelet decomposition layers, and then based on the coefficient of the finest layer D1, the noise standard deviation ε is estimated by the median absolute error to calculate the universal threshold λ. The specific formula is as follows: Where L is the length of the partial discharge signal. The noise-reduced partial discharge signal is reconstructed by inverse wavelet transform. The specific formula is as follows: is the partial discharge signal function after processing, To use threshold processing D J The detail coefficient obtained.

3. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 1, characterized in that: The step 3 is specifically as follows: S301: The unit of the partial discharge signal amplitude is pC, the phase of a complete voltage cycle is 0° to 360°, and a scatter plot is drawn with the phase as the independent variable and the product of the amplitude and polarity as the dependent variable to obtain a partial discharge phase spectrum; S302: After grayscale processing is performed on the partial discharge phase spectrum, a white image with a black background is presented, and the image is filtered using a filter factor; S303: The default grid matrix size is set to 360*100. The horizontal axis takes each degree of phase as one grid, and the vertical axis takes each 0.01 after normalization as one grid, with 50 positive and negative grids. The number of partial discharge signal points in the grid at row a and column b is the size of the matrix at row a and column b, and the partial discharge phase spectrum matrix is ​​obtained.

4. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 1, characterized in that: Step 4 specifically includes: S401: The Inception layer contains several convolutions that are performed simultaneously, including 1*1 convolution, 1*1 convolution and 3*3 convolution in series, 1*1 convolution and 5*5 convolution in series, and pooling layer and 1*1 convolution in series. The Inception layer is connected to itself in series multiple times. The convolution size B and the number of series connections C are appropriately adjusted according to the actual situation to improve recognition accuracy. S402: The time-frequency parameters are obtained during the partial discharge signal acquisition. The characteristic parameters of the feature set are calculated, including the mean value, variance VA, skewness Sk, kurtosis Ku, and wavelet parameters. The specific formula is as follows: Skewness is defined using the third-order moment. The calculation formula for skewness is: In the formula, S k is the skewness; μ3 is the third-order central moment; σ is the standard deviation; When the statistical data is right-skewed, S k >0,S k The larger the value, the higher the degree of right skewness; when the statistical data is left-skewed, S k <0,S k The smaller the value, the higher the degree of left skewness; when the statistical data is symmetrically distributed, there is S k =0; In actual calculations, although the definition formula is used for calculation, it is necessary to traverse each sample multiple times to calculate the mean and variance, which is very time-consuming when the sample size is large. Therefore, the 1st to 3rd order origin moments are often used for calculation: μ=EX σ 2 =E(X-EX) 2 =EX 2 -μ 2 μ is the mean, E is the calculation symbol, EX is the mean of the calculated data X, σ is the standard deviation, S k is the skewness; Taking the normal distribution as a reference, the kurtosis describes the steepness of the distribution shape. If b k <3, the distribution is said to have insufficient kurtosis. If b k >3, the distribution is said to have excessive kurtosis. If it is known that the distribution may deviate from the normal distribution in terms of kurtosis, the kurtosis can be used to test the normality of the distribution. The expression of kurtosis is as follows: b k is the kurtosis, n is the total number of data, m4 is the fourth-order sample central moment, m2 is the second-order central moment, and x is the sample variance. i is the i-th data value, is the sample mean.

5. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 1, characterized in that: The step 5 specifically includes: S501: The following events occur in sequence during the simulated evolution of the feature set: (1) Elimination: The container can only hold a maximum of M features. Features with importance greater than M will be eliminated directly. If the total number of features exceeds M after the multiplication in the step, this does not conflict with the elimination rule. Elimination only requires that the number of features cannot exceed M when the evolution starts, and these features will enter the elimination phase together after the next round of feature optimization. (2) Reproduction: The top k most important features will be combined in pairs to produce their offspring, with k being 0.2*M. The combination methods include but are not limited to multiplication, division, or arithmetic mean, and the probability of occurrence of these methods is specified; a pair of combinations can only produce one offspring, and the offspring's name is determined by the combination method of the parents, and the offspring will not disappear after the parents are combined; (3) Mutation: Features that have not been eliminated but whose importance is in the latter t% cannot reproduce. They try to increase their own importance through mutation and will transform themselves into the pth power of the original feature, where p is a random number between 0 and 2. The name of the mutated feature is represented by the original feature name plus the pth power. The mutated feature will replace itself. S502: When the features obtained after each round of evolution in the step are subjected to random forest learning, the prediction accuracy of these features for the type of partial discharge will be tested synchronously. If the prediction accuracy is found to have decreased, the parameters of the simulated evolution model will be automatically changed, including the proportion of the number of reproduction and mutation in the total features, the default probability of the three combination methods of reproduction, and the default range of variation of the parameter p during mutation, to ensure that the evolutionary environment will adapt to different feature sets and enhance the versatility of the method.

6. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 1, characterized in that: The step 6 specifically includes: S601: The optimized feature set obtained above is randomly divided into a training set and a test set in a ratio of 5:

3. The training target is set as the partial discharge defect type, which is represented by a number. The partial discharge defect type number of the test set is used to calculate the classification error. S602: Use a one-dimensional improved convolutional neural network for pattern recognition, and improve the neural network structure based on GoogLeNet. During training, the digital label of the partial discharge defect type is used as the training target. The training set in the optimized feature set is used as the training data. An appropriate number of training times is set to stabilize the recognition rate. The training of the improved convolutional neural network model based on GoogLeNet is completed, and a trained improved convolutional neural network model based on GoogLeNet is obtained. When the test set data of the optimized feature set is input, the trained improved convolutional neural network model based on GoogLeNet will give the corresponding digital label of the partial discharge defect type of the data set as the result. Add several Inception layers to the convolutional neural network, and the number of layers can be changed as needed. After extracting features from the Inception layer, input the activation function layer, and finally input the fully connected layer for pattern recognition. The obtained classification error is fed back to the spectrum matrix feature extraction and construction. Change the convolution size B and the number of series C in the Inception1 layer to improve the feature extraction efficiency, and change the grid accuracy parameters of the gridding of the partial discharge phase spectrum image. To improve the overall accuracy of the improved convolutional neural network model based on GoogLeNet.

7. The method for identifying partial discharge based on feature set simulation evolutionary optimization according to claim 5, characterized in that: The value of M is: 50 to 100, and the default value is 50.