Fault identification method, device and system for partial discharge signal of ring main unit
Through the Stacking integrated learning method, combined with 1DCNN-ResNet, XGBoost, RandomForest and SVM models, the problem of low accuracy in local discharge signal fault recognition of ring network cabinets is solved, and higher accuracy in fault recognition and power supply safety are achieved.
Patent Information
- Application Number
- CN202411445500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The prior art has low accuracy in the identification of local discharge signal faults in ring network cabinets, making it difficult to effectively identify the type and severity of the fault.
Using Stacking integrated learning method, 1DCNN-ResNet, XGBoost, and RandomForest are used as the base learners, and SVM is used as the meta learners to train and test the data through K-fold cross-validation to improve the accuracy of fault identification.
Through the combination of multiple differentiated models, the local discharge signal fault type of the ring network cabinet can be more accurately identified, which improves the accuracy of fault identification and enhances the guarantee of the power supply safety of the ring network cabinet.
Smart Images

Figure CN118965138B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of detection technology of ring main units, and in particular to a method, device, medium and system for fault identification of partial discharge signals of a ring main unit. Background Art
[0002] Ring main unit is an important part of urban and rural power grids. Its operating status directly determines the power consumption of urban and rural power users in my country. It is also widely used in distribution stations and box-type substations in load centers such as high-rise buildings, urban residential areas, large public buildings, factories and enterprises. It is a key equipment to ensure the reliability and stability of power supply in the power grid. However, if the internal electrical equipment of the ring main unit works in a high temperature and harsh environment for a long time, it is easy to age the insulation and produce partial discharge signals. The deterioration of insulation performance is the main cause of its failure, which may cause large-scale power outages in the power grid in severe cases. The partial discharge signal is the key factor causing the insulation failure of the ring main unit. Its fault type is closely related to the severity of the insulation failure, which is mainly divided into three types of faults: surface discharge, air gap discharge and corona discharge. By identifying and monitoring the fault type of partial discharge signal, the insulation status of the ring main unit equipment can be effectively evaluated, providing a reliable basis for fault diagnosis and equipment maintenance, which helps to eliminate the hidden dangers of insulation failure, curb the occurrence of insulation accidents from the source, and ensure the stability and reliability of power supply.
[0003] Among the existing schemes, the partial discharge signal pattern recognition schemes can be mainly divided into the following two categories: 1) Directly extracting the partial discharge signal fault feature information through feature extraction methods as the discriminant feature parameters for fault type identification. This type of method requires the extraction of more feature parameters and often depends on the on-site experience of experts. It is highly subjective and has great limitations. It also requires manual feature extraction, making it difficult to solve the problem of end-to-end fault diagnosis, and the recognition results have high uncertainty; 2) Using traditional single model classifiers for fault classification methods. This type of method has fast training speed, fewer parameters, and can overcome the small sample problem well. However, the fault diagnosis results of partial discharge signals they obtain are all based on a single classification model detection, and cannot learn the characteristics of data in different spaces, so the recognition accuracy is not high.
[0004] That is, the existing solution has poor accuracy in fault identification of partial discharge signals of ring main units. Summary of the invention
[0005] The main purpose of the present application is to provide a method, device, medium and system for fault identification of partial discharge signals of a ring main unit, so as to at least solve the problem that the existing solutions have poor accuracy in fault identification of partial discharge signals of a ring main unit.
[0006] In order to achieve the above object, according to one aspect of the present application, a method for fault identification of partial discharge signals of a ring main unit is provided, the method comprising:
[0007] Get the original partial discharge fault dataset , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, wherein, is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples;
[0008] Using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and using the SVM model as a meta-learner, and using the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and dividing the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner;
[0009] Using the base training set and the base test set to train and test each of the base learners respectively to obtain a first base learner, a second base learner, and a third base learner, and using the meta training set and the meta test set to train and test the meta learner to obtain a trained meta learner;
[0010] Using the trained meta-learner to identify the fault type of newly collected partial discharge signal fault data to obtain a final identification result, and using the first base learner, the second base learner and the third base learner to verify the final identification result;
[0011] Optionally, in the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes:
[0012] The XGBoost algorithm model is constructed as follows:
[0013] ;
[0014] in, is the model prediction value of the i-th sample, is the structure of the t-th independent tree, is the i-th data input, F is the i-th data input, and K is the number of trees.
[0015] Optionally, in the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes:
[0016] The loss function is constructed as:
[0017] ;
[0018] Among them, Loss is the loss value, is the kth model parameter, is a regularization function used to constrain the complexity of the model. is the model prediction value of the i-th sample, is the target true value of the i-th sample, is the loss function, which represents the error between the true value and the predicted value of the i-th sample, n is the total number of data, and K is the total number of trees.
[0019] Optionally, before constructing the loss function, the method further includes:
[0020] The regularization term is determined as:
[0021] ;
[0022] in, is a regularization function used to constrain the complexity of the model. To control the number of leaf nodes, T is the number of leaf nodes, is the preset coefficient, is the score of the leaf node of the j node.
[0023] Optionally, dividing the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner includes:
[0024] The original training set is divided into K-fold subsets, where the number of training examples in D is m, and each subset has m / k training examples. , and meet , ;
[0025] From the divided subsets, take the i-th fold as the test set and the other K-1 folds as the training set.
[0026] Optionally, the base training set and the base test set are used to train and test the base learners 1DCNN-ResNet, the RandomForest, and the XGBoost, respectively, to obtain a first base learner, a second base learner, and a third base learner, including:
[0027] Selection step: select 1 fold of the K-fold data as the base test set, and the remaining K-1 folds as the base training set, and train to obtain a 1DCNN-ResNet1 model;
[0028] Prediction step: using the trained 1DCNN-ResNet1 model to predict the data in the base test set to obtain a first prediction result matrix;
[0029] Using a K-fold cross validation method, repeat the selection step and the prediction step until all K-fold predictions are completed, and obtain a first final prediction result matrix;
[0030] Optionally, using the meta-training set and the meta-testing set to train and test the meta-learner to obtain a trained meta-learner includes:
[0031] Determine that the training data of the meta-learner is a set of final prediction result matrices of all the base learners on the meta-training set, and use the set of final prediction result matrices to train the SVM model;
[0032] Using the 1DCNN-ResNet1 model, 1DCNN-ResNet2 model, 1DCNN-ResNet3 model, 1DCNN-ResNet4 model, and 1DCNN-ResNet5 model to predict all sample data in the meta-test set, five prediction result matrices are obtained, and the prediction results of the five models are averaged to obtain a first average prediction result matrix; and the 1DCNN-ResNet is replaced with XGBboost and RandomForest, respectively, to obtain a second average prediction result matrix and a third average prediction result matrix;
[0033] Determine that the average prediction result matrix set is a set of the first average prediction result matrix, the second average prediction result matrix and the third average prediction result matrix;
[0034] The trained SVM model is used to predict the data in the average prediction result matrix set to obtain an SVM prediction result.
[0035] According to another aspect of the present application, a fault identification device for a partial discharge signal of a ring main unit is provided, the device comprising:
[0036] Acquisition unit, used to acquire the original partial discharge fault data set , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, wherein, is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples;
[0037] A first processing unit is used to use 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and use the SVM model as a meta-learner, and use the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and divide the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner;
[0038] a second processing unit, configured to respectively train and test each of the base learners using the base training set and the base test set to obtain a first base learner, a second base learner, and a third base learner, and to train and test the meta learner using the meta training set and the meta test set to obtain a trained meta learner;
[0039] The third processing unit is used to use the trained meta-learner to identify the fault type of the newly collected partial discharge signal fault data to obtain a final identification result, and use the first base learner, the second base learner and the third base learner to verify the final identification result.
[0040] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described.
[0041] According to another aspect of the present application, a fault identification system for partial discharge signals of a ring main unit is provided, the system comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of the methods described.
[0042] By applying the technical solution of the present application, each base learner is trained and tested respectively by using a base training set and a base test set to obtain a first base learner, a second base learner and a third base learner, and a meta-training set and a meta-test set are used to train and test the meta-learner to obtain a trained meta-learner, thereby utilizing a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, and effectively avoiding the occurrence of insulation failures in ring network cabinets, which has important practical significance for ensuring the safety and reliability of power supply of ring network cabinets, and at the same time improves the accuracy of fault identification of local discharge signals of ring network cabinets, thereby solving the problem of poor accuracy of fault identification of local discharge signals of ring network cabinets in existing solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings constituting part of the present application are used to provide a further understanding of the present application. The exemplary embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0044] Figure 1 A schematic flow chart of a method for fault identification of a partial discharge signal of a ring main unit provided in accordance with an embodiment of the present application is shown;
[0045] Figure 2 A schematic diagram of the ring main unit PD fault diagnosis architecture of Stacking integrated learning is shown;
[0046] Figure 3 A schematic diagram of the 1DCNN-ResNet network structure is shown;
[0047] Figure 4 A schematic diagram of the RandomForest model structure is shown;
[0048] Figure 5 Shown is a schematic diagram of the SVM algorithm structure;
[0049] Figure 6 A schematic diagram of the comparison curve between 1DCNN and 1DCNN-ResNet accuracy algorithms is shown;
[0050] Figure 7 A structural block diagram of a fault identification device for partial discharge signals of a ring main unit provided according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0051] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0054] For the convenience of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0055] 1DCNN-ResNet: 1DCNN-ResNet is a deep learning model that combines 1D Convolutional Neural Network (1D Convolutional Neural Network) and ResNet (Residual Neural Network). 1D Convolutional Neural Network is usually used to process time series data, such as audio, text, etc., while ResNet is a deep residual network structure that can help solve the gradient vanishing and gradient exploding problems in deep network training. 1DCNN-ResNet combines the advantages of these two models and can be used to process tasks such as classification and regression of time series data.
[0056] XGBoost: XGBoost is a machine learning algorithm based on Gradient Boosting DecisionTree, which is widely used in machine learning tasks such as classification, regression, and sorting. XGBoost trains weak classifiers through continuous iterations and adjusts model parameters based on the results of the previous iteration, thereby gradually improving the accuracy of the model. XGBoost has performed well in data science competitions such as Kaggle and is a very powerful machine learning algorithm.
[0057] RandomForest: RandomForest is an ensemble learning method based on the Random Forest algorithm, which is used to solve classification and regression problems. Random Forest is an ensemble learning method that improves the accuracy and generalization ability of the model by training multiple decision trees and combining their prediction results. RandomForest performs well in processing high-dimensional data and large-scale data sets, and is widely used in data mining, financial risk control, medical diagnosis and other fields.
[0058] SVM (Support Vector Machine) is a supervised learning algorithm used for classification and regression analysis. In classification problems, SVM separates data points of different categories by finding an optimal hyperplane in the feature space. In regression problems, SVM tries to find an optimal hyperplane that minimizes the distance between the data points and the plane. SVM performs well in solving small sample, nonlinear and high-dimensional data classification problems and is widely used in machine learning and data mining.
[0059] As introduced in the background technology, in the existing schemes, the schemes for partial discharge signal pattern recognition can be mainly divided into the following two categories: 1) Directly extracting partial discharge signal fault feature information through feature extraction methods as discriminant feature parameters for fault type identification. This type of method requires the extraction of more feature parameters, often depends on the on-site experience of experts, is highly subjective and has great limitations, and requires manual feature extraction, making it difficult to solve the problem of end-to-end fault diagnosis, and the recognition results have high uncertainty; 2) Using traditional single model classifiers for fault classification methods. This type of method has fast training speed, fewer parameters, and can overcome the small sample problem well, but the fault diagnosis results of the partial discharge signals they obtain are all based on a single classification model detection, and the characteristics of the data in different spaces cannot be learned, and the recognition accuracy is not high. In order to solve the problem of poor accuracy of fault identification of partial discharge signals of ring network cabinets in existing schemes, the embodiments of the present application provide a method, device, medium and system for fault identification of partial discharge signals of ring network cabinets.
[0060] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0061] In this embodiment, a method for fault identification of partial discharge signals of a ring main unit is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0062] Figure 1 FIG. 1 is a flow chart of a method for identifying a fault of a partial discharge signal of a ring main unit provided in accordance with an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:
[0063] Step S101, obtaining the original partial discharge fault data set , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, where: is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples;
[0064] Step S102, using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and using the SVM model as a meta-learner, and using the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and dividing the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner;
[0065] Specifically, by utilizing a variety of differentiated models to observe the data space and structure from different angles, the advantages of different models can be fully utilized.
[0066] like Figure 2 As shown in the figure, the ring main unit PD fault diagnosis architecture of Stacking ensemble learning is presented. In the first layer of this architecture, the 1DCNN-ResNet model based on deep learning and the XGBoost and RandomForest models based on traditional machine learning are used as base learners to learn the diversified fault feature expressions of partial discharge. In the 1DCNN-ResNet model structure, the MaxPooling1D and GlobalAveragePooling1D modules are responsible for performing feature dimensionality reduction and global average pooling on the output of the one-dimensional convolution layer to reduce the data dimension and computational complexity; the fully connected layer (FC) is responsible for outputting the final classification results. In addition, the XGBoost and RandomForest models use different ensemble learning strategies, namely Boosting and Bagging, to enhance the classification performance. XGBoost gradually improves the accuracy of the model by sequentially adding weak learners and weighting the errors; while RandomForest reduces overfitting and improves generalization ability by building multiple decision trees and voting. These traditional machine learning models have shown good performance in processing structured data, providing effective base learner selection for fault diagnosis tasks, further enriching the diversity of the Stacking model, and helping to improve the robustness and accuracy of the final integrated model. In the second layer of the Stacking model, the support vector machine (SVM) is selected as the meta-learner to effectively handle nonlinear, small sample, and high-dimensional classification problems, thereby further improving the overall performance of the Stacking integrated model. In addition, since SVM has shown good performance in processing sample data and noise, its application in the Stacking model helps to enhance the model's robustness to outliers and noise;
[0067] To verify the effectiveness of the 1DCNN-ResNet algorithm, we can compare 1DCNN-ResNet with the original 1DCNN algorithm. The details will not be repeated here.
[0068] In one embodiment of the present application, in the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes:
[0069] The XGBoost algorithm model is constructed as follows:
[0070] ;
[0071] in, is the model prediction value of the i-th sample, is the structure of the t-th independent tree, is the i-th data input, F is the i-th data input, and K is the number of trees.
[0072] Specifically, the strategy of additive training is a training method that helps learners gradually master knowledge and skills by gradually increasing the difficulty and complexity of training. This strategy can help learners gradually build confidence and self-confidence, so that they can better cope with challenges and difficulties. By gradually increasing the difficulty of training, learners can gradually improve their skill levels and achieve higher learning goals. The XGBoost algorithm optimizes the model by adopting the strategy of additive training, further reducing the size of the loss function until all K trees are optimized.
[0073] In one embodiment of the present application, in the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes:
[0074] The loss function is constructed as:
[0075] ;
[0076] Among them, Loss is the loss value, is the kth model parameter, is a regularization function used to constrain the complexity of the model. is the model prediction value of the i-th sample, is the target true value of the i-th sample, is the loss function, which represents the error between the true value and the predicted value of the i-th sample, n is the total number of data, and K is the total number of trees.
[0077] Specifically, the loss value is set to optimize the corresponding model with the goal of minimizing the loss value.
[0078] In one embodiment of the present application, before constructing the loss function, the method further includes:
[0079] Regularization is a widely used technique in machine learning and statistical modeling to prevent model overfitting and improve the generalization ability of the model. Regularization is usually achieved by adding a penalty term to the loss function, which is related to the parameters of the model. The regularization term is determined as:
[0080] ;
[0081] in, is a regularization function used to constrain the complexity of the model. To control the number of leaf nodes, T is the number of leaf nodes, is the preset coefficient, is the score of the leaf node of the j node.
[0082] Specifically, Used to ensure that the score of leaf nodes is not too large.
[0083] In one embodiment of the present application, the meta-test set in the meta-learner is divided into a base training set and a base test set in a K-fold cross-validation manner, including:
[0084] The original training set is divided into K-fold subsets, where the number of training examples in D is m, and each subset has m / k training examples. , and meet , ;
[0085] From the divided subsets, take the i-th fold as the test set and the other K-1 folds as the training set.
[0086] Specifically, the training set in the meta-learner is divided into a base training set and a base test set according to the K-fold cross-validation method, and K is taken as 5. That is, 1500 samples are selected as the test set each time, and the remaining 6000 sample data are used as the training set.
[0087] K-fold cross validation is a commonly used model evaluation method. It divides the data set into K parts, and then selects one of them as the validation set and the remaining K-1 parts as the training set. The model is then trained and the model performance is evaluated on the validation set. This process is repeated K times, with a different validation set selected each time. The results of the K validations are finally averaged as the final validation result. This can more accurately evaluate the performance of the model and reduce the risk of overfitting and underfitting.
[0088] The K-fold cross validation algorithm process is as follows:
[0089] Divide the original training set D into K-fold subsets of similar size and non-overlapping. The number of training examples in D is m, then each subset has m / k training examples, and the corresponding subset , and meet , .
[0090] From the divided subsets, take the i-th fold as the test set and the other K-1 folds as the training set.
[0091] According to the training set, train the i-th base learner model h i (x).
[0092] The trained h i (x) The base learner is placed on the test set and the generalization error and classification rate of the model are calculated.
[0093] Calculate the average of the classification rates obtained K times as the final true classification rate of the model, where the K value range is generally [2,10].
[0094] The network structure of 1DCNN-ResNet is as follows Figure 3 As shown in the figure, the network model structure mainly includes input layer, convolution layer (Conv), multiple channel-level threshold residual shrinkage blocks (RBU) connected in series, batch normalization (BN), activation function (ReLU), global average pooling layer (GlobalAveragePooling1D, GAP) and fully connected layer (FC). The GlobalAveragePooling1D module is used to perform global average pooling on the input feature map to reduce the dimension and computational complexity of the data. The FC module outputs the final classification result for the fully connected layer. " / 2" means moving the convolution kernel with a stride of 2 to reduce the width of the output feature map. K is the number of convolution kernels in the convolution layer, M is the number of neurons in the FC network, and C, W and 1 in C×W×1 are indicators of the number of channels, width and height of the feature map respectively.
[0095] like Figure 4As shown in the figure, the RandomForest model structure is shown. This figure shows a typical random forest model, which contains multiple independently constructed decision trees (marked as decision trees A, B, C, and D in the figure). Each decision tree is generated based on a random subset of the input data during model training. This randomness helps to reduce the overfitting problem of the model. During the prediction process, the input data will be classified or regressed through multiple decision trees at the same time, and each tree will give a prediction result. Finally, these prediction results will be summarized through a "combiner" (for example, by majority voting or average) to generate the final output of the model;
[0096] Step S103, using the base training set and the base test set to train and test each of the base learners, respectively, to obtain a first base learner, a second base learner, and a third base learner, and using the meta training set and the meta test set to train and test the meta learners, to obtain a trained meta learner;
[0097] In one embodiment of the present application, the base training set and the base test set are used to train and test the base learners 1DCNN-ResNet, the RandomForest, and the XGBoost, respectively, to obtain a first base learner, a second base learner, and a third base learner, including:
[0098] Selection step: select 1 fold of the K-fold data as the base test set, and the remaining K-1 folds as the above-mentioned base training set, and train to obtain the 1DCNN-ResNet1 model;
[0099] Prediction step: Use the trained 1DCNN-ResNet1 model to predict the data in the base test set to obtain a first prediction result matrix;
[0100] The K-fold cross validation method is adopted to repeat the above selection step and the above prediction step until all K-fold predictions are completed, and the first final prediction result matrix is obtained.
[0101] Specifically, similarly, RandomForest will obtain the second prediction result matrix and the second final prediction result matrix, and XGBoost will obtain the third prediction result matrix and the third final prediction result matrix.
[0102] 1) Select 1 fold (1500) of the K-fold data as the base test set, and the remaining 4 folds (6000) as the base training set, and train to obtain the 1DCNN-ResNet1 model;
[0103] 2) Use the 1DCNN-ResNet1 model trained in 1) to predict the data in the base test set and obtain a 1500x7 (because there are 7 categories) prediction result matrix;
[0104] 3) Use K-fold cross validation and repeat 1) and 2) until all 5 folds are completed. This is how we get 1DCNN-ResNet i (i=1,2…,5)These 5 models, and obtained a complete prediction result for 7500 samples, that is, a 7500×7 prediction result matrix, recorded as Tr_Pred 1DCNN-ResNet (i.e. the first final prediction result matrix);
[0105] 4) Similarly, for XGBoost and RandomForest base learners, we followed steps 1), 2), and 3) to obtain the five models XGBoosti, RandomForesti (i=1,2,…,5), and the prediction result matrices for these two base learners, with dimensions of 7500×7, respectively denoted as Tr_Pred XGBoost and Tr_Pred RandomForest .
[0106] In one embodiment of the present application, the meta-learner is trained and tested using the meta-training set and the meta-test set to obtain a trained meta-learner, including:
[0107] Determine that the training data of the meta-learner is a set of final prediction result matrices of all the base learners on the meta-training set, and use the set of final prediction result matrices to train the SVM model;
[0108] Use 1DCNN-ResNet1 model, 1DCNN-ResNet2 model, 1DCNN-ResNet3 model, 1DCNN-ResNet4 model, and 1DCNN-ResNet5 model to predict all sample data in the above meta-test set, and obtain 5 prediction result matrices, and take the average of the prediction results of these 5 models to obtain the first average prediction result matrix; and replace 1DCNN-ResNet with XGBboost and RandomForest respectively, to obtain the second average prediction result matrix and the third average prediction result matrix respectively;
[0109] Determine that the average prediction result matrix set is a set of the first average prediction result matrix, the second average prediction result matrix, and the third average prediction result matrix;
[0110] The trained SVM model is used to predict the data in the above average prediction result matrix set to obtain the SVM prediction result.
[0111] Specifically, the specific steps of the meta-learner SVM training and testing process are as follows:
[0112] (1) Meta-learner training data preparation: The training data of the meta-learner is the concatenation of the prediction results of each base learner on the meta-training set;
[0113] That is (Tr_Pred1DCNN-ResNet, Tr_PredXGBoost, Tr_PredRandomForest), which is a 7500×21 matrix, denoted as Tr_Data.
[0114] (2) Meta-learner training: Use Tr_Data to train an SVM model. The SVM algorithm structure is shown in Figure 5 As shown in Figure 1, the algorithm realizes mapping 21-dimensional data into a 7-dimensional target space (there are 7 fault categories in the ring main unit).
[0115] (3) Meta-learner test data preparation: The test data acquisition method of the meta-learner is slightly different from the training data, and an additional operation of averaging the prediction results is required. First, 1DCNN-ResNet is used. 1 , 1DCNN-ResNet 2 , 1DCNN-ResNet 3 , 1DCNN-ResNet 4 , 1DCNN-ResNet 5 These five models make predictions on 2500 sample data in the meta-test set, and obtain five 2500×7 prediction result matrices;
[0116] Next, the prediction results of these five models are averaged to obtain a 2500×7 prediction result matrix, recorded as Ts_Pred 1DCNN-ResNet .
[0117] (4) According to the steps in (3), 1DCNN-ResNet is replaced by XGBboost and RandomForest respectively, and two 2500×7 prediction result matrices Ts_Pred are obtained respectively. XGBoos t and Ts_Pred RandomForest .
[0118] (5) Connect the data obtained in steps (3) and (4) to obtain the test data of the meta-learner: (Ts_Pred 1DCNN-ResNet , Ts_Pred XGBoost , Ts_Pred RandomForest ), a matrix with a dimension of 2500x21, denoted as Ts_Data.
[0119] (6) Use the SVM model trained in (2) to predict the data in Ts_Data and obtain a result of 2500×7, which is the final fault type classification result of the Stacking method on the test set. The effectiveness of the Stacking method can be verified by checking whether the final classification result is better than the best result in Ts_Pred1DCNN-ResNet, Ts_PredXGBoost, and Ts_PredRandomForest (the Stacking method is an integrated learning method that combines multiple different basic models and obtains the final prediction result by weighted averaging or voting their prediction results. In the Stacking method, the original partial discharge fault data set is first divided into a training set and a test set, and then multiple different basic models are trained on the training set. Each basic model will obtain a prediction result. Then, these prediction results are used as input data, and a meta-model is trained to obtain the final prediction result. The Stacking method can improve the generalization ability and prediction performance of the model and is applicable to various types of data sets and machine learning problems).
[0120] The Stacking ensemble learning method effectively solves the problems that traditional methods rely on a single classification model for PD fault type identification, have difficulty learning the expression of multiple feature spaces of data, and have limited recognition accuracy.
[0121] like Figure 5 The figure shows the structure of the SVM algorithm. This figure shows the basic concept of SVM in the binary classification problem, especially when the data is linearly separable. The solid line in the middle of the figure represents the decision boundary, which can separate the two types of data points (black points and white points). The two dotted lines around the decision boundary represent the support vector hyperplanes, which correspond to the classification conditions and The points on these dotted lines are support vectors, which play a key role in determining the decision boundary. The goal of SVM is to maximize the interval between these two dotted lines, which can improve the generalization ability of the model. The arrows in the figure represent the weight vectors , its direction determines the direction of the decision boundary, and the size of the interval is related to is inversely proportional to the norm (length);
[0122] like Figure 6 As shown in the figure, the accuracy comparison curves of 1DCNN and 1DCNN-ResNet are shown. It can be observed that the convergence performance and recognition accuracy of the proposed 1DCNN-ResNet are far better than those of 1DCNN, proving that 1DCNN-ResNet can more effectively learn the fault feature expression of partial discharge;
[0123] Step S104, using the trained meta-learner to identify the fault type of the newly collected partial discharge signal fault data to obtain a final identification result, and using the first base learner, the second base learner and the third base learner to verify the final identification result.
[0124] In the above steps, the base learners are trained and tested respectively by using the base training set and the base test set to obtain the first base learner, the second base learner and the third base learner, and the meta-training set and the meta-test set are used to train and test the meta-learner to obtain the trained meta-learner, thereby utilizing a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, and effectively avoiding the occurrence of insulation failures in the ring main unit, which has important practical significance for ensuring the power supply safety and reliability of the ring main unit, and at the same time improves the accuracy of fault identification of the local discharge signal of the ring main unit, thereby solving the problem of poor accuracy of fault identification of the local discharge signal of the ring main unit in the existing scheme.
[0125] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the fault identification method of the partial discharge signal of the ring main unit of the present application will be described in detail below in combination with specific embodiments.
[0126] This embodiment relates to a specific method for identifying a fault of a partial discharge signal of a ring main unit, including:
[0127] Get the original partial discharge fault dataset , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, where: is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples;
[0128] 1DCNN-ResNet, XGBoost, and RandomForest are used as base learners of the Stacking ensemble model, and the SVM model is used as the meta-learner. The training set and test set of the original partial discharge fault dataset D are used as the meta-training set and meta-test set of the meta-learner, and the meta-test set in the meta-learner is divided into a base training set and a base test set by K-fold cross-validation.
[0129] Selection step: select 1 fold of the K-fold data as the base test set, and the remaining K-1 folds as the base training set, and train to obtain the 1DCNN-ResNet1 model;
[0130] Prediction step: Use the trained 1DCNN-ResNet1 model to predict the data in the base test set to obtain the first prediction result matrix;
[0131] Using the K-fold cross-validation method, the selection step and the prediction step are repeated until all K folds are predicted, and the first final prediction result matrix is obtained.
[0132] Determine that the training data of the meta-learner is the set of final prediction result matrices of all base learners on the meta-training set, and use the set of final prediction result matrices to train the SVM model;
[0133] Use 1DCNN-ResNet1 model, 1DCNN-ResNet2 model, 1DCNN-ResNet3 model, 1DCNN-ResNet4 model, and 1DCNN-ResNet5 model to predict all sample data in the meta-test set, and obtain 5 prediction result matrices. Take the average of the prediction results of these 5 models to obtain the first average prediction result matrix; and replace 1DCNN-ResNet with XGBboost and RandomForest respectively to obtain the second average prediction result matrix and the third average prediction result matrix respectively;
[0134] Determine that the average prediction result matrix set is a set of a first average prediction result matrix, a second average prediction result matrix, and a third average prediction result matrix;
[0135] Use the trained SVM model to predict the data in the average prediction result matrix set to obtain the SVM prediction result;
[0136] The trained meta-learner is used to identify the fault type of the newly collected partial discharge signal fault data to obtain the final identification result, and the first base learner, the second base learner and the third base learner are used to verify the final identification result.
[0137] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0138] The embodiment of the present application also provides a fault identification device for a partial discharge signal of a ring main unit. It should be noted that the fault identification device for a partial discharge signal of a ring main unit in the embodiment of the present application can be used to execute the fault identification method for a partial discharge signal of a ring main unit provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and those that have been described will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0139] The following introduces a fault identification device for partial discharge signals of a ring main unit provided in an embodiment of the present application.
[0140] Figure 7 1 is a structural block diagram of a fault identification device for partial discharge signals of a ring main unit provided according to an embodiment of the present application. Figure 7 As shown, the device comprises:
[0141] The acquisition unit 71 is used to acquire the original partial discharge fault data set , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, where: is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples;
[0142] The first processing unit 72 is used to use 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and use the SVM model as a meta-learner, and use the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and divide the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner;
[0143] The second processing unit 73 is used to respectively train and test each of the base learners using the base training set and the base test set to obtain a first base learner, a second base learner and a third base learner, and to train and test the meta learners using the meta training set and the meta test set to obtain a trained meta learner;
[0144] The third processing unit 74 is used to use the trained meta-learner to identify the fault type of the newly collected partial discharge signal fault data to obtain a final identification result, and use the first base learner, the second base learner and the third base learner to verify the final identification result.
[0145] In the above-mentioned device, the above-mentioned base learners are respectively trained and tested by using the above-mentioned base training set and the above-mentioned base test set to obtain the first base learner, the second base learner and the third base learner, and the above-mentioned meta-training set and the above-mentioned meta-test set are used to train and test the above-mentioned meta-learner to obtain the trained meta-learner, thereby utilizing a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, and effectively avoiding the occurrence of insulation failures in the ring main unit, which has important practical significance for ensuring the safety and reliability of the power supply of the ring main unit, and at the same time improves the accuracy of fault identification of the partial discharge signal of the ring main unit, thereby solving the problem of poor accuracy of fault identification of the partial discharge signal of the ring main unit in the existing scheme.
[0146] In one embodiment of the present application, the first processing unit includes a first building module. In the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the first building module is used to construct the XGBoost algorithm model as follows:
[0147] ;
[0148] in, is the model prediction value of the i-th sample, is the structure of the t-th independent tree, is the i-th data input, F is the i-th data input, and K is the number of trees.
[0149] In one embodiment of the present application, the first processing unit includes a second building module. In the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the second building module is used to construct a loss function as follows:
[0150] ;
[0151] Among them, Loss is the loss value, is the kth model parameter, is a regularization function used to constrain the complexity of the model. is the model prediction value of the i-th sample, is the target true value of the i-th sample, is the loss function, which represents the error between the true value and the predicted value of the i-th sample, n is the total number of data, and K is the total number of trees.
[0152] In one embodiment of the present application, the first processing unit includes a determination module, and before constructing the loss function, the determination module is used to determine the regularization term as:
[0153] ;
[0154] in, is a regularization function used to constrain the complexity of the model. To control the number of leaf nodes, T is the number of leaf nodes, is the preset coefficient, is the score of the leaf node of node j.
[0155] In one embodiment of the present application, the first processing unit includes a first processing module and a second processing module.
[0156] The first processing module is used to divide the original training set into K-fold subsets, where the number of training samples in D is m, each subset has m / k training samples, and the corresponding subsets , and meet , ;
[0157] The second processing module is used to sequentially select the i-th fold from the divided subsets as the test set and the other K-1 folds as the training sets.
[0158] In one embodiment of the present application, the second processing unit includes a third processing module, a fourth processing module and a fifth processing module.
[0159] The third processing module is used for the selection step: selecting 1 fold of the K-fold data as the base test set, and the remaining K-1 folds as the above-mentioned base training set, and training to obtain the 1DCNN-ResNet1 model;
[0160] The fourth processing module is used for the prediction step: using the trained 1DCNN-ResNet1 model to predict the data in the base test set to obtain a first prediction result matrix;
[0161] The fifth processing module is used to adopt a K-fold cross-validation method to repeat the above selection step and the above prediction step until all K-fold predictions are completed to obtain a first final prediction result matrix.
[0162] In one embodiment of the present application, the second processing unit includes a sixth processing module, a seventh processing module, an eighth processing module and a ninth processing module.
[0163] The sixth processing module is used to determine that the training data of the meta-learner is a set of final prediction result matrices of all the base learners on the meta-training set, and use the set of the final prediction result matrices to train the SVM model;
[0164] The seventh processing module is used to use the 1DCNN-ResNet1 model, the 1DCNN-ResNet2 model, the 1DCNN-ResNet3 model, the 1DCNN-ResNet4 model, and the 1DCNN-ResNet5 model to predict all the sample data in the above meta-test set, and obtain 5 prediction result matrices, and average the prediction results of these 5 models to obtain a first average prediction result matrix; and replace 1DCNN-ResNet with XGBboost and RandomForest respectively, to obtain a second average prediction result matrix and a third average prediction result matrix respectively;
[0165] The eighth processing module is used to determine that the average prediction result matrix set is a set of the first average prediction result matrix, the second average prediction result matrix and the third average prediction result matrix;
[0166] The ninth processing module is used to use the trained SVM model to predict the data in the above average prediction result matrix set to obtain the SVM prediction result.
[0167] The fault identification device for partial discharge signal of the ring main unit includes a processor and a memory. The acquisition unit, the first processing unit, the second processing unit and the third processing unit are all stored in the memory as program units, and the processor executes the program units stored in the memory to realize corresponding functions. The modules are all located in the same processor; or, the modules are located in different processors in any combination.
[0168] The processor includes a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and the problem of poor accuracy of fault identification of partial discharge signals of ring main cabinets in the existing solution can be solved by adjusting kernel parameters.
[0169] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0170] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the fault identification method of the partial discharge signal of the ring main unit.
[0171] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein the fault identification method of the partial discharge signal of the ring main unit is executed when the program is run.
[0172] An embodiment of the present invention provides a device, comprising a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, at least the following steps are implemented: obtaining an original partial discharge fault data set; , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, where: is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples; 1DCNN-ResNet, XGBoost, and RandomForest are used as the base learners of the Stacking ensemble model, and the SVM model is used as the meta-learner, and the training set and test set of the original partial discharge fault data set D are used as the meta-training set and meta-test set of the meta-learner, and the meta-test set in the meta-learner is divided into a base training set and a base test set according to the K-fold cross-validation method; the base training set and the base test set are used to train and test each of the base learners to obtain the first base learner, the second base learner, and the third base learner, and the meta-training set and the meta-test set are used to train and test the meta learner to obtain the trained meta learner; the trained meta learner is used to identify the fault type of the newly collected partial discharge signal fault data to obtain the final identification result, and the first base learner, the second base learner, and the third base learner are used to verify the final identification result. The device in this article can be a server, PC, PAD, mobile phone, etc.
[0173] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that is initialized to have at least the following method steps: obtaining an original partial discharge fault data set , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, where: is the feature vector of the nth sample, is the predicted value corresponding to the nth sample, and N is the total number of samples; 1DCNN-ResNet, XGBoost, and RandomForest are used as base learners of the Stacking ensemble model, and the SVM model is used as a meta-learner, and the training set and test set of the original partial discharge fault data set D are used as the meta-training set and meta-test set of the meta-learner, and the meta-test set in the meta-learner is divided into a base training set and a base test set according to the K-fold cross-validation method; the base training set and the base test set are used to train and test each of the base learners to obtain a first base learner, a second base learner, and a third base learner, and the meta-training set and the meta-test set are used to train and test the meta-learner to obtain a trained meta-learner; the trained meta-learner is used to identify the fault type of the newly collected partial discharge signal fault data to obtain a final identification result, and the first base learner, the second base learner, and the third base learner are used to verify the final identification result.
[0174] The present application also provides a fault identification system for partial discharge signals of a ring network cabinet, the system comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for executing any of the above methods. The first base learner, the second base learner, and the third base learner are obtained by respectively training and testing the above base learners using the above base training set and the above base test set, and the meta-training set and the above meta-test set are used to train and test the above meta-learner to obtain the trained meta-learner, thereby using a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, thereby effectively avoiding the occurrence of insulation faults in the ring network cabinet, which has important practical significance for ensuring the safety and reliability of the power supply of the ring network cabinet, and at the same time improving the accuracy of fault identification of the partial discharge signal of the ring network cabinet, thereby solving the problem that the existing scheme has poor accuracy in fault identification of the partial discharge signal of the ring network cabinet.
[0175] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0176] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0177] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0178] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0180] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0181] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0182] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0183] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0184] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0185] 1) The method for fault identification of partial discharge signals of a ring main unit of the present application, by using the above-mentioned base training set and the above-mentioned base test set to respectively train and test each of the above-mentioned base learners, to obtain a first base learner, a second base learner and a third base learner, and using the above-mentioned meta-training set and the above-mentioned meta-test set to train and test the above-mentioned meta-learner, to obtain a trained meta-learner, thereby utilizing a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, thereby effectively avoiding the occurrence of insulation faults in the ring main unit, which has important practical significance for ensuring the power supply safety and reliability of the ring main unit, and at the same time improves the accuracy of fault identification of partial discharge signals of the ring main unit, thereby solving the problem of poor accuracy of fault identification of partial discharge signals of the ring main unit in the existing scheme.
[0186] 2) The fault identification device for partial discharge signals of the ring main unit of the present application, by adopting the above-mentioned base training set and the above-mentioned base test set to respectively train and test each of the above-mentioned base learners, obtains a first base learner, a second base learner and a third base learner, and adopts the above-mentioned meta-training set and the above-mentioned meta-test set to train and test the above-mentioned meta-learner, obtains a trained meta-learner, thereby utilizing a variety of differentiated models to observe the data space and structure from different angles, giving full play to the advantages of different models, thereby effectively avoiding the occurrence of insulation faults in the ring main unit, which has important practical significance for ensuring the power supply safety and reliability of the ring main unit, and at the same time improves the accuracy of fault identification of partial discharge signals of the ring main unit, thereby solving the problem of poor accuracy of fault identification of partial discharge signals of the ring main unit in the existing scheme.
[0187] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for identifying faults of partial discharge signals of a ring main unit, characterized in that: include: Get the original partial discharge fault dataset , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, wherein, For the n The feature vector of the samples, For the n The predicted value corresponding to the sample is N is the total number of samples; Using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and using the SVM model as a meta-learner, and using the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and dividing the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner; Using the base training set and the base test set to train and test each of the base learners respectively to obtain a first base learner, a second base learner, and a third base learner, and using the meta training set and the meta test set to train and test the meta learner to obtain a trained meta learner; Using the trained meta-learner to identify the fault type of newly collected partial discharge signal fault data to obtain a final identification result, and using the first base learner, the second base learner and the third base learner to verify the final identification result; The meta-learner is trained and tested using the meta-training set and the meta-test set to obtain a trained meta-learner, including: Determine that the training data of the meta-learner is a set of final prediction result matrices of all the base learners on the meta-training set, and use the set of final prediction result matrices to train the SVM model; Using the 1DCNN-ResNet1 model, 1DCNN-ResNet2 model, 1DCNN-ResNet3 model, 1DCNN-ResNet4 model, and 1DCNN-ResNet5 model to predict all sample data in the meta-test set, five prediction result matrices are obtained, and the prediction results of the five models are averaged to obtain a first average prediction result matrix; and the 1DCNN-ResNet is replaced with XGBboost and RandomForest, respectively, to obtain a second average prediction result matrix and a third average prediction result matrix; Determine that the average prediction result matrix set is a set of the first average prediction result matrix, the second average prediction result matrix and the third average prediction result matrix; The trained SVM model is used to predict the data in the average prediction result matrix set to obtain a final partial discharge fault type identification result.
2. The method according to claim 1, characterized in that In the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes: The XGBoost algorithm model is constructed as follows: ; in, is the model prediction value of the i-th sample, is the structure of the t-th independent tree, is the i-th data input, F is the i-th data input, and K is the number of trees.
3. The method according to claim 1, characterized in that: In the process of using 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, the method further includes: The loss function is constructed as: ; Among them, Loss is the loss value, is the kth model parameter, is a regularization function used to constrain the complexity of the model. is the model prediction value of the i-th sample, is the target true value of the i-th sample, is the loss function, which represents the error between the true value and the predicted value of the i-th sample, n is the total number of data, and K is the total number of trees.
4. The method according to claim 3, characterized in that Before constructing the loss function, the method further includes: The regularization term is determined as: ; in, is a regularization function used to constrain the complexity of the model. To control the number of leaf nodes, T is the number of leaf nodes, is the preset coefficient, for j The fraction of nodes that are leaf nodes.
5. The method according to claim 1, characterized in that Dividing the meta-test set in the meta-learner into a base training set and a base test set according to a K-fold cross-validation method, comprising: The original training set is divided into K-fold subsets, where the number of training examples in D is m, and each subset has m / k training examples. , and meet , ; From the divided subsets, take the i-th fold as the test set and the other K-1 folds as the training set.
6. The method according to claim 1, characterized in that The base training set and the base test set are used to train and test the base learners 1DCNN-ResNet, the RandomForest, and the XGBoost, respectively, to obtain a first base learner, a second base learner, and a third base learner, including: Selection step: select 1 fold of the K-fold data as the base test set, and the remaining K-1 folds as the base training set, and train to obtain a 1DCNN-ResNet1 model; Prediction step: using the trained 1DCNN-ResNet1 model to predict the data in the base test set to obtain a first prediction result matrix; The K-fold cross validation method is adopted to repeat the selection step and the prediction step until all K-fold predictions are completed, thereby obtaining a first final prediction result matrix.
7. A fault identification device for partial discharge signals of a ring main unit, characterized in that: include: Acquisition unit, used to acquire the original partial discharge fault data set , the original partial discharge fault data set D is divided into a training set and a test set according to a preset ratio, wherein, For the n The feature vector of the samples, For the n The predicted value corresponding to the sample is N is the total number of samples; A first processing unit is used to use 1DCNN-ResNet, XGBoost, and RandomForest as base learners of the Stacking ensemble model, and use the SVM model as a meta-learner, and use the training set and the test set of the original partial discharge fault data set D as the meta-training set and the meta-test set of the meta-learner, and divide the meta-test set in the meta-learner into a base training set and a base test set in a K-fold cross-validation manner; a second processing unit, configured to respectively train and test each of the base learners using the base training set and the base test set to obtain a first base learner, a second base learner, and a third base learner, and to train and test the meta learner using the meta training set and the meta test set to obtain a trained meta learner; a third processing unit, configured to use the trained meta-learner to identify the fault type of the newly collected partial discharge signal fault data to obtain a final identification result, and use the first base learner, the second base learner and the third base learner to verify the final identification result; The second processing unit includes a sixth processing module, a seventh processing module, an eighth processing module and a ninth processing module, The sixth processing module is used to determine that the training data of the meta-learner is a set of final prediction result matrices of all the base learners on the meta-training set, and use the set of the final prediction result matrices to train the SVM model; The seventh processing module is used to use the 1DCNN-ResNet1 model, the 1DCNN-ResNet2 model, the 1DCNN-ResNet3 model, the 1DCNN-ResNet4 model, and the 1DCNN-ResNet5 model to predict all the sample data in the meta-test set, and obtain 5 prediction result matrices, and average the prediction results of the 5 models to obtain a first average prediction result matrix; and replace the 1DCNN-ResNet with XGBboost and RandomForest, respectively, to obtain a second average prediction result matrix and a third average prediction result matrix, respectively; The eighth processing module is used to determine that the average prediction result matrix set is a set of the first average prediction result matrix, the second average prediction result matrix and the third average prediction result matrix; The ninth processing module is used to use the trained SVM model to predict the data in the average prediction result matrix set to obtain the SVM prediction result.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 6.
9. A fault identification system for partial discharge signals of a ring main unit, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 6.
Citation Information
Patent Citations
Short-term wind power integrated prediction method and system based on error correction
CN113361761A
Method for identifying ground glass pulmonary nodules
CN114692748A
Online monitoring method for partial discharge of ring main unit
CN117849552A