A Method for Detecting Charging Pile Faults and an Oversampling Algorithm

Through the improved Borderline-SMOTE oversampling algorithm, the problem of unbalanced positive and negative samples in charging pile fault detection is dealt with, new samples are synthesized and their authenticity is improved through the weight coefficient allocation mechanism, which solves the problem of low classification prediction accuracy in charging pile fault detection, and achieves higher detection accuracy.

CN114881166BActive Publication Date: 2025-06-27JIANGSU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210568605.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-06-27
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

In the fault detection of existing charging piles, due to the imbalance of positive and negative samples, the accuracy of classification prediction is low, especially when a few samples are close to or at the classification boundary, which increases the difficulty of classification tasks.

Method used

An improved Borderline-SMOTE oversampling algorithm is proposed. By identifying dangerous samples and synthesizing new samples in a few types of samples, the weight coefficient allocation mechanism is used to improve the authenticity of the new samples, thereby balancing the positive and negative samples.

Benefits of technology

It effectively enhances the data volume of a few types of samples in the training set, balances the sample proportion, improves the quality of machine learning and the performance of classifiers, and improves the accuracy of charging pile fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881166B_ABST
    Figure CN114881166B_ABST
Patent Text Reader

Abstract

The present invention provides an improved Borderline-SMOTE oversampling algorithm. Based on the Borderline-SMOTE algorithm, oversampling is performed on the sample set C. After identifying the danger samples, by introducing a weight coefficient allocation mechanism, the weights are redistributed to make the generated data more authentic, thereby improving the classification accuracy. Using the oversampling algorithm of the present invention can effectively enhance the data volume of the minority class samples in the training set, keep the balance between the minority class samples and the majority class samples, thereby improving the quality of subsequent machine learning and ensuring the performance of the original machine classification prediction. The present invention also provides a method for detecting charging pile faults. Before training the classifier, the collected training set samples are oversampled by the improved oversampling algorithm of the present invention, effectively ensuring the fault detection and recognition accuracy of the classifier used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of oversampling algorithms, and particularly relates to a method for detecting charging pile faults and an oversampling algorithm. Background Art

[0002] With the popularization of electric vehicles, the supporting charging piles have been gradually laid out in various parts of the city. Since most charging piles are installed outdoors, they are easily damaged by external environments such as weather. At the same time, due to the fact that the vast majority of charging piles are in an unattended state for a long time, it is difficult to detect charging pile faults in a timely manner.

[0003] By using relevant monitoring data and cooperating with corresponding machine learning algorithms to achieve remote intelligent fault classification detection based on the monitored data is a feasible means to solve the problem of charging pile fault monitoring. However, existing machine learning algorithms all require a large amount of effective data to be collected and trained in advance to achieve a high recognition accuracy. But due to the small amount of current charging pile fault data, the ratio of healthy data to fault data is seriously unbalanced, which seriously affects the accuracy of classification prediction. At the same time, some minority class samples are close to or on the classification boundary, increasing the difficulty of the classification task and thus affecting the performance of the classifier.

[0004] To overcome the imbalance problem between fault data and healthy data in the dataset, that is, the problem of imbalance between positive and negative samples, it is necessary to enhance the data of minority class samples that are close to or located on the classification boundary so that the number of the two types of samples can be kept as balanced as possible. In the prior art, algorithms such as Borderline - SMOTE have been used to perform oversampling on the training set to achieve data balance. However, in the process of expanding minority class samples by the existing Borderline - SMOTE algorithm, the newly synthesized samples still lack improvement in the actual training effect and are difficult to meet the requirements. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a method for detecting charging pile faults and an oversampling algorithm to solve the problem of imbalance between positive and negative samples and improve the accuracy of charging pile fault detection.

[0006] The present invention realizes the above technical objectives through the following technical means.

[0007] An improved Borderline - SMOTE oversampling algorithm: perform oversampling on the sample set C. After identifying danger samples, synthesize new samples through the following steps:

[0008] S1, for any sample M in the danger class sample set, find k neighboring samples in the minority class sample set and randomly select one of them as sample N;

[0009] S2. In the sample set C, find k nearest neighbors of sample M and sample N respectively. The numbers of minority class samples among the k nearest neighbors are slm and sln respectively, and calculate the ratio ratio = slm / sln.

[0010] S3. Synthesize a new sample S from sample M and sample N:

[0011]

[0012] where gap is the weight coefficient, and the value of gap is assigned according to ratio, and dif is the feature difference between sample N and sample M.

[0013] Furthermore, the value of the weight coefficient gap is:

[0014]

[0015] where rand(a, b) represents selecting a random number between a and b. When ratio = ∞ and slm = 0, then sample S is not synthesized.

[0016] Furthermore, the identification steps of the danger sample are: for each minority class sample, find h nearest neighbors from the sample set C. The number of majority class samples among the h nearest neighbors is h'. When h / 2 ≤ h' < h, this minority class sample is identified as a danger sample.

[0017] A method for detecting charging pile faults by applying the above oversampling algorithm: use a classifier to detect faults. Before training the classifier, use the oversampling algorithm to oversample the training set.

[0018] Furthermore, the classifier is a LightGBM integrated learning classifier.

[0019] Furthermore, the fault detection data of the classifier includes K1K2 drive signals, electronic lock drive signals, emergency stop signals, access control signals, voltage total harmonic distortion, and current total harmonic distortion.

[0020] Furthermore, perform five-fold cross-validation on the hyperparameters of the LightGBM integrated learning classifier to select the optimal hyperparameters.

[0021] Furthermore, the five-fold cross-validation is specifically: adopt the grid search cross-validation method, and through manual specification, perform an exhaustive search on a subset of the hyperparameter space to select the optimal hyperparameters.

[0022] Furthermore, the hyperparameters of the LightGBM classifier are set as follows: max_depth takes values from [3, 4, 5, 6, 7, 8], num_leaves takes values from [5, 10, 15, 20, …, 100], max_bin takes values from [5, 15, 25, …, 256], min_data_in_leaf takes values from [1, 11, 21, …, 102], Feature_fraction takes values from [0.6, 0.7, 0.8, 0.9, 1.0], Bagging_fraction takes values from [0.6, 0.7, 0.8, 0.9, 1.0], Bagging_freq takes values from [0, 10, 20, …, 81], Lambda_l1 takes values from [10 -5 , 10 -3 , 0.1, 0, 0.1, 0.3, 0.5, 0.7, 0.9, 1.0], Lambda_l2 takes values from [10 -5 , 10 -3 , 0.1, 0, 0.1, 0.3, 0.5, 0.7, 0.9, 1.0], min_split_gain takes values from [0.0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0].

[0023] The beneficial effects of the present invention are as follows:

[0024] (1) The present invention provides an improved Borderline - SMOTE oversampling algorithm, which can effectively increase the data volume of minority class samples in the training set, balance the minority class samples and the majority class samples, thereby improving the quality of subsequent machine learning and ensuring the performance of the original machine classification prediction.

[0025] (2) In the improved Borderline - SMOTE oversampling algorithm of the present invention, by introducing a weight coefficient distribution mechanism, the weights in the Borderline - SMOTE algorithm are redistributed to make the generated data more authentic, thereby improving the classification accuracy.

[0026] (3) The present invention also provides a method for detecting charging pile faults. Before training the classifier, the training set samples collected are oversampled by the improved oversampling algorithm of the present invention, effectively ensuring the fault detection and recognition accuracy of the classifier used. Description of the Drawings

[0027] Figure 1 is the flowchart of the method for detecting charging pile faults of the present invention;

[0028] Figure 2 is the comparison diagram of the effects after processing charging pile data with different oversampling algorithms;

[0029] Figure 3 It is a comparison chart of the effects after applying different oversampling algorithms to IMS bearing life data;

[0030] Figure 4 It is the ROC curve of the charging pile data processed by the oversampling algorithm of the present invention;

[0031] Figure 5 It is the ROC curve of the IMS bearing life data processed by the oversampling algorithm of the present invention. Detailed implementation manners

[0032] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0033] I. Improved Borderline-SMOTE oversampling algorithm

[0034] For the sample set C used for machine learning, the class with a smaller number in the sample set C is the minority class sample, and the class with a larger number is the majority class sample. Let the set of minority class samples be D and the set of majority class samples be E, D = {D1, D2,..., D d}, E = {E1, E2,..., E e}, and d and e are the numbers of minority class samples and majority class samples respectively.

[0035] For the above sample set C, the present invention performs oversampling based on the Borderline-SMOTE algorithm, so that the sample data volumes of the two classes reach a 1:1 balance. The specific steps are as follows:

[0036] Step 1, for each sample D i (i = 1, 2,..., d) in the minority class sample set D, h neighboring samples are found from the sample set C respectively, where the number of majority class samples among these h neighboring samples is h′ (0 ≤ h′ ≤ h).

[0037] Step 2, judge the attribute of the sample D i through h′, where:

[0038] 1) When h′ = h, all h neighboring samples of D i are majority class samples, so the sample D i is identified as a noise sample and no subsequent operation is performed;

[0039] 2) When h / 2 ≤ h′ < h, the sample D i is identified as a danger sample;

[0040] 3) When 0 ≤ h′ < h / 2, sample D i is identified as a safe sample and no subsequent operations are performed.

[0041] Step 3. After identifying all danger samples, the original Borderline - SMOTE algorithm synthesizes multiple new samples by linearly interpolating danger samples and minority class samples through the following formula:

[0042] synthetic j = danger + rand(0,1)*dif j , (j = 1,2,…,s)

[0043] where rand(0,1) represents a random number selected from 0 to 1, and the definition of dif j is as follows: For a danger sample, find the s nearest samples in the minority class sample set D and calculate the distances between this danger sample and these s nearest samples respectively. However, in the above synthesis algorithm, the weight coefficient rand(0,1) is only randomly selected and lacks a corresponding judgment mechanism, resulting in poor effects of the synthesized new samples. Therefore, the present invention introduces the following weight coefficient distribution mechanism:

[0044] Step 3.1. Select an arbitrary sample M from the danger class sample set (i.e., the set composed of all danger samples); for sample M, find k nearest samples in the minority class sample set D and randomly select one sample N from these k nearest samples.

[0045] Step 3.2. In the sample set C, find the k nearest samples of sample M, and the number of minority class samples among these k nearest samples is slm; in the sample set C, find the k nearest samples of sample N, and the number of minority class samples among these k nearest samples is sln; calculate the ratio ratio = slm / sln.

[0046] Step 3.3. Synthesize a new sample S from sample M and sample N:

[0047]

[0048] where dif is the feature difference between sample N and sample M, gap is the weight coefficient, and the value of gap is determined according to the ratio ratio:

[0049] 1) When ratio = ∞ and slm = 0, both sample M and sample N are identified as noise samples and no synthesis of sample S is performed;

[0050] 2) When ratio = ∞ and slm ≠ 0, sample N is identified as a noise sample. Therefore, when synthesizing, it is necessary to be far away from sample N, so gap = 0;

[0051] 3) When ratio = 0, the sample M is identified as a noise sample. Therefore, when synthesizing, it is necessary to stay away from sample M, so gap = 1;

[0052] 4) When 0 < ratio < 0.5, the security level of sample N is higher than that of sample M. When synthesizing, it is necessary to get closer to sample N, so gap takes a random number between (1 - ratio, 1);

[0053] 5) When 0.5 ≤ ratio < 1, the security level of sample N is higher than that of sample M. When synthesizing, it is necessary to get closer to sample N, so gap takes a random number between (ratio, 1);

[0054] 6) When ratio = 1, the security levels of sample M and sample N are the same. At this time, gap takes a random number between (0, 1);

[0055] 7) When 1 < ratio < 2, the security level of sample M is higher than that of sample N. When synthesizing, it is necessary to get closer to sample M, so gap takes a random number between (0, 1 - 1 / ratio);

[0056] 8) When ratio ≥ 2, the security level of sample M is higher than that of sample N. When synthesizing, it is necessary to get closer to sample M, so gap takes a random number between (0, 1 / ratio).

[0057] In summary, the value formula of gap is as follows:

[0058]

[0059] In addition, when ratio = ∞ and slm = 0, the sample S is not synthesized.

[0060] Finally, new samples are synthesized by the above method to expand the number of minority class samples, making the number of minority class samples equal to that of majority class samples.

[0061] II. Charging Pile Fault Detection Method

[0062] For the charging pile fault detection, the present invention selects six characteristic information items, namely ① K1K2 drive signal, ② electronic lock drive signal, ③ emergency stop signal, ④ access control signal, ⑤ voltage total harmonic distortion, and ⑥ current total harmonic distortion, as the fault detection data; and uses the LightGBM integrated learning classifier to perform the corresponding charging pile fault detection. To ensure that the LightGBM classifier has a high recognition accuracy, five-fold cross-validation is performed on the hyperparameters of the LightGBM classifier. Specifically, the grid search cross-validation method is adopted, and through manual specification, an exhaustive search is performed on a subset of the hyperparameter space to select the optimal hyperparameters; finally, the relevant hyperparameter settings are shown in Table 1 below:

[0063] Table 1: Classifier Parameter Settings

[0064]

[0065] As Figure 1 shown in the flowchart of the charging pile fault detection method of the present invention, before using the LightGBM integrated learning classifier, a training set is first constructed by collecting charging pile data containing the above six features to train the classifier. Since the amount of charging pile fault data is small and the proportion of positive and negative samples in the original training set is unbalanced, before training, it is necessary to first oversample the training set through the above improved Borderline - SMOTE to balance the positive and negative samples.

[0066] III. Effect Test

[0067] The healthy data and fault data are used to construct a training set at a ratio of 2.5:1. Then, the above training set is oversampled using ① the oversampling algorithm of the present invention, ② the SMOTE algorithm, ③ the Borderline - SMOTE algorithm, ④ the Safe - Level - SMOTE algorithm, ⑤ the MWMOTE algorithm, ⑥ the ADASYN algorithm, and ⑦ the ASUWO algorithm respectively, so that the positive and negative samples in it reach a 1:1 balance. Finally, the LightGBM integrated learning classifier is trained using the above oversampled training sets respectively, and the corresponding F1 - score (F1 value), Recall (recall rate), and G - mean are calculated to verify the effect of the present invention.

[0068] To fully verify the effectiveness of the method of the present invention, in addition to collecting charging pile data, the existing publicly available IMS bearing life data is directly used for the above - mentioned comparative test. The results are respectively as Figure 2 and Figure 3 shown. It can be seen from the comparison of the two groups of tests that after oversampling the training set with unbalanced positive and negative samples using the algorithm of the present invention, the effect on machine learning training is significantly improved, and the processing effect of the algorithm of the present invention is also more excellent compared with the same type of oversampling algorithms.

[0069] As Figure 4 and Figure 5 shown, they are respectively the ROC curves of the training samples of electric vehicle charging pile data and IMS bearing life data after oversampling processing by the present invention. The more right - angled the ROC curve is and the closer the corner point is to (0,1), the better the effect of this training sample for fault classification; thus, it can be directly seen from Figure 3 and Figure 4 that the training sample data after oversampling by the present invention has a good classification effect.

[0070] The present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions or modifications that those skilled in the art can make fall within the protection scope of the present invention.

Claims

1. A method for detecting charging pile faults, characterized in that: Use a classifier to detect faults. Before training the classifier, use the following oversampling algorithm to oversample the training set: Oversample the sample set C. For each minority-class sample, find h neighboring samples from the sample set C. The number of majority-class samples among the h neighboring samples is h. ′ , when h / 2 ≤ h ′ < h, this minority-class sample is identified as a danger sample; after identifying the danger samples, synthesize new samples through the following steps: S1. For any sample M in the danger class sample set, find k neighboring samples in the minority class sample set and randomly select one of them as sample N; S2. In the sample set C, find the k neighboring samples of sample M and sample N respectively. The number of minority class samples in the k neighboring samples are slm and sln respectively, and calculate the ratio ratio = slm / sln; S3. Synthesize a new sample S from sample M and sample N: where gap is the weight coefficient, and the value of gap is assigned according to ratio, and dif is the feature difference between sample N and sample M; The value of the weight coefficient gap is: where rand(a, b) represents selecting a random number between a and b. When ratio = ∞ and slm = 0, the sample S is not synthesized.

2. The charging pile fault detection method according to claim 1, wherein: The classifier is a LightGBM ensemble learning classifier.

3. The method for detecting charging pile faults according to claim 1, wherein: The fault detection data of the classifier includes K1K2 drive signals, electronic lock drive signals, emergency stop signals, access control signals, voltage total harmonic distortion, and current total harmonic distortion.

4. The method for detecting charging pile faults according to claim 2, characterized in that: Perform five-fold cross-validation on the hyperparameters of the LightGBM ensemble learning classifier and select the optimal hyperparameters.

5. The method for detecting charging pile faults according to claim 4, wherein: The specific process of the five-fold cross-validation is as follows: adopt the grid search cross-validation method, and through manual specification, perform an exhaustive search on a subset of the hyperparameter space to select the optimal hyperparameters.

6. The method for detecting charging pile faults according to claim 5, wherein: The hyperparameters of the LightGBM integrated learning classifier are set as follows: max_depth takes values from [3, 4, 5, 6, 7, 8], num_leaves takes values from [5, 10, 15, 20, …, 100], max_bin takes values from [5, 15, 25, …, 256], min_data_in_leaf takes values from [1, 11, 21, …, 102], Feature_fraction takes values from [0.6, 0.7, 0.8, 0.9, 1.0], Bagging_fraction takes values from [0.6, 0.7, 0.8, 0.9, 1.0], Bagging_freq takes values from [0, 10, 20, …, 81], Lambda_l1 takes values from [10 -5 , 10 -3 , 0.1, 0, 0.1, 0.3, 0.5, 0.7, 0.9, 1.0], Lambda_l2 takes values from [10 -5 , 10 -3 , 0.1, 0, 0.1, 0.3, 0.5, 0.7, 0.9, 1.0], min_split_gain takes values from [0.0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0].

Citation Information

Patent Citations

  • Power transformer fault sample equalization and fault diagnosis method based on Borderline SMOTE

    CN111832664A

  • Oversampling method based on spectral clustering

    CN112418352A