A Linear Motor Air Gap Fault Classification Method Based on Two-Level Reinforced Decision Boundaries
Through the two-level method of strengthening decision boundaries, combined with OVA, OVO decomposition and resampling strategies, a balanced data set is formed, which solves the problem of imbalanced multi-class data set classification in linear motor air gap fault diagnosis, and improves the comprehensive performance and fault diagnosis efficiency of the classifier.
Patent Information
- Application Number
- CN202210655023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The existing technology has an unbalanced multi-class data set classification problem in linear motor air gap fault diagnosis, resulting in insufficient global fault diagnosis and processing capabilities of classifiers, making it difficult to effectively improve classification efficiency and accuracy.
Using a two-level enhanced decision boundary method, data is collected through ranging and temperature measurement sensors, OVA and OVO decomposition strategies are used, and Borderline-SMOTE oversampling and RUS undersampling are combined to strengthen the classification decision boundary to form a balanced data set, and SVM is used as the base classifier for results integration.
The classifier performance in multiple categories of imbalanced scenarios has been improved, the classification efficiency and accuracy of linear motor air gap faults has been improved, and the global fault diagnosis capability has been enhanced.
Smart Images

Figure CN115099314B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data mining and fault detection, and particularly to a linear motor air-gap fault classification method based on a two-level enhanced decision boundary. Background Art
[0002] At present, in the urban rail transit subway system, the linear motor is the mainstream traction drive engine. The operating energy consumption of the linear motor is closely related to the size of the air-gap spacing of the linear motor. If the air-gap spacing of the linear motor is too large, a larger excitation current is required under the same train power; while if the air-gap spacing is too small, events such as motor scratching are likely to occur, resulting in motor burnout and even safety accidents, affecting the normal operation of the subway. Therefore, the daily detection and maintenance of the linear motor air-gap is an important link to ensure the safe and stable operation of the subway.
[0003] The data obtained from the linear motor air-gap detection system mainly relates to the states of the air-gap, slot wedge, and slot depth of the linear motor. By analyzing the offline data, it can be known that the faults associated with the linear motor air-gap are divided into the following five categories: slot wedge sinking, slot wedge loss, foreign object adhesion, circular runout, and single-end sinking. The occurrence probabilities of these five types of faults are different. The data set formed by combining the data of normal operation without faults and the data of the five types of faults in the subway is an imbalanced multi-class data set. Therefore, the task of statistically analyzing the data to obtain whether there is a fault and which fault actually occurs can be abstractly transformed into the task of classifying an imbalanced multi-class data set.
[0004] The solutions to the multi-class imbalanced data set classification problem are mainly divided into two types. One is the direct strategy, which modifies the existing algorithm to directly construct a more complex classifier; the other is the decomposition strategy, which decomposes the multi-classification problem to extend the solution method of binary imbalanced classification, solves it sequentially, and finally integrates the results. Since the latter is relatively simple, easy to implement, and the related theoretical technologies are mature, it is widely used.
[0005] The decomposition strategy consists of two parts: problem splitting in the training stage and result integration in the prediction stage. The decomposition methods for multi-class problems mainly include four types: OVO decomposition, OVA decomposition, ECOC decomposition, and hierarchical classification method. Among them, although the number of sub-classifiers in OVO decomposition is relatively large, its training speed is better than other decomposition methods, and the problem scale is smaller.
[0006] After decomposing the multi-class imbalance problem, the solutions for binary-class imbalance are used in sequence. Improvements in this part can be considered at the algorithm level and the data level. Among them, at the data level, resampling strategies are commonly used means to balance the dataset and improve the classification training performance of classifiers, including oversampling, undersampling, and synthetic sampling. Synthetic sampling combines the advantages of oversampling and undersampling and can effectively reduce the imbalance degree of the dataset. In the multi-class imbalance classification scenario, class overlap has always been a key issue, and whether the boundaries between classes can be clearly defined has a great impact on the classification results. In addition, a model with stronger classification performance can be constructed in a balanced data environment, and resampling, especially the synthetic sampling strategy that can improve the classification boundary, is particularly important.
[0007] Most of the methods currently used in the field of linear motor air-gap fault diagnosis are to collect on-site data by combining advanced sensor devices and apply data processing algorithms. For a certain type of fault, targeted methods are usually used for measurement and processing. Patent 1, "Linear Motor Stator Slot Wedge Fault Diagnosis Method Based on Statistical Analysis" (Application No.: CN202010456608.0, Publication No.: CN111596210A) and Patent 2, "Linear Motor Stator Foreign Object Adhesion Fault Diagnosis Method Based on Statistical Analysis" (Application No.: CN202010456610.8, Publication No.: CN111649706A) use distance sensors to measure the air-gap value and slot-wedge value of the motor stator, and conduct statistical analysis on the collected data to respectively judge the slot-wedge sinking fault and the stator foreign object adhesion fault. The disadvantage is that no unified dataset is formed, and only the occurrence or non-occurrence of a certain type of fault is judged in the article, lacking the ability to diagnose and process global faults. Summary of the Invention
[0008] The purpose of the present invention is to provide a linear motor air-gap fault classification method based on a two-level enhanced decision boundary, to improve the comprehensive performance of the classifier in the multi-classification of linear motor air-gap faults in a multi-class imbalance scenario, thereby improving the classification efficiency and accuracy of linear motor air-gap faults.
[0009] The technical solution to achieve the purpose of the present invention is: a linear motor air-gap fault classification method based on a two-level enhanced decision boundary, including the following steps:
[0010] Step 1: Use distance and temperature sensors to collect linear motor air-gap slot-wedge data, form an unbalanced multi-classification dataset of linear motor air-gap faults, and set the original dataset T=(A1, A2,..., A N ), A i represents the i-th class, i∈N, N is the total number of sample categories, and the number of samples in each category is respectively expressed as C1, C2,..., C N, the original dataset T is randomly divided into 70% as the training set T1 and 30% as the test set T2 according to categories; the training set T1 contains a total of C Total samples. Using the OVA decomposition strategy for the original dataset, it is completely split into m1 = N pairs of binary sub-datasets (Pure(i), Impure(i)), where pure(i) represents pure class samples and Impure(i) represents non-pure class samples. Calculate the total proportion of the number of pure class samples where i = 1, 2,..., m1, C Pure (i) represents the number of pure class samples, and C Impure (i) represents the number of non-pure class samples;
[0011] Step 2: Calculate the numerical ratio of the pure class samples to the non-pure class samples in each group Perform the first-level decision boundary strengthening, and introduce a preset control factor Calculate the synthetic sample allocation weights and allocation sequences. When τ i < 1, perform improved Borderline-SMOTE oversampling on the pure class samples; when τ i > 1, perform RUS undersampling on the non-boundary sample set of the pure class samples, where i = 1, 2,..., m1;
[0012] Step 3: Update the number of all pure class samples obtained in Step 2, and combine all pure class samples to form a new multi-class imbalanced dataset T1′, which contains a total of C Total ′ samples. Use the OVO decomposition strategy for it and completely split it into pairs of binary sub-datasets (Minority(i), Majority(i)), and calculate the total proportion of the number of minority class samples where i = 1, 2,..., m2, C Minority (i) represents the number of minority class samples, and C Majority (i) represents the number of majority class samples;
[0013] Step 4: Perform the second-level decision boundary strengthening on each pair of binary sub-datasets in turn, that is, introduce a preset control factor Set the loop process L and the predetermined imbalance degree π preset : In the hth loop, calculate the synthetic sample allocation weights and allocation sequences, perform improved Borderline-SMOTE oversampling on the minority class samples, update the number of minority class samples, and then perform RUS undersampling on the non-boundary sample set of the majority class samples, update the number of majority class samples, and calculate the imbalance degree π h ; After multiple loops, a relatively balanced dataset is obtained, that is, π h = π preset, terminate the above loop L, and perform the above loop operation on the next group of secondary sub-datasets until the second-level decision boundary strengthening of all secondary sub-datasets is completed;
[0014] Step 5: Conduct classification training and testing on the balanced multi-class dataset after the two-level decision boundary strengthening in Step 4. Use SVM as the base classifier, train the base classifier with the balanced binary sub-datasets, use the trained classifier to predict the test set T2, adopt a weighted voting strategy for result integration to obtain the final prediction result of the linear motor air-gap fault classification, and calculate macro-F1 as the evaluation index.
[0015] Compared with the prior art, the significant advantages of the present invention are as follows: (1) By effectively balancing the training samples and strengthening the classification decision boundary, the comprehensive classification performance of the classifier is improved; (2) The classification efficiency of the linear motor air-gap fault is improved. Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the main process of the linear motor air-gap fault classification method based on a two-level strengthened decision boundary of the present invention.
[0017] Figure 2 It is a comparison chart of the classification prediction effects of TL-SDB and RUS-OVO, SMOTE-OVO, and BSMOTE-OVO on 4 groups of UCI public multi-class unbalanced datasets in the embodiments of the present invention.
[0018] Figure 3 It is a comparison chart of the classification prediction effects of TL-SDB and RUS-OVO, SMOTE-OVO, and BSMOTE-OVO on the linear motor air-gap multi-class unbalanced dataset in the embodiments of the present invention. Detailed Embodiments
[0019] Starting from the dataset perspective, the present invention uses distance measurement and temperature measurement sensors to collect the data of the linear motor air-gap slot wedges, summarizes typical features, forms an unbalanced multi-class dataset of linear motor air-gap faults, and applies a classification method based on the decision boundary strengthening method to obtain a well-performing multi-classifier from offline data, which can handle various linear motor air-gap fault diagnosis tasks including slot wedge sinking and dropping, motor temperature rise, foreign object adhesion, air-gap height increase, etc.
[0020] A linear motor air-gap fault classification method based on a two-level strengthened decision boundary of the present invention includes the following steps:
[0021] Step 1: Use distance measurement and temperature measurement sensors to collect the data of the linear motor air-gap slot wedges, form an unbalanced multi-class dataset of linear motor air-gap faults, and set the original dataset T = (A1, A2,..., A N ), A iDenote the \(i\)-th class, where \(i\in N\) and \(N\) is the total number of sample categories. The number of samples in each category is denoted as \(C_1, C_2, \ldots, C\) N , randomly divide the original dataset \(T\) into 70% as the training set \(T_1\) and 30% as the test set \(T_2\) according to the categories; the training set \(T_1\) contains a total of \(C\) Total samples. Use the OVA decomposition strategy for the original dataset, and completely split it into \(m_1 = N\) groups of binary sub-datasets pairs \((Pure(i), Impure(i))\), where \(Pure(i)\) represents pure class samples and \(Impure(i)\) represents impure class samples. Calculate the total proportion of pure class samples where \(i = 1, 2, \ldots, m_1\), \(C\) Pure (i) represents the number of pure class samples, and \(C\) Impure (i) represents the number of impure class samples;
[0022] Step 2: Calculate the numerical ratio of pure class samples to impure class samples in each group Perform the first-level decision boundary strengthening, introduce a preset control factor , calculate the synthetic sample allocation weights and allocation sequences. When \(\tau\) i < 1, perform improved Borderline-SMOTE oversampling on pure class samples; when \(\tau\) i > 1, perform RUS undersampling on the non-boundary sample set of pure class samples, where \(i = 1, 2, \ldots, m_1\);
[0023] Step 3: Update the number of all pure class samples obtained in Step 2, combine all pure class samples to form a new multi-class imbalanced dataset \(T_1'\), which contains a total of \(C\) Total ' samples. Use the OVO decomposition strategy for it and completely split it into groups of binary sub-datasets pairs \((Minority(i), Majority(i))\), calculate the total proportion of minority class samples where \(i = 1, 2, \ldots, m_2\), \(C\) Minority (i) represents the number of minority class samples, and \(C\) Majority (i) represents the number of majority class samples;
[0024] Step 4: Perform the second-level decision boundary strengthening on each group of binary sub-datasets in turn, that is, introduce a preset control factor Set the loop process \(L\) and the predetermined imbalance degree \(\pi\) preset : In the \(h\)-th loop, calculate the synthetic sample allocation weights and allocation sequences, perform improved Borderline-SMOTE oversampling on minority class samples, update the number of minority class samples, then perform RUS undersampling on the non-boundary sample set of majority class samples, update the number of majority class samples, and calculate the imbalance degree \(\pi\) h; After multiple cycles, a relatively balanced data set, i.e., π, is obtained h = π preset , terminate the above loop L, and perform the above loop operation on the next group of two-class sub-data sets until the second-level decision boundary reinforcement of all two-class sub-data sets is completed;
[0025] Step 5: Perform classification training and testing on the balanced multi-class data set after the two-level decision boundary reinforcement in Step 4. Use SVM as the base classifier, train the base classifier with the balanced two-class sub-data set, predict the test set T2 with the trained classifier, adopt a weighted voting strategy for result integration to obtain the final prediction result of the linear motor air gap fault classification, and calculate macro-F1 as the evaluation index.
[0026] As a specific implementation manner, the preset control factor described in Step 2 is specifically as follows:
[0027] Preset control factor is used to control the number of newly added and subtracted samples during the sampling process. The sampling ratios of simple class samples and non-simple class samples are μ i , μ i ′ respectively. The total number of newly added samples by Borderline-SMOTE oversampling during the first-level decision boundary reinforcement is: The total number of newly subtracted samples by RUS undersampling is: where is an intermediate quantity, where C Impure represents the number of non-simple class samples, C Pure represents the number of simple class samples, P represents the set of simple class samples, P′ represents the set of samples to be undersampled, and θ i is the total proportion of the number of simple class samples.
[0028] As a specific implementation manner, the allocation sequence described in Step 2 is specifically as follows:
[0029] Calculate the criticality of each boundary simple class sample Sort all criticalities from small to large to obtain the allocation sequence Φ i , and sequentially allocate the synthetic sample numbers to the boundary samples corresponding to the positions from front to back according to α i (j) until the upper limit value α i of the total number of synthetic new samples is reached, and then stop the allocation;
[0030] where D sum represents the sum of the Euclidean distances from the current boundary sample to all non-simple class samples in the K1 nearest neighbors, W sum represents the total number of all non-simple class samples in the K1 nearest neighbors, and Φi Represents a list after sorting the boundary simplex class samples.
[0031] As a specific implementation, the improved Borderline - SMOTE oversampling described in step 2 is as follows:
[0032] When performing SMOTE oversampling to synthesize new samples for the simplex class samples in the boundary region, an allocation weight is introduced. The allocation weight is defined as the proportion ω of the number of non - simplex class samples in the K1 - nearest neighbors of each simplex class sample in the boundary region to the total number of non - simplex class samples in the boundary region. i , ω i Calculation formula: The number of new samples synthesized for each boundary sample is
[0033] where P j is the number of non - simplex class samples in the K1 - nearest neighbors, v i is the number of boundary simplex class samples, and α i is the total number of new samples synthesized.
[0034] As a specific implementation, the preset control factor described in step 4 is as follows:
[0035] In a single - loop, the preset control factor is used to control the number of newly added / subtracted samples and the convergence speed of the imbalance degree π h during the sampling process. When the second - level decision boundary is strengthened, the total number of new samples added by Borderline - SMOTE oversampling is: The total number of new samples subtracted by RUS undersampling is: , where is an intermediate quantity,
[0036] where h represents the current loop number, the subscript i represents that the current two - class sub - dataset group number is the i - th group, C Majority (i) represents the number of majority - class samples, C Minority (i) represents the number of minority - class samples, θ i ′ represents the proportion of minority - class samples recalculated each time entering the loop, C Minority ′(i) represents the number of minority - class samples updated in the h - th loop, M i represents the current majority - class sample set, and M i ′ represents the majority - class sample set in the boundary region.
[0037] As a specific implementation, the allocation sequence described in step 4 is as follows:
[0038] Calculate the critical degree of each boundary minority class sample Sort all the critical degrees from small to large to obtain the allocation sequence Φ i ′, and sequentially allocate the synthetic sample numbers to the boundary samples corresponding to the positions from front to back according to α i ′ (j) until the upper limit value α i ′ of the total number of synthetic new samples is reached, and then stop the allocation;
[0039] where D sum ′ represents the sum of the Euclidean distances from the current boundary minority class sample to all non-simple class samples in its K2-nearest neighbors, and W sum ′ represents the total number of all non-simple class samples in the K2-nearest neighbors of the current boundary minority class sample, and Φ i ′ represents the list after sorting the boundary minority class samples.
[0040] As a specific implementation manner, the improved Borderline-SMOTE oversampling described in step 4 is as follows:
[0041] When performing SMOTE oversampling to synthesize new samples for the minorities in the boundary region in a single loop, an allocation weight is introduced. The allocation weight is defined as the proportion λ i of the number of majority class samples in the K2-nearest neighbors of each minority class sample in the boundary region to the total number of all majority class samples in the boundary region i . The calculation formula of λ The number of new samples synthesized for each boundary sample is
[0042] where Q j represents the number of majority class samples in the K2-nearest neighbors, and v i ′ represents the total number of majority class samples in the boundary region.
[0043] As a specific implementation manner, the loop process L described in step 4 is as follows:
[0044] Each time entering the loop, the minority class samples and majority class samples need to be updated, and the allocation weight and allocation sequence of the synthetic samples are recalculated, as well as the imbalance degree π h in the h-th loop; the termination condition of loop L is π h = π preset ;
[0045] where π preset is the preset imbalance degree to be achieved.
[0046] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] Embodiment
[0048] Combined with Figure 1 , the linear motor air gap fault classification method based on the two-stage enhanced decision boundary in this embodiment includes the following steps:
[0049] Step 1: Use distance measurement and temperature measurement sensors to collect linear motor air gap slot wedge data, form an unbalanced multi-class dataset of linear motor air gap faults, and set the original dataset T = (A1, A2,..., A N ), A i represents the i-th class, i ∈ N, N is the total number of sample categories, and the number of samples in each category is respectively expressed as C1, C2,..., C N , randomly divide 70% of the original dataset T into the training set T1 and 30% into the test set T2 according to categories; the training set T1 contains a total of C Total samples. Use the OVA decomposition strategy for the original dataset and completely split it into m1 = N groups of binary sub-dataset pairs (Pure(i), Impure(i)), where Pure(i) represents pure class samples and Impure(i) represents non-pure class samples, and calculate the total proportion of pure class sample quantities where i = 1, 2,..., m1, C Pure (i) represents the number of pure class samples, and C Impure (i) represents the number of non-pure class samples;
[0050] Step 2: Calculate the numerical ratio of each group of pure class samples to non-pure class samples Perform the first-level decision boundary enhancement, introduce a preset control factor Calculate the synthetic sample allocation weight and allocation sequence. When τ i < 1, perform improved Borderline-SMOTE oversampling on pure class samples; when τ i > 1, perform RUS undersampling on the non-boundary sample set of pure class samples, where i = 1, 2,..., m1;
[0051] Step 3: Update the number of all pure class samples obtained in Step 2, combine all pure class samples to form a new multi-class unbalanced dataset T1′, which contains a total of C Total ′ samples. Use the OVO decomposition strategy for it and completely split it into groups of binary sub-dataset pairs (Minority(i), Majority(i)), and calculate the total proportion of minority class sample quantities where i = 1, 2,..., m2, C Minority (i) represents the number of minority class samples, and C Majority (i) represents the number of majority class samples;
[0052] Step 4. Perform second-level decision boundary strengthening on each group of secondary sub-datasets in turn, that is, introduce a preset control factor Set the loop process L and the predetermined imbalance degree π preset : In the h-th loop, calculate the allocation weights and allocation sequences of the synthetic samples, perform improved Borderline-SMOTE oversampling on the minority class samples, update the number of minority class samples, then perform RUS undersampling on the non-boundary sample set of the majority class samples, update the number of majority class samples, and calculate the imbalance degree π h ; After multiple loops, obtain a relatively balanced dataset, that is, π h =π preset , terminate the above loop L, and perform the above loop operation on the next group of secondary sub-datasets until the second-level decision boundary strengthening of all secondary sub-datasets is completed;
[0053] Step 5. Perform classification training and testing on the balanced multi-class dataset after the two-level decision boundary strengthening in Step 4. Use SVM (RBF kernel) as the base classifier, train the base classifier with the balanced secondary sub-datasets, use the trained classifier to predict the test set T2, adopt a weighted voting strategy for result integration to obtain the final prediction result of the linear motor air gap fault classification, and calculate macro-F1 as the evaluation index.
[0054] Furthermore, when τ i <1 in Step 2, perform Borderline-SMOTE oversampling, specifically as follows:
[0055] Record each sample in the i-th simple class Obtain the K1 nearest neighbors by calculating the Euclidean distance from the sample to all other samples, and calculate the proportion of the number of non-simple class samples in the K1 nearest neighbors P j is the number of non-simple class samples in the K1 nearest neighbors. If 0.5 ≤ ω p(i) (j) <1, divide the samples of this simple class into boundary samples (taking one point x P(i) (j) as an example). If P0′ exists, add all the samples belonging to P in the boundary samples to the set P0′ without repetition to obtain P′. Let the number of boundary simple class samples be v i , and the sampling magnification be μ i , when performing SMOTE oversampling on the boundary samples, the total number of increased samples is:
[0056]
[0057] Derive from it Let the proportion of the number of non-simple-class samples among the K1 nearest neighbors of each simple-class sample in the boundary region to the number of non-simple-class samples in the boundary region be the weight coefficient ω assigned. i , The number of new samples assigned to each boundary sample for synthesis is Subsequently, calculate the criticality ζ of each boundary sample i , where D sum represents the sum of the Euclidean distances from the current boundary sample to all non-simple-class samples in the K1 nearest neighbors, and W sum represents the total number of all non-simple-class samples in the K1 nearest neighbors. Sort all criticalities from smallest to largest to obtain the assignment sequence Φ i , and sequentially assign the synthesized sample numbers to the corresponding boundary samples from front to back according to α i (j) . After each assignment, calculate the current total number of assigned samples: If then continue to assign α i (g) ; if reset the current assigned quantity to and stop the assignment. Finally, the g-th simple-class sample point to be assigned selects the α i (g) -th nearest neighbor simple-class sample for linear interpolation to synthesize the simple-class sample , r j is a random number between (0, 1), and l s is the Euclidean distance between x P(i) (j) and its s-th simple-class sample neighbor (s = 1, 2,..., α i (j) ).
[0058] Furthermore, when τ i > 1 in step 2, perform RUS undersampling, specifically as follows:
[0059] For the non-boundary region of the current class samples, that is, P - P′, the total number of samples is sum(P - P′), and μ i ′ is the sampling ratio. Randomly delete β i samples, and the intermediate quantity β i is calculated as follows:
[0060]
[0061] From β i = C Pure (i)×(1 - μ i ′), we can obtain
[0062] Further, for the h-th cycle described in step 4, taking the i-th group of Minority(i) (h) , Majority(i) (h) ) as an example, Borderline-SMOTE oversampling is performed on the minority class. Variables with subscript (h) indicate that the variable is in the h-th cycle, specifically as follows:
[0063] Set each sample in the i-th minority class Obtain the K2 nearest neighbors by calculating the Euclidean distance from the sample to all other samples, and calculate the proportion of the number of majority-class samples in the K2 nearest neighbors is the number of majority-class samples in the K2 nearest neighbors. If , divide the simple-class sample into a boundary sample (taking one point as an example), then add all the samples belonging to in the boundary sample to the set without repetition to obtain . Let the number of boundary minority-class samples be , and the sampling magnification be Recalculate the proportion of the minority class: When performing SMOTE oversampling on the boundary samples, the increase in the total number of samples is:
[0064]
[0065] where h (h ∈ N + ) is the number of the current cycle, used to control the convergence speed of the imbalance degree. From this, Let the proportion of the number of majority-class samples in the K2 nearest neighbors of each minority-class sample in the boundary region to the number of majority-class samples included in the boundary region be the allocation weight coefficient The number of new samples allocated to each boundary sample is Subsequently, calculate the critical degree of each boundary sample , where represents the sum of the Euclidean distances from the current boundary sample to all non-simple-class samples in the K2 nearest neighbors, represents the total number of all non-simple-class samples in the K2 nearest neighbors. Sort all the critical degrees from small to large to obtain the allocation sequence From the front to the back, sequentially allocate the synthetic sample numbers to the boundary samples corresponding to the positions according to . After each allocation, calculate the current total number of allocated samples: If Then continue the allocation ; If reset the current allocation quantity to and stop the allocation. Finally, the g-th sample point of the simplex class to be allocated selects the nearest neighbor simplex class samples for linear interpolation to synthesize the minority class samples is a random number between (0, 1), is the Euclidean distance to its t-th minority class sample neighbor .
[0066] Furthermore, the h-th loop described in step 4 performs RUS undersampling on the majority class, specifically as follows:
[0067] First, update the sample quantity of the minority class , and randomly delete samples with a quantity of from the non-boundary region of the majority class samples, that is to obtain a balanced data set. Therefore, there is: the intermediate quantity β i 's calculation formula is is the sampling magnification, and then from it can be obtained that:
[0068]
[0069] Update the majority class sample quantity, , calculate the imbalance degree The h-th loop ends. When π h = π preset , terminate the loop process L.
[0070] So far, the method of the present invention can obtain a total of m2 balanced two-class sub-data sets, which can be used for subsequent classifier testing and comparison.
[0071] Furthermore, the training and testing described in step 5:
[0072] The basic classifier uses SVM, and the kernel function selects RBF. The SVM parameters, the kernel function coefficient γ i , the penalty coefficient C i and the parameters required by this method, are obtained through grid search technology. The classification prediction result is generated using the voting strategy. By obtaining the confidence (indicating the probability of predicting the test sample as class t1 relative to class t2) of each two-class SVM predicting the t-th test sample, calculate the prediction result - class:
[0073]
[0074] Further, the calculation of macro-F1 in step 5 is used as an evaluation index:
[0075] According to the confusion matrix constructed in Table 1, as well as the above prediction results and the actual categories of the test set, for each category, the precision Precision and recall Recall are calculated respectively, then the F1 value of each group is calculated, and finally the macro-F1 is calculated. The formula is as follows:
[0076]
[0077] Among them, Precision i represents the precision of the i-th category, and Recall i represents the recall of the i-th category; FP i represents the number of samples where the actual category and the predicted category are the same as the i-th category; FP i represents the number of samples where the predicted category is the same as the i-th category and the actual category is other categories; FN i represents the number of samples where the actual category is the same as the i-th category and the predicted category is other categories. Macro-F1 is the harmonic mean of precision and recall, taking into account both the precision and recall of the classification model, and can better judge the performance of the classifier. The higher the macro-F1 score, the better the prediction performance of the classification method used.
[0078] Based on the linear motor air gap fault classification method with a two-level enhanced decision boundary in this embodiment, the following experimental simulations are carried out:
[0079] Experiment 1: Select 4 groups of UCI public multi-class imbalanced datasets: Robot-Execution-Failures.lp1, Robot-Execution-Failures.lp4, Mechanical-Analysis, Steel-Plates-Faults. Use TL-SDB and RUS-OVO, SMOTE-OVO, BSMOTE-OVO to perform classification training on these 4 groups of datasets in turn, and compare the test effects. Set π preset = 1. Repeat the experimental steps of the method of the present invention 5 times for each group of datasets, and finally take the average value of the macro-F1 scores as the final result of the experiment on the current dataset.
[0080] Experiment 2: Select the multi-classification dataset of linear motor air gap faults, perform data preprocessing in advance, use TL-SDB and RUS-OVO, SMOTE-OVO, BSMOTE-OVO to perform classification training on this group of datasets respectively, and compare the test effects. Set π preset= {1.8, 1.6, 1.4, 1.2, 1.0}. A total of 5 sets of small experiments were conducted, and the experimental steps of the method of the present invention were repeated 5 times for each small experiment. Finally, the average value of the macro-F1 scores was taken as the final result of the current small experiment.
[0081] Confusion matrix constructed by the classification in Table 1
[0082]
[0083] Comparison list of the results of Experiment 1 in Table 2 (π preset = 1)
[0084]
[0085] (Note: The above evaluation index is macro-F1 (%))
[0086] The effect of Experiment 1 of the present invention is shown in Figure 2 and Table 2. Figure 2 On the horizontal axis in, Dataset1 to Dataset4 respectively correspond to 4 sets of multi-class imbalanced datasets: Robot-Execution-Failures.lp1, Robot-Execution-Failures.lp4, Mechanical-Analysis, Steel-Plates-Faults, and the vertical axis corresponds to the score value of the evaluation index macro-F1; from Figure 2 and Table 2, it can be seen that the performance of the comprehensive sampling method of the present invention on 4 sets of multi-class imbalanced datasets is slightly improved in terms of classification performance compared with BSMOTE-OVO and RUS-OVO, and is generally better than the SMOTE-OVO method. Thus, it can be seen that the present invention has good application effects in various imbalanced classification scenarios.
[0087] Comparison list of the results of Experiment 2 in Table 3
[0088]
[0089] (Note: The above evaluation index is macro-F1 (%))
[0090] The effect of Experiment 2 of the present invention is shown in Figure 3 and Table 3. Figure 3 On the horizontal axis, it respectively corresponds to the values of the predetermined imbalance degree, and the vertical axis corresponds to the score value of the evaluation index macro-F1; from Figure 3 and Table 3, it can be seen that in terms of the performance on the multi-classification dataset of the linear motor air gap fault, compared with RUS-OVO, the comprehensive sampling method of the present invention is more stable, and compared with SMOTE-OVO and BSMOTE-OVO, the classification performance of the method of the present invention is slightly improved.
[0091] In summary, the present invention reduces the imbalance of the data set through more meticulous comprehensive sampling, reduces the influence of sample imbalance on constructing the optimal classification hyperplane, and strengthens and improves the decision classification boundary. The experimental results show that the comprehensive classification performance of the method of the present invention is improved.
Claims
1. A linear motor air-gap fault classification method based on a two-level enhanced decision boundary, characterized in that It includes the following steps: Step 1: Use a ranging and temperature measuring sensor to collect linear motor air-gap slot wedge data, forming an unbalanced multi-class dataset of linear motor air-gap faults. Set the original dataset T = (A1, A2, …, A N ), where A i represents the i-th class, i ∈ N, and N is the total number of sample classes. The number of samples in each class is represented as C1, C2, …, C N . Randomly divide the original dataset T into 70% as the training set T1 and 30% as the test set T2 according to the classes. The training set T1 contains a total of C Total samples. Use the OVA decomposition strategy for the original dataset and completely split it into m1 = N groups of binary sub-datasets pairs (Pure(i), Impure(i)), where Pure(i) represents pure class samples and Impure(i) represents non-pure class samples. Calculate the total proportion of pure class sample numbers where i = 1, 2, …, m1, C Pure (i) represents the number of pure class samples, and C Impure (i) represents the number of non-pure class samples; Step 2: Calculate the numerical ratio of each group of simple class samples to non-simple class samples Perform first-level decision boundary strengthening and introduce a preset control factor Calculate the synthetic sample allocation weight and allocation sequence. When τ i < 1, perform improved Borderline-SMOTE oversampling on the simple class samples; when τ i > 1, perform RUS undersampling on the non-boundary sample set of the simple class samples, where i = 1, 2, …, m1; Step 3: Update the number of all simple class samples obtained in Step 2, combine all simple class samples to form a new multi-class imbalanced dataset T1′, which contains C Total ′ samples in total. Apply the OVO decomposition strategy to it and completely split it into groups of binary sub-dataset pairs (Minority(i), Majority(i)), and calculate the total proportion of the number of minority class samples where i = 1, 2, …, m2, C Minority (i) represents the number of minority class samples, and C Majority (i) represents the number of majority class samples; Step 4. Perform secondary decision boundary strengthening on each group of secondary sub-datasets in turn, that is, introduce a preset control factor Set the loop process L and the predetermined imbalance degree π preset : In the h-th loop, calculate the allocation weights and allocation sequences of the synthetic samples, perform improved Borderline-SMOTE oversampling on the minority class samples, update the number of minority class samples, and then perform RUS undersampling on the non-border sample set of the majority class samples, update the number of majority class samples, and calculate the imbalance degree π h ; After multiple loops, a relatively balanced dataset is obtained, that is, π h = π preset , terminate the above loop L, and perform the above loop operation on the next group of secondary sub-datasets until the secondary decision boundary strengthening of all secondary sub-datasets is completed; Step 5: Conduct classification training and testing on the balanced multi-class dataset after strengthening the two-level decision boundary in Step 4. Use SVM as the base classifier, train the base classifier with the balanced binary sub-dataset, predict the test set T2 with the trained classifier, adopt a weighted voting strategy for result integration to obtain the final prediction result of the linear motor air-gap fault classification, and calculate macro-F1 as the evaluation index.
2. The linear motor air-gap fault classification method based on a two-stage enhanced decision boundary according to claim 1, wherein The preset control factor described in Step 2 Specifically as follows: Preset control factor Used to control the number of newly added and reduced samples during sampling. The sampling multiples of simple class samples and non-simple class samples are μ i and μ i ′ respectively. When the first-level decision boundary is strengthened, the total number of newly added samples by Borderline-SMOTE oversampling is: The total number of newly reduced samples by RUS undersampling is: Where is an intermediate quantity, Where c Impure represents the number of non-simple class samples, C Pure represents the number of simple class samples, P represents the set of simple class samples, P′ represents the set of samples to be undersampled, and θ i is the total proportion of the number of simple class samples.
3. The linear motor air gap fault classification method based on a two-level enhanced decision boundary according to claim 1, wherein The allocation sequence described in Step 2 is specifically as follows: Calculate the critical degree of each boundary simplex class sample Sort all critical degrees from smallest to largest to obtain the allocation sequence Φ i , and sequentially allocate the synthetic sample numbers to the boundary samples corresponding to the positions from front to back according to α i (j) until the upper limit value α of the total number of synthetic new samples is reached i ; stop the allocation when this occurs Among which D sum represents the sum of the Euclidean distances from the current boundary sample to all non-simple-class samples of its K1 nearest neighbors, and W sum represents the total number of all non-simple-class samples of its K1 nearest neighbors, and Φ i represents the list after sorting the boundary simple-class samples.
4. The linear motor air gap fault classification method based on a two-stage enhanced decision boundary according to claim 1, wherein The improved Borderline-SMOTE oversampling described in Step 2 is specifically as follows: When introducing the allocation weight during the SMOTE oversampling of simplex class samples in the boundary region to synthesize new samples, the allocation weight is defined as the proportion ω of the number of non-simplex class samples among the K1 nearest neighbors of each simplex class sample in the boundary region to the total number of all non-simplex class samples in the boundary region. i , ω i The calculation formula of The number of new samples synthesized for each boundary sample is Where P j is the number of non-simple class samples among the K1 nearest neighbors, v i is the number of boundary simple class samples, and α i is the total number of newly synthesized samples.
5. The linear motor air-gap fault classification method based on a two-stage enhanced decision boundary according to claim 1, wherein The preset control factor described in step 4 is specifically as follows: In a single loop, a preset control factor is used to control the number of newly added / subtracted samples and the imbalance degree π during the sampling process h convergence rate When the second-level decision boundary is strengthened, the total number of newly added samples by Borderline-SMOTE oversampling is: The total number of newly subtracted samples by RUS undersampling is: wherein is an intermediate quantity where h represents the current number of iterations, the subscript i represents the i-th group of the current binary sub-dataset group, C Majority (i) represents the number of majority-class samples, C Minority (i) represents the number of minority-class samples, θ i ' represents the recalculation of the proportion of minority-class samples each time entering the loop, C Minority '(i) represents the updated number of minority-class samples in the h-th iteration, M i represents the current set of majority-class samples, M i ' represents the set of majority-class samples in the boundary region.
6. The linear motor air gap fault classification method based on a two-level enhanced decision boundary according to claim 1, characterized in that, The allocation sequence described in Step 4 is specifically as follows: Calculate the criticality of each boundary minority-class sample Sort all criticalities from smallest to largest to obtain the allocation sequence Φ i ', and sequentially allocate the synthetic sample numbers to the boundary samples corresponding to the positions from front to back according to α i ' (j) until the upper limit value α i ' of the total number of synthetic new samples is reached, and then stop the allocation; Among them, D sum ′ represents the sum of the Euclidean distances from the current boundary minority class samples to all non-simple class samples of the K2 nearest neighbors, and W sum ′ represents the total number of all non-simple class samples among the K2 nearest neighbors of the current boundary minority class samples, and Φ i ′ represents the list after sorting the boundary minority class samples.
7. The linear motor air-gap fault classification method based on a two-stage enhanced decision boundary according to claim 1, wherein The improved Borderline-SMOTE oversampling described in Step 4 is specifically as follows: When introducing the allocation weight during the SMOTE oversampling of the minority in the boundary region in a single loop, the allocation weight is defined as the proportion λ of the number of majority-class samples among the K2 nearest neighbors of each minority-class sample in the boundary region to the total number of majority-class samples in the boundary region i , λ i The calculation formula of The number of new samples synthesized for each boundary sample is Among which Q j represents the number of samples of the majority class among the K2 nearest neighbors, and v i ' represents the number of samples of the majority class contained in the boundary region.
8. The linear motor air gap fault classification method based on a two-stage enhanced decision boundary according to claim 1, characterized in that The loop process L described in Step 4 is specifically as follows: Each time the loop is entered, the minority-class samples and majority-class samples need to be updated, the allocation weights and allocation sequences of the synthetic samples are recalculated, and the imbalance degree π in the h-th round of the loop is calculated. h ; The termination condition of loop L is π h = π preset ; where π preset is the preset unbalance degree to be achieved.
Citation Information
Patent Citations
Linear motor stator slot wedge fault diagnosis method based on statistical analysis
CN111596210A
A Statistical Analysis-Based Method for Fault Diagnosis of Stator Slot Wedges in Linear Motors
CN111596210B
Linear motor stator foreign matter adhesion fault diagnosis method based on statistical analysis
CN111649706A
A multi-classification method based on adaptive balanced integration and dynamic hierarchical decision-making
CN109359704A
A circuit breaker imbalance monitoring data set oversampling method
CN112800917A