Fault diagnosis method and device based on sample active marking and deep forest

By using the fault diagnosis methods of sample active labeling and deep forests in complex mechanical equipment, samples are selected for annotation using uncertainty, marginal sampling or entropy sampling strategies, and the problems of scarcity and high labeling cost are solved by cascading forest optimization, and efficient fault diagnosis is achieved.

CN120277558AInactive Publication Date: 2025-07-08NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510769212.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the fault diagnosis of complex mechanical equipment, the scarcity of sample data and high labeling cost lead to limited diagnostic performance of deep learning models.

Method used

The fault diagnosis method based on sample active marking and deep forest is adopted. By collecting one-dimensional vibration signals of the gearbox, the samples are selected from the pool dataset for annotation using uncertainty sampling, marginal sampling or entropy sampling strategies, and the model features are progressively optimized through cascading forests to achieve effective training of the model.

Benefits of technology

Under small sample conditions, the accuracy and efficiency of fault diagnosis are significantly improved, the labeling workload and time cost are reduced, the dependence on large-scale labeling samples is reduced, and the fault diagnosis performance of complex equipment is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277558A_ABST
    Figure CN120277558A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a fault diagnosis method and device based on sample active marking and deep forest, and relates to the technical field of intelligent diagnosis of industrial equipment. According to the method, the importance of unlabeled samples is evaluated through an algorithm, and the samples beneficial to improving the model prediction precision are preferentially selected for manual labeling, so that the model can intelligently select and label the sample with the most information gain of the current model in the training process, and therefore, effective training of the model is realized at the least labeling cost; the problems of sample scarcity and high marking cost in fault diagnosis of complex equipment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent diagnosis for industrial equipment, and particularly to a fault diagnosis method and device based on sample active marking and deep forest. Background Art

[0002] Complex mechanical equipment usually operates in complex and harsh environments, and its fault characteristics are diverse and difficult to accurately capture. Taking the fault diagnosis of a gearbox as an example, when performing intelligent fault diagnosis, the performance of the model depends on a large number of high-quality labeled samples. However, it is difficult to obtain the actual sample data of equipment faults. Coupled with the fact that the sample annotation process requires professional knowledge support and often needs to be manually annotated by experts, it is not only time-consuming and laborious, but also extremely costly. In addition, due to the low fault incidence rate, the problem of scarce sample data is particularly prominent, seriously restricting the diagnostic performance of data-driven methods such as deep learning.

[0003] Therefore, how to effectively train the model under the condition of small samples and solve the pain points of sample scarcity and high annotation cost in complex equipment fault diagnosis has become a topic that needs to be further studied and improved. Summary of the Invention

[0004] Embodiments of the present invention provide a fault diagnosis method and device based on sample active marking and deep forest, which can effectively train the model under the condition of small samples and solve the problems of sample scarcity and high annotation cost in complex equipment fault diagnosis.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0006] In the first aspect, the method provided by the embodiments of the present invention includes:

[0007] S1. Collect one-dimensional vibration signals of the gearbox as sample data, and store the unlabeled sample data in the pool data set;

[0008] S2. Execute a query strategy in each round of active learning, annotate the sample data in the pool data set, and update the training set with the annotated sample data;

[0009] S3. After each update of the training set, retrain the deep forest model, where the features of the input samples in the model are progressively optimized through a cascaded forest;

[0010] Repeat S2 - S3 until it is detected that the performance of the deep forest model reaches the target accuracy, or all the sample data in the pool data set is marked;

[0011] S4. Use the deep forest model that has completed the iteration to detect the one-dimensional vibration signal during the actual operation of the gearbox.

[0012] S1 includes: constructing a training set and a test set by using the collected sample data, determining the number of samples in the initial training set and the number of samples selected for each round of active learning; normalizing the sample data in the training set and then performing multi-granularity scanning, extracting features of different granularities and performing feature transformation by using a random forest, and then splicing the transformed features to obtain a high-dimensional transformed feature vector as the initial input of the deep forest model.

[0013] Specifically, in S2, executing a query strategy to select sample data from the pool dataset for annotation, including: selecting the sample data with the lowest confidence from the pool dataset by using an uncertainty sampling strategy; and / or, selecting the sample data with the smallest difference in class probabilities from the pool dataset by using a margin sampling strategy; and / or, selecting the sample data with the highest prediction entropy value from the pool dataset by using an entropy sampling strategy.

[0014] Among them, the uncertainty sampling strategy includes: , where represents the input sample, that is, the original data predicted by the model; represents the true label, that is, the actual category to which the sample belongs; represents the predicted label, that is, the judgment result of the model on the sample category; represents the optimal sample selected by uncertainty sampling, and the model is most uncertain about its prediction; represents selecting the sample that makes the largest; represents selecting the sample that makes the smallest; represents the conditional probability when the model parameter is , that is, the confidence of the model predicting .

[0015] The margin sampling strategy includes: , where represents the sample selected by margin sampling, and the difference between the two categories with the highest prediction probabilities ( and ) is the smallest, indicating that the sample is near the classification boundary; and represent the category with the highest prediction probability and the second most likely category corresponding to .

[0016] The entropy sampling includes: , where represents the label that has a chance to be traversed, represents the sample selected by entropy sampling, and its information entropy is the largest, that is, the model is most uncertain about its classification; An index indicating traversal of all possible categories, used to calculate the entropy value.

[0017] Specifically, in S3, the progressive optimization of the features of the input samples in the model by the cascade forest includes:

[0018] The processing method for each layer of the cascade forest is: , where Indicates the output generated by the -th layer of the cascade forest, is the classification probability output of the -th decision tree for the input sample , indicates the number of trees in this layer, Indicates the processing function of the -th layer of the cascade forest, which generates a probability vector output for the input sample through decision trees, for feature transformation and classification probability representation; the prediction probability of the -th layer is , Indicates the prediction probability that the sample in the -th layer belongs to the category ;

[0019] The final decision result is obtained using the outputs of each layer and used as the result of progressive optimization, where , indicates the decision result, is the weight of the -th layer.

[0020] In a second aspect, the apparatus provided by the embodiments of the present invention includes:

[0021] A data partitioning module, configured to collect one-dimensional vibration signals of a gearbox as sample data, where the unlabeled sample data is stored in a pool dataset;

[0022] A sample updating module, configured to execute a query strategy in each round of active learning, label the sample data in the pool dataset, and update the training set with the labeled sample data;

[0023] A model optimization module, configured to retrain the deep forest model each time the training set is updated, where the features of the input samples in the model are progressively optimized by a cascade forest; until it is detected that the performance of the deep forest model reaches the target accuracy, or all the sample data in the pool dataset is labeled;

[0024] An operation testing module, configured to use the deep forest model that has completed iteration to detect the one-dimensional vibration signals during the actual operation of the gearbox.

[0025] Specifically, the data partitioning module is specifically configured to construct a training set and a test set by using the collected sample data, determine the initial number of training set samples and the number of samples selected for each round of active learning; perform normalization processing on the sample data in the training set, then perform multi-granularity scanning, extract features of different granularities, perform feature transformation using a random forest, and then splice the transformed features to obtain a high-dimensional transformed feature vector as the initial input of the deep forest model.

[0026] The sample updating module is specifically configured to select the sample data with the lowest confidence from the pool dataset through an uncertainty sampling strategy;

[0027] and / or, select the sample data with the smallest difference in class probabilities from the pool dataset through a margin sampling strategy;

[0028] and / or, select the sample data with the highest predicted entropy value from the pool dataset through an entropy sampling strategy.

[0029] In the embodiment of the present invention, the importance of unlabeled samples is evaluated by an algorithm, and samples that are helpful for improving the model prediction accuracy are preferentially selected for manual annotation, so that the model can intelligently select samples with the most information gain for the current model during the training process for annotation, thereby realizing the effective training of the model at the lowest annotation cost and solving the problems of sample scarcity and high annotation cost in complex equipment fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0031] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention;

[0032] Figure 2 It is a schematic framework diagram of the deep forest intelligent diagnosis model based on the sample active marking strategy provided by the embodiment of the present invention;

[0033] Figure 3 It is a flowchart of the active learning algorithm provided by the embodiment of the present invention;

[0034] Figure 4 It is a classification decision framework diagram of the deep forest provided by the embodiment of the present invention;

[0035] Figure 5 It is a physical and structural diagram of the gearbox provided by the embodiment of the present invention;

[0036] Figure 6 In the experimental cases provided by the embodiments of the present invention, the comparison of the diagnostic accuracy of the deep forest diagnostic method based on the sample active labeling strategy under the training-total data ratio of 30%-70%;

[0037] Figure 7 In the experimental cases provided by the embodiments of the present invention, the comparison of the initial diagnostic accuracy and the diagnostic accuracy after active learning;

[0038] Figure 8 In the experimental cases provided by the embodiments of the present invention, the confusion matrix of the deep forest intelligent diagnostic method based on the sample active labeling strategy under the training-total data ratio of 50%;

[0039] Figure 9 In the experimental cases provided by the embodiments of the present invention, the dot-line graph of the influence of the initial training set sample ratio on the classification accuracy;

[0040] Figure 10 In the experimental cases provided by the embodiments of the present invention, the comparison histogram of the influence of the initial training set sample ratio on the classification accuracy;

[0041] Figure 11 In the experimental cases provided by the embodiments of the present invention, the dot-line graph of the influence of the sample ratio of each round of learning on the classification accuracy;

[0042] Figure 12 In the experimental cases provided by the embodiments of the present invention, the comparison histogram of the influence of the sample ratio of each round of learning on the classification accuracy;

[0043] Figure 13 In the experimental cases provided by the embodiments of the present invention, the dot-line graph of the combined influence of the initial training set and the sample ratio of each round of learning;

[0044] Figure 14 In the experimental cases provided by the embodiments of the present invention, the comparison histogram of the combined influence of the initial training set and the sample ratio of each round of learning;

[0045] Figure 15 In the experimental cases provided by the embodiments of the present invention, the dot-line graph of the influence of the sample active labeling strategy based on uncertainty on the classification accuracy;

[0046] Figure 16 In the experimental cases provided by the embodiments of the present invention, the comparison histogram of the influence of the sample active labeling strategy based on uncertainty on the classification accuracy;

[0047] Figure 17 In the experimental cases provided by the embodiments of the present invention, the time taken for the sample active labeling strategy based on uncertainty to reach a specific classification accuracy;

[0048] Figure 18In the experimental case provided by the embodiment of the present invention, a line graph showing the influence of different sample active marking strategies on the classification accuracy;

[0049] Figure 19 In the experimental case provided by the embodiment of the present invention, a comparison histogram showing the influence of different sample active marking strategies on the classification accuracy. Detailed implementation manners

[0050] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. The implementation manners of the present invention will be described in detail below. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention. Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more of the associated listed items. Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless defined as herein.

[0051] The design concept of this embodiment is that through the introduction of a "human-machine collaboration" mechanism, active learning enables the model to intelligently select the samples with the most information gain for the current model during the training process, thereby significantly improving the diagnostic performance of the model at the lowest annotation cost. Its core idea is to evaluate the importance of unlabeled samples through algorithms and preferentially select the samples that contribute to improving the model's prediction accuracy for manual annotation, thus avoiding the ineffective annotation of a large number of redundant or unrepresentative samples. The main advantages of the sample labeling strategy based on active learning are as follows: First, it significantly reduces the workload and time cost of sample annotation and alleviates the problem of sample scarcity caused by difficult data annotation in complex mechanical equipment; second, it improves the training efficiency and performance of the model with small samples through efficient sample selection; third, it reduces the dependence on a large number of labeled samples, making the application of intelligent diagnosis technology in actual engineering more feasible and economical. Therefore, the sample labeling strategy based on active learning can well address the problems of sample scarcity and high annotation cost in complex equipment fault diagnosis with its high efficiency and adaptability. Its efficient learning ability in a small-sample data environment provides important technical support for improving the fault diagnosis accuracy of complex systems.

[0052] An embodiment of the present invention provides a fault diagnosis method based on sample active labeling and deep forest, as Figure 1 、 2 shown. The method framework of the deep forest intelligent diagnosis model based on the sample active labeling strategy is as Figure 2 shown, mainly including 4 links:

[0053] S1. Collect the one-dimensional vibration signal of the gearbox as sample data, and store the unlabeled sample data in the pool dataset;

[0054] S2. Execute the query strategy in each round of active learning, annotate the sample data in the pool dataset, and update the training set with the annotated sample data;

[0055] In supervised learning problems, there are problems that the labeling cost is relatively expensive and it is difficult to obtain a large number of labels. For some specific tasks, only industry experts can accurately label the samples. In this problem context, active learning (AL) attempts to train a better-performing model by selectively labeling less data. The active learning strategy attempts to overcome the annotation bottleneck by posing queries in the form of unlabeled instances and having them annotated by experts. In this way, the goal of active learning is to achieve high accuracy using as few labeled instances as possible, thereby minimizing the cost of obtaining labeled data. In this embodiment, as Figure 3As shown, the active learning algorithm model can be formalized as: Active learning result A = (C, L, S, Q, U). Among them, C represents one or more classifiers. When there are multiple models, an ensemble model can be used; L represents the labeled training sample set (Labeled Dataset), that is, the data for which label information has been obtained; U represents the entire unlabeled sample set (Unlabeled Dataset), that is, the set of samples that have not been labeled; S is the supervisor, responsible for manually labeling the samples in the unlabeled sample set U; Q is the query function or query strategy (Query Function or Query Strategy), used to query samples with a large amount of information from the unlabeled sample set U and submit these samples to the supervisor S for labeling. The working mechanism of active learning is an iterative process of training classifiers. The pseudo-code of the active learning algorithm framework designed in this embodiment is shown in Table 1 and mainly includes the following steps:

[0056] (1) Select a training model and train the model using a training set of a small number of labeled samples. At this time, the performance of the model is not high;

[0057] (2) Use the trained model to predict unlabeled samples;

[0058] (3) Define a query strategy (including measuring the uncertainty of the prediction and the query strategy applied to request labeling), return the priority score of the unlabeled samples according to the strategy, select the data that needs to be labeled, and perform manual labeling;

[0059] (4) Add the newly selected data to the training set to update the training set, and use the updated training set to train the model;

[0060] (5) Determine whether the model reaches the stopping criterion? If the stopping criterion is not reached, continue to use the query strategy to select the samples that need to be labeled and perform manual labeling, and loop steps 3 to 5 until the stopping criterion is reached;

[0061] (6) End the algorithm when the conditions are met.

[0062] Table 1 Pseudo-code of the active learning algorithm framework

[0063]

[0064] S3. After each update of the training set, retrain the deep forest model.

[0065] The deep forest model runs on a deep forest classifier. As an efficient ensemble learning method, deep forest can achieve high classification performance without extensive parameter tuning through multi-granularity scanning and feature extraction and optimization of cascaded forests. Combining active learning with deep forest can further improve the efficiency and accuracy of fault diagnosis tasks. Specifically, it includes:

[0066] Multi-granularity scanning feature extraction and sample information gain evaluation: The multi-granularity scanning module of the deep forest can scan one-dimensional vibration signals with different window sizes to extract local and global features. For example, in fault vibration signals, different windows can capture short-term impact features (such as periodic impact signals caused by rolling element defects) and long-term trend features (such as the cumulative effect of wear). When using the active learning strategy, each unlabeled sample extracts multi-granularity features through the initial stage of the deep forest, and the information gain of each sample is calculated in combination with the prediction probability. The sample with the highest information gain is preferentially labeled, and the model optimizes its feature extraction ability through these newly added samples. For example, samples of rare fault patterns will significantly change the multi-granularity feature extraction ability of the model, enabling the model to better capture similar fault features.

[0067] Progressive optimization of cascaded forests and dynamic sample update: Based on multi-granularity scanning, the deep forest model uses cascaded forests to progressively optimize features. The input of each level of cascaded forest is composed of the initial transformed features concatenated with the features extracted during the active learning process. During the training process of each level, cross-validation is used to evaluate the diagnostic performance: if the diagnostic performance does not reach the convergence standard, the model automatically grows a new level of cascaded forest; PCA is used for feature dimensionality reduction to alleviate the problem of high-dimensional feature redundancy and improve the training efficiency and diagnostic accuracy. During the active learning process, each newly added high-value sample is dynamically added to the training set after being labeled. In the cascaded structure of the deep forest, in each progressive optimization level, the features of the newly added samples are concatenated with the features of the original training data and input, significantly enhancing the classification ability of the model. There are advantages in hierarchical optimization. The features of different fault patterns in fault data may have complex hierarchies. For example, the features of rolling element faults and inner race faults may be similar in low-level features but gradually show obvious differences in high-level features. The cascaded forest can capture this feature hierarchy through progressive optimization. It also has the effect of dynamic sample update. The active learning strategy preferentially selects samples of rare fault patterns or near the decision boundary. These samples will significantly affect the high-level feature expression ability of the model during the optimization process of the cascaded forest. For example, after adding rare fault samples, the high-level features of the model will tend to reduce the confusion between the healthy state and the fault state, thereby improving the diagnostic accuracy.

[0068] Synergy between Feature Dimensionality Reduction and Information Gain Maximization: In each level of the cascading process of the deep forest, PCA is used to reduce the dimensionality of features, reduce the impact of redundant features, and improve computational efficiency. In the framework of active learning, PCA dimensionality reduction can also help highlight the key features of newly added samples. For example, when active learning selects a set of samples with high information gain, the features of these samples may contain the most discriminative parts of the fault patterns. Through PCA dimensionality reduction, redundant features unrelated to the newly added samples can be effectively removed, enabling the model to focus more on the high-value information of the current samples.

[0069] In the design of this embodiment, the combination of active learning and deep forest can, in the fault diagnosis task of the gearbox: (1) Improve data utilization efficiency: In the case of limited annotation resources, active learning can significantly reduce the number of samples that need to be annotated. For example, for rare rolling element fault data, the active learning strategy can quickly find those samples that are most helpful for the model to distinguish fault categories without having to annotate the entire dataset.

[0070] (2) Accelerate model convergence: Since active learning preferentially selects samples with high information gain, the deep forest can achieve a high diagnostic accuracy in fewer training rounds. For example, in fault data, the model can quickly optimize the decision boundary through a small number of annotated boundary samples, avoiding overlearning of healthy state data.

[0071] (3) Improve fault diagnosis accuracy: The rare pattern samples and boundary samples selected by active learning can significantly improve the classification performance of the deep forest, especially in the case of class imbalance. For example, active learning can increase the classification accuracy of the model in rare fault patterns from the original 60%-70% to over 90%, while maintaining a high classification accuracy for the healthy state.

[0072] Repeat S2 - S3 until the performance of the deep forest model reaches the target accuracy or all the sample data in the pool dataset are marked; where the model continuously obtains high-value samples from the pool data through active learning and iteratively optimizes until the performance converges. The specific process is as follows: Train the deep forest model using the initial training set and test the basic performance of the model; after each round of active learning, the model evaluates the accuracy of the test set and records the performance change curve; stop the iteration when the model performance reaches the target accuracy or all the data in the sample pool are marked; finally, use the iteratively completed model for classification prediction and analyze the diagnostic ability of the model.

[0073] S4. Use the iteratively completed deep forest model to detect the one-dimensional vibration signal during the actual operation of the gearbox.

[0074] In this embodiment, S1 includes: First, collect one-dimensional vibration signals and construct a training set and a test set. Determine the number of samples in the initial training set and the number of samples selected for each round of active learning, and use the remaining unlabeled samples as the pool data set. After normalizing the training data, input it into the multi-granularity scanning model. Extract features of different granularities through multi-window scanning and use random forests for feature transformation. Finally, splice and generate a high-dimensional transformed feature vector as the initial input of the deep forest.

[0075] In S2, execute the query strategy to select sample data from the pool data set for annotation, including: Active learning selects the most informative samples from the sample data in the pool data set through strategies such as uncertainty sampling, margin sampling, or entropy sampling, significantly improving the data utilization efficiency. For example: Uncertainty sampling selects the samples with the lowest prediction probability, margin sampling selects the samples with the smallest difference in class probabilities, and entropy sampling selects the samples with the highest prediction entropy value. The annotated samples will be dynamically added to the initial training set, and the deep forest model will be retrained after updating.

[0076] The query strategy in active learning is the core mechanism for the algorithm to select the most informative unlabeled samples. Its main role is to reduce the quantity requirement of labeled data and improve the learning efficiency by selecting the samples that are most helpful for improving the model performance. The design of the query strategy determines how the model identifies the samples that are most crucial for model optimization from the unlabeled data, enabling each annotation to maximize the improvement of the model. In tasks with a large amount of data but high annotation costs, this selection process can not only save annotation costs but also accelerate the learning speed of the model. Commonly used typical query strategies and sampling marking methods can be divided into three major categories: query strategies based on informativeness, query strategies based on representativeness, and query strategies based on expected improvements. The following research will be carried out separately from these three categories. The query strategy based on informativeness is a query strategy in active learning, aiming to improve the learning efficiency of the model by selecting the samples that can bring the most information to the model. The core idea of this strategy is that the model may show greater uncertainty or error when facing certain samples. Therefore, by preferentially labeling these samples, the learning effect of the model can be maximized. Informativeness is usually related to the current state of the model, and the selection of samples often depends on their importance or uncertainty for the model prediction results. Common query strategies based on informativeness include methods such as uncertainty-based sampling, disagreement-based sampling, and model-change-based methods, which accelerate the convergence of the model by reducing the uncertainty or error of the model.

[0077] The goal of the query strategy is to select the samples that are most uncertain about the current model. For example, when using a probability model for binary classification, the uncertainty sampling strategy only needs to query the samples whose posterior probability of the positive class is closest to 0.5. Usually, this method includes the following three measurement criteria:

[0078] Entropy

[0079] A common uncertainty sampling strategy uses entropy as the uncertainty measurement, and its formula is:

[0080]

[0081] where Traverse all possible labels. Entropy is an information - theoretic measure that represents the amount of information required to "encode" a distribution. Therefore, in machine learning, entropy is usually considered a way to measure uncertainty or purity. For binary classification problems, entropy - based uncertainty sampling is the same as selecting the samples with posterior probabilities closest to 0.5. Entropy - based methods can be easily extended to multi - label classifiers and probability models for more complex structures such as sequences.

[0082] Least Confidence

[0083] In more complex scenarios, in addition to entropy, another alternative is to query the optimal label samples with the lowest confidence. In the strategy using least confidence as the uncertainty measure, select the samples with the smallest maximum probability for annotation. That is, sort the maximum values of the selection probabilities for each data point from smallest to largest, and its formula is:

[0084]

[0085] where is the class label with the highest probability. For binary classification problems, this method is equivalent to the entropy - based strategy.

[0086] Margin

[0087] The third uncertainty measure method is margin sampling. This method is similar to the least - confidence method, but the difference is that this method considers the difference between the maximum probability and the second - maximum probability. Margin sampling selects the sample data that are very likely to be classified into two categories, and the probabilities of these data being classified into two categories are not very different, that is, select the samples with the smallest difference between the maximum and the second - maximum probabilities predicted by the model, and its formula can be expressed as:

[0088]

[0089] where and for are the class with the highest predicted probability and the second - most likely class of the model prediction, respectively.

[0090] In the fault diagnosis task, active learning based on the strategy of maximizing information gain significantly improves the learning efficiency and diagnostic performance of the deep forest classifier by preferentially selecting the most valuable samples for annotation. The multi - granularity scanning module dynamically optimizes the feature extraction ability through active learning, the cascaded forest module continuously improves the classification performance through the progressive optimization of new samples, and PCA dimensionality reduction further highlights the feature contribution of samples with high information gain. Finally, this combination can achieve efficient and high - precision intelligent fault diagnosis with limited annotation resources.

[0091] To achieve effective fault diagnosis under data complexity, the sample active labeling strategy must maximize the information gain. The active learning strategy needs to select those samples that can provide the most information for annotation. In fault diagnosis, some samples may be more critical for improving the classification accuracy, especially when the sample classes are imbalanced or there are rare fault patterns. These samples usually contain the features that the model is currently most uncertain about or most challenging. By maximizing the information gain brought by each newly added labeled data, the learning efficiency of the model under limited labeled data can be significantly improved. This not only helps to accelerate the model convergence but also achieve a high fault diagnosis accuracy with fewer labeled samples.

[0092] In fault diagnosis, the active learning strategy that maximizes the information gain effectively solves the learning efficiency problem under limited labeling resources by selecting the samples that are most valuable for the model performance. The fault data in actual industrial applications usually has the following characteristics: (1) Class imbalance: The healthy state data often accounts for the vast majority, while the data volume of minor faults or rare fault patterns is small, resulting in the model being prone to bias towards the classification of the healthy state. (2) Feature complexity: The vibration signals contain a large number of nonlinear and non-stationary features, and the feature differences between different fault classes may be very subtle. (3) High labeling cost: The actual acquisition and labeling of fault data require human or expert participation, especially the labeling cost of rare fault patterns is higher.

[0093] In such a context, the core goal of the active learning strategy is to maximize the model performance by selecting the samples with the largest information gain for annotation. These samples usually meet the following conditions: (1) High uncertainty in model prediction: The current model has a low classification confidence in these samples. (2) Boundary samples: These samples are close to the decision boundaries of different fault classes and have significant potential for improving the model's classification ability. (3) Rare pattern samples: These samples may belong to the fault classes with a small number in the dataset, and labeling these samples helps to alleviate the class imbalance problem.

[0094] Active learning quantifies the information gain through the following strategies: (1) Uncertainty sampling: Select the samples with the lowest prediction probabilities, that is, the model has the lowest confidence in these samples. (2) Margin sampling: Select the samples with the smallest difference in class prediction probabilities. These samples are near the classification boundary, and the decision boundary of the updated model will be more accurate. (3) Entropy sampling: Select the samples with the highest prediction entropy values, which reflects the most chaotic prediction distribution of the samples.

[0095] In this embodiment, the Deep Forest (DF) technology is proposed to address the problems of complex parameter tuning and high computational costs in deep neural networks (DNNs) when dealing with small and medium-sized datasets. Based on the idea of ensemble learning, the DF stacks decision tree models through a hierarchical structure, extracting features layer by layer and enhancing them. At each layer, multiple random forests or completely random tree forests output rich feature representations of the input data through classification probabilities and then pass them layer by layer to the next-level forest until the final layer gives the classification result through weighted voting or probability synthesis. The process is as Figure 4 shown. The DF does not require complex parameter tuning, can automatically determine the number of model layers, has strong generalization ability and robustness, and is especially suitable for small and medium-sized dataset classification tasks. While maintaining high accuracy, the DF greatly reduces the computational complexity and performs particularly well when the data volume is small compared to deep neural networks. The following details the algorithm process of DF:

[0096] (1) Input Features and Layer-by-Layer Feature Enhancement

[0097] Assume the input sample is , and its dimension is . The cascaded forest at each layer learns features based on the input and outputs classification probabilities or enhanced feature representations. For the -th layer, the output of the cascaded forest can be expressed as:

[0098]

[0099] where is the classification probability output of the -th decision tree for the input sample , and represents the number of trees in this layer. This output will be used as the input for the next layer, enhancing the input features.

[0100] (2) Layer-by-Layer Classification in the Cascaded Structure

[0101] Each layer of the forest makes predictions based on the current input data and the output of the previous layer, and the final decision is aggregated from the results of all layers. Assume the output generated by the cascaded forest of the -th layer is , then the prediction probability of this layer is:

[0102]

[0103] where represents the prediction probability that the sample belongs to the class at the -th layer, is the number of decision trees in this layer. The entire cascade structure will continue to stack until the model converges or achieves optimal performance on the validation set.

[0104] (3) Final Decision

[0105] When the feature extraction and classification processes of all layers are completed, the final classification decision is based on weighted voting or direct selection of the outputs of each layer. Suppose there are layers in the cascade structure, then the final decision can be expressed as:

[0106]

[0107] where, is the weight of the -th layer, which is used to integrate the prediction probabilities of each layer and finally determine which class the sample belongs to .

[0108] DeepForest has the ability to automatically determine the number of layers. By gradually increasing the forest layer by layer, the performance of the model on the validation set gradually improves until convergence. During this process, the number of layers is adaptively adjusted to ensure that the model is not overfitted or underfitted.

[0109] The following combines specific experimental cases to demonstrate the actual application effect of the solution of this embodiment:

[0110] The data is from the 2009 PHM Data Challenge. The data was collected from a two-stage standard cylindrical spur gear reducer. The reducer includes an input shaft, an idler shaft, and an output shaft, and its physical and structure are as Figure 5 shown.

[0111] The first-stage reduction ratio of the gearbox is 1.5, and the second-stage reduction ratio is 1.667. The input shaft speed used for data collection is 30 Hz. The sampling frequency is 66.7 kHz, the sampling time is 4 s, and vibration signals of 8 health states are collected in total. In this experiment, the dataset contains the health states of 8 types of gearboxes as shown in Table 2. The 8 health states include the normal state and seven types of fault states. Among them, the seven types of fault states are mixed faults of parts such as gears, bearings, and shafts. For example, health state 2 is a mixed fault of a broken 32-tooth gear and an eccentric 48-tooth gear, with the characteristics of mixed faults and multiple fault modes. The dataset has a total of 4000 samples, that is, each state contains 500 samples, and the length of each sample is 1024.

[0112] Table 2 Health States of Two-Stage Spur Gear Reducer

[0113]

[0114] An ablation experiment was conducted to compare the gearbox fault diagnosis method based on deep forest and the deep forest fault diagnosis method based on the sample active labeling strategy, and to explore the necessity and effectiveness of improving the gearbox fault diagnosis method based on the original deep forest by integrating the part of the sample active labeling strategy and the deep forest classifier. Since the sample length of each fault sample is 1024, that is, the original feature dimension is 1024 dimensions, the multi-granularity scanning window sizes were selected to be 64 dimensions, 128 dimensions, and 256 dimensions respectively. The number of random forests in the multi-granularity scanning structure is 2, and the number of decision trees in each random forest is 10. The number of random forests in the cascaded forest structure is 4, and the number of decision trees in each random forest is 100.

[0115] First, the effectiveness of the sample active labeling strategy under different training-total data ratios was verified. Experiments were conducted on the deep forest fault diagnosis method integrating the active labeling strategy when the training-total data ratios were 30%, 40%, 50%, 60%, and 70%. The number of samples in the initial training set and the number of samples learned in each round were 4% of the total number of samples. The active learning strategy adopted was the uncertainty-based sampling method, and the metric was margin sampling. Figure 6 It shows the variation of the diagnostic accuracy with the iteration of the active learning rounds under 5 training-total data ratios.

[0116] From Figure 6 it can be seen that under 5 different training-total data ratios, since the number of samples in the initial training set only accounts for 4% of all samples, that is, only 4% of the samples are labeled, when using the deep forest classifier to classify the fault samples, the classification accuracies under the 5 data ratios are only 62.54%, 65.09%, 71.63%, 74.81%, and 80.42%. This indicates that when a large number of samples lack correct labels and the effective information content of the samples included in the training set is small, it is difficult to meet the requirements of sample-labeled gearbox fault diagnosis only using the deep forest model as the classifier. After adding the active learning strategy, the uncertainty-based sampling method selects 4% of the samples that are most uncertain about the current model from the unlabeled pool data each time. For the margin sampling method used in the experiment, the selected sample data are those that are very easy to be judged as two classes. It can be seen from the figure that the diagnostic accuracy increases with the iteration of active learning, and the growth rate is relatively fast at the beginning when the amount of information is small, and gradually slows down when the amount of information is large, and finally reaches a diagnostic accuracy of more than 95%. For example, under the 10% training-total data ratio, the initial diagnostic accuracy was only 62.54%. After 1 iteration of sample active labeling, the diagnostic accuracy reached 71.09%, an increase of 8.55%. Finally, after 10 iterations of sample active labeling, the diagnostic accuracy reached 96.34%, with a total increase of 33.8%.

[0117] To more clearly show the comparison of diagnostic accuracy before and after active learning, as Figure 7 shown, the comparison of the initial diagnostic accuracy and the diagnostic accuracy after active learning is presented. It can be seen from the figure that after 10 rounds of active learning, the final accuracies are respectively improved to 96.34%, 97.05%, 98.15%, 98.62% and 98.92%. This shows that the active learning strategy can significantly improve the diagnostic accuracy under different data ratios. Further analysis reveals that when the initial training-total data ratio is relatively low (such as 30% and 40%), the initial accuracy is relatively low, but the accuracy improvement brought by active learning is relatively larger; while at higher ratios (such as 60% and 70%), although the initial accuracy is relatively high, the improvement amplitude after active learning is relatively small. This indicates that active learning has a stronger improvement effect when the data labeling is relatively scarce, and as the training set ratio increases, the relative gain of active learning gradually decreases due to the improvement of the initial performance of the model. Overall, the active learning strategy can generally improve the diagnostic accuracy to a high level greater than 95% in most cases, verifying its effectiveness and robustness in the fault diagnosis task.

[0118] Figure 8 is the confusion matrix of the classification results after adding active learning under the condition of 50% training-total data ratio. Among them, "Predicted labels" is the predicted health status category, and "True labels" is the actual health status category. The numbers "0 - 7" in the figure respectively represent the 1st - 8th health status categories in Table 3. It can be seen that after adding the sample active labeling strategy, the diagnostic accuracy of each health status category exceeds 95%.

[0119] Combined with the above analysis, it can be concluded that the deep forest intelligent diagnosis model based on the sample active labeling strategy can effectively improve the accuracy of fault diagnosis according to a small number of labeled samples. Especially in practical application scenarios where the data labeling cost is high and the sample data volume is limited, this method shows significant advantages. Through experiments, it can be observed that under different training-total data ratios, the active learning strategy can effectively improve the model performance, and it is particularly prominent when the initial information is less. This shows that the active learning strategy can address the problem of insufficient information in the existing labeled samples, and by using the uncertainty sampling strategy for unlabeled samples, it preferentially selects samples with high information gain for model training, thereby improving the diagnostic efficiency and accuracy.

[0120] Further analysis reveals that the improvement in diagnostic accuracy is related to the training-total data ratio to a certain extent. At a relatively high training-total data ratio, due to the relatively large number of initially labeled samples, the initial performance of the model is better, so the performance gain after active learning iteration is relatively low; while at a relatively low training-total data ratio, although the initial diagnostic accuracy is low, the model performance can be significantly improved through active learning iteration. This indicates that the active learning strategy has greater application potential in the case of scarce sample labeling.

[0121] In addition, from the trend of the improvement in diagnostic accuracy in the experiment, it can be seen that in the initial stage of active learning, the model can quickly master the key features of diagnosis by learning a small number of high-value samples, thus rapidly improving the diagnostic accuracy. However, as the active learning iteration progresses, the information gain available for the model to learn in the unlabeled samples gradually decreases, and the improvement speed of the diagnostic accuracy tends to level off, which is consistent with the marginal effect of active learning. Finally, with the support of a small number of labeled samples, the model can still achieve a diagnostic accuracy of over 95%, which indicates that the sample active labeling strategy has broad application value in the field of intelligent diagnosis.

[0122] Secondly, study the influence of the number of samples in the initial training set and the number of samples learned in each round on the improvement of the model performance. At a training-total data ratio of 50%, the experiment is carried out in the following 3 steps.

[0123] (1) When the number of samples learned in each round is 2% of the total training data, study the diagnostic accuracy when the number of samples in the initial training set is 2%, 4%, 6%, 8% and 10% of the total training data respectively.

[0124] From the dotted line Figure 9 and Figure 10 it can be observed that: the larger the proportion of samples in the initial training set, the faster the classification accuracy improves and the higher the final accuracy. When the initial training set is 2%, the accuracy starts from 44.75% and reaches 94.45% after 10 rounds. When the initial training set is 10%, the accuracy starts from 82.85% and finally reaches 97.95%. A relatively small proportion of the initial training set will significantly reduce the initial performance of the model, but it can still be significantly improved after multiple rounds of learning. In addition, there is a phenomenon of diminishing marginal benefit of accuracy improvement, that is, when the initial training set is relatively large, the contribution of each subsequent round of learning to the accuracy weakens.

[0125] (2) When the number of samples in the initial training set is 2% of the total training data, study the diagnostic accuracy when the number of samples learned in each round is 2%, 4% and 6% of the total training data respectively.

[0126] From Figure 11 and Figure 12It can be seen that the larger the proportion of the learning samples in each round, the more significant the improvement in the classification accuracy. When the proportion of the learning samples in each round is 2%, the accuracy after 10 rounds is 94.45%; when it is 4% and 6% in each round, they are 96.05% and 96.65% respectively. The performance differences in the initial stage are not obvious, but the final accuracy is proportional to the proportion of the learning samples in each round. This shows that a larger proportion of the learning samples can provide more information, thus improving the model's fitting ability to the global data distribution.

[0127] (3) Respectively study the diagnostic accuracies when the number of learning samples in each round and the number of samples in the initial training set are 2%, 4% and 6% of the total training data.

[0128] From Figure 13 and Figure 14 it can be known that: The proportion of the initial training set and the proportion of the learning samples in each round jointly affect the final classification accuracy. When both are 2%, the initial accuracy is only 44.75%, and it is 94.45% after 10 rounds of iteration. When both are 6%, the initial accuracy is 78.3%, and it is 98.1% after 10 rounds of iteration. The proportion of the initial training set has a greater impact on the initial accuracy, and the proportion of the learning samples in each round has a more significant impact on the subsequent accuracy improvement. When the number of learning rounds is relatively large (such as after the 6th round), the impacts of both on the accuracy gradually tend to be stable.

[0129] Based on the above experiments and analyses, the following conclusions can be drawn: First, the number of samples in the initial training set and the number of learning samples in each round have a significant impact on the model performance. A larger number of samples in the initial training set can improve the initial accuracy of the model, providing a good foundation for subsequent learning. A larger number of learning samples in each round can accelerate the convergence speed of the model and further improve the final accuracy. Second, the joint optimization of the two can significantly improve the model performance, and increasing both the proportion of the initial training set and the proportion of the learning samples in each round can achieve balanced performance improvement in the initial stage and the learning process. Finally, there is a phenomenon of diminishing marginal benefits in the improvement of the model accuracy. Whether it is the proportion of the initial training set or the proportion of the learning samples in each round, when it is too large, the improvement effect on the accuracy tends to be flat. When applying the sample active marking strategy, in the case of limited data, the proportion of the initial training set should be preferentially increased to ensure the initial accuracy. For long-term learning tasks, the proportion of the learning samples in each round can be appropriately increased to accelerate the accuracy improvement.

[0130] Comparison of the sample active marking strategies based on uncertainty:

[0131] Selecting an appropriate sample active labeling strategy is crucial for improving the diagnostic accuracy and accelerating the speed of sample feature learning. In this section, by comparing the effects of three sample active labeling strategies based on uncertainty on the classification accuracy of fault diagnosis, the optimal strategy is expected to be found. The experiment adopted three different sample active labeling strategies: minimum confidence sampling, margin sampling, and entropy sampling. The training-total data ratio was 50%, the initial training sample percentage and the percentage of samples learned in each round were 4%, and the number of active learning rounds was 10 rounds.

[0132] The experimental results are as Figure 15 shown. It can be observed that as the number of learning rounds increases, the classification accuracy of all strategies shows an upward trend. In the initial stage, the classification accuracy of the minimum confidence sampling strategy is the lowest, but as the training progresses, its accuracy increases rapidly and is similar to that of other strategies after the 4th round of learning. The margin sampling strategy has a relatively high classification accuracy in the initial stage and maintains a stable growth throughout the training process, ultimately reaching a classification accuracy close to 100%. The entropy sampling strategy has a slightly lower classification accuracy than the margin sampling strategy in the initial stage, but also shows a good growth trend during the training process.

[0133] Figure 16 Further shows the comparison of the classification accuracy of different sample active labeling strategies in the initial stage and after 10 rounds of iteration. It can be seen from the figure that the margin sampling strategy has the highest classification accuracy in the initial stage, reaching 71.6%, while the minimum confidence sampling strategy has the lowest initial accuracy, only 65.7%. After 10 rounds of iteration, the classification accuracy of the margin sampling strategy increases to 98.4%, and the minimum confidence sampling strategy also increases to 98.1%, and their accuracies are very close. The classification accuracy of the entropy sampling strategy is 70.9% in the initial stage and increases to 97.5% after 10 rounds of iteration.

[0134] Based on the above analysis, the following conclusions can be drawn. First, among all the sample active labeling strategies, the margin sampling strategy shows the highest classification accuracy. As Figure 15 can be seen, the accuracy of the margin sampling strategy remains leading in all learning rounds, indicating that the margin sampling strategy has good performance in fault diagnosis, especially in application scenarios with extremely high accuracy requirements. Second, although the classification accuracy of the minimum confidence strategy is the lowest in the initial stage, its accuracy improvement speed is very fast and can quickly match the accuracy of other strategies. In Figure 16Among them, after 10 rounds of iteration, the accuracy of the minimum confidence strategy reached 98.1%, which is very close to the accuracy of the marginal strategy. This shows that the minimum confidence strategy has the characteristic of fast convergence, which is an important advantage for a fault diagnosis system that requires quick response. Finally, in the initial stage, the classification accuracy of the entropy strategy is lower than that of the marginal strategy and the minimum confidence strategy, as Figure 15 shown, its initial accuracy is 70.9%. Although the accuracy of the entropy strategy increased to 97.5% after 10 rounds of iteration, it is still lower than the final accuracy of the marginal strategy and the minimum confidence strategy. In addition, from Figure 16 it can be seen that the convergence speed of the entropy strategy is not as fast as that of the minimum confidence strategy. Therefore, considering the final classification accuracy and convergence speed, the entropy strategy performs relatively poorly among these three strategies.

[0135] Considering that the three sampling strategies have different iterative convergence speeds and the time taken to reach the established diagnostic accuracy target is also different, and the computing time is a key indicator of the model computing cost, a comparative study on the computing time of the three strategies is carried out below. The time taken for the model classification accuracy to reach 90%, 92%, 94% and 96% under the three sampling strategies is respectively counted, and the results are as Figure 17 shown.

[0136] Observing Figure 17 the computing time of different sampling strategies when reaching 90%, 92%, 94% and 96% classification accuracy, it is not difficult to find that the computing time of the marginal sampling strategy is significantly lower than that of the other two strategies when reaching the same accuracy. Specifically, the time taken for the marginal sampling strategy to reach the four accuracies is 666.5 seconds, 788.3 seconds, 923.1 seconds and 1124.5 seconds respectively, which is 76.7%, 80%, 80.2% and 83.6% of the time taken for the minimum confidence strategy to reach the four accuracies, and is 67.1%, 70.1, 74% and 77.6% of the time taken for the entropy strategy to reach the four accuracies. The result is less than 90% of the computing time of the other two strategies, indicating that the marginal sampling strategy has achieved the goal of saving no less than 10% of the computing cost.

[0137] Comparison of active labeling strategies for different categories of samples:

[0138] To verify the superiority of the uncertainty-based sample active labeling strategy selected in this paper, different types of labeling strategies were selected for further comparative experiments. The experiments compared the diagnostic performances of the sampling methods based on disagreement (information-theoretic query strategy), density-based sampling methods (representativeness-based query strategy), error-reduction sampling methods (expected improvement-based query strategy), and three sampling methods based on uncertainty. Through the active learning method, in each round, 4% of the unlabeled samples in the total samples were selected for labeling, gradually improving the performance of the classification model. In the experiment, the classification accuracy of each strategy increased with the increase in the number of learning rounds, and the differences in the initial classification accuracy and the final classification accuracy after 10 rounds of active labeling of each strategy were compared.

[0139] As Figure 18 shown, with the increase in the number of learning rounds, the classification accuracies of all sampling strategies increased significantly, indicating that active learning can effectively improve the classification performance of the model. However, there are differences in the classification accuracies of different strategies. Judging from Figure 19 the final results, margin sampling performed the best, reaching an accuracy of 98.4%; least confidence and entropy sampling followed, with accuracies of 98.1 and 97.5 respectively; the final accuracies of error-reduction, disagreement, and density sampling were lower than those of the three strategies based on uncertainty sampling, but reached relatively high levels of 97.3%, 97, and 96.69% respectively.

[0140] In the initial stage, margin sampling and entropy sampling had relatively high accuracies, both exceeding 70%, indicating that they could identify information beneficial to the model earlier. The initial accuracies of density sampling and error-reduction sampling were relatively low, especially error-reduction sampling (64%), probably because in the initial stage, the distribution of samples with large errors in the unlabeled data might not cover the overall data distribution well.

[0141] The significant increase in accuracy (by about more than 12%) of margin sampling in the first few rounds was due to the fact that the samples it focused on were usually near the decision boundary and had a greater impact on the model's decision. The improvement amplitudes of least confidence and entropy sampling were relatively uniform, reflecting the coverage ability of these two strategies for the overall data distribution.

[0142] The error reduction sampling strategy mainly selects samples with relatively large current prediction errors of the model for labeling. Its core idea is to reduce the prediction error of the model, so as to enable the model to converge quickly and improve the classification performance. First, in the initial stage of the experiment, the prediction error of the model is relatively large, and the error reduction sampling preferentially selects those samples with relatively high computational difficulty. These samples may be too complex or lack sufficient information, resulting in a sharp increase in computational cost and slow improvement of the model. Second, samples with large errors often lie on the edge of the decision boundary. These samples usually have a greater impact on the classifier. However, if the distribution of these samples is uneven, the features learned by the classifier may not be representative, thus affecting the final classification accuracy. Finally, the samples selected by the error reduction sampling strategy are not the types that the model currently urgently needs, but instead lead to excessive focus on local errors during the model training process, affecting the learning of global features.

[0143] The disagreement sampling strategy mainly selects samples with relatively large disagreements during the classification of the model for labeling, that is, the model's prediction results for some samples are inconsistent. Samples with relatively large disagreements are usually difficult to classify, and labeling these samples can help improve the model performance. The disagreement sampling emphasizes samples that show disagreements among multiple models or multiple iterations. In the experiment, the large disagreement of the model may be because the samples are near the decision boundary, and these samples often have less information. The disagreement of the model for these samples may not necessarily effectively improve the stability of the decision boundary. Therefore, labeling these samples may not effectively improve the classification accuracy. In addition, the disagreement of the model may only be concentrated in certain categories, resulting in the labeling of these samples not being able to effectively improve the performance of other categories, thus affecting the accuracy of the overall model.

[0144] The density sampling strategy selects samples distributed in the low-density regions of the data space for labeling, believing that the samples in these low-density regions are relatively rare and they have relatively large amounts of information. Labeling these samples can help the classifier better learn the distribution of the data. However, although the samples in the low-density regions are rare in space, they may not represent the core features of the data, or their distribution overlaps more with the samples of other categories. The model may have a high bias on these samples, resulting in slow progress during the training process. Second, the density sampling usually selects samples in the low-density regions for labeling, but these samples are not common in the data space. In the initial training, it is difficult for the model to obtain sufficient effective information. For datasets with complex distributions and significant overlaps between categories, the samples in the low-density regions cannot provide sufficient discrimination, resulting in an inefficient learning process for the model. At this time, labeling the samples in the low-density regions may reduce the classification accuracy of the model rather than improve the learning ability of the model.

[0145] Based on the above analysis, selecting the samples with the highest uncertainty of the model or the samples that are "ambiguous" between two classes for annotation according to the three active marking strategies based on uncertainty can more directly focus on the current learning state of the model. The selected samples contain relatively large amounts of information, directly reflecting the deficiencies of the model at a specific stage, and can effectively improve the training efficiency and classification accuracy.

[0146] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A fault diagnosis method based on sample active marking and deep forest, characterized in that It includes: S1. Collect the one-dimensional vibration signal of the gearbox as sample data. Among them, the unlabeled sample data is stored in the pool dataset; S2. Execute a query strategy in each round of active learning to label the sample data in the pool dataset, and update the training set with the labeled sample data; S3. After each update of the training set, retrain the deep forest model. Among them, the features of the input samples in the model are progressively optimized through a cascaded forest; Repeat S2 - S3 until it is detected that the performance of the deep forest model reaches the target accuracy, or all the sample data in the pool dataset is labeled; S4. Use the deep forest model that has completed the iteration to detect the one-dimensional vibration signal during the actual operation of the gearbox.

2. The method according to claim 1, wherein S1 includes: Use the collected sample data to construct a training set and a test set, and determine the number of initial training set samples and the number of samples selected in each round of active learning; After normalizing the sample data in the training set, perform multi-granularity scanning, extract features of different granularities and perform feature transformation using a random forest. Then, splice the transformed features to obtain a high-dimensional transformed feature vector as the initial input of the deep forest model.

3. The method according to claim 1, wherein In S2, when executing the query strategy to select sample data from the pool dataset for labeling, it includes: Select the sample data with the lowest confidence from the pool dataset through the uncertainty sampling strategy; And / or, select the sample data with the smallest difference in class probabilities from the pool dataset through the margin sampling strategy; And / or, select the sample data with the highest predicted entropy value from the pool dataset through the entropy sampling strategy.

4. The method according to claim 3, wherein The uncertainty sampling strategy includes: , where represents the input sample, represents the true label, represents the predicted label, represents the optimal sample selected through uncertainty sampling, for which the model's prediction is the most uncertain, represents selecting the sample that makes the largest, represents selecting the sample that makes the smallest, represents the conditional probability when the model parameters are .

5. The method according to claim 4, characterized in that, The marginal sampling strategy includes: , where represents the sample selected by marginal sampling, and the difference between the two classes with the highest predicted probabilities ( and ) is the smallest, reflecting that the sample is near the classification boundary; and represent the class with the highest predicted probability and the second most likely class corresponding to .

6. The method according to claim 5, characterized in that The entropy sampling includes: , where represents the label that has a probability of being traversed, represents the sample selected by entropy sampling, represents the information entropy, represents the index for traversing all possible categories for calculating the entropy value.

7. The method according to claim 1, wherein In S3, the progressive optimization of the features of the input samples in the model through the cascaded forest includes: The processing method for each layer of the cascade forest is as follows: , where represents the output generated by the cascade forest of the -th layer, is the classification probability output of the -th decision tree for the input sample , represents the number of trees in this layer, represents the processing function of the cascade forest of the -th layer, which generates a probability vector output for the input sample through decision trees for feature transformation and classification probability representation; the prediction probability of the -th layer is , represents the prediction probability that the sample of the -th layer belongs to the category ; Obtain the final decision result using the outputs of each layer, and use it as the result of progressive optimization, where represents the decision result, is the weight of the layer.

8. A fault diagnosis device based on sample active labeling and deep forest, characterized in that A data division module, which is used to collect the one-dimensional vibration signal of the gearbox as sample data. Among them, the unlabeled sample data is stored in the pool dataset; A sample update module, which is used to execute a query strategy in each round of active learning to label the sample data in the pool dataset, and update the training set with the labeled sample data; A model optimization module, which is used to retrain the deep forest model after each update of the training set. Among them, the features of the input samples in the model are progressively optimized through a cascaded forest; until it is detected that the performance of the deep forest model reaches the target accuracy, or all the sample data in the pool dataset is labeled; An operation test module, which is used to use the deep forest model that has completed the iteration to detect the one-dimensional vibration signal during the actual operation of the gearbox.

9. The device according to claim 8, wherein, The data division module is specifically used to use the collected sample data to construct a training set and a test set, determine the number of initial training set samples and the number of samples selected in each round of active learning; normalize the sample data in the training set and then perform multi-granularity scanning, extract features of different granularities and perform feature transformation using a random forest. Then, splice the transformed features to obtain a high-dimensional transformed feature vector as the initial input of the deep forest model.

10. The device according to claim 8 or 9, characterized in that, The sample update module is specifically used to select the sample data with the lowest confidence from the pool dataset through the uncertainty sampling strategy; And / or, the marginal sampling strategy selects the sample data with the smallest difference in class probabilities from the pool dataset; And / or, the entropy sampling strategy selects the sample data with the highest predicted entropy value from the pool dataset.

Citation Information

Patent Citations

  • Radar interference category recognition method and system

    CN112560596A