Online fault diagnosis method for industrial streaming data based on deep knowledge distillation network

Through deep knowledge distillation network and self-supervised learning methods, the problems of data privacy protection and diagnostic performance degradation in industrial streaming data are solved, and the accuracy and rapid generalization of online fault diagnosis are achieved, which is suitable for intelligent fault diagnosis of mechanical equipment.

CN117113169BActive Publication Date: 2025-09-26SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310918470.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2025-09-26
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

When dealing with industrial streaming data, existing technologies do not fully consider the need for data privacy protection, and the diagnostic performance of the model deteriorates in new and old scenarios, making it impossible to achieve effective fault identification and diagnosis.

Method used

A method based on deep knowledge distillation network is adopted. By maximizing information entropy technology and self-supervised learning, unlabeled data is used for model adaptation, and the knowledge distillation mechanism of confidence voting is combined to update the model to achieve online fault diagnosis of streaming data.

Benefits of technology

Without infringing on data privacy, the accuracy of fault identification and the rapid generalization ability of the model are improved, real-time diagnosis of new working conditions is achieved, the diagnostic accuracy on old data is maintained, and the problem of real-time matching between data collection and streaming computing is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117113169B_ABST
    Figure CN117113169B_ABST
Patent Text Reader

Abstract

The present invention discloses an online fault diagnosis method for industrial streaming data based on a deep knowledge distillation network, comprising the following steps: (1) collecting raw vibration signals of mechanical equipment under multiple different operating conditions, dividing them into a source domain and multiple target domains, and intercepting data of the same length to form training and test samples; (2) designing and constructing a network structure, wherein the overall network framework includes: a feature extractor and a fault classifier; (3) using labeled source domain samples to perform supervised training on the model to obtain an initial source model; (4) combining a knowledge distillation algorithm based on probability confidence and a maximum information entropy algorithm to train the model; (5) inputting the test samples into the trained model to obtain the fault classification results of the model in each domain (each task). The present invention can achieve effective diagnosis of incremental faults in industrial streaming data while protecting data privacy, and has high industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mechanical fault diagnosis, and in particular to an online fault diagnosis method for industrial streaming data based on a deep knowledge distillation network. Background Art

[0002] With the rapid development of next-generation artificial intelligence (AI) technologies, my country's machinery manufacturing industry has begun its intelligent transformation and is gradually entering the digital age. During the ongoing production and operation of machinery and equipment, key components are often prone to failure. Due to the unpredictable and sporadic nature of equipment failures, these failures pose a serious threat to production safety. Therefore, intelligent fault diagnosis of machinery and equipment is crucial for reducing equipment downtime, developing planned maintenance, increasing economic benefits, and avoiding tragic disasters.

[0003] In recent years, intelligent fault diagnosis methods based on deep learning have been widely used in many practical industrial scenarios. Deep neural networks with powerful feature learning capabilities can be trained with large amounts of data to quickly extract fault features and achieve accurate diagnosis and classification. However, with the continuous production and operation of mechanical equipment, a large amount of data is constantly generated and needs to be stored and processed. This data is usually transmitted in the form of streams and is called industrial streaming data. Industrial streaming data is usually diverse, complex, continuous, real-time, and high-capacity. Data at different stages are usually collected from different operating conditions. When faced with large amounts of streaming data, the model may experience "catastrophic forgetting" of learned knowledge, resulting in serious performance degradation and the inability to achieve effective intelligent fault diagnosis in both new and old scenarios.

[0004] Furthermore, with the widespread adoption and application of the Internet of Things (IoT) technology, machinery and equipment have entered an era of shared data. This shared data may contain sensitive information such as the equipment's technical parameters, operating status, and even trade secrets. If leaked without protection, it could significantly impact the safety and stable operation of the equipment. Therefore, implementing data privacy policies is crucial for both protecting corporate trade secrets and ensuring the safe and stable operation of machinery and equipment. However, existing fault diagnosis methods typically require data samples from various sources to simultaneously train or fine-tune the model to ensure diagnostic performance.

[0005] Zhu Haiping et al. disclosed in Chinese invention patent CN113281048A a "rolling bearing fault diagnosis method and system based on relational knowledge distillation." This patented technology improves the real-time response efficiency and accuracy of the fault diagnosis model. However, this method requires newly generated fault data to be labeled and does not fully consider the privacy protection needs of industrial practice.

[0006] Therefore, it is necessary to study the fault diagnosis model update method based on the real-time stream computing framework under data privacy protection, so as to further improve the diagnostic accuracy of the model in dealing with the continuously emerging streaming data in industrial scenarios. Summary of the Invention

[0007] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides an online fault diagnosis method for industrial streaming data based on a deep knowledge distillation network. This method can effectively refine key diagnostic knowledge from a locally pre-trained model, so that the model can be quickly generalized to unknown new working conditions, thereby realizing continuous fault diagnosis of incremental sample sources in industrial streaming data.

[0008] To achieve the purpose of the present invention, the present invention provides an online fault diagnosis method for industrial streaming data based on a deep knowledge distillation network, which specifically comprises the following steps:

[0009] S1. Data collection: Multiple sensors are used to continuously collect original mechanical vibration signals of multiple working conditions from mechanical equipment. The samples collected under different working conditions are divided into a source domain D s and multiple target domains Where N represents the number of target domains, and the original vibration signal is intercepted by the sliding window sampling method to obtain data of the same length to form training and test samples.

[0010] Among them, the source domain It consists of labeled samples, where Represents source domain samples and follows the distribution X s represents the data distribution of source domain samples, Represents the fault category label corresponding to the source domain sample and follows the distribution Y s represents the data distribution of source domain sample labels, n s The number of samples representing the source domain; the target domain It consists of unlabeled samples, where Represents the sample of the jth target domain and obeys the data distribution P t j represents the data distribution of the j-th target domain sample, represents the number of samples in the jth target domain. The data distribution between the source domain and multiple target domains is different, that is, X s ≠P t 1 ≠P t 2 ≠...≠P t N , P t N Represents the data distribution of the Nth target domain sample.

[0011] S2. Model Construction: The network structure is constructed based on the specific diagnostic task and dataset information. Based on the characteristics of the one-dimensional vibration signal, a feature extractor f0 based on a one-dimensional convolutional neural network is first designed for feature extraction. Then, a fault classifier g0 with a softmax function is constructed to classify the fault type.

[0012] S3. Local model pre-training: First, use the Xavier initializer to initialize the weights and biases of the initial model M0. Then, input the labeled source domain samples into the designed model, and obtain the initial model M0 through supervised training.

[0013] Furthermore, the feature extractor f0 and the fault classifier g0 are pre-trained using labeled source domain samples. The standard cross entropy loss function is used as the loss function for model pre-training. The initial model is optimized by minimizing the fault classification error. The initial model pre-training loss function L s The calculation is as follows:

[0014]

[0015] Where C is the number of health status categories in the source domain.

[0016] S4, incremental learning stage 1: After the initial model M0 is trained, it is first uploaded to the first target domain At the same time, the Xavier initializer is used to initialize the weights and biases of the first stage diagnostic model M1; in the incremental learning stage, the maximum information entropy technology is first used to assist the model in adapting to the data distribution of the target domain; then the knowledge distillation method is used to obtain the initial model M0 in the target domain Soft labels on the data; finally, the maximum information entropy loss and knowledge distillation loss are combined to optimize the first stage diagnostic model M1;

[0017] S5. Testing phase 1: After training is completed, the data from the source domain and the target domain are tested. The test samples are simultaneously input into the trained model M1, and the fault diagnosis results in the two domains are automatically output.

[0018] S6, incremental learning stage j: the model {M0,...,M j-1}Upload to the jth target domain At the same time, use the Xavier initializer to initialize and construct the j-th stage diagnosis model M j In the incremental learning phase, we first use the information entropy maximization technique to help the model adapt to the data distribution of the target domain; then we use the knowledge distillation method based on probability confidence to obtain the model {M0,...,Mj-1 In the target domain Consistent diagnostic knowledge on the data; finally, the maximum information entropy loss and knowledge distillation loss are combined to optimize the j-th stage diagnostic model M j ;

[0019] S7, testing phase j: after training is completed, the data from the source domain and multiple target domains are tested. The test samples are simultaneously input into the trained model M j and automatically outputs the fault diagnosis results on all domains.

[0020] Furthermore, in step S4, due to data privacy restrictions, the source domain samples are inaccessible during the fault classification process of the unlabeled target domain samples. The target domain samples can only be classified using the network parameters of the initial model trained and uploaded locally in the source domain. Therefore, a self-supervised learning method based on knowledge distillation is used to process the unlabeled target domain samples, specifically including the following:

[0021] S41, calculate the initial model M0 in the target domain sample Soft label on

[0022]

[0023] in, Represents source domain samples The high-dimensional features obtained after the feature extractor f0, Represents source domain samples Output probability after passing through feature extractor f0 and fault classifier g0.

[0024] S42. Calculate the knowledge distillation loss of the first-stage diagnostic model M1 by measuring the KL distance between the prediction result of the first-stage diagnostic model M1 and the soft label obtained by knowledge distillation:

[0025]

[0026] D KL (A||B) represents the KL distance between A and B, represents the knowledge distillation loss of the first stage diagnostic model M1, Indicates that the initial model M0 is in the target domain sample The soft label on Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1.

[0027] S43. Calculate the maximum information entropy loss of the first-stage diagnostic model M1:

[0028]

[0029] represents the maximum information entropy loss of the first stage diagnostic model M1, Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1.

[0030] S44. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

[0031] Furthermore, step S6 specifically includes the following contents:

[0032] S61, calculation model {M0,...,M j-1 In the target domain sample The probability distribution output on

[0033]

[0034] in, Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0035] S62. Use the confidence voting strategy to select models with high confidence probability output:

[0036]

[0037] Where β is the confidence threshold, which is used to evaluate the confidence of the model probability output. The selected models with high confidence are uploaded to the model library M.

[0038] S63, according to the output probability distribution of the models in the model library M, the average output probability distribution p of all models can be calculated. i,j and the number of M models in the model library s i,j According to the above confidence voting strategy, the consistent knowledge representation can be obtained

[0039]

[0040] S64, through the measurement model M j The KL distance between the prediction result and the consistent knowledge obtained by knowledge distillation is used to calculate the j-th stage diagnosis model M j The knowledge distillation loss of the model:

[0041]

[0042] in, Represents the j-th stage diagnosis model M j The knowledge distillation loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0043] S65. Calculate the maximum information entropy loss:

[0044]

[0045] in, Represents the j-th stage diagnosis model M j The maximum information entropy loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0046] S66. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0048] 1. Existing deep transfer learning methods typically use metric-based or adversarial strategies to match the source and target domains. This requires access to source domain data and does not meet the requirements for data privacy protection in practical industrial applications. The proposed method introduces information entropy maximization technology and utilizes self-supervised learning to enable the model to capture more key fault features from unlabeled data, thereby improving the accuracy of target domain fault identification.

[0049] 2. The present invention designs a knowledge distillation mechanism based on confidence voting. It obtains consistent diagnostic knowledge on the target domain data by simply evaluating the output confidence of models from different stages without accessing the source data, fully protecting the privacy of data from different clients; and realizes dynamic real-time update of the new model based on the consistent diagnostic knowledge, so that the model can be quickly generalized to new streaming data while maintaining the diagnostic accuracy on the old data.

[0050] 3. The present invention fully considers the real-time and diverse characteristics of big data streams in actual industrial applications. The constructed intelligent fault diagnosis model can realize online updates and accurate diagnosis, solves the problem of the inability to match data collection and streaming computing in real time, and greatly improves the diagnostic performance of the fault diagnosis model for real-time monitored industrial streaming data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of the online fault diagnosis method for industrial streaming data based on deep knowledge distillation network of the present invention.

[0052] Figure 2 yes Figure 1 Schematic diagram of the deep knowledge distillation network model feature extractor and fault classifier in the method.

[0053] Figure 3 3 is a comparison chart of the fault diagnosis results of the method of the present invention and other methods in the experimental example. DETAILED DESCRIPTION

[0054] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0055] The present invention provides an online fault diagnosis method for industrial streaming data based on a deep knowledge distillation network. The overall flow chart of the implementation of the method can be found in Figure 1 , specifically including the following steps:

[0056] S1. Data acquisition: Multiple sensors are used to continuously collect raw mechanical vibration signals from mechanical equipment under multiple working conditions. The samples collected under different working conditions are divided into a source domain and multiple target domains. Where j∈{1,2,…,N}, N represents the number of target domains, and the original vibration signal is intercepted by the sliding window sampling method to extract data of the same length to form training and test samples.

[0057] In this step, the source domain It consists of labeled samples, where Represents source domain samples and follows the distribution X s represents the data distribution of source domain samples, Represents the fault category label corresponding to the source domain sample and follows the distribution Y s represents the data distribution of source domain sample labels, n s The number of samples representing the source domain; the target domain It consists of unlabeled samples, where Represents the sample of the jth target domain and obeys the data distribution P t j represents the data distribution of the j-th target domain sample, represents the number of samples in the jth target domain. The data distribution between the source domain and multiple target domains is different, that is, X s ≠P t 1 ≠P t 2 ≠...≠P t N , P t N Represents the data distribution of the Nth target domain sample.

[0058] S2. Model Construction: The network structure is constructed based on the specific diagnostic task and dataset information. Based on the characteristics of the one-dimensional vibration signal, a feature extractor f0 based on a one-dimensional convolutional neural network is first designed for feature extraction. Then, a fault classifier g0 with a softmax function is constructed to classify the fault type.

[0059] Among them, the detailed structural parameters of the feature extractor f0 and fault classifier g0 based on one-dimensional convolutional neural network are as follows: Figure 2As shown in the figure, the input signal to the network is of length 1×2048. The feature extractor f0 includes four convolutional modules. The first three convolutional modules contain a convolutional layer, a maximum pooling layer, and a LeakyReLU activation function, and the last convolutional module contains a convolutional layer and a LeakyReLU activation function. The fault classifier g0 includes two fully connected modules, each of which contains a fully connected layer and a ReLU activation function. The convolutional layer extracts features of different scales and locations from the input data through convolution operations, achieving feature extraction and transformation of the data. The main function of the maximum pooling layer is to downsample the input feature map, reducing the size of the feature map and extracting the main feature information. The activation function enables the neural network to learn and represent nonlinear relationships by introducing nonlinear transformations, which can help gradients propagate better and enhance the network's expressive power. The main function of the fully connected layer is to fully connect all neurons in the previous layer with all neurons in the current layer, achieving linear transformation and feature combination from input features to output.

[0060] S3. Local model pre-training: First, use the Xavier initializer to initialize the weights and biases of the initial model M0. Then, input the labeled source domain samples into the model in step S2, and obtain the initial model M0 through supervised training.

[0061] In some embodiments of the present invention, the feature extractor f0 and the fault classifier g0 are pre-trained using labeled source domain samples. The standard cross entropy loss function is used as the loss function for model pre-training. The initial model is optimized by minimizing the fault classification error. The initial model pre-training loss function is as follows:

[0062]

[0063] Where C is the number of health status categories in the source domain, Represents source domain samples The high-dimensional features obtained after the feature extractor f0, Represents source domain samples Output probability after passing through feature extractor f0 and fault classifier g0.

[0064] S4, incremental learning stage 1: After the initial model M0 is trained, it is first uploaded to the first target domain At the same time, the Xavier initializer is used to initialize the weights and biases of the model for constructing the first stage diagnostic model M1. In the incremental learning stage, the information entropy maximization technique is first used to help the model adapt to the data distribution of the target domain. Information entropy is used in information theory to measure the uncertainty of a random variable. The larger the value of information entropy, the higher the uncertainty of the random variable. When the information entropy is maximized, it can provide the most information, and the features learned by the model are also the most comprehensive. Then, the knowledge distillation method is used to obtain the initial model M0 in the first target domain. Finally, the maximum information entropy loss and knowledge distillation loss are combined to optimize the first stage diagnosis model M1.

[0065] Due to data privacy restrictions in step S4, the source domain samples are inaccessible during the fault classification process of the unlabeled target domain samples. The only way to guide the fault classification of the target domain samples is to use the network parameters of the initial model trained and uploaded locally in the source domain. Therefore, a self-supervised learning method based on knowledge distillation is used to process the unlabeled target domain samples, specifically including the following:

[0066] S41, calculate the initial model M0 in the target domain sample Soft label on

[0067]

[0068] in, Represents source domain samples The high-dimensional features obtained after the feature extractor f0, Represents source domain samples Output probability after passing through feature extractor f0 and fault classifier g0.

[0069] S42. Calculate the knowledge distillation loss of the first-stage diagnostic model M1 by measuring the KL distance between the prediction result of the first-stage diagnostic model M1 and the soft label obtained by knowledge distillation:

[0070]

[0071] D KL (·||·) represents the KL distance between A and B, represents the knowledge distillation loss of the first stage diagnostic model M1, Indicates that the initial model M0 is in the target domain sample The soft label on Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1.

[0072] S43. Calculate the maximum information entropy loss of the first-stage diagnostic model M1:

[0073]

[0074] represents the maximum information entropy loss of the first stage diagnostic model M1, represents the number of samples in the first target domain, Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1.

[0075] S44. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

[0076] S5. Testing phase 1: After training is completed, the data from the source domain and the target domain are tested. The test samples are simultaneously input into the trained diagnosis model M1, and the fault diagnosis results in the two domains are automatically output.

[0077] S6, incremental learning stage j: the model {M0,...,M j-1}Upload to the jth target domain At the same time, use the Xavier initializer to initialize and construct the j-th stage diagnosis model M j In the incremental learning phase, we first use the information entropy maximization technique to help the model adapt to the data distribution of the target domain; then we use the knowledge distillation method based on probability confidence to obtain the model {M0,...,M j-1 In the target domain Consistent diagnostic knowledge on the data; finally, the maximum information entropy loss and knowledge distillation loss are combined to optimize the j-th stage diagnostic model M j .

[0078] Step S6 specifically includes the following contents:

[0079] S61, calculation model {M0,...,M j-1 In the target domain sample The probability distribution output on

[0080]

[0081] in, Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0082] S62. Use the confidence voting strategy to select models with high confidence probability output:

[0083]

[0084] Where β is the confidence threshold, which is used to evaluate the confidence of the model probability output. The selected models with high confidence are uploaded to the model library M.

[0085] S63, according to the output probability distribution of the models in the model library M, the average output probability distribution p of all models can be calculated. i,j and the number of M models in the model library s i,j According to the above confidence voting strategy, the consistent knowledge representation can be obtained

[0086]

[0087] p i,j Represents a sample The corresponding average output probability distribution, s i,j Represents a sample The corresponding number of M models in the obtained model library.

[0088] S64, through the measurement model M j The KL distance between the prediction result and the consistent knowledge obtained by knowledge distillation is used to calculate the j-th stage diagnosis model M j The knowledge distillation loss of the model:

[0089]

[0090] in, Represents the j-th stage diagnosis model M j The knowledge distillation loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0091] S65. Calculate the maximum information entropy loss:

[0092]

[0093] in, Represents the j-th stage diagnosis model M j The maximum information entropy loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after .

[0094] S66. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

[0095] S7, testing phase j: after training is completed, the data from the source domain and the target domain are The test samples are simultaneously input into the trained model M j and automatically outputs the fault diagnosis results on all domains, where S represents the source domain.

[0096] Experimental case:

[0097] Gearboxes are common, key components in mechanical equipment. Through the meshing and rotation of gears, different transmission ratios can be achieved, thereby regulating the speed and power of vehicles such as automobiles and aircraft. To verify the effectiveness of the proposed method, a continuous fault diagnosis and classification experiment was designed using a three-axis, five-speed transmission gearbox.

[0098] 1. Experimental Dataset

[0099] The test platform for a three-shaft, five-speed transmission gearbox consists of a drive motor, a load motor, a transmission support, an SG135-2 transmission, a universal drive shaft, and an accelerometer. The operating principle of the automotive five-speed transmission test bench is that power generated by the drive motor is input into the transmission and then output to the load motor via a universal drive shaft. By adjusting the output torques of the drive and load motors, the test bench simulates different operating conditions. Fault simulation experiments were conducted at three different speeds: 750 rpm, 1000 rpm, and 1250 rpm. The load torques were set to 0 Nm and 50 Nm, respectively. Since bearings and gears are key components of the gear transmission system, faults were implanted in the bearings of the gears and output shaft in the transmission's five-speed transmission path. The faulty bearings were NUP311EN. Using wire-cut machining, two different inner race faults with a depth of 1 mm and widths of 0.2 mm and 2 mm were machined into two different NUP311EN bearings. Three types of gear faults were simulated: mild tooth breakage, moderate tooth breakage, and complete tooth breakage. The PAK data acquisition system of German Müller-BBM company was used to collect the operating data of automobile transmission under different health states. The vibration acceleration sensor was fixed on the housing of the output shaft bearing seat of the automobile transmission, and the sampling frequency was set to 24kHz.

[0100] Six passive incremental transfer diagnosis experiments were designed on the aforementioned transmission gearbox dataset, as shown in Table 1, where "750-0" indicates an operating speed of 750 rpm and a load of 0 Nm. Each passive incremental transfer diagnosis experiment consists of multiple stages, as shown in Table 2. Taking Experiment 1 as an example, in the local model pre-training stage, the model is trained using labeled data from the source domain. In incremental learning stages 1 to 4, the model is trained using unlabeled data from the target domain {750-50, 1000-0, 1000-50, 1250-0, 1250-50}, respectively. Model performance is then tested on both the source domain and the previously trained target domain. As the number of stages increases, the number of target domains also increases, which means that the diagnostic difficulty continues to increase.

[0101] In this experiment, each sample length is 2048 data points, and each fault category is 2500 samples. There are 10 different fault states in the transmission gearbox data set, so there are 25,000 samples for each operating condition. The final test result uses the average of ten test results to avoid the influence of random errors.

[0102] Table 1 Passive incremental migration diagnosis experiment designed based on transmission gearbox dataset

[0103]

[0104] Table 2 Detailed description of Experiment 1

[0105]

[0106] 2. Algorithm parameter setting and evaluation indicators

[0107] The model parameter settings were determined in Experiment 1 using a grid search method. The maximum number of training times was set to 60, the learning rate for model training was 0.001, and the batch size was 64. The detailed structural parameters of the feature extractor f0 and the fault classifier g0 based on the one-dimensional convolutional neural network are shown in Figure 2. Figure 2 shown.

[0108] In order to fully evaluate the performance of the model on the incremental migration task, the harmonic mean is used Forward propagation and backpropagation Three evaluation indicators are used to show the experimental results, among which Acc t represents the average value of the test results in all target domains, Acc s Represents the test results of the source domain, R j,k Represents model M j The test results in the k-th target domain, R 0,k represents the test result of the pre-trained model M0 in the kth target domain, R k,k Represents model M k The test results in the kth target domain. The HM indicator represents the performance of the model M at the jth stage. j The global performance of the model M in the jth stage is represented by FT. j The degree of performance improvement after training on all target domains; BT represents the j-th stage model M j The degree of forgetting of learned diagnostic knowledge.

[0109] 3. Analysis of experimental results

[0110] To demonstrate the superiority of our proposed method, we compared it with classic algorithms in the field. These included SHOT, a typical passive domain adaptation algorithm that can perform fault diagnosis in the target domain using only the source domain model, and USF, a passive universal domain adaptation algorithm that can address fault diagnosis while protecting data privacy. To ensure fairness in the comparative experiments, the feature extractors of SHOT and USF were replaced with the same ones used in the proposed method. Each incremental transfer experiment was repeated ten times, and the average value was used as the final test result.

[0111] The experimental results are shown in Table 3. The proposed method showed the best classification performance in seven incremental migration fault diagnosis experiments, with an overall fault identification accuracy of up to 92.35%. The average degradation of the model fault diagnosis performance was only -5.77%, indicating that after the parameter update of multiple streaming data stages, the model can still maintain accurate diagnosis of old data; the average performance improvement of the model is as high as 33.63%, indicating that the diagnostic accuracy of the trained model under various working conditions has been greatly improved compared with the initial model. Comparison of fault diagnosis results of the proposed method with other methods Figure 3 As shown in the figure (each experimental section shows SHOT, USF, and the present invention method from left to right), the diagnostic accuracy of the present invention method is significantly improved compared with SHOT and USF. Obviously, the present invention method can effectively realize incremental migration fault diagnosis under industrial streaming data.

[0112] Table 3 Comparison of diagnostic results between the proposed method and the classical algorithm

[0113]

[0114]

[0115] This paper aims to solve the problems of data privacy protection and catastrophic forgetting of knowledge when performing fault diagnosis on streaming data from different working conditions. Taking the transmission gearbox as the research object, this paper designs a fault streaming diagnosis method based on a deep knowledge distillation network, which realizes the rapid generalization and fault diagnosis of the model in multiple target domains and has high industrial application value.

[0116] Finally, it should be noted that although the implementation of the present invention has been described in detail with reference to examples, it is easy for those skilled in the art to understand that any modifications, substitutions and improvements made without departing from the spirit and principles of the present invention as described in the appended claims should be included in the scope of protection of the present invention.

Claims

1. An online fault diagnosis method for industrial streaming data based on deep knowledge distillation network, characterized by: The following steps are involved: S1. Collect the original mechanical vibration signals of multiple working conditions. The samples collected under different working conditions are divided into a source domain D s and multiple target domains Where j∈{1,2,…,N}, N represents the number of target domains, and constructs training and test samples; S2. Based on the characteristics of the one-dimensional vibration signal, a feature extractor f0 based on a one-dimensional convolutional neural network is constructed to extract features, and a fault classifier g0 with a Softmax function is constructed to classify the fault type; S3. Initialize the weights and biases of the initial model M0, then input the labeled source domain samples into the model constructed in step S2, and obtain the initial model M0 through supervised training; S4, incremental learning stage 1: After the initial model M0 is trained, it is first uploaded to the first target domain At the same time, the weights and biases of the first stage diagnostic model M1 are initialized; in the incremental learning stage, the maximum information entropy technology is first used to assist the model to adapt to the data distribution of the target domain, and then the knowledge distillation method is used to obtain the initial model M0 in the first target domain. The soft labels on the data are finally optimized by combining the maximum information entropy loss and knowledge distillation loss to optimize the first stage diagnostic model M1; S5. Testing phase 1: After training is completed, the data from the source domain and the target domain are tested. The test samples are simultaneously input into the trained diagnosis model M1, and the fault diagnosis results in the two domains are automatically output; S6, incremental learning stage j: the model {M0,...,M j-1 }Upload to the jth target domain At the same time, initialize and build the j-th stage diagnosis model M j The weights and biases of the model; in the incremental learning stage, the information entropy maximization technique is first used to help the model adapt to the data distribution of the target domain; then the knowledge distillation method based on probability confidence is used to obtain the model {M0,...,M j-1 In the target domain Consistent diagnostic knowledge on the data; finally, the maximum information entropy loss and knowledge distillation loss are combined to optimize the j-th stage diagnostic model M j ; S7, testing phase j: after training is completed, the data from the source domain and multiple target domains are tested. The test samples are simultaneously input into the trained model M j And output the fault diagnosis results on all domains, S represents the source domain.

2. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S1, data of the same length are intercepted from the original vibration signal by a sliding window sampling method to construct training and test samples.

3. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S1, the source domain It consists of labeled samples, where Represents source domain samples and follows the distribution X s represents the data distribution of source domain samples, Represents the fault category label corresponding to the source domain sample and follows the distribution Y s represents the data distribution of source domain sample labels, n s The number of samples representing the source domain; the target domain It consists of unlabeled samples, where Represents the sample of the jth target domain and obeys the data distribution P t j represents the data distribution of the j-th target domain sample, represents the number of samples in the jth target domain. The data distribution between the source domain and multiple target domains is different, that is, X s ≠P t 1 ≠P t 2 ≠...≠P t N , P t N Represents the data distribution of the Nth target domain sample.

4. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S3, the Xavier initializer is used for initialization.

5. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: The feature extractor f0 is used to obtain high-dimensional features and includes multiple convolution modules connected in sequence. The last convolution module includes a convolution layer and an activation function layer. Other convolution modules include convolution layers, maximum pooling layers and activation function layers. The convolution layer is used to extract features of different scales and positions. The maximum pooling layer is used to downsample the input feature map. The activation function layer is used to learn and represent nonlinear relationships.

6. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: The fault classifier g0 includes a fully connected module, which includes a fully connected layer and an activation function layer. The fully connected layer is used to fully connect all neurons in the previous layer with all neurons in the current layer to achieve linear transformation and feature combination from input features to output.

7. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S3, the feature extractor f0 and the fault classifier g0 are pre-trained using labeled source domain samples. The standard cross entropy loss function is used as the loss function for model pre-training. The initial model is optimized by minimizing the fault classification error. The initial model pre-training loss function is as follows: Where C is the number of health status categories in the source domain, n s The number of samples representing the source domain, Represents source domain samples The high-dimensional features obtained after the feature extractor f0, Represents source domain samples Output probability after passing through feature extractor f0 and fault classifier g0.

8. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S4, the network parameters of the initial model obtained by local training and uploading in the source domain are used to guide the target domain samples for fault classification.

9. The method for online fault diagnosis of industrial streaming data based on deep knowledge distillation network according to claim 1 is characterized in that: In step S4, a self-supervised learning method based on knowledge distillation is used to process unlabeled target domain samples, including the following steps: S41, calculate the initial model M0 in the target domain sample Soft label on in, Represents source domain samples The high-dimensional features obtained after the feature extractor f0, Represents source domain samples Output probability after feature extractor f0 and fault classifier g0; S42. Calculate the knowledge distillation loss of the first-stage diagnostic model M1 by measuring the KL distance between the prediction result of the first-stage diagnostic model M1 and the soft label obtained by knowledge distillation: Where D KL (·||·) represents the KL distance between A and B, represents the knowledge distillation loss of the first stage diagnostic model M1, Indicates that the initial model M0 is in the target domain sample The soft label on Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1; S43. Calculate the maximum information entropy loss of the first-stage diagnostic model M1: Where, represents the maximum information entropy loss of the first stage diagnostic model M1, represents the number of samples in the first target domain, Represents source domain samples The high-dimensional features obtained after the feature extractor f1, Represents source domain samples Output probability after passing through feature extractor f1 and fault classifier g1; S44. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

10. The method for online fault diagnosis of industrial streaming data based on a deep knowledge distillation network according to any one of claims 1 to 9, characterized in that: Step S6 includes: S61, calculation model {M0,...,M j-1 In the target domain sample The probability distribution output on in, Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after ; S62. Use the confidence voting strategy to select models with high confidence probability output: Where β is the confidence threshold, which is used to evaluate the confidence of the model probability output. The selected models with high confidence are uploaded to the model library M. S63, calculate the average output probability distribution p of all models based on the output probability distribution of the models in the model library M. i,j and the number of M models in the model library s i,j , according to the above confidence voting strategy, we can obtain consistent knowledge representation S64, through the measurement model M j The KL distance between the prediction result and the consistent knowledge obtained by knowledge distillation is used to calculate the j-th stage diagnosis model M j The knowledge distillation loss of the model: in, Represents the j-th stage diagnosis model M j The knowledge distillation loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after ; S65. Calculate the maximum information entropy loss: in, Represents the j-th stage diagnosis model M j The maximum information entropy loss is Represents source domain samples After the feature extractor f j The high-dimensional features obtained later are Represents source domain samples After the feature extractor f j and fault classifier g j The output probability after ; S66. Calculate the final loss function of the training process And the gradient-compensated back-propagation algorithm is used to update all parameters of each layer in the model.

Citation Information

Patent Citations

  • Rolling bearing fault diagnosis method and system based on relational knowledge distillation

    CN113281048A

  • Rolling bearing fault diagnosis method for improving model migration strategy

    CN111721536A

  • Knowledge distillation-based equipment fault diagnosis method for underground coal mining machine in coal mine

    CN113283386A