Industrial equipment residual life prediction method based on classification and error correction knowledge distillation and medium
Through the distillation method based on classification and error correction knowledge, the remaining life prediction network of industrial equipment is trained, which solves the computing resource and memory limitation problems of traditional deep learning networks in edge device deployment, and achieves a better balance between accuracy and inference memory.
Patent Information
- Application Number
- CN202510342881.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional deep learning networks require a large amount of computing resources and memory in the remaining life prediction of industrial equipment, limiting their deployment on edge devices. How to achieve a better balance between accuracy and inference memory has become an urgent problem.
Using the method of distillation based on classification and error correction knowledge, we train the neural network model based on classification and error correction knowledge distillation, and combine the loss function to train the student network model to achieve the remaining life prediction of industrial equipment. The method includes obtaining parameter time series, inputting multiple teacher network models, performing linear fusion, dividing simple and difficult samples, and constructing a distillation loss function for self-reflection and self-correction.
Improves the prediction performance and adaptability of student network models, reduces inference memory and improves prediction accuracy, and achieves a better balance between accuracy and inference memory, making it easier to deploy on mobile devices.
Smart Images

Figure CN119939357A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial equipment life prediction, and in particular to a method and medium for predicting the remaining life of industrial equipment based on classification and error correction knowledge distillation. Background Art
[0002] Health management (PHM) is an important research field in the industrial field, emphasizing the real-time monitoring and analysis of equipment conditions to predict health status and potential failures. Remaining useful life (RUL) prediction is a core aspect of PHM and plays a vital role in improving equipment reliability, optimizing maintenance plans and reducing operating costs. At present, the methods used to predict RUL mainly include physical model-based methods, statistical model-based methods and data-driven methods. Physics-based models predict RUL based on the working principle and wear mechanism of the equipment through physical equations and models. These methods require a deep understanding of the physical characteristics and working environment of the equipment. Statistical-based models require the use of historical failure data and life distribution models to estimate the RUL of the equipment. These models usually require a large amount of historical data to produce reliable results. Data-driven models use machine learning and deep learning algorithms to predict RUL by analyzing sensor data (such as vibration, temperature and sound), which is a hot direction in recent years.
[0003] However, traditional deep learning networks usually require a lot of computing resources and memory, which limits their deployment on edge devices. Therefore, how to achieve a better balance between accuracy and inference memory for easier deployment on mobile devices has become an urgent problem to be solved. Summary of the invention
[0004] The technical problem to be solved by the present application is to provide a method and medium for predicting the remaining life of industrial equipment based on classification and error correction knowledge distillation, which has the characteristics of achieving a better balance between accuracy and inference memory for easier deployment on mobile devices.
[0005] In a first aspect, an embodiment provides a method for predicting the remaining life of industrial equipment based on classification and error correction knowledge distillation, including: combining a loss function, training an industrial equipment remaining life prediction network based on classification and error correction knowledge distillation, and predicting the remaining life of industrial equipment based on the industrial equipment remaining life prediction network; the industrial equipment remaining life prediction network is a neural network model obtained by training based on classification and error correction knowledge distillation, and the training method includes: Obtaining a parameter time series for predicting the remaining life of industrial equipment and labeling the true value of the remaining life to obtain a labeled parameter time series, and preprocessing the labeled parameter time series to obtain a training sample; Input the training samples to multiple teacher network models respectively to obtain the corresponding output prediction values; the multiple teacher network models are neural network models with complementary advantages; Linearly fuse the corresponding output prediction values to obtain the teacher output prediction value; Based on a preset similarity threshold, the teacher output prediction value at each divided time point is classified into a simple sample and a difficult sample; the simple sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value reaches the preset similarity threshold, and the difficult sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value is less than the preset similarity threshold; According to the time points, the simple samples and difficult samples are mapped to the teacher output prediction value time series, the true value time series and the student network model output prediction value time series, so as to obtain the corresponding time series marked with difficult samples and simple samples; Based on the difficult samples of the teacher's output prediction value time series, the true value time series and the student network model's output prediction value time series, a first distillation loss function is constructed; Based on a simple sample of the teacher output prediction value time series and the student network model output prediction value time series, a second distillation loss function is constructed; Obtaining a third distillation loss function based on the first distillation loss function and the second distillation loss function; Constructing a total loss function based on the third distillation loss function and the loss function between the predicted value and the true value of the student network model; The student network model is trained in combination with the total loss function, and the trained student network model is used as the remaining life prediction network for industrial equipment.
[0006] In one embodiment, the method of classifying the teacher output prediction value at each divided time point into a simple sample and a difficult sample based on a preset similarity threshold includes: The teacher output prediction value time series and the true value time series are constructed into a A two-dimensional time series, where n is the length of the time series; The Transpose the two-dimensional time series to generate a The feature matrix of Setting a local sliding window, and sliding the local sliding window according to the time sequence and the preset step size; For each sliding window, the correlation between the two column vectors in the sliding window is calculated, including: generating a Gaussian kernel matrix based on the two column vectors, and decentralizing the Gaussian kernel matrix to obtain a Gaussian kernel correlation matrix, obtaining the characteristic entropy corresponding to each local sliding window based on the Gaussian kernel correlation matrix, and based on a preset entropy threshold, defining the data corresponding to the sliding window whose characteristic entropy is less than the entropy threshold as a simple sample, and defining the data corresponding to the sliding window whose characteristic entropy is greater than or equal to the entropy threshold as a difficult sample.
[0007] In one embodiment, the generating of a Gaussian kernel matrix based on two column vectors includes: Where K represents the Gaussian kernel matrix, k represents the Gaussian kernel function, x1, x2, and x3 represent the time series values of the first column vector of the two column vectors, y1, y2, and y3 represent the time series values of the second column vector of the two column vectors, p represents the index of the time series value in the first column vector, and q represents the index of the time series value in the second column vector. represents the Euclidean distance calculation, represents the kernel width of the Gaussian kernel function, and exp represents the natural exponential function; The Gaussian kernel matrix is decentralized to obtain a Gaussian kernel correlation matrix, including: in, represents the decentralized Gaussian kernel function, represents a square matrix with the same dimension as the Gaussian kernel matrix K and all the elements are the same, m represents the dimension of the square matrix; C represents the Gaussian kernel correlation matrix, and T represents transpose; The characteristic entropy corresponding to each local sliding window obtained based on the Gaussian kernel correlation matrix includes: Among them, FE represents characteristic entropy, r represents any eigenvalue index of the three eigenvalues decomposed by the singular value of the Gaussian kernel correlation matrix C, represents the eigenvalues of the singular value decomposition of the Gaussian kernel correlation matrix C, and log represents the logarithmic function with base 2.
[0008] In a second aspect, an embodiment provides a computer-readable storage medium, wherein a program is stored in the medium, and the program can be loaded by a processor to execute the method for predicting the remaining life of industrial equipment described in any one of the above embodiments.
[0009] The beneficial effects of the present invention are: Since the outputs of multiple teacher network models with complementary advantages are linearly fused to obtain the teacher output prediction value, the student network model can better learn the teacher's characteristics and learning ability, thereby improving the prediction performance of the student network model; since the difficult samples and simple samples of each time series are obtained based on the similarity between the teacher output prediction value and the corresponding true value, the distillation loss function can be constructed based on the obtained difficult samples and simple samples, thereby obtaining the total loss function, so that the student network model can perform knowledge self-reflection during the training process in combination with the loss function instead of directly replacing the true value, emphasizing self-correction, and enhancing the adaptability of the student model to real data, so that the prediction accuracy can be improved while reducing the inference memory, thereby achieving a better balance between accuracy and inference memory, so as to facilitate easier deployment on mobile devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a flow chart of a training method of an industrial equipment remaining life prediction network according to an embodiment of the present application; Figure 2 This application Figure 1 A method flow diagram of an embodiment of step S04; Figure 3 This application Figure 2 A method flow chart of an embodiment of step S0404 in FIG. DETAILED DESCRIPTION
[0011] The present invention is further described in detail below by specific embodiments in conjunction with the accompanying drawings. Wherein similar elements in different embodiments adopt associated similar element numbers. In the following embodiments, many detailed descriptions are for making the present application better understood. However, those skilled in the art can easily recognize that some features can be omitted in different situations, or can be replaced by other elements, materials, methods. In some cases, some operations related to the present application are not shown or described in the specification, this is to avoid the core part of the present application being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and the general technical knowledge in the art.
[0012] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various implementations. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a required sequence, unless otherwise specified that a certain sequence must be followed.
[0013] The serial numbers assigned to the components in this article, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning.
[0014] To facilitate the description of the inventive concept of this application, the knowledge distillation technology is briefly described below.
[0015] In addition, the stringent requirements of edge devices for data protection and operational efficiency further highlight the need for lightweight models. Recent advances in lightweight models focus on model compression techniques, including knowledge distillation (KD), quantization, pruning, and neural architecture search. Among these methods, KD has attracted much attention due to its flexibility and efficiency. As an effective model compression technique, knowledge distillation extracts knowledge from a trained large teacher network model and transfers it to a small, lightweight student network model, thereby greatly reducing the number of model parameters and computational complexity without significantly losing predictive performance. Knowledge distillation was originally applied in fields such as computer vision and natural language processing, but it is still in its infancy in the application of industrial equipment RUL. Therefore, how to effectively and reliably deploy industrial equipment on the edge has become a major difficulty at present.
[0016] In view of this, an embodiment of the present application provides an industrial equipment remaining life prediction method and medium based on classification and error correction knowledge distillation. In the training method of the industrial equipment remaining life prediction network for remaining life prediction, the outputs of multiple teacher network models with complementary advantages are linearly fused to obtain the teacher output prediction value, so that the student network model can better learn the characteristics and learning ability of the teacher, thereby improving the prediction performance of the student network model; since the difficult samples and simple samples of each time series are obtained based on the similarity between the teacher output prediction value and the corresponding true value, a distillation loss function can be constructed based on the obtained difficult samples and simple samples to obtain a total loss function, so that the student network model can perform knowledge self-reflection in the training process in combination with the loss function instead of directly replacing the true value, emphasizing self-correction, and enhancing the adaptability of the student model to real data, so that the accuracy of the prediction can be improved while reducing the inference memory, thereby achieving a better balance between accuracy and inference memory, so as to facilitate easier deployment on mobile devices.
[0017] In one embodiment of the present application, a method for predicting the remaining life of industrial equipment based on classification and error correction knowledge distillation is provided, comprising: combining a loss function, training an industrial equipment remaining life prediction network based on classification and error correction knowledge distillation, and predicting the remaining life of industrial equipment based on the industrial equipment remaining life prediction network. The industrial equipment remaining life prediction network is a neural network model obtained by training based on classification and error correction knowledge distillation. Please refer to Figure 1 , the training method of the industrial equipment remaining life prediction network includes: Step S01, obtaining a parameter time series for predicting the remaining life of industrial equipment and annotating the true value of the remaining life to obtain a labeled parameter time series, and preprocessing the labeled parameter time series to obtain a training sample.
[0018] In one embodiment, the acquired parameter time series may include vibration, temperature and sound time series of industrial equipment collected by sensors. The parameter time is labeled to obtain parameter time series training samples with labels of true values of the remaining life.
[0019] Those skilled in the art will appreciate that, in order to obtain the required training samples, the preprocessing method may adopt the method of the prior art, which will not be described in detail here.
[0020] Step S02: input the training samples to multiple teacher network models respectively to obtain their corresponding output prediction values.
[0021] Among them, multiple teacher network models are neural network models with complementary advantages.
[0022] In one embodiment, the plurality of teacher network models include a first teacher network model and a second teacher network model, that is, the training samples are respectively input to the first teacher network model and the second teacher network model to obtain the first output prediction value and the second output prediction value corresponding to each other. The first teacher network model and the second teacher network model are neural network models with complementary advantages. For example, the advantage of one teacher network model is fast inference speed, and the advantage of another teacher network model is high accuracy. Thus, the two teacher network models have complementary advantages.
[0023] Step S03, linearly fuse the corresponding output prediction values to obtain the teacher output prediction value.
[0024] In one embodiment, step S03 includes: in, represents the teacher output prediction value, represents the first output prediction value, represents the second output prediction value, Represents the preset hyperparameter for adjusting the contribution of the first teacher network model and the second teacher network model, 0< <1.
[0025] Through the complementary advantages of the teacher network model, the student network model can better learn the characteristics and learning ability of the teacher, thereby improving the prediction performance of the student network model.
[0026] In one embodiment, to more fully reflect the performance of the teacher network model, you can select =0.5.
[0027] Step S04, based on a preset similarity threshold, classify the teacher output prediction value at each divided time point into simple samples and difficult samples.
[0028] Among them, the simple sample indicates that the similarity between the teacher's output prediction value and the corresponding true value reaches the preset similarity threshold, and the difficult sample indicates that the similarity between the teacher's output prediction value and the corresponding true value is less than the preset similarity threshold.
[0029] In one embodiment, please refer to Figure 2 , step S04 may include: Step S0401: construct the teacher output prediction value time series and the true value time series into a Two-dimensional time series.
[0030] Where n is the length of the time series.
[0031] Step S0402: Transpose the two-dimensional time series to generate a The feature matrix of .
[0032] Step S0403: set a local sliding window, and slide the local sliding window according to the time sequence and the preset step size.
[0033] In one embodiment, a The sliding window has a step size of 3. Therefore, the sliding window will slide in chronological order. Each sliding window contains two columns of data, one column contains 3 teacher output prediction value data, and one column contains 3 real value data, a total of 6 data. The specific sliding window W can be expressed as: Among them, X represents the teacher output prediction value column vector, Y represents the true value column vector, x1, x2 and x3 represent the time series values of the X column vector, and y1, y2 and y3 represent the time series values of the Y column vector.
[0034] Since the data inside the window are all continuous time points, the causal relationship of the time series can be guaranteed.
[0035] Step S0404: for each sliding window, calculate the correlation between two column vectors in the sliding window.
[0036] Based on the sliding window given above, the correlation between two column vectors in the sliding window can be calculated. Figure 3 , step S0404 may include: Step S100, generating a Gaussian kernel matrix based on two column vectors.
[0037] In order to analyze the correlation between two column vectors, in one implementation, we apply a Gaussian kernel function to calculate the correlation between each element of the two column vectors of the matrix. Combined with step S0403, step S100 may include: Where K represents the Gaussian kernel matrix, k represents the Gaussian kernel function, x1, x2, and x3 represent the time series values of the first column vector of the two column vectors, y1, y2, and y3 represent the time series values of the second column vector of the two column vectors, p represents the index of the time series value in the first column vector, and q represents the index of the time series value in the second column vector. represents the Euclidean distance calculation, represents the kernel width of the Gaussian kernel function, and exp represents the natural exponential function.
[0038] is a hyperparameter, and in one embodiment, it can be preset to the average Euclidean distance of the data in each corresponding window.
[0039] Combined with step S0403, x1, x2 and x3 here represent the time series values of the teacher's output prediction value column vector, and y1, y2 and y3 represent the time series values of the true value column vector.
[0040] Based on the above process, the mapping calculation of the Gaussian kernel matrix can be performed on each local window.
[0041] Step S200, decentralizing the Gaussian kernel matrix to obtain a Gaussian kernel correlation matrix.
[0042] By decentralizing each Gaussian kernel correlation matrix, the global bias can be removed to prevent the influence of individual points with large deviations on the overall situation. In one embodiment, step S200 may include: in, represents the decentralized Gaussian kernel function, represents a square matrix with the same dimension as the Gaussian kernel matrix K and all the elements are the same, m represents the dimension of the square matrix; C represents the Gaussian kernel correlation matrix, and T represents the transpose.
[0043] The Gaussian kernel correlation matrix C is obtained through matrix operations, and the correlation can be further extracted.
[0044] Step S300, obtaining the characteristic entropy corresponding to each local sliding window based on the Gaussian kernel correlation matrix.
[0045] In one embodiment, the singular values of the generated Gaussian kernel correlation matrix C are decomposed into three eigenvalues, and then a characteristic entropy value is generated using the characteristic entropy formula, and finally m / 3 characteristic entropy values are generated. Then step S300 may include: Among them, FE represents characteristic entropy, r represents any eigenvalue index of the three eigenvalues decomposed by the singular value of the Gaussian kernel correlation matrix C, represents the eigenvalues of the singular value decomposition of the Gaussian kernel correlation matrix C, and log represents the logarithmic function with base 2.
[0046] Step S400, based on a preset entropy threshold, defines the data corresponding to the sliding window whose characteristic entropy is less than the entropy threshold as a simple sample, and defines the data corresponding to the sliding window whose characteristic entropy is greater than or equal to the entropy threshold as a difficult sample.
[0047] We arrange the generated feature entropy values in the order of previous time series, and divide simple samples and difficult samples based on the preset entropy threshold.
[0048] In one embodiment, we can preset the entropy threshold based on the distribution of feature entropy and the performance of the required student network model on a specific data set. In one embodiment, the preset entropy threshold makes the simple sample larger than the difficult sample. In one embodiment, the entropy threshold can be preset to 0.3. Since the larger the feature, the smaller the similarity, and the smaller the feature entropy, the greater the similarity, based on the preset feature entropy, the data corresponding to the sliding window whose feature entropy is less than the entropy threshold is defined as a simple sample, and the data corresponding to the sliding window whose feature entropy is greater than or equal to the entropy threshold is defined as a difficult sample. In this way, the simple sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value reaches the preset similarity threshold, and the difficult sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value is less than the preset similarity threshold.
[0049] Since the predicted value and the true value in the local window are continuous. Therefore, through the above method, we can not only focus on the relationship between the predicted value and the true value at the same time point, but also focus on the relationship between the predicted value and the true value at adjacent time points. Mathematically, the Gaussian kernel function is used to measure the correlation and construct the Gaussian kernel matrix. Then the correlation between the predicted value and the true value is converted into an entropy-based expression, and the information classification is obtained through the dispersion and correlation matrix. In this way, the information classification helps to prevent incorrect teacher knowledge from being transferred to the student network model and improve the performance of the student network model.
[0050] Step S05, according to the time point, the simple samples and difficult samples are mapped to the teacher output prediction value time series, the true value time series and the student network model output prediction value time series, so as to obtain the corresponding time series marked with difficult samples and simple samples.
[0051] According to the classified simple sample and difficult sample time points, they are mapped to the teacher output prediction value time series, the true value time series and the student network model output prediction value time series in order, and a simple sample of the teacher output prediction value time series can be obtained. and difficult samples , a simple sample of the true valued time series and difficult samples , a simple sample of the time series of predicted values output by the student network model and difficult samples . Where i represents the index of the difficult sample and j represents the index of the easy sample.
[0052] The applicant found in the research that in the current method, the predictions of the teacher network model at the time points corresponding to difficult samples are not accurate enough to be transmitted to students. Therefore, the difficult samples are directly replaced with the true values. However, this method of improving the performance of the student network model will hinder the ability to learn to reflect on the incorrect predictions of the teacher network model.
[0053] In view of this, in one embodiment of the present application, the following new method is used to construct a loss function to overcome the defects caused by directly replacing difficult samples with true values.
[0054] Since the samples are divided into difficult samples and simple samples, the distillation loss we get is divided into two parts, one is the first distillation loss for difficult samples, and the other is the second distillation loss for simple samples.
[0055] Step S06, constructing a first distillation loss function based on the difficult samples of the teacher's output prediction value time series, the true value time series and the student network model's output prediction value time series.
[0056] Set the Hub, Positive and Negative objects, then you can Set as Hub, Set to Negative, Set to Positive. Through dynamic learning, the hub will gradually adjust its position in the horizontal or vertical direction to reduce the Manhattan distance with the Positive and increase the Manhattan distance with the Negative. Step S06 can be expressed as: in, represents the first distillation loss function, i represents the index of the difficult sample, I represents the number of difficult samples, 1≤i≤I, represents the i-th difficult sample in the output prediction value of the student network model, represents the i-th difficult sample in the true value, represents the i-th difficult sample in the teacher's output prediction value, represents the margin enforced between positive and negative training sample pairs, represents the Manhattan distance.
[0057] In this way, through the strategy of knowledge self-reflection of Manhattan distance, through the knowledge self-reflection in the student network model, instead of directly replacing the true value, emphasizing self-correction, the adaptability of the student network model to real data is enhanced, and the prediction accuracy is improved.
[0058] Among them, The constraint ensures that the Hub sample and the Positive sample are closer.
[0059] Step S07, constructing a second distillation loss function based on a simple sample of the teacher output prediction value time series and the student network model output prediction value time series.
[0060] In one embodiment, step S07 includes: in, represents the distillation loss function, j represents the index of a simple sample, g represents the number of simple samples, 1≤j≤g, represents the jth simple sample in the output prediction value of the student network model, represents the jth simple sample in the teacher's output prediction value.
[0061] Step S08: obtaining a third distillation loss function based on the first distillation loss function and the second distillation loss function.
[0062] In one embodiment, step S08 includes: in, represents the third distillation loss function.
[0063] Step S09, obtaining a total loss function based on the third distillation loss function and the loss function between the predicted value and the true value of the student network model.
[0064] In one embodiment, step S09 includes: in, represents the total loss function, Represents loss weight, 0< <1, Represents the loss function between the predicted value and the true value of the student network model.
[0065] In one embodiment, The value of is 0.8.
[0066] Those skilled in the art can understand that the loss function between the predicted value and the true value of the student network model can be obtained using existing technical methods, which will not be repeated here.
[0067] Step S10, training the student network model in combination with the total loss function, and using the trained student network model as the remaining life prediction network for industrial equipment.
[0068] Based on the trained student network model, the data to be predicted after preprocessing can be input into the student network model, and the final output is the RUL prediction value.
[0069] In the method for predicting the remaining life of industrial equipment based on any of the above-mentioned embodiments, since the difficult samples and simple samples of each time series are obtained by dividing the similarity between the teacher's output prediction value and the corresponding true value, a distillation loss function can be constructed based on the obtained difficult samples and simple samples, thereby obtaining a total loss function, so that the student network model can perform knowledge self-reflection during the training process in combination with the loss function instead of directly replacing the true value, emphasizing self-correction, and enhancing the adaptability of the student model to real data, so that the prediction accuracy can be improved while reducing the inference memory, thereby achieving a better balance between accuracy and inference memory, so as to facilitate easier deployment on mobile devices.
[0070] In one embodiment of the present application, a computer-readable storage medium is provided, on which a program is stored. The stored program includes a method that can be loaded by a processor and process any of the above embodiments.
[0071] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above-mentioned embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above-mentioned functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above-mentioned functions can be implemented. In addition, when all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and can be downloaded or copied and saved in the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.
[0072] The above specific examples are used to illustrate the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art, according to the concept of the present invention, some simple deductions, modifications or substitutions can be made.
Claims
1. A method for predicting the remaining life of industrial equipment based on classification and error correction knowledge distillation, characterized in that: include: Combined with the loss function, train the industrial equipment remaining life prediction network based on classification and error correction knowledge distillation, and perform the remaining life prediction of industrial equipment based on the industrial equipment remaining life prediction network; The industrial equipment remaining life prediction network is a neural network model obtained by knowledge distillation training based on classification and error correction, and the training method includes: Obtaining a parameter time series for predicting the remaining life of industrial equipment and labeling the true value of the remaining life to obtain a labeled parameter time series, and preprocessing the labeled parameter time series to obtain a training sample; Input the training samples to multiple teacher network models respectively to obtain the corresponding output prediction values; the multiple teacher network models are neural network models with complementary advantages; Linearly fuse the corresponding output prediction values to obtain the teacher output prediction value; Based on a preset similarity threshold, the teacher output prediction value at each divided time point is classified into a simple sample and a difficult sample; the simple sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value reaches the preset similarity threshold, and the difficult sample indicates that the similarity between the obtained teacher output prediction value and the corresponding true value is less than the preset similarity threshold; According to the time points, the simple samples and difficult samples are mapped to the teacher output prediction value time series, the true value time series and the student network model output prediction value time series, so as to obtain the corresponding time series marked with difficult samples and simple samples; Based on the difficult samples of the teacher's output prediction value time series, the true value time series and the student network model's output prediction value time series, a first distillation loss function is constructed; Based on a simple sample of the teacher output prediction value time series and the student network model output prediction value time series, a second distillation loss function is constructed; Obtaining a third distillation loss function based on the first distillation loss function and the second distillation loss function; Constructing a total loss function based on the third distillation loss function and the loss function between the predicted value and the true value of the student network model; The student network model is trained in combination with the total loss function, and the trained student network model is used as the remaining life prediction network for industrial equipment.
2. The method for predicting the remaining life of industrial equipment according to claim 1, characterized in that: The step of inputting the training samples to multiple teacher network models to obtain the corresponding output prediction values includes: Input the training samples to the first teacher network model and the second teacher network model respectively, and obtain the first output prediction value and the second output prediction value corresponding to each of them; The linear fusion of the respective corresponding output prediction values to obtain the teacher output prediction value includes: linear fusion of the first output prediction value and the second output prediction value to obtain the teacher output prediction value.
3. The method for predicting the remaining life of industrial equipment according to claim 2, characterized in that: The linear fusion of the first output prediction value and the second output prediction value to obtain the teacher output prediction value includes: in, represents the teacher output prediction value, represents the first output prediction value, represents the second output prediction value, Represents the preset hyperparameters for adjusting the contribution of the first teacher network model and the second teacher network model.
4. The method for predicting the remaining life of industrial equipment according to claim 1, characterized in that: The method of classifying the teacher output prediction value at each divided time point into simple samples and difficult samples based on a preset similarity threshold includes: The teacher output prediction value time series and the true value time series are constructed into a A two-dimensional time series, where n is the length of the time series; The Transpose the two-dimensional time series to generate a The feature matrix of Setting a local sliding window, and sliding the local sliding window according to the time sequence and the preset step size; For each sliding window, the correlation between the two column vectors in the sliding window is calculated, including: generating a Gaussian kernel matrix based on the two column vectors, and decentralizing the Gaussian kernel matrix to obtain a Gaussian kernel correlation matrix, obtaining the characteristic entropy corresponding to each local sliding window based on the Gaussian kernel correlation matrix, and based on a preset entropy threshold, defining the data corresponding to the sliding window whose characteristic entropy is less than the entropy threshold as a simple sample, and defining the data corresponding to the sliding window whose characteristic entropy is greater than or equal to the entropy threshold as a difficult sample.
5. The method for predicting the remaining life of industrial equipment according to claim 4, characterized in that: The generating of the Gaussian kernel matrix based on two column vectors includes: Where K represents the Gaussian kernel matrix, k represents the Gaussian kernel function, x1, x2, and x3 represent the time series values of the first column vector of the two column vectors, y1, y2, and y3 represent the time series values of the second column vector of the two column vectors, p represents the index of the time series value in the first column vector, and q represents the index of the time series value in the second column vector. represents the Euclidean distance calculation, represents the kernel width of the Gaussian kernel function, and exp represents the natural exponential function; The Gaussian kernel matrix is decentralized to obtain a Gaussian kernel correlation matrix, including: in, represents the decentralized Gaussian kernel function, represents a square matrix with the same dimension as the Gaussian kernel matrix K and all the elements are the same, m represents the dimension of the square matrix; C represents the Gaussian kernel correlation matrix, and T represents transpose; The characteristic entropy corresponding to each local sliding window obtained based on the Gaussian kernel correlation matrix includes: Among them, FE represents characteristic entropy, r represents any eigenvalue index of the three eigenvalues decomposed by the singular value of the Gaussian kernel correlation matrix C, represents the eigenvalues of the singular value decomposition of the Gaussian kernel correlation matrix C, and log represents the logarithmic function with base 2.
6. The method for predicting the remaining life of industrial equipment according to claim 1, characterized in that: The difficult samples based on the teacher output prediction value time series, the true value time series and the student network model output prediction value time series, construct the first distillation loss function, including: in, represents the first distillation loss function, i represents the index of the difficult sample, I represents the number of difficult samples, 1≤i≤I, represents the i-th difficult sample in the output prediction value of the student network model, represents the i-th difficult sample in the true value, represents the i-th difficult sample in the teacher's output prediction value, represents the margin enforced between positive and negative training sample pairs, represents the Manhattan distance.
7. The method for predicting the remaining life of industrial equipment according to claim 6, characterized in that: The second distillation loss function is constructed based on a simple sample of the teacher output prediction value time series and the student network model output prediction value time series, including: in, represents the second distillation loss function, j represents the index of a simple sample, g represents the number of simple samples, 1≤j≤g, represents the jth simple sample in the output prediction value of the student network model, represents the jth simple sample in the teacher's output prediction value.
8. The method for predicting the remaining life of industrial equipment according to claim 7, characterized in that: The obtaining of the third distillation loss function based on the first distillation loss function and the second distillation loss function includes: in, represents the third distillation loss function.
9. The method for predicting the remaining life of industrial equipment according to claim 8, characterized in that: The total loss function obtained based on the third distillation loss function and the loss function between the predicted value and the true value of the student network model includes: in, represents the total loss function, Represents the loss weight, 0< <1, Represents the loss function between the predicted value and the true value of the student network model.
10. A computer-readable storage medium, characterized in that: The medium stores a program, which can be loaded by a processor and execute the method for predicting the remaining life of industrial equipment as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Industrial process optimization decision model migration optimization method based on knowledge distillation
CN112529188A
Data-driven electric vehicle battery remaining life prediction method
CN117805658A
Self-distillation method and system based on mixed sample, electronic equipment and medium
CN118051848A
Fish identification method, system and equipment based on knowledge distillation and medium
CN118212457A
Method, device, and recording medium for image encoding / decoding
WO2024043760A1