Optimization Method for Device Status Prediction Model, Prediction Method and System for Device Status

By updating feature weights in the device state prediction model and compressing the model using the knowledge distillation algorithm, the problem that the device state prediction model is difficult to adapt is solved, and high-precision and time-efficient device state prediction is achieved.

CN120068312BActive Publication Date: 2025-07-25ZHONGKE YUNGU TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510528261.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-25
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing equipment state prediction model has complex structure and fixed, and it is difficult to adaptively adjust with changes in the working conditions and environment, resulting in low timeliness and accuracy of prediction results.

Method used

By receiving the working condition data at the edge end and the equipment state prediction results, the importance score of each working condition is calculated, the feature weight of the equipment state prediction model is updated, the feature matrix is constructed, and the optimized model is compressed into a simple distillation model using the knowledge distillation algorithm and deployed to the edge end.

Benefits of technology

The adaptability optimization of the equipment state prediction model is realized, ensuring that the edge-end model always adapts to the latest working conditions, and improving prediction accuracy and timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068312B_ABST
    Figure CN120068312B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of construction machinery, and provide an optimization method for an equipment state prediction model, a prediction method for equipment state, and a system. Applied to the server side, the optimization method includes: receiving the working condition data of the equipment, the equipment state prediction result, and the importance score of each working condition sent by the edge side, where the equipment state prediction result is obtained by the edge side based on the working condition data through the first distillation model, and the importance score of each working condition is obtained by the edge side based on the working condition data and the equipment state prediction result; updating the feature weight corresponding to each working condition in the equipment state prediction model based on the importance score of each working condition; constructing a feature matrix based on the equipment state prediction result, the historical equipment state prediction result, and the working condition data; optimizing the model parameters of the equipment state prediction model based on the updated feature weight and the feature matrix; compressing the optimized equipment state prediction model into a second distillation model; and deploying the second distillation model to the edge side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of construction machinery, and particularly relates to an optimization method for an equipment status prediction model, a prediction method for equipment status, and a system therefor. Background Art

[0002] Engineering machinery equipment is an important tool for modern social infrastructure construction and industrial production, and its reliability and safety are crucial. Equipment failures not only cause production stagnation but may also result in casualties and property losses. Therefore, effective predictive maintenance of engineering machinery equipment, early identification of potential failures, and taking preventive measures are of great significance for ensuring the safe operation of equipment, improving production efficiency, and reducing maintenance costs.

[0003] In the prior art, when predicting the operating state of equipment, usually, the operating state data or equipment working condition data of the equipment are first collected through sensors or equipment terminals, etc., and data analysis and data preprocessing are performed, such as data cleaning, noise reduction, and feature extraction, to ensure the quality and effectiveness of the data. Secondly, the preprocessed data is input into a trained prediction model, such as a machine learning or deep learning model, to predict the operating state of the equipment in a certain period in the future through the prediction model. Finally, according to the prediction results, corresponding maintenance strategies are formulated, such as arranging maintenance or replacing components in advance, to improve the reliability and availability of the equipment. However, the prediction models required by such predictive maintenance means usually need a large amount of historical data for training. For new equipment or equipment with a large change in the operating environment, there is a lack of effective historical data, making it difficult for the prediction model to achieve iterative optimization and resulting in low prediction accuracy. At the same time, such prediction models often have a fixed structure and cannot meet the adaptive adjustment of the model with changes in the working condition environment. When the working condition environment changes, the deep learning model needs to consume a large amount of computing resources for training, making it difficult to be deployed at the edge end and the timeliness of the prediction is difficult to guarantee. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide an optimization method for an equipment status prediction model, a prediction method for equipment status, and a system therefor, so as to solve the technical defect that the structure of the status prediction model in the prior art is complex and fixed, and it is difficult to make adaptive adjustments with changes in the working condition environment, resulting in low timeliness and accuracy of the equipment status prediction results.

[0005] To achieve the above purpose, in the first aspect of this application, an optimization method for an equipment status prediction model is provided, which is applied to the server side. The optimization method includes:

[0006] Receive the operating condition data of the device, the device state prediction result, and the importance score of each operating condition sent by the edge device. The device state prediction result is obtained by the edge device through the first distillation model based on the operating condition data, and the importance score of each operating condition is obtained by the edge device based on the operating condition data and the device state prediction result acquired from the device side;

[0007] Update the feature weights corresponding to each operating condition in the device state prediction model based on the importance score of each operating condition;

[0008] Construct a feature matrix based on the device state prediction result, the historical device state prediction result, and the operating condition data;

[0009] Optimize the model parameters of the device state prediction model based on the updated feature weights and the feature matrix to obtain an optimized device state prediction model;

[0010] Compress the optimized device state prediction model into a second distillation model;

[0011] Deploy the second distillation model to the edge device.

[0012] In the embodiment of the present application, updating the feature weights corresponding to each operating condition in the device state prediction model based on the importance score of each operating condition includes updating the feature weights corresponding to each operating condition in the device state prediction model according to formula (1):

[0013] (1),

[0014] where, is the feature weight corresponding to the th operating condition, is the importance score corresponding to the th operating condition, is the sum of the importance scores corresponding to each operating condition, and N is the number of types of operating conditions included in the operating condition data.

[0015] In the embodiment of the present application, compressing the optimized device state prediction model into a second distillation model includes: using the optimized device state prediction model as the teacher model; using the operating condition data and the historical operating condition data as sample data; training the student model based on the sample data and the teacher model to minimize the total loss function to obtain the second distillation model.

[0016] In the embodiment of the present application, the total loss function is obtained based on the focal loss function and the distance loss function.

[0017] In the embodiment of the present application, the total loss function is determined according to formula (2):

[0018] (2),

[0019] Among them, is the total loss, is the focal loss function, is the distance loss function, is the weighted hyperparameter of different optimization terms, is the probability distribution output by the student model, is the hard label, is the soft label.

[0020] In the embodiments of the present application, the device state prediction model is an LSTM network trained based on historical working condition data. The device state prediction model includes an input layer, an LSTM layer, an attention module, and an output layer. The attention module is used to perform weighted calculation on each hidden state sequence output by the LSTM layer to determine the importance degree of each hidden state sequence.

[0021] The second aspect of the present application provides a prediction method for device state, which is applied to the edge side. The prediction method includes:

[0022] Obtain the working condition data of the device;

[0023] Input the working condition data into the first distillation model to obtain the device state prediction result of the device;

[0024] Determine the importance score of each working condition based on the working condition data and the device state prediction result;

[0025] Send the working condition data, the device state prediction result, and the importance score to the server side, where the server side uses the optimization method for the device state prediction model in any one of the above to generate a second distillation model according to the received working condition data, the device state prediction result, and the importance score;

[0026] Receive the second distillation model and use the second distillation model to replace the first distillation model.

[0027] In the embodiments of the present application, determining the importance score of each working condition based on the working condition data and the device state prediction result includes: calculating the importance score of each working condition based on the working condition data and the device state prediction result through the LIME algorithm.

[0028] The third aspect of the present application provides a server configured to execute the above optimization method for the device state prediction model.

[0029] The fourth aspect of the present application provides an edge node configured to execute the above prediction method for device state.

[0030] The fifth aspect of the present application provides a prediction system for device state, including:

[0031] The above-mentioned server;

[0032] An edge side, configured to communicate with the server, the edge side includes at least one of the above-mentioned edge nodes.

[0033] In an embodiment of the present application, a prediction system for device status further includes:

[0034] At least one device side, configured to communicate with at least one edge node.

[0035] A sixth aspect of the present application provides a machine-readable storage medium, and when the instructions are executed by a processor, the processor is configured to perform the above-mentioned optimization method for a device status prediction model or the above-mentioned prediction method for device status.

[0036] A seventh aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned optimization method for a device status prediction model or the above-mentioned prediction method for device status.

[0037] The above technical solution, by receiving the working condition data of the device, the device status prediction result and the importance score of each working condition sent by the edge side, where the device status prediction result is obtained by the edge side based on the working condition data through the first distillation model, and the importance score of each working condition is obtained by the edge side based on the working condition data and the device status prediction result obtained from the device side; updating the feature weight corresponding to each working condition in the device status prediction model based on the importance score of each working condition; constructing a feature matrix based on the device status prediction result, the historical device status prediction result and the working condition data; optimizing the model parameters of the device status prediction model based on the updated feature weight and the feature matrix to obtain an optimized device status prediction model; compressing the optimized device status prediction model into a second distillation model; deploying the second distillation model to the edge side. This method compresses the device status prediction model into a distillation model with fewer parameters and a simple structure, enabling it to make adaptive optimizations according to changes in the working condition environment, and can ensure that the distillation model at the edge side always adapts to the latest working condition environment, so as to obtain a device status prediction result with higher accuracy in subsequent prediction processes.

[0038] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the embodiments of the present application together with the following specific implementation manners, but do not constitute a limitation to the embodiments of the present application. In the drawings:

[0040] Figure 1Schematically shows a flowchart of an optimization method for a device status prediction model according to an embodiment of the present application;

[0041] Figure 2 Schematically shows a structural diagram of an LSTM network according to an embodiment of the present application;

[0042] Figure 3 Schematically shows a flowchart of a prediction method for device status according to an embodiment of the present application;

[0043] Figure 4 Schematically shows a system architecture diagram of a prediction system for device status according to an embodiment of the present application;

[0044] Figure 5 Schematically shows the internal structure diagram of a computer device according to an embodiment of the present application. Detailed implementation manners

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0046] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, such descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0047] Figure 1 Schematically shows a flowchart of an optimization method for a device status prediction model according to an embodiment of the present application. As Figure 1 shown, the embodiments of the present application provide an optimization method for a device status prediction model. The method is applied to the server side, and the optimization method may include the following steps:

[0048] Step 101: Receive the working condition data of the device, the device status prediction result, and the importance score of each working condition sent by the edge device. The device status prediction result is obtained by the edge device through the first distillation model based on the working condition data, and the importance score of each working condition is obtained by the edge device based on the working condition data and the device status prediction result acquired from the device side.

[0049] In the embodiment of the present application, it should be noted that the working condition data of the device may include features of multiple dimensions, such as various working conditions like fuel consumption, engine speed, motor pressure, and fuel temperature. The device status prediction result may include, but is not limited to, normal and boom fault, etc. The importance score of each working condition may refer to the influence degree of each working condition in the working condition data on the device status prediction result. After the edge device obtains the working condition data of the device, it can input the working condition data into the first distillation model to obtain the device status prediction result of the device. At the same time, the importance score of each working condition in the working condition data is calculated through the working condition data and the device status prediction result. Secondly, the edge device sends the acquired working condition data, the obtained device status prediction result, and the importance score of each working condition to the server side, and the server side then receives the working condition data of the device, the device status prediction result, and the importance score of each working condition sent by the edge device.

[0050] Step 102: Update the feature weight corresponding to each working condition in the device status prediction model based on the importance score of each working condition.

[0051] In the embodiment of the present application, it should be noted that the device status prediction model may refer to an LSTM network. Among them, the LSTM network refers to a long short-term memory network, which is a special recurrent neural network (RNN) designed to solve the problems of gradient disappearance and explosion in traditional RNNs when processing long sequence data. LSTM can effectively capture long-term dependencies in time series by introducing memory units and gating mechanisms. After the server side receives the working condition data of the device, the device status prediction result, and the importance score of each working condition sent by the edge device, it can further update the feature weight corresponding to each working condition in the device status prediction model according to the importance score of each working condition. For example, taking the case of including N working conditions, the set composed of the importance scores of these N working conditions can be , according to the feature weight corresponding to each working condition in the device status prediction model can be updated.

[0052] In the embodiment of the present application, the device status prediction model is an LSTM network trained based on historical working condition data. The device status prediction model includes an input layer, an LSTM layer, an attention module, and an output layer. The attention module is used to perform weighted calculation on each hidden state sequence output by the LSTM layer to determine the importance degree of each hidden state sequence.

[0053] In this embodiment, it should be noted that the equipment status prediction model may refer to an LSTM network trained based on historical operating condition data. In this technical solution, since the historical operating condition data may include features in multiple dimensions, such as fuel consumption, engine speed, motor pressure, fuel temperature, and other operating conditions. In the scenario of predictive maintenance of engineering machinery equipment, the historical operating condition data of the equipment is usually collected over a long period, which may cover data for several months or years and belongs to data with a long time span. In view of the characteristics of this scenario data, an attention mechanism can be introduced into the standard LSTM network in this technical solution. The attention mechanism is a deep learning technology, and its core idea is to enable the model to dynamically focus on different parts of the input sequence when processing the input sequence, thereby improving the performance of the model. In this way, the equipment status prediction model can capture the long-term dependencies in the time series effectively and at the same time pay attention to the important data information in the multi-dimensional features to improve the prediction performance of the equipment status prediction model.

[0054] As Figure 2 shown, a structural schematic diagram of an LSTM network is provided. As Figure 2 shown, the LSTM network may include an input layer, an LSTM layer, an attention module (Attention layer), and an output layer. The attention module is used to perform weighted calculations on each hidden state sequence output by the LSTM layer to determine the importance of each hidden state sequence. Specifically, first, the input data passes through the LSTM layer. The LSTM layer will learn the long-term dependencies in the time series data and output the hidden state sequence. At the same time, Dropout can be combined to randomly discard neurons to solve the problem of overfitting in training. After that, in the Attention layer, based on the current LSTM hidden state and the input sequence, the attention weights for each time step are calculated, and the hidden states for each time step are weighted and summed to generate a weighted representation. Finally, it is sent to the fully connected layer for the prediction of the equipment status.

[0055] In this technical solution, by adding the corresponding equipment status under different historical operating condition conditions, the construction of the training set and the validation set can be completed. An example of the equipment status label is shown in Table 1.

[0056] Table 1 Equipment Status Label Description

[0057]

[0058] It should be noted that since the prediction of the device state belongs to a multi-classification task, the F1-score can be selected as the evaluation metric for the device state prediction model. The F1-score is an evaluation metric that comprehensively considers precision and recall. It is usually used to measure the performance of a classifier in a multi-classification task, and its value ranges from 0 to 1. The closer the value is to 1, the better the performance of the classifier. The calculation formula is as follows:

[0059] ,

[0060] Among them, Precision is the proportion of samples predicted as positive examples by the model that are actually positive examples, that is, the prediction accuracy of the model. Recall represents the proportion of samples that are actually positive examples and are correctly predicted as positive examples by the model, that is, how many positive examples the model has discovered.

[0061] In the embodiment of the present application, updating the feature weights corresponding to each working condition in the device state prediction model based on the importance scores of each working condition includes updating the feature weights corresponding to each working condition in the device state prediction model according to formula (1):

[0062] (1),

[0063] Among them, is the feature weight corresponding to the th working condition, is the importance score corresponding to the th working condition, is the sum of the importance scores corresponding to each working condition, and N is the number of types of working conditions included in the working condition data.

[0064] In this embodiment, it should be noted that after the server receives the importance scores of each working condition, for each working condition, the feature weights corresponding to each working condition can be calculated according to the above formula (1).

[0065] Step 103, construct a feature matrix based on the device state prediction result, the historical device state prediction result, and the working condition data.

[0066] In the embodiment of the present application, it should be noted that after the server receives the working condition data, the device state prediction result, and the importance score of each working condition sent by the edge side, it can further construct a feature matrix according to the device state prediction result, the historical device state prediction result, and the working condition data. Specifically, taking the working condition data as , the device state prediction result as , and the historical device state prediction result as as an example, the feature matrix constructed according to the device state prediction result, the historical device state prediction result, and the working condition data can be , where k represents the number of historical moments used.

[0067] Step 104: Optimize the model parameters of the device state prediction model based on the updated feature weights and feature matrix to obtain an optimized device state prediction model.

[0068] In the embodiment of this application, it should be noted that when the server updates the feature weights corresponding to each working condition in the device state prediction model according to the importance score of each working condition , and constructs a feature matrix based on the device state prediction result, the historical device state prediction result, and the working condition data , it can further optimize the model parameters of the device state prediction model with the updated feature weights and feature matrix to obtain an optimized device state prediction model. Specifically, it can use and the corresponding working condition data to retrain the device state prediction model after updating the feature weights to optimize the device state prediction model parameters :

[0069] ,

[0070] where is the learning rate, is the true label at the moment.

[0071] Step 105: Compress the optimized device state prediction model into a second distillation model.

[0072] In the embodiment of this application, it should be noted that after the server obtains the optimized device state prediction model, due to the characteristics of the model itself, the interpretability of its prediction results is poor, and the model complexity is high. Therefore, in order to meet the requirements of real-time prediction and formulate a reasonable and effective maintenance strategy, a knowledge distillation algorithm can be introduced to further optimize the optimized device state prediction model. The knowledge distillation algorithm is a model compression technology that can transfer the knowledge of a large and complex teacher model to a small and simple student model, so that the small model has performance equivalent to that of the large model. Its core idea is to use the output probability distribution of the teacher model, that is, the hard label distribution , to guide the learning process of the student model, and finally enable the student model to obtain a performance level similar to that of the teacher model under the premise of a smaller scale. The main core optimization items in this process include a data fitting item and a knowledge distillation item, which are used to ensure that the student model can be in the soft label On the basis of good performance above, correctly obtain the knowledge of the teacher model. Therefore, this technical solution can further compress the optimized device status prediction model into a second distilled model with a simple structure and less computational complexity through the knowledge distillation algorithm, so as to be deployed to the edge side to realize real-time prediction of the device status and formulation of the optimal maintenance strategy.

[0073] In the embodiment of the present application, compressing the optimized device status prediction model into a second distilled model includes: using the optimized device status prediction model as the teacher model; using the working condition data and historical working condition data as sample data; training the student model based on the sample data and the teacher model to minimize the total loss function, so as to obtain the second distilled model.

[0074] In this embodiment, it should be noted that this technical solution can further compress the optimized device status prediction model into a second distilled model with a simple structure and less computational complexity through the knowledge distillation algorithm. Specifically, the optimized device status prediction model can be used as the teacher model, and secondly, the working condition data and historical working condition data can be used as sample data, and the student model is trained based on the sample data and the teacher model to minimize the total loss function, so as to obtain the compressed second distilled model.

[0075] In the embodiment of the present application, the total loss function is obtained based on the focal loss function and the distance loss function.

[0076] In this embodiment, it should be noted that in the knowledge distillation algorithm, usually the prediction performance of the student model on the hard labels uses the cross-entropy function, and the knowledge distillation loss uses the KL divergence. However, the output of the teacher model is usually of high confidence, while the output of the student model may have a large error. Therefore, using the cross-entropy loss will cause the student model to overly focus on the easy-to-classify samples of the teacher model and ignore the difficult-to-classify samples of the teacher model, thus affecting the effect of knowledge distillation. For example, when there is a class imbalance problem in the dataset, the cross-entropy loss will tend to bias the prediction result of the model towards the majority class, resulting in a decrease in the model's recognition ability for the minority class. At the same time, since the KL divergence is asymmetric and sensitive to the difference between discrete and continuous distributions, this will cause the difference between the output distribution of the student model and the output distribution of the teacher model to be amplified in knowledge distillation, thus affecting the effect of knowledge distillation.

[0077] Therefore, to address the class imbalance problem and the KL divergence asymmetry problem existing in traditional knowledge distillation algorithms, this technical solution uses the Focal Loss function and the Wasserstein distance loss function as the total loss function of the knowledge distillation algorithm to calculate the overall loss. Specifically, the Focal Loss function can effectively solve the class imbalance problem. It downweights the loss of easy-to-classify samples, thereby focusing the model's attention on difficult-to-classify samples. Its formula is as follows:

[0078] ,

[0079] where, is the predicted probability of the correct class by the student model, is a hyperparameter used to control the degree of downweighting of easy-to-classify samples.

[0080] The distance loss function is a metric used to measure the difference between two probability distributions and can better handle the difference between discrete and continuous distributions. Its formula is as follows:

[0081] ,

[0082] where, is the set of all coupling distributions from to , is the probability density of the coupling distribution at , is and The distance between them, inf represents the global optimal solution obtained among all γ satisfying the constraints to ensure that the distance metric reflects the minimum cost.

[0083] In the embodiments of this application, the total loss function is determined according to formula (2):

[0084] (2),

[0085] where, is the total loss, is the Focal Loss function, is the distance loss function, is the weighted hyperparameter of different optimization terms, is the probability distribution output by the student model, is the hard label, is the soft label.

[0086] In this embodiment, the total loss function constructed by the Focal Loss function and the distance loss function in this technical solution is as shown in the above formula (2).

[0087] Step 106: Deploy the second distilled model to the edge side.

[0088] In the embodiment of the present application, it should be noted that after the server side further compresses the optimized device state prediction model into a second distilled model with a simple structure and less computational complexity through the knowledge distillation algorithm, the server side can deploy the second distilled model to the edge side. Since the obtained second distilled model has fewer model parameters and a simple structure, the second distilled model can be quickly deployed to the edge side for real-time prediction of the device state. At the same time, its prediction results and key feature sequences can also be uploaded to the server side again to complete the iterative optimization of the prediction model on the server side and further improve the prediction accuracy.

[0089] In the above technical solution, by compressing the device state prediction model into a distilled model with fewer parameters and a simple structure, it can make adaptive optimizations according to the changes in the working conditions, ensuring that the distilled model at the edge side always adapts to the latest working conditions, so as to obtain higher-precision device state prediction results in subsequent prediction processes.

[0090] In the embodiment of the present application, as Figure 3 shown, the embodiment of the present application provides a prediction method for device state. This method is applied to the edge side and may include the following steps:

[0091] Step 301: Obtain the working condition data of the device;

[0092] Step 302: Input the working condition data into the first distilled model to obtain the device state prediction result of the device;

[0093] Step 303: Determine the importance score of each working condition based on the working condition data and the device state prediction result;

[0094] Step 304: Send the working condition data, the device state prediction result, and the importance score to the server side, where the server side uses the optimization method for the device state prediction model in the above embodiment to generate a second distilled model according to the received working condition data, the device state prediction result, and the importance score;

[0095] Step 305: Receive the second distilled model and use the second distilled model to replace the first distilled model.

[0096] In this embodiment, it should be noted that the operating condition data of the device can include features in multiple dimensions, such as various operating conditions like fuel consumption, engine speed, motor pressure, and fuel temperature. After the edge side obtains the operating condition data of the device, it can input the operating condition data into the first distilled model to obtain the device state prediction result of the device. Among them, the device state prediction result can include, but is not limited to, normal and boom failure, etc. The first distilled model is a model with a simple structure and less computational complexity obtained by compressing the device state prediction model optimized by the server side within the historical time period.

[0097] It should be noted that the importance score of each operating condition can refer to the degree of influence of each operating condition in the operating condition data on the device state prediction result. After the edge side obtains the device state prediction result of the device through the first distilled model, it can further calculate the importance score of each operating condition in the operating condition data based on the operating condition data and the device state prediction result. Secondly, the edge side sends the obtained operating condition data, the device state prediction result, and the importance score of each operating condition to the server side. In this way, after the server side receives the operating condition data, the device state prediction result, and the importance score of each operating condition sent by the edge side, it can first update the feature weight corresponding to each operating condition in the device state prediction model according to the importance score of each operating condition. Secondly, construct a feature matrix based on the device state prediction result, the historical device state prediction result, and the operating condition data, and optimize the model parameters of the device state prediction model according to the updated feature weight and the feature matrix to obtain an optimized device state prediction model. Finally, the optimized device state prediction model can be compressed into a second distilled model through the knowledge distillation algorithm, and the second distilled model can be sent to the edge side. After the edge side receives the second distilled model, it uses the second distilled model to replace the first distilled model. In this way, it can ensure that the distilled model at the edge side always adapts to the latest operating condition scenario, so as to obtain a better device state prediction result in the subsequent prediction process.

[0098] In the embodiment of the present application, determining the importance score of each operating condition based on the operating condition data and the device state prediction result includes: calculating the importance score of each operating condition based on the operating condition data and the device state prediction result through the LIME algorithm.

[0099] In this embodiment, it should be noted that the LIME algorithm refers to a local interpretability machine learning algorithm. The main idea of the LIME algorithm is to construct a linear model around the input data to approximate the decision-making process of the deep learning model. This algorithm uses local weighted regression to determine the influence of each feature and generates a set of local explanations. In this technical solution, since the interpretability of deep learning models such as the LSTM network is poor, the LIME algorithm can be selected to calculate the importance score of each operating condition based on the operating condition data and the device state prediction result.

[0100] Based on the above server - side model optimization strategy, for the newly generated working condition data each time, through the processes of server - side model retraining, distilled model redeployment, reporting of equipment status prediction results and key feature sequences, and server - side model optimization, it can be ensured that the distilled model at the edge side always adapts to the current working condition scenario and obtains the optimal equipment status prediction result.

[0101] In the embodiment of the present application, as Figure 4 shown, the embodiment of the present application provides a system architecture diagram for a prediction system of equipment status. As Figure 4 shown, the system may include a server side, an edge side and a device side. The edge side may be composed of one or more edge nodes. Each edge node can communicate with one or more device sides, and the edge side communicates with the server side. The edge side obtains the working condition data of the device, inputs the working condition data into the first distilled model to obtain the equipment status prediction result of the device, and then calculates the importance score of each working condition in the working condition data based on the working condition data and the equipment status prediction result through the LIME algorithm. At this time, the edge side will report data to the server side, and send the working condition data, the equipment status prediction result and the importance score of each working condition to the server side.

[0102] After receiving the working condition data, the equipment status prediction result and the importance score of each working condition sent by the edge side, the server side updates the feature weights corresponding to each working condition in the equipment status prediction model according to the importance score of each working condition. Among them, the equipment status prediction model is an LSTM network introducing an attention module, which may include an input layer, an LSTM layer, an attention module and an output layer. The server side can construct a feature matrix based on the equipment status prediction result, the historical equipment status prediction result and the working condition data, and optimize the model parameters of the equipment status prediction model based on the updated feature weights and the feature matrix to obtain an optimized equipment status prediction model. Finally, the optimized equipment status prediction model is compressed into a second distilled model through the knowledge distillation algorithm, and the second distilled model is sent to the edge side.

[0103] The edge side replaces the first distilled model with the received second distilled model, so that it can be ensured that the distilled model at the edge side always adapts to the latest working condition scenario, so as to obtain a better equipment status prediction result in the subsequent prediction process.

[0104] Based on knowledge graphs and knowledge distillation, this technical solution proposes an innovative method for the predictive maintenance of engineering machinery and equipment. The prediction model that combines LSTM and Attention can not only capture long-term dependencies but also enable the model to focus on key information, thereby improving the accuracy and generalization ability of the model. The proposed improved knowledge distillation algorithm fully considers the problem of class imbalance and the difficulty of accurately identifying the difference in output distribution between the student model and the teacher model, and is more suitable for the scenario of predictive maintenance of engineering machinery and equipment with multiple device status categories. Under the cloud-edge collaborative architecture, through continuous iterative optimization, real-time adaptive predictive maintenance of the equipment can be achieved.

[0105] In the embodiments of this application, it should be noted that knowledge distillation, as a model compression technology, is usually used to transfer the knowledge of a large model (teacher model) to a small model (student model). In this way, the student model can maintain relatively low computational resource requirements while retaining the performance of the teacher model as much as possible. Embodied intelligence refers to the ability of an intelligent agent to learn and complete tasks through interaction with the physical environment, emphasizing the perception and action of the body in the environment. In an embodied intelligence system, it is usually necessary to process complex sensor data (such as vision, touch, sound, etc.) and make decisions in real-time or in resource-constrained environments. At this time, knowledge distillation can be used to transfer the knowledge of a complex model to a more lightweight model, optimizing the learning efficiency and deployment ability of embodied intelligence in the real environment for efficient operation on robots or embedded devices.

[0106] Knowledge distillation provides an efficient path for embodied intelligence from simulation to reality and from complex models to lightweight deployment. By designing an appropriate model architecture, multimodal knowledge transfer strategy, and Sim2Real optimization method, the practicality and robustness of embodied intelligent agents in real scenarios can be significantly improved. Specifically, taking a service robot as an example, applying knowledge distillation can transfer the complex grasping decision trained in simulation to a real robot; taking autonomous driving as an example, applying knowledge distillation can distill a large perception model such as 3D object detection to an in-vehicle embedded system; taking unmanned aerial vehicle navigation as an example, applying knowledge distillation can lightweight the student model to avoid obstacles in real-time in a dynamic environment.

[0107] The embodiments of this application provide a server configured to execute the above-mentioned optimization method for the device status prediction model.

[0108] The embodiments of this application provide an edge node configured to execute the above-mentioned prediction method for the device status.

[0109] The embodiments of this application provide a prediction system for the device status, including:

[0110] The above-mentioned server;

[0111] An edge side, configured to communicate with the server, and the edge side includes at least one of the above-mentioned edge nodes.

[0112] In an embodiment of the present application, a prediction system for device status further includes:

[0113] At least one device side, configured to communicate with at least one of the above-mentioned edge nodes.

[0114] The embodiment of the present application provides a machine-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor is configured to execute the above-mentioned optimization method for the device status prediction model or the above-mentioned prediction method for the device status.

[0115] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as Figure 5 shown. The computer device includes a processor A01, a network interface A02, a memory (not shown in the figure), and a database (not shown in the figure) connected through a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown in the figure). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A04. The database of the computer device is used to store data for the prediction method of device status. The network interface A02 of the computer device is used to communicate with an external terminal through a network connection. When the computer program B02 is executed by the processor A01, it realizes a prediction method for device status.

[0116] Those skilled in the art can understand that Figure 5 the structure shown in

[0117] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0118] The present application also provides a computer program product, which is suitable for executing a program for initializing the steps of the prediction method for device status when executed on a data processing device.

[0119] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0120] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0123] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0124] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0125] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0126] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0127] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An optimization method for a device status prediction model, characterized in that, Applied to the server side, the optimization method includes: Receiving the working condition data of the device, the device state prediction result and the importance score of each working condition sent by the edge side, where the device state prediction result is obtained by the edge side through the first distillation model based on the working condition data, and the importance score of each working condition is obtained by the edge side based on the working condition data and the device state prediction result acquired from the device side; Updating the feature weights corresponding to each working condition in the device state prediction model based on the importance score of each working condition. The device state prediction model is an LSTM network trained based on historical working condition data. The device state prediction model includes an input layer, an LSTM layer, an attention module and an output layer. The attention module is used to perform weighted calculation on each hidden state sequence output by the LSTM layer to determine the importance degree of each hidden state sequence; Constructing a feature matrix based on the device state prediction result, the historical device state prediction result and the working condition data; Retraining the device state prediction model with updated feature weights using the feature matrix and the corresponding working condition data to optimize the device state prediction model parameters and obtain an optimized device state prediction model; Compressing the optimized device state prediction model into a second distillation model; Deploying the second distillation model to the edge side.

2. The optimization method for the device status prediction model according to claim 1, wherein The updating the feature weights corresponding to each working condition in the device state prediction model based on the importance score of each working condition includes updating the feature weights corresponding to each working condition in the device state prediction model according to formula (1): Among them, W i is the characteristic weight corresponding to the i-th working condition, I i is the importance score corresponding to the i-th working condition, is the sum of the importance scores corresponding to each working condition, and N is the number of types of working conditions included in the working condition data.

3. The optimization method for the device status prediction model according to claim 1, characterized in that The compressing the optimized device state prediction model into a second distillation model includes: Regarding the optimized device state prediction model as the teacher model; Regarding the working condition data and the historical working condition data as sample data; Training the student model based on the sample data and the teacher model to minimize the total loss function to obtain the second distillation model.

4. The optimization method for the device status prediction model according to claim 3, characterized in that The total loss function is obtained based on the focal loss function and the distance loss function.

5. The optimization method for the device state prediction model according to claim 3, wherein The total loss function is determined according to formula (2): Among them, L is the total loss, L F is the focal loss function, L W is the distance loss function, λ is the weighted hyperparameter of different optimization terms, is the probability distribution output by the student model, C i (j) is the hard label, is the soft label.

6. A prediction method for device status, characterized in that, Applied to the edge side, the prediction method includes: Obtaining the working condition data of the device; Inputting the working condition data into the first distillation model to obtain the device state prediction result of the device; Determining the importance score of each working condition based on the working condition data and the device state prediction result; Sending the working condition data, the device state prediction result and the importance score to the server side, where the server side uses the optimization method for the device state prediction model according to any one of claims 1 to 5 to generate a second distillation model based on the received working condition data, the device state prediction result and the importance score; Receiving the second distillation model and using the second distillation model to replace the first distillation model.

7. The prediction method according to claim 6, wherein The determining the importance score of each working condition based on the working condition data and the device state prediction result includes: Calculating the importance score of each working condition based on the working condition data and the device state prediction result through the LIME algorithm.

8. A server, characterized in that, configured to execute the optimization method for the device state prediction model according to any one of claims 1 to 5.

9. An edge node, characterized in that, configured to execute the prediction method for the device state according to claim 6 or 7.

10. A prediction system for device status, characterized in that, comprising: a server according to claim 8; an edge side, configured to communicate with the server, the edge side comprising at least one edge node according to claim 9.

11. The prediction system for device status according to claim 10, characterized in that, further comprising: at least one device side, configured to communicate with the at least one edge node.

12. A machine-readable storage medium having instructions stored thereon, characterized in that, When executed by a processor, the instruction causes the processor to be configured to execute the optimization method for the device state prediction model according to any one of claims 1 to 5 or the prediction method for the device state according to any one of claims 6 to 7.

13. A computer program product, characterized in that, comprising a computer program which, when executed by a processor, implements the optimization method for the device state prediction model according to any one of claims 1 to 5 or the prediction method for the device state according to any one of claims 6 to 7.

Citation Information

Patent Citations

  • Data classification method and device based on neural network feature selection enhancement

    CN117763393A

  • Wind turbine generator operation performance detection method and system based on power generation working condition

    CN119712453A