Deep learning complex unmanned platform life prediction method based on edge unloading
By adopting a deep learning method based on edge offload on complex unmanned platforms, combining cloud and edge models, high-precision and high-real-time life prediction are achieved, solving the problem that the accuracy and real-time life prediction of complex unmanned platforms in the existing technology is difficult to take into account.
Patent Information
- Application Number
- CN202510217091.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to achieve high-precision and high-real-time life prediction on complex unmanned platforms, requiring both processing complex data to obtain high precision and fast response to achieve real-time.
Using a deep learning method based on edge offload, combining cloud-based high-precision deep learning model and edge-end lightweight deep learning model, the high-precision and high-real-time life prediction of complex unmanned platforms are achieved through the collaborative work of edge-end feature extraction, cloud prediction and dynamic weight generation models.
While ensuring prediction accuracy, it significantly improves the real-time prediction. It is suitable for tasks with high requirements for equipment life prediction accuracy and real-time prediction in industrial scenarios, reducing the system's calculation and transmission pressure.
Smart Images

Figure CN120067721A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of predicting the remaining service life of equipment, and relates to a method for predicting the life of a complex unmanned platform based on deep learning with edge offloading. Background Art
[0002] With the rise of Industry 4.0 and intelligent manufacturing, the prediction of the remaining service life of equipment has shifted from traditional regular equipment inspection and maintenance methods to predictive maintenance methods based on machine learning and deep learning. Algorithms such as support vector machines, random forests, and neural networks have been widely applied to the prediction of the remaining service life of equipment.
[0003] Complex unmanned platforms are increasingly widely used in fields such as industrial inspection, logistics distribution, and disaster relief. To ensure the high performance and reliability of the platform during task execution, it is crucial to perform high-precision real-time life prediction on it. Current complex deep learning algorithms, such as deep convolutional neural networks and Transformers, can extract accurate temporal features and the coupling relationships between complex equipment subsystems from large amounts of data through complex network structures, thereby achieving accurate life prediction, but they require a large amount of computing resources and time. While simple-structured deep learning algorithms, such as lightweight AlexNet and lightweight long short-term memory networks, can achieve real-time response to the life prediction of unmanned platforms, their ability to process complex data is weak, and the accuracy of life prediction for complex systems is low.
[0004] Therefore, there is an urgent need for a life prediction method for complex unmanned platforms that can achieve high-precision and high-real-time life prediction for complex unmanned platforms. Summary of the Invention
[0005] In view of this, the present invention provides a method for predicting the life of a complex unmanned platform based on deep learning with edge offloading, which can combine the prediction results of a complex high-precision deep learning model in the cloud and a lightweight deep learning model at the edge, greatly improving the real-time performance of the prediction while ensuring the prediction accuracy, and is particularly suitable for tasks with high requirements for the accuracy and real-time performance of equipment life prediction in industrial scenarios.
[0006] To solve the above technical problems, the present invention is implemented as follows.
[0007] A method for predicting the life of a complex unmanned platform based on deep learning with edge offloading includes:
[0008] Deploying a cloud prediction model to a cloud server; deploying an edge feature extraction model, an edge prediction model, and a dynamic weight generation model to edge devices; the cloud prediction model uses a high-precision model, and the edge feature extraction model and the edge prediction model use lightweight models;
[0009] The edge feature extraction model performs dimensionality reduction on multi-sensor time series data, extracts feature Z, and sends it to the cloud prediction model, the edge prediction model, and the dynamic weight generation model;
[0010] The cloud prediction model outputs the cloud life prediction value according to the input feature Z and transmits it back to the edge device;
[0011] The edge prediction model outputs the edge life prediction value according to the input feature Z
[0012] The dynamic weight generation model calculates the edge-cloud prediction fusion weight w according to the input feature Z and the distance from the current clustering center t ;
[0013] The edge device fuses the edge and cloud life prediction results according to the edge-cloud prediction fusion weight w t and and outputs the final life prediction value
[0014] Among them, when the dynamic weight generation model is trained, it performs clustering and iterative update according to the historical data of feature Z to obtain the historical clustering center; during online prediction, when the data distribution change of the real-time multi-sensor time series data corresponding to feature Z exceeds the set expectation, the clustering center is updated to dynamically adjust the edge-cloud prediction fusion weight.
[0015] Preferably, the edge feature extraction model is composed of a lightweight Transformer encoder, and performs dimensionality reduction on multi-sensor time series data in the time and sensor dimensions through sparse attention and multi-head attention mechanisms, and extracts the key features Z of time series and sensors.
[0016] Preferably, the pre-training method of the edge feature extraction model is: optimizing the edge feature extraction model by using the continuous contrast learning loss method; optimizing the model by comparing the difference in life labels between samples and the Euclidean distance of the dimensionality-reduced feature Z.
[0017] Preferably, the cloud prediction model adopts a Transformer encoder-decoder structure;
[0018] The pre-training method of the cloud prediction model is: connecting the edge feature extraction model with frozen parameters and the cloud prediction model in series, inputting the multi-sensor time series data sample into the edge feature extraction model to generate feature Z, inputting it into the cloud prediction model, and the cloud prediction model outputs the cloud life prediction value
[0019] Based on the cloud life prediction value Calculate the loss with the life label y and optimize the cloud prediction model.
[0020] Preferably, the method further includes: when the cloud prediction model makes an online prediction, adaptively adjusting the parameters of the cloud prediction model:
[0021] Store the predicted life prediction values of T clouds The life prediction value of the edge device And the final prediction value According to the life prediction value of the edge device And the life prediction value of the cloud Judge whether the current data distribution has changed according to the degree of difference. If it has changed, then according to the life prediction value of the cloud And the final prediction value Online update the parameters of the cloud prediction model.
[0022] Preferably, the judging whether the current data distribution has changed according to the life prediction value of the edge device And the life prediction value of the cloud Is:
[0023] Calculate the life prediction value of the edge device And the life prediction value of the cloud Error Δ:
[0024]
[0025] Where Δ represents the average error between the life prediction value of the cloud and the life prediction value of the edge device in T predictions;
[0026] When Δ exceeds the set threshold δ, it is determined that the current data distribution has changed, triggering the cloud adaptive adjustment mechanism to online update the parameters of the cloud prediction model; the update process is optimized through the online learning loss function, and the current final prediction value Is used as a pseudo label to adjust the weights of the cloud prediction model part; the online learning loss function Is:
[0027] Preferably, the edge prediction model is composed of an LSTM decoder;
[0028] The pre-training method of the edge prediction model is: connect the edge feature extraction model with frozen parameters and the edge prediction model in series; the trained cloud prediction model is used as the teacher model, and the edge prediction model is used as the student model, and the distillation learning method is used for training; the multi-sensor time-series data sample is input into the edge feature extraction model to generate feature Z, and is input into the cloud prediction model and the edge prediction model, and the edge prediction model outputs the life prediction value of the edge device The predicted value output by the cloud prediction model is used as a soft label The predicted value finally output by the system is used as the true label and is optimized through a weighted distillation loss function.
[0029] Preferably, the method further includes: when the edge prediction model makes an online prediction, online learning is performed: at each time step, the edge prediction model uses the current final life prediction value of the system as a pseudo label, and the edge life prediction value to calculate the current loss, and based on the current loss, the weights of the LSTM are adjusted online.
[0030] Preferably, for the dynamic weight generation model, during the online prediction, when the data distribution change of the feature Z corresponding to the real-time multi-sensor time series data exceeds the set expectation, the clustering center is dynamically updated as:
[0031] Calculate the average distance between the feature Z of the real-time multi-sensor time series data and the current clustering center. When the average distance exceeds the set range threshold, the clustering center is updated according to the following formula:
[0032]
[0033] where μ k is the k-th clustering center before update, is the k-th clustering center after update, and η is the update step size.
[0034] Preferably, the sensing device collects multi-sensor data and transmits it to the edge device through the lightweight communication protocol MQTT; the edge device sends the feature Z to the cloud server through the gRPC protocol; the cloud server sends the cloud prediction result back to the edge device through the HTTP / 2 protocol.
[0035] Beneficial effects:
[0036] (1) Taking into account the computing power of the device and the communication cost between devices: Through the distributed processing design of feature extraction, prediction, and dynamic weight calculation, the resource utilization rate of the edge and the cloud is optimized. The edge undertakes real-time and lightweight tasks, and the cloud is responsible for high-precision modeling and collaborative optimization. The organic combination of the two realizes an intelligent monitoring system with edge autonomy and cloud efficiency. While meeting real-time, accuracy, and efficiency, it reduces the computing and transmission pressure of the system and provides a high-quality solution for industrial scenario applications.
[0037] (2) With strong adaptability: The present invention uses a clustering method to determine the edge-side clustering model. By clustering the historical data of the device, representative features of the device state are extracted, improving the device's understanding ability of data distribution at the edge side and providing effective dynamic weights for downstream tasks. At the same time, during online prediction, through the edge-side dynamic weight generation model, according to the real-time changes in the device operation state, the fusion weight of the edge-side and cloud-side prediction results is adaptively adjusted. When it is found that the data distribution has changed significantly, the clustering center is dynamically adjusted to always be able to better describe the clustering structure of the data, ensuring the effectiveness of the edge-cloud prediction fusion weight.
[0038] Furthermore, the edge-side and cloud-side prediction models dynamically optimize parameters through pseudo-label online learning and cloud-side adaptive adjustment mechanisms, continuously improving the system's adaptability to the complex operation state of the device.
[0039] (3) Balancing high accuracy and high real-time performance: By combining the prediction results of the cloud side and the edge side through the dynamic weight model, the high-precision model on the cloud side ensures the prediction accuracy of the device life, and the lightweight model on the edge side significantly improves the real-time performance of the device life prediction. At the same time, the edge-side prediction model is trained by the cloud-side prediction model as the teacher model through the knowledge distillation method, improving the prediction accuracy of the lightweight model while ensuring the real-time performance of the edge-side lightweight model to meet the requirements of high accuracy and high real-time performance for device life prediction in complex unmanned platforms in industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the overall framework diagram of the method for predicting the life of a complex unmanned platform based on edge offloading provided by the present invention.
[0041] Figure 2 It is the schematic diagram of the knowledge distillation method used in the model training of the present invention.
[0042] Figure 3 It is the schematic diagram of the method for online fine-tuning of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To more clearly illustrate the technical solution of the present invention, the following further describes the present invention in detail with reference to the drawings and embodiments.
[0044] In this embodiment, taking the device life prediction for multi-sensor data as an example, the method of the present invention is specifically described. This method uses the method of offloading some cloud tasks by edge devices, making full use of the real-time performance of the edge side and the high computing power of the cloud side, aiming to efficiently extract the correlation characteristics and temporal dependence characteristics between multi-sensor data, providing support for the accurate and real-time prediction of the device life. The overall architecture of this method is as Figure 1 shown.
[0045] The deep learning complex unmanned platform life prediction system based on edge offloading constructed by the present invention includes four core models, namely: an edge feature extraction model, a cloud prediction model, an edge prediction model, and a dynamic weight generation model.
[0046] The edge feature extraction model is used to extract and perform dimensionality reduction processing on multi-sensor time series data to extract feature Z. In this embodiment, the edge feature extraction model is composed of a lightweight Transformer Encoder, aiming to perform dimensionality reduction processing on multi-sensor time series data in the time and sensor dimensions through sparse attention and multi-head attention mechanisms, extract the key features of time series and sensors, and compress them into a compact representation. Through this dimensionality reduction operation, both the computing burden of edge devices and the communication burden between different devices are reduced, and the main time series patterns and relationships between sensors are retained, providing high-quality feature inputs for the subsequent high-precision Transformer prediction model in the cloud and the lightweight LSTM prediction model at the edge.
[0047] The cloud prediction model outputs the life prediction value of the cloud according to the input feature Z It aims to deeply model the compact feature representation transmitted from the edge, further extract complex time-dependent and sensor relationship features, and finally achieve high-precision device life prediction. In this embodiment, the cloud prediction model is composed of a Transformer Encoder and a Transformer Decoder. It utilizes high computing power resources to extract deep features across time steps, combines the global interaction relationship between sensors, and realizes accurate device life prediction through the multi-head attention mechanism and the specific sequence modeling ability of the decoder.
[0048] The edge prediction model outputs the life prediction value of the edge according to the input feature Z It aims to capture the time series dynamic characteristics of the device operation state from the compact feature representation extracted by the edge feature extraction model and quickly generate the prediction result of the device life. In this embodiment, the edge prediction model is composed of an LSTM Decoder, and extracts knowledge from the high-precision cloud prediction model through the method of knowledge distillation. By combining the lightweight design of the edge model and the powerful ability of LSTM in time series modeling, this model can achieve efficient and real-time device life prediction on edge devices with limited computing resources.
[0049] The dynamic weight generation model is designed based on the edge feature extraction model and the clustering network. Its purpose is to dynamically adjust the weights of the prediction results of the edge and cloud models according to the differences between the current state and the historical state of the device, so as to generate a more accurate device life prediction value. The dynamic weight generation model dynamically generates the fusion weights of the edge and cloud prediction results by calculating the distances between the current data features and the historical feature clustering centers, enabling the model to dynamically adjust the prediction weights of the edge and cloud according to the real-time environment, and ensuring that the system can achieve the optimal device life prediction performance in different scenarios.
[0050] During the actual operation of the complex unmanned platform, the device collects multi-sensor time-series data of the operating states such as rotational speed, displacement, and force in real time, and transmits it to the edge through the lightweight communication protocol MQTT. After preliminary cleaning and normalization of the data, the edge uses the pre-trained edge feature extraction model to extract the key features of the current device operation, combines with the edge prediction model to predict the real-time device life, and transmits the features to the cloud through the gRPC protocol. The cloud prediction model receives the feature data from the edge, combines its own rich computing resources, deeply models the device operation state, further captures the complex time-series dependencies across time steps and the global relationships between multi-sensors, generates a high-precision prediction result of the device life, and transmits the cloud prediction result back to the edge through the HTTP / 2 protocol. The edge calculates the final device life prediction result according to the edge and cloud prediction results and the edge-cloud prediction fusion weights generated by the edge dynamic weight generation model.
[0051] The following describes in detail the model training and prediction processes of the method for predicting the life of a complex unmanned platform based on edge offloading according to the present invention. The method includes the following steps:
[0052] Step 1: Model pre-training
[0053] After preliminary cleaning and normalization of the historical data, historical data samples {X 1 ,X 2 ,...,X N} and the corresponding health index sequence Y = {y 1 ,y 2 ,...,y N} are obtained, where N is the total number of samples. Each input sample X contains the time-series data collected by S sensors within the time step length T, that is
[0054] Step 1.1: Pre-training of the edge feature extraction model
[0055] The edge feature extraction model consists of a lightweight Transformer Encoder. To capture the correlations between time and sensors simultaneously, the model uses a time dimensionality reduction module to capture key information across time steps through sparse attention, and a sensor dimensionality reduction module to capture the key correlations between sensors through multi-head attention.
[0056] In the time dimensionality reduction module, the sparse attention mechanism reduces the computational complexity through the Top-k attention selection strategy, only focusing on the key time steps in the input sequence, thereby efficiently extracting information across time steps. The formula for sparse attention is as follows:
[0057] Q = XW Q , K = XW K , V = XW V
[0058]
[0059] where Q is the query matrix, K is the key matrix, V is the value matrix, W Q , W K and W V are the corresponding linear transformation weight matrices, d k is the feature dimension of the query and key, and SparseMask(·) is the sparsity mask function used to retain the Top-k attention scores. For any input matrix A, the specific calculation process of the sparsity mask function is as follows:
[0060]
[0061] SparseMask(A) = Softmax(A ⊙ M)
[0062] The sparsity mask function uses the mask M to find the k positions with the largest scores for each row of the input matrix A, i.e., the Top-k strategy, assigns 0 to all non-Top-k positions through the element-wise Hadamard product ⊙, and finally performs matrix normalization. Through sparse attention, the model extracts key features from T time steps to obtain the time feature matrix after dimensionality reduction where t < T.
[0063] In the sensor dimensionality reduction module, the multi-head attention mechanism is used to learn the complex interaction relationships between sensors and extract the global representation of sensors. The formula for the multi-head attention mechanism is as follows:
[0064] MultiHead(Q, K, V) = Concat(head 1 , head 2 ,..., head H )W o
[0065] Among them, H is the number of attention heads, and head i is the output of each attention head, and W o is the output projection weight, Concat(·) is the concatenation function, and the formula for each attention head is:
[0066]
[0067] Through multi-head attention, the model extracts global interaction features from S sensors to obtain a reduced-dimensional sensor feature matrix where s < S.
[0068] Finally, the edge-side feature extraction model outputs a compact feature representation matrix as the key features of device operation, and serves as the input for the cloud prediction model, edge-side prediction model, and clustering network of the dynamic weight generation model.
[0069] In the training stage, this design imitates the idea of the Contrastive Loss to design a Continuous_Contrastive_Loss function to optimize the edge-side feature extraction model, making the reduced-dimensional features with small health state differences closer and those with large health state differences farther apart, thus ensuring that the features extracted by the edge-side feature extraction model are more discriminative and robust.
[0070] The continuous contrast loss L CCL is defined as:
[0071]
[0072] Among them, Δy ij = |y i - y j | is the difference between the i-th health indicator and the j-th health indicator, representing the health state difference; d ij = ||Z i - Z j || 2 is the Euclidean distance between the i-th reduced-dimensional feature vector Z i and the j-th reduced-dimensional feature vector Z j .
[0073] Step 1.2: Pre-training of the cloud prediction model
[0074] The cloud prediction model adopts a Transformer Encoder-Decoder structure, uses historical data to extract deep features of robot operation, and predicts the device life. The cloud prediction model inputs the two-dimensional vector of the compact feature representation output by the edge-side feature extraction model
[0075] Perform pre-training on the cloud model with the corresponding life label y, and output the predicted life on the cloud sequence
[0076] Adopt a decoupled training strategy for the pre-training of the cloud prediction model. First, train the edge feature extraction model, then freeze the parameters of the edge feature extraction model, connect the frozen edge feature extraction model and the cloud prediction model in series. Input the multi-sensor time series data sample into the edge feature extraction model to generate a fixed feature representation Z, and the cloud prediction model outputs the life prediction value The cloud prediction model learns and optimizes based on this
[0077] The Transformer Encoder of the cloud prediction model uses the multi-head attention mechanism to calculate the correlation between different elements in the input sequence
[0078] Q enc =ZW Q ,K enc =ZW K ,V enc =ZW V
[0079] MultiHead(Q enc ,K enc ,V enc )=Concat(head 1 ,head 2 ,...,head H )W o
[0080] where Q enc is the query matrix, K enc is the key matrix, V enc is the value matrix, W Q ,W K and W V are the corresponding linear transformation weight matrices, head i is the output of each attention head, W o is the output projection weight, Concat(·) is the concatenation function
[0081] The output result of the attention mechanism is fed into a feed-forward neural network (FFN), which can further extract features and enhance the non-linear expression ability of the model. FFN contains two linear transformations and a non-linear activation function
[0082] H′=σ(W 1 Attention(Qenc , K enc , V enc ) + b 1 )
[0083] H FFN = σ(W 2 H′ + b 2 )
[0084] where H′ is the intermediate result after the first linear transformation and non - linear activation processing of the feed - forward neural network, W 1 and W 2 are weight matrices, b 1 and b 2 are bias terms, and σ is a non - linear activation function.
[0085] The output H FFN of the FFN is added to the output of the multi - head attention mechanism to obtain the feature H enc extracted by the encoder:
[0086] H enc = H FFN + MultiHead(Q enc , K enc , V enc )
[0087] The Transformer Decoder uses the sequence already generated by the Decoder to calculate the self - attention weights:
[0088]
[0089] Next, the Transformer Decoder combines the hidden state H enc extracted by the Encoder and the decoder hidden state through the cross - attention mechanism to capture the correlation between the input sequence and the target output sequence:
[0090]
[0091] Each layer of the decoder contains residual connections and layer normalization (Layer Normalization) to stabilize the training process, and the formula is:
[0092]
[0093] H dec = LayerNorm(H dec + Attention cross )
[0094] where, Represents the hidden state of the lth layer Transformer Decoder. The LayerNorm layer normalization operation is used to standardize the hidden state, improve gradient flow and accelerate convergence.
[0095] After the attention mechanism, the hidden state H of the Decoder dec It is fed into a feed-forward neural network to further extract features:
[0096] H FFN =σ(W 2 (σ(W 1 H dec +b 1 ))+b 2 )
[0097] Finally, the hidden state is mapped to the predicted health indicator value through a fully connected layer (Feed-Forward Network, FFN):
[0098]
[0099] The cloud model uses the mean square error (MSE) suitable for life prediction tasks as the objective function to optimize the error between the predicted value and the true value:
[0100]
[0101] Among them, n is the number of samples for a single training, y i is the true lifespan value of the i-th sample, is the lifespan value predicted by the cloud model for the i-th sample, The loss function is used to measure the difference between the predicted value and the true value. Back propagation is performed based on the loss value to optimize the network parameters.
[0102] Step 1.3: Pre-training the edge prediction model
[0103] The edge prediction model is composed of LSTM, and the edge prediction model and the cloud prediction model share the same edge feature extraction model. Similar to the cloud prediction model training, when pre-training the edge prediction model, it is necessary to pre-connect the frozen edge feature extraction model in series to avoid affecting the accuracy of the trained edge feature extraction model and the cloud model.
[0104] The edge prediction model input is the compact feature representation matrix output by the edge feature extraction model Output predicted lifespan The model extracts time series features through LSTM, and the hidden state of LSTM is calculated recursively. The formula is as follows:
[0105]
[0106] Among them, represents the hidden state at time t, which depends on the previous hidden state and the current input Z. The hidden state is processed through a fully connected feed-forward neural network to obtain a health prediction index:
[0107]
[0108] The pre-training of the edge lightweight model adopts the distillation learning method, aiming to extract knowledge from a complex teacher model with higher performance and transfer it to the edge lightweight prediction model, so as to reduce the occupation of computing resources while ensuring the performance of the edge lightweight prediction model. The specific process is as Figure 2 shown.
[0109] The core of distillation learning training lies in using the soft labels generated by the teacher model as the learning target of the student model to further improve the generalization ability of the student model to the data distribution. In the present invention, the edge prediction model is used as the student model, and the cloud prediction model is used as the teacher model, combined with the true label y i and the predicted value extracted from the teacher model as the soft label is optimized through a weighted distillation loss function:
[0110]
[0111] Among them, β is a weight hyperparameter that controls the balance between soft labels and hard labels and can be adjusted during training. n is the number of samples in a single training. The above formulas and both calculate the average error of T predictions.
[0112] Step 1.4: Pre-training of the dynamic weight generation model
[0113] The dynamic weight generation model clusters the device historical data to extract representative features of the device state, improves the understanding ability of the edge device to the data distribution, and provides effective dynamic weights for downstream tasks.
[0114] The dynamic weight generation model receives a series of features extracted by the edge feature extraction model and uses the unsupervised clustering algorithm K-means to divide the input features into K clustering centers. The Euclidean distance between each input feature Z and the clustering center μ k is defined as follows:
[0115] d(Z,μ k )=||Z-μ k || 2
[0116] Among them, μ k is the k-th clustering center, and ||·|| 2 represents the Euclidean distance. The K-Means algorithm is used to update the clustering center during the training process, and the update method is as follows:
[0117]
[0118] Among them, is the k-th clustering center after each update during training; represents the weight of the i-th sample with respect to the clustering center k, and the calculation formula is as follows:
[0119]
[0120] Among them, d(·) is a selectable distance function.
[0121] Step 2: Prediction of the edge-cloud collaboration framework
[0122] Step 2.1: Deploy the cloud prediction model to the cloud server, and deploy the edge feature extraction model, edge prediction model, and dynamic weight generation model to the edge device.
[0123] Step 2.2: The device collects multi-sensor data and transmits it to the edge server through the lightweight communication protocol MQTT.
[0124] Step 2.3: After the edge end performs data cleaning and normalization on the received raw data, the edge feature extraction model extracts features from the data to generate a feature vector Z. The feature information is directly transmitted to the edge prediction model for real-time prediction of the device life, and at the same time, the feature vector Z is sent to the cloud prediction model through the gRPC protocol to achieve edge-cloud collaborative computing.
[0125] Step 2.4: The edge prediction model outputs the edge device life prediction value according to the feature vector Z and historical time series information
[0126] Step 2.5: The cloud receives the feature vector Z transmitted from the edge end, and uses the cloud prediction model to predict the device life to generate the cloud life prediction value Then, the prediction result is transmitted back to the edge end through the HTTP / 2 protocol.
[0127] Step 2.6: The dynamic weight generation model dynamically calculates the edge-cloud prediction fusion weight w k according to the Euclidean distance between the feature vector Z at the current time step and the historical clustering center μ t to measure the importance of edge prediction and cloud prediction.
[0128]
[0129] Among them, K is the number of clustering centers, and μ k is the k-th clustering center. The edge terminal calculates the edge-cloud prediction fusion weight w t according to the clustering network, fuses the prediction results of the edge terminal and the cloud, and outputs the final system life prediction value
[0130]
[0131] Among them, α is the weight scaling coefficient.
[0132] Step 3: Online fine-tuning of the system model
[0133] Online fine-tuning is to cope with the change of data distribution during the operation of the device, so as to ensure that the prediction performance of the model can be continuously maintained at a high level. In the design, online fine-tuning includes two core mechanisms: the adaptive adjustment mechanism of the cloud prediction model and the online learning of the edge prediction model. Combining the dynamic update of the clustering center to adapt to the change of data distribution, the modules of the online fine-tuning of the system model are as Figure 3 shown.
[0134] Step 3.1: Adaptive adjustment mechanism of the cloud prediction model
[0135] The adaptive adjustment mechanism of the cloud prediction model stores the T cloud prediction results predicted by the system Edge prediction results and system prediction results At the same time, the cloud detects the edge prediction results and cloud prediction results to judge whether the current data distribution has changed. The error detection formula is:
[0136]
[0137] Among them, Δ represents the average error between the cloud prediction result and the edge prediction result in T predictions. When Δ exceeds a certain threshold δ, it is determined that the current data distribution has changed, triggering the cloud adaptive adjustment mechanism to update the parameters of the cloud prediction model.
[0138] The update process is optimized through the online learning loss function, and the current system life prediction value is used as the pseudo label to adjust the weights of the cloud prediction model part. The online learning loss function is defined as follows:
[0139]
[0140] Step 3.2: Online learning of the edge prediction model
[0141] The purpose of online learning of the edge prediction model is to enable the lightweight model to adapt to the changes in the data distribution at the device end. At each time step, the edge prediction model uses the current system prediction result as a pseudo-label, and dynamically updates the online learning loss function by adjusting the weights of the LSTM as follows:
[0142]
[0143] This process can slightly update the parameters of the model according to the data collected in real time at the device end, enabling it to quickly respond to the changes in the real-time data distribution, thus ensuring the accuracy and robustness of edge prediction.
[0144] Step 3.3: Dynamic update of the clustering center
[0145] To ensure the effectiveness of the edge-cloud prediction fusion weights, an update mechanism for the clustering center is introduced in the design. When a significant change in the data distribution is detected, the clustering center μ is dynamically adjusted k so that it can always better describe the clustering structure of the data. The specific update formula is as follows:
[0146]
[0147] where η is the update step size. This formula gradually adjusts the clustering center to the new data distribution center through a weighted update method.
[0148] Among them, whether the clustering center needs to be updated is determined by the following conditions:
[0149]
[0150] When the average distance Threshold (i.e., the threshold) between the new data points and the existing clustering center exceeds a certain range, the update process of the clustering center is triggered. The dynamic update of the clustering center further improves the model's ability to represent the changing data distribution, thus achieving a more robust device life prediction.
[0151] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the spirit and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.
Claims
1. A deep learning complex unmanned platform life prediction method based on edge offloading, characterized in that: include: Deploy cloud prediction models to cloud servers; deploy edge feature extraction models, edge prediction models, and dynamic weight generation models to edge devices; use high-precision models for cloud prediction models, and lightweight models for edge feature extraction models and edge prediction models; The edge feature extraction model performs dimensionality reduction processing on multi-sensor time series data, extracts feature Z, and sends it to the cloud prediction model, edge prediction model, and dynamic weight generation model; The cloud prediction model outputs the cloud life prediction value based on the input feature Z. Pass it back to the edge device; The edge prediction model outputs the predicted lifespan of the edge based on the input feature Z. The dynamic weight generation model calculates the edge-cloud prediction fusion weight w based on the input feature Z and the distance from the current cluster center. t ; The edge device predicts the fusion weight w according to the edge cloud t , integrating the lifespan prediction results of edge and cloud and Output the final life expectancy prediction value Among them, during training, the dynamic weight generation model performs clustering and iterative updates based on the historical data of feature Z to obtain the historical cluster center; during online prediction, when the data distribution change of the real-time multi-sensor time series data corresponding to feature Z exceeds the set expectation, the cluster center is updated to dynamically adjust the edge-cloud prediction fusion weight.
2. The method according to claim 1, characterized in that The edge feature extraction model is composed of a lightweight Transformer encoder, which reduces the dimension of multi-sensor time series data in time and sensor dimensions through sparse attention and multi-head attention mechanisms, and extracts the key features Z of time series and sensors.
3. The method according to claim 1, characterized in that The pre-training method of the edge feature extraction model is: using a continuous contrastive learning loss method to optimize the edge feature extraction model; and optimizing the model by comparing the lifespan label differences between samples and the Euclidean distance of the feature Z after dimensionality reduction.
4. The method according to claim 1, characterized in that The cloud prediction model adopts a Transformer encoder-decoder structure; The pre-training method of the cloud prediction model is: the edge feature extraction model with frozen parameters and the cloud prediction model are connected in series, multi-sensor time series data samples are input into the edge feature extraction model to generate feature Z, which is input into the cloud prediction model, and the cloud prediction model outputs the cloud life prediction value Cloud-based lifespan prediction Calculate the loss with the lifetime label y and optimize the cloud-based prediction model.
5. The method according to claim 1 or 4, characterized in that The method further includes: when the cloud prediction model is predicting online, adaptively adjusting the cloud prediction model parameters: Store the predicted life expectancy values of T clouds Life expectancy of the edge and the final predicted value According to the predicted value of edge life and cloud life expectancy The difference between the current data distribution and the lifespan prediction value of the cloud is used to determine whether the current data distribution has changed. If so, and the final predicted value Update the cloud prediction model parameters online.
6. The method according to claim 5, characterized in that The edge life prediction value and cloud life expectancy The degree of difference between the current data distribution and the change of the current data distribution is as follows: Calculate the predicted edge life and cloud life expectancy The error Δ: Where Δ represents the average error between the cloud life prediction value and the edge life prediction value in T predictions; When Δ exceeds the set threshold δ, it is determined that the current data distribution has changed, triggering the cloud-based adaptive adjustment mechanism to update the cloud-based prediction model parameters online; the update process is optimized through the online learning loss function, using the current final prediction value As a pseudo-label, adjust the weight of the cloud prediction model; online learning loss function for:
7. The method according to claim 1, characterized in that The edge prediction model is composed of an LSTM decoder; The pre-training method of the edge prediction model is: connecting the edge feature extraction model with frozen parameters and the edge prediction model in series; the trained cloud prediction model is used as the teacher model, and the edge prediction model is used as the student model, and the training is performed using a distillation learning method; multi-sensor time series data samples are input into the edge feature extraction model to generate feature Z, which is input into the cloud prediction model and the edge prediction model, and the edge prediction model outputs the life prediction value of the edge The predicted value output by the cloud prediction model is used as a soft label The predicted value of the final output of the system As the true label, it is optimized through the weighted distillation loss function.
8. The method according to claim 7, characterized in that The method further includes: when the edge prediction model is online predicting, online learning is performed: at each time step, the edge prediction model uses the final life prediction value of the current system As pseudo labels, and edge life prediction values Calculate the current loss and adjust the LSTM weights online based on the current loss.
9. The method according to claim 1, characterized in that For the dynamic weight generation model, during the online prediction, when the data distribution change of the feature Z corresponding to the real-time multi-sensor time series data exceeds the set expectation, the cluster center is dynamically updated as follows: Calculate the average distance between the feature Z of the real-time multi-sensor time series data and the current cluster center. When the average distance exceeds the set range threshold, update the cluster center according to the following formula: Among them, μ k is the kth cluster center before updating, is the updated k-th cluster center, and η is the update step size.
10. The method according to claim 1, characterized in that The sensor device collects multi-sensor data and transmits it to the edge device through the lightweight communication protocol MQTT; the edge device sends feature Z to the cloud server through the gRPC protocol; the cloud server transmits the cloud prediction results back to the edge device through the HTTP / 2 protocol.
Citation Information
Cited By
Kitchen air conditioner filtering device service life prediction method and system based on deep learning
CN120950908A
A Deep Learning-Based Method and System for Predicting the Lifespan of Kitchen Air Conditioner Filter Devices
CN120950908B
Cloud edge collaborative health state diagnosis method and system for multiple coal mining machines
CN121052339A