Space-Time Prediction Algorithm Based on Federated Learning on Industrial Internet of Things Edge Devices

By combining the TCN-GCN model of expansion time convolution network and dynamic graph convolution network on industrial IoT edge devices, and using the improved FedAVG algorithm for federated learning, the privacy protection and data silos problems are solved, and efficient space-time data prediction is achieved.

CN114265913BActive Publication Date: 2025-07-01INNER MONGOLIA UNIVERSITY +1

Patent Information

Application Number
CN202111654558.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-01
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing space-time prediction algorithm cannot effectively adapt to distributed scenarios on industrial IoT edge devices, and due to privacy protection issues, data silos are serious, affecting prediction performance.

Method used

The TCN-GCN model based on expansion time convolution network and dynamic graph convolution network is used for local training, and the weighted aggregation of model parameters is carried out in the cloud through the improved FedAVG algorithm to realize federated learning, reduce communication overhead and protect privacy.

Benefits of technology

Under the premise of privacy protection, high-precision space-time data prediction is achieved, suitable for large-scale IIoT networks, improving prediction performance and reducing communication overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114265913B_ABST
    Figure CN114265913B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-temporal prediction algorithm based on federated learning for industrial Internet of Things (IIoT) edge devices, belonging to the field of industrial Internet of Things. Specifically, it includes: First, an application scenario including an industry set and a customer set is built; Then, for a single client, the sensing data monitored by its device forms a database, and a TCN-GCN deep model is constructed in this client, and the parameters wC of the TCN-GCN deep model are updated using the data in the database; Similarly, each client trains its own TCN-GCN model locally and obtains its own model parameters wC; Finally, the model parameters wC of each client are uploaded to the cloud for weighted aggregation using an improved FedAVG algorithm to form a new global model, realizing spatio-temporal prediction based on federated learning. The present invention randomly samples the customers participating in federated learning and the devices within the customers, reducing the communication overhead of the algorithm, and is particularly suitable for large-scale IIoT networks and distributed prediction; on the basis of protecting privacy, it has good prediction performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial Internet of things, and particularly relates to a spatio-temporal prediction algorithm based on federated learning on industrial Internet of things edge devices. Background Art

[0002] The wide deployment of industrial Internet of things (IIoT) edge devices has given rise to various emerging applications with edge computing, such as intelligent manufacturing, intelligent logistics, and factory monitoring and management. Edge devices provide powerful computing resources, enabling IIoT applications to make decisions in real time, flexibly, and quickly, and greatly promoting the development of Industry 4.0.

[0003] In addition, with the popularization and application of information systems, low-latency and efficient networking make it possible to share a large amount of data in the industrial Internet of things, and improve the user experience or system performance through predictive analysis of a large amount of shared data. Therefore, accurate prediction of IIoT sensing layer monitoring data is crucial, providing guarantees for intelligent monitoring, intelligent diagnosis, intelligent response, and control devices, and related research has attracted extensive attention from scholars and the industrial field.

[0004] Currently, edge devices (such as industrial robots) usually collect sensing data from IIoT nodes, analyze and capture the behaviors and operating conditions of IIoT nodes through edge computing [1], achieving effective coverage and efficient measurement and control. However, IIoT applications face serious security risks caused by node anomalies, hindering the rapid development of IIoT. For example, in an intelligent manufacturing scenario, an engine with sensors as a node device of IIoT, if abnormal behaviors occur (such as abnormal flow rate, abnormal reporting frequency), may cause industrial production interruption, resulting in huge economic losses to the factory. Therefore, it is very necessary and challenging to monitor and highly accurately detect sensing layer data.

[0005] With the combination of artificial intelligence (AI) technology and IIoT big data, deep learning (DL) has become an effective solution for realizing the analysis and accurate prediction of sensing layer monitoring data.

[0006] In the prior art, Reference [2] proposed to use a Convolutional Long Short-Term Memory (ConvLSTM) network for prediction. First, two-dimensional convolution is used to capture relevant features in the surrounding area, and then LSTM is used to extract features in the time dimension. Reference [3] proposed a unified framework that integrates a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) network for multi-node prediction, effectively extracting the time-varying features of the IIoT network. Reference [4] proposed a prediction mechanism using multi-task learning, combined with a deep architecture based on the LSTM model to achieve high-precision prediction. Reference [5] proposed to use a Temporal Graph Convolutional Network (T-GCN) to solve the constraint problem of the topological structure. Reference [6] proposed a ResLSTM deep learning architecture that combines a Residual Network (ResNet), a Graph Convolutional Network (GCN), and an LSTM to predict short-term passenger flow. Reference [7] proposed a spatio-temporal prediction method based on an attention mechanism. First, the attention mechanism is used to extract the global features of the target point, and then the extracted spatial features are input into the LSTM network to obtain the long-term state information of the spatial factors. Reference [8] proposed a spatio-temporal deep learning framework that accurately predicts spatio-temporal data by combining ConvLSTM and GCN. To capture spatial relationships more comprehensively, Reference [9] established a dynamic graph network to learn the complex spatial and dynamic time correlations presented in spatio-temporal data by combining GCN and LSTM.

[0007] Although the existing spatio-temporal prediction algorithms have improved the prediction accuracy by combining spatial features and time features and achieved phased progress. However, the above-mentioned literature cannot be directly applied to the IIoT scenario with distributed edge devices: First, most spatio-temporal prediction models in deep learning methods are not flexible enough, and edge devices lack dynamic and automatically updated prediction models for different scenarios, resulting in the inability to accurately predict frequently updated time series data. Second, due to privacy issues, edge devices cannot share the collected sequence data, so the data exists in the form of "islands". Data islands seriously reduce the performance of data prediction. In addition, complex neural networks use a large amount of data to train and optimize models, which may lead to potential privacy leakage.

[0008] Therefore, to flexibly deploy the spatio-temporal prediction model at the edge, solve the problem of insufficient training sets caused by privacy protection, and thus degrade the performance of the spatio-temporal prediction model, Federated Learning (FL), as a promising distributed machine learning paradigm, is proposed to meet the requirements of privacy protection. Each distributed terminal collaboratively trains a globally shared model through its local data, while keeping the training data set locally and uploading the locally updated model weights without exchanging the original data, effectively preventing the leakage of privacy.

[0009] References:

[0010] [1] Y. Liu et al., "Deep anomaly detection for time-series data in industrial iot: A communication-efficient on-device federated learning approach," IEEE Internet of Things J., vol. 8, no. 8, pp. 6348 - 6358, Apr. 15, 2021.

[0011] [2] X. Shi, Z. Chen, H. Wang, D. Yeung, W. Wong, and W. Woo, “Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting” Adv. neural inf. proces. syst., Montreal, QC, Canada, Jan. 2015, pp. 802 - 810.

[0012] [3] Q. Zhu, J. Chen, D. Shi, L. Zhu, X. Bai, X. Duan, and Y. Liu, "Learning temporal and spatial correlations jointly: A unified framework for wind speed prediction," IEEE Trans. Sustainable Energy, vol. 11, no. 1, pp. 509 - 523, Jan. 2020.

[0013] [4]L. Nie, X. Wang, S. Wang, Z. Ning, and S. Li, “Network traffic prediction in industrial internet of things backbone networks: a multi-task learning mechanism.” IEEE Trans. Ind. Inf., doi:10.1109 / TII.2021.3050041.

[0014] [5]L. Zhao, Y. Song, C. Zhang, Y. Liu, P. Wang, T. Lin, M. Deng and H. Li, "T-GCN: A temporal graph convolutional network for traffic prediction," in IEEE Trans. Intell. Transp. Syst., vol. 21, no. 9, pp. 3848-3858, Sept. 2020.

[0015] [6]J. Zhang, F. Chen, Z. Cui, Y. Guo and Y. Zhu, "Deep learning architecture for short-term passenger flow forecasting in urban rail transit," IEEE Trans. Intell. Transp. Syst., doi:10.1109 / TITS.2020.3000761.

[0016] [7]S. Duan, W. Yang, X. Wang, S. Mao, and Y. Zhang, “Temperature forecasting for stored grain: A deep spatio-temporal attention approach.” IEEE Internet Things J., doi:10.1109 / JIOT.2021.3078332.

[0017] [8]F.Dai,P.Huang,X.Xu,L.Qi,and M.R.Khosravi,“Spatio-temporal deeplearning framework for traffic speed forecasting in iot.”IEEE Internet ofThings J.,vol.3,no.4,pp.66-69,2020.

[0018] [9]Z.Cui,K.Henrickson,R.Ke and Y.Wang,"Traffic graph convolutionalrecurrent neural network:A deep learning framework for network-scale trafficlearning and forecasting,"IEEE Trans.Intell.Transp.Syst.,vol.21,no.11,pp.4883-4894,Nov.2020. Summary of the Invention

[0019] In order to accurately predict spatio-temporal data under the constraints of privacy protection, the present invention proposes a spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things (IIoT) edge devices. By aggregating model parameters from different geographical locations and different institutions, a global deep learning model is constructed under the condition of privacy protection, which has good prediction performance.

[0020] The spatio-temporal prediction algorithm based on federated learning on the industrial Internet of Things edge devices is specifically as follows:

[0021] Step 1: Build an application scenario including an industry set and a customer set;

[0022] The industry set is O = {O1, O2,..., O n ,..., O N}, each industry has I clients, and each client has K sensor devices; the customer set is C = {C1, C2,..., C i ,..., C I}

[0023] Step 2: For client C i , the sensing data monitored by its K devices constitutes database D i , and a TCN-GCN deep model is constructed in this client C i , and the parameters w of the TCN-GCN deep model are updated using the data in database D i ​C ;

[0024] The TCN-GCN deep model is formed by alternately combining the dilated temporal convolutional network DTCN and the dynamic graph convolutional network DGCN;

[0025] The specific update process is as follows:

[0026] Step 201, Client C i Use K sensors to collect two-dimensional data containing time steps and spatial nodes respectively, and construct an initial graph adjacency matrix to represent the correlation relationship of spatial nodes;

[0027] Step 202, Send the two-dimensional data into the TCN-GCN deep model in sequence, and extract temporal and spatial features respectively in combination with the graph adjacency matrix;

[0028] Specifically:

[0029] First, given the input two-dimensional data x in , then the output x out of passing through the DTCN is:

[0030] x out = tanh(f1 * x in ) × sigmoid(f2 * x in )

[0031] Among them, f1 represents the filtering convolution function, f2 represents the gating convolution function. sigmoid(·) represents the S-shaped activation function, and tanh(·) represents the hyperbolic tangent activation function.

[0032] Then, send the features captured by the DTCN into the DGCN module, and its information propagation layer is:

[0033]

[0034]

[0035]

[0036] Among them, H l represents the propagation layer of the l-th layer; σ1 and σ2 are different activation functions, is the dynamic graph adjacency matrix obtained by random dynamic sampling, and W l-1 represents the network weight of the (l - 1)-th layer; represents the propagation layer after skip connection; β is a hyperparameter that controls the ratio of retaining the original state of the root node; H (l) is the propagation layer of the l-th layer where the node state is continuously updated as the depth of the graph convolution increases, and H (l-1) represents the propagation layer of the previous node state retained; Hout It is the output layer after the jump layer superposition.

[0037] Step 203: Use the TCN-GCN model to update the associations of each spatial node in the graph adjacency matrix;

[0038] Step 204: Perform a convolution operation on the updated graph adjacency matrix and the extracted spatial feature vectors to continuously update the spatial features of the moving nodes;

[0039] Step 205: Achieve high-precision prediction of two-dimensional data according to the captured temporal and spatial correlations;

[0040] The specific prediction process is as follows:

[0041] The two-dimensional data X input at time step t is expressed as:

[0042] X = {z1[i], z2[i], …, z t [i]}

[0043] where z t [i] represents the value of the i-th sensor at time step t, i ∈ K;

[0044] Then the predicted value at the next time step is expressed as:

[0045] Y = {z t+1 [i]}

[0046] Step 206: By continuously fitting the predicted value with the label value of the real data, obtain the parameters w of the TCN-GCN deep model updated locally at client C i . C .

[0047] Step Three: Each client trains its corresponding TCN-GCN model locally to obtain its respective model parameters w C , w C which contains the spatio-temporal characteristics mined by the TCN-GCN model from the original data;

[0048] Step Four: Use the improved FedAVG algorithm to upload the TCN-GCN model parameters w of each client C to the cloud for weighted aggregation into a new global model to achieve spatio-temporal prediction based on federated learning.

[0049] The improved FedAVG algorithm includes the following steps:

[0050] Step I: At the beginning of each round, the server where industry O is located selects volunteers from the clients to participate in this round of training and broadcasts the globally initialized model parameters of TCN-GCN to the selected clients.

[0051] Step II: Randomly sample the participating customers and all the sensor devices included in the customers;

[0052] For each randomly sampled client, the model parameter F l (w) that minimizes its loss is:

[0053]

[0054] where w ∈ R d is the local model parameter of each sensor device, h(·) is the regularization term, and D k represents the k samples after device sampling that the client needs to learn in the local model.

[0055] In the cloud, the average model parameter F g (w) obtained after client sampling of the global prediction model is expressed as

[0056]

[0057] n = β × N

[0058] where N is the total number of all customers in the industry, and β represents the sampling ratio of customers.

[0059] Therefore, after double sampling of the customers and devices, the model parameter obtained by the client after weighted averaging by the central cloud is:

[0060]

[0061] The device sampling of the client is reflected in the node sampling of the graph adjacency matrix, specifically:

[0062] For a client with K devices, the mathematical representation of its randomly initialized graph adjacency matrix A is:

[0063]

[0064] where A ij represents the relationship between the i,j nodes in the graph adjacency matrix, is the set of vertices, (v i ,v j ) represents the relationship between a pair of nodes, and the edge connecting the vertices is c represents a non-zero real number.

[0065] To reduce the communication overhead, sample the K devices of each client, and the node index idx after taking the first k relevant devices is expressed as

[0066] idx = arg topk(A[i,:])

[0067] Among them, arg topk(·) returns the indices of the top k maximum values of the vector; for each device, its top k nearest edge computing nodes are selected as its neighbors, that is, the top k most relevant subsets among multiple devices are filtered out; A[i,:] represents the node relationship between the i-th node and all other nodes in the graph adjacency matrix.

[0068] Step III: Initialize the set number of rounds E and the training batch B, and update the TCN-GCN model parameters for the sampled clients on the local training dataset

[0069] Step IV: The server weights and averages the parameters of each TCN-GCN model through a parameter aggregation mechanism and updates and applies it to its global state, repeating until a new global model is synthesized, realizing spatio-temporal prediction based on federated learning.

[0070] The advantages of the present invention are as follows:

[0071] 1), A spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things edge devices, a spatio-temporal prediction algorithm considering privacy protection, combines the emerging FL with the spatio-temporal prediction algorithm based on TCN-GCN, provides reliable data privacy protection through local training models, and does not require the exchange of raw data.

[0072] 2), A spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things edge devices. To improve the scalability of FL in spatio-temporal prediction, the present invention adopts an improved federated averaging (FedAVG) algorithm, randomly samples the clients participating in federated learning and the devices within the clients to reduce the communication overhead of the algorithm, and is particularly suitable for large-scale IIoT networks and distributed prediction.

[0073] 3), A spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things edge devices, uses the FedTCN-GCN algorithm to conduct simulation verification on public datasets, and the results show that it has good prediction performance on the basis of protecting privacy.

[0074] 4), A spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things edge devices, the FedTCN-GCN prediction algorithm based on federated learning is used to predict the sensing layer monitoring data on industrial edge devices. Taking into account the needs of saving device resources and application performance, a federated learning strategy is developed to allow clients to locally train raw data and perform weighted averaging of model gradient updates on the server to improve the prediction performance of the client models deployed on edge devices.

[0075] 5) The spatio-temporal prediction algorithm based on federated learning on industrial Internet of Things (IIoT) edge devices evaluates the spatio-temporal prediction TCN-GCN algorithm before and after federated learning on the dataset monitored by multiple sensors. The experimental results show that although the FedTCN-GCN prediction algorithm is slightly less accurate than the TCN-GCN in terms of prediction accuracy, it achieves better prediction performance than existing advanced spatio-temporal algorithms on the premise of privacy protection. In addition, to more flexibly meet the requirements of large-scale customers sharing model parameters for communication overhead, the proposed FedTCN-GCN prediction algorithm effectively reduces the communication overhead generated by uploading large-scale model parameters to the server through double sampling of devices and customers. Description of the Drawings

[0076] Figure 1 Block diagram of the IIoT prediction system in the prior art shown in the present invention;

[0077] Figure 2 Flowchart of the spatio-temporal prediction algorithm based on federated learning on IIoT edge devices of the present invention;

[0078] Figure 3 Federated learning framework diagram in the IIoT shown in the present invention.

[0079] Figure 4 Spatio-temporal prediction teacher network model diagram of the sensing layer data shown in the present invention.

[0080] Figure 5 Comparison diagram of prediction curves before and after federated learning shown in the present invention.

[0081] Figure 6 Comparison diagram of training losses before and after federated learning shown in the present invention.

[0082] Figure 7 Diagram of the changes in model errors MAE and RMSE with the number of sampled devices shown in the present invention. Detailed Description of the Invention

[0083] The present invention will be further described in detail below in conjunction with the drawings and embodiments.

[0084] The present invention proposes a time-space prediction algorithm based on federated learning on industrial Internet of Things (IIoT) edge devices (Time-Space Convolution Networks Prediction Algorithm Based on Federated Learning, FedTCN-GCN), which is used to accurately predict the sensing layer monitoring data on industrial edge devices under the constraints of privacy protection. Aiming at the requirements of both saving device resources and application performance, a federated learning strategy is developed to allow customers to locally train the original data and implement the weighted average of model gradient updates on the server, so as to improve the prediction performance of the customer models deployed on edge devices. The time-space prediction before and after federated learning is evaluated on the multi-sensor monitoring dataset, which will play an important role in many industrial applications that require privacy protection.

[0085] The traditional TCN-GCN is a high-precision time-space data prediction algorithm for multi-dimensional time-space sequences by aggregating information from different spatial nodes. The FedTCN-GCN of the present invention aggregates the model parameters from different geographical locations and different institutions to construct a global deep learning model under the condition of privacy protection. By combining the emerging FL with the time-space prediction algorithm based on TCN-GCN, reliable data privacy protection is provided through local training of the model without exchanging the original data.

[0086] To improve the scalability of FL in time-space prediction, the FedTCN-GCN algorithm designed in the present invention adopts an improved federated average (FedAVG) algorithm, which randomly samples the customers participating in federated learning and the devices within the customers to reduce the communication overhead of the algorithm, and is especially suitable for large-scale IIoT networks and distributed prediction.

[0087] As Figure 1 shown, the IIoT prediction system integrates the latest technologies such as cloud computing, big data, and mobile Internet with the control terminal to realize data acquisition, data communication, data analysis, and monitoring. With the characteristics of low power consumption and multi-function, it is suitable for intelligent industrial systems in various environments and provides an effective solution. The IIoT time-space data prediction system is mainly divided into three modules, including the sensor data acquisition layer, the edge intelligent prediction layer, and the data monitoring and analysis layer;

[0088] Intelligent sensors have been successfully applied to sense data in scenarios such as industrial control, environmental monitoring, smart grid, digital oilfield, and intelligent industry. By sending the sensing layer monitoring data with spatial correlation into the time-space network, features are extracted respectively to achieve accurate prediction. Finally, anomaly judgment and joint intelligent decision-making are carried out.

[0089] As a key component of the IIoT, the sensor layer consists of a large number of sensors deployed in various industrial environments to provide low-cost monitoring services. The aggregation sensor nodes analyze and predict the data collected by the sensors regularly for final decision-making. Currently, the sensor data acquisition layer based on cloud computing preliminarily screens a large amount of data collected in the monitoring area and then transmits it to the cloud for analysis and calculation. With the rise of edge computing, the edge devices deployed at the bottom effectively make up for the deficiency of the sensor computing power. At the same time, part of the work of the sensor nodes and cloud services is offloaded to the edge services, which can effectively save bandwidth and prevent the nodes from being attacked during the data uploading process, improving security.

[0090] To ensure the prediction performance and achieve privacy protection, the present invention applies the strategy of federated learning to the spatio-temporal prediction of industrial Internet of Things. By aggregating the model parameters updated by multiple clients in their local datasets, resource sharing is realized. While effectively ensuring privacy security, it meets the requirements of deep learning models for a large amount of data, and then ensures the performance of the proposed FedTCN-GCN prediction algorithm; the present invention provides a safe and high-quality analysis basis for intelligent decision-making; updating the model parameters on edge devices without handing over a large amount of raw data to the cloud greatly improves the processing efficiency and reduces the load on the cloud, and has obvious advantages in terms of security, energy consumption and network lifetime. Since FL makes the model update closer to the users and aggregates the parameters in the cloud, it provides the basic needs in terms of security and privacy protection for users.

[0091] The spatio-temporal prediction algorithm based on federated learning on the industrial Internet of Things edge devices is as Figure 2 shown, and the specific steps are as follows:

[0092] Step 1: Build an application scenario including an industry set and a client set;

[0093] The present invention uses "organization" to describe different application scenarios in the prediction of sensor layer monitoring data, such as intelligent animal husbandry, intelligent manufacturing, etc. "Client" describes different entities, such as different enterprises and factories. And "device" is used to describe the edge computing nodes corresponding to one or more sensors in FL.

[0094] The industry set is O = {O1, O2,..., O n ,..., O N}, each industry has I clients, and each client has K sensor devices; the client set is C = {C1, C2,..., C i ,..., C I}.

[0095] In the context of industrial Internet of Things data prediction, strict privacy protection needs to be achieved among clients; each client updates the gradient and trains its local model by using the local dataset instead of sharing data.

[0096] Step 2. For client C i , the sensing data monitored by its K devices constitutes database D i , and in this client C i , construct a TCN-GCN deep model and use the local training data of database D i to update the parameters w of the TCN-GCN deep model C ;

[0097] In order to accurately predict the spatio-temporal sequence of multi-sensor nodes, a deep neural network that can accurately model the temporal correlation and spatial correlation of multi-dimensional sensing data needs to be designed. The present invention uses a prediction algorithm combining a dilated time convolutional network (DTCN) and a dynamic graph convolutional network (DGCN) to achieve spatio-temporal prediction, as Figure 3 shown.

[0098] First, send the two-dimensional data containing time information and spatial information into the network, and randomly initialize the graph adjacency matrix according to its spatial feature dimension. Then, send the input data into the DTCN module and the DGCN module in sequence to extract features, and use the graph learning module to update the node information. After that, send the adjacency matrix updated in the graph learning process and the relevant nodes into the graph convolution module for operation, so as to capture the spatial correlation.

[0099] Perform DTCN and DGCN operations alternately according to the number of model layers to learn the temporal and spatial correlations multiple times.

[0100] In addition, before the start of the time convolution and after the end of the graph convolution in the present invention, residual connections are developed to avoid the problem of gradient disappearance. At the same time, layer regularization is designed to prevent the model from being overly complex, and the captured hidden features are mapped to the required output size according to the learning objective.

[0101] Finally, according to the implicit spatial relationship, move the multi-sensor time series to complement each other's key information, and achieve high-precision prediction of the two-dimensional spatio-temporal sequence of multiple nodes simultaneously.

[0102] As Figure 4 shown, the specific update process is as follows:

[0103] Step 201. Client C iUse K sensors to collect two-dimensional data containing time steps and spatial nodes respectively, and construct an initial graph adjacency matrix to represent the association relationship of spatial nodes;

[0104] Step 202: Send the two-dimensional data into the TCN-GCN deep model in sequence, and extract time and spatial features by combining the graph adjacency matrix;

[0105] The DTCN module includes two dilated convolutional layers, activated by the tangent hyperbolic function and the sigmoid function respectively. Given the input two-dimensional data x in , then the output x out after passing through the DTCN is:

[0106] x out = tanh(f1 * x in ) × sigmoid(f2 * x in )

[0107] Where f1 represents the filtering convolutional function, f2 represents the gating convolutional function. sigmoid(·) represents the S-shaped activation function, and tanh(·) represents the tangent hyperbolic activation function.

[0108] Then, send the features captured by the DTCN into the DGCN module, and its information propagation layer is:

[0109]

[0110]

[0111]

[0112] Where H l represents the propagation layer of the l-th layer; σ1 and σ2 are different activation functions, is the dynamic graph adjacency matrix obtained by random dynamic sampling, and W l-1 represents the network weight of the (l - 1)-th layer; represents the propagation layer after skip connection; β is a hyperparameter that controls the ratio of retaining the original state of the root node; H (l) is the propagation layer of the l-th layer where the node state is continuously updated as the depth of the graph convolution increases, and H (l-1) represents the propagation layer of the previous node state retained; H out is the output layer after the skip layers are stacked.

[0113] By stacking different node state information, prevent the information from the higher layer from having a negative impact on the overall performance. At the same time, retain the feature information of the nodes and the propagation information of the previous layer to prevent overfitting and improve the prediction performance of the DGCN.

[0114] Step 203: Update the associations of each spatial node in the graph adjacency matrix using the TCN-GCN model;

[0115] Step 204: Perform a convolution operation on the updated graph adjacency matrix and the extracted spatial feature vectors to continuously update the spatial features of the moving nodes;

[0116] Step 205: Achieve high-precision prediction of two-dimensional data based on the captured temporal and spatial correlations;

[0117] Without sharing the original data and privacy leakage, the goal of the present invention is to predict the monitoring values of the sensing layer at the next moment from different clients; the specific prediction process is as follows:

[0118] For client C i , the two-dimensional data X input at the given historical time step t is expressed as:

[0119] X = {z1[i], z2[i], …, z t [i]}

[0120] where z t [i] represents the value of the i-th sensor at the time step t, and i ∈ K;

[0121] Then the predicted value at the next time step of the multi-sensor is expressed as:

[0122] Y = {z t+1 [i]}

[0123] Step 206: By continuously fitting the predicted value with the label value of the real data, obtain the parameter w i of the locally updated TCN-GCN deep model of client C C .

[0124] Step 3: Each client locally trains its corresponding TCN-GCN model respectively to obtain its own model parameter w C ;

[0125] w C contains the spatio-temporal characteristics mined by the TCN-GCN model from the original data;

[0126] Step 4: Use the improved FedAVG algorithm to upload the TCN-GCN model parameters w C of each client to the cloud for weighted aggregation to form a new global model, and achieve spatio-temporal prediction based on federated learning.

[0127] To reduce the communication cost consumed when customers upload and update model weights, improve the efficiency of data prediction to ensure the real-time performance of the system, the present invention proposes to use an improved FedAVG algorithm as a spatio-temporal prediction model parameter aggregation mechanism to collect gradient update information from different customers. At the same time, by randomly sampling the devices K included in the client and the participating clients to reduce the communication overhead of the algorithm, the proposed FedTCN-GCN can achieve high-performance prediction on a large scale and in a distributed manner under the condition of protecting privacy.

[0128] Use the improved FedAVG algorithm as the aggregation mechanism of multiple local client model parameters to collect gradient update information from different customers. In the improved FedAVG algorithm, to reduce the communication overhead, randomly sample the K devices included in the client and the participating clients, and achieve high-performance prediction on a large scale and in a distributed manner under the condition of protecting privacy.

[0129] Each client has local training data; at the beginning of each round, the algorithm randomly samples a part of the clients and devices, and the server where industry O is located sends the current global algorithm state to each sampled client. Then, each selected client performs local calculations according to the global state and its local dataset, and sends updates to the server. The server applies these updates to its global state and repeats the process.

[0130] The improved FedAVG algorithm includes the following steps:

[0131] Step I: Each client has a local database. At the beginning of each round, the server where industry O is located selects volunteers from the clients to participate in this round of training, and broadcasts the globally initialized model parameters of TCN-GCN to the selected clients.

[0132] Step II: Randomly sample the participating clients and all the sensor devices included in the clients;

[0133] To further reduce the computing resources spent by the client when uploading to the server, the present invention also randomly samples the clients within industry O. For each client trained on the local dataset, the model parameter F l (w) is:

[0134]

[0135] where w ∈ R d is the local model parameter of each sensor device, h(·) is the regularization term, and D k represents the k samples after device sampling that the client needs to learn in the local model.

[0136] On the cloud, the average model parameter F obtained after the global prediction model is sampled by the client g is represented as

[0137]

[0138] n = β × N

[0139] where N is the total number of all customers in the industry, and β represents the customer sampling ratio.

[0140] Therefore, after double sampling of customers and devices, the model parameters obtained by the client after weighted averaging by the central cloud are:

[0141]

[0142] Sampling the devices included in the client can be achieved by sampling the nodes of the graph adjacency matrix; specifically:

[0143] For a client with K devices, the mathematical representation of its randomly initialized graph adjacency matrix A is:

[0144]

[0145] where A ij represents the relationship between the i,j-th nodes in the graph adjacency matrix, is the set of vertices, (v i , v j ) represents the relationship between a pair of nodes, and the edge connecting the vertices is c represents a non-zero real number.

[0146] To reduce the communication overhead, sample the K devices of each client, and the node index idx after taking the first k relevant devices is represented as

[0147] idx = arg topk(A[i,:])

[0148] where arg topk(·) returns the indices of the first k maximum values of the vector; for each device, select its first k nearest edge computing nodes as its neighbors, that is, filter out the first k most relevant subsets among multiple devices; A[i,:] represents the node relationship between the i-th node and all other nodes in the graph adjacency matrix.

[0149] By sampling the devices of client C, the number of model parameters updated by each client for local calculation is reduced.

[0150] Step III. Initialize the number of communication rounds E and training batches B with the server, and update the TCN - GCN model parameters on the local training dataset of the sampled clients

[0151] Step Ⅳ: Each selected client performs local computing based on the global state and its local dataset and sends updates to the server. The server, through the parameter aggregation mechanism, weights and averages the parameters of each TCN-GCN model. And updates are applied to its global state, repeating until a new global model is synthesized, achieving spatio-temporal prediction based on federated learning.

[0152] Through the double sampling of devices and clients in the present invention, the amount of parameters for model update and upload is greatly reduced. For clients that are not sampled, their model parameters are continuously updated with the global model updated by the server. In model aggregation, the parallel computing of multiple clients uploading in each round, the parameters uploaded by the previous-round clients are released after achieving global update. Through simulation verification on a public dataset, the results show that on the basis of protecting privacy, the FedTCN-GCN proposed in the present invention has good prediction performance.

[0153] To verify that the FedTCN-GCN prediction framework proposed in the present invention can effectively achieve privacy-preserving distributed prediction, this embodiment selects a sensor array sensing dataset for simulation verification. The dataset records the time series obtained by the sensors and the measured values of CO concentration, humidity, and temperature in the air chamber, with data recorded every 5 seconds. A total of 28,800 time steps of samples collected by 14 sensors are used for training.

[0154] To avoid numerical problems for gradient update and accelerate finding the optimal solution, it is necessary to normalize the data and scale different types of data to the same range [0, 1] proportionally. After normalizing the dataset, the training set, test set, and validation set are divided into 80%, 10%, and 10% respectively.

[0155] In the model setting, the time step is set to 7, the number of epochs is set to 5, the batch_size is set to 16, the learning rate of the batch gradient descent algorithm is set to 0.0001, and the L2 regularization penalty is 10-4. The number of clients N is set to 100, and the default sampling ratio β of clients is set to 0.5, that is, 28,800 * 0.8 groups of samples are divided into 100 parts to simulate the local datasets respectively owned by 100 clients. Each time the model parameters updated by the clients are uploaded to the server, 50 users are randomly selected for average weighting. Since the dataset has the values monitored by 14 sensors, the number of devices K is 14, and the default sampling number of devices k in this embodiment is set to 3.

[0156] This embodiment uses six evaluation metrics to evaluate the performance of the model: Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Accuracy, Coefficient of Determination (R 2 )、Explained Variance Score (Var)

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164] Among them, y ij and represent the true value and predicted value of the i-th node at time j, M represents the time length, and N represents the number of devices. Y and represent the sets of y ij and respectively, which are two-dimensional arrays of size M×N, represents the average value of Y.

[0165] MSE and RMSE are obtained by calculating the values of individual nodes and then summing them to find the average. RMSE and MAE are used to measure the prediction error. The smaller the value, the better the prediction effect. ||Y|| F is the Frobenius norm of the matrix, defined as the square root of the sum of the absolute values squared of the elements of the matrix. The larger the value of the accuracy Accuracy defined by the F-norm, the better the prediction effect. R 2 and Var are used to calculate the correlation coefficient. The larger the value, the stronger the ability of the prediction result to represent the actual data.

[0166] The FedTCN-GCN prediction algorithm proposed by the present invention is compared with three traditional well-performing spatio-temporal prediction algorithms:

[0167] 1) CNN-based spatio-temporal prediction algorithm: ConvLSTM and CNN-LSTM realize the feature extraction of spatio-temporal data through the reconstruction and combination of CNN and LSTM.

[0168] 2) GCN-based spatio-temporal prediction algorithm: GCN and T-GCN capture the topological relationship between multiple devices through a predefined static graph adjacency matrix to achieve good prediction of spatial and temporal data.

[0169] 3) Residual-based spatio-temporal data prediction algorithm: The ResLSTM algorithm effectively learns time series by combining residual networks and LSTM. Graph-ResLSTM adds a graph learning module on this basis to achieve deep mining of device topological relationships. The two effectively avoid the overfitting problem of deep learning models by combining residual networks.

[0170] The prediction accuracy of TCN-GCN before and after federated learning is compared with the existing cutting-edge space-time prediction algorithms. The comparison of various evaluation indicators is shown in Table 1:

[0171] Table 1

[0172]

[0173]

[0174] It can be seen that the TCN-GCN algorithm has good MAE, RMSE, and R 2 The best results were obtained in both the MSE and Var indicators. The differences between the MSE and Accuracy indicators and those of CNN-LSTM were only 0.10137 and 0.000381, respectively, and the prediction performance did not decrease significantly.

[0175] Compared with the ConvLSTM and CNN-LSTM algorithms, the TCN-GCN algorithm solves the defect that the CNN network is only suitable for extracting Euclidean structure data features. It uses GCN to learn the topological structure of multiple devices and effectively achieves accurate prediction of different devices by obtaining neighborhood information.

[0176] Compared with the GCN and T-GCN algorithms, the TCN module used by the TCN-GCN algorithm can effectively capture the temporal correlation in the sequence, thereby improving the prediction effect.

[0177] Compared with ResLSTM and Graph-ResLSTM, the TCN-GCN algorithm adopts an adaptive graph learning process, which overcomes the limitations caused by the pre-defined adjacency matrix in the Graph-ResLSTM algorithm and fully utilizes the complementary relationship between devices by dynamically searching for the optimal graph adjacency matrix.

[0178] For the FedTCN-GCN algorithm, when only 50% of the clients are selected for sampling, it outperforms the GCN, T-GCN, ResLSTM, and Graph-ResLSTM algorithms in all metrics and has good prediction performance.

[0179] Compared with the ConvLSTM algorithm and the CNN-LSTM algorithm, it only differs by 0.00693 and 0.01873 in the Accuracy metric, and only by 0.00261 and 0.00943 in the Var metric. Therefore, the FedTCN-GCN algorithm proposed in the present invention has good capture ability for time dependence and spatial correlation.

[0180] As Figure 5 shown, it shows the prediction effects of the TCN-GCN algorithm on Device 1 and Device 14 before and after federated learning. It can be seen that the prediction effect of the FedTCN-GCN algorithm of the present invention is very close to that of the TCN-GCN algorithm. The backbone network predicted by the FedTCN-GCN algorithm is the TCN-GCN structure. Both of them develop a TCN model to learn the long-term and short-term features of the time series, and adopt a GCN model to capture the spatial relationship of device nodes to improve the prediction performance of time series data. Although the sensor monitoring data changes continuously over time due to environmental impacts, by combining TCN and GCN, the TCN-GCN prediction algorithm can fully exploit the spatio-temporal characteristics. In each communication round of the FedTCN-GCN algorithm, the sampled clients parallelly train a model with the TCN-GCN structure, and upload the updated model parameters for weighted averaging, so that each client finally obtains optimized model parameters.

[0181] In Figure 5 (b), there will be slightly more spikes in the prediction curve of FedTCN-GCN for Device 14, and its performance is slightly worse than that of the TCN-GCN algorithm. This is because each client of the FedTCN-GCN algorithm only trains on 1 / 100 of the dataset locally. Therefore, the FedTCN-GCN algorithm proposed in the present invention effectively protects local data security while ensuring prediction performance.

[0182] To analyze the convergence of the TCN-GCN algorithm and the FedTCN-GCN algorithm during model training, Figure 6The training losses of the TCN-GCN algorithm and the FedTCN-GCN algorithm before and after federated learning are compared. At the beginning of training, the training loss of the FedTCN-GCN algorithm is significantly greater than that of the TCN-GCN algorithm because the training dataset is scattered among different users. As the number of batches increases, due to the continuous update of the model parameters trained on the local datasets of different clients and the realization of the aggregation of model parameters, the loss of the FedTCN-GCN algorithm gradually approaches that of the TCN-GCN algorithm. At the same time, it can be seen from Figure 6 that the FedTCN-GCN algorithm has good convergence and stability. Therefore, the FedTCN-GCN algorithm can achieve high-precision spatio-temporal data prediction while effectively protecting privacy.

[0183] To reduce the communication overhead, the FedTCN-GCN prediction algorithm proposed in the present invention samples K devices and adaptively learns the graph adjacency matrix of k relevant device nodes. To evaluate the impact of the number of sampled devices on the model performance, the number of sampled devices is selected from 3 to 13, and the performance comparisons of MAE and RMSE are as Figure 7 shown.

[0184] As the number of sampled devices increases, the MAE and RMSE curves of the FedTCN-GCN algorithm tend to be stable, and the change in prediction accuracy is very small. This is because in the prediction of multiple spatially correlated devices, the monitored value of a certain device K1 is related to a limited number of neighboring devices, and the monitored value of device K1 can be mined through the monitored values of neighboring devices. In addition, as the number of devices increases, mining the spatial relationships of relevant devices will also introduce irrelevant features and noises into the model.

[0185] Therefore, sampling a small number of devices will not affect the accuracy of the model. To reduce the communication overhead during the federated learning process, the present invention only selects the three most relevant devices for learning among multiple devices in model training to ensure the prediction accuracy. The impact of different client sampling ratios (i.e., β = 0.1, 0.3, 0.5, 0.7, 1) on the performance of the FedTCN-GCN algorithm is shown in Table 2, and the total number of clients is 100.

[0186] Table 2

[0187] Customer ratio MAE MSE RMSE Accuracy R2 Var 0.1 2.76476 38.13665 6.03257 0.80961 0.88743 0.92345 0.3 2.02676 23.61766 4.74250 0.85017 0.93534 0.95048 0.5 1.80213 20.04014 4.35404 0.86199 0.94431 0.95708 0.7 1.81874 19.08044 4.24116 0.86533 0.94628 0.95987 1 1.64665 17.52992 4.05548 0.87092 0.95170 0.96203

[0188] It can be seen that the client sampling ratio has a great impact on the performance of the FedTCN-GCN algorithm. The higher the client sampling ratio, the higher the prediction accuracy of the model. Specifically, when the client sampling ratio is 1, the MAE is reduced by 9.4%, the MSE is reduced by 14.3%, and the RMSE is reduced by 7.4% compared with when it is 0.5; when the client sampling ratio is 0.5, the MAE is reduced by 53.3%, the MSE is reduced by 90.3%, and the RMSE is reduced by 38.6% compared with when it is 0.1. The reason is that for spatio-temporal data, in the communication process between the client and the server in each round, increasing the client sampling ratio will increase the proportion of the original data used, providing more model parameters of the client for the server. By increasing the proportion of the client uploading model parameters, the cloud can simultaneously aggregate the gradient information of multiple clients, thereby improving the accuracy of the global model.

[0189] In addition, in real life, the scale of clients participating in spatio-temporal prediction is large. FedTCN-GCN effectively ensures that when some clients have communication failures or cannot upload gradient information, the global model can still be effectively shared by sampling the clients.

Claims

1. Space-time prediction algorithm based on federated learning on industrial Internet of Things edge devices, characterized in that It includes the following steps: First, build an application scenario including an industry set and a customer set; each industry has I clients, and each client has K sensor devices; for a single client C i , the sensing data monitored by its K devices constitutes a database D i , on the client C i , build a TCN-GCN deep model and use the data in the database D i to update the parameters w of the TCN-GCN deep model C ; Similarly, each client locally trains its corresponding TCN-GCN deep model respectively to obtain its own model parameters w C ; w C It contains the spatio-temporal features mined by the TCN-GCN deep model from the original data; Finally, use the improved FedAVG algorithm to upload the TCN-GCN deep model parameters w of each client C , to the cloud for weighted aggregation to form a new global model, realizing spatio-temporal prediction based on federated learning; The TCN-GCN deep model is formed by alternately combining the dilated temporal convolutional network DTCN and the dynamic graph convolutional network DGCN; the specific update process is as follows: Step 201, Client C i Use K sensors to collect two-dimensional data containing time steps and spatial nodes respectively, and construct an initial graph adjacency matrix to represent the correlation relationship of spatial nodes; Step 202: Send the two-dimensional data into the TCN-GCN deep model in sequence, and extract temporal and spatial features by combining with the graph adjacency matrix; Specifically: First, given the input two-dimensional data x in , the output x out after passing through DTCN is as follows: x out = tanh(f1 * x in ) × sigmoid(f2 * x in ) Among them, f1 represents the filtering convolutional function, f2 represents the gated convolutional function, sigmoid(·) represents the S-shaped activation function, and tanh(·) represents the hyperbolic tangent activation function; Then, send the features captured by DTCN into the DGCN module, and its information transfer layer is: Among them, H l represents the propagation layer of the l-th layer; σ1 and σ2 are different activation functions, is the dynamic graph adjacency matrix obtained by random dynamic sampling, and W l-1 represents the network weight of the (l - 1)-th layer; represents the propagation layer after skip connection; β is a hyperparameter for controlling the ratio of retaining the original state of the root node; H (l) is the l-th layer propagation layer whose node state is continuously updated as the depth of graph convolution increases, and H (l - 1) represents the propagation layer of the previously retained node state; H out is the output layer after the skip layers are stacked; Step 203: Use the TCN-GCN model to update the association of each spatial node in the graph adjacency matrix; Step 204: Perform a convolutional operation on the updated graph adjacency matrix and the extracted spatial feature vectors to continuously update the spatial features of the moving nodes; Step 205: Achieve high-precision prediction of the two-dimensional data according to the captured temporal and spatial correlations; The specific prediction process is as follows: The two-dimensional data X input at the time step t is expressed as: X = {z1[i], z2[i], ···, z t [i]} where z t [i] represents the value of the i-th sensor at time step t, where i ∈ K; Then the predicted value at the next time step is expressed as: Y = {z t+1 [i]} Step 206: By continuously fitting the predicted value with the label value of the real data, the client C is obtained i The parameters w of the locally updated TCN-GCN deep model C .

2. The spatio-temporal prediction algorithm based on federated learning on the industrial Internet of Things edge device according to claim 1, characterized in that The improved FedAVG algorithm includes the following steps: Step I: At the beginning of each round, the server where industry O is located selects volunteers from the clients to participate in this round of training, and broadcasts the globally initialized model parameters of TCN-GCN to the selected clients; Step II: Randomly sample the participating clients and all sensor devices included in the clients; For each randomly sampled client, the model parameter F that minimizes its loss l (w) is as follows: where \(w\in R\) d are the local model parameters of each sensor device, \(h(\cdot)\) is the regularization term, and \(D\) k represents that the client needs to learn \(k\) samples after device sampling in the local model; On the cloud side, the average model parameter F obtained after the client sampling of the global prediction model g (w) is expressed as n = β × N Among them, N is the total number of all clients included in the industry, and β represents the sampling ratio of the clients; Therefore, after double sampling of the clients and devices, the model parameters obtained by the clients after weighted averaging by the central cloud are: The sampling of the devices on the client side is reflected in the node sampling of the graph adjacency matrix, specifically: For a client with K devices, the mathematically expressed randomly initialized graph adjacency matrix A is: Among them, A ij represents the relationship between the i-th and j-th nodes in the graph adjacency matrix, is the set of vertices, (v i , v j ) represents the relationship of a pair of nodes, and the edge connecting the vertices is c represents a non-zero real number; To reduce the communication overhead, sample the K devices of each client, and the node index idx after taking the first k relevant devices is expressed as idx = argtopk(A[i,:]) Among them, argtopk(·) returns the indices of the first k maximum values of the vector; for each device, select its first k nearest edge computing nodes as its neighbors, that is, filter out the first k most relevant subsets among multiple devices; A[i,:] represents the node relationship between the i-th node and all other nodes in the graph adjacency matrix; Step III. Initialize the number of rounds E and the number of training batches B, and update the TCN-GCN model parameters for the sampled clients on their local training datasets. Step IV. The server uses a parameter aggregation mechanism to sample the parameters of each TCN-GCN model by weighted average and updates and applies them to its global state, repeating until a new global model is synthesized, thereby achieving spatio-temporal prediction based on federated learning.

Citation Information

Patent Citations

  • Graph convolutional networks with motif-based attention

    US20200285944A1

  • A negative differential resistance based memory

    WO2019066821A1

Cited By

  • Memory dynamic adjustment method and system, electronic equipment and storage medium

    CN121187767A

  • A memory dynamic adjustment method and system, electronic equipment and storage medium

    CN121187767B