A Digital Twin-Assisted Federated Learning Freshness Optimization Method and System

By building a digital twin model in the industrial Internet of Things and optimizing bandwidth and data frequency, the federated learning delay and reverse optimization problems are solved, model performance and learning efficiency are improved, and data security and privacy protection are ensured.

CN115481748BActive Publication Date: 2025-07-04GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211051083.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-04
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The prior art has high latency and is prone to reverse optimization when training federated learning models in the industrial Internet of Things, resulting in a degradation of model performance and failing to effectively consider the data reference value and the impact of communication environment.

Method used

Build an industrial IoT federated learning model, generate a digital twin for each smart device, calculate energy consumption, data freshness and model parameter freshness, optimize bandwidth allocation and data frequency through deep reinforcement learning networks, establish Markov decision-making process, and obtain the optimal scheduling strategy.

Benefits of technology

It reduces the delay of federated learning, improves learning efficiency, eliminates reverse optimization, enhances data security and privacy protection, and realizes real-time low-power consumption and high-quality services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481748B_ABST
    Figure CN115481748B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for optimizing freshness in federated learning assisted by digital twins, which relates to the technical field of industrial Internet of Things. The method includes constructing a federated learning model for industrial Internet of Things, calculating the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning; taking the minimization of the sum of energy consumption and freshness as the goal, establishing an optimization problem for joint bandwidth allocation, data collection frequency, and data calculation frequency, and transforming it into a Markov decision process, defining the state space, action space, and reward function; establishing a deep reinforcement learning network and training it, using the trained deep reinforcement learning network for resource scheduling to obtain the optimal scheduling strategy and applying it to the corresponding intelligent devices. The present invention effectively reduces the latency of federated learning, improves the learning efficiency of federated learning, eliminates the occurrence of reverse optimization, and improves the performance of the federated learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial Internet of Things, and more specifically, to a freshness optimization method and system for federated learning assisted by digital twin. Background Art

[0002] Industry is an important field for the application of the Internet of Things. With the continuous development of science and technology, the industrial Internet of Things has become synonymous with intelligent manufacturing and Industry 4.0. Various advanced intelligent technologies are constantly integrated into all aspects of industrial production, such as artificial intelligence (AI), machine learning, augmented / virtual reality (AR / VR), digital twin / thread, cloud / edge computing and other intelligent technologies. Industrial Internet of Things intelligent devices need to perform federated learning with other intelligent devices to improve their own model performance. However, considering the privacy of data, users are reluctant to provide personal data. However, it is not easy for users to perceive that intelligent devices collect data without their consent and upload it to the server. Therefore, at present, a "data island" phenomenon has formed in the Internet, and data from all parties cannot be directly shared or exchanged. The goal of federated learning is to achieve joint modeling on the basis of ensuring data privacy, security, legality and compliance. Under the coordination of a central server or service provider, each participating party transmits and aggregates specific intermediate operation results to achieve the purpose of machine learning model training. No raw data is exchanged or transmitted between participating parties, ensuring the security of local private data. Federated learning can combine the experiences of multiple participating parties to optimize a common model. Analyzing from the sample size of data, this model is necessarily better than the models trained by any single participating party. However, at the same time, the reference value of data also greatly affects the performance of the model. In addition to considering the freshness of local samples, federated learning also needs to consider the impact of the user communication environment on the model upload delay. A poor communication environment will cause users to drop out. After a user drops out, they cannot upload their model. As time goes by, this model also loses its value. Participating in model training with outdated data will only make the model performance worse and cannot give full play to the advantage of the data volume. Correctly considering the freshness of the model can greatly improve the performance and utilization value of the final model, eliminate the potential reverse optimization risk of federated learning, and give full play to the advantages of federated learning. The rapid development of communication technology has made the interaction delay between industrial Internet of Things intelligent devices extremely low, making it possible to apply a series of delay-sensitive technologies in the Internet of Things. Digital twin technology can construct a digital twin in the digital world that is highly similar to the physical entity; this digital twin can simulate the actions and movement change laws of the physical entity in the real world. The digital twin interacts with the physical entity in real time, updates the state and environment in real time, has a high degree of fidelity, and at the same time, the simulation and action simulation speculation on the digital twin can also act on the physical entity.

[0003] The prior art provides a federated learning method and system based on edge digital twin association, including that users participating in federated learning respectively generate digital twins; using a many-to-one matching algorithm to pair the digital twins with edge servers; the server constructs tasks and publishes them to the edge servers, and the digital twins utilize the resources of the edge servers paired with them to perform federated learning, and transmit the models obtained by their local iterative training to the edge servers; the edge servers aggregate the models obtained by training the digital twins paired with them to obtain local models, and upload the local models to the server; the server aggregates the local models received by it to obtain a global model, checks whether the accuracy of the global model reaches a preset threshold, if not, then distributes the current global model to the edge servers for training; if so, then the current federated learning is completed. Although this method solves the defect of high latency during the training of the federated learning model to a certain extent, it does not consider the reference value of the training data, and there is a high risk of reverse optimization, resulting in worse model performance. Summary of the Invention

[0004] To overcome the above-mentioned defects of high latency and easy reverse optimization in the prior art during the training of the federated learning model, the present invention provides a method and system for optimizing the freshness of federated learning assisted by digital twins, which can reduce the time delay of federated learning, improve the learning efficiency of federated learning, eliminate the occurrence of reverse optimization, and improve the performance of the federated learning model.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] The present invention provides a method for optimizing the freshness of federated learning assisted by digital twins, including:

[0007] S1: Construct an industrial Internet of Things federated learning model, the model includes a central server and several intelligent devices, and generate digital twins corresponding to each intelligent device, and store them in the central server;

[0008] S2: According to the industrial Internet of Things federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all the digital twins of the intelligent devices for one round of federated learning;

[0009] S3: Taking the minimization of the sum of the energy consumption, local data freshness, and model parameter freshness of all the digital twins of the intelligent devices for one round of federated learning as the goal, establish an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency;

[0010] S4: Convert the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the industrial Internet of Things federated learning model;

[0011] S5: Establish a deep reinforcement learning network based on the Proximal Policy Optimization (PPO) algorithm, and train the deep reinforcement learning network using the state space, action space, and reward function to obtain a trained deep reinforcement learning network;

[0012] S6: Use the trained deep reinforcement learning network for resource scheduling to obtain an optimal scheduling policy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device;

[0013] S7: Apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

[0014] The present invention first constructs an industrial Internet of Things (IIoT) federated learning model, which includes a central server and several intelligent devices. Each intelligent device generates a corresponding digital twin and stores it in the central server. According to the IIoT federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all the digital twins of the intelligent devices for one round of federated learning, and establish an optimization problem for joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the three. By introducing local data freshness and model parameter freshness, reasonably adjust the allocated bandwidth, maximize the freshness of the local model, reduce the straggler probability of intelligent devices, and thus effectively reduce the latency of federated learning, and finally improve the performance of the global model. Generate corresponding digital twins for the intelligent devices in the real space in the virtual space, and the operation modes of the two are completely synchronized. By establishing and optimizing a deep reinforcement learning network, obtain the optimal scheduling policy on the digital twins and then apply it to the corresponding intelligent devices in the real space, which not only improves the learning efficiency but also reduces the physical resource consumption of intelligent devices. The Proximal Policy Optimization algorithm can adapt to and eliminate the noise interference of the virtual-real mapping of digital twins during the network optimization process, ensuring the correctness of the optimization direction and results. At the same time, it ensures the privacy and security of each intelligent device, enhances the usability and security of data, reduces the time delay and energy consumption of intelligent devices participating in federated learning, and realizes real-time, low-power, and high-quality services.

[0015] Preferably, in the step S1, the constructed industrial Internet of Things federated learning model is specifically:

[0016] The industrial Internet of Things federated learning model includes a central server and K intelligent devices. Each intelligent device generates a corresponding digital twin, and a total of K digital twins are stored in the central server. The central server sends a federated learning request to the digital twins. After receiving the federated learning request, the digital twins start federated learning, collect data and calculate, train the local model, and update the trained local model to the central server. The central server aggregates the received local models into a global model;

[0017] Set the maximum bandwidth of the central server to B m , and the bandwidth allocated to the k-th digital twin is B k , and the data collection frequency of the k-th digital twin is The data calculation frequency of the k-th digital twin is k = 1, 2, …, K.

[0018] There are mainly two factors affecting the delay and energy consumption of a round of federated learning. One is the delay from when the intelligent device receives the federated learning request to when it sends the local model to the central server, and the other is the delay from the start of uploading to the completion of uploading by the intelligent device. In federated learning, the training of intelligent devices requires a large number of data samples. Stale data has little effect on the training of intelligent devices and may even reduce the final performance of intelligent devices. In addition to considering the freshness of local data, the present invention also considers the impact of the communication environment on the upload delay of the local model: a poor communication environment will cause the intelligent device to fall behind. After falling behind, it cannot upload the local model. As time goes by, this local model also loses its value, that is, the freshness of model parameters. During the process of federated learning, intelligent devices achieve global model download and local model upload through wireless communication. In addition to maintaining and updating the global model, the central server also needs to allocate appropriate bandwidth to each intelligent device and guide it to adjust its own data collection frequency and calculation frequency; the computing resources of the central server are much more powerful than those of intelligent devices. The digital twins of intelligent devices are all stored in the central server. By taking advantage of the fact that making decisions in the virtual space does not consume physical entity computing and communication resources, the scheduling strategy is executed in the virtual space, which not only saves the computing time and energy consumption of intelligent devices during the execution of the scheduling strategy, but also eliminates the long delay caused by the possible falling-behind problem during the communication process.

[0019] Preferably, in the step S2, the specific method for calculating the energy consumption of all the digital twins of intelligent devices for a round of federated learning according to the industrial Internet of Things federated learning model is as follows:

[0020] The number of floating-point operations per CPU cycle of the k-th digital twin is C k , and the number of floating-point operations required to collect a set of data is Then the time taken to collect a set of data is:

[0021]

[0022] In the formula, represents the time taken for the k-th digital twin to collect a set of data;

[0023] The time interval from when the k-th digital twin completes one round of federated learning to receiving the next federated learning request is Then the energy consumption of the data collection process is:

[0024]

[0025] In the formula, represents the energy consumption of the data collection process of the k-th digital twin, and P k represents the energy consumption per unit time of the k-th digital twin;

[0026] The number of floating-point operations required for the k-th digital twin to complete the calculation and update of a set of collected data is Then the local latency for one round of data calculation is:

[0027]

[0028] In the formula, represents the local latency of the k-th digital twin for one round of data calculation;

[0029] The upload time for the k-th digital twin to upload the trained local model to the central server is:

[0030]

[0031] In the formula, D represents the data volume of the trained local model, and τ k represents the signal-to-noise ratio of the k-th digital twin to the central server;

[0032] Then the energy consumption of all digital twins of intelligent devices for one round of federated learning is:

[0033]

[0034] In the formula, E represents the energy consumption of all digital twins of intelligent devices for one round of federated learning.

[0035] Preferably, in step S2, the specific method for calculating the local data freshness according to the industrial Internet of Things federated learning model is:

[0036] It is set that the digital twin of the intelligent device needs to collect N groups of data for training the local model, and the moment when the digital twin starts training the local model after receiving the federated learning request is recorded as TrainTime;

[0037] When the digital twin is in an idle state when receiving the federated learning request, the data freshness is:

[0038]

[0039] In the formula, represents the data freshness of the n-th group of data of the k-th digital twin in the idle state, Denote the saving time of the n-th group of data of the k-th digital twin.

[0040] If the digital twin is collecting the i-th group of data when receiving a federated learning request, the data consists of the i-th group of data and the previous N - i groups of data, and the data freshness is:

[0041]

[0042] where i < N.

[0043] Preferably, in the step S2, according to the industrial Internet of Things federated learning model, the specific method for calculating the freshness of model parameters is:

[0044] Set the waiting time threshold T of the central server astrict , if and only if , the trained local model uploaded by the k-th digital twin is received by the central server, otherwise it is regarded as the digital twin falling behind; all K digital twins complete the upload of the trained local model parameters within T astrict , and the freshness of the model parameters at the completion time of upload is:

[0045]

[0046] where denotes the freshness of the model parameters at the completion time of upload of the m-th trained local model, denotes the upload time of the m-th trained local model to the central server;

[0047] Denote the time when the central server saves the model parameters of the m-th trained local model as t sc (t), and denote the time when the central server starts global model aggregation as AggregateTime, then the freshness of the model parameters at the aggregation time is:

[0048]

[0049] where denotes the freshness of the model parameters at the aggregation time of the m-th trained local model.

[0050] Preferably, in the step S3, aiming at minimizing the sum of the energy consumption, local data freshness and model parameter freshness of all digital twins of intelligent devices in one round of federated learning, an optimization problem of joint bandwidth allocation, data collection frequency and data calculation frequency is established, specifically:

[0051] Calculate the total freshness A according to the local data freshness and the model parameter freshness:

[0052]

[0053] Then the optimization problem of jointly allocating bandwidth, data collection frequency, and data calculation frequency is expressed as:

[0054]

[0055]

[0056]

[0057]

[0058]

[0059] C5: λ + μ = 1

[0060] In the formula, M agg represents the number of digital twins participating in federated learning, and M upd represents the number of digital twins that have successfully uploaded their local models within T astrict , represents the maximum available calculation frequency of the kth digital twin, λ represents the freshness weight, and μ represents the energy consumption weight.

[0061] Preferably, in the step S4, the state space, action space, and reward function defined for the industrial Internet of Things federated learning model are specifically as follows:

[0062] In the industrial Internet of Things federated learning model, the central server observes the current state s of the digital twin from the environment t to form the state space. The current state s t corresponds to the current action a t . The digital twin of the intelligent device executes the current action a in the action space t , interacts with the central server, and returns the current reward r t and the new state s t+1 ;

[0063] In the state space, the current state expression is s t = {the amount of data that still needs to be collected by the digital twin of each intelligent device when it receives a federated learning request, the amount of data collected by the digital twin of each intelligent device during the gap between the latest two rounds of federated learning requests, the time interval between two rounds of federated learning requests, the number of stragglers of the digital twin of the intelligent device in each round of federated learning};

[0064] In the action space, the current action expression

[0065] The current reward r tEqual to executing the current action a t Between the time after executing the current action a and the next round of federated learning request is sent, the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning, that is, r t = λA + μE.

[0066] Preferably, in the step S5, the established deep reinforcement learning network includes an experience buffer, an Actor network, a Critic network, and an oldActor network;

[0067] At each moment, the input of the Actor network is the current state s t , and it generates and outputs the current action a t = μ θ (s t ); The Critic network evaluates the value of the current action a t according to the current state s t ; After the digital twin executes the current action a t , a new state s t+1 and the current reward r t are generated, and [s t , a t , r t , s t+1 is stored in the experience buffer; The network parameters of the oldActor network are updated at intervals. Every few rounds of federated learning, the network parameters θ of the Actor network are copied as the network parameters θ of the oldActor network * , and the output of the oldActor network is used to compare the change amplitude of the current action a t ;

[0068] The loss function for updating the network parameters θ of the Actor network is:

[0069]

[0070] In the formula, represents the current action output by the oldActor network, and Adv t (s t , a t ) represents the advantage function, represents restricting the upper and lower bounds of to [1 - ∈, 1 + ∈], where ∈ represents the upper and lower bound limit value, and ∈ ∈ (0, 1);

[0071] The loss function for updating the network parameters θ of the Critic network is: c

[0072] L(θ c ) = MSE(v t , r t )

[0073] wherein, MSE(v t , r t ) represents the mean square error function;

[0074] Using the policy gradient method to iteratively train the Actor network and the Critic network, when the loss function value converges, the optimal network parameters θ, θ c and θ * are obtained, and a trained deep reinforcement learning network is obtained.

[0075] Preferably, the specific method of step S6 is as follows:

[0076] The central server observes the current states of the digital twins of all current intelligent devices, inputs them into the trained deep reinforcement learning network, and generates the current actions; the in the current actions is used as the optimal scheduling policy, and the included in the current actions B k is used as the optimal data collection frequency, the optimal data calculation frequency, and the allocated optimal bandwidth for the digital twin of each intelligent device.

[0077] The present invention also provides a digital twin-assisted federated learning freshness optimization system. Based on the above digital twin-assisted federated learning freshness optimization method, the system includes:

[0078] A model construction module for constructing an industrial Internet of Things federated learning model, the model includes a central server and several intelligent devices, and generates digital twins corresponding to each intelligent device, which are stored in the central server;

[0079] A calculation module for calculating the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning according to the industrial Internet of Things federated learning model;

[0080] An optimization problem establishment module for establishing an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning;

[0081] An optimization problem transformation module for transforming the optimization problem into a Markov decision process, and defining the state space, action space, and reward function of the industrial Internet of Things federated learning model;

[0082] A network construction training module is used to establish a deep reinforcement learning network based on the Proximal Policy Optimization (PPO) algorithm, and train the deep reinforcement learning network using the state space, action space, and reward function to obtain a trained deep reinforcement learning network;

[0083] A resource scheduling module uses the trained deep reinforcement learning network for resource scheduling to obtain an optimal scheduling policy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device;

[0084] A scheduling application module is used to apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

[0085] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0086] The present invention first constructs an industrial Internet of Things federated learning model, including a central server and several intelligent devices. Each intelligent device generates a corresponding digital twin and stores it in the central server. According to the industrial Internet of Things federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all the digital twins of the intelligent devices for one round of federated learning, and establish an optimization problem for joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the three. By introducing local data freshness and model parameter freshness, reasonably adjust the allocated bandwidth, maximize the freshness of the local model, reduce the straggler probability of intelligent devices, and thus effectively reduce the delay of federated learning, and finally improve the performance of the global model. Generate corresponding digital twins for the intelligent devices in the real space in the virtual space, and their operation modes are completely synchronized. By establishing a deep reinforcement learning network and optimizing it, obtain the optimal scheduling policy on the digital twin and then apply it to the corresponding intelligent devices in the real space, which not only improves the learning efficiency but also reduces the physical resource consumption of intelligent devices. The Proximal Policy Optimization algorithm can adapt to and eliminate the noise interference of the virtual-real mapping of digital twins during the network optimization process, ensuring the correctness of the optimization direction and results. At the same time, it ensures the privacy and security of each intelligent device, enhances the usability and security of data, reduces the time delay and energy consumption of intelligent devices participating in federated learning, and realizes real-time, low-power, and high-quality services. Description of the Drawings

[0087] Figure 1 It is a flowchart of a method for optimizing freshness of federated learning assisted by digital twins according to Embodiment 1.

[0088] Figure 2 It is a schematic diagram of the industrial Internet of Things federated learning model according to Embodiment 1.

[0089] Figure 3 Schematic diagram of the deep reinforcement learning network described in Embodiment 2.

[0090] Figure 4 Schematic diagram of the structure of a digital twin-assisted federated learning freshness optimization system described in Embodiment 3. Detailed implementation manners

[0091] The drawings are only for illustrative purposes and should not be construed as limitations on this patent.

[0092] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product.

[0093] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0094] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0095] Embodiment 1

[0096] This embodiment provides a digital twin-assisted federated learning freshness optimization method, as Figure 1 shown, including:

[0097] S1: Construct an industrial Internet of Things federated learning model, the model includes a central server and several intelligent devices, and generate digital twins corresponding to each intelligent device, and store them in the central server;

[0098] S2: According to the industrial Internet of Things federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning;

[0099] S3: With the goal of minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning, establish an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency;

[0100] S4: Convert the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the industrial Internet of Things federated learning model;

[0101] S5: Establish a deep reinforcement learning network based on the proximal policy optimization algorithm, and use the state space, action space, and reward function to train the deep reinforcement learning network to obtain a trained deep reinforcement learning network;

[0102] S6: Use the trained deep reinforcement learning network for resource scheduling to obtain the optimal scheduling strategy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device;

[0103] S7: Apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

[0104] In the specific implementation process, this embodiment first constructs an industrial Internet of Things federated learning model. As Figure 2 shown, it includes a central server and several intelligent devices. Each intelligent device generates a corresponding digital twin and stores it in the central server. According to the industrial Internet of Things federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all the digital twins of intelligent devices for one round of federated learning, and establish an optimization problem for joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the three. By introducing local data freshness and model parameter freshness, reasonably adjust the allocated bandwidth, maximize the freshness of the local model, reduce the straggler probability of intelligent devices, thereby effectively reducing the latency of federated learning, and ultimately improving the performance of the global model. Generate corresponding digital twins for the intelligent devices in the real space in the virtual space, and their operation modes are completely synchronized. By establishing and optimizing a deep reinforcement learning network, obtain the optimal scheduling strategy on the digital twin, and then apply it to the corresponding intelligent device in the real space, which not only improves the learning efficiency but also reduces the physical resource consumption of intelligent devices. The proximal policy optimization algorithm can adapt to and eliminate the noise interference of the virtual-real mapping of digital twins during the network optimization process, ensuring the correctness of the optimization direction and results. At the same time, it ensures the privacy security of each intelligent device, enhances the usability and security of data, reduces the time delay and energy consumption of intelligent devices participating in federated learning, and realizes real-time, low-power, and high-quality services.

[0105] Embodiment 2

[0106] This embodiment provides a method for optimizing the freshness of federated learning assisted by digital twins, including:

[0107] S1: Construct an industrial Internet of Things federated learning model, the model includes a central server and several intelligent devices, and generate corresponding digital twins for each intelligent device, and store them in the central server; specifically:

[0108] The industrial Internet of Things (IIoT) federated learning model includes a central server and K intelligent devices. Each intelligent device generates a digital twin, and a total of K digital twins are generated and stored in the central server. The central server sends a federated learning request to the digital twins. After receiving the federated learning request, the digital twins start federated learning, collect data, calculate, train the local model, and update the trained local model to the central server. The central server aggregates the received local models into a global model.

[0109] Assume the maximum bandwidth of the central server is B m , and the bandwidth allocated to the k-th digital twin is B k , and the data collection frequency of the k-th digital twin is The data calculation frequency of the k-th digital twin is k = 1, 2, …, K.

[0110] S2: According to the IIoT federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all the digital twins of the intelligent devices for one round of federated learning. Specifically:

[0111] The number of floating-point operations per CPU cycle of the k-th digital twin is C k , and the number of floating-point operations required to collect a set of data is Then the time taken to collect each set of data is:

[0112]

[0113] In the formula, represents the time taken for the k-th digital twin to collect each set of data;

[0114] The time interval from when the k-th digital twin completes one round of federated learning to when it receives the next federated learning request is Then the energy consumption during the data collection process is:

[0115]

[0116] In the formula, represents the energy consumption of the k-th digital twin during the data collection process, and P k represents the energy consumption per unit time of the k-th digital twin;

[0117] The number of floating-point operations required for the k-th digital twin to complete the calculation and update of a set of collected data is Then the local latency for one round of data calculation is:

[0118]

[0119] In the formula, Denote the local latency for the k-th digital twin to perform a round of data calculation;

[0120] The upload time for the k-th digital twin to upload the trained local model to the central server is:

[0121]

[0122] In the formula, D represents the data volume of the trained local model, and τ k denotes the signal-to-noise ratio from the k-th digital twin to the central server;

[0123] Then the energy consumption for all digital twins of intelligent devices to perform a round of federated learning is:

[0124]

[0125] In the formula, E represents the energy consumption for all digital twins of intelligent devices to perform a round of federated learning;

[0126] Assume that the digital twin of an intelligent device needs to collect N groups of data for local model training. The moment when the digital twin starts training the local model after receiving the federated learning request is denoted as TrainTime;

[0127] When the digital twin is in an idle state when receiving the federated learning request, the data freshness is:

[0128]

[0129] In the formula, denotes the data freshness of the n-th group of data of the k-th digital twin in the idle state, denotes the storage moment of the n-th group of data of the k-th digital twin;

[0130] When the digital twin is collecting the i-th group of data when receiving the federated learning request, the data consists of the i-th group of data and the previous N - i groups of data, and the data freshness is:

[0131]

[0132] In the formula, i < N;

[0133] Assume the waiting time threshold T of the central server astrict , and only when , the trained local model uploaded by the k-th digital twin is received by the central server, otherwise it is regarded as the digital twin falling behind; All K digital twins complete the upload of the parameters of the trained local model within T astrict , and the freshness of the model parameters at the completion of the upload moment is:

[0134]

[0135] wherein, represents the freshness of the model parameters at the moment when the m-th trained local model finishes uploading, represents the upload time when the m-th trained local model is uploaded to the central server;

[0136] The moment when the central server saves the model parameters of the m-th trained local model is denoted as t sc (t), and the moment when the central server starts global model aggregation is denoted as AggregateTime. Then, the freshness of the model parameters at the aggregation moment is:

[0137]

[0138] wherein, represents the freshness of the model parameters at the aggregation moment of the m-th trained local model;

[0139] S3: Taking the minimization of the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices in a round of federated learning as the goal, an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency is established; specifically:

[0140] Calculate the total freshness A based on the local data freshness and model parameter freshness:

[0141]

[0142] Then, the optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency is expressed as:

[0143]

[0144]

[0145]

[0146]

[0147]

[0148] C5: λ + μ = 1

[0149] wherein, M agg represents the number of digital twins participating in federated learning, M upd represents the number of digital twins that successfully upload their local models within T astrict , represents the maximum available calculation frequency of the k-th digital twin, λ represents the freshness weight, and μ represents the energy consumption weight.

[0150] S4: Transform the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the industrial Internet of Things federated learning model;

[0151] In the industrial Internet of Things federated learning model, the central server observes the current state s of the digital twin from the environment t to form the state space, and the current state s t corresponds to the current action a t . The digital twin of the intelligent device executes the current action a in the action space t , interacts with the central server, and returns the current reward r t and the new state s t+1 ;

[0152] In the state space, the current state expression is s t = {the amount of data that still needs to be collected when the digital twin of each intelligent device receives a federated learning request and is being executed, the amount of data collected by the digital twin of each intelligent device during the gap between the latest two rounds of federated learning requests, the time interval between two rounds of federated learning requests, the number of stragglers of the digital twin of the intelligent device in each round of federated learning};

[0153] In the action space, the current action expression

[0154] The current reward r t is equal to the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning between the execution of the current action a t and the issuance of the next round of federated learning requests, that is, r t = λA + μE.

[0155] S5: Establish a deep reinforcement learning network based on the proximal policy optimization algorithm, and train the deep reinforcement learning network using the state space, action space, and reward function to obtain a trained deep reinforcement learning network;

[0156] As Figure 3 shown, the established deep reinforcement learning network includes an experience buffer, an Actor network, a Critic network, and an oldActor network;

[0157] At each moment, the input of the Actor network is the current state s t , generates and outputs the current action a t = μ θ (s t ); The Critic network evaluates the value of the current action a t according to the current state s t ; The digital twin executes the current action a t and generates a new state s t+1 and the current reward r t , and stores [s t , a t , r t , s t+1 in the experience buffer; the network parameters of the oldActor network are updated at intervals. Every few rounds of federated learning, the network parameters θ of the Actor network are copied as the network parameters θ of the oldActor network * , and the output of the oldActor network is used to compare the change amplitude of the current action a t ;

[0158] The loss function for updating the network parameters θ of the Actor network is:

[0159]

[0160] In the formula, represents the current action output by the oldActor network, and Adv t (s t , a t ) represents the advantage function, represents restricting the upper and lower bounds of to [1 - ∈, 1 + ∈], where ∈ represents the upper and lower bound limit value, and ∈ ∈ (0, 1);

[0161] The loss function for updating the network parameters θ of the Critic network is: c

[0162] L(θ c ) = MSE(v t , r t )

[0163] In the formula, MSE(v t , r t ) represents the mean squared error function;

[0164] Using the policy gradient method to iteratively train the Actor network and the Critic network, when the loss function value converges, the optimal network parameters θ, θ c and θ * are obtained, and a trained deep reinforcement learning network is obtained.

[0165] S6: Use the trained deep reinforcement learning network for resource scheduling to obtain the optimal scheduling strategy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twins of each intelligent device; specifically: ​

[0166] The central server observes the current states of the digital twins of all current intelligent devices, inputs them into the trained deep reinforcement learning network, and generates the current actions; takes the current actions as the optimal scheduling strategy, where the current actions include B k as the optimal data collection frequency, optimal data calculation frequency, and allocated optimal bandwidth for the digital twin of each intelligent device.

[0167] S7: Apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

[0168] In the specific implementation process, first, an enterprise group is formed by enterprises with common optimization services and mutual promotion needs. These enterprises have common optimization goals, such as maximizing the fault prediction accuracy. These enterprises reach a consensus to select a central server (such as a third-party service provider, government agency, or group leader) to provide global model aggregation and broadcasting services. To ensure the performance and security of the global model, the central server will initiate federated learning tasks irregularly, and will also initiate FL tasks by the central server when individual enterprises have special needs. Taking a commercial brewery as an example, the increase in the distillate temperature and the decrease in the alcohol conversion rate in the mixture may be caused by the reduction of the cooling water flow or the blockage of the secondary distillation tray, or even may be due to related failures of the steam valve. The intelligent device managed by the k-th enterprise collects data at a frequency during the gap between two adjacent FL tasks, such as the increase in the distillate temperature and the decrease in the alcohol conversion rate corresponding to the reduction of the cooling water flow, the occurrence probability of the secondary distillation tray blockage fault and the operating duration of the distiller, ambient temperature and humidity and other state information when this fault occurs, the sensor data of each component before the steam valve fails and the accompanying faults, the sensor data of each component on the device when the operator cleans and maintains the device, etc. The occurrence of device failures is uncertain, and there may be a situation where no failure occurs during data collection. In this case, the sensor data of each component of the current device is used as a sample and classified as no failure. In this way, the enterprise's data samples may be sparse, and the need for federated learning is more obvious. The data collected by the enterprise is saved to the cache pool by the method proposed in the present invention. When receiving a federated learning task, the device responsible for data processing in the k-th enterprise adjusts the calculation frequency to and trains and generates a local model using the data in the cache pool on this basis. The bandwidth allocated by the central server to the k-th enterprise is B k , and the k-th enterprise will upload its local model under such communication conditions. The central server aggregates and broadcasts the global model according to the method proposed in this design. At this time, the PPO agent deployed by the central server will also make a decision a according to the current state space s t ​t and apply it to each enterprise involved.

[0169] Embodiment 3

[0170] This embodiment provides a freshness optimization system for federated learning assisted by digital twins, based on the freshness optimization method for federated learning assisted by digital twins described in Embodiment 1 or 2. As Figure 4 shown, the system includes:

[0171] A model construction module, used to construct an industrial Internet of Things federated learning model, the model includes a central server and several intelligent devices, and generate a digital twin corresponding to each intelligent device, which is stored in the central server;

[0172] A calculation module, used to calculate the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning according to the industrial Internet of Things federated learning model;

[0173] An optimization problem establishment module, used to establish an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices for one round of federated learning;

[0174] An optimization problem transformation module, used to transform the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the industrial Internet of Things federated learning model;

[0175] A network construction and training module, used to establish a deep reinforcement learning network based on the proximal policy optimization algorithm, and train the deep reinforcement learning network using the state space, action space, and reward function to obtain a trained deep reinforcement learning network;

[0176] A resource scheduling module, using the trained deep reinforcement learning network for resource scheduling to obtain an optimal scheduling strategy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device;

[0177] A scheduling application module, used to apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

[0178] The same or similar reference numerals correspond to the same or similar components;

[0179] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0180] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A digital twin-assisted federated learning freshness optimization method, characterized in that Including: S1: Construct an industrial Internet of Things (IIoT) federated learning model, which includes a central server and several intelligent devices, and generate a digital twin corresponding to each intelligent device, and save it in the central server; S2: According to the IIoT federated learning model, calculate the energy consumption, local data freshness, and model parameter freshness of all intelligent device digital twins for one round of federated learning; S3: With the goal of minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all intelligent device digital twins for one round of federated learning, establish an optimization problem for joint bandwidth allocation, data collection frequency, and data calculation frequency; S4: Transform the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the IIoT federated learning model; S5: Based on the proximal policy optimization algorithm, establish a deep reinforcement learning network, and use the state space, action space, and reward function to train the deep reinforcement learning network to obtain a trained deep reinforcement learning network; S6: Use the trained deep reinforcement learning network for resource scheduling to obtain the optimal scheduling strategy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device; S7: Apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twin of each intelligent device to the corresponding intelligent device.

2. The freshness optimization method for federated learning assisted by digital twin according to claim 1, wherein In step S1, the constructed IIoT federated learning model is specifically: The IIoT federated learning model includes a central server and K intelligent devices. Each intelligent device generates a corresponding digital twin, and a total of K digital twins are generated and saved in the central server; The central server sends a federated learning request to the digital twins. After receiving the federated learning request, the digital twins start federated learning, collect data and calculate, train the local model, and update the trained local model to the central server; The central server aggregates the received local models into a global model; Set the maximum bandwidth of the central server to B m , and the bandwidth allocated to the k-th digital twin is B k , the data collection frequency of the k-th digital twin is The data calculation frequency of the k-th digital twin is k = 1, 2, …, K.

3. The freshness optimization method for federated learning based on digital twin assistance according to claim 2, wherein In step S2, according to the IIoT federated learning model, the specific method for calculating the energy consumption of all intelligent device digital twins for one round of federated learning is: The number of floating-point operations per CPU cycle of the k-th digital twin is C k , and the number of floating-point operations required to collect a set of data is Then the time taken to collect each set of data is: Wherein, represents the time taken for the k-th digital twin to collect a set of data; The time interval from when the k-th digital twin completes one round of federated learning to when it receives the next federated learning request is Then the energy consumption of the data collection process is: wherein, represents the energy consumption of the k-th digital twin data collection process, and P k represents the energy consumption per unit time of the k-th digital twin; The number of floating-point operations required for the k-th digital twin to complete a set of data collection calculations and updates is Then the local latency for one round of data calculation is: wherein, represents the local delay of the k-th digital twin for one round of data calculation; The upload time for the k-th digital twin to upload the trained local model to the central server is: where D represents the amount of data of the trained local model, and τ k represents the signal-to-noise ratio of the k-th digital twin to the central server; Then the energy consumption of all intelligent device digital twins for one round of federated learning is: In the formula, E represents the energy consumption of all intelligent device digital twins for one round of federated learning.

4. The freshness optimization method for federated learning assisted by digital twin according to claim 3, characterized in that, In step S2, according to the IIoT federated learning model, the specific method for calculating the local data freshness is: It is set that the digital twin of the intelligent device needs to collect N groups of data to train the local model. The moment when the digital twin starts training the local model after receiving the federated learning request is recorded as TrainTime; When the digital twin is in an idle state when it receives the federated learning request, the data freshness is: In the formula, represents the data freshness of the nth group of data of the kth digital twin in the idle state, represents the storage time of the nth group of data of the kth digital twin; When the digital twin is collecting the i-th group of data when it receives the federated learning request, the data consists of the i-th group of data and the previous N - i groups of data, and the data freshness is: In the formula, i < N.

5. The freshness optimization method for federated learning assisted by digital twin according to claim 3, wherein In the step S2, the specific method for calculating the freshness of model parameters according to the industrial Internet of Things federated learning model is as follows: Set the waiting time threshold T of the central server astrict , if and only if , the trained local model uploaded by the k-th digital twin is received by the central server, otherwise the digital twin is considered to fall behind; all K digital twins complete the upload of the trained local model parameters within T astrict . The freshness of the model parameters at the moment of completing the upload is as follows: In the formula, represents the freshness of the model parameters at the moment when the m-th trained local model finishes uploading, represents the upload time when the m-th trained local model is uploaded to the central server; The moment when the central server saves the model parameters of the m-th trained local model is denoted as t sc (t). The moment when the central server starts global model aggregation is denoted as AggregateTime. Then the freshness of the model parameters at the aggregation moment is as follows: wherein, represents the freshness of the model parameters at the aggregation moment of the m-th trained local model.

6. The freshness optimization method for federated learning assisted by digital twin according to claim 4 or 5, characterized in that In the step S3, aiming at minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all smart device digital twins in a round of federated learning, an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency is established, specifically as follows: Calculate the total freshness A according to the local data freshness and the model parameter freshness: Then the optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency is expressed as: C5: λ + μ = 1 Where, M agg represents the number of digital twins participating in federated learning, N upd represents the number of digital twins that successfully upload their local models within T astrict , represents the maximum available computing frequency of the k-th digital twin, λ represents the freshness weight, and μ represents the energy consumption weight.

7. The freshness optimization method for federated learning assisted by digital twin according to claim 2, characterized in that, In the step S4, the state space, action space, and reward function of the industrial Internet of Things federated learning model defined are specifically as follows: In the industrial Internet of Things federated learning model, the central server observes the current state s of the digital twin from the environment t to form the state space, and the current state s t corresponds to the current action a t . The digital twin of the intelligent device executes the current action a in the action space t , interacts with the central server, and returns the current reward r t and the new state s t+1 ; In the state space, the current state expression is s t = {the amount of data that still needs to be collected when the digital twin of each intelligent device receives a federated learning request, the amount of data collected by the digital twin of each intelligent device during the gap between the latest two rounds of federated learning requests, the time interval between two rounds of federated learning requests, the number of stragglers of the digital twin of the intelligent device in each round of federated learning}; In the action space, the current action expression Current reward r t equals the sum of the energy consumption, local data freshness, and model parameter freshness of all digital twins of intelligent devices during one round of federated learning after performing the current action a t until the next round of federated learning request is issued, i.e., r t = λA + μE.

8. The freshness optimization method for federated learning assisted by digital twin according to claim 7, wherein In the step S5, the established deep reinforcement learning network includes an experience buffer, an Actor network, a Critic network, and an oldActor network; At each moment, the input of the Actor network is the current state s t , and the corresponding current action a is output t = μ θ (s t ); The Critic network evaluates the value of the current action a t according to the current state s t ; After the digital twin executes the current action a , a new state s t and the current reward r t+1 are generated. [s t , a t , r t , s t is stored in the experience buffer; The network parameters of the old Actor network are updated at intervals. After several rounds of federated learning at intervals, the network parameters θ of the Actor network are copied and used as the network parameters θ t+1 of the old Actor network. The output of the old Actor network is * used to compare the change range of the current action a ; t ​ The loss function for the Actor network to update the network parameter θ is: In the formula, represents the current action output by the old Actor network, and Adv t (s t ,a t ) represents the advantage function, represents restricting the upper and lower bounds to [1 - ∈, 1 + ∈], where ∈ represents the upper and lower bound limit value, and ∈ ∈ (0, 1); The Critic network updates the network parameters θ c The loss function is as follows: L(θ c ) = MSE(v t , r t ) where MSE(v t , r t ) represents the mean square error function; Iteratively train the Actor network and the Critic network using the policy gradient method. When the loss function value converges, obtain the optimal network parameters θ, θ c and θ * , and obtain a trained deep reinforcement learning network.

9. The freshness optimization method for federated learning assisted by digital twin according to claim 8, characterized in that, The specific method of the step S6 is: The central server observes the current states of the digital twins of all current intelligent devices, inputs them into the trained deep reinforcement learning network, and generates the current actions; the current actions are used as the optimal scheduling strategy, and the B k serve as the optimal data collection frequency, the optimal data calculation frequency, and the allocated optimal bandwidth for the digital twin of each intelligent device.

10. A digital twin-assisted federated learning freshness optimization system, characterized in that, Based on the federated learning freshness optimization method assisted by digital twins according to any one of claims 1-9, the system includes: A model construction module, which is used to construct an industrial Internet of Things federated learning model. The model includes a central server and several smart devices, and generates digital twins corresponding to each smart device, which are stored in the central server; A calculation module, which is used to calculate the energy consumption, local data freshness, and model parameter freshness of all smart device digital twins in a round of federated learning according to the industrial Internet of Things federated learning model; An optimization problem establishment module, which is used to establish an optimization problem of joint bandwidth allocation, data collection frequency, and data calculation frequency with the goal of minimizing the sum of the energy consumption, local data freshness, and model parameter freshness of all smart device digital twins in a round of federated learning; An optimization problem transformation module, which is used to transform the optimization problem into a Markov decision process, and define the state space, action space, and reward function of the industrial Internet of Things federated learning model; A network construction and training module, which is used to establish a deep reinforcement learning network based on the proximal policy optimization algorithm, and train the deep reinforcement learning network using the state space, action space, and reward function to obtain a trained deep reinforcement learning network; A resource scheduling module, which uses the trained deep reinforcement learning network for resource scheduling to obtain an optimal scheduling strategy, that is, the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twins of each smart device; A scheduling application module, which is used to apply the optimal bandwidth, optimal data collection frequency, and optimal data calculation frequency allocated to the digital twins of each smart device to the corresponding smart device.

Citation Information

Patent Citations

  • Federal learning method and system based on edge digital twin association

    CN113419857A

  • Configurable IoT device data collection

    US20180176663A1