A Low-Latency Hierarchical Computation Offloading Method for Digital Twins in Smart Grids
By constructing a three-layer collaborative architecture and using deep reinforcement learning to optimize the drone trajectory and hierarchical unloading strategy, the problems of limited coverage, high data acquisition latency, and computational bottlenecks in the smart grid were solved. This enabled low-latency and efficient data processing of the digital twin model, ensuring real-time monitoring and status updates of the power grid.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies struggle to effectively address issues such as limited coverage, high data acquisition latency, computational bottlenecks, and the high coupling between the trajectory of drones combined with edge computing and multi-layer unloading in smart grids, resulting in insufficient real-time performance and accuracy of digital twin models.
A three-layer collaborative architecture of drone, virtual service provider, and edge computing node is constructed. Combining deep reinforcement learning and convex optimization algorithms, the drone trajectory and hierarchical task offloading strategy are optimized. Through multi-agent deep deterministic policy gradient algorithm and equilibrium-driven minimax optimization strategy, the complex mixed integer nonlinear programming problem is decomposed to minimize the end-to-end latency of the system.
It significantly reduced end-to-end system latency, improved the real-time performance and accuracy of digital twin models, broke through geographical coverage limitations, alleviated computing resource bottlenecks, and ensured the ability to perceive and process massive heterogeneous data across the entire domain.
Smart Images

Figure CN121367958B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the cross-technical fields of mobile communication, edge computing, and smart grids. Specifically, it relates to a low-latency hierarchical computing offloading method for digital twins in smart grids. More specifically, it relates to a method that uses drones and edge computing technology to reduce end-to-end system latency in a smart grid scenario oriented towards digital twins by jointly optimizing flight trajectories and hierarchical computing offloading strategies. Background Technology
[0002] With the rapid convergence of advanced information and communication technologies (ICT), modern power systems are gradually evolving into smart grids, aiming to achieve more efficient, reliable, and sustainable energy distribution. However, this convergence brings a major challenge: how to handle the massive amounts of heterogeneous data from real-time sensors and Internet of Things (IoT) devices. This influx of data places extremely high demands on the grid's situational awareness and real-time monitoring capabilities.
[0003] Digital twin technology has emerged as a key enabling technology to address the aforementioned challenges. By constructing a high-fidelity virtual copy of the physical power grid, digital twins can map, simulate, and analyze the grid's operational status in real time, thereby facilitating accurate system situational awareness and real-time monitoring. However, the effectiveness of digital twins largely depends on their "real-time feedback" capability, meaning the virtual model must remain highly synchronized with the physical entity. If the end-to-end latency from data acquisition to virtual state updates is too high, the digital twin model will become an outdated and inaccurate mirror of the physical world, rendering all subsequent decisions based on it invalid.
[0004] Traditional smart grid data acquisition and processing architectures are unable to meet the stringent low-latency requirements of digital twins, mainly due to the following technical bottlenecks: (1) Data acquisition coverage and latency issues: Traditional data acquisition relies on fixed ground infrastructure, such as Advanced Metering Architecture (AMI) or Wireless Neighborhood Networks (NANs). Such fixed infrastructures have limitations in terms of deployment flexibility and wide-area coverage, especially in geographically dispersed remote areas, where deployment costs are high and it is difficult to cover all assets. In addition, ground data transmission usually requires long-distance non-line-of-sight (NLoS) links, resulting in non-negligible acquisition latency and signal attenuation. (2) Bottleneck of centralized computing: Traditional models tend to transmit sensor data to a remote centralized cloud platform or Power System Control Center (PSCC) for unified processing. This centralized approach is prone to overloading computing resources and introduces significant data transmission latency and queuing latency, which cannot meet the real-time requirements of digital twin applications. (3) The complexity of combining edge computing with UAVs: Although introducing UAVs as mobile data collectors can solve coverage and link quality problems, and introducing edge computing can sink computing power to solve computing latency problems, combining the two into a digital twin architecture faces great challenges. In the three-layer architecture assisted by digital twins, the end-to-end latency of the system is a highly coupled complex function, which is simultaneously affected by the UAV flight trajectory, the first-stage offloading decision (UAV to VSP), and the second-stage offloading decision (VSP to EC). This coupling causes the optimization problem to become a non-convex mixed-integer nonlinear programming problem with NP-hard characteristics. Traditional optimization methods are difficult to find the optimal solution in the complex dynamic environment of multiple UAVs, multiple sensors, and multiple edge nodes.
[0005] Therefore, existing technologies lack a comprehensive solution that can simultaneously address issues such as limited coverage, high acquisition latency, computational bottlenecks, and the high coupling between trajectory and multi-layer unloading. There is an urgent need for a latency-optimized joint trajectory and hierarchical unloading method to support the practical application of digital twins in smart grids. Summary of the Invention
[0006] To address the aforementioned issues, this invention discloses a low-latency hierarchical computation offloading method for digital twins in smart grids. By constructing a three-layer collaborative architecture of "drone-virtual service provider-edge computing node" and utilizing an algorithm framework combining deep reinforcement learning and convex optimization, the method jointly optimizes the drone trajectory and the two-stage task offloading strategy, thereby minimizing the end-to-end latency of the system and ensuring the high fidelity and real-time performance of the digital twin model.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] A low-latency hierarchical computing offloading method for digital twins in smart grids includes:
[0009] Step S1: Construct a three-layer network system model for digital twin (DT) of smart grids; the system model includes a data acquisition and transmission model and a UAV (unmanned aerial vehicle) flight trajectory model;
[0010] Step S2: Based on the system model, quantify the hierarchical task computation offloading model of data from sensor acquisition, UAV relay transmission to Virtual Service Provider (VSP), VSP local and edge computing EC node collaborative computation, as well as the end-to-end full-link latency, and define the latency constraints of digital twin applications;
[0011] Step S3: Considering the strong coupling characteristics of UAV flight trajectory, first-stage unloading decision (UAV to VSP) and second-stage unloading decision (VSP to EC), construct a joint trajectory planning and hierarchical unloading optimization problem model (P) with the goal of minimizing the average end-to-end total system delay.
[0012] Step S4: Taking into account the mixed-integer nonlinear programming characteristics of problem (P), the original problem is decomposed into two independent and parallel subproblems by utilizing the additive structure of the objective function: namely, the non-convex UAV trajectory and the first-stage unloading subproblem (P1), and the convex second-stage unloading subproblem (P2).
[0013] Step S5: For subproblem (P1), model it as a multi-agent Markov decision process, and propose a delayed-aware multi-agent deep deterministic policy gradient algorithm. Iteratively solve the UAV continuous flight control strategy and the first-stage task unloading ratio through a centralized training distributed execution architecture.
[0014] Step S6: For subproblem (P2), using the principle of balanced computing resources between VSP and EC nodes, a minimax optimization strategy based on balanced drive is proposed, and the closed-form solution of the unloading ratio in the second stage is derived and calculated.
[0015] Step S7: Combine the optimization results of steps S5 and S6 to perform data processing, and based on the processed new insight data, perform real-time construction and status update of the digital twin model on the DT server to achieve accurate mapping of the physical power grid to the virtual space.
[0016] Further, in step S1, assume the system includes Sensor devices (SEs) and A drone.
[0017] Data acquisition and transmission model:
[0018] Define sensor To drones Data acquisition latency for
[0019] (1);
[0020] in, It is a binary variable. This indicates that data from the sensor has been collected; otherwise, it is 0. Indicates sensor The amount of data generated; The wireless transmission rate is calculated using the following formula:
[0021] (2);
[0022] in, The bandwidth allocated to the sensor; This refers to the sensor's transmit power. and Sensor devices With drones The noise power and channel gain between them follow the Friis free-space path loss model:
[0023] (3);
[0024] in, and Sensors and drones Antenna gain; The wavelength of the signal; The path loss index; For sensors With drones Euclidean distance between:
[0025] (4);
[0026] in Coordinates of the drone These are the sensor coordinates.
[0027] Unmanned aerial vehicle (UAV) flight trajectory model:
[0028] Drone at a fixed altitude Flight, its horizontal position coordinates The update formula is
[0029] (5);
[0030] (6);
[0031] in, For drones in time slots The flight heading angle; The flight distance within each time slot must be considered. Simultaneously, multi-drone collision avoidance constraints must be met, i.e., the distance between any two drones. It must be greater than the minimum safe distance :
[0032] (7);
[0033] Two-stage hierarchical task computation unloading model:
[0034] (1) First-stage offloading (UAV to VSP) model: The UAV acts as a relay to transfer the collected data proportionally Uninstall to VSP Transmission delay for
[0035] (8);
[0036] in, For drones Total amount of data collected; For the transmission rate of the drone VSP:
[0037] (9);
[0038] in, Provide bandwidth for drone transmission; This refers to the drone's transmission power. The noise power between the drone and the VSP; This represents the channel gain between the UAV and the VSP.
[0039] (2) Second-stage unloading (VSP to EC) model: VSP Receive data Then, proportionally Perform local calculations, remaining Uninstall to Each edge computing node. VSP local computing latency. Defined as
[0040] (10);
[0041] in, Indicates the computational complexity of the task (cycles / bit); This indicates the CPU frequency (cycles / s) of the VSP server. Edge computing latency includes the time it takes for the VSP to transmit data to the EC node. Transmission delay :
[0042] (11);
[0043] in For VSP To EC node The transmission rate is calculated as follows:
[0044] (12);
[0045] in, Bandwidth allocated to VSP For the VSP's transmit power, The coordinates of the VSP. For shared channel interference, This represents the background noise power. EC node. computation delay :
[0046] (13);
[0047] in, To be allocated to EC nodes The proportion; The transmission rate from VSP to EC; This refers to the CPU frequency of the EC node. Edge computing latency. Determined by the slowest branch in parallel processing
[0048] (14);
[0049] VSP total processing latency and total end-to-end system latency:
[0050] VSP final processing latency The larger of the local and edge latency values:
[0051] (15);
[0052] drones Total end-to-end system latency for:
[0053] (16);
[0054] Furthermore, in step S3, the specific formulation of the joint trajectory planning and hierarchical unloading optimization problem model (P) is as follows:
[0055] The objective function is to minimize the average end-to-end total latency of all drone missions.
[0056] (17);
[0057] in, This is a set of drone location sequences. For collecting the set of states, This is the set of unloading strategies for the first phase. This is the set of unloading strategies for the second phase. Furthermore, the optimization problem is subject to the following constraints:
[0058] C1: In other words, the total latency must not exceed the maximum latency required by the application. ;
[0059] C2: That is, the unloading ratio in the first phase needs to be between 0 and 1;
[0060] C3: That is, the unloading ratio in the second phase needs to be between 0 and 1;
[0061] C4: That is, the acquisition status is a binary variable;
[0062] C5: This means that drones must fly within a designated area;
[0063] C6: This means that drones must maintain a safe distance from each other.
[0064] Furthermore, in step S4, the specific process of decomposing the original problem using the additive structure of the objective function is as follows:
[0065] (1) Objective function reconstruction and separation: based on the total time delay formula defined in steps S2 and S3. , the original objective function Expand:
[0066] (18);
[0067] By utilizing the linear property of summation, the objective function can be reorganized into a sum of two parts:
[0068] (19);
[0069] in Only includes drone trajectories Data collection status and the first stage of uninstallation Related latency items:
[0070] (20);
[0071] and Only includes variables related to the second-level unloading. Related VSP processing latency items:
[0072] (twenty one);
[0073] (2) Complete decoupling of decision variables and constraints:
[0074] Analysis of the constraints revealed that the variable set Subject only to constraints C2, C4, C5, and C6; variable set Only constrained by constraint C3; therefore, the optimal solution to the primal problem is... This is equivalent to the superposition of the optimal solutions to the two subproblems:
[0075] (twenty two);
[0076] Furthermore, in step S5, the specific modeling and definition of the LA-MADDPG algorithm are as follows:
[0077] (1) Observation space: each unmanned aerial vehicle agent At any moment Observation status Defined as:
[0078] (twenty three);
[0079] in, The coordinates of the drone itself; The relative distance to other drones; The relative distance to VSP; The relative distance to the sensor; This indicates the sensor data acquisition status. This represents the current communication rate.
[0080] (2) Action Space: Intelligent Agent action Includes continuous flight control and unloading control variables:
[0081] (twenty four);
[0082] in, For flight heading angle; This represents the uninstallation ratio for the first phase.
[0083] (3) Reward function: Define a composite reward function for delayed perception. for:
[0084] (25);
[0085] in, For drones The total system latency is calculated, and the rewards include positive and negative rewards, as detailed below:
[0086] Positive incentive rewards include: Data collection rewards: ,in This represents the amount of data collected. Total data volume; Reward for efficient uninstallation: ,in Ideal unloading ratio; proximity to reward: This encourages drones to approach the target.
[0087] Penalties include: Latency exceeding limits penalty: Collision penalty: Penalty for crossing the boundary: ,in This refers to the distance beyond the boundary.
[0088] Furthermore, the specific algorithm flow of the LA-MADDPG algorithm is as follows:
[0089] Step S5-1: Initialize each agent Actor Network and Critic Network and the corresponding target network and Initialize the experience replay pool. .
[0090] Step S5-2: At each time step The agent generates actions based on the Actor network. And add exploration noise , For Actor networks,
[0091] (26);
[0092] Step S5-3: Next, perform the combined action to obtain the reward. And transition to the next state Next, and Perform normalization to make model training more stable:
[0093] (27);
[0094] Then the normalized transformation tuple Store in the experience replay pool .
[0095] Step S5-4: Next, sample small batches of data from the experience replay pool and use the target network to calculate the target Q value. :
[0096] (28);
[0097] in, This is the discount factor.
[0098] Step S5-5: Next, minimize the loss function Update Critic network parameters:
[0099] (29);
[0100] Steps S5-6: Then, update the Actor network parameters using the deterministic policy gradient:
[0101] (30);
[0102] Step S5-7: Finally, use the soft update coefficient. Update the target network parameters to complete one iteration.
[0103] (31);
[0104] (32);
[0105] Furthermore, in step S6, the specific algorithm flow of the EDMO strategy is as follows:
[0106] Step S6-1: For sub-problem (P2), ignore the transmission latency from VSP to EC nodes, and treat all EC nodes within the VSP coverage area as a parallel edge computing resource pool. Calculate the total computing power of this resource pool. :
[0107] (33);
[0108] in, The computation frequency for a single EC node. This is the set of EC nodes.
[0109] Step S6-2: Next, based on the convex optimization property, when the local computation latency of the VSP layer equals the computation latency of the edge resource pool, the total latency... The minimum value is reached. Establish the equilibrium equation:
[0110] (34);
[0111] Step S6-3: Next, substitute the local and edge computing models into the equilibrium equation:
[0112] (35);
[0113] in, The local calculation ratio for the VSP to be solved. This represents the proportion of the load that is unloaded to the edge side.
[0114] Step S6-4: Finally, solve the above equations to obtain the optimal unloading ratio for the second layer. Closed-form solution:
[0115] (36);
[0116] Output this ratio This is the optimal task unloading strategy for VSP.
[0117] Furthermore, in step S7, the construction and updating process of the digital twin model is as follows:
[0118] Step S7-1: After calculation and processing by the VSP and EC nodes, a power grid entity is generated. At the present moment New insights into data, denoted as ;
[0119] Step S7-2: Next, the server uses the state update function. Combined with the virtual state of the previous moment With newly arrived data , computational entity Current virtual state :
[0120] (37);
[0121] Step S7-3: Finally, aggregate the virtual states of all monitored entities, construct and maintain the timeline. Global high-fidelity digital twin model :
[0122] (38);
[0123] This enables real-time simulation and power grid status monitoring based on the global model.
[0124] The beneficial effects of this invention are as follows:
[0125] First, by leveraging the high mobility of UAVs as mobile relays, this invention overcomes the geographical coverage limitations of traditional fixed ground infrastructure, solving the data acquisition blind spot problem in remote areas and non-line-of-sight scenarios. Simultaneously, by layering and offloading computational tasks to virtual service providers and edge computing nodes, it alleviates the resource bottleneck of a single centralized computing system, significantly improving the system's ability to perceive and process massive amounts of heterogeneous data concurrently. Second, addressing the strong coupling and non-convexity of UAV trajectory planning and task offloading decisions, this invention designs a composite reward mechanism including latency exceedance penalties and guidance rewards, and adopts a centralized training distributed execution architecture. This algorithm enables the UAV agent to autonomously learn the optimal flight trajectory and first-stage offloading strategy in a dynamic environment, effectively avoiding collisions and interference between multiple agents, minimizing data acquisition and relay transmission latency while ensuring safe flight, and significantly reducing the overall system latency. Finally, this invention creatively decomposes the complex mixed-integer nonlinear programming problem into two sub-problems. For the second-stage offloading sub-problem, a closed-form solution based on balanced computational resources is derived using convex optimization properties. This method not only avoids complex iterative calculations and significantly shortens decision-making time, but also provides stable environmental feedback for upper-level reinforcement learning algorithms, ensuring that the data flow from physical entities to virtual space can be processed in real time and efficiently, thereby guaranteeing the real-time performance and high fidelity of digital twin model state updates. Attached Figure Description
[0126] Figure 1 This is a system architecture diagram of hierarchical computing offloading for digital twins in a smart grid according to an embodiment of the present invention;
[0127] Figure 2 This is a specific flowchart in an embodiment of the present invention;
[0128] Figure 3 This is an architecture diagram of the LA-MADDPG algorithm proposed in this invention;
[0129] Figure 4 This is a three-dimensional flight trajectory simulation diagram of the UAV in this embodiment of the invention;
[0130] Figure 5 This is a comparison diagram of the latency of UAV1 (left) and UAV2 (right) under a two-stage unloading strategy and two single unloading strategies in an embodiment of the present invention;
[0131] Figure 6 This is a comparison chart of the convergence performance of the LA-MADDPG algorithm proposed in this embodiment of the invention and existing benchmark algorithms;
[0132] Figure 7 This is a comparison chart showing the effect of introducing the EDMO strategy on the algorithm training convergence in this embodiment of the invention;
[0133] Figure 8 This is a comparison diagram of the latency of UAV1 (left) and UAV2 (right) under different unloading strategies in the second stage of unloading in an embodiment of the present invention. Detailed Implementation
[0134] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0135] like Figure 1 and Figure 2 As shown, the present invention provides a low-latency hierarchical computing offloading method for digital twins in smart grids, comprising:
[0136] Step S1: Construct a three-layer network system model for digital twin (DT) of smart grids; the system model includes a data acquisition and transmission model and a UAV (unmanned aerial vehicle) flight trajectory model;
[0137] like Figure 1 As shown, this paper considers how to minimize the latency of UAV data acquisition and hierarchical computation offloading to meet the real-time construction requirements of digital twin applications and power grid status monitoring. Assume the system includes... Individual sensor devices and Deploy a drone. Define the sensors. To drones Data acquisition latency for
[0138] (1);
[0139] in, It is a binary variable. This indicates that data from the sensor has been collected; otherwise, it is 0. Indicates sensor The amount of data generated; The wireless transmission rate is calculated using the following formula:
[0140] (2);
[0141] in, The bandwidth allocated to the sensor; This refers to the sensor's transmit power. and Sensor devices With drones The noise power and channel gain between them follow the Friis free-space path loss model:
[0142] (3);
[0143] in, and Sensors and drones Antenna gain; The wavelength of the signal; The path loss index; For sensors With drones Euclidean distance between:
[0144] (4);
[0145] in Coordinates of the drone These are the sensor coordinates.
[0146] Drone at a fixed altitude Flight, its horizontal position coordinates The update formula is
[0147] (5);
[0148] (6);
[0149] in, For drones in time slots The flight heading angle; The flight distance within each time slot must be considered. Simultaneously, multi-drone collision avoidance constraints must be met, i.e., the distance between any two drones. It must be greater than the minimum safe distance :
[0150] (7);
[0151] Step S2: Based on the system model, quantify the hierarchical task computation offloading model of data from sensor acquisition, UAV relay transmission to Virtual Service Provider (VSP), VSP local and edge computing EC node collaborative computation, as well as the end-to-end full-link latency, and define the latency constraints of digital twin applications;
[0152] First, the drone acts as a relay to transmit the collected data. proportionally Uninstall to VSP Transmission delay during this process for
[0153] (8);
[0154] in, For drones Total amount of data collected; For the transmission rate of the drone VSP:
[0155] (9);
[0156] in, Provide bandwidth for drone transmission; This refers to the drone's transmission power. The noise power between the drone and the VSP; This represents the channel gain between the UAV and the VSP.
[0157] Then, VSP Receive data Then, proportionally Perform local calculations, remaining Uninstall to Each edge computing node. VSP local computing latency. Defined as
[0158] (10);
[0159] in, Indicates the computational complexity of the task (cycles / bit); This indicates the CPU frequency (cycles / s) of the VSP server. Edge computing latency includes the time it takes for the VSP to transmit data to the EC node. Transmission delay :
[0160] (11);
[0161] in For VSP To EC node The transmission rate is calculated as follows:
[0162] (12);
[0163] in, For the VSP's transmit power, The coordinates of the VSP. For shared channel interference, This represents the background noise power. EC node. computation delay :
[0164] (13);
[0165] in, To be allocated to EC nodes The proportion; The transmission rate from VSP to EC; This refers to the CPU frequency of the EC node. Edge computing latency. Determined by the slowest branch in parallel processing
[0166] (14);
[0167] Next, the final processing latency of VSP. The larger of the local and edge latency values:
[0168] (15);
[0169] Finally, drones Total end-to-end system latency for:
[0170] (16);
[0171] Step S3: Considering the strong coupling characteristics of UAV flight trajectory, first-stage unloading decision (UAV to VSP) and second-stage unloading decision (VSP to EC), construct a joint trajectory planning and hierarchical unloading optimization problem model (P) with the goal of minimizing the average end-to-end total system delay.
[0172] The objective function is to minimize the average end-to-end total latency of all drone missions.
[0173] (17);
[0174] in, This is a set of drone location sequences. For collecting the set of states, This is the set of unloading strategies for the first phase. This is the set of unloading strategies for the second phase. Furthermore, the optimization problem is subject to the following constraints:
[0175] C1: In other words, the total latency must not exceed the maximum latency required by the application. ;
[0176] C2: That is, the unloading ratio in the first phase needs to be between 0 and 1;
[0177] C3: That is, the unloading ratio in the second phase needs to be between 0 and 1;
[0178] C4: That is, the acquisition status is a binary variable;
[0179] C5: This means that drones must fly within a designated area;
[0180] C6: This means that drones must maintain a safe distance from each other.
[0181] Step S4: Taking into account the mixed-integer nonlinear programming characteristics of problem (P), the original problem is decomposed into two independent and parallel subproblems by utilizing the additive structure of the objective function: namely, the non-convex UAV trajectory and the first-stage unloading subproblem (P1), and the convex second-stage unloading subproblem (P2).
[0182] The specific problem breakdown process is as follows:
[0183] First, based on the total delay formula defined in steps S2 and S3. , the original objective function Expand:
[0184] (18);
[0185] Secondly, by utilizing the linear property of summation, the objective function is reorganized into a sum of two parts:
[0186] (19);
[0187] in Only includes drone trajectories Data collection status and the first stage of uninstallation Related latency items:
[0188] (20);
[0189] and Only includes variables related to the second-level unloading. Related VSP processing latency items:
[0190] (twenty one);
[0191] Then, analyzing the constraints revealed that the variable set Subject only to constraints C2, C4, C5, and C6; variable set Only constrained by constraint C3; therefore, the optimal solution to the primal problem is... This is equivalent to the superposition of the optimal solutions to the two subproblems:
[0192] (twenty two);
[0193] Step S5: For sub-problem (P1), it is modeled as a multi-agent Markov decision process, and a delayed-aware multi-agent deep deterministic policy gradient algorithm is proposed. This algorithm iteratively solves the UAV continuous flight control strategy and the first-stage task offloading ratio using a centralized training and distributed execution architecture. The specific algorithm architecture is as follows: Figure 3 As shown.
[0194] (1) Observation space: each unmanned aerial vehicle agent At any moment Observation status Defined as:
[0195] (twenty three);
[0196] in, The coordinates of the drone itself; The relative distance to other drones; The relative distance to VSP; The relative distance to the sensor; This indicates the sensor data acquisition status. This represents the current communication rate.
[0197] (2) Action Space: Intelligent Agent action Includes continuous flight control and unloading control variables:
[0198] (twenty four);
[0199] in, For flight heading angle; This represents the uninstallation ratio for the first phase.
[0200] (3) Reward function: Define a composite reward function for delayed perception. for:
[0201] (25);
[0202] in, For drones The total system latency is calculated, and the reward function includes both positive and negative rewards, as detailed below:
[0203] Positive incentive rewards include: Data collection rewards: ,in This represents the amount of data collected. Total data volume; Reward for efficient uninstallation: ,in Ideal unloading ratio; proximity to reward: This encourages drones to approach the target.
[0204] Penalties include: Latency exceeding limits penalty: Collision penalty: Penalty for crossing the boundary: ,in This refers to the distance beyond the boundary.
[0205] The LA-MADDPG algorithm flow is as follows:
[0206] Step S5-1: Initialize each agent Actor Network and Critic Network and the corresponding target network and Initialize the experience replay pool. .
[0207] Step S5-2: At each time step The agent generates actions based on the Actor network. And add exploration noise :
[0208] (26);
[0209] Step S5-3: Next, perform the combined action to obtain the reward. And transition to the next state Next, and Perform normalization to make model training more stable:
[0210] (27);
[0211] Then the normalized transformation tuple Store in the experience replay pool .
[0212] Step S5-4: Next, sample small batches of data from the experience replay pool and use the target network to calculate the target Q value. :
[0213] (28);
[0214] in, This is the discount factor.
[0215] Step S5-5: Next, minimize the loss function Update Critic network parameters:
[0216] (29);
[0217] Steps S5-6: Then, update the Actor network parameters using the deterministic policy gradient:
[0218] (30);
[0219] Step S5-7: Finally, use the soft update coefficient. Update the target network parameters to complete one iteration.
[0220] (31);
[0221] (32);
[0222] Step S6: For subproblem (P2), using the principle of balanced computing resources between VSP and EC nodes, a minimax optimization strategy based on balanced drive is proposed, and the closed-form solution of the unloading ratio in the second stage is derived and calculated.
[0223] The specific algorithm flow of the EDMO strategy is as follows:
[0224] Step S6-1: For sub-problem (P2), ignore the transmission latency from VSP to EC nodes, and treat all EC nodes within the VSP coverage area as a parallel edge computing resource pool. Calculate the total computing power of this resource pool. :
[0225] (33);
[0226] in, The computation frequency for a single EC node. This is the set of EC nodes.
[0227] Step S6-2: Next, based on the convex optimization property, when the local computation latency of the VSP layer equals the computation latency of the edge resource pool, the total latency... The minimum value is reached. Establish the equilibrium equation:
[0228] (34);
[0229] Step S6-3: Next, substitute the local and edge computing models into the equilibrium equation:
[0230] (35);
[0231] in, The local calculation ratio for the VSP to be solved. This represents the proportion of the load that is unloaded to the edge side.
[0232] Step S6-4: Finally, solve the above equations to obtain the optimal unloading ratio for the second layer. Closed-form solution:
[0233] (36);
[0234] Output this ratio This is the optimal task unloading strategy for VSP.
[0235] Step S7: Combine the optimization results of steps S5 and S6 to perform data processing, and based on the processed new insight data, perform real-time construction and status update of the digital twin model on the DT server to achieve accurate mapping of the physical power grid to the virtual space.
[0236] Step S7-1: After calculation and processing by the VSP and EC nodes, a power grid entity is generated. At the present moment New insights into data, denoted as ;
[0237] Step S7-2: Next, the server uses the state update function. Combined with the virtual state of the previous moment With newly arrived data , computational entity Current virtual state :
[0238] (37);
[0239] Step S7-3: Finally, aggregate the virtual states of all monitored entities, construct and maintain the timeline. Global high-fidelity digital twin model :
[0240] (38);
[0241] This enables real-time simulation and power grid status monitoring based on the global model.
[0242] To verify the effectiveness of the proposed low-latency hierarchical computing offloading method for digital twins in smart grids, this invention conducted numerous experiments in a simulation environment and compared it with existing technologies.
[0243] like Figure 4 As shown, the three-dimensional flight trajectories of two drones generated by the method of the present invention are displayed in an area containing 15 sensor devices and 3 virtual service providers. Figure 4The solid line represents the drone's flight path, the red squares represent VSPs (Vehicle Support Modules), and the points at different altitudes represent distributed sensors. From Figure 4 As can be clearly seen, the UAV can autonomously plan its path from the starting point to fly over and cover sensor devices distributed in different locations for data collection, while effectively avoiding each other's flight paths and meeting the minimum safe distance constraint. Furthermore, the UAV's trajectory shows a trend towards the VSP (Virtual Private Space), facilitating higher-rate data offloading. This verifies the effectiveness of the method of this invention in handling complex joint trajectory planning and collision avoidance tasks.
[0244] like Figure 5 As shown, the performance advantages of the proposed two-stage computation offloading strategy based on balanced-driven minimax optimization are verified. The two-stage offloading strategy is compared with two benchmark strategies: "Virtual Service Provider Only" and "Edge Node Only". Experimental results show that as the task data volume increases from 10 Gbits to 30 Gbits, the local offloading strategy of this invention consistently maintains the lowest total latency. This is because a single computation mode is limited by the computational capacity bottleneck of a single node, easily leading to resource overload. In contrast, the EDMO strategy of this invention, through the optimal offloading ratio derived analytically, can dynamically balance the load between the VSP local and edge resource pools, achieving parallel computation and thus significantly reducing processing latency.
[0245] like Figure 6 As shown, the reward convergence curves of the LA-MADDPG algorithm proposed in this invention are compared with those of the traditional MADDPG, MASAC, and MAPPO algorithms during the training process. From Figure 6 As can be seen, under the same number of training rounds, the LA-MADDPG algorithm proposed in this invention exhibits the best convergence performance. Specifically, the LA-MADDPG algorithm can rapidly increase the reward value in the early stages of training and tends to stabilize after about 800 rounds, with the final convergence reward value being significantly higher than the other three benchmark algorithms. In contrast, the MASAC algorithm, although showing an upward trend, fluctuates greatly and lacks stability; while the MAPPO algorithm failed to converge effectively in this scenario, even experiencing policy collapse in the later stages of training; the unimproved MADDPG algorithm has a slow convergence speed and a lower final reward. This fully demonstrates that by introducing state normalization and a delayed-aware reward mechanism, combined with the CTDE architecture, this invention can effectively overcome the non-stationarity of multi-agent environments and learn a better cooperative policy.
[0246] like Figure 7As shown, the impact of using the offloading decision as a learnable parameter versus employing a fixed average policy and a random policy on the reward acquisition of a UAV agent is illustrated. The curves show that the UAV agent assisted by the EDMO policy of this invention (purple curve) obtains significantly higher reward values than those using the average or random policies, and exhibits extremely high stability. In contrast, the average policy experiences a sharp performance drop in certain round intervals, while the random policy remains in a low-reward oscillation state. This demonstrates that the closed-form solution for the second-stage offloading obtained through the EDMO policy can provide stable environmental feedback for the upper-level reinforcement learning algorithm, avoiding policy learning failures caused by chaotic lower-level offloading decisions, thereby ensuring the overall robustness of the system.
[0247] like Figure 8 As shown, the performance comparison of the proposed method with the average allocation strategy and the random allocation strategy in terms of total system latency is illustrated as the amount of task data increases. Experimental results show that, for both UAV1 and UAV2, the proposed method achieves the lowest total system latency at all data volume test points. Especially under high load conditions (such as data volumes exceeding 25 Gbits), the advantages of the proposed method are more pronounced, with a lower latency growth slope than the other two benchmark strategies. This further demonstrates that the proposed method, through joint trajectory optimization, first-stage offloading, and second-stage offloading, can intelligently schedule resources based on real-time channel conditions and computational load, thereby maximizing the system's data processing efficiency while meeting the real-time requirements of digital twins.
[0248] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.
Claims
1. A low-latency hierarchical computing offloading method for digital twins in smart grids, characterized in that, The method includes the following steps: Step S1: Construct a three-layer network system model for digital twin (DT) of smart grids; the system model includes a data acquisition and transmission model and a UAV (unmanned aerial vehicle) flight trajectory model; Step S2: Based on the system model, quantify the hierarchical task computation offloading model of data from sensor acquisition, UAV relay transmission to Virtual Service Provider (VSP), VSP local and edge computing EC node collaborative computation, as well as the end-to-end full-link latency, and define the latency constraints of digital twin applications; The specific definitions of the hierarchical task computation offloading model and the end-to-end full-link latency are as follows: (1) First stage: Unloading UAV to VSP model: The UAV acts as a relay to transfer the collected data. proportionally Uninstall to VSP Transmission delay for (1); in, For drones Total amount of data collected; For the transmission rate between the drone and the VSP: (2); in, Provide bandwidth for drone transmission; This refers to the drone's transmission power. The noise power between the drone and the VSP; Channel gain between the UAV and VSP; (2) Second phase: Unloading VSP to EC model: VSP Receive data Then, proportionally Perform local calculations, remaining Uninstall to One edge computing node; VSP local computing latency Defined as (3); in, Indicates the computational complexity of the task; This indicates the CPU frequency of the VSP server; edge computing latency includes: the time it takes for the VSP to transmit data to the EC node. Transmission delay : (4); in For VSP To EC node The transmission rate is calculated as follows: (5); in, Bandwidth allocated to VSP This refers to the transmit power of the VSP. The coordinates of the VSP. For shared channel interference, Background noise power; EC node computation delay : (6); in, To be allocated to EC nodes The proportion; The transmission rate from VSP to EC; CPU frequency of EC nodes; edge computing latency Determined by the slowest branch in parallel processing (7); (3) Total VSP processing latency and total system latency: The final processing latency of the VSP layer The larger of the local and edge latency values is used to determine: (8); Therefore, drones Total end-to-end system latency for: (9); in For sensors To drones Data acquisition latency; Step S3: Considering the strong coupling between the UAV flight trajectory, the first-stage unloading decision from UAV to VSP, and the second-stage unloading decision from VSP to EC, construct a joint trajectory planning and hierarchical unloading optimization problem model P with the goal of minimizing the average end-to-end total system delay; Step S4: Based on the mixed integer nonlinear programming (MINLP) characteristics of problem P, the original problem is decomposed into two independent and parallel subproblems using the additive structure of the objective function: namely, the non-convex UAV trajectory and the first-stage unloading subproblem P1, and the convex second-stage unloading subproblem P2. Step S5: For subproblem P1, model it as a multi-agent Markov decision process (MAMDP) and propose a delayed-aware multi-agent deep deterministic policy gradient (LA-MADDPG) algorithm. Iteratively solve the UAV continuous flight control strategy and the first-stage task unloading ratio through a centralized training and distributed execution (CTDE) architecture. The specific modeling and definition of the LA-MADDPG algorithm are as follows: (1) Observation space: each unmanned aerial vehicle agent At any moment Observation status Defined as: (10); in, The coordinates of the drone itself; The relative distance to other drones; The relative distance to VSP; The relative distance to the sensor; This indicates the sensor data acquisition status. This represents the current communication rate. (2) Action Space: Intelligent Agent action Includes continuous flight control and unloading control variables: (11); in, For the flight heading angle, This represents the uninstallation ratio for the first phase. (3) Reward function: Define a composite reward function for delayed perception. for: (12); in, For drones The total system latency is calculated, and the rewards include both positive and negative rewards, which will be detailed below: Positive incentive rewards include: Data collection rewards: ,in This represents the amount of data collected. Total data volume; Reward for efficient uninstallation: ,in Ideal unloading ratio; proximity to reward: Encourage drones to approach targets; Penalties include: Latency exceeding limits penalty: Collision penalty: Penalty for crossing the boundary: ,in This refers to the distance beyond the boundary. The specific execution flow of the LA-MADDPG algorithm includes: Step S5-1: Initialize each agent Actor Network and Critic Network and the corresponding target network and Initialize the experience replay pool ; Step S5-2: At each time step The agent generates actions based on the Actor network. And add exploration noise : (13); Step S5-3: Next, perform the combined action to obtain the reward. And transition to the next state Next, and Perform normalization to make model training more stable: (14); Then the normalized transformation tuple Store in the experience replay pool ; Step S5-4: Next, sample small batches of data from the experience replay pool and use the target network to calculate the target Q value. : (15); in, Discount factor; Step S5-5: Next, minimize the loss function Update Critic network parameters: (16); Steps S5-6: Then, update the Actor network parameters using the deterministic policy gradient: (17); Step S5-7: Finally, use the soft update coefficient. Update the target network parameters to complete one iteration; (18); (19); Step S6: For subproblem P2, using the principle of balancing computational resources between VSP and EC nodes, a minimum-maximum optimization EDMO strategy based on balance-driven optimization is proposed, and the closed-form solution of the unloading ratio in the second stage is derived and calculated. The specific execution process of the EDMO strategy includes: Step S6-1: For subproblem P2, ignore the transmission latency from VSP to EC nodes, and treat all EC nodes within the VSP coverage area as a parallel edge computing resource pool; calculate the total computing power of this resource pool. : (20); in, The computation frequency for a single EC node. For the set of EC nodes; Step S6-2: Next, based on the convex optimization property, when the local computation latency of the VSP equals the computation latency of the edge resource pool, the total latency... Reaching the minimum value; establishing the equilibrium equation: (21); Step S6-3: Next, substitute the local and edge computing models into the equilibrium equation: (22); in, The local calculation ratio for the VSP to be solved. This represents the proportion of the material unloaded to the edge side. Step S6-4: Finally, solve the above equations to obtain the optimal unloading ratio for the second stage. Closed-form solution: (23); Output this ratio As the optimal task unloading strategy for VSP; Step S7: Combine the optimization results of steps S5 and S6 to perform data processing, and based on the processed new insight data, perform real-time construction and status update of the digital twin model on the DT server to achieve accurate mapping of the physical power grid to the virtual space; The construction and updating process of the digital twin model is as follows: Step S7-1: After calculation and processing by the VSP and EC nodes, a power grid entity is generated. At the present moment New insights into data, denoted as ; Step S7-2: Next, the server uses the state update function. Combined with the virtual state of the previous moment With newly arrived data , computational entity Current virtual state : (24); Step S7-3: Finally, aggregate the virtual states of all monitored entities, construct and maintain the timeline. Global high-fidelity digital twin model : (25); This enables real-time simulation and power grid status monitoring based on the global model.
2. The low-latency hierarchical computing offloading method for digital twins in a smart grid according to claim 1, characterized in that: In step S1, the specific definitions of the data acquisition and transmission model and the UAV flight trajectory model are as follows: (1) Data Acquisition and Transmission Model: Assume the system includes Individual sensor devices and Deploy drones; define sensors To drones Data acquisition latency for (26); in, It is a binary variable. This indicates that data from the sensor has been collected; otherwise, it is 0. Indicates sensor The amount of data generated; The wireless transmission rate is calculated using the following formula: (27); in, The bandwidth allocated to the sensor; This refers to the sensor's transmit power. and Sensor devices With drones The noise power and channel gain between them follow the Friis free-space path loss model: (28); in, and Sensors and drones Antenna gain; The wavelength of the signal; The path loss index; For sensors With drones Euclidean distance between: (29); in Coordinates of the drone Sensor coordinates; (2) Unmanned Aerial Vehicle (UAV) flight trajectory model: The UAV at a fixed altitude Flight, its horizontal position coordinates The update formula is (30); (31); in, For drones in time slots The flight heading angle; The flight distance within each time slot must be considered; simultaneously, multi-drone collision avoidance constraints must be met, i.e., the distance between any two drones. It must be greater than the minimum safe distance : (32)。 3. A low-latency hierarchical computing offloading method for digital twins in a smart grid according to claim 2, characterized in that: In step S3, the specific formulation of the joint trajectory planning and hierarchical unloading optimization problem model P is as follows: The objective function is to minimize the average end-to-end total latency of all drone missions. (33); in, This is a set of drone location sequences. For collecting the set of states, This is the set of unloading strategies for the first phase. This is the set of unloading strategies for the second phase; furthermore, the optimization problem must be subject to the following constraints: C1: In other words, the total latency must not exceed the maximum latency required by the application. ; C2: That is, the unloading ratio in the first phase needs to be between 0 and 1; C3: That is, the unloading ratio in the second phase needs to be between 0 and 1; C4: That is, the acquisition status is a binary variable; C5: This means that drones must fly within a designated area; C6: This means that drones must maintain a safe distance from each other.
4. A low-latency hierarchical computing offloading method for digital twins in a smart grid according to claim 3, characterized in that: In step S4, the specific process of decomposing the original problem using the additive structure of the objective function is as follows: (1) Objective function reconstruction and separation: based on the total time delay formula defined in formula (9) , the original objective function Expand: (34); By utilizing the linear property of summation, the objective function can be reorganized into the sum of two parts: (35); in Only includes drone trajectories Data collection status and the first stage of uninstallation Related latency items: (36); and Includes only variables related to the second phase of unloading. Related VSP processing latency items: (37); (2) Complete decoupling of decision variables and constraints: Analysis of the constraints revealed that the variable set Constrained by constraints C2, C4, C5, and C6; variable set Only constrained by constraint C3; therefore, the optimal solution to the primal problem is... This is equivalent to the superposition of the optimal solutions to the two subproblems: (38)。
Citation Information
Patent Citations
Unmanned aerial vehicle digital energy calculation joint resource allocation method based on digital twinning
CN115988543A
CNN (Convolutional Neural Network)-oriented mobile edge computing dynamic unloading method and equipment medium
CN119697701A