Heterogeneous task and resource end-edge collaborative scheduling method based on digital twinning
By combining digital twin technology and multi-agent deep reinforcement learning, the problems of resource competition and fragmentation of heterogeneous tasks on 5G networks are solved, and the total latency of heterogeneous tasks is minimized and the quality of service is improved.
Patent Information
- Application Number
- CN202310046985.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-31
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-01-31
AI Technical Summary
Existing methods have failed to effectively address the resource contention and fragmentation issues of heterogeneous tasks accessing 5G networks with high concurrency, especially the mutual interference issues of heterogeneous computing resource type matching and device computing migration, leading to a decline in service quality.
Digital twin technology is used to virtualize and model heterogeneous computing resources. Combined with multi-agent deep reinforcement learning, an Actor-Critic neural network model is constructed to achieve collaborative scheduling of heterogeneous tasks and resources. Resource allocation and task migration are carried out through multi-agent Markov decision process to optimize the use of computing and communication resources.
It minimizes the total latency of heterogeneous tasks, meets the quality of service requirements of heterogeneous tasks, supports real-time collaborative processing of computationally intensive and latency-sensitive tasks, and reduces task processing latency.
Smart Images

Figure CN116156563B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless networks, specifically to a method for heterogeneous task and resource edge-to-edge collaborative scheduling based on digital twins. Background Technology
[0002] With the rapid development of 5G, more and more devices are connecting to the internet, linking people, machines, and things to achieve the Internet of Everything. Therefore, the transmission of large-scale heterogeneous tasks on 5G networks is explosive. Typical heterogeneous tasks include: video / media tasks requiring high-bandwidth communication, sensing / measurement tasks requiring low-power communication, and industrial control tasks requiring real-time computation and deterministic communication. To complete a complex task, multiple heterogeneous tasks need to be coordinated. When these heterogeneous tasks access the network concurrently, they must compete for limited "spatiotemporal frequency" domain communication resources, causing transmission conflicts and leading to a decline in service quality.
[0003] To improve service quality, multi-access edge computing can be applied to assist in processing tasks on terminal devices, thereby reducing task processing latency. Typically, by deploying edge servers on base stations, the base stations can perform some network management functions while providing computing resources to terminal devices to assist in computation and reduce task processing latency. However, adopting multi-access edge computing may further exacerbate the problem of competition for communication resources. In particular, heterogeneous industrial tasks have different requirements for computing and communication resources, which will lead to resource fragmentation. Therefore, rationally allocating computing and communication resources between terminals and edge servers according to the needs of heterogeneous tasks is a core challenge currently faced.
[0004] Existing methods address different multi-access edge computing scenarios, employing various optimization algorithms or theories to achieve different objectives such as minimizing latency, minimizing energy consumption, and maximizing throughput by optimizing different parameters. However, these methods do not address issues such as matching heterogeneous computing resource types, mutual interference during device migration, and resource estimation errors. Summary of the Invention
[0005] This invention addresses a broad scenario involving a single cloud server, multiple edge servers, and multiple terminal devices. It employs digital twin technology to virtualize and model heterogeneous computing resources, enabling collaborative scheduling of heterogeneous tasks, computing resources, and communication resources. To this end, the invention fully considers the deadlines of heterogeneous tasks, the types of computing resources, the types and maximum computing capabilities of terminal devices and edge servers, the estimation bias of computing resources in digital twins, the maximum transmit power of terminal devices, and the peak interference power of terminal devices. It proposes an edge-end collaborative scheduling method for heterogeneous tasks and network computing and communication resources based on multi-agent deep reinforcement learning. This method solves the problem of state space explosion in dynamic network environments, which traditional scheduling methods struggle to address, and minimizes the total processing latency of heterogeneous tasks. It supports real-time collaborative processing of computationally intensive, latency-sensitive, and other heterogeneous high-concurrency tasks.
[0006] The technical solution adopted by this invention to achieve the above objectives is: a heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins, which realizes the collaborative scheduling of heterogeneous tasks and heterogeneous computing and communication resources based on multi-agent deep reinforcement learning, including the following steps:
[0007] 1) Establish an edge wireless network based on digital twins;
[0008] 2) Based on the deadline requirements of heterogeneous tasks and the constraints of heterogeneous computing and communication resources, construct the end-edge collaborative scheduling problem of heterogeneous tasks and resources.
[0009] 3) Transform the scheduling problem into a multi-agent Markov decision process problem;
[0010] 4) Construct an Actor-Critic neural network model based on multi-agent deep reinforcement learning to solve the multi-agent Markov decision process problem;
[0011] 5) The digital twin is trained offline in a centralized manner to obtain the experience pool and neural network parameters;
[0012] 6) Terminal devices perceive the environmental status online and perform distributed task migration and allocation of computing and communication resources based on the centrally trained Actor-Critic neural network model to collaboratively process heterogeneous tasks and minimize the total task processing latency.
[0013] The digital twin-based edge wireless network includes: N base stations configured with edge servers and M terminal devices;
[0014] The base station is configured with an edge server to provide computing resources for multiple terminal devices and to support the scheduling of terminal devices within its coverage area.
[0015] The terminal device is used to perform local computation on heterogeneous tasks and supports migrating heterogeneous tasks to edge servers for edge computation via wireless channels.
[0016] The digital twin, placed on the network's cloud server, is represented as a virtualized model of the network containing base stations and terminal devices. It is used to evaluate the operating status, computing resource types, and quantity of computing and communication resources of base stations, edge servers, and terminal devices, and supports the training of deep reinforcement learning methods to carry out network edge-end collaborative scheduling.
[0017] For a single terminal device, computation is performed by migrating tasks, either not, partially, or entirely, to one or more edge servers.
[0018] The transmission rate of the terminal device during task migration is:
[0019]
[0020] Among them, W m,n This represents the bandwidth between terminal device m and edge server n. G represents the noise at edge server n. m,n and g m',n Let p represent the channel power gain from terminal device m and terminal device m' to edge server n, respectively. m p m' These represent the transmit power of terminal device m and terminal device m', respectively.
[0021] The problem of edge-to-edge collaborative scheduling of heterogeneous tasks and resources is as follows:
[0022]
[0023]
[0024] C2:0≤p m ≤P max m=1,...M,
[0025]
[0026]
[0027] C5:0≤f m,n +Δf m,n ≤F max,n ,
[0028]
[0029] C7:T m ≤T max,m
[0030] in, Let T be the objective of the problem, representing minimizing the total task processing latency. m This represents the task processing latency of terminal device m. Let be the set of variables to be optimized in the problem, representing the computing resource type matching decision, task migration ratio, terminal device transmit power, and edge server computing resource allocation, respectively.
[0031] C1 is the task migration ratio constraint; where v m,n ∈[0,1] is the proportion of tasks migrated from terminal device m to edge server n, v m,n =0 indicates that the terminal device m has not migrated the task to the edge server n, v m,n =1 indicates a task migrating from terminal device m to edge server n, v m,0 =0 indicates that terminal device m does not perform local calculations, v m,0 =1 indicates that terminal device m performs local calculations;
[0032] C2 and C3 are the transmit power constraints for the terminal equipment; where, P max I represents the maximum transmit power of the terminal device. p This indicates the peak interference power that the terminal device can tolerate. and Representing the connection from terminal device m and terminal device m' to terminal device m respectively. * The channel gain; where m * =argmax g m,m' It is the terminal device that causes the most interference to terminal device m;
[0033] C4 represents the matching decision constraint for heterogeneous computing resource types; where o m With o n These represent the computing resource types of the terminal device m and the edge server n, respectively. Represents the XOR operation; u m,n =1 indicates that the terminal device m and the edge server n have the same type of computing resources; u m,n =0 indicates that the terminal device m and the edge server n have different types of computing resources;
[0034] C5 and C6 are computational resource constraints; where f m,n Represents the edge computing resources estimated by the digital twin; Δf m,n F represents the bias in the computational resource estimation of a digital twin; max,n This represents the maximum computing speed of edge server n;
[0035] C7 is the task deadline constraint; where T max,mThis represents the deadline for the task to be executed by terminal device m, which is the longest task processing delay that terminal device m can accept.
[0036] The task processing latency of the terminal device is determined by the edge computing latency. and local computing latency The decision and calculation method are as follows:
[0037]
[0038] The edge computing latency Calculated as
[0039]
[0040] in, This represents the edge computing latency of the task migrated from edge server n to terminal device m, and the communication latency caused by the task migration. Computational latency for task processing Decision, calculated as
[0041]
[0042] Communication latency of the task migration The amount and rate of task migration determined by the terminal device are calculated as follows:
[0043]
[0044] Among them, D m Indicates the task size of terminal device m;
[0045] The computational delay of the task processing The task migration amount of terminal device m and the computing resources f allocated by edge server n to terminal device m. m,n Decision, calculated as
[0046]
[0047] The edge computing latency estimated by the digital twin is calculated as follows:
[0048]
[0049] Among them, C m This indicates the computation cycle required to compute a 1-byte task;
[0050] It is the deviation between the actual calculation delay and the estimated calculation delay, calculated as...
[0051]
[0052] The local computation latency Calculated as
[0053]
[0054] It is the local computation latency estimated by the digital twin, calculated as
[0055]
[0056] Among them, f m =F max,m -Δf m Indicates local computing resources;
[0057] This is the local calculation delay deviation, calculated as follows:
[0058]
[0059] The process of transforming the optimization scheduling problem into a multi-agent Markov decision process problem includes the following steps:
[0060] a) Establish a multi-agent Markov decision model, including the agent set, state space, action space, state transition probability, and reward function;
[0061] The intelligent ensemble is an intelligent ensemble formed by M terminal devices.
[0062] The state space is the state of agent m at time t, denoted as:
[0063]
[0064] Among them, D m (t) represents the task size of terminal device m; C m (t) represents the number of computation cycles required for terminal device m; T max,m (t) represents the task deadline of terminal device m; Δf m (t) represents the estimation bias of the local computing resources of the terminal device m; W represents the estimation bias of computing resources for N edge servers of terminal device m; m (t)={W m,1 (t),...,W m,N (t)} and G m (t)={g m,1 (t),...,g m,Ns(t) and s(t) represent the bandwidth and channel gain between the terminal device m and the N edge servers, respectively; the total state space of all agents at time t is s(t) = {s1(t), ... s2(t)}. M (t)};
[0065] The action space refers to the actions performed by agent m at time t, denoted as:
[0066] a m (t)={u m (t),v m (t),p m (t),f m (t)}
[0067] Among them, u m (t)={u m,1 (t),...,u m,N (t)} represents the computing resource type matching decision, determining whether the computing resource type of the edge server is consistent with that of the terminal device m; v m (t)={v m,0 (t),v m,1 (t),...,v m,N (t)} represents the proportion of tasks migrated between terminal device m and N edge servers; p m (t) represents the transmit power of terminal device m used for mission migration; f m (t)={f m,1 (t),...,f m,N {a(t)} represents the computing resources allocated to terminal device m by N edge servers; the total action space of all agents at time t is a(t) = {a1(t), ... a2(t)}. M (t)};
[0068] The state transition probability is given when agent m executes action a. m When (t), s m (t) is transferred to s m The probability of (t+1), i.e., z m (s m (t+1); s m (t),a m (t));
[0069] The reward function is the reward or punishment for an agent taking an action in a certain set state, denoted as r. m (t); where the individual reward obtained by agent m is ρ m This represents the weight parameters set based on the deadline requirements of heterogeneous tasks; the latency reward is... Deadline reward is
[0070] b) Determine the long-term cumulative reward function as follows
[0071]
[0072] Where t represents the current time, t0 represents the previous time, and γ m ∈[0,1] represents the discount factor, indicating the influence of past rewards on the current reward of agent m;
[0073] c) Transform the problem into
[0074] max R m (t)
[0075] stC1,C2,C3,C4,C5,C6
[0076] Under the constraints C1-C6, the strategy is to maximize the long-term cumulative reward to obtain the optimal state transition probability, thereby minimizing the total task processing latency.
[0077] The Actor-Critic neural network model, constructed based on multi-agent deep reinforcement learning, includes an Actor network and a Critic network.
[0078] The Actor network employs a policy-based deep neural network, comprising an estimating Actor network for training and a target Actor network for executing actions to generate agent actions.
[0079] The Critic network employs a value-based deep neural network, including an estimation Critic network and a target Critic network, to evaluate the Actor's actions and guide the Actor to produce better actions.
[0080] The offline centralized training of the digital twin neural network model includes the following steps:
[0081] a) Input s m (t) to estimate the Actor network, obtain Where, π m Indicates taking action a m The strategy of (t), This indicates the estimation of the parameters of the Actor network;
[0082] b) In state s m (t) Execute action a m (t), calculate reward r m (t), to obtain s m (t+1);
[0083] c)(sm (t),a m (t),r m (t),s m (t+1) is stored as an experience in the experience pool and used for experience replay;
[0084] d) Randomly draw experience from the experience pool and input it. and To estimate the Critic network, calculate the Q-value of agent m. enter and Calculate the Q-value of agent m at the next time step in the target Critic network. in, and These represent the states of all agents and their states at the next moment, respectively. and These represent the actions of all agents and their actions at the next moment, respectively. and These represent the parameters of the estimated Critic network and the target Critic network, respectively.
[0085] e) Calculate the timing difference error δ and the loss function
[0086] f) Calculate Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation. This indicates the calculation of the expected value;
[0087] g) Input s m (t) is obtained from the estimated Actor network. Input s m (t+1) to the target Actor network, obtain Where, π' m Indicates taking action a m The strategy of (t+1) Represents the parameters of the target Actor network;
[0088] h) Calculation Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation;
[0089] i) According to and Update separately and Where η∈[0,1] represents the parameter update rate;
[0090] j) Repeat steps a)-i) until the preset number of training iterations are reached, to obtain the trained experience pool and the parameters of the neural network model. and This is the result of offline centralized training of the digital twin.
[0091] The terminal device senses the environmental status online and performs distributed task migration and allocation of computing and communication resources based on a centrally trained Actor-Critic neural network model, including the following steps:
[0092] a) All agents download the offline centralized training results of the digital twin;
[0093] b) All agents perceive the environment, acquire their respective states, and calculate their respective rewards based on the trained neural network parameters, then execute actions in an online distributed manner; where agent m's state s m (t) After being input into its target Actor network, according to the reward r m (t) Output action a m (t), namely: the decision result of the computing type matching between terminal device m and N edge servers, the task migration ratio, the device transmission power and the computing resource allocation result;
[0094] c) All terminal devices perform task migration and collaborative computing based on the output actions of their respective neural networks, i.e., the scheduling results of heterogeneous tasks and resources.
[0095] The present invention has the following beneficial effects and advantages:
[0096] 1. This invention addresses the problem of edge-end collaborative processing of high-concurrency heterogeneous tasks. It fully considers the deadline requirements of heterogeneous tasks, the types of computing resources, the maximum computing power of terminal devices and edge servers, the estimation bias of computing resources in digital twins, the maximum transmit power of terminal devices, and the peak interference power of terminal devices. It proposes an edge-end collaborative scheduling method for heterogeneous tasks and resources based on digital twins, which can meet the service quality requirements of heterogeneous tasks and support edge-end collaborative processing of heterogeneous tasks.
[0097] 2. This invention addresses the challenges of modeling difficulties and algorithm state space explosion caused by the complex coupling of multidimensional resources in heterogeneous computing and communication. It proposes a digital twin-based edge-to-edge collaborative scheduling method for heterogeneous tasks and resources using multi-agent deep reinforcement learning. This method enables centralized offline training and online distributed execution of the scheduling algorithm, minimizing the total processing latency of tasks while meeting the different deadline requirements of heterogeneous tasks. Attached Figure Description
[0098] Figure 1 This is a flowchart of the method of the present invention;
[0099] Figure 2 This is a schematic diagram illustrating a scenario based on digital twins for a single cloud server, multiple edge servers, and multiple terminal devices.
[0100] Figure 3 This is a diagram of the Actor network structure used in this invention;
[0101] Figure 4 This is a diagram of the Critic network structure used in this invention;
[0102] Figure 5 This is a flowchart illustrating the deep reinforcement learning training process for this invention. Detailed Implementation
[0103] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0104] This invention addresses edge-end collaborative processing of large-scale heterogeneous tasks in a single cloud server-multiple edge server-multiple terminal device scenario. It proposes an edge-end collaborative scheduling method for heterogeneous tasks, network communication, and computing resources based on multi-agent deep reinforcement learning. Using this method, on-demand migration of heterogeneous tasks can be supported, enabling edge-end resource collaboration. Under the premise of meeting constraints such as deadlines for heterogeneous tasks, types of computing resources, maximum computing power of terminal devices and edge servers, estimation bias of computing resources in digital twins, maximum transmit power of terminal devices, and peak interference power of terminal devices, the total processing latency of heterogeneous tasks is minimized, supporting edge-end collaborative processing of heterogeneous tasks.
[0105] The proposed method for edge-to-edge collaborative scheduling of heterogeneous tasks and resources based on digital twins includes the following steps: 1) Establishing an edge wireless network based on digital twins; 2) Constructing an edge-to-edge collaborative scheduling problem for heterogeneous tasks and resources according to the deadline requirements of heterogeneous tasks and the constraints of heterogeneous computing and communication resources; 3) Transforming the scheduling problem into a multi-agent Markov decision process problem; 4) Constructing an Actor-Critic neural network model based on multi-agent deep reinforcement learning to solve the multi-agent Markov decision process problem; 5) Offline centralized training of the Actor-Critic neural network model by the digital twin to obtain an experience pool and neural network parameters; 6) Online distributed execution of task migration and allocation of computing and communication resources by terminal devices to collaboratively process heterogeneous tasks and minimize the total task processing latency. The overall process of this invention is as follows: Figure 1 As shown.
[0106] 1) Establish an edge wireless network based on digital twins. For example... Figure 2 As shown, in a physical space, there is one cloud server, N base stations configured with edge servers, and M terminal devices. The cloud server deploys a digital twin, which can mirror all network elements, model the network space, perceive heterogeneous tasks, measure heterogeneous computing and communication resources, train scheduling algorithms, and schedule tasks and resources. The base stations are configured with edge servers to provide computing resources to multiple terminal devices and support scheduling of terminal devices within their coverage area. The terminal devices perform local computation on heterogeneous tasks and support migrating heterogeneous tasks to the edge servers for edge computing via wireless channels.
[0107] A single terminal device performs computations by migrating tasks, either entirely or without migration, to one or more edge servers; the transmission rate during task migration by the terminal device is...
[0108]
[0109] Among them, W m,n This represents the bandwidth between terminal device m and edge server n. G represents the noise at edge server n. m,n and g m',n Let p represent the channel power gain from terminal device m and terminal device m' to edge server n, respectively. m p m' These represent the transmit power of terminal device m and terminal device m', respectively.
[0110] 2) Constructing edge-end collaborative scheduling for heterogeneous tasks and resources
[0111] Terminal device task processing latency T m Delay calculated from the edge and local computing latency The decision and calculation method are as follows:
[0112]
[0113] a) Edge computing latency Calculated as
[0114]
[0115] in, This represents the edge computing latency of edge server n to terminal device m subtasks, which is determined by communication latency. and computational delay Decision, calculated as
[0116]
[0117] Communication latency in edge computing The amount and rate of task migration determined by the terminal device are calculated as follows:
[0118]
[0119] Among them, D m Indicates the task size of terminal device m;
[0120] computation latency of edge computing The task migration amount of terminal device m and the computing resources f allocated by edge server n to terminal device m. m,n Decision, calculated as
[0121]
[0122] The edge computing latency estimated by the digital twin is calculated as follows:
[0123]
[0124] Among them, C m This indicates the computation cycle required to compute a 1-byte task;
[0125] It is the deviation between the actual calculation delay and the estimated calculation delay, calculated as...
[0126]
[0127] b) Local computation latency Calculated as
[0128]
[0129] It is the local computation latency estimated by the digital twin, calculated as
[0130]
[0131] Among them, f m =F max,m -Δf m Indicates local computing resources;
[0132] This is the local calculation delay deviation, calculated as follows:
[0133]
[0134] Based on the deadline requirements of heterogeneous tasks and the constraints of heterogeneous computing and communication resources, and with the objective of minimizing the total latency of task processing, the joint scheduling problem of heterogeneous tasks and network computing and communication resources is constructed as follows:
[0135]
[0136]
[0137] C2:0≤p m ≤P max m=1,...M,
[0138]
[0139]
[0140] C5:0≤f m,n +Δf m,n ≤F max,n ,
[0141]
[0142] C7:T m ≤T max,m
[0143] in, Let be the set of variables to be optimized in the problem, representing the computing resource type matching decision, task migration ratio, terminal device transmit power, and edge server computing resource allocation, respectively. The goal of the problem is to minimize the total task completion time.
[0144] C1 is the task migration ratio constraint; where v m,n ∈[0,1] is the proportion of tasks migrated from terminal device m to edge server n, v m,n =0 indicates that the terminal device m has not migrated the task to the edge server n, v m,n =1 indicates a task migrating from terminal device m to edge server n, v m,0 =0 indicates that terminal device m does not perform local calculations, v m,0 =1 indicates that terminal device m performs local calculations;
[0145] C2 and C3 are transmit power constraints; where P max I represents the maximum transmit power of the terminal device. p This indicates the peak interference power that the terminal device can tolerate. and Representing the connection from terminal device m and terminal device m' to terminal device m respectively. * The channel gain. Where, m * =argmax g m,m' It is the terminal device that is given the maximum interference constraint m;
[0146] C4 represents the matching decision constraint for heterogeneous computing resource types; where o n With o m These represent the computing resource types of the terminal device m and the edge server n, respectively. Indicates the XOR operation; u m,n =1 indicates that the terminal device m and the edge server n have the same type of computing resources; u m,n =0 indicates that the terminal device m and the edge server n have different types of computing resources;
[0147] C5 and C6 are computational resource constraints; where f m,n Represents the edge computing resources estimated by the digital twin; Δf m,n This indicates that digital twins may have biases in their estimation of computational resources; F max,n This represents the maximum computing speed of edge server n;
[0148] C7 represents the task deadline constraint; where T... max,m This indicates the deadline for the task to be executed by terminal device m, that is, the longest task processing time that terminal device m can accept.
[0149] 3) Problem transformation based on multi-agent Markov decision process
[0150] a) Establish a multi-agent Markov decision model, including the agent set, state space, action space, state transition probability, and reward function;
[0151] An intelligent body set is an intelligent body set formed by M terminal devices.
[0152] The state space is the state of agent m at time t, denoted as:
[0153]
[0154] Among them, D m (t) represents the task size of terminal device m; C m (t) represents the number of computation cycles required for terminal device m; T max,m (t) represents the task deadline of terminal device m; Δf m (t) represents the estimation bias of the local computing resources of the terminal device m; W represents the estimation bias of computing resources for N edge servers of terminal device m; m (t)={W m,1 (t),...,W m,N (t)} and G m (t)={g m,1 (t),...,g m,Ns(t) and s(t) represent the bandwidth and channel gain between the terminal device m and the N edge servers, respectively; the total state space of all agents at time t is s(t) = {s1(t), ... s2(t)}. M (t)};
[0155] The action space is the action performed by agent m at time t, denoted as:
[0156] a m (t)={u m (t),v m (t),p m (t),f m (t)}
[0157] Among them, u m (t)={u m,1 (t),...,u m,N (t)} represents the computing resource type matching decision, determining whether the computing resource type of the edge server is consistent with that of the terminal device m; v m (t)={v m,0 (t),v m,1 (t),...,v m,N (t)} represents the proportion of tasks migrated between terminal device m and N edge servers; p m (t) represents the transmit power of terminal device m used for mission migration; f m (t)={f m,1 (t),...,f m,N {a(t)} represents the computing resources allocated to terminal device m by N edge servers; the total action space of all agents at time t is a(t) = {a1(t), ... a2(t)}. M (t)};
[0158] The state transition probability is when agent m executes action a. m When (t), s m (t) is transferred to s m The probability of (t+1), i.e., z m (s m (t+1); s m (t),a m (t));
[0159] The reward function is the reward or penalty for an agent to take an action in a given state, denoted as r. m (t); where the individual reward obtained by agent m is ρ m This represents the weight parameters set based on the deadline requirements of heterogeneous tasks; the latency reward is... Deadline reward is
[0160] b) Determine the long-term cumulative reward function as follows
[0161]
[0162] Where t represents the current time, t0 represents the previous time, and γ m ∈[0,1] represents the discount factor, indicating the influence of past rewards on the current reward of agent m;
[0163] c) Transform the problem into
[0164] max R m (t)
[0165] stC1,C2,C3,C4,C5,C6
[0166] Under the constraints C1-C6, the strategy is to maximize the long-term cumulative reward to obtain the optimal state transition probability, thereby minimizing the total task processing latency.
[0167] 4) Constructing an Actor-Critic neural network model based on multi-agent deep reinforcement learning.
[0168] The Actor network and the Critic network are respectively as follows: Figure 3 and Figure 4 As shown. The Actor is used to generate agent actions, and the Critic is used to guide the Actor to generate better actions; the Actor includes an estimated Actor network for training and a target Actor network for performing actions; the Critic includes an estimated Critic network and a target Critic network for evaluating the Actor's actions.
[0169] The Actor network employs a policy-based deep neural network, while the Critic network uses a value-based deep neural network. The Actor network consists of one input layer, three fully connected layers, one softmax layer, and one output layer. For the first two hidden layers, the ReLU function is used as a non-linear approximate activation function. For the last hidden layer, Tanh is used as the activation function to constrain the actions. After passing through the softmax layer, the sum of the output probabilities for each action is 1. Then, an action is selected as the final output action 'a'. m (t). The Critic network consists of one input layer, three fully connected layers, and one output layer, with the activation function of the first two hidden layers being ReLU.
[0170] 5) Offline centralized training of digital twin neural network models
[0171] To obtain a strategy that minimizes the total processing latency of the task, such as Figure 5 As shown, the centralized offline training of the Actor-Critic neural network model for digital twins includes the following steps:
[0172] a) Input s m (t) to estimate the Actor network, obtain Where, π m Indicates taking action a m The strategy of (t), This indicates the estimation of the parameters of the Actor network;
[0173] b) In state s m (t) Execute action a m (t), calculate reward r m (t), to obtain s m (t+1);
[0174] c)(s m (t),a m (t),r m (t),s m (t+1) is stored as an experience in the experience pool and used for experience replay;
[0175] d) Randomly draw experience from the experience pool and input it. and To estimate the Critic network, calculate the Q-value of agent m. enter and Calculate the Q-value of agent m at the next time step in the target Critic network. in, and These represent the states of all agents and their states at the next moment, respectively. and These represent the actions of all agents and their actions at the next moment, respectively. and These represent the parameters of the estimated Critic network and the target Critic network, respectively.
[0176] e) Calculate the timing difference error δ and the loss function
[0177] f) Calculate Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation. This indicates the calculation of the expected value;
[0178] g) Input sm (t) is obtained from the estimated Actor network. Input s m (t+1) to the target Actor network, obtain Where, π' m Indicates taking action a m The strategy of (t+1) Represents the parameters of the target Actor network;
[0179] h) Calculation Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation;
[0180] i) According to and Update separately and Where η∈[0,1] represents the parameter update rate;
[0181] j) Repeat steps a)-i) until the preset number of training iterations are reached, to obtain the trained experience pool and the parameters of the neural network model. and This is the result of offline centralized training of the digital twin.
[0182] 6) Online distributed execution of tasks on terminal devices, including migration and allocation of computing and communication resources.
[0183] Based on the strategy of minimizing total task latency, terminal devices perform distributed online wireless communication and task migration to collaboratively process heterogeneous tasks, including the following steps:
[0184] a) All agents download the training results of the digital twin and input them into their own neural networks;
[0185] b) All agents perceive the environment, acquire their respective states, calculate their respective rewards based on the parameters of the trained neural network model, and execute actions in an online distributed manner; where the state s of agent m is... m (t) After being input into its target Actor network, according to the reward r m (t) Output action a m (t), which is: the result of the computing type matching decision between terminal device m and N edge servers, the task migration ratio, the device transmission power and the result of computing resource allocation.
[0186] c) All terminal devices perform task migration and collaborative computing based on the output actions of their respective neural networks, i.e., the scheduling results of heterogeneous tasks and resources.
Claims
1. A heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins, characterized in that, The collaborative scheduling of heterogeneous tasks and heterogeneous computing and communication resources based on multi-agent deep reinforcement learning includes the following steps: 1) Establish an edge wireless network based on digital twins; 2) Based on the deadline requirements of heterogeneous tasks and the constraints of heterogeneous computing and communication resources, construct the end-edge collaborative scheduling problem of heterogeneous tasks and resources. 3) Transform the edge-end collaborative scheduling problem into a multi-agent Markov decision process problem; 4) Construct an Actor-Critic neural network model based on multi-agent deep reinforcement learning to solve the multi-agent Markov decision process problem; 5) The digital twin is trained offline in a centralized manner to obtain the experience pool and neural network parameters; 6) Terminal devices perceive the environmental status online and perform distributed task migration and allocation of computing and communication resources based on the centrally trained Actor-Critic neural network model to collaboratively process heterogeneous tasks and minimize the total task processing latency; The problem of edge-to-edge collaborative scheduling of heterogeneous tasks and resources is as follows: in, Let T be the objective of the problem, representing minimizing the total task processing latency. m This represents the task processing latency of terminal device m. Let be the set of variables to be optimized in the problem, representing the computing resource type matching decision, task migration ratio, terminal device transmit power, and edge server computing resource allocation, respectively. C1 is the task migration ratio constraint; where v m,n ∈[0,1] is the proportion of tasks that terminal device m migrates to edge server n; C2 and C3 are the transmit power constraints for the terminal equipment; where, P max I represents the maximum transmit power of the terminal device. p This indicates the peak interference power that the terminal device can tolerate. and Representing the connection from terminal device m and terminal device m' to terminal device m respectively. * The channel gain; where m * =argmaxg m,m' It is the terminal device that causes the most interference to terminal device m; C4 represents the matching decision constraint for heterogeneous computing resource types; where o m With o n These represent the computing resource types of the terminal device m and the edge server n, respectively. Represents the XOR operation; u m,n =1 indicates that the terminal device m and the edge server n have the same type of computing resources; u m,n =0 indicates that the terminal device m and the edge server n have different types of computing resources; C5 and C6 are computational resource constraints; where f m,n Represents the edge computing resources estimated by the digital twin; Δf m,n F represents the bias in the estimation of computational resources for a digital twin; max,n This represents the maximum computing speed of edge server n; C7 is the task deadline constraint; where T max,m This represents the deadline for the task to be executed by terminal device m, which is the longest task processing delay that terminal device m can accept.
2. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The digital twin-based edge wireless network includes: N base stations configured with edge servers and M terminal devices; The base station is configured with an edge server to provide computing resources for multiple terminal devices and to support the scheduling of terminal devices within its coverage area. The terminal device is used to perform local computation on heterogeneous tasks and supports migrating heterogeneous tasks to edge servers for edge computation via wireless channels. The digital twin, placed on the network's cloud server, is represented as a virtualized model of the network containing base stations and terminal devices. It is used to evaluate the operating status, computing resource types, and quantity of computing and communication resources of base stations, edge servers, and terminal devices, and supports the training of deep reinforcement learning methods to carry out network edge-end collaborative scheduling.
3. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 2, characterized in that, For a single terminal device, computation is performed by migrating tasks, either not, partially, or entirely, to one or more edge servers. The transmission rate of the terminal device during task migration is: Among them, W m,n This represents the bandwidth between terminal device m and edge server n. G represents the noise at edge server n. m,n and g m',n Let p represent the channel power gain from terminal device m and terminal device m' to edge server n, respectively. m p m' These represent the transmit power of terminal device m and terminal device m', respectively.
4. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The task processing latency of the terminal device is determined by the edge computing latency. and local computing latency The decision and calculation method are as follows: The edge computing latency Calculated as in, This represents the edge computing latency of the task migrated from edge server n to terminal device m, and the communication latency caused by the task migration. Computational latency for task processing Decision, calculated as Communication latency of the task migration The amount and rate of task migration determined by the terminal device are calculated as follows: Among them, D m Indicates the task size of terminal device m; The computational delay of the task processing The task migration amount of terminal device m and the computing resources f allocated by edge server n to terminal device m. m,n Decision, calculated as The edge computing latency estimated by the digital twin is calculated as follows: Among them, C m This indicates the computation cycle required to compute a 1-byte task; It is the deviation between the actual calculation delay and the estimated calculation delay, calculated as... The local computation latency Calculated as It is the local computation latency estimated by the digital twin, calculated as Among them, f m =F max,m -Δf m Indicates local computing resources; This is the local calculation delay deviation, calculated as follows:
5. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The process of transforming the edge-end collaborative scheduling problem into a multi-agent Markov decision process problem includes the following steps: a) Establish a multi-agent Markov decision model, including the agent set, state space, action space, state transition probability, and reward function; The intelligent ensemble is an intelligent ensemble formed by M terminal devices. The state space is the state of agent m at time t, denoted as: Among them, D m (t) represents the task size of terminal device m; C m (t) represents the number of computation cycles required for terminal device m; T max,m (t) represents the task deadline of terminal device m; Δf m (t) represents the estimation bias of the local computing resources of the terminal device m; W represents the estimation bias of computing resources for N edge servers of terminal device m; m (t)={W m,1 (t),...,W m,N (t)} and G m (t)={g m,1 (t),...,g m,N s(t) and s(t) represent the bandwidth and channel gain between the terminal device m and the N edge servers, respectively; the total state space of all agents at time t is s(t) = {s1(t), ... s2(t)}. M (t)}; The action space refers to the actions performed by agent m at time t, denoted as a. m (t)={u m (t),v m (t),p m (t),f m (t)} Among them, u m (t)={u m,1 (t),...,u m,N (t)} represents the computing resource type matching decision, determining whether the computing resource type of the edge server is consistent with that of the terminal device m; v m (t)={v m,0 (t),v m,1 (t),...,v m,N (t)} represents the proportion of tasks migrated between terminal device m and N edge servers; p m (t) represents the transmit power of terminal device m used for mission migration; f m (t)={f m,1 (t),...,f m,N {a(t)} represents the computing resources allocated to terminal device m by N edge servers; the total action space of all agents at time t is a(t) = {a1(t), ... a2(t)}. M (t)}; The state transition probability is given when agent m executes action a. m When (t), s m (t) is transferred to s m The probability of (t+1), i.e., z m (s m (t+1); s m (t),a m (t)); The reward function is the reward or punishment for an agent taking an action in a certain set state, denoted as r. m (t); where the individual reward obtained by agent m is ρ m This represents the weight parameters set based on the deadline requirements of heterogeneous tasks; the latency reward is... Deadline reward is b) Determine the long-term cumulative reward function as follows Where t represents the current time, t0 represents the previous time, and γ m ∈[0,1] represents the discount factor, indicating the influence of past rewards on the current reward of agent m; c) Transform the problem into max R m (t) stC1,C2,C3,C4,C5,C6 Under the constraints C1-C6, the strategy is to maximize the long-term cumulative reward to obtain the optimal state transition probability, thereby minimizing the total task processing latency.
6. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The Actor-Critic neural network model, constructed based on multi-agent deep reinforcement learning, includes an Actor network and a Critic network. The Actor network employs a policy-based deep neural network, comprising an estimating Actor network for training and a target Actor network for executing actions to generate agent actions. The Critic network employs a value-based deep neural network, including an estimation Critic network and a target Critic network, to evaluate the Actor's actions and guide the Actor to produce better actions.
7. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The offline centralized training of the digital twin neural network model includes the following steps: a) Input s m (t) to estimate the Actor network, obtain Where, π m Indicates taking action a m The strategy of (t), This indicates the estimation of the parameters of the Actor network; b) In state s m (t) Execute action a m (t), calculate reward r m (t), to obtain s m (t+1); c)(s m (t),a m (t),r m (t),s m (t+1) is stored as an experience in the experience pool and used for experience replay; d) Randomly draw experience from the experience pool and input it. and To estimate the Critic network, calculate the Q-value of agent m. enter and Calculate the Q-value of agent m at the next time step in the target Critic network. in, and These represent the states of all agents and their states at the next moment, respectively. and These represent the actions of all agents and their actions at the next moment, respectively. and These represent the parameters of the estimated Critic network and the target Critic network, respectively. e) Calculate the timing difference error δ and the loss function f) Calculate Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation. This indicates the calculation of the expected value; g) Input s m (t) is obtained from the estimated Actor network. Input s m (t+1) to the target Actor network, obtain Where, π' m Indicates taking action a m The strategy of (t+1) Represents the parameters of the target Actor network; h) Calculation Update parameters in, Indicates in the parameter Below, regarding the loss function Perform stochastic gradient descent calculation; i) According to and Update separately and Where η∈[0,1] represents the parameter update rate; j) Repeat steps a)-i) until the preset number of training iterations are reached, to obtain the trained experience pool and the parameters of the neural network model. and This is the result of offline centralized training of the digital twin.
8. The heterogeneous task and resource edge-to-edge collaborative scheduling method based on digital twins according to claim 1, characterized in that, The terminal device senses the environmental status online and performs distributed task migration and allocation of computing and communication resources based on a centrally trained Actor-Critic neural network model, including the following steps: a) All agents download the offline centralized training results of the digital twin; b) All agents perceive the environment, acquire their respective states, and calculate their respective rewards based on the trained neural network parameters, then execute actions in an online distributed manner; where agent m's state s m (t) After being input into its target Actor network, according to the reward r m (t) Output action a m (t), namely: the decision result of the computing type matching between terminal device m and N edge servers, the task migration ratio, the device transmission power and the computing resource allocation result; c) All terminal devices perform task migration and collaborative computing based on the output actions of their respective neural networks, i.e., the scheduling results of heterogeneous tasks and resources.
Citation Information
Patent Citations
Wireless sensor network resource scheduling method and system based on digital twinning
CN113810953A
Calculation and communication resource joint allocation method of industrial wireless network
CN115413044A