Wireless computing power network dynamic task intelligent scheduling method
By using the SAC algorithm that minimizes task completion delay in wireless computing power network, the task scheduling strategy is optimized, and the task scheduling problems caused by uncertainty in the wireless communication environment and node mobility are solved, and efficient completion of computing tasks and rational utilization of resources are achieved.
Patent Information
- Application Number
- CN202510544609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Due to the uncertainty of the wireless communication environment and node mobility in the wireless computing power network, task scheduling is difficult to meet the real-time requirements, resource utilization is insufficient, and the calculation task completion time is long.
The Soft Actor-Critic (SAC) algorithm based on task completion delay minimization is adopted, combined with deep reinforcement learning, the task scheduling strategy in the wireless computing power network is optimized. By constructing the Markov decision-making process, considering the wireless channel capacity and node mobility, an optimization model is established to achieve efficient scheduling of computing tasks.
It effectively reduces the average completion delay of computing tasks in wireless computing power networks, improves resource utilization efficiency, and adapts to task scheduling needs in dynamic environments.
Smart Images

Figure CN120417092A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computing power resource scheduling, and particularly relates to a computing power scheduling technology. Background Art
[0002] With the rapid development of artificial intelligence technology, new services such as intelligent healthcare, autonomous driving, and image recognition have generated computing requirements with higher demands for the computing power of nodes and the real-time completion of tasks. In traditional networks, it is easy to encounter situations where the computing resources of nodes are in short supply and the computing requirements cannot be met. At the same time, the generation of these computing requirements is often accompanied by suddenness in time distribution and sparsity in spatial distribution. In the case of task blocking at some nodes, some computing nodes are idle, resulting in insufficient resource utilization and longer task completion times. The continuous innovation of 5G advanced communication technology makes it possible to achieve dynamic scheduling of computing tasks between nodes through communication. In order to make full use of computing power and network resources, major operators have put forward the concept of Computing Power Network (CPN), deeply integrating the computing power resources of computing, network, cloud, edge, and terminal, aiming to achieve on-demand scheduling of computing and network resources and comprehensively improve resource utilization efficiency and task completion performance.
[0003] Meanwhile, the number of global wireless terminal accesses has increased sharply. It is expected that by 2025, the number of global mobile device accesses will reach 8.5 billion. However, most current research on computing power networks is for large-scale computing power networks in wired data centers and is difficult to cope with the computing requirements in the current widespread wireless environment. Therefore, it is necessary to study the task scheduling problem of computing power networks in the wireless environment. Against this background, the Wireless Computing Power Network (WCPN) has emerged.
[0004] In a wireless computing power network, due to the uncertainty of the wireless communication environment and the limitation of wireless channel capacity, the transmission performance of node computing tasks is significantly affected, resulting in difficulty in accurately predicting the transmission time in task scheduling. If only nodes with larger computing power are selected based on node computing power, it may cause a shortage of communication resources and further affect the task completion efficiency. There is little existing research on wireless computing power networks, and the standard community has not yet formed a unified concept definition. There is also little academic research on it, and the number of literature is at a single-digit level. The impact of node mobility on task scheduling in the wireless environment has not been considered in existing research. However, the mobility of nodes in the wireless environment may cause dynamic changes in the position and number of wireless access points, resulting in poor performance of the scheduling scheme generated in the static environment in the dynamic environment and difficulty in meeting actual application requirements. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a dynamic task intelligent scheduling method for a wireless computing power network, which realizes the optimization of the node computing task scheduling strategy in the wireless computing power network and effectively ensures the high efficiency of computing task completion.
[0006] The technical solution adopted by the present invention is: a dynamic task intelligent scheduling method for a wireless computing power network, and the application scenarios include the cloud, several base stations, edge servers configured for each base station, and several user terminals; the cloud is connected to the base stations through a wide area network or a backbone network, the base stations are connected through a wired channel, and the edge server is connected to its corresponding server;
[0007] The cloud, each edge server, and each user terminal all have computing capabilities. When processing a computing task, the cloud, each edge server, and each user terminal are all regarded as computing nodes;
[0008] The specific scheduling method includes;
[0009] S1. Establish an optimization model for minimizing the completion delay of computing tasks;
[0010] Divide a continuous time period into Τ time slots, and the length of a single time slot is denoted as Δτ. The user terminal generates a computing task at the beginning of each time slot. During the generation of the computing task, a subtask is generated in each time slot. The subtasks of the same computing task have the same requirements for computing and communication resources, and the requirements between the subtasks of different computing tasks are different. The scheduling goal of the subtasks of the same task is the same, and the scheduling goal of the subtask is a certain computing node;
[0011] According to the communication delay, queuing delay, computing delay, and backhaul delay of a single subtask, calculate the overall completion delay of the single subtask;
[0012] Calculate the completion delay of each computing task according to the overall completion delay of a single subtask;
[0013] Considering that the sum of the completion delays of all computing tasks is as small as possible, and taking the task scheduling scheme as the optimization variable, construct an optimization model for minimizing the completion delay of computing tasks;
[0014] S2. Model the optimization model for minimizing the completion delay of computing tasks in step S1 as a Markov decision process and solve it based on a deep reinforcement learning algorithm.
[0015] The beneficial effects of the present invention: Aiming at the task scheduling problem in the wireless computing power network, the present invention considers the finiteness of the wireless channel capacity and the mobility of nodes, and takes them into account in the communication model, and accordingly establishes an optimization model. Thus, it can guide the task scheduling in the wireless computing power network to be more reasonable and the task completion to be more efficient.
[0016] For the established optimization model of minimizing the task completion delay in the wireless computing power network, the present invention also proposes a Soft Actor-Critic (SAC) algorithm based on minimizing the task completion delay to solve the model. This algorithm solves the problems that it is difficult for the wireless computing power network to obtain effective immediate rewards and the action space dimension disaster with the increase in the number of nodes, and effectively reduces the average completion delay of the computing tasks in the wireless computing power network. Description of the Drawings
[0017] Figure 1 is a scenario diagram of task scheduling in the wireless computing power network;
[0018] Figure 2 is the SAC algorithm framework based on minimizing the task completion delay;
[0019] Figure 3 is a schematic diagram of task generation at user nodes;
[0020] Figure 4 is a graph of the reward convergence process during the training process;
[0021] Figure 5 is a graph of the scheduling delay of the proposed algorithm and the comparison algorithm in the wireless computing power network. Detailed Implementation Manner
[0022] To facilitate those skilled in the art to understand the technical content of the present invention, the content of the present invention will be further explained below with reference to the accompanying drawings.
[0023] Considering that in the wireless computing power network, the wireless communication environment is uncertain and nodes move, and due to the suddenness of the time distribution and the sparsity of the spatial distribution of computing tasks, the uneven computing resources of nodes cause waste of computing resources and low efficiency of task completion. The present invention designs a task scheduling method based on the SAC algorithm for minimizing the task completion delay to optimize the average delay of computing task completion. The following is a specific description of the technical solution:
[0024] 1. Scenario Description
[0025] In the current and foreseeable future 5G era, mobile devices mainly access cellular base stations wirelessly, while base stations rely on high-bandwidth and low-latency wired channels such as optical fibers for interconnection to achieve high-speed and stable data transmission, and support efficient network resource management and scheduling. Under this architecture, the wireless computing power network not only endows base stations with traditional access and management functions, but also gives them the capabilities of computing power perception, resource scheduling, and intelligent collaboration. Base stations can efficiently manage the information of access devices and perform real-time data interaction through low-latency wired communication to achieve intelligent allocation and scheduling optimization of computing tasks. Thus, the constructed wireless computing power network scenario and its key components are as follows Figure 1 shown.
[0026] As Figure 1 shown, the wireless computing power network consists of key components such as the cloud, edge servers, and user nodes. Among them, user nodes include terminal devices such as smartphones, autonomous driving vehicles, and drones. These nodes are not only the initiators of computing tasks but also have certain computing, storage, and communication capabilities. Their spatial locations are usually mobile, and they are connected to communication base stations in the corresponding areas through wireless communication. Base stations are deployed at fixed locations and are equipped with edge servers. Compared with user nodes, edge servers have stronger computing, storage, and communication capabilities and can be used as the offloading targets of computing tasks to relieve the computing pressure of terminal devices and improve task processing efficiency. Edge servers are connected to base stations, and base stations are usually connected through wired channels to achieve coordinated scheduling and optimized utilization of computing power resources within the region. In addition, the system also includes a cloud computing center, which is denoted as the cloud in this embodiment. The cloud has ultra-large-scale computing and storage resources and can provide high-performance computing capabilities. However, since the cloud is usually deployed in locations far from users and edge servers, data transmission needs to pass through a wide area network (WAN) or a backbone network. Therefore, communication with the cloud is usually accompanied by high latency and bandwidth consumption. Under the wireless computing power network architecture, how to efficiently coordinate the computing resources among the cloud, edge servers, and user nodes and optimize the task scheduling strategy to improve computing performance and reduce latency is one of the key issues in current research.
[0027] In this scenario, the task scheduling of the current wireless computing power network usually assumes that tasks are completely generated at a certain point in time. In fact, in applications, there is often another typical scenario, that is, only a small part of a task is generated when the task is created, and the subsequent content of the task is gradually generated over time. For example, in an intelligent monitoring and security system, a camera periodically collects environmental image information, and this image information needs to be further analyzed to meet the monitoring requirements. Due to the front-back correlation of the image information, each image cannot be analyzed independently during the analysis. Therefore, the analysis of each image can be regarded as a subtask, and the serial analysis of a series of images that are related to each other before and after can be regarded as a complete task composed of several subtasks.
[0028] In this type of scenario, subtasks are not generated simultaneously at the initial moment, but are gradually generated over time. If the computing resources of the camera are insufficient, these subtasks need to be scheduled to other nodes for calculation. This poses three new requirements for the task scheduling of the wireless computing power network: (1) The completion of subtasks is an ordered serial sequence, and the order cannot be disrupted; (2) Each subtask under the same task should be scheduled to the same target device to ensure the front-back correlation of subtasks and avoid frequent communication between subtasks. In a mobile network environment, this is a great challenge to maintaining the efficiency of scheduling; (3) There needs to be a good match between the time interval when subtasks are generated and the time length for subtasks to be completed.
[0029] Based on the above scenario, considering that there are N user nodes in the wireless computing power network, denoted as U = {u1, u2, ···, u N}, any user node is represented as where represents the CPU computing power of the user node, represents the location of the current user node. Usually, the location of the user node will change, represents the rated transmission power of the user node.
[0030] Starting at time 0, select a continuous time period from 0 to t end where t end is large enough to cover various typical states of network operation, such as dynamic changes in network topology, sudden traffic surges, load imbalance, etc. Divide this time period into Τ time slots, and the length of a single time slot is t end / Τ, denoted as Δτ. The time slot set can be expressed as {1, 2, …, Τ}. User nodes generate computing tasks at the beginning of each time slot, and the generation of this computing task will last for time slots, that is, a subtask is generated every Δτ, for a total of subtasks. User node u iThe task generated at the beginning of time slot τ can be expressed as The k-th sub-task can be expressed as The sub-tasks of each task have the same requirements for computing and communication resources, while the sub-tasks of different tasks have different requirements. That is, the computing and communication requirements of the sub-tasks are independent of k. Therefore, is used to represent the number of floating-point computations required to complete each sub-task; is used to represent the size of the communication transmission volume of each sub-task, with the unit of byte, which affects its transmission delay and bandwidth requirements in the wireless channel; The computing task can be scheduled to any node, and the sub-tasks of the same task have the same scheduling goal, which can be represented by to represent the computing task of the scheduling object.
[0031] The computing tasks generated by the user nodes can be completed not only on the nodes themselves, but also on the edge servers and in the cloud. Considering that there are M edge servers and 1 cloud in the scenario. The edge servers are represented as E = {e1, e2, ···, e M}, where any one of the edge servers can be represented as e j , j ∈ {1, 2, …, M}, represents the CPU computing power of the edge server, represents the location of the edge server. Usually, the edge server is close to the base station and its location does not change. Each edge server is connected to a base station, denoted as BS j . In the present invention, it is assumed that any edge server is connected to a base station in a one-to-one relationship, and it is considered that the edge server and its connected base station have the same spatial location, that is, the base station BS j is located at . And B j is used to represent the bandwidth corresponding to the base station BS j .
[0032] The cloud located at the far end is denoted as C = {c}, and c c is used to represent the computing power of the cloud server, which is usually much higher than that of the edge server and the user node, and is used to represent the typical uplink communication rate of the cloud server, and is used to represent the typical downlink communication rate of the cloud server. Since the cloud computing center is usually deployed in an area far from the users and edge servers, when the task is uploaded to the cloud for computing, it will be affected by a large communication delay.
[0033] 2. Optimization model for minimizing the task completion delay in the wireless computing power network under the node movement scenario
[0034] (1) Communication model
[0035] In a wireless computing power network, when a user node performs task scheduling and selects a target node to complete the task, it may choose to complete the task by itself or by other nodes. The communication cost of the former is 0, while the latter will introduce additional communication delay. During this process, the transmission of the computing task will introduce additional communication delay. Assume that the computing task generated by node i is transferred to the computing node for computing, and the user node u i is connected to the base station BS j . The communication rate between the user equipment and the base station will directly affect the transmission efficiency of task offloading. Considering the Orthogonal Frequency Division Multiple Access (OFDMA) technology widely used in wireless communication systems, the transmission rate between the user equipment u i and the base station BS j can be expressed as:
[0036]
[0037] where B ij (t) is the bandwidth allocated to the user equipment u j by the base station BS i in the t time slot, I represents the interference from other base stations, and σ 2 represents the noise power. Since the bandwidth resources in the wireless communication system are limited and multiple user equipments share the bandwidth of the base station, let B j represent the total bandwidth of the base station BS j . Therefore, there are the following constraints on bandwidth allocation:
[0038]
[0039] In formula (1), is the transmit power adopted by the user equipment u i in the t time slot. In the present invention, power control of the user equipment is not performed, and the user equipment communicates with a rated power ; h ij (t) is the channel gain between the current user equipment u i and the base station BS j in the t time slot; dis ij (t) -α represents the path loss, which is obtained according to the position of the user equipment and the position of the base station . Among them, α represents the path loss exponent, usually taking values between 2 and 6. The specific calculation formula of the path loss is as follows:
[0040]
[0041] The calculation formula of I is as follows:
[0042]
[0043] Among them, represents the rated transmission power of user node u i′ ; h i′j represents the channel gain between node u i′ and base station BS j ; δ(i, i′) is used to represent whether user node u i and user node u i′ use the same subcarrier. If they share the subcarrier, the value of δ(i, i′) is 1; otherwise, the value of δ(i, i′) is 0.
[0044] When the computing task is scheduled in the wireless computing power network, the task may be scheduled within the coverage area of the base station accessed by the user node, or the task may be scheduled to a node outside the coverage area of the base station. When the task is scheduled to a node outside the coverage area of the base station, the transmission path of the task includes the wired link between the base stations. Let φ be the communication transmission rate factor of the wired link, then the transmission rate of the task data from base station BS j to base station BS j′ during the transmission process can be expressed as:
[0045] r j′j (t) = f(φ, dis jj′ (t)) (5)
[0046] Among them, dis jj′ (t) represents the length of the wired link between base station BS j and BS j′ . The specific form of f(·) is affected by physical characteristics such as channel bandwidth and signal quality.
[0047] Formula (1) and formula (5) respectively represent the communication rates of the computing task transmitted through the wireless channel and the wired link in the network. Based on this, according to the computing task scheduling target node type and the base station access relationship, the transmission time for scheduling any subtask to a non-self target node can be expressed in the following three cases:
[0048] (a) Scheduled to a user node or edge server under the same base station
[0049] For any computing subtask generated by user node u i whose scheduling target is If the scheduling target is u and the accessed base station is BS i j The corresponding edge server, whose computing tasks can be completed only by wireless channel transmission; if the scheduled target is another user node and is connected to the same base station as u i Considering that the rate of the base station transmitting tasks to the user node is much greater than the rate of the user node transmitting tasks to the base station, the present invention only considers the time for tasks to be transmitted from the user node to the base station. At this time, at time t, the transmission time required for any sub-task to be scheduled to the target node can be expressed as:
[0050]
[0051] wherein, represents the communication data volume size of each sub-task of the task , and r ij (t) represents the wireless communication rate between the user node u i and the base station BS j .
[0052] (b) Scheduled to user nodes or edge servers under different base stations
[0053] If the sub-task is scheduled to an edge server, but not the edge server corresponding to the base station BS i accessed by u j , or the scheduling target of the sub-task is a user node but is not connected to the same base station as the user node u i , then the sub-task needs to be communicated and transmitted through the wired link between the base stations. At this time, BSj′ is uniformly used to represent the base station corresponding to the target edge server or the base station accessed by the target user node. Then, at time t, the transmission time required for any sub-task to be scheduled to the target node can be expressed as:
[0054]
[0055] wherein, represents the communication transmission size of each sub-task of the task; r ij (t) represents the wireless communication rate between the user node u i and the base station BS j , which can be calculated by formula (1); r jj′ (t) represents the wired link communication rate between the base station BS j and the base station BS j′ , which can be calculated by formula (5).
[0056] (c) Scheduling the cloud server
[0057] If this subtask is scheduled to the remote cloud, in the scenario of the present invention, the communication between different nodes and the cloud is not greatly affected by the spatial location, and the communication rate can be considered a fixed value. Denote the typical uplink communication rate for communicating with the cloud, then for any subtask the transmission time required to be scheduled to the target node can be expressed as:
[0058]
[0059] The above three cases cover all cases of the transmission delay of the task when it is scheduled to a non-self node. And when the task is calculated locally, the transmission time is 0. Then, for any subtask at time t the transmission time required to be scheduled to the target node is uniformly expressed as:
[0060]
[0061] Among them, denotes the scheduling target of the task , and since the scheduling targets of all subtasks are the same, is the scheduling target of subtask ; C represents the set of cloud servers, used to indicate that the scheduling target is the cloud, otherwise it is not; BS j denotes the base station to which the user node is connected, and BS' j denotes the corresponding base station when the target is the edge server or the base station to which the user node is connected when the target is the user node.
[0062] After the task is scheduled, it queues at the destination node until the calculation is completed, and after the calculation is completed, the calculation result is sent back to the original node. In the traditional research on task scheduling in the computing power network, considering that the communication environment is relatively stable and the node mobility is small, it is considered that the ratio of the calculation result to the original task size is very small and the communication downlink rate is much greater than the uplink rate, and the return time of the calculation result is often ignored. However, in the research scenario of the wireless computing power network, the wireless communication environment is unstable and the nodes move relatively frequently, and the return time of the task cannot be ignored. Define the ratio of the calculation result to the original size of the calculation task as β. Then, for any subtask starting to return the calculation result at time t, the time for the calculation result to return is:
[0063]
[0064] In the above formula, denotes the target to which the task is scheduled; U, E, and C respectively represent the set of user nodes, the set of edge servers, and the set of cloud servers in the wireless computing power network; BSj Denote the base station BS′ accessed by user node u at time t i Denote the base station BS′ accessed by user node u at time t j Denote the target node at time t the accessed base station; Denote the target node and base station BS j the communication rate between them, r jj′ (t) denotes the communication rate between base station BS j and base station BS′ j Formula (9) is divided into five cases according to the node type of the target node and whether the task generation node and the target node access the same base station. When the scheduled target node is itself or the edge server corresponding to the base station it accesses, the calculation result is directly sent to the task generation node by the base station. Considering that the downlink communication rate is relatively large, this part of the delay is ignored; when the target node is another user node accessing the same base station, the return delay considers the time for the target node to upload the calculation result to the base station; when the target node is a user node other than itself but the base station it accesses is not the base station of the task generation node, the return delay considers the upload delay of the target node and the transmission delay of the wired link between the base stations; when the target node is an edge server but the corresponding base station is not the base station accessed by the task generation node, the return delay considers the transmission delay of the wired link between the base stations; when the target node is a cloud server, the return delay is affected by the typical downlink communication rate of the cloud server.
[0065] (2) Calculation model
[0066] When a computing task is scheduled to a target computing node, the computing node may still be processing other computing tasks. Therefore, the newly arrived task will be added to the task queue and executed in sequence according to a certain scheduling strategy. Since computing nodes usually adopt First Come First Served (FCFS), the present invention also uses the FCFS method to complete tasks. Therefore, there may be a queuing waiting time before formal calculation, that is, additional queuing delay needs to be introduced.
[0067] For any computing task its subtask needs to be scheduled to the target node at time t for calculation completion. Let the subtask being calculated and processed by the target node at time t be Denote as the start time of the execution of this subtask, as the arrival time of this subtask at the computing node (if the computing queue is idle, it represents the arrival time of the next task). Denote indicating that the subtask arrives at the computing node through communication At this moment, combined with the analysis of the task communication delay in formula (8), and the moment t when this subtask starts to be scheduled has the following relationship:
[0068]
[0069] For the target node Use to represent the set of all subtasks waiting to be completed by calculation at this node. For any subtask d waiting to be processed and completed at the node wait , the present invention assumes that the computing node can obtain and store the moment when each subtask arrives at itself. Denote the moment when subtask d wait arrives at the computing node as t arrive (d wait ), and the number of floating-point calculations required to complete the subtask is ζ(d wait ). Then for any subtask that arrives at the computing node at a moment between and , subtask needs to wait for these tasks to be processed and completed. Therefore, the amount of calculation that subtask needs to wait for the processed and completed tasks is:
[0070]
[0071] Then the queuing delay of this task can be expressed as:
[0072]
[0073] Among them, represents the computing power of the computing node . Before the minus sign in formula (12), represents the execution time required for the total amount of calculation of the tasks to be completed at the node , and after the minus sign, represents the time difference from the start of execution of the current task to the arrival of the new task, that is, the completed part of the tasks being executed in the queue, so it should be deducted. When defining , it is set that when the queue is idle (that is, ), it represents the arrival time of the next task at this node. Substituting it into formula (12) gives the waiting time as 0. Therefore, formula (12) satisfies the case when the queue is idle.
[0074] The calculation time required for the computing task to complete the calculation after arriving at the computing node is as follows:
[0075]
[0076] (3) Average Delay Optimization Model
[0077] Based on the above analysis of the communication model and task computing model during the task completion process, equations (8), (12), (13), and (8) respectively give the communication delay, queuing delay, computing delay, and feedback delay of a single subtask Therefore, the overall completion delay of a single subtask scheduled at time t is as follows:
[0078]
[0079] where t′ represents the time when the calculation result starts to be fed back, and there is the following relationship:
[0080]
[0081] t represents the moment when the subtask starts to be scheduled. The start scheduling times of each subtask of a task are different. Specifically, for the k-th subtask of the task, the start scheduling moment t k can be expressed as:
[0082] t k = τΔτ + (k - 1)Δτ (16)
[0083] In equation (16), t k represents the start scheduling time of the k-th subtask, which has the same meaning as t in equation (14), both representing the start scheduling time of the subtask. Among them, equation (14) reflects the relationship between the start scheduling moment of any subtask and the task completion delay, while equation (16) further specifically represents the start scheduling moment of any subtask, that is, the specific representation of t for a specific subtask in equation (14). Therefore, it can be substituted into equation (14) to obtain the overall completion delay of each subtask, and the completion delays of each subtask are summed to obtain the task completion delay as follows:
[0084]
[0085] Based on the above analysis of the task completion time, within the time slot range of [1, T], the total time for the completion of the computing tasks generated by N user devices can be expressed as:
[0086]
[0087] Considering that the latency of task completion in the wireless computing power network is as small as possible, taking the task scheduling scheme as the optimization variable, the following optimization model with the minimum latency is proposed:
[0088]
[0089] In the above optimization problem, constraint (19a) is the value constraint of the task scheduling expression, indicating that the k-th sub-task of the computing task generated by user node i at the beginning of time slot τ is scheduled to node j for computing completion, and the overall scheduling scheme S is composed of them; constraint (19b) means that all sub-tasks must be scheduled and completed without omission; constraint (19c) further means that all sub-tasks must be scheduled to the same computing node for completion; constraint (19d) represents the bandwidth resource constraint, indicating that the bandwidth resources allocated by the base station to each node cannot exceed its total bandwidth resources; constraint (19e) is the sub-task completion time constraint, indicating that the time for sub-task completion cannot exceed the interval time when the sub-task is generated.
[0090] 3. SAC Scheduling Algorithm Based on Minimizing Task Completion Latency
[0091] The present invention designs an SAC algorithm based on minimizing task completion latency to solve the above problems.
[0092] (1) Basic Idea and Framework
[0093] To solve the above challenges with lower complexity, the present invention models the problem as a Markov decision process with a large state space and solves the problem based on a deep reinforcement learning algorithm. However, traditional reinforcement learning algorithms still face some key difficulties when applied to such problems.
[0094] First, the traditional SAC algorithm gives an immediate reward after the agent makes a decision to optimize the policy during the learning process. In the wireless computing power network, the latency needs to be used to measure the quality of the scheduling action. After the central scheduler gives an action, the environment cannot immediately give the size of the latency at this time, and an effective reward can only be obtained after the task is executed. However, in the initial stage of training, the scheduling strategy of the agent cannot be effectively optimized, which may lead to a large time required for task completion, and there will be a significant delay in the feedback of the reward, which may affect the effect in the initial stage of reinforcement learning training as well as the overall stability and training efficiency. For this reason, the present invention designs a reward truncation mechanism, that is, when the processing latency of the task exceeds the latency of local direct calculation or the interval time when the sub-task is generated, an immediate reward feedback is given to avoid the negative impact of reward delay on training.
[0095] Secondly, the intelligent agent policy network of the traditional SAC algorithm outputs a probability distribution and obtains specific actions through sampling. When making task decisions in the wireless computing power network scenario, the decision-maker gives the scheduling object of any user node. Assuming there are N user nodes, M edge servers, and a cloud in the scenario, that is, there are N + M + 1 computing nodes. Therefore, the size of the scheduling space for any node is N + M + 1, and the size of the action space is N N+M+1 , which will exponentially increase with the increase in the number of nodes. This exponential increase leads to a serious curse of dimensionality, not only increasing the computational complexity of the reinforcement learning model but also affecting the decision-making effect, making it difficult for the intelligent agent to effectively learn in the high-dimensional action space. To address this, the present invention proposes a multi-dimensional action decomposition strategy based on the task generator, which divides the original large-scale actions into multiple small actions, reducing the size of the action space to N(N + M + 1), changing the exponential increase to multiplicative growth, reducing the training difficulty, and improving the efficiency of decision-making.
[0096] To solve the above problems, the present invention proposes a reinforcement learning algorithm based on minimizing the task completion delay of SAC. The framework of this algorithm is as Figure 2 shown. Each node in the wireless computing power network constitutes the environment of reinforcement learning, and the scheduler, as the intelligent agent, gives the scheduling scheme of each node's task in each time slot. Define the state space as S, the action space as A, and the reward function as R. In each training episode, the scheduler observes the current environment to obtain the state s, selects the action a, gives the reward r, and enters the next state s'. The following will introduce the design of the specific state space, action space, and reward function in detail.
[0097] (a) State space
[0098] The state space of the system consists of the states of each node in the wireless computing power network. The types of nodes can be divided into three types: user nodes, edge servers, and the cloud. The initial positions of user nodes are randomly distributed in space and move randomly according to the Markov model, and at the same time generate computing tasks according to the ON - OFF model; edge servers are distributed near the base station and are relatively uniform in space. Any user node can communicate with edge servers through the base station, and edge servers are connected by wire; the cloud is far from user nodes and edge servers, and there is a large communication delay when communicating with the cloud. Therefore, the state space of the system is jointly composed of the states of the three types of nodes, and the specific information is as follows:
[0099] i. User nodes
[0100] The state information of user nodes includes the task load, resource capabilities, and spatial motion characteristics of mobile terminal devices, including:
[0101] · Task characteristics: triple {ki ,ζ i ,D i}, respectively representing the number of task subtasks, the CPU cycle requirement of a single subtask, and the amount of communication data;
[0102] Resource status: Covered transmit power Local computing power and the remaining time of the task to be executed
[0103] Movement state: {x i ,y i ,v i ,θ i}, describing the node plane coordinates (x, y), moving speed v and direction angle θ.
[0104] ii. Edge Server
[0105] Reflects the resource availability and spatial distribution of edge computing nodes, including:
[0106] Resource status: Including communication bandwidth Computing power and the remaining time of the task to be executed
[0107] Position status: {x i ,y i}, describes the node plane coordinates (x, y), and the edge server location is fixed.
[0108] iii. Remote Cloud
[0109] Simplified to typical communication uplink rate Downlink rate And the computing power c c The triples formed Due to the resource pooling characteristics of the cloud, it is assumed that its computing resources are sufficient and the task queuing delay is ignored.
[0110] (b) Action space
[0111] In this invention, when any user node generates a computing task, the central scheduler outputs a corresponding scheduling plan based on the current system state, that is, the actions determined by the agent according to its strategy. Considering that there are N user nodes, M edge servers, and 1 cloud server in the scenario, a total of N+M+1 computing nodes, all computing nodes are numbered from 1 to N+M+1.
[0112] However, directly encoding the scheduling actions for all computing nodes will lead to an exponential growth in the dimensionality of the action space, greatly increasing the solution complexity. Therefore, the present invention adopts a strategy of action decomposition: the agent outputs scheduling sub-actions for each user node respectively, which are called node sub-actions. For each user node, the action space of its sub-action is denoted as a i , all of size N + M + 1, and the values range from 1 to N + M + 1, indicating that the computing tasks generated by this user node are scheduled to the computing node with the corresponding number.
[0113] Therefore, at time t, the overall action given by the agent is the set of sub-actions of each user node, that is:
[0114] a t = {a1, a2, …, a i , …, a N}; a i ∈ {1, 2, …, N + M + 1} (20)
[0115] (c) Reward function
[0116] In the process of reinforcement learning, the behavior of the agent is driven by the reward signal. As the core component of the reinforcement learning algorithm, the reward function plays a decisive role in the update and convergence of the agent's strategy. Considering the characteristics of the wireless computing power network, the system performance is closely related to the delay of computing tasks, and in the actual environment, the delay often has the problem of latency. Therefore, the present invention designs a reward truncation mechanism, that is, when the processing delay of the task exceeds the local direct computing delay or the time interval between subtasks, the reward feedback is given immediately to avoid the negative impact of reward delay on training.
[0117] Therefore, based on the reward truncation mechanism, the present invention designs a comprehensive reward function, which not only reflects the impact of the task completion delay on the system performance, but also takes the local completion delay of the task as a benchmark to accurately evaluate the quality of the scheduling strategy. Specifically, the reward function will fully consider the full-process delay of the computing task from scheduling, transmission, computing to feedback, and encourage the agent to adopt a reasonable truncation strategy while optimizing the overall task completion delay through the construction of delayed rewards to avoid the excessive negative impact of long delays on the reward signal. The reward for taking an action at time slot τ is determined by the delays of all tasks scheduled at this moment, that is:
[0118]
[0119] where The reward value obtained from the k-th subtask of the task generated by node i at time slot τ, which is determined by the completion time of each subtask and is determined by generating an interval Δτ, specifically as follows:
[0120]
[0121] (2) Algorithm flow
[0122] Based on the above analysis of the algorithm framework and the state space, action space, and reward function in reinforcement learning, the present invention optimizes the training process on the basis of the traditional SAC algorithm, improves the problems of too large action space and reward delay, and proposes an optimized SAC algorithm. The algorithm consists of five core networks. Among them, the policy network (Actor Network) is responsible for generating the actions of the agent, enabling it to select the optimal policy as much as possible while exploring the environment; the evaluation network (Critic Network) calculates the Q value of the state-action pair, and two independent Q networks are used to alleviate the problem of overestimation of the Q value, thereby improving the stability and convergence of learning; the target network (Target Network) has the same structure as the evaluation network, mainly used to calculate the stable target Q value, avoid the oscillation phenomenon during the training process, and smooth the parameter changes through the soft update mechanism to further enhance the stability of training.
[0123] In addition, SAC has the dual goals of maximizing rewards and maximizing the actual entropy, and introduces a temperature coefficient (Temperature Coefficient) α t , which is used to control the exploration degree of the policy. It determines the randomness of the policy and is dynamically adjusted by optimizing the objective function during the training process to maintain a balance between exploration and exploitation. Specifically, in the SAC algorithm, the policy objective is defined as follows:
[0124]
[0125] Among them, represents sampling the state s from the experience replay pool t , and the experience replay pool stores the historical state transition records; a t ~π θ represents randomly sampling the action a θ according to the current action distribution π t (·|s t ), and the distribution of π θ (·|s t ) is given by the current policy network according to the current state s t ; in the first term α t in the expectation (-logπ θ (a t |s t)) Encourages high entropy values, i.e., encourages more exploration, while the second term Encourages high returns, i.e., more exploitation. In the SAC algorithm, the weight ratio of the two terms in the policy objective can be adjusted by tuning the temperature coefficient α t to balance between exploration and exploitation. The algorithm first needs to input the maximum number of training epochs and the maximum number of time steps per epoch, and then perform necessary parameter initialization, including five neural network parameters (policy network, two Q evaluation networks, two target networks), temperature coefficient, etc., and empty the experience replay pool. In addition, the environmental parameters of the node need to be set, including the position distribution of the node, motion state parameters, resource status, and task model.
[0126] After initialization, the algorithm enters the training phase. In each training round, the agent samples actions from the policy network (Actor Network) according to the environmental state parameters at the beginning of each time slot and makes scheduling decisions accordingly. Before executing the scheduling action, the time for local computation of the task is calculated and recorded for subsequent decision-making and reward calculation. Then, the agent performs task allocation and computation operations according to the scheduling action output by the policy network. After the task is executed, it is judged whether there is a previously computed task that has been completed or whether the processing time of the current task exceeds the local processing delay or the subtask generation interval. If the above conditions are met, the agent obtains the corresponding reward and records the state transition (s t , a t , r t , s t+1 , γ) into the experience replay pool D for subsequent training. If the computational task does not meet the above conditions, the computational task continues to be processed until it meets the above conditions and is then put into the experience replay pool. Among them, s t represents the environmental state at the current moment, a t represents the currently selected action, r t represents the obtained reward, s t+1 represents the state transferred to after taking the action, and γ represents the discount factor, reflecting the importance of future rewards; to avoid training instability caused by using a small number of samples at the beginning of training, neural network parameter updates do not start until the records in the experience replay pool reach the minimum number required for starting training. Generally, the starting training number is taken as 10% of the experience replay pool capacity. After starting training, the Critic network, Actor network, and Target network are updated in sequence, and the temperature coefficient a t is adjusted to complete one training.
[0127] In the present invention, to reduce the number of hyperparameters, an adaptive adjustment strategy is adopted for the temperature coefficient to make the policy entropy (i.e., the actual entropy) adaptively approach the target entropy The target entropy is usually set to -dim(A), which is the dimension of the action space. To make the actual entropy close to the target entropy define the temperature loss function as:
[0128]
[0129] To minimize the temperature loss function, perform gradient descent on α t i.e., α t is updated as follows:
[0130]
[0131] where λ is the learning rate, is the temperature loss function, denotes the gradient estimate of the loss function at α t When all time slots are executed, the algorithm enters the next round of training, continuously optimizing the scheduling strategy until the agent converges to the optimal decision-making scheme.
[0132] The specific process of the algorithm is as follows:
[0133]
[0134] 4. Implementation Case
[0135] To verify the performance of the SAC algorithm proposed in the present invention, the present invention uses the Python language to write node events under the Simpy framework and uses the Pytorch framework for neural network construction and training. This embodiment will describe the simulation scenario and parameters in detail, and then display and analyze the simulation results.
[0136] (1) Implementation Scenario and Parameter Settings
[0137] To verify the application effect of the SAC algorithm in the wireless computing power network task scheduling scenario, the present invention conducts a simulation modeling of the wireless computing power network scenario. The simulation environment is set in a square area with a side length of L, and the set simulation area size is 1600m × 1600m. There are 20 user devices, 5 edge servers and a cloud in the area, and the user devices and edge servers are evenly distributed in the simulation area.
[0138] Considering that in actual application scenarios, the generation of computing tasks is often sudden, and the sudden traffic has a certain duration. At the same time, the computing requirements are often unevenly distributed in space. Therefore, from the perspective of the entire wireless computing power network, the spatio-temporal distribution of tasks shows dispersion and sparsity. The present invention uses the ON / OFF model to describe the task situation of any computing node, as Figure 3 shown. This model can effectively characterize the dynamic characteristics of user nodes generating computing tasks, that is, in the ON state, the node is in the active period of computing tasks, and at the beginning of each time slot, it will generate computing tasks with a certain probability pr i while in the OFF state, the node has no task generation and is in a silent state.
[0139] In the simulation scenario, each time slot is set to 1 s. At the beginning of the time slot, when the user equipment is in the ON state, computing tasks are randomly generated. The computing tasks can be scheduled to any node to complete the calculation, and the completion time of the task is calculated after the task is completed. The parameter settings of the simulation environment are shown in Table 1 below:
[0140] Table 1 Parameter settings of the wireless computing power network simulation environment
[0141]
[0142] During the simulation process, the actions of the nodes are given by the SAC algorithm. The training method of the neural network is shown in Algorithm 1. Both the policy network and the evaluation network are set to two fully connected hidden layers. The number of neurons in the hidden layer close to the input layer is set to 256, and the number of neurons in the hidden layer close to the output layer is set to 128. The number of training rounds is set to 800 Episodes, and each Episode contains 1000 time slots. Other specific training parameters are shown in Table 2.
[0143] Table 2 SAC algorithm parameter table
[0144]
[0145] (2) Analysis of implementation effectiveness
[0146] Experimental simulations are carried out according to the above simulation environment and network parameters, and the reward value is output during the training process. Its change trend with the training rounds is as Figure 4 shown. In the initial stage, due to the randomization of network parameters, the agent mainly relies on random exploration to select the scheduling object. During this process, affected by the randomness of the scheduling strategy, computing tasks may be assigned to user equipment with limited computing resources, or multiple computing tasks may be centrally scheduled to the same computing node, resulting in the blocking of computing tasks, so that the task completion time is longer. Therefore, in the first 100 training iterations, the reward value feedback by the environment is lower, and the scheduling strategy of the agent has not been effectively optimized.
[0147] As the training progresses, the experience replay pool of reinforcement learning gradually accumulates enough samples, and the formal training of the neural network starts. Guided by the environmental reward feedback, the agent gradually adjusts its decision-making strategy and is more inclined to select scheduling schemes that can obtain higher rewards. After 100 rounds of training iterations, the reward value starts to increase significantly, indicating that the scheduling strategy of the agent is gradually optimized. After further training, when the training iteration reaches about 200 rounds, the reward value tends to converge, and the agent has basically converged to a stable scheduling strategy. At this time, the decision-making scheme given by the neural network can better match the computing and communication resources with the task requirements to ensure the execution of the task. After the training process converges and reaches a stable state, the parameters of the policy network are fixed and tested under the same environmental settings. The test duration is set to 1000 seconds, and the latency values of each computing task are recorded during the test, and the average completion latency of a single task is calculated. Under the same environmental parameters, the common task scheduling schemes are compared. The compared scheduling strategies include the SAC algorithm based on minimizing the task completion latency proposed in the present invention, the greedy scheduling algorithm based on the node load status, the fixed edge server offloading scheme, and the local computing scheme. The average task completion latency of each scheduling strategy obtained is as Figure 5 shown.
[0148] From Figure 5 it can be seen that among the four scheduling schemes, the algorithm proposed in the present invention performs the best, and can control the average completion latency of a single task within 1 s. The average latency of task completion in the test is 0.73 s. The average latency of task completion of other scheduling methods is higher than that of the algorithm proposed in the present invention. Among them, the average latency of task completion of local computing is 2.98 s, and the task completion latencies of the greedy algorithm and the fixed offloading are 1.25 s and 1.57 s respectively. And from the trend of the curve, it can be seen that there is a certain high latency period in the task completion process of the other three scheduling methods. This is because the ON-OFF model is adopted for task generation of nodes in the simulation test of the present invention, and these scheduling schemes fail to efficiently adapt to the dynamics of the distribution of computing tasks in time and space, resulting in an increase in latency for a period of time. However, the algorithm proposed in the present invention can make good use of task dynamics, fully utilize the resources in the wireless computing power network, and efficiently complete computing tasks.
[0149] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A dynamic task intelligent scheduling method for a wireless computing power network, characterized in that The application scenarios include the cloud, several base stations, edge servers configured for each base station, and several user terminals; the cloud is connected to the base stations through a wide area network or a backbone network, the base stations are connected through a wired channel, and the edge server is connected to its corresponding server; The cloud, each edge server, and each user terminal have computing capabilities. When processing a computing task, the cloud, each edge server, and each user terminal are all regarded as computing nodes; The specific scheduling method includes; S1. Establish an optimization model for minimizing the completion delay of computing tasks; Divide a continuous time period into Τ time slots, and denote the length of a single time slot as Δτ. The user terminal generates a computing task at the beginning of each time slot. During the generation of the computing task, a subtask is generated in each time slot. The subtasks of the same computing task have the same requirements for computing and communication resources, while the subtasks of different computing tasks have different requirements. The scheduling goal of the subtasks of the same task is the same, and the scheduling goal of the subtasks is a certain computing node; According to the communication delay, queuing delay, computing delay, and backhaul delay of a single subtask, calculate the overall completion delay of a single subtask; Calculate the completion delay of each computing task according to the overall completion delay of a single subtask; Considering that the sum of the completion delays of all computing tasks is as small as possible, and taking the task scheduling scheme as the optimization variable, construct an optimization model for minimizing the completion delay of computing tasks; S2. Model the optimization model for minimizing the completion delay of computing tasks in step S1 as a Markov decision process and solve it based on a deep reinforcement learning algorithm.
2. The intelligent scheduling method for dynamic tasks of a wireless computing power network according to claim 1, wherein, The optimization model for minimizing the completion delay of computing tasks constructed in step S1 is expressed as: where, represents the total time for the completion of the computing tasks generated by N clients, represents the computing task completion delay, represents the overall completion delay of a single subtask, represents the number of time slots that the generation process of a computing task lasts, represents the scheduling result of the k-th subtask of the computing task generated by the i-th client at the start of the τ time slot being scheduled to the j-th computing node, B j represents the base station BS j total bandwidth, B ij (t) is the bandwidth allocated by the base station BS j to the i-th client at the t time slot, and M represents the number of base stations.
3. The intelligent dynamic task scheduling method for a wireless computing power network according to claim 2, wherein The calculation formula is as follows: Among them, represents the transmission time required for the subtask to be scheduled to the target node, represents the queuing delay of the subtask ; represents the computing time required for the subtask to complete the computation after arriving at the computing node; represents the transmission time required for the subtask to transmit the computation result back from the target node.
4. The intelligent scheduling method for dynamic tasks in a wireless computing power network according to claim 3, characterized in that, The state space of the deep reinforcement learning algorithm in step S2 includes: user terminal state, edge server state, and cloud state; The user terminal state includes task characteristics, user terminal resource state, and motion state. The task characteristics include the number of subtasks, the CPU cycle requirement of a single subtask, and the communication data volume. The user terminal resource state includes the transmission power, the local computing ability of the user terminal, and the remaining time of the task to be executed by the user terminal. The motion state includes the planar coordinates of the user terminal, the moving speed, and the direction angle; The edge server state includes the edge server resource state and the location state. The edge server resource state includes the communication bandwidth, the computing ability of the edge server, and the remaining time of the task to be executed in the edge server; The cloud state includes the uplink rate, downlink rate, and cloud computing power.
5. The intelligent dynamic task scheduling method for a wireless computing power network according to claim 4, wherein In the deep reinforcement learning algorithm in step S2, the agent outputs a scheduling sub-action for each user terminal. Then, at time t, the overall action given by the agent is the set of scheduling sub-actions for each user, that is: a t = {a1, a2, …, a i , …, a N}; a i ∈ {1, 2, …, N + M + 1} Among them, a i represents the action space of the scheduling sub-action of the i-th client.
6. The intelligent dynamic task scheduling method for a wireless computing power network according to claim 5, characterized in that, The optimization objectives in the deep reinforcement learning algorithm in step S2 include maximizing the reward and maximizing the actual entropy.
7. The intelligent scheduling method for dynamic tasks in a wireless computing power network according to claim 6, wherein In the deep reinforcement learning algorithm in step S2, the reward for taking an action in the τ time slot is determined by the delays of all tasks scheduled at this moment, that is: Among them, represents the reward value obtained for the k-th sub-task of the computing task generated by the i-th client in time slot τ, and the calculation formula of Among them, represents the completion time of the subtask , represents the local computing time.
8. The intelligent scheduling method for dynamic tasks of a wireless computing power network according to claim 7, characterized in that, The actual entropy is expressed as: Among them, denotes the expectation, and π θ (·|s t ) is the current action distribution given by the current policy network according to the current state s t . π θ (a t |s t ) represents the probability value corresponding to the action a t given by the current policy network according to the current state s t . a t ~π θ means randomly sampling the action a θ (·|s t ) according to the current action distribution π t .
9. The intelligent dynamic task scheduling method for a wireless computing power network according to claim 8, characterized in that It also includes a temperature coefficient. By adopting an adaptive adjustment strategy for the temperature coefficient, the actual entropy can adaptively approach the target entropy. Define a temperature loss function. It is expressed as: where α t represents the temperature coefficient, and dim(A) represents the dimension of the action space A.
Citation Information
Patent Citations
In-network pooling resource allocation optimization method based on contribution perception in computing power network
CN114401532A
Cloud edge computing power network system computing unloading method and system
CN117749796A
Task decomposition optimization method and device facing end-side computing power network, and storage medium
CN119718657A
Digital twin-based edge-end collaborative scheduling method for heterogeneous tasks and resources
US20250086005A1
Task offloading and resource allocation method based on mobile edge computing
WO2024174426A1