Dynamic Task Intelligent Scheduling Method for Wireless Computing Networks

By employing the Soft Actor-Critic (SAC) algorithm based on minimizing task completion latency in wireless computing networks, the scheduling strategy for computing tasks is optimized, solving the problems of uncertainty in wireless communication environment and channel capacity limitation, and realizing the optimization of task completion latency and efficient utilization of resources.

CN120417092BActive Publication Date: 2025-11-14YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544609.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-11-14
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In wireless computing networks, the transmission performance of node computing tasks is affected by the uncertainty of the wireless communication environment and the limitation of wireless channel capacity. This makes it difficult to accurately predict task scheduling. Static scheduling schemes perform poorly in dynamic environments, resulting in insufficient resource utilization and long task completion times.

Method used

A dynamic intelligent task scheduling method for wireless computing networks is proposed. It adopts the Soft Actor-Critic (SAC) algorithm based on minimizing task completion latency, combined with deep reinforcement learning, to optimize the scheduling strategy of computing tasks. Considering wireless channel capacity and node mobility, a Markov decision process model is established, and task completion latency is optimized through intelligent scheduling.

Benefits of technology

It effectively reduces the average completion latency of computing tasks in wireless computing networks, improves the efficiency of task completion and resource utilization, and solves the impact of node mobility on task scheduling in wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120417092B_ABST
    Figure CN120417092B_ABST
Patent Text Reader

Abstract

This invention discloses a dynamic intelligent task scheduling method for wireless computing networks, applied to the field of computing resource scheduling. Addressing the issue that current research on computing networks largely focuses on large-scale computing networks in wired data centers, which struggles to meet the computing demands of today's widespread wireless environments, this invention constructs an average latency optimization model based on the uplink and downlink transmission times, computation waiting times, and computation times of tasks in the wireless computing network. This optimization problem is then modeled as a Markov decision process with a large state space, and solved using a deep reinforcement learning algorithm. The scheduling scheme obtained by this invention effectively ensures the high efficiency of computing task completion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computing power resource scheduling, and specifically relates to a computing power scheduling technology. Background Technology

[0002] With the rapid development of artificial intelligence technology, emerging services such as smart healthcare, autonomous driving, and image recognition have generated higher computing demands on node computing power and real-time task completion. Traditional networks are prone to shortages of node computing resources and unmet computing needs. Furthermore, these computing demands are often accompanied by sudden temporal distribution and spatial sparsity. While some nodes are experiencing task congestion, others remain idle, leading to insufficient resource utilization and prolonged task completion times. The continuous innovation of 5G advanced communication technology makes it possible to dynamically schedule computing tasks between nodes through communication. To fully utilize computing and network resources, major operators have proposed the concept of Computing Power Network (CPN), which deeply integrates computing, network, cloud, edge, and terminal computing resources. This aims to achieve on-demand scheduling of computing and network resources, comprehensively improving resource utilization efficiency and task completion performance.

[0003] At the same time, the number of global wireless terminal accesses has surged, and it is estimated that by 2025, the number of global mobile device accesses will reach 8.5 billion. However, current research on computing power networks is mostly focused on large-scale computing power networks for wired data centers, which is difficult to cope with the computing needs of the now widespread wireless environment. Therefore, it is necessary to study the task scheduling problem of computing power networks in wireless environments. Against this background, Wireless Computing Power Network (WCPN) has emerged.

[0004] In wireless computing networks, the transmission performance of node computing tasks is significantly affected by the uncertainty of the wireless communication environment and the limitations of wireless channel capacity, making it difficult to accurately predict transmission time in task scheduling. Selecting nodes with higher computing power solely based on their computing power may lead to a shortage of communication resources, further impacting task completion efficiency. Existing research on wireless computing networks is limited; a unified conceptual definition has not yet been established in the standards community, and academic research on the subject is scarce, with only a handful of publications. Current research has not considered the impact of node movement on task scheduling in the wireless environment. Furthermore, the mobility of nodes in the wireless environment means that the location and number of wireless access points may change dynamically, causing scheduling schemes generated in static environments to perform poorly in dynamic environments, failing to meet practical application requirements. Summary of the Invention

[0005] To address the aforementioned technical issues, this invention proposes a dynamic intelligent task scheduling method for wireless computing networks. This method optimizes the scheduling strategy for node computing tasks in wireless computing networks, effectively ensuring the high efficiency of task completion.

[0006] The technical solution adopted in this invention is: a dynamic task intelligent scheduling method for wireless computing networks. The application scenarios include the cloud, several base stations, edge servers configured in each base station, and several user terminals. The cloud is connected to the base stations through a wide area network or backbone network, the base stations are connected to each other through wired channels, and the edge servers are connected to their corresponding servers.

[0007] The cloud, edge servers, and user terminals all have computing capabilities. When processing computing tasks, the cloud, edge servers, and user terminals are all recorded as computing nodes.

[0008] Specific scheduling methods include;

[0009] S1. Establish an optimization model that minimizes the completion latency of computation tasks;

[0010] A continuous time period is divided into T time slots, and the length of a single time slot is denoted as Δτ. The user terminal generates a computing task at the beginning of each time slot. During the continuous generation of computing tasks, a subtask is generated in each time slot. Subtasks of the same computing task have the same requirements for computing and communication resources. Subtasks of different computing tasks have different requirements. Subtasks of the same task have the same scheduling target, which is a certain computing node.

[0011] Calculate the overall completion delay of a single subtask based on its communication delay, queuing delay, computation delay, and transmission delay.

[0012] The completion latency of each computation task is calculated based on the overall completion latency of a single subtask;

[0013] To minimize the sum of the completion times of all computational tasks, an optimization model is constructed that minimizes the completion times of computational tasks, using the task scheduling scheme as the optimization variable.

[0014] S2. The optimization model that minimizes the completion delay of the computation task in step S1 is modeled as a Markov decision process and solved based on a deep reinforcement learning algorithm.

[0015] The beneficial effects of this invention are as follows: This invention addresses the task scheduling problem in wireless computing networks. It considers the limited capacity of wireless channels and the mobility of nodes, incorporating these factors into the communication model and establishing an optimization model accordingly. This leads to more rational task scheduling and more efficient task completion in wireless computing networks.

[0016] This invention addresses the established optimization model for minimizing task completion latency in wireless computing networks and proposes a Soft Actor-Critic (SAC) algorithm based on minimizing task completion latency to solve the model. This algorithm solves the problems of wireless computing networks struggling to obtain effective immediate rewards and the curse of dimensionality in the action space as the number of nodes increases, and effectively reduces the average completion latency of computational tasks in wireless computing networks. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating a task scheduling scenario in a wireless computing network.

[0018] Figure 2 It is a SAC algorithm framework based on minimizing task completion latency;

[0019] Figure 3 This is a diagram illustrating the generation of user node tasks;

[0020] Figure 4 This is a graph showing the reward convergence process during training;

[0021] Figure 5 It is a scheduling delay diagram of the proposed algorithm and the comparison algorithm in wireless computing networks. Detailed Implementation

[0022] To facilitate understanding of the technical content of this invention by those skilled in the art, the following description, in conjunction with the accompanying drawings, further illustrates the invention.

[0023] Considering the uncertainty of the wireless communication environment and the movement of nodes in wireless computing networks, and the uneven distribution of computing resources due to the burstiness of computing tasks in time and the sparsity of their spatial distribution, leading to wasted computing resources and inefficient task completion, this invention designs a task scheduling method based on the SAC algorithm for minimizing task completion latency, thereby optimizing the average latency of computing task completion. The technical solution is described in detail below:

[0024] 1. Scene Description

[0025] In the current and foreseeable 5G era, mobile devices primarily access cellular base stations wirelessly. Base stations, in turn, rely on high-bandwidth, low-latency wired channels such as fiber optics for interconnection, enabling high-speed, stable data transmission and supporting efficient network resource management and scheduling. Under this architecture, wireless computing networks not only endow base stations with traditional access and management functions but also with capabilities for computing power awareness, resource scheduling, and intelligent collaboration. Base stations can efficiently manage information from accessing devices and engage in real-time data exchange through low-latency wired communication, achieving intelligent allocation and scheduling optimization of computing tasks. Therefore, the constructed wireless computing network scenario and its key components are as follows: Figure 1 As shown.

[0026] like Figure 1 As shown, the wireless computing network consists of key components such as the cloud, edge servers, and user nodes. User nodes include terminal devices such as smartphones, autonomous vehicles, and drones. These nodes not only initiate computing tasks but also possess certain computing, storage, and communication capabilities. Their spatial locations are typically mobile, and they are connected to communication base stations in their respective areas via wireless communication. Base stations are deployed in fixed locations and equipped with edge servers. Compared to user nodes, edge servers have stronger computing, storage, and communication capabilities, and can serve as offloading targets for computing tasks, alleviating the computing pressure on terminal devices and improving task processing efficiency. Edge servers are connected to base stations, which are typically connected via wired channels to achieve coordinated scheduling and optimized utilization of computing resources within the region. Furthermore, the system also includes a cloud computing center (referred to as the cloud in this embodiment). The cloud possesses massive computing and storage resources, providing high-performance computing capabilities. However, since the cloud is usually deployed far from users and edge servers, data transmission needs to traverse a wide area network (WAN) or backbone network, resulting in high latency and bandwidth consumption for communication with the cloud. In the context of wireless computing network architecture, how to efficiently coordinate computing resources among cloud, edge servers, and user nodes, and optimize task scheduling strategies to improve computing performance and reduce latency, is one of the key research issues.

[0027] In this scenario, current wireless computing networks typically assume that tasks are generated completely at a specific point in time during task scheduling. However, in practice, another common scenario exists where only a small portion of a task is generated initially, with subsequent content appearing gradually over time. For example, in intelligent monitoring and security systems, cameras periodically collect environmental images, which require further analysis to meet monitoring needs. Due to the correlation between images, each image cannot be analyzed in isolation. Therefore, the analysis of each image can be considered a sub-task, while the sequential analysis of a series of related images can be viewed as a complete task composed of several sub-tasks.

[0028] In such scenarios, subtasks are not generated simultaneously at the initial moment, but gradually over time. If the camera's computing resources are insufficient, these subtasks need to be scheduled to other nodes for computation. This poses three new requirements for task scheduling in wireless computing networks: (1) The completion of subtasks must be a sequential, serial sequence, and the order cannot be disordered; (2) Subtasks under the same task should be scheduled to the same target device to ensure the correlation between subtasks and avoid frequent communication between subtasks. In a mobile network environment, maintaining the efficiency of scheduling is a great challenge; (3) The time interval between the generation of subtasks and the length of time it takes for the subtasks to be completed need to be well matched.

[0029] Based on the above scenario, consider the existence of N user nodes in the wireless computing network, denoted as U = {u1, u2, ..., u...} N Any user node is represented as} in This indicates the CPU computing power of the user node. This indicates the current location of the user node. The location of a user node typically changes. This indicates the rated transmission power of the user node.

[0030] Starting at time 0, select a time interval from 0 to t. end A continuous time period, where t end It is large enough to cover various typical network operating states, such as dynamic changes in network topology, sudden surges in traffic, and load imbalance. This time period is divided into T time slots, each with a length of t. end / Τ, denoted as Δτ, and the time slot set can be represented as {1,2,…,Τ}. User nodes generate computational tasks at the beginning of each time slot, and the generation of these tasks continues. There are 1 time slot, meaning that a subtask is generated every Δτ, for a total of 100 time slots. One. User node u iThe task generated at the beginning of time slot τ can be represented as: The k-th subtask can be represented as Each task's subtasks have the same computational and communication resource requirements, but the requirements differ between subtasks of different tasks. In other words, the computational and communication requirements of subtasks are independent of k, therefore, k is available. This represents the number of floating-point calculations required to complete each subtask; using This indicates the size of the communication transmission volume of each subtask, measured in bytes, which affects its transmission latency and bandwidth requirements in the wireless channel; computational tasks can be scheduled to any node, and subtasks of the same task have the same scheduling objective. Represents computational task The scheduling object.

[0031] The computational tasks generated by user nodes can be completed not only by the node itself, but also by edge servers and the cloud. Consider a scenario with M edge servers and 1 cloud server. The edge servers are represented as E = {e1, e2, ..., e...}. M}, where any one of the edge servers can be represented as e j ,j∈{1,2,…,M}, This indicates the CPU computing power of the edge server. This indicates the location of the edge server. Typically, the edge server and the base station are close to each other and their locations do not change. Each edge server is connected to one base station, denoted as BS. j This invention assumes that any edge server corresponds to one base station, forming a one-to-one relationship, and that the edge server and its connected base station are located in the same spatial location, i.e., the base station BS. j lie in Location. And use B. j Indicates base station BS j The corresponding bandwidth.

[0032] The cloud located at the remote end is denoted as C = {c}, and c is used as the reference point. c This indicates that the computing power of cloud servers typically far exceeds that of edge servers and user nodes. The typical uplink communication rate of a cloud server is represented by... This represents the typical downlink communication rate of a cloud server. Since cloud computing centers are usually deployed in areas far from users and edge servers, tasks uploaded to the cloud for computing are subject to significant communication latency.

[0033] 2. Optimization Model for Minimizing Task Completion Latency in Wireless Computing Networks under Node Mobility Scenarios

[0034] (1) Communication model

[0035] In wireless computing networks, when a user node schedules tasks and selects a target node to complete a task, it may choose to complete the task on its own node or choose another node. The former has zero communication cost, while the latter introduces additional communication latency. During this process, the transmission of the computing task introduces additional communication latency. Assume that the computing task generated by node i is transferred to the computing node... Perform calculations, and user node u i Connected to base station BS j In this context, the communication rate between the user equipment and the base station directly affects the transmission efficiency of task offloading. Considering the Orthogonal Frequency Division Multiple Access (OFDMA) technology widely used in wireless communication systems, the communication rate between the user equipment and the base station directly impacts the transmission efficiency of task offloading. i and base station BS j The transmission rate between them can be expressed as:

[0036]

[0037] Among them B ij (t) is the base station BS in time slot t. j Assigned to user equipment u i The bandwidth, I represents interference from other base stations, σ 2 This represents noise power. Because bandwidth resources are limited in wireless communication systems, multiple user equipments share the base station's bandwidth, represented by B. j Indicates base station BS j The total bandwidth, therefore bandwidth allocation is subject to the following constraints:

[0038]

[0039] In formula (1), In the t-slot user equipment u i The transmit power used in this invention does not involve power control of the user equipment; the user equipment uses its rated power. To communicate; h ij (t) is the current user equipment u in time slot t. i and base station BS j The gain of the channel between; dis ij (t) -α Indicates path loss, depending on the location of the user equipment. and the location of the base station We obtain the path loss exponent, where α represents the path loss exponent, typically ranging from 2 to 6. The specific formula for calculating path loss is as follows:

[0040]

[0041] The formula for calculating I is as follows:

[0042]

[0043] in, Represents user node u i′ Rated transmit power; h i′j Represents node u i′ and base station BS j The channel gain between; δ(i,i′) is used to represent the user node u i and user node u i′ Whether the same subcarrier is used between them, if they share the same subcarrier, then δ(i,i′) takes the value of 1, otherwise δ(i,i′) takes the value of 0.

[0044] When computing tasks are scheduled in a wireless computing network, tasks may be scheduled within the coverage area of ​​the base station to which the user node is connected, or they may be scheduled to nodes outside the coverage area of ​​the base station. When a task is scheduled to a node outside the coverage area of ​​the base station, the transmission path of the task includes a wired link between the base stations. Let φ be the communication transmission rate factor of the wired link, then the task data is transmitted from the base station BS via the wired link. j Transmitted to base station BS j′ The transmission rate during the process can be expressed as:

[0045] r j′j (t)=f(φ,dis jj′ (t)) (5)

[0046] Among them, dis jj′ (t) represents the base station BS j and BS j′ The length of the wired link between them, and the specific form of f(·), are affected by physical characteristics such as channel bandwidth and signal quality.

[0047] Formulas (1) and (5) represent the communication rates of computational tasks transmitted over the network via wireless channels and wired links, respectively. Based on this, the transmission time of any subtask scheduled to a non-self target node can be represented in the following three ways, according to the type of target node for the computational task scheduling and the base station access relationship:

[0048] (a) User nodes or edge servers scheduled to the same base station

[0049] For any user node u i The generated computational subtasks Its scheduled target is If the scheduling target is u i The accessed base station BS jThe corresponding edge server's computing tasks can be completed solely through wireless channel transmission; if the scheduled target is another user node and is related to u i The connected base stations are the same. Considering that the rate at which the base station transmits the task downlink to the user node is much greater than the rate at which the user node transmits the task uplink to the base station, this invention only considers the time it takes for the task to be transmitted from the user node to the base station. At this time, any subtask at time t... Transmission time required to schedule to the target node It can be represented as:

[0050]

[0051] in, Indicates task The size of the communication data in each subtask, r ij (t) represents user node u i and base station BS j The wireless communication rate between them.

[0052] (b) Dispatch to user nodes or edge servers under different base stations

[0053] If subtask The target of the scheduling is the edge server, but not the u. i The accessed base station BS j The corresponding edge server, or the scheduling target of the subtask is the user node but not the user node u i If the accessed base stations are not the same, then the subtask needs to communicate via a wired link between the base stations. In this case, let BSj′ uniformly represent the base station corresponding to the target edge server or the base station accessed by the target user node. Then, at time t, any subtask... Transmission time required to schedule to the target node It can be represented as:

[0054]

[0055] in, This indicates the communication transmission size of each subtask within the task; r ij (t) represents user node u i and base station BS j The wireless communication rate between them can be calculated using formula (1); r jj′ (t) represents the base station BS j and base station BS j′ The wired link communication rate between them can be calculated using formula (5).

[0056] (c) Scheduling cloud servers

[0057] If the subtask The target of the scheduling is the remote cloud. In the scenario of this invention, the communication between different nodes and the cloud is not significantly affected by spatial location, and the communication rate can be considered a fixed value. Representing the typical uplink communication rate for communicating with the cloud, any subtask... Transmission time required to schedule to the target node It can be represented as:

[0058]

[0059] The above three scenarios cover all cases where task scheduling to a non-local node results in task transmission latency, and when the task is computed locally, the transmission time is 0. Therefore, at time t, any subtask... Transmission time required to schedule to the target node The unified representation is:

[0060]

[0061] in, Indicates task The scheduling objective is the same for all subtasks. For subtasks The scheduling target; C represents the set of cloud servers. Used to indicate that the target of scheduling is the cloud, otherwise it is not; BS j Indicates the base station to which the user node is connected, BS′ j This indicates the base station corresponding to the target being an edge server, or the base station accessed when the target is a user node.

[0062] After task scheduling is completed, it queues at the destination node until computation is finished. Once computation is complete, the result is sent back to the original node. In traditional computing network task scheduling research, considering a relatively stable communication environment and low node mobility, the ratio of the computation result to the original task size is assumed to be small, and the downlink rate is much higher than the uplink rate, often neglecting the return time of the computation result. However, in wireless computing network research scenarios, the wireless communication environment is unstable, and node movement is more frequent, making the task return time non-negligible. Let β be the ratio of the computation result to the original task size. Then, for any subtask... The calculation results are transmitted back starting at time t, and the time when the calculation results are returned is:

[0063]

[0064] In the above formula, The target of the task is indicated; U, E, and C represent the set of user nodes, the set of edge servers, and the set of cloud servers in the wireless computing network, respectively; BSj Indicates user node u at time t i The base station that is accessed, BS′ j Represents the target node at time t The base station that is connected; Represents the target node and base station BS j The communication rate between them, r jj′ (t) represents the base station BS j and base station BS′ j The communication rate between them. Formula (9) is divided into five cases according to the node type of the target node and whether the task generating node and the target node are connected to the same base station. When the scheduling target node is itself or the edge server corresponding to the base station it is connected to, the calculation result is directly sent to the task generating node through the base station. Considering that the downlink communication rate is large, this part of the delay is ignored. When the target node is another user node connected to the same base station, the return delay considers the time for the target node to upload the calculation result to the base station. When the target node is not its own user node but the base station it is connected to is not the base station of the task generating node, the return delay considers the upload delay of the target node and the delay of the wired link transmission between the base stations. When the target node is an edge server but the corresponding base station is not the base station connected to the task generating node, the return delay considers the delay of the wired link transmission between the base stations. When the target node is a cloud server, the return delay is affected by the typical downlink communication rate of the cloud server.

[0065] (2) Computational Model

[0066] When a computation task is scheduled to the target computing node, that node may still be processing other computation tasks. Therefore, newly arrived tasks are added to the task queue and executed sequentially according to a certain scheduling strategy. Since computing nodes typically use First Come First Served (FCFS), this invention also uses FCFS to complete tasks. Therefore, a queuing waiting time may be required before the actual computation, which means additional queuing latency needs to be introduced.

[0067] For any computational task subtasks At time t, it needs to be scheduled to the target node. The calculation is complete. Let the subtask being processed by the target node at time t be... remember This is the moment when the subtask begins execution. This represents the time when the subtask arrives at the computation node (or, if the computation queue is idle, the arrival time of the next task). Let this be the time. Subtasks Reaching the computing node via communication At that moment, combined with the analysis of task communication delay using formula (8), The following relationship exists between the time t when the subtask is scheduled and the time t when it begins to be scheduled:

[0068]

[0069] For the target node use This represents the set of all subtasks waiting to be computed at this node. For any given node... Subtasks d awaiting completion wait This invention assumes that computing nodes can obtain and store the arrival times of each subtask, and denoted as subtask d. wait The time when the computation node is reached is t arrive (d wait The number of floating-point calculations required to complete the subtask is ζ(d). wait Then for any node reaching the computation node... The moment at and Subtasks between, subtasks We need to wait for these tasks to be processed and completed, therefore subtasks The computational workload that needs to wait for the task to complete is:

[0070]

[0071] The queuing delay for this task can then be expressed as:

[0072]

[0073] in, Represents a computing node The computational ability. In formula (12), before the minus sign... Represents a node Total computational load of tasks to be completed Required execution time, minus sign This represents the time difference between the start of the current task and the arrival of the new task; that is, the portion of the tasks currently being executed in the queue that have been completed, and therefore should be deducted. (In the definition...) When the queue is idle (i.e., ...), set the time when the queue is empty (i.e., ...) When ), it indicates the arrival time of the next task at that node. Substituting this into formula (12) yields the waiting time. Since the value is 0, formula (12) satisfies the condition when the queue is idle.

[0074] The computation task completes its computation after arriving at the computation node, and the required computation time is as follows:

[0075]

[0076] (3) Average delay optimization model

[0077] Based on the above analysis of the communication model and task calculation model in the task completion process, formulas (8), (12), (13) and (8) respectively give the individual subtasks. The communication delay, queuing delay, computation delay, and return delay are included, thus determining the overall completion delay of a single subtask scheduled starting at time t. for:

[0078]

[0079] Where t′ represents the time when the calculation results begin to be transmitted back, and the following relationship applies:

[0080]

[0081] t represents the moment when the subtask begins scheduling. The scheduling time for each subtask of a task is different. Specifically, the task... The start scheduling time t of the kth subtask k It can be represented as:

[0082] t k =τΔτ+(k-1)Δτ (16)

[0083] In formula (16), t k The time at which the k-th subtask begins scheduling is consistent with the meaning of t in formula (14), both representing the time at which the subtask begins scheduling. Formula (14) reflects the relationship between the start time of any subtask scheduling and the task completion delay, while formula (16) further specifies the start time of any subtask scheduling, that is, the specific representation of t for a specific subtask in formula (14). Therefore, it can be substituted into formula (14) to obtain the overall completion delay of each subtask, and the completion delay of each subtask can be summed to obtain the task completion delay. Task completion delay for:

[0084]

[0085] Based on the above analysis of task completion time, within the time slot range of [1,T], the total time for N user devices to complete their computational tasks can be expressed as:

[0086]

[0087] To minimize the latency of task completion in wireless computing networks, and using the task scheduling scheme as the optimization variable, the proposed optimization model for minimizing latency is as follows:

[0088]

[0089] In the above optimization problem, constraint (19a) is a value constraint on the task scheduling expression. The k-th subtask of the computation task generated by user node i at the beginning of time slot τ is scheduled to be computed by node j, and together they form the overall scheduling scheme S; constraint (19b) indicates that all subtasks must be scheduled to be completed without omission; constraint (19c) further indicates that all subtasks must be scheduled to be completed by the same computing node; constraint (19d) indicates a bandwidth resource constraint, indicating that the bandwidth resources allocated by the base station to each node cannot exceed its total bandwidth resources; constraint (19e) is a subtask completion time constraint, indicating that the completion time of a subtask cannot exceed the interval between the subtasks.

[0090] 3. SAC scheduling algorithm based on minimizing task completion delay

[0091] This invention designs a SAC algorithm based on minimizing task completion delay to solve the above problems.

[0092] (1) Basic Ideas and Framework

[0093] To address the aforementioned challenges with lower complexity, this invention models the problem as a Markov decision process with a large state space and solves the problem based on a deep reinforcement learning algorithm. However, traditional reinforcement learning algorithms still face some key difficulties when applied to such problems.

[0094] Firstly, traditional SAC algorithms provide rewards immediately after the agent makes a decision to optimize the policy during the learning process. However, in wireless computing networks, latency is used to measure the quality of scheduling actions. After the central scheduler issues an action, the environment cannot immediately provide the latency value; the task must be completed before a valid reward is given. However, in the early stages of training, if the agent's scheduling strategy is not effectively optimized, the task completion time may be very long, resulting in a significant delay in reward feedback. This could affect the initial performance of reinforcement learning training, as well as its overall stability and training efficiency. Therefore, this invention designs a reward truncation mechanism: when the processing latency of a task exceeds the latency of local direct computation or the interval between subtasks, reward feedback is given immediately to avoid the negative impact of reward delay on training.

[0095] Secondly, the traditional SAC algorithm's agent policy network outputs a probability distribution and obtains specific actions through sampling. In wireless computing network scenarios, when making task decisions, the decision-maker assigns scheduling objects to any user node. Assuming there are N user nodes, M edge servers, and the cloud in the scenario, there are N+M+1 computing nodes. Therefore, the scheduling space for any node is N+M+1, and the action space is N. N+M+1 As the number of nodes increases, the action space grows exponentially. This exponential growth leads to a severe curse of dimensionality, which not only increases the computational complexity of reinforcement learning models but also affects decision-making performance, making it difficult for agents to learn effectively in high-dimensional action spaces. To address this, this invention proposes a multi-dimensional action decomposition strategy based on the task generator. This strategy divides the original large-scale actions into multiple smaller actions, reducing the size of the action space to N(N+M+1). It transforms the exponential increase into a multiplicative growth, reducing training difficulty and improving decision-making efficiency.

[0096] To address the aforementioned problems, this invention proposes a reinforcement learning algorithm based on minimizing task completion latency (SAC). The framework of this algorithm is as follows: Figure 2 As shown, the nodes in the wireless computing network constitute the reinforcement learning environment, and the scheduler, acting as the intelligent agent, provides a scheduling scheme for the tasks of each node in each time slot. The state space is defined as S, the action space as A, and the reward function as R. In each training round, the scheduler observes the current environment to obtain state s, selects action a, gives reward r, and enters the next state s′. The design of the specific state space, action space, and reward function will be described in detail below.

[0097] (a) State space

[0098] The system's state space consists of the states of each node in the wireless computing network. Nodes can be categorized into three types: user nodes, edge servers, and cloud nodes. User nodes are initially randomly distributed in space and move randomly according to a Markov model, while generating computational tasks according to an ON-OFF model. Edge servers are distributed close to base stations and are relatively evenly distributed in space. Any user node can communicate with an edge server through a base station, and edge servers are connected to each other via wired connections. The cloud is far from user nodes and edge servers, resulting in significant communication latency. Therefore, the system's state space is jointly composed of the states of these three types of nodes, as detailed below:

[0099] i. User Node

[0100] User node status information includes the mobile terminal device's workload, resource capabilities, and spatial motion characteristics, including:

[0101] • Task characteristics: Triple {ki ,ζ i D i} represent the number of subtasks, the CPU cycle requirement per subtask, and the amount of communication data, respectively;

[0102] • Resource status: Covering transmission power Local computing power and remaining time for tasks to be performed

[0103] • Motion state: {x i ,y i ,v i ,θ i} describes the node's planar coordinates (x, y), movement speed v, and direction angle θ.

[0104] ii. Edge server

[0105] Reflecting the resource availability and spatial distribution of edge computing nodes, including:

[0106] • Resource status: Including communication bandwidth computing power and remaining time for tasks to be performed

[0107] • Position state: {x i ,y i} describes the planar coordinates (x, y) of the node, and the edge server location is fixed.

[0108] iii. Remote Cloud

[0109] Simplified to typical communication uplink rate downlink speed And computing power size c c The triplet Due to the pooling nature of cloud resources, we assume that there are sufficient computing resources and ignore task queuing latency.

[0110] (b) Action space

[0111] In this invention, when any user node generates a computing task, the central scheduler will output a corresponding scheduling scheme based on the current system state, that is, the action decided by the agent according to its strategy. Considering that there are N user nodes, M edge servers, and 1 cloud server in the scenario, totaling N+M+1 computing nodes, all computing nodes are numbered sequentially from 1 to N+M+1.

[0112] However, directly encoding the scheduling actions of all computing nodes leads to an exponential increase in the dimension of the action space, greatly increasing the solution complexity. Therefore, this invention employs an action decomposition strategy: the agent outputs a scheduling sub-action for each user node, referred to as a node sub-action. For each user node, the action space of its sub-actions is denoted as a. i The size of each is N+M+1, and the value is an integer from 1 to N+M+1, indicating that the computing tasks generated by the user node are scheduled to the computing node with the corresponding number.

[0113] Therefore, at time t, the overall action given by the agent is the set of sub-actions of each user node, which is:

[0114] a t ={a1,a2,…,a i ,…,a N};a i ∈{1,2,…,N+M+1} (20)

[0115] (c) Reward Function

[0116] In reinforcement learning, the agent's behavior is driven by reward signals, and the reward function, as a core component of the reinforcement learning algorithm, plays a decisive role in the updating and convergence of the agent's policy. Considering the characteristics of wireless computing networks, system performance is closely related to the latency of computation tasks, and latency in real-world environments often presents a delay issue. Therefore, this invention designs a reward truncation mechanism: when the processing latency of a task exceeds the latency of local direct computation or the time interval between subtasks, reward feedback is immediately provided to avoid the negative impact of reward delay on training.

[0117] Therefore, this invention designs a comprehensive reward function based on a reward truncation mechanism. This function not only reflects the impact of task completion latency on system performance but also uses the local completion latency of the task as a benchmark, thereby accurately evaluating the merits of the scheduling strategy. Specifically, the reward function fully considers the entire process latency of the computation task from scheduling, transmission, computation to feedback. By constructing delayed rewards, it encourages the agent to adopt a reasonable truncation strategy while optimizing the overall task completion latency, so as to avoid excessive negative impact of long latency on the reward signal. The reward for an action performed in time slot τ is determined by the latency of all tasks scheduled at that time, i.e.:

[0118]

[0119] in The reward value obtained from the k-th subtask of the task generated at node i in time slot τ is determined by the completion time of each subtask. and local computing time And the generation of the interval Δτ is determined as follows:

[0120]

[0121] (2) Algorithm Flow

[0122] Based on the above analysis of the algorithm framework and the state space, action space, and reward function in reinforcement learning, this invention optimizes the training process based on the traditional SAC algorithm, improves the problems of excessive action space and reward delay, and proposes an optimized SAC algorithm. This algorithm consists of five core networks: the policy network (Actor Network) is responsible for generating the agent's actions, enabling it to select the optimal policy while exploring the environment; the evaluation network (Critic Network) calculates the Q-value of the state-action pair, using two independent Q-networks to alleviate the problem of Q-value overestimation, thereby improving the stability and convergence of learning; the target network (Target Network) has the same structure as the evaluation network, mainly used to calculate stable target Q-values ​​to avoid oscillations during training, and uses a soft update mechanism to smooth parameter changes, further enhancing training stability.

[0123] Furthermore, SAC has a dual objective of maximizing reward and maximizing actual entropy, and introduces a temperature coefficient α. t The objective function, used to control the degree of exploration in the policy, determines the randomness of the policy and is dynamically adjusted during training by optimizing the objective function to maintain a balance between exploration and exploitation. Specifically, in the SAC algorithm, the policy objective is defined as follows:

[0124]

[0125] in, Indicates the experience replay pool The state s is obtained by sampling. t Experience replay pool It stores historical state transition records; a t ~π θ This indicates that the current action distribution π is... θ (·|s t Action a is obtained by random sampling t , π θ (·|s t The distribution is determined by the current policy network based on the current state s. t Given; the first term α in the expectation is calculated. t (-logπ θ (a t |s tEncouraging high entropy values, that is, encouraging more exploration, while the second item... To encourage higher returns, i.e., greater utilization, the SAC algorithm can adjust the temperature coefficient α. t The algorithm influences the weight ratio of the two items in the policy objective, thus balancing exploration and exploitation. First, it requires inputting the maximum number of training epochs and the maximum number of time steps per epoch. Then, it performs necessary parameter initialization, including parameters for five neural networks (policy network, two Q-evaluation networks, and two objective networks), temperature coefficients, etc., and clears the experience replay pool. Furthermore, it needs to set the environmental parameters of the nodes, including node location distribution, motion state parameters, resource status, and task model.

[0126] After initialization, the algorithm enters the training phase. In each training round, the agent samples actions from the Actor Network at the beginning of each time slot based on environmental state parameters and makes scheduling decisions accordingly. Before executing scheduling actions, the agent calculates and records the local computation time of the task for subsequent decision-making and reward calculation. Then, the agent executes task allocation and computation operations according to the scheduling actions output by the policy network. After task execution, the agent checks whether a previously computed task has been completed or whether the processing time of the current task exceeds the local processing latency or the interval between subtasks. If the above conditions are met, the agent receives the corresponding reward and records the state transition (s). t ,a t ,r t ,s t+1 ,γ) are stored in the experience replay pool D for subsequent training. If the computation task does not meet the above conditions, the computation task continues to be processed until the above conditions are met and then placed into the experience replay pool, where s t a represents the current state of the environment. t Indicates the currently selected action, r t Indicates the reward received, s t+1 The state transitioned to after an action is performed is represented by γ, which represents the discount factor reflecting the importance of future rewards. To avoid training instability due to starting with a small number of samples, neural network parameters are updated only after the number of records in the experience replay pool reaches the minimum required for training. Generally, the initial training sample size is 10% of the experience replay pool capacity. After training begins, the Critic network, Actor network, and Target network are updated sequentially, and the temperature coefficient α is applied. t Adjustments are made to complete a training session.

[0127] In this invention, to reduce the number of hyperparameters, an adaptive adjustment strategy is adopted for the temperature coefficient, so that the strategy entropy (i.e., the actual entropy) adaptively approaches the target entropy. The target entropy is usually set to -dim(A), which is the dimension of the action space. To make the actual entropy Approaching target entropy Define temperature loss function for:

[0128]

[0129] To minimize the temperature loss function, for α t Perform gradient descent, i.e., α t The following updates will be made:

[0130]

[0131] Where λ is the learning rate. Let the temperature loss function be... This indicates that the loss function at α t Gradient estimation at each time slot. After all time slots have been executed, the algorithm enters the next round of training, continuously optimizing the scheduling strategy until the agent converges to the optimal decision scheme.

[0132] The specific algorithm flow is as follows:

[0133]

[0134] 4. Implementation Cases

[0135] To verify the performance of the SAC algorithm proposed in this invention, the present invention uses the Python language, writes node events in the Simpy framework, and uses the PyTorch framework to build and train the neural network. This embodiment will describe the simulation scenario and parameters in detail, and then present and analyze the simulation results.

[0136] (1) Implementation scenarios and parameter settings

[0137] To verify the application effect of the SAC algorithm in the task scheduling scenario of wireless computing power network, a simulation model of the wireless computing power network scenario was performed in this invention. The simulation environment was set in a square area with a side length of L, and the size of the simulation area was set to 1600m×1600m. There were 20 user devices, 5 edge servers and one cloud in the area, and the user devices and edge servers were evenly distributed in the simulation area.

[0138] Considering that in real-world applications, computational tasks are often generated in bursts, and these bursts of traffic have a certain duration, while computational demands are often spatially unevenly distributed, the spatiotemporal distribution of tasks in the entire wireless computing network exhibits dispersion and sparsity. This invention uses an ON / OFF model to describe the task status of any computing node, such as... Figure 3 As shown. This model can effectively characterize the dynamic characteristics of user nodes generating computing tasks. That is, in the ON state, the node is in the active period of computing tasks, and at the beginning of each time slot, it will generate tasks according to a certain probability pr. i In the OFF state, the node generates computational tasks and remains silent with no tasks generated.

[0139] The simulation scenario is designed with each time slot lasting 1 second. At the beginning of each time slot, when the user equipment is in the ON state, a computation task is randomly generated. The computation task can be scheduled to any node for completion. After the task is completed, the completion time is calculated. The parameter settings for the simulation environment are shown in Table 1 below:

[0140] Table 1. Wireless Computing Network Simulation Environment Parameter Settings

[0141]

[0142] During the simulation, the actions of the nodes are given by the SAC algorithm. The training method of the neural network is shown in Algorithm 1. Both the policy network and the evaluation network are set to two fully connected hidden layers. The number of neurons in the hidden layer near the input layer is set to 256, and the number of neurons near the output layer is set to 128. The training rounds are set to 800 episodes, and each episode contains 1000 time slots. Other specific training parameters are shown in Table 2.

[0143] Table 2 SAC Algorithm Parameter Table

[0144]

[0145] (2) Implementation effectiveness analysis

[0146] Experimental simulations were conducted using the aforementioned simulation environment and network parameters. The reward value was output during training, and its trend with training rounds is shown below. Figure 4 As shown. In the initial stage, due to the randomization of network parameters, the agent mainly relies on random exploration to select scheduling objects. During this process, affected by the randomness of the scheduling strategy, computing tasks may be assigned to user devices with limited computing resources, or multiple computing tasks may be concentrated on the same computing node, causing computing tasks to be blocked, thus resulting in longer task completion times. Therefore, in the first 100 training iterations, the reward value of environmental feedback is low, and the agent's scheduling strategy has not yet been effectively optimized.

[0147] As training progresses, the reinforcement learning experience replay pool gradually accumulates sufficient samples, and neural network training officially begins. Guided by environmental reward feedback, the agent gradually adjusts its decision-making strategy, tending to choose scheduling schemes that yield higher rewards. After 100 training iterations, the reward value begins to rise significantly, indicating that the agent's scheduling strategy is gradually optimizing. After further training, when the training iterations reach approximately 200 rounds, the reward value tends to converge, and the agent has essentially converged to a stable scheduling strategy. At this point, the decision scheme provided by the neural network can better match computational and communication resources with task requirements, ensuring task execution. After the training process converges and reaches a stable state, the policy network parameters are fixed, and testing is conducted under the same environmental settings. The test duration is set to 1000 seconds, and the latency values ​​of each computational task are recorded during the test. The average completion latency of a single task is calculated, and common task scheduling schemes are compared under the same environmental parameters. The compared scheduling strategies include the SAC algorithm based on minimizing task completion latency proposed in this invention, the greedy scheduling algorithm based on node load status, the fixed edge server offloading scheme, and the local computing scheme. The average task completion latency of each scheduling strategy is obtained as follows: Figure 5 As shown.

[0148] from Figure 5 As can be seen, among the four scheduling schemes, the algorithm proposed in this invention performs best, controlling the average completion latency of a single task to within 1 second, with an average task completion latency of 0.73 seconds in the test. The average completion latency of other scheduling methods is higher than that of the algorithm proposed in this invention. Specifically, the average completion latency of locally computed tasks is 2.98 seconds, while the greedy algorithm and fixed offloading methods have completion latencies of 1.25 seconds and 1.57 seconds, respectively. Furthermore, the trend of the curves shows that the other three scheduling methods all exhibit significant latency during task completion. This is because the task generation of nodes in this invention uses an ON-OFF model during simulation testing. These scheduling schemes fail to efficiently adapt to the dynamic distribution of computational tasks in time and space, resulting in increased latency over a period of time. In contrast, the algorithm proposed in this invention effectively utilizes task dynamics, fully leverages the resources in the wireless computing network, and efficiently completes computational tasks.

[0149] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of the claims of the invention.

Claims

1. A method for dynamic task intelligent scheduling in wireless computing networks, characterized in that, Application scenarios include the cloud, several base stations, edge servers configured in each base station, and several user terminals; the cloud is connected to the base stations through a wide area network or backbone network, the base stations are connected to each other through wired channels, and the edge servers are connected to their corresponding servers; The cloud, edge servers, and user terminals all have computing capabilities. When processing computing tasks, the cloud, edge servers, and user terminals are all recorded as computing nodes. Specific scheduling methods include; S1. Establish an optimization model that minimizes the completion latency of computation tasks; Divide a continuous time period into There are 1 time slot, and the length of a single time slot is denoted as . At the beginning of each time slot, the user terminal generates a computing task. During the continuous generation of computing tasks, a subtask is generated in each time slot. Subtasks of the same computing task have the same requirements for computing and communication resources. Subtasks of different computing tasks have different requirements. Subtasks of the same task are scheduled to the same target node, which is a certain computing node. Calculate the overall completion delay of a single subtask based on its transmission time, queuing delay, computation time, and return delay. The completion latency of each computation task is calculated based on the overall completion latency of a single subtask; To minimize the sum of the completion times of all computational tasks, an optimization model is constructed that minimizes the completion times of computational tasks, using the task scheduling scheme as the optimization variable. S2. The optimization model that minimizes the completion delay of the computation task in step S1 is modeled as a Markov decision process and solved based on a deep reinforcement learning algorithm. In step S2 of the deep reinforcement learning algorithm, the agent outputs scheduling sub-actions for each user terminal, and then in the time slot... The overall action given by the intelligent agent is a set of user-scheduled sub-actions, namely: ; in, Indicates the first The action space of scheduling sub-actions for each user client; In step S2, the optimization objectives of the deep reinforcement learning algorithm include maximizing the reward and maximizing the actual entropy; In step S2, the deep reinforcement learning algorithm, in The reward for performing an action in a time slot is determined by the latency of all tasks scheduled at that moment, i.e.: ; in, Indicates in time slot No. The first computing task generated by the user terminal The reward value obtained from each sub-task The formula for calculation is: ; in, Subtasks Completion time, Indicates the local calculation time.

2. The method for dynamic task intelligent scheduling in a wireless computing network according to claim 1, characterized in that, The optimization model for minimizing the completion latency of the computation task constructed in step S1 is expressed as follows: ; ; ; ; ; ; Where S represents the overall scheduling scheme, express The total time for each user client to complete its computing tasks. , Represents computational task The completion delay , This represents the overall completion latency of a single subtask. This represents the number of time slots that constitute the process of generating a computational task. Indicates the first individual user terminals The computational task generated at the start of the time slot The subtask is scheduled to the first The scheduling results of each computing node Indicates base station Total bandwidth, Is Time-slot base station Assigned to the The bandwidth of each user terminal, where M represents the number of base stations.

3. The method for dynamic task intelligent scheduling in a wireless computing network according to claim 2, characterized in that, The formula for calculation is: ; in, Subtasks The transmission time required to schedule to the target node. Subtasks Queuing delay; Subtasks The computation time required to complete the calculation after reaching the target node; To indicate the return latency, subtasks are represented. The transmission time required to send the calculation results back from the target node.

4. The method for dynamic task intelligent scheduling in a wireless computing network according to claim 3, characterized in that, The state space of the deep reinforcement learning algorithm in step S2 includes: user-side state, edge server state, and cloud state; The user terminal status includes task characteristics, user terminal resource status, and motion status. Task characteristics include the number of subtasks, CPU cycle requirements of a single subtask, and communication data volume. User terminal resource status includes transmission power, user terminal local computing power, and remaining time for tasks to be executed on the user terminal. Motion status includes user terminal planar coordinates, movement speed, and direction angle. The status of an edge server includes the resource status and location status of the edge server. The resource status of the edge server includes communication bandwidth, computing power, and remaining time for tasks to be executed on the edge server. Cloud status includes communication uplink speed, downlink speed, and cloud computing power.

5. The method for dynamic task intelligent scheduling in a wireless computing network according to claim 4, characterized in that, The actual entropy is expressed as: ; in, Indicates the expectation. For the current policy network, based on the current state Given the current action distribution, This indicates that the current policy network is based on the current state. The given action The corresponding probability value, Indicates the distribution of current actions Actions obtained through random sampling .

6. The method for dynamic task intelligent scheduling in a wireless computing network according to claim 5, characterized in that, It also includes the temperature coefficient, which uses an adaptive adjustment strategy to make the actual entropy adaptively approach the target entropy. Define the temperature loss function. Represented as: ; in, Indicates the temperature coefficient. Representing the action space Dimensions.