Task-dependent long-slot general computing resource allocation method
By employing Lyapunov optimization and the Actor-Critic reinforcement learning framework, the complexity of resource scheduling and allocation in cellless networks is addressed, achieving efficient resource utilization and optimization of task dependencies over long timescales, thereby improving system stability and computational offloading efficiency.
Patent Information
- Application Number
- CN202511355240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-12
AI Technical Summary
In user-centric cellless networks, existing methods struggle to coordinate communication and computing resources effectively over long timescales, resulting in low resource utilization, increased task latency, and failure to effectively handle task dependencies, thus impacting system stability and user experience.
The Lyapunov optimization method is used to transform the long time-slot task scheduling optimization problem into a single time-slot optimization subproblem that can be solved sequentially. Combined with the Actor-Critic reinforcement learning framework, an adaptive scheduling strategy is designed. By using a deep neural network to perceive the channel state and the availability of computing resources, the selection of access points and resource allocation are optimized, and a parallel offloading model that supports dynamic task dependencies is constructed.
It improved resource utilization, reduced task execution latency, enhanced system stability and computational unloading efficiency, optimized the parallel unloading strategy under task dependencies, and improved overall task processing capabilities.
Smart Images

Figure CN121126541A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, specifically relating to a method for allocating long-slot computing resources under task-dependent conditions. Background Technology
[0002] In the 6G era, computationally intensive, low-latency applications such as intelligent transportation, healthcare, and the Industrial Internet place higher demands on the coordinated scheduling of communication and computing resources. In highly dynamic network environments, ensuring stable system operation and optimizing computing performance have become major concerns in recent years. Particularly in user-centric, cell-free (UC) networks, where communication and computing resources are distributed widely, efficient resource scheduling to improve overall performance has become a research hotspot.
[0003] Existing research has proposed online computation offloading algorithms to maximize network data processing capacity under long-term data queue stability and average power constraints. These methods typically employ dynamic optimization strategies to maintain high computational efficiency in changing environments. However, these methods primarily focus on resource scheduling and allocation over short timescales, often neglecting the profound impact of dynamic changes in task queues on system performance over long timescales. For example, some studies only explore resource scheduling and allocation strategies within a single time slot, or fail to adequately consider the stability of task queues over long timescales, resulting in a failure to effectively capture the impact of task queue accumulation and consumption across multiple time slots on overall system performance.
[0004] Other existing research focuses on the latency and energy consumption tradeoffs of task offloading in multi-user scenarios. In these studies, the optimization objective is usually to reduce task processing latency or power consumption to improve user experience. However, in resource-constrained, user-centric cellless networks, due to the varying computing capabilities of access points (APs), such methods struggle to achieve efficient resource scheduling and allocation in dynamic network environments.
[0005] In user-centric cellless networks, there are numerous access points with varying computing server load capacities. As tasks arrive dynamically, the load capacity of different access points constantly changes, requiring highly adaptable allocation of computing resources. Furthermore, task dependencies necessitate specific execution sequences, further increasing the complexity of resource scheduling and allocation. Traditional single-timeslot optimization methods may get stuck in local optima when facing these dynamic changes, limiting the improvement of overall network performance. Therefore, how to rationally coordinate communication and computing resources within long timeslots in access point environments with limited computing resources, and optimize resource scheduling and allocation strategies, remains a pressing problem to be solved.
[0006] In fact, long-term resource scheduling and allocation strategies can more comprehensively optimize system performance, especially in dynamic and complex network environments. These strategies can better adapt to changes in resource status and significantly improve task processing efficiency and resource utilization. However, long-term resource scheduling and allocation also face greater complexity, requiring a fine balance between task completion latency on short timescales and system stability on long timescales. Specifically, short-timescale resource scheduling and allocation strategies may lead to uneven resource allocation due to an excessive pursuit of immediate performance, thus affecting the stability of long-term data queues; while long-term resource scheduling and allocation strategies need to comprehensively consider factors such as task arrival rate, changes in resource status, and queue backlog to achieve globally optimal system performance.
[0007] While current research on computation offloading, resource scheduling, and allocation has made some progress, it still has many limitations. Existing methods mainly focus on resource scheduling and allocation within short timescales, lacking a systematic consideration of long-term resource scheduling, and are difficult to adapt to the dynamic environment of user-centric cellless networks with limited computing resources. Due to the random arrival of tasks and the heterogeneity of computing resources, the accumulation and consumption of task queues across multiple time slots have a profound impact on system performance. However, existing methods usually do not fully consider this factor, resulting in a lack of global optimization capabilities in resource scheduling and allocation. Locally optimal scheduling strategies within short timescales may actually reduce the overall utilization of computing resources.
[0008] Meanwhile, in dynamic business request environments, computing resource scheduling needs to be highly adaptive. However, existing methods mostly rely on static or semi-static optimization strategies, making it difficult to perceive changes in business needs in real time and flexibly adjust resource allocation. When business requests fluctuate drastically, this lack of adaptability can lead to wasted computing resources and even increase task execution latency due to scheduling delays, severely impacting system service quality and user experience.
[0009] Furthermore, existing methods still have significant shortcomings in the parallel offloading of complex tasks. Many existing methods assume that tasks can be offloaded and executed independently, ignoring the dependencies between subtasks. In practical applications, tasks often consist of multiple subtasks with different sequential constraints. Especially in multi-stage computational tasks, if computational resources are not scheduled properly, some subtasks may experience additional latency due to incomplete preceding tasks or resource contention, thereby reducing the overall computational offloading efficiency and affecting the overall task execution. Therefore, optimizing parallel scheduling under task dependencies while ensuring efficient utilization of computational resources remains an urgent problem to be solved.
[0010] In summary, the existing methods mainly have the following technical problems:
[0011] (1) The randomness of user tasks, the dynamic changes in network status, and the dependence of tasks make it difficult for traditional static scheduling methods to adapt to complex environments, resulting in low resource utilization, increased task latency, and even affecting network stability.
[0012] (2) Most existing methods rely on fixed resource allocation, which is difficult to adapt and adjust, resulting in low short-term scheduling efficiency and difficulty in ensuring long-term system stability.
[0013] (3) Existing methods have high computational complexity, are difficult to converge in real time, and cannot meet the requirements of high-time-efficiency computational unloading.
[0014] Therefore, there is an urgent need for a collaborative scheduling scheme for communication and computing resources that can sense network status, optimize task dependencies, and take into account both system stability and execution efficiency. Summary of the Invention
[0015] To address the challenges of limited computing resources, complex task dependencies, and fluctuating dynamic service requests in user-centric cellless networks, this invention provides a long-slot computing resource allocation method under task-dependent conditions. By coordinating the computing and communication resources of multiple access points, it improves task processing efficiency, reduces overall system latency, and thus enhances resource utilization and service quality.
[0016] This invention aims to improve system stability and computational resource utilization efficiency, and optimize parallel offloading strategies under task dependencies. By constructing a long-slot task scheduling optimization model, this invention employs the Lyapunov optimization method to transform the long-slot task scheduling optimization problem into a serially solvable single-slot optimization subproblem, thereby improving scheduling stability and resource utilization. Simultaneously, this invention designs an adaptive scheduling strategy based on the Actor-Critic reinforcement learning framework, utilizing deep neural networks to perceive channel states and computational resource availability, dynamically optimizing access point selection decisions, achieving intelligent allocation of computational resources, and reducing task execution latency. Furthermore, to address complex dependencies between subtasks, this invention proposes a parallel offloading model that supports dynamic adaptation to different task dependencies, optimizing task execution order, coordinating computational and communication resource scheduling, reducing additional latency, and improving computational offloading efficiency.
[0017] The technical solution adopted by this invention to solve the technical problem is as follows:
[0018] The present invention provides a collaborative scheduling algorithm for general computing resources, which specifically includes the following steps:
[0019] (1) The CPU calculates the data upload rate;
[0020] At the beginning of each time slot, the CPU calculates the user's uplink signal-to-noise ratio (SNR) based on the user's uplink signal, the channel state between the access point and the user, the channel estimate, and the receiver combining vector selected by the access point. Then, the CPU calculates the user's data transmission rate based on the uplink SNR.
[0021] (2) The CPU calculates the latency consumption of each user task;
[0022] Each user's task is divided into multiple subtasks with data dependencies. The CPU first calculates the transmission latency when the user transmits the task, and then assigns the subtasks to the access point for parallel offloading. The CPU calculates the task processing latency based on the processing latency of the subtasks, and then calculates the user's total processing latency. At the same time, the CPU models the dynamic queue based on the total processing latency.
[0023] (3) The CPU constructs a long time slot task scheduling optimization model and uses the Lyapunov optimization method to decouple the long time slot task scheduling optimization problem into a single time slot optimization subproblem that can be solved sequentially. The optimal long time slot task scheduling scheme is obtained by using reinforcement learning methods.
[0024] (4) The CPU sends the optimal long time slot task scheduling scheme to the requesting user and each access point that provides the service.
[0025] Furthermore, the goal of the Lyapunov optimization method is to minimize the average total processing latency of long-term user tasks by jointly optimizing user transmission power, access point selection, task allocation, and computing resource allocation.
[0026] Furthermore, the mathematical expression of the long time-slot task scheduling optimization model is:
[0027]
[0028] Where G represents the total time, K represents the number of users, and M represents the number of access points; D, P, Q, and F all represent decision matrices, D represents the AP connection matrix, P represents the user upload power matrix, Q represents the subtask allocation matrix, and F represents the computing resource allocation matrix. This represents the total processing latency of the k-th user task; Indicates user data transmission power; Q m (t) represents the time required for the queue of tasks to be processed at time t on the m-th access point; This indicates that the m-th access point is assigned to the subtask. The computing resources; ζ represents the maximum number of access points that each user can connect to; constraint C1 indicates that the elements of the AP connection matrix D are binary variables. This indicates that the m-th access point is connected to the k-th user at time t. This indicates no connection; constraint C2 limits the maximum number of access points each user can connect to; constraint C3 indicates the maximum user data transmission power is p. max Constraint C4 indicates that the elements of the subtask assignment matrix Q are binary variables. Subtasks It is assigned to be processed by the m-th access point. Subtasks The processing does not occur at the m-th access point; constraint C5 indicates that tasks can only be assigned to connected access points; constraint C6 indicates that subtasks can only be processed on one access point; constraint C7 indicates the stability of the queue at each access point; constraint C8 indicates that the resources allocated to the m-th access point do not exceed its upper limit f. m .
[0029] Furthermore, in the Lyapunov optimization method, a drift-penalty minimization method is used to stabilize the queue Z(t) while minimizing the average total processing latency of long-term user tasks; this is achieved by introducing a Lyapunov function. Drifting with Lyapunov Minimize the upper bound of the drift plus penalty expression in each time frame.
[0030] Furthermore, the drift penalty expression is as follows:
[0031]
[0032] Where V is the weighting variable, used to adjust the proportion of the penalty term; Q m (t) represents the time required for the queue of tasks to be processed at time t on the m-th access point; L(Z(t+1) represents the total processing delay of the k-th user task; L(Z(t+1) represents the Lyapunov function expression at time t+1.
[0033] Furthermore, the upper bound of the drift plus penalty expression is:
[0034]
[0035] Where θ is a constant term; t g Indicates the interval between each time slot; Subtasks It is assigned to be processed by the m-th access point. Subtasks The processing does not occur at the m-th access point; Indicates computation delay; This represents the total processing latency of the k-th user task;
[0036] Starting from the t-th time frame, remove the constant term θ and determine the action by minimizing the following:
[0037]
[0038] Considering the constraints of each frame, solve the following deterministic per-frame subproblem in the t-th time frame:
[0039]
[0040] Furthermore, in the reinforcement learning method, the main problem of the long time-slot task scheduling optimization problem is represented as an access point selection problem, and the intelligent access point selection strategy of Actor-Critic reinforcement learning is used to solve the main problem; the subproblems of the long time-slot task scheduling optimization problem are decoupled into a power allocation subproblem and a computing resource allocation subproblem, which are solved by fractional programming and linear programming methods, respectively.
[0041] Furthermore, the specific implementation process of the intelligent access point selection strategy for Actor-Critic reinforcement learning is as follows:
[0042] The Actor module is responsible for policy learning, selecting the optimal action based on the current state; the policy is to dynamically adjust the number of candidate decisions; the candidate decisions are generated using a noise scrambling method.
[0043] The Critic module is responsible for value function estimation, evaluating the value of the current state or state-action pair, and feeding it back to the Actor module to guide policy improvement.
[0044] Furthermore, the Actor module consists of a deep neural network and an action quantizer. The input to the deep neural network is the channel estimate and the queue state, and the output of the deep neural network is the offloading decision. The parameter α of the deep neural network... t By minimizing the cross-entropy loss function PL(α) t The Actor module is updated using the updated parameters and trained using the Adam optimization algorithm. After training, the updated parameters are used to update the Actor module in the next time frame.
[0045] Furthermore, in the aforementioned noise scrambling-based generation method, the continuous variable is first processed... Perform direct quantization to generate the first set of binary actions. By continuous variables Random Gaussian noise n ~ N(0, I) is introduced above. N ), thus obtaining the perturbed variables:
[0046]
[0047] Where Sigmoid(·) is a mapping function used to map each component to the range (0,1);
[0048] Then based on the perturbed variables Regenerate the remaining M t -1 group of candidate uninstall actions.
[0049] This invention provides a method for allocating long-slot general computing resources under task dependence, implemented using the aforementioned general computing resource collaborative scheduling algorithm. The method includes the following steps:
[0050] Step 1: After generating a computing or communication service, the user will first send a service request message to the main access point;
[0051] Step 2: Each access point forwards service request information and sends load status information to the CPU when the load status information changes;
[0052] Step 3: After the CPU collects the load status information of each access point, it executes the computing resource collaborative scheduling algorithm, and formulates a multi-access point joint service strategy by comprehensively considering factors such as the communication link status, computing power, and user-side constraints of the access points.
[0053] Step 4: The CPU notifies each access point and the corresponding user of the resource allocation decision information;
[0054] Step 5: The user uploads the task through the designated access point service cluster;
[0055] Step 6: Each access point service cluster forwards user tasks to the CPU;
[0056] Step 7: The CPU breaks down the original task into subtasks based on resource allocation decision information and distributes them to different access points; different subtasks will be distributed to multiple cooperating access points for execution;
[0057] Step 8: After processing the user task, the access point will send the task result back to the corresponding user.
[0058] Step 9: The access point feeds back the task completion information and load status information to the CPU. The CPU adjusts the computing resource collaborative scheduling algorithm based on the task completion information and load status information.
[0059] The beneficial effects of this invention are:
[0060] (1) This invention designs a resource scheduling mechanism that can dynamically adapt to changes in business needs in access point environments with limited computing resources. It can achieve coordinated scheduling of communication and computing resources in long time slots under a user-centric cellless network framework. By combining the status of communication and computing resources, tasks are reasonably allocated to different access points. By adjusting the task allocation ratio in real time, the system can maximize resource utilization, reduce the congestion risk of high-load nodes, and improve the overall task processing capacity while ensuring system stability.
[0061] (2) Compared with the traditional single-slot optimization method, the present invention adopts a long-slot task scheduling optimization strategy, which can better adapt to dynamic task arrival and changes in computing resources, avoid local optima, thereby reducing task execution latency and improving task unloading success rate over a long time scale.
[0062] (3) The parallel unloading model proposed in this invention, which supports dynamic adaptation to different task dependencies, is more flexible than existing methods. It can reasonably split tasks under complex task dependencies, improve the utilization of computing resources, and effectively reduce the additional latency caused by the limited execution order of tasks. Attached Figure Description
[0063] Figure 1 This is a hybrid scheduling model for multi-user communication and computing resources.
[0064] Figure 2 This is a flowchart for resource scheduling and allocation decision-making. Detailed Implementation
[0065] The present invention will be further described in detail below with reference to the accompanying drawings.
[0066] This invention provides a method for allocating long-slot general computing resources under task-dependent conditions, considering, for example... Figure 1 The user-centric, cell-free multi-user radio offloading scenario shown depicts a hybrid scheduling model for multi-user communication and computing resources, comprising single-antenna users, single-antenna access points (APs), and servers (CPUs). Access points possess computing resources. Assuming a Ricean fading channel between the access point and the user, the access point and user are connected via a data offloading link. The access point is connected to the CPU via an error-free ideal fronthaul link, and data exchange between access points is achieved through X2 links. Based on the basic flow of the 3GPP TS 38.331 radio resource control protocol, the hybrid scheduling model for multi-user communication and computing resources executes the following at the beginning of each time slot: Figure 2 The resource scheduling and allocation decision-making process is shown below:
[0067] Step 1: Request information is generated and uploaded.
[0068] After generating a computing or communication service, the user will first send a service request message to the main access point.
[0069] Specifically, the service request information mainly includes: service type, estimated data size, required communication and computing resources, and its own transmit power constraints.
[0070] Step 2: Forward the request and upload load information;
[0071] Each access point forwards service request information and sends load status information to the CPU when the load status information changes.
[0072] Specifically, the load status information mainly includes: the duration required for the current pending task and the remaining amount of computing resources that can be allocated.
[0073] Step 3: Resource allocation decision-making;
[0074] After the CPU collects the load status information of each access point, it executes the collaborative scheduling algorithm for computing resources, i.e., the resource scheduling and allocation decision-making process. This process comprehensively considers factors such as the communication link status, computing power, and user-side constraints of the access points to formulate a joint service strategy for multiple access points.
[0075] Step 4: Feedback of decision-making information;
[0076] The CPU will notify each access point and the corresponding user of the resource allocation decision information.
[0077] Specifically, the resource allocation decision information mainly includes the composition of the access point service cluster and the task allocation relationship.
[0078] Step 5: Upload the task;
[0079] Users upload tasks through a defined access point service cluster;
[0080] Step Six: Forward the task;
[0081] Each access point service cluster forwards user tasks to the CPU;
[0082] Step 7: CPU allocation of subtasks and task processing;
[0083] The CPU breaks down the original task into subtasks based on resource allocation decision information and distributes them to different access points. Different subtasks will be distributed to multiple cooperating access points for execution, and the forwarding of task results between access points is completed through the X2 link.
[0084] Step 8: Task Result Feedback;
[0085] After processing the user's task, the access point will send the task result back to the corresponding user.
[0086] Step 9: Task and load information feedback;
[0087] The access point feeds back task completion information and load status information to the CPU, and the CPU adjusts the computing resource collaborative scheduling algorithm based on the task completion information and load status information.
[0088] Regarding the resource scheduling and allocation decision-making process described above, this invention focuses on steps three, seven, and nine, namely the general computing resource collaborative scheduling algorithm executed in the CPU. The specific implementation process of this general computing resource collaborative scheduling algorithm is as follows:
[0089] S1: The CPU collects the maximum transmission power p uploaded by user n. n Information such as task size, combined with the real-time load status information of each access point, serves as known inputs to the collaborative scheduling algorithm for computing resources.
[0090] S2: The CPU invokes the general computing resource collaborative scheduling algorithm to complete the resource scheduling and allocation decision-making process; its specific implementation process is as follows:
[0091] S2.1: CPU calculates data upload rate;
[0092] At the beginning of each time slot, Let represent the uplink signal of the k-th user at time t, with power . This indicates that the mean is 0 and the variance is 0. The channel state follows a complex Gaussian distribution. Under the block fading assumption, the channel state remains constant over a certain time interval but varies independently between different frames. Let... This represents the channel state between the m-th access point and the k-th user at time t. This represents the channel estimate based on least squares, where the selector-combiner vector for the m-th access point is... And there are The uplink signal-to-noise ratio of the k-th user at time t can be expressed as:
[0093]
[0094] in, This represents the access point selection and receiving merge vector matrix for user k at time t. express The conjugate transpose of the matrix. This represents the access point selection matrix for the k-th user at time t. express The conjugate transpose of the matrix. This represents the uplink transmission power of the k-th user at time t. Let σ represent the channel state matrix of user i at time t. 2The received noise power is represented by K, the number of users is represented by T, and the superscript H represents the conjugate transpose.
[0095] In the connection matrix D t In, elements This indicates that the m-th access point is connected to the k-th user at time t. This indicates no connection. Let B represent the frequency bandwidth. The data transmission rate of the k-th user at time t can be expressed as:
[0096]
[0097] Where, τ c τ represents the channel coherence period. p This indicates the pilot transmission time.
[0098] S2.2: The CPU calculates the latency consumption of each user task;
[0099] This invention designs a parallel offloading mechanism oriented towards task dependencies. For complex task dependencies, a parallel offloading mechanism supporting different task topologies is proposed, enabling tasks to be rationally split and distributed to multiple access points for execution according to their dependencies. This reduces the additional latency caused by constraints on task execution order and improves the processing efficiency of computational tasks.
[0100] Specifically, assuming the arrival amount It follows a general independent and identically distributed distribution with bounded second moments, i.e.: Assume η k The value is known, for example, it can be estimated through past observations. Then, each user's task will be... It can be divided into I subtasks with data dependencies.
[0101] First, the transmission delay of the k-th user to the CPU at time t can be expressed as:
[0102]
[0103] Then, at time t, the CPU assigns the subtask to the access point for parallel offloading.
[0104] Specifically, let Q m (t) represents the time required for the queue of tasks to be processed at time t on the m-th access point. For the subtask of the k-th user... make Subtasks It is assigned to be processed by the m-th access point. Subtasks The processing does not occur at the m-th access point. This indicates that the m-th access point is assigned to the subtask. If the computing resources are limited, then the computation latency can be expressed as:
[0105]
[0106] Where υ represents the conversion factor between the task data size and the required number of cycles. Subtasks Assignment matrix, Indicates server-side subtasks Calculate the transpose of the resource allocation matrix.
[0107] make This indicates that subtask x requires the output of subtask y as input data. Considering task queuing delay, this invention assumes that each access point processes the tasks left over from the previous time slot before continuing to process the tasks in the current time slot, and that Q... t ={Q1(t),Q2(t),…,Q m If (t)}, then the processing delay of the subtask can be expressed as Therefore, the task processing latency can be expressed as:
[0108]
[0109] in, This means that user k's subtask x requires the output of subtask y as input data. This indicates that user k's subtask x does not require the output of subtask y as input data. The time matrix Q represents the time required for the task queues to be processed at each access point. t The transpose of .
[0110] Therefore, at time t, the total processing latency of the k-th user task can be expressed as:
[0111]
[0112] Considering that transmission latency is much smaller than task processing latency, when calculating queue changes, the impact of transmission latency is ignored. Therefore, a dynamic queue can be modeled as follows:
[0113]
[0114] Among them, Q m (1) = 0, and t g This indicates the interval between each time slot.
[0115] S2.3: Based on the data processed above, the CPU constructs a long time slot task scheduling optimization model and solves it to obtain the optimal long time slot task scheduling scheme.
[0116] This invention proposes a long-timeslot optimized task scheduling strategy based on Lyapunov optimization, aiming to minimize the average total processing latency of long-running user tasks. It jointly optimizes user transmit power, access point selection, task allocation, and computing resource allocation. In environments with heterogeneous computing resources across multiple access points and significant load fluctuations, the long-timeslot optimized task scheduling strategy based on Lyapunov optimization enhances the global optimization capability of task scheduling while ensuring system stability. This avoids the problem of getting trapped in local optima in single-timeslot optimization and improves the success rate of task offloading and the long-term performance of the system through queue stability constraints.
[0117] Specifically, the long time-slot task scheduling optimization model is expressed in the following form:
[0118]
[0119] Where G represents the total time, K represents the number of users, and M represents the number of access points; ζ represents the maximum number of access points each user can connect to; D, P, Q, and F all represent decision matrices, where D represents the AP connection matrix, P represents the user upload power matrix, Q represents the subtask allocation matrix, and F represents the computing resource allocation matrix. Constraint C1 indicates that the elements of the AP connection matrix D are binary variables. Constraint C2 limits the maximum number of access points each user can connect to. Constraint C3 indicates that the maximum user data transmission power is p. max Constraint C4 states that the elements of the subtask allocation matrix Q are binary variables. Constraint C5 states that tasks can only be assigned to connected access points. Constraint C6 states that subtasks can only be processed on one access point. Constraint C7 states the stability of the queue at each access point. Constraint C8 states that the resources allocated to the m-th access point do not exceed its upper limit f. m .
[0120] The aforementioned problem is a multi-stage mixed-integer nonconvex optimization problem, which is extremely challenging to solve because it involves both discrete and continuous variables. Therefore, this invention first utilizes Lyapunov optimization to decouple and decompose the original long-slot task scheduling optimization problem into single-slot optimization subproblems.
[0121] Specifically, this invention introduces Lyapunov functions. Drifting with Lyapunov
[0122] To minimize the average total processing latency of long-running user tasks while stabilizing the queue Z(t), a drift-penalty minimization method is employed. Specifically, in each time frame, an attempt is made to minimize the upper bound of the following drift-penalty expression:
[0123]
[0124] Here, V serves as a weighting variable, used to adjust the proportion of the penalty term.
[0125] make We can obtain:
[0126]
[0127] Summing the M queues on both sides of the above equation, we get:
[0128]
[0129] Because the amount of work reached meets the requirements It is bounded, therefore we can obtain:
[0130]
[0131] Where θ is a constant.
[0132] Therefore, the upper bound of the drift plus penalty expression can be derived as follows:
[0133]
[0134] By removing constant terms from the observations starting at time frame t, the algorithm determines the action by minimizing the following:
[0135]
[0136] Considering the constraints of each frame, the following deterministic per-frame subproblem is solved in the t-th time frame:
[0137]
[0138] The objective function of the above problem is non-convex, making it difficult to directly obtain the optimal solution. By decoupling the original long time slot task scheduling optimization problem into a main problem and a series of sub-problems, and then solving it through reinforcement learning and other methods.
[0139] Specifically, for the deterministic sub-problem of each frame, the original problem can be decoupled and decomposed into a main problem and a series of sub-problems. The main problem can be represented as the access point selection problem:
[0140]
[0141] The subproblem can be represented as:
[0142]
[0143] The subproblem can be decoupled into a power allocation subproblem and a computational resource allocation subproblem, which can be solved by fractional programming and linear programming methods, respectively.
[0144] For the main problem, an intelligent access point selection strategy using Actor-Critic (AC) reinforcement learning can be employed. Specifically, the Actor module is responsible for policy learning, selecting the optimal action based on the current state. It is typically implemented using a deep neural network, taking the environment state as input and outputting an action distribution or a specific action decision. The Actor module optimizes the policy using the policy gradient method to maximize the expected reward of the selected action. The Critic module is responsible for value function estimation, evaluating the value of the current state or state-action pair and providing feedback to the Actor module to guide policy improvement. The Critic module optimizes the value function by calculating the Temporal Difference Error (TD) to measure the difference between the actual reward and the estimated value, thus making it more accurately assess the long-term reward of the state.
[0145] Specifically, the Actor module consists of a deep neural network and an action quantizer. Using α... t This represents the parameter values of the deep neural network at time slot t, where α is at t=1. t Initialized as random values following a standard normal distribution. Channel estimation and queue state are used as observations. Input a deep neural network, and the deep neural network outputs an unloading decision. Right now
[0146] To adapt to different network environments and computing resource allocation situations, this invention chooses to dynamically adjust the number L of candidate decisions. t The specific update strategy is as follows:
[0147] L t =min(max({l i mod L max |i∈[t-Δt,t]}+2,L max )
[0148] Among them, l i The number of candidate decisions from time t-Δt to time t is traversed, L max As a constant, the maximum number of candidate decisions is limited. This method selects the number of candidate decisions within a certain time interval and increases it appropriately to increase the exploratory nature while limiting the maximum number of selections, thereby balancing exploration and computational overhead.
[0149] Based on noise scrambling, multiple candidate decision-making methods are generated. Specifically, the OPN method first applies noise scrambling to continuous variables. Perform direct quantization to generate the first set of binary actions. The calculation method is as follows:
[0150]
[0151] Where i = 1, ..., M·K. Next, through continuous variables... Random Gaussian noise n ~ N(0, I) is introduced above. N ), thus obtaining the perturbed variables:
[0152]
[0153] Sigmoid(·) is a mapping function used to map each component to the range (0,1).
[0154] Then, based on the perturbed variables Regenerate the remaining M t -1 group of candidate unloading actions. Specifically, for the perturbed variables... The distances of each element to 0.5 are sorted in ascending order, and these distances are used sequentially as the new quantization thresholds.
[0155] For m = 2, ..., M t The second group and subsequent actions Calculate using the following method:
[0156]
[0157] The Critic module selects the best action x t The strategy can be expressed as Ω t This represents the alternative access point selection scheme at time t. When training the deep neural network, the selection is made from (∈...) t ,x t The data is used as labeled input and output samples, and an empirical replay with a replay capacity of u is maintained. Since the initial memory is empty, the deep neural network is trained after collecting more than u / 2 data samples, and a batch of data samples is randomly selected for training in every δ time slots to avoid overfitting.
[0158] By minimizing the cross-entropy loss function PL(α) t To update the parameters of the deep neural network, use the Adam optimization algorithm for training:
[0159]
[0160] Among them, S t |S represents the set of time indices of the selected samples. t | indicates the sample batch size.
[0161] Once training is complete, update the Actor module using the updated parameters in the next time frame: α t+1 ←αt .
[0162] The algorithm described above can be used to solve the deterministic sub-problem for each frame.
[0163] S3: The CPU sends the optimal long time slot task scheduling scheme parameters to the requesting user and each access point providing the service.
[0164] All users are unloaded according to the above optimal long time slot task scheduling scheme. All service access points receive the optimal long time slot task scheduling scheme parameters, complete the calculation, and return the results. Each access point updates the load status information to the CPU.
[0165] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A collaborative scheduling algorithm for general computing resources, characterized in that, Includes the following steps: (1) The CPU calculates the data upload rate; At the beginning of each time slot, the CPU calculates the user's uplink signal-to-noise ratio (SNR) based on the user's uplink signal, the channel state between the access point and the user, the channel estimate, and the receiver combining vector selected by the access point. Then, the CPU calculates the user's data transmission rate based on the uplink SNR. (2) The CPU calculates the latency consumption of each user task; Each user's task is divided into multiple subtasks with data dependencies; the CPU first calculates the transmission latency when the user transmits the task, and then assigns the subtasks to the access point for parallel offloading. The CPU calculates the task processing latency based on the processing latency of the subtasks, and then calculates the total processing latency for the user. At the same time, it models the dynamic queue based on the total processing latency. (3) The CPU constructs a long time slot task scheduling optimization model and uses the Lyapunov optimization method to decouple the long time slot task scheduling optimization problem into a single time slot optimization subproblem that can be solved sequentially. The optimal long time slot task scheduling scheme is obtained by using reinforcement learning methods. (4) The CPU sends the optimal long time slot task scheduling scheme to the requesting user and each access point that provides the service.
2. The general computing resource collaborative scheduling algorithm according to claim 1, characterized in that, The goal of the Lyapunov optimization method is to minimize the average total processing latency of long-term user tasks by jointly optimizing user transmission power, access point selection, task allocation, and computing resource allocation.
3. The general computing resource collaborative scheduling algorithm according to claim 1, characterized in that, The mathematical expression for the long time-slot task scheduling optimization model is: Where G represents the total time, K represents the number of users, and M represents the number of access points; D, P, Q, and F all represent decision matrices, D represents the AP connection matrix, P represents the user upload power matrix, Q represents the subtask allocation matrix, and F represents the computing resource allocation matrix. This represents the total processing latency of the k-th user task; Indicates user data transmission power; Q m (t) represents the time required for the queue of tasks to be processed at time t on the m-th access point; This indicates that the m-th access point is assigned to the subtask. The computing resources; ζ represents the maximum number of access points that each user can connect to; constraint C1 indicates that the elements of the AP connection matrix D are binary variables. This indicates that the m-th access point is connected to the k-th user at time t. This indicates no connection; constraint C2 limits the maximum number of access points each user can connect to; constraint C3 indicates the maximum user data transmission power is p. max Constraint C4 indicates that the elements of the subtask assignment matrix Q are binary variables. Subtasks It is assigned to be processed by the m-th access point. Subtasks The processing does not occur at the m-th access point; constraint C5 indicates that tasks can only be assigned to connected access points; constraint C6 indicates that subtasks can only be processed on one access point; constraint C7 indicates the stability of the queue at each access point; constraint C8 indicates that the resources allocated to the m-th access point do not exceed its upper limit f. m .
4. The general computing resource collaborative scheduling algorithm according to claim 3, characterized in that, The Lyapunov optimization method employs a drift-penalty minimization approach to stabilize the queue Z(t) while minimizing the average total processing latency of long-term user tasks. This is achieved by introducing a Lyapunov function. Drifting with Lyapunov Minimize the upper bound of the drift plus penalty expression in each time frame.
5. The general computing resource collaborative scheduling algorithm according to claim 4, characterized in that, The drift penalty expression is: Where V is the weighting variable, used to adjust the proportion of the penalty term; Q m (t) represents the time required for the queue of tasks to be processed at time t on the m-th access point; Let L(Z(t+1)) represent the total processing latency of the k-th user task; L(Z(t+1)) represents the Lyapunov function expression at time t+1. The upper bound of the drift plus penalty expression is: Where θ is a constant term; t g Indicates the interval between each time slot; Subtasks It is assigned to be processed by the m-th access point. Subtasks The processing does not occur at the m-th access point; Indicates computation delay; This represents the total processing latency of the k-th user task; Starting from the t-th time frame, remove the constant term θ and determine the action by minimizing the following: Considering the constraints of each frame, solve the following deterministic per-frame subproblem in the t-th time frame: stC1~C6,C8.
6. The general computing resource collaborative scheduling algorithm according to claim 1, characterized in that, In the reinforcement learning method, the main problem of the long time-slot task scheduling optimization problem is represented as the access point selection problem, and the intelligent access point selection strategy of Actor-Critic reinforcement learning is used to solve the main problem; the subproblems of the long time-slot task scheduling optimization problem are decoupled into power allocation subproblems and computing resource allocation subproblems, which are solved by fractional programming and linear programming methods, respectively.
7. The general computing resource collaborative scheduling algorithm according to claim 6, characterized in that, The specific implementation process of the intelligent access point selection strategy of Actor-Critic reinforcement learning is as follows: The Actor module is responsible for policy learning, selecting the optimal action based on the current state; the policy is to dynamically adjust the number of candidate decisions; the candidate decisions are generated using a noise scrambling method. The Critic module is responsible for value function estimation, evaluating the value of the current state or state-action pair, and feeding it back to the Actor module to guide policy improvement.
8. The general computing resource collaborative scheduling algorithm according to claim 7, characterized in that, The Actor module consists of a deep neural network and an action quantizer. The input to the deep neural network is the channel estimate and the queue state, and the output of the deep neural network is the offloading decision. The parameter α of the deep neural network... t By minimizing the cross-entropy loss function PL(α) t The Actor module is updated using the updated parameters and trained using the Adam optimization algorithm. After training, the updated parameters are used to update the Actor module in the next time frame.
9. The general computing resource collaborative scheduling algorithm according to claim 7, characterized in that, In the aforementioned noise scrambling-based generation method, the continuous variable is first processed... Perform direct quantization to generate the first set of binary actions. By continuous variables Random Gaussian noise n ~ N(0, I) is introduced above. N ), thus obtaining the perturbed variables: Here, Sigmoid(·) is a mapping function used to map each component to the range (0,1); then, based on the perturbated variables... Regenerate the remaining M t -1 group of candidate uninstall actions.
10. A method for allocating long-slot general computing resources under task dependence, characterized in that, The method is implemented using the collaborative scheduling algorithm for general computing resources as described in any one of claims 1-9, and includes the following steps: Step 1: After generating a computing or communication service, the user will first send a service request message to the main access point; Step 2: Each access point forwards service request information and sends load status information to the CPU when the load status information changes; Step 3: After the CPU collects the load status information of each access point, it executes the computing resource collaborative scheduling algorithm, and formulates a multi-access point joint service strategy by comprehensively considering factors such as the communication link status, computing power, and user-side constraints of the access points. Step 4: The CPU notifies each access point and the corresponding user of the resource allocation decision information; Step 5: The user uploads the task through the designated access point service cluster; Step 6: Each access point service cluster forwards user tasks to the CPU; Step 7: The CPU breaks down the original task into subtasks based on resource allocation decision information and distributes them to different access points; different subtasks will be distributed to multiple cooperating access points for execution; Step 8: After processing the user task, the access point will send the task result back to the corresponding user. Step 9: The access point feeds back the task completion information and load status information to the CPU. The CPU adjusts the computing resource collaborative scheduling algorithm based on the task completion information and load status information.