Adaptive CPU priority scheduling method based on near-end strategy optimization

By adopting an adaptive CPU priority scheduling method based on near-end policy optimization, this paper solves the problem of balancing multi-dimensional performance indicators of traditional CPU scheduling algorithms in multi-task environments. It realizes the dynamic adaptability and multi-objective optimization of the system, improves system performance and scheduling fairness, and is suitable for multi-core heterogeneous architectures and resource-constrained platforms.

CN120973515APending Publication Date: 2025-11-18SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511032199.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing CPU scheduling algorithms struggle to balance performance metrics such as system throughput, response time, and task starvation in multi-task concurrent environments. Furthermore, traditional reinforcement learning models are difficult to implement in operating systems while meeting the requirements of real-time performance, reliability, and lightweight design.

Method used

An adaptive CPU priority scheduling method based on near-end policy optimization is adopted. Through state acquisition and temporal feature extraction, a policy network and a value network are constructed, a composite reward function is designed, and a course-based training mechanism is adopted to realize dynamic perception of task status and intelligent adjustment of scheduling strategy. Combined with max-heap optimization, efficient scheduling is achieved.

Benefits of technology

It achieves dynamic adaptability and multi-objective collaborative optimization of the system in complex task scenarios, improves the overall system performance and scheduling fairness, is suitable for resource-constrained platforms, and has significant performance advantages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973515A_ABST
    Figure CN120973515A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive CPU (central processing unit) priority scheduling method based on near-end strategy optimization, which comprises the following steps of: firstly, establishing a task scheduling model in a simulation environment, performing matrix processing on a current task queue state and a task characteristic into an environment state for input, and extracting a task sequence time sequence characteristic by introducing a gating circulation unit; and respectively outputting task priority strategy distribution and state value estimation by using a dual-network architecture of PPO. In the training process, a course learning mechanism is adopted to gradually transit from a short task load scene to a complex mixed load. During operation, the trained scheduling agent is deployed to an operating system kernel scheduler, the priority of each task is output in real time according to the current state, and a rapid scheduling decision is realized by means of a priority heap. Through a multi-target reward function, the intelligent agent is guided to optimize the system throughput, the task average response time and the hunger prevention index at the same time, and extreme unfair scheduling is avoided. Experiments show that the method shows relatively high performance, stability and robustness in heterogeneous task, high concurrency and long-tail load scenes, and the system resource utilization rate and the user experience are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer operating system resource management and task scheduling technology, and in particular to an adaptive CPU priority scheduling method based on Proximity Policy Optimization (PPO), belonging to the cross-technical field of computer operating system scheduling strategies and intelligent scheduling methods. Background Technology

[0002] With the increasing diversification of computing resources and the widespread application of heterogeneous systems, task scheduling problems in modern operating systems have become increasingly complex. Traditional CPU scheduling strategies, such as Round Robin (RR), Multi-Level Feedback Queue (MLFQ), and Completely Fair Scheduler (CFS), while possessing certain scheduling efficiency in general scenarios, commonly suffer from problems such as response lag, rigid scheduling strategies, and static priority configuration when facing complex scenarios with high task density, diverse load types, or frequent dynamic changes in resources.

[0003] Furthermore, existing schedulers mostly employ rule-based or heuristic methods, with optimization objectives often singular, such as optimizing only task turnaround time or response time, neglecting the trade-off between overall system throughput and task fairness. In the context of multi-core, heterogeneous architectures, schedulers not only need to address the differences in execution time, computational intensity, and resource requirements among different tasks, but also must achieve rapid adaptation and generalization to different execution environments. Traditional algorithms often face performance degradation issues in such environments.

[0004] In recent years, reinforcement learning (RL) technology has made significant progress in fields such as intelligent decision-making and resource scheduling. In particular, proximal policy optimization (PPO) algorithms, due to their good convergence stability and high sample efficiency, have become one of the most widely used policy gradient methods in reinforcement learning. However, directly applying RL algorithms to schedulers at the operating system level still faces several challenges, such as difficulties in state space modeling, high real-time requirements, complex reward design, and limited model deployment overhead.

[0005] Some studies have attempted to incorporate algorithms such as Q-Learning and DQN into scheduling optimization, but these generally suffer from insufficient state abstraction capabilities, unstable model training, and inability to generalize to unknown task distributions. Furthermore, many scheduling reinforcement learning models are too large or heavily dependent on training samples, making it difficult to meet the real-time, reliability, and lightweight requirements of operating system scheduling. Therefore, there is still an urgent need for an efficient adaptive scheduling method that combines real-time system state awareness, task historical behavior feedback, and multi-objective scheduling optimization capabilities. Summary of the Invention

[0006] In view of this, the present invention proposes an adaptive CPU priority scheduling method based on near-end policy optimization. The purpose of the present invention is to overcome the problem that existing scheduling algorithms are difficult to balance multiple performance indicators such as system throughput, response time and task starvation in multi-task concurrent environments. The present invention provides a scheduling method that can dynamically sense task status, intelligently adjust scheduling policy, and achieve online adaptive optimization through reinforcement learning, thereby achieving high efficiency, stability and generalization of the scheduling policy.

[0007] The technical solution adopted by the present invention to achieve the above objectives is as follows:

[0008] An adaptive CPU priority scheduling method based on near-end policy optimization, the method includes the following steps:

[0009] 1) Status Acquisition and Temporal Feature Extraction: Collect the status information of the current queue of tasks to be scheduled, including but not limited to the following key scheduling features of each task: unique task identifier, arrival timestamp, total number of task instructions, number of remaining execution instructions, and priority assigned during the last scheduling; normalize the collected status information through a feature encoder to convert the multi-dimensional status data into a unified status vector representation.

[0010] 2) Construction of Policy Network and Value Network: A policy network and a value network are constructed to predict the distribution of action policies and evaluate the value of states based on state information, respectively. The policy network outputs the selection probability distribution of each task to be scheduled under different priorities, and the value network evaluates the long-term expected return corresponding to the current scheduling state. Both networks adopt a neural network structure to form a dual-network architecture based on the near-end policy optimization (PPO) algorithm.

[0011] 3) Design a composite reward function: Construct a reward function by combining multiple performance indicators, including a positive incentive for task throughput, a penalty term for average response time, and an adjustment factor for the starvation prevention indicator. This function is used to adjust the weight parameters of the policy network and the value network through backpropagation based on the composite reward function, so as to guide the dual-network architecture to achieve scheduling behavior that balances fairness and efficiency under different task structures and load conditions.

[0012] 4) Training strategy and curriculum-based training mechanism: A curriculum-based training mechanism is introduced to construct a phased distribution of training tasks and enhance the generalization ability of the model in various scheduling scenarios. During the training process, a method based on the proximal policy optimization algorithm PPO-Clip combined with generalized advantage estimation (GAE) is adopted, and the policy network and value network are optimized independently.

[0013] 5) Strategy Deployment and Real-time Decision-Making: The ideal policy network obtained after final training is integrated into the CPU scheduler framework as a scheduler. It receives task status information fed back by the operating system in real time, executes the calculation output corresponding priority adjustment strategy, and realizes fast scheduling decision-making with the help of priority heap, thereby driving the system to complete efficient dynamic scheduling.

[0014] The task status information collected in step 1) is represented as a fixed-dimensional two-dimensional matrix with a size of P×Q, where P represents the maximum number of tasks supported in the ready queue and Q represents the number of key scheduling features for each task. After normalization, the two-dimensional matrix is ​​reshaped into a three-dimensional tensor with a shape of (batch_size, P, Q) before being input into the policy network and the value network, in order to preserve the sequence structure and temporal features of the tasks.

[0015] The policy network and value network constructed in step 2) have the following structure:

[0016] The input layer receives a three-dimensional tensor of shape (batch_size, P, Q);

[0017] The sequence coding layer uses a gated recurrent unit (GRU) with a hidden state dimension of 64 to mine the temporal correlation and dynamic evolution features of the task sequence.

[0018] The hidden state output by GRU is used as a feature representation of the overall task sequence and input into a multilayer perceptron (MLP). The MLP contains a linear transformation layer (64-dimensional input and output), layer normalization (LayerNorm), ReLU activation, and Dropout to improve nonlinear expressive power and model generalization performance.

[0019] The policy network Actor outputs action logits of length 5, which are then transformed into an action probability distribution using softmax.

[0020] The value network Critic outputs a scalar value representing the expected reward for the current state.

[0021] The policy network and value network are trained independently to ensure the stability and convergence of the training process.

[0022] The composite reward function described in step 3) includes three sub-items: throughput incentive, response efficiency penalty, and task starvation prevention, as expressed below:

[0023]

[0024] The throughput incentive measures the system's processing capacity by the number of tasks completed per unit time. The use of a logarithmic function ensures the incentive effect for task completion while avoiding numerical instability in high-frequency scenarios.

[0025] The response efficiency penalty term constrains the system response latency based on the arithmetic mean of task turnaround times;

[0026] The task starvation prevention measure introduces task priority weights and dynamic thresholds, and applies secondary penalties to high-priority tasks that exceed the threshold waiting time to prevent critical tasks from being shelved for a long time.

[0027] The three hyperparameters w1, w2, and w3 control the contribution intensity of each sub-item, and the overall reward function takes into account throughput, response efficiency, and task hunger, which can effectively guide the reinforcement learning agent to achieve robust and efficient policy optimization in a multi-objective scheduling environment.

[0028] The course-based training mechanism sets different task distributions for different stages: the initial stage is characterized by short, intensive tasks; the middle stage gradually introduces a higher proportion of long tasks; and the final stage introduces non-uniform task distribution and resource bottlenecks to improve the model's adaptability and generalization ability to diverse scheduling scenarios.

[0029] The method employs a proximal policy optimization algorithm (PPO-Clip) combined with generalized advantage estimation (GAE) for training. This method limits the policy update magnitude to prevent performance fluctuations and improves training stability and sample utilization efficiency. The policy network and value network adopt a dual-network architecture and are independently optimized to ensure the stability and convergence of the learning process.

[0030] Based on the priority adjustment strategy output by the aforementioned strategy network, a max-heap priority queue structure is used to manage and select tasks to be scheduled:

[0031] Each task is inserted into the corresponding max-heap according to the priority value output by the policy model;

[0032] The scheduler accesses the top task of the max heap in each round of scheduling, and the time complexity of the scheduling selection operation is O(1).

[0033] When a task's status changes or it completes, the heap structure is automatically updated and the priority order is adjusted to ensure that high-priority tasks are executed first.

[0034] An adaptive CPU priority scheduling system based on near-end policy optimization includes the following modules built into the hardware platform:

[0035] The status acquisition and temporal feature extraction module is used to collect the status information of the current queue of tasks to be scheduled, including but not limited to the following key scheduling features of each task: unique task identifier, arrival timestamp, total number of task instructions, number of remaining execution instructions, and priority assigned during the last scheduling. The collected status information is normalized by a feature encoder to convert the multi-dimensional status data into a unified status vector representation.

[0036] The dual-network architecture module based on the Proximal Policy Optimization (PPO) algorithm includes a policy network and a value network, which are used to output action policy distribution prediction and state value evaluation based on state information, respectively. The policy network outputs the selection probability distribution of each task to be scheduled under different priorities, and the value network evaluates the long-term expected return corresponding to the current scheduling state. Both networks adopt a neural network structure.

[0037] The composite reward function module uses a composite reward function that combines positive incentives for task throughput, penalties for average response time, and adjustment factors for starvation prevention indicators to guide the dual network architecture module to obtain scheduling strategy behavior that balances fairness and efficiency under different task structures and load conditions.

[0038] The training module has a built-in course-based training mechanism and constructs a phased training task distribution to enhance the model's generalization ability in diverse scheduling scenarios. During the training process, a method based on the proximal policy optimization algorithm PPO-Clip combined with generalized advantage estimation (GAE) is used to independently optimize and train the policy network and value network.

[0039] The strategy deployment and real-time decision-making module controls the coordinated work of the above modules to form a closed-loop adaptive scheduling system. It completes iterative optimization training to obtain an ideal policy network, which is integrated into the CPU scheduler framework as a scheduler. It receives task status information fed back by the operating system in real time, executes calculations to output the corresponding priority adjustment policy, and realizes fast scheduling decisions with the help of priority heap, thereby driving the system to complete efficient dynamic scheduling.

[0040] The scheduling executor in the strategy deployment and real-time decision-making module: based on the task priority output by the strategy network, it adopts a maximum heap priority queue structure to achieve efficient task classification and scheduling, ensure that high-priority tasks are executed first, support dynamic task status updates and heap structure adjustments, and meet the needs of large-scale concurrent scheduling.

[0041] It also includes an evaluation and feedback module, which is used to monitor and evaluate the performance indicators of the scheduling system in real time, and use the feedback results to guide policy updates and scheduling optimization; the performance indicators include, but are not limited to, task throughput, response time, and turnaround time.

[0042] This invention achieves a CPU priority scheduling method with adaptive scheduling capabilities by integrating the Proximal Policy Optimization (PPO) reinforcement learning algorithm with multidimensional state modeling, a composite reward mechanism, and a curriculum-based training method. This method has the following significant advantages:

[0043] 1. Strong dynamic adaptability: The PPO algorithm can automatically adjust the scheduling strategy under different system load changes, so that the system has good real-time response capability and can adapt to complex task scenarios such as short task density, long tail task delay and resource fluctuation.

[0044] 2. Multi-objective collaborative optimization: By using a composite reward function to simultaneously consider multiple key indicators such as task throughput, response time, and turnaround time, the strategy is guided to improve the overall performance of the system while maintaining scheduling fairness and efficiency.

[0045] 3. Simple structure and easy deployment: The designed strategy network structure is lightweight, and the scheduling execution module adopts max heap optimization to achieve task scheduling operation with O(1) time complexity, which is suitable for resource-constrained platforms such as embedded and edge computing nodes.

[0046] 4. Significantly superior performance compared to traditional schedulers: Experimental results show that, on typical scheduling datasets, this method has significant advantages over traditional scheduling algorithms such as CFS and RR in terms of response time and throughput. Attached Figure Description

[0047] Figure 1 This is a network structure diagram for the PPO strategy.

[0048] Figure 2 This is a training flowchart.

[0049] Figure 3 This is a schematic diagram of the system architecture.

[0050] Figure 4 A comparison chart of turnaround times for the Blue Bar test.

[0051] Figure 5 A comparison chart of turnaround times for scalability testing.

[0052] Figure 6 This is a comparison chart of turnaround times for length sensitivity testing. Detailed Implementation

[0053] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0055] This invention provides an adaptive CPU priority scheduling method based on proximal policy optimization (PPO), aiming to improve system throughput, task response speed, and scheduling fairness in multi-process environments. This method utilizes the proximal policy optimization (PPO) algorithm in reinforcement learning, combined with task feature modeling, composite reward function design, and a course training mechanism to construct an intelligent scheduling model, which is then deployed online in an operating system or virtual scheduler.

[0056] 1) State acquisition and temporal feature extraction

[0057] In the implementation of this invention, it is first necessary to collect the state information of the current queue of tasks to be scheduled for the reinforcement learning model to perceive. The construction of this state space mainly includes the following steps:

[0058] Status acquisition includes, but is not limited to, the following key scheduling features for each task: unique task identifier, arrival timestamp, total number of instructions for the task, number of remaining instructions to be executed, and priority assigned during the last scheduling.

[0059] Feature extraction: State information is organized into a fixed-dimensional two-dimensional matrix with a size of 31×5, where 31 represents the maximum number of tasks supported in the ready queue, and 5 represents the key scheduling features of each task. This two-dimensional matrix is ​​then reshaped into a three-dimensional tensor with a shape of (batch_size, 31, 5) before being input into the policy network and value network to preserve the sequence structure and temporal features of the tasks.

[0060] 2) Scheduling strategy modeling and policy network design

[0061] In this invention, the scheduling strategy is modeled as a Markov decision process in reinforcement learning, including a state space S, an action space A, and a reward function R.

[0062] Action space definition: The action space is a discrete set of priority levels, divided into N discrete levels. The goal of the policy network is to select the appropriate priority action for each task.

[0063] like Figure 1 As shown, the policy network (Actor) and value network (Critic) both employ neural network structures, with the first few layers being identical and the last layer differing. Both take task sequence features of shape (batch_size, 31, 5) as input. First, a gated recurrent unit (GRU) extracts the temporal dynamic features of the task sequence, with a hidden state dimension of 64. Subsequently, the hidden states output by the GRU undergo non-linear mapping via a multilayer perceptron (MLP), including linear transformation, layer normalization, ReLU activation, and Dropout, enhancing the model's expressive power and generalization performance. The last layer of the policy network (Actor) outputs 5-dimensional action logits, which are transformed into an action probability distribution via softmax; the last layer of the value network (Critic) outputs a scalar of the expected reward value for the current state. The two networks are trained independently to ensure training stability and convergence.

[0064] The PPO-Clip policy optimization mechanism uses the PPO-Clip algorithm to limit the policy update magnitude and avoid drastic policy fluctuations. Generalized advantage estimation (GAE) is employed to calculate the advantage function, reducing the variance of gradient estimation and improving learning stability.

[0065] 3) Reward function design

[0066] To achieve policy optimization in multi-objective scheduling tasks, this invention proposes a structured and scalable multi-objective reward function based on traditional reinforcement learning reward mechanisms. This function effectively balances multi-dimensional optimization objectives such as system throughput, response efficiency, and task fairness through the synergistic effect of three key components. Specifically, during the scheduling process, after each task scheduling decision by the agent, the environment calculates a composite reward function based on the current system state and task queue status.

[0067] The composite reward function consists of three sub-items: throughput incentive, response efficiency penalty, and task starvation prevention, expressed as follows:

[0068]

[0069] Throughput incentive term is the number of tasks N completed per unit time. completed To measure the system's processing capacity, the use of a logarithmic function not only ensures the incentive effect for task completion but also avoids numerical instability in high-frequency scenarios.

[0070] Response efficiency penalty is based on task turnaround time Tturnaround (i) is the arithmetic mean, which constrains the system response delay;

[0071] Task-based starvation prevention introduces task priority weights w i With dynamic threshold T th High-priority tasks that exceed the threshold waiting time will be penalized a second time to prevent critical tasks from being shelved for a long time.

[0072] The three hyperparameters w1, w2, and w3 control the contribution intensity of each sub-item, and the overall reward function takes into account throughput, response efficiency, and task hunger, which can effectively guide the reinforcement learning agent to achieve robust and efficient policy optimization in a multi-objective scheduling environment.

[0073] 4) Training strategies and curriculum training mechanisms

[0074] like Figure 2 The diagram shows the training process. To enhance the model's generalization ability under diverse scheduling scenarios, a course-based training mechanism is introduced, with a phased distribution of training tasks: the initial training phase uses only short, intensive tasks; the middle phase introduces long tasks and dependencies between tasks; and the later training phase introduces complex environments such as unbalanced loads and high-concurrency queues. The training strategy gradually progresses from simple to complex to improve generalization ability.

[0075] During training, a method combining Proximal Policy Optimization (PPO-Clip) with Generalized Advantage Estimation (GAE) is employed to limit the policy update magnitude to prevent performance fluctuations and improve training stability and sample utilization efficiency. The policy network and value network are optimized independently to ensure the stability of the training process and the effective improvement of the policy.

[0076] 5) Scheduler Design and Model Deployment

[0077] Model Deployment: The trained policy network model is deployed in the core scheduler module of the scheduling system. The scheduler inputs the current state vector into the policy network based on the real-time state to obtain the optimal priority action.

[0078] Scheduler architecture design: Use a max-heap structure to store tasks, with the top of the heap always being the highest priority task, ensuring O(1) priority access time complexity.

[0079] The scheduling process is as follows:

[0080] The system collects the status of tasks to be scheduled;

[0081] Input the state vector into the Actor network, and output the priority action;

[0082] Insert into the max-heap according to priority;

[0083] Schedule tasks from the top of the heap to run and update the system status;

[0084] Repeat the above process.

[0085] 6) Experimental verification and performance evaluation

[0086] Experimental environment setup: The experiment is based on real scheduling datasets Google Cluster Trace and Alibaba Trace; in the scheduling simulation platform, the scheduling behavior of 1000+ task queues and multi-core CPU resources is simulated;

[0087] Compared with existing algorithms, including CFS (Completely Fair Scheduler), RR (Round-Robin Scheduler), and MLQ (Multi-Level Feedback Queue).

[0088] Evaluation metrics: average task response time; task turnaround time; system throughput (number of tasks completed per unit time).

[0089] Experimental results:

[0090] In a chi-square distributed load scenario, the average turnaround time is reduced by approximately 20% compared to CFS. Figure 4 Even in a large-scale test set with a task size of 10,000, it maintained stable performance, demonstrating good scalability and generalization ability. Figure 5 ).

[0091] In the task length sensitivity test, the method of this invention has a better turnaround time, which can ensure the continuous execution of long tasks, such as... Figure 6 For short tasks (such as 20 instructions), the response time is shown in Table 1 below (key performance indicators for a maximum of 20 instructions). While enhancing the stability of long task execution, it does not significantly affect the scheduling efficiency of short tasks.

[0092] Table 1

[0093] Scheduling Algorithm Turnaround time (time slice) Response time (time slice) Total running time (seconds) RR 1.278 0.170 0.0159 MLQ 1.246 0.225 0.0284 CFS 1.305 0.12 0.0391 ML_Prio 1.208 0.231 0.1269

[0094] 7) System Integration and Application Scenarios

[0095] like Figure 3 As shown, the system architecture constructed according to the method of this invention includes: a data generation layer, responsible for providing data for model training and performance validation. It is connected to the PPO training environment layer via a down arrow.

[0096] The PPO training environment consists of three key components forming a closed-loop system: (1) Gym simulation environment: generates a "state space" output, which serves as the input source for the PPO policy network. (2) PPO policy network (core AI module): receives the state space input, outputs the "priority decision" to the scheduler, and feeds it back to the reward calculator through the "gradient signal". (3) Reward calculator: receives the gradient signal, calculates the reward value, and feeds it back. Finally, the ML-Priority scheduler is obtained through training.

[0097] The scheduling execution layer has a "Scheduler" control module at the top, which integrates four major scheduling algorithms: MLQ (Multilevel Queuing), CFS (Completely Fair Scheduler), FIFO (First In First Out), and RR (Round-Robin Scheduler). The lower layer connects to the "Performance Metrics Calculation" module, which evaluates the performance of each scheduling algorithm and outputs the results to the "Visual Analysis" interface.

[0098] The scheduling method proposed in this invention can be integrated into various practical application environments, including but not limited to:

[0099] Operating system scheduling module: replaces the default static priority allocation mechanism to achieve dynamic scheduling;

[0100] Container orchestration platforms (such as Kubernetes): combine Pod-level task characteristics to improve task scheduling efficiency through intelligent priority control;

[0101] Edge computing nodes: Enhancing the responsiveness of AI inference / edge computing applications in resource-constrained environments;

[0102] Virtualization scheduler / cloud platform: Supports multi-tenant resource allocation and dynamically assigns priority weights among different tenants.

[0103] The above descriptions are merely specific embodiments of this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An adaptive CPU priority scheduling method based on near-end strategy optimization, characterized in that, The method includes the following steps: 1) Status Acquisition and Temporal Feature Extraction: Collect the status information of the current queue of tasks to be scheduled, including but not limited to the following key scheduling features of each task: unique task identifier, arrival timestamp, total number of task instructions, number of remaining execution instructions, and priority assigned during the last scheduling; normalize the collected status information through a feature encoder to convert the multi-dimensional status data into a unified status vector representation. 2) Construction of Policy Network and Value Network: A policy network and a value network are constructed to predict the distribution of action policies and evaluate the value of states based on state information, respectively. The policy network outputs the selection probability distribution of each task to be scheduled under different priorities, and the value network evaluates the long-term expected return corresponding to the current scheduling state. Both networks adopt a neural network structure to form a dual-network architecture based on the near-end policy optimization (PPO) algorithm. 3) Construct a composite reward function: Combine multiple performance metrics to construct a reward function, including a positive incentive for task throughput, a penalty term for average response time, and an adjustment factor for starvation prevention metrics. This function is used to adjust the weight parameters of the policy network and the value network through backpropagation based on the composite reward function, thereby guiding the dual-network architecture to achieve scheduling behavior that balances fairness and efficiency under different task structures and load conditions. 4) Training strategy and curriculum-based training mechanism: A curriculum-based training mechanism is introduced to construct a phased distribution of training tasks and enhance the generalization ability of the model in various scheduling scenarios. During the training process, a method based on the proximal policy optimization algorithm PPO-Clip combined with generalized advantage estimation (GAE) is adopted, and the policy network and value network are optimized independently. 5) Strategy Deployment and Real-time Decision-Making: The ideal policy network obtained after final training is integrated into the CPU scheduler framework as a scheduler. It receives task status information fed back by the operating system in real time, executes the calculation output corresponding priority adjustment strategy, and realizes fast scheduling decision-making with the help of priority heap, thereby driving the system to complete efficient dynamic scheduling.

2. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, The task status information collected in step 1) is represented as a fixed-dimensional two-dimensional matrix with a size of P×Q, where P represents the maximum number of tasks supported in the ready queue and Q represents the number of key scheduling features for each task. After normalization, the two-dimensional matrix is ​​reshaped into a three-dimensional tensor with shape (batch_size, P, Q) before being input into the policy network and value network, in order to preserve the sequence structure and temporal features of the task.

3. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, The policy network and value network constructed in step 2) have the following structure: The input layer receives a three-dimensional tensor of shape (batch_size, P, Q); The sequence coding layer uses a gated recurrent unit (GRU) with a hidden state dimension of 64 to mine the temporal correlation and dynamic evolution features of the task sequence. The hidden state output by GRU is used as a feature representation of the overall task sequence and input into a multilayer perceptron (MLP). The MLP contains a linear transformation layer (64-dimensional input and output), layer normalization (LayerNorm), ReLU activation, and Dropout to improve nonlinear expressive power and model generalization performance. The policy network Actor outputs action logits of length 5, which are then transformed into an action probability distribution using softmax. The value network Critic outputs a scalar value representing the expected reward for the current state. The policy network and value network are trained independently to ensure the stability and convergence of the training process.

4. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, The composite reward function described in step 3) includes three sub-items: throughput incentive, response efficiency penalty, and task starvation prevention, as expressed below: The throughput incentive measures the system's processing capacity by the number of tasks completed per unit time. The use of a logarithmic function ensures the incentive effect for task completion while avoiding numerical instability in high-frequency scenarios. The response efficiency penalty term constrains the system response latency based on the arithmetic mean of task turnaround times; The task starvation prevention measure introduces task priority weights and dynamic thresholds, and applies secondary penalties to high-priority tasks that exceed the threshold waiting time to prevent critical tasks from being shelved for a long time. The three hyperparameters w1, w2, and w3 control the contribution intensity of each sub-item, and the overall reward function takes into account throughput, response efficiency, and task hunger, which can effectively guide the reinforcement learning agent to achieve robust and efficient policy optimization in a multi-objective scheduling environment.

5. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, The course-based training mechanism sets different task distributions for different stages: the initial stage is characterized by short, intensive tasks; the middle stage gradually introduces a higher proportion of long tasks; and the final stage introduces non-uniform task distribution and resource bottlenecks to improve the model's adaptability and generalization ability to diverse scheduling scenarios.

6. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, The method employs a proximal policy optimization algorithm (PPO-Clip) combined with generalized advantage estimation (GAE) for training. This method limits the policy update magnitude to prevent performance fluctuations and improves training stability and sample utilization efficiency. The policy network and value network adopt a dual-network architecture and are independently optimized to ensure the stability and convergence of the learning process.

7. The adaptive CPU priority scheduling method based on near-end strategy optimization according to claim 1, characterized in that, Based on the priority adjustment strategy output by the aforementioned strategy network, a max-heap priority queue structure is used to manage and select tasks to be scheduled: Each task is inserted into the corresponding max-heap according to the priority value output by the policy model; The scheduler accesses the top task of the max heap in each round of scheduling, and the time complexity of the scheduling selection operation is O(1). When a task's status changes or it completes, the heap structure is automatically updated and the priority order is adjusted to ensure that high-priority tasks are executed first.

8. An adaptive CPU priority scheduling system based on near-end strategy optimization, characterized in that, This includes the following modules built into the hardware platform: The status acquisition and temporal feature extraction module is used to collect the status information of the current queue of tasks to be scheduled, including but not limited to the following key scheduling features of each task: unique task identifier, arrival timestamp, total number of task instructions, number of remaining execution instructions, and priority assigned during the last scheduling. The collected status information is normalized by a feature encoder to convert the multi-dimensional status data into a unified status vector representation. The dual-network architecture module based on the Proximal Policy Optimization (PPO) algorithm includes a policy network and a value network, which are used to output action policy distribution prediction and state value evaluation based on state information, respectively. The policy network outputs the selection probability distribution of each task to be scheduled under different priorities, and the value network evaluates the long-term expected return corresponding to the current scheduling state. Both networks adopt a neural network structure. The composite reward function module uses a composite reward function that combines positive incentives for task throughput, penalties for average response time, and adjustment factors for starvation prevention indicators to guide the dual network architecture module to obtain scheduling strategy behavior that balances fairness and efficiency under different task structures and load conditions. The training module has a built-in course-based training mechanism and constructs a phased training task distribution to enhance the model's generalization ability in diverse scheduling scenarios. During the training process, a method based on the proximal policy optimization algorithm PPO-Clip combined with generalized advantage estimation (GAE) is used to independently optimize and train the policy network and value network. The strategy deployment and real-time decision-making module controls the coordinated work of the above modules to form a closed-loop adaptive scheduling system. It completes iterative optimization training to obtain an ideal policy network, which is integrated into the CPU scheduler framework as a scheduler. It receives task status information fed back by the operating system in real time, executes calculations to output the corresponding priority adjustment policy, and realizes fast scheduling decisions with the help of priority heap, thereby driving the system to complete efficient dynamic scheduling.

9. The adaptive CPU priority scheduling system based on near-end strategy optimization according to claim 8, characterized in that, The scheduling executor in the strategy deployment and real-time decision-making module: based on the task priority output by the strategy network, it adopts a maximum heap priority queue structure to achieve efficient task classification and scheduling, ensuring that high-priority tasks are executed first, supporting dynamic task status updates and heap structure adjustments, and meeting the needs of large-scale concurrent scheduling.

10. An adaptive CPU priority scheduling system based on near-end strategy optimization according to claim 8, characterized in that, It also includes an evaluation and feedback module, which is used to monitor and evaluate the performance indicators of the scheduling system in real time, and use the feedback results to guide policy updates and scheduling optimization; the performance indicators include, but are not limited to, task throughput, response time, and turnaround time.

Citation Information

Cited By

  • Automatic customer complaint work order circulation system and method based on composite intention disassembly

    CN121810297A

  • An intelligent customer service system and method based on a reinforcement learning model and composite intent disassembly

    CN121810297B