Multi-core processor task scheduling method under NUMA architecture based on reinforcement learning

Through the method based on reinforcement learning, the task scheduling optimization of multi-core processors under the NUMA architecture is solved, and the problem of insufficient adaptability of task scheduling methods in the existing technology is achieved, shorter task execution time and higher resource utilization, and improved system performance and throughput.

CN120335953APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510384008.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The task scheduling method of multi-core processors under the NUMA architecture is difficult to flexibly cope with dynamically changing workloads and complex system environments, with high computational complexity, and failing to fully consider inter-core communication, data transmission delay between NUMA nodes and hardware resource constraints, resulting in long task execution time and low resource utilization.

Method used

Using reinforcement learning-based method, the Markov decision-making process model is constructed through directed acyclic graph modeling tasks and processor resources, a deep Q network is used to optimize tasks scheduling, reasonable state space, action space and reward functions are designed, and the allocation of tasks between cores is dynamically adjusted to reduce execution delay and improve resource utilization.

Benefits of technology

It realizes better task scheduling adaptability in a dynamic environment, reduces task execution time, improves system throughput and resource utilization, and optimizes the system performance of multi-core processors under the NUMA architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335953A_ABST
    Figure CN120335953A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based multi-core processor task scheduling method under an NUMA architecture, and the method comprises the steps: carrying out the modeling of a task inputted into user equipment through a directed acyclic graph, and obtaining a task model; according to the NUMA architecture of the user equipment, constructing a processor resource model and a task scheduling model of the user equipment; constructing a task scheduling decision optimization problem into a Markov decision process model according to the task model, the processor resource model and the task scheduling model, modeling a state space, an action space and a reward function in the Markov decision process model, and performing reinforcement learning algorithm model training; and obtaining Q values of different actions in the current state by using the trained reinforcement learning algorithm model, and selecting the action with the highest Q value as a task mobilization decision result. According to the method, the task execution time can be reduced as much as possible to meet the strict real-time requirement, the throughput of the system is improved, and processor resources are fully utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of task scheduling, and particularly relates to a task scheduling method for a multi-core processor under a NUMA architecture based on reinforcement learning. Background Art

[0002] With the rapid development of data-intensive applications, multi-core central processing units (CPUs) based on the non-uniform memory access (NUMA) architecture are increasingly widely used in the field of high-performance computing. The NUMA architecture divides multiple CPU cores into different nodes, and each node is equipped with an independent memory module, thereby achieving higher parallelism and lower memory access conflicts in large-scale multi-core systems. However, due to the complexity of this system, especially the non-uniform memory access characteristics in the NUMA architecture, the task scheduling problem has become extremely challenging.

[0003] Under this background, it is of great significance to study the task scheduling scheme for CPU systems based on the NUMA architecture. There are significant access latency differences between different nodes in the NUMA architecture. How to reasonably allocate tasks according to the computing requirements and data access patterns of tasks to reduce memory access overhead and improve computing efficiency is an urgent problem to be solved. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a task scheduling method for a multi-core processor under a NUMA architecture based on reinforcement learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] In a first aspect, the present invention provides a task scheduling method for a multi-core processor under a NUMA architecture based on reinforcement learning, including:

[0006] Modeling the tasks input to the user device using a directed acyclic graph to obtain a task model;

[0007] Constructing a processor resource model and a task scheduling model for the user device according to the NUMA architecture of the user device;

[0008] Constructing the task scheduling decision optimization problem as a Markov decision process model according to the task model, the processor resource model, and the task scheduling model, modeling the state space, action space, and reward function therein, and training the reinforcement learning algorithm model;

[0009] Using the trained reinforcement learning algorithm model, obtaining the Q values of different actions in the current state, and selecting the action with the highest Q value as the task scheduling decision result.

[0010] In a second aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-provided multi-core processor task scheduling method based on reinforcement learning under the NUMA architecture are implemented.

[0011] In a third aspect, the present invention further provides a computer device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; where:

[0012] The memory is used to store a computer program;

[0013] The processor is used to execute the steps of the above-provided multi-core processor task scheduling method based on reinforcement learning under the NUMA architecture by running the program stored on the memory.

[0014] Advantages of the present invention:

[0015] A multi-core processor task scheduling method based on reinforcement learning provided by the present invention realizes effective management and allocation of tasks on a multi-core processor, improving system performance. By systematically modeling the data transfer time and hardware resources between cores, between nodes, and between boards, and using a clever task scheduling algorithm, the system can more effectively organize and arrange the execution order of various tasks, thereby minimizing the task execution time as much as possible to meet strict real-time requirements, improving the system throughput, and making full use of the processor resources.

[0016] The following will further elaborate on the present invention in detail with reference to the drawings and embodiments. Description of the Drawings

[0017] Figure 1 is a schematic diagram of a multi-core processor task scheduling method based on reinforcement learning under the NUMA architecture provided by an embodiment of the present invention;

[0018] Figure 2 is a flowchart of a multi-core processor task scheduling method based on reinforcement learning under the NUMA architecture provided by an embodiment of the present invention;

[0019] Figure 3 is a schematic diagram of a task model provided by an embodiment of the present invention;

[0020] Figure 4 is a schematic diagram of a multi-core processor system including NUMA nodes provided by an embodiment of the present invention;

[0021] Figure 5 is a schematic diagram of an inter-core / inter-node test block diagram provided by an embodiment of the present invention;

[0022] Figure 6 It is a schematic diagram of the inter-board test block diagram provided by the embodiment of the present invention;

[0023] Figure 7 It is a schematic diagram of the Euclidean distance model provided by the embodiment of the present invention;

[0024] Figure 8 It is a schematic diagram of the reinforcement learning algorithm provided by the embodiment of the present invention. Detailed implementation manners

[0025] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0026] In the prior art, the task scheduling problem is NP-hard, especially when considering multiple constraints. To solve the task scheduling problem, researchers have proposed some methods. The existing task scheduling methods based on multi-core systems mainly include: list-based scheduling algorithms and meta-heuristic scheduling algorithms. Tian Q et al. proposed Heterogeneous Earliest Finish Time (HEFT) and Critical Path on Processor (CPOP). M.S. Arif, Z et al. proposed a scheduling algorithm based on Parent Priority Earliest Finish Time (PPEFT) for heterogeneous environments. Tasks are arranged in a parent priority queue (PPQ) according to the downward rank and parent priority. Simulation results show that this algorithm is superior to HEFT and CPOP in terms of cost and schedule duration. These methods determine the optimal order of task execution and prioritize tasks on the fastest processor according to the characteristics of related tasks. Tian Q et al. proposed that in a heterogeneous multi-core processor system, with the goal of minimizing task execution latency, task clustering is first performed, and then a priority calculation method combining task communication overhead and task execution time is proposed. Finally, each subtask is scheduled by combining the task replication method. The proposed hybrid task scheduling algorithm based on task clustering has better performance than the traditional HEFT algorithm and the critical path algorithm on the processor. Although the list-based heuristic algorithm is easy to implement, it is difficult to guarantee the scheduling performance. Meta-heuristic algorithms, such as Genetic Algorithm (GA) and Particle Swarm Optimization (PSO), are evaluation-based algorithms. They use random search techniques but require a large number of iterations during the evolution process. Compared with heuristic algorithms, meta-heuristic algorithms usually have better scheduling performance but also have higher computational complexity.

[0027] In summary, there are still many deficiencies in the current task scheduling method. First, traditional scheduling strategies usually rely on fixed priority lists or heuristic algorithms calculated in advance through professional knowledge, and it is difficult to flexibly cope with dynamically changing workloads and complex system environments. Second, existing meta-heuristic methods, such as genetic algorithms and particle swarm optimization, have large computational overheads, long computational processes, and may not be able to effectively search for the globally optimal scheduling scheme under complex constraint conditions. Finally, existing scheduling methods still have deficiencies in modeling complex parameters such as inter-core communication, data transfer latency between NUMA nodes, and NUMA architecture characteristics, and the consideration of hardware resource constraints is also not comprehensive enough. These factors have an important impact on the execution time of tasks and the utilization rate of system resources. If not fully modeled and optimized, it may lead to low task scheduling efficiency, uneven resource allocation, and even system performance bottlenecks.

[0028] In view of this, a multi-core processor task scheduling method based on reinforcement learning provided by the present invention has the following advantages. First, for the task scheduling of multi-core processors based on reinforcement learning, it has better dynamic adaptability. Traditional list-based scheduling algorithms and meta-heuristic algorithms usually rely on static scheduling rules and are difficult to cope with dynamically changing tasks and system states. The DQN algorithm can dynamically adjust the task scheduling strategy according to the real-time feedback of the environment through reinforcement learning. The agent continuously adjusts the strategy through interaction with the environment, so as to be able to cope with the changing workload and resource conditions and is more adaptable. Second, most existing scheduling methods are limited in dealing with the dependencies between tasks and resource constraints. The DQN algorithm can design the state space, establish a comprehensive state space by modeling the CPUs under the NUMA architecture, comprehensively consider the execution time of tasks, dependencies, resource usage, and the real-time state of the system, and make more flexible decisions in task scheduling. Third, the DQN algorithm can provide a more efficient, flexible, and adaptable task scheduling scheme through reinforcement learning, showing stronger optimization ability under the NUMA architecture in a multi-core processor system. Especially under dynamic task scheduling and complex resource constraint conditions, it can significantly improve system performance and reduce task execution latency.

[0029] The present invention designs a task scheduling method for a CPU multi-core system under the NUMA architecture based on reinforcement learning. Please refer to Figure 1 , Figure 1It is a schematic diagram of a multi-core processor task scheduling method based on reinforcement learning provided by an embodiment of the present invention, including a task model, a processor resource model, and a task scheduling model. Before performing task scheduling, it is first necessary to model the data transfer time within and between cores and the hardware architecture, and use relevant methods to measure data for subsequent modules. After training the agent, it is possible to make predictions based on the input task flow graph to obtain a deployment plan with a shorter overall task execution delay.

[0030] Figure 1 As shown, it can be divided into three main parts, namely module construction, framework setting, and training. Each part contains specific steps and content. Module construction constructs the basic modules required for multi-core processor task scheduling under the NUMA architecture, mainly including task-related modeling modules, which define the basic attributes of tasks and the overall framework of the task scheduling problem, and also include the modeling of hardware computing resources and communication delays under the NUMA architecture; the particularity of the NUMA architecture lies in the difference in memory access time between different nodes, and the modeling of communication delays is very important.

[0031] After the task and resource models are constructed, the framework setting needs to define the entire system scheduling framework. According to the defined model and framework, set the initial state and prepare to enter the training stage.

[0032] The training process mainly uses the Deep Q-Network (DQN) to optimize task scheduling. Finally, DQN will learn how to optimally allocate tasks in a multi-core processor system with a NUMA architecture, reduce task execution time, and improve the resource utilization rate of the system.

[0033] Please refer to Figure 2 , Figure 2 It is a flowchart of a multi-core processor task scheduling method based on reinforcement learning provided by an embodiment of the present invention. A multi-core processor task scheduling method based on reinforcement learning provided by the present invention includes:

[0034] S101. Model the tasks input to the user device using a directed acyclic graph to obtain a task model.

[0035] Specifically, please refer to Figure 3 , Figure 3 It is a schematic diagram of a task model provided by an embodiment of the present invention. In this embodiment, the tasks input to the user device are modeled using a directed acyclic graph to obtain a task model, including:

[0036] Model the tasks on the user device as a directed acyclic graph, denoted as G=(ν,ε), where ν represents the set of tasks, denoted as v i Through {b i,c i ,a i ,d i Description, b i Represents task v i Size (KB), w i Indicates completion of task v i The number of processor cycles required, a i Represents task v i The task scheduling decision d i represents bandwidth usage; ε represents the set of directed edges between tasks, expressed as ε={e(v i ,v j )|v i ,v j ∈ν}, directed edge e(v i ,v j ) represents task v i With task v j The dependency relationship between tasks v j In task v i After the execution is completed and the execution result is received, the execution begins. i ,v j ) weight g ij Represents task v i With task v j The communication volume between tasks (in bytes), C represents the task v i With task v j The communication volume set between tasks v i The predecessor tasks constitute the set pred(v i ), task v i The successor tasks constitute the set succ(v i ), where task v i With task v j The communication volume set C between them is expressed as:

[0037]

[0038] Among them, c ij represents the communication volume between task i and task j.

[0039] S102: Construct a processor resource model and a task scheduling model of the user device according to the NUMA architecture of the user device.

[0040] Specifically, in this embodiment, there are multiple processors on the user equipment FT2000+ processing system, each processor may include one or more cores, and 8 cores constitute a NUMA node.

[0041] It should be noted that to obtain the overall execution delay of a task, it is necessary to accurately know the data transmission time and system overhead time of the task. Therefore, it is necessary to further study the impact of the number of task modules within a single core and the communication volume between tasks on the system performance.

[0042] Specifically, a hardware computing resource model and a communication time model under the NUMA architecture are respectively constructed. Among them,

[0043] Constructing the hardware computing resource model includes respectively constructing a user device model and a NUMA node model;

[0044] The user device model is described as: The user device includes α groups of processor boards, denoted as S α ={S1, S2, …, S α}, the user device includes M cores, denoted as P = {P1, P2, …, P M}, the user device includes H network cards, denoted as Net = {Net1, Net2, …, Net H}, each group of processor boards includes at least one core, and multiple cores form a node; there is also a data transmission unit on the user device, which is responsible for transmitting tasks and processing results.

[0045] In this device configuration, each processor can contain one or more cores, and several cores form a node, which depends on the specific architecture and configuration of the processor; the data transmission unit is responsible for efficiently managing and transmitting task data and processing results to ensure smooth interaction and data transmission between processors or between processors and external systems.

[0046] Please refer to Figure 4 , Figure 4 is a schematic diagram of a multi-core processor system including NUMA nodes provided by an embodiment of the present invention. In the figure, there is a processor with 2 NUMA nodes and 6 cores. Each core has local memory and remote memory, and the latency caused by the core accessing different memories is also different.

[0047] The NUMA node model is described as: Each group of processors of the user device includes N nodes, denoted as N = {N1, N2,..., N N}, each node includes X cores, denoted as P i ={P1, P2,..., P X}, each node includes N local memories, denoted as M = {M1, M2, …, M N};

[0048] When task v i is executed on node N1, the successor task v i of task v jExecute on node N2. Nodes N1 and N2 are remote nodes, then for task v i The bandwidth used to transmit the execution result of to node N2 is expressed as:

[0049]

[0050] where g ij represents the traffic volume between the i-th task v i and the j-th task v j and t ij represents the communication time between the i-th task v i and the j-th task v j ;

[0051] When task v i executes on node N1, the successor task v i of task v j executes on node N1, and task v i and the successor task v j execute on different cores of node N1, then the bandwidth for accessing local memory is expressed as:

[0052]

[0053] where g ij represents the traffic volume between the i-th task v i and the j-th task v j and t ij represents the communication time between the i-th task v i and the j-th task v j ;

[0054] Construct a communication time model under the NUMA architecture, including:

[0055] First, please refer to Figure 5 , Figure 5 which is a schematic diagram of an inter-core / inter-node test block diagram provided by an embodiment of the present invention. Modules 1 and 2 represent tasks. To study the influence of the transmission volume between cores / NUMA nodes on the transmission time, the data transmission time will increase as the traffic volume between tasks increases. Therefore, to study the influence of the traffic volume on the data transmission time, first, the influence of the number of modules on the system overhead time and the data transmission time needs to be removed. Set the running time of a single task to 10 μs, and test the data transmission time when the traffic volume is 500, 1K, 10K, 50K, 100K, 200K Byte respectively.

[0056] When multiple tasks are assigned to the same core, different cores, or different nodes, the data transmission time Tpa q is expressed as:

[0057] Tpa q = EST(v2) - EFT(v1) - Tra2;

[0058] Among them, EST(v2) represents the start time of task v2, EFT(v1) represents the end time of task v1, and Tra2 represents the system overhead time when there are multiple subtasks to be processed in the nucleus;

[0059] Secondly, please refer to Figure 6 , Figure 6 is a schematic diagram of the inter-board test block diagram provided by the embodiment of the present invention. Modules 1, 2, and 3 represent tasks. To study the impact of the communication volume between tasks on the inter-board / intra-chip transfer time, it is necessary to consider that it is difficult to unify the timestamps between different cores during the test. Therefore, three serial task modules are used for testing. The running time of task module 3 is set to 0 us, and the running times of the other two tasks are set to 10 us. The data transfer times are respectively tested when the communication volumes are 1, 500, 1K, 10K, 50K, 100K, 1M, and 10 MBytes. During the test, the deployment scheme is set such that task module 1 and task module 3 are deployed on CPU1, and task module 2 is deployed on CPU2, as Figure 6 shown.

[0060] When multiple tasks are assigned to the same processor or different processors, the data transfer time is expressed as:

[0061]

[0062] Among them, EST(v3) represents the start time of task v3, EFT(v1) represents the end time of task v1. T v2 represents the calculation time of task v2.

[0063] In this embodiment, a task scheduling model is constructed. The task scheduling model consists of an application program, a target computing environment, and scheduling performance criteria. An application program is represented by a directed acyclic graph G = (ν, ε). The target computing environment consists of Q multi-core processors, and these processors are connected in a fully i-connected topological structure. In this topological structure, it is assumed that all communications between processors are executed without contention. In addition, it is assumed that the task execution of the given application program is non-preemptive.

[0064] In this embodiment, it includes: respectively constructing a time prediction model and a resource constraint model; among them,

[0065] The time prediction model is described as:

[0066] Let W be a computing cost matrix of size ν×M, and the parameter w in the computing cost matrix ij represents on processor pj Complete the task v i The estimated execution time, i.e., the computational cost; before scheduling, the task is marked with the average execution cost.

[0067] The average execution cost of the task Is expressed as:

[0068]

[0069] Where, w ij Represents the computational cost, M represents the number of cores;

[0070] The data transfer rate between processors is stored in a matrix B of size M×M, the communication startup cost of the processor is a vector L of dimension M, and the communication cost c of edge (i,k) ik For transferring data from task v i To task v k Is expressed as:

[0071]

[0072] Where, B mn Represents the average transfer rate from core n to core m, L m Represents the communication startup time, C ik Represents the communication volume from task i to task k; when task v i And task v k Are scheduled on the same processor, the communication cost c ik Is 0 because it is assumed that the communication cost within the processor is negligible compared to the communication cost between processors.

[0073] Before representing the objective function, it is necessary to define the EST and EFT attributes. EST(v i ,p j ) and EFT(v i ,p j ) are the earliest execution start time and the earliest execution completion time of task v i On processor p j .

[0074] For task v of the input user device entry , EST(n entry ,p j ) = 0,

[0075] For other tasks in the graph, the EFT and EST values are calculated recursively starting from the entry task. To calculate the EFT of task v i , all direct predecessor tasks of v i Must have been scheduled.

[0076] Task v i The earliest execution start time on processor p j is expressed as:

[0077]

[0078] Task v i The earliest execution completion time on processor p j is expressed as:

[0079] EFT(v i ,p j ) = w ij + EST(v i ,p j );

[0080] Where pred(v i ) represents the set of predecessor tasks of task v i , avail(j) represents the earliest time when processor p j is ready to execute a task. When task v m is the last task assigned to processor p j , then avail(j) is the time when processor p j completes the execution of task v m and is ready to execute another task (when having a non-insertion-based scheduling policy). The internal max block in the EST equation returns the ready time, that is, the time when all the data required by n v arrives at processor p j .

[0081] After scheduling task v j on processor p m , the earliest start time and earliest completion time of task v m on processor p j are respectively equal to the actual start time AST(v m ) and actual completion time AFT(v m ). The total completion time of all tasks is the actual completion time of the exit task v m . If there are multiple exit tasks and the convention of inserting pseudo-exit tasks is not applied, the total delay time is expressed as: exit makespan = max{AFT(v

[0082] )}; exit

[0083] The task scheduling decision optimization problem is described as minimizing the total delay time. It can be understood that the objective function of the task scheduling problem is to determine the task assignment of a given application to processors such that its total delay time is minimized.​

[0084] To better allocate tasks to processors, each task has a specific set of resource requirements, and in a NUMA architecture, each NUMA node has a set of available resources. Given these resource requirements and resource availability, how to create a schedule for all tasks so as to meet each resource requirement if possible. Therefore, the problem becomes finding a mapping from tasks to processors such that each resource requirement is satisfied while no processor exceeds its resource availability.

[0085] The resource constraint model is described as follows:

[0086] Each node includes a set of available resources. There are three types of resource constraints, namely processor utilization rate, available memory percentage, and bandwidth usage. The processor utilization rate and bandwidth usage are soft constraints, and the available memory percentage is a hard constraint. The processor utilization rate and bandwidth usage can be overloaded and are considered soft constraints, while memory is considered a hard constraint and cannot exceed the total available memory on the system, marked as a hard constraint. Resources marked as hard constraints must be satisfied, while resources marked as soft constraints may not be fully satisfied.

[0087] The resource requirements of a task can be represented as a set or a three-dimensional vector. The resource requirements A of a task vi are represented as:

[0088]

[0089] The resource availability of a NUMA node can be represented as a set or a three-dimensional vector. The resource availability of a node is represented as:

[0090]

[0091] Obtain the minimum Euclidean distance between the resource requirements of a task and the resource availability of a node , and ensure that all types of resources in the resource requirements of the task are less than the resources in the resource availability of the node ;

[0092] Select n as the node for scheduling task v, and update the remaining available resources on the node until all tasks are allocated to nodes.

[0093] S103. According to the task model, processor resource model, and task scheduling model, construct the task scheduling decision optimization problem as a Markov decision process model, model the state space, action space, and reward function therein, and perform training on the reinforcement learning algorithm model.

[0094] Specifically, in this embodiment, modeling the state space includes:

[0095] According to the definition of the Markov property, the task allocation decision considers the state of the user equipment at the current time t, and the state space S of the Markov decision process model t is expressed as:

[0096] S t = [EST1, …, EST M ,T i,1 , …, T i,M ;

[0097] where the user equipment includes M cores, and EST M represents the earliest start time when task v i is deployed on core m, and T i,M represents the execution time of task v i on core m.

[0098] In this embodiment, modeling the action space includes:

[0099] Each task needs to be assigned to a core of the processor of the user equipment for execution, and the action space Ω is expressed as:

[0100] Ω = {(S1, N1, P1, Net1), (S2, N2, P2, Net2),...(S i ,N n ,P m ,Net j )};

[0101] where (S i ,N n ,P m ,Net j ) represents the jth network card of the mth core of the nth node of the ith processor in the user equipment, and the network card is only selected for tasks across devices. In each decision step t, there is one and only one task that can be determined, and the action At at the current time t is expressed as:

[0102] At = (S i ,N n ,P m ,Net j ).

[0103] In this embodiment, modeling the reward function includes:

[0104] According to the task scheduling model, obtain the Euclidean distance L between the resource requirements of the task and the resource availability of the node τθ 2, the shorter the Euclidean distance between two points, the greater the reward, which is expressed as:

[0105] L vN 2 =(m N -m) 2 +(c N -c) 2 +(b N -b) 2 ;

[0106] Among them, A τi represents the resource requirements of the task, and A θi represents the resource availability of the node. m, c, and b respectively represent the available memory percentage, processor utilization rate, and bandwidth usage, which are mapped to the x, y, and z coordinate axes in three-dimensional coordinates. Each time a task is scheduled, the resources in the node need to be updated once. The resource update in the node is expressed as:

[0107] A θ =A θ -A τ , which means subtracting the resources occupied by the task from the node resources to obtain the current available resources of the node;

[0108] At the t-th time step, the agent makes a decision on the execution location of the task, and this task is the q-th task in the task sequence TS, that is, the priority of this task is ranked at the q-th position; if the agent selects an action (S i , N n , P m , Net j ) from the action space Ω and maps the task to the processor (S i , N n , P m , Net j ), then the reward number is expressed as:

[0109] R t =γR1 t -(1-γ)R2 t ;

[0110]

[0111] R2 t =makespan;

[0112] Among them, R1 t represents using the Euclidean distance between the resources required by the task and the available resources of the node as the reward criterion. Under the requirement of meeting the resource constraints for task operation, the shorter the distance, the higher the reward. R2 t represents the total delay time of the task, and the shorter the time, the higher the reward. R t represents the total reward of DQN scheduling, Lτθ The Euclidean distance between the task resource requirements and the node resources is denoted as, and the discount factor is denoted as γ.

[0113] In this embodiment, training the reinforcement learning algorithm model includes:

[0114] Determining the resource requirements of the current task and the resource availability of the node according to the environmental state;

[0115] Given an action, calculating the Q-value of the given action in the current state, and taking the maximum Q-value as the task scheduling decision result;

[0116] Calculating the reward according to the task scheduling decision result of the given action;

[0117] After executing the given action, updating the state space;

[0118] Putting the updated state space and the reward into the experience replay unit, and performing the next task scheduling.

[0119] In this embodiment, please refer to Figure 8 , Figure 8 is a schematic diagram of the reinforcement learning algorithm provided by the embodiment of the present invention. The figure shows the scheduling algorithm framework of DAG tasks in a multi-core processor system with a NUMA architecture. The scheduling algorithm first obtains the state information of the current NUMA architecture multi-core processor system from the environmental state in the state module, including the load conditions of each CPU core, the communication delay between NUMA nodes, and the deployment status, etc.

[0120] The deep Q-network combines the information from the environmental state to make a decision. Calculate the Q-values of different actions according to the current state, and select the action with the highest Q-value as the decision result. The selected action is executed in the environmental module, and the environmental module gives a reward according to the execution result of the action. For example, if the action can effectively reduce the task execution time or improve the overall system performance, the reward may be a positive value; otherwise, if the action causes a performance decline, the reward may be a negative value. After executing the action, the state space is updated, and the new state is fed back to the deep Q-network for the next round of decision-making.

[0121] S104. Using the trained reinforcement learning algorithm model, obtaining the Q-values of different actions in the current state, and selecting the action with the highest Q-value as the task scheduling decision result. Optionally, the Q-value is calculated through the model and the policy, and represents the expected cumulative reward after taking a certain action in the current state.

[0122] In summary, a multi-core processor task scheduling method based on reinforcement learning provided by the present invention has the following beneficial effects:

[0123] First, the present invention uses the DQN algorithm to optimize the task scheduling of multi-core processors under the NUMA architecture. Through reinforcement learning, DQN can learn the optimal task scheduling strategy in a dynamic environment. By adjusting the allocation of tasks among different cores, it minimizes the execution delay and improves the utilization rate of system resources. The key lies in combining the characteristics of the NUMA architecture. By designing a reasonable state space, action space, and reward function, the agent can adapt to various resource constraints in complex task scheduling problems and optimize the overall performance of task execution.

[0124] Second, the multi-core processor system analysis and training model based on reinforcement learning designed by the present invention comprehensively considers parameters such as the data transfer time between cores, between NUMA nodes, and between different devices. The scheduling task is transformed into a DAG (Directed Acyclic Graph) model, and a corresponding state space is designed according to the node characteristics of the task. This state space serves as the input to the neural network. A reward function is designed to accelerate the convergence speed during training, thereby saving training time. In addition, this reward function ensures that each computing node is fully utilized, optimizes the response speed of the entire system, and significantly improves the throughput of the system.

[0125] Based on the same inventive concept, the present invention also discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the above-mentioned task scheduling method for multi-core processors under the NUMA architecture based on reinforcement learning.

[0126] The present invention also discloses a computer device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. Among them:

[0127] The memory is used to store the computer program;

[0128] The processor is used to execute the steps of the above-mentioned task scheduling method for multi-core processors under the NUMA architecture based on reinforcement learning by running the program stored on the memory.

[0129] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising said element. Similar words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0130] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0131] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A task scheduling method for multi-core processors under the NUMA architecture based on reinforcement learning, characterized in that, Including: Modeling the tasks of the input user device using a directed acyclic graph to obtain a task model; Constructing a processor resource model and a task scheduling model for the user device according to the NUMA architecture of the user device; According to the task model, the processor resource model, and the task scheduling model, constructing the task scheduling decision optimization problem as a Markov decision process model, modeling the state space, action space, and reward function therein, and training a reinforcement learning algorithm model; Using the trained reinforcement learning algorithm model to obtain the Q-values of different actions in the current state, and selecting the action with the highest Q-value as the task scheduling decision result.

2. The method for scheduling tasks of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein The modeling the tasks of the input user device using a directed acyclic graph to obtain a task model includes: Model the tasks on the user equipment as a directed acyclic graph, denoted as G = (ν, ε), where ν represents the set of tasks, denoted as ν = {v1, v2, …, v N1}, v i is described by {b i , w i , a i , d i}, where b i represents the size of task v i , w i represents the number of processor cycles required to complete task v i , a i represents the task scheduling decision of task v i , and d i represents the bandwidth usage; ε represents the set of directed edges between tasks, denoted as ε = {e(v i , v j ) | v i , v j ∈ ν}, and the directed edge e(v i , v j ) represents the dependency relationship between task v i and task v j . Task v j starts to execute after task v i ends and the execution result is received. The weight g ij of the directed edge e(v i , v j ) represents the communication volume between task v i and task v j . C represents the set of communication volumes between task v i and task v i . The set of predecessor tasks of task v i forms the set pred(v i ), and the set of successor tasks of task v i forms the set succ(v j ); among them, the set of communication volumes C between task v ij and task v α is expressed as: Among them, c ij represents the traffic volume between task i and task j.

3. The method for scheduling tasks of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein Constructing a processor resource model for the user device, including: respectively constructing a hardware computing resource model and a communication time model under the NUMA architecture; wherein, Constructing the hardware computing resource model includes respectively constructing a user device model and a NUMA node model; The user equipment model is described as follows: The user equipment includes α groups of processor boards, denoted as S α ={S1, S2, …, S α}, the user equipment includes M cores, denoted as P = {P1, P2, ..., P M}, the user equipment includes H network cards, denoted as Net = {Net1, Net2, ..., Net H}, each group of the processor boards includes at least one core, and multiple cores form a node; The NUMA node model is described as follows: Each set of the processor boards of the user equipment includes N nodes, denoted as N = {N1, N2,..., N N}, and each of the nodes includes X cores, denoted as P i = {P1, P2,..., P X}, and each of the nodes includes N local memories, denoted as M = {M1, M2,..., M N}; When task v i is executed on node N1, the successor task v i of task v j is executed on node N2. If node N1 and node N2 are remote nodes, then the bandwidth used to transmit the execution result of task v i to node N2 is expressed as: Among them, g ij represents the traffic volume between the i-th task v i and the j-th task v j and t ij represents the communication time between the i-th task v i and the j-th task v j ; When task v i is executed on node N1, the successor task v i of task v j is executed on node N1, and task v i and the successor task v j are executed on different cores of node N1, then the bandwidth for accessing local memory is expressed as: Among them, g ij represents the traffic volume between the i-th task v i and the j-th task v j ; t ij represents the communication time between the i-th task v i and the j-th task v j ; Constructing the communication time model under the NUMA architecture includes: When multiple tasks are assigned to the same core, different cores, or different nodes, the data transfer time Tpa q is expressed as: Tpa q = EST(v2) - EFT(v1) - Tra2; Wherein, EST(v2) represents the start time of task v2, EFT(v1) represents the end time of task v1, and Tra2 represents the system overhead time when there are multiple sub-tasks to be processed within the core; When multiple tasks are assigned to the same processor or different processors, the data transfer time is expressed as: Among them, EST(v3) represents the start time of task v3, and EFT(v1) represents the end time of task v1; represents the computing time of task v2.

4. The method for scheduling tasks of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein Constructing a task scheduling model includes: respectively constructing a time prediction model and a resource constraint model; wherein, The time prediction model is described as: Let W be a computational cost matrix of size ν×M, where the parameter w in the computational cost matrix ij represents the estimated execution time, i.e., the computational cost, of completing task v j on processor p i ​ Average execution cost of the task Expressed as: where, w ij represents the computational cost, and M represents the number of kernels; The data transfer rate between the processors is stored in a matrix B of size M×M, and the communication startup cost of the processors is a vector L of M dimensions. The communication cost c i for transferring data from task v k to task v ik is expressed as: Among them, B mn represents the average transmission rate from nucleus n to nucleus m, and L m represents the communication startup time, and C ik represents the traffic from task i to task k; when task v i and task v k are scheduled on the same processor board, the communication cost c ik is 0; Task v for the input user device entry , EST(n entry , p j ) = 0, Task v i The earliest execution start time on processor p j is represented as: Task v i The earliest execution completion time on processor p j is expressed as: EFT(v i ,p j ) = w ij + EST(v i ,p j ); Among them, pred(v i ) represents the set of predecessor tasks of task v i , avail(j) represents the earliest time when the processor p j is ready to execute a task. When the task v m is the last task assigned on the processor p j , then avail(j) is the time when the processor p j completes the execution of task v m and is ready to execute another task; Schedule task v on processor p j After that, the earliest start time and earliest completion time of task v m on processor p m are equal to the actual start time AST(v j ) and the actual completion time AFT(v m ) of task v respectively. The total completion time of all tasks is the actual completion time of the exiting task v m . If there are multiple exiting tasks and the convention of inserting pseudo-exiting tasks is not applied, the total delay time is expressed as: m exit ​​ makespan = max{AFT(v exit )}; The task scheduling decision optimization problem is described as the shortest total delay time; The resource constraint model is described as: Each node includes a set of available resources, and there are three resource constraints, namely the available memory percentage, the processor utilization rate, and the bandwidth usage. The processor utilization rate and the bandwidth usage are soft constraints, and the available memory percentage is a hard constraint; Resource requirements of the task Expressed as: Resource availability of nodes It is expressed as: Obtain the task resource requirements and the resource availability of the node The minimum Euclidean distance between them, and ensure that the task resource requirements All types of resources in are less than the resource availability of the node Resources in; Select n as the node to schedule task v and update the resource availability of the node The remaining available resources until all tasks are scheduled on the nodes.

5. The method for scheduling tasks of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein Modeling the state space includes: According to the definition of the Markov property, the task allocation decision considers the state of the user equipment at the current moment t, and the state space S of the Markov decision process model t is expressed as: S t = [EST1,…, EST M , T i,1 ,…, T i,M ; Among them, the user equipment includes M cores, EST M represents the earliest start time when task v i is deployed on core m, and T i,M represents task v i The execution time on core m.

6. The method for scheduling tasks of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein Modeling the action space includes: Each task needs to be assigned to a core of the processor of the user device for execution, and the action space Ω is expressed as: Ω = {(S1, N1, P1, Net1), (S2, N2, P2, Net2), … (S i , N n , P m , Net j )}; Among them, (S i , N n , P m , Net j ) represents the j-th network card of the m-th core of the n-th node of the i-th processor in the user equipment, and the action At at the current time t is expressed as: At=(S i ,N n ,P m ,Net j )。 7. The method for multi-core processor task scheduling under the NUMA architecture based on reinforcement learning according to claim 1, wherein Modeling the reward function includes: According to the task scheduling model, obtain the Euclidean distance L between the resource requirements of the task and the resource availability of the node vN 2 , which is expressed as: L vN 2 = (m N - m) 2 + (c N - c) 2 + (b N - b) 2 ; Among them, represents the resource requirements of the task, represents the resource availability of the node. m, c, and b respectively represent the available memory percentage, processor utilization rate, and bandwidth usage, which are mapped to the x, y, and z coordinate axes in three-dimensional coordinates; every time a task is scheduled, the resources within the node need to be updated once, and the resource update within the node is expressed as: A N = A N -A v ; At the t-th time step, the agent is at the execution location of the decision-making task, and this task is the q-th task in the task sequence TS, that is, the priority of this task is ranked at the q-th position; if the agent selects an action (S i ,N n ,P m ,Net j ) from the action space Ω and maps the task onto the processor (S i ,N n ,P m ,Net j ), the reward number is expressed as: R t = γR1 t - (1 - γ)R2 t ; R2 t = makespan; Among them, R1 t represents using the Euclidean distance between the resources required for the task and the available resources of the node as the reward criterion. Under the requirement of meeting the resource constraints for task operation, the shorter the distance, the higher the reward. R2 t represents the total delay time of the task. The shorter the time, the higher the reward. R t represents the total reward of DQN scheduling. L τθ represents the Euclidean distance between the task resource requirements and the node resources. γ represents the discount factor.

8. The method for task scheduling of a multi-core processor under a NUMA architecture based on reinforcement learning according to claim 1, wherein The training of the reinforcement learning algorithm model includes: Determining the resource requirements of the current task and the resource availability of the node according to the environmental state; Given an action, calculating the Q-value of the given action in the current state, and taking the maximum Q-value as the task scheduling decision result; Calculating the reward according to the task scheduling decision result of the given action; After executing the given action, updating the state space; Putting the updated state space and the reward into the experience replay unit, and performing the next task scheduling.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the multi-core processor task scheduling method based on reinforcement learning according to any one of claims 1 to 8 are implemented.

10. A computer device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; wherein: The memory is used to store a computer program; The processor is configured to execute the steps of the multi-core processor task scheduling method under the NUMA architecture based on reinforcement learning according to any one of claims 1 to 8 by running the program stored on the memory.

Citation Information

Cited By

  • Task scheduling method, system and equipment of data stream processor and storage medium

    CN120909744A