Task Scheduling Method, Device, Medium and Product Based on Reinforcement Learning
Through the task scheduling method based on reinforcement learning, the processor node, scheduling domain and core are accurately selected, which solves the refinement problem of task scheduling to the CPU core, improves task execution efficiency and system performance, and avoids the performance losses caused by cache contention.
Patent Information
- Application Number
- CN202510053188.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Under a specific processor hardware architecture, it is difficult for the prior art to refine task scheduling to specific CPU cores, resulting in a degradation of cache performance.
Using a task scheduling method based on reinforcement learning, through the action experience value table and preset scheduling constraints, the target processor node, scheduling domain and processor core are accurately selected, task scheduling is optimized, resources are avoided idle or overuse, and the advantages of shared cache are used to reduce data access delays.
It realizes refined scheduling on the core granularity of the CPU, improves task execution efficiency and system performance, avoids performance reduction caused by cache contention, and dynamically adjusts task allocation to adapt to system changes.
Smart Images

Figure CN119473565B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a task scheduling method, device, medium, and product based on reinforcement learning. Background Art
[0002] In an operating system of the Linux kernel, task scheduling and allocation on a specific processor hardware architecture (multiple CPU cores sharing an L2 cache) can improve the utilization rate of the cache. However, when multiple cores simultaneously access the L2 shared cache, it also brings the risk of cache overuse, which may further lead to a decline in cache performance.
[0003] In the prior art, in the case of a specific processor hardware architecture, when there are task processes, the L3 cache is scheduled by a scheduler to achieve load balancing, or the scheduler will preferentially schedule a process to a completely idle scheduling domain, so that the process can use the entire L2 shared cache in the scheduling domain where it is located. These two technologies can only refine the task scheduling within the original CPU node range to the granularity of the scheduling domain.
[0004] In summary, in the case of a specific processor hardware architecture, how to refine task scheduling to specific CPU cores is a technical problem to be solved in this field. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a task scheduling method, device, medium, and product based on reinforcement learning, which can solve the task scheduling problem of processor cores sharing an L2 cache.
[0006] In a first aspect, the present invention discloses a task scheduling method based on reinforcement learning, including:
[0007] Determine a task scheduling action at the current time step according to the current action experience value table and a preset scheduling constraint condition; the task scheduling action includes a first action, a second action, and a third action determined based on a target experience value corresponding to the current time step in the current action experience value table; the first action is used to select a target processor node for executing a target task, the second action is used to select a target scheduling domain in the target processor node, and the third action is used to select a processor core in the target scheduling domain; several processor cores capable of sharing a preset secondary cache are configured in each scheduling domain;
[0008] Update the current system environment information and the current task load information of the embedded operating system based on the current reward information after the task scheduling action is executed by the embedded operating system;
[0009] Update the experience values in the current action experience value table based on the current system environment information, current task load information, and current reward information, and jump to the step of determining the task scheduling action for the current time step according to the current action experience value table and the preset scheduling constraint conditions, until the target task is scheduled to the target processor core that meets the preset conditions for task execution.
[0010] Optionally, before updating the experience values in the current action experience value table based on the current system environment, current task load information, and current reward information, it further includes:
[0011] Obtain the current processor hardware architecture information and current task information of the embedded operating system to obtain the current system environment information;
[0012] Obtain the current hardware utilization status information and current task running status information of the embedded system to obtain the current task load information.
[0013] Optionally, before determining the task scheduling action for the current time step according to the current action experience value table and the preset scheduling constraint conditions, it further includes:
[0014] Set the first affinity rule for the processor nodes, the second affinity rule for the scheduling domains, and the third affinity rule for the processor cores in the processor hardware architecture information level by level based on the preset scheduling target; the preset scheduling target is the scheduling target determined based on having the optimal system performance during the scheduling process;
[0015] Construct the preset scheduling constraint conditions for the scheduling task process based on the first affinity rule, the second affinity rule, and the third affinity rule.
[0016] Optionally, determining the task scheduling action for the current time step according to the current action experience value table and the preset scheduling constraint conditions includes:
[0017] Determine the processor node selection action experience value corresponding to the current time step in the current action experience value table, and based on the processor node selection action experience value, select the target processor node that meets the first affinity rule from the local multi-processor nodes for executing the target task;
[0018] Determine the scheduling domain selection action experience value corresponding to the current time step in the current action experience value table, and based on the scheduling domain selection action experience value, select the target scheduling domain that meets the second affinity rule from each scheduling domain in the target processor node;
[0019] Determine the processor core selection action experience value corresponding to the current time step in the current action experience value table, and based on the processor core selection action experience value, select the processor core that meets the third affinity rule from each processor core in the scheduling domain.
[0020] Optionally, set the first affinity rule for the processor nodes, the second affinity rule for the scheduling domains, and the third affinity rule for the processor cores in the processor hardware architecture information level by level based on a preset scheduling target, including:
[0021] Set the utilization threshold of the processor nodes in the processor hardware architecture information based on a preset scheduling target, so as to construct the first affinity rule for the task scheduling process according to the magnitude relationship between the current processor node utilization and the processor node utilization threshold;
[0022] Set the second affinity rule for scheduling the execution task threads of the same task and the process tasks with data sharing to the same scheduling domain based on a preset scheduling target;
[0023] Set the third affinity rule for scheduling tasks to the processor core with the lowest utilization rate in the scheduling domain based on a preset scheduling target.
[0024] Optionally, obtain the current processor hardware architecture information and the current task information of the embedded system, including:
[0025] Obtain the associated information of the hardware information of the current system environment of the embedded operating system, so as to obtain the current processor hardware architecture information constructed based on each physical structure of the processor, the interconnection information between each physical structure and the preset secondary cache, where the physical structures include processor nodes, scheduling domains, and processor cores;
[0026] Obtain the task priority, task cycle information, execution task threads, and data dependency relationships between tasks in the current task of the embedded operating system, so as to obtain the corresponding current task information.
[0027] Optionally, before obtaining the current processor hardware architecture information and the current task information of the embedded operating system, it further includes:
[0028] Divide the processor cores on the embedded operating system according to whether they share the preset secondary cache to obtain scheduling domains;
[0029] Divide each scheduling domain according to whether it accesses the local memory of the same node to obtain the corresponding processor nodes;
[0030] Determine the physical structure of the processor according to the architectural relationship between the processor nodes, scheduling domains, and processor cores.
[0031] Optionally, obtain the current hardware utilization status information of the embedded system, including:
[0032] Obtain the utilization rate of the processor cores and the usage of the shared preset secondary cache during the current system operation in the embedded operating system, so as to obtain the corresponding current hardware utilization status information.
[0033] Optionally, determine the processor node selection action experience value corresponding to the current time step in the current action experience value table, and select a target processor node that meets the first affinity rule from the local multi-processor nodes for executing the target task based on the processor node selection action experience value, including:
[0034] Determine the maximum processor node selection action experience value in the current action experience value table, and select a processor node with a current processor node utilization rate less than the processor node utilization rate threshold and an existing task load from the local multi-processor nodes as the target processor node for executing the target task based on the maximum processor node selection action experience value.
[0035] Optionally, select a processor node with a current processor node utilization rate less than the processor node utilization rate threshold and an existing task load from the local multi-processor nodes as the target processor node for executing the target task, including:
[0036] Obtain the current processor node utilization rate of each processor node recorded in the task load information;
[0037] Select a processor node with a current processor node utilization rate less than the processor node utilization rate threshold as the initial processor node;
[0038] Determine whether there is a processor node with an existing task load among the initial processor nodes;
[0039] If there is only one processor node with an existing task load among the initial processor nodes, determine the processor node with the existing task load as the target processor node for executing the target task.
[0040] Optionally, after determining whether there is a processor node with an existing task load among the initial processor nodes, it further includes:
[0041] If there are multiple processor nodes with existing task loads among the initial processor nodes, determine the processor node with the existing task load that is physically closest as the target processor node for executing the target task.
[0042] Optionally, the task scheduling method based on reinforcement learning of the present application further includes:
[0043] Based on each state parameter value in the state parameter change range generated by each scheduling action, determine the task scheduling actions that can be executed at the current time step to construct a target action set; the state parameter change range is the state parameter change range corresponding to the current system environment information and the current task load information at the current time step;
[0044] Obtain the experience values corresponding to each action in the target action set from the current action experience value table to obtain an experience value set corresponding to the current time step;
[0045] Determine the target experience value corresponding to the current time step from the experience value set according to the highest value selection rule.
[0046] In a second aspect, the present invention discloses an electronic device, including:
[0047] A memory for storing a computer program;
[0048] A processor for executing the computer program to implement the steps of the aforementioned task scheduling method based on reinforcement learning.
[0049] In a third aspect, the present invention discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned task scheduling method based on reinforcement learning are implemented.
[0050] In a fourth aspect, the present invention discloses a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the aforementioned task scheduling method based on reinforcement learning are implemented.
[0051] It can be seen that the present invention discloses a task scheduling method based on reinforcement learning, including: determining the task scheduling action of the current time step according to the current action experience value table and the preset scheduling constraint conditions; the task scheduling action includes a first action, a second action, and a third action determined based on the target experience value corresponding to the current time step in the current action experience value table; the first action is used to select a target processor node for executing the target task, the second action is used to select a target scheduling domain in the target processor node, and the third action is used to select a processor core in the target scheduling domain; several processor cores capable of sharing a preset secondary cache are configured in each scheduling domain; updating the current system environment information and the current task load information of the embedded operating system based on the current reward information after executing the task scheduling action based on the embedded operating system; updating the experience values in the current action experience value table based on the current system environment information, the current task load information, and the current reward information, and jumping to the step of determining the task scheduling action of the current time step according to the current action experience value table and the preset scheduling constraint conditions until the target task is scheduled to a target processor core that meets the preset conditions for task execution.
[0052] As can be seen from the above technical solutions, by comprehensively considering system environment information and task load information, and through reinforcement learning, tasks can be accurately assigned to appropriate hardware resources, avoiding resource idleness or overuse. Considering the task scheduling optimization problem on the processor architecture with a preset secondary cache, the reduction of application performance caused by cache contention is avoided. According to the task association relationship and task load information of the target processor node, etc., the target task is scheduled to run on the processor core with the lowest utilization rate in the appropriate scheduling domain. The data sharing advantage brought by sharing the preset secondary cache can be fully utilized, reducing the latency of data access during task execution, thereby accelerating the task execution speed, improving the execution efficiency of a single task, and the overall performance of the system for multitasking processing. Optimize the core selection scheduling of ready tasks to achieve fine-grained scheduling at the CPU core level and maximize the guarantee of the running performance of application tasks. Finally, as the embedded operating system runs, tasks are continuously generated and ended, and the system environment and task load conditions are in dynamic change. By updating the system environment information and task load information, these changes can be tracked in real time, enabling subsequent task scheduling decisions to always be adjusted based on the latest and accurate system state. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a flowchart of a task scheduling method based on reinforcement learning provided by an embodiment of the present invention;
[0055] Figure 2 It is a module diagram of a task scheduling system based on reinforcement learning provided by an embodiment of the present invention;
[0056] Figure 3 It is a flowchart of a specific task scheduling method based on reinforcement learning provided by an embodiment of the present invention;
[0057] Figure 4 It is a structure diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0059] The terms "including" and "having" in the specification of the present invention and the accompanying drawings above, as well as any variations related to "including" and "having", are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units not listed.
[0060] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0061] In an operating system of the Linux kernel, task scheduling and allocation on a specific processor hardware architecture (multiple CPU cores sharing an L2 cache) can improve the utilization rate of the cache. However, when multiple cores access the L2 shared cache simultaneously, it also brings the risk of cache overrun, which may lead to a decline in cache performance.
[0062] In the prior art, in the case of a specific processor hardware architecture, when there is a task process, the L3 cache is scheduled by the scheduler to achieve load balancing, or the scheduler will preferentially schedule the process to a completely idle scheduling domain, so that the process can use the entire L2 shared cache in the scheduling domain where it is located. These two technologies can only refine the task scheduling within the original CPU node range to the granularity of the scheduling domain.
[0063] To this end, the present invention provides a task scheduling scheme based on reinforcement learning, which can refine task scheduling to specific CPU cores under a specific processor hardware architecture.
[0064] Refer to Figure 1 As shown, the present invention discloses a task scheduling method based on reinforcement learning, including:
[0065] Step S11: Determine the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions; the task scheduling action includes a first action, a second action, and a third action determined based on the target experience value corresponding to the current time step in the current action experience value table; the first action is used to select the target processor node for executing the target task, the second action is used to select the target scheduling domain in the target processor node, and the third action is used to select the processor core in the target scheduling domain; several processor cores capable of sharing a preset secondary cache are configured in each scheduling domain.
[0066] In this embodiment, before determining the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions, it further includes: hierarchically setting the first affinity rule of the processor node, the second affinity rule of the scheduling domain, and the third affinity rule of the processor core in the processor hardware architecture information based on the preset scheduling target; the preset scheduling target is a scheduling target determined based on having the optimal system performance during the scheduling process; constructing the preset scheduling constraint conditions for the scheduling task process based on the first affinity rule, the second affinity rule, and the third affinity rule. It can be understood that before determining the task scheduling action at the current time step, the affinity rules are hierarchically set in advance according to the scheduling target to obtain the preset scheduling constraint conditions constructed by the affinity rules at different levels. After determining the L2 cache architecture, the affinity for task scheduling at three levels is constructed, namely the affinity on the central processor node, the affinity at the scheduling domain granularity, and the affinity on each central processor core within the scheduling domain. When the goal is to optimize system performance, the affinity rule for task scheduling on the central processor core is: 1) Ensure that tasks with shared data are on the scheduling domain that shares the preset secondary cache to avoid performance loss caused by cache thrashing; 2) When there are multiple central processor nodes or central processor cores to choose from, select the node / core with the lowest utilization rate to avoid performance loss caused by central processor preemption. The affinity rules are set as follows:
[0067] Set the utilization rate threshold of the processor node in the processor hardware architecture information based on the preset scheduling target to construct the first affinity rule for the task scheduling process according to the magnitude relationship between the current processor node utilization rate and the processor node utilization rate threshold; set the second affinity rule for scheduling the execution task thread and the process task with data sharing of the same task to the same scheduling domain based on the preset scheduling target; set the third affinity rule for scheduling the task to the processor core with the lowest utilization rate in the scheduling domain based on the preset scheduling target.
[0068] It is understandable that for the central processing unit (CPU) node affinity, that is, the first affinity rule of the CPU node, for a system with multiple CPU nodes in the system, first set the upper limit of the utilization rate of the CPU node (the average utilization rate on all cores of the CPU node). If it is lower than this processor node utilization threshold, the task is scheduled to the same CPU node; otherwise, it is scheduled to other CPU nodes.
[0069] Regarding the affinity within the scheduling domain, since multiple CPU cores within each scheduling domain (these cores are called a scheduling domain) share an L2 cache, it is necessary to set the second affinity rule for task scheduling to address the possible performance loss caused by contention for the L2 cache among tasks during operation. These second affinity rules constitute the constraint conditions for task scheduling based on reinforcement learning. It includes the following situations: 1) Threads under the same task need to be scheduled to the Cluster sharing the L2 cache; 2) Processes with data sharing need to be scheduled to the Cluster sharing the L2 cache; 3) Mutually exclusive access processes are scheduled to different Clusters; 4) Other tasks are scheduled to low-load Clusters according to CPU utilization.
[0070] Regarding the third affinity rule for each CPU core, first set the upper limit of the utilization rate of each CPU core, and schedule the task to the CPU core with the lowest current utilization rate within the Cluster range.
[0071] In this embodiment, based on the reinforcement learning algorithm and according to the current action experience value table and preset scheduling constraint conditions, determine the task scheduling action at the current time step. Among them, the task scheduling action at the current time step specifically includes the first action, the second action, and the third action. It should be noted that the selection of the first action, the second action, and the third action obtains the target experience value at the current time step from the current action experience value table, and under the constraint of the preset scheduling constraint conditions, uses the scheduling action corresponding to the target experience value as the task scheduling action at the current time step.
[0072] Specifically, based on the values of each state parameter in the range of state parameter changes generated by each scheduling action, determine the task scheduling actions that can be executed at the current time step to construct a target action set; the range of state parameter changes is the range of state parameter changes corresponding to the current system environment information and the current task load information at the current time step; obtain the experience value corresponding to each action in the target action set from the current action experience value table to obtain an experience value set corresponding to the current time step; determine the target experience value corresponding to the current time step from the experience value set according to the highest value selection rule. It can be understood that, based on the values of each state parameter generated by each scheduling action, determine the task scheduling actions that can be executed at the current time step, where each scheduling action can specifically include: the action of scheduling a task to a processor node, the action of scheduling a task to a scheduling domain, and the action of scheduling a task to a processor core. It can be understood that define an action space, so that the task selects a certain CPU node on multiple CPU nodes, a certain Custer scheduling domain on the CPU node, and a specific CPU core on the Cluster scheduling domain as the platform for task execution during operation. According to the preset scheduling constraint conditions defined by the present invention, the incentive function corresponds to three types according to the three types of operations on the action space: 1) When selecting a CPU node for scheduling, the incentive for each selected CPU node action is that, on the premise of ensuring that the set CPU node utilization threshold is not exceeded (this threshold indicates that when the CPU node is below the upper limit, task execution will not be blocked), the CPU nodes used in the system are as few as possible. If there are multiple eligible CPU nodes in the system, preferentially select the node that is physically closest to the CPU node with existing load, because usually the closer the two CPU nodes are, the faster the speed when accessing each other's memory; 2) When selecting a scheduling domain, the incentive for each selected scheduling domain action is to ensure that the threads within the same task and multiple tasks with data sharing are scheduled to the same scheduling domain, and multiple tasks with mutually exclusive data access are scheduled to different scheduling domains; 3) When selecting a specific CPU core, the incentive for each selected CPU core action is that, on the premise of ensuring that the set CPU core utilization threshold is not exceeded, the CPU cores used in the scheduling domain are as few as possible.
[0073] Specifically, the actions of scheduling tasks to processor nodes include: the scheduling action of scheduling a task to the CPU node labeled 1, the scheduling action of scheduling a task to the CPU node labeled 2, …, the scheduling action of scheduling a task to the CPU node labeled N, where N represents the total number of CPU nodes in the current embedded system; the actions of scheduling tasks to scheduling domains include: the scheduling action of scheduling a task from a CPU node to the scheduling domain labeled 1.1, the scheduling action of scheduling a task from a CPU node to the scheduling domain labeled 1.2, …, the scheduling action of scheduling a task from a CPU node to the scheduling domain labeled M, where M represents the total number of scheduling domains; the actions of scheduling tasks to processor cores include: the scheduling action of scheduling a task from a certain scheduling domain to the CPU core labeled 2.1 inside, the scheduling action of scheduling a task from a certain scheduling domain to the CPU core labeled 2.2 inside, …, the scheduling action of scheduling a task from a certain scheduling domain to the CPU core labeled 2.Z inside, where Z represents the total number of CPU cores in the same scheduling domain. In this way, a set of target actions can be constructed, and based on the experience values corresponding to each scheduling action in the set of target actions recorded in the current action experience value table, a set of experience values at the current time step can be obtained, and finally the target experience value corresponding to the highest value in the set of experience values is selected.
[0074] In this embodiment, the processor node selection action experience value corresponding to the current time step in the current action experience value table is determined to select, based on the processor node selection action experience value, a target processor node that satisfies the first affinity rule from the local multi-processor nodes for executing the target task; the scheduling domain selection action experience value corresponding to the current time step in the current action experience value table is determined to select, based on the scheduling domain selection action experience value, a target scheduling domain that satisfies the second affinity rule from each scheduling domain in the target processor node; the processor core selection action experience value corresponding to the current time step in the current action experience value table is determined to select, based on the processor core selection action experience value, a processor core that satisfies the third affinity rule from each processor core in the scheduling domain. It can be understood that after determining the selected target experience value, the execution actions (the first action, the second action, and the third action) corresponding to the target experience value are determined. When executing the first action, the following steps are specifically executed:
[0075] Determine the maximum processor node selection action experience value in the current action experience value table, and based on the maximum processor node selection action experience value, select a processor node from the local multi-processor nodes whose current processor node utilization rate is less than the processor node utilization rate threshold and which already has a task load as the target processor node for executing the target task. It can be understood that when there is a target task, according to the current processor hardware architecture information of the system environment information, each CPU node is determined, and according to the selected experience value in the current action experience value table, further, the CPU nodes that meet the foregoing definitions and conditions recorded in the task load information are determined as the target CPU nodes for executing the target task.
[0076] Specifically, obtain the current processor node utilization rate of each processor node recorded in the task load information; select the processor nodes whose current processor node utilization rate is less than the processor node utilization rate threshold as the initial processor nodes; determine whether there are processor nodes with existing task loads among the initial processor nodes; if there is only one processor node with an existing task load among the initial processor nodes, then determine the processor node with the existing task load as the target processor node for executing the target task; if there are multiple processor nodes with existing task loads among the initial processor nodes, then determine the processor node with the existing task load that is physically closest as the target processor node for executing the target task. It can be understood that according to the action definition, when selecting CPU nodes, if there are CPU nodes whose utilization rate of the current processor node is less than the processor node utilization rate threshold, then the eligible CPU nodes are used as the initial processor nodes, and then further determine whether there are CPU nodes with existing task loads among the initial CPU nodes. If there is exactly one CPU node that meets this condition among the initial CPU nodes, then the initial CPU node is directly used as the target CPU node for executing the target task. If there are multiple CPU nodes that meet this condition among the initial CPU nodes, then the CPU node that is physically closest is preferentially selected as the target CPU node to save the time for accessing each other's memory.
[0077] Specifically, the algorithms for selecting CPU nodes, scheduling domains, and CPU cores are the Q-learning reinforcement learning algorithm. Perform the first task scheduling work, and obtain the corresponding current reward information through a preset incentive function. The Q-learning reinforcement learning algorithm adopted in the present invention is as follows:
[0078] ;
[0079] Among them, 、 respectively represent the learning rate and the discount factor. is used to control the amplitude of each Q-value update and determines the convergence speed of the algorithm; is a discount factor with a value between 0 and 1; It means that at the time step of the state, taking the task scheduling action can obtain the expected value of the reward, that is, the Q value. The system environment will feedback the corresponding incentive according to the task scheduling action . Therefore, the main idea of the algorithm is to construct a two-dimensional table of the state space and the action space to store the Q value, and then select the task scheduling action that can obtain the maximum incentive in combination with the Q value.
[0080] Step S12: Update the current system environment information and the current task load information of the embedded operating system based on the current reward information after executing the task scheduling action.
[0081] In this embodiment, after executing the task scheduling action, the generated operation information is used for incentive feedback, that is, the current system environment information and the current task load information of the embedded operating system are updated according to the returned operation information. Among them, during the information update process, the corresponding information that generates state fluctuations in the system environment information and the task load information is mainly updated, while the static information therein does not need to be updated.
[0082] Step S13: Update the experience value in the current action experience value table based on the current system environment information, the current task load information, and the current reward information, and jump to the step of determining the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions until the target task is scheduled to the target processor core that meets the preset conditions for task execution.
[0083] In this embodiment, before updating the experience value in the current action experience value table based on the current system environment information, the current task load information, and the current reward information, it also includes: obtaining the current processor hardware architecture information and the current task information of the embedded system to obtain the current system environment information; obtaining the current hardware utilization status information and the current task running status information of the embedded system to obtain the current task load information. It can be understood that before task scheduling, it is first necessary to obtain the current system environment and the current task load information to construct the environment and state in reinforcement learning, which includes two types of information: static information and runtime dynamic information.
[0084] Static information includes hardware information and task information in the system. For the hardware information part, that is, the processor hardware architecture information, it shows the physical structure of the processor. Therefore, when obtaining the current system environment information and current task load information, including: obtaining the associated information of the hardware information of the current system environment of the embedded operating system to obtain the current processor hardware architecture information constructed based on each physical structure of the processor, the interconnection information between each physical structure and the preset secondary cache, where the physical structure includes processor nodes, scheduling domains, and processor cores; obtaining the task priority, task cycle information, executing task threads, and data dependency relationships between tasks in the embedded operating system to obtain the corresponding current task information.
[0085] Before obtaining the current processor hardware architecture information, it also includes: dividing the processor cores on the embedded operating system according to whether they share the preset secondary cache to obtain scheduling domains; dividing each scheduling domain according to whether it accesses the local memory of the same node to obtain the corresponding processor nodes; determining the physical structure of the processor according to the architecture relationship of the processor nodes, scheduling domains, and processor cores.
[0086] It can be understood that the processor hardware architecture information includes which processor nodes the processor chip consists of, which scheduling domains each processor node contains, and which central processing unit cores each scheduling domain contains; it also includes the interconnection information between these hardware components and between the above hardware components and the L2 cache, for example: which central processing unit cores share an L2 cache. For task information, it describes information such as task priority and periodicity, which threads the task consists of, and the data dependency relationship between this task and other tasks in the system.
[0087] In this embodiment, obtaining the current hardware utilization status information of the embedded system includes: obtaining the utilization rate of the processor cores during the current system operation in the embedded operating system and the usage of the shared preset secondary cache to obtain the corresponding current hardware utilization status information. It can be understood that the dynamic information includes hardware utilization status information and task running status information. The hardware utilization status information mainly describes the current utilization rate on the central processing unit cores and the usage of the L2 cache; the task running status information mainly describes the current state of the task, including information such as the ready state and the running state.
[0088] In this way, by obtaining the current system environment information, current task load information, and current reward information, the experience values in the current action experience value table are updated, so that the action experience value table updates the experience values according to each execution of the action. When selecting the next execution action, the experience values of the scheduling actions of the optimal processor node, optimal scheduling domain, and optimal processor core determined from the current action experience value table can be determined, and the corresponding scheduling action can be selected for execution until the selected CPU core meets the end requirement.
[0089] As Figure 2 shown, the module structure diagram of the preset task scheduling system is disclosed. Among them, Figure 2 For the specific implementation processes of each module, please refer to the content disclosed above, and details will not be elaborated here.
[0090] To further improve the overall energy consumption efficiency of the system, when the system receives a target task, the task energy consumption characteristics of the target task are analyzed in advance to determine whether the energy consumption type of the target task belongs to the low-energy consumption type or the high-energy consumption type. The specific judgment criteria can be set according to historical data, and details will not be elaborated here. When it is determined that the energy consumption type of the target task is the low-energy consumption type, when selecting the target CPU node, in addition to selecting the node according to the above scheduling conditions and load information, the scheduling condition of the low-energy consumption core condition of each core in the CPU node is added. In this way, during task scheduling, tasks with low energy consumption are preferentially assigned to the low-power CPU cores. For the case where it is determined that the energy consumption type of the target task is the high-energy consumption type, when selecting the target CPU node, in addition to selecting the node according to the above scheduling conditions and load information, the scheduling condition of the high-energy consumption core condition of each core in the CPU node is added. In this way, during task scheduling, tasks with high energy consumption are preferentially assigned to the high-power CPU cores, optimizing the overall energy consumption of the system. For example: For some simple sensor data acquisition tasks (target tasks of the low-power consumption type), they are scheduled to run on the CPU cores operating in the energy-saving mode; for computationally intensive tasks that are not very sensitive to energy consumption, such as scientific computing simulation tasks, they can be assigned to the cores with high performance but high energy consumption, optimizing the overall energy consumption of the system in this way.
[0091] Refer to Figure 3As shown in the figure, according to the defined action space and incentive function, the execution process of the task scheduling method based on reinforcement learning is as follows: First, obtain the system environment information and task load information, and establish a corresponding Q-learning reinforcement learning algorithm according to the action space on the core selection scheduling and the state space on the system environment model. For the scheduling at three levels, set the utilization thresholds on the CPU nodes and each CPU core. Based on the reinforcement learning algorithm, first select the optimal CPU node and schedule the task to this node, and then sequentially select the optimal Cluster and the CPU cores within the Cluster, and finally accurately schedule the task to run on a specific CPU core. Specifically, this embodiment takes a platform with a 2-CPU-node architecture as an example to illustrate the specific implementation process of the method proposed by the present invention. Each CPU node has 4 Cluster scheduling domains, and there are 32 CPU cores inside each Cluster, and these cores share an L2 cache. The implementation process is as follows:
[0092] 1) Obtain the system environment information, including the information of the CPU, Cluster, and CPU cores and the cache information, and set the utilization thresholds on each CPU node and CPU core. For example, set the upper limit of the utilization rate of each CPU node to 50%, and the upper limit of the utilization rate of each core to 60%.
[0093] 2) Obtain the task load information. When there are two tasks A and B sharing data in the system, mainly obtain several aspects such as the status of the tasks (whether they are ready), the scheduled cores, and the association relationships with other tasks. Here, set task A to the ready state, the scheduled core to NULL, and there is shared data with task B; task B is set to the running state, the scheduled core is the 0th core on the 0th Cluster on CPU 0, and the utilization rate of the 0th core is 10%.
[0094] 3) First perform the optimization operation on the CPU node through the reinforcement learning method (such as using Q-learning). Calculate that the average utilization rate of CPU 0 in the current system is 3%, and task A also selects CPU 0; if there is no suitable node, repeat step 3).
[0095] 4) Perform the optimization operation on the Cluster based on Q-learning. Task B, which has shared data with task A, is currently running on the 0th Cluster, so task A also selects the 0th Cluster;
[0096] 5) If no suitable Cluster is found in other cases, first repeat step 4). If it is still not successful, fallback to repeat step 3), and at this time, it is necessary to select another CPU node.
[0097] 6) Perform optimization operations on the CPU core based on Q-learning. Task B is currently running on core 0, and the current utilization rate of this core is 10%, which is lower than 60%. Then task A also selects core 0;
[0098] 7) If no suitable core is found in other cases, first repeat step 6). If it still fails, fallback to repeating step 4), and at this time, another Cluster needs to be selected.
[0099] 8) After selecting the core, run task A, and update the status of task A to the running state. The scheduled core is (0 - 0 - 0), that is, CPU 0, Cluster 0, and core 0. At the same time, update the resource utilization rates on the CPU cores and nodes in the system.
[0100] In this way, for the processor architecture with shared L2 cache, based on the real-time perception of the system environment and task load, precise task scheduling on the CPU core for performance optimization is achieved, fully considering the shared data characteristics between tasks, and avoiding performance losses caused by shared L2 cache; at the same time, the CPU node and core utilization rates in task scheduling are considered to reduce performance losses caused by CPU overload.
[0101] It can be seen that the present invention discloses a task scheduling method based on reinforcement learning, including: determining the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions; the task scheduling action includes a first action, a second action, and a third action determined according to the target experience value corresponding to the current time step in the current action experience value table; the first action is used to select the target processor node for executing the target task, the second action is used to select the target scheduling domain in the target processor node, and the third action is used to select the processor core in the target scheduling domain; several processor cores capable of sharing the preset secondary cache are configured in each scheduling domain; updating the current system environment information and the current task load information of the embedded operating system based on the current reward information after performing the task scheduling action by the embedded operating system; updating the experience values in the current action experience value table based on the current system environment information, the current task load information, and the current reward information, and jumping to the step of determining the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions until the target task is scheduled to the target processor core that meets the preset conditions for task execution.
[0102] As can be seen from the above technical solutions, by comprehensively considering the system environment information and task load information, and through reinforcement learning, tasks can be accurately assigned to appropriate hardware resources, avoiding resource idleness or overuse. Considering the task scheduling optimization problem on the processor architecture with a preset secondary cache, it avoids the reduction of application performance caused by cache contention. According to the task association relationship and task load information of the target processor node, etc., the target task is scheduled to the processor core with the lowest utilization rate in the appropriate scheduling domain for execution. It can make full use of the data sharing advantage brought by sharing the preset secondary cache, reduce the latency of data access during task execution, thereby accelerating the task execution speed, improving the execution efficiency of a single task, and the overall performance of the system for multitasking processing. Optimize the core selection scheduling of ready tasks to achieve fine-grained scheduling at the CPU core level and maximize the guarantee of the running performance of application tasks. Finally, as the embedded operating system runs, tasks are continuously generated and ended, and the system environment and task load situation are in dynamic change. Through the link of updating the system environment information and task load information, these changes can be tracked in real time, so that subsequent task scheduling decisions can always be adjusted based on the latest and accurate system state.
[0103] Furthermore, an embodiment of the present invention also discloses an electronic device, Figure 4 which is a structural diagram of an electronic device shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation to the scope of use of the present invention. The electronic device may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the task scheduling method based on reinforcement learning disclosed in any of the foregoing embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.
[0104] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present invention, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0105] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0106] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the task scheduling method based on reinforcement learning executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.
[0107] Furthermore, the present invention also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the task scheduling method based on reinforcement learning disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0108] Furthermore, the present invention also discloses a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps of the method disclosed in any of the foregoing embodiments.
[0109] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0110] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0111] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0112] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0113] The technical solutions provided by the present invention have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A task scheduling method based on reinforcement learning, characterized in that Including: Construct a first affinity rule for the task scheduling process based on the size relationship between the current processor node utilization rate and the processor node utilization rate threshold in the processor hardware architecture information set based on a preset scheduling target; construct a second affinity rule for scheduling the execution task threads and process tasks with data sharing of the same task to the same scheduling domain based on the preset scheduling target setting; construct a third affinity rule for scheduling tasks to the processor core with the lowest utilization rate in the scheduling domain based on the preset scheduling target setting; construct a preset scheduling constraint condition for the scheduling task process based on the first affinity rule, the second affinity rule, and the third affinity rule. Determine the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint condition; the task scheduling action includes a first action, a second action, and a third action determined based on the target experience value corresponding to the current time step in the current action experience value table; the first action is used to select the target processor node for executing the target task, the second action is used to select the target scheduling domain in the target processor node, and the third action is used to select the processor core in the target scheduling domain; several processor cores capable of sharing a preset secondary cache are configured in each scheduling domain. Obtain the current reward information after the embedded operating system executes the task scheduling action through a preset incentive function, and update the current system environment information and the current task load information of the embedded operating system; the incentive function includes: when selecting a CPU node for scheduling, the incentive for each selected CPU node action is that, on the premise of ensuring that the set CPU node utilization rate threshold is not exceeded, as few CPU nodes as possible are used in the system; if there are multiple eligible CPU nodes in the system, preferentially select the node with the physical distance as close as possible to the CPU node with existing load, because usually the closer the two CPU nodes are, the faster the speed when accessing each other's memory; when selecting a scheduling domain, the incentive for each selected scheduling domain action is to ensure that the threads within the same task and multiple tasks with data sharing are scheduled to the same scheduling domain, and multiple tasks with mutually exclusive data access are scheduled to different scheduling domains; when selecting a specific CPU core, the incentive for each selected CPU core action is that, on the premise of ensuring that the set CPU core utilization rate threshold is not exceeded, as few CPU cores as possible are used in the scheduling domain. Based on the current system environment information, the current task load information, and the current reward information, update the experience values in the current action experience value table, so that when making the next execution action selection, the experience values of the scheduling actions of the optimal processor node, the optimal scheduling domain, and the optimal processor core determined from the current action experience value table can be obtained, and the corresponding scheduling actions are selected and executed.
2. The task scheduling method based on reinforcement learning according to claim 1, wherein Before updating the experience values in the current action experience value table based on the current system environment information, the current task load information, and the current reward information, it further includes: Obtain the current processor hardware architecture information and the current task information of the embedded system to obtain the current system environment information. Obtain the current hardware utilization status information and the current task running status information of the embedded system to obtain the current task load information.
3. The task scheduling method based on reinforcement learning according to claim 2, wherein The preset scheduling objective is a scheduling objective determined based on having the optimal system performance during the scheduling process.
4. The task scheduling method based on reinforcement learning according to claim 3, wherein, The determining the task scheduling action at the current time step according to the current action experience value table and the preset scheduling constraint conditions includes: Determine the processor node selection action experience value corresponding to the current time step in the current action experience value table, and based on the processor node selection action experience value, select a target processor node that meets the first affinity rule from the local multi-processor nodes for executing the target task; Determine the scheduling domain selection action experience value corresponding to the current time step in the current action experience value table, and based on the scheduling domain selection action experience value, select a target scheduling domain that meets the second affinity rule from each of the scheduling domains in the target processor node; Determine the processor core selection action experience value corresponding to the current time step in the current action experience value table, and based on the processor core selection action experience value, select a processor core that meets the third affinity rule from each of the processor cores in the scheduling domain.
5. The task scheduling method based on reinforcement learning according to claim 2, wherein The obtaining the current processor hardware architecture information and the current task information of the embedded system includes: Obtain the associated information of the hardware information of the current system environment of the embedded operating system to obtain the current processor hardware architecture information constructed based on each physical structure of the processor, the interconnection information between each physical structure and the preset secondary cache, where the physical structure includes a processor node, a scheduling domain, and a processor core; Obtain the task priority, task period information, execution task thread, and data dependency relationship between tasks of the current task in the embedded operating system to obtain the corresponding current task information.
6. The task scheduling method based on reinforcement learning according to claim 5, wherein, Before the obtaining the current processor hardware architecture information and the current task information of the embedded system, it further includes: Divide the processor cores on the embedded operating system according to whether they share the preset secondary cache to obtain scheduling domains; Divide each of the scheduling domains according to whether they access the local memory of the same node to obtain the corresponding processor nodes; Determine the physical structure of the processor according to the architecture relationship between the processor node, the scheduling domain, and the processor core.
7. The task scheduling method based on reinforcement learning according to claim 2, wherein The obtaining the current hardware utilization status information of the embedded system includes: Obtain the utilization rate of the processor cores during the current system operation in the embedded operating system and the usage situation of the shared preset secondary cache to obtain the corresponding current hardware utilization status information.
8. The task scheduling method based on reinforcement learning according to claim 4, wherein, The determining the processor node selection action experience value corresponding to the current time step in the current action experience value table, and based on the processor node selection action experience value, select a target processor node that meets the first affinity rule from the local multi-processor nodes for executing the target task, includes: Determine the maximum processor node selection action experience value in the current action experience value table, and based on the maximum processor node selection action experience value, select from the local multi-processor nodes the processor nodes whose current processor node utilization rate is less than the processor node utilization rate threshold and that already have task loads, and determine them as the target processor nodes for executing the target task.
9. The task scheduling method based on reinforcement learning according to claim 8, wherein The step of selecting from the local multi-processor nodes the processor nodes whose current processor node utilization rate is less than the processor node utilization rate threshold and that already have task loads, and determining them as the target processor nodes for executing the target task, includes: Obtain the current processor node utilization rate of each of the processor nodes recorded in the task load information; Select the processor nodes whose current processor node utilization rate is less than the processor node utilization rate threshold as the initial processor nodes; Judge whether there are processor nodes with existing task loads among the initial processor nodes; If there is only one processor node with an existing task load among the initial processor nodes, then determine the processor node with the existing task load as the target processor node for executing the target task.
10. The task scheduling method based on reinforcement learning according to claim 9, wherein, After judging whether there are processor nodes with existing task loads among the initial processor nodes, it further includes: If there are multiple processor nodes with existing task loads among the initial processor nodes, then determine the processor node with the existing task load that is physically closest as the target processor node for executing the target task.
11. The task scheduling method based on reinforcement learning according to any one of claims 1 to 10, characterized in that, It further includes: Based on each state parameter value in the state parameter change range generated by each scheduling action, determine the task scheduling actions that can be executed at the current time step to construct a target action set; the state parameter change range is the state parameter change range corresponding to the current system environment information and the current task load information at the current time step; Obtain the experience value corresponding to each action in the target action set from the current action experience value table to obtain an experience value set corresponding to the current time step; Determine the target experience value corresponding to the current time step from the experience value set according to the highest value selection rule.
12. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the reinforcement learning-based task scheduling method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the steps of the reinforcement learning-based task scheduling method according to any one of claims 1 to 11.
14. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps of the reinforcement learning-based task scheduling method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method for implementing load equalization of multicore processor operating system
CN101256515A
Task scheduling method and device, equipment and storage medium
CN117519919A