Scheduling method and system for Direct IO intensive tasks under NUMA architecture

By using user-state scheduler and bpf technology under the NUMA architecture, dynamically adjusting the scheduling and memory distribution of Direct IO-intensive tasks, the performance problems of CFS schedulers in this scenario are solved, and efficient task scheduling and performance improvement are achieved.

CN119473564BActive Publication Date: 2025-05-06KYLIN CORP

Patent Information

Application Number
CN202510052540.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Under the NUMA architecture, the CFS scheduler cannot effectively schedule when facing Direct IO-intensive tasks, resulting in extended task execution and degrading system performance.

Method used

By creating a user-state scheduler, using the Linux kernel's sched_ext extensible scheduler and bpf technology, efficient scheduling for Direct IO-intensive tasks is achieved. The specific steps include: scanning and recording the NUMA node where each disk device is located, establishing a scheduling domain, tracking the Direct IO amount of tasks, and dynamically adjusting the task scheduling range and memory migration.

Benefits of technology

This method can minimize the overhead of Direct IO-intensive tasks accessing disks across nodes, achieve efficient scheduling, improve task performance, and facilitate deployment without modifying the kernel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119473564B_ABST
    Figure CN119473564B_ABST
Patent Text Reader

Abstract

The present invention provides a scheduling method and system for Direct IO intensive tasks under NUMA architecture, which creates a user-mode scheduler; establishes a scheduling domain, initializes CPU utilization, and implements a bpf callback function; takes over the scheduling strategy of the sched_ext extensible scheduler; interacts with the callback function through the bpf map to implement task queuing and dispatching; marks Direct IO intensive tasks according to whether the IO volume exceeds the IO volume threshold; schedules tasks in corresponding disk nodes; dynamically adjusts the task scheduling range if it exceeds the set utilization threshold; and determines whether it is necessary to migrate the memory to the disk node being accessed according to the set threshold. The present invention reduces the overhead caused by accessing disks across NUMA nodes, implements efficient scheduling, and improves task performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of operating systems, and in particular relates to a scheduling method and system for Direct IO intensive tasks under a NUMA architecture. Background Art

[0002] At present, most Linux systems use CFS scheduler (fair scheduler) as the default task scheduling method. CFS is known for its fairness and efficiency. By introducing the concept of virtual time, it ensures that all tasks can obtain relatively fair system resource allocation.

[0003] However, although CFS can meet the performance requirements of most application scenarios, its scheduling strategy seems to be inadequate in some special scenarios. In particular, when there are a large number of Direct IO intensive tasks in the Linux system under the NUMA architecture, the scheduling effect of CFS is not ideal. Such tasks usually involve a large number of data read and write operations, which often need to wait for the response of external devices (such as hard disks), thus consuming a lot of waiting time. Under the NUMA architecture, due to the non-uniformity of memory access, such tasks may further extend the execution time because of waiting for data to be transferred from remote memory nodes to local cores, thereby reducing the overall performance of the system.

[0004] In addition, the scheduling strategy of CFS is fixed and cannot be adjusted according to different task types. This means that in application scenarios dominated by Direct IO intensive tasks, CFS cannot provide targeted scheduling methods, and thus cannot fully utilize the potential of the NUMA architecture. Therefore, an efficient scheduling method for such application scenarios is needed to improve the performance of such application scenarios. Summary of the invention

[0005] The purpose of the present invention is to provide a scheduling method and system for Direct IO intensive tasks under NUMA architecture, reduce the number of tasks accessing disks across nodes, achieve efficient scheduling, and improve the performance of such tasks.

[0006] In order to achieve the above object, the technical solution of the present invention is as follows:

[0007] A scheduling method for Direct IO intensive tasks under NUMA architecture, comprising:

[0008] S1. Create a user-mode scheduler;

[0009] S2, scan and record the NUMA nodes where each disk device is located, establish a scheduling domain, initialize CPU utilization, create a thread pool, and assign threads to regularly update CPU utilization;

[0010] S3, implement bpf callback function based on sched_ext extensible scheduler of Linux kernel;

[0011] S4. Attach the bpf callback function to each callback point of the sched_ext extensible scheduler through the libbpf library, and take over the scheduling strategy of the sched_ext extensible scheduler;

[0012] S5. Interact with the callback function through bpf map to realize task queuing and dispatching;

[0013] S6. Track the amount of Direct IO of the task in unit time, accumulate and decay the IO amount, and mark the Direct IO intensive task according to whether the IO amount exceeds the IO amount threshold;

[0014] S7, checking the disk device being accessed by the Direct IO intensive task, and scheduling the task in the corresponding disk node;

[0015] S8, regularly checking the CPU utilization within the Direct IO intensive task scheduling range, and dynamically adjusting the task scheduling range if the utilization exceeds a set threshold;

[0016] S9. Regularly check the distribution of the private memory of the Direct IO intensive task in each NUMA node, and decide whether to migrate the memory to the disk node being accessed according to a set threshold.

[0017] Further, step S2 includes:

[0018] Open the / sys / class / block directory, read the soft link of each file, find the PCI device file and record the PCI device number; then read / sys / bus / pci / devices / PCI device number / numa_node, and use the global structure variable to record each PCI device and the corresponding NUMA node;

[0019] Read the / proc / schedstat file, which contains information about the scheduling domain level of each CPU in the system. Parse the file content and record the cpumask mask of each scheduling domain level of the CPU in the global array variable.

[0020] Read the / proc / stat file, which contains information about the time of each CPU. Parse the file content and record the time of each CPU in the global variable, including the user state time, system state time, idle time, and interrupt time of each CPU.

[0021] Call the C library function pthread_create() to apply for working threads in advance. Each working thread calls pthread_cond_wait() to block and wait for tasks to be issued. Call pthread_cond_signal() to wake up a working thread, which updates the CPU utilization once a second.

[0022] Further, step S3 includes:

[0023] A bpf program is written based on the sched_ext extensible scheduler of the Linux kernel to implement related bpf callback functions, which include: init callback function, init_task callback function, select_cpu callback function, enqueue callback function, and dispatch callback function.

[0024] Further, step S6 includes:

[0025] Take the task out of the bpf enqueued map and queue it. If the set time has passed since the last sampling of the task, update the last sampling time point and the accumulated IO amount. The accumulated IO amount is calculated according to the weighted moving average algorithm EWMA. Each time, compare whether the accumulated IO amount exceeds the set IO amount threshold. If so, mark the task as a Direct IO intensive task.

[0026] Further, step S7 includes:

[0027] Check whether the disk device and NUMA node that the Direct IO intensive task is accessing have been set. If not, read the file descriptor in the / proc / PID / fd directory to find the disk device file, match it with the previously recorded disk device and NUMA node, and record it; at the same time, initialize the scheduling domain of the task to the lowest level, and pass the lowest level cpumask mask to the bpf callback function dispatch through the dispatched map.

[0028] Further, step S8 includes:

[0029] Periodically traverse the utilization of all CPUs within the current scheduling domain of Direct IO intensive tasks, calculate whether the average CPU utilization exceeds the first threshold, and if so, increase the scheduling domain level; if the average CPU utilization is lower than the second threshold, decrease the scheduling domain level; Direct IO intensive tasks dynamically adjust the CPU operation range according to the CPU busy situation. When the scheduling domain is under high load, increase the scheduling domain level to reduce the scheduling delay caused by task queuing; when the CPU in the scheduling domain is idle, decrease the scheduling domain level to improve the cache hit rate of Direct IO intensive tasks.

[0030] Further, step S9 includes:

[0031] A regular time is set to check the distribution of the private memory of the Direct IO intensive task in each NUMA node. When the time arrives, the / proc / PID / numa_maps file is opened, the private memory pages in the file are parsed, and the size of the private pages in each NUMA node is counted respectively; if it exceeds the set threshold, the C library function migrate_pages() is called to perform memory migration, and the memory migration task is assigned to each thread for processing in a thread pool manner, and the memory is migrated to the disk node being accessed by the Direct IO intensive task.

[0032] On the other hand, the present invention also proposes a scheduling system for Direct IO intensive tasks under NUMA architecture, comprising:

[0033] Program module: create a user-mode scheduler;

[0034] Initialization module: scans and records the NUMA nodes where each disk device is located, establishes a scheduling domain, initializes CPU utilization, creates a thread pool, and assigns threads to regularly update CPU utilization;

[0035] Callback function module: implement bpf callback function based on sched_ext extensible scheduler of Linux kernel;

[0036] Policy module: attaches the bpf callback function to each callback point of the sched_ext extensible scheduler through the libbpf library, and takes over the scheduling policy of the sched_ext extensible scheduler;

[0037] Interaction module: interacts with the callback function through bpf map to realize task queuing and dispatching;

[0038] Marking module: tracks the amount of Direct IO in a task per unit time, accumulates and decays the amount of IO, and marks Direct IO-intensive tasks based on whether the amount of IO exceeds the IO threshold;

[0039] Node scheduling module: checks the disk device being accessed by the Direct IO intensive task and schedules the task in the corresponding disk node;

[0040] Adjustment module: regularly checking the CPU utilization within the Direct IO intensive task scheduling range, and dynamically adjusting the task scheduling range if the utilization exceeds a set threshold;

[0041] Migration module: regularly checks the distribution of the private memory of the Direct IO intensive task in each NUMA node, and determines whether the memory needs to be migrated to the disk node being accessed according to a set threshold.

[0042] The present invention also provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned scheduling method for Direct IO intensive tasks under the NUMA architecture.

[0043] The present invention also provides a computer program product, including a computer program, which implements the above-mentioned scheduling method for Direct IO intensive tasks under the NUMA architecture when executed by a processor.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. The scheduling method and system for Direct IO intensive tasks under the NUMA architecture proposed in the present invention do not require modification of the kernel in a production environment where downtime for maintenance is not possible. The pluggable loading scheduling method replaces the original scheduling method of the system, which is convenient for deployment without restarting.

[0046] 2. The present invention optimizes the scheduling logic for Direct IO intensive tasks, is more flexible, and can adjust various parameters and thresholds at any time according to different on-site production environments to meet needs.

[0047] 3. The present invention provides a specific scheduling method for Direct IO intensive tasks, which can minimize the overhead caused by accessing disks across NUMA nodes, achieve efficient scheduling, and improve task performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a schematic diagram of the overall principle framework of Embodiment 1 of the present invention.

[0049] Figure 2 This is a basic flow chart of the first embodiment of the present invention.

[0050] Figure 3This is a flow chart of Direct IO intensive task identification according to the first embodiment of the present invention.

[0051] Figure 4 This is a flow chart of Direct IO intensive task scheduling according to the first embodiment of the present invention.

[0052] Figure 5 This is a flow chart of a domain-wide dynamic update of Direct IO intensive task scheduling according to Embodiment 1 of the present invention.

[0053] Figure 6 This is a flow chart of private memory scanning and memory migration for Direct IO intensive tasks according to the first embodiment of the present invention. DETAILED DESCRIPTION

[0054] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0055] Related technical terms:

[0056] NUMA: Non Uniform Memory Access architecture.

[0057] Direct IO: The application bypasses the operating system's caching mechanism and directly transfers data to the underlying storage device (such as a hard disk or SSD).

[0058] CFS: Completely Fair Scheduler in the Linux kernel scheduling subsystem.

[0059] BPF: Berkeley Packet Filter, originally designed as an efficient network packet filtering mechanism. However, over time, BPF has evolved into a more powerful technology that is widely used in the Linux kernel and user space to perform various tasks, including network monitoring, performance analysis, and system security.

[0060] sched-ext: BPF extensible scheduler class.

[0061] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0062] Embodiment 1:

[0063] like Figure 1The figure shows the overall principle framework diagram provided by the embodiment of the present invention. The present invention is based on the sched_ext extensible scheduler and implements the customization of the scheduling strategy through the bpf callback interface. At the same time, the queuing and dispatching of tasks are completed through the interaction of bpfmap and the user-mode scheduler to complete the reading, dispatching and a series of checks of tasks, including: task identification, CPU utilization tracking, private memory scanning, memory migration, and dynamic scheduling domain. Through the sched_ext extensible scheduler, the present invention can be conveniently deployed for the Direct IO intensive task scheduling method.

[0064] like Figure 2 This is a basic flow chart provided by an embodiment of the present invention, which specifically includes the following steps:

[0065] S1: Create a user-mode scheduler.

[0066] This embodiment is implemented in C language, and other languages ​​such as Rust and C++ may also be used.

[0067] S2: Scan and record the NUMA nodes where each disk device is located, establish a scheduling domain, initialize CPU utilization, create a thread pool, and assign threads to regularly update CPU utilization.

[0068] This step is the initialization process of the user-mode scheduler, which needs to collect necessary system information for subsequent task scheduling and memory migration and provide some basic support.

[0069] Among them, the scanning and recording of the NUMA node where the system disk device is located is exemplified as follows: opening the / sys / class / block directory, reading the soft link of each file therein, finding the PCI device file therein and recording the PCI device number; then reading / sys / bus / pci / devices / PCI device number / numa_node, using the global structure variable to record each PCI device and the corresponding NUMA node, in order to facilitate the subsequent search task to find the node where the disk being accessed is located.

[0070] Among them, the establishment of the scheduling domain of each CPU is exemplified by: reading the / proc / schedstat file, which contains scheduling domain level information about each CPU in the system, parsing the content of the file, and recording the cpumask mask of each scheduling domain level of the CPU in a global array variable, in order to prepare for the subsequent implementation of the dynamic scheduling range of tasks.

[0071] Among them, the initialization of each CPU utilization is exemplified by: reading the / proc / stat file, which contains information about the time of each CPU, parsing the content of the file, and recording the time of each CPU in a global variable, including the user state time, system state time, idle time, and interrupt time of each CPU, in order to provide a basic value for realizing real-time update of CPU utilization.

[0072] The thread pool is created and threads are assigned to update CPU utilization periodically. For example, the C library function pthread_create() is called to apply for a worker thread in advance, and each worker thread calls pthread_cond_wait() to block and wait for tasks to be issued. pthread_cond_signal() is called to wake up a worker thread, and the worker thread updates CPU utilization once per second.

[0073] S3: Write a bpf program based on the sched_ext extensible scheduler of the Linux kernel to implement related bpf callback functions.

[0074] Related bpf callback functions include: init callback function, init_task callback function, select_cpu callback function, enqueue callback function, and dispatch callback function.

[0075] The sched_ext extensible scheduler was introduced in Linux kernel 6.12-rc1. The scheduler adds a user-defined extension interface definition and defines a set of BPF-based extension functions. Take select_task_rq_scx as an example. During operation, it will determine whether the select_cpu callback interface (generally referred to as ops.select_cpu) in the corresponding sched_ext_ops structure is registered. If the loaded program defines the callback interface, it will be called and executed. If it is not defined, the original process will continue. Therefore, the scheduling policy can be customized through the user-defined BPF callback interface, and due to the characteristics of BPF technology, it can be easily pluggable without modifying the kernel.

[0076] As described in step S3, the init callback function is implemented, including obtaining the process PID of the user-mode scheduler. Exemplarily, in the bpf program, the global variables are defined as follows:

[0077] int usr_sched;

[0078] In the user-mode scheduler, when loading the bpf program through the libbpf-bootstrap framework, assign usr_sched:

[0079] skel->rodata->usr_sched = getpid();

[0080] As described in step S3, the init_task callback function is implemented. For example, an entry is added to the global bpfmap: task_ctx_stor for each newly created process. The map type is BPF_MAP_TYPE_TASK_STORAGE. Each entry in the map corresponds to a process. Each entry is represented by a structure to record relevant parameters of the process, including: a flag indicating whether the task is a kernel thread.

[0081] The function calls the bpf helper function bpf_task_storage_get() to add a bpf entry for each newly created process p: bpf_task_storage_get(&task_ctx_stor,p,0,BPF_LOCAL_STORAGE_GET_F_CREATE);

[0082] In step S3, the select_cpu callback function is implemented. Exemplarily, if the task is a kernel thread, the task is marked, and other tasks directly return to prev_cpu.

[0083] As described in step S3, the enqueue callback function is implemented. For example, if the task is marked as a kernel thread, it is directly enqueued into the local queue. For other tasks, the task is enqueued into the globally created bpf map: enqueued, where: the map type is BPF_MAP_TYPE_QUEUE, each value in the map corresponds to an enqueued process, and the value uses the structbpf_enqueued_task structure to record the relevant parameters of each enqueued process, including: task PID, actual task execution time, and task weight. Setting the global flag indicates that the user-mode scheduler needs to run.

[0084] As described in step S3, the dispatch callback function is implemented. For example, if the user-mode scheduler running flag is set in the enqueue callback interface, the program is dispatched to the global queue, so that it can be scheduled and run first, and then the task to be dispatched is taken from the globally created bpf map: dispatched, where: the map type is BPF_MAP_TYPE_QUEUE, each value in the map corresponds to a process to be dispatched, and the value uses the struct bpf_dispatched_task structure to record the relevant parameters of each process to be dispatched, including: task PID, the NUMA node where the task is accessing the disk, and the CPU range where the task can run. Once the NUMA node where the disk is being accessed is set, the kfunc function scx_bpf_pick_any_cpu() provided by the kernel is called to select an idle and online CPU from the CPU range where the task can run to dispatch the task, otherwise the task is dispatched to the global queue.

[0085] S4: Attach the callback function to each callback point of the kernel's sched_ext scheduler through the libbpf library and take over the scheduling policy of the sched_ext scheduler.

[0086] Exemplary: call bpf_object__open_skeleton() to open each callback function implemented in the S3 step; call bpf_object__load_skeleton() to load each callback function into the kernel; call bpf_object__attach_skeleton() to attach each callback function to each callback point of the sched_ext scheduler.

[0087] S5: The bpf map interacts with the callback function, and the user-mode scheduler implements task queuing and dispatching.

[0088] For example: in the enqueue callback function in S3, the task will be enqueued to the globally created bpf map: enqueued; when the user-mode scheduler runs, it takes out the tasks in the map, queues the tasks according to the virtual running time, and then queues the queued tasks one by one to the globally created bpf map: dispatched; in the dispatch callback function in S3, the tasks in the dispatched map are taken out and dispatched to the CPU one by one.

[0089] S6: Track the Direct IO amount of the task in unit time, accumulate and decay the IO amount, and mark the Direct IO intensive task according to whether the IO amount exceeds the IO amount threshold.

[0090] Exemplary: Figure 3 As shown, the user-mode scheduler takes tasks from bpf map: enqueued and queues them. If the task has passed the set time threshold since the last sampling, the last sampling time point and the accumulated IO volume are updated. The accumulated IO volume is calculated according to the weighted moving average algorithm (EWMA), and the formula is as follows:

[0091] ;

[0092] in:

[0093] It is the weighted moving average of the IO volume at time t;

[0094] is the attenuation factor, the smaller its value is, the faster it decreases;

[0095] is the average IO volume sampled in the time period (t-1, t];

[0096] The formula can be simplified to:

[0097] ;

[0098] in: It is the weighted moving average of the IO volume at time t-1.

[0099] exist Figure 3 In this example Set to 0.7, so if the result of a single sampling fluctuates greatly, it will not have a big impact on the growth of the cumulative result, so that the curve can be smoothed. Each time, compare whether the cumulative value exceeds the set IO volume threshold. If so, mark the task as a Direct IO intensive task;

[0100] S7: Check which disk device the Direct IO intensive task is accessing and schedule the task in the corresponding disk node.

[0101] Exemplary: Figure 4As shown, in step S6, after the task is identified as a Direct IO intensive task, further check whether the disk device being accessed by the task and the NUMA node where it is located have been set. If not, read the file descriptor in the / proc / PID / fd directory to find the disk device file therein, match it with the disk device and NUMA node recorded in step S2 and record it. At the same time, initialize the scheduling domain of the task to the lowest level, and pass the lowest level cpumask mask to the bpf callback function dispatch through the dispatched map.

[0102] S8: Regularly check the CPU utilization within the Direct IO intensive task scheduling range, and dynamically adjust the task scheduling range if it exceeds a certain threshold.

[0103] Exemplary: Figure 5 As shown, the utilization of all CPUs within the current scheduling domain of the task is traversed, and the average CPU utilization is calculated to see if it exceeds the first threshold (set to 80% in this embodiment). If so, the scheduling domain level is raised. If the average CPU utilization is lower than the second threshold (set to 50% in this embodiment), the scheduling domain level is lowered. This implementation method can ensure that the task dynamically adjusts the CPU operating range according to the CPU busyness. Under high load conditions in the scheduling domain, raising the scheduling domain level helps reduce the scheduling delay caused by task queuing. When the CPU in the scheduling domain is idle, lowering the scheduling domain level helps improve the cache hit rate of the task.

[0104] S9: Regularly check the distribution of private memory of Direct IO intensive tasks on each NUMA node, and decide whether to migrate the memory to the disk node being accessed based on the set threshold.

[0105] Exemplary: Figure 6 As shown, for example, the scanning interval is set to 1 second. After the time is reached, the / proc / PID / numa_maps file is opened, the private memory pages in the file are parsed, and the size of the private pages in each NUMA node is counted. If the set threshold is exceeded, such as more than 20 pages, the C library function migrate_pages() is called to perform memory migration. Since migrating memory is time-consuming, it is time-consuming to perform such an operation in the user-state scheduler program, and it will be aggravated when there are many tasks. Therefore, the present invention adopts a thread pool method to assign the memory migration task to each thread for processing, so as not to affect the performance of the user-state scheduler.

[0106] The beneficial effects of this embodiment are as follows:

[0107] Based on the sched_ext scheduler, it can be easily deployed without restarting in a production environment where downtime for maintenance is not possible.

[0108] Customized scheduling strategies are more flexible and can adjust parameters and thresholds at any time to meet needs according to different on-site production environments.

[0109] The specific scheduling method provided for Direct IO intensive tasks can minimize the overhead caused by cross-node disk access for such tasks and improve task performance.

[0110] Embodiment 2:

[0111] This second embodiment proposes a scheduling system for Direct IO intensive tasks under the NUMA architecture, including:

[0112] Program module: create a user-mode scheduler;

[0113] Initialization module: scans and records the NUMA nodes where each disk device is located, establishes a scheduling domain, initializes CPU utilization, creates a thread pool, and assigns threads to regularly update CPU utilization;

[0114] Callback function module: implement bpf callback function based on sched_ext extensible scheduler of Linux kernel;

[0115] Policy module: attaches the bpf callback function to each callback point of the sched_ext extensible scheduler through the libbpf library, and takes over the scheduling policy of the sched_ext extensible scheduler;

[0116] Interaction module: interacts with the callback function through bpf map to realize task queuing and dispatching;

[0117] Marking module: tracks the amount of Direct IO in a task per unit time, accumulates and decays the amount of IO, and marks Direct IO-intensive tasks based on whether the amount of IO exceeds the IO threshold;

[0118] Node scheduling module: checks the disk device being accessed by the Direct IO intensive task and schedules the task in the corresponding disk node;

[0119] Adjustment module: regularly checking the CPU utilization within the Direct IO intensive task scheduling range, and dynamically adjusting the task scheduling range if the utilization threshold is exceeded;

[0120] Migration module: regularly checks the distribution of the private memory of the Direct IO intensive task in each NUMA node, and determines whether the memory needs to be migrated to the disk node being accessed according to a set threshold.

[0121] In this system, the initialization module includes:

[0122] Open the / sys / class / block directory, read the soft link of each file, find the PCI device file and record the PCI device number; then read / sys / bus / pci / devices / PCI device number / numa_node, and use the global structure variable to record each PCI device and the corresponding NUMA node;

[0123] Read the / proc / schedstat file, which contains information about the scheduling domain level of each CPU in the system. Parse the file content and record the cpumask mask of each scheduling domain level of the CPU in the global array variable.

[0124] Read the / proc / stat file, which contains information about the time of each CPU. Parse the file content and record the time of each CPU in the global variable, including the user state time, system state time, idle time, and interrupt time of each CPU.

[0125] Call the C library function pthread_create() to apply for working threads in advance. Each working thread calls pthread_cond_wait() to block and wait for tasks to be issued. Call pthread_cond_signal() to wake up a working thread, which updates the CPU utilization once a second.

[0126] In this system, the callback function module includes:

[0127] A bpf program is written based on the sched_ext extensible scheduler of the Linux kernel to implement related bpf callback functions, which include: init callback function, init_task callback function, select_cpu callback function, enqueue callback function, and dispatch callback function.

[0128] In this system, the marking module includes:

[0129] Take the task out of the bpf enqueued map and queue it. If the set time has passed since the last sampling of the task, update the last sampling time point and the accumulated IO amount. The accumulated IO amount is calculated according to the weighted moving average algorithm EWMA. Each time, compare whether the accumulated IO amount exceeds the set IO amount threshold. If so, mark the task as a Direct IO intensive task.

[0130] In this system, the node scheduling module includes:

[0131] Check whether the disk device and NUMA node that the Direct IO intensive task is accessing have been set. If not, read the file descriptor in the / proc / PID / fd directory to find the disk device file, match it with the previously recorded disk device and NUMA node, and record it; at the same time, initialize the scheduling domain of the task to the lowest level, and pass the lowest level cpumask mask to the bpf callback function dispatch through the dispatched map.

[0132] In this system, the adjustment module includes:

[0133] Periodically traverse the utilization of all CPUs within the current scheduling domain of Direct IO intensive tasks, calculate whether the average CPU utilization exceeds the first threshold, and if so, increase the scheduling domain level; if the average CPU utilization is lower than the second threshold, decrease the scheduling domain level; Direct IO intensive tasks dynamically adjust the CPU operation range according to the CPU busy situation. When the scheduling domain is under high load, increase the scheduling domain level to reduce the scheduling delay caused by task queuing; when the CPU in the scheduling domain is idle, decrease the scheduling domain level to improve the cache hit rate of Direct IO intensive tasks.

[0134] In this system, the migration module includes:

[0135] A regular time is set to check the distribution of the private memory of the Direct IO intensive task in each NUMA node. When the time arrives, the / proc / PID / numa_maps file is opened, the private memory pages in the file are parsed, and the size of the private pages in each NUMA node is counted respectively; if it exceeds the set threshold, the C library function migrate_pages() is called to perform memory migration, and the memory migration task is assigned to each thread for processing in a thread pool manner, and the memory is migrated to the disk node being accessed by the Direct IO intensive task.

[0136] The scheduling system for Direct IO intensive tasks under the NUMA architecture proposed in the second embodiment can implement the scheduling method for Direct IO intensive tasks under the NUMA architecture proposed in the first embodiment, and has the same technical effect as the first embodiment.

[0137] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that instructions executed by the processor of a computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 Process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0138] The above-mentioned embodiments are only preferred implementations of the present invention, and are only used to help understand the method and core ideas of the present invention. The protection scope of the present invention is not limited to the above-mentioned embodiments. All technical solutions under the idea of ​​the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A scheduling method for Direct IO intensive tasks under NUMA architecture, characterized in that: include: S1. Create a user-mode scheduler; S2, scan and record the NUMA nodes where each disk device is located, establish a scheduling domain, initialize CPU utilization, create a thread pool, and assign threads to regularly update CPU utilization; S3, implement bpf callback function based on sched_ext extensible scheduler of Linux kernel; S4. Attach the bpf callback function to each callback point of the sched_ext extensible scheduler through the libbpf library, and take over the scheduling strategy of the sched_ext extensible scheduler; S5. Interact with the callback function through bpf map to realize task queuing and dispatching; S6. Track the amount of Direct IO of the task in unit time, accumulate and decay the IO amount, and mark the Direct IO intensive task according to whether the IO amount exceeds the IO amount threshold; S7. Check whether the disk device and NUMA node that the Direct IO intensive task is accessing have been set. If not, read the file descriptor in the / proc / PID / fd directory to find the disk device file, match it with the previously recorded disk device and NUMA node, and record it, and schedule the task in the corresponding NUMA node; initialize the task's scheduling domain to the lowest level, and pass the lowest level cpumask mask to the bpf callback function dispatch through the dispatched map; S8, regularly checking the CPU utilization of the Direct IO intensive task within the current scheduling domain, and dynamically adjusting the scheduling domain level if the utilization exceeds a set threshold; S9, regularly checking the distribution of the private memory of the Direct IO intensive task in each NUMA node, and determining whether to migrate the memory to the NUMA node being accessed according to a set threshold; Step S9 includes: setting a regular time for checking the distribution of the private memory of the Direct IO intensive task in each NUMA node, and after the time is reached, opening the / proc / PID / numa_maps file, parsing the private memory pages in the file, and counting the number of private pages in each NUMA node; If it exceeds the set threshold, the C library function migrate_pages() is called to perform memory migration. The thread pool method is used to assign the memory migration task to each thread to handle it, and the memory is migrated to the NUMA node that the Direct IO intensive task is accessing.

2. The scheduling method for Direct IO intensive tasks under NUMA architecture according to claim 1, characterized in that: Step S2 includes: Open the / sys / class / block directory, read the soft link of each file, find the PCI device file and record the PCI device number; then read / sys / bus / pci / devices / PCI device number / numa_node, and use the global structure variable to record each PCI device and the corresponding NUMA node; Read the / proc / schedstat file, which contains information about the scheduling domain level of each CPU in the system. Parse the file content and record the cpumask mask of each scheduling domain level of the CPU in the global array variable. Read the / proc / stat file, which contains information about the time of each CPU. Parse the file content and record the time of each CPU in the global variable, including the user state time, system state time, idle time, and interrupt time of each CPU. Call the C library function pthread_create() to apply for working threads in advance. Each working thread calls pthread_cond_wait() to block and wait for tasks to be issued. Call pthread_cond_signal() to wake up a working thread, which updates the CPU utilization once a second.

3. The scheduling method for Direct IO intensive tasks under NUMA architecture according to claim 1, characterized in that: Step S3 includes: A bpf program is written based on the sched_ext extensible scheduler of the Linux kernel to implement related bpf callback functions, which include: init callback function, init_task callback function, select_cpu callback function, enqueue callback function, and dispatch callback function.

4. The scheduling method for Direct IO intensive tasks under NUMA architecture according to claim 1, characterized in that: Step S6 includes: Take the task out of the bpf enqueued map and queue it. If the set time has passed since the last sampling of the task, update the last sampling time point and the accumulated IO amount. The accumulated IO amount is calculated according to the weighted moving average algorithm EWMA. Each time, compare whether the accumulated IO amount exceeds the set IO amount threshold. If so, mark the task as a Direct IO intensive task.

5. The scheduling method for Direct IO intensive tasks under NUMA architecture according to claim 1, characterized in that: Step S8 includes: Periodically traverse the utilization of all CPUs within the current scheduling domain of Direct IO intensive tasks, calculate whether the average CPU utilization exceeds the first threshold, and if so, increase the scheduling domain level; if the average CPU utilization is lower than the second threshold, decrease the scheduling domain level; Direct IO intensive tasks dynamically adjust the CPU operation range according to the CPU busy situation. When the scheduling domain is under high load, increase the scheduling domain level to reduce the scheduling delay caused by task queuing; when the CPU in the scheduling domain is idle, decrease the scheduling domain level to improve the cache hit rate of Direct IO intensive tasks.

6. A scheduling system for Direct IO intensive tasks under NUMA architecture, characterized in that: include: Program module: create a user-mode scheduler; Initialization module: scans and records the NUMA nodes where each disk device is located, establishes a scheduling domain, initializes CPU utilization, creates a thread pool, and assigns threads to regularly update CPU utilization; Callback function module: implement bpf callback function based on sched_ext extensible scheduler of Linux kernel; Policy module: attaches the bpf callback function to each callback point of the sched_ext extensible scheduler through the libbpf library, and takes over the scheduling policy of the sched_ext extensible scheduler; Interaction module: interacts with the callback function through bpf map to realize task queuing and dispatching; Marking module: tracks the amount of Direct IO in a task per unit time, accumulates and decays the amount of IO, and marks Direct IO-intensive tasks based on whether the amount of IO exceeds the IO threshold; Node scheduling module: Checks whether the disk device being accessed by the Direct IO intensive task and the NUMA node where it is located have been set. If not, it reads the file descriptor in the / proc / PID / fd directory to find the disk device file, matches it with the previously recorded disk device and NUMA node and records it, and schedules the task in the corresponding NUMA node; initializes the task's scheduling domain to the lowest level, and passes the lowest level cpumask mask to the bpf callback function dispatch through the dispatched map; Adjustment module: regularly checks the CPU utilization of the Direct IO intensive task within the current scheduling domain, and dynamically adjusts the scheduling domain level if the utilization exceeds a set threshold; Migration module: regularly checks the distribution of the private memory of the Direct IO intensive task in each NUMA node, and determines whether the memory needs to be migrated to the NUMA node being accessed according to the set threshold; including: setting a regular time to check the distribution of the private memory of the Direct IO intensive task in each NUMA node, and after the time arrives, opening the / proc / PID / numa_maps file, parsing the private memory pages in the file, and counting the number of private pages in each NUMA node; If it exceeds the set threshold, the C library function migrate_pages() is called to perform memory migration. The thread pool method is used to assign the memory migration task to each thread to handle it, and the memory is migrated to the NUMA node that the Direct IO intensive task is accessing.

7. A computer-readable storage medium storing a computer program, characterized in that: The computer program is used to execute the scheduling method for Direct IO intensive tasks under the NUMA architecture as described in any one of claims 1-5.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the scheduling method for Direct IO intensive tasks under a NUMA architecture is implemented as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • NUMA (Non Uniform Memory Access) perceived load balancing scheduling method and system and medium

    CN116126525A

  • Multi-task scheduling management system under NUMA architecture

    CN116820750A

Cited By

  • Method and system for optimizing direct I / O read performance under Linux system

    CN121979461A

  • Method and system for optimizing direct I / O read performance under Linux system

    CN121979461B