A method, device, product, server and medium for controlling an operation process
By allocating tokens to job processes and combining them with resource usage updates, combined with namespace and control group management, the fairness issue when multiple processes share resources is solved, refined control and efficient utilization of resources are achieved, and load fluctuations in high-performance computing are adapted.
Patent Information
- Application Number
- CN202510919781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
When multiple processes share resources, the fairness of the processes in using the resources is poor, and existing technologies cannot effectively solve the problem of fair resource allocation.
By allocating tokens to each job process to represent resource usage and updating the token count based on resource usage, resource usage by the job process is controlled. Furthermore, namespaces and control groups are created, and job processes are managed through daemons, ensuring isolation between different job processes and refined resource control.
It achieves fairness in resource usage among processes, improves resource utilization, enhances system stability and security, avoids resource waste and performance loss, and adapts to load fluctuations in high-performance computing.
Smart Images

Figure CN120429091B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing clusters, and in particular to a method, device, product, server and medium for controlling an operation process. Background Art
[0002] With the widespread adoption of the Internet of Things (IoT), data volumes are exploding. To meet these massive computing demands, multiple high-performance servers are interconnected via high-speed networks to form a high-performance computing cluster, processing massive amounts of data in parallel at extremely high speeds.
[0003] To control resource usage, related technologies monitor resource usage and restrict resource usage when resource usage is high. If one process occupies a large amount of resources, other processes cannot use the resources, which means that multiple processes cannot share resources fairly.
[0004] Therefore, when multiple processes share resources, how to improve the fairness of processes in using resources is a technical problem that people in this field urgently need to solve. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, device, product, server and medium for controlling a job process, which are used to solve the problem of poor fairness in the use of resources by multiple processes when multiple processes share resources.
[0006] To solve the above technical problems, the present invention provides a method for controlling a job process, which is applied to a computing node. The method comprises:
[0007] Controlling the job process to call a preset file library and obtaining a context corresponding to the job process recorded in the preset file library and used to characterize resource usage; wherein different job processes in the preset file library correspond to different contexts;
[0008] Allocate a token for representing resource usage to the job process based on the context corresponding to the job process;
[0009] Obtain resource usage and update the number of tokens corresponding to the job process based on the resource usage;
[0010] The resource usage of the job process is controlled based on the updated token number.
[0011] On the one hand, before the control operation process calls the preset file library and obtains the context for representing resource usage corresponding to the operation process recorded in the preset file library, the method further includes:
[0012] Receiving a job submitted by a user through a scheduling system; wherein the job submitted by the user includes a resource model and a resource quantity to be used; the resource quantity is an integer or a non-integer;
[0013] After receiving the job submitted by the user, the scheduling system further includes:
[0014] A computing node that meets preset conditions is selected based on the job submitted by the user, and after resources on the computing node are allocated to the job, the resource usage is marked; wherein the preset condition is that the number of resources on the computing node is greater than or equal to the number of resources in the job submitted by the user; resource usage includes full use or partial use.
[0015] On the other hand, after receiving a job submitted by a user through the scheduling system, the control job process calls a preset file library and obtains the context corresponding to the job process recorded in the preset file library for representing resource usage, and further includes:
[0016] Creating a daemon process and having the daemon process create a namespace and a control group for representing and limiting resource usage by a job process;
[0017] Receiving the job process to be started sent by the scheduling system through the daemon process, and controlling the startup of the job process to be started;
[0018] Adding the job process to the namespace and the control group;
[0019] The control operation process calls the preset file library including:
[0020] The job processes in the namespace and the control group are controlled to call a preset file library.
[0021] On the other hand, the controlling operation process calling the preset file library includes:
[0022] When it is detected that the number of resources allocated to the job process is an integer, the directory of the job process is placed in the original preset file library to control the job process to call the preset file library; wherein the original preset file library allows the job process to directly access and manage resources;
[0023] When it is detected that the number of resources allocated to the job process is non-integer, the files in the original preset file library are modified to obtain a new preset file library, and the directory of the job process is placed in the new preset file library to control the job process to call the preset file library; wherein, the new preset file library allows some job processes to directly access and manage resources.
[0024] On the other hand, obtaining resource usage includes:
[0025] Create a monitoring process for characterizing and monitoring the resource usage of each job process;
[0026] Obtaining a resource usage history curve of the job process through the monitoring process;
[0027] The average resource usage is determined based on the resource usage history curve of the job process to serve as the resource usage.
[0028] On the other hand, the resource is a hardware device for parallel computing, and before allocating a token for representing the use of the resource to the job process according to the context corresponding to the job process, the method further includes:
[0029] Obtaining the computing capability of a hardware device for parallel computing and setting a first preset number of tokens in a token bucket according to the computing capability of the hardware device for parallel computing;
[0030] The step of allocating a token for representing resource usage to the job process according to the context corresponding to the job process includes:
[0031] If it is determined based on the context corresponding to the job process that the job process is prohibited from submitting computing tasks to the target resource, the number of tokens allocated to the job process for representing the use of resources is 0;
[0032] If it is determined based on the context corresponding to the job process that the computing task is allowed to be submitted to the target resource, a second preset number of tokens for representing the use of resources is allocated to the job process based on the context corresponding to the job process; wherein the first preset number is greater than the second preset number.
[0033] On the other hand, updating the number of tokens corresponding to the job process based on resource usage includes:
[0034] Determining the current number of tokens corresponding to the current resource utilization rate according to a preset relationship between the utilization rate of the hardware device for parallel computing and the number of tokens;
[0035] Tokens of the current number of tokens are removed from the first preset number of tokens to update the number of tokens corresponding to the job process; wherein the updated number of tokens is the difference obtained by subtracting the current number of tokens from the first preset number.
[0036] On the other hand, controlling the resource usage of the job process based on the updated token number includes:
[0037] When detecting that the updated number of tokens is greater than a first preset value, allowing the job process to submit computing tasks to use resources;
[0038] When it is detected that the updated number of tokens is equal to the first preset value, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
[0039] On the other hand, the control method of the operation process also includes:
[0040] When it is detected that the updated number of tokens is a second preset value, predicting the usage trend of the hardware device for parallel computing by a time series prediction algorithm; wherein the second preset value is greater than the first preset value, and the difference between the second preset value and the first preset value is less than a preset difference;
[0041] If it is detected that the usage trend is an upward trend, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
[0042] On the other hand, the control method of the operation process also includes:
[0043] Pre-create computing tasks to store blocked data;
[0044] After determining that the job process is prohibited from submitting computing tasks to the target resource according to the context corresponding to the job process, or after prohibiting the job process from submitting computing tasks, the prohibited computing tasks are placed in the buffer instruction queue.
[0045] On the other hand, the control method of the operation process also includes:
[0046] Obtaining the status of computing tasks in the buffer instruction queue;
[0047] If it is detected that the computing task occupies the entire buffer instruction queue, the operation process in the central processing unit is suspended from the operating system.
[0048] On the other hand, the control method of the operation process also includes:
[0049] If it is detected that there is a target job process with a resource usage rate greater than a preset usage rate, the context of the target job process is forced to be switched out and the execution of the context of the target job process is suspended;
[0050] The context of controlling other job processes is switched into the execution of the hardware device for parallel computing; wherein the other job processes are job processes other than the target job process in all job processes.
[0051] On the other hand, the context of controlling other job processes to be switched into the execution of the hardware device for parallel computing includes:
[0052] Get the priority order of all other job processes;
[0053] The job process with the highest priority is selected, and the context of the job process with the highest priority is controlled to be switched into the execution of the hardware device for parallel computing.
[0054] On the other hand, the control method of the operation process also includes:
[0055] After detecting that the updated token quantity is equal to 0, if it is detected that the resource usage rate of the job process prohibited from submitting computing tasks decreases, the token is reissued to the job process prohibited from submitting computing tasks;
[0056] When it is detected that the number of tokens is greater than 0, the context execution of the job process that is prohibited from submitting computing tasks is resumed.
[0057] On the other hand, resuming the execution of the context of the job process that is prohibited from submitting the computing task includes:
[0058] If it is detected that the context of the job process that is prohibited from submitting computing tasks is in a suspended state, resuming execution of the context of the job process that is prohibited from submitting computing tasks on the hardware device for parallel computing;
[0059] Alternatively, if it is detected that there is remaining space in the instruction queue, the computing task in the buffer instruction queue is migrated to the instruction queue; wherein the instruction queue is used to store computing tasks that are allowed to be submitted;
[0060] If it is detected that the job process migrated to the instruction queue is in a state where the operation of the job process in the central processing unit is suspended in the operating system, the operation of the job process in the central processing unit is resumed from the operating system.
[0061] On the other hand, the control method of the operation process also includes:
[0062] When the execution of the detection job process ends, the resources occupied by the job process are released;
[0063] The daemon process is controlled to release the namespace and the control group used to represent and restrict the use of resources by the job process.
[0064] In order to solve the above technical problems, the present invention further provides a control device for a job process, which is applied to a computing node, and the control device includes:
[0065] A first control module is configured to control a job process to call a preset file library and obtain a context corresponding to the job process recorded in the preset file library and used to represent resource usage; wherein different job processes in the preset file library correspond to different contexts;
[0066] An allocation module, configured to allocate a token for representing resource usage to a job process according to a context corresponding to the job process;
[0067] The acquisition and update module is used to obtain resource usage and update the number of tokens corresponding to the job process according to the resource usage;
[0068] The second control module is used to control the use of resources by the job process according to the updated token quantity.
[0069] In order to solve the above technical problems, the present invention also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned operation process control method when executed by a processor.
[0070] In order to solve the above technical problems, the present invention further provides a server, comprising:
[0071] memory for storing computer programs;
[0072] The processor is used to implement the steps of the above-mentioned method for controlling the operation process when executing the computer program.
[0073] In order to solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for controlling the operation process are implemented.
[0074] The beneficial effects of the present invention are that, in this method, a token for representing resource use is allocated to each process, and the number of tokens is updated in combination with the resource utilization rate, and the job process uses resources according to the updated number of tokens. Compared with the method in the related art that directly restricts the use of resources when the resource utilization rate is high, the method provided by the present invention ensures that each process is eligible to use resources by allocating a token for representing resource use to each process and using resources according to the token, thereby improving the fairness of resource use; secondly, in this method, the use of resources by the process is controlled by combining the number of tokens updated based on the resource utilization rate, rather than relying solely on the resource utilization rate to limit resource use, thereby achieving refined control of resources; thirdly, before allocating a token for representing resource use to the job process, the job process is controlled to call a preset file library and obtain the context for representing resource use corresponding to the job process recorded in the preset file library. Compared with the method of placing all processes in the same context, in the method provided by the present invention, different job processes correspond to different contexts in the preset file library, ensuring that the job processes do not interfere with each other and achieving process isolation.
[0075] In addition, the computing node receives jobs submitted by users through the scheduling system. The jobs submitted by users include the resource model and resource quantity to be used; the resource quantity is an integer or non-integer. That is, not only can an integer number of resources be requested, but non-integer resources can also be requested and used in this method, thereby improving resource utilization. Moreover, after receiving the job submitted by the user, the scheduling system selects a computing node whose resource quantity is greater than or equal to the resource quantity in the job submitted by the user, thereby ensuring that the job can be processed. After allocating resources on the computing node to the job, the scheduling system marks the resource usage, such as full use or partial use, so that the resource usage can be intuitively understood.
[0076] A daemon process is created, and the daemon creates a namespace and a control group that represents resource restrictions for job processes. The daemon receives pending job processes from the scheduling system and controls their launch. Job processes are added to the namespace and control group, and the job processes in the namespace and control group are controlled to call pre-set file repositories. Because the namespace provides an isolation mechanism, the control group that represents resource restrictions for job processes can also limit the resources used by the process. This not only limits resources but also enables collaborative isolation of different resources, forming a complete job isolation system. Furthermore, unified management through the daemon enhances the stability and security of multi-resource environments.
[0077] If a job process is detected to have an integer number of allocated resources, it is directly controlled to access the original preset file repository (allowing the job process to directly access and manage resources). If a job process is detected to have a non-integer number of allocated resources, the job process's directory is moved to a new preset file repository to control access to the preset file repository. The new preset file repository allows some job processes to directly access and manage resources. This method distinguishes between integer and non-integer resource allocation scenarios. For jobs that exclusively use all resources, the original preset file repository is used to avoid the additional performance loss caused by hijacking. For jobs that share resources, the modified preset file repository is used to implement resource restrictions. This "on-demand hijacking" design maximizes performance while ensuring isolation. For compute-intensive jobs (such as deep learning training), exclusive resource allocation eliminates the additional overhead, while still achieving efficient isolation in shared scenarios, balancing performance and fairness.
[0078] Create a monitoring process to monitor the resource usage of each job process. Obtain a historical resource usage curve for each job process through the monitoring process. Determine the average resource usage based on the historical resource usage curve for each job process, and use it as the resource usage rate. This ensures that the obtained resource usage rate is useful as a reference.
[0079] A first preset number of tokens is set in a token bucket based on the computing capabilities of the hardware device used for parallel computing, ensuring that the number of tokens placed in the token bucket is appropriate. Furthermore, if the context corresponding to the job process determines that submission of the computing task to the target resource is permitted, a second preset number of tokens representing resource usage is allocated to the job process based on the context corresponding to the job process. Because the first preset number is greater than the second preset number, excessive load on the hardware device used for parallel computing is avoided, thereby improving system reliability.
[0080] When updating the number of tokens corresponding to a job process based on resource usage, the current number of tokens corresponding to the current resource usage is determined based on a pre-set relationship between the usage of the hardware device used for parallel computing and the number of tokens. Tokens corresponding to the current number of tokens are removed from a first pre-set number of tokens to update the number of tokens corresponding to the job process. The updated number of tokens is the difference between the first pre-set number and the current number of tokens. This ensures accurate determination of the number of removed tokens, allowing the job process to use resources based on a more accurate number of tokens, thereby achieving refined control over resource usage.
[0081] When it is detected that the number of updated tokens is greater than the first preset value, the job process is allowed to submit computing tasks. When it is detected that the number of updated tokens is equal to the first preset value, the job process is prohibited from submitting computing tasks, thereby avoiding the job process occupying more resources.
[0082] When it is detected that the updated number of tokens is at a second preset value, a time series prediction algorithm is used to predict the usage trend of the hardware device used for parallel computing; wherein the second preset value is greater than the first preset value, and the difference between the second preset value and the first preset value is less than the preset difference; if an upward trend is detected, the job process is prohibited from submitting computing tasks, thereby prohibiting resource use. This method uses a time series prediction algorithm to predict resource usage trends, and when the number of tokens reaches the second preset value, prohibits the job process from submitting computing tasks, thereby achieving early blocking of the process's use of resources.
[0083] After determining, based on the context of the job process, that a job process is prohibited from submitting computing tasks to the target resource, or after prohibiting a job process from submitting computing tasks, the prohibited computing tasks are placed in a buffered instruction queue. This buffered instruction queue allows for the storage of blocked job processes. If it is detected that a computing task occupies the entire buffered instruction queue, the operating system suspends the execution of the job process on the CPU. This suspends the process, ensuring that no new instructions are submitted to the resource for execution, thereby preventing the buffered instruction queue from overflowing.
[0084] If a target job process is detected whose resource usage exceeds a preset rate, the target job process's context is forced out and its execution is suspended. Instead, the contexts of other job processes are controlled to be switched into the execution of the hardware devices used for parallel computing. This prevents further increases in resource usage by the target job process, while ensuring that other job processes can use resources, thus achieving dynamic and flexible resource allocation. Furthermore, the highest-priority job process is selected and its context is controlled to be switched into the execution of the hardware devices used for parallel computing. This ensures that high-priority job processes can be executed promptly.
[0085] After detecting that the updated number of tokens equals 0, if the resource usage of the job process prohibited from submitting computing tasks is detected to have decreased, tokens are reissued to the job process prohibited from submitting computing tasks; if the number of tokens is detected to be greater than 0, the context execution of the job process prohibited from submitting computing tasks is resumed. This method provides a system that replenishes tokens in a timely manner, and the previous method of removing tokens is used to adjust the resource usage of the process. This dynamic nature adapts to common load fluctuations in high-performance computing (such as iterative changes in artificial intelligence (AI) model training), prevents resource overuse, and quickly resumes job execution after resource release, improving the system's responsiveness.
[0086] When a job process is detected to have finished executing, the resources occupied by the job process are released. The control daemon process releases the namespace and the control group used to represent and restrict the resource usage of the job process, ensuring the integrity of resource management.
[0087] In addition, the present invention also provides a control device for an operation process, a computer program product, a server, and a computer-readable storage medium, which have the same or corresponding technical features as the above-mentioned control method for the operation process and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0089] Figure 1 A flowchart of a method for controlling a job process provided by an embodiment of the present invention;
[0090] Figure 2 A schematic diagram of resource usage by a job process provided by an embodiment of the present invention;
[0091] Figure 3 An overall schematic diagram of a GPU sharing and restriction solution based on CUDA hijacking in a high-performance computing scenario provided by an embodiment of the present invention;
[0092] Figure 4 A structural diagram of a server provided in another embodiment of the present invention. DETAILED DESCRIPTION
[0093] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0094] The core of the present invention is to provide a method, device, product, server and medium for controlling an operation process, which are used to solve the problem of poor fairness in the use of resources by multiple processes when multiple processes share resources.
[0095] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods. Figure 1 The present invention provides a flowchart of a method for controlling a job process. The method is applied to a computing node. For example, the computing node is a server including a graphics processing unit (GPU). Figure 1 As shown, the method includes:
[0096] S10: Controlling the job process to call a preset file library and obtain a context corresponding to the job process recorded in the preset file library, which is used to characterize resource usage; wherein different job processes in the preset file library correspond to different contexts;
[0097] S11: Allocate a token for representing resource usage to the job process according to the context corresponding to the job process;
[0098] S12: Obtain resource usage and update the number of tokens corresponding to the job process according to the resource usage;
[0099] S13: Control the resource usage of the job process according to the updated token quantity.
[0100] In this context, a job refers to a computing task performed in a high-performance computing environment, typically involving large-scale data processing, complex numerical simulations, or parallel computing. It leverages the computing resources of a high-performance computing cluster (including multi-core central processing units (CPUs), GPU accelerators, or other specialized hardware) to solve complex computational problems in science, engineering, data analysis, and other fields.
[0101] In order to obtain the job, before controlling the job process to call the preset file library and obtaining the context corresponding to the job process recorded in the preset file library and used to represent the resource usage, the following is also included:
[0102] The scheduling system receives jobs submitted by users. The jobs submitted by users include resource models and resource quantities to be used. The resource quantities are integers or non-integers.
[0103] In a method for controlling resource usage by job processes using MIG (Multi-Instance GPU), hardware-level multi-instance technology is employed to partition a single resource (e.g., a GPU) into up to seven independent instances, each with independent video memory and computing resources, supporting multi-user or multi-job sharing. MIG achieves isolation through hardware partitioning, with a fixed granularity (preset instance size, such as 1 / 7 or 2 / 7), and cannot support flexible allocation of non-integer values (e.g., 0.5 GPUs). This results in resource waste for job processes that do not require the entire resource. To reduce resource waste, the scheduling system of the present invention, after receiving a job submitted by a user, further includes:
[0104] Based on the job submitted by the user, a computing node that meets the preset conditions is selected. After allocating resources on the computing node to the job, the resource usage is marked. The preset condition is that the number of resources on the computing node is greater than or equal to the number of resources in the job submitted by the user. The resource usage includes full use or partial use.
[0105] In this method, after receiving a job submitted by a user, the scheduling system selects a compute node with a resource count equal to or greater than the resource count in the job, ensuring that the job can be processed. After allocating resources on a compute node to the job, the scheduling system also marks the resource usage, such as fully or partially used, to provide a visual overview of resource usage.
[0106] To avoid interference between multiple job processes, in implementation, after receiving a job submitted by a user through the scheduling system, before controlling the job process to call the preset file library and obtain the context corresponding to the job process recorded in the preset file library and used to represent resource usage, the following steps are also included:
[0107] Create a daemon process and use it to create a namespace (such as Namespace) and a control group (such as Cgroup) used to represent and limit the resources used by the job process;
[0108] Receive the job process to be started sent by the scheduling system through the daemon process, and control the startup of the job process to be started;
[0109] Add the job process to the namespace and control group.
[0110] Controlling a job process to call a preset file library includes: controlling a job process in a namespace and a control group to call a preset file library.
[0111] Namespaces are an isolation mechanism used to partition system resources (such as process IDs, network interfaces, and file system mount points) into independent, non-interfering logical areas. Each namespace maintains its own view of resources, allowing resources in different namespaces to have the same identifier (such as the same process ID or network IP address) without conflict. Cgroups are a kernel feature that allow the operating system to manage system resources by controlling and limiting resource usage by process groups.
[0112] In this method, a daemon process is created, which then creates a namespace and a control group used to represent resource restrictions on job processes. The daemon process receives job processes to be started from the scheduling system and controls their startup. The job processes are added to the namespace and control group, and the job processes in the namespace and control group are controlled to call a preset file repository. Because the namespace provides an isolation mechanism, the control group used to represent resource restrictions on job processes can limit the resources used by the processes. This not only limits resources but also enables collaborative isolation of different resources, forming a complete job isolation system. Furthermore, unified management through the daemon process enhances the stability and security of multi-resource environments.
[0113] There is no limit on the number of job processes, which is determined based on actual conditions. The job process calls a preset file library. The preset file library records the context used to characterize the use of resources corresponding to the job process. Specifically, the preset file library is the Compute Unified Device Architecture (CUDA) library. CUDA is a general-purpose parallel computing architecture that enables GPUs to solve complex computing problems. It includes the CUDA instruction set architecture (ISA) and the parallel computing engine inside the GPU. The context used to characterize the use of resources recorded in the preset file library is, for example, the GPU context. The GPU context is an abstract concept used to manage and isolate GPU resources and execution environments. It provides an independent execution space for each program running on the GPU, ensuring that different programs do not interfere with each other. The context saves the operating state of the GPU, including registers, memory mapping, and configuration parameters.
[0114] In a method using MPS (Multi-Process Service) to control resource usage by job processes, MPS allows multiple processes to share the same GPU, centrally managing kernel submission and video memory allocation through a server-client model, providing certain resource isolation and scheduling capabilities. In client mode, all job processes are placed in the same context. A context error in one process can also cause errors in the contexts of other processes. Therefore, the present invention utilizes a preset file library (such as a CUDA library) that records different contexts corresponding to different job processes. Each job process's context is independent, rather than all job processes sharing the same context, thus achieving isolation of the job processes.
[0115] In order to allocate integer or non-integer resources, in implementation, the control job process calls the preset file library including:
[0116] When it is detected that the number of resources allocated to the job process is an integer, the directory of the job process is placed in the original preset file library to control the job process to call the preset file library; wherein the original preset file library allows the job process to directly access and manage resources;
[0117] When it is detected that the number of resources allocated to the job process is non-integer, the files in the original preset file library are modified to obtain a new preset file library, and the directory of the job process is placed in the new preset file library to control the job process to call the preset file library; wherein, the new preset file library allows some job processes to directly access and manage resources.
[0118] This method distinguishes between integer and non-integer resource allocation scenarios. For jobs that exclusively use all resources, the original preset file library is used, avoiding the additional performance loss caused by hijacking. For jobs that share resources, the modified preset file library is used to implement resource restrictions. This "on-demand hijacking" design maximizes performance while ensuring isolation. For compute-intensive jobs (such as deep learning training), exclusive resource ownership eliminates the additional overhead, while still achieving efficient isolation in shared scenarios, balancing performance and fairness.
[0119] After obtaining the resource usage context corresponding to the job process recorded in the preset file library, a token representing resource usage is allocated to the job process based on the job process context. To ensure an appropriate number of allocated tokens, in implementation, the resource is a hardware device for parallel computing (such as a GPU). Before allocating the resource usage token based on the job process context, the following steps are also performed:
[0120] The computing capability of the hardware device used for parallel computing is obtained, and a first preset number of tokens is set in a token bucket based on the computing capability of the hardware device used for parallel computing. The first preset number is not limited. For example, if the computing capability of a GPU is 100, 100 tokens (denoted as tokens) are set in the corresponding token bucket.
[0121] The tokens allocated to the job process based on the context of the job process and used to represent the resources used include:
[0122] If it is determined based on the context corresponding to the job process that the job process is prohibited from submitting computing tasks to the target resource, the number of tokens allocated to the job process for representing the use of resources is 0;
[0123] If it is determined based on the context corresponding to the job process that the computing task is allowed to be submitted to the target resource, a second preset number of tokens for representing the use of resources is allocated to the job process based on the context corresponding to the job process; wherein the first preset number is greater than the second preset number.
[0124] In this method, since the first preset number is greater than the second preset number, excessive load on the hardware device used for parallel computing is avoided, thereby improving the reliability of the system.
[0125] After assigning tokens to job processes to represent resource usage, the number of tokens can be further updated based on resource usage to prevent some job processes from constantly occupying resources. First, resource usage needs to be obtained. In practice, obtaining resource usage includes:
[0126] Create a monitoring process for characterizing and monitoring the resource usage of each job process;
[0127] Obtain the resource usage history curve of the job process through the monitoring process;
[0128] The average resource usage is determined based on the resource usage history curve of the job process to serve as the resource usage.
[0129] For example, a monitoring process is started on a compute node. This process uses the NVML library to monitor the GPU usage of each GPU job process. A monitoring process must be run on each compute node that runs a GPU job. The monitoring process operates in an endless loop, continuously detecting and recording GPU usage. This generates a historical GPU usage curve for each process, which can then be used to calculate the average GPU usage.
[0130] In implementation, updating the number of tokens corresponding to the job process according to resource usage includes:
[0131] Determining the current number of tokens corresponding to the current resource utilization rate according to a preset relationship between the utilization rate of the hardware device for parallel computing and the number of tokens;
[0132] Tokens of the current number of tokens are removed from the first preset number of tokens to update the number of tokens corresponding to the job process; wherein the updated number of tokens is the difference obtained by subtracting the current number of tokens from the first preset number.
[0133] If the average GPU usage of a job process is 20%, the tokens used by the job are calculated as 100 * 20% = 20 according to the aforementioned relationship between GPU usage and tokens. Since the number of tokens initially allocated to the job process is 50, after 20 tokens are used, the number of tokens remaining is 30. The number of tokens remaining available for the process in the token bucket is updated to 30.
[0134] After obtaining the updated token quantity, controlling the resource usage of the job process according to the updated token quantity includes:
[0135] When detecting that the updated number of tokens is greater than a first preset value, allowing the job process to submit computing tasks to use resources;
[0136] When it is detected that the updated number of tokens is equal to the first preset value, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
[0137] The corresponding first preset value is not limited, such as 0. That is, when it is detected that the updated number of tokens is greater than 0, the job process is allowed to submit computing tasks to use resources; when it is detected that the updated number of tokens is equal to 0, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
[0138] Since the scheme of limiting the resource usage of job processes through token buckets has a certain lag, it is possible that the resource usage of a job is high for a short period of time, resulting in a negative token in the token bucket. To avoid this situation, the control method of the job process also includes:
[0139] When it is detected that the updated number of tokens is a second preset value, predicting the usage trend of the hardware device for parallel computing by a time series prediction algorithm; wherein the second preset value is greater than the first preset value, and the difference between the second preset value and the first preset value is less than the preset difference;
[0140] If it is detected that the usage trend is an upward trend, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
[0141] There is no limitation on the second preset value and the preset difference. If the first preset value is 0, the second preset value is a value close to 0.
[0142] The method adopts a time series prediction algorithm to predict the resource usage trend, and when the number of tokens reaches a second preset value, the job process is prohibited from submitting computing tasks, thereby achieving early blocking of the process's use of resources.
[0143] The computing task may be any parallel computing task such as matrix multiplication, convolution operation, vector addition, etc. The computing task is also called the Kernel function. The Kernel function is one of the core concepts of CUDA programming and is a function that runs specifically on the GPU. It is written by the programmer using CUDA syntax and accelerates computing-intensive tasks by executing a large number of threads in parallel. Unlike ordinary C / C++ functions, the Kernel function needs to be declared using the _global_ keyword, called by the host (CPU) code, and executed on the device (GPU). Its design purpose is to make full use of the parallel computing capabilities of the GPU to achieve efficient data processing. In order to be able to continue processing the blocked computing tasks in the future, in implementation, the control method of the job process also includes:
[0144] Pre-create computing tasks to store blocked data;
[0145] After determining that the job process is prohibited from submitting computing tasks to the target resource according to the context corresponding to the job process, or after prohibiting the job process from submitting computing tasks, the prohibited computing tasks are placed in the buffer instruction queue.
[0146] Furthermore, to prevent the buffered instruction queue from overflowing, in some embodiments, the method for controlling the job process further includes:
[0147] Get the status of computing tasks in the buffer instruction queue;
[0148] If it is detected that the computing task occupies the entire buffer instruction queue, the operating system will suspend the operation of the job process in the central processing unit.
[0149] That is, this method suspends the process, ensuring that no new instructions are submitted to the resource for execution, thereby ensuring that the buffered instruction queue does not overflow.
[0150] To prevent some job processes from using relatively high resources and further increasing resource usage if they are still executed on hardware devices used for parallel computing, and to ensure that other job processes can use resources, in some embodiments, the method for controlling job processes further includes:
[0151] If it is detected that there is a target job process with a resource usage rate greater than a preset usage rate, the context of the target job process is forced to be switched out and the execution of the context of the target job process is suspended;
[0152] The context of controlling other job processes is switched into the execution of the hardware device for parallel computing; wherein other job processes are job processes other than the target job process among all job processes.
[0153] There is no limit on the preset usage rate and other selected operation processes, which will be determined based on actual conditions.
[0154] In implementation, the context of controlling other job processes to be switched into the execution of the hardware device for parallel computing includes:
[0155] Get the priority order of all other job processes;
[0156] A job process with the highest priority is selected, and the context of the job process with the highest priority is controlled to enter the execution of the hardware device for parallel computing.
[0157] In practice, other scheduling algorithms can also be used to select the job process to use. This method prevents the target job process from further increasing its resource usage while ensuring that other jobs can use resources, thus achieving dynamic and flexible resource allocation. Furthermore, by selecting the highest-priority job process and controlling its contextual entry into the parallel computing hardware, this ensures that the high-priority job process is executed promptly.
[0158] In practice, load fluctuations may occur. To meet the needs of the job process, in some embodiments, the job process control method further includes:
[0159] After detecting that the updated token quantity is equal to 0, if it is detected that the resource usage rate of the job process prohibited from submitting computing tasks decreases, the token is reissued to the job process prohibited from submitting computing tasks;
[0160] When it is detected that the number of tokens is greater than 0, the context execution of the job process that is prohibited from submitting computing tasks is resumed.
[0161] Specifically, resuming the context execution of the job process that is prohibited from submitting computing tasks includes:
[0162] If it is detected that the context of the job process prohibited from submitting computing tasks is in a suspended state, then resuming execution of the context of the job process prohibited from submitting computing tasks on the hardware device used for parallel computing;
[0163] Alternatively, if it is detected that there is remaining space in the instruction queue, the computing task in the buffer instruction queue is migrated to the instruction queue; wherein the instruction queue is used to store computing tasks that are allowed to be submitted;
[0164] If it is detected that the job process migrated to the instruction queue is in a state where the operation of the job process in the central processing unit is suspended in the operating system, the operation of the job process in the central processing unit is resumed from the operating system.
[0165] This method provides a system that replenishes tokens in a timely manner, and the previous method of removing tokens allows adjustments to be made to the resource usage of the process. This dynamic nature adapts to common load fluctuations in high-performance computing (such as iterative changes in AI model training), preventing resource overuse and quickly resuming job execution after resource release, thereby improving the system's responsiveness.
[0166] To ensure the integrity of resource management, the control method of the operation process also includes:
[0167] When the execution of the detection job process ends, the resources occupied by the job process are released;
[0168] The control daemon releases the namespace and the control group used to represent and restrict the resources used by the job process.
[0169] In the control method of the operation process provided by the present invention:
[0170] 1) Support flexibility in non-integer resource allocation.
[0171] Taking GPU resources as an example, this invention supports user-specified allocation of non-integer numbers of GPU resources (such as 0.5 or 2.5 GPUs), breaking the limitations of traditional allocation of entire GPUs. By marking partially used GPUs in the scheduling system and dynamically allocating them, resource utilization is significantly improved to accommodate diverse workload demands.
[0172] Compared with the fixed instance partitioning of MIG or the static quota of MPS, the present invention provides more fine-grained and flexible resource partitioning.
[0173] 2) Quantitative combination of token bucket algorithm and resource utilization.
[0174] The token bucket algorithm is introduced, which uses GPU utilization as a quantitative indicator of computing resource consumption and dynamically limits the GPU usage of jobs through token allocation, deduction, and replenishment.
[0175] Proportional allocation and real-time adjustment of computing resources are achieved to ensure fairness and resource utilization efficiency.
[0176] Different from the coarse-grained restrictions of traditional Cgroups, the present invention achieves fine control of GPU computing capabilities.
[0177] 3) Comprehensive design of hysteresis compensation mechanism.
[0178] To address the lag in token bucket control (such as short-term overuse resulting in negative tokens), a compensation mechanism combining timing prediction, buffered instruction queue and GPU context switching is designed.
[0179] By predicting blocking in advance, buffering to avoid hanging, and switching forced intervention, the accuracy and smoothness of resource control are significantly improved.
[0180] 4) Dual-mode strategy of on-demand hijacking and performance optimization.
[0181] Depending on the number of GPUs allocated to the job (integer or non-integer), the original CUDA library or the modified CUDA library is selectively used to avoid unnecessary hijacking overhead.
[0182] It maintains optimal performance in exclusive GPU scenarios and implements strict isolation in shared scenarios, taking into account both performance and isolation requirements.
[0183] Compared with the unified processing method of MPS or MIG, the dual-mode design of the present invention is smarter and more efficient.
[0184] 5) Systematic implementation of collaborative isolation of multiple resources.
[0185] Combining Linux Cgroup and namespaces not only limits GPU computing resources, but also collaboratively isolates CPU, memory and other resources to form a complete job isolation system.
[0186] Unified management through daemon processes enhances the stability and security of multi-resource environments.
[0187] While existing technologies mostly focus on a single resource (such as GPU or CPU), the present invention achieves systematic management of multi-dimensional resources.
[0188] In order to enable those skilled in the art to better understand the above method, the entire process will be described below in conjunction with the accompanying drawings and specific embodiments. Figure 2 A schematic diagram of resource usage by a job process provided by an embodiment of the present invention. Figure 2 In the example, the resource is a graphics processing unit, and two processes are process one and process two, respectively, to illustrate the process of using the graphics processing unit by a process. Process one is located in namespace one, and process two is located in namespace two. First, the Compute Unified Device Architecture library (CUDA library) is called, and then after being driven by the graphics processing unit driver, the graphics processing unit hardware is used to process the process. In the method provided by the present invention, the refined allocation and dynamic management of GPU computing resources are achieved through software-level CUDA call hijacking, combined with a token bucket mechanism and dynamic monitoring. For non-integer GPU allocation scenarios, Kernel submission is hijacked to limit computing power; for exclusive GPU scenarios, the native CUDA library is used to avoid performance loss. The solution uses GPU utilization as a resource consumption indicator, quantifies the allocation ratio through token buckets, and introduces a prediction and compensation mechanism to solve control lag, thereby ensuring isolation, fairness and high utilization when multiple jobs share the GPU.
[0189] Figure 3 The overall schematic diagram of a GPU sharing and restriction solution based on CUDA hijacking in a high-performance computing scenario provided by an embodiment of the present invention is as follows: Figure 3 As shown, a user submits a job (a user submits a GPU job, specifying the GPU model and quantity); the scheduling system allocates compute nodes and marks the GPU resource status; in the compute node layer, the compute node daemon receives scheduling requests, creates a namespace and control group, and starts the job process. The job process runs the user job and loads the Compute Unified Device Architecture library (original or modified version); in the Compute Unified Device Architecture library hijacking and control layer, the Compute Unified Device Architecture library intercepts calls and hijacks the submission of compute task functions (kernel functions), dynamically controls GPU resources, and quantifies GPU resources using token buckets. Furthermore, in the monitoring and optimization layer, the monitoring process uses the NVML library to monitor GPU utilization; the lag compensation layer performs lag compensation through timing prediction, buffered instruction queues, and context switching; finally, the compute task function is executed in the GPU hardware and the calculation results are output.
[0190] The present invention provides a sharing and restriction scheme based on CUDA hijacking in high-performance computing scenarios, which aims to achieve flexible allocation and strict isolation of GPU resources. Users can submit jobs and specify the required integer or non-integer number of GPUs. The scheduling system allocates GPUs according to the resource status and marks the usage. In order to limit the job's use of GPU, CPU and memory resources, the scheme isolates processes through Linux Cgroup and namespace, and uses the modified CUDA library to hijack Kernel function calls, combined with the token bucket algorithm to dynamically control GPU computing resources. The monitoring process detects GPU usage in real time, adjusts the number of tokens, blocks or pauses job execution when tokens are exhausted, optimizes resource contention through buffered instruction queues and context switching, and achieves efficient sharing and isolation.
[0191] In practice, the specific implementation process of the GPU sharing and restriction solution based on CUDA hijacking in high-performance computing scenarios is as follows:
[0192] 1. When a user submits a GPU job, they need to specify the GPU model and quantity. The quantity can be an integer or a non-integer. For example, you can specify that you need 0.5 or 2.5 GPUs of a certain model.
[0193] 2. The scheduling system performs scheduling and finds compute nodes that meet the requirements. Each time a job is scheduled and allocated GPU resources, the scheduling system marks the allocated GPU as (fully or partially) used. Fully used means the entire GPU is allocated to a job, while partially used means a portion of the GPU is allocated to a job (for example, 0.5 of a GPU is allocated to a job). A fully used GPU cannot be used by other jobs, while a partially used GPU can be allocated to other jobs (if sufficient resources are available).
[0194] 3. To more strictly limit the resource usage of the job process (not only GPU resource usage but also CPU, memory, and other resources), use the Cgroup group provided by the Linux system to limit it. The specific process is as follows:
[0195] 1) A daemon process runs on each compute node, called the compute node daemon. When the scheduling system schedules a job to run, it sends a request to the compute node daemon, providing it with information about the job process to be started, and instructs it to start the job process.
[0196] 2) The compute node daemon receives the request, creates a command space and Cgroup group on the compute node, and sets limits on resources such as CPU and memory for the Cgroup group.
[0197] 3) Start the user job process and add the job process to the namespace and Cgroup group created previously.
[0198] 4) Switch the job process's root directory to a specified directory containing the modified CUDA library. This directory will then use the modified CUDA library when the job process makes CUDA calls, thereby intercepting the user job process's calls to the CUDA library. This distinguishes whether the number of allocated GPUs is an integer or a non-integer number. In some cases, users have high demand for GPU computing resources and need to allocate an entire GPU. In this case, because the job process has exclusive access to the GPU and does not need to share GPU resources with other processes, there is no need to consider GPU resource isolation and restriction, and CUDA hijacking is unnecessary. Instead of placing the original CUDA library in the process directory, using the modified CUDA library can avoid the performance loss caused by CUDA hijacking. The modified CUDA library is only required when the number of GPUs allocated to the job process is a non-integer number.
[0199] 4. GPU jobs use the GPU by submitting one computing task after another to the GPU. These tasks may be any parallel computing tasks such as matrix multiplication, convolution operations, vector addition, etc. The computing tasks are also known as kernel functions. The process of submitting computing tasks is called kernel launch.
[0200] 5. Start a process (hereafter referred to as the monitoring process) on the compute node. This process uses NVIDIA's NVML library to monitor the GPU usage of each GPU job process. A monitoring process must run on each compute node running a GPU job. The monitoring process is an endless loop that continuously checks and records GPU usage, generating a historical GPU usage curve for each process, which can then be used to calculate the average GPU usage.
[0201] 6. Token bucket control scheme for GPU computing resource usage of job processes.
[0202] The monitored GPU usage represents the usage of the GPU's computing resources. The higher the GPU usage, the more computing resources are being used. We designed a scheme that controls the GPU usage of job processes through a token bucket. Assume that the computing power of a GPU is 100 (this setting is set to 100 for simplicity of calculation), which corresponds to 100 tokens in the token bucket. Initially, a certain number of tokens are allocated to the user's GPU job process. For example, allocating 50 tokens to a job process means that 50% of the GPU's computing resources can be used. As the job process uses the GPU, the GPU usage increases. When the GPU usage is detected to be increased, the corresponding tokens are deducted from the token bucket. When the tokens are reduced to 0, the job process is not allowed to continue submitting kernel functions to the GPU hardware for execution.
[0203] 7. Token calculation method. The following is a specific example to illustrate.
[0204] Assume that the computing power of a GPU corresponds to 100 tokens in the token bucket, and the number of tokens allocated to a GPU job process is 50.
[0205] 1) Time point 1: The initial start of the job process.
[0206] The average GPU usage of the job process is 20%. According to the relationship between GPU usage and tokens, the tokens used by the job are calculated as 100 × 20% = 20. Since the number of tokens initially allocated to the job process is 50, after using 20, the number of tokens remaining is 30. The number of tokens remaining available for the process in the token bucket is updated to 30.
[0207] If the job process continues to submit the Kernal function at this time, because the number of remaining available tokens in the token bucket is greater than 0, that is, there are still available tokens for the job process, so submission is allowed.
[0208] 2) Time point 2: Load increases.
[0209] The average GPU utilization of this job process reaches 50%. Based on the aforementioned relationship between GPU utilization and tokens, the tokens used by the job are calculated as 100 × 50% = 50. Since the initial number of tokens allocated to the job process is 50, after 50 tokens are used, the remaining number of tokens is 0. Therefore, the number of available tokens in the token bucket for this process is updated to 0. If the job process continues to submit the kernel function at this time, because the number of available tokens in the token bucket is equal to 0, there are no available tokens for this job process. Therefore, the kernel function submission is blocked and the kernel function is submitted only when there are available tokens in the token bucket.
[0210] 8. Because it is difficult to accurately estimate the GPU computing power consumed by each kernel function submitted by a job process, the solution of limiting the GPU usage of a job process through a token bucket has a certain lag. It is possible that the GPU usage rate of a job is high for a short period of time, resulting in a negative token in the token bucket. To compensate for this problem, the following methods are used:
[0211] 1) Method 1: When the token bucket is less than or equal to 0, the job process is blocked from submitting kernel functions, further increasing GPU utilization. For more precise control, a time series prediction algorithm can be used to predict the future GPU utilization of the process based on its historical GPU utilization. If the token bucket drops close to 0 and GPU utilization is predicted to increase, kernel function submission can be blocked. In other words, kernel function submission can be blocked in advance when token depletion is predicted.
[0212] Kernel functions are submitted by submitting instructions to the instruction queue. The GPU hardware will then retrieve instructions from the instruction queue for execution. When a kernel function submits an instruction, if the kernel function submission should be blocked at this time, the instruction should not be placed directly into the instruction queue. Instead, the instruction can be temporarily stored in a buffer memory area (called the buffer instruction queue). Because the buffer instruction queue has a limited space, if a process continuously submits kernel functions, the buffer instruction queue may run out of space. In this case, the process can only be suspended, that is, the operating system suspends the process from running on the CPU (to distinguish it from the GPU, this state is called the CPU pause state). When the process is suspended, no new instructions will be submitted to the GPU for execution, thus ensuring that the buffer instruction queue does not overflow.
[0213] The purpose of using a buffered instruction queue is to minimize the performance loss caused by directly suspending the process.
[0214] 2) Method 2: When the monitoring process detects excessive GPU usage for a process, it forces a GPU context switch. This process's context is switched out, and the GPU context of another process is switched in for execution (this can be done by selecting the GPU context of a higher-priority process or using other scheduling algorithms). The GPU context of the process is then suspended, meaning that the context will not be scheduled for further execution on the GPU hardware until sufficient tokens are available in the token bucket. This method prevents GPU context execution from continuing on the GPU hardware, thus preventing further increases in GPU usage.
[0215] 3) You can combine Method 1 and Method 2. Method 1 focuses on controlling the process in advance when the GPU usage of the process is about to reach the upper limit; Method 2 focuses on controlling the process after the GPU usage of the process has reached or exceeded the upper limit.
[0216] 9. When the process uses GPU resources using the above method, the GPU utilization rate of the process will decrease. The monitoring process will update the number of tokens in the token bucket after monitoring the change in GPU utilization rate. When the number of tokens is greater than 0, the context execution of the process can be restored. The restoration process is as follows:
[0217] 1) If the GPU context of the process has been suspended, resume the GPU context and execute it on the GPU hardware.
[0218] 2) If there are instructions to be executed in the buffered instruction queue, wait for the execution in the instruction queue. Because the GPU context has been restored and executed on the GPU hardware, the instructions in the instruction queue will be gradually consumed. When there is enough space in the instruction queue, the instructions in the buffered instruction queue will be moved into the instruction queue. If the process is still in the CPU pause state at this time, it will be resumed and allowed to submit the kernel function.
[0219] 10. After the user job process ends, the relevant resources are released, and the computing node daemon process deletes the previously created namespace and Cgroup group.
[0220] The GPU isolation solution based on CUDA hijacking provided by this invention is indeed an innovative approach to high-performance computing resource management, particularly unique in terms of GPU resource sharing and refined control. The following are the potential benefits of this technical solution:
[0221] 1. Improve GPU resource utilization.
[0222] Effect Description: By supporting allocation of non-integer numbers of GPUs (such as 0.5 or 2.5 GPUs), this solution allows for more flexible allocation of GPU computing resources, avoiding the resource idleness problem caused by traditional "whole GPU allocation." For example, a job that only requires 50% of the computing power does not need to monopolize the entire GPU; the remaining resources can be allocated to other jobs.
[0223] In multi-user or multi-tasking high-performance computing clusters, it can significantly improve GPU utilization and reduce resource waste. It is especially suitable for small-scale job scenarios with light loads but large numbers.
[0224] 2. Enhance operation isolation and stability.
[0225] By limiting CPU, memory, and other resources through Cgroups and combining CUDA hijacking and token bucket mechanisms to finely control GPU computing resources, this solution ensures that job processes do not interfere with each other. Even if a job attempts to exceed GPU resource usage, it will be restricted to the allocated range.
[0226] This prevents a job from over-occupying the GPU and causing other jobs to stall, improving the overall stability of the cluster. It is particularly suitable for high-performance computing clusters or cloud GPU sharing scenarios.
[0227] 3. Flexibility to reduce performance overhead.
[0228] The solution distinguishes between integer and non-integer GPU allocation scenarios. For jobs that exclusively use the entire GPU, the original CUDA library is used to avoid the additional performance loss caused by hijacking. For jobs that share the GPU, the modified CUDA library is used to implement resource restrictions.
[0229] This "on-demand hijacking" design maximizes performance while ensuring isolation. For compute-intensive tasks (such as deep learning training), exclusive GPU use eliminates the overhead, while still achieving efficient isolation in shared scenarios, balancing performance and fairness.
[0230] 4. Adaptability of dynamic resource management.
[0231] By monitoring processes and detecting GPU utilization in real time, combined with a token bucket and context switching mechanism, the solution dynamically adjusts resource allocation. When workloads change (for example, from 20% to 50% utilization), the system promptly deducts or replenishes tokens, and blocks or pauses jobs when tokens are depleted.
[0232] This dynamism adapts to common load fluctuations in high-performance computing (such as iterative changes in AI model training), preventing resource overuse and quickly resuming job execution after resource release, thereby improving the system's responsiveness.
[0233] 5. Reduce the lag effect of resource contention.
[0234] To address token bucket lag (a short-term GPU usage exceeding the limit, resulting in negative tokens), the solution introduces compensation mechanisms such as timing prediction, buffered instruction queues, and context switching. Prediction of blocking is used in advance, buffered queues prevent direct suspension, and context switching provides forcible intervention in overuse.
[0235] These measures reduce system jitter or job interruptions caused by resource overuse, especially in high-concurrency scenarios, and can handle sudden loads more smoothly, ensuring service quality.
[0236] 6. Support efficient execution of diverse computing tasks.
[0237] Whether it is a kernel function such as matrix multiplication, convolution operation or vector addition, this solution can ensure that the task runs efficiently within the allocated resources by hijacking CUDA calls and token bucket control.
[0238] It is suitable for a variety of GPU-intensive application scenarios such as deep learning, scientific computing, and image processing, enhancing the versatility and practicality of the solution.
[0239] 7. Simplify cluster management and maintenance.
[0240] The compute node daemon process uniformly manages job startup, resource limits, and releases, and works with the monitoring process to provide usage history curves and average values, making it easier for administrators to analyze resource usage and optimize scheduling strategies.
[0241] The complexity of cluster operation and maintenance is reduced. Administrators can adjust token allocation or scheduling rules based on monitoring data to further optimize resource allocation efficiency.
[0242] 8. Improve user experience and fairness.
[0243] Users can flexibly specify the amount of GPU resources (integer or non-integer) based on actual needs, while the token bucket and isolation mechanism ensure fair resource usage. All jobs compete fairly for tokens, preventing strong jobs from squeezing out weaker jobs.
[0244] Users no longer have to worry about resource preemption, and job submissions can be more accurately matched to demand, improving the user experience and making it especially competitive in commercial shared GPU clusters. This also facilitates accurate metering and billing.
[0245] While the above embodiments describe a method for controlling a job process in detail, the present invention also provides corresponding embodiments of a device and server for controlling a job process. It should be noted that the present invention describes embodiments of the device from two perspectives: one based on functional modules and the other based on hardware.
[0246] An embodiment of the present invention provides a control device for a job process, which includes, based on the perspective of functional modules:
[0247] A first control module is configured to control a job process to call a preset file library and obtain a context corresponding to the job process recorded in the preset file library and used to represent resource usage; wherein different job processes in the preset file library correspond to different contexts;
[0248] An allocation module, configured to allocate a token for representing resource usage to a job process according to a context corresponding to the job process;
[0249] The acquisition and update module is used to obtain resource usage and update the number of tokens corresponding to the job process according to the resource usage;
[0250] The second control module is used to control the use of resources by the job process according to the updated token quantity.
[0251] In some embodiments, the control device for the operation process further includes:
[0252] The receiving module is used to receive the job submitted by the user through the scheduling system; wherein the job submitted by the user includes the resource model and resource quantity to be used; the resource quantity is an integer or a non-integer;
[0253] The allocation and marking module is used to schedule the system to select computing nodes that meet preset conditions based on the jobs submitted by the user, and after allocating resources on the computing nodes to the jobs, mark the resource usage; the preset condition is that the number of resources on the computing nodes is greater than or equal to the number of resources in the jobs submitted by the user; the resource usage includes full use or partial use.
[0254] In some embodiments, the control device for the operation process further includes:
[0255] A first creation module is used to create a daemon process and use the daemon process to create a namespace and a control group for representing and restricting resource usage of the job process;
[0256] The receiving and controlling module is used to receive the job process to be started sent by the scheduling system through the daemon process, and control the startup of the job process to be started;
[0257] The join module is used to join the job process to the namespace and control group.
[0258] In some embodiments, the first control module includes:
[0259] The first placement and control module is configured to place the directory of the job process into an original preset file library when detecting that the number of resources allocated to the job process is an integer, so as to control the job process to call the preset file library; wherein the original preset file library allows the job process to directly access and manage resources;
[0260] The second placement and control module is used to modify the files in the original preset file library to obtain a new preset file library when it is detected that the number of resources allocated to the job process is non-integer, and place the directory of the job process in the new preset file library to control the job process to call the preset file library; wherein, the new preset file library allows some job processes to directly access and manage resources.
[0261] In some embodiments, the acquisition submodule in the acquisition and update module is used to obtain resource usage.
[0262] Get submodules include:
[0263] The second creation module is used to create a monitoring process for characterizing and monitoring resource usage of each job process;
[0264] The first acquisition module is used to obtain the resource usage history curve of the job process through the monitoring process;
[0265] The first determining module is configured to determine an average resource utilization rate according to a resource utilization rate history curve of the job process as the resource utilization rate.
[0266] In some embodiments, the control device for the operation process further includes:
[0267] a second acquisition module, configured to acquire the computing power of the hardware device for parallel computing and set a first preset number of tokens in the token bucket according to the computing power of the hardware device for parallel computing;
[0268] In some embodiments, the allocation module includes:
[0269] A first allocation submodule is configured to allocate to the job process 0 the number of tokens representing resource usage if it is determined based on the context corresponding to the job process that the job process is prohibited from submitting computing tasks to the target resource;
[0270] The second allocation submodule is used to allocate a second preset number of tokens for representing the use of resources to the job process according to the context corresponding to the job process if it is determined that the computing task is allowed to be submitted to the target resource according to the context corresponding to the job process; wherein the first preset number is greater than the second preset number.
[0271] In some embodiments, the updating module in the acquisition and updating module is used to update the number of tokens corresponding to the job process according to the resource usage rate.
[0272] The update modules include:
[0273] A second determining module is used to determine the current number of tokens corresponding to the current resource utilization rate according to a preset relationship between the utilization rate of the hardware device for parallel computing and the number of tokens;
[0274] The elimination module is used to eliminate the tokens of the current number of tokens from the first preset number of tokens to update the number of tokens corresponding to the job process; wherein the updated number of tokens is the difference obtained by subtracting the current number of tokens from the first preset number.
[0275] In some embodiments, the second control module includes:
[0276] a detection and permission module, configured to allow the job process to submit computing tasks to use resources when detecting that the updated number of tokens is greater than a first preset value;
[0277] The detection and prohibition module is used to prohibit the job process from submitting computing tasks so as to prohibit the use of resources when it is detected that the updated token quantity is equal to the first preset value.
[0278] In some embodiments, the control device for the operation process further includes:
[0279] a prediction module configured to predict, when detecting that the updated number of tokens is a second preset value, a usage trend of the hardware device for parallel computing using a time series prediction algorithm; wherein the second preset value is greater than the first preset value, and a difference between the second preset value and the first preset value is less than a preset difference;
[0280] The prohibition module is used to prohibit the job process from submitting computing tasks if it is detected that the usage trend is an upward trend, so as to prohibit the use of resources.
[0281] In some embodiments, the control device for the operation process further includes:
[0282] The third creation module is used to pre-create computing tasks for storing blocked data;
[0283] The placement module is used to place the prohibited computing tasks into the buffer instruction queue after determining that the job process is prohibited from submitting computing tasks to the target resource according to the context corresponding to the job process, or after prohibiting the job process from submitting computing tasks.
[0284] In some embodiments, the control device for the operation process further includes:
[0285] The third acquisition module is used to obtain the status of computing tasks in the buffer instruction queue;
[0286] The pause module is used to pause the operation of the job process in the central processing unit from the operating system if it is detected that the computing task occupies the entire buffer instruction queue.
[0287] In some embodiments, the control device for the operation process further includes:
[0288] A switch-out and pause module, configured to force the context of the target job process to be switched out and to pause execution of the context of the target job process if it is detected that the target job process has a resource usage rate greater than a preset usage rate;
[0289] The cut-in module is used to control the context of other job processes to cut into the execution of the hardware device for parallel computing; wherein, other job processes are all job processes except the target job process.
[0290] In some embodiments, the cut-in module includes:
[0291] The fourth acquisition module is used to obtain the priority order of all other job processes;
[0292] The switching submodule is used to select the job process with the highest priority and control the context of the job process with the highest priority to switch to the execution of the hardware device for parallel computing.
[0293] In some embodiments, the control device for the operation process further includes:
[0294] The issuing module is used to reissue tokens to the job process prohibited from submitting computing tasks if it is detected that the resource usage rate of the job process prohibited from submitting computing tasks decreases after detecting that the updated token number is equal to 0;
[0295] The recovery module is used to resume the context execution of the job process that is prohibited from submitting computing tasks when it is detected that the number of tokens is greater than 0.
[0296] In some embodiments, the recovery module includes:
[0297] A first recovery submodule is configured to recover execution of the context of the job process that is prohibited from submitting computing tasks on the hardware device for parallel computing if it is detected that the context of the job process that is prohibited from submitting computing tasks is in a suspended state;
[0298] or, a migration module, configured to migrate the computing tasks in the buffered instruction queue to the instruction queue if it is detected that there is remaining space in the instruction queue; wherein the instruction queue is used to store computing tasks that are allowed to be submitted;
[0299] The second recovery submodule is configured to recover the operation of the job process in the CPU from the operating system if it is detected that the job process migrated to the instruction queue is in a state where the operation of the job process in the CPU is suspended in the operating system.
[0300] In some embodiments, the control device for the operation process further includes:
[0301] A first releasing module is used to release the resources occupied by the job process when detecting that the job process has finished executing;
[0302] The second release module is used to control the daemon process to release the namespace and the control group used to represent the resource usage restriction of the job process.
[0303] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0304] Figure 4 This is a structural diagram of a server provided by another embodiment of the present invention. This embodiment is based on the hardware perspective, such as Figure 4 As shown, the server includes:
[0305] Memory 20, for storing computer programs;
[0306] The processor 21 is configured to implement the steps of the method for controlling the operation process as described in the above embodiment when executing a computer program.
[0307] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented in at least one hardware form selected from the group consisting of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array. The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU, which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI processor for processing computing operations related to machine learning.
[0308] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201, wherein, after the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the method for controlling the job process disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include but is not limited to the data involved in the above-mentioned method for controlling the job process, etc.
[0309] In some embodiments, the server may further include a display screen 22 , an input / output interface 23 , a communication interface 24 , a power supply 25 , and a communication bus 26 .
[0310] Those skilled in the art will understand that Figure 4 The structure shown in the figure does not constitute a limitation to the server, and may include more or fewer components than shown in the figure.
[0311] The server provided by the embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the following method: a method for controlling a job process, with the same effect as above.
[0312] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned method for controlling the operation process when executed by a processor.
[0313] Finally, the present invention also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps described in the above method embodiment.
[0314] It is understood that if the methods in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0315] The computer-readable storage medium provided by the present invention includes the above-mentioned method for controlling the operation process, and the effect is the same as above.
[0316] The above is a detailed introduction to the control method, device, product, server and medium for an operation process provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.
[0317] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A method for controlling a work process, characterized in that: Applied to a computing node, the method includes: Controlling the job process to call a preset file library and obtaining a context corresponding to the job process recorded in the preset file library and used to characterize resource usage; wherein different job processes in the preset file library correspond to different contexts; Allocate a token for representing resource usage to the job process based on the context corresponding to the job process; Obtain resource usage and update the number of tokens corresponding to the job process based on the resource usage; Control the resource usage of the job process based on the updated token number; The control operation process calls the preset file library including: When it is detected that the number of resources allocated to the job process is an integer, the directory of the job process is placed in the original preset file library to control the job process to call the preset file library; wherein the original preset file library allows the job process to directly access and manage resources; When it is detected that the number of resources allocated to the job process is non-integer, the files in the original preset file library are modified to obtain a new preset file library, and the directory of the job process is placed in the new preset file library to control the job process to call the preset file library; wherein, the new preset file library allows some job processes to directly access and manage resources.
2. The method for controlling the operation process according to claim 1, wherein: Before the control operation process calls the preset file library and obtains the context corresponding to the operation process recorded in the preset file library and used to characterize the use of resources, the method further includes: Receiving a job submitted by a user through a scheduling system; wherein the job submitted by the user includes a resource model and a resource quantity to be used; the resource quantity is an integer or a non-integer; After receiving the job submitted by the user, the scheduling system further includes: A computing node that meets preset conditions is selected based on the job submitted by the user, and after resources on the computing node are allocated to the job, the resource usage is marked; wherein the preset condition is that the number of resources on the computing node is greater than or equal to the number of resources in the job submitted by the user; resource usage includes full use or partial use.
3. The method for controlling the operation process according to claim 2, wherein: After receiving a job submitted by a user through the scheduling system, the control job process calls a preset file library and obtains a context corresponding to the job process recorded in the preset file library and used to characterize resource usage, and further includes: Creating a daemon process and having the daemon process create a namespace and a control group for representing and limiting resource usage by a job process; Receiving the job process to be started sent by the scheduling system through the daemon process, and controlling the startup of the job process to be started; Adding the job process to the namespace and the control group; The control operation process calls the preset file library including: The job processes in the namespace and the control group are controlled to call a preset file library.
4. The method for controlling the operation process according to claim 1, wherein: The obtaining of resource usage includes: Create a monitoring process for characterizing and monitoring the resource usage of each job process; Obtaining a resource usage history curve of the job process through the monitoring process; The average resource usage is determined based on the resource usage history curve of the job process to serve as the resource usage.
5. The method for controlling the operation process according to claim 3, wherein: The resource is a hardware device used for parallel computing. Before allocating a token for representing the use of the resource to the job process according to the context corresponding to the job process, the method further includes: Obtaining the computing capability of a hardware device for parallel computing and setting a first preset number of tokens in a token bucket according to the computing capability of the hardware device for parallel computing; The step of allocating a token for representing resource usage to the job process according to the context corresponding to the job process includes: If it is determined based on the context corresponding to the job process that the job process is prohibited from submitting computing tasks to the target resource, the number of tokens allocated to the job process for representing the use of resources is 0; If it is determined based on the context corresponding to the job process that the computing task is allowed to be submitted to the target resource, a second preset number of tokens for representing the use of resources is allocated to the job process based on the context corresponding to the job process; wherein the first preset number is greater than the second preset number.
6. The method for controlling the operation process according to claim 5, wherein: Updating the number of tokens corresponding to the job process based on resource usage includes: Determining the current number of tokens corresponding to the current resource utilization rate according to a preset relationship between the utilization rate of the hardware device for parallel computing and the number of tokens; Tokens of the current number of tokens are removed from the first preset number of tokens to update the number of tokens corresponding to the job process; wherein the updated number of tokens is the difference obtained by subtracting the current number of tokens from the first preset number.
7. The method for controlling the operation process according to claim 6, wherein: Controlling the resource usage of the job process based on the updated token number includes: When detecting that the updated number of tokens is greater than a first preset value, allowing the job process to submit computing tasks to use resources; When it is detected that the updated number of tokens is equal to the first preset value, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
8. The method for controlling the operation process according to claim 7, wherein: The method further comprises: When it is detected that the updated number of tokens is a second preset value, predicting the usage trend of the hardware device for parallel computing by a time series prediction algorithm; wherein the second preset value is greater than the first preset value, and the difference between the second preset value and the first preset value is less than a preset difference; If it is detected that the usage trend is an upward trend, the job process is prohibited from submitting computing tasks to prohibit the use of resources.
9. The method for controlling the operation process according to claim 8, wherein: Also includes: Pre-create computing tasks to store blocked data; After determining that the job process is prohibited from submitting computing tasks to the target resource according to the context corresponding to the job process, or after prohibiting the job process from submitting computing tasks, the prohibited computing tasks are placed in the buffer instruction queue.
10. The method for controlling the operation process according to claim 9, wherein: Also includes: Obtaining the status of computing tasks in the buffer instruction queue; If it is detected that the computing task occupies the entire buffer instruction queue, the operation process in the central processing unit is suspended from the operating system.
11. The method for controlling the operation process according to claim 10, wherein: Also includes: If it is detected that there is a target job process with a resource usage rate greater than a preset usage rate, the context of the target job process is forced to be switched out and the execution of the context of the target job process is suspended; The context of controlling other job processes is switched into the execution of the hardware device for parallel computing; wherein the other job processes are job processes other than the target job process in all job processes.
12. The method for controlling the operation process according to claim 11, wherein: The control of switching the context of other job processes into the execution of the hardware device for parallel computing includes: Get the priority order of all other job processes; The job process with the highest priority is selected, and the context of the job process with the highest priority is controlled to be switched into the execution of the hardware device for parallel computing.
13. The method for controlling the operation process according to claim 12, wherein: Also includes: After detecting that the updated token quantity is equal to 0, if it is detected that the resource usage rate of the job process prohibited from submitting computing tasks decreases, the token is reissued to the job process prohibited from submitting computing tasks; When it is detected that the number of tokens is greater than 0, the context execution of the job process that is prohibited from submitting computing tasks is resumed.
14. The method for controlling the operation process according to claim 13, wherein: The process of resuming the execution of the context of the job process that is prohibited from submitting computing tasks includes: If it is detected that the context of the job process that is prohibited from submitting computing tasks is in a suspended state, resuming execution of the context of the job process that is prohibited from submitting computing tasks on the hardware device for parallel computing; Alternatively, if it is detected that there is remaining space in the instruction queue, the computing task in the buffer instruction queue is migrated to the instruction queue; wherein the instruction queue is used to store computing tasks that are allowed to be submitted; If it is detected that the job process migrated to the instruction queue is in a state where the operation of the job process in the central processing unit is suspended in the operating system, the operation of the job process in the central processing unit is resumed from the operating system.
15. The method for controlling the operation process according to claim 3, wherein: Also includes: When the execution of the detection job process ends, the resources occupied by the job process are released; The daemon process is controlled to release the namespace and the control group used to represent and restrict the use of resources by the job process.
16. A control device for an operation process, characterized in that: Applied to a computing node, the control device includes: A first control module is configured to control a job process to call a preset file library and obtain a context corresponding to the job process recorded in the preset file library and used to represent resource usage; wherein different job processes in the preset file library correspond to different contexts; An allocation module, configured to allocate a token for representing resource usage to a job process according to a context corresponding to the job process; The acquisition and update module is used to obtain resource usage and update the number of tokens corresponding to the job process according to the resource usage; A second control module is used to control the use of resources by the job process according to the updated number of tokens; The first control module includes: The first placement and control module is configured to place the directory of the job process into an original preset file library when detecting that the number of resources allocated to the job process is an integer, so as to control the job process to call the preset file library; wherein the original preset file library allows the job process to directly access and manage resources; The second placement and control module is used to modify the files in the original preset file library to obtain a new preset file library when it is detected that the number of resources allocated to the job process is non-integer, and place the directory of the job process in the new preset file library to control the job process to call the preset file library; wherein, the new preset file library allows some job processes to directly access and manage resources.
17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method for controlling the operation process according to any one of claims 1 to 15 are implemented.
18. A server, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for controlling the operation process according to any one of claims 1 to 15 when executing the computer program.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for controlling the operation process according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
GPU sharing method and device, electronic equipment and storage medium
CN117827423A
Task scheduling method and device, electronic equipment and storage medium
CN118132217A