Resource adjustment method and apparatus, electronic device, storage medium and training platform

By building process group configuration files on the AI ​​training platform and dynamically adjusting the central processor core binding, the problem that AI training platform is difficult to achieve optimal strategies in CPU resource allocation is solved, and efficient execution of AI training tasks and optimized use of CPU resources is achieved.

WO2025112885A1PCT designated stage expired Publication Date: 2025-06-05INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/122146
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-09-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing AI training platform is difficult to implement optimal strategies in CPU resource allocation, resulting in inefficient execution of AI training tasks.

Method used

By pre-constructing the process group configuration file, obtain the container set of computing nodes that perform the task to be trained and its associated containers. According to the resource configuration information of the container set, the resource occupation information of the computing nodes and the user resource requirements, the target graphics processor and the target central processor core are preferred. The target central processor core is selected for the container, and the configuration information of the corresponding process group of the container is updated to complete the binding between the central processor core and the container.

Benefits of technology

It achieves the maximum efficiency of AI training tasks, improves data access speed and data processing efficiency, and meets the optimization requirements of AI training tasks for CPU resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024122146_05062025_PF_FP_ABST
    Figure CN2024122146_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a resource adjustment method and apparatus, an electronic device, a storage medium and a training platform, which are applied to the technical field of artificial intelligence. The method comprises: acquiring a container set of computing nodes executing a task to be trained and a container associated with same; on the basis of resource configuration information of the container set, resource occupation information of the computing nodes and user resource requirements, and in a mode that target graphics processing unit and target central processing unit cores executing said task are preferentially located in the same non-uniform memory access group, selecting target central processing unit cores needing to be bound for the container, and updating configuration information of a process group corresponding to the container, so as to complete binding of the central processing unit cores and the container. The present application can solve the problem that AI training tasks cannot be efficiently completed in the related art, and can maximally ensure the efficient completion of AI training tasks by means of resource adjustment during the execution of the AI training tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Resource adjustment method, device, electronic device, storage medium and training platform

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311597253.7 and application name “Resource Adjustment Method, Device, Electronic Device, Storage Medium and Training Platform”, all contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of artificial intelligence technology, and in particular to a resource adjustment method, device, electronic device, readable storage medium, and artificial intelligence training platform. Background Art

[0004] With the rapid development of cloud technology, AI (Artificial Intelligence) training currently typically involves packaging training tasks into containers and building multi-task AI training platforms based on Kubernetes (an open-source container cluster management system). The AI ​​training process requires significant CPU (Central Processing Unit) resources to execute various tasks. To ensure efficient completion of AI training tasks, AI training platforms must provide optimal CPU allocation strategies for these tasks.

[0005] Summary of the Invention

[0006] This application provides the following technical solutions:

[0007] On one hand, the present application provides a resource adjustment method, comprising:

[0008] Pre-build a process group configuration file; the process group configuration file stores configuration information of the process group corresponding to the container, and the configuration information at least records the central processing unit core binding information of the corresponding container;

[0009] Get the container set of the computing node that executes the task to be trained and its associated containers;

[0010] Selecting a target CPU core for binding to the container based on the resource configuration information of the container set, the resource usage information of the computing node, and the user's resource requirements, such that the target GPU and the target CPU core executing the task to be trained are preferentially located in the same NUMAG; and

[0011] Update the configuration information of the process group corresponding to the container to complete the binding between the CPU core and the container;

[0012] Among them, the resource configuration information includes the graphics processors allocated to the container set, the user resource requirements include the required number of graphics processors and the required number of central processing unit cores required by the user to execute the task to be trained, the target number of central processing unit cores is the same as the required number of central processing unit cores, and the target number of graphics processors is the same as the required number of graphics processors.

[0013] In a first exemplary embodiment, selecting a target CPU core to be bound to a container in a manner that a target GPU and a target CPU core that execute a task to be trained are preferentially located in the same NUMAG includes:

[0014] Obtain the number of idle CPU cores and idle GPU cores of the compute node in the same target NUMAG at the current moment; and

[0015] In response to the number of idle GPU cores being greater than or equal to the required GPU number and the number of idle CPU cores being greater than or equal to the required CPU core number, a target GPU core and a target CPU core are selected from the target NUMAM group.

[0016] In a second exemplary embodiment, in response to the number of idle GPU cores being greater than or equal to the required number of GPU cores and the number of idle CPU cores being less than the required number of CPU cores, a target GPU is selected from the target NUMAR group, all CPU cores are selected from the target NUMAR group, and the remaining CPU cores are selected from the candidate NUMAR group to collectively form the target CPU cores.

[0017] In a third exemplary embodiment, obtaining the number of idle CPU cores and idle GPU cores of a computing node in the same target NUMAG at a current moment includes:

[0018] Obtaining the CPU core-GPU topology of the compute node; and

[0019] Based on the CPU core-GPU topology relationship, the number of idle CPU cores and idle GPU cores of the computing node in the same target NUMAG at the current moment is obtained.

[0020] In a fourth exemplary embodiment, if there are multiple target CPU cores, all target GPUs are located in the same target NUMAR group, and at least one target CPU core does not belong to the target NUMAR group, after selecting a target CPU core to bind to the container, the method further includes:

[0021] monitoring a usage status of at least one candidate CPU core that is within the target NUM group but does not belong to the target CPU core; and

[0022] In response to detecting the presence of an idle target candidate CPU core, while the container corresponding to the task to be trained maintains the current running state, the target candidate CPU core is adjusted to the target CPU core, and the first target CPU core that does not belong to the target non-uniform memory access group is released.

[0023] In a fifth exemplary embodiment, updating configuration information of a process group corresponding to a container includes:

[0024] The number of the first target CPU core in the configuration information of the process group corresponding to the container is replaced with the number of the target candidate CPU core.

[0025] In a sixth exemplary embodiment, obtaining a container set of computing nodes that execute a task to be trained and its associated containers includes:

[0026] After the training task is created on the business platform, the computing nodes allocated to the training task based on user resource requirements are obtained; the computing nodes have completed the creation of a container set with a medium-priority service quality based on resource configuration information;

[0027] Obtain resource configuration information corresponding to the container set based on the topological relationship of the computing nodes; and

[0028] Find the container corresponding to the container set based on resource configuration information.

[0029] In a seventh exemplary embodiment, after obtaining the container set of the computing nodes that execute the task to be trained and the associated containers thereof, the following steps are included:

[0030] In response to the container being a newly created container, based on the resource configuration information of the container set, the resource usage information of the computing node, and the user resource requirements, a target CPU core to be bound to the container is selected in such a manner that the target graphics processor and the target CPU core that execute the task to be trained are preferentially located in the same non-uniform memory access group.

[0031] In an eighth exemplary embodiment, after obtaining the container set of the computing node executing the task to be trained and its associated containers, the following steps are included:

[0032] In response to the container not being a newly created container, determining whether the target GPU and the target CPU of the container belong to the same NUMAG; and

[0033] In response to the target graphics processor and the target central processor of the container not belonging to the same non-uniform memory access group, a target processor core dynamic adjustment instruction is generated.

[0034] In a ninth exemplary embodiment, after obtaining the container set of the computing nodes that execute the task to be trained and the associated containers thereof, the following steps are included:

[0035] In response to the container not being a newly created container, when a CPU core quantity adjustment instruction is received, configuration information of a process group corresponding to the container is updated, and the updated configuration information of the process group is uploaded to the service platform.

[0036] In a tenth exemplary embodiment, updating configuration information of a process group corresponding to a container includes:

[0037] Get the directory of the host where the container is located; and

[0038] The configuration information of the process group corresponding to the container is read from the directory to update the number of the target CPU core to the corresponding configuration information.

[0039] In an eleventh exemplary embodiment, after updating the configuration information of the process group corresponding to the container, the method further includes:

[0040] monitoring utilization of at least one target central processing unit core of the container;

[0041] Automatically adjust the number of target CPU cores bound to the container based on the utilization of each target CPU core and preset resource dynamic adjustment conditions; and

[0042] In response to a change in the number of target CPU cores bound to the container, the configuration information of the process group corresponding to the container is updated again.

[0043] In a twelfth exemplary embodiment, before dynamically adjusting the conditions according to the utilization of each target CPU core and the preset resources, the method further includes:

[0044] Obtaining a maximum allowed CPU utilization rate and a minimum allowed CPU utilization rate; and

[0045] A preset resource dynamic adjustment condition is generated according to a numerical relationship between the utilization of the target CPU core and the minimum allowable CPU utilization and the maximum allowable CPU utilization.

[0046] In a thirteenth exemplary embodiment, automatically adjusting the number of target CPU cores bound to a container based on utilization of each target CPU core and a preset resource dynamic adjustment condition includes:

[0047] In response to detecting that the utilization of the current target CPU core is greater than or equal to the maximum allowed CPU utilization, at least one idle CPU core is selected from the NUMAG group to which the target GPU belongs as a new target CPU core.

[0048] In a fourteenth exemplary embodiment, automatically adjusting the number of target CPU cores bound to a container based on utilization of each target CPU core and a preset resource dynamic adjustment condition includes:

[0049] In response to detecting that the utilization of the current target CPU core is greater than or equal to the maximum allowable CPU utilization, and in response to the absence of an idle CPU core in the non-uniform memory access group to which the target graphics processor belongs, at least one idle CPU core is selected from other non-uniform memory access groups having idle CPU cores to serve as a new target CPU core.

[0050] In a fifteenth exemplary embodiment, automatically adjusting the number of target CPU cores bound to a container based on the utilization of each target CPU core and a preset resource dynamic adjustment condition includes:

[0051] In response to detecting that the utilization of the current target CPU core is less than the minimum allowable CPU utilization, the current target CPU core is unbound from the container. In a sixteenth exemplary embodiment, monitoring the utilization of at least one target CPU core of the container includes:

[0052] In response to receiving a utilization monitoring frequency configuration instruction, obtaining a utilization monitoring frequency; and

[0053] The utilization of at least one target central processing unit core of the container is monitored according to the utilization monitoring frequency.

[0054] In a seventeenth exemplary embodiment, the task to be trained is a multi-machine distributed training task, and after updating the configuration information of the process group corresponding to the container, the following steps are further included:

[0055] Obtaining the expected reduction in the number of CPU cores adjusted for each execution unit of the task to be trained, and using the minimum number of CPU cores adjusted as the target number of CPU cores adjusted; and

[0056] The CPU cores of each execution unit are adjusted accordingly according to the target CPU core adjustment number, and the target CPU core adjustment number is sent to the business platform so that the business platform updates the resource configuration of the task to be trained.

[0057] Another aspect of the present application provides a resource adjustment device, comprising:

[0058] A configuration file pre-building module is used to pre-build a process group configuration file; the process group configuration file stores configuration information of the process group corresponding to the container, and the configuration information records at least the CPU core binding information of the corresponding container;

[0059] The container information acquisition module is used to obtain the container set of the computing node that executes the training task and its associated containers;

[0060] a CPU core selection module for selecting a target CPU core to be bound to the container based on resource configuration information of the container set, resource usage information of the computing node, and user resource requirements, in a manner such that the target GPU and the target CPU core that execute the task to be trained are preferentially located in the same non-uniform memory access group; wherein the resource configuration information includes the GPUs allocated to the container set, the user resource requirements include the required number of GPUs and the required number of CPU cores required by the user for executing the task to be trained, the target number of CPU cores being the same as the required number of CPU cores, and the target number of GPUs being the same as the required number of GPU cores; and

[0061] The binding module is used to update the configuration information of the process group corresponding to the container to complete the binding of the CPU core and the container.

[0062] The present application also provides an electronic device, comprising a processor, wherein the processor is configured to implement the steps of any of the above resource adjustment methods when executing computer-readable instructions stored in a memory.

[0063] The present application also provides one or more non-volatile computer-readable storage media storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of any of the above resource adjustment methods.

[0064] Finally, this application also provides an artificial intelligence training platform, including a k8s cluster and multiple resource adjustment devices; at least one computing node of the k8s cluster is simultaneously deployed with a resource adjustment device.

[0065] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0067] FIG1 is a flow chart of a resource adjustment method provided in one or more embodiments of the present application;

[0068] FIG2 is a schematic diagram of binding a container to a CPU core according to one or more embodiments of the present application;

[0069] FIG3 is another schematic diagram of binding a container to a CPU core according to one or more embodiments of the present application;

[0070] FIG4 is a schematic diagram of binding an update container to a CPU core according to one or more embodiments of the present application;

[0071] FIG5 is a flow chart of another resource adjustment method provided by one or more embodiments of the present application;

[0072] FIG6 is a schematic diagram of a dynamic adjustment process of a CPU core according to one or more embodiments of the present application;

[0073] FIG7 is a schematic diagram of a control flow of dynamic adjustment of CPU cores according to one or more embodiments of the present application;

[0074] FIG8 is a schematic diagram of a framework of an exemplary application scenario provided by one or more embodiments of the present application;

[0075] FIG9 is a flow chart of another resource adjustment method provided in one or more embodiments of the present application;

[0076] FIG10 is a structural framework diagram of a specific implementation of a resource adjustment device provided by one or more embodiments of the present application;

[0077] FIG11 is a structural diagram of a specific implementation of an electronic device provided by one or more embodiments of the present application;

[0078] FIG12 is a structural framework diagram of a specific implementation of the artificial intelligence training platform provided by one or more embodiments of the present application. DETAILED DESCRIPTION

[0079] In order to enable those skilled in the art to better understand the technical solution of the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, the terms "first", "second", "third", "fourth", etc. in the specification, claims, and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. The term "exemplary" means "used as an example, embodiment, or illustrative." Any embodiment described here as "exemplary" is not necessarily to be interpreted as being superior or better than other embodiments.

[0080] With the development of AI technology, artificial intelligence (AI) has been widely applied across various industries. The demand for AI training tasks, which use large amounts of data and AI algorithms to train AI models and enable AI systems to perform specific tasks, is increasing. This has led to the emergence of AI training platforms that integrate project management, dataset management, data labeling, and model training, covering the entire process. Among these platforms, AI training platforms built on Kubernetes are widely used.

[0081] AI training tasks typically place high demands on computing resources such as GPUs and CPUs, storage resources, memory, and inter-host network bandwidth. CPU resources are particularly important. AI training tasks require CPU resources to perform tasks such as reading and writing disk datasets and forwarding network data. Furthermore, in some large-model scenarios, some training frameworks offload model parameters that cannot be stored in GPU memory to internal memory or utilize the CPU for calculations. Data transfer between internal memory and GPU memory also consumes significant CPU resources, requiring AI training platforms to provide optimal CPU allocation for training tasks.

[0082] Without CPU limits for containers, multiple container processes running on the same node will compete for CPU resources, leading to unstable training speeds during distributed training tasks. Simply limiting the number of CPU cores a container can use without CPU binding means that processes within the container will still randomly run on multiple CPUs based on Linux's native CPU scheduling mechanism. This only guarantees a certain percentage of CPU usage, but cannot guarantee efficient AI training tasks. To address this, related technologies not only limit the number of CPU cores but also bind containers to CPU cores. However, these technologies require that each container in a pod have the same CPU and memory request (i.e., the startup resources that Kubernetes must guarantee) and limit (i.e., the upper limit on the resources that the container can potentially use in the future) for CPU core binding to work. This is not suitable for pod-based AI training scenarios. Furthermore, while processes within a container are limited to a fixed number of CPU cores, there is no guarantee that the allocated CPUs and the container's GPUs will be in the same NUMA group, resulting in inefficient AI training tasks.

[0083] In view of this, in order to solve the above-mentioned drawbacks of the related art, this application selects the central processing unit core bound to the container and the graphics processor allocated to the container set according to the resource configuration information of the container set, the resource occupancy information of the computing node and the user's resource requirements, so as to ensure that they belong to the same non-uniform memory access group to the greatest extent possible, thereby adjusting the resources required for the AI ​​training task to achieve the maximum guarantee of efficient completion of the AI ​​training task. After introducing the technical solution of the present application, various non-restrictive implementation methods of the present application are described in detail below. In order to better illustrate the present application, many specific details are given in the specific implementation methods below. Those skilled in the art should understand that the present application can also be implemented without these specific details. In other examples, the methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.

[0084] First, please refer to Figure 1, which is a flow chart of a resource adjustment method provided in this embodiment. This application is applicable to the adjustment of computing resources in the process of executing training tasks on an AI training platform built based on Kubernetes. This embodiment may include the following contents:

[0085] S101: Pre-build a process group configuration file.

[0086] The process group configuration file of this embodiment is used to store configuration information for the process group corresponding to each container. This configuration information at least records the CPU core binding information for the corresponding container. The process group configuration file may include multiple subfiles, each of which uniquely corresponds to a container and is used only to record data related to the CPU core binding of that container, such as the CPU core number. Of course, the process group configuration file may include multiple subfiles, each of which records data related to the CPU core binding of multiple containers. The CPU core binding information for each container can be indexed using the container's unique identification information, such as the container ID, to facilitate data retrieval. The process group configuration file supports updating the CPU core binding information of a container during operation. Update operations include, but are not limited to, deletion, addition, and modification. When the CPU core binding information of a container in the process group configuration file changes, the CPU core bound to the container is adjusted accordingly to maintain consistency. Because the process group configuration file is updated during normal container operation, the container does not need to be rebuilt and will not affect container operation. In other words, dynamic adjustment and binding of the container's CPU core is achieved without terminating the corresponding pending tasks of the container.

[0087] S102: Obtain a container set of computing nodes that execute the task to be trained and their associated containers.

[0088] In this step, the task to be trained is sent to the AI ​​training platform by the user through the business platform or automatically triggered by the business platform through an automated program. The task to be trained can be any training task in the field of artificial intelligence, such as distributed training tasks, multi-machine multi-card training tasks, single-machine multi-card training tasks, single-machine single-card tasks, multi-level single-card tasks, etc. Among them, multi-machine means that the task to be trained is executed by multiple workers (that is, execution units), and multi-card refers to multiple CPUs. When the business platform sends the task to be trained, it will carry the user resource requirements of the task to be trained. The so-called user resource requirements refer to the computing resources required when the user wants to execute the task to be trained. For example, the business platform creates a distributed training task. The job (that is, the task) contains 2 workers, and each worker requires 4 GPUs and 40 CPUs. The following resource configuration is usually made for AI training tasks: There is no limit on memory usage, only the use of CPU resources is limited. At this time, the Qos of the Pod is Burstable, and the built-in container binding CPU core function of k8s cannot be used as shown below:

[0089] The platform scheduler of the k8s cluster completes node-level resource allocation. When based on a scheduling strategy such as spread (also known as tiling), the execution units of the training tasks, such as two workers, will be allocated to the corresponding computing nodes. By scheduling these two computing nodes to execute the tasks to be processed, the computing nodes can be identified after the Pod creation and startup is completed based on the k8s mechanism. At the same time, the resource configuration information of the container set can also be obtained. The resource configuration information of the container set includes but is not limited to the number of CPUs configured for the Pod and the GPU information allocated to the Pod. Pod is the smallest unit of k8s and is also the resource object that minimizes the running of containerized applications. It contains a group of containers. A Pod represents a process running in a k8s cluster.

[0090] S103: According to the resource configuration information of the container set, the resource usage information of the computing node and the user resource requirements, the target CPU core to be bound to the container is selected in such a manner that the target GPU and the target CPU core that execute the task to be trained are preferentially located in the same non-uniform memory access group.

[0091] After determining the computing node, Pod (container set) and obtaining the task to be trained in the above steps, the resource configuration information of the container set, the user resource requirements and the resource occupancy information of the computing node, especially the CPU core resource usage, can be obtained. When this step manages the CPU core bound to the container, the CPU core in the same NUMA group as the GPU is preferably bound. Multiple GPUs in the same group can be connected to each CPU core through a PCIe (peripheral component interconnect express) switch (also known as an interface chip). For ease of description, this step defines the GPU used to execute the task to be trained as the target graphics processor, and the CPU core used as the target central processing unit core, that is, the central processing unit core bound to the container. Since the resource configuration information includes the graphics processors allocated to the container set, the target graphics processor and the target central processing unit core that execute the task to be trained are preferentially located in the same non-uniform memory access group, so as to determine whether the GPU and the CPU core belong to the same NUMA group. Based on the user resource requirements including the user requirements for the required number of graphics processors and the required number of central processing unit cores for executing the task to be trained, it can be determined that the final selected target central processing unit core number is the same as the required number of central processing unit cores, and the target graphics processor number is the same as the required number of graphics processors.

[0092] S104: Update the configuration information of the process group corresponding to the container to complete the binding between the CPU core and the container.

[0093] It is understandable that when a container is created, processes within the container are restricted to running only on a few fixed CPU cores. Using the memory and GPU of the NUMA group containing the CPU core can improve performance within the container. After determining the target CPU core to bind the container to in the previous step, the binding operation is required. Cgroup (control group) is a mechanism in Linux that manages processes by group. It is used to limit, control, and count system resource usage by process group. Binding a CPU core to a container can be achieved by updating the corresponding CPU configuration file data recorded in the configuration information of the process group corresponding to the container.

[0094] In the technical solution provided in this embodiment, the central processing unit core and the graphics processor allocated to the container set can be selected and bound to the container according to the resource configuration information of the container set, the resource occupancy information of the computing node and the user resource requirements, so as to ensure that they belong to the same non-uniform memory access group to the greatest extent, greatly improve the data access speed, thereby improving the data processing efficiency, and thus completing the AI ​​training task efficiently to the greatest extent; in addition, the container binding CPU core does not limit the highest priority of its service quality level, and there is no need to set the same restrictions on the CPU and memory of each container in the Pod, which is more in line with the Pod usage scenario of AI training tasks and is more practical.

[0095] It should be noted that there is no strict order of execution between the steps in this application. As long as they comply with the logical order, these steps can be executed simultaneously or in a preset order. Figure 1 is only a schematic diagram and does not mean that this is the only execution order.

[0096] In the above embodiment, when managing the CPU cores bound to the container, priority is given to binding the CPU cores in the same NUMA group as the GPU. However, it is understandable that there are inevitably application scenarios where this requirement cannot be met. Based on this, taking into account both practicality and task execution efficiency, this application also provides an exemplary CPU core management method, which may include the following:

[0097] Obtain the number of idle CPU cores and idle GPU cores of the computing node in the same target NUMAG at the current moment;

[0098] If the number of idle GPU cores is greater than or equal to the required GPU core number, and the number of idle CPU cores is greater than or equal to the required CPU core number, a target GPU core and a target CPU core are selected from the target NUMAM group.

[0099] If the number of idle GPU cores is greater than or equal to the required number of GPU cores, and the number of idle CPU cores is less than the required number of CPU cores, a target GPU is selected from the target NUMAR group, all CPU cores are selected from the target NUMAR group, and the remaining CPU cores are selected from the candidate NUMAR groups to form the target CPU cores.

[0100] This embodiment selects bound CPU cores for the container based on the usage of the computing resources of the computing node. For example, when the computing node assigned to the Pod runs and executes the task to be trained, 4 GPUs and 60 CPU cores are required to execute the task to be trained. The computing nodes each include 8 GPUs and 120 CPUs. The NUMA0 resource group contains GPUs numbered 0, 1, 2, and 3, and CPUs numbered 0-59. The NUMA1 group contains GPUs numbered 4, 5, 6, and 7, and CPUs numbered 60-119. When the idle resources of the GPU and CPU cores of the computing node are sufficient, priority is given to binding the CPU cores in the same NUMA group as the container GPU. For example, all GPUs in the NUMA0 group are selected and the 40 CPU cores from CPU0 to CPU39 are bound, as shown in Figure 2. When the compute node has ample GPU resources but limited CPU resources, such as GPU4, GPU5, GPU6, GPU7, and CPU20 to CPU59 in the compute node are already occupied by other tasks, CPU resources across NUMA groups and GPU resources within the same NUMA group are allocated to the container. For example, 4 GPUs in the NUMA0 group, 20 CPUs in the NUMA0 group, and 20 CPUs in the NUMA1 group are selected, as shown in Figure 3. After the node's CPU cores are selected, the CPU core binding configuration file in the Cgroup associated with the container is dynamically searched. The file directory can be located in the container's host directory: / sys / fs / Cgroup / CPUset / kubePods.slice / kubePods-PodID / containerID / CPUset.CPUs. The CPU core allocation result is updated to this file, completing the binding of the specified CPU core to the container.

[0101] To further improve the execution efficiency of the training task, the CPU cores allocated to the training task container in the initial stage are in different NUMA groups. When other CPUs in the NUMA group of the computing node are released, this application also provides an implementation method for dynamically adjusting the CPU cores bound to the container, which may include the following:

[0102] In this embodiment, when it is not possible to bind a CPU core in the same NUMA group as the GPU to a container, a CPU core in a different NUMA group may be selected to be bound to the container. To ensure that CPU cores in the same NUMA group as the GPU are preferentially bound to the container, the system monitors the usage status of each CPU core on the compute node in real time. When a CPU core in the same NUMA group is released by another task, the CPU core bound to the container is automatically updated to a CPU in the same NUMA group. In other words, if there are multiple target CPU cores, all target GPUs are in the same target NUMAM group, but at least one target CPU core does not belong to the target NUMAM group, during the execution of the task to be trained, the usage status of at least one candidate CPU core in the target NUMAM group but not in the target CPU core is monitored. When an idle target candidate CPU core is detected, the target candidate CPU core may be directly adjusted to the target CPU core while the container corresponding to the task to be trained maintains its current running state, i.e., without terminating the task to be trained, and the first target CPU core that does not belong to the target NUMAM group is released. For example, the CPU cores bound to the container in the initial state are 20 CPUs in the NUMA0 group and 20 CPUs in the NUMA1 group. Through the method steps of this embodiment, the container and CPU core binding relationship configuration file is automatically updated without restarting the container, adjusting the binding relationship to 40 CPUs in NUMA0, as shown in Figure 4. The target non-uniform memory access group is used to refer to the NUMA where the target GPU is located. The candidate CPU cores can be randomly selected from one or more CPU cores currently in use in that NUMA, or all CPU cores currently in use in that NUMA. If multiple target candidate CPU cores are released at the same time, and their number is greater than the number of target CPU cores, then the same number of target candidate CPU cores can be randomly selected to bind to the container. If multiple target candidate CPU cores are released at the same time, but their number is less than the number of target CPU cores, then the same number of target candidate CPU cores can be randomly selected to unbind from the container, so that all target candidate CPU cores can be bound. After replacing the target CPU core with a target candidate CPU core, the configuration information of the corresponding process group needs to be updated. For the sake of ease of description, this embodiment can define the target central processing unit core unbound from the container as the first target central processing unit core, that is, replace the number of the first target central processing unit core in the configuration information of the process group corresponding to the container with the number of the target candidate central processing unit core, thereby realizing dynamic adjustment of CPU resources and further improving the execution efficiency of the task to be trained.

[0103] As an exemplary embodiment, when updating the configuration information of a process group, the directory of the host machine where the container resides can be first obtained; the configuration information of the process group corresponding to the container can be read from the directory to update the target CPU core number in the configuration information. In the above embodiment, the number of the first target CPU core in the configuration information of the process group corresponding to the container can be replaced with the number of the target candidate CPU core.

[0104] As a simple implementation method, to determine whether the GPU and CPU cores belong to the same NUMAR group, we can first obtain the compute node's CPU-GPU topological relationship. A topological relationship refers to the relationship between spatial data that satisfies the principles of topological geometry. Based on this CPU-GPU topological relationship, we can obtain the number of idle CPU and GPU cores in the same target NUMAR group at the current moment.

[0105] The above embodiment does not limit how to perform S102. This application also provides an exemplary method for obtaining a container set of computing nodes for a task to be trained and their associated containers, and provides corresponding subsequent binding implementation methods based on different application scenarios of the containers, which may include the following:

[0106] As shown in Figure 5, after the business platform creates a task to be trained, K8s allocates a compute node to the task based on user resource requirements. The compute node has already created a container set with a service quality of medium or higher priority based on resource configuration information. The platform then obtains the compute node and, based on the topological relationship of the compute node, obtains the resource configuration information corresponding to the container set. The platform then searches for the container corresponding to the container set based on the resource configuration information. If the container is a newly created container, the platform selects a target CPU core to bind to the container based on the resource configuration information of the container set, the resource usage information of the compute node, and the user's resource requirements, ensuring that the target GPU and target CPU core executing the task to be trained are preferentially located in the same non-uniform memory access group. If the container is a newly created container, the platform selects a CPU core to bind to the newly created container in accordance with the methods described in the above embodiment, and updates the Cgroup configuration information corresponding to the container. If the container is not a newly created container, the platform dynamically adjusts the CPU core based on the NUAM group, determining whether the target GPU and target CPU of the container belong to the same non-uniform memory access group. If the target GPU and target CPU of the container do not belong to the same non-uniform memory access group, a target CPU core dynamic adjustment instruction is generated. When receiving a command to dynamically adjust the number of target CPU cores, the system monitors the usage status of the CPU cores in the same NUAM group as the target GPU in real time. If any idle CPU cores are found in the same NUAM group, they are bound and the container's corresponding Cgroup configuration is updated. If the container is not a newly created container, upon receiving a command to adjust the number of CPU cores, the system updates the configuration of the container's corresponding process group and uploads the updated process group configuration to the service platform.

[0107] Kubernetes provides different QoS (Quality of Service) levels for pods, determined by the resource requests and limits specified when creating a Kubernetes pod. These levels include BestEffort (lowest priority), Burstable (medium priority), and Guaranteed (highest priority). In this embodiment, Kubernetes allocates CPUs and binds CPU cores to pods with Guaranteed (highest priority) and Burstable (medium priority) QoS based on NUMA (Non-Uniform Memory Access) affinity. NUMA affinity refers to the ability to associate a task or process with a specific NUMA node. By setting NUMA affinity, you can specify that a task run on a specific NUMA node to minimize remote memory access and improve performance. When a task is associated with a specific NUMA node, it is more likely to use the local memory associated with that node. Local memory refers to the memory associated with the CPU on the same NUMA node as the task, significantly improving data access speed and data processing efficiency.

[0108] As can be seen from the above, this embodiment provides different container binding methods based on the status of different containers, which is more flexible and can also realize dynamic adjustment of CPU resources, further improving the execution efficiency of the training task.

[0109] It is understandable that when users submit tasks to be trained to the business platform, they are not sure how many CPU cores are needed, and they may apply for more CPU cores. In this case, the task may not be scheduled or CPU resources may be wasted due to the lack of idle resources in the cluster. Of course, there will also be situations where fewer CPU cores are applied, which may cause the task to be trained to become a bottleneck in the stage with greater CPU demand, such as loading data sets. Based on this, the application can also support monitoring the CPU usage in each container of the computing node, and dynamically increase or decrease the number of CPU cores bound to the container based on the container CPU utilization, so as to maximize the effective utilization of node resources. This embodiment may include the following:

[0110] Monitor the utilization of at least one target CPU core of the container; automatically adjust the number of target CPU cores bound to the container based on the utilization of each target CPU core and preset resource dynamic adjustment conditions; when the number of target CPU cores bound to the container changes, update the configuration information of the process group corresponding to the container again.

[0111] Among them, the preset resource dynamic adjustment condition is a condition for measuring whether the user demand resources of the current task to be trained need to be adjusted. It can be a fixed condition or a condition that can be flexibly configured according to the actual scenario. Exemplarily, the maximum allowable CPU utilization and the minimum allowable CPU utilization can be obtained; according to the numerical relationship between the utilization of the target CPU core and the minimum allowable CPU utilization and the maximum allowable CPU utilization, the preset resource dynamic adjustment condition is generated. The maximum allowable CPU utilization and the minimum allowable CPU utilization can be flexibly selected according to the actual application scenario. This embodiment can automatically complete the dynamic increase and decrease of CPU cores according to the CPU utilization of the task, the business dynamically updates the CPU core of the task according to the resource usage, monitors the utilization of each CPU core, and dynamically increases or reduces the CPU core binding. For the adjustment of the CPU core bound to a container, there are the following two processing methods for application scenarios:

[0112] When it is detected that the utilization of the current target CPU core is greater than or equal to the maximum allowable CPU utilization, in order to improve the task execution efficiency, at least one idle CPU core is selected from the non-uniform memory access group to which the target GPU belongs as a new target CPU core. In order to ensure the task execution efficiency, when it is detected that the utilization of the current target CPU core is greater than or equal to the maximum allowable CPU utilization, if there are no idle CPU cores in the non-uniform memory access group to which the target GPU belongs, at least one idle CPU core is selected from other non-uniform memory access groups with idle CPU cores as a new target CPU core. When it is detected that the utilization of the current target CPU core is less than the minimum allowable CPU utilization, the current target CPU core is unbound from the container.

[0113] For example, the maximum allowed CPU utilization is 80%, the minimum allowed CPU utilization is 20%, and the cluster CPU resources are relatively idle. The user applies for 4 CPUs for each worker to be trained. The system can quickly complete resource scheduling for the task. Although the CPUs allocated to the container are relatively small, the distributed training task can be quickly started. After the training task runs for a period of time, an idle CPU appears on the node where the container is located. This embodiment recognizes that the worker has CPU pressure. If the current CPU utilization is above 80%, multiple CPU cores are automatically added to the container and bound to reduce the CPU utilization of the container. As shown in Figure 6, the utilization of the four CPU cores (i.e., CPU0, CPU1, CPU2, and CPU3) initially bound to the Pod is 90%. After dynamic adjustment through the above embodiment, a new CPU core is bound to the Pod, and the utilization of the five CPU cores (i.e., CPU0, CPU1, CPU2, CPU3, and CPU4) currently bound to the Pod is 72%. The user has applied for a large amount of CPU resources for the training task. After the training task has run for a period of time, the container CPU utilization is very low. When the CPU utilization in the container is identified to be less than 20%, the system can automatically reduce the number of CPU cores bound to the container to improve the container's CPU utilization.

[0114] As an exemplary implementation of this embodiment, in order to maximize the effective utilization of resources and avoid the CPU core utilization monitoring occupying too many resources and affecting the normal operation of the entire system, this embodiment flexibly configures or modifies the utilization monitoring frequency according to the actual application scenario. When the utilization monitoring frequency configuration instruction is received, the utilization monitoring frequency is obtained; according to the utilization monitoring frequency, the utilization of at least one target central processing unit core of the container is monitored.

[0115] As can be seen from the above, this embodiment dynamically increases or decreases the number of CPU cores bound to the container based on the container CPU utilization, thereby maximizing the effective utilization of computing node resources, which is conducive to improving the execution efficiency of the training task.

[0116] It is understandable that when the training task issued by the user is a multi-machine distributed training task, the training task will be executed by multiple execution units. In order to ensure that the correct number of CPU cores is used when the Pod is rebuilt or the task is resubmitted after the number of CPU cores of the container is dynamically adjusted, based on the above embodiment, this application also provides an embodiment to ensure that the correct number of CPU cores is used when the Pod is rebuilt or the task is resubmitted, which may include the following content:

[0117] Obtain the expected reduction in the number of CPU cores adjusted for each execution unit of the task to be trained, and use the minimum number of CPU cores adjusted as the target number of CPU cores adjusted; adjust the CPU cores of each execution unit according to the target number of CPU cores adjusted, and send the target number of CPU cores adjusted to the business platform so that the business platform updates the resource configuration of the task to be trained.

[0118] In this embodiment, the resource configuration of each worker of the multi-machine distributed training task is the same. When the number of CPU cores of the container is dynamically adjusted, it is necessary to ensure that the resource configuration of each worker is consistent. The Pod managed by k8s will be rebuilt due to an exception. At this time, when the CPU cores used by the underlying container are increased or decreased, it is necessary to simultaneously notify the business system to update the resource configuration of the AI ​​training task to ensure that the correct number of CPU cores is used when the Pod is rebuilt or the task is resubmitted. Taking Figure 7 as an example, after the CPU cores are dynamically adjusted according to the above embodiment, the number of CPUs expected to be reduced for the three execution units executing the multi-machine distributed training task is respectively 2 CPUs for execution unit 0, 4 CPUs for execution unit 1, and 3 CPUs for execution unit 2. The expected number of CPUs to be reduced for each execution unit is obtained, and the minimum number of CPUs to be reduced is determined. The number of CPU cores of each execution unit is updated based on the minimum number of CPUs to be reduced. As shown in Figure 7, the updated number of CPUs for execution unit 0 is expected to be reduced by 2, the number of CPUs for execution unit 1 is expected to be reduced by 2, and the number of CPUs for execution unit 2 is expected to be reduced by 2. In the case of adding bound CPU cores, the expected number of CPUs to be added to each execution unit is obtained, and the maximum number of CPUs to be added is determined. The number of CPU cores of each execution unit is updated based on the maximum number of CPUs to be added.

[0119] As can be seen from the above, this embodiment can ensure that after the number of CPU cores of the container is dynamically adjusted, the correct number of CPU cores is used when the Pod is rebuilt or the task is resubmitted, which is conducive to improving task execution efficiency.

[0120] Finally, based on the above-mentioned technical solution of the present application, some possible application scenarios involved in the technical solution of the present application are introduced below with reference to FIG8. FIG8 is a schematic diagram of a hardware composition framework applicable to a resource adjustment method provided by the present application, which may include the following contents:

[0121] The hardware framework may include a first electronic device 81 and a second electronic device 82, which are connected via a network 83. The first electronic device 81 deploys a Kubernetes cluster, which manages the memory, GPU, CPU, and network resources within the cluster. A processor is deployed on each compute node to execute the resource adjustment method described in any of the above embodiments. The processor monitors the CPU utilization of each container on each compute node within the cluster, uses information about idle CPU cores, the relationship between CPU cores and NUMA groups, and GPUs to dynamically identify the pod configuration running on each compute node, locates the container corresponding to the pod, searches the Cgroup configuration file for the container on the host based on the container ID, and updates the CPU core binding configuration file therein. For pods with Guaranteed and Burstable QoS, the appropriate CPU core is selected based on the CPU core distribution of the compute node and updated to the container CPU core binding configuration file. When a compute node CPU core is freed by other tasks, the container's bound CPU core is dynamically updated based on the CPU's NUMA grouping. Based on the container CPU utilization, the container's CPU cores are dynamically increased or decreased based on the CPU configuration file update scheme in the Cgroup, and corresponding binding and unbinding operations are performed. Taking the binding process shown in Figure 9 as an example to illustrate the process of binding a container to a CPU core in this embodiment, the user creates a multi-level multi-card task through the second electronic device 82 and sends it to the first electronic device 81. The scheduler of the k8s cluster of the first electronic device 81 completes the node scheduling, and the device plug-in of the node completes the GPU allocation. The topological relationship between the CPU core and the GPU of the node is obtained, and it is determined whether there are enough CPU cores in the same NUMA group as the GPU. If so, the CPU cores of the same NUMA group are selected for binding. If not, the CPU cores of different NUMA groups are selected for binding. The configuration information of the process group of the Pod is searched and the number of the selected CPU core is written to the configuration file. For the cross-group binding scenario, the CPU core in the same NUMA group as the GPU is monitored to see if it is idle. If so, the number of the idle CPU core is obtained, the configuration information of the process group of the Pod is obtained, and the number of the CPU core in the configuration information that is not in the same NUMA group as the GPU is replaced. The first electronic device 81 can be a server. The second electronic device 82 deploys a business platform that can be used to provide a human-computer interaction interface.

[0122] Based on the above-mentioned technical solution of the present application, one of the application scenarios of the embodiment of the present application can be realized through the interaction between the second electronic device 82 and the user. In this application scenario, the user sends tasks to be executed through the second electronic device 82 and can also access data. The data access can be accessing information on the first electronic device 81 through interaction between the second electronic device 82 and the first electronic device 81, or it can be used to directly access the information of the second electronic device 82 itself. This embodiment does not limit this.

[0123] It should be noted that the above application scenarios are only shown to facilitate understanding of the ideas and principles of this application, and the embodiments of this application are not limited in this respect. On the contrary, the embodiments of this application can be applied to any applicable scenario.

[0124] As can be seen from the above, this embodiment manages AI training tasks based on the Pod orchestration management mechanism of k8s, increases the container binding to the CPU core in the same NUMA group as the GPU, and monitors the CPU utilization in the container at the same time, so as to dynamically increase or decrease the CPU cores bound to the container and improve resource utilization; when the CPU of a single NUMA group cannot meet the CPU core number requirement of the container during the container creation phase, the CPUs of the two NUMA groups are allocated to the same container. When the CPU core of a NUMA group is released, all the CPU cores of the container are dynamically bound to the same NUMA group to improve the CPU and GPU utilization.

[0125] The present application also provides a corresponding device for the resource adjustment method, which further makes the method more practical. Among them, the device can be described from the perspective of functional modules and hardware. The resource adjustment device provided by the present application is introduced below. The device is used to implement the resource adjustment method provided by the present application. In this embodiment, the resource adjustment device may include or be divided into one or more program modules. The one or more program modules are stored in a storage medium and executed by one or more processors to complete the resource adjustment method disclosed in Example 1. The program module referred to in this embodiment refers to a series of computer-readable instruction segments that can complete specific functions, which is more suitable for describing the execution process of the resource adjustment device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module of this embodiment. The resource adjustment device described below and the resource adjustment method described above can be referenced to each other.

[0126] From the perspective of functional modules, see FIG10 , which is a structural diagram of a resource adjustment device provided in this embodiment in a specific implementation manner. The device may include:

[0127] The configuration file pre-construction module 101 is used to pre-construct a process group configuration file; the process group configuration file stores configuration information of the process group corresponding to the container, and the configuration information records at least the CPU core binding information of the corresponding container;

[0128] The container information acquisition module 102 is used to obtain the container set of the computing node that executes the task to be trained and its associated containers;

[0129] The CPU core selection module 103 is configured to select a target CPU core to be bound to the container based on the resource configuration information of the container set, the resource usage information of the computing node, and the user resource requirements, in a manner such that the target GPU and the target CPU core that perform the task to be trained are preferentially located in the same non-uniform memory access group. The resource configuration information includes the GPUs allocated to the container set, the user resource requirements include the required number of GPUs and the required number of CPU cores required by the user for performing the task to be trained, the target number of CPU cores being the same as the required number of CPU cores, and the target number of GPUs being the same as the required number of GPU cores.

[0130] The binding module 104 is configured to update the configuration information of the process group corresponding to the container to complete the binding between the CPU core and the container.

[0131] For example, in some implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0132] Obtain the number of idle CPU cores and idle GPU cores of the computing node in the same target NUMAG at the current moment;

[0133] If the number of idle GPU cores is greater than or equal to the required GPU core number, and the number of idle CPU cores is greater than or equal to the required CPU core number, a target GPU core and a target CPU core are selected from the target NUMAM group.

[0134] As an implementation method parallel to the above embodiment, the CPU core selection module 103 may also be used to:

[0135] If the number of idle GPU cores is greater than or equal to the required number of GPU cores, and the number of idle CPU cores is less than the required number of CPU cores, a target GPU is selected from the target NUM group, all CPU cores are selected from the target NUM group, and the remaining CPU cores are selected from the candidate NUM group to form the target CPU cores. As an exemplary implementation of the above embodiment, the CPU core selection module 103 may also be used to:

[0136] Obtaining a CPU core-GPU topology relationship of the computing node; based on the CPU core-GPU topology relationship, obtaining the number of idle CPU cores and idle GPU cores of the computing node in the same target NUMAG at the current moment.

[0137] For example, in some other implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0138] If there are multiple target CPU cores, all target GPUs are located in the same target non-uniform memory access group, and there is at least one target CPU core that does not belong to the target non-uniform memory access group. During the execution of the task to be trained, the usage status of at least one candidate CPU core that is located in the target non-uniform memory access group but does not belong to the target CPU core is monitored; when it is detected that there is an idle target candidate CPU core, the target candidate CPU core is adjusted to the target CPU core while the container corresponding to the task to be trained maintains the current running state, and the first target CPU core that does not belong to the target non-uniform memory access group is released.

[0139] As an exemplary implementation of the above embodiment, the binding module 104 may also be used to:

[0140] The number of the first target CPU core in the configuration information of the process group corresponding to the container is replaced with the number of the target candidate CPU core.

[0141] For example, in some further implementations of this embodiment, the container information acquisition module 102 may also be used to:

[0142] After the training task is created on the business platform, the computing nodes allocated to the training task are obtained based on user resource requirements. The computing nodes have completed the creation of a container set with a medium-priority service quality based on the resource configuration information. Based on the topological relationship of the computing nodes, the resource configuration information corresponding to the container set is obtained. The container corresponding to the container set is found based on the resource configuration information.

[0143] For example, in some further implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0144] If the container is a newly created container, the target CPU core to be bound to the container is selected based on the resource configuration information of the container collection, the resource usage information of the computing node, and the user's resource requirements, with the target GPU and target CPU core executing the training task preferentially located in the same non-uniform memory access group.

[0145] For example, in some further implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0146] If the container is not a newly created container, determine whether the target graphics processor and the target central processor of the container belong to the same non-uniform memory access group; if the target graphics processor and the target central processor of the container do not belong to the same non-uniform memory access group, generate a target processor core dynamic adjustment instruction.

[0147] For example, in some further implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0148] If the container is not a newly created container, when a CPU core number adjustment instruction is received, the configuration information of the process group corresponding to the container is updated, and the updated configuration information of the process group is uploaded to the business platform.

[0149] For example, in some further implementations of this embodiment, the binding module 104 may also be used to:

[0150] Obtain the directory of the host machine where the container is located; read the configuration information of the process group corresponding to the container from the directory to update the number of the target central processing unit core to the corresponding configuration information.

[0151] For example, in some further implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0152] Monitor the utilization of at least one target CPU core of the container; automatically adjust the number of target CPU cores bound to the container based on the utilization of each target CPU core and preset resource dynamic adjustment conditions; when the number of target CPU cores bound to the container changes, update the configuration information of the process group corresponding to the container again.

[0153] As an exemplary implementation of the above embodiment, the CPU core selection module 103 may also be used to:

[0154] Obtaining a maximum allowable CPU utilization rate and a minimum allowable CPU utilization rate; generating a preset resource dynamic adjustment condition based on a numerical relationship between the utilization rate of the target CPU core and the minimum allowable CPU utilization rate and the maximum allowable CPU utilization rate.

[0155] As another implementation of the above embodiment, the CPU core selection module 103 may also be used to:

[0156] When it is detected that the utilization of the current target CPU core is greater than or equal to the maximum allowed CPU utilization, at least one idle CPU core is selected from the non-uniform memory access group to which the target GPU belongs as a new target CPU core.

[0157] As another implementation of the above embodiment, the CPU core selection module 103 may also be used to:

[0158] When it is detected that the utilization of the current target CPU core is greater than or equal to the maximum allowable CPU utilization, if there are no idle CPU cores in the non-uniform memory access group to which the target GPU belongs, at least one idle CPU core is selected from other non-uniform memory access groups that have idle CPU cores to serve as a new target CPU core.

[0159] As another implementation of the above embodiment, the CPU core selection module 103 may also be used to:

[0160] When it is detected that the utilization rate of the current target CPU core is less than the minimum value allowed by the CPU utilization rate, the current target CPU core is unbound from the container. As another implementation of the above embodiment, the CPU core selection module 103 can also be used to:

[0161] When a utilization monitoring frequency configuration instruction is received, the utilization monitoring frequency is obtained; and the utilization of at least one target central processing unit core of the container is monitored according to the utilization monitoring frequency.

[0162] For example, in some further implementations of this embodiment, the CPU core selection module 103 may also be used to:

[0163] Obtain the expected reduction in the number of CPU cores adjusted for each execution unit of the task to be trained, and use the minimum number of CPU cores adjusted as the target number of CPU cores adjusted; adjust the CPU cores of each execution unit according to the target number of CPU cores adjusted, and send the target number of CPU cores adjusted to the business platform so that the business platform updates the resource configuration of the task to be trained.

[0164] The functions of the functional modules of the resource adjustment device of this embodiment can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0165] From the above, it can be seen that this embodiment can ensure the efficient completion of AI training tasks to the greatest extent by adjusting the resources in the AI ​​training task execution process.

[0166] The resource adjustment device mentioned above is described from the perspective of a functional module. Furthermore, the present application also provides an electronic device, which is described from the perspective of hardware. Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application under one embodiment. As shown in Figure 11, the electronic device includes a memory 110 for storing computer-readable instructions; a processor 111 for implementing the steps of the resource adjustment method mentioned in any of the above embodiments when executing computer-readable instructions.

[0167] The processor 111 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 111 may also be a controller, microcontroller, microprocessor, or other data processing chip. The processor 111 may be implemented in at least one hardware form: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor 111 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 111 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, the processor 111 may also include an AI (Artificial Intelligence) processor, which is used to handle computing operations related to machine learning.

[0168] The memory 110 may include one or more non-volatile computer-readable storage media storing computer-readable instructions, and the computer-readable storage medium may be non-transitory. The memory 110 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the memory 110 may be an internal storage unit of an electronic device, such as a hard disk of a server. In other embodiments, the memory 110 may also be an external storage device of an electronic device, such as a plug-in hard disk equipped on a server, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Furthermore, the memory 110 may also include both an internal storage unit of an electronic device and an external storage device. The memory 110 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the program in the process of executing the resource adjustment method, but can also be used to temporarily store data that has been output or is to be output. In this embodiment, memory 110 is configured to store at least the following computer-readable instructions 1101. When loaded and executed by processor 111, these computer-readable instructions can implement the relevant steps of the resource adjustment method disclosed in any of the aforementioned embodiments. Furthermore, resources stored in memory 110 may also include an operating system 1102 and data 1103, which may be stored in either a temporary or permanent manner. Operating system 1102 may include Windows, Unix, Linux, and the like. Data 1103 may include, but is not limited to, data corresponding to resource adjustment results.

[0169] In some embodiments, the electronic device may further include a display screen 112, an input / output interface 113, a communication interface 114 or a network interface, a power supply 115, and a communication bus 116. The display screen 112 and the input / output interface 113, such as a keyboard, are user interfaces. Exemplary user interfaces may also include standard wired interfaces, wireless interfaces, and the like. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or display unit, and is used to display information processed in the electronic device and to display a visual user interface. The communication interface 114 may exemplarily include a wired interface and / or a wireless interface, such as a Wi-Fi interface or a Bluetooth interface, and is typically used to establish a communication connection between the electronic device and other electronic devices. The communication bus 116 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. This bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, FIG11 shows only one thick line, but this does not mean that there is only one bus or one type of bus.

[0170] Those skilled in the art will appreciate that the structure shown in FIG11 does not limit the electronic device and may include more or fewer components than shown in the figure, for example, may also include a sensor 117 for implementing various functions.

[0171] The functions of the functional modules of the electronic device of this embodiment can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0172] From the above, it can be seen that this embodiment can ensure the efficient completion of AI training tasks to the greatest extent by adjusting the resources in the AI ​​training task execution process.

[0173] It is understandable that if the resource adjustment method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium and executes all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable ROM, a register, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a removable disk, a CD-ROM, a magnetic disk or an optical disk, and other media that can store program code.

[0174] Based on this, the present application also provides a readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the resource adjustment method in any of the above embodiments are performed.

[0175] This application also provides an artificial intelligence training platform, see Figure 12, which may include the following:

[0176] The AI ​​training platform may include a Kubernetes cluster 121, which includes a master node and nodes. The master, serving as the overall control center for Kubernetes, is responsible for managing and scheduling all resources in the cluster. The components running on each node are responsible for managing the lifecycle of the pods on the compute nodes and implementing service proxy functionality.

[0177] The artificial intelligence training platform also includes multiple resource adjustment devices 122, each of which is deployed on a computing node of the k8s cluster. A resource adjustment device can be deployed on each computing node of the k8s cluster, or multiple computing nodes can be selected and resource adjustment devices can be deployed for these multiple computing nodes, that is, at least one computing node in the k8s cluster is deployed with a resource adjustment device at the same time.

[0178] The functions of the various functional modules of the resource adjustment system in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0179] From the above, it can be seen that this embodiment can ensure the efficient completion of AI training tasks to the greatest extent by adjusting the resources in the AI ​​training task execution process.

[0180] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. The hardware disclosed in the embodiments, including devices and electronic devices, is described briefly because it corresponds to the methods disclosed in the embodiments. For relevant details, refer to the method description.

[0181] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0182] The above is a detailed introduction to a resource adjustment method, device, electronic device, readable storage medium and artificial intelligence training platform provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that based on the embodiments in the present application, for ordinary technicians in this technical field, all other embodiments obtained without creative work are within the scope of protection of the present application. Without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A resource adjustment method, characterized in that: include: Pre-built process group configuration files; The process group configuration file stores configuration information of the process group corresponding to the container, and the configuration information records at least the central processing unit core binding information of the corresponding container; Get the container set of the computing node that executes the task to be trained and its associated containers; According to the resource configuration information of the container set, the resource occupation information of the computing node and the user resource demand, in a manner that the target graphics processor and the target central processing unit core executing the task to be trained are preferentially located in the same non-uniform memory access group, selecting a target central processing unit core to be bound for the container; and Updating configuration information of the process group corresponding to the container to complete the binding of the central processing unit core and the container; Among them, the resource configuration information includes the graphics processor allocated to the container set, the user resource requirement includes the required number of graphics processors and the required number of central processing unit cores required by the user for executing the task to be trained, the target number of central processing unit cores is the same as the required number of central processing unit cores, and the target number of graphics processors is the same as the required number of graphics processors.

2. The resource adjustment method according to claim 1, characterized in that: The method of selecting a target CPU core to be bound for the container in a manner that a target GPU and a target CPU core executing the task to be trained are preferentially located in the same non-uniform memory access group includes: Obtaining the number of idle central processing unit cores and idle graphics processing unit cores of the computing node in the same target non-uniform memory access group at the current moment; and In response to the number of idle GPU cores being greater than or equal to the required GPU core number, and the number of idle CPU cores being greater than or equal to the required CPU core number, a target GPU core and a target CPU core are selected from the target NUMAM group.

3. The resource adjustment method according to claim 2, characterized in that: The method of selecting a target CPU core to be bound for the container in a manner that a target GPU and a target CPU core executing the task to be trained are preferentially located in the same non-uniform memory access group includes: In response to the number of idle GPU cores being greater than or equal to the required number of GPU cores and the number of idle CPU cores being less than the required number of CPU cores, a target GPU is selected from the target NUMAM group, all CPU cores are selected from the target NUMAM group, and remaining CPU cores are selected from the candidate NUMAM group to form a target CPU core.

4. The resource adjustment method according to claim 2, characterized in that: The obtaining of the number of idle central processing unit cores and the number of idle graphics processing unit cores of the computing node in the same target non-uniform memory access group at the current moment includes: Obtaining a CPU core-GPU topology relationship of the computing node; and Based on the CPU core-GPU topology relationship, the number of idle CPU cores and the number of idle GPU cores of the computing node in the same target NUMAG at the current moment are obtained.

5. The resource adjustment method according to claim 1, characterized in that: If there are multiple target CPU cores, all target GPUs are located in the same target NUMAG, and at least one target CPU core does not belong to the target NUMAG, after selecting the target CPU core to be bound for the container, the method further includes: monitoring a usage status of at least one candidate CPU core that is within the target NUMAG group but does not belong to the target CPU core; and In response to detecting that there is an idle target candidate CPU core, the target candidate CPU core is adjusted to the target CPU core while maintaining the container corresponding to the task to be trained in the current running state, and the CPU core that does not belong to the target inconsistency is released. The first target CPU core of the memory access group.

6. The resource adjustment method according to claim 5, characterized in that: The updating of the configuration information of the process group corresponding to the container includes: The number of the first target central processing unit core in the configuration information of the process group corresponding to the container is replaced with the number of the target candidate central processing unit core.

7. The resource adjustment method according to claim 1, characterized in that: The step of obtaining a container set of computing nodes that execute the task to be trained and their associated containers includes: Obtaining a computing node allocated to the task to be trained based on user resource requirements after the task to be trained is created on the business platform; the computing node has completed the creation of a container set with a medium-priority quality of service based on resource configuration information; Acquiring resource configuration information corresponding to the container set according to the topological relationship of the computing nodes; and A container corresponding to the container set is searched according to the resource configuration information.

8. The resource adjustment method according to claim 1, characterized in that: After obtaining the container set of the computing nodes that execute the task to be trained and the containers associated therewith, the following steps are included: In response to the container being a newly created container, according to the resource configuration information of the container set, the resource occupancy information of the computing node and the user resource requirements, a target central processing unit core to be bound to the container is selected in such a manner that the target graphics processor and the target central processing unit core that execute the task to be trained are preferentially located in the same non-uniform memory access group.

9. The resource adjustment method according to claim 1, characterized in that: After obtaining the container set of the computing nodes that execute the task to be trained and the containers associated therewith, the following steps are included: In response to the container not being a newly created container, determining whether a target graphics processor and a target central processor of the container belong to the same non-uniform memory access group; and In response to the target graphics processor and the target central processor of the container not belonging to the same non-uniform memory access group, a target processor core dynamic adjustment instruction is generated.

10. The resource adjustment method according to claim 1, characterized in that: After obtaining the container set of the computing nodes that execute the task to be trained and the containers associated therewith, the following steps are included: In response to the container not being a newly created container, when a CPU core quantity adjustment instruction is received, configuration information of a process group corresponding to the container is updated, and the updated configuration information of the process group is uploaded to the business platform.

11. The resource adjustment method according to claim 1, characterized in that: The updating of the configuration information of the process group corresponding to the container includes: Obtaining a directory of the host machine where the container is located; and The configuration information of the process group corresponding to the container is read from the directory to update the number of the target central processing unit core into the corresponding configuration information.

12. The resource adjustment method according to any one of claims 1 to 11, characterized in that: After the configuration information of the process group corresponding to the container is updated, the method further includes: monitoring utilization of at least one target central processing unit core of the container; Automatically adjusting the number of target CPU cores bound to the container according to the utilization rate of each target CPU core and a preset resource dynamic adjustment condition; and In response to a change in the number of target CPU cores bound to the container, the configuration information of the process group corresponding to the container is updated again.

13. The resource adjustment method according to claim 12, characterized in that: Before dynamically adjusting the conditions according to the utilization rate of each target CPU core and the preset resources, the method further includes: Obtaining a maximum value allowed for CPU utilization and a minimum value allowed for CPU utilization; and A preset resource dynamic adjustment condition is generated according to a numerical relationship between the utilization of the target CPU core and the minimum allowable CPU utilization and the maximum allowable CPU utilization.

14. The resource adjustment method according to claim 13, characterized in that: The method of automatically adjusting the number of target CPU cores bound to the container according to the utilization rate of each target CPU core and the preset resource dynamic adjustment condition includes: In response to detecting that the utilization of the current target CPU core is greater than or equal to the maximum allowed CPU utilization, at least one idle CPU core is selected from the non-uniform memory access group to which the target GPU belongs as a newly added target CPU core.

15. The resource adjustment method according to claim 13, characterized in that: The method of automatically adjusting the number of target CPU cores bound to the container according to the utilization rate of each target CPU core and the preset resource dynamic adjustment condition includes: In response to detecting that the utilization of the current target CPU core is greater than or equal to the maximum allowed CPU utilization, and in response to the non-uniform memory access group to which the target graphics processor belongs having no idle CPU cores, at least one idle CPU core is selected from other non-uniform memory access groups having idle CPU cores to serve as a newly added target CPU core.

16. The resource adjustment method according to claim 13, characterized in that: The method of automatically adjusting the number of target CPU cores bound to the container according to the utilization rate of each target CPU core and the preset resource dynamic adjustment condition includes: In response to detecting that the utilization of the current target CPU core is less than the minimum allowable CPU utilization, the current target CPU core is unbound from the container.

17. The resource adjustment method according to claim 12, characterized in that: The monitoring of the utilization of at least one target central processing unit core of the container includes: In response to receiving a utilization monitoring frequency configuration instruction, obtaining a utilization monitoring frequency; and The utilization of at least one target central processing unit core of the container is monitored according to the utilization monitoring frequency.

18. The resource adjustment method according to any one of claims 1 to 11, characterized in that: The task to be trained is a multi-machine distributed training task, and after the configuration information of the process group corresponding to the container is updated, the method further includes: Obtaining the expected reduction in the number of CPU cores adjusted for each execution unit of the task to be trained, and taking the minimum number of CPU cores adjusted as the target number of CPU cores adjusted; and The CPU cores of each execution unit are correspondingly adjusted according to the target CPU core adjustment number, and the target CPU core adjustment number is sent to the business platform so that the business platform updates the resource configuration of the task to be trained.

19. A resource adjustment device, characterized in that: include: Profile pre-building module, used to pre-build process group profiles; The process group configuration file stores configuration information of the process group corresponding to the container, and the configuration information records at least the central processing unit core binding information of the corresponding container; A container information acquisition module is used to obtain a container set of computing nodes that execute the task to be trained and its associated containers; a CPU core selection module, configured to select a target CPU core for binding for the container according to the resource configuration information of the container set, the resource occupancy information of the computing node and the user resource requirements, in a manner that the target GPU and the target CPU core for executing the task to be trained are preferentially located in the same non-uniform memory access group; wherein the resource configuration information includes the GPU allocated to the container set, the user resource requirements include the required number of GPUs and the required number of CPU cores required by the user for executing the task to be trained, the target number of CPU cores is the same as the required number of CPU cores, and the target number of GPUs is the same as the required number of GPUs; and The binding module is used to update the configuration information of the process group corresponding to the container to complete the binding of the central processing unit core and the container.

20. An electronic device, characterized in that: It comprises a processor and a memory, wherein the processor is used to implement the steps of the resource adjustment method according to any one of claims 1 to 18 when executing computer-readable instructions stored in the memory.

21. One or more non-volatile computer-readable storage media storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the resource adjustment method according to any one of claims 1 to 18.

22. An artificial intelligence training platform, characterized in that: Comprising a k8s cluster and a plurality of resource adjustment devices as described in claim 19; At least one computing node of the k8s cluster simultaneously deploys the resource adjustment device.

Citation Information

Patent Citations

  • Process access method and device based on NUMA node

    CN107436798A

  • Business scheduling method, apparatus and device, and readable storage medium

    CN110389843A

  • Method and device for binding kernel of CPU of Kubernetes container platform

    CN112052068A

  • Resource adjustment method and device, electronic equipment, storage medium and training platform

    CN117311990A

  • Hybrid aggregation for deep learning neural networks

    US20180253646A1

Cited By

  • Task scheduling method, device, equipment, medium, product, system and platform

    CN120540823A

  • Resource cross-layer collaborative management method and substrate manager

    CN120849140A

  • Training task scheduling method, server, storage medium and program product

    CN120851249A

  • Ollama reasoning optimization method for domestic platform

    CN121706995A

  • Data management method and device, computer equipment and storage medium

    CN121743312A