NUMA scheduling method, device and equipment in large model training scene and medium

By collecting the affiliation between NUMA nodes and the graphics processor in the large-scale training scenario, generating topology configuration files, and filtering and scoring the target NUMA nodes according to requirements, the problem of inefficient communication between NUMA nodes is solved, which improves training efficiency and reduces costs.

CN120560811APending Publication Date: 2025-08-29SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510706034.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In large-scale deep learning model training, the low communication efficiency between NUMA nodes leads to a decrease in bandwidth utilization and an increase in latency. The existing scheduling mechanism cannot effectively utilize hardware topology information, resulting in low training efficiency and increased cost.

Method used

Acquire the affiliation between each processor node in the target cluster and the graphics processor, generate topological relationship configuration files, filter candidate NUMA nodes according to the training task requirements, and determine the target NUMA nodes through performance communication scores to ensure that the training container is bound to the target node to avoid communication across NUMA nodes.

Benefits of technology

It improves the affinity of NUMA scheduling, improves the efficiency of large model training, reduces costs, and ensures communication efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560811A_ABST
    Figure CN120560811A_ABST
Patent Text Reader

Abstract

The invention discloses an NUMA scheduling method and device in a large model training scene, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: collecting a target cluster topological relation configuration file; obtaining a target affinity strategy corresponding to the graphics processor demand of the current large model training task, and if the target affinity strategy is a first affinity strategy, screening out candidate processor nodes including candidate NUMA nodes from the processor nodes based on the topological relation configuration file, the candidate NUMA nodes are the NUMA nodes of which the idle number of the graphics processor meets the requirements of the graphics processor under a single NUMA node, and determining a target NUMA node from the candidate NUMA nodes according to the performance communication score of each candidate NUMA node; and when the training container is started, scheduling each graphics processor under the target NUMA node to complete a current large model training task. The efficiency of large model training is improved, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a NUMA scheduling method, apparatus, device, and medium in a large model training scenario. Background Art

[0002] Large-scale deep learning model training has entered the era of hundreds of billions of parameters, and collective communication efficiency has become a key bottleneck restricting training performance. Gradient synchronization algorithms, such as AllReduce, require full parameter aggregation across graphics processing units (GPUs). This communication overhead accounts for 30%-50% of the iteration time in typical training tasks. In heterogeneous computing clusters based on the PCIe (Peripheral Component Interconnect Express) 4.0 architecture, inter-GPU communication across NUMA (Non-Uniform Memory Access) nodes suffers from significant performance degradation. This is because the communication path must traverse PCIe switches and transit through the central processing unit (CPU), resulting in reduced effective bandwidth utilization, nonlinear increases in transmission latency, and increased contention for interconnect resources.

[0003] On the other hand, Kubernetes clusters have become the de facto standard for distributed training orchestration of large models. However, the default Kubernetes scheduler uses a coarse-grained resource allocation strategy based on the number of GPUs, unaware of the physical constraints of the hardware topology. A typical scenario involves a dual-socket CPU server architecture with eight GPUs assigned to two physical nodes, NUMA0 and NUMA1. However, existing device plugins (such as the NVIDIA Device Plugin) only report the discrete number of GPU resources and lack key topology information such as NUMA affinity and PCIe topology. This topology-agnostic scheduling mechanism results in inefficient training tasks.

[0004] The shortcomings of the existing scheduling mechanism directly lead to bandwidth degradation. Due to complex paths, protocol conversion and other overheads, cross-NUMA communication has lower bandwidth utilization than the same-NUMA nodes, and latency accumulates. Taking the Ring-AllReduce algorithm as an example, if the task GPUs are distributed across NUMA, the number of communication link hops increases, and the latency of a single iteration increases.

[0005] In summary, how to improve the affinity of NUMA scheduling to improve the efficiency and reduce the cost of large model training is a problem to be solved in this field. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a NUMA scheduling method, apparatus, device, and medium for large-model training scenarios, which improves the affinity of NUMA scheduling to improve the efficiency and reduce the cost of large-model training. The specific solution is as follows:

[0007] In a first aspect, the present application discloses a NUMA scheduling method for large model training scenarios, including:

[0008] Collect the affiliation relationships between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file;

[0009] Get the target affinity policy corresponding to the GPU requirements of the current large model training task;

[0010] If the target affinity policy is the first affinity policy, candidate processor nodes including a candidate NUMA node that satisfies the first affinity policy are screened from the processor nodes based on the topology relationship profile; wherein the candidate NUMA node that satisfies the first affinity policy is a NUMA node in which the number of idle graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor demand;

[0011] Determine a target NUMA node from each of the candidate NUMA nodes according to the performance communication score of each of the candidate NUMA nodes;

[0012] The training container is bound to the target NUMA node so that when the training container is started, each graphics processor under the target NUMA node is scheduled to complete the current large model training task.

[0013] Optionally, collecting the affiliation relationships between different NUMA nodes under each processor node in the target cluster and each graphics processor to generate a topology relationship configuration file includes:

[0014] Deploy the Device Plugin application to the GPU of the target cluster as a DaemonSet, and use the NVML library to obtain the affiliation between different NUMA nodes and each GPU under each processor node.

[0015] According to the affiliation relationship, each graphics processor number and the corresponding NUMA node number are encoded into a key-value pair to generate a topology relationship configuration file.

[0016] Optionally, determining the target NUMA node from the candidate NUMA nodes according to the performance communication score of each candidate NUMA node includes:

[0017] Scoring each candidate NUMA node based on the continuity of the graphics processor numbers corresponding to each candidate NUMA node to obtain a communication score for each candidate NUMA node;

[0018] Scoring each of the candidate NUMA nodes based on the load balancing status of each of the candidate NUMA nodes to obtain a performance score of each of the candidate NUMA nodes;

[0019] The sum of the communication score and the performance score is determined as the performance communication score of each candidate NUMA node, and a target NUMA node is determined from each candidate NUMA node according to the performance communication score.

[0020] Optionally, the NUMA scheduling method in the large model training scenario further includes:

[0021] If the target affinity policy is the second affinity policy, the total number of idle graphics processors under all NUMA nodes in each processor node is determined based on the topology relationship configuration file, and candidate processor nodes whose total idle number meets the graphics processor requirement are screened out from each processor node.

[0022] Optionally, the NUMA scheduling method in the large model training scenario further includes:

[0023] Arrange the NUMA nodes in the current candidate processor nodes in descending order according to the number of idle GPUs under each NUMA node to obtain a NUMA node sequence, and determine the current NUMA node from the NUMA node sequence;

[0024] Determine whether the total number of idle graphics processors from the first NUMA node to the current NUMA node in the NUMA node sequence meets the graphics processor requirement;

[0025] If the total number of idle GPUs does not meet the GPU requirement, a new current NUMA node is determined from the NUMA node sequence, and the process jumps again to the step of determining whether the total number of idle GPUs from the first NUMA node to the current NUMA node in the NUMA node sequence meets the GPU requirement.

[0026] If the total number of idle nodes meets the graphics processor requirement, the first NUMA node to the current NUMA node in the NUMA node sequence are determined as candidate NUMA nodes that meet the second affinity policy.

[0027] Optionally, the NUMA scheduling method in the large model training scenario further includes:

[0028] The target processor node monitors the scheduling request of the training container based on the ListWatch mechanism, and binds the training container to the target NUMA node to limit the training container to access only the graphics processor bound to the target NUMA node; wherein, the target processor node is the processor node corresponding to the target NUMA node.

[0029] Optionally, the process of scheduling each graphics processor under the target NUMA node to complete the current large model training task further includes:

[0030] Sending the number list of the graphics processors bound to the target NUMA node to the container runtime through the target processor node;

[0031] The device file of the graphics processor bound to the target NUMA node is mounted to the training container during container runtime, and the graphics processor visibility environment variable is set according to the number list.

[0032] In a second aspect, the present application discloses a NUMA scheduling device for large model training scenarios, comprising:

[0033] The configuration generation module is used to collect the affiliation relationship between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file;

[0034] A strategy acquisition module is used to obtain the target affinity strategy corresponding to the GPU requirements of the current large model training task;

[0035] a first node determination module configured to, if the target affinity policy is a first affinity policy, filter out candidate processor nodes from each of the processor nodes based on the topology relationship configuration file, including a candidate NUMA node that satisfies the first affinity policy; wherein the candidate NUMA node that satisfies the first affinity policy is a NUMA node having a number of idle graphics processors under a single NUMA node that is not less than the number of graphics processors corresponding to the graphics processor demand;

[0036] A second node determination module is configured to determine a target NUMA node from each of the candidate NUMA nodes according to a performance communication score of each of the candidate NUMA nodes;

[0037] The scheduling module is used to bind the training container to the target NUMA node so that when the training container is started, the graphics processors under the target NUMA node are scheduled to complete the current large model training task.

[0038] In a third aspect, the present application discloses an electronic device, comprising:

[0039] Memory, used to store computer programs;

[0040] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed NUMA scheduling method in the large model training scenario.

[0041] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned NUMA scheduling method in the large model training scenario are implemented.

[0042] The beneficial effects of the present application are as follows: the present application collects the affiliation between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship profile; obtains a target affinity strategy corresponding to the graphics processor requirement of the current large model training task; if the target affinity strategy is a first affinity strategy, then based on the topology relationship profile, candidate processor nodes including candidate NUMA nodes that meet the first affinity strategy are screened out from each of the processor nodes; wherein the candidate NUMA node that meets the first affinity strategy is a NUMA node whose idle number of graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor requirement; determines the target NUMA node from each of the candidate NUMA nodes based on the performance communication score of each of the candidate NUMA nodes; binds the training container to the target NUMA node so that when the training container is started, the graphics processors under the target NUMA node are scheduled to complete the current large model training task.It can be seen that the present application collects the affiliation between NUMA nodes and graphics processors, and then selects suitable candidate processor nodes from each processor node according to the target affinity strategy corresponding to the graphics processor requirements of the current large model training task. Specifically, if the target affinity strategy is the first affinity strategy, suitable candidate processor nodes are selected from each processor node based on the topology relationship configuration file, that is, candidate processor nodes of NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements. In other words, the candidate processor nodes include NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements, and the NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements are determined as candidate NUMA nodes, that is, graphics processors that need to cross NUMA nodes are excluded. Then, the graphics processors under the target NUMA node finally determined will not cross NUMA nodes. In the case of a point, the training container is bound to the target NUMA node. In this way, when the training container is started, there will be no cross-NUMA node situation when scheduling the various graphics processors under the target NUMA node to complete the current large model training task, thereby improving the affinity of NUMA scheduling. That is to say, because there is no need to communicate across NUMA nodes in the process of completing the large model training task, each graphics processor under the target NUMA node belongs to the target NUMA node, and the communication efficiency is higher, avoiding the overhead of complex paths and protocol conversion caused by cross-NUMA communication, and reducing costs; in addition, the present application determines the target NUMA node through two-level screening, that is, first screening according to the graphics processor requirements, so that each candidate NUMA node meets the graphics processor requirements, and secondly screening according to the performance communication score. In this way, the target NUMA node finally determined not only meets the graphics processor requirements, but also ensures good performance communication during large model training, which can further improve the efficiency of large model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0044] Figure 1 This is a flowchart of a NUMA scheduling method for large model training scenarios disclosed in this application;

[0045] Figure 2 A schematic diagram of a specific NUMA affinity scheduling architecture disclosed in this application;

[0046] Figure 3 This is a schematic diagram of the structure of a NUMA scheduling device in a large model training scenario disclosed in this application;

[0047] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] Large-scale deep learning model training has entered the era of hundreds of billions of parameters, and collective communication efficiency has become a key bottleneck restricting training performance. Gradient synchronization algorithms, such as AllReduce, require full parameter aggregation across GPUs, and their communication overhead accounts for 30%-50% of the iteration time in typical training tasks. In heterogeneous computing clusters based on the PCIe 4.0 architecture, inter-GPU communication across NUMA nodes suffers from significant performance degradation. This is because the communication path must traverse PCIe switches and transit through the central processing unit (CPU), resulting in reduced effective bandwidth utilization, nonlinear increases in transmission latency, and increased competition for interconnect resources.

[0050] Kubernetes clusters have become the de facto standard for distributed training of large models. However, the default Kubernetes scheduler uses a coarse-grained resource allocation strategy based on the number of GPUs, unaware of the physical constraints of the hardware topology. A typical scenario involves a dual-socket CPU server architecture with eight GPUs assigned to two physical nodes, NUMA0 and NUMA1. However, the existing Device Plugin only reports the discrete number of GPU resources, lacking key topology information such as NUMA affinity and PCIe topology levels. This topology-agnostic scheduling mechanism results in inefficient training tasks.

[0051] The shortcomings of the existing scheduling mechanism directly lead to bandwidth degradation. Due to complex paths, protocol conversion and other overheads, cross-NUMA communication has lower bandwidth utilization than the same-NUMA nodes, and latency accumulates. Taking the Ring-AllReduce algorithm as an example, if the task GPUs are distributed across NUMA, the number of communication link hops increases, and the latency of a single iteration increases.

[0052] To this end, this application provides a NUMA scheduling solution for large model training scenarios, which improves the affinity of NUMA scheduling to improve the efficiency and reduce the cost of large model training.

[0053] See also Figure 1 As shown, the embodiment of the present application discloses a NUMA scheduling method in a large model training scenario, including:

[0054] Step S11: collecting the affiliation relationships between different NUMA nodes under each processor node in the target cluster and each graphics processor to generate a topology relationship configuration file.

[0055] In this embodiment, collecting the affiliation between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file includes: deploying the Device Plugin application to the graphics processor of the target cluster in the form of a DaemonSet, and obtaining the affiliation between different NUMA nodes and each graphics processor under each processor node through the NVML library; encoding each graphics processor number and the corresponding NUMA node number into a key-value pair based on the affiliation to generate the topology relationship configuration file.

[0056] A processor node includes multiple NUMA nodes, and different graphics processors may belong to different NUMA nodes. Before training large models, it is necessary to collect the topological relationship between GPUs and NUMA nodes to complete the node annotation initialization of the processor nodes. Based on the NVIDIA Device Plugin, a custom Device Plugin module is implemented. This plugin is deployed as a DaemonSet on each GPU node in the target cluster (specifically, a Kubernetes cluster) and registers the nvidia.com / gpu resource with the Kubelet device manager. The custom extension module is started as a "Goroutine" in the NVIDIA Device Plugin main method. After the custom module is started, it calls the NVML Go bindings library (Go bindings) to obtain the relationship between the node's GPU and NUMA nodes and generate a node-level NUMA topology configuration file. After the custom extension module generates the node GPU and NUMA node topology configuration, the topology configuration file is written into the node annotations of each processor node through the PATCH interface of the API (Application Programming Interface) Server in the target cluster in annotation mode. The NUMA node number in the topology configuration file is used as the node annotation key, and the list of graphics processor numbers under different NUMA nodes is used as the node annotation value. The example is as follows:

[0057] Annotations:

[0058] nvidia.com / gpu-topology.NUMA0: "gpu0,gpu1,gpu2,gpu3";

[0059] nvidia.com / gpu-topology.NUMA1: "gpu4,gpu5,gpu6,gpu7";

[0060] After the Device Plugin reports the NUMA topology annotation, it monitors GPU status changes (such as device offline or hot plug) in real time and uses an event-driven batch update mechanism to reduce the API server load.

[0061] Step S12: Obtain a target affinity strategy corresponding to the graphics processor requirements of the current large model training task.

[0062] Configure the GPU affinity scheduling policy, also known as the target affinity policy. Declare the NUMA scheduling policy (i.e., target affinity policy) in the metadata.annotations of the pod (container). The annotation key is NUMA.affinity.policy (i.e., target affinity policy). The corresponding values ​​include "strict (i.e., first affinity policy)" and "besteffort (i.e., second affinity policy)." When the value is strict, it means that the pod is strictly bound to the same NUMA node. If all GPU nodes in the cluster cannot meet the resource allocation requirement, the pod is placed in the Pending state until the resource allocation requirement is met. When the value is besteffort, it prioritizes scheduling to nodes that can meet the scheduling requirement. If no node meets the resource allocation requirement, but a node with the required number of GPUs exists, the pod is scheduled to that node. Configure the resource request and specify the desired number of GPUs in the Pod's spec.containers.resources.limits (container resource limit field). When requesting 4 GPUs, set the value of spec.containers.resources.limits.nvidia.com / gpu to 4.

[0063] Step S13: If the target affinity policy is the first affinity policy, candidate processor nodes including candidate NUMA nodes that meet the first affinity policy are screened out from each of the processor nodes based on the topology relationship profile; wherein the candidate NUMA node that meets the first affinity policy is a NUMA node in which the number of idle graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor demand.

[0064] The configured target affinity policy can be the first affinity policy or the second affinity policy. Regardless of which affinity policy is used, candidate NUMA nodes that meet the GPU requirements must be screened. Under the first affinity policy, candidate NUMA nodes are NUMA nodes where the number of idle GPUs under a single NUMA node is not less than the number of GPUs corresponding to the GPU requirement. Under the second affinity policy, candidate NUMA nodes are NUMA nodes where the total number of idle GPUs under all candidate NUMA nodes meets the GPU requirement.

[0065] In the first specific embodiment, the target affinity policy is the first affinity policy, which is the strict policy, indicating strict binding to the same NUMA node. When all GPU nodes in the cluster cannot meet the resource allocation requirements, the Pod is in a pending state until the resource allocation requirements are met. For example, the affiliation in the topology configuration file indicates that gpu0, gpu1, gpu2, and gpu3 belong to the NUMA0 node, while gpu4, gpu5, and gpu6 belong to the NUMA1 node. The number of idle graphics processors in the NUMA0 node is 4, and the number of idle graphics processors in the NUMA1 node is 2. The graphics processor requirement indicates that the current large model training task requires 3 idle graphics processors. Then the NUMA0 node is a candidate NUMA node that meets the first affinity policy. For another example, the target cluster includes processor node A and processor node B, where processor node A includes A-NUMA0 node and A-NUMA1 node, and both A-NUMA0 node and A-NUMA1 node include 2 idle graphics processors. Processor node B includes B-NUMA0 node and B-NUMA1 node, and B-NUMA0 node includes 4 idle graphics processors and B-NUMA1 node includes 2 idle graphics processors. The graphics processor requirement indicates that the current large model training task requires 3 idle graphics processors. Then only processor node B is the candidate processor node, and only B-NUMA0 node is the candidate NUMA node that meets the first affinity policy. In this way, the target NUMA node is subsequently screened out from the candidate NUMA nodes, and the graphics processor in the target NUMA node can be used to complete the current large model training task, without the need for graphics processors across NUMA nodes to complete the training task. It should be noted that under the first affinity policy, when all GPU nodes in the cluster cannot meet the resource allocation requirements, the Pod is in the Pending state. Therefore, in order to prevent this situation, it is necessary to obtain the target affinity policy corresponding to the graphics processor requirements of the current large model training task. If the target affinity policy is the first affinity policy, that is, the graphics processor requirements of the current large model training task can be met under the first affinity policy, avoiding the situation where the Pod is in the Pending state after binding.

[0066] In a second specific embodiment, it also includes: if the target affinity policy is the second affinity policy, determining the total number of idle graphics processors under all NUMA nodes in each of the processor nodes based on the topology relationship profile, and screening out candidate processor nodes from each of the processor nodes whose total idle number meets the graphics processor requirement.

[0067] It is understandable that after considering the GPU requirements of the current large model training task, the target affinity strategy is determined to be the second affinity strategy. This may mean that if the first affinity strategy is used, the Pod will be in a pending state. Therefore, in order to avoid this situation, the target affinity strategy is determined to be the second affinity strategy. The second affinity strategy indicates that priority is given to scheduling to nodes that can meet the scheduling requirements. When there are no nodes that meet the resource allocation requirements, but there are nodes that meet the GPU quantity requirements, they are scheduled to the corresponding nodes. In other words, if the target affinity strategy is the second affinity strategy, the total number of idle GPUs under all NUMA nodes in each processor node is determined based on the topology relationship configuration file, and candidate processor nodes whose total number of idle GPUs meets the GPU requirements are screened from each processor node. That is, the second affinity strategy is not as strict as the first affinity strategy. As long as the total number of all idle GPUs in the candidate processor nodes meets the GPU requirements, the idle GPUs can be from different NUMA nodes. Idle graphics processors that meet the graphics processor requirements are filtered out from the candidate processor nodes, and then the NUMA nodes to which these idle graphics processors belong are determined as candidate NUMA nodes. For example, the candidate processor nodes include the B-NUMA0 node and the B-NUMA1 node. The B-NUMA0 node includes 2 idle graphics processors, and the B-NUMA1 node includes 2 idle graphics processors. The graphics processor requirement indicates that the current large model training task requires 3 idle graphics processors, so 2 idle graphics processors in the B-NUMA0 node and 1 idle graphics processor in the B-NUMA1 node are required.

[0068] If the target affinity policy is the second affinity policy, this embodiment further includes: arranging each NUMA node in the current candidate processor node in descending order according to the number of idle graphics processors under a single NUMA node to obtain a NUMA node sequence, and determining the current NUMA node from the NUMA node sequence; determining whether the total number of idle graphics processors from the first NUMA node to the current NUMA node in the NUMA node sequence meets the graphics processor requirement; if the total number of idle graphics processors does not meet the graphics processor requirement, determining a new current NUMA node from the NUMA node sequence, and jumping back to the step of determining whether the total number of idle graphics processors from the first NUMA node to the current NUMA node in the NUMA node sequence meets the graphics processor requirement; if the total number of idle graphics processors meets the graphics processor requirement, determining the first NUMA node to the current NUMA node in the NUMA node sequence as candidate NUMA nodes that meet the second affinity policy. That is to say, although the second affinity policy is not as strict as the first affinity policy, as long as the total number of all idle GPUs in the candidate processor node meets the GPU demand, the NUMA node with more idle GPUs is preferred when determining the candidate NUMA node. That is, the NUMA nodes are sorted in descending order according to the number of idle GPUs to obtain a NUMA node sequence. For example, the NUMA node sequence is [NUMA1, NUMA0, NUMA2], where NUMA1 includes 4 idle GPUs, NUMA0 includes 3 idle GPUs, and NUMA2 includes 2 idle GPUs. The graphics processor requirement indicates that the current large model training task requires 6 idle graphics processors. Although the 2 idle graphics processors in NUMA2 plus the 3 idle graphics processors in NUMA0 and the 1 idle graphics processor in NUMA1 can meet the graphics processor requirement, the graphics processors need to cross 3 NUMA nodes. If only the 4 idle graphics processors in NUMA1 and the 2 idle graphics processors in NUMA0 are used as graphics processors for processing training tasks, the number of cross-NUMA nodes will be less, and the processing efficiency will be higher.

[0069] Step S14: determining a target NUMA node from the candidate NUMA nodes according to the performance communication score of each candidate NUMA node.

[0070] In this embodiment, determining the target NUMA node from each of the candidate NUMA nodes based on the performance communication score of each of the candidate NUMA nodes includes: scoring each of the candidate NUMA nodes based on the continuity of the graphics processor numbers corresponding to each of the candidate NUMA nodes to obtain a communication score of each of the candidate NUMA nodes; scoring each of the candidate NUMA nodes based on the load balancing of each of the candidate NUMA nodes to obtain a performance score of each of the candidate NUMA nodes; determining the sum of the communication score and the performance score as the performance communication score of each of the candidate NUMA nodes, and determining the target NUMA node from each of the candidate NUMA nodes based on the performance communication score.

[0071] The selected candidate NUMA nodes are scored, including a communication score and a performance score. The communication score reflects the continuity of the GPU numbers corresponding to each candidate NUMA node. More consecutive GPU numbers within the same candidate NUMA node result in fewer PCIe sublink switches, leading to a higher communication score. For example, 40 points are added for every four consecutive GPUs. The score increases with more consecutive GPU numbers within the same NUMA node, and the score increases by a weighted value related to n for each consecutive allocation of n GPUs. The performance score reflects the load balance of each candidate NUMA node, determined by combining CPU or memory utilization within the NUMA node. Lower CPU or memory utilization indicates less load, resulting in a higher performance score for the candidate NUMA node. The introduction of a continuity score (which prioritizes consecutively numbered GPUs to reduce PCIe sublink switches) and a load balance score (which dynamically adjusts priorities based on CPU / memory utilization within the NUMA node) can address the resource fragmentation and load imbalance issues inherent in traditional scheduling strategies. Furthermore, the sum of the communication score and the performance score is determined as the performance communication score of each candidate NUMA node, and the target NUMA node is determined from the candidate NUMA nodes according to the performance communication score, that is, the node with the highest score is selected as the target NUMA node.

[0072] Step S15: Bind the training container to the target NUMA node so that when the training container is started, each graphics processor under the target NUMA node is scheduled to complete the current large model training task.

[0073] Write the target NUMA node information (such as allocated-numa: "0") into the annotation of the training container to bind the training container to the target NUMA node.

[0074] In this embodiment, it also includes: monitoring the scheduling request of the training container based on the ListWatch mechanism through the target processor node, and binding the training container to the target NUMA node to limit the training container to access only the graphics processor bound to the target NUMA node; wherein, the target processor node is the processor node corresponding to the target NUMA node.

[0075] The Pod's spec.nodeName field is updated through the API Server, triggering the Kubelet to perform resource binding. The processor node to which the target NUMA node belongs is the target processor node. The target processor node (i.e., the target Kubelet node) monitors scheduling requests for training containers based on the ListWatch mechanism. After the Kubelet detects that the Pod has been scheduled to this node based on the ListWatch mechanism, it calls the Allocate interface of the custom DevicePlugin based on the number of nvidia.com / gpu resources requested by the container in the Pod to allocate GPU resources to the Pod container. The training container is then bound to the target NUMA node, limiting the training container to access only the GPUs bound to the target NUMA node. It is understandable that if the target affinity policy is the first affinity policy, the GPUs accessed by the training container all belong to the same NUMA node, preventing cross-NUMA node access and significantly reducing communication efficiency.

[0076] In this embodiment, the process of scheduling each graphics processor under the target NUMA node to complete the current large model training task also includes: sending the number list of the graphics processors bound to the target NUMA node to the container runtime through the target processor node; mounting the device file of the graphics processor bound to the target NUMA node to the training container through the container runtime, and setting the graphics processor visibility environment variable according to the number list.

[0077] After receiving the Allocate request, the custom Device Plugin returns a list of GPUs in the target NUMA node (such as gpu0, gpu1, gpu2, and gpu3). That is, the target processor node sends the numbered list of graphics processors bound to the target NUMA node to the container runtime. The container runtime automatically sets the NVIDIA_VISIBLE_DEVICES environment variable (that is, the graphics processor visibility environment variable) based on the GPU list allocated by the Device Plugin to ensure that the container only recognizes the target GPU.

[0078] Bind the training container to the target NUMA node so that when the training container starts, the GPUs under the target NUMA node are scheduled to complete the current large model training task. The specific process is as follows:

[0079] First, Kubelet applies for the resources required by the Pod container. The implementation details include the following:

[0080] 1) Listening for Pod Scheduling Events: After Kubelet detects that a Pod is being scheduled to this node through the ListWatch mechanism, it calls the Allocate interface of the custom Device Plugin to allocate GPU resources to the Pod container based on the number of nvidia.com / gpu resources requested by the container in the Pod.

[0081] 2) Device allocation and binding: After receiving the Allocate request, the custom Device Plugin returns a list of GPUs in the target NUMA node (such as gpu0, gpu1, gpu2, gpu3).

[0082] 3) Call the container runtime and start the container: Based on the container resource request of the Pod, obtain the resources required by the container and call the container runtime to start the container declared by the Pod.

[0083] Next, the container runtime starts the container. The implementation details include the following:

[0084] 1) The container runtime (such as containerd) calls NVIDIA Container Runtime based on the device information delivered by Kubelet and mounts the GPU device file (such as / dev / nvidia0-3) into the container.

[0085] 2) The Allocate interface of the NVIDIA Device Plugin returns a list of devices, and the container runtime automatically sets the NVIDIA_VISIBLE_DEVICES environment variable based on the GPU list allocated by the Device Plugin to ensure that the container only recognizes the target GPU.

[0086] The beneficial effects of the present application are as follows: the present application collects the affiliation between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship profile; obtains a target affinity strategy corresponding to the graphics processor requirement of the current large model training task; if the target affinity strategy is a first affinity strategy, then based on the topology relationship profile, candidate processor nodes including candidate NUMA nodes that meet the first affinity strategy are screened out from each of the processor nodes; wherein the candidate NUMA node that meets the first affinity strategy is a NUMA node whose idle number of graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor requirement; determines the target NUMA node from each of the candidate NUMA nodes based on the performance communication score of each of the candidate NUMA nodes; binds the training container to the target NUMA node so that when the training container is started, the graphics processors under the target NUMA node are scheduled to complete the current large model training task.It can be seen that the present application collects the affiliation between NUMA nodes and graphics processors, and then selects suitable candidate processor nodes from each processor node according to the target affinity strategy corresponding to the graphics processor requirements of the current large model training task. Specifically, if the target affinity strategy is the first affinity strategy, suitable candidate processor nodes are selected from each processor node based on the topology relationship configuration file, that is, candidate processor nodes of NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements. In other words, the candidate processor nodes include NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements, and the NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements are determined as candidate NUMA nodes, that is, graphics processors that need to cross NUMA nodes are excluded. Then, the graphics processors under the target NUMA node finally determined will not cross NUMA nodes. In the case of a point, the training container is bound to the target NUMA node. In this way, when the training container is started, there will be no cross-NUMA node situation when scheduling the various graphics processors under the target NUMA node to complete the current large model training task, thereby improving the affinity of NUMA scheduling. That is to say, because there is no need to communicate across NUMA nodes in the process of completing the large model training task, each graphics processor under the target NUMA node belongs to the target NUMA node, and the communication efficiency is higher, avoiding the overhead of complex paths and protocol conversion caused by cross-NUMA communication, and reducing costs; in addition, the present application determines the target NUMA node through two-level screening, that is, first screening according to the graphics processor requirements, so that each candidate NUMA node meets the graphics processor requirements, and secondly screening according to the performance communication score. In this way, the target NUMA node finally determined not only meets the graphics processor requirements, but also ensures good performance communication during large model training, which can further improve the efficiency of large model training.

[0087] Below is Figure 2 The application is described using a specific NUMA affinity scheduling architecture diagram as an example. The overall architecture includes a custom device plugin module, a Kubelet module, a Kube-apiserver module, a custom scheduler plugin module, and a container runtime module.

[0088] The custom Device Plugin module extends the NVIDIA DevicePlugin plug-in to detect GPU NUMA topology and generate a hardware configuration description on the node. It then registers device resources with the Kubelet module and, based on the mapping between physical GPU locations and NUMA nodes, updates the Kube-apiserver module with node annotations that include the relationship between NUMA nodes and GPU numbers. This allows the custom scheduler plug-in module to implement GPU hardware resource topology awareness and resource allocation based on the annotations.

[0089] The Kubelet (processor node in the target cluster) module is responsible for implementing NUMA resource binding and isolation policies. Based on Kubernetes' ListWatch mechanism, it observes the load of pods scheduled to this node. Based on the custom scheduler's decision and the NUMA topology configuration file, it calls the custom Device Plugin interface to request target GPU resources. Furthermore, through the cgroup mechanism, the pod's CPU, memory, and GPU are confined to the same NUMA node, ensuring physical consistency of resource access at runtime.

[0090] The Kube-apiserver (API server) module serves as the global control center and unified access point for the Kubernetes cluster, exposing cluster operation interfaces via a RESTful API. It synchronizes resource status change events in real time using a Watch mechanism, driving the collaboration between components such as the scheduler (kube-scheduler), controller (kube-controller-manager), and node Kubelet to ensure that the cluster state is consistent with the user's desired state. It receives and stores GPU topology information reported by custom device plugins and Kubelet modules, and simultaneously synchronizes resource information and change events to custom scheduler plugins and Kubelet modules in real time.

[0091] A custom scheduler plug-in module that parses NUMA topology configuration files and executes affinity scheduling policies. Based on the NUMA constraints declared by the pod (such as strict binding and cross-node fault tolerance) and the GPU topology status of the target node, it screens candidate nodes that meet the conditions, calculates their priorities, and ultimately determines the optimal node for pod scheduling.

[0092] A container runtime module that loads NUMA resource isolation configuration during container startup. Based on the device mount instructions and cgroup policy files issued by Kubelet, it binds GPU devices, CPU cores, and memory regions to the container process.

[0093] See also Figure 3 As shown, the embodiment of the present application discloses a NUMA scheduling device in a large model training scenario, including:

[0094] The configuration generation module 11 is used to collect the affiliation relationship between different NUMA nodes under each processor node in the target cluster and each graphics processor to generate a topology relationship configuration file;

[0095] A strategy acquisition module 12 is used to obtain a target affinity strategy corresponding to the GPU requirements of the current large model training task;

[0096] A first node determination module 13 is configured to, if the target affinity policy is a first affinity policy, filter out candidate processor nodes including a candidate NUMA node that satisfies the first affinity policy from each of the processor nodes based on the topology relationship profile; wherein the candidate NUMA node that satisfies the first affinity policy is a NUMA node having a number of idle graphics processors under a single NUMA node that is not less than the number of graphics processors corresponding to the graphics processor demand;

[0097] A second node determination module 14 is configured to determine a target NUMA node from each of the candidate NUMA nodes according to the performance communication score of each of the candidate NUMA nodes;

[0098] The scheduling module 15 is used to bind the training container to the target NUMA node so that when the training container is started, each graphics processor under the target NUMA node is scheduled to complete the current large model training task.

[0099] The beneficial effects of the present application are as follows: the present application collects the affiliation between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship profile; obtains a target affinity strategy corresponding to the graphics processor requirement of the current large model training task; if the target affinity strategy is a first affinity strategy, then based on the topology relationship profile, candidate processor nodes including candidate NUMA nodes that meet the first affinity strategy are screened out from each of the processor nodes; wherein the candidate NUMA node that meets the first affinity strategy is a NUMA node whose idle number of graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor requirement; determines the target NUMA node from each of the candidate NUMA nodes based on the performance communication score of each of the candidate NUMA nodes; binds the training container to the target NUMA node so that when the training container is started, the graphics processors under the target NUMA node are scheduled to complete the current large model training task.It can be seen that the present application collects the affiliation between NUMA nodes and graphics processors, and then selects suitable candidate processor nodes from each processor node according to the target affinity strategy corresponding to the graphics processor requirements of the current large model training task. Specifically, if the target affinity strategy is the first affinity strategy, suitable candidate processor nodes are selected from each processor node based on the topology relationship configuration file, that is, candidate processor nodes of NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements. In other words, the candidate processor nodes include NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements, and the NUMA nodes where the idle number of graphics processors under a single NUMA node meets the graphics processor requirements are determined as candidate NUMA nodes, that is, graphics processors that need to cross NUMA nodes are excluded. Then, the graphics processors under the target NUMA node finally determined will not cross NUMA nodes. In the case of a point, the training container is bound to the target NUMA node. In this way, when the training container is started, there will be no cross-NUMA node situation when scheduling the various graphics processors under the target NUMA node to complete the current large model training task, thereby improving the affinity of NUMA scheduling. That is to say, because there is no need to communicate across NUMA nodes in the process of completing the large model training task, each graphics processor under the target NUMA node belongs to the target NUMA node, and the communication efficiency is higher, avoiding the overhead of complex paths and protocol conversion caused by cross-NUMA communication, and reducing costs; in addition, the present application determines the target NUMA node through two-level screening, that is, first screening according to the graphics processor requirements, so that each candidate NUMA node meets the graphics processor requirements, and secondly screening according to the performance communication score. In this way, the target NUMA node finally determined not only meets the graphics processor requirements, but also ensures good performance communication during large model training, which can further improve the efficiency of large model training.

[0100] Furthermore, an embodiment of the present application also provides an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.

[0101] Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Specifically, the device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the NUMA scheduling method for large model training scenarios performed by the electronic device as disclosed in any of the aforementioned embodiments.

[0102] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0103] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0104] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.

[0105] Among them, the operating system 221 is used to manage and control the various hardware devices and computer programs 222 on the electronic device to enable the processor 21 to calculate and process the massive data 223 in the memory 22. It can be Windows, Unix, Linux, etc. In addition to including a computer program that can be used to complete the NUMA scheduling method disclosed in any of the aforementioned embodiments and executed by the electronic device in the large model training scenario, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to including data transmitted from an external device to the electronic device, the data 223 can also include data collected by its own input and output interface 25.

[0106] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned NUMA scheduling method for large-model training scenarios. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.

[0107] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0108] Professionals may further appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in random access memory (RAM), memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), registers, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.

[0109] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0110] The above is a detailed introduction to the NUMA scheduling method, device, equipment and medium provided by the present invention in a large model training scenario. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A NUMA scheduling method for large model training scenarios, characterized in that: include: Collect the affiliation relationships between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file; Get the target affinity policy corresponding to the GPU requirements of the current large model training task; If the target affinity policy is the first affinity policy, candidate processor nodes including a candidate NUMA node that satisfies the first affinity policy are screened from the processor nodes based on the topology relationship profile; wherein the candidate NUMA node that satisfies the first affinity policy is a NUMA node in which the number of idle graphics processors under a single NUMA node is not less than the number of graphics processors corresponding to the graphics processor demand; Determine a target NUMA node from each of the candidate NUMA nodes according to the performance communication score of each of the candidate NUMA nodes; The training container is bound to the target NUMA node so that when the training container is started, each graphics processor under the target NUMA node is scheduled to complete the current large model training task.

2. The NUMA scheduling method in a large model training scenario according to claim 1, characterized in that: The collecting of the affiliation between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file includes: Deploy the Device Plugin application to the GPU of the target cluster as a DaemonSet, and use the NVML library to obtain the affiliation between different NUMA nodes and each GPU under each processor node. According to the affiliation relationship, each graphics processor number and the corresponding NUMA node number are encoded into a key-value pair to generate a topology relationship configuration file.

3. The NUMA scheduling method in a large model training scenario according to claim 2, characterized in that: Determining a target NUMA node from each of the candidate NUMA nodes according to the performance communication score of each of the candidate NUMA nodes includes: Scoring each candidate NUMA node based on the continuity of the graphics processor numbers corresponding to each candidate NUMA node to obtain a communication score for each candidate NUMA node; Scoring each of the candidate NUMA nodes based on the load balancing status of each of the candidate NUMA nodes to obtain a performance score of each of the candidate NUMA nodes; The sum of the communication score and the performance score is determined as the performance communication score of each candidate NUMA node, and a target NUMA node is determined from each candidate NUMA node according to the performance communication score.

4. The NUMA scheduling method in a large model training scenario according to any one of claims 1 to 3, characterized in that: Also includes: If the target affinity policy is the second affinity policy, the total number of idle graphics processors under all NUMA nodes in each processor node is determined based on the topology relationship configuration file, and candidate processor nodes whose total idle number meets the graphics processor requirement are screened out from each processor node.

5. The NUMA scheduling method in a large model training scenario according to claim 4, characterized in that: Also includes: Arrange the NUMA nodes in the current candidate processor nodes in descending order according to the number of idle GPUs under each NUMA node to obtain a NUMA node sequence, and determine the current NUMA node from the NUMA node sequence; Determine whether the total number of idle graphics processors from the first NUMA node to the current NUMA node in the NUMA node sequence meets the graphics processor requirement; If the total number of idle GPUs does not meet the GPU requirement, a new current NUMA node is determined from the NUMA node sequence, and the process jumps again to the step of determining whether the total number of idle GPUs from the first NUMA node to the current NUMA node in the NUMA node sequence meets the GPU requirement. If the total number of idle nodes meets the graphics processor requirement, the first NUMA node to the current NUMA node in the NUMA node sequence are determined as candidate NUMA nodes that meet the second affinity policy.

6. The NUMA scheduling method in a large model training scenario according to claim 1, characterized in that: Also includes: The target processor node monitors the scheduling request of the training container based on the ListWatch mechanism, and binds the training container to the target NUMA node to limit the training container to access only the graphics processor bound to the target NUMA node; wherein, the target processor node is the processor node corresponding to the target NUMA node.

7. The NUMA scheduling method in a large model training scenario according to claim 6, characterized in that: The process of scheduling each graphics processor under the target NUMA node to complete the current large model training task also includes: Sending the number list of the graphics processors bound to the target NUMA node to the container runtime through the target processor node; The device file of the graphics processor bound to the target NUMA node is mounted to the training container during container runtime, and the graphics processor visibility environment variable is set according to the number list.

8. A NUMA scheduling device for large model training scenarios, characterized in that: include: The configuration generation module is used to collect the affiliation relationship between different NUMA nodes and each graphics processor under each processor node in the target cluster to generate a topology relationship configuration file; A strategy acquisition module is used to obtain the target affinity strategy corresponding to the GPU requirements of the current large model training task; a first node determination module configured to, if the target affinity policy is a first affinity policy, filter out candidate processor nodes from each of the processor nodes based on the topology relationship configuration file, including a candidate NUMA node that satisfies the first affinity policy; wherein the candidate NUMA node that satisfies the first affinity policy is a NUMA node having a number of idle graphics processors under a single NUMA node that is not less than the number of graphics processors corresponding to the graphics processor demand; A second node determination module is configured to determine a target NUMA node from each of the candidate NUMA nodes according to a performance communication score of each of the candidate NUMA nodes; The scheduling module is used to bind the training container to the target NUMA node so that when the training container is started, the graphics processors under the target NUMA node are scheduled to complete the current large model training task.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the NUMA scheduling method in the large model training scenario as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the NUMA scheduling method in a large model training scenario as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Task scheduling method and device for training cluster, electronic equipment, computer readable storage medium and computer program product

    CN120872552A

  • Memory management method and device, electronic equipment, storage medium and program product

    CN122195869A

  • Memory management methods, devices, electronic devices, storage media, and program products

    CN122195869B