GPU topology awareness scheduling method, electronic equipment and medium

Through the GPU topology-aware scheduling method, the GPU allocation scheme with the highest communication efficiency is selected, which solves the problems of low GPU usage and communication efficiency in computer clusters, and achieves more balanced load and higher resource utilization.

CN119938289APending Publication Date: 2025-05-06MUXI LINGZHI TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202311457304.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the usage rate of GPUs and the communication efficiency between multiple GPUs in computer clusters, resulting in resource fragmentation and load imbalance.

Method used

The GPU topology-aware scheduling method is adopted to receive task processing requests, analyze the number of target GPUs, select computer nodes, obtain available GPU lists and topological relationships, generate candidate allocation schemes, calculate weights based on topological relationships, and select the GPU allocation scheme with the highest communication efficiency.

Benefits of technology

It improves the usage rate of GPUs in the computer cluster and the communication efficiency between multiple GPUs, making the load of the entire cluster more balanced, and avoids resource fragmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938289A_ABST
    Figure CN119938289A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chips, in particular to a GPU topology sensing scheduling method, electronic equipment and a medium, and the method comprises the following steps: S1, receiving a target task processing request and analyzing the target task processing request to obtain a target GPU number M; s2, selecting a target computer node from a computer cluster based on M; s3, obtaining a current available GPU list of the target computer node and a topological relation between the current available GPUs; s4, generating all candidate allocation schemes based on the current available GPU list and the target GPU number M; s5, obtaining a weight AXn corresponding to the An and a weight BXn corresponding to the Bn based on the topological relation between the currently available GPUs; and step S6, determining the maximum An of the corresponding (AXn + BXn) as a target GPU for processing the target task. According to the method, the utilization rate maximization of the GPUs in the computer cluster and the communication efficiency among the GPUs are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chip technology, and in particular to a GPU topology-aware scheduling method, electronic equipment and medium. Background Art

[0002] As the demand for computing increases, multiple graphics processing units (GPUs) are usually needed to jointly perform tasks, but the number of GPUs corresponding to a computer node is limited. Therefore, it is necessary to build a computer cluster consisting of multiple computer nodes to process tasks. When a task is sent to a computer cluster, the computer cluster needs to complete two-dimensional scheduling. The first dimension is to select a computer node in the computer cluster, and the second dimension is to select a target number of GPUs from multiple GPUs in the selected computing node to perform the task.

[0003] The existing technology can usually complete the scheduling of the first dimension efficiently and accurately, while the scheduling of the second dimension usually directly allocates GPUs based on the K8s architecture or allocates GPUs based on a greedy strategy. However, directly allocating GPUs based on the K8s architecture will result in the inability of tasks to use the optimal GPU performance when there are sufficient GPU resources, resulting in low communication efficiency between multiple GPUs. Allocating GPUs based on a greedy strategy will lead to resource fragmentation, and it is impossible to give the best GPU allocation solution based on the fragmentation, resulting in low GPU utilization and low communication efficiency between multiple GPUs. In the computer cluster mode, it is necessary to maximize the utilization of GPUs and ensure that the communication efficiency between multiple GPUs is as high as possible, so that the load of the entire computer cluster is more balanced. It can be seen from this that how to maximize the utilization of GPUs in computer clusters and the communication efficiency between multiple GPUs has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The present invention aims to provide a GPU topology-aware scheduling method, an electronic device and a medium, which maximize the utilization rate of GPUs in a computer cluster and improve the communication efficiency between multiple GPUs.

[0005] According to a first aspect of the present invention, a GPU topology-aware scheduling method is provided, comprising:

[0006] Step S1, receiving a target task processing request and parsing it to obtain the target GPU quantity M corresponding to the target task;

[0007] Step S2, selecting a target computer node from the computer cluster based on the target number M of GPUs;

[0008] Step S3, obtaining a list of currently available GPUs of the target computer node and a topological relationship between currently available GPUs;

[0009] Step S4: Generate all candidate allocation schemes {(A1, B1), (A2, B2), ..., (A n ,B n ),…,(A N ,B N )}}, where (A n ,B n ) is the nth group of candidate allocation schemes, A n is the candidate target GPU list, B n Remove A from the list of currently available GPUs n The GPU list after the GPU in A n There are M GPUs in it, and the value of n ranges from 1 to N;

[0010] Step S5: Obtain A based on the topological relationship between currently available GPUs n The corresponding weight AX n and B n The corresponding weight BX n , AX n With A n The corresponding GPU communication efficiency is proportional to BX n With B n The corresponding GPU communication efficiency is proportional;

[0011] Step S6: The corresponding (AX n +BX n )The largest A n Determine the target GPU to be used to process the target task.

[0012] According to a second aspect of the present invention, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to execute the method described in the first aspect of the present invention.

[0013] According to a third aspect of the present invention, there is provided a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the first aspect of the present invention.

[0014] Compared with the prior art, the present invention has obvious advantages and beneficial effects. By means of the above technical solution, the GPU topology-aware scheduling method, electronic device and medium provided by the present invention can achieve considerable technical advancement and practicality, and have wide industrial utilization value, and at least have the following beneficial effects:

[0015] When selecting a target GPU for processing a target task, the present invention considers both the weight of the GPU to be selected and the weight of the remaining GPUs after the selection, thereby taking into account both the communication efficiency of the GPU allocated for the current task and the communication efficiency of the remaining GPU resources when being allocated for the next task, thereby making the load of the entire computer cluster more balanced and improving the utilization rate of the GPUs in the computer cluster and the communication efficiency between multiple GPUs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A flowchart of a GPU topology-aware scheduling method provided by an embodiment of the present invention;

[0018] Figure 2 A schematic diagram of a method for dividing target task 2 into two groups provided in an embodiment of the present invention;

[0019] Figure 3 A schematic diagram of another method of dividing target task 2 into two groups provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0021] The embodiment of the present invention provides a GPU topology-aware scheduling method, such as Figure 1 As shown, including:

[0022] Step S1, receiving a target task processing request and parsing it to obtain the target GPU quantity M corresponding to the target task.

[0023] The target task processing request includes the target GPU quantity M, so the target GPU quantity M corresponding to the target task can be directly obtained by parsing the target task processing request.

[0024] Step S2: Select a target computer node from the computer cluster based on the target number M of GPUs.

[0025] It is understandable that when a target task processing request is received, there may be multiple computer nodes in the computer cluster for processing the target task. Therefore, the best computer node can be selected from the computer cluster for processing the target task. As an example, the computer cluster is specifically a K8s cluster. K8s is short for Kubernetes, which is an open source system for automatically deploying, scaling and managing containerized applications.

[0026] Step S3: Obtain a list of currently available GPUs of the target computer node and a topological relationship between currently available GPUs.

[0027] It should be noted that each computer node is provided with an acquisition component, which can directly obtain the currently available GPU list and the topological relationship between the currently available GPUs from the device file in real time through the acquisition component, which will not be repeated here.

[0028] Step S4: Generate all candidate allocation schemes {(A1, B1), (A2, B2), ..., (A n ,B n ),…,(A N ,B N )}}, where (A n ,B n ) is the nth group of candidate allocation schemes, A n is the candidate target GPU list, n ranges from 1 to N, N is the total number of candidate allocation schemes, B n Remove A from the list of currently available GPUs n The GPU list after the GPU in A n There are M GPUs in B n The number of GPUs included is (VM), and V is the total number of GPUs currently available in the target computer node.

[0029] It should be noted that, when selecting a target GPU, the embodiment of the present invention not only needs to consider the current task, but also needs to consider the communication efficiency when the remaining GPU resources are allocated to the next task. In step S4, all possible solutions need to be obtained as candidate allocation solutions.

[0030] Step S5: Obtain A based on the topological relationship between currently available GPUsn The corresponding weight AX n and B n The corresponding weight BX n , AX n With A n The corresponding GPU communication efficiency is proportional to BX n With B n The corresponding GPU communication efficiency is proportional.

[0031] It should be noted that different topological relationships between GPUs have different corresponding communication efficiencies. Therefore, step S5 needs to obtain the corresponding weights based on the topological relationships between currently available GPUs. n and BX n Selecting the target GPU can improve the utilization of the GPU and help improve the overall load balancing of the GPU in the computer cluster.

[0032] Step S6: The corresponding (AX n +BX n )The largest A n Determine the target GPU to be used to process the target task.

[0033] When selecting a target GPU, it is necessary to comprehensively consider the communication efficiency of the target GPU selected for executing the current task, as well as the communication efficiency of the remaining GPU resources when allocated to the next task, so as to improve the utilization of the GPU of the entire computer cluster and the overall load balancing.

[0034] As an embodiment, step S2 includes:

[0035] Step S21: Obtain candidate computer nodes in the current computer cluster whose available GPU quantity is greater than the target GPU quantity.

[0036] It is understandable that the number of available GPUs must be greater than the target number of GPUs to be used as target computer nodes.

[0037] Step S22: Determine the candidate computer node with the largest current available space as the target computer node.

[0038] It should be noted that in order to improve the utilization of the GPU of the entire computer cluster and the overall load balancing, it is necessary to select the candidate computer node with the largest available space as the target computer node. The candidate computer node with the largest available space refers to the target computer node with the largest load that is currently light and can be added.

[0039] As an embodiment, step S5 includes:

[0040] Step S51: Get En All GPU combinations with topological relationships between them {(G 11 n ,G 12 n ),(G 21 n ,G 22 n ),…,(G i1 n ,G i2 n ),…,(G f(n)1 n ,G f(n)2 n )}, where (G i1 n ,G i2 n ) is E n The i-th GPU combination with topological relationship in the above example, i ranges from 1 to f(n), and f(n) is E n The total number of GPU combinations with topological relationships in E n A n or B n ,Topological relationships include direct connection relationships and indirect connection relationships.

[0041] It should be noted that the GPUs having a topological relationship between each other refer to the GPUs having a direct connection relationship or an indirect connection relationship between each other. n A n or B n , that is, for A n and B n A can be obtained through steps S51 to S54. n The corresponding weight AX n and B n The corresponding weight BX n It is understandable that if E n A n , then the corresponding EX in step S54 n AX n If E n For B n , then the corresponding EX in step S54 n For BX n .

[0042] Step S52: Get (G i1 n ,G i2 n ) corresponds to the topological relationship {L1 in ,L2 in ,…,Lj in ,…,L h(in) in}, where L j in For (G i1 n ,G i2 n ) corresponds to the jth topological relationship, the value range of j is 1 to h(in), h(in) is (G i1 n ,G i2 n ) corresponds to the total number of topological relationships.

[0043] Among them, h(in)≥1, that is, (G i1 n ,G i2 n ) may include only one topological relationship or multiple topological relationships, which may be a direct connection relationship or an indirect connection relationship, among which there may be multiple indirect connection situations.

[0044] Step S53: Based on (G i1 n ,G i2 n ) corresponding to all topological relationships (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n .

[0045] Step S54: Get E n Corresponding weight EX n :

[0046]

[0047] As an embodiment, step S53 includes:

[0048] Step S531, obtain (G i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in .

[0049] Among them, the communication weight is related to the topological relationship.

[0050] Step S532: i1n ,G i2 n ) corresponding to all L j in The corresponding communication weight E j in The sum is determined as (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n .

[0051] Step S531-step S532, based on (G i1 n ,G i2 n ) corresponding to each L j in To determine (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n On this basis, in order to further improve the acquisition (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n To improve the accuracy of the connection, we can further add the physical distance factor and adjust the L of each indirect connection by physical distance. j in The corresponding communication weight E j in , because in actual selection, when there is a direct connection relationship between the GPUs, the direct connection will be preferred, therefore, only the weight of the inter-connection needs to be adjusted. As an embodiment, the step S53 includes:

[0052] Step R531, obtain (G i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in .

[0053] Step R532: Obtain each L with an indirect connection relationship j in The corresponding physical distance d j in .

[0054] Step R533, based on d j in Set L with indirect connection relationship j in The weight adjustment coefficient T j in , T j in With d j in Inversely proportional, each L with a direct connection relationship j in The weight adjustment coefficient T j in Set to 1.

[0055] Step R534: Based on T j in and E j in Determined as (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n :

[0056]

[0057] By adding the physical distance factor with indirect connection relationship to adjust the weight to select the target GPU, the utilization rate of the GPU of the entire computer cluster and the overall load balancing are further improved.

[0058] As an embodiment, the topological relationship between any two GPUs includes {C1, C2, C3, C4, C5, C6}, where C1 represents a topological relationship based on a direct connection via a high-speed connection channel between the GPU and the CPU (Central Processing Unit), and C2 represents a topological relationship indirectly connected via a PCIe Switch. PCIe is a universal bus specification, and devices connected via the PCIe bus are called PCIe devices. PCIe Switch provides the ability to interconnect PCIe devices and is used as a packet router. C3 represents a topological relationship indirectly connected via a root component (Root Complex, RC for short) of a computer node. The PCIE architecture generally includes a root component RC, a switch, a terminal device EP (End Point) and other types of PCIE devices. There is only one RC in the bus architecture, which is used for the connection between the processor and memory subsystem and the I / O device. C4 represents a topological relationship indirectly connected through multiple PCIe switches, C5 represents a topological relationship indirectly connected through a non-uniform memory access architecture (NUMA) node. NUMA is a computing system composed of multiple nodes, each of which has its own independent memory space, CPU core and PCIe bus system. C6 represents a topological relationship indirectly connected through multiple NUMA nodes. i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in ,include:

[0059] Step S5311: If L j in The corresponding topological relationship is C1, then L j in The corresponding communication weight E j in Set to D1; if L j in The corresponding topological relationship is C2, then L j in The corresponding communication weight E j in Set to D2; if L j in The corresponding topological relationship is C3, then L j in The corresponding communication weight E j inSet to D3; if L j in The corresponding topological relationship is C4, then L j in The corresponding communication weight E j in Set to D4; if L j in The corresponding topological relationship is C5, then L j in The corresponding communication weight E j in Set to D5; if L j in The corresponding topological relationship is C6, then L j in The corresponding communication weight E j in Set to D6, where D1>D2>D3>D4>D5>D6.

[0060] For topological relationships that are indirectly connected through multiple PCIe Switches, the number of PCIe Switches passed through is different, and the corresponding communication weight values ​​are also different. As an example, in step S5311, if there are more than two D4s, the number of PCIe Switches corresponding to each D4 is obtained as Y and Z respectively, and the communication weights are set to D4(Y) and D4(Z) respectively, where D4(Y)>D4(Z), Z>Y>2.

[0061] For topological relationships that are indirectly connected through multiple NUMA nodes, the number of NUMA nodes passed through is different, and the corresponding communication weight values ​​are also different. As an example, in step S5311, if there are more than two D6s, the number of NUMA nodes corresponding to each D6 is obtained as P and Q respectively, and the communication weights are set to D6(P) and D6(Q) respectively, where D6(P)>D6(Q), Q>P>2.

[0062] As an example, the high-speed connection channel between the GPU and the CPU is named Metalinks. Assume that there are 4 PCIe Switches and 4 GPUs connected by Metalinks in the target computer node, where 4 PCIe Switches means that there are 4 GPUs connected by PCIe Switches, and 4 Metalinks means that 4 GPUs are connected by Metalinks. Based on the method described in the embodiment of the present invention, the following examples of some GPU allocation scenarios are obtained:

[0063] Scenario 1: No GPU is currently occupied, and there are 4 PCIe Switches and 4 GPUs connected by Metalinks:

[0064] If the target task requires two GPUs, the GPUs connected to the two PCIe switches are prioritized.

[0065] If the target task requires three GPUs, the GPUs connected to the three PCIe switches are prioritized.

[0066] If the target task requires 4 GPUs, 4 GPUs connected by Metalinks are prioritized.

[0067] Scenario 2: When two GPUs connected to PCIe switches are currently occupied:

[0068] If the target task requires two GPUs, the GPUs connected to the two PCIe switches are prioritized.

[0069] If the target task requires 3 GPUs, the 3 GPUs connected by Metalinks are allocated first.

[0070] If the target task requires 4 GPUs, 4 GPUs connected by Metalinks are prioritized.

[0071] Scenario 3: When two GPUs connected by Metalinks are currently occupied:

[0072] If the target task requires 2 GPUs, the 2 GPUs connected by Metalinks are prioritized.

[0073] If the target task requires three GPUs, the GPUs connected to the three PCIe switches are prioritized.

[0074] If the target task requires 4 GPUs, the GPUs connected to the 4 PCIe switches are prioritized.

[0075] The present invention is further illustrated below by a specific example. Assume that there are 8 GPUs, namely GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, and GPU7. The 8 GPUs are divided into two groups. The first group includes GPU0, GPU1, GPU2, and GPU3, and the second group includes GPU3, GPU4, GPU5, GPU6, and GPU7. GPU0, GPU1, GPU2, and GPU3 are connected through PCIe Switch1, and are also connected through Metalink1. GPU3, GPU4, GPU5, GPU6, and GPU7 are connected through PCIe Switch2, and are also connected through Metalink2. For target task 1 (Task1), two GPUs need to be allocated, and the 8 GPUs are divided into two groups (SubGraph1 and Sub Graph2), one group includes two GPUs, and the other group includes 6 GPUs. Through the method described in the embodiment of the present invention, the left and right possibilities are traversed, and a group of optimal points is selected, and GPU6 and GPU7 are allocated to target task 1.

[0076] Next, target task 2 (Task2) also requires two GPUs. The remaining 6 GPUs are divided into two groups, one group includes two GPUs, and the other group includes four GPUs. Figure 2 The example shown, Figure 2 In , if GPU0 and GPU1 are selected for target task 2, the weights of GPU0 and GPU1 are 20, and the weights of GPU2, GPU3, GPU4, and GPU5 are 80. Figure 3 In the example shown, if GPU4 and GPU5 are selected for target task 2, the weights of GPU4 and GPU5 are 20, and the weights of GPU0, GPU1, GPU2, and GPU3 are 120. In the existing method, usually only the target task 2 to be assigned is considered, and GPU0 and GPU1 that are not in the same PCIe Switch as target task 1 are likely to be selected and assigned to target task 2, which will lead to GPU resource fragmentation. In the embodiment of the present invention, Figure 2 The total weight of the corresponding division method is 100. Figure 3 The total weight of the corresponding partitioning method is 140, because based on the embodiment of the present invention, GPU4 and GPU5 in the same PCIe Switch as target task 1 are selected to be allocated to target task 2. For target task 3, 4 GPUs are required, and GPU0, GPU1, GPU2, and GPU3 can be directly allocated to target task 3.

[0077] It should be noted that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0078] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.

[0079] The embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.

[0080] When selecting a target GPU to process a target task, the embodiment of the present invention considers both the weight of the GPU to be selected and the weight of the remaining GPUs after the selection, thereby taking into account both the communication efficiency of the GPU allocated for the current task and the communication efficiency of the remaining GPU resources when allocated for the next task, thereby making the load of the entire computer cluster more balanced and improving the utilization rate of the GPUs in the computer cluster and the communication efficiency between multiple GPUs.

[0081] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A GPU topology-aware scheduling method, characterized in that: include: Step S1, receiving a target task processing request and parsing it to obtain the target GPU quantity M corresponding to the target task; Step S2, selecting a target computer node from the computer cluster based on the target number M of GPUs; Step S3, obtaining a list of currently available GPUs of the target computer node and a topological relationship between currently available GPUs; Step S4: Generate all candidate allocation schemes {(A1, B1), (A2, B2), ..., (A n ,B n ),…,(A N ,B N )}}, where (A n ,B n ) is the nth group of candidate allocation schemes, A n is the candidate target GPU list, B n Remove A from the list of currently available GPUs n The GPU list after the GPU in A n There are M GPUs in it, and the value of n ranges from 1 to N; Step S5: Obtain A based on the topological relationship between currently available GPUs n The corresponding weight AX n and B n The corresponding weight BX n , AX n With A n The corresponding GPU communication efficiency is proportional to BX n With B n The corresponding GPU communication efficiency is proportional; Step S6: The corresponding (AX n +BX n )The largest A n Determine the target GPU to be used to process the target task.

2. The method according to claim 1, characterized in that: The step S2 comprises: Step S21, obtaining candidate computer nodes in the current computer cluster whose available GPU number is greater than the target GPU number; Step S22: Determine the candidate computer node with the largest current available space as the target computer node.

3. The method according to claim 1, characterized in that The step S5 comprises: Step S51: Get E n All GPU combinations with topological relationships between them {(G 11 n ,G 12 n ),(G 21 n ,G 22 n ),…,(G i1 n ,G i2 n ),…,(G f(n)1 n ,G f(n)2 n )}, where (G i1 n ,G i2 n ) is E n The i-th GPU combination with topological relationship in the above example, i ranges from 1 to f(n), and f(n) is E n The total number of GPU combinations with topological relationships in E n A n or B n ,Topological relations include direct connection relations and indirect connection relations; Step S52: Get (G i1 n ,G i2 n ) corresponds to the topological relationship {L1 in ,L2 in ,…,L j in ,…,L h(in) in }, where L j in For (G i1 n ,G i2 n ) corresponds to the jth topological relationship, the value range of j is 1 to h(in), h(in) is (G i1 n ,G i2 n )The total number of topological relationships corresponding to; Step S53: Based on (G i1 n ,G i2 n ) corresponding to all topological relationships (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n ; Step S54: Get E n Corresponding weight EX n :

4. The method according to claim 3, characterized in that The step S53 comprises: Step S531, obtain (G i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in ; Step S532: i1 n ,G i2 n ) corresponding to all L j in The corresponding communication weight E j in The sum is determined as (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n .

5. The method according to claim 3, characterized in that: The step S53 comprises: Step R531, obtain (G i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in . Step R532: Obtain each L with an indirect connection relationship j in The corresponding physical distance d j in . Step R533, based on d j in Set L with indirect connection relationship j in The weight adjustment coefficient T j in , T j in With d j in Inversely proportional, each L with a direct connection relationship j in The weight adjustment coefficient T j in Set to 1. Step R534: Based on T j in and E j in Determined as (G i1 n ,G i2 n ) The sum of the corresponding communication weights U i n :

6. The method according to claim 4 or 5, characterized in that: The topological relationship between any two GPUs includes {C1, C2, C3, C4, C5, C6}, where C1 represents a topological relationship of direct connection based on a high-speed connection channel between the GPU and the CPU, C2 represents a topological relationship of indirect connection through a PCIe Switch, C3 represents a topological relationship of indirect connection through an RC of a computer node, C4 represents a topological relationship of indirect connection through multiple PCIe Switches, C5 represents a topological relationship of indirect connection through a NUMA node, and C6 represents a topological relationship of indirect connection through multiple NUMA nodes. The acquisition (G i1 n ,G i2 n ) corresponding to each L j in The corresponding communication weight E j in ,include: Step S5311: If L j in The corresponding topological relationship is C1, then L j in The corresponding communication weight E j in Set to D1; If L j in The corresponding topological relationship is C2, then L j in The corresponding communication weight E j in Set to D2; If L j in The corresponding topological relationship is C3, then L j in The corresponding communication weight E j in Set to D3; If L j in The corresponding topological relationship is C4, then L j in The corresponding communication weight E j in Set to D4; If L j in The corresponding topological relationship is C5, then L j in The corresponding communication weight E j in Set to D5; If L j in The corresponding topological relationship is C6, then L j in The corresponding communication weight E j in Set to D6, Among them, D1>D2>D3>D4>D5>D6.

7. The method according to claim 6, characterized in that In step S5311, if there are more than two D4s, the number of PCIe switches corresponding to each D4 is obtained as Y and Z respectively, and the communication weights are set to D4(Y) and D4(Z) respectively, where D4(Y)>D4(Z), Z>Y>2.

8. The method according to claim 6, characterized in that In step S5311, if there are more than two D6s, the number of NUMA nodes corresponding to each D6 is obtained as P and Q respectively, and the communication weights are set to D6(P) and D6(Q) respectively, where D6(P)>D6(Q), Q>P>2.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, wherein the instructions are configured to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute the method of any one of the preceding claims 1-8.

Citation Information

Cited By

  • Intelligent topology awareness-based hundred-thousand-card cluster communication acceleration system

    CN120321246A

  • Accelerator card deployment method and device, equipment, storage medium and program product

    CN120407200A

  • Accelerator card deployment method, device, equipment, storage medium and program product

    CN120407200B

  • Method and system for scheduling heterogeneous GPU (Graphics Processing Unit) by cross-cluster management software

    CN121542052A

  • GPU allocation method, GPU allocation apparatus, electronic device, storage medium, and program product

    WO2026170848A1