A hardware bandwidth utilization optimization system based on communication algorithm

By initializing the number of slicing in the AllReduce communication algorithm and determining the slicing data and its corresponding communication links based on the bandwidth of the heterogeneous communication link, the problem of insufficient bandwidth utilization under the heterogeneous communication link is solved, and a more efficient hardware bandwidth utilization is achieved.

CN119814647BActive Publication Date: 2025-05-16METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280275.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-16
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

During the distributed training of artificial intelligence models, the AllReduce communication algorithm is insufficient in bandwidth utilization under heterogeneous communication links, resulting in some communication links being idle for a long time.

Method used

By initializing the number of slicing k=2, the number of slicing data y1 and y2 of the heterogeneous communication link and its corresponding communication links is determined according to the bandwidths w1 and w2 of the heterogeneous communication link, the time-consuming interval sequences T and S, and the overall time-consuming Q are calculated, and the number of slicing and the corresponding bandwidth utilization Vk are updated until the number of slicing threshold U is reached.

Benefits of technology

By adding the time difference of the slicing data to fill the communication link, the time difference of the heterogeneous communication link is reduced, and the hardware bandwidth utilization is improved so that it can meet expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814647B_ABST
    Figure CN119814647B_ABST
Patent Text Reader

Abstract

The present application relates to the field of hardware optimization technology, and in particular to a hardware bandwidth utilization optimization system based on a communication algorithm. The system determines the split amount corresponding to the first split data and the second split data according to the bandwidth of the first communication link and the second communication link, and then determines the total time consumption T corresponding to the first communication link and the total time consumption S corresponding to the second communication link, thereby calculating the reference bandwidth utilization corresponding to the number of splits, adding split data and updating T and S, and then updating the corresponding relationship between the number of splits and the reference bandwidth utilization. When processing the target data, the appropriate target split number is selected to split the target data and then execute the communication algorithm. It can be seen that the time difference between the first communication link and the second communication link is filled by adding split data, so that as the number of new split data gradually increases, the time difference between the first communication link and the second communication link gradually decreases, thereby improving the hardware bandwidth utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hardware optimization, and in particular to a hardware bandwidth utilization optimization system based on a communication algorithm. Background Art

[0002] When artificial intelligence models, especially large language models, are trained in a distributed manner, the AllReduce communication algorithm usually accounts for a high proportion of the training process. The AllReduce communication algorithm may refer to a reduction operation on the data in all GPU chips used for training, so that each GPU chip contains the result of the reduction operation. The reduction operation may refer to operations such as taking the maximum value and summing. In the first solution proposed in the prior art, the AllReduce communication algorithm is usually split into a ReduceScatter operation and an AllGather operation. The ReduceScatter operation is first performed on each GPU chip, and then the AllGather operation is performed on each GPU chip.

[0003] However, in the scenario of distributed training of artificial intelligence, different GPU interconnection topologies may be involved. In the GPU interconnection topology of multiple servers, there are not only GPU chips interconnected within the server, but also GPU chips interconnected between servers. The bandwidth of the communication link of the GPU interconnection within the server may be different from the bandwidth of the communication link of the GPU interconnection between servers, that is, the GPU interconnection topology under heterogeneous communication links. The implementation of the above AllReduce communication algorithm will result in the inability to enable all communication links to communicate in the scenario where the bandwidth of heterogeneous communication links is different.

[0004] In response to the above problem, the prior art proposes a second solution, that is, first performing ReduceScatter operations separately in the server, then performing ReduceScatter operations between servers, then performing AllGather operations between servers, and finally performing AllGather operations separately in the server, thereby ensuring that each GPU chip is always in a busy state, utilizing all communication links for communication, and improving the utilization of hardware bandwidth to a certain extent.

[0005] However, although the second solution mentioned above ensures that each GPU chip is always busy, it does not guarantee that the bandwidth of the communication link is fully utilized. When the GPU chip has heterogeneous communication links, some communication links will be idle for a long time waiting for transmission from other communication links. Compared with the first solution, although the utilization rate of the hardware bandwidth has been improved, it still fails to fully utilize the hardware bandwidth. Therefore, how to improve the utilization rate of the hardware bandwidth has become an urgent problem to be solved. Summary of the invention

[0006] In view of the above technical problems, the technical solution adopted by the present invention is:

[0007] A hardware bandwidth utilization optimization system based on a communication algorithm, the system comprising: D GPU chips {a1, a2, ..., a d , …, a D}, a processor and a memory storing a computer program, wherein a d is the dth GPU chip, d is an integer in the range of [1, D], the GPU chip includes a first communication port and a second communication port, a d The M-1 GPU chips are interconnected through the first communication ports respectively included to form a first communication sub-link, and all the first communication sub-links form a first communication link, a d The second communication sub-link is formed by interconnecting N-1 GPU chips through the second communication ports respectively included therein, and all the second communication sub-links form a second communication link, 2≤M≤D, 2≤N≤D, the communication bandwidth of a single GPU chip in the first communication link is w1, and the communication bandwidth of a single GPU chip in the second communication link is w2. When the computer program is executed by the processor, the following steps are implemented:

[0008] S101, initialize the number of segments k=2.

[0009] S102, according to w1 and w2, determine the segmentation component y1 corresponding to the first segmentation data x1 and the segmentation component y2 corresponding to the second segmentation data x2, and record the communication links corresponding to x1 and x2 respectively, wherein the communication link is the first communication link or the second communication link.

[0010] S103: Determine a time-consuming interval sequence T corresponding to the first communication link, a time-consuming interval sequence S corresponding to the second communication link, and a total time consumption Q according to y1, y2, M, N, w1, and w2.

[0011] S104, determining a reference bandwidth utilization V according to T, S, w1 and w2 k , record k and V k The corresponding relationship.

[0012] S105, update k=k+1.

[0013] S106, determine the kth segmentation data x according to T and S k The corresponding tangent value y k and its corresponding communication links.

[0014] S107, according to y k , M, N, w1 and w2, update T and S.

[0015] S108, returning to execute steps S104 to S107 until k=U, where U is the segmentation number threshold.

[0016] S109, when executing the preset communication algorithm on the target data, according to k and V k The target segmentation number e is determined based on all corresponding relationships between the target data and the bandwidth utilization threshold corresponding to the target data.

[0017] S110, dividing the target data according to the target segmentation number e, and dividing the target data according to y1, y2, ..., y e and their respective corresponding communication links execute the communication algorithm.

[0018] Compared with the prior art, the present invention has obvious beneficial effects. By means of the above technical solution, the hardware bandwidth utilization optimization system based on the communication algorithm provided by the present invention can achieve considerable technical progress and practicality, and has wide industrial utilization value, and has at least the following beneficial effects:

[0019] The present invention provides a hardware bandwidth utilization optimization system based on a communication algorithm, the system comprising: D GPU chips {a1, a2, ..., a d , …, a D}, a processor and a memory storing a computer program, wherein a d is the dth GPU chip, d is an integer in the range of [1, D], the GPU chip includes a first communication port and a second communication port, a d The M-1 GPU chips are interconnected through the first communication ports respectively included to form a first communication sub-link, and all the first communication sub-links form a first communication link, a dThe second communication sub-link is formed by interconnecting with N-1 GPU chips through the second communication ports respectively included, and all the second communication sub-links form a second communication link, 2≤M≤D, 2≤N≤D, the communication bandwidth of a single GPU chip in the first communication link is w1, and the communication bandwidth of a single GPU chip in the second communication link is w2. When the computer program is executed by the processor, the following steps are implemented: S101, initializing the number of splits k=2, S102, according to w1 and w2, determining the split component y1 corresponding to the first split data x1 and the split component y2 corresponding to the second split data x2, recording the communication links corresponding to x1 and x2 respectively, the communication link is the first communication link or the second communication link, the communication link is the first communication link or the second communication link, S103, according to y1, y2, M, N, w1 and w2, determining the time-consuming interval sequence T corresponding to the first communication link and the time-consuming interval sequence S corresponding to the second communication link, as well as the overall time consumption Q, S104, according to T, S, w1 and w2, determining the reference bandwidth utilization rate V k , record k and V k , S105, update k=k+1, S106, determine the kth segmentation data x according to T and S k The corresponding tangent value y k and its corresponding communication link, S107, according to y k , M, N, w1 and w2, update T and S, S108, return to execute steps S104 to S107 until k=U, U is the segmentation number threshold, S109, when executing the preset communication algorithm on the target data, according to k and V k The target segmentation number e is determined based on all corresponding relationships between the target data and the bandwidth utilization threshold corresponding to the target data. S110: segment the target data based on the target segmentation number e. e and their respective corresponding communication links execute the communication algorithm.

[0020] It can be seen that the time difference between the first communication link and the second communication link is filled by adding new segmented data, so that as the number of new segmented data gradually increases, the time difference between the first communication link and the second communication link gradually decreases. Before the actual execution of the communication algorithm, the corresponding relationship between the number of segments and the reference bandwidth utilization is predetermined. When the communication algorithm is actually executed on the target data, the target data can be segmented with a suitable number of segments before the communication algorithm is executed, thereby ensuring that the hardware bandwidth utilization can meet expectations, that is, the hardware bandwidth utilization is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 A schematic diagram of a flow chart of a computer program being executed by a processor in a hardware bandwidth utilization optimization system based on a communication algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0024] This embodiment provides a hardware bandwidth utilization optimization system based on a communication algorithm, the system comprising: D GPU chips {a1, a2, ..., a d , …, a D}, a processor and a memory storing a computer program, wherein a d is the dth GPU chip, d is an integer in the range of [1, D], the GPU chip includes a first communication port and a second communication port, a d The M-1 GPU chips are interconnected through the first communication ports respectively included to form a first communication sub-link, and all the first communication sub-links form a first communication link, a d The second communication sub-link is formed by interconnecting N-1 GPU chips through the second communication ports respectively included therein, and all the second communication sub-links form a second communication link, 2≤M≤D, 2≤N≤D, the communication bandwidth of a single GPU chip in the first communication link is w1, and the communication bandwidth of a single GPU chip in the second communication link is w2. When the computer program is executed by the processor, the following steps are implemented:

[0025] S101, initialize the number of segments k=2;

[0026] S102, according to w1 and w2, determine the segmentation component y1 corresponding to the first segmentation data x1 and the segmentation component y2 corresponding to the second segmentation data x2, and record the communication links corresponding to x1 and x2 respectively, the communication link is the first communication link or the second communication link, and the communication link is the first communication link or the second communication link;

[0027] S103, determining a time-consuming interval sequence T corresponding to the first communication link, a time-consuming interval sequence S corresponding to the second communication link, and a total time consumption Q according to y1, y2, M, N, w1, and w2;

[0028] S104, determining a reference bandwidth utilization V according to T, S, w1 and w2 k , record k and V k The corresponding relationship;

[0029] S105, update k=k+1;

[0030] S106, determine the kth segmentation data x according to T and S k The corresponding tangent value y k and its corresponding communication links;

[0031] S107, according to y k , M, N, w1 and w2, update T and S;

[0032] S108, returning to step S104 to step S107, until k=U, where U is the segmentation number threshold;

[0033] S109, when executing the preset communication algorithm on the target data, according to k and V k All corresponding relationships between the target data and the bandwidth utilization threshold corresponding to the target data are used to determine the target segmentation number e;

[0034] S110, dividing the target data according to the target segmentation number e, and dividing the target data according to y1, y2, ..., y e and their respective corresponding communication links execute the communication algorithm.

[0035] In one implementation, the hardware bandwidth utilization optimization system of this embodiment is applied to an interconnected topology of M servers, a single server includes N GPU chips, the GPU chips in the server are interconnected via a first communication link, and the GPU chips between servers are interconnected via a second communication link.

[0036] In another implementation, the hardware bandwidth utilization optimization system of this embodiment can also be applied to a GPU interconnection topology structure in a server, where the server includes D GPU chips, and there are at most D GPU chips interconnected by a first communication link through a first communication port, and there are at most D GPU chips interconnected by a second communication link through a second communication port.

[0037] A single GPU chip may include C interconnection ports, and the interconnection port may be a first communication port or a second communication port, wherein the first communication port may be used for a first communication link, and the number of the first communication ports may be c1, and the second communication port may be used for a second communication link, and the number of the second communication ports may be c2, 1≤c1<C, 1≤c2<C, and c1+c2=C.

[0038] The time-consuming interval sequence T corresponding to the first communication link may include several time-consuming intervals corresponding to the first communication link, and a single time-consuming interval may be a range represented by a left boundary value and a right boundary value. The time-consuming interval sequence S corresponding to the second communication link may include several time-consuming intervals corresponding to the second communication link.

[0039] For a single GPU chip, the input bandwidth and output bandwidth of the GPU chip are the same by default, which are both its communication bandwidth.

[0040] Specifically, the number of splits is initially set to 2, corresponding to two split data x1 and x2. According to the split component y1 corresponding to the first split data x1 and the split component y2 corresponding to the second split data x2, the first split data x1 performs a ReduceScatter operation through the first communication link, and then performs a ReduceScatter operation and an AllGather operation through the second communication link, and then performs an AllGather operation through the first communication link to complete the AllReduce operation of the first split data. The second split data x2 performs a ReduceScatter operation through the second communication link, and then performs a ReduceScatter operation and an AllGather operation through the first communication link, and then performs an AllGather operation through the second communication link to complete the AllReduce operation of the second split data. It should be noted that the split data refers to the data obtained by splitting the data in a single GPU chip, and the AllReduce operations of the split data are independent of each other. When the AllReduce operations of all the split data are completed, it is equivalent to that the AllReduce operations of each GPU chip are completed.

[0041] In recording k and V k When the corresponding relationship is determined, and the tangent component y is determined k When the corresponding communication link is recorded, the start execution time of each segmented data can also be recorded to indicate the subsequent segmentation of the target data into y1, y2, ..., y e Afterwards, according to y1, y2, ..., y e And their corresponding communication links and starting execution times, the process of executing the communication algorithm.

[0042] In a specific embodiment, step S102 includes the following steps:

[0043] S1021, if w1 < w2, then let y1 = P, y2 = (w2 / w1)×P, where P is the unit data volume;

[0044] S1022, if w1 > w2, then let y2 = P, y1 = (w1 / w2)×P.

[0045] Among them, the interval length r1 of the time-consuming interval for the first communication link to perform the first preset operation on the first split data is r1 = (y1 / w1)×((N - 1) / N), which can be approximated as r1 = y1 / w1. The interval length r2 of the time-consuming interval for the second communication link to perform the first preset operation on the second split data is r2 = (y2 / w2)×((M - 1) / M), which can be approximated as r2 = y2 / w2. To ensure that the heterogeneous communication link is always busy, r1 = r2 can be set. In this embodiment, y1 / w1 = y2 / w2 is adopted, and then the proportional relationship between y1 and y2 is determined.

[0046] The unit data volume is a variable parameter, which can be determined according to the data volume Y of the target data and the split volume of each split data. For example, when the number of splits is 2 and w1 < w2, y1 + y2 = (1 + w2 / w1)×P = Y, and at this time, P = (w1 + w2)×Y / w1.

[0047] In a specific embodiment, step S102 includes the following steps:

[0048] S1023, if w1 < w2, then let y1 = P, y2 = (w2 / w1)×((M(N - 1)) / (N(M - 1)))×P, where P is the unit data volume;

[0049] S1024, if w1 > w2, then let y2 = P, y1 = (w1 / w2)×((N(M - 1)) / (M(N - 1)))×P.

[0050] Among them, on the basis of r1 = r2, this embodiment adopts (y1 / w1)×((N - 1) / N) = (y2 / w2)×((M - 1) / M), and then determines the proportional relationship between y1 and y2.

[0051] In a specific embodiment, the communication algorithm includes a first preset operation, a second preset operation, and a third preset operation. Step S103 includes the following steps:

[0052] S1031, determine the time-consuming interval t1 corresponding to the first preset operation performed by the first communication link according to y1 and w1;

[0053] S1032, determining, based on y1, w2, M, N, and t1, a time interval s2 corresponding to x1 performing the second preset operation through the second communication link;

[0054] S1033, determining, based on y1, w1, and s2, a time interval t3 corresponding to x1 performing the third preset operation through the first communication link, where the lengths of the intervals corresponding to t1 and t3 are both y1 / w1;

[0055] S1034, determining, based on y2 and w2, a time interval s1 corresponding to when x2 performs the first preset operation through the second communication link;

[0056] S1035, determining, based on y2, w1, N, M, and s1, a time interval t2 corresponding to x2 performing the second preset operation through the first communication link;

[0057] S1036, determining, based on y2, w2, and t2, a time interval s3 corresponding to x2 performing the third preset operation through the second communication link, where the lengths of the intervals corresponding to s1 and s3 are both y2 / w2;

[0058] S1037, determining a time-consuming interval sequence T corresponding to the first communication link according to t1, t2, and t3;

[0059] S1038, determining a time-consuming interval sequence S corresponding to the second communication link according to s1, s2, and s3;

[0060] S1039: The larger value between the maximum value in T and the maximum value in S is taken as the total time consumption Q.

[0061] In this embodiment, for any segmented data, the communication link that starts to execute the first preset operation of the segmented data is used as the initial communication link of the segmented data. Then, the segmented data executes the first preset operation through the initial communication link, and then executes the second preset operation through another communication link, and then executes the third preset operation through the initial communication link.

[0062] The first preset operation may refer to a ReduceScatter operation, the second preset operation may refer to a ReduceScatter operation and an AllGather operation, and the third preset operation may refer to an AllGather operation.

[0063] Specifically, in this embodiment, the interval lengths corresponding to t1 and t3 are both y1 / w1. Using an approximate calculation method, the implementer may also adopt the interval lengths corresponding to t1 and t3 to be (y1 / w1)×((N-1) / N). Similarly, the interval lengths corresponding to s1 and s3 may also be (y2 / w2)×((M-1) / M). The priori assumption here is that in the AllReduce operation, the time consumption corresponding to the ReduceScatter operation and the AllGather operation are consistent.

[0064] The interval length corresponding to s2 is 2×y1 / (w2×N), and the interval length corresponding to t2 is 2×y2 / (w1×M). The time taken for x1 to execute the ReduceScatter operation through the second communication link is y1 / (w2×N), and the time taken to execute the AllGather operation is also y1 / (w2×N). Therefore, the time taken for x1 to execute the second preset operation through the second communication link is 2×y1 / (w2×N). Similarly, the interval length corresponding to t2 is 2×y2 / (w1×M).

[0065] For example, taking y2=P, y1=(w1 / w2)×P and M=N as an example, the first time-consuming intervals contained in the time-consuming interval sequence T corresponding to the first communication link and the time-consuming interval sequence S corresponding to the second communication link are consistent, both of which are [0, P / w2]. Therefore, the second time-consuming interval contained in T is [P / w2, P / w2+2P / (M×w1)], and the second time-consuming interval contained in S is [P / w2, P / w2+(w1 / w2)×2P / (N×w2)]. Since w1>w2, when x1 performs the second preset operation through the second communication link, the second preset operation of x2 may have been completed on the first communication link. Since the first communication link needs to use x1 to perform the result of the second preset operation through the second communication link, after the second preset operation of x2 on the first communication link has been completed, the first communication link is in an idle state, waiting for the second preset operation of x1 on the second communication link to be completed. Then the third time-consuming interval included in T is [P / w2+(w1 / w2)×2P / (N×w2), 2P / w2+(w1 / w2)×2P / (N×w2)], and the third time-consuming interval included in S is also [P / w2+(w1 / w2)×2P / (N×w2), 2P / w2+(w1 / w2)×2P / (N×w2)].

[0066] Furthermore, this embodiment fills the idle time of the communication link by adding newly segmented data. Continuing with the above example, the interval length of the second time-consuming interval of T is less than the interval length of the second time-consuming interval of S. The corresponding idle interval of the first communication link is [P / w2+2P / (M×w1), P / w2+(w1 / w2)×2P / (N×w2)].

[0067] It should be noted that, according to actual needs, implementers can introduce an offset in the time consumption calculation. The offset can be used to describe the time consumption of data synchronization between GPU chips and the delay of the communication link itself, thereby further improving the accuracy of subsequent reference bandwidth utilization calculation.

[0068] In a specific implementation, step S104 includes the following steps:

[0069] S1041, according to T and S, determine the overall idle interval set α=T∪ST∩S;

[0070] S1042, determining a first idle interval set β1 corresponding to the first communication link according to α-α∩T;

[0071] S1043, determining a second idle interval set β2 corresponding to the second communication link according to α-α∩S;

[0072] S1044, determining a usage duration γ1 corresponding to the first communication link according to the sum of the interval lengths corresponding to each idle interval in Q and β1;

[0073] S1045, determining a usage duration γ2 corresponding to the second communication link according to the sum of the interval lengths corresponding to each idle interval in Q and β2;

[0074] S1046, determine the reference bandwidth utilization V according to Q, w1, w2, γ1, γ2 k =((γ1 / Q)×w1+(γ2 / Q)×w2) / (w1+w2), record k and V k The corresponding relationship.

[0075] Among them, T∪S can represent the intersection of T and S, and T∩S can represent the union of T and S. According to T and S, the overall idle interval set formed by all idle intervals in the first communication link and the second communication link can be determined, and then the first idle interval set β1 corresponding to the first communication link and the second idle interval set β2 corresponding to the second communication link can be determined. The ratio of the usage time of the first communication link to the overall time consumption is γ1 / Q, and the ratio of the usage time of the second communication link to the overall time consumption is γ2 / Q. Then, the reference bandwidth (γ1 / Q)×w1 of the first communication link and the reference bandwidth (γ2 / Q)×w2 of the second communication link can be determined, thereby determining the overall reference bandwidth (γ1 / Q)×w1+(γ2 / Q)×w2, and comparing it with the theoretical bandwidth w1+w2, the reference bandwidth utilization can be obtained.

[0076] In a specific implementation, step S106 includes the following steps:

[0077] S1061, determining the interval length θ corresponding to the idle interval to which the minimum value in α belongs;

[0078] S1062: If the free interval to which the minimum value in α belongs is included in β1, determine the kth segmentation data x k The corresponding tangent value y k =θ×w1;

[0079] S1063: If the free interval to which the minimum value in α belongs is included in β2, determine the kth segmentation data x k The corresponding tangent value y k =θ×w2.

[0080] The minimum value may refer to the minimum value in the left boundaries of all space intervals in α, and the idle interval to which the minimum value in α belongs is the time period that needs to be filled.

[0081] In a specific implementation, step S107 includes the following steps:

[0082] S1071, if the idle interval to which the minimum value in α belongs is included in β1, then determine the kth segmentation data x k The time interval I1, x corresponding to executing the first preset operation via the first communication link k The time interval I2 corresponding to executing the second preset operation via the second communication link, x k The time interval I3 corresponding to the third preset operation is executed through the first communication link, I1 and I3 are added to T, and I2 is added to S, wherein the interval length of I1 and I3 is y k / w1, I2 interval length is y k / (N×w2);

[0083] S1072: If the idle interval to which the minimum value in α belongs is included in β2, then determine the kth segmentation data x k The time interval J1, x corresponding to executing the first preset operation via the second communication link k The time interval J2 corresponding to executing the second preset operation through the first communication link, x k The time interval J3 corresponding to the third preset operation is executed through the second communication link, J1 and J3 are added to S, and J2 is added to T, wherein the interval length of J1 and J3 is y k / w2, the interval length of J2 is y k / (M×w1).

[0084] Among them, taking the idle interval of the minimum value in α as an example, if the maximum value of the time-consuming interval in S is less than or equal to the minimum value in I1, then the interval length y of the minimum value in I1 and I2 is k / (N×w2) determines I2. If the maximum time-consuming interval in S is greater than the minimum time-consuming interval in I1, the maximum time-consuming interval in S and the interval length y of I2 are used. k / (N×w2) determines I2, using the maximum value of the time-consuming interval in T and the interval length y of I3 k / w1 confirm I3.

[0085] In a specific implementation, step S109 includes the following steps:

[0086] S1091, when executing the communication algorithm on the target data, from k and V k Among all corresponding relationships between the target data and the target data, determine all the numbers of partitions whose bandwidth utilization is greater than or equal to the bandwidth utilization threshold corresponding to the target data as reference numbers of partitions;

[0087] S1092: Determine the minimum value of all reference segmentation numbers as the target segmentation number e.

[0088] In one embodiment, the target data volume corresponding to the target data can also be obtained, and the minimum segmentation amount of the segmented data under each segmentation number can be determined based on the target data volume. The reference segmentation number whose minimum segmentation amount is greater than a preset segmentation amount threshold is determined from all reference segmentation numbers as a temporary segmentation number, and then the minimum value of all temporary segmentation numbers is determined as the target segmentation number e.

[0089] In this embodiment, the time difference between the first communication link and the second communication link is filled by adding new segmented data, so that as the number of new segmented data gradually increases, the time difference between the first communication link and the second communication link gradually decreases. Before the actual execution of the communication algorithm, the corresponding relationship between the number of segments and the reference bandwidth utilization is predetermined. When the communication algorithm is actually executed on the target data, the target data can be segmented with a suitable number of segments before the communication algorithm is executed, thereby ensuring that the hardware bandwidth utilization can meet expectations, that is, the hardware bandwidth utilization is improved.

[0090] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A hardware bandwidth utilization optimization system based on communication algorithm, characterized in that: The system includes: a first communication link, a second communication link, a processor, and a memory storing a computer program, wherein the first communication link includes a plurality of first communication sub-links formed by M GPU chips, the second communication link includes a plurality of second communication sub-links formed by N GPU chips, the communication bandwidth of a single GPU chip in the first communication link is w1, the communication bandwidth of a single GPU chip in the second communication link is w2, and when the computer program is executed by the processor, the following steps are implemented: S101, Initialize the number of splits k = 2; S102, Determine the split quantity y1 and the communication link corresponding to the first split data x1, and the split quantity y2 and the communication link corresponding to the second split data x2 according to w1 and w2; S103, Determine the time-consuming interval sequence T corresponding to the first communication link and the time-consuming interval sequence S corresponding to the second communication link according to y1, y2, M, N, w1 and w2; S104, determining a reference bandwidth utilization V according to T, S, w1 and w2 k , record k and V k The communication algorithm includes a first preset operation, a second preset operation and a third preset operation, and step S104 includes the following steps: S1041, Determine the overall idle interval set α = T ∪ S - T ∩ S according to T and S; S1042, Determine the first idle interval set β1 corresponding to the first communication link according to α - α ∩ T; S1043, Determine the second idle interval set β2 corresponding to the second communication link according to α - α ∩ S; S105, Update k = k + 1; S106, determine the kth segmentation data x according to T and S k The corresponding tangent value y k and communication links; S107, according to y k , M, N, w1 and w2, update T and S, wherein step S107 includes the following steps: S1071, if the idle interval to which the minimum value in α belongs is included in β1, then determine the kth segmentation data x k The time interval I1, x corresponding to executing the first preset operation via the first communication link k The time interval I2 corresponding to executing the second preset operation via the second communication link, x k The time interval I3 corresponding to the third preset operation is executed through the first communication link, I1 and I3 are added to T, and I2 is added to S, wherein the interval length of I1 and I3 is y k / w1, I2 interval length is y k / (M×w2); S1072: If the idle interval to which the minimum value in α belongs is included in β2, then determine the kth segmentation data x k The time interval J1, x corresponding to executing the first preset operation via the second communication link k The time interval J2 corresponding to executing the second preset operation via the first communication link, x k The time interval J3 corresponding to the third preset operation is executed through the second communication link, J1 and J3 are added to S, and J2 is added to T, wherein the interval length of J1 and J3 is y k / w2, the interval length of J2 is y k / (N×w1); S108, Return to execute S104 until k = U, where U is the split number threshold; S109, according to k and V k All corresponding relationships between them and the bandwidth utilization threshold corresponding to the target data are used to determine the target segmentation number e; S110, dividing the target data according to e, and dividing the target data according to y1, y2, ..., y e and their corresponding communication links execute communication algorithms.

2. The hardware bandwidth utilization optimization system according to claim 1, characterized in that: Step S102 includes the following steps: S1021, If w1 < w2, then let y1 = P, y2 = (w2 / w1) × P, where P is the unit data volume; S1022, If w1 > w2, then let y2 = P, y1 = (w1 / w2) × P.

3. The hardware bandwidth utilization optimization system according to claim 1, characterized in that: Step S102 includes the following steps: S1023, If w1 < w2, then let y1 = P, y2 = (w2 / w1) × ((M(N - 1)) / (N(M - 1))) × P, where P is the unit data volume; S1024, If w1 > w2, then let y2 = P, y1 = (w1 / w2) × ((N(M - 1)) / (M(N - 1))) × P.

4. The hardware bandwidth utilization optimization system according to claim 1, characterized in that: Step S103 includes the following steps: S1031, Determine the time-consuming interval t1 for x1 to execute the first preset operation through the first communication link according to y1 and w1; S1032, Determine the time-consuming interval s2 for x1 to execute the second preset operation through the second communication link according to y1, w2, M, N, and t1; S1033, Determine the time-consuming interval t3 for x1 to execute the third preset operation through the first communication link according to y1, w1, and s2. The interval lengths corresponding to t1 and t3 are both y1 / w1; S1034, Determine the time-consuming interval s1 for x2 to execute the first preset operation through the second communication link according to y2 and w2; S1035, Determine the time-consuming interval t2 for x2 to execute the second preset operation through the first communication link according to y2, w1, N, M, and s1; S1036, determining, based on y2, w2, and t2, a time interval s3 corresponding to x2 performing the third preset operation through the second communication link, where the lengths of the intervals corresponding to s1 and s3 are both y2 / w2; S1037, determining a time-consuming interval sequence T corresponding to the first communication link according to t1, t2, and t3; S1038: Determine the time-consuming interval sequence S corresponding to the second communication link according to s1, s2, and s3.

5. The hardware bandwidth utilization optimization system according to claim 4, characterized in that: Step S104 also includes the following steps: S1044, taking the larger value between the maximum value in T and the maximum value in S as the total time consumption Q, and determining the usage time γ1 corresponding to the first communication link according to the sum of the interval lengths corresponding to each idle interval in Q and β1; S1045, determining a usage duration γ2 corresponding to the second communication link according to the sum of the interval lengths corresponding to each idle interval in Q and β2; S1046, determine the reference bandwidth utilization V according to Q, w1, w2, γ1, γ2 k =((γ1 / Q)×w1+(γ2 / Q)×w2) / (w1+w2), record k and V k The corresponding relationship.

6. The hardware bandwidth utilization optimization system according to claim 5, characterized in that: Step S106 includes the following steps: S1061, determining the interval length θ corresponding to the idle interval to which the minimum value in α belongs; S1062: If the free interval to which the minimum value in α belongs is included in β1, determine the kth segmentation data x k The corresponding tangent value y k =θ×w1; S1063: If the free interval to which the minimum value in α belongs is included in β2, determine the kth segmentation data x k The corresponding tangent value y k =θ×w2.

7. The hardware bandwidth utilization optimization system according to claim 1, characterized in that: Step S109 includes the following steps: S1091, from k and V k Among all corresponding relationships between the target data and the target data, determine all the numbers of splits whose bandwidth utilization is greater than or equal to the bandwidth utilization threshold corresponding to the target data as reference numbers of splits; S1092: Determine the minimum value of all reference segmentation numbers as the target segmentation number e.

Citation Information

Patent Citations

  • Method and device for acquiring bandwidth utilization rate and storage medium

    CN108304288A

  • Communication method, apparatus and device, and computer readable storage medium

    CN115687233A