Cross-GPU efficient synchronization method
By setting flag bits in the GPU global memory, the information synchronization process between GPUs is optimized, and the problem of long response time of PCIe communication is solved, which significantly improves communication efficiency.
Patent Information
- Application Number
- CN202311389857.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, information synchronization between GPUs depends on the long response time of PCIe communication, which affects communication efficiency.
By setting the flag bit in the GPU's global memory, the flag bit value of the received GPU is first obtained. If a data packet has been received, the target data packet will be sent. Otherwise, it will not be sent. The flag bit will be updated immediately after receiving the data packet, eliminating the request and waiting for a response step.
The communication efficiency between GPUs is improved, and the communication speed can be increased several times to dozens of times.
Smart Images

Figure CN120301894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital information transmission, and particularly to an efficient cross-GPU synchronization method. Background Art
[0002] In the prior art, the method for information synchronization between two GPUs is as follows: The first GPU sends a data packet to the second GPU. After the second GPU receives the data packet sent by the first GPU, the second GPU sets the flag bit used to represent whether the data packet is received in its global memory from 0 to 1. The first GPU sends a request to read the flag bit to the second GPU, and the second GPU responds to the request of the first GPU to read the flag bit. Thus, the first GPU can read the value of the flag bit, and further determine whether the second GPU has successfully received the data packet. In the above method, the first GPU determines whether the second GPU has successfully received the data packet depending on the second GPU's response to the first GPU's request to read the flag bit. When communicating between the first GPU and the second GPU through PCIe, the process of the second GPU responding to the first GPU's request to read the flag bit takes a relatively long time, which affects the communication efficiency between GPUs. Summary of the Invention
[0003] The purpose of the present invention is to provide an efficient cross-GPU synchronization method to improve the communication efficiency between GPUs.
[0004] According to the present invention, an efficient cross-GPU synchronization method is provided. The synchronization method is applied to a GPU group, and the GPU group includes N GPUs, and the GPUs communicate with each other through PCIe. The communication list of the i-th GPU in the GPU group is L i , L i = {rec1, rec2,..., rec mi ,..., rec Mi}, where rec mi is the mi-th receiving GPU corresponding to the i-th GPU in the GPU group. The value range of mi is from 1 to Mi, and Mi is the number of receiving GPUs corresponding to the i-th GPU in the GPU group. The value range of i is from 1 to N, and N is the number of GPUs included in the GPU group. The receiving GPU corresponding to the i-th GPU in the GPU group is the GPU that serves as the recipient of the data packet sent by the i-th GPU. The synchronization method is executed by the i-th GPU, and the synchronization method includes the following steps:
[0005] S100, traverse L i , and obtain the value of the first flag bit comfla mi of rec mi , comfla miStored in the global memory of the i-th GPU, comfla mi has a value of 0 or comfla mi has a value of 1, and the comfla mi has a value of 0 to indicate that rec mi has not received the latest data packet sent by the i-th GPU. The comfla mi has a value of 1 to indicate that rec mi has received the latest data packet sent by the i-th GPU.
[0006] S200, if comfla mi has a value of 1, then execute S300.
[0007] S300, set both comfla mi and sefla mi to 0. sefla mi is set in the global memory of rec mi and is a flag bit used to indicate whether rec mi has received the latest data packet sent by the i-th GPU. sefla mi has a value of 0 or sefla mi has a value of 1. The sefla mi has a value of 0 to indicate that rec mi has not received the latest data packet sent by the i-th GPU. The sefla mi has a value of 1 to indicate that rec mi has received the latest data packet sent by the i-th GPU.
[0008] S400, send the target data packet to rec mi so that rec mi will set both sefla mi and comfla mi to 1 after receiving the target data packet.
[0009] The present invention has at least the following beneficial effects compared with the prior art:
[0010] Before the $i$-th GPU in the GPU group sends a target data packet to its corresponding $m_i$-th receiving GPU in the present invention, it first obtains the value of the first flag bit of the $m_i$-th receiving GPU placed in the global memory of the $i$-th GPU. If the value is 1 (indicating that the $m_i$-th receiving GPU has set the value to 1), it means that the $m_i$-th receiving GPU has received the latest data packet sent by the $i$-th GPU. In this case, the $i$-th GPU then performs the step of sending the target data packet to its corresponding $m_i$-th receiving GPU; otherwise (indicating that the $m_i$-th receiving GPU has not set the value to 1 and the $m_i$-th receiving GPU has not received the latest data packet), the $i$-th GPU does not perform the step of sending the target data packet to its corresponding $m_i$-th receiving GPU.
[0011] After the $i$-th GPU in the present invention sends a target data packet to its corresponding $m_i$-th receiving GPU, if the $m_i$-th receiving GPU receives the target data packet sent by the $i$-th GPU, then the $m_i$-th receiving GPU sets both the flag bit in its global memory used to represent whether a data packet has been received and the corresponding flag bit in the global memory of the $i$-th GPU to 1. Thus, the $i$-th GPU does not need to send a request to read the flag bit to the above-mentioned $m_i$-th receiving GPU, and correspondingly, does not need to wait for the response of the above-mentioned $m_i$-th receiving GPU to this request. The $i$-th GPU can directly read the value of the flag bit in its global memory that represents whether the $m_i$-th receiving GPU has received the data packet. Since the present invention eliminates the time for the $i$-th GPU to send a request and wait for the request response, it only needs to place the corresponding flag bit in the global memory of the $i$-th GPU and have the $m_i$-th receiving GPU write 1 to this flag bit after successfully receiving the data packet. Since the writing time of the $m_i$-th receiving GPU is shorter than the time for the $i$-th GPU to send a request and wait for the request response when communicating between GPUs through PCIe, therefore, the present invention improves the communication efficiency between the $i$-th GPU and the $m_i$-th receiving GPU. In the present invention, any two communicating GPUs in the GPU group use the above synchronization method for communication. Therefore, the communication efficiency between the GPUs in the GPU group is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a flowchart of an efficient cross-GPU synchronization method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0015] Embodiment 1
[0016] According to this embodiment, an efficient cross-GPU synchronization method is provided. The synchronization method is applied to a GPU group, and the GPU group includes N GPUs, and the GPUs communicate through PCIe; the communication list of the i-th GPU in the GPU group is L i , L i = {rec1, rec2,..., rec mi ,..., rec Mi}, rec mi is the mi-th receiving GPU corresponding to the i-th GPU in the GPU group. The value range of mi is from 1 to Mi, and Mi is the number of receiving GPUs corresponding to the i-th GPU in the GPU group. The value range of i is from 1 to N, and N is the number of GPUs included in the GPU group. The receiving GPU corresponding to the i-th GPU in the GPU group is the GPU that serves as the recipient of the data packet sent by the i-th GPU.
[0017] In this embodiment, there is no limitation on the communication connection relationship between the GPUs in the GPU group. Optionally, the communication connection relationship between the GPUs in the GPU group is any one of the existing seven communication connection relationships. The existing seven communication connection relationships are: broadcast, scatter, gather, all-gather, all-to-all, reduce, all-reduce.
[0018] The synchronization method is executed by the i-th GPU. The synchronization method includes the following steps, as Figure 1 shown:
[0019] S100, traverse L i , and obtain the value of the first flag bit comfla mi of rec mi . Comfla mi is placed in the global memory of the i-th GPU. The value of comfla mi is 0 or the value of comfla mi is 1. The value of comfla mi being 0 is used to represent recmi The latest data packet sent by the i-th GPU is not received, and the value of comfla mi is 1, which is used to represent rec mi The latest data packet sent by the i-th GPU is received.
[0020] In this embodiment, the first flag bit of the corresponding receiving GPU is placed in the global memory of the i-th GPU, enabling the corresponding receiving GPU to perform a write operation on the first flag bit. Optionally, the function of writing in the global memory of the other GPU can be achieved by using the corresponding API control permission.
[0021] S200, if the value of comfla mi is 1, then execute S300.
[0022] S300, set both comfla mi and sefla mi to 0. sefla mi is a flag bit set in the global memory of rec mi and used to represent whether rec mi has received the latest data packet sent by the i-th GPU. The value of sefla mi is 0 or the value of sefla mi is 1. The value of sefla mi being 0 is used to represent that rec mi has not received the latest data packet sent by the i-th GPU, and the value of sefla mi being 1 is used to represent that rec mi has received the latest data packet sent by the i-th GPU.
[0023] S400, send the target data packet to rec mi so that after rec mi receives the target data packet, it sets both sefla mi and comfla mi to 1.
[0024] In this embodiment, after rec mi receives the target data packet, it sets both sefla mi and comfla mi to 1. Before rec mi receives the target data packet, it will not rewrite the value of sefla mi and the value of comfla mi . Thus, the i-th GPU can determine whether rec mi has received the data packet by reading the value of comfla mi stored in its global memory.Whether the target data packet is received. Specifically, if the value of comfla mi is 0, it is determined that rec mi has not received the target data packet; if the value of comfla mi is 1, it is determined that rec mi has received the target data packet.
[0025] Small-scale experiments show that, compared with the synchronization method in the prior art where the ith GPU sends a read flag bit and waits for a response to the corresponding receiving GPU, the communication efficiency of the synchronization method in this embodiment has been significantly improved, and the communication speed can be increased by about one order of magnitude, that is, increased by several times to dozens of times.
[0026] In this embodiment, the target data packet is the next data packet after the latest data packet that the ith GPU intends to send to the mith receiving GPU. The latest data packet is the data packet with the latest sending time sent by the ith GPU at a historical moment, and the historical moment is the set of moments before the current time point.
[0027] In this embodiment, if the ith GPU does not send data packets to other GPUs in the GPU group, that is, L i is an empty set, then the ith GPU does not execute S100 - S400; the synchronization method of this embodiment is applicable to the ith GPU in the GPU group for which the corresponding L i is not an empty set. Before S100 in this embodiment, the following steps are further included:
[0028] S001, determine whether L i is an empty set.
[0029] S002, if L i is not an empty set, then execute S100.
[0030] In this embodiment, before the ith GPU in the GPU group sends a target data packet to its corresponding mith receiving GPU, it will first obtain the value of the first flag bit of the mith receiving GPU placed in the global memory of the ith GPU. If the value is 1 (indicating that the mith receiving GPU has set this value to 1), it means that the mith receiving GPU has received the latest data packet sent by the ith GPU. In this case, the ith GPU will execute the step of sending the target data packet to its corresponding mith receiving GPU; otherwise (indicating that the mith receiving GPU has not set this value to 1 and the mith receiving GPU has not received the latest data packet), the ith GPU will not execute the step of sending the target data packet to its corresponding mith receiving GPU.
[0031] After the $i$-th GPU in this embodiment sends the target data packet to its corresponding $m_i$-th receiving GPU, if the $m_i$-th receiving GPU receives the target data packet sent by the $i$-th GPU, the $m_i$-th receiving GPU sets both the flag bit in its global memory for indicating whether a data packet has been received and the corresponding flag bit in the global memory of the $i$-th GPU to 1. Thus, the $i$-th GPU does not need to send a request to read the flag bit to the above-mentioned $m_i$-th receiving GPU anymore. Correspondingly, it also does not need to wait for the response of the $m_i$-th receiving GPU to this request. The $i$-th GPU can directly read the value of the flag bit in its global memory that indicates whether the $m_i$-th receiving GPU has received the data packet. Since this embodiment eliminates the time for the $i$-th GPU to send a request and wait for the request response, it only needs to place the corresponding flag bit in the global memory of the $i$-th GPU and have the $m_i$-th receiving GPU write 1 to this flag bit after successfully receiving the data packet. Since the writing time of the $m_i$-th receiving GPU is shorter than the time for the $i$-th GPU to send a request and wait for the request response when communicating between GPUs through PCIe, therefore, this embodiment improves the communication efficiency between the $i$-th GPU and the $m_i$-th receiving GPU. In this embodiment, any two communicating GPUs in the GPU group communicate using the above synchronization method. Therefore, the communication efficiency between the GPUs in the GPU group is improved.
[0032] Embodiment 2
[0033] Compared with Embodiment 1, in this Embodiment 2, the $i$-th GPU sends data packets to the GPUs in $L$ by broadcasting. i Only when all the GPUs in $L$ i have received the latest data packet sent by the $i$-th GPU, will it send the next data packet (i.e., the target data packet) to all the GPUs in $L$. i
[0034] The efficient cross-GPU synchronization method in this embodiment includes the following steps:
[0035] S100. Traverse $L$ i , obtain the value of the first flag bit $comfla$ mi in $rec$ mi , append the value of $comfla$ mi to the preset first flag bit sequence $val$, and get $val=(val1,val2,\cdots,val$ mi ,\cdots,val$ Mi ), where $val$ mi is the value of $comfla$ mi , and the initial value of $val$ is an empty value; $comfla$ mi is placed in the global memory of the $i$-th GPU, $comfla$mi The value is 0 or comfla mi The value is 1, and the comfla mi The value of 0 is used to represent rec mi has not received the latest data packet sent by the i-th GPU, and the comfla mi The value of 1 is used to represent rec mi has received the latest data packet sent by the i-th GPU.
[0036] In this embodiment, traverse L i After completion, val includes the first flag bit comfla mi for each rec mi value.
[0037] S200, if the value of comfla mi is 1, then execute S210.
[0038] S210, obtain the number num of elements with the value of 1 in val.
[0039] S220, if num = Mi, then execute S300.
[0040] S300, set both comfla mi and sefla mi to 0. sefla mi is set in the global memory of rec mi and is used to represent whether rec mi has received the latest data packet sent by the i-th GPU. The value of sefla mi is 0 or the value of sefla mi is 1. The value of sefla mi of 0 is used to represent that rec mi has not received the latest data packet sent by the i-th GPU, and the value of sefla mi of 1 is used to represent that rec mi has received the latest data packet sent by the i-th GPU.
[0041] S400, send a target data packet to rec mi so that rec mi will set both sefla mi and comfla mi to 1 after receiving the target data packet.
[0042] In this embodiment, S210 is the step executed when the value of comfla mi is 1. When comfla miWhen the value is 1, it is determined that the $m_i$-th receiving GPU of the $i$-th GPU can receive the latest data packet. In this case, the number of elements with a value of 1 in val is obtained. If the number of elements with a value of 1 in val is equal to the number of GPUs in L i , it is determined that all GPUs in L i have received the latest data packet. In this case, S300 and S400 are only executed to avoid the situation where the data packets received among the GPUs in L i are inconsistent, thus improving the synchronization of data reception among the GPUs in L i .
[0043] In this embodiment, S200 further includes: if the value of comfla mi is 0 and ceil(Mi / 2) ≤ num < Mi, then rec mi is appended to the preset first GPU sequence B to obtain B = (rec’1, rec’2, …, rec’ r , …, rec’ R ), and then enter S201. rec’ r is the $r$-th GPU appended to B, where the value range of $r$ is from 1 to R, and R is the number of GPUs appended to B. The initial value of B is null, and ceil() is the ceiling function.
[0044] S201, obtain the time interval $t$ between the time point when the $i$-th GPU sends the latest data packet and the current time point mi .
[0045] S202, if $t$ mi > t0, then re-send the latest data packet to the GPUs in B. t0 is a preset time interval.
[0046] In this embodiment, t0 is an empirical value, corresponding to the time usually required for the GPUs in L i to receive the latest data packet sent by the $i$-th GPU in the case of no communication failure.
[0047] Optionally, if $t$ mi > t0, in addition to re-sending the latest data packet to the GPUs in B, a preset message for indicating that the GPUs in B fail to receive the latest data packet is also displayed on the user interface. Optionally, the preset message is a text message.
[0048] In this embodiment, appending rec mi to the preset first GPU sequence B and S201 - S202 are steps executed when the value of comfla mi is 0 and ceil(Mi / 2) ≤ num < Mi. When comflami When the value is 0 and ceil(Mi / 2) ≤ num < Mi, it is determined that at the current moment, more GPUs have received the latest data packet and a small number of GPUs have not received the latest data packet. The reason for this situation may be that the time interval between the current time point and the time point when the ith GPU sent the latest data packet is short, or the time interval between the current time point and the time point when the ith GPU sent the latest data packet has met the requirements for the GPUs in L to receive the latest data packet, but due to some reasons, some GPUs in L cannot receive the latest data packet. In this embodiment, when the value is 0 and ceil(Mi / 2) ≤ num < Mi, the step of appending rec i to the preset first GPU sequence B and S201 - S202 can obtain in time the GPUs in L that have not received the latest data packet, and further improve the probability that the GPUs in L that have not received the latest data packet receive the latest data packet by specifically resending the latest data packet to them, so as to achieve the purpose of enabling all GPUs in L to receive the latest data packet; moreover, this embodiment takes num ≥ ceil(Mi / 2) as the condition for executing the step of appending rec i to the preset first GPU sequence B and S201, which can avoid the waste of resources caused by the ith GPU starting to execute the step of appending rec mi to the preset first GPU sequence B and S201 in the short term after sending the latest data packet, saving computing resources. mi to the preset first GPU sequence B and S201 - S202, which can timely obtain the GPUs in L that have not received the latest data packet, and further improve the probability that the GPUs in L that have not received the latest data packet receive the latest data packet by specifically resending the latest data packet to them, so as to achieve the purpose of enabling all GPUs in L to receive the latest data packet; moreover, this embodiment takes num ≥ ceil(Mi / 2) as the condition for executing the step of appending rec i to the preset first GPU sequence B and S201, which can avoid the waste of resources caused by the ith GPU starting to execute the step of appending rec i to the preset first GPU sequence B and S201 in the short term after sending the latest data packet, saving computing resources. i In this embodiment, if the value of comfla mi is 0 but does not meet the condition of ceil(Mi / 2) ≤ num < Mi, it is determined that at the current moment, many GPUs in L have not received the latest data packet, and it is determined that the current time point is a time point relatively close to the time point when the ith GPU sent the latest data packet. In this case, this embodiment does not execute the step of appending rec mi to the preset first GPU sequence B and S201 - S202.
[0049] In this embodiment, if comfla mi When the value is 0 but does not meet the condition of ceil(Mi / 2) ≤ num < Mi, it is determined that at the current moment, many GPUs in L have not received the latest data packet, and it is determined that the current time point is a time point relatively close to the time point when the ith GPU sent the latest data packet. In this case, this embodiment does not execute the step of appending rec i to the preset first GPU sequence B and S201 - S202. mi to the preset first GPU sequence B and S201 - S202.
[0050] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. An efficient synchronization method across GPUs, characterized in that, The synchronization method is applied to a GPU cluster, which includes N GPUs. The GPUs communicate with each other via PCIe. The communication list of the i-th GPU in the GPU cluster is L i , L i = {rec1, rec2, …, rec mi , …, rec Mi}, where rec mi is the mi-th receiving GPU corresponding to the i-th GPU in the GPU cluster. The value range of mi is from 1 to Mi, and Mi is the number of receiving GPUs corresponding to the i-th GPU in the GPU cluster. The value range of i is from 1 to N, and N is the number of GPUs included in the GPU cluster. The receiving GPU corresponding to the i-th GPU in the GPU cluster is the GPU that acts as the recipient of the data packet sent by the i-th GPU. The synchronization method is executed by the i-th GPU, and the synchronization method includes the following steps: S100, traverse L i , obtain rec mi 's first flag bit comfla mi 's value. Place comfla mi in the global memory of the i-th GPU. The value of comfla mi is 0 or the value of comfla mi is 1. The value of the said comfla mi being 0 is used to represent that rec mi has not received the latest data packet sent by the i-th GPU. The value of the said comfla mi being 1 is used to represent that rec mi has received the latest data packet sent by the i-th GPU; S200, if comfla mi has a value of 1, then execute S300; S300, set comfla mi and sefla mi to 0. sefla mi is set in the global memory of rec mi and is a flag bit used to represent whether rec mi has received the latest data packet sent by the i-th GPU. The value of sefla mi is 0 or the value of sefla mi is 1. The value of the said sefla mi being 0 is used to represent that rec mi has not received the latest data packet sent by the i-th GPU, and the value of the said sefla mi being 1 is used to represent that rec mi has received the latest data packet sent by the i-th GPU; S400, send a target data packet to rec mi so that rec mi sets both sefla mi and comfla mi to 1 after receiving the target data packet.
2. The efficient synchronization method across GPUs according to claim 1, wherein The i-th GPU sends data packets to the GPUs in L in a broadcast manner. S100 further includes: appending the value of comfla i to a preset first flag bit sequence val to obtain val = (val1, val2,..., val mi ,…, val mi ,…, val Mi ), where val mi is the value of comfla mi , and the initial value of val is an empty value; in S200, if the value of comfla mi is 1, then S210 is executed before S300: S210, obtain the number num of elements with the value of 1 in val; S220, if num = Mi, then execute S300; otherwise, do not execute S300.
3. The efficient synchronization method across GPUs according to claim 2, wherein S200 also includes: If the value of comfla mi is 0 and ceil(Mi / 2) ≤ num < Mi, then append rec mi to the preset first GPU sequence B to obtain B = (rec’1, rec’2, …, rec’ r , …, rec’ R ), and enter S201, where rec’ r is the r-th GPU appended to B, the value range of r is from 1 to R, R is the number of GPUs appended to B, the initial value of B is null, and ceil() is the ceiling function; S201, obtain the time interval t between the time point when the i-th GPU sends the latest data packet and the current time point mi ; S202, if t mi > t0, then resend the latest data packet to the GPU in B, where t0 is a preset time interval.
4. The efficient synchronization method across GPUs according to claim 1, wherein The communication connection relationship between GPUs in the GPU group is any one of the following communication connection relationships: broadcast, scatter, gather, all-gather, all-to-all, reduce, all-reduce.
5. The efficient synchronization method across GPUs according to claim 1, wherein Before S100, the synchronization method further includes the following steps: S001, Determine L i Is it an empty set? S002, if L i is not an empty set, then execute S100.
6. The efficient synchronization method across GPUs according to claim 1, wherein In S202, if t mi > t0, then preset information for indicating that the GPU in B fails to receive the latest data packet is further displayed on the user interface.
7. The efficient synchronization method across GPUs according to claim 6, wherein The preset information is text information.