Full-protocol method, device, system, equipment and medium based on switch connection
By actively sending blocked data between computing nodes connected to the switch, the existing full reduction operation is solved, and more efficient distributed system data synchronization and parameter update are achieved.
Patent Information
- Application Number
- CN202510387484.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-05-13
AI Technical Summary
The existing full reduction operations based on on-network computing are complex and have a large delay, making it difficult to efficiently synchronize data or update parameters in distributed systems.
A fully reduction method based on switch connection is proposed. Each computing node actively sends the block data required by other computing nodes. The switch does not require multicast specification requests, reducing operation steps and reducing delays.
It significantly reduces the complexity and delay of full reduction operations, reduces the processing pressure of the switch, and improves the efficiency of distributed computing.
Smart Images

Figure CN119996354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a full protocol method, device, system, equipment and medium based on switch connection. Background Art
[0002] When performing distributed computing, data synchronization or parameter updates are required through communication links between computing nodes (such as GPUs or CPUs, etc.). AllReduce is a key technology for implementing data synchronization or parameter updates in distributed systems, and can be used to synchronize gradients between computing nodes during large model training. For example, in an allreduce operation, a reduction operation is performed on the data held by all computing nodes in the distributed system, and the reduction results are multicast to all computing nodes through communication links to ensure information consistency between computing nodes.
[0003] In-Network Computing is a technology that offloads computing tasks to network devices (such as switches, smart network cards, etc.). The purpose is to reduce data transmission volume, reduce latency, and improve the overall efficiency of distributed systems by directly processing data streams in the network. The core idea of implementing full reduction based on in-network computing is to use the computing power of network hardware to offload part or all of the full reduction operations to network devices, thereby reducing the data transmission and computing burden between computing nodes and improving the efficiency of distributed computing.
[0004] In the current full-reduce operation based on network computing, more steps are involved, which has the disadvantages of complex implementation and high latency. Summary of the invention
[0005] The present invention proposes a full protocol method, device, system, equipment, storage medium and product based on switch connection, which helps to reduce complexity and delay.
[0006] The technical solution of the embodiment of the present invention is as follows:
[0007] A full reduction method based on switch connection, the method is applicable to an m-th computing node among M computing nodes connected to the switch, the m-th computing node contains N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data; the method comprises:
[0008] Sending the nth block data except the mth block data among the N block data to the nth computing node among the M computing nodes, and receiving the block data numbered m among the M-1 computing nodes from the M-1 computing nodes except the mth computing node;
[0009] Performing a reduction operation on the block data numbered m in the N block data and the block data numbered m in the M-1 computing nodes except the m-th computing node to obtain a first reduction result;
[0010] The first protocol result is sent to the switch, so that the first protocol result is multicast by the switch.
[0011] In one embodiment, the multicasting of the first protocol result includes:
[0012] Multicasting the first reduction result to the M computing nodes; or
[0013] The first reduction result is multicasted to M-1 computing nodes except the mth computing node.
[0014] In one embodiment, the N is equal to the M.
[0015] In one embodiment, N is K times M, where K is a positive integer of at least 2; the method comprises:
[0016] Sending the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the nth computing node among the M computing nodes, where the value range of k is [1, K-1];
[0017] Receive, from M-1 computing nodes other than the m-th computing node, block data numbered (m+k*M) in the M-1 computing nodes;
[0018] Performing a reduction operation on the block data numbered (m+k*M) in the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result;
[0019] The second protocol result is sent to the switch, so that the second protocol result is multicast by the switch.
[0020] In one embodiment, the multicasting the second protocol result includes:
[0021] Multicasting the second reduction result to the M computing nodes; or
[0022] The second reduction result is multicasted to M-1 computing nodes except the mth computing node.
[0023] A full protocol device based on switch connection, the device is applicable to the mth computing node among M computing nodes connected to the switch, the mth computing node contains N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data; the device comprises:
[0024] A first sending module, configured to send the nth block data except the mth block data among the N block data to the nth computing node among the M computing nodes, and receive the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the mth computing node;
[0025] A reduction module, configured to perform a reduction operation on the block data numbered m in the N block data and the block data numbered m in the M-1 computing nodes except the mth computing node, so as to obtain a first reduction result;
[0026] The second sending module is used to send the first protocol result to the switch, so that the switch multicasts the first protocol result.
[0027] In one embodiment, N is K times M, where K is a positive integer of at least 2;
[0028] The first sending module is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the nth computing node among the M computing nodes, where the value range of k is [1, K-1]; and receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the mth computing node;
[0029] The reduction module is used to perform a reduction operation on the block data numbered (m+k*M) in the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result;
[0030] The second sending module is used to send the second protocol result to the switch, so that the switch multicasts the second protocol result.
[0031] A full protocol system based on switch connection, the system comprising:
[0032] switch;
[0033] M computing nodes, connected to the switch;
[0034] The mth computing node among the M computing nodes contains N block data, M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data;
[0035] The m-th computing node is used to send the n-th block data among the N block data except the m-th block data to the n-th computing node among the M computing nodes; receive the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the m-th computing node; perform a reduction operation on the block data numbered m among the N block data and the block data numbered m in the M-1 computing nodes except the m-th computing node to obtain a first reduction result; and send the first reduction result to the switch so that the switch multicasts the first reduction result.
[0036] In one embodiment, N is K times M, where K is a positive integer of at least 2;
[0037] The m-th computing node is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the n-th computing node among the M computing nodes, where the value range of k is [1, K-1]; receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the m-th computing node; perform a reduction operation on the block data numbered (m+k*M) among the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result; and send the second reduction result to the switch so that the switch multicasts the second reduction result.
[0038] An electronic device, comprising:
[0039] Memory;
[0040] processor;
[0041] The memory stores an application program executable by the processor, which is used to enable the processor to execute any of the above-mentioned full protocol methods based on switch connection.
[0042] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, cause the processor to execute any of the above-described full protocol methods based on switch connection.
[0043] A program product includes a computer program, wherein when the computer program is executed by a processor, the full protocol method based on switch connection as described above is implemented.
[0044] It can be seen from the above technical scheme that in the implementation mode of the present invention, the nth block data except the mth block data among the N block data is sent to the nth computing node among the M computing nodes, and the block data numbered m in the M-1 computing nodes except the mth computing node is received; the block data numbered m in the N block data and the block data numbered m in the M-1 computing nodes except the mth computing node are subjected to reduction operation to obtain the first reduction result; the first reduction result is sent to the switch, so that the switch multicasts the first reduction result. It can be seen from the above technical scheme that in the implementation mode of the present invention, each computing node actively sends the block data to be reduced required by other computing nodes to the corresponding computing node, and the switch does not need to multicast the reduction request of each computing node, which significantly reduces the operation steps and reduces the delay. Moreover, the implementation mode of the present invention performs respective reduction calculations in respective computing nodes, which can reduce the processing pressure of the switch as a bottleneck. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is an exemplary schematic diagram of the first process of the full protocol operation in the related art.
[0046] Figure 2 It is an exemplary schematic diagram of the second process of the full protocol operation in the related art.
[0047] Figure 3 The figure is an exemplary flow chart of a full protocol method based on switch connection according to an embodiment of the present invention.
[0048] Figure 4 FIG. 4 is an exemplary structural diagram of a full protocol system based on switch connection according to an embodiment of the present invention.
[0049] Figure 5 FIG. 1 is an exemplary schematic diagram of a first process of a full protocol operation based on a switch connection according to an embodiment of the present invention.
[0050] Figure 6 FIG. 1 is an exemplary schematic diagram of a second process of a full protocol operation based on a switch connection according to an embodiment of the present invention.
[0051] Figure 7 FIG. 4 is an exemplary structural diagram of a full protocol device based on switch connection according to an embodiment of the present invention.
[0052] Figure 8 is an exemplary structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings.
[0054] For the sake of brevity and intuitiveness in description, the scheme of the present invention is explained below by describing several representative implementations. A large number of details in the implementations are only used to help understand the scheme of the present invention. However, it is obvious that the technical scheme of the present invention may not be limited to these details when implemented. In order to avoid unnecessarily obscuring the scheme of the present invention, some implementations are not described in detail, but only a framework is given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the number of a component is not specifically specified below, it means that the component can be one or more, or can be understood as at least one.
[0055] The core functions of the full reduction operation mainly include: (1) Data aggregation: Each computing node (such as GPU, CPU or server) independently calculates the local gradient or parameter of its data slice; then, these local data are aggregated (that is, reduced, such as summing or averaging, etc.). (2) Data multicast: The aggregated global data is multicast to all computing nodes to ensure that each computing node obtains the same update result. Latency is one of the key indicators to measure the performance of the full reduction, especially in large-scale distributed training, where low latency is crucial to improving the overall training efficiency. Latency is the total time required to complete a full reduction operation, including the time overhead of steps such as data transmission, reduction calculation and multicast.
[0056] The current full protocol operation based on network computing includes more execution steps, which has the disadvantages of high complexity and large delay.
[0057] Taking the four GPUs (for example, GPU0 to GPU3) connected to the switch as an example, the full reduction operation in the related art is explained. Assume that GPU0 to GPU3 respectively contain their own four block data (for example, block data 0 to block data 3). The full reduction operation in the related art may include a first process and a second process. In the first process: GPU0 to GPU3 respectively send their own reduction requests to the switch; the switch multicasts the reduction requests sent by each GPU to their own GPUs; each GPU sends its own data to be reduced associated with the reduction request to the switch; the switch performs reduction processing based on the data to be reduced sent by each GPU, and sends the reduction result to the GPU that initiated the reduction request. In the second process: GPU0 to GPU3 respectively send their own reduction results received from the switch to the switch, so that the switch multicasts the reduction results sent by each GPU to all GPUs.
[0058] Figure 1 : is an exemplary schematic diagram of the first process of the full protocol operation in the related art. For example, taking GPU0 as an example, Figure 1 The first process shown includes the following steps:
[0059] Step S1: GPU0 sends a protocol request to the switch.
[0060] Step S2: The switch multicasts the protocol request sent by GPU0 to GPU0-GPU3.
[0061] Step S3: GPU0-GPU3 respectively send the to-be-protocol data (eg, respective block data 0) associated with the protocol request sent by GPU0 to the switch.
[0062] Step S4: the switch performs reduction processing on the to-be-reduced data received from GPU0 to GPU3 (ie, the four block data 0 received from GPU0 to GPU3 respectively), and returns the reduction result to GPU0.
[0063] GPU1~GPU3 respectively synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the first process includes: GPU1 sends a protocol request to the switch; the switch multicasts the protocol request sent by GPU1 to GPU0~GPU3; GPU0~GPU3 respectively sends the data to be protocolized (for example, their respective block data 1) associated with the protocol request sent by GPU1 to the switch; the switch performs protocol processing on the data to be protocolized received from GPU0~GPU3 (that is, the 4 block data 1 received from GPU0~GPU3 respectively), and returns the protocol result to GPU1.
[0064] Similarly, GPU2-GPU3 respectively and synchronously execute the above steps to complete the first process.
[0065] After the first process is completed, the second process is executed. In the second process, GPU0-GPU3 respectively send the protocol results received from the switch to the switch, so that the switch multicasts the protocol results sent by GPU0-GPU3 to GPU0-GPU3.
[0066] Figure 2 This is an exemplary schematic diagram of the second process of the full protocol operation in the related art. For example, taking GPU0 as an example, the second process includes the following steps:
[0067] Step S5: GPU0 sends the protocol result received from the switch in step S4 (ie, the protocol result of block data 0 in GPU0, block data 0 of GPU1, block data 0 of GPU2 and block data 0 of GPU3) to the switch.
[0068] Step S6: The switch multicasts the reduction result sent by GPU0 to GPU0 to GPU3. Therefore, GPU0 to GPU3 can obtain the reduction results of block data 0 in GPU0, block data 0 in GPU1, block data 0 in GPU2 and block data 0 in GPU3.
[0069] In the second process, GPU1~GPU3 respectively synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the second process includes: GPU1 sends the protocol result received from the switch (that is, the protocol result of block data 1 in GPU0, block data 1 of GPU1, block data 1 of GPU2 and block data 1 of GPU3) to the switch; the switch multicasts the protocol result sent by GPU1 to GPU0~GPU3. Therefore, GPU0~GPU3 can also obtain the protocol results of block data 1 in GPU0, block data 1 of GPU1, block data 1 of GPU2 and block data 1 of GPU3 respectively. Similarly, GPU2~GPU3 respectively synchronously execute the above steps, so that GPU0~GPU3 can also obtain the protocol results of block data 2 in GPU0, block data 2 of GPU1, block data 2 of GPU2 and block data 2 of GPU3, as well as the protocol results of block data 3 in GPU0, block data 3 of GPU1, block data 3 of GPU2 and block data 3 of GPU3. At this point, the second process is completed.
[0070] It can be seen that the full protocol operation process of the relevant technology is complicated, and many operation steps are required to complete a full protocol operation, resulting in a long total time and thus having the disadvantage of large delay.
[0071] In the implementation of the present invention, each computing node actively sends the block data to be reduced that is needed by other computing nodes to the corresponding computing node, and the switch does not need to multicast the reduction request of each computing node, thereby significantly reducing the operation steps and reducing the delay. Moreover, unlike the related art in which the reduction processing is performed at the switch, the implementation of the present invention performs the respective reduction calculations in the respective computing nodes, which can reduce the processing pressure at the switch as a bottleneck.
[0072] The above disclosure details the technical defects existing in the related art, the causes of the technical defects, and the thinking and analysis process of overcoming the technical defects. In fact, the cognition of the above technical defects is not common knowledge in the field, but a novel discovery made by the inventor in his research. In addition, the cause tracing of the technical defects and the thinking and analysis process of overcoming the technical defects are also the gradual analysis results of the inventor in the actual research process, and are not common knowledge in the field.
[0073] Figure 3 The figure is an exemplary flow chart of a full protocol method based on switch connection according to an embodiment of the present invention. Figure 3 The method shown can be performed by a processor. The processor can be implemented as any one of a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processor (NPU), a deep learning processor (DPU), an accelerated processing unit (APU), and a general-purpose graphics processing unit (GPGPU).
[0074] Figure 3 The method shown is applicable to the mth computing node (i.e., any computing node) among the M computing nodes connected to the switch, the mth computing node contains N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data. Based on the data forwarding function of the switch, data can be transmitted between the computing nodes.
[0075] like Figure 3 As shown, the method includes:
[0076] Step 101: Send the nth block data among N block data except the mth block data to the nth computing node among M computing nodes, and receive the block data numbered m among M-1 computing nodes from M-1 computing nodes except the mth computing node.
[0077] Here, the mth computing node sends the nth block data except the mth block data to the nth computing node among the M computing nodes via the switch. It can be seen that each computing node actively sends the block data to be reduced required by other computing nodes to the corresponding computing node, and the switch does not need to multicast the reduction request of each computing node, thereby significantly reducing the operation steps and reducing the delay.
[0078] For example, assume that M is 4 (eg, computing node 0 to computing node 3, that is, the value range of m is [0,3]), and N is 4 (eg, block data 0 to block data 3, that is, the value range of n is [0,3]).
[0079] For the computing node numbered m 0 (ie, computing node 0): computing node 0 sends block data 1 of computing node 0 to computing node 1, block data 2 of computing node 0 to computing node 2, and block data 3 of computing node 0 to computing node 3.
[0080] For the computing node numbered m 1 (ie, computing node 1): computing node 1 sends block data 0 of computing node 1 to computing node 0, block data 2 of computing node 1 to computing node 2, and block data 3 of computing node 1 to computing node 3.
[0081] For the computing node numbered m as 2 (ie, computing node 2): computing node 2 sends block data 0 of computing node 2 to computing node 0, block data 1 of computing node 2 to computing node 1, and block data 3 of computing node 2 to computing node 3.
[0082] For the computing node numbered m 3 (ie, computing node 3): computing node 3 sends block data 0 of computing node 3 to computing node 0, block data 1 of computing node 3 to computing node 1, and block data 2 of computing node 3 to computing node 2.
[0083] Here, the m-th computing node receives the block data numbered m in the M-1 computing nodes from the M-1 computing nodes other than the m-th computing node.
[0084] For example, following the above example:
[0085] For the computing node numbered m 0 (ie, computing node 0): computing node 0 receives block data 0 of computing node 1 from computing node 1, receives block data 0 of computing node 2 from computing node 2, and receives block data 0 of computing node 3 from computing node 3.
[0086] For the computing node numbered m 1 (ie, computing node 1): computing node 1 receives block data 1 of computing node 0 from computing node 0, receives block data 1 of computing node 2 from computing node 2, and receives block data 1 of computing node 3 from computing node 3.
[0087] For the computing node numbered m as 2 (ie, computing node 2): computing node 2 receives block data 2 of computing node 0 from computing node 0, receives block data 2 of computing node 1 from computing node 1, and receives block data 2 of computing node 3 from computing node 3.
[0088] For the computing node numbered m 3 (ie, computing node 3): computing node 3 receives block data 3 of computing node 0 from computing node 0, receives block data 3 of computing node 1 from computing node 1, and receives block data 3 of computing node 2 from computing node 2.
[0089] Step 102: Perform a reduction operation on the block data numbered m among the N block data and the block data numbered m in the M-1 computing nodes except the mth computing node to obtain a first reduction result.
[0090] Here, the mth computing node performs a reduction operation on the block data numbered m. It can be seen that, unlike the related art in which the reduction processing is performed at the switch, the embodiments of the present invention perform respective reduction calculations in respective computing nodes, which can reduce the processing pressure at the switch as a bottleneck. Among them: the reduction operation can include: summing, averaging, maximum value, minimum value, product value, and logical operation (such as logical AND, logical OR) value, etc.
[0091] For example, following the above example:
[0092] For the computing node numbered m as 0 (i.e., computing node 0): computing node 0 performs a reduction operation on block data 0 of computing node 1, block data 0 of computing node 2, block data 0 of computing node 3, and block data 0 of computing node 0 to obtain the reduction results of 4 block data 0.
[0093] For the computing node numbered m as 1 (i.e., computing node 1): computing node 1 performs a reduction operation on block data 1 of computing node 0, block data 1 of computing node 2, block data 1 of computing node 3, and block data 1 of computing node 1 to obtain the reduction results of 4 block data 1s.
[0094] For the computing node numbered m as 2 (i.e., computing node 2): computing node 2 performs a reduction operation on the block data 2 of computing node 0, the block data 2 of computing node 1, the block data 2 of computing node 3, and the block data 2 of computing node 2 to obtain the reduction results of 4 block data 2.
[0095] For the computing node numbered m as 3 (i.e., computing node 3): computing node 3 performs a reduction operation on the block data 3 of computing node 0, the block data 3 of computing node 1, the block data 3 of computing node 2, and the block data 3 in computing node 3 to obtain the reduction results of 4 block data 3.
[0096] Step 103: Send the first protocol result to the switch, so that the switch multicasts the first protocol result.
[0097] Here, the mth computing node sends its own first protocol result to the switch, so that the switch multicasts the first protocol result, so that each computing node can obtain the first protocol result, thereby ensuring data synchronization or parameter update.
[0098] In one embodiment, multicasting the first protocol result includes the following methods:
[0099] (1) Multicast the first reduction result to M computing nodes.
[0100] In method (1), the first reduction result is multicast to all computing nodes (including the source computing node that provides the first reduction result).
[0101] (2) Multicast the first reduction result to M-1 computing nodes except the mth computing node.
[0102] In method (2), the first reduction result is multicasted to all computing nodes except the source computing node that provides the first reduction result.
[0103] For example, following the above example:
[0104] For the computing node with number m being 0 (i.e., computing node 0): computing node 0 sends the reduction result of the 4 block data 0 to the switch, and the switch then multicasts the reduction result of the 4 block data 0 to computing nodes 0 to 3.
[0105] For the computing node with number m being 1 (i.e., computing node 1): computing node 1 sends the reduction result of the 4 block data 1 to the switch, and the switch then multicasts the reduction result of the 4 block data 1 to computing nodes 0 to 3.
[0106] For the computing node numbered m as 2 (i.e., computing node 2): computing node 2 sends the reduction results of the 4 block data 2 to the switch, and the switch then multicasts the reduction results of the 4 block data 2 to computing nodes 0 to 3.
[0107] For the computing node numbered m as 3 (i.e., computing node 3): computing node 3 sends the reduction result of the 4 block data 3 to the switch, and the switch then multicasts the reduction result of the 4 block data 3 to computing nodes 0 to 3.
[0108] In one embodiment, N is equal to M. That is, the number of computing nodes is equal to the number of partitions.
[0109] In one embodiment, N is K times M, where K is a positive integer of at least 2. Figure 3 The method shown includes: sending the (n+k*M)th block data except the (m+k*M)th block data among N block data to the nth computing node among M computing nodes, wherein the value range of k is [1, K-1]; receiving the block data numbered (m+k*M) among M-1 computing nodes from M-1 computing nodes except the mth computing node; performing a reduction operation on the block data numbered (m+k*M) among N block data and the block data numbered (m+k*M) among M-1 computing nodes except the mth computing node to obtain a second reduction result; sending the second reduction result to the switch so that the switch multicasts the second reduction result. It can be seen that the embodiment of the present invention also proposes a full reduction operation mode when N is K times of M, which improves applicability.
[0110] In one embodiment, multicasting the second protocol result includes the following methods:
[0111] (1) Multicast the second reduction result to M computing nodes.
[0112] In method (1), the second reduction result is multicast to all computing nodes (including the source computing node that provides the second reduction result).
[0113] (2) Multicast the second reduction result to M-1 computing nodes except the mth computing node.
[0114] In method (2), the second reduction results are multicast to all computing nodes except the source computing node that provides the second reduction results.
[0115] For example, assume that M is 4 (for example, computing node 0 to computing node 3, that is, the value range of m is [0,3]), and N is 8 (for example, block data 0 to block data 7, that is, the value range of n is [0,7], that is, k is equal to 2).
[0116] For the computing node with number m being 0 (i.e., computing node 0): computing node 0 sends block data 1 of computing node 0 to computing node 1, block data 2 of computing node 0 to computing node 2, and block data 3 of computing node 0 to computing node 3. Moreover, computing node 0 sends block data 5 of computing node 0 to computing node 1, block data 6 of computing node 0 to computing node 2, and block data 7 of computing node 0 to computing node 3.
[0117] For the computing node with number m being 1 (i.e., computing node 1): computing node 1 sends block data 0 of computing node 1 to computing node 0, block data 2 of computing node 1 to computing node 2, and block data 3 of computing node 1 to computing node 3. Moreover, computing node 1 sends block data 4 of computing node 1 to computing node 0, block data 6 of computing node 1 to computing node 2, and block data 7 of computing node 1 to computing node 3.
[0118] For the computing node with number m being 2 (i.e., computing node 2): computing node 2 sends block data 0 of computing node 2 to computing node 0, block data 1 of computing node 2 to computing node 1, and block data 3 of computing node 2 to computing node 3. Moreover, computing node 2 sends block data 4 of computing node 2 to computing node 0, block data 5 of computing node 2 to computing node 1, and block data 7 of computing node 2 to computing node 3.
[0119] For the computing node with number m being 3 (i.e., computing node 3): computing node 3 sends block data 0 of computing node 3 to computing node 0, block data 1 of computing node 3 to computing node 1, and block data 2 of computing node 3 to computing node 2. Moreover, computing node 3 sends block data 4 of computing node 3 to computing node 0, block data 5 of computing node 3 to computing node 1, and block data 6 of computing node 3 to computing node 2.
[0120] At this point, each computing node actively sends the block data to be reduced that is needed by other computing nodes to the corresponding computing node.
[0121] Then, each computing node performs its own data reception and protocol operation process, including:
[0122] For the computing node with number m being 0 (i.e., computing node 0): computing node 0 performs a reduction operation on block data 0 of computing node 1, block data 0 of computing node 2, block data 0 of computing node 3, and block data 0 of computing node 0 to obtain the reduction results of 4 block data 0. Moreover, computing node 0 performs a reduction operation on block data 4 of computing node 1, block data 4 of computing node 2, block data 4 of computing node 3, and block data 4 of computing node 0 to obtain the reduction results of 4 block data 4.
[0123] For the computing node with number m being 1 (i.e., computing node 1): computing node 1 performs a reduction operation on block data 1 of computing node 0, block data 1 of computing node 2, block data 1 of computing node 3, and block data 1 of computing node 1 to obtain the reduction results of 4 block data 1. Moreover, computing node 1 performs a reduction operation on block data 5 of computing node 0, block data 5 of computing node 2, block data 5 of computing node 3, and block data 5 of computing node 1 to obtain the reduction results of 4 block data 5.
[0124] For the computing node with number m being 2 (i.e., computing node 2): computing node 2 performs a reduction operation on block data 2 of computing node 0, block data 2 of computing node 1, block data 2 of computing node 3, and block data 2 of computing node 2 to obtain the reduction results of 4 block data 2. Moreover, computing node 2 performs a reduction operation on block data 6 of computing node 0, block data 6 of computing node 1, block data 6 of computing node 3, and block data 6 of computing node 2 to obtain the reduction results of 4 block data 6.
[0125] For the computing node with number m being 3 (i.e., computing node 3): computing node 3 performs a reduction operation on block data 3 of computing node 0, block data 3 of computing node 1, block data 3 of computing node 2, and block data 3 of computing node 3 to obtain the reduction results of 4 block data 3. Moreover, computing node 3 performs a reduction operation on block data 7 of computing node 0, block data 7 of computing node 1, block data 7 of computing node 2, and block data 7 of computing node 3 to obtain the reduction results of 4 block data 7.
[0126] At this point, each computing node completes its own data reception and protocol operation process.
[0127] Then, each node executes the multicast process of the protocol result, which includes:
[0128] For the computing node numbered m=0 (i.e., computing node 0): computing node 0 sends the protocol results of the 4 block data 0 and the protocol results of the 4 block data 4 to the switch, and the switch then multicasts the protocol results of the 4 block data 0 and the protocol results of the 4 block data 4 to computing nodes 0 to 3 respectively.
[0129] For the computing node numbered m=1 (i.e., computing node 1): computing node 1 sends the protocol results of the 4 block data 1 and the protocol results of the 4 block data 5 to the switch, and the switch then multicasts the protocol results of the 4 block data 1 and the protocol results of the 4 block data 5 to computing nodes 0 to 3 respectively.
[0130] For the computing node numbered m as 2 (i.e., computing node 2): computing node 2 sends the protocol results of the 4 block data 2 and the protocol results of the 4 block data 6 to the switch, and the switch then multicasts the protocol results of the 4 block data 2 and the protocol results of the 4 block data 6 to computing nodes 0 to 3 respectively.
[0131] For the computing node numbered m as 3 (i.e., computing node 3): computing node 3 sends the protocol results of the 4 block data 3 and the protocol results of the 4 block data 7 to the switch, and the switch then multicasts the protocol results of the 4 block data 3 and the protocol results of the 4 block data 7 to computing nodes 0 to 3 respectively.
[0132] In the above description, the example where k is equal to 2 is used for description. Those skilled in the art can appreciate that the value of k can also be other positive integers, and those skilled in the art have no limitation on this.
[0133] The embodiment of the present invention also proposes a full protocol system based on switch connection. The system includes: a switch; M computing nodes connected to the switch; wherein the mth computing node among the M computing nodes contains N block data, M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of each computing node among the M computing nodes, and n is the number of each block data among the N block data; the mth computing node is used to send the nth block data among the N block data except the mth block data to the nth computing node among the M computing nodes; receive the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the mth computing node; perform a protocol operation on the block data numbered m among the N block data and the block data numbered m in the M-1 computing nodes except the mth computing node to obtain a first protocol result; send the first protocol result to the switch so that the switch multicasts the first protocol result.
[0134] In one embodiment, N is K times M, where K is a positive integer of at least 2; the mth computing node is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the nth computing node among the M computing nodes, where the value range of k is [1, K-1]; receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the mth computing node; perform a reduction operation on the block data numbered (m+k*M) among the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the mth computing node to obtain a second reduction result; and send the second reduction result to the switch so that the switch multicasts the second reduction result.
[0135] Figure 4 FIG. 1 is an exemplary structural diagram of a full protocol system based on switch connection according to an embodiment of the present invention. Figure 4 In the example, the computing node includes four GPUs connected to the switch, namely GPU0 to GPU3. Each GPU contains its own block data to be reduced.
[0136] The full reduction operation of the embodiment of the present invention may include a first process and a second process. In the first process: GPU0-GPU3 respectively send the block data to be reduced required by other computing nodes to the corresponding computing nodes, and each GPU calculates its own data to be reduced. In the second process: GPU0-GPU3 respectively send their own reduction results to the switch, so that the switch multicasts the reduction results sent by each GPU to all GPUs.
[0137] Figure 5 FIG. 1 is an exemplary schematic diagram of a first process of a full protocol operation based on a switch connection according to an embodiment of the present invention.
[0138] exist Figure 5 In the example, GPU0 contains block data C numbered 0. 0_0 , block data C with number 1 0_1 , block data C numbered 2 0_2 and block data C numbered 3 0_3 GPU1 contains block data C numbered 0 1_0 , block data C with number 1 1_1 , block data C numbered 2 1_2 and block data C numbered 3 1_3 GPU2 contains block data C numbered 0 2_0 , block data C with number 1 2_1 , block data C numbered 2 2_2 and block data C numbered 30_3 GPU3 contains block data C numbered 0 3_0 , block data C with number 1 3_1 , block data C numbered 2 3_2 and block data C numbered 3 3_3 .
[0139] In the first process, each GPU sends the block data to be reduced that is needed by other GPUs to the corresponding computing nodes, and each GPU calculates its own data to be reduced.
[0140] Taking GPU0 as an example, the first process includes:
[0141] Step S1: GPU0 sends block data C numbered 1 0_1 To GPU1, block data C numbered 2 0_2 To GPU2, block data C numbered 3 0_3 To GPU3. In the parallel process of GPU0 executing step S1: GPU1 sends block data C numbered 0 1_0 To GPU0, block data C numbered 2 1_2 To GPU2, block data C numbered 3 1_3 To GPU3; GPU2 sends block data C numbered 0 2_0 To GPU0, block data C numbered 1 2_1 To GPU1, block data C numbered 3 2_3 To GPU3; GPU3 sends block data C numbered 0 3_0 To GPU0, block data C numbered 1 3_1 To GPU1, block data C numbered 2 3_2 In step S1, GPU0 also processes the block data C 0_0 , block data C 1_0 , block data C 2_0 and block data C 1_0 Execute the reduction process and obtain the reduction result 0.
[0142] GPU1 to GPU3 respectively perform the above steps similar to GPU0 synchronously, so that GPU1 obtains the block data C 0_1 , block data C 1_1 , block data C 2_1 and block data C 3_1 The result of the reduction is 1, GPU2 obtains the block data C 0_2 , block data C 1_2 , block data C 2_2 and block data C 3_2The result of the reduction 2, GPU3 obtains the block data C 0_3 , block data C 1_3 , block data C 2_3 and block data C 3_3 The result of the regulation is 3.
[0143] In the second process: each GPU sends its own protocol result to the switch, so that the switch multicasts the protocol result sent by each GPU to all GPUs.
[0144] Figure 6 FIG. 1 is an exemplary schematic diagram of a second process of a full protocol operation based on a switch connection according to an embodiment of the present invention.
[0145] Taking GPU0 as an example, the second process includes:
[0146] Step S2: GPU0 carries the protocol result 0 in the multicast request and sends the multicast request to the switch.
[0147] Step S3: GPU0 multicasts the reduction result 0 to switches GPU0-GPU3, so that each GPU can obtain the reduction result 0.
[0148] GPU1 to GPU3 respectively and synchronously execute the above steps similar to GPU0, so that each GPU can also obtain reduction results 1 to 3.
[0149] The embodiments of the present invention can be applied to a variety of scenarios. For example, specific application scenarios may include:
[0150] (1) Distributed Deep Learning
[0151] Full reduction is the core algorithm for gradient synchronization in distributed deep learning. It ensures that the model parameters of all nodes remain consistent by efficiently aggregating gradients between multiple computing nodes (such as GPUs or CPUs). For example, in large-scale distributed training, each node independently calculates the gradient of its data shard, and the gradient can be synchronized by the full reduction method of the embodiment of the present invention. For another example, in a model parallel scenario, the full reduction method of the embodiment of the present invention can be used to synchronize model parameters on different nodes.
[0152] (2) High Performance Computing (HPC)
[0153] In high-performance computing, the full reduction method of the embodiment of the present invention can be used to synchronize and aggregate computing results, ensuring efficient execution of large-scale parallel tasks. For example: (2.1) Scientific computing: In simulation computing in the fields of physics, chemistry, and biology, the full reduction method of the embodiment of the present invention is used to synchronize computing results. (2.2) Matrix operations: In matrix transposition and matrix operations, the full reduction method of the embodiment of the present invention is used to efficiently synchronize and aggregate data.
[0154] (3) Large model training
[0155] As the scale of models continues to increase, full reduction algorithms play a key role in large model training. For example, in a multi-node environment, the full reduction method of the embodiment of the present invention aggregates gradients from different nodes to ensure global consistency.
[0156] (4) Reinforcement Learning
[0157] Through the full reduction method of the implementation mode of the present invention, the learning results of multiple agents can be synchronized and aggregated to ensure the consistency of the global strategy.
[0158] (5) Federated Learning
[0159] In a distributed environment, through the full reduction method of the implementation mode of the present invention, model updates on different devices can be synchronized and aggregated while protecting data privacy.
[0160] (6) Deep Learning Framework Integration
[0161] The full reduction method of the embodiment of the present invention can be widely integrated into various deep learning frameworks, such as TensorFlow, PyTorch, and Horovod, etc.
[0162] The above exemplary description describes the use scenario of the full specification method of the embodiment of the present invention. Those skilled in the art can appreciate that such description is only exemplary and is not used to limit the protection scope of the embodiment of the present invention.
[0163] Figure 7FIG. 4 is an exemplary structural diagram of a full protocol device based on switch connection according to an embodiment of the present invention. The device is suitable for an m-th computing node among M computing nodes connected to a switch, wherein the m-th computing node contains N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data; the device comprises: a first sending module 701, used for sending the n-th block data among the N block data except the m-th block data to the n-th computing node among the M computing nodes, and receiving the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the m-th computing node; a reduction module 702, used for performing a reduction operation on the block data numbered m among the N block data and the block data numbered m in the M-1 computing nodes except the m-th computing node, so as to obtain a first reduction result; and a second sending module 703, used for sending the first reduction result to the switch, so that the switch multicasts the first reduction result.
[0164] In one embodiment, multicasting the first reduction result includes: multicasting the first reduction result to M computing nodes; or, multicasting the first reduction result to M-1 computing nodes except the mth computing node. In one embodiment, N is equal to M.
[0165] In one embodiment, N is K times M, where K is a positive integer of at least 2; a first sending module 701 is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the nth computing node among the M computing nodes, where the value range of k is [1, K-1]; and receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the mth computing node; a reduction module 702 is used to perform a reduction operation on the block data numbered (m+k*M) among the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the mth computing node to obtain a second reduction result; a second sending module 703 is used to send the second reduction result to the switch so that the switch multicasts the second reduction result.
[0166] In one embodiment, multicasting the second reduction result includes: multicasting the second reduction result to M computing nodes; or multicasting the second reduction result to M-1 computing nodes except the mth computing node.
[0167] The embodiment of the present invention further provides an electronic device having a processor-memory architecture. Figure 8 is a structural diagram of an electronic device according to an embodiment of the present invention. Figure 8As shown, the electronic device includes a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, any of the above switch-connected full protocol methods is implemented. Among them, the memory 802 can be specifically implemented as a variety of storage media such as an electrically erasable programmable read-only memory (EEPROM), a flash memory (Flash memory), and a programmable program read-only memory (PROM). The processor 801 can be implemented as including one or more central processing units or one or more field programmable gate arrays, wherein the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU, a GPU, a GPGPU, an MCU or a DSP, and the like.
[0168] It should be noted that not all steps and modules in the above processes and structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The division of each module is only for the convenience of describing the functional division adopted. In actual implementation, a module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0169] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module may include specially designed permanent circuits or logic devices (such as dedicated processors, such as FPGA or ASIC) to perform specific operations. For example, specific operations can be performed in various types of chips (for example, artificial intelligence chips). The hardware module may also include programmable logic devices or circuits (such as general-purpose processors or other programmable processors) temporarily configured by software to perform specific operations. As for whether to implement the hardware module by mechanical means, or by using a dedicated permanent circuit, or by using a temporarily configured circuit (such as configured by software), it can be decided based on cost and time considerations.
[0170] The present invention also provides a machine-readable storage medium, storing instructions for making a machine perform a method as described in the present application. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code for realizing the function of any of the embodiments in the above-mentioned embodiments is stored, and the computer (or CPU or MPU) of the system or device reads out and executes the program code stored in the storage medium. In addition, the operating system etc. operated on the computer can also be completed part or all of the actual operation by instructions based on the program code. The program code read out from the storage medium can also be written to the memory provided in the expansion board inserted into the computer or to the memory provided in the expansion unit connected to the computer, and then the CPU etc. installed on the expansion board or the expansion unit are made to perform part and all of the actual operation based on the instructions of the program code, so as to realize the function of any of the embodiments in the above-mentioned embodiments. The storage medium implementation for providing the program code includes a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card and a ROM. Alternatively, the program code can be downloaded from a server computer or a cloud by a communication network.
[0171] In this article, "schematic" means "serving as an example, instance or explanation", and any diagram or implementation method described as "schematic" in this article should not be interpreted as a more preferred or more advantageous technical solution. In order to make the drawings concise, only the parts related to the present invention are schematically shown in each figure, and do not represent the actual structure of the product. In addition, in order to make the drawings concise and easy to understand, in some figures, only one of the parts with the same structure or function is schematically drawn, or only one of them is marked. In this article, "one" does not mean that the number of the relevant parts of the present invention is limited to "only one", and "one" does not mean that the number of the relevant parts of the present invention is "more than one". In this article, "upper", "lower", "front", "back", "left", "right", "inside", "outside", etc. are only used to indicate the relative position relationship between the relevant parts, rather than to limit the absolute position of these relevant parts.
[0172] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A full protocol method based on switch connection, characterized in that: The method is applicable to an m-th computing node among M computing nodes connected to a switch, wherein the m-th computing node contains N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is a number of a computing node among the M computing nodes, and n is a number of a block data among the N block data; The method comprises: Sending the nth block data except the mth block data among the N block data to the nth computing node among the M computing nodes, and receiving the block data numbered m among the M-1 computing nodes from the M-1 computing nodes except the mth computing node; Performing a reduction operation on the block data numbered m in the N block data and the block data numbered m in the M-1 computing nodes except the m-th computing node to obtain a first reduction result; The first protocol result is sent to the switch, so that the first protocol result is multicast by the switch.
2. The method according to claim 1, characterized in that The first multicast protocol result includes: Multicasting the first reduction result to the M computing nodes; or The first reduction result is multicasted to M-1 computing nodes except the mth computing node.
3. The method according to claim 1 or 2, characterized in that: The N is equal to the M.
4. The method according to claim 1, characterized in that: N is K times M, where K is a positive integer of at least 2; The method comprises: sending (n+k*M)th block data except (m+k*M)th block data among the N block data to an nth computing node among the M computing nodes, wherein the value range of k is [1, K-1]; receiving block data numbered (m+k*M) among the M-1 computing nodes from M-1 computing nodes except the mth computing node; Performing a reduction operation on the block data numbered (m+k*M) in the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result; The second protocol result is sent to the switch, so that the second protocol result is multicast by the switch.
5. The method according to claim 4, characterized in that The multicast second protocol result includes: Multicasting the second reduction result to the M computing nodes; or The second reduction result is multicasted to M-1 computing nodes except the mth computing node.
6. A full protocol device based on switch connection, characterized in that: The device is applicable to an m-th computing node among M computing nodes connected to a switch, the m-th computing node comprising N block data, wherein M is a positive integer of at least 2, N is a positive integer of at least 1, m is a number of a computing node among the M computing nodes, and n is a number of a block data among the N block data; The device comprises: A first sending module, configured to send the nth block data except the mth block data among the N block data to the nth computing node among the M computing nodes, and receive the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the mth computing node; A reduction module, configured to perform a reduction operation on the block data numbered m in the N block data and the block data numbered m in the M-1 computing nodes except the mth computing node, so as to obtain a first reduction result; The second sending module is used to send the first protocol result to the switch, so that the switch multicasts the first protocol result.
7. The device according to claim 6, characterized in that N is K times M, where K is a positive integer of at least 2; The first sending module is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the nth computing node among the M computing nodes, where the value range of k is [1, K-1]; and receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the mth computing node; The reduction module is used to perform a reduction operation on the block data numbered (m+k*M) in the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result; The second sending module is used to send the second protocol result to the switch, so that the switch multicasts the second protocol result.
8. A full protocol system based on switch connection, characterized in that: The system comprises: switch; M computing nodes, connected to the switch; The mth computing node among the M computing nodes contains N block data, M is a positive integer of at least 2, N is a positive integer of at least 1, m is the number of the computing node among the M computing nodes, and n is the number of the block data among the N block data; The m-th computing node is used to send the n-th block data among the N block data except the m-th block data to the n-th computing node among the M computing nodes, and receive the block data numbered m in the M-1 computing nodes from the M-1 computing nodes except the m-th computing node; perform a reduction operation on the block data numbered m among the N block data and the block data numbered m in the M-1 computing nodes except the m-th computing node to obtain a first reduction result; and send the first reduction result to the switch so that the switch multicasts the first reduction result.
9. The system according to claim 8, characterized in that N is K times M, where K is a positive integer of at least 2; The m-th computing node is used to send the (n+k*M)th block data among the N block data except the (m+k*M)th block data to the n-th computing node among the M computing nodes, where the value range of k is [1, K-1]; receive the block data numbered (m+k*M) in the M-1 computing nodes from the M-1 computing nodes except the m-th computing node; perform a reduction operation on the block data numbered (m+k*M) among the N block data and the block data numbered (m+k*M) in the M-1 computing nodes except the m-th computing node to obtain a second reduction result; and send the second reduction result to the switch so that the switch multicasts the second reduction result.
10. An electronic device, characterized in that: include: Memory; processor; The memory stores an application program executable by the processor, which is used to enable the processor to execute the full protocol method based on switch connection as described in any one of claims 1 to 5.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the full protocol method based on switch connection according to any one of claims 1 to 5.
12. A program product comprising a computer program, characterized in that When the computer program is executed by a processor, the full protocol method based on switch connection according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Global reduction method, apparatus, device, and storage medium
CN122601665A