Data processing device and data processing method
By realizing data multicasting in the switching device, integrating tensor parallelism and pipeline parallelism operations, the problem of tensor parallelism and pipeline parallelism lacks parallelism in artificial intelligence computing is solved, and the computing efficiency and parallelism of data processing devices are improved.
Patent Information
- Application Number
- CN202510387718.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, tensor parallelism and pipeline parallelism lack parallelism in artificial intelligence computing, resulting in reduced computing performance, serial execution between parallel operations, and AllGather operations occupy a large amount of computing node cores and bandwidth.
By realizing data multicasting in switching devices, integrating tensor parallelism and pipeline parallel operations, the number of switching plane cores in pipeline parallelism is reduced, data exchange traffic and bandwidth occupancy is reduced, and parallelism is improved.
The computing efficiency of the data processing device is improved, the calculation unit occupies and data exchange traffic of pipeline parallel operation are reduced, and the efficient integration of tensor parallelism and pipeline parallelism is realized.
Smart Images

Figure CN120335954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a data processing device and a method thereof. Background Art
[0002] In artificial intelligence technology, a large number of computing nodes are required to perform parallel operations to jointly execute computing tasks. For example, in the training or inference of large models, the tensor processing for in-layer parameter processing in large models is executed in parallel, which is usually called tensor parallelism. Similarly, the pipelining processing for inter-layer parameter processing in large models is also executed in parallel, which is usually called pipeline parallelism.
[0003] Improving the parallelism of data processing devices is one of the many technical issues in this field. Summary of the Invention
[0004] The present invention provides a data processing device and an implementation method thereof to improve the computing efficiency of the data processing device.
[0005] In a first aspect of the present invention, a data processing device is provided, and the data processing device includes:
[0006] A plurality of computing nodes coupled to a switching device for jointly executing a computing task,
[0007] One of the plurality of computing nodes multicasts the data for the computing task to each computing node in the multicast group through the switching device, so that the pipelining operation for receiving and sending the data for the computing task is integrated into the multicast of the data for the computing task.
[0008] As a possible implementation manner, the plurality of computing nodes are coupled to each computing node in the multicast group by the same local switching device, and the pipelining operation is locally integrated into the local multicast of the data for the computing task.
[0009] As a possible implementation manner, one of the plurality of computing nodes sends a request carrying the data for the computing task and the switch multicast information to the local switching device, so that: the local switching device determines the switch connection information corresponding to the carried switch multicast information according to the pre-stored mapping relationship, and multicasts the carried data to each computing node in the local multicast group according to the determined switch connection information.
[0010] Wherein,
[0011] The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of its corresponding local switching device.
[0012] As a possible implementation, the multiple computing nodes are coupled to a first switching device.
[0013] Each computing node in the multicast group includes: one of the multiple computing nodes, and each computing node coupled to a second switching device.
[0014] The first switching device exchanges data with the second switching device.
[0015] The pipelining operation performed by the second switching device is remotely integrated into the multicast of the data for the computing task performed by the first switching device.
[0016] As a possible implementation, one of the multiple computing nodes sends a request carrying data for a computing task and switch multicast information to the first switching device, such that: the first switching device transparently forwards the request to the second switching device, and the second switching device determines, according to a mapping relationship prestored therein, switch connection information corresponding to the carried switch multicast information, and multicasts the carried data to each computing node in the multicast group that is coupled to the second switching device according to the determined switch connection information.
[0017] Wherein,
[0018] The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of the second switching device corresponding to it one by one.
[0019] As a possible implementation, the request carrying data for a computing task and switch multicast information is obtained in the following manner:
[0020] One of the multiple computing nodes generates a request carrying computing node multicast information and data for a computing task, and a mapping execution unit in the one of the multiple computing nodes converts the computing node multicast information into switch multicast information.
[0021] Wherein,
[0022] The switch multicast information corresponds to more than one computing node multicast information of one of the multiple computing nodes, and corresponds one by one to the operation result of the computing node multicast information of one of the multiple computing nodes and the same connection information of each computing node in the multicast group.
[0023] As a possible implementation, the request carrying data for a computing task and switch multicast information further carries multicast selection information for indicating local multicast or remote multicast, and the multicast selection information causes the switching device to perform the following processing:
[0024] In the case where the multicast selection information indicates local multicast, the switching device looks up the mapping relationship stored in itself.
[0025] In the case where the multicast selection information indicates remote multicast, the switching device transparently forwards the request.
[0026] The second aspect of the present invention provides a data processing method, which includes:
[0027] On the side of any one of the multiple computing nodes coupled to the switching device to jointly execute a computing task,
[0028] Multicast the data for the computing task to each computing node in the multicast group through the switching device, so that the pipelined parallel operation for parallel reception and parallel transmission of the data for the computing task is integrated into the multicast of the data for the computing task.
[0029] As a possible implementation manner, the multiple computing nodes and each computing node in the multicast group are coupled to the same local switching device.
[0030] The multicasting the data for the computing task to each computing node in the multicast group includes:
[0031] Sending a request carrying the data for the computing task and the switching device multicast information to the local switching device, so that: the local switching device determines the switching device connection information corresponding to the carried switching device multicast information according to the pre-stored mapping relationship, and multicasts the carried data for the computing task to each computing node in the local multicast group according to the determined switching device connection information.
[0032] Wherein,
[0033] The mapping relationship includes: the corresponding relationship between each switching device multicast information and the switching device connection information of its corresponding local switching device.
[0034] As a possible implementation manner, the multiple computing nodes are coupled to the first switching device.
[0035] Each computing node in the multicast group includes: the any one computing node and each computing node coupled to the second switching device.
[0036] The first switching device exchanges data with the second switching device.
[0037] The multicasting the data for the computing task to each computing node in the multicast group includes:
[0038] Send a request carrying data for a computing task and switch multicast information to a first switching device, such that: the first switching device transparently transmits the request to a second switching device, and the second switching device determines, according to a pre-stored mapping relationship, switching connection information corresponding to the carried switch multicast information, and multicasts the carried data for the computing task to each computing node coupled to the second switching device in a multicast group according to the determined switching connection information.
[0039] Wherein,
[0040] The mapping relationship includes: the corresponding relationship between each switch multicast information and the switching connection information of the second switching device corresponding to it one by one.
[0041] As a possible implementation manner, the request carrying data for a computing task and switch multicast information further carries multicast selection information for indicating local multicast or remote multicast, and this multicast selection information enables the switching device to perform the following processing:
[0042] In the case where the multicast selection information indicates local multicast, the switching device searches for the mapping relationship stored in itself.
[0043] In the case where the multicast selection information indicates remote multicast, the switching device transparently transmits the request. A data processing device provided by the present invention multicasts the data for a computing task to each computing node in a multicast group through a switching device, so that the pipeline operation for receiving and sending the data for the computing task is integrated into the multicast of the data for the computing task, thereby reducing the number of cores for pipeline parallelism, reducing both the occupation of the computing unit by the all-gather operation and the traffic inside the computing node. Description of the Drawings
[0044] Figure 1 It is a schematic diagram of data in each core in the case of tensor parallelism and pipeline parallelism with a GPU as an example in the related art.
[0045] Figure 2 It is another schematic diagram of tensor parallelism and pipeline parallelism in the related art.
[0046] Figure 3 It is a schematic diagram of the data processing device in the embodiment of the present application.
[0047] Figure 4a It is a schematic diagram of a data processing device in Embodiment 1 of the present application.
[0048] Figure 4b It is a schematic diagram of data in each core with a GPU as an example of a computing node in Embodiment 1 of the present application.
[0049] Figure 5a Another schematic diagram of the data processing device according to the first embodiment of the present application.
[0050] Figure 5b Another schematic diagram of the data in each kernel with the computing node being a GPU as an example in the first embodiment.
[0051] Figure 6 A schematic diagram of the data processing device according to the second embodiment of the present application.
[0052] Figure 7 A schematic diagram of the data in each kernel with the computing node being a GPU as an example in the second embodiment.
[0053] Figure 8 A schematic flowchart of the data processing method according to the embodiment of the present application. Detailed implementation manners
[0054] In order to make the purpose, technical means and advantages of the present application clearer and more understandable, the following further describes the present application in detail with reference to the accompanying drawings.
[0055] The applicant found that in the related art of artificial intelligence, tensor parallelism and pipeline parallelism are separately processed by two computing node kernels. Among them, the kernel for tensor parallelism performs tensor calculations, such as ReduceScatter operations, AllGather operations, matrix multiplication operations, etc., and the kernel for pipeline parallelism receives and sends data.
[0056] For ease of understanding tensor parallelism and pipeline parallelism in the related art, reference can be made to Figure 1 as shown in Figure 1 A schematic diagram of the data in each kernel during tensor parallelism and pipeline parallelism with the computing node GPU as an example in the related art. In the figure, GPU0 to GPU7 are connected through a switching plane (switch) to achieve data exchange between GPUs. Among them, GPU0 to GPU3 are one layer of tensor parallelism, GPU4 to GPU7 are another layer of tensor parallelism, and the switching plane forwards data for pipeline parallelism. Each layer of tensor parallelism occupies a switching plane kernel (not shown in the figure); the data to be reduced is stored in GPU0 to GPU3 for tensor parallelism respectively; GPU0 to GPU3 perform reduction operations respectively to obtain reduced data 1 to 4 respectively to achieve the ReduceScatter operation; each of GPU0 to GPU3 sends the reduced data to GPU4 to GPU7 through the switch, such as Figure 1Among them, GPU0 sends the reduced data 1 to GPU4, GPU1 sends the reduced data 2 to GPU5, GPU2 sends the reduced data 3 to GPU6, and GPU3 sends the reduced data 4 to GPU7; each of GPUs 4 to 7 respectively sends the received reduced data to GPUs 0 to 3 or GPUs 4 to 7 via a switch, as Figure 1 shown in Figure 1 . Among them, GPU4 sends the reduced data 1 to GPUs 0 to 3 or GPUs 4 to 7 respectively, GPU5 sends the reduced data 2 to GPUs 0 to 3 or GPUs 4 to 7 respectively, GPU6 sends the reduced data 3 to GPUs 0 to 3 or GPUs 4 to 7 respectively, and GPU7 sends the reduced data 4 to GPUs 0 to 3 or GPUs 4 to 7 respectively to perform the AllGather operation.
[0057] See Figure 2 as shown Figure 2 in Figure 2 . It is another schematic diagram in the related art. GPUs 0 to 7 are respectively connected to a switching device such as a switch (switching plane). Among them, GPUs 0 to 7 are used for tensor parallelism, as shown by the differently colored blocks in the figure, and the switch performs data forwarding for pipeline parallelism. Taking the processing of GPU0 and GPU4 as an example.
[0058] GPU0 performs a reduction operation on the data to be reduced from GPUs 0 to 3 to obtain the reduced data, as shown by the red block in the figure,
[0059] GPU0 sends the obtained reduced data to GPU4 via the switch, as shown by step 1 of the red arrow in the figure,
[0060] After receiving the reduced data from the switch, GPU4 sends the received reduced data to GPUs 0 to 3 via the switch respectively, as shown by step 2 of the blue arrow in the figure.
[0061] It can be seen that in the related technologies of artificial intelligence, tensor parallelism and pipeline parallelism are executed serially. That is to say, the ReduceScatter operation, the sending and receiving operations, and the AllGather operation are executed in sequence. This not only results in a lack of parallelism between tensor parallelism and pipeline parallelism and reduces the artificial intelligence computing performance, but also the AllGather operation occupies a large number of computing node cores; the switch only performs data forwarding, and the receiving and sending operations during the forwarding process occupy a large amount of bandwidth and traffic due to tensor parallelism.
[0062] The present application provides a data processing device that realizes the fusion of tensor parallelism and pipeline parallelism based on a switching network to improve the parallelism of the data processing device.
[0063] See Figure 3 as described Figure 3 This is a schematic diagram of a data processing device according to an embodiment of the present application. The data processing device includes:
[0064] n computing nodes coupled to a switching device for jointly executing a computing task, where n is a natural number greater than 1,
[0065] One of the n computing nodes multicasts the data for the computing task to each computing node in the multicast group through the switching device, so that the pipelining operation for receiving and sending the data for the computing task is integrated into the multicast operation. In this way, when the n computing nodes are parallel, tensor parallelism and pipelining parallelism are integrated.
[0066] Among them, the multicast operation is an operation in which one of the computing nodes multicasts the data for the computing task to each computing node in the multicast group. The multicast group can be configured according to service needs, and each computing node in the multicast group has the same connection information. For example, it has the same connection identification information.
[0067] As an example,
[0068] The n computing nodes are coupled to each computing node in the multicast group through the same local switching device, and the pipelining operation is locally integrated into the local multicast operation.
[0069] For example, one of the n computing nodes sends a request carrying multicast data for the computing task and switching device multicast information to the local switching device. The local switching device determines the switching device connection information corresponding to the carried switching device multicast information according to the pre-stored mapping relationship, and multicasts the carried multicast data to each computing node in the local multicast group according to the determined switching device connection information.
[0070] Among them,
[0071] The mapping relationship includes: the corresponding relationship between each switching device multicast information and the switching device connection information of its corresponding local switching device.
[0072] As another example,
[0073] The n computing nodes are coupled to a first switching device. Each computing node in the multicast group includes: one of the n computing nodes and each computing node coupled to a second switching device. The first switching device exchanges data with the second switching device, and the remote pipelining operation is integrated into the local multicast operation.
[0074] For example, one of the n computing nodes sends a request carrying multicast data for a computing task and switch multicast information to a first switching device. The first switching device transparently transmits the request to a second switching device. The second switching device determines switch connection information corresponding to the carried switch multicast information according to the mapping relationship stored in advance, and multicasts the carried multicast data to each computing node coupled to the second switching device in the multicast group according to the determined switch connection information.
[0075] Among them,
[0076] The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of the second switching device corresponding to it one by one.
[0077] As an example, the request carrying multicast data for a computing task and switch multicast information also carries multicast selection information for indicating local multicast or remote multicast. This multicast selection information enables the switching device to perform the following processing:
[0078] In the case where the multicast selection information indicates local multicast, the switching device searches the mapping relationship stored in itself.
[0079] In the case where the multicast selection information indicates remote multicast, the switching device transparently transmits the request.
[0080] Among them, the request carrying multicast data for a computing task and switch multicast information is obtained in the following manner:
[0081] One of the n computing nodes generates a request carrying computing node multicast information and multicast data. This request is converted into switch multicast information by the mapping execution unit of this computing node.
[0082] Among them,
[0083] The switch multicast information corresponds to more than one computing node multicast information of this computing node, and corresponds one by one to the operation result of the same connection information of this computing node multicast information and each computing node in the multicast group.
[0084] The computing node multicast information corresponds to the multicast group one by one. For example, the computing node multicast identification information corresponding to the multicast group one by one, which is used to characterize the identification information of the computing node requesting multicast.
[0085] The same connection information of each computing node in the multicast group includes: the computing node connection identification information corresponding one by one to the computing node connection information of the multicast group, which is used to characterize the identification information of the computing node connected to the switching device.
[0086] The switch multicast information includes: switch multicast identification information that corresponds one-to-one with the connection identification information of the computing nodes in the multicast group. This information is used to represent the fusion information of the computing node multicast identification information and the computing node connection identification information, and this fusion information is the operation result.
[0087] The switch connection information includes: switch connection identification information that corresponds one-to-one with the switch multicast identification information. This information is used to represent the identification information of the connection between the switching device and the computing nodes.
[0088] The data processing device provided by the embodiments of the present application, based on the data multicast of the switching network, reduces the number of switching plane cores used for pipeline parallelism in the related art, enabling the fusion of data transceiver and tensor parallelism in pipeline parallel operations. This not only reduces the number of computing node cores for all-gather operations, reduces the traffic and bandwidth occupied by data exchange between the switching device and the computing nodes, but also avoids the serial processing between pipeline parallelism and tensor parallelism, which is beneficial to improving the parallelism of the data processing device and thus improving the computing efficiency.
[0089] To facilitate the understanding of the present application, the following takes the GPU as an example for illustration. It should be understood that the computing nodes in the present application are not limited to GPUs, and can be any one of a tensor processing unit (TPU), a neural network processing unit (NPU), a deep learning processing unit (DPU), an accelerated processing unit (APU), and a general-purpose graphics processing unit (GPGPU). The present application does not make any restrictions in this regard.
[0090] Embodiment 1
[0091] See Figure 4a as shown Figure 4a This is a schematic diagram of the data processing device according to Embodiment 1 of the present application. In this embodiment, each GPU in the multicast group is connected to the same switching device to fuse pipeline parallelism and tensor parallelism locally.
[0092] Each GPU used to perform computing tasks is connected to different switch connection information ports (e.g., different identification ports) of the same switching device through connection information ports with the same connection information (e.g., the same identification port), so that each GPU can be configured into the same multicast group. For example, in the figure, GPU0 is connected to port 0 of the switch through its port 0, GPU1 is connected to port 1 of the switch through its port 0, GPU2 is connected to port 2 of the switch through its port 0, GPU3 is connected to port 3 of the switch through its port 0, and GPU0 to GPU3 form a multicast group.
[0093] For each GPU, its connection configuration information is configured according to the connection information between it and the switching device. Each GPU configures its multicast group and its multicast identifier according to its connection configuration information. For example, GPU0 configures its multicast group 0, which includes GPU0 to GPU3, and the multicast identifier information of GPU0 is multicast identifier 0. Similarly, GPU1 configures its multicast group 1, which includes GPU0 to GPU3, and the multicast identifier information of GPU1 is multicast identifier 1. GPU2 configures its multicast group 2, which includes GPU0 to GPU3, and the multicast identifier information of GPU2 is multicast identifier 2. GPU3 configures its multicast group 3, which includes GPU0 to GPU3, and the multicast identifier information of GPU3 is multicast identifier 3.
[0094] It should be understood that the connection information ports with the same connection information may not be unique, that is, there may be multiple, and the specific quantity can be configured according to service requirements. For example, GPU0 can also be connected to port 4 of the switch through its port 1, GPU1 is connected to port 5 of the switch through its port 1, GPU2 is connected to port 6 of the switch through its port 1, GPU3 is connected to port 7 of the switch through its port 1. Similarly, GPU0 to GPU3 form a multicast group.
[0095] On the switch side, according to the multicast identifier information of each GPU and the same connection information (e.g., connection identifier information) of each GPU in its multicast group, a mapping relationship representing the one-to-one correspondence between the switch multicast identifier information of each GPU and the switch connection information (e.g., port identifier information) connected by each GPU is configured and stored in the switch. Among them, the switch multicast identifier of each GPU is determined according to the operation rule of the multicast identifier of this GPU and the same port identifier of each GPU in its multicast group. The operation rules of each GPU can be the same or different, and this embodiment does not limit this.
[0096] As shown in Table 1, Table 1 is a schematic illustration of the mapping relationship.
[0097] Table 1
[0098]
[0099] Each GPU sends a multi - memory load reduction instruction to the switch, and this instruction carries the data to be reduced and is sent to the switch; in response to the multi - memory load reduction instructions from each GPU, the switching device performs reduction operations respectively through in - network computing, and returns the reduced data obtained from the reduction operations to the corresponding GPU to implement the ReduceScatter operation. This step is not shown in the figure. It should be understood that the reduction operation can also be performed on the GPU itself, and the present application does not limit this.
[0100] The on - chip networks in each GPU respectively generate multi - memory store instructions. This instruction carries GPU multicast identification information and multicast data (such as reduced data) for multicasting to a multicast group. The generated multi - memory store instructions are sent to the switching device after being mapped by the mapping execution unit. The multi - memory store instructions sent carry switch multicast identification and multicast data. Only the multi - memory store instruction sent by GPU0 is shown in FIG. 4, as indicated by the red directed line segment. Among them, the mapping execution unit is used to convert GPU multicast identification information into switch multicast identification information. One GPU multicast identification information corresponds to more than one switch multicast identification information. The operation result obtained by a GPU multicast identification and the same port identification of each GPU in its multicast group according to the operation rules corresponds to a switch multicast identification. The operation rules include but are not limited to operation operations and their combinations such as string concatenation, summation, AND operation, OR operation, etc.
[0101] Refer to Table 2. Table 2 shows the switch multicast identification information sw - mcid, GPU multicast identification information gpu - mcid, and GPU port identification information corresponding to each multicast group. The case where there are multiple GPU port identification information is given in the table.
[0102] Table 2
[0103]
[0104] The switching device respectively responds to the multi - memory store instructions from each GPU, determines the switch port identification corresponding one - to - one with the switch multicast identification carried in the multi - memory store instruction according to the mapping relationship stored in the switching device, and sends the reduced data carried in the multi - memory store instruction to each GPU through the switch physical port corresponding to the determined switch port identification to implement the AllGather operation. Only the process of GPU0 multicasting the multicast data represented by the red block to GPU0 - GPU3 is shown in FIG. 4. The blue directed line segment in the figure represents the switching device for multicasting.
[0105] In this embodiment, by swapping the mapping relationships stored in the switching device, the reduced data is multicast to each GPU in the multicast group. In this way, not only is the operation of sending and receiving the reduced data required between tensor parallelism and pipeline parallelism omitted, but also the number of switching device cores used for pipeline parallelism is reduced, enabling the integration of pipeline parallelism and tensor parallelism on the same switching device.
[0106] See Figure 4b as shown in Figure 4b Figure 7 is a schematic diagram of the data in each core with a computing node being a GPU as an example in this embodiment. Each data to be reduced in GPU0 to GPU3 respectively obtains the reduced data in GPU0 to GPU3 based on in-network computing to implement the ReduceScatter operation. For example, Figure 4b in Figure 7, the reduced data 1 in GPU0, the reduced data 2 in GPU1, the reduced data 3 in GPU2, and the reduced data 4 in GPU3; after these reduced data are respectively multicast by the switch, GPU0 to GPU3 respectively obtain the reduced data 1 to 4, thus implementing the AllGather operation.
[0107] See Figure 5a as shown in Figure 5a Figure 16 is another schematic diagram of the data processing device in this embodiment. Different from Figure 4a the implementation manner, the multicast groups respectively configured for GPU0 to GPU3 are GPU4 to GPU7, and the reduced data in GPU0 to GPU3 are respectively multicast to GPU4 to GPU7. As shown in the figure, the reduced data in GPU0 is multicast to GPU4 to GPU7. In this way, not only is the operation of sending and receiving the reduced data required between tensor parallelism and pipeline parallelism omitted, but also the GPU cores used for pipeline parallelism are removed, enabling the integration of pipeline parallelism and tensor parallelism on the same switching device. Moreover, through the multicast group configuration, the reduced data can be aggregated in the required target computing nodes, improving the flexibility of data processing.
[0108] See Figure 5b as shown in Figure 5b Figure 25 is another schematic diagram of the data in each core with a computing node being a GPU as an example in this embodiment. Each data to be reduced in GPU0 to GPU3 respectively obtains the reduced data in GPU0 to GPU3 based on in-network computing to implement the ReduceScatter operation. For example, Figure 5b in Figure 25, the reduced data 1 in GPU0, the reduced data 2 in GPU1, the reduced data 3 in GPU2, and the reduced data 4 in GPU3; these reduced data are respectively multicast to GPU4 to GPU7 by the switch, thus implementing the AllGather operation.
[0109] Embodiment 2
[0110] Refer to Figure 6 as shown Figure 6 which is a schematic diagram of the data processing device of this Embodiment 2.
[0111] GPU0 to GPU3 are connected to the first switching device, and GPU4 to GPU7 are connected to the second switching device. Any one of GPU0 to GPU3, such as GPU0, and at least GPU4 to GPU7 are respectively connected to different switching device identification ports of the second switch through ports with the same port identification, so that each GPU can be configured into the same multicast group. For example, GPU0 is connected to port 0 of the first switching device through its port 0, GPU4 is connected to port 0 of the second switch through its port 0, GPU5 is connected to port 1 of the second switch through its port 0, GPU6 is connected to port 2 of the second switch through its port 0, and GPU7 is connected to port 3 of the second switch through its port 0. GPU0, GPU4 to GPU7 form a multicast group. Relatively speaking, the first switching device is a local switching device, and the second switching device is a remote local switching device.
[0112] The first switching device and the second switching device can be connected through a communication network for data interaction. The second switching device is configured and stores a mapping relationship representing the one-to-one correspondence between the switching device multicast identification information of each GPU and the port identification information of each port of the second switching device connected to each GPU in the multicast group. Among them, the switching device multicast identification is determined according to the operation result of the source GPU multicast identification and the same port identification of the GPUs in the multicast group. The source GPU is the GPU that generates the multi-memory storage instruction. In this embodiment, it includes one of GPU0 to GPU3.
[0113] After the to-be-reduced data in GPU0 to GPU3 are respectively subjected to reduction operations by the first switch based on in-network computing, the reduced data in GPU0 to GPU3 are respectively obtained to implement the ReduceScatter operation.
[0114] GPU0~GPU3 respectively generate multi-memory storage instructions, which carry a source GPU multicast identifier and reduced data multicasted to the multicast group. The generated multi-memory storage instructions are sent to the first switching device after being mapped by a mapping execution unit in the source GPU. The sent multi-memory storage instructions carry switch multicast identifier information and multicast data, wherein the mapping execution unit is used to convert the source GPU multicast identifier information into switch multicast identifier information, one source GPU multicast identifier information corresponds to more than one switch multicast identifier information, and a calculation result obtained according to a calculation rule of one source GPU multicast identifier information and the same port identifier information of each GPU in the multicast group corresponds to one switch multicast identifier information.
[0115] The first switching device transparently transmits the received multi-memory storage instruction to the second switching device,
[0116] The second switching device responds to the received multi-memory storage instruction, determines the switch port identification information corresponding to the switch multicast identification information carried in the multi-memory storage instruction according to the mapping relationship stored in the second switching device, and sends the reduced data carried in the multi-memory storage instruction to each GPU in the multicast group, such as GPU4 to GPU7 in this embodiment, through the switch physical port corresponding to the determined switch port identification information to implement the AllGather operation.
[0117] See also Figure 7 As shown, Figure 7 This is a schematic diagram of the data in each kernel in the second embodiment, taking the computing node as a GPU as an example. The data to be reduced in GPU0~GPU3 are obtained based on the network calculation to implement the ReduceScatter operation, such as Figure 7 In the figure, the reduced data 1 is in GPU0, the reduced data 2 is in GPU1, the reduced data 3 is in GPU2, and the reduced data 4 is in GPU3. After these reduced data are multicasted by the switch, GPU4 to GPU7 obtain reduced data 1 to 4 respectively, thereby realizing the AllGather operation.
[0118] This embodiment realizes multicast of multi-memory storage instructions across switches and also realizes remote AllGather operations, which not only omits the protocol data sending and receiving operations required between tensor parallelism and pipelines, but also reduces the switching device cores used for pipeline parallelism, allowing pipeline parallelism and tensor parallelism to be remotely integrated.
[0119] In view of the flexibility of local fusion of pipeline parallelism and tensor parallelism and remote fusion of pipeline parallelism and tensor parallelism required in practical applications, the multi-memory storage instructions generated in the first embodiment and the second embodiment may also carry a multicast selection identifier for characterizing local multicast or remote multicast. For example, a multicast selection identifier of 0 indicates local multicast, and an identifier of 1 indicates remote multicast. The switch performs corresponding processing according to the multicast selection identifier carried in the multi-memory storage instruction:
[0120] When the multicast selection identifier indicates local multicast, the local switching device looks up the mapping relationship it stores and performs multicast.
[0121] When the multicast selection identifier indicates remote multicast, the local switching device transparently transmits the multi-memory storage instruction to the remote switching device so that the remote switching device looks up the mapping relationship it stores and performs multicast.
[0122] See Figure 8 as shown Figure 8 is a schematic flowchart of a data processing method according to an embodiment of the present application. The method includes:
[0123] One of the multiple computing nodes coupled to the switching device to jointly execute a computing task multicasts the data for the computing task to each computing node in the multicast group through the switching device, so that the pipeline parallel operation for parallel reception and parallel transmission of the data for the computing task is fused into the multicast operation for multicasting the data for the computing task.
[0124] As an example,
[0125] One of the multiple computing nodes sends a request carrying switch multicast information and the data for the computing task to the switching device.
[0126] In response to the request, the switching device determines the switch connection information corresponding to the switch multicast information carried in the request according to the mapping relationship stored in the switching device, and multicasts the data for the computing task to each computing node in the multicast group through the corresponding port of the switch according to the determined switch connection information.
[0127] As another example, the request carrying the data for the computing task and the switch multicast information also carries multicast selection information for indicating local multicast or remote multicast.
[0128] In response to the request, the switching device determines whether it is local multicast or remote multicast according to the multicast selection information carried in the request.
[0129] In the case where the multicast selection information indicates local multicast, the switching device looks up the mapping relationship stored in itself, determines the switching device connection information corresponding to the switching device multicast information carried in the request, and according to the determined switching device connection information, multicasts the data for the computing task to each computing node in the multicast group through the corresponding port of the switching device.
[0130] In the case where the multicast selection information indicates remote multicast, the switching device transparently transmits the request to a remote switching device. The remote switching device looks up the mapping relationship stored in itself, determines the switching device connection information corresponding to the switching device multicast information carried in the request, and according to the determined switching device connection information, multicasts the data for the computing task to each computing node in the multicast group through the corresponding port of the remote switching device.
[0131] For the apparatus / network-side device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments.
[0132] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.
[0133] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A data processing device, characterized in that, The data processing device includes: A plurality of computing nodes coupled to a switching device, for jointly executing computing tasks. One of the plurality of computing nodes multicasts the data for the computing task to each computing node in the multicast group through the switching device, so that the pipelining operation for receiving and sending the data for the computing task is integrated into the multicast of the data for the computing task.
2. The data processing device according to claim 1, wherein The plurality of computing nodes and each computing node in the multicast group are coupled to the same local switching device, and the pipelining operation is locally integrated into the multicast of the data for the computing task locally.
3. The data processing device according to claim 2, wherein One of the plurality of computing nodes sends a request carrying the data for the computing task and switch multicast information to the local switching device, so that: the local switching device determines the switch connection information corresponding to the carried switch multicast information according to the pre-stored mapping relationship, and multicasts the carried data to each computing node in the local multicast group according to the determined switch connection information. Wherein, The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of its corresponding local switching device.
4. The data processing device according to claim 1, wherein The plurality of computing nodes are coupled to a first switching device. Each computing node in the multicast group includes: one of the plurality of computing nodes, and each computing node coupled to a second switching device. The first switching device exchanges data with the second switching device. The pipelining operation performed by the second switching device is remotely integrated into the multicast of the data for the computing task performed by the first switching device.
5. The data processing device according to claim 4, wherein One of the plurality of computing nodes sends a request carrying the data for the computing task and switch multicast information to the first switching device, so that: the first switching device transparently transmits the request to the second switching device, and the second switching device determines the switch connection information corresponding to the carried switch multicast information according to the pre-stored mapping relationship, and multicasts the carried data to each computing node in the multicast group coupled to the second switching device according to the determined switch connection information. Wherein, The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of its corresponding second switching device.
6. The data processing device according to claim 3 or 5, characterized in that The request carrying the data for the computing task and switch multicast information is obtained in the following manner: One of the plurality of computing nodes generates a request carrying computing node multicast information and the data for the computing task, and the mapping execution unit in one of the plurality of computing nodes converts the computing node multicast information into switch multicast information. Wherein, The switch multicast information corresponds to one or more computing node multicast information of one of the plurality of computing nodes, and corresponds one-to-one to the operation result of the computing node multicast information of one of the plurality of computing nodes and the same connection information of each computing node in the multicast group.
7. The data processing device according to claim 6, characterized in that, The request carrying the data for the computing task and switch multicast information also carries multicast selection information for indicating local multicast or remote multicast, and the multicast selection information enables the switching device to perform the following processing: In the case where the multicast selection information indicates local multicast, the switching device looks up the mapping relationship stored in itself. In the case where the multicast selection information indicates remote multicast, the switching device forwards the request transparently.
8. A data processing method, characterized in that, The method includes: On any one of a plurality of computing nodes coupled to a switching device to jointly perform a computing task, Data for the computing task is multicast to each computing node in the multicast group through the switching device, so that the pipeline parallel operations for parallel reception and parallel transmission of the data for the computing task are integrated into the multicast of the data for the computing task.
9. The data processing method according to claim 8, wherein The plurality of computing nodes are coupled to the same local switching device as each computing node in the multicast group. The multicasting of the data for the computing task to each computing node in the multicast group includes: Sending a request carrying the data for the computing task and the switch multicast information to the local switching device, so that: the local switching device determines the switch connection information corresponding to the carried switch multicast information according to the pre-stored mapping relationship, and multicasts the carried data for the computing task to each computing node in the local multicast group according to the determined switch connection information. Wherein, The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of its corresponding local switching device.
10. The data processing method according to claim 8, wherein The plurality of computing nodes are coupled to a first switching device. Each computing node in the multicast group includes: the any one computing node and each computing node coupled to a second switching device. The first switching device exchanges data with the second switching device. The multicasting of the data for the computing task to each computing node in the multicast group includes: Sending a request carrying the data for the computing task and the switch multicast information to the first switching device, so that: the first switching device forwards the request transparently to the second switching device, and the second switching device determines the switch connection information corresponding to the carried switch multicast information according to the pre-stored mapping relationship, and multicasts the carried data for the computing task to each computing node in the multicast group coupled to the second switching device according to the determined switch connection information. Wherein, The mapping relationship includes: the corresponding relationship between each switch multicast information and the switch connection information of its corresponding second switching device.
11. The data processing method according to claim 9 or 10, characterized in that The request carrying the data for the computing task and the switch multicast information further carries multicast selection information for indicating local multicast or remote multicast, and this multicast selection information enables the switching device to perform the following processing: In the case where the multicast selection information indicates local multicast, the switching device looks up the mapping relationship stored in itself. In the case where the multicast selection information indicates remote multicast, the switching device forwards the request transparently.