Communication system, communication method and related device
By offloading the computational tasks of the processing unit to the transmission unit and performing data computation directly on the transmission unit, the resource waste and latency problems caused by inter-node memory operations and RDMA module interaction are solved, thus improving the efficiency of aggregated communication.
Patent Information
- Application Number
- CN202510686603.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-10-28
AI Technical Summary
In multi-process collaborative collective communication, existing technologies involve numerous memory operations and RDMA module interactions between nodes, leading to wasted transmission resources and increased computational latency, thus affecting the efficiency of computational operations and collective communication.
The computing power of the transmission unit is expanded so that it is not only responsible for data transmission, but also for performing calculations based on the received requests from the processing unit. This reduces memory write and read operations, and allows data calculations to be performed directly on the transmission unit, reducing the number of memory operations and interactive operations.
It reduces the number of memory operations and RDMA module interactions, lowers computational latency and transmission resource consumption, and improves the efficiency of computational operations and aggregated communication.
Smart Images

Figure CN120849334A_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202410350153.2 and the original application date is March 22, 2024. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computers, and more specifically, to a communication system, communication method and related apparatus. Background Technology
[0003] With the rapid development of high-performance computing (HPC) clusters and artificial intelligence (AI), an increasing amount of data requires more processes to collaborate and process in parallel. In scenarios involving multiple processes working together, collective communication is necessary. In collective communication, the topology typically involves a large proportion of reduction computation (MPI_reduce) or reduction-and-broadcast computation (MPI_allreduce), and computational operations are a crucial part of MPI_reduce or MPI_allreduce. The execution efficiency of computational operations significantly impacts the execution efficiency of collective communication.
[0004] The data required for a node to perform computational operations, either all or part of it, typically comes from multiple other nodes. Nodes generally transmit data to other nodes via a Remote Direct Memory Access (RDMA) module. The RDMA module writes the received data into memory, and then the node reads the data from memory to perform computational operations. It is evident that the process of aggregated communication involves a large number of memory operations and interactions between RDMA modules, which not only wastes significant transmission resources but also increases the latency of computational operations, thus hindering the improvement of computational efficiency. Summary of the Invention
[0005] This application provides a communication system, communication method, and related apparatus, which helps to reduce the number of memory operations and the number of interaction operations between RDMA modules involved in the collective communication process. This not only saves transmission resources but also helps to reduce the latency of performing computational operations, thereby improving the execution efficiency of computational operations and the execution efficiency of collective communication.
[0006] In a first aspect, this application provides a communication system that may include multiple nodes, two of which are referred to as a first node and a second node. The second node may include a transmission unit and one or more processing units, the one or more processing units including a first processing unit. The first node is configured to send a message to the transmission unit, the message encapsulating first data to be passed to the first processing unit. The transmission unit is configured to receive the message and obtain a first reception request generated by the first processing unit, the first reception request including first address information describing a first storage location. The transmission unit is further configured to read second data from the first storage location according to the first address information, and perform calculation operations on the first data and the second data to obtain a calculation result.
[0007] The transmission unit is a unit with computing capabilities. This application proposes that the transmission unit not only performs data transmission tasks, such as receiving messages, but also performs computational operations on the received data according to the first receiving request generated by the first processing unit, instead of writing the received data into the memory of the first processing unit and then having the first processing unit read the data from memory again for computation. By offloading the computational tasks of the first processing unit to the transmission unit, it is beneficial to avoid or reduce memory write and read operations on the received first data in the second node, thereby reducing the number of memory operations in the system, reducing the latency and memory bandwidth consumption of computational operations, improving the execution efficiency of computational operations, and improving the execution efficiency of aggregated communication.
[0008] Furthermore, since the transmission unit determines the storage location (i.e., the first storage location) of the local operand (i.e., the second data) of the computational operation from the reception request generated by the first processing unit, which is the receiving end, rather than determining the first storage location from the message sent by the sending end, this not only helps to avoid the message carrying the first location information of the first storage location, reducing the size of the message, but also helps to avoid the first processing unit announcing the first location information to the first node before the first node sends the message, reducing the interaction between the first processing unit and the first node, saving transmission resources between nodes, reducing the latency of executing computational operations, improving the execution efficiency of computational operations, and improving the execution efficiency of collective communication.
[0009] Optionally, before sending a message, the first node can negotiate with the second node to confirm that the data sent to the first processing unit needs to be sent to the transmission unit, and to confirm the address of the transmission unit.
[0010] Optionally, the message further encapsulates first indication information, which is used to determine the receiving request generated by the first processing unit from the receiving requests generated by the one or more processing units; the transmission unit is specifically used to obtain the first receiving request according to the first indication information in the message.
[0011] Thus, when the second node also includes other processing units besides the processing unit, the transmission unit is also used to obtain the receiving request generated by the other processing units. The transmission unit obtains the first receiving request according to the first instruction information, which is beneficial to obtain the receiving request generated by the processing unit from the receiving requests generated by each processing unit, improve the accuracy of obtaining the first receiving request, and thus improve the accuracy of communication.
[0012] Furthermore, when the instruction to be executed by the first processing unit not only instructs to use the data in the first storage location and the first data of the first node to perform the calculation operation, but also instructs to use other data sent by other nodes and the data in the first storage location to perform other calculation operations, the transmission unit can obtain multiple receive requests generated sequentially by the first processing unit. When the multiple calculation operations satisfy the commutative law, after receiving the first data, the transmission unit obtains the receive request generated by the first processing unit according to the first instruction information, and uses the obtained receive request to perform the calculation operation on the first data. This is beneficial for the transmission unit to perform calculation operations on the data that arrives first, thereby reducing the execution delay of the instruction and improving the efficiency of the collective communication.
[0013] Optionally, the communication system further includes nodes other than the first node and the second node. The receiving requests generated by the first processing unit include receiving requests for receiving data from the first node and receiving requests for receiving data from the other nodes. The first indication information is also used to determine the receiving request for receiving data from the first node from the receiving requests generated by the first processing unit. In this way, even if the multiple computational operations indicated by the instruction to be executed by the first processing unit do not satisfy the commutative law, after the transmission unit receives the first data transmitted from the first node to the first processing unit, it obtains the receiving request generated by the first processing unit for receiving data from the first node according to the first indication information, and uses the receiving request to execute the corresponding computational operation. This helps ensure the correct execution of the instruction and improves the accuracy of the collective communication.
[0014] Optionally, the first data is a portion of the target data transmitted from the first node to the first processing unit. The receive request generated by the first processing unit for receiving data from the first node includes a receive request for receiving the first data and a receive request for receiving other data in the target data besides the first data. The first indication information is further used to determine the receive request for receiving the first data from the receive requests generated by the first processing unit for receiving data from the first node. Thus, even if the first node transmits data to the first processing unit through multiple messages, and the first processing unit generates multiple receive requests for receiving different data from the first node, the transmission unit, after receiving the first data, obtains the receive request for receiving the first data from the multiple receive requests according to the first indication information. This improves the accuracy of the transmission unit in processing the first data and enhances the accuracy of the aggregated communication.
[0015] Optionally, the first receiving request further includes second indication information, which is used to instruct the computation operation to be performed on the second data and the first data. By using the second indication information, this application can flexibly instruct whether to process each message sent by the first node to the transmission unit according to the receiving request.
[0016] Optionally, the first receiving request further includes third indication information, which indicates the type of computational operation. By using this third indication information in the receiving request, this application can flexibly indicate the type of computational operation performed on each message sent by the first node to the transmission unit.
[0017] Optionally, the transmission unit is further configured to preprocess the first data and / or the second data before performing the calculation operation on the first data and the second data.
[0018] Optionally, the first receiving request further includes second address information describing a second storage location, the second storage location being used to store the calculation result; the transmission unit is further configured to write the calculation result into the second storage location according to the second address information. This allows for flexible indication of the storage location of the calculation result.
[0019] Optionally, the transmission unit is further configured to write the calculation result into the first storage location according to the first address information. This helps to reduce the length of the first receiving request and saves storage resources used to store the receiving request.
[0020] Optionally, the first node and the second node can perform scheduling synchronization to ensure that the first node sends the message after the first processing unit in the second node submits the first receive request. For example, after the first processing unit submits a receive request to the transmission unit, it notifies the first node that it has submitted the first receive request. Only after receiving this notification can the first node send the message to the transmission unit. This allows the transmission unit to successfully query the first processing unit's first receive request when it receives the message, so as to read the second data and perform calculation operations on the first and second data.
[0021] Optionally, the first node and the second node may not perform scheduling synchronization. That is, the first node sending the message is independent of whether the first processing unit has submitted the first receive request. Regardless of whether the first processing unit has submitted the first receive request, the first node can still send the message. In this way, if the first node prepares the data before the first processing unit, the first node can send the message before the first processing unit submits the first receive request. This is beneficial for the transmission unit to receive the first data with less latency after receiving the first receive request, and may even receive the first data at the same time as receiving the first receive request, thereby reducing the latency of completing the calculation operation.
[0022] Optionally, the message instructing the transmission unit to obtain the first receive request can be a message retransmitted by the first node once or multiple times. For example, if the first node sends the message to the transmission unit before sending the retransmitted message, but the transmission unit cannot obtain the first receive request from the message, the transmission unit can discard the message and instruct the first node to retransmit the message. This facilitates scheduling-free synchronization between the first and second nodes, allowing the transmission unit to receive the first data with less latency after receiving the first receive request, or even receive the first data simultaneously with the receive request. This reduces the latency of completing computational operations and improves the execution efficiency of computational operations.
[0023] Optionally, the computation operation is a collective communication computation. For example, the first node and the second node, or the processing unit in the first node (such as the source of the first data, referred to as the second processing unit) and the processing unit in the second node (such as the first processing unit) are different members in a group participating in collective communication, and the computation operation is a computation operation in collective communication. Since the progress of different members in the group performing collective communication is generally relatively consistent, for example, the time when the second processing unit prepares the first data and the time when the first processing unit prepares the second data are relatively close, even if the second processing unit submits a send request to send the first data after preparing the first data, when the first data arrives at the transmission unit, the first processing unit usually also prepares the second data and submits a receive request. This helps to reduce the number of times the first node retransmits the message, thereby helping to reduce the latency of collective communication and save transmission resources between the first node and the second node.
[0024] Optionally, the set communication computation includes set communication protocol computation, or set communication protocol and broadcast computation. Optionally, the computation operation includes at least one of the following: addition operation, maximum value operation, sum operation, OR operation, XOR operation, and minimum value operation.
[0025] Optionally, the message is a Remote Memory Direct Access Protocol (RDMA) message or a Transmission Control Protocol (TCP) message.
[0026] Optionally, based on the fact that the message is an RDMA message, the transmission unit includes an RDMA module. The RDMA module can specifically be an RDMA engine or an RDMA network interface controller (RNIC). The RNIC can also be called an RDMA network card.
[0027] Optionally, based on the fact that the message is a TCP message, the transmission unit may include a TCP protocol stack.
[0028] Optionally, in addition to the RDMA module or TCP protocol stack, the transmission unit may also include other modules, which are referred to as home agents (HA) in this application. After the RDMA module or TCP protocol stack of the transmission unit receives the message and queries the storage address described in the reception request, it can send a calculation command and first data to the HA. The calculation command is used to instruct the HA to read second data from the storage address, perform calculation operations on the first and second data, and write the calculation result into memory.
[0029] HA can communicate with RDMA modules or TCP protocol stacks via the bus, as well as with memory via a non-bus interface. This helps avoid RDMA modules or TCP protocol stacks reading secondary data from memory via the bus and writing computation results to memory modules via the bus, thereby reducing latency and bandwidth consumption in collective communication.
[0030] In one possible design, in another implementation of the first aspect of this application, the memory may include dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), or double data rate synchronous dynamic random access memory (DDR SDRAM), etc., without limitation.
[0031] In one possible design, in another implementation of the first aspect of this application, the communication system may also include memory.
[0032] In one possible design, in another implementation of the first aspect of this application, the transmission unit is configured to receive a reception request from the first processing unit and store the reception request in memory, for example, in a reception queue in memory. The transmission unit is configured to query the reception request in memory or the reception queue in memory after receiving the first message. The transmission unit is also configured to delete the reception request after querying it or after receiving and processing the first data according to the instructions of the reception request.
[0033] Optionally, the transmission unit is further configured to notify the processing unit of the completion of the calculation operation after obtaining the calculation result.
[0034] In one possible design, the transmission unit is also configured to report an interrupt to the first processing unit after writing the calculation result into memory, the interrupt being used to indicate that the receiving request has been completed.
[0035] In one possible design, the transmission unit is also configured to receive polling from the first processing unit to return completion of the receiving request.
[0036] Secondly, this application provides a communication method applicable to a communication system. The communication system may include a first node and a second node. The second node may include a transmission unit and one or more processing units, with the one or more processing units including the first processing unit. The method may include: the first node sending a message to the transmission unit, the message encapsulating first data to be passed to the first processing unit; the transmission unit receiving the message and obtaining a first reception request generated by the first processing unit, the first reception request including first address information describing a first storage location; and the transmission unit reading second data from the first storage location according to the first address information and performing calculation operations on the first data and the second data to obtain a calculation result.
[0037] Optionally, before sending a message, the first node can negotiate with the second node to confirm that the data sent to the first processing unit needs to be sent to the transmission unit, and to confirm the address of the transmission unit.
[0038] Optionally, the message further encapsulates first indication information, which is used to determine the reception request generated by the first processing unit from the reception requests generated by the one or more processing units; the transmission unit obtains the first reception request according to the first indication information in the message.
[0039] Optionally, the communication system further includes nodes other than the first node and the second node. The receiving requests generated by the first processing unit include receiving requests for receiving data from the first node and receiving requests for receiving data from the other nodes. The first indication information is also used to determine the receiving request for receiving data from the first node from the receiving requests generated by the first processing unit. In this way, even if the multiple computational operations indicated by the instruction to be executed by the first processing unit do not satisfy the commutative law, after the transmission unit receives the first data transmitted from the first node to the first processing unit, it obtains the receiving request generated by the first processing unit for receiving data from the first node according to the first indication information, and uses the receiving request to execute the corresponding computational operation. This helps ensure the correct execution of the instruction and improves the accuracy of the collective communication.
[0040] Optionally, the first data is a portion of the target data transmitted from the first node to the first processing unit. The receive request generated by the first processing unit for receiving data from the first node includes a receive request for receiving the first data and a receive request for receiving other data in the target data besides the first data. The first indication information is further used to determine the receive request for receiving the first data from the receive requests generated by the first processing unit for receiving data from the first node. Thus, even if the first node transmits data to the first processing unit through multiple messages, and the first processing unit generates multiple receive requests for receiving different data from the first node, the transmission unit, after receiving the first data, obtains the receive request for receiving the first data from the multiple receive requests according to the first indication information. This improves the accuracy of the transmission unit in processing the first data and enhances the accuracy of the aggregated communication.
[0041] Optionally, the first receiving request further includes second indication information, which is used to instruct the computation operation to be performed on the second data and the first data. By using the second indication information, this application can flexibly instruct whether to process each message sent by the first node to the transmission unit according to the receiving request.
[0042] Optionally, the first receiving request further includes third indication information, which indicates the type of computational operation. By using this third indication information in the receiving request, this application can flexibly indicate the type of computational operation performed on each message sent by the first node to the transmission unit.
[0043] Optionally, the transmission unit preprocesses the first data and / or the second data before performing the calculation operation on the first data and the second data.
[0044] Optionally, the first receiving request further includes second address information describing a second storage location, the second storage location being used to store the calculation result; the transmission unit writes the calculation result into the second storage location according to the second address information. This allows for flexible indication of the storage location of the calculation result.
[0045] Optionally, the transmission unit writes the calculation result to the first storage location according to the first address information. This helps to reduce the length of the first receiving request and saves storage resources used to store the receiving request.
[0046] Optionally, the first node and the second node can perform scheduling synchronization to ensure that the first node sends the message after the first processing unit in the second node submits the first receive request. For example, after the first processing unit submits a receive request to the transmission unit, it notifies the first node that it has submitted the first receive request. Only after receiving this notification can the first node send the message to the transmission unit. This allows the transmission unit to successfully query the first processing unit's first receive request when it receives the message, so as to read the second data and perform calculation operations on the first and second data.
[0047] Optionally, the first node and the second node may not perform scheduling synchronization. That is, the first node sending the message is independent of whether the first processing unit has submitted the first receive request. Regardless of whether the first processing unit has submitted the first receive request, the first node can still send the message. In this way, if the first node prepares the data before the first processing unit, the first node can send the message before the first processing unit submits the first receive request. This is beneficial for the transmission unit to receive the first data with less latency after receiving the first receive request, and may even receive the first data at the same time as receiving the first receive request, thereby reducing the latency of completing the calculation operation.
[0048] Optionally, the message instructing the transmission unit to obtain the first receive request can be a message retransmitted by the first node once or multiple times. For example, if the first node sends the message to the transmission unit before sending the retransmitted message, but the transmission unit cannot obtain the first receive request from the message, the transmission unit can discard the message and instruct the first node to retransmit the message. This facilitates scheduling-free synchronization between the first and second nodes, allowing the transmission unit to receive the first data with less latency after receiving the first receive request, or even receive the first data simultaneously with the receive request. This reduces the latency of completing computational operations and improves the execution efficiency of computational operations.
[0049] Optionally, the computation operation is a collective communication computation. For example, the first node and the second node, or the processing unit in the first node (such as the source of the first data, referred to as the second processing unit) and the processing unit in the second node (such as the first processing unit) are different members in a group participating in collective communication, and the computation operation is a computation operation in collective communication. Since the progress of different members in the group performing collective communication is generally relatively consistent, for example, the time when the second processing unit prepares the first data and the time when the first processing unit prepares the second data are relatively close, even if the second processing unit submits a send request to send the first data after preparing the first data, when the first data arrives at the transmission unit, the first processing unit usually also prepares the second data and submits a receive request. This helps to reduce the number of times the first node retransmits the message, thereby helping to reduce the latency of collective communication and save transmission resources between the first node and the second node.
[0050] Optionally, the set communication computation includes set communication protocol computation, or set communication protocol and broadcast computation. Optionally, the computation operation includes at least one of the following: addition operation, maximum value operation, sum operation, OR operation, XOR operation, and minimum value operation.
[0051] Optionally, the message is a Remote Memory Direct Access Protocol (RDMA) message or a Transmission Control Protocol (TCP) message.
[0052] Optionally, based on the fact that the message is an RDMA message, the transmission unit includes an RDMA module. The RDMA module can specifically be an RDMA engine or an RDMA network interface controller (RNIC). The RNIC can also be called an RDMA network card.
[0053] Optionally, based on the fact that the message is a TCP message, the transmission unit may include a TCP protocol stack.
[0054] Optionally, in addition to the RDMA module or TCP protocol stack, the transmission unit may also include other modules, which are referred to as home agents (HA) in this application. After the RDMA module or TCP protocol stack of the transmission unit receives the message and queries the storage address described in the reception request, it can send a calculation command and first data to the HA. The calculation command is used to instruct the HA to read second data from the storage address, perform calculation operations on the first and second data, and write the calculation result into memory.
[0055] HA can communicate with RDMA modules or TCP protocol stacks via the bus, as well as with memory via a non-bus interface. This helps avoid RDMA modules or TCP protocol stacks reading secondary data from memory via the bus and writing computation results to memory modules via the bus, thereby reducing latency and bandwidth consumption in collective communication.
[0056] In one possible design, in another implementation of the first aspect of this application, the memory may include dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), or double data rate synchronous dynamic random access memory (DDR SDRAM), etc., without limitation.
[0057] In one possible design, in another implementation of the first aspect of this application, the communication system may also include memory.
[0058] In one possible design, in another implementation of the first aspect of this application, the transmission unit receives a reception request from the first processing unit and stores the reception request in memory, for example, in a reception queue in memory. After receiving the first message, the transmission unit queries the memory or the reception queue in memory for the reception request. After retrieving the reception request or receiving and processing the first data according to the instructions in the reception request, the transmission unit deletes the reception request.
[0059] Optionally, after obtaining the calculation result, the transmission unit may also notify the processing unit of the completion of the calculation operation.
[0060] In one possible design, the transmission unit also reports an interrupt to the first processing unit after writing the calculation result into memory. The interrupt is used to indicate that the receiving request has been completed.
[0061] In one possible design, the transmission unit also receives polling from the first processing unit to return completion of the reception request.
[0062] The transmission unit in the second aspect can be understood with reference to the transmission unit introduced in the first aspect. The methods and steps provided by the second aspect and its possible implementations, as well as the beneficial effects obtained, can be understood with reference to the corresponding actions (or functions) performed by the transmission unit in the first aspect and its corresponding implementations, and the effects of the solution in the first aspect. Other implementations of the second aspect and the beneficial effects obtained can be understood with reference to the other actions and effects that the transmission unit may perform, as introduced above, and will not be elaborated here.
[0063] Thirdly, this application provides a communication method, the method comprising: receiving a message sent by a first node, the message encapsulating first data to be passed to a first processing unit among one or more processing units; obtaining a first receiving request generated by the first processing unit, the first receiving request including first address information describing a first storage location; reading second data in the first storage location according to the first address information, and performing a calculation operation on the first data and the second data to obtain a calculation result.
[0064] Optionally, the one or more processing units are deployed on a second node other than the first node, and the method is executed by the second node or by a unit in the second node (such as the transmission unit described above).
[0065] The optional implementation methods and effects of the communication method provided in the third aspect can be understood by referring to any possible implementation method and effect performed by the transmission unit in the method provided in the second aspect, and will not be elaborated here.
[0066] Fourthly, this application provides a communication device, the communication device comprising: a communication module for receiving a message sent by a first node, the message encapsulating first data to be passed to a first processing unit among one or more processing units; a processing module for acquiring a first receiving request generated by the first processing unit, the first receiving request including first address information describing a first storage location; the processing module is further configured to read second data in the first storage location according to the first address information, and perform calculation operations on the first data and the second data to obtain a calculation result.
[0067] Optionally, the one or more processing units are deployed on a second node other than the first node, and the communication device is the second node or is deployed on the second node.
[0068] The optional steps and effects performed by the communication device can be understood by referring to the optional steps and effects performed by the processing unit in the first or second aspect, and will not be repeated here.
[0069] The communication device can be a communication equipment, or a communication module installed in or used in conjunction with a communication equipment. The communication module can be a hardware module, such as a network interface controller (or network card) or a chip. Alternatively, the communication module can be a software virtual module.
[0070] Communication equipment can also be called host equipment. This application does not limit the type of communication equipment. For example, communication equipment can be a terminal or a server. According to the function of the communication equipment, it can be a computing device, a storage device, or a network device, etc.
[0071] Communication devices and network interface cards (NICs) can both be considered computer devices. Optionally, the computer device may include a processor and a memory, the memory and the processor being coupled, the processor being used to execute the method described in the second aspect or any optional manner of the second aspect. The software virtual module may be generated by the processor through executing the method.
[0072] Optionally, these instructions are stored in external memory of the computer device. When these instructions are decoded and executed by the processor of the computer device, some or all of the contents of the instructions are temporarily stored in the internal memory of the computer device. Optionally, some of the contents of these instructions are stored in external memory of the computer device, and other parts of the contents of these instructions are stored in the internal memory of the computer device.
[0073] The chip or chip system may include one or more logic circuits to implement the method described in the second aspect or any alternative manner of the second aspect.
[0074] The fifth aspect of this application provides a computer-readable storage medium storing program code that, when run on a computing device, causes the computer device to perform the methods described in the second aspect of this application or any alternative method of the second aspect.
[0075] The sixth aspect of this application provides a computer program product containing program code that, when executed by a computing device, implements the method described in the second aspect of this application or any optional method of the second aspect.
[0076] In this application, the terms "first," "second," and "third" are used only for descriptive convenience. It should be understood that such terms can be used interchangeably where appropriate. This is merely a way of distinguishing objects with the same attributes in the description of this application. Attached Figure Description
[0077] Figure 1 This schematically illustrates the system architecture corresponding to existing aggregated communication schemes;
[0078] Figures 2-4 The system architecture corresponding to the collection communication scheme of this application is illustrated in the illustrations;
[0079] Figure 5-1 and Figure 5-2 The possible processes for node S1 and node D to perform set communication are illustrated separately.
[0080] Figure 6 and Figure 7 The possible processes for nodes S1, S2, and D to perform set communication are illustrated separately.
[0081] Figure 8 This illustration shows one possible structure of the communication device of this application. Detailed Implementation
[0082] This application's solution can be applied to a communication system comprising multiple nodes connected to each other. This application does not limit the scale of the communication system. For example, the communication system can be a single circuit board, and the multiple nodes can be multiple modules on that single board. For example, the communication system can be a rack-mounted server, and at least two of the multiple nodes can be different circuit boards within the rack-mounted server or modules deployed on different circuit boards. For example, the communication system can be a server cluster, and at least two of the multiple nodes can be different servers within the server cluster or modules deployed on different servers. When a node is a circuit board or a module on a server, the module can be a physical module or a virtual module. When a module is a virtual module, the deployment location of the module can refer to the deployment location of the physical resource on which the virtual module is based. This application does not limit the type of server; for example, the server can be a high-performance computing (HPC) server or an AI training center server, etc.
[0083] This communication system can be a system for performing collective communication, where nodes can perform collective communication computations. This application does not limit the mode of collective communication; the following description uses multi-point interface (MPI) reduction operations or integration as examples. This application does not limit the functionality of this communication system; for example, the system can perform distributed training of artificial intelligence (AI) models through collective communication.
[0084] Specifically, multiple nodes in this communication system can train the same model through a global reduction operation. During a single training iteration, different nodes can use different training data (mini-batch) to train the model and obtain the gradient of the loss function. Afterward, each node not only sends its own gradient of the loss function to other nodes but also receives gradients of the loss function from other nodes. It then integrates its own gradient with the received gradients and uses the integrated gradient to update the model. This integration includes, but is not limited to, at least one of the following operations: add, max, and, or, XOR, and min.
[0085] Figure 1 This schematically illustrates one possible structure of the communication system, such as... Figure 1 As shown, the communication system may include node S1 and node D. Node S1 and node D respectively include a processing element (PE), an interconnect controller (IC), a memory controller (MC), and dynamic random access memory (DRAM).
[0086] The PE (Process Provider) is used to train the model using training data. This application does not limit the type of PE. The PE can be a hardware module; this application does not limit the type of hardware module, for example, the PE can be a graphics processing unit (GPU), a neural processing unit (NPU), or a central processing unit (CPU). Alternatively, the PE can be a software virtual module; this application does not limit the type of software virtual module, for example, the PE can be a process or application running on a node.
[0087] The IC is used to transmit data to the PE. For example, the IC can support the Remote Direct Memory Access (RDMA) protocol, and can transmit data to the PE via RDMA messages. Specifically, the IC can be an RDMA engine or an RDMA network interface controller (NIC), which is not limited in this application. This application does not limit the transmission protocols supported by the IC. For example, the IC can support other types of transmission protocols, such as the Transmission Control Protocol (TCP), and can transmit data to the PE via TCP messages. Specifically, the IC can be a TCP protocol stack. This application does not limit the type of messages or the protocol used for transmission between ICs on different nodes; hereinafter, RDMA messages will be used as an example. This application does not limit the implementation method of the IC; for example, the IC can be a hardware module or a software virtual module.
[0088] The MC (Memory Management Module) is used to manage DRAM and is responsible for coordinating data transfer between the PE (Preinstallation Environment) and memory, and / or between the IC (Integrated Circuit) and memory. This application does not limit the implementation of the MC; for example, the MC can be a hardware module or a software virtual module.
[0089] DRAM is used to store node data (such as model training data and gradients of the loss function). This application does not limit the type of memory in the node. Figure 1 Taking DRAM as an example of memory in a node, optionally, the DRAM can be replaced with other types of memory. For example, the memory in a node can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), flash memory, or optical memory, etc. This application does not limit the number of memories in a node; more of the same or different types of memory can be configured in a node.
[0090] This application does not limit the deployment method of nodes S1 and D. For example, nodes S1 and D can be deployed on the same board, or on different boards in the same rack, or in different racks.
[0091] The following is an introduction Figure 1 The system shown illustrates the method of set communication.
[0092] After training the model with the training data, the PE (denoted as PE-D) in node D obtains the gradient of the loss function (denoted as data 0), and then writes data 0 into the DRAM (denoted as DRAM-D) of node D. After training the model with the training data, the PE (denoted as PE-S1) in node S1 obtains the gradient of the loss function (denoted as data 1), and then writes data 1 into the DRAM (denoted as DRAM-S1) of node S1.
[0093] After PE-S1 writes data 1 into DRAM-S1, assuming it reaches an instruction or calculation operation, such as "data N = data 0 + data 1", PE-S1 can send data 1 to the IC of node D (denoted as IC-D) through the IC of node S1 (denoted as IC-S1). After receiving data 1, IC-D can write data 1 into DRAM-D. This process can be referenced. Figure 1 The process corresponding to the arrow marked 1 in the middle.
[0094] After PE-D prepares data 0 and data 1 (i.e., both are written to DRAM-D), assuming it reaches an instruction or calculation operation, such as "data N = data 0 + data 1", PE-D can read data 1 from DRAM-D (this process can be found in [reference]). Figure 1 The process corresponding to the arrow marked 2 in the middle), and, read data 0 from DRAM-D (this process can be referred to...). Figure 1 (The process corresponding to the arrow marked 3 in the middle) is as follows: Then, data 1 and data 0 are integrated to obtain data 1', and data 1' is written into DRAM-D. This process can be found in [reference needed]. Figure 1 The process indicated by arrow 4 in the diagram. Afterwards, PE-D can use data 1' to update the model parameters.
[0095] exist Figure 1 In the process of the set communication shown, for node D, a total of 2 write operations are required to DRAM-D (i.e., Figure 1 The processes corresponding to arrows marked 1 and 4) and two read operations (i.e. Figure 1 The delay in the process by which node D obtains data 1' (corresponding to arrows marked 2 and 3 in the middle) includes... Figure 1 The delays of the processes corresponding to the arrows marked 1 to 4, and the calculation delay of PE-D, are shown below. Figure 1 Each arrow in the diagram represents a process that requires a read or write operation to the DRAM-D.
[0096] The bandwidth between ICs at different nodes is also called interconnect bandwidth or input / output (IO) bandwidth. Currently, interconnect bandwidth is approaching DRAM bandwidth, and the two can reach the same order of magnitude. For example, David's UB interconnect bandwidth has reached 896GB / s (unidirectional) and 1792GB / s (bidirectional), while the bandwidth of Custom-DRAM is 1638GB / s. If the data coming from the UB operates on the DRAM multiple times, then the DRAM bandwidth will become a bottleneck, which will lead to latency and bandwidth consumption. Therefore, the technical problem to be solved in this application is to minimize the amount of data that operates on the DRAM when coming from the UB, especially to avoid multiple copies of data, in order to reduce the latency and bandwidth consumption of the aggregated communication.
[0097] Figure 2 This schematically illustrates another example of a collection communication method in the system. Figure 2 The system shown can be referenced in the previous text. Figure 1 The system shown is explained in the diagram. Figure 2 and Figure 1 The difference in the communication process shown is that, Figure 2 In the illustrated collection communication process, to reduce the number of DRAM operations, IC-D receives an RDMA message (i.e., Figure 2 After the process indicated by the arrow marked 1, instead of writing data 1 to DRAM-D, data 0 is read from DRAM-D (i.e., ...). Figure 2 The process indicated by the arrow marked 2 in the diagram involves integrating data 0 and data 1 to obtain data 1', and then writing the result into DRAM-D (i.e.,...). Figure 2 (The process corresponding to the arrow marked 3 in the middle).
[0098] After IC-D writes data 1' to DRAM, PE-D no longer needs to read data 0 and data 1 separately from DRAM, nor does it need to perform an integration calculation before writing data 1' to DRAM. For node D, a total of one read operation is required for DRAM-D (i.e., Figure 2 The process corresponding to the arrow marked 2 in the middle) and one write operation (i.e. Figure 2 The delay in the process by which node D obtains data 1' (the process corresponding to arrow 3 in the middle) includes... Figure 2 The delays of the processes corresponding to the arrows marked 1 to 3, and the calculation delay of IC-D, among which only Figure 2 The processes corresponding to arrows 2 and 3 require reading or writing to the DRAM-D, which helps to reduce the number of memory copies, thereby reducing the latency and bandwidth consumption of the aggregate communication.
[0099] Figure 3 This schematically illustrates another structure of the system. Figure 3 The system shown can be used as a reference. Figure 1 or Figure 2 The system shown is understood, but, and Figure 1 or Figure 2 The systems shown are different. Figure 3 The system shown also includes a home agent (HA). The HA can be a hardware module, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., without limitation.
[0100] As the core of computational processing, the PE (Processing Module) differs significantly from memory in terms of frequency and communication / processing speed. Therefore, they generally do not communicate directly, but rather perform read or write operations on memory through their own Layer 1 (L1) cache, Layer 2 (L2) cache, and a shared Layer 3 (L3) cache. The HA (High-Availability Module), as a dedicated computing module, can connect directly to memory through a non-bus interface (such as a private or dedicated interface), matching its frequency to its communication / processing speed. This allows it to perform read or write operations on memory without needing a cache or bus.
[0101] Figure 3 Another example of a collection communication method in the system is also illustrated. Figure 2 The methods shown are different, in Figure 3 In the method shown, after receiving the RDMA message, IC-D does not read data 0 from DRAM itself and perform integration operations on data 0 and data 1. Instead, it offloads the task of accessing DRAM and performing integration operations to HA. For example, after receiving the RDMA message, IC-D can send data 1 and the calculation command to the HA of its local node (denoted as HA-D), that is... Figure 3 The process corresponding to the arrow marked 1 in the diagram. The calculation command can be used to instruct HA to read data 0 from DRAM, perform calculations on data 1 and data 0. Afterwards, HA caches data 1 according to the calculation command and reads data 0 from DRAM (i.e.,... Figure 3 The process indicated by the arrow marked 2 in the diagram involves integrating data 0 and data 1 to obtain data 1', and then writing the result into DRAM-D (i.e.,...). Figure 3 (The process corresponding to the arrow marked 3 in the middle).
[0102] The HA and DRAM can communicate via a non-bus interface, which can be a proprietary (or dedicated) interface. Therefore, when the HA accesses the DRAM, data does not need to be transmitted through the bus, thus further saving bus bandwidth. Furthermore, the HA can be positioned close to the DRAM, which helps to shorten the latency of read or write operations on the DRAM, thereby further reducing the latency of aggregated communication.
[0103] For ease of description, this application denotes the storage location of data 0 in DRAM-D as addr0. Before reading data 0, IC-D or HA-D needs to determine addr0 to integrate the received data 1 with the data 0 in addr0. The following describes the possible ways in which IC-D determines addr0.
[0104] Optionally, data 1 is transmitted between node S1 and node D using unilateral semantics. IC-S1 carries addr0 in the message, and IC-D determines addr0 from the message.
[0105] For example, PE-S1 can obtain addr0 in advance and prepare data 1. Based on data 1 and addr0, it can submit a transmission request to IC-S1. IC-S1 then sends data 1 and addr0 to IC-D via an RDMA message. After receiving the message, IC-D can determine addr0 from the message and can then read data 0 from addr0 itself or control HA to perform integrated operations on data 0 and the received data 1.
[0106] Currently, neural networks increasingly utilize dynamic graph frameworks. In dynamic graphs, the entire computation graph structure can be dynamically built and modified during program execution. This flexibility gives dynamic graphs significant advantages in the training, debugging, and deployment of deep learning models. Taking PyTorch as an example, because its data interaction address is dynamic, PE-S1 cannot statically obtain addr0 in advance; instead, it needs to obtain addr0 dynamically. For instance, after PE-D prepares data 0, PE-D needs to notify PE-S1 of addr0, i.e., perform address synchronization. This process can be referenced [reference needed]. Figure 2 or Figure 3 The arrow marked 0. Address synchronization increases the number of interactions between PE-D and PE-S1, wasting the computing resources of the PE and the network resources between the IC.
[0107] Figure 4 This schematically illustrates another example of a collection communication method in the system. Figure 4 The system shown can be referenced in the previous text. Figure 3 To understand the system, please refer to the introduction shown. However, and Figure 3 The methods shown are different, in Figure 4 In the method shown, data 1 is transmitted between node S1 and node D using bilateral semantics. That is, in order to transmit data 1, not only does PE-S1 submit a transmission request (SR) to IC-S1, but PE-D also submits a transmission request (RR) to IC-D. Furthermore, the receive request is described as addr0. This process can be referenced... Figure 4 The process corresponding to the arrow marked 0 in the diagram. Accordingly, IC-D can determine addr0 by receiving a request, which helps avoid address synchronization between PE-D and PE-S1, saving computing resources of PE and network resources between ICs, and reducing the latency of aggregated communication.
[0108] For example, after receiving a receive request, IC-D can save the receive request in its own receive queue (RQ). The receive request saved in the RQ can also be called a registered receive queue element (RQE). This process can be referenced. Figure 4 The process corresponding to the arrow marked 0 in the middle. Unlike existing RQEs that describe the write location of data to be received in node D (e.g., data 1), this application proposes an RQE that can describe the storage location (e.g., addr0) of local data (e.g., data 0) in node D.
[0109] After receiving a transmission request, IC-S1 can save the transmission request in its own transmission queue (SQ). The transmission request saved in the SQ is also called a registered transmission queue element (SQE). The SQE points to data 1 in DRAM-S1. IC-S1 schedules to this SQE, reads data 1 from DRAM-S1, encapsulates data 1 into an RDMA message, and sends the RDMA message to IC-D.
[0110] After data 1 arrives at IC-D, the aforementioned RQE is consumed. Unlike existing RQE consumption processes where IC-D writes data 1 to the storage location described by the RQE, this application proposes that after IC-D determines addr0 in the RQE, it sends data 1 and a calculation command to HA. The calculation command may include address information describing addr0. This process can be referenced... Figure 4 The process corresponding to the arrow marked 1 in the middle.
[0111] After HA receives data 1 and the calculation command, it can read data 0 from addr0 (i.e., Figure 4 (The process corresponding to the arrow marked 2 in the middle) is then performed, followed by the integration operation of data 0 and data 1 to obtain data 1', and data 1' is written to DRAM-D (i.e., Figure 4 (The process corresponding to the arrow marked 3 in the middle).
[0112] As mentioned above, Figure 4 The node shown may not include HA. Accordingly, after IC-D receives the RDMA message and obtains RQE, it reads data 0 from addr0 described by RQE, performs an integration operation on the read data 0 and the received data 1 to obtain data 1', and writes data 1' into DRAM-D.
[0113] To ensure that IC-D can successfully process data 1 carried in the received RDMA message, PE-S1 and PE-D can adopt a synchronous execution method to submit transmission requests, so that PE-S1 submits a transmission request only after PE-D submits a receive request. For example, they can perform scheduled synchronization. After submitting a receive request, PE-D sends an announcement message to PE-S1. After receiving the announcement message, PE-S1 submits a transmission request. In this way, after IC-D receives data 1, it can look up RQE in RQ, read data 0 from addr0 described by RQE, perform an integration operation on data 0 and data 1, and write the result (i.e., data 1') into DRAM.
[0114] However, assuming PE-S1 has prepared data 1 before receiving the announcement message, PE-S1 needs to wait to receive the announcement message before submitting a send request to send data 1. Thus, after receiving the receive request, IC-D can only receive data 1 after the announcement is completed and data 1 has been transmitted. This long delay in receiving data 1, coupled with the need for IC-D to use the received data 1 and data 0 read according to the receive request to perform integration, increases the latency of the aggregation communication and reduces its efficiency.
[0115] This application proposes that PE-S1 and PE-D submit transmission requests in an asynchronous, non-synchronous manner. That is, PE-S1's submission of a send request is independent of the submission of a receive request; PE-S1 can submit a send request regardless of whether PE-D has submitted a receive request. Thus, if PE-S1 prepares its data before PE-D, PE-S1 can send data 1 before PE-D submits its receive request. This allows IC-D to receive data 1 with less latency after receiving PE-D's receive request, and may even receive data 1 simultaneously with the receive request, thereby reducing the latency of the aggregated communication.
[0116] When IC-D receives a message carrying data 1, if PE-D has not yet submitted a receive request, IC-D can discard the message and send a message to IC-S1 indicating reception failure or retransmission. The following discussion uses a receive not ready (RNR) message as an example. After receiving the RNR, IC-S1 can retransmit the message immediately or after a period of silence. This allows IC-D to receive the receive request before receiving one or more retransmitted messages, thus querying the receive request upon receiving a retransmitted message, reading local data (i.e., data 0) from addr0 described in the receive request, and performing integrated operations on data 1 in the retransmitted message and local data 0, thereby resolving the problems caused by asynchronous processing.
[0117] This application does not limit the message type used by IC-D to notify IC-S1 to retransmit the message, as long as IC-S1 can retransmit the message after receiving it.
[0118] Alternatively, this application proposes that if the PE-D cannot obtain the corresponding receive request after receiving the RDMA message, it can cache the RDMA message. When the receive request is obtained, it can perform calculation operations on data 1 in the RDMA message and the local data (i.e. data 0) indicated by the receive request.
[0119] The following section provides an example of the communication method of this application, based on the scheme described above. Figure 5-1 An example of one possible approach is illustrated. Figure 5-1 The transmission element (TE) shown can be understood as the IC introduced earlier, or as including both the IC and HA in the same node introduced earlier. For example... Figure 5-1 As shown, the method provided in this application may include S501 to S514.
[0120] S501, TE-S1 and TE-D establish a connection between SQ and RQ, which is used to transmit data from PE-S1 to PE-D;
[0121] TE-S1 and TE-D can establish a connection for transmitting data from PE-S1 to PE-D. This application refers to the connection used for transmitting data from PE-S1 to PE-D as Connection 1, and the data transmitted from PE-S1 to PE-D as Data of Connection 1. Establishing Connection 1 between TE-S1 and TE-D can mean that TE-S1 and TE-D obtain / negotiate the information required to transmit the data of Connection 1 (referred to as connection establishment information), and then transmit the data from PE-S1 to PE-D according to the connection establishment information.
[0122] Considering that TE-S1 may establish connections with other TEs besides TE-D, in order to facilitate TE-S1 sending connection 1 data to TE-D, the connection establishment information obtained by TE-S1 may include the identifier of TE-D. This application does not limit the type of the TE-D identifier, as long as TE-S1 can send the connection data to TE-D based on the TE-D identifier. For example, the TE-D identifier can be a globally unique identifier (GID) or an Internet protocol address (IP) address.
[0123] Optionally, TE-S1 and TE-D can transmit data for Connection 1 via SQ and RQ. Accordingly, during the establishment of Connection 1, TE-S1 creates an SQ (referred to as the SQ of Connection 1) for Connection 1, which stores a send request indicating the transmission of data for Connection 1. TE-D creates an RQ (referred to as the RQ of Connection 1) for Connection 1, which stores a receive request indicating the reception of data for Connection 1. For ease of description, the SQ and RQ of Connection 1 will be referred to as SQ1 and RQ1, respectively, below.
[0124] Considering that TE-S1 may create other SQs for other connections, and TE-S may create other RQs for other connections, to facilitate TE-S1's accurate identification of the SQ1 and RQ1 created for connection 1, the connection establishment information obtained by TE-S1 may, for example, include the queue information of connection 1. The queue information of connection 1 can indicate the identifiers of SQ1 and RQ1. Then, TE-S1 can store the queue information of connection 1 in the queue pair context (QPC).
[0125] To facilitate accurate identification of SQ1 and RQ1 created for connection 1 by TE-D, the connection establishment information obtained by TE-D may, for example, include the queue information of connection 1. TE-D can then store the queue information of connection 1 in a queue pair context (QPC).
[0126] Optionally, to facilitate TE-S1 and TE-D in accurately identifying the chain establishment information of connection 1, TE-S1 and TE-D can associate and save the chain establishment information of connection 1 with the identifier of connection 1. This application does not limit the type of the identifier of connection 1. For example, connection 1 may include the identifier of PE-S1 and the identifier of PE-D.
[0127] S502 and PE-S1 write data 1 into DRAM-S1;
[0128] PE-S1 can write data 1 to DRAM-S1, assuming the write location is addr1. For example, data 1 could be the gradient of the loss function obtained by PE-S1 using training data to train an AI model.
[0129] S503, PE-S1 submits SR1 to TE-S1, SR1 is used to request that data 1 be sent to PE-D;
[0130] After PE-S1 writes data 1 into addr1, it can submit an SR (denoted as SR1) to TE-S1. SR1 is used to request that data 1 be sent to PE-D. For example, SR1 includes the identifier of connection 1 and the address information of addr1.
[0131] S504, TE-S1 creates SQE1 in SQ1 based on SR1;
[0132] After receiving SR1, TE-S1 can determine that the SQ of connection 1 is SQ1 based on the connection establishment information of connection 1. Then, an SQE (denoted as SQE1) is created in this SQ1. SQE1 can describe addr1 or point to the data 1 in addr1.
[0133] S505. When TE-S1 is scheduled to SQE1, the data 1 pointed to by SQE1 is encapsulated into message 1.
[0134] When TE-S1 is scheduled to SQE1, for example, when the pointer CI of SQ1 points to SQE1, TE-S1 can encapsulate the data 1 pointed to by SQE1 into a message (called message 1). Here, the pointer CI of SQ1 is used to point to the SQE to be consumed in SQ1.
[0135] S506, TE-S1 sends message 1 to TE-D;
[0136] After receiving message 1, TE-S1 can send message 1 to TE-D. The destination address of message 1 is the address of TE-D. As mentioned earlier, TE-S1 can determine the destination address of message 1 based on the connection establishment information of connection 1.
[0137] This application does not limit the timing between steps S502 and S501, as long as both S501 and S502 are executed before S503. Furthermore, step S502 is an optional step, and this application does not limit the source of data 1 in addr1. For example, data 1 can be written to addr1 by other processing units other than PE-S1.
[0138] S507, PE-D writes data 0 to addr0 in DRAM-D;
[0139] PE-D can write data 0 to DRAM-D, assuming the write location is addr0. For example, data 0 could be the gradient of the loss function obtained by PE-D using training data to train an AI model.
[0140] S508 and PE-D submit RR1 to TE-D. RR1 is used to describe addr0 in DRAM-D.
[0141] Based on the instructions executed by PE-D to perform computational operations, such as "data N = data 0 + data 1", and data 0 is ready (i.e. data 0 is stored in addr0), PE-D can submit RR (denoted as RR1) to TE-D. RR1 can include the identifier of connection 1 and the address information describing / indexing addr0. This RR1 is used to request the reception of data 1 and to perform computational operations on data 1 and the data in addr0 (i.e. data 0).
[0142] S509, TE-D creates RQE1 in RQ1 based on RR1;
[0143] After receiving RR1, TE-D can determine that the RQ of connection 1 is RQ1 based on the connection establishment information of connection 1. Then, it can create an RQE (denoted as RQE1) in RQ1. RQE1 can include the address information of addr0.
[0144] Optionally, RQE1 may include indication information (or indication field) 1, which indicates that after receiving data 1, data (i.e. data 0) is read from addr0 described by RQE1, and calculation operations are performed on data 1 and data 0.
[0145] Alternatively, message 1 may also encapsulate indication information 1. After receiving message 1 and obtaining RQE1 for receiving data 1 in message 1, TE-D determines to read data (i.e. data 0) from addr0 described by RQE1 according to indication information 1, and performs calculation operations on data 1 and data 0.
[0146] The following section will describe how TE-D determines RQE1 for receiving data 1, or in other words, how TE-D obtains RQE1 after receiving message 1. This will not be discussed in detail here.
[0147] This application does not limit the timing between steps S507 and S501, as long as both S501 and S507 are executed before S508. Furthermore, step S507 is an optional step, and this application does not limit the source of data 0 in addr0. For example, data 0 can be written to addr0 by other processing units besides PE-D.
[0148] After receiving message 1, S510 and TE-D retrieve RQE1 from RQ1;
[0149] After receiving message 1, TE-D needs to obtain RQE1 from RQ1. In addition to encapsulating data 1, message 1 may also encapsulate indication information. TE-D obtains RQE1 based on the indication information in message 1.
[0150] Considering that node D may include multiple PEs, TE-D can receive and store the RRs generated by each of the multiple PEs, and data 1 is used to transmit to PE-D. Therefore, the indication information in message 1 can be used to determine the RR generated by PE-D from the RRs generated by the multiple PEs. For example, the indication information in message 1 may include the identifier of connection 1 or the identifier of PE-D. Alternatively, TE-D can create different RQs for different PEs on node D, and the indication information in message 1 may include the identifier of RQ1. The RQE obtained by TE-D in RQ1 is the RR generated by PE-D.
[0151] Considering that PE-S1 may need to transmit n data points to PE-D sequentially, where n is a positive integer greater than 1, and data 1 is only a part of the data in this connection, correspondingly, TE-S1 receives n SRs from connection 1 sequentially, and TE-D receives n RRs from connection 1 sequentially. TE-S1 can store n SQEs sequentially in SQ1, and TE-D can store n RQEs sequentially in RQ1. The i-th SQE in SQ1 and the i-th RQE in RQ1 are used to transmit the same data, where i is any positive integer less than or equal to n. Furthermore, TE-S1 and TE-D consume elements in the queue in sequence. In this way, TE-D can identify the RQE1 to be consumed in RQ1 as the RQE used to receive data 1. As shown in Figure 5, the pointer CI of RQ1 can be used to point to the storage location of the RQE to be consumed in RQ1, and TE-D can retrieve the RQE to be consumed in RQ1 according to the storage location pointed to by the pointer CI.
[0152] Considering that TE-S1 may consume other SQEs (denoted as SQE0) before consuming SQE1 in SQ1, and sends a message (denoted as message 0) based on SQE0, and that message 1 may arrive at TE-D before message 0, in this case, the RQE pointed to by the CI of RQ1 when TE-D receives message 1 is the RQE used to receive message 0. To improve the accuracy of the calculation operation performed by TE-D on data 1, the indication information in message 1 can also be used to determine the RQE used to receive data 1 from RQ1. For example, the indication information in message 1 may include a sequence identifier, which indicates the order of SQE1 in SQ1. TE-D can determine whether the RQE to be consumed in RQ1 is used to receive data 1 in message 1 based on the sequence identifier in message 1. Assuming the sequence identifier indicates the order as j, where j is a positive integer less than or equal to n, when the RQE to be consumed in RQ1 is the j-th RQE in RQ1, or in other words, when the order of the RQEs to be consumed in RQ1 is j, TE-D can determine that the RQE to be consumed in RQ1 is used to receive data 1 in message 1; otherwise, it determines that the RQE to be consumed in RQ1 is not used to receive data 1 in message 1. When TE-D determines that the RQE to be consumed in RQ1 is not used to receive data 1 in message 1, TE-D can buffer message 1. When the RQE to be consumed in RQ1 is the j-th RQE, TE-D can use that RQE to process data 1, or TE-D can discard message 1 and notify TE-S1 to retransmit message 1.
[0153] S511, TE-D reads data 0 pointed to by RQE1;
[0154] After TE-D obtains RQE1, it can read data 0 from addr0 according to the instructions of RQE1.
[0155] S512 and TE-D perform calculations on the read data 0 and data 1 in message 1 to obtain the calculation result;
[0156] After the TE-D reads data 0, it can perform calculation operations on the read data 0 and data 1 in message 1 according to the instructions of RQE1 to obtain the calculation result. This application does not limit the type of calculation operation performed by the TE-D on data 0 and data 1. As described above, the calculation operation can be a set communication calculation or a calculation operation in a set communication process or algorithm. For example, the calculation operation can include the integration operation described above. For example, the calculation operation can include, but is not limited to, at least one of the following: add operation, max operation, and operation, or operation, XOR operation, and min operation.
[0157] This application does not limit the method by which the TE-D determines the type of the computation operation. Optionally, RQE1 may also include indication information 2, which indicates the type of the computation operation. After obtaining RQE1, the TE-D can perform the computation operation of the type indicated by indication information 2 on data 1 and data 0. Alternatively, message 1 may also encapsulate indication information 2, and the TE-D can determine the type of the computation operation based on indication information 2 in message 1.
[0158] Optionally, before performing calculations on the read data 0 and data 1 in message 1, the TE-D can preprocess data 0 and / or data 1. Afterwards, the TE-D performs calculations on data 1 and the preprocessed data 0, or on data 0 and the preprocessed data 1, or on the preprocessed data 0 and the preprocessed data 1. This application does not limit the type of preprocessing; for example, preprocessing may include one or more operations such as encryption, decryption, and format conversion.
[0159] S513 and TE-D write the calculation results to addr0;
[0160] After obtaining the calculation result, TE-D can save the calculation result, for example, by writing the calculation result to addr0. That is to say, the storage location addr0 described by the location information in RQE1 is not only used to store the data 0 that TE-D needs to read, but also used to write the calculation result of data 0 and data 1.
[0161] This application does not limit the location where the TE-D writes the calculation result. For example, the TE-D can write the calculation result to a storage location other than addr0 (denoted as addr2). addr2 can be located in a storage location other than addr0 in DRAM-D, or in a memory other than DRAM-D.
[0162] This application does not limit the method by which TE-D determines addr2. For example, message 1 may encapsulate location information describing addr2, or the QPC may store the location information of addr2, or RQE1 may include the location information of addr2.
[0163] Step S513 is an optional step. For example, after TE-D obtains the calculation result, it may not write the calculation result to the memory, but instead send the calculation result to PE-D.
[0164] S514, TE-D notifies PE-D of the completion of the calculation operation;
[0165] After writing the calculation result to addr0, TE-D can notify PE-D of the completion of the calculation operation (i.e., completion of RQE1 or RR1). This application does not limit the way TE-D notifies PE-D of the completion of the calculation operation. For example, TE-D can indicate the completion of RQE1 or RR1 by reporting an interruption to PE-D.
[0166] Step S514 is an optional step. For example, PE-D can determine whether RQE1 or RR1 has been completed by polling TE-D.
[0167] Step S501 is an optional step, as long as TE-S1 can pass data 1 to TE-D, and TE-D can obtain RQE1 for receiving data 1.
[0168] S502–S506 represent the data transmission process executed by node S1, and S507–S509 represent the data reception preparation process executed by node D. These two processes are executed asynchronously; that is, this application does not limit the temporal order between any step in the data transmission process and any step in the data reception preparation process. Therefore, in one possible implementation, when TE-D receives message 1, S509 has not yet been completed. In this case, TE-D's query for RQE1 based on message 1 fails. Optionally, after TE-D receives message 1 and fails to query RQE1, it can buffer message 1 and execute S510 after executing S509. Alternatively, to conserve cache, as mentioned earlier, after TE-D receives message 1 and fails to query RQE1, it can discard message 1 and instruct TE-S1 to retransmit message 1. This allows TE-D to complete S509 before receiving one or more retransmitted messages of message 1, thereby obtaining RQE1 based on the received retransmitted message 1 and then performing the calculation operation.
[0169] Communication methods based on this retransmission mechanism can be as follows: Figure 5-2 As shown, Figure 5-2 The method shown is the same as Figure 5-1The difference in the method shown is that TE-D executes step S507 only after receiving the first message 1 sent by TE-S1 according to SQE1, and after TE-D receives the first message 1, Figure 5-2 The method shown may also include S515 to S517. Figure 5-2 The steps S501 to S514 of the method shown can be understood by referring to the introduction of the corresponding steps in the previous text, and will not be repeated here. S515 to S517 are introduced below.
[0170] When S515 and TE-D receive message 1, they cannot obtain RQE1 in RQ1;
[0171] When TE-D receives message 1, it cannot find RQE1 in RQ1 because RQE1 has not yet been created in RQ1.
[0172] S516. If the query based on RQE1 fails, TE-D sends RNR to TE-S1;
[0173] As described above, if TE-D cannot obtain the RQE (i.e., RQE1) used to receive the data (i.e., data 1) in message 1 after receiving message 1, TE-D can send an RNR to TE-S1 to instruct TE-S1 to retransmit message 1. As described above, this application does not limit the type of message sent by TE-D to TE-S1 to instruct retransmission; RNR is merely an example.
[0174] S517, TE-S1 retransmits message 1 to TE-D;
[0175] After TE-D sends an RNR to TE-S1, TE-S1 can receive the RNR. TE-S1 can then retransmit message 1 to TE-D. TE-S1 can retransmit message 1 immediately or after a period of silence.
[0176] like Figure 5-2 As shown, before TE-D receives the retransmitted message 1, it can complete S507 to S509. Therefore, after TE-D receives the retransmitted message 1, it can execute S510 to S514 in sequence.
[0177] Figure 5-2Taking the example that TE-D can query RQE1 upon receiving the first retransmitted message 1, Optionally, TE-D can query RQE1 only after receiving multiple retransmitted messages 1. In other words, if S515-S517 is referred to as a single retransmission process, after TE-D receives the first message 1, before S510, this method can include one or more retransmission processes. However, for different members participating in the collective communication, their progress in executing the collective communication is generally relatively consistent; that is, the time when PE-S1 prepares data 1 and the time when PE-D prepares data 0 are relatively close. Therefore, in the scenario of collective communication, the number of retransmissions is usually small, even as... Figure 5-1 As shown, when data 1 arrives at TE-D, TE-D has already created RQE1, which helps reduce the latency of aggregated communication and saves transmission resources between TE-S1 and TE-D.
[0178] Upon receiving any message, TE-D can determine whether the data in the message needs to be processed according to RQE in the RQ, as exemplified by the example where TE-D uses the opcode type "send" in the header of the RDMA message. Optionally, TE-D can also determine whether the data in the message needs to be processed according to RQE in the RQ by determining that the opcode type is another semantic type. That is, TE-S1 can send data 1 using other types of messages. Alternatively, TE-D can determine whether the data in the message needs to be processed according to RQE in the RQ by the values of other fields in the header of the RDMA message.
[0179] The preceding example illustrates how PE-S1 and PE-D use bilateral semantics to request data transfer, where PE-S1 submits an SR and PE-D submits an RR to achieve data transfer between them. Optionally, the communication system can support more data transfer semantics; for example, PE-S1 can transfer data between them by submitting an SR or PE-D can transfer data by submitting an RR. To facilitate TE-D's identification of data transfer semantics, at least one of message 1, SQE1, and QPC may include indication information 3, which indicates the type of data transfer semantics. For example, indication information 3 can be an opcode type field carried in message 1, which can indicate send, send with invalidate, send with immediate, RDMA write, RDMA write with immediate, or RDMA read, etc. This embodiment does not limit this. TE-D uses this field to identify whether to process data 1 in message 1 according to RQE1. For example, when this field in message 1 indicates a send operation, the TE-D can obtain RQE1 from RQ1 and process data 1 according to RQE1. This process can be referred to in S510 to S513. When this field in message 1 indicates an RDMA write operation, the TE-D can no longer obtain RQE1 from RQ1, but instead process data 1 in message 1 according to the indication in message 1.
[0180] The above example uses node S1 as the data sender and node D as the data receiver. Optionally, the roles of node S1 and node D can be interchanged. For example, node S1 can receive data 0 from node D and perform calculations using its local data 1 and the received data 0. The method flow executed by node S1 as the receiver is the same as or similar to the method flow executed by node D as described above, and will not be repeated here. Similarly, the method flow executed by node D as the sender is the same as or similar to the method flow executed by node S1 as described above, and will not be repeated here.
[0181] The above describes the communication method between two nodes in a communication system. Optionally, the communication system may include more nodes. The method flow executed by other nodes is the same as or similar to that executed by node S1 and / or node D.
[0182] like Figure 6As shown, the communication system may also include node S2, which can be used to send data 2 to node D. Assume that the operands of the instruction to be executed by PE-D include data 0, data 1, and data 2. For example, the instruction might instruct the execution of "data N = data 0 + data 1 + data 2". After PE-D prepares data 0, it can generate RR1 and RR2 respectively. RR1 instructs the receiving of data 1, performing an addition operation on data 0 and data 1 to obtain the result (denoted as data 1'). RR2 instructs the receiving of data 2, performing an addition operation on data 1' and data 2 to obtain the result (denoted as data N). Figure 6 As shown, another method example provided in this application may include S601 to S622.
[0183] S601, TE-D and node S1 establish connection 1;
[0184] S602, TE-D and node S2 establish connection 2;
[0185] The process of establishing a connection between TE-D and a node can be understood with reference to S501. For example, the process of TE-D and node S1 establishing connection 1 may include the two negotiating the chain establishment information for connection 1, and TE-D creating the RQ for connection 1 (denoted as RQ1). The process of TE-D and node S2 establishing connection 2 may include the two negotiating the chain establishment information for connection 2, and TE-D creating the RQ for connection 2 (denoted as RQ2).
[0186] When node S1 includes multiple PEs, establishing connection 1 between TE-D and node S1 can refer to establishing connection 1 between TE-D and a specific PE in node S1 (e.g., PE-S1). Similarly, when node S2 includes multiple PEs, establishing connection 2 between TE-D and node S2 can refer to establishing connection 2 between TE-D and a specific PE in node S2 (e.g., PE-S2).
[0187] This application does not limit the timing or sequence of S601 and S602.
[0188] S603, PE-D writes data 0 to addr0 in DRAM-D;
[0189] S604, PE-D submits RR1 to TE-D, RR1 includes address information describing addr0;
[0190] S605, TE-D creates RQE1 in RQ1 based on RR1;
[0191] S603 to S605 can be understood by referring to S507 to S509, and will not be elaborated here.
[0192] S606, Node S2 sends message 2 to TE-D;
[0193] After node S2 has prepared data 2, it can send data 2 to TE-D via message 2. See S606 for reference. Figure 5-1 The concepts shown in S503 to S506 will be understood and will not be elaborated upon here.
[0194] S607 and TE-D cannot obtain the RQE (denoted as RQE2) used to receive data 2 in message 2;
[0195] After receiving message 2, TE-D can query RQE2, which is used to receive data 2 in message 2. RQE2 can be RR2 stored in RQ2. As mentioned earlier, RR2 is used to perform calculation operations on data 2 and data 1'. Since data 1' is not yet ready, TE-D can temporarily not generate RR2. Consequently, TE-D cannot obtain RQE2 from RQ2.
[0196] The method of TE-D recognizing RQE2 can be understood by referring to the method of TE-D recognizing RQE1 mentioned above, and will not be repeated here.
[0197] S608 and TE-D discard message 2 and send RNR to node S2;
[0198] S609, Node S2 retransmits message 2 according to the instructions of RNR;
[0199] S607~S609 can be understood by referring to the retransmission process of S515~S517 introduced above. For example, after TE-D receives message 2, TE-D and node S3 can execute one or more retransmission processes, which will not be elaborated further.
[0200] S610, Node S1 sends message 1 to TE-D;
[0201] After node S1 has prepared data 1, it can send data 1 to TE-D via message 1. S610 can be understood by referring to S503 to S506.
[0202] S611 and TE-D obtain RQE1 in RQ1;
[0203] Since the RR generated by PE-D includes requests to receive data from different senders, the indication information in message 1 described above can also be used to determine the receive request for receiving data from node S1 from the receive request generated by PE-D, in order to facilitate the differentiation of different senders. For example, the indication information in message 1 includes the identifier of RQ1 or the identifier of PE-S1.
[0204] S612 and TE-D read data 0 from addr0 according to the instruction of RQE1;
[0205] S613 and TE-D perform addition on data 0 and data 1 to obtain the calculation result (i.e. data 1');
[0206] S614, TE-D writes data 1' to addr0;
[0207] S615, TE-D notifies PE-D of the completion of RR1;
[0208] S611 to S614 can be understood by referring to S510 to S514 mentioned above, and will not be repeated here.
[0209] S616, PE-D submits RR2 to TE-D, RR2 describes addr0 in DRAM-D;
[0210] After S615, PE-D determines that data 1' is ready and can submit RR2, as described above, to TE-D. RR2 describes addr0 in DRAM-D. Optionally, as described in S513, TE-D can write data 1' to addr2. In this case, RR2 may include address information describing addr2, rather than address information describing addr0.
[0211] S617, TE-D creates RQE2 in RQ2 based on RR2;
[0212] S616 and S617 can be understood by referring to S508 and S509.
[0213] S618 and TE-D obtain RQE2 in RQ2;
[0214] Assuming that TE-D executes S617 before receiving message 2 from node S2 after one or more retransmissions, then after receiving message 2, TE-D can query RQ2 to obtain RQE2.
[0215] Since the RR generated by PE-D includes requests to receive data from different senders, the indication information in message 2 can also be used to determine the receive request for receiving data from node S2 from the receive request generated by PE-D in order to facilitate the differentiation of different senders. For example, the indication information in message 2 includes the identifier of RQ2 or the identifier of PE in node 2.
[0216] S619 and TE-D read data 1' from addr0 according to the instruction of RQE2;
[0217] S620 and TE-D perform addition on data 1' and data 2 to obtain the calculation result (i.e. data N);
[0218] S621, TE-D writes data N to addr0;
[0219] S622, TE-D notifies PE-D of the completion of RR2.
[0220] S618 to S622 can be understood by referring to S510 to S514, and will not be elaborated here.
[0221] like Figure 6 As shown, for node D in the system to receive data, other different nodes in the system can correspond to different RQs. After TE-D receives an RR for receiving data sent by a certain node or a PE on that node, it creates an RQE in the corresponding RQ of that node. Furthermore, after receiving a message from that node or PE, it queries the RQE in the corresponding RQ of that node. Thus, as... Figure 6 As shown, even if message 2 arrives at message 1 before message 1, it still helps ensure that TE-D performs calculations on each operand according to the order of operations indicated by the instruction. Even when multiple calculation operations in an instruction do not satisfy the commutative law, it still helps ensure the correct execution of the instruction.
[0222] The execution of S610 by node S1 and the execution of S606 by node S2 are asynchronous; that is, this application does not limit the timing between S610 and S606. S603–S605 are the preparation process for node D to receive data 1. As described above, S610 and this process are asynchronous, and S606 and this process are also asynchronous. S616–S617 are the preparation process for node D to receive data 2. As described above, S606 and the preparation process for receiving data 2 are asynchronous.
[0223] Assuming the instruction to be executed by PE-D still indicates "data N = data 0 + data 1 + data 2", considering that the calculation operation in this instruction satisfies the commutative law, if data 2 arrives at TE-D before data 1, TE-D first calculates data 0 and data 2 to obtain the result (denoted as data 2'). Then, when data 1 arrives, it calculates data 1 and data 2' to obtain the result (i.e., data N). This improves the computational efficiency of the set communication. To achieve the first-come, first-served (FFS) calculation, this application proposes that TE-D place RR1 and RR2 in the same RQ. After PE-D prepares data 0, it places RR1 and RR2 in the same RQ. This allows RR1 to be used to receive data 1 or data 2, and similarly, RR2 to be used to receive data 1 or data 2. Thus, when data 2 arrives before data 1, TE-D uses RR1 to calculate data 2, and then uses RR2 to calculate data 1, thereby reducing instruction latency and improving the efficiency of the set communication. Based on this concept, as follows... Figure 7 As shown, another method example provided in this application may include S701 to S719.
[0224] S701, TE-D and node S1 establish connection 1;
[0225] S702, TE-D and node S2 establish connection 2;
[0226] The process of establishing a connection between TE-D and a node can be understood with reference to S501. For example, the process of TE-D establishing connection 1 with node S1 may include negotiating the connection establishment information for connection 1, and TE-D creating the RQ for connection 1. The process of TE-D establishing connection 2 with node S2 may include negotiating the connection establishment information for connection 2, and TE-D creating the RQ for connection 2. As mentioned earlier, RR1 and RR2 share the same RQ; therefore, TE-D creates the same RQ (denoted as SRQ) for connection 1 and connection 2.
[0227] When node S1 includes multiple PEs, establishing connection 1 between TE-D and node S1 can refer to establishing connection 1 between TE-D and a specific PE in node S1 (e.g., PE-S1). Similarly, when node S2 includes multiple PEs, establishing connection 2 between TE-D and node S2 can refer to establishing connection 2 between TE-D and a specific PE in node S2 (e.g., PE-S2).
[0228] This application does not limit the timing or sequence of S701 and S702.
[0229] S703, PE-D writes data 0 to addr0 in DRAM-D;
[0230] S704, PE-D submits RR1 to TE-D, RR1 includes address information describing addr0;
[0231] S705 and TE-D create RQE1 in SRQ based on RR1;
[0232] S703 to S705 can be understood by referring to S507 to S509, and will not be elaborated here.
[0233] S706, Node S2 sends message 2 to TE-D;
[0234] S706 can be understood by referring to S503~S506 and S606, which will not be elaborated here.
[0235] S707 and TE-D obtain RQE1 in SRQ;
[0236] As mentioned earlier, Connection 1 and Connection 2 can share the same SRQ. Therefore, after TE-D receives data 2 from Connection 2, it can obtain RQE1 from the SRQ. In other words, TE-D does not need to distinguish whether RQE1 is used to receive data 1 from Connection 1 or data 2 from Connection 2. That is, unlike in S618 where TE-D queries RQE2 according to the indication information in message 2, the indication information in message 2 does not need to be used to determine the receive request for receiving data from node S2 from the receive request generated by PE-D.
[0237] S708 and TE-D read data 0 from addr0 according to the instruction of RQE1;
[0238] S709 and TE-D perform addition on data 0 and data 2 to obtain the calculation result (i.e. data 2');
[0239] S710 and TE-D write data 2' to addr0;
[0240] S711, TE-D notifies PE-D of the completion of RR1;
[0241] S707~S711 can be understood by referring to S510~S514 mentioned above, and will not be repeated here.
[0242] S712, PE-D submits RR2 to TE-D, RR2 describes addr0 in DRAM-D;
[0243] After S711, PE-D determines that the result of a computation operation instructed by the instruction has been written to addr0, and can submit RR2 to TE-D. RR2 describes addr0 in DRAM-D.
[0244] S713, TE-D creates RQE2 in SRQ based on RR2;
[0245] S712 and S713 can be understood by referring to S508 and S509.
[0246] S714, Node S1 sends message 1 to TE-D;
[0247] After node S1 has prepared data 1, it can send data 1 to TE-D via message 1. S714 can be understood by referring to S503 to S506.
[0248] S715 and TE-D obtain RQE2 in SRQ;
[0249] As mentioned earlier, Connection 1 and Connection 2 can share the same SRQ. Therefore, after TE-D receives data 1 from Connection 1, it can obtain the RQE from the SRQ. Since data 1 arrives at TE-D later than data 2, the RQE obtained by TE-D for data 1 is RQE2. In other words, TE-D does not need to distinguish whether RQE2 is used to receive data 1 from Connection 1 or data 2 from Connection 2. Unlike S611, where TE-D queries RQE1 according to the indication information in message 1, the indication information in message 1 may not be used to determine the receive request for receiving data from node S1 from the receive request generated by PE-D.
[0250] S716 and TE-D read data 2' from addr0 according to the instruction of RQE2;
[0251] S717 and TE-D perform addition operations on data 2' and data 1 to obtain the calculation result (i.e. data N);
[0252] S718 and TE-D write data N to addr0;
[0253] S719, TE-D notifies PE-D of the completion of RR2.
[0254] S715 to S719 can be understood by referring to S510 to S514, and will not be elaborated here.
[0255] When computational operations satisfy the commutative law, this application proposes that for a node D in the system to receive data, other different nodes in the system correspond to the same RQ, which is referred to as a shared receive queue (SRQ). After TE-D receives an RR, regardless of which node the RR is used to receive data from, TE-D creates an RQE in the SRQ. Furthermore, upon receiving a message, regardless of which node the message originates from, TE-D queries the RQE in the SRQ. This allows TE-D to immediately execute computational operations upon receiving a message from any node, thereby reducing the latency of TE-D in completing computational operations and improving the execution efficiency of aggregated communication.
[0256] The above examples illustrate that SQ and RQ are merely one possible way for the TE to store and query requests generated by the PE. In practical applications, the TE can use other methods to store requests generated by the PE and retrieve the corresponding requests after receiving a message. Figure 5-1 For example, S504 and S505 are optional methods and can be replaced with other steps to achieve the saving and consumption of SR1. Similarly, S509 and S510 are optional methods and can be replaced with other steps to achieve the saving and consumption of RR1.
[0257] The above describes the communication method and system provided in this application using a set communication scenario as an example. This application does not limit the application scenario of this solution, and correspondingly, this application does not limit the computational operation to be performed by node D to be a computational operation in set communication.
[0258] The previous example used addition as an example of the operation between the local data of node D and the received data. In practical applications, addition can be replaced with other types of operations. Similarly, the previous example used the operation performed on data 1 and data 2 as the same type of operation. In practical applications, the operation performed on each data 1 can be of different types.
[0259] The previous example used PE-D to write data 0 to DRAM-D and PE-S1 to write data 1 to DRAM-S1. In practical applications, data can be written to DRAM by other devices (such as TE).
[0260] It should be understood that, in the various method examples of this application, the sequence number of each step does not imply the order of execution; the execution order of each step should be determined by its function and internal logic. It should also be understood that such terminology can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of the scheme of this application.
[0261] In the examples of this application, the various numerical designations are only for the convenience of description and are not intended to limit the scope of the examples of this application.
[0262] In the system provided in this application, the deployment method of each unit in the node is not limited. For example, all or some units in the same node can be integrated into the same chip. Figure 5-1 or Figure 5-2 or Figure 6 or Figure 7 As shown, the PE and TE on the same node can be integrated into the same processor, or they can be physically separate. For example, the PE is a processor or integrated into the processor, while the TE is located externally to the processor, such as in a separate network interface card (NIC). Figure 4 In the system shown, the PE, HA, IC and MC in the same node can be integrated into the same processor, or the PE, HA and MC in the same node can be integrated into the same processor, while the IC is a separate network card or is set in a separate network card.
[0263] This application does not limit the implementation method of TE performing the above-described corresponding method steps. The above-described corresponding method steps may include, for example, the following: Figure 5-1 or Figure 5-2 or Figure 6 or Figure 7 The method steps performed by TE-D.
[0264] Optionally, the TE can execute the corresponding method steps described above via software. For example, the TE may include a processor and a memory, the memory storing instructions, and the processor executing the corresponding method steps by executing the instructions.
[0265] Optionally, the TE can execute the corresponding method steps described above via hardware. For example, the TE may include one or more circuits for executing the corresponding method steps described above. These one or more circuits may be integrated into the same chip or different chips.
[0266] Optionally, the TE can execute the corresponding method steps described above using a combination of hardware and software. For example, the TE may include a processor, a memory, and one or more circuits. The processor executes a portion of the corresponding method steps by executing instructions stored in the memory, and the one or more circuits are used to execute the remaining steps of the corresponding method steps.
[0267] In other words, those skilled in the art, combining the units and algorithm steps of the examples disclosed herein, can implement these functions using electronic hardware, computer software, or a combination of both. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0268] The method provided in this application has been described above; the apparatus provided in this application will be described below. Figure 8 The diagram illustrates a possible structure of a communication device. For example... Figure 8 As shown, the communication device may include a communication module and a processing module.
[0269] The communication module is used to receive messages sent by a first node, the messages encapsulating first data to be passed to a first processing unit among one or more processing units. The processing module is used to obtain a first receive request generated by the first processing unit, the first receive request including first address information describing a first storage location. The processing module is further used to read second data from the first storage location according to the first address information, and perform calculation operations on the first data and the second data to obtain a calculation result.
[0270] Optionally, the message further encapsulates first indication information, which is used to determine the receiving request generated by the first processing unit from the receiving requests generated by the one or more processing units; the transmission unit is specifically used to obtain the first receiving request according to the first indication information in the message.
[0271] Optionally, the communication system further includes other nodes besides the first node and the second node. The receiving request generated by the first processing unit includes a receiving request for receiving data from the first node and a receiving request for receiving data from the other nodes. The first indication information is also used to determine the receiving request for receiving data from the first node from the receiving request generated by the first processing unit.
[0272] Optionally, the first data is a portion of the target data transmitted from the first node to the first processing unit. The receiving request generated by the first processing unit for receiving data from the first node includes a receiving request for receiving the first data and a receiving request for receiving other data in the target data besides the first data. The first indication information is also used to determine the receiving request for receiving the first data from the receiving request for receiving data from the first node generated by the first processing unit.
[0273] Optionally, the processing module is further configured to notify the processing unit of the completion of the calculation operation after obtaining the calculation result.
[0274] The first receiving request further includes second indication information, which is used to instruct the calculation operation to be performed on the second data and the first data.
[0275] Optionally, the first receiving request may further include third indication information, which indicates the type of the computational operation.
[0276] Optionally, the processing module is further configured to preprocess the first data and / or the second data before performing the calculation operation on the first data and the second data.
[0277] Optionally, the first receiving request further includes second address information describing the second storage location, the second storage location being used to store the calculation result; the processing module is further configured to write the calculation result into the second storage location according to the second address information.
[0278] Optionally, the second transmission unit is further configured to write the calculation result into the first storage location according to the first address information.
[0279] Optionally, the message is a Remote Memory Direct Access Protocol (RDMA) message or a Transmission Control Protocol (TCP) message.
[0280] Optionally, the computation operation is a set communication computation.
[0281] The communication device can be TE-D or Node D as described above. The message can be message 1 or message 2. The first processing unit can be, for example, PE-D, and the first node can be, for example, node S1, node S2, or PE-S1 within node S1. The first data can be data 1, and the second data can be data 0 or data 2'. Alternatively, the first data can be data 2, and the second data can be data 0 or data 1'. The first receive request can be, for example, RR1 (or RQE1) or RR2 (or RQE2), the first storage location can be addr0, the second storage location can be addr2, and the calculation operation can include any one or more operations described above. The first indication information can be the indication information in the message described above, the second indication information can be indication information 1 described above, and the third indication information can be indication information 2 described above.
[0282] For example, the communication module can be used to execute S506, and the processing module can be used to execute all or some of the steps in S508 to S514. Optionally, the communication module can also be used to execute all or some of the steps in S516 and S517, and the processing module can also be used to execute S515.
[0283] For example, the communication module can be used to execute all or some of the steps in S606, S608 and S609, and the processing module can be used to execute all or some of the steps in S604 to S605, S607, S611 to S622.
[0284] For example, the communication module can be used to execute S706 and / or S714, and the processing module can be used to execute all or part of the steps in S704-S705, S707-S713 and S715-S719.
[0285] As mentioned above, Figure 8 The modules shown can be virtual functional modules generated by the processor by executing instructions in memory, or... Figure 8 The modules shown can be hardware circuits, or... Figure 8 Some of the modules shown are virtual functional modules generated by the memory through the execution of instructions in the memory, while the remaining modules are hardware circuits.
[0286] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatus, and methods can be implemented in other ways. For example, Figure 8 The division of modules in the described device is merely illustrative and can be understood as a logical functional division. In actual implementation, there may be other division methods. For example, a module can be split into multiple units, multiple modules can be integrated together or integrated into another system, or some features can be ignored or not executed.
[0287] When various aspects of the embodiments of this application, or possible implementations of various aspects, are implemented using software, all or part of the aforementioned aspects or possible implementations may be implemented in the form of a computer program product. A computer program product refers to instructions (or computer-readable instructions, computer program instructions, functional programs, or program code) stored in a computer-readable medium. When these instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.
[0288] Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM).
[0289] The terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a list of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to those processes, methods, products, or apparatuses. The term “a plurality of” as used in embodiments of this application refers to two or more.
[0290] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A second node, characterized in that, The second node includes a transmission unit and one or more processing units, wherein the one or more processing units include a first processing unit; The transmission unit is configured to receive a message sent by the first node and obtain a first receiving request generated by the first processing unit. The first receiving request includes first address information describing a first storage location, and the message includes first data. The transmission unit is further configured to read the second data in the first storage location according to the first address information, and perform calculation operations on the first data and the second data to obtain the calculation result.
2. The second node according to claim 1, characterized in that, The message also encapsulates first indication information, which is used to determine the receiving request generated by the first processing unit from the receiving requests generated by the one or more processing units. The transmission unit is specifically used to obtain the first receiving request according to the first indication information in the message.
3. The second node according to claim 2, characterized in that, The receiving requests generated by the first processing unit include receiving requests for receiving data from the first node and receiving requests for receiving data from other nodes. The first indication information is also used to determine the receiving request for receiving data from the first node from the receiving requests generated by the first processing unit, wherein the other nodes are nodes other than the first node and the second node.
4. The second node according to claim 3, characterized in that, The first data is a portion of the target data transmitted from the first node to the first processing unit. The receiving request generated by the first processing unit for receiving data from the first node includes a receiving request for receiving the first data and a receiving request for receiving other data in the target data besides the first data. The first indication information is also used to determine the receiving request for receiving the first data from the receiving request for receiving data from the first node generated by the first processing unit.
5. The second node according to any one of claims 1-4, characterized in that, The first receiving request further includes second indication information, which is used to instruct the calculation operation to be performed on the second data and the first data.
6. The second node according to any one of claims 1-5, characterized in that, The first receiving request also includes third indication information, which is used to indicate the type of the computational operation.
7. The second node according to any one of claims 1-6, characterized in that, The transmission unit is further configured to preprocess the first data and / or the second data before performing the calculation operation on the first data and the second data.
8. The second node according to any one of claims 1-7, characterized in that, The first receiving request also includes second address information describing the second storage location, which is used to store the calculation result; The transmission unit is further configured to write the calculation result into the second storage location according to the second address information.
9. The second node according to any one of claims 1-8, characterized in that, The transmission unit is further configured to write the calculation result into the first storage location according to the first address information.
10. The second node according to any one of claims 1-9, characterized in that, The message is either a Remote Memory Direct Access Protocol (RDMA) message or a Transmission Control Protocol (TCP) message.
11. The second node according to any one of claims 1-10, characterized in that, The computational operation is a set communication computation.
12. A communication method, characterized in that, The communication method is applied to a second node. The second node includes a transmission unit and one or more processing units, wherein the one or more processing units include a first processing unit. The method includes: The transmission unit receives the message sent by the first node and obtains the first receiving request generated by the first processing unit. The first receiving request includes first address information describing the first storage location, and the message contains first data. The transmission unit reads the second data from the first storage location according to the first address information, and performs calculation operations on the first data and the second data to obtain the calculation result.
13. A communication method, characterized in that, The method includes: Receive a message sent by a first node, the message encapsulating first data to be passed to a first processing unit in one or more processing units; Obtain a first receiving request generated by the first processing unit, wherein the first receiving request includes first address information describing a first storage location; The second data in the first storage location is read according to the first address information, and a calculation operation is performed on the first data and the second data to obtain the calculation result.
14. The method according to claim 13, characterized in that, The message further encapsulates first indication information, which is used to determine the receiving request generated by the first processing unit from the receiving requests generated by the one or more processing units. Obtaining the first receiving request generated by the first processing unit includes: The first receiving request is obtained according to the first indication information in the message.
15. The method according to claim 14, characterized in that, The receiving requests generated by the first processing unit include receiving requests for receiving data from the first node and receiving requests for receiving data from other nodes besides the first node and the nodes where the one or more processing units are located. The first indication information is also used to determine the receiving request for receiving data from the first node from the receiving requests generated by the first processing unit.
16. The method according to claim 15, characterized in that, The first data is a portion of the target data transmitted from the first node to the first processing unit. The receiving request generated by the first processing unit for receiving data from the first node includes a receiving request for receiving the first data and a receiving request for receiving other data in the target data besides the first data. The first indication information is also used to determine the receiving request for receiving the first data from the receiving request for receiving data from the first node generated by the first processing unit.
17. The method according to any one of claims 13-16, characterized in that, The first receiving request further includes second indication information, which is used to instruct the calculation operation to be performed on the second data and the first data.
18. The method according to any one of claims 13-17, characterized in that, The message is either a Remote Memory Direct Access Protocol (RDMA) message or a Transmission Control Protocol (TCP) message.
19. The method according to any one of claims 13-18, characterized in that, The computational operation is a set communication computation.
20. The method according to any one of claims 13-19, characterized in that, The one or more processing units are deployed on a second node other than the first node, and the method is executed by the second node or by a transmission unit in the second node.
21. A communication device, characterized in that, The communication device includes: A communication module is used to receive messages sent by a first node, wherein the messages encapsulate first data to be passed to a first processing unit in one or more processing units; The processing module is configured to acquire a first receiving request generated by the first processing unit, wherein the first receiving request includes first address information describing a first storage location; The processing module is further configured to read the second data in the first storage location according to the first address information, and perform calculation operations on the first data and the second data to obtain the calculation result.
22. The communication device according to claim 21, characterized in that, The one or more processing units are deployed on a second node other than the first node, and the communication device is the second node or is deployed on the second node.
23. The communication device according to claim 22, characterized in that, The communication device is a chip, a network interface controller, or a host device.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code that, when executed by a processor in a computer device, implements the method as described in any one of claims 13-20.
25. A computer program product, characterized in that, When the program code contained in the computer program product is executed by a processor in a computer device, it implements the method as described in any one of claims 13-20.
Citation Information
Patent Citations
Data operation method and device
CN115989478A
Communication system, communication method and related device
CN120687405A