Switch, method and system for implementing collective communication based on a multicast replication engine
By introducing a multicast replication engine (SMRE) into the switch, the problems of redundant transmission and latency under the PCIe architecture are solved, efficient collective communication is achieved, and the training performance of multi-chip systems is improved.
Patent Information
- Application Number
- CN202511113712.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-11
AI Technical Summary
In existing PCIe architectures, point-to-point transmission leads to redundant data transmission and communication delays in multi-chip systems, becoming a performance bottleneck in multi-chip parallel training, especially inefficient during broadcast operations.
The Multicast Replication Engine (SMRE) is used to implement collective communication in the switch. The broadcast request is determined by the group matcher and the broadcast module, and the data is broadcast directly to multiple ports, avoiding repeated transmission of TLP packets.
It improves the efficiency of broadcast operations, reduces data transmission latency, reduces redundant transmission, and enhances communication performance.
Smart Images

Figure CN120614327B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer and communication technology, and in particular to a switch, method and system for realizing collective communication based on a multicast replication engine. Background Technology
[0002] With the widespread application of large-scale artificial intelligence models (such as GPT, BERT, and ViT), the number of parameters in a single neural network has grown to tens of billions or even trillions. To support the training of such massive models, the industry generally adopts a multi-chip parallel training strategy, distributing model parameters and data across multiple computing units (such as ASICs, GPUs, and TPUs) for collaborative processing. Common parallel training methods include data parallelism, model parallelism, and hybrid parallelism.
[0003] In this type of training process, the performance of communication operations directly determines the overall training speed. The most critical and frequent communication modes include:
[0004] Broadcast: Used to synchronize the latest parameters on the master node (or a master parameter server) to all compute nodes;
[0005] Gather: Used to collect the results (such as gradients) calculated from multiple child nodes.
[0006] All-Reduce: Used for multiple nodes to collaboratively complete reduction operations (such as parameter summation or averaging).
[0007] These communication operations often occupy 30%-60% of the training cycle, especially in multi-chip systems, where insufficient communication link efficiency can become a performance bottleneck. In traditional PCIe architectures, data transmission is typically performed in a point-to-point manner.
[0008] Under this architecture, such as Figure 1 The process of broadcasting in the traditional PCIe architecture shown below requires the following steps when GPU 0 wants to send data to GPU 1, GPU 2, and GPU 3.
[0009] Single transmission of data: GPU 0 produces and initiates a separate data transmission to each target GPU, i.e., GPU 0 initiates independent PCIe Memory Write operations to GPU 1, GPU 2, and GPU 3, respectively. A PCIe Transaction Layer Packet (TLP) needs to be constructed for each GPU data transmission, and each target GPU needs to receive the data once;
[0010] PCIe switch relay: Data is transmitted from GPU 0 to a PCIe switch, which then forwards the data to the target GPUs. Each data transmission passes through the arbitration and routing of the switch, which can increase latency.
[0011] As can be seen, based on the existing point-to-point transmission of the PCIe architecture, at least the following deficiencies exist in the broadcast operation process:
[0012] (1) Redundant transmission: GPU 0 needs to generate and send a separate transaction for each target GPU, which means that if GPU 0 wants to send the same data to multiple GPUs, it needs to send duplicate data packets for each GPU, and each GPU also receives the same data, resulting in redundant data transmission;
[0013] (2) Bandwidth and latency issues: As the number of target GPUs increases, the redundancy and bandwidth consumption of data transmission will significantly increase, especially when the number of target nodes is large. This approach can cause excessive occupation of PCIe bus bandwidth resources and cause cumulative communication latency, which becomes a bottleneck for training performance.
[0014] In summary, current multi-ASIC computing systems are mostly based on PCIe bus topology, and multiple computing chips are interconnected through a PCIe switch (PCIe Switch). However, the PCIe protocol is designed for point-to-point transmission and lacks native support for collective communication, resulting in serious performance obstacles when performing high-frequency communication operations such as broadcasting.
[0015] The disclosure of the above background art is only used to assist in understanding the inventive concept and technical solutions of the present application, and does not necessarily belong to the prior art of the present application, nor does it necessarily provide technical teaching. In the absence of explicit evidence that the above content was disclosed before the filing date of the present application, the above background art should not be used to evaluate the novelty and inventiveness of the present application. SUMMARY
[0016] The purpose of the present application is to provide a switch, method and system based on a multicast replication engine to implement collective communication, which can improve the efficiency of broadcast operations between multiple devices and reduce the latency of data transmission.
[0017] To achieve the above object, the technical scheme adopted by the present application is as follows:
[0018] A switch for realizing collective communication based on a multicast replication engine, comprising a plurality of ports and a multicast replication engine electrically connected with each of the ports respectively, the multicast replication engine comprising a group matcher and a broadcast module;
[0019] The multicast replication engine is configured to determine whether a communication request transmitted by a request port is a broadcast request, the request port being one of the ports, and if so, the multicast replication engine sends the communication request to the group matcher and receives broadcast data targeted by the request port for transmission;
[0020] The group matcher, in response to receiving the broadcast request, determines a broadcast port and a corresponding broadcast address according to a preset broadcast group table and transmits them to the broadcast module, the broadcast group table comprising memory addresses configured for each of the ports, the broadcast port being a port other than the request port, and the broadcast address being one or more of the memory addresses;
[0021] The broadcast module, in response to receiving the broadcast port and the broadcast address, controls each of the broadcast ports to receive the broadcast data through the broadcast address.
[0022] Further, any of the above technical solutions or a combination of multiple technical solutions, the switch is provided with a shared memory, the shared memory comprising a plurality of memory addresses, each of the ports being configured with a corresponding memory address, and each of the ports directly communicating with its corresponding memory address;
[0023] The number of broadcast addresses is multiple, and the memory address corresponding to each of the broadcast ports is configured as the broadcast address;
[0024] The broadcast module controls each of the broadcast ports to receive the broadcast data through the broadcast address, comprising: the broadcast module replicates the broadcast data into multiple copies corresponding to the number of broadcast ports and transmits them to each of the broadcast addresses.
[0025] Further, any of the above technical solutions or a combination of multiple technical solutions, each of the ports is configured with a port ID, the broadcast group table comprises port IDs and memory addresses having a corresponding relationship, and the broadcast port and the corresponding broadcast address are determined according to the broadcast group table in the following manner:
[0026] The broadcast request includes a request port ID, the request port ID is a port ID corresponding to the request port, a port ID in the broadcast group table different from the request port ID is determined as a target ID, and a port corresponding to the target ID is the broadcast port;
[0027] The target ID corresponding memory address is determined as the broadcast address.
[0028] Further, any of the above technical solutions or combinations thereof, the switch is provided with a shared memory, the shared memory includes a plurality of memory addresses, each port is configured with a corresponding memory address, and the port is configured to directly communicate with any memory address;
[0029] The number of broadcast addresses is one, and the memory address corresponding to each request port is configured as the broadcast address.
[0030] The broadcast module controls each broadcast port to receive the broadcast data through the broadcast address, including: the broadcast module stores the broadcast data to the broadcast address and notifies each broadcast port to obtain the broadcast data in the broadcast address.
[0031] Further, any of the above technical solutions or combinations thereof, each port is configured with a port ID, the broadcast group table includes port IDs and memory addresses having a corresponding relationship, and the broadcast port and the corresponding broadcast address are determined according to the broadcast group table in the following manner:
[0032] The broadcast request includes a request port ID, the request port ID is a port ID corresponding to the request port, a port ID in the broadcast group table different from the request port ID is determined as a target ID, and a port corresponding to the target ID is the broadcast port;
[0033] The request port ID corresponding memory address is determined as the broadcast address.
[0034] Further, any of the above technical solutions or combinations thereof, each port is configured to connect one device to be communicated, the memory address is a DMA address, and the DMA address is configured to directly communicate with a physical memory address of the device; and / or,
[0035] If the multicast replication engine determines that the communication request is a non-broadcast request, the switch controls the request port to communicate with a target port based on the communication request.
[0036] Further, any one of the above technical solutions or a combination of the above technical solutions, the communication request comprises a broadcast item, the broadcast item comprises a broadcast valid bit, when the broadcast valid bit is a target field, the communication request is the broadcast request, when the broadcast valid bit is a non-target field, the communication request is a non-broadcast request.
[0037] The multicast replication engine is configured to determine whether the broadcast valid bit in the communication request is the target field, if yes, the multicast replication engine determines that the communication request is the broadcast request, otherwise, the multicast replication engine determines that the communication request is a non-broadcast request.
[0038] Further, any one of the above technical solutions or a combination of the above technical solutions, the group matcher, in response to receiving the broadcast request, first determines whether the broadcast valid bit is a target field, if yes, determines a broadcast port and a broadcast address according to a preset broadcast group table, otherwise, sends an error message.
[0039] Further, any one of the above technical solutions or a combination of the above technical solutions, the broadcast item further comprises a target port ID, if the broadcast valid bit is a target field, the group matcher determines that a port corresponding to the target port ID is the broadcast port.
[0040] Further, any one of the above technical solutions or a combination of the above technical solutions, the broadcast item further comprises a broadcast start physical address and a broadcast data length, the broadcast data is data with the start physical address as a start address and with the broadcast data length as a data length.
[0041] Further, any one of the above technical solutions or a combination of the above technical solutions, the switch is configured as a downstream switch of a plurality of devices to be communicated; or,
[0042] The switch is configured as an upstream switch of a plurality of devices to be communicated.
[0043] According to another aspect of the present application, a multicast replication engine-based collective communication method is provided, based on the switch for realizing collective communication based on the multicast replication engine according to any one of the above technical solutions or a combination of the above technical solutions, comprising the following steps:
[0044] The multicast replication engine is configured to determine whether the communication request transmitted by the request port is a broadcast request, the request port being one of the ports, if yes, the multicast replication engine sends the communication request to the group matcher and receives broadcast data targeted by the request port;
[0045] The group matcher determines the broadcast port and the corresponding broadcast address according to a preset broadcast group table in response to receiving the broadcast request, the broadcast group table comprises memory addresses configured for each of the ports, the broadcast port is a port other than the request port, and the broadcast address is one or more of the memory addresses.
[0046] The broadcast module controls each of the broadcast ports to receive the broadcast data through the broadcast address in response to receiving the broadcast port and the broadcast address.
[0047] According to another aspect of the present application, a communication system is provided, comprising the switch for realizing collective communication based on the multicast replication engine according to any one or a combination of the technical solutions.
[0048] The technical solution provided by the present application has the following beneficial effects:
[0049] a. The multicast replication engine is arranged in the switch for multicast transmission, the multicast replication engine comprises a group matcher and a broadcast module, when the multicast replication engine monitors that the communication request transmitted by the request port is a broadcast request, the multicast replication engine is started, the group matcher in the multicast replication engine determines the broadcast port and the broadcast address, and the broadcast module broadcasts the broadcast data to each of the broadcast ports, so that the request port does not need to copy multiple TLP data packets and transmit each of the TLP data to each of the ports, thereby effectively improving the efficiency of the broadcast operation and reducing the latency of data transmission.
[0050] b. After the multicast replication engine (SMRE) is activated by the switch when the communication request is monitored to be a broadcast request, the group matcher further checks the broadcast signal again, thereby reducing the misjudgment rate of the broadcast operation and improving the accuracy and reliability of data transmission. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0052] Figure 1 The principle diagram of the existing PCIe architecture in the broadcast operation;
[0053] Figure 2 The principle diagram of the switch provided for an exemplary embodiment of the present application in the broadcast operation;
[0054] Figure 3 A schematic diagram of a first broadcast operation provided for an exemplary embodiment of the present application;
[0055] Figure 4 A schematic diagram of a second broadcast operation provided for an exemplary embodiment of the present application;
[0056] Figure 5 A schematic diagram of a PCIe architecture of a switch as an uplink switch provided for an exemplary embodiment of the present application;
[0057] Figure 6 A schematic diagram of a PCIe architecture of a switch as a downlink switch provided for an exemplary embodiment of the present application;
[0058] Figure 7 A workflow diagram of a switch based on a multicast replication engine to implement collective communication provided for an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the personnel in the art better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without making creative efforts should belong to the scope of protection of the present application.
[0060] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product, or apparatus that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatuses.
[0061] In an embodiment of the present application, a switch based on a multicast replication engine to implement collective communication is provided, as shown in Figures 2 to 4 and Figure 7 which comprises a plurality of ports and a multicast replication engine (SMRE) electrically connected to each of the ports, the multicast replication engine comprising a group matcher and a broadcast module;
[0062] The multicast replication engine is configured to determine whether the communication request requesting port transmission is a broadcast request, the request port being one of the ports, and if so, the multicast replication engine sends the communication request to the group matcher and receives broadcast data of the request port target transmission;
[0063] The group matcher, in response to receiving the broadcast request, determines a broadcast port and a corresponding broadcast address according to a preset broadcast group table and transmits to the broadcast module, the broadcast group table including memory addresses configured for each of the ports, the broadcast port being a port other than the request port, and the broadcast address being one or more of the memory addresses;
[0064] The broadcast module, in response to receiving the broadcast port and the broadcast address, controls each of the broadcast ports to receive the broadcast data through the broadcast address.
[0065] In the embodiment, the multicast replication engine is configured in the switch for multicast transmission.
[0066] In the embodiment, the communication request is configured with a broadcast entry to determine a broadcast range and a target port and to activate a multicast replication engine (SMRE) to perform multicast operation. As shown in Table 1, the broadcast entry includes a target port ID, a broadcast valid bit, a broadcast start physical address, and a broadcast data length. The broadcast data is data with the start physical address as the start address and the broadcast data length as the data length.
[0067] Table 1 Broadcast entry
[0068]
[0069] As shown in Table 1, the broadcast valid bit is a target field, i.e. VALID = 1, and the communication request is the broadcast request. When the broadcast valid bit is a non-target field, the communication request is a non-broadcast request. In the embodiment, the multicast replication engine is configured to determine whether the broadcast valid bit in the communication request is the target field, i.e. VALID = 1, and if so, the multicast replication engine determines that the communication request is the broadcast request.
[0070] Suppose in a round of model training, a 2GB parameter weight data of a master node (such as GPU0) needs to be broadcast to other three ports (such as GPU1, GPU2, and GPU3), then:
[0071] Pre-set broadcast data address: select a physical memory address segment as a broadcast entry, for example:
[0072] BROADCAST_ADDR_BASE = 0xC000_0000
[0073] BROADCAST_LEN = 0x8000_0000 (2GB)
[0074] Map this address segment to the master node PCIe BAR, which is convenient for DMA initiation;
[0075] Each receiving port is uniformly configured with a target buffer for carrying reception.
[0076] The multicast replication engine determines that the communication request is the broadcast request to enable the broadcast operation, specifically, all TLP data packets satisfying Attr[1] == 1 are sent to the group matcher for broadcast group table matching judgment to determine the broadcast port and the broadcast address, and then the broadcast module starts the copy generation process according to the broadcast port and the broadcast address to execute the broadcast operation.
[0077] Preferably, the group matcher, in response to receiving the broadcast request, first confirms whether the broadcast valid bit is the target field, if yes, determines the broadcast port and the broadcast address according to the pre-set broadcast group table, otherwise, issues an error message. That is, after the switch detects that the communication request is a broadcast request and activates the multicast replication engine (SMRE), the group matcher will further check the broadcast signal to reduce the misjudgment rate of the broadcast operation and improve the accuracy and reliability of data transmission.
[0078] If VALID = 0, it is determined that the communication request is a non-broadcast request, and the switch controls the request port and the target port to communicate based on the communication request. That is, when the switch confirms that the communication request is a non-broadcast request, the multicast replication engine is not enabled, and the TLP data packet initiated by the request port is directly sent to the target port.
[0079] In an embodiment of the present application, a shared storage space is configured in the switch, the shared storage space is divided into a plurality of memory units, each memory unit is configured with a corresponding memory address, that is, one memory address corresponds to one memory unit. Preferably, the number of memory units is not less than the number of ports, that is, each port is configured with at least one corresponding memory address, for example, each port is configured with one memory address, and the memory address is configured to communicate with the physical memory address of the device.
[0080] In an embodiment of the present application, as Figure 3As shown, the memory address is a DMA address, and each port is configured with a corresponding DMA controller configured to control communication between the DMA address corresponding to the same port and the physical memory address of the device connected thereto. For example, port 1 is connected to device 1 and is configured with DMA controller 1 and DMA address 1, and the DMA controller 1 transmits the DMA address 1 to the physical memory address of the device 1, or the DMA controller 1 transmits the physical memory address of the device 1 to the DMA address 1.
[0081] In the embodiment, each port directly communicates with the memory address corresponding thereto, and the port does not communicate with the memory address corresponding to other ports. The advantage of such data transmission is that different ports cannot directly communicate with each other, and the security of data transmission is higher. Each port is configured with a port ID, and the broadcast group table includes port IDs and the memory addresses having a corresponding relationship. The broadcast port and the corresponding broadcast address are determined according to the broadcast group table in the following manner: the broadcast request includes a request port ID, the request port ID is the port ID corresponding to the request port, the port ID different from / different from the request port ID in the broadcast group table is determined as a target ID, the port corresponding to the target ID is the broadcast port, and the memory address corresponding to the target ID is determined as the broadcast address.
[0082] As shown in Figure 3 The switch monitors that the communication request transmitted by the port 1 is a broadcast request, and then starts the multicast replication engine. The DMA controller 1 transmits the broadcast data to the DMA address 1, the broadcast module replicates the data in the DMA address 1 to other DMA addresses except the DMA address 1, and notifies each broadcast port to obtain the data in the DMA address corresponding thereto. In this way, for the broadcast operation, the port 1 does not need to replicate multiple TLP data packets and transmit them to each port, thereby effectively improving the efficiency of the broadcast operation and reducing the latency of data transmission.
[0083] In one specific embodiment of the application, the master node (such as GPU0, ASIC0 or Host CPU) broadcasts a certain segment of parameters (such as 2GB of weights) to multiple slave nodes in a training iteration, and only needs to be written once. The data is replicated in the multicast replication engine through the switch, and finally reaches multiple target ports at the same time. The specific steps are as follows.
[0084] Step 1: Switch initialization configuration.
[0085] Objective: Configure the broadcast group table (BG Table) entry, define the broadcast range and target port, and activate the multicast replication engine to realize the multicast replication function.
[0086] Preparation: broadcast task parameter preset.
[0087] Suppose in this round of model training, a 2GB parameter weight data of the master node (such as ASIC0) needs to be broadcast to the other three receiving ends (such as ASIC1, ASIC2, and ASIC3), then:
[0088] Preset broadcast data address: select a physical memory address segment as the broadcast entry, for example:
[0089] BROADCAST_ADDR_BASE=0xC000_0000;
[0090] BROADCAST_LEN=0x8000_0000 (2GB).
[0091] Map this address segment to the master node PCIe BAR for easy DMA initiation; each receiving end uniformly configures the target buffer for receiving.
[0092] Step two: the master node triggers broadcast write. First, the master node GPU0 prepares to write broadcast data, which includes:
[0093] Data volume: for example, 2GB model parameters;
[0094] Storage address: CPU, ASIC0, GPU0 internal memory;
[0095] Target address: can be any PCIe address (does not have to match the BASE_ADDR in the BG table);
[0096] TLP type: Memory Write (PCIe standard transaction);
[0097] TLP special flag: set Attr[1] to 1 in Header as broadcast identification. Header is a standard PCIe header for user extensible options, and the present application is easy to implement by configuring the broadcast effective bit as Attr=1 in Header.
[0098] Then, TLP data packet is sent, and the write operation triggers DMA or CPU PCIe packet sending;
[0099] Construct a complete PCIe TLP data packet:
[0100] Type field: memory write;
[0101] Address field: write target address;
[0102] Length field: 2GB;
[0103] Broadcast Valid Bit Attr field: Attr[1]=1;
[0104] TLP is sent to the switch up / down port.
[0105] Then, the TLP packet arrives at the switch, and the TLP enters the switch entry port, i.e., port 1.
[0106] The TLP packet is sent to the standard forwarding path before the multicast replication engine.
[0107] The multicast replication engine identifies Attr[1] in the TLP Header, and the identification and master node trigger broadcast write execution process as follows.
[0108] Step three: the master node triggers broadcast write.
[0109] Trigger entry: TLP arrives at the switch; TLP (Memory Write, with Attr[1]=1) sent by the master node (such as ASIC0 or Host) arrives at the switch through the PCIe link; the TLP is received at the entry of the switch, and enters the forwarding control module.
[0110] Enter the multicast replication engine and trigger the identification logic: the multicast replication engine captures the TLP packet with the broadcast valid bit as the target field; the SMRE reads the Attr field in the TLP Header to determine whether it is a broadcast request / broadcast packet; if Attr[1]=1, the communication request is a broadcast request, and the SMRE starts the group matcher and the broadcast module to perform broadcast operation; otherwise, the TLP is forwarded according to the normal path and does not participate in replication. It should be noted that in the present embodiment, the SMRE is used as a monitoring module, and only the TLP packet with Attr[1]=1 is intercepted to enter the SMRE; in other embodiments, all TLP packets entering the switch can flow through the multicast replication engine, and only the TLP packet with Attr[1]=1 can trigger the group matcher to identify it.
[0111] Query the broadcast group table (BG Table): the SMRE starts the matching logic to traverse all entries in the BG table for the group matcher, and only considers the entries with VALID=1; since the receiving end no longer depends on address matching, the matching logic only needs to satisfy Attr[1]=1 and the entry is valid, which can quickly identify the broadcast request; select the matching entry: by default, select the first valid entry that hits (which can be extended to support multiple entries); read the broadcast port ID to determine the target port / broadcast port: get the target port mask, for example, 0b00000111 indicates replication to Port0 / 1 / 2; the target port is the control basis for subsequent copy generation.
[0112] Port ID (PORT_MASK) is input to the broadcast module: the group matcher sends the selected broadcast port ID and the original TLP data packet to the broadcast module; the broadcast module performs pipeline processing as shown in Table 2, traverses the mask bit by bit, generates multiple TLP copies, each TLP copy includes an independent Header (including target port identification), shares the original data Payload pointer, and recalculates the CRC check (optional, or can not include the recalculated CRC check); the broadcast module outputs each TLP copy to the corresponding port FIFO, ready to be sent to the receiving node.
[0113] Table 2: Broadcast module performs pipeline processing phase
[0114]
[0115] In another embodiment of the present application, the difference from the above-mentioned embodiment is that the port is configured to directly communicate with any of the memory addresses. In this embodiment, each of the ports is configured with a port ID, and the broadcast group table includes the port ID and the memory address with a corresponding relationship. The broadcast port and the corresponding broadcast address are determined according to the broadcast group table in the following manner: the broadcast request includes a request port ID, the request port ID is the port ID corresponding to the request port, the port ID in the broadcast group table that is different from the request port ID is determined as a target ID, and the port corresponding to the target ID is the broadcast port. Alternatively, the broadcast port and the broadcast address are directly determined by the broadcast entry.
[0116] In this embodiment, as shown in Figure 4 , the switch monitors that the communication request initiated by port 1 is a broadcast request, and then starts the multicast replication engine. The DMA controller 1 transmits the broadcast data to the DMA address 1, and the broadcast module notifies each of the broadcast ports of the broadcast notification of receiving the broadcast data. The DMA controller corresponding to the broadcast port transmits the data in the DMA address 1 to the physical memory address of the device connected to the port in response to receiving the broadcast notification.
[0117] In this embodiment, for the broadcast operation, not only does port 1 not need to replicate multiple TLP data packets and transmit them to each port, but the broadcast port also does not need to replicate the broadcast data into multiple copies, which can further improve the efficiency of the broadcast operation and reduce the latency of data transmission.
[0118] For any of the above embodiments, the switch can be configured as an upper switch of multiple devices to be communicated in a PCIe architecture as shown in Figure 5 .
[0119] More preferably, the switch is configured as an upper switch of multiple devices to be communicated in a PCIe architecture as shown in Figure 6 .Downstream switches of a plurality of devices to be communicated in the PCIe architecture shown.
[0120] In one embodiment of the present application, a collective communication method based on a multicast replication engine is provided, a switch for realizing collective communication based on the multicast replication engine according to any one of the above embodiments or a combination of multiple embodiments, the collective communication method based on the multicast replication engine comprising the following steps:
[0121] The multicast replication engine is used to determine whether a communication request transmitted by a request port is a broadcast request, the request port being one of the ports, if yes, the multicast replication engine sends the communication request to the group matcher and receives broadcast data targeted by the request port;
[0122] The group matcher, in response to receiving the broadcast request, determines a broadcast port and a corresponding broadcast address according to a preset broadcast group table and transmits to the broadcast module, the broadcast group table comprising memory addresses configured for each of the ports, the broadcast port being a port other than the request port, and the broadcast address being one or more of the memory addresses;
[0123] The broadcast module, in response to receiving the broadcast port and the broadcast address, controls each of the broadcast ports to receive the broadcast data through the broadcast address.
[0124] In one embodiment of the present application, a communication system is also provided, the communication system comprising the switch for realizing collective communication based on the multicast replication engine according to any one of the above embodiments or a combination of multiple embodiments.
[0125] It should be noted that the collective communication method based on the multicast replication engine and the communication system embodiments provided by the present application have the same inventive concept as the above-mentioned switch for realizing collective communication based on the multicast replication engine, and the entire content of the switch for realizing collective communication based on the multicast replication engine is incorporated into the collective communication method based on the multicast replication engine and the communication system by way of introduction.
[0126] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The above description is only a specific embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A switch implementing collective communication based on a multicast replication engine, characterized in that, The switch is a PCIe switch, comprising a plurality of ports and a multicast replication engine electrically connected with each of the ports respectively, the multicast replication engine comprising a group matcher and a broadcast module; The multicast replication engine is configured to determine whether a communication request transmitted by a request port is a broadcast request, the request port being one of the ports, if yes, the multicast replication engine sends the communication request to the group matcher and receives broadcast data targeted by the request port; the communication request comprises a broadcast entry, the broadcast entry comprising a broadcast valid bit, when the broadcast valid bit is a target field, the communication request is the broadcast request, when the broadcast valid bit is a non-target field, the communication request is a non-broadcast request; The multicast replication engine is configured to determine whether the broadcast valid bit in the communication request is the target field, if yes, the multicast replication engine determines that the communication request is the broadcast request, otherwise, the multicast replication engine determines that the communication request is a non-broadcast request; The group matcher is configured to, in response to receiving the broadcast request, determine a broadcast port and a corresponding broadcast address according to a preset broadcast group table and transmit to the broadcast module, the broadcast group table comprising memory addresses configured for each of the ports, the broadcast port being a port other than the request port, the broadcast address being one or more of the memory addresses; The broadcast module is configured to, in response to receiving the broadcast port and the broadcast address, control each of the broadcast ports to receive the broadcast data through the broadcast address; The switch is provided with a shared memory, the shared memory comprising a plurality of the memory addresses, each of the ports being configured with a corresponding memory address, the memory address being a DMA address, each of the ports directly communicating with the corresponding DMA address; If the number of the broadcast addresses is a plurality, the memory address corresponding to each of the broadcast ports is configured as the broadcast address; The broadcast module controls each of the broadcast ports to receive the broadcast data through the broadcast address, comprising: the broadcast module replicates the broadcast data into a plurality of copies corresponding to the number of broadcast ports and transmits to each of the broadcast addresses.
2. The switch implementing collective communication based on a multicast replication engine of claim 1, wherein, Each of the ports is configured with a port ID, the broadcast group table comprising port IDs and the memory addresses having a corresponding relationship, the broadcast port and the corresponding broadcast address being determined according to the broadcast group table in the following manner: The broadcast request comprises a request port ID, the request port ID being a port ID corresponding to the request port, a port ID different from the request port ID in the broadcast group table being determined as a target ID, the port corresponding to the target ID being the broadcast port; The memory address corresponding to the target ID is determined as the broadcast address.
3. The switch implementing collective communication based on a multicast replication engine of claim 1, wherein, The switch is provided with a shared memory, the shared memory comprising a plurality of the memory addresses, each of the ports being configured with a corresponding memory address, the port being configured to directly communicate with any of the memory addresses; If the number of the broadcast addresses is one, the memory address corresponding to each of the request ports is configured as the broadcast address; The broadcast module controls each of the broadcast ports to receive the broadcast data through the broadcast address, including that the broadcast module stores the broadcast data into the broadcast address and informs each of the broadcast ports to acquire the broadcast data in the broadcast address.
4. The switch implementing collective communication based on a multicast replication engine of claim 3, wherein, Each of the ports is configured with a port ID, and the broadcast group table includes the port ID and the memory address with a corresponding relationship, and the broadcast port and the corresponding broadcast address are determined according to the broadcast group table in the following manner: The broadcast request includes a request port ID, which is the port ID corresponding to the request port, and the port ID in the broadcast group table that is different from the request port ID is determined as a target ID, and the port corresponding to the target ID is the broadcast port; The memory address corresponding to the request port ID is determined as the broadcast address.
5. The switch based on the multicast replication engine to realize collective communication according to claim 1, wherein If the multicast replication engine determines that the communication request is a non-broadcast request, the switch controls the request port and the target port to communicate based on the communication request.
6. The switch implementing collective communication based on a multicast replication engine of claim 1, wherein, The group matcher, in response to receiving the broadcast request, first determines whether the broadcast valid bit is the target field, and if so, determines the broadcast port and the broadcast address according to a preset broadcast group table, and otherwise, issues an error message.
7. The multicast replication engine based switch to enable collective communication according to claim 1, wherein, The broadcast entry further includes a target port ID, and if the broadcast valid bit is the target field, the group matcher determines that the port corresponding to the target port ID is the broadcast port.
8. The switch implementing collective communication based on a multicast replication engine of claim 1, wherein, The broadcast entry further includes a broadcast start physical address and a broadcast data length, and the broadcast data is data with the start physical address as the start address and the broadcast data length as the data length.
9. The switch implementing collective communication based on a multicast replication engine of claim 1, wherein, The switch is configured as a downstream switch of a plurality of devices to be communicated; or The switch is configured as an upstream switch of a plurality of devices to be communicated.
10. A method of collective communication based on a multicast replication engine, characterized in that, The switch based on the multicast replication engine to realize collective communication according to any one of claims 1-9, comprising the following steps: Determining, by the multicast replication engine, whether a communication request transmitted by a request port is a broadcast request, the request port being one of the ports, and if so, the multicast replication engine sends the communication request to the group matcher and receives broadcast data targeted by the request port; The group matcher, in response to receiving the broadcast request, determines a broadcast port and a corresponding broadcast address according to a preset broadcast group table and transmits them to the broadcast module, the broadcast group table including a memory address configured for each of the ports, the broadcast port being a port other than the request port, and the broadcast address being one or more of the memory addresses; The broadcast module, in response to receiving the broadcast port and the broadcast address, controls each of the broadcast ports to receive the broadcast data through the broadcast address.
11. A communication system, characterized by A switch including a multicast replication engine as claimed in any of claims 1-9 to enable collective communication.
Citation Information
Patent Citations
Switching method and apparatus of real-time multicastpacket stream, and ethernet switching system using thesame
KR1020070053502A