Data exchange system and method
By providing a configurable data switching system within the computing node and supporting the shared use of multiple data switching types, it solves the problem of power consumption and cost increase caused by multiple switching chips in the computing node, achieving higher access flexibility and lower system power consumption and cost.
Patent Information
- Application Number
- PCT/CN2024/141032
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
In distributed computing, there are multiple data exchange types of switching chips in a single computing node, resulting in an increase in the number of chips inside the computing node, increasing power consumption and cost. At the same time, the physical ports are limited, which cannot meet the needs of high throughput communication, limiting the access flexibility of the computing node.
It provides a data exchange system, including a configurable physical port and a switching unit, and distributes and aggregates messages according to the data exchange type of the physical port, supports the shared use of multiple data exchange types, and improves the utilization rate of the physical port.
By supporting the shared use of multiple data exchange types, the number of switching chips or devices within the computing node is reduced, the system power consumption and cost are reduced, and the access flexibility of the computing node is improved.
Smart Images

Figure CN2024141032_26062025_PF_FP_ABST
Abstract
Description
System and method for data exchange
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent applications filed with the China Patent Office on December 21, 2023, with application number 202311775785.5 and disclosed name “System and method for data exchange” and filed with the China Patent Office on August 5, 2024, with application number 202411064478.0 and disclosed name “System and method for data exchange”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to distributed computing, and in particular to a system, method, electronic device, and medium for data exchange. Background Art
[0004] In distributed computing, such as that used for AI (Artificial Intelligence) model training, a computing node is usually composed of multiple different chips, typically including CPUs, GPUs, network cards, and various application accelerator chips. These chips require communication with each other, and for scalability, they are generally interconnected using switching chips. Different chips have different interconnection requirements, and a computing node may contain multiple types of switching chips for data exchange. For example, Intel's CPUs typically only provide PCIe interfaces to the outside world and can only be interconnected with GPUs, network cards, application accelerators, and other chips through PCIe switching chips. Nvidia's GPUs, because PCIe bandwidth does not meet the high-throughput performance requirements of AI training, have developed their own NV Switch chips for interconnection between multiple GPUs. Multiple application accelerator chips may also be interconnected using Ethernet data exchange-type switching chips due to their communication characteristics.
[0005] See Figure 1, which shows a network topology for deploying AI accelerators based on a general-purpose computing cluster interconnection network. This network, comprised of switches and routers, connects to multiple compute nodes. Each compute node primarily consists of a CPU, network interface card (NIC), and AI accelerator. CPUs are interconnected using a proprietary cache coherence protocol implemented by the CPU vendor, while CPUs, NICs, and AI accelerators from different vendors are interconnected using standard protocols, such as PCIe. Within a compute node, CPUs and AI accelerators, and AI accelerators, typically communicate via PCIe switches. Across compute nodes, CPUs and AI accelerators, and AI accelerators, communicate via PCIe switches and then network interfaces (e.g., Ethernet cards).
[0006] As AI models and training data grow in size, communication throughput requirements between AI accelerators within a single compute node are also increasing. The PCIe switch in Figure 1 uses a single-root model (i.e., communication between any two nodes must pass through the CPU), resulting in low communication efficiency and insufficient support for the high-throughput communication demands between AI accelerators. Furthermore, PCIe protocol updates and upgrades are completely controlled by Intel. Therefore, some accelerator vendors use standard Ethernet protocols or customized protocols to interconnect their accelerators. For example, NVIDIA has defined NVLink / NVSwitch for interconnecting its own GPUs, enabling a single compute node to support up to eight GPUs. As shown in Figure 2, AI accelerators within a compute node communicate using another switch. At the same time, there is also high-throughput communication traffic between AI accelerators across computing nodes. If the network card is connected through a PCIe switch and then passes through a general computing cluster interconnection network, it still cannot meet the demand for high-throughput communication, and it will also cause the communication traffic of the AI accelerator and the general computing communication traffic to affect each other. Therefore, a dedicated AI accelerator cluster interconnection network is generally established across computing nodes. Therefore, based on the network topology of deploying AI accelerators based on the general computing cluster interconnection network, ports are set on the switches that interconnect AI accelerators to connect to the AI accelerator cluster interconnection network to achieve interconnection between AI accelerators across computing nodes. See Figure 2.
[0007] A single compute node contains two types of switches, and the mixed use of multiple switch chips not only increases the number of chips within the compute node but also increases the node's power consumption and cost. Furthermore, a switch chip has a fixed number of access physical ports. In practice, it's very likely that the physical ports of a switch chip of a certain data exchange type will be used up, making it impossible to expand that switch chip's physical ports. However, physical ports of other switch chips of other data exchange types may still be available, limiting the compute node's access flexibility. Summary of the Invention
[0008] Based on the above situation, the main purpose of this application is to provide a system, method, electronic device and medium for data exchange.
[0009] To achieve the above objectives, the technical solutions adopted in this application are as follows:
[0010] A first aspect of the present application provides a data exchange system for processing message transmission of multiple data exchange types, the system comprising:
[0011] A plurality of physical ports, used for receiving messages on the input side and / or sending messages on the output side, wherein the data exchange type of each physical port can be configured as any one of the multiple data exchange types;
[0012] a plurality of switching units, configured to parse and process the message, wherein each of the switching units can be configured as one of the multiple data exchange types;
[0013] The distribution and aggregation unit is used to distribute the message on the input side to the corresponding switching unit according to the data exchange type of each physical port, and send the message to the corresponding physical port for sending the message on the output side according to the processing result of the switching unit.
[0014] Preferably, each physical port is provided with a configuration unit for configuring the data exchange type of the physical port.
[0015] Preferably, the distribution and aggregation unit is used to receive a message from a first physical port, send the message to a switching unit of a corresponding data exchange type according to the data exchange type of the first physical port, receive a processing result of the switching unit, and send the message to a second physical port of a corresponding data exchange type according to the processing result.
[0016] Preferably, the system further comprises:
[0017] The receiving side physical layer module is used to perform receiving side physical layer processing on the messages received from the physical port;
[0018] The sending side physical layer module is used to perform sending side physical layer processing on the message sent by the distribution and aggregation unit to the physical port.
[0019] Preferably, the switching unit includes:
[0020] A receiving-side link layer module, configured to perform receiving-side link layer processing on messages received from the distribution and aggregation unit;
[0021] The routing module is used to query the routing table according to the parsing result of the message and determine the physical port used to send the message;
[0022] The sending side link layer module is used to perform sending side link layer processing on the parsed and processed messages.
[0023] Preferably, the switching unit further includes:
[0024] A message parsing module, used to parse the message processed by the receiving side link layer module;
[0025] The message editing module is used to edit the message that needs to be sent to the sending side link layer module.
[0026] Preferably, the system further comprises:
[0027] A congestion management unit is used to perform queue management on messages of multiple data exchange types from the input side for use in message sending on the output side.
[0028] Preferably, the data exchange type includes CXL, PCIe, and Ethernet exchange.
[0029] A second aspect of the present application provides a data exchange method for processing message transmission of multiple data exchange types, the method comprising the following steps:
[0030] The input side of the message transmission receives the message;
[0031] Obtaining a data exchange type of the physical port that receives the message, determining a switching unit for the message according to the data exchange type, and sending the message to the switching unit;
[0032] The switching unit parses and processes the message;
[0033] The message is sent to the corresponding physical port according to the processing result of the switching unit for message sending on the output side.
[0034] Preferably, the method further comprises the steps of:
[0035] Pre-configure the data switching type for each physical port.
[0036] Preferably, the input side of the message transmission receives the message, comprising the following steps:
[0037] Perform receiving-side physical layer processing on the packets received by each physical port.
[0038] Preferably, the acquiring the data exchange type of the physical port receiving the message, determining the switching unit of the message according to the data exchange type, and sending the message to the switching unit comprises the following steps:
[0039] A message is received from a first physical port, and the message is sent to a switching unit corresponding to a data exchange type according to a data exchange type of the first physical port.
[0040] Preferably, the switching unit parses and processes the message, comprising the following steps:
[0041] Performing receiving-side link layer processing on the message received from the input side;
[0042] Parsing the message processed by the receiving side link layer, querying the routing table according to the parsing result and determining the physical port for sending the message;
[0043] The parsed message is edited and processed at the sending side link layer.
[0044] Preferably, the step of sending the message to the corresponding physical port according to the processing result of the switching unit for sending the message on the output side comprises the following steps:
[0045] Receive the processing result of the switching unit, and send the message to the second physical port corresponding to the data exchange type according to the processing result.
[0046] Preferably, the step of sending the message to the corresponding physical port according to the processing result of the switching unit for sending the message on the output side comprises the following steps:
[0047] Perform sending-side physical layer processing on the message sent to the corresponding physical port.
[0048] Preferably, the method further comprises the steps of:
[0049] Queue management is performed on messages of multiple data exchange types from the input side for use in message sending on the output side.
[0050] Preferably, the data exchange type includes CXL, PCIe, and Ethernet exchange.
[0051] A third aspect of the present application provides a switching chip, wherein the switching chip includes the system as described in the first aspect above.
[0052] Preferably, the switching chip is used in a distributed computing node.
[0053] A third aspect of the present application further provides a switching chip, comprising:
[0054] programmable processing devices; and
[0055] A memory communicatively connected to the programmable processing device; wherein the memory stores a computer program executable by the programmable processing device, and the computer program is executed by the programmable processing device to enable the programmable processing device to perform the method as described in the second aspect above.
[0056] Preferably, the switching chip is used in a distributed computing node.
[0057] The fourth aspect of the present application provides an electronic device, comprising: a processor; and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it can implement the method described in the second aspect above.
[0058] A fifth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is configured to be run to implement the method described in the second aspect above.
[0059] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] FIG1 is a schematic diagram of a computing node network topology based on a general computing cluster interconnection network in the prior art;
[0061] FIG2 is a schematic diagram of a computing node network topology based on a general computing cluster interconnection network and an AI accelerator cluster interconnection network in the prior art;
[0062] FIG3 is a flow chart of a preferred embodiment of the data exchange method of the present application;
[0063] FIG4 is a flow chart of another preferred embodiment of the data exchange method of the present application;
[0064] FIG5 is a schematic diagram of a computing node network topology according to a preferred embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to further illustrate the technical means and effects adopted by this application to achieve the intended application purpose, the following, in combination with the accompanying drawings and preferred embodiments, describes in detail the specific implementation methods, methods, steps, features and effects of the methods, systems, electronic devices and computer-readable storage media proposed in this application.
[0066] A data exchange system for processing message transmission of multiple data exchange types, the system comprising:
[0067] A plurality of physical ports, used for receiving messages on the input side and / or sending messages on the output side, wherein the data exchange type of each physical port can be configured as any one of the multiple data exchange types;
[0068] a plurality of switching units, configured to parse and process the message, wherein each of the switching units can be configured as one of the multiple data exchange types;
[0069] The distribution and aggregation unit is used to distribute the message on the input side to the corresponding switching unit according to the data exchange type of each physical port, and send the message to the corresponding physical port for sending the message on the output side according to the processing result of the switching unit.
[0070] Specifically, a data exchange system, such as a switching chip, usually has multiple physical ports for transmitting messages. Generally, these physical ports can be divided into input physical ports and output physical ports according to the direction of the business flow. When a physical port is used as an input physical port, it is used to receive messages on the input side, and when a physical port is used as an output physical port, it is used to send messages on the output side.
[0071] Regarding the data exchange type, multiple chips in a computing node, such as CPU, AI accelerators (including GPU / GPGPU, TPU, NPU, etc.), DPU (including NIC / SNIC, etc.) are interconnected. The data exchange type is generally divided from the perspective of the cache access model, which can be divided into Cache consistency interconnection and Cache non-consistent interconnection. Several switching units for different data exchange types are set in the switching chip, and the distribution and aggregation units can distribute the messages received on the input side to the switching units of the corresponding data exchange type according to the data exchange type configured on the physical port. At the same time, arbitration is performed between switching units of different data exchange types, and then the messages are output to the corresponding physical port for sending messages on the output side.
[0072] Regarding the distribution and aggregation unit, the message data of different data exchange types on the input side can be distributed to the corresponding switching unit for link layer processing after processing, and the message parsed and processed by the switching unit is sent to the physical port of the corresponding output side.
[0073] Therefore, the data exchange system of the present application can handle the transmission of messages of multiple data exchange types, wherein the data exchange type of the physical port can be configured according to actual needs to access devices of different data exchange types, thereby improving the access flexibility of the computing node, and the distribution and aggregation unit can distribute the message to the switching unit of the corresponding data exchange type according to the physical port of different data exchange types to meet the needs of message transmission. Thus, the multiple chips or devices of different data exchange types in the computing node are normalized, so that the communication transmission of multiple data exchange types can share the physical port, significantly improving the utilization rate of the physical port, effectively reducing the number of switching chips or devices inside the computing node, and helping to reduce system power consumption and cost. In addition, the switching chip that can adapt to multiple data exchange types can avoid the development of different switching chips for different data exchange types, effectively reducing the R&D costs of chip manufacturers.
[0074] As an optional embodiment, each physical port is provided with a configuration unit for configuring the data exchange type of the physical port.
[0075] Specifically, the configuration unit can configure the data exchange type of each physical port according to actual needs to meet the access needs of different chips or devices. For example, if a physical port is used for interconnection between CPUs, the data exchange type of the physical port can be configured as a CXL type that supports cache consistency; for another example, if a physical port is used for interconnection between CPUs and GPUs, the data exchange type of the physical port can be configured as a PCIe type that supports cache inconsistency. If the bandwidth requirement for cache consistency interconnection is high, the number of physical ports for cache inconsistency interconnection can be reduced, and more physical ports can be configured as data exchange types that support cache consistency interconnection.
[0076] In addition, if a physical port is used for interconnection between GPUs, the data exchange type of the physical port can be configured as a peer-to-peer protocol Ethernet type. It should be noted that the peer-to-peer protocol is for the PCIe / CXL protocol. This is because there is a root node in the PCIe / CXL protocol. Communication between any two nodes will pass through the root node, that is, all nodes in the network are not in a peer-to-peer relationship. Therefore, PCIe / CXL is a non-peer protocol network model, while communication between any two nodes in an Ethernet network does not require transit through a third node. It is a peer-to-peer protocol network model. Therefore, the data exchange type of the physical port can be configured through the configuration unit to meet the message transmission requirements of different data exchange types, effectively improving the utilization rate of the physical port, so that the physical port can be shared in the transmission of multiple data exchange type messages, which helps to improve the access flexibility of the data exchange system.
[0077] As an optional embodiment, the distribution and aggregation unit receives a message from the first physical port, sends the message to a switching unit of a corresponding data exchange type according to the data exchange type of the first physical port, receives a processing result of the switching unit, and sends the message to a second physical port of a corresponding data exchange type according to the processing result.
[0078] Specifically, on the input side of message transmission, multiple physical ports receive message data of different data exchange types, and the distribution and aggregation unit distributes the corresponding messages of each physical port to the switching unit of the corresponding data exchange type according to the data exchange type for receiving side link layer processing; on the output side of message transmission, the distribution and aggregation unit arbitrates between multiple switching units of different data exchange types, and sends each message to the corresponding physical port for output side message sending.
[0079] Therefore, through the distribution and aggregation unit, different data exchange types of message transmission can be realized within the same data exchange system, without the need to set up different data exchange systems to support different data exchange types of message transmission. It is used for interconnection between different chips in the computing node and can support multiple data exchange types at the same time.
[0080] As an optional embodiment, the system further includes:
[0081] The receiving side physical layer module is used to perform receiving side physical layer processing on the messages received from the physical port;
[0082] The sending side physical layer module is used to perform sending side physical layer processing on the message sent by the distribution and aggregation unit to the physical port.
[0083] Specifically, the message received by the physical port on the input side is processed by the physical layer on the receiving side and then sent to the switching unit via the distribution and aggregation unit for subsequent link layer processing on the receiving side; accordingly, the message is parsed and processed by the switching unit (including link layer processing on the sending side) and is also sent to the output side via the distribution and aggregation unit, and is sent out by the corresponding physical port after being processed by the physical layer on the sending side.
[0084] As an optional embodiment, the switching unit includes:
[0085] A receiving-side link layer module, configured to perform receiving-side link layer processing on messages received from the distribution and aggregation unit;
[0086] The routing module is used to query the routing table according to the parsing result of the message and determine the physical port used to send the message;
[0087] The sending side link layer module is used to perform sending side link layer processing on the parsed and processed messages.
[0088] Specifically, the message is processed by the receiving side physical layer on the input side and sent to its corresponding switching unit by the distribution and aggregation unit. The switching unit processes the message through the receiving side link layer module and the sending side link layer module, and obtains the next physical port for message transmission by parsing the message and querying the routing table.
[0089] As a further improvement of the above embodiment, the switching unit further includes:
[0090] A message parsing module, used to parse the message processed by the receiving side link layer module;
[0091] The message editing module is used to edit the message that needs to be sent to the sending side link layer module.
[0092] Specifically, the routing table is queried according to the result of the message parsing to determine the physical port for sending the message.
[0093] As a further improvement to the above embodiment, the system further includes:
[0094] A congestion management unit is used to perform queue management on messages of multiple data exchange types from the input side for use in message sending on the output side.
[0095] Specifically, a data exchange system includes multiple input physical ports and multiple output physical ports. During packet transmission, packets from multiple input physical ports may simultaneously be directed to the same output physical port, or packets from a high-speed input physical port may need to be directed to a low-speed output physical port. This can cause congestion, and a congestion management unit can be used to handle such situations during packet transmission. The congestion management unit typically places packets into different queues based on information such as packet priority, input physical port, and output physical port. A scheduler then selects between the queues based on preset scheduling parameters, such as priority and weight, and the bandwidth capacity of the output physical ports, determining the order in which packets are sent from the output physical ports. Generally speaking, the congestion management unit only needs to understand the length of the packet and is not concerned with the specific data exchange type being transmitted. Therefore, this embodiment allows for multiplexing of the congestion management unit in a system that handles multiple data exchange types, rather than simply combining multiple systems for different data exchange types, resulting in a system implementation cost of 1+1<2.
[0096] Therefore, the congestion management unit manages the message transmission queue, and the congestion management unit is generalized so that the unit can receive messages to be transmitted from switching units of different data exchange types, and send messages according to the scheduling parameters of message transmission, so as to optimize the message processing performance within the same data exchange system. At the same time, it can realize the sharing and reuse of general units or modules between different data exchange types within the system, which is conducive to reducing implementation costs.
[0097] As an optional embodiment, the data exchange type includes CXL, PCIe, and Ethernet exchange.
[0098] Specifically, the data exchange types that the data exchange system can adapt to include but are not limited to CXL, PCIe, Ethernet exchange, etc.
[0099] Among them, the PCIe (PCI Express) protocol has replaced all internal buses including AGP and PCI with its high performance, high scalability, high reliability and excellent compatibility. However, in response to the terabyte-level growth of data and the trend of heterogeneous computing, PCIe has faced pressure in terms of memory utilization efficiency, latency and data throughput.
[0100] The PCIe-based Compute Express Link (CXL) protocol, a new open interconnect technology standard, enables high-speed and efficient interconnection between CPUs and GPUs, FPGAs, and other accelerators, meeting the requirements of high-performance heterogeneous computing. It also maintains coherence between the CPU memory space and the memory / cache of connected devices. Overall, its advantages lie in its extremely high compatibility and memory / cache coherence.
[0101] Therefore, by supporting the transmission of different data exchange types of messages within the same data exchange system, for example, the use requirements of cache consistency and cache inconsistency can be met at the same time, and the utilization rate of physical ports can be improved, effectively reducing the number of switching chips or devices inside the computing node, which is conducive to reducing system power consumption and cost.
[0102] Referring to FIG3 , a data exchange method is provided for processing message transmissions of multiple data exchange types. The method comprises the following steps:
[0103] S100, the input side of the message transmission receives a message;
[0104] S200, obtaining a data exchange type of the physical port that receives the message, determining a switching unit for the message according to the data exchange type, and sending the message to the switching unit;
[0105] S300, the switching unit parses and processes the message;
[0106] S400: Send the message to the corresponding physical port according to the processing result of the switching unit for message sending on the output side.
[0107] Through the above steps, the transmission path of each message is determined according to the data exchange type of the physical port to realize the transmission of messages of multiple data exchange types, so that communication transmissions of multiple data exchange types can share physical ports, effectively reducing the number of switching chips or devices inside the computing node, which is conducive to reducing system power consumption and cost.
[0108] As an optional embodiment, the method further includes the following steps:
[0109] Pre-configure the data switching type for each physical port.
[0110] Through the above steps, the data exchange type of the physical port is configured to meet the message transmission requirements of different data exchange types, effectively improving the utilization of the physical port, allowing the physical port to be shared in the transmission of messages of multiple data exchange types, and helping to improve the access flexibility of the data exchange system.
[0111] As an optional embodiment, referring to FIG4 , the input side of the message transmission receives a message, including the following steps:
[0112] Perform receiving-side physical layer processing on the packets received by each physical port.
[0113] Through the above steps, the message received by the input side is processed by the physical layer of the receiving side and then sent to the switching unit via the distribution and aggregation unit for subsequent link layer processing on the receiving side.
[0114] As an optional embodiment, obtaining the data exchange type of the physical port receiving the message, determining the switching unit of the message according to the data exchange type, and sending the message to the switching unit includes the following steps:
[0115] A message is received from a first physical port, and the message is sent to a switching unit corresponding to a data exchange type according to a data exchange type of the first physical port.
[0116] Through the above steps, on the input side of the message transmission, multiple physical ports receive message data of different data exchange types, and distribute the corresponding messages to the switching units of the corresponding data exchange types according to the data exchange type of each physical port for receiving side link layer processing, thereby enabling message transmission of different data exchange types to be realized in the same data exchange system without setting up different data exchange systems to support message transmission of different data exchange types. It is used for interconnection between different chips in a computing node and can support multiple data exchange types at the same time.
[0117] As an optional embodiment, referring to FIG4 , the switching unit parses and processes the message, including the following steps:
[0118] Performing receiving-side link layer processing on the message received from the input side;
[0119] Parsing the message processed by the receiving side link layer, querying the routing table according to the parsing result and determining the physical port for sending the message;
[0120] The parsed message is edited and processed at the sending side link layer.
[0121] Through the above steps, the message is processed by the receiving side physical layer on the input side and sent to its corresponding switching unit. The switching unit processes the message through the receiving side link layer module and the sending side link layer module, and obtains the next physical port for message transmission by parsing the message and querying the routing table.
[0122] As an optional embodiment, the sending of the message to the corresponding physical port according to the processing result of the switching unit for sending the message on the output side includes the following steps:
[0123] Receive the processing result of the switching unit, and send the message to the second physical port corresponding to the data exchange type according to the processing result.
[0124] Through the above steps, on the output side of message transmission, the distribution and aggregation unit arbitrates between multiple switching units of different data exchange types and sends each message to the corresponding physical port for output side message sending.
[0125] As an optional embodiment, referring to FIG4 , the step of sending the message to the corresponding physical port for sending the message on the output side according to the processing result of the switching unit includes the following steps:
[0126] Perform sending-side physical layer processing on the message sent to the corresponding physical port.
[0127] Through the above steps, the message is parsed and processed by the switching unit and processed by the link layer on the sending side and sent to the output side. After being processed by the physical layer on the sending side, the message is sent out by the corresponding physical port.
[0128] As an optional embodiment, referring to FIG4 , the method further includes the following steps:
[0129] Queue management is performed on messages of multiple data exchange types from the input side for use in message sending on the output side.
[0130] Through the above steps, congestion management of internal message transmission can receive messages to be transmitted from switching units of different data exchange types, and send messages according to the scheduling parameters of message transmission, so as to optimize the message processing performance within the system of the same data exchange. At the same time, it can realize the sharing and multiplexing of common units or modules between different data exchange types within the system, which is conducive to reducing implementation costs.
[0131] As an optional embodiment, the data exchange type includes CXL, PCIe, and Ethernet exchange.
[0132] Therefore, by supporting the transmission of different data exchange types of messages within the same data exchange system, for example, the use requirements of cache consistency and cache inconsistency can be met at the same time, and the utilization rate of physical ports can be improved, effectively reducing the number of switching chips or devices inside the computing node, which is conducive to reducing system power consumption and cost.
[0133] The present application also provides a switching chip, which includes the system described in the above embodiment.
[0134] Referring to Figure 5, the switching chip described above is used in the computing node. The switching chip includes but is not limited to a PCIe Switch module and an Ethernet Switch module. The CPU and AI accelerator can be interconnected through the PCIe Switch module of the switching chip. The AI accelerators across the computing nodes can be interconnected through the Ethernet Switch module of the switching chip and then through the AI accelerator cluster interconnection network. In addition, the AI accelerators within the computing node can be interconnected through the PCIe Switch module and / or the Ethernet Switch module. As mentioned above, PCIe is a non-peer protocol network model. Communication between any two nodes needs to go through the root node, while Ethernet is a peer protocol network model. Communication between any two nodes does not need to be transferred through a third node.
[0135] It can be seen that the number of switching chips or devices in the computing node is significantly reduced, and at the same time, the communication traffic of the AI accelerator and the general computing communication traffic can be prevented from affecting each other.
[0136] As an optional embodiment, the switching chip is used in a distributed computing node.
[0137] The present application also provides a switching chip, comprising:
[0138] programmable processing devices; and
[0139] A memory communicatively connected to the programmable processing device; wherein the memory stores a computer program executable by the programmable processing device, and the computer program is executed by the programmable processing device to enable the programmable processing device to perform the method described in the above embodiment.
[0140] As an optional embodiment, the switching chip is used in a distributed computing node.
[0141] The present application also provides an electronic device, comprising: a processor; and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the data exchange method as described in the above embodiment can be implemented.
[0142] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is configured to be run to implement the data exchange method as described in the above embodiment.
[0143] The switching chip, electronic device and computer-readable storage medium of the present application, by adopting the above-mentioned system or method, enable the physical port used for message transmission to be shared, which not only helps to improve the access flexibility of the computing node, but also can effectively reduce the number of switching chips or devices inside the computing node.
[0144] The above description is merely a preferred embodiment of the present application and does not constitute any form of limitation to the present application. Although the present application has been disclosed as a preferred embodiment, it is not intended to limit the present application. Any technician familiar with the present profession can make slight changes or modifications to equivalent embodiments of the technical contents disclosed above without departing from the scope of the technical solution of the present application. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A data exchange system for processing message transmission of multiple data exchange types, characterized in that: The system comprises: A plurality of physical ports, used for receiving messages on the input side and / or sending messages on the output side, wherein the data exchange type of each physical port can be configured as any one of the multiple data exchange types; A plurality of switching units, used for parsing and processing the message, wherein each of the switching units can be configured as one of the multiple data exchange types; The distribution and aggregation unit is used to distribute the message on the input side to the corresponding switching unit according to the data exchange type of each physical port, and send the message to the corresponding physical port for sending the message on the output side according to the processing result of the switching unit.
2. The system according to claim 1, characterized in that Each physical port is provided with a configuration unit for configuring the data exchange type of the physical port.
3. The system according to claim 2, characterized in that The distribution and aggregation unit is used to receive a message from a first physical port, send the message to a switching unit of a corresponding data exchange type according to a data exchange type of the first physical port, receive a processing result of the switching unit, and send the message to a second physical port of a corresponding data exchange type according to the processing result.
4. The system according to claim 1, characterized in that The system further comprises: The receiving side physical layer module is used to perform receiving side physical layer processing on the message received from the physical port; The sending side physical layer module is used to perform sending side physical layer processing on the message sent by the distribution and aggregation unit to the physical port.
5. The system according to claim 1, wherein: The switching unit comprises: A receiving-side link layer module, used for performing receiving-side link layer processing on the message received from the distribution and aggregation unit; A routing module, used to query the routing table according to the parsing result of the message and determine the physical port used to send the message; The sending side link layer module is used to perform sending side link layer processing on the parsed and processed messages.
6. The system according to claim 5, characterized in that The switching unit also includes: A message parsing module, used to parse the message processed by the receiving side link layer module; The message editing module is used to edit the message that needs to be sent to the sending side link layer module.
7. The system according to claim 1, characterized in that The system further comprises: The congestion management unit is used to perform queue management on messages of multiple data exchange types from the input side for use in message sending on the output side.
8. The system according to any one of claims 1 to 7, characterized in that: The data exchange types include CXL, PCIe, and Ethernet exchange.
9. A data exchange method for processing message transmission of multiple data exchange types, characterized in that: The method comprises the following steps: The input side of the message transmission receives the message; Acquire a data exchange type of a physical port that receives the message, determine a switching unit for the message according to the data exchange type, and send the message to the switching unit; The switching unit parses and processes the message; The message is sent to the corresponding physical port according to the processing result of the switching unit for message sending on the output side.
10. The method according to claim 9, characterized in that The method further comprises the steps of: Preconfigure the data switching type for each physical port.
11. The method according to claim 9, characterized in that The input side of the message transmission receives the message, including the following steps: Perform receiving-side physical layer processing on the messages received by each physical port.
12. The method according to claim 9, characterized in that The step of obtaining the data exchange type of the physical port receiving the message, determining the switching unit of the message according to the data exchange type, and sending the message to the switching unit comprises the following steps: A message is received from a first physical port, and according to a data exchange type of the first physical port, the message is sent to a switching unit corresponding to the data exchange type.
13. The method according to claim 9, characterized in that The switching unit parses and processes the message, including the following steps: Performing receiving-side link layer processing on the message received from the input side; Parsing the message processed by the receiving side link layer, querying the routing table according to the parsing result and determining the physical port used to send the message; The parsed message is edited and processed at the sending side link layer.
14. The method according to claim 9, characterized in that The step of sending the message to the corresponding physical port according to the processing result of the switching unit for sending the message on the output side comprises the following steps: A processing result of the switching unit is received, and the message is sent to a second physical port corresponding to the data exchange type according to the processing result.
15. The method according to claim 9, characterized in that The step of sending the message to the corresponding physical port according to the processing result of the switching unit for sending the message on the output side comprises the following steps: Perform sending-side physical layer processing on the message sent to the corresponding physical port.
16. The method according to claim 9, characterized in that The method further comprises the steps of: Queue management is performed on messages of multiple data exchange types from the input side for use in message sending on the output side.
17. The method according to any one of claims 9 to 16, characterized in that: The data exchange types include CXL, PCIe, and Ethernet exchange.
18. A switching chip, characterized in that: The switching chip comprises the system as claimed in any one of claims 1 to 8.
19. The switching chip according to claim 18, characterized in that: The switching chip is used in a distributed computing node.
20. A switching chip, characterized in that: The switching chip comprises: A programmable processing device; and A memory communicatively connected to the programmable processing device; wherein the memory stores a computer program executable by the programmable processing device, the computer program being executed by the programmable processing device so that the programmable processing device can perform the method as claimed in any one of claims 9 to 17.
21. The switching chip according to claim 20, characterized in that: The switching chip is used in a distributed computing node.
22. An electronic device, characterized in that: include: processor; as well as A memory having a computer program stored thereon, wherein the computer program, when executed by the processor, can implement the method according to any one of claims 9 to 17.
23. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is used to run to implement the method according to any one of claims 9 to 17.
Citation Information
Patent Citations
Message forwarding method and device
CN111278059A
Apparatuses for Hybrid Wired and Wireless Universal Access Networks
US20100085948A1
Multi-CPU packet processing method and system, switching unit and board
WO2014023023A1