Data communication method and system, electronic equipment, computer readable storage medium and computer program product
By mapping switch and host resource information and using switch routing information for port forwarding, the problem of low communication efficiency between computing units is solved, efficient and flexible point-to-point communication is achieved, and system scale expansion is supported.
Patent Information
- Application Number
- CN202510304329.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The prior art has additional delays and hardware overhead when implementing point-to-point communication between computing units, low communication efficiency and difficult to scale the system to improve communication efficiency.
By corresponding to the resource information allocated by the switches for each connected computing unit, the resource information allocated by the host for all the computing units within it, and the routing information between the switches is used for port forwarding, point-to-point communication between the computing units is realized without the need to forward with the upstream port.
It realizes efficient point-to-point communication between computing units, avoids additional hardware overhead and address conversion delay, improves communication efficiency, and supports system scale expansion.
Smart Images

Figure CN120017619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a data communication method, system, electronic equipment, computer-readable storage medium and computer program product. Background Art
[0002] As the scale of user data and the demand for high-speed data processing increase, different types of computing units are currently used in the same computing system to complete tasks together. Related technologies use PCIe (Peripheral Component Interconnect Express) NTB (Non-Transparent Bridge) technology to achieve P2P (Point-to-Point) communication between computing units. This method relies on the complex logic of the NTB controller of the PCIeSwitch to map the remote node address space to the local address space through the NTB to open a point-to-point access path. This cross-domain address mapping and interrupt mechanism will introduce additional delays and hardware overhead, and the communication efficiency is low. Summary of the invention
[0003] The present invention provides a data communication method, system, electronic device, computer-readable storage medium and computer program product, which not only do not need to introduce additional delays and hardware overheads, but also can effectively improve the efficiency of point-to-point communication.
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] In one aspect, the present invention provides a data communication method, which is applied to a data communication system including at least two computing nodes, wherein at least two computing nodes include at least one group of computing units, and the same group of computing units is connected to the same switch, comprising:
[0006] When the source computing unit receives an access request to the target computing unit, the resource allocation information of the target switch to which the target computing unit belongs is determined according to the routing mapping relationship and the target resource allocation information of the target computing unit carried in the access request; the source computing unit and the target computing unit belong to different computing nodes; an intermediate forwarding port is determined according to the resource allocation information and the switch routing information of the source switch corresponding to the source computing unit, and the access request is forwarded to the target communication port of the target switch through the intermediate forwarding port;
[0007] forwarding the access request to the target computing unit again via the target communication port and the target resource allocation information;
[0008] Among them, the routing mapping relationship is the correspondence between the host resource allocation information and the switch resource allocation information of the same computing unit, the host resource allocation information is the resource information allocated by the host to the computing unit, the switch resource allocation information is the resource information allocated by the switch to the computing unit, and the switch routing information includes port information corresponding to the data communication path between different switches.
[0009] The present invention also provides an electronic device, comprising a memory and a processor, wherein the processor is used to implement the steps of any of the above-mentioned data communication methods when executing a computer program stored in the memory.
[0010] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned data communication methods are implemented.
[0011] The present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above data communication methods when executed by a processor.
[0012] Finally, the present invention also provides a data communication system, comprising a host having a first computing node, a second computing node and a routing controller, wherein the host stores host resource allocation information;
[0013] The first computing node includes at least a first switch and a second switch; the first switch is connected to at least a first computing unit and a second computing unit, and the second switch is connected to at least a third computing unit and a fourth computing unit; the first switch and the second switch both store their own switch resource allocation information and switch routing information;
[0014] The second computing node includes at least a third switch and a fourth switch; the third switch is connected to at least the fifth computing unit and the sixth computing unit, and the fourth switch is connected to at least the seventh computing unit and the eighth computing unit; the third switch and the fourth switch both store their respective switch resource allocation information and switch routing information;
[0015] The routing controller is used to implement the steps of any one of the above data communication methods when executing a computer program when the first computing node accesses each computing unit, or when the computing units access each other.
[0016] The technical solution provided by the present invention has the advantage that the resource information allocated by the switch to the computing units connected to each other is matched with the resource information allocated by the host to all the computing units inside it, and the information of the switch connected to the computing unit to be accessed can be quickly located through this address mapping. There is routing information that can communicate between each switch, so that the computing units connected to different switches can realize fast port forwarding through the switch routing information, thereby realizing point-to-point communication between any computing units through single hop or multi-hop, without the need for upstream port forwarding, and without introducing additional hardware overhead and address conversion delay, effectively improving the efficiency of point-to-point communication. Further, by increasing the number of switches, not only the scale of the computing units owned by the computing nodes can be increased, but also the number of computing nodes contained in the host can be expanded. On the basis of the expansion of the scale of the data communication system, the efficient routing path determination between any computing units can still be realized, thereby improving the point-to-point communication of large-scale data communication systems, and providing a more efficient, lower latency, flexible and scalable routing deployment and management method for expanding the PCIe Switch system topology of multiple cards, thereby improving the communication efficiency of the entire system, and can also solve the problem of bandwidth limitation of parallel communication requirements between computing units and hosts, and computing units to a certain extent. In addition, the present invention also provides a corresponding implementation system, electronic device, computer-readable storage medium and computer program product for the data communication method, which further makes the method more practical, and the system, electronic device, computer-readable storage medium and computer program product have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a schematic diagram of a data communication method in the related art; Figure 2 A schematic diagram of a system framework in an exemplary application scenario provided by the present invention; Figure 3 A schematic diagram of a flow chart of a data communication method provided by the present invention; Figure 4 A schematic diagram of data access in an exemplary application scenario provided by the present invention; Figure 5 A schematic diagram of data access in another exemplary application scenario provided by the present invention; Figure 6A schematic diagram of data access in another exemplary application scenario provided by the present invention; Figure 7 A structural framework diagram of an exemplary implementation of a data communication device provided by the present invention; Figure 8 A structural framework diagram of an exemplary implementation of a data communication system provided by the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations of the two are intended to cover non-exclusive inclusions. The term "exemplary" means "used as an example, embodiment or illustrative". Any embodiment described here as "exemplary" is not necessarily interpreted as being superior or better than other embodiments.
[0020] As the scale of user data and the demand for high-speed data processing are increasing, different types of computing units are currently used in the same computing system to complete tasks together, such as by inserting multiple GPUs (Graphics Processing Units) into the server to complete tasks together with the CPU (Central Processing Unit). Taking the field of artificial intelligence as an example, traditional artificial intelligence servers usually only support small-scale (e.g., 8-card) GPU module expansion, and memory data can be exchanged between computing nodes through DMA (Direct Memory Access). In order to meet the growing training and reasoning needs of large artificial intelligence models, artificial intelligence servers will interconnect more GPUs, such as through switches, and interconnect multiple GPUs through Scale-up (vertical expansion), such as allowing the system topology to expand to 16 cards or even more nodes, greatly improving the overall computing power, enhancing the resources of artificial intelligence servers, and thus improving the performance and capacity of artificial intelligence servers. Point-to-point communication between computing units can be achieved through Ethernet and RDMA (Remote Direct Memory Access). RMDA technology allows data to be directly transmitted between the memories of different computing nodes without going through the traditional network protocol stack. However, building a multi-card point-to-point interconnection topology based on Ethernet inevitably introduces additional switch nodes, which will cause additional time overhead in terms of performance, and will also have certain limitations in terms of congestion control, space, cost, scalability, etc.
[0021] As the number of computing units increases, the demand for low-cost and low-latency point-to-point communication between computing units is getting higher and higher. Related technologies will expand computing units, such as GPUs, through PCIe Switches for data exchange. PCIeSwitch can provide a more direct data transmission path without the need for complex network protocols or multi-layer network equipment forwarding, which to a certain extent reduces the delay overhead of data transmission, as well as the risk of congestion and bandwidth bottlenecks that may exist under high load conditions. In order to meet the frequent data exchange between computing units, it also has good communication efficiency. Based on this, related technologies use NTB in PCIe interconnection systems to achieve multi-host communication. Figure 1 As shown in the figure, the system includes two HOSTs, namely the local host HOST0 and the remote host HOST1. The local host and the remote host are connected to PCIe Switch respectively, and GPU heterogeneous communication software is integrated in the host respectively. The local host and the remote host each have a GPU, namely the local GPU and the remote GPU. The two GPUs exchange data with the host through DSP (Digital Signal Processing). The two PCIe Switches, namely PCIe Switch0 and PCIe Switch1, are connected through Crosslink. The connection port of PCIe Switch is set to NTB mode. The local GPU and the remote GPU are interconnected and communicated through NTB. The hosts are interconnected through PCIe Crosslink, and NTB is used to achieve cross-host memory access, and then the DMA inside PCIeSwitch is used to move data. The base address register space of the remote GPU is windowed in the local NTB, and the local GPU initiates access to the base address register space of the remote GPU through DMA, that is, by mapping the remote node address space to the local address space through NTB, the access path between the two GPUs is opened, thereby realizing P2P transmission.
[0022] However, this method is implemented through Virtual Switch, and the address domains at both ends are completely isolated, so it needs to rely on the complex logic of the NTB controller. In addition, the topology of cross-HOST communication requires the NTB ports of the two switches to be interconnected, and the cross-domain address mapping and interrupt mechanism will introduce additional delays and hardware overhead. The control logic of NTB has a cascade depth limit. As the number of PCIe Switch levels on the inter-GPU communication link increases, the delay overhead will also increase accordingly. Usually only 2-level Switch links are supported. Multi-level Switches will significantly reduce communication efficiency, which limits the scale of the system.
[0023] In view of this, in order to solve the problem of cascade depth limitation in realizing inter-GPU communication through PCIe non-transparent bridge, the present invention maps the resource information allocated by the switch to the computing units connected to each other with the resource information allocated by the host to all the computing units inside it, and there is routing information that enables communication between the switches. Through port forwarding, P2P communication can be realized between any computing nodes in the PCIe Switch interconnection system through a single hop or multiple hops, without the need for upstream forwarding, with good flexibility and scalability, and can improve the communication efficiency of the entire system.
[0024] Based on the technical solution of the present invention, in combination with the specific application environment architecture or specific hardware architecture on which the execution of the data communication method depends, the specific application environment architecture or specific hardware architecture is described here:
[0025] One of the application scenarios of the embodiment of the present invention is as follows: Figure 2 As shown in the figure, in this application scenario, the server acts as a host and has two CPUs, namely CPU0 and CPU1. One CPU can be used as a super node. Each CPU includes four root nodes, such as RC (root complex) 0, RC1, RC2 and RC3. Each switch is connected to the corresponding RC. The entire system includes eight switches, namely Switch01, Switch02, Switch03, Switch04, Switch05, Switch06, Switch07 and Switch08. The switches under the same CPU are connected, and the switches of different CPUs are connected to two switches at symmetrical positions, thus forming a fully interconnected data communication system including 32 GPUs. Each switch is connected to multiple GPUs, and it locally stores the resource information allocated to the connected GPUs, namely the switch resource allocation information, and also stores the communication routing information for data communication with other switches, namely the switch routing information. The server stores the resource information allocated to all internal GPUs, that is, the host resource allocation information, which can be shared by CPU0 and CPU1. In order to achieve point-to-point communication, a mapping relationship between the switch resource allocation information and the host resource allocation information, that is, a routing mapping relationship, is established.
[0026] Based on the above data communication system, when CPU0 and CPU1 access any GPU in the host, they will send down the host address and host bus number assigned to the GPU, and map the host resource allocation information to the switch resource allocation information corresponding to the GPU according to the host address, host bus number and routing mapping relationship. Then, according to the switch resource allocation information, the GPU can be accessed. The access process between GPUs in different switches of the same CPU can be as follows: the source GPU will send down an access request carrying at least the switch address and bus information of the GPU to be accessed, and according to the information and routing mapping relationship, determine the identification information of the target switch to which the GPU to be accessed belongs, and according to the switch routing information of the source switch to which the source computing unit belongs, determine the communication port for the source switch to communicate with the target switch, and then forward the access request to the target switch through the communication port, and the target switch can send the access request to the GPU to be accessed according to the switch address of the GPU to be accessed. The access process between GPUs located in different CPUs can be as follows: the source GPU will send an access request that carries at least the switch address and bus information of the GPU to be accessed, determine the identification information of the target switch to which the GPU to be accessed is connected based on this information and the routing mapping relationship, and determine the first communication port that communicates with the target switch based on the switch routing information of the source switch to which the source GPU belongs; the first communication port is the port of the intermediate switch connected to the source switch, the intermediate switch and the source switch belong to the same CPU, and the intermediate switch is connected to the target switch. The access request is forwarded to the intermediate switch through the first communication port, the second communication port that communicates with the target switch is determined based on the switch routing information of the intermediate switch, the access request is forwarded to the target switch through the second communication port, and the second target switch forwards the access request to the GPU to be accessed through the switch address of the GPU to be accessed, thereby realizing the shortest path point-to-point communication through a three-level access method.
[0027] It should be noted that the above application scenarios are only shown to facilitate understanding of the ideas and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention are described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0028] First see Figure 3 , Figure 3A flow chart of a data communication method provided in this embodiment is provided. The data communication method is applied to a data communication system. The data communication system in this embodiment includes at least one host, multiple computing nodes are deployed on the host, and each computing node is connected to multiple computing units through at least one switch, that is, each computing unit is deployed on the host, each computing unit in the same group is connected to the same switch, and each computing unit in different groups is connected to different switches. Each switch can be connected through an actual physical link or communicate through a predefined port. As long as data communication can be achieved between each switch, this embodiment may include the following contents:
[0029] S101: When a source computing unit receives an access request to a target computing unit, the source computing unit determines resource allocation information of a target switch to which the target computing unit belongs according to a routing mapping relationship and target resource allocation information of the target computing unit carried in the access request.
[0030] In this embodiment, the source computing unit is the computing unit that sends the access request, and the target computing unit is the computing unit to which the access request is to reach, that is, the computing unit that currently needs to be accessed. The source computing unit and the target computing unit belong to different computing nodes, that is, the access request in this embodiment refers to a computing unit of a computing node of a data communication system accessing a computing unit of another computing node, and the switches to which these two computing units belong are located on different computing nodes. The routing mapping relationship is the correspondence between the host resource allocation information and the switch resource allocation information of the same computing unit. The host resource allocation information is the resource information allocated by the host to the computing unit. The resource information may at least include a unique bus number and a unique address. The switch resource allocation information is the resource information allocated by the switch to the computing unit. The resource information may at least include a unique bus number and a unique address. For the same computing unit, the resource information allocated by the host to it and the resources allocated by the switch to it are mapped. No matter what kind of information is carried in the access request for identifying the computing unit, the switch connected to the computing unit can be uniquely and quickly located through address mapping. For the convenience of description, the switch connected to the target computing unit is defined as the target switch. After the switch is determined, the relevant information identifying the target switch can be obtained. For the convenience of description, this embodiment defines it as resource allocation information.
[0031] S102: Determine an intermediate forwarding port according to the resource allocation information and the switch routing information of the source switch corresponding to the source computing unit, forward the access request to the target communication port of the target switch through the intermediate forwarding port, and forward the access request to the target computing unit again through the target communication port and the target resource allocation information.
[0032] In this embodiment, the source switch is the switch to which the source computing unit is connected. The switch stores switch routing information. The switch routing information includes communication routing information between different switches. The communication routing information at least includes identification information for each switch and port information corresponding to the data communication path between different switches. Each switch in the data communication system stores not only the information allocated to the connected computing unit, but also the routing information for data communication with other switches. This embodiment defines this information as switch routing information. After the resource allocation information of the target switch is determined in the previous step, the path to the target switch can be determined by querying the switch routing information. Since the present embodiment is the access to computing units of different computing nodes, at least one port for forwarding the access request and a communication port for reaching the target switch from the port can be determined. For ease of description, the port for forwarding the access request to the target switch is defined as an intermediate forwarding port, and the access request of a computing node is sent to another computing node through the intermediate forwarding port. The port of the target switch receiving the access request is defined as a target communication port. The access request can be sent to the target switch through the target communication port. The target switch can locate the target computing unit through the target resource allocation information of the target computing unit, thereby forwarding the access request to the target computing unit to achieve access to the target computing unit.
[0033] The present invention can realize point-to-point communication between any computing units through address mapping and port forwarding. Compared with the method that only relies on the address mapping technology inside the Switch and the communication requests across switches need to use the upstream port, such as the CPU RC level, forwarding, it can not only realize the P2P communication between computing units belonging to the same Switch with the shortest route, but also reduce the communication delay because the communication requests across switches do not need the upstream port to assist in forwarding. Figure 4As shown, the process of forwarding through the host is, for example: GPU0 of Switch01 accesses GPU3 of Switch04, GPU0 of Switch01 initiates a communication request containing GPU3 of Switch04, and queries whether the identification information of the upstream switch, i.e., Switch01, is the same as the identification information of the target switch, i.e., Switch04. If they are the same, then according to the address assigned to the target computing unit in the routing table of Switch01 containing the address information of all computing units, routing forwarding is completed inside the Switch to establish communication. If the two are not the same, the switch address of the data packet is mapped to the host switch address, routed to the upstream port RC1 of the Switch, reported to the Host, the target Switch is located through the Host and the request is forwarded, and then forwarded to the target computing unit according to the routing table information in the target Switch. Forwarding through the upstream port will greatly increase the data transmission overhead, reduce the communication efficiency and the computing efficiency of the entire system.
[0034] In the technical solution provided in this embodiment, the resource information allocated by the switch to the computing units connected to each other is matched with the resource information allocated by the host to all the computing units inside it, and the information of the switch connected to the computing unit to be accessed can be quickly located through this address mapping. There is routing information that can communicate between each switch, so that the computing units connected to different switches can realize fast port forwarding through the switch routing information, thereby realizing point-to-point communication between any computing units through single hop or multi-hop, without the need for upstream port forwarding, and without introducing additional hardware overhead and address conversion delay, effectively improving the efficiency of point-to-point communication. Further, by increasing the number of switches, not only the scale of the computing units owned by the computing nodes can be increased, but also the number of computing nodes contained in the host can be expanded. On the basis of the expansion of the scale of the data communication system, the efficient routing path determination between any computing units can still be realized, thereby improving the point-to-point communication of large-scale data communication systems, and providing a more efficient, lower latency, flexible and scalable routing deployment and management method for expanding the PCIeSwitch system topology of multiple cards, thereby improving the communication efficiency of the entire system, and can also solve the problem of bandwidth limitation of parallel communication requirements between computing units and hosts, and computing units to a certain extent.
[0035] It should be noted that there is no strict order of execution between the steps in the present invention. As long as they comply with the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 3 This is just a schematic and does not mean that this is the only execution order.
[0036] In the above embodiment, there is no limitation on how to generate the switch resource allocation information. Based on the above embodiment, the present invention also provides an exemplary method for generating the switch resource allocation information, which may include the following contents:
[0037] A switch resource allocation table can be pre-constructed, and the switch resource allocation table includes at least switch identification information, bus identification information of the computing unit, and switch address of the computing unit. Each switch shares the switch resource allocation table, and each switch will enumerate all the computing units connected to it. When enumerating each computing unit, an identification information will be assigned to each computing unit, so that the switch bus identification information of each computing unit can be obtained. An address will also be assigned to each computing unit, so that the switch address of each computing unit can be obtained. To facilitate communication, each switch has unique switch identification information. The switch identification information, the bus identification information of each computing unit, and the switch address are filled into the corresponding position of the switch resource allocation table to generate the switch resource allocation information, that is, the switch resource allocation information is a switch resource allocation table filled with correct data, and the correctly filled switch resource allocation table is used as the switch resource allocation information and stored locally.
[0038] In order to make the process more clear to those skilled in the art, this embodiment also takes four switches Switch01, Switch02, Switch03 and Switch04, each switch includes four computing units, namely GPU00, GPU01, GPU02, GPU03, the computing unit is GPU, the switch identification information is GPU Switch ID, the computing unit identification information is GPUBUS, and the computing unit switch address is GPU Address as an example to illustrate the generation process of switch resource allocation information, wherein the switch resource allocation information can be expressed as GPU Switch Info: Before starting, the HOST will establish an address domain number for all its internal PCIeSwitch numbers, namely GPU Switch ID, each switch enumerates and allocates resources to the downstream GPU, and the BUS number and address of each Switch are independent, and the format is as follows: Switch01: #GPU Switch Info { GPU00: "GPU Switch ID" = "0x01"; "GPU BUS" = "0x01"; "GPU Address" = "0x1_0000_0000"; }, … GPU03: "GPU Switch ID" = "0x01"; "GPU BUS" = "0x04"; "GPU Address" = "0x4_0000_0000"; } } … Switch04: #GPU Switch Info { GPU00: "GPU Switch ID" = "0x04"; "GPU BUS" = "0x01"; "GPU Address" = "0x1_0000_0000"; }, … GPU03: "GPU Switch ID" = "0x04"; "GPU BUS" = "0x04"; "GPU Address" = "0x4_0000_0000"; } }.
[0039] Among them, GPU Switch ID is the number of each switch interconnected in the data communication system, GPU BUS is the BUS (bus) number enumerated by the switch for the GPU, and GPU Address is the address assigned by the switch to the GPU. For the convenience of description, it is defined as the switch address.
[0040] As can be seen from the above, this embodiment can quickly generate switch resource allocation information by filling in the template. By using the bus number and address as the information allocated by the switch, the switch can quickly and accurately locate the computing unit to which it is connected, which is conducive to improving data communication efficiency.
[0041] In the above embodiment, there is no limitation on how to generate the host resource allocation information. Based on the above embodiment, the present invention further provides an exemplary method for generating the host resource allocation information, which may include the following contents:
[0042] The host can pre-build a host resource allocation table, which includes at least the host bus identification information of the computing unit and the host address of the computing unit. During the startup process, the host can enumerate the internal computing units. When enumerating each computing unit, a bus identification information will be assigned to all the computing units it owns. For the convenience of description, it can be defined as the main line identification information, so that the main line bus identification information of each computing unit can be obtained. The host will also assign an address to each computing unit, so that the address of each computing unit can be obtained. For the convenience of description, it is defined as the host address. Finally, the host bus identification information and host address of each computing unit are filled into the corresponding position of the host resource allocation table, and the host resource allocation information can be generated, that is, the host resource allocation information is a host resource allocation table filled with correct data, and the correctly filled host resource allocation table is used as the host resource allocation information, and is stored locally for each CPU to share information.
[0043] In order to make the process more clear to those skilled in the art, this embodiment also takes four switches, each of which includes four computing units, and a total of 16 GPUs, namely GPU00, GPU01, GPU02, ..., GPU15, the computing unit is GPU, the computing unit bus identification information is GPU HOST BUS, and the host address is GPU HOST Address as an example to illustrate the generation process of the host resource allocation information, wherein the host resource allocation information can be expressed as GPU HOST Info: During the startup process, the HOST can enumerate and allocate resources to all GPU devices through the basic input and output system, and its format is as follows: #GPU HOST Info { GPU00: "GPU HOST BUS" = "0x50"; "GPU HOST Address" = "0x1_0000_0000"; }, GPU01: "GPU HOST BUS" = "0x51";
[0044] "GPU HOST Address" = "0x2_0000_0000"; }, … GPU15: "GPU HOST BUS" = "0x5F"; "GPU HOST Address" = "0x10_0000_0000"; }, }.
[0045] Among them, GPU HOST BUS is the BUS number enumerated by the host for each GPU, and GPU HOST Address is the address assigned by the HOST to the GPU, that is, the host address.
[0046] After obtaining GPU Switch Info and GPU HOST Info, the switch can establish a one-to-one correspondence between the host resource information and switch resource information of each GPU in order to perform route mapping. When the HOST initiates access to the GPU, it can use the BUS number and host address in the GPU HOST info, and convert the address into the address in the GPU Switch Info according to the route mapping relationship in the switch, so as to access the GPU. For example: HOST-GPU15: "GPU HOST BUS" = "0x5F"; "GPU HOST Address" = "0xF_0000_0000"; }, Maps to: Switch04-GPU03: "GPU Switch ID" = "0x04"; "GPU BUS" = "0x04"; "GPU Address" = "0x4_0000_0000"; }.
[0047] As can be seen from the above, this embodiment can quickly generate host resource allocation information by filling in the template. By using the bus number and address as the information allocated by the host, the host can quickly and accurately locate the computing unit to which it is connected, and directly communicate data with the connected computing unit. In the above embodiment, there is no limitation on how to generate switch routing information. Based on the above embodiment, the present invention also provides an exemplary method for generating switch routing information, which may include the following content:
[0048] The host includes a computing node, which includes multiple switches. The switch routing information stored in each switch is the actual communication route between the switch and other switches. If the host includes two computing nodes, when the switch of each computing node communicates with another switch, it can be forwarded through the intermediate forwarding port. The intermediate forwarding port can be pre-bound to the switch. For the convenience of description, it is defined as an intermediate switch, that is, forwarding through the intermediate switch, that is, an intermediate switch can be pre-specified for forwarding. The shortest path P2P communication between any computing units is realized through the communication path of the three-level switch cascade. In this implementation, for any switch, for the convenience of description, it is defined as the first switch, the first switch and the second switch belong to the same CPU, the third switch communication routing information and the fourth switch belong to another CPU, then the switch routing information of the first switch includes the second switch communication routing information, the third switch communication routing information and the fourth switch communication routing information; wherein the second switch communication routing information at least includes the second switch identification information and the port information for the first switch and the second switch to communicate; the third switch communication routing information at least includes the third switch identification information and the first intermediate conversion port information; the fourth switch communication routing information at least includes the fourth switch identification information and the second intermediate conversion port information; wherein the switch routing information of the first target intermediate switch to which the first intermediate port corresponding to the first intermediate conversion port information belongs at least includes the communication routing information between the third switch and the target intermediate switch; the switch routing information of the second target intermediate switch to which the second intermediate port corresponding to the second intermediate conversion port information belongs at least includes the communication routing information between the fourth switch and the second target intermediate switch.
[0049] In order to make the technical solution of the present invention more clearly understood by those skilled in the art, the present invention takes the switch port as the Fabric (port name) type, establishes switch routing information inside each switch of the data communication system, and the switch routing information includes other switches except itself. For the scenario of the same computing node, the computing node includes 4 switches, and the format of the switch routing information Switch Routing Info of the first switch can be as follows: Switch01: #Switch Routing Info { Switch02: "Switch ID" = "0x02"; "Egress Fabric Port" = "0x02"; }, … Switch04: "Switch ID" = "0x04"; "Egress Fabric Port" = "0x04"; }, } Switch02: #Switch Routing Info { Switch01: "Switch ID" = "0x01"; "Egress Fabric Port" = "0x01"; }, … Switch04: "Switch ID" = "0x04"; "Egress Fabric Port" = "0x04"; }, }.
[0050] In the scenario of two computing nodes, each computing node includes four switches, the middle switch is Switch04, and the format of the switch routing information Switch Routing Info of the first switch can be as follows: Switch01: #Switch Routing Info { Switch02: "Switch ID" = "0x02"; "Egress Fabric Port" = "0x02"; }, Switch03: "Switch ID" = "0x03"; "Egress Fabric Port" = "0x03"; }, Switch04: "Switch ID" = "0x04"; "Egress Fabric Port" = "0x04"; }, Switch05: "Switch ID" = "0x05"; "Egress Fabric Port" = "0x04"; }, Switch06: "Switch ID" = "0x06"; "Egress Fabric Port" = "0x04"; }, Switch07: "Switch ID" = "0x07"; "Egress Fabric Port" = "0x04"; }, Switch08: "Switch ID" = "0x08"; "Egress Fabric Port" = "0x04"; } }.
[0051] As can be seen from the above, for scenarios involving multiple computing nodes, this embodiment will route and forward communications across computing nodes through an intermediate switch, and implement the shortest path P2P communication between any computing units through a three-stage switch cascade communication path, effectively improving the efficiency of point-to-point communication between computing units across computing nodes.
[0052] It is understandable that each switch has limited ports. As the number of computing units included in the data communication system increases, if all switches are connected, it will not only cause difficulty in physical link wiring, but also have certain requirements on the physical space of the data communication system. In view of this, the present invention provides a variety of switch routing information determination methods based on different scenarios, which may include the following:
[0053] For a small-scale data communication system, or an application scenario that allows complex wiring, all switches of the data communication system can be interconnected, that is, switches of different groups of computing units are connected through a target bus, and the target bus can be, for example, a PCIe bus, that is, each switch is connected through a real physical link, and the first port and the second port on the physical link between the target switch and the source switch are obtained, and the first port and the source switch belong to the same computing node; the first port is used as an intermediate forwarding port, and the second port is used as a target communication port; and the access request is forwarded to the second port through the first port.
[0054] For a large-scale data communication system, such as a topology containing 16 computing units or more, there may be more than 4 switches in the system. Since the number of ports of a single Switch is limited, the physical link cannot achieve full interconnection between the switches. The communication port between each switch can be pre-specified. The communication port is used to specify the routing path when the switch communicates with the target Switch without being directly connected to the target Switch, that is, the switch is not connected through a real physical link, that is, the target bus. For the source switch and the target switch that are not connected through the target bus, the routing path for communicating with other switches in the communication system can be pre-configured for the switch. For the current access request, for the convenience of description, the routing path between the source switch and the target switch is defined as the target routing path. The target routing path includes at least a first port and a second port. The first port receives the request sent by the source switch, and the second port sends the request of the first port to the target switch. The switch routing exchange information is generated according to each routing path of the source switch. In this way, the routing path corresponding to the source switch and the target switch can be determined according to the switch routing information, and the access request is forwarded to the target computing unit through the routing path. That is, according to the switch routing information, it can be determined that the intermediate forwarding port is the first port, and the target communication port is the second port; the access request is forwarded to the second port through the first port. As can be seen from the above, this embodiment provides different switch routing information determination methods according to the scale of the data communication system, which neither increases the difficulty of physical link wiring nor is limited by the physical space of the data communication system. On the basis of improving the overall communication efficiency of the data communication system, it is also conducive to supporting the continuous expansion of the scale of the data communication system.
[0055] Based on the above embodiments, the present invention further provides various data access implementation methods of the data communication system in different application scenarios. The various data access methods are listed in parallel and may include the following contents:
[0056] As an exemplary application scenario, for the point-to-point communication of each computing unit within the same switch, each switch will form the switch allocation information of each computing unit connected to it into internal routing information. The internal routing information is shared by all ports of the switch. Data access between different computing units connected to the same switch can be directly communicated through the internal routing information, effectively improving the point-to-point communication efficiency of the data communication system.
[0057] As another exemplary application scenario, a data communication system includes a computing node CPU0, which includes 4 root nodes, all switches, such as Figure 5As shown, Switch01, Switch02, Switch03 and Switch04 are all connected to the same computing node CPU0, each switch is connected to at least one computing unit, and the computing units of the same switch are a group of computing units. Accordingly, each group of computing units is deployed on the first computing node.
[0058] In this scenario, an exemplary scenario is that the host accesses each computing unit, that is, the access request is issued by the host. Accordingly, the host resource allocation information of the target computing unit is defined as the target host resource allocation information. The target resource allocation information is part of the parameters or all of the parameters in the target host resource allocation information of the target computing unit. For example, the host resource allocation information includes the host bus identification information and the host address. The target resource allocation information can be the host bus identification information and the host address, or can only include the host address. After receiving the access request to the target computing unit, the target host resource allocation information is mapped to the target switch resource allocation information according to the target resource allocation information and the routing mapping relationship; according to the target switch resource allocation information, the access request is forwarded to the target computing unit through the target switch.
[0059] In this scenario, another exemplary scenario is mutual access between computing units of different switches. For ease of description, the computing unit that issues the access request is defined as the source computing unit, the switch to which the source computing unit is connected is the source switch, and the computing unit to which the access request arrives is defined as the target computing unit. The source computing unit and the target computing unit are not connected to the same switch. The switch resource allocation information of the target computing unit is defined as the target switch resource allocation information, and the switch to which the target computing unit is connected is the target switch. The source computing unit sends an access request, which carries target resource allocation information. The target resource allocation information is part or all of the parameters in the target switch resource allocation information of the target computing unit. For example, the switch resource allocation information includes the switch identification information, the computing unit bus identification information and the switch address. The target resource allocation information may include the computing unit bus identification information and the switch address. After receiving the access request to the target computing unit, the resource allocation information of the target switch to which the target computing unit belongs is determined based on the target switch resource allocation information and the routing mapping relationship. For example, it may be the unique identification information of the target switch. Based on the switch routing information of the source switch to which the source computing unit belongs, the target communication port for communication between the source switch and the target switch is determined. The access request is forwarded to the target switch through the target communication port, so that the target switch sends the access request to the target computing unit based on the target switch resource allocation information. Figure 5For example, when GPU00 of Switch01 initiates a communication request to GPU03 of Switch04, the GPU Address is used as the index in Switch01 to find the correspondence between GPU HOST Info and GPU Switch Info, and obtain the request routing target Switch ID = 0x04; according to Switch Routing Info of Switch01, the Fabric port 0x04 used for communication with Switch04 is obtained, and the request is forwarded to Switch04; Switch04 finds GPU Switch Info and routes to the specified GPU03 according to GPU Address, 0x4_0000_0000.
[0060] As another exemplary application scenario, the data communication system includes at least two computing nodes, that is, there are at least a first computing node and a second computing node, the first computing node and the second computing node are both deployed with at least one switch, each switch is connected to at least one computing unit, and the computing units of the same switch are a group of computing units, that is, the first computing node and the second computing node are both deployed with at least one group of computing units. Figure 6 As shown, CPU0 and CPU1, one CPU can be used as a super node, each CPU includes 4 root nodes, such as RC0, RC1, RC2 and RC3, each Switch is connected to the corresponding RC, and the entire system includes 8 switches, namely Switch01, Switch02, Switch03, Switch04, Switch05, Switch06, Switch07 and Switch08, each Switch under the same CPU is connected, and the Switches of different CPUs are connected to two Switches located in symmetrical positions, that is, Switch01 is connected to Switch07, Switch02 is connected to Switch08, Switch03 is connected to Switch05, and Switch04 is connected to Switch06. With the increase of computing nodes, the number of switches in the data communication system increases, and the intermediate forwarding port in the switch routing information of the above embodiment can adjust the position of the switch. In other words, for a large-scale data communication system, a switch can be selected from the data communication system based on experience or preset conditions, such as the one with the least running services and the best physical performance. For the sake of ease of description, it is defined as an intermediate switch, and a port is selected from each port of the intermediate switch as an intermediate forwarding port. As a simple implementation method, the port identification number of the port selected by the intermediate switch can be directly changed to the port identification information of the intermediate forwarding port in the switch routing information.
[0061] In this scenario, an exemplary scenario is that the first computing node accesses each computing unit inside the first computing node, or the first computing node accesses each computing unit inside the second computing node, and the access request is issued by the first node. Similarly, the host resource allocation information of the target computing unit is defined as the target host resource allocation information, and the target resource allocation information is part of the parameters or all of the parameters in the target host resource allocation information of the target computing unit. For example, the host resource allocation information includes the host bus identification information and the host address, and the target resource allocation information can be the host bus identification information and the host address. After receiving the access request to the target computing unit, according to the target resource allocation information of the target computing unit and the routing mapping relationship, the target host resource allocation information is mapped to the target switch resource allocation information; according to the target switch resource allocation information, the access request is forwarded to the target computing unit through the target switch.
[0062] In this scenario, another exemplary scenario is a computing unit of a first computing node, which is defined as a first target computing unit for ease of description. The first target computing unit accesses each computing unit of a second computing node, which is defined as a second target computing unit for ease of description. The access request is an access request for the first target computing unit of the first computing node to access the second target computing unit of the second computing node, and the target resource allocation information is part or all of the parameters in the second target host resource allocation information of the second target computing unit. For example, the switch resource allocation information includes switch identification information, computing unit bus identification information and switch address, and the second target host resource allocation information may include the computing unit bus identification information and switch address corresponding to the second target computing unit. An intermediate switch that meets a preset condition is determined among the switches of the first computing node, and a port of the intermediate switch is adjusted to the port identification information corresponding to the intermediate forwarding port in the switch routing information; after receiving an access request to a target computing unit, the resource allocation information of a second target switch to which the second target computing unit is connected is determined according to the resource allocation information of the second target switch and the routing mapping relationship; a first target communication port that communicates with the second target switch is determined according to the switch routing information of the first source switch to which the first target computing unit belongs; the first target communication port is a port of an intermediate switch connected to the first source switch, the intermediate switch belongs to the first computing node, and the intermediate switch is connected to the second target switch; the access request is forwarded to the intermediate switch through the first target communication port, and a second target communication port that communicates with the second target switch is determined through the switch routing information of the intermediate switch, and the access request is forwarded to the second target switch through the second target communication port, so that the second target switch forwards the access request to the target computing unit through the second target switch resource allocation information. Take the intermediate switch as Switch04. When GPU00 of Switch01 initiates a communication request to GPU03 of Switch08, the GPU Address is used as the index to find the correspondence between GPU HOST Info and GPU Switch Info in Switch01, and the target SwitchID=0x08 of the request route Switch08 is obtained. According to the Switch Routing Info of Switch01, the Fabric port 0x04 used by the intermediate switch Switch04 for communication is determined, and the request is forwarded to Switch04.The request data packet is routed to Switch04 through Fabric port 0x04 of Switch01. The target Switch is not matched. Switch04 determines the communication port 0x08 between it and Switch08 according to Switch Routing Info of Switch04, and sends the request data packet to Switch08 through communication port 0x08. Switch08 looks up GPU Switch Info according to GPUAddress and routes the request data packet to the specified GPU on it.
[0063] As can be seen from the above, this embodiment provides different point-to-point communication implementation methods for different application scenarios, allowing the shortest path P2P communication between any computing units in a single-host extended data communication system without the need for upstream port forwarding, and without introducing additional hardware overhead and address conversion delay. For point-to-point communication across computing nodes, a three-level switch cascade communication path is implemented by specifying an intermediate switch to achieve the shortest path P2P communication, thereby improving the communication efficiency of the entire data communication system. Furthermore, a computing node is used as a super node, such as Figure 6 As shown in the figure, the 16-card Switch interconnection system can be regarded as a super node, that is, all computing units of the same computing node are interconnected through the bus and appear to be a larger server. It can support a single host to be expanded into a large-scale computing system through the Switch, realize flexible expansion, and build a super node system of 32 cards or more.
[0064] It is understandable that, in the above-mentioned embodiment, when computing units across computing nodes perform point-to-point communication, an intermediate switch is pre-specified in the switch routing information, thereby determining an optimal communication path for point-to-point communication of computing units across computing nodes. Although this method can well avoid the problem of slow system response due to optimal path optimization when system resources are limited, it can also ensure that the entire system has a high communication efficiency to a certain extent, such as the intermediate switch has fewer tasks assigned, or the hardware resource configuration and its own performance are the best among all switches in the system. However, this method obviously needs to rely on human experience and requires subsequent business configuration. Once the intermediate switch is incorrectly specified or the switch fails or the switch has huge business pressure, this will lead to a decrease in the overall communication efficiency. Based on this, the present invention also provides another embodiment. This embodiment can store the switch routing information to the target register, and the target register is connected to a pre-specified user port, which may include the following contents:
[0065] When it is detected that the communication time of the computing units across the computing nodes exceeds the preset time threshold, the switch performance detection thread is started to detect whether the intermediate switch is faulty or whether the business data is excessive. When it is detected that the switch is faulty or the business pressure is high, the intermediate switch redetermination thread is started. The intermediate switch redetermination thread can call the optimal path optimization algorithm, and at the same time obtain the physical parameters and business operation data of other switches in the current data communication system, determine the new intermediate switch according to the optimal path optimization algorithm, and at the same time, call multiple register configuration threads through the pre-set user port to adjust the relevant data of the intermediate switch of the switch routing information of each switch.
[0066] From the above, it can be seen that this embodiment can ensure that the data communication system can always use the optimal switch to perform port forwarding tasks without affecting the user's business operation by real-time monitoring of the operation status of the intermediate switch, thereby ensuring the high efficiency of point-to-point communication of computing units across computing nodes.
[0067] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0068] The present invention also provides a corresponding device for the data communication method, which further makes the method more practical. The following description will specifically introduce the functions of each program module of this embodiment. The data communication device described below and the data communication method described above can be referred to in correspondence with each other. Based on the perspective of the functional module, see Figure 7 This embodiment provides a data communication device for a data communication system including multiple groups of computing units, where each computing unit in the same group is connected to the same switch, and the device may include:
[0069] The switch determination module 701 is used to determine the resource allocation information of the target switch to which the target computing unit belongs, based on the routing mapping relationship and the target resource allocation information of the target computing unit carried in the access request, when receiving an access request to the target computing unit, wherein the routing mapping relationship is the correspondence between the host resource allocation information and the switch resource allocation information of the same computing unit, the host resource allocation information is the resource information allocated by the host to the computing unit, and the switch resource allocation information is the resource information allocated by the switch to the computing unit.
[0070] The routing determination module 702 is used to determine the intermediate forwarding port according to the resource allocation information and the switch routing information of the source switch corresponding to the source computing unit, and forward the access request to the target communication port of the target switch through the intermediate forwarding port; forward the access request to the target computing unit again through the target communication port and the target resource allocation information; wherein the source computing unit and the target computing unit belong to different computing nodes; and the switch routing information includes port information corresponding to the data communication path between different switches.
[0071] Exemplarily, in some implementations of this embodiment, the above-mentioned routing determination module 702 can also be used for: for a host accessing a computing unit within the same computing node, the access request is issued by the host, and the target resource allocation information is part of the parameters or all of the parameters in the target host resource allocation information of the target computing unit. After receiving the host's access request to the target computing unit, the target host resource allocation information is mapped to the target switch resource allocation information according to the target resource allocation information and the routing mapping relationship; and the access request is forwarded to the target computing unit through the target switch according to the target switch resource allocation information.
[0072] Exemplarily, in some other implementations of this embodiment, the above-mentioned routing determination module 702 can also be used for: for access between different computing units in the same computing node, the access request is issued by the source computing unit, the target resource allocation information is part of the parameters or all of the parameters in the target switch resource allocation information of the target computing unit, the source computing unit and the target computing unit are not connected to the same switch, and the resource allocation information of the target switch to which the target computing unit belongs is determined according to the target switch resource allocation information and the routing mapping relationship; according to the switch routing information of the source switch to which the source computing unit belongs, the target communication port for communication between the source switch and the target switch is determined; and the access request is forwarded to the target switch through the target communication port, so that the target switch sends the access request to the target computing unit according to the target switch resource allocation information.
[0073] Illustratively, in some other implementations of this embodiment, the switch determination module 701 may also be used to: connect switches of different groups of computing units through a target bus, obtain a first port and a second port on a physical link between the target switch and the source switch, the first port and the source switch belong to the same computing node; use the first port as an intermediate forwarding port and the second port as a target communication port; and forward an access request to the second port through the first port.
[0074] Exemplarily, in some other implementations of this embodiment, the switch determination module 701 may also be used for: switches of different groups of computing units are not connected through a target bus, the source switch and the target switch are not connected through a target bus, a target routing path for communicating with the target switch is pre-configured for the source switch, the target routing path includes at least a first port and a second port, the first port receives a request sent by the source switch, and the second port sends the request of the first port to the target switch; switch routing information is generated according to the target routing path, and according to the switch routing information, the intermediate forwarding port is determined to be the first port, and the target communication port is determined to be the second port; and the access request is forwarded to the second port through the first port.
[0075] Exemplarily, in some other implementations of the present embodiment, the above-mentioned routing determination module 702 can also be used for: the data communication system has at least a first computing node and a second computing node, the first computing node and the second computing node are both deployed with at least one group of computing units, the access request is an access request of the first computing node to access the target computing unit of the second computing node, the target resource allocation information is part of the parameters or all of the parameters in the target host resource allocation information of the target computing unit, and according to the target resource allocation information of the target computing unit and the routing mapping relationship, the target host resource allocation information is mapped to the target switch resource allocation information; according to the target switch resource allocation information, the access request is forwarded to the target computing unit through the target switch.
[0076] Exemplarily, in some other implementations of this embodiment, the above-mentioned routing determination module 702 can also be used: the data communication system has at least a first computing node and a second computing node, the first computing node and the second computing node are both deployed with at least one group of computing units, the access request is an access request of a first target computing unit of the first computing node to access a second target computing unit of the second computing node, the target resource allocation information is part of the parameters or all of the parameters in the second target host resource allocation information of the second target computing unit, an intermediate switch that meets the preset conditions is determined in each switch of the first computing node, and a port of the intermediate switch is adjusted to the port identification information corresponding to the intermediate forwarding port in the switch routing information; according to the resource allocation information of the second target switch and the routing mapping relationship, determine resource allocation information of a second target switch connected to the second target computing unit; determining a first target communication port communicating with the second target switch according to switch routing information of a first source switch to which the first target computing unit belongs; the first target communication port is a port of an intermediate switch connected to the first source switch, the intermediate switch belongs to the first computing node, and the intermediate switch is connected to the second target switch; forwarding the access request to the intermediate switch through the first target communication port, determining a second target communication port communicating with the second target switch through the switch routing information of the intermediate switch, and forwarding the access request to the second target switch through the second target communication port, so that the second target switch forwards the access request to the target computing unit through the second target switch resource allocation information.
[0077] For the description of the features in the embodiment corresponding to the data communication device, reference can be made to the relevant description of the embodiment corresponding to the data communication method, which will not be repeated here.
[0078] The data communication device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. FIG9 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention under an implementation mode. The electronic device includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above data communication method embodiments.
[0079] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data communication method embodiments when running.
[0080] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0081] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above data communication method embodiments are implemented.
[0082] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data communication method embodiments are implemented.
[0083] Finally, the present invention also provides a data communication system, see Figure 8 The data communication system includes at least one host 800, which may include at least a first computing node 801, a routing controller 802, and a second computing node 803, and the host stores the host resource allocation information.
[0084] At least a first switch and a second switch are included; the first switch is connected to at least the first computing unit and the second computing unit, and the second switch is connected to at least the third computing unit and the fourth computing unit; the first switch and the second switch both store their respective switch resource allocation information and switch routing information. The second computing node 803 includes at least a third switch and a fourth switch; the third switch is connected to at least the fifth computing unit and the sixth computing unit, and the fourth switch is connected to at least the seventh computing unit and the eighth computing unit; the third switch and the fourth switch both store their respective switch resource allocation information and switch routing information. The routing controller 802 is used to implement the steps of the data communication method recorded in any of the above method embodiments when the first computing node 801 accesses each computing unit, or each computing unit accesses each other, and when executing a computer program.
[0085] Exemplarily, the above switch can generate switch resource allocation information and store it locally. For ease of description, this embodiment takes the first switch as an example to describe the process. The generation concept of switch resource allocation information of other switches is the same as that of the first switch. The first switch is also configured to: enumerate the connected computing units to obtain the bus identification information of each computing unit; assign addresses to each computing unit to obtain the switch address of each computing unit; fill the switch identification information of the first switch, the bus identification information of each computing unit and the switch address into the corresponding position of the switch resource allocation table, and store it locally as the switch resource allocation information; wherein the switch resource allocation table includes at least the switch identification information, the computing unit bus identification information and the computing unit switch address.
[0086] Exemplarily, the host is also configured to: during the startup process, enumerate each internal computing unit to obtain the main bus identification information of each computing unit; assign an address to each computing unit to obtain the host address of each computing unit; fill the host bus identification information and host address of each computing unit into the corresponding position of the host resource allocation table, and store it as the host resource allocation information; wherein the host resource allocation table includes at least the computing unit host bus identification information and the computing unit host address.
[0087] In order to achieve flexible expansion of the data communication system and build a large-scale data communication system, a computing node can be regarded as a super node to achieve flexible expansion and build a super node system. Based on this, the host also includes at least a second computing node; the first computing node 801 and the second computing node 803 are each a super node; the first computing node 801 includes at least a first root node and a second root node, the first root node is connected to the first switch, the second root node is connected to the second switch, and the first switch is connected to the second switch; the second computing node 802 includes at least a third root node and a third root node, the third root node is connected to the third switch, the fourth root node is connected to the fourth switch, the fourth switch is connected to the third switch, the third switch is connected to the first switch, and the second switch is connected to the fourth switch; the third switch is connected to at least the fifth computing unit and the sixth computing unit, and the fourth switch is connected to at least the seventh computing unit and the eighth computing unit; the third switch and the fourth switch both store their respective switch resource allocation information and switch routing information. Figure 6 For example, a 16-card Switch interconnection system can be regarded as a super node, which can be flexibly expanded to build a 32-card or higher super node system. On this basis, a single-machine expansion Switch fully interconnected 32-card system can also be regarded as an overall host node to further build a larger-scale super node system.
[0088] Exemplarily, the above-mentioned switch can generate switch routing information and store it locally. For the convenience of description, this embodiment takes the first switch as an example to describe the process, and the generation concept of switch routing information of other switches is the same as that of the first switch. The switch routing information of the first switch includes the communication routing information of the second switch, the communication routing information of the third switch, and the communication routing information of the fourth switch; wherein the communication routing information of the second switch at least includes the second switch identification information and the port information for the first switch to communicate with the second switch; the communication routing information of the third switch at least includes the third switch identification information and the first intermediate conversion port information; the communication routing information of the fourth switch at least includes the fourth switch identification information and the second intermediate conversion port information; wherein the switch routing information of the first target intermediate switch to which the first intermediate port corresponding to the first intermediate conversion port information belongs at least includes the communication routing information between the third switch and the target intermediate switch; the switch routing information of the second target intermediate switch to which the second intermediate port corresponding to the second intermediate conversion port information belongs at least includes the communication routing information between the fourth switch and the second target intermediate switch.
[0089] As can be seen from the above, the data communication system of this embodiment has efficient, low-latency, flexible and scalable routing management, allowing computing units to achieve P2P communication with the shortest path through a single hop or multiple hops, and can also be flexibly expanded to build a large-scale data communication system.
[0090] The above is a detailed introduction to a data communication method, system, electronic device, computer-readable storage medium and computer program product provided by the present invention. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can refer to each other. The units and algorithm steps of each example described in each disclosed embodiment are executed in electronic hardware or computer software depending on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods for each specific application to implement the described functions, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A data communication method, characterized in that: The invention is applied to a data communication system including at least two computing nodes, wherein at least two computing nodes include at least one group of computing units, and the same group of computing units is connected to the same switch, comprising: When the source computing unit receives an access request to the target computing unit, the resource allocation information of the target switch to which the target computing unit belongs is determined according to the routing mapping relationship and the target resource allocation information of the target computing unit carried in the access request; the source computing unit and the target computing unit belong to different computing nodes; an intermediate forwarding port is determined according to the resource allocation information and the switch routing information of the source switch corresponding to the source computing unit, and the access request is forwarded to the target communication port of the target switch through the intermediate forwarding port; forwarding the access request to the target computing unit again via the target communication port and the target resource allocation information; Among them, the routing mapping relationship is the correspondence between the host resource allocation information and the switch resource allocation information of the same computing unit, the host resource allocation information is the resource information allocated by the host to the computing unit, the switch resource allocation information is the resource information allocated by the switch to the computing unit, and the switch routing information includes port information corresponding to the data communication path between different switches.
2. The data communication method according to claim 1, characterized in that: When a host accesses a computing unit in the same computing node, the access request is issued by the host, and the target resource allocation information is part or all of the parameters in the target host resource allocation information of the target computing unit, and further includes: After receiving the access request of the host to the target computing unit, mapping the target host resource allocation information to the target switch resource allocation information according to the target resource allocation information and the routing mapping relationship; According to the target switch resource allocation information, the access request is forwarded to the target computing unit through the target switch.
3. The data communication method according to claim 1, characterized in that: For access between different computing units in the same computing node, the target resource allocation information is part or all of the parameters in the target switch resource allocation information of the target computing unit, the source computing unit and the target computing unit belong to the same computing node but are not connected to the same switch, and further includes: Determine the resource allocation information of the target switch to which the target computing unit belongs according to the target switch resource allocation information and the routing mapping relationship; Determine, according to switch routing information of a source switch to which the source computing unit belongs, a target communication port for communication between the source switch and the target switch; The access request is forwarded to the target switch through the target communication port, so that the target switch sends the access request to the target computing unit according to the target switch resource allocation information.
4. The data communication method according to claim 1, characterized in that: The switches of different groups of computing units are connected through a target bus, and the intermediate forwarding port is determined according to the resource allocation information and the switch routing information of the source switch corresponding to the source computing unit, and the access request is forwarded to the target communication port of the target switch through the intermediate forwarding port, including: Acquire a first port and a second port on a physical link between the target switch and the source switch, where the first port and the source switch belong to the same computing node; Using the first port as an intermediate forwarding port and using the second port as a target communication port; The access request is forwarded to the second port through the first port.
5. The data communication method according to claim 1, characterized in that: There are at least two switches that are not connected through a target bus, and determining an intermediate forwarding port according to the resource allocation information and switch routing information of a source switch corresponding to the source computing unit, and forwarding the access request to a target communication port of the target switch through the intermediate forwarding port, including: The source switch and the target switch are not connected via a target bus, a target routing path for communicating with the target switch is pre-configured for the source switch, the target routing path includes at least a first port and a second port, the first port receives a request sent by the source switch, and the second port sends the request of the first port to the target switch; Generate switch routing information according to the target routing path, and determine, according to the switch routing information, that the intermediate forwarding port is the first port and the target communication port is the second port; The access request is forwarded to the second port through the first port.
6. The data communication method according to any one of claims 1 to 5, characterized in that: The data communication system has at least a first computing node and a second computing node, the first computing node and the second computing node are both deployed with at least one group of computing units, the access request is an access request for the first computing node to access a target computing unit of the second computing node, the target resource allocation information is part of or all of the parameters in the target host resource allocation information of the target computing unit, and after receiving the access request to the target computing unit, the method further includes: According to the target resource allocation information and the routing mapping relationship of the target computing unit, the target host resource allocation information is mapped into the target switch resource allocation information; According to the target switch resource allocation information, the access request is forwarded to the target computing unit through the target switch.
7. The data communication method according to any one of claims 1 to 5, characterized in that: The data communication system has at least a first computing node and a second computing node, the first computing node and the second computing node are both deployed with at least one group of computing units, the access request is an access request of a first target computing unit of the first computing node to access a second target computing unit of the second computing node, the target resource allocation information is part of or all of the parameters in the second target host resource allocation information of the second target computing unit, and after receiving the access request to the target computing unit, the method further includes: Determine an intermediate switch that meets a preset condition among the switches of the first computing node, and adjust a port of the intermediate switch to the port identification information corresponding to the intermediate forwarding port in the switch routing information; Determine the resource allocation information of the second target switch to which the second target computing unit is connected according to the resource allocation information of the second target switch and the routing mapping relationship; Determine, according to switch routing information of a first source switch to which the first target computing unit belongs, a first target communication port for communicating with the second target switch; the first target communication port is a port of an intermediate switch connected to the first source switch, the intermediate switch belongs to the first computing node, and the intermediate switch is connected to the second target switch; The access request is forwarded to the intermediate switch through the first target communication port, and a second target communication port communicating with the second target switch is determined through the switch routing information of the intermediate switch, and the access request is forwarded to the second target switch through the second target communication port, so that the second target switch forwards the access request to the target computing unit through the second target switch resource allocation information.
8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data communication method according to any one of claims 1 to 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data communication method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the data communication method according to any one of claims 1 to 7 are implemented.
11. A data communication system, characterized in that: The method comprises a host having a first computing node, a second computing node and a routing controller, wherein the host stores host resource allocation information; The first computing node includes at least a first switch and a second switch; the first switch is connected to at least a first computing unit and a second computing unit, and the second switch is connected to at least a third computing unit and a fourth computing unit; the first switch and the second switch both store their own switch resource allocation information and switch routing information; The second computing node includes at least a third switch and a fourth switch; the third switch is connected to at least the fifth computing unit and the sixth computing unit, and the fourth switch is connected to at least the seventh computing unit and the eighth computing unit; the third switch and the fourth switch both store their respective switch resource allocation information and switch routing information; The routing controller is used to implement the steps of the data communication method according to any one of claims 1 to 7 when executing a computer program when the first computing node accesses each computing unit, or when the computing units access each other.
12. The data communication system according to claim 11, characterized in that: The first switch is further configured to: enumerate the connected computing units to obtain bus identification information of each computing unit; assign an address to each computing unit to obtain a switch address of each computing unit; fill the switch identification information of the first switch, the bus identification information of each computing unit and the switch address into the corresponding position of the switch resource allocation table, and store it locally as switch resource allocation information; The switch resource allocation table at least includes switch identification information, computing unit bus identification information and computing unit switch address.
13. The data communication system according to claim 11, characterized in that: The host is further configured to: during the startup process, enumerate each internal computing unit to obtain the main bus identification information of each computing unit; allocate an address to each computing unit to obtain the host address of each computing unit; fill the host bus identification information and host address of each computing unit into the corresponding position of the host resource allocation table, and store it as the host resource allocation information; The host resource allocation table at least includes computing unit host bus identification information and computing unit host address.
14. A communication system according to any one of claims 11 to 13, characterized in that: The first computing node and the second computing node each serve as a super node; The first computing node includes at least a first root node and a second root node, the first root node is connected to the first switch, the second root node is connected to the second switch, and the first switch is connected to the second switch; The second computing node includes at least a third root node and a third root node, the third root node is connected to the third switch, the fourth root node is connected to the fourth switch, the fourth switch is connected to the third switch, the third switch is connected to the first switch, and the second switch is connected to the fourth switch.
15. The data communication system according to claim 14, characterized in that: The switch routing information of the first switch includes second switch communication routing information, third switch communication routing information and fourth switch communication routing information; The communication routing information of the second switch includes at least the identification information of the second switch and the port information for the first switch to communicate with the second switch; the communication routing information of the third switch includes at least the identification information of the third switch and the first intermediate conversion port information; the communication routing information of the fourth switch includes at least the identification information of the fourth switch and the second intermediate conversion port information; Among them, the switch routing information of the first target intermediate switch to which the first intermediate port corresponding to the first intermediate conversion port information belongs at least includes the communication routing information between the third switch and the target intermediate switch; the switch routing information of the second target intermediate switch to which the second intermediate port corresponding to the second intermediate conversion port information belongs at least includes the communication routing information between the fourth switch and the second target intermediate switch.
Citation Information
Patent Citations
System and method of High-Speed Data Communication Frame For Cloud Gaming Data Storage and Retrieval
CN112422606A
Data processing method, device and system
CN113098773A
Resource allocation method and device, electronic equipment and computer readable storage medium
CN113452731A
Abnormality detection method and device, electronic equipment, distributed computing system and storage medium
CN118012680A
Host interconnection access method, system and device, computer equipment and storage medium
CN118519931A
Cited By
Node communication method and device, storage medium and electronic equipment
CN121210381A
Cooperative management system and method based on multi-card interconnection
CN122160348A