Cabinet-type server and communication method
By employing a rack-mount server architecture with converged nodes and switching nodes in the computer cluster, combined with orthogonal connections and various network topologies, the problem of limited switching capacity of switches is solved, realizing a high-performance, high-bandwidth, low-latency computing system, and improving the performance and scalability of the computer cluster.
Patent Information
- Application Number
- CN202411436253.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2022-08-30
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-08-30
AI Technical Summary
The limited switching capacity of switches in computer clusters leads to bandwidth limitations in computing chips, becoming a performance bottleneck and failing to meet the needs of bandwidth- and latency-sensitive applications such as high-performance computing and artificial intelligence.
The rack-mount server architecture employs multiple converged nodes and switching nodes. Through orthogonal connections and two-level data switching, combined with backplane-less orthogonal connectors and optical blind-mating connectors, it enables the close deployment of computing chips and switching chips, enhances the network topology, including dragonfly, ring, and fat tree networks, and increases the number of switching chips to meet bandwidth requirements.
It improves the computing performance of computer clusters, expands the scale of computer clusters, reduces cable connection error rate and loss, and enables high-bandwidth, low-latency communication in high-density computing systems.
Smart Images

Figure CN119520443B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202211049150.2 and the original application date is August 30, 2022. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computers, and more particularly to a rack-mounted server and a communication method thereon. Background Technology
[0003] Bandwidth and latency are two key metrics in network communication. For computing clusters consisting of multiple computing devices, bandwidth and latency determine the performance of the computing cluster system. Especially for bandwidth- and latency-sensitive applications such as high-performance computing (HPC) and artificial intelligence (AI), the various computing chips in the computing devices need to collaborate to complete computing tasks. There is a large amount of communication between these computing chips. Therefore, high-bandwidth, low-latency communication can effectively improve the processing performance of the computer cluster.
[0004] Typically, compute-intensive applications such as HPC and AI require multiple computing devices to work together on data processing. These devices communicate via switches. Due to the large number of computing devices, actual deployments usually require multiple switches, with long physical network cables connecting different devices to the switches. In practice, the more powerful the switch and the shorter the communication distance between devices, the easier it is to achieve high-bandwidth, low-latency communication in a computer cluster. Switch performance is often measured by its switching capacity; a higher switching capacity means greater data processing capability, but also higher design costs. Switch switching capacity, also known as backplane bandwidth or switching bandwidth, is the maximum amount of data that can be processed between the switch's interface processor or interface card and the data bus. Its unit is Gbps or Tbps, where Gbps is gigabits and Tbps is terabit-per-seconds.
[0005] However, due to the limited switching capacity of switches, the bandwidth they can provide to computing chips is also limited. For example, a switch with a switching capacity of 12.8Tbps can provide 64 200Gb ports. Assuming 32 ports are used for downstream connections, when one switch connects 128 computing chips, each chip can only be allocated 50Gb of bandwidth. For some powerful computing chips, 50Gb of bandwidth will limit their performance, becoming a bottleneck. Reducing the number of computing chips connected to the switch to achieve the required bandwidth will also decrease the overall parallel computing capability of the computer cluster. In fact, with the continuous development of HPC and AI technologies, a computer cluster with 128 computing chips is no longer sufficient. Therefore, with the increasing demand for computing chips, the limited bandwidth provided by switches creates a bottleneck in the computing power of the computer cluster. Summary of the Invention
[0006] This application provides a rack-mount server and a communication method to solve the problem of computing power bottlenecks in computer clusters.
[0007] In a first aspect, a rack-mount server is provided. The computing system includes multiple converged nodes and multiple switching nodes. Among the multiple converged nodes, a first converged node includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to realize communication connections between the multiple computing chips. Furthermore, among the multiple switching nodes, the first switching node is coupled to the first converged node through a connector. The first switching node is used to realize communication connections between the first converged node and other converged nodes among the multiple converged nodes.
[0008] The first switching node in a multi-node system is coupled to the first fusion node via a connector. Multiple computing chips within the first fusion node are connected to the first switching chip. The switching and computing chips are deployed close to each other. The multiple fusion nodes and switching nodes employ orthogonal connections and a two-level data switching mechanism to construct a high-performance, high-bandwidth, low-latency, high-density computing system. Furthermore, to enhance the computing power of the system, the number of first switching chips can be increased to meet the bandwidth demands resulting from the increased number of computing chips. This increases the available bandwidth of each computing chip, further improving the overall performance of the computing system.
[0009] Optionally, the computing system can be a high-density rack-and-rack integrated architecture. The rack housing the computing system not only houses converged nodes and switching nodes but also facilitates communication connections between different converged nodes and switching nodes. Specifically, the computing system includes rack-mount servers and cabinet servers. Rack-mount servers can include rack servers, blade servers, and high-density servers. Rack-mount servers typically have an open structure, allowing converged nodes and switching nodes within each computing system to be placed in the same rack. Blade servers and high-density servers offer higher computing density and save more hardware space compared to rack servers, but have lower storage capacity; therefore, the required server type can be selected based on the actual application scenario. Cabinet servers can be full-rack servers, and are fully enclosed or semi-enclosed structures. A single cabinet server can include one or more computing systems, including converged nodes and switching nodes.
[0010] In some embodiments, chassis servers can also be stored in rack servers; in other words, multiple chassis servers can be placed in a rack server.
[0011] In one possible implementation, the connection between the first fusion node and the first switching node is an orthogonal connection, and the connector includes a backplane-less orthogonal connector or an optical blind mating connector.
[0012] Optionally, backplane-less orthogonal connectors allow switching nodes and convergence nodes to be directly orthogonally connected without cables or optical fibers. This direct mating avoids the problems of high cable connection error rates and high cable or fiber loss caused by excessive cables or optical fibers. Optical blind-mating connectors achieve orthogonal connection through optical signals. These optical signals can be further enhanced through dense wavelength division multiplexing (DWDM) and high-density fiber optic connectors to increase the communication bandwidth between switching and convergence nodes. This further increases the bandwidth available to each computing chip, enabling the computer cluster to scale to a larger scale and improve its performance.
[0013] In one possible implementation, the connector includes a high-speed connector that achieves an orthogonal connection between the first fusion node and the first switching node by twisting a 90-degree angle.
[0014] Flexible connections between switching nodes, convergence nodes, and high-speed connectors are achieved using cables. The high-speed connectors are then twisted 90 degrees, making the switching nodes and convergence nodes almost directly orthogonal. This requires very little cable length. Because of the reduced cable length, the problems of high cable connection error rate and high cable or fiber loss caused by excessive cable or fiber length can be mitigated.
[0015] In one possible implementation, the network topology between multiple computing chips, at least one first switching chip, and multiple first switching nodes includes a network topology without a central switching node and a network topology with a central switching node.
[0016] Optionally, when the network topology is a network topology without a central switching node, the switching node is used to realize the communication connection between multiple computing systems; when the network topology is a network topology with a central switching node, the communication connection between multiple computing systems is realized through the central switching node.
[0017] Optionally, network topologies without a central switching node may include, but are not limited to, dragonfly (or dragonfly+) networks, torus networks, etc., while network topologies with a central switching node may include, but are not limited to, fat-tree networks.
[0018] Dragonfly comprises multiple groups, each essentially a sub-network enabling intra-group communication. Each group can be represented by the first rack mentioned above. Groups are connected via links, and these inter-group links are relatively narrow, resembling the wide body and narrow wings of a dragonfly, hence the name "Dragonfly Network." In other words, the aforementioned computing systems are connected via links, and data communication between the computing systems is achieved through switching nodes within each system. It should be understood that because the Dragonfly Network uses narrow links between its groups, it significantly reduces the number of global links, lowering networking costs.
[0019] Furthermore, the groups in a dragonfly network can be connected in an all-to-all manner. This type of network is also called a dragonfly+ network. All-to-all means that each group has at least one direct link to every other group in the dragonfly network. Dragonfly+ networks have larger group sizes and support more connected groups. Within a group, a connection method with as few hops as possible is used, usually an all-to-all or flat butterfly structure. This can further shorten the link length, thereby providing low-latency communication.
[0020] A Torus network consists of multiple nodes, each of which can be the aforementioned computing system. There are warap around links (or toroidal boundaries) between each node. The existence of warap around links means that there is no longer a distinction between nodes located at the topology center or edge in the Torus network. Each node has two adjacent nodes in each dimension, which allows for multiple forwarding paths between any two nodes. The network has high reliability, switching capacity, and scalability.
[0021] In specific implementations, in the above-mentioned network topology without a central switching node, namely the dragonfly network, the dragonfly+ network, and the torus network, the forwarding of data packets between computing systems needs to go through the switching nodes in each computing system. The forwarding path of the data packets may pass through one or more switching nodes, which is not specifically limited in this application.
[0022] A fat tree network is a large-scale, non-blocking network constructed using numerous low-performance switches and multi-layered networks. Typically, a fat tree network includes multiple clusters (pods) and multiple core switches connecting the pods. Each pod includes multiple Layer 1 switches and multiple Layer 2 switches. Each Layer 1 switch connects to one or more compute chips, handling data communication between these chips. Each Layer 2 switch connects to one or more Layer 1 switches, handling data communication between Layer 1 switches. Each core switch connects to one or more pods, handling data communication between them. Furthermore, core switches can form multiple sub-networks, with data communication between each sub-network handled by the layer above it, and so on. This allows for multiple parallel paths between any two compute nodes in the fat tree network, resulting in good fault tolerance and efficient traffic distribution within pods to prevent overload. It should be understood that fat tree network is a non-blocking network technology that uses a large number of switching chips to build a large-scale non-blocking network. The bandwidth of this network does not converge from bottom to top, and there are many data communication paths. There is always a path that can make the communication bandwidth reach the required bandwidth, which further solves the problem that the performance of computing chips is limited by network bandwidth and improves the performance of computer clusters.
[0023] In practice, in the network topology with a central switching node, i.e., in a fat tree network, each computing system can establish a communication connection with one or more central switching nodes, and the forwarding of data packets between computing systems needs to be achieved through one or more of these central switching nodes.
[0024] It should be understood that the network topology composed of multiple converged nodes and multiple switching nodes can also include other commonly used large-scale network topologies that are sensitive to bandwidth and latency. These will not be listed here, and this application does not impose any specific limitations.
[0025] The above implementation method allows multiple fusion nodes and multiple switching nodes in the computing system provided by this application to select a suitable network topology to establish communication connections between fusion nodes and switching nodes according to the business needs in the actual application scenario. This makes the computing system provided by this application simple to implement, highly feasible, and applicable to a wide range of scenarios.
[0026] In one possible implementation, the number of multiple converged nodes and multiple switching nodes in the computing system is determined based on at least one of the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
[0027] Optionally, the converged node can be deployed in the first box and the switching node can be deployed in the second box. The box can refer to a chassis or other box that has the function of placing and fixing accessories, supporting and protecting the various accessories inside the chassis. The box may include the outer shell, bracket, various switches, indicator lights, etc. on the panel. The box can be made of a combination of steel plate and plastic. This application does not make specific limitations.
[0028] Furthermore, after determining the number of fusion nodes and switching nodes in the computing system, the number of second computing chips in each switching node and the number of computing chips and first switching chips in each fusion node can be determined by considering various environmental conditions in the actual application scenario, such as the dimensions of the first and second enclosures, the number of enclosures that can be installed in the computing system, and the parallel processing capabilities required by the user. The number of enclosures that can be installed in the computing system refers to the number of enclosures that can be placed in the rack or cabinet where the computing system is located.
[0029] The above implementation determines the number of converged nodes and switching nodes in the computing system based on the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node. The network scale formed in this way can not only guarantee the available bandwidth of each computing chip and the scale of the entire computing system, but also have multiple parallel paths between each computing chip. Therefore, the network has good fault tolerance performance and can reasonably distribute traffic within the pod to avoid overload problems.
[0030] In one possible implementation, at least one first switching chip includes a first switching chip of a first switching plane and a first switching chip of a second switching plane, the first switching node includes a plurality of second switching chips, the plurality of second switching chips include a second switching chip of a first switching plane and a second switching chip of a second switching plane, and the first switching plane and the second switching plane undertake different services.
[0031] In simple terms, the network composed of the first and second switching chips within the computing system can include multiple switching planes. A portion of the first and second switching chips are responsible for data communication on the first switching plane, while another portion of the first and second switching chips are responsible for data communication on the second switching plane. It should be understood that the aforementioned first and second switching planes do not imply that the technical solution of this application only supports two switching planes, but rather that the technical solution of this application supports different switching planes. In practice, it can support more than two switching planes, and the specific number of switching planes can be determined according to the actual application scenario.
[0032] Optionally, the first and second switching planes can handle different services. For example, user service 1 uses the first switching plane, while service 2 uses the second switching plane, thus avoiding network latency caused by service conflicts. By using multiple switching planes, the computer cluster can provide users with greater bandwidth capabilities, improving the user experience.
[0033] Optionally, the first and second switching planes can be network-isolated switching planes. Specifically, for business security considerations, to achieve network isolation between user service 1 and service 2 and avoid network security issues, a layer 2 or higher fat-tree network can be formed using the cabinet provided in this application, and a corresponding switching plane can be assigned to each switching chip. This allows the user to use the first switching plane for data communication of service 1 and the second switching plane for data communication of service 2. No data communication occurs between the two switching planes; that is, the first and second switching chips of the first switching plane will not process data packets from the second switching plane, and vice versa, thereby achieving network isolation between the first and second switching planes. It should be understood that the above examples are for illustrative purposes only and this application does not impose specific limitations.
[0034] Optionally, the first switching plane and the second switching plane use different communication protocols. For example, service 1 requires communication protocol 1, and service 2 requires communication protocol 2. Using the cabinet provided in this application, and allocating corresponding switching planes to each switching chip, the first switching plane uses communication protocol 1 to process service 1, and the second switching plane uses communication protocol 2 to process service 2. It should be understood that the above examples are for illustration only, and this application does not make any specific limitations.
[0035] Understandably, computing chips can connect to first switching chips on different switching planes according to business needs. There is no communication connection between the first switching chips on different switching planes, nor is there a communication connection between the second switching chips on different switching planes. This can achieve network isolation between different switching planes, allowing different switching planes to handle different services and use different communication protocols, thus better meeting the diverse needs of users and improving the user experience.
[0036] Optionally, depending on user needs, the computing system may include more switching planes, such as a third switching plane, a fourth switching plane, etc. Each switching plane may also include more first switching chips and second switching chips. This application does not make specific limitations.
[0037] In the above implementation, multiple first switching chips and multiple second switching chips in the computing system can be assigned their respective corresponding switching planes, enabling the computing system to handle network communication across multiple switching planes. This satisfies users' network isolation requirements, multiple business processing requirements, and requirements for various communication protocols, making the computing system provided in this application applicable to a wider range of scenarios and improving the user experience.
[0038] In one possible implementation, after the computing chip, the first switching chip, and the second switching chip establish a communication connection, the first switching chip and the second switching chip can generate a routing table or a MAC address table according to their respective connection relationships through a forwarding algorithm. The routing table or MAC address table may include multiple entries, each entry may represent a forwarding path, and each entry may include at least a source address, a destination address, and a corresponding next-hop address. After receiving a data packet sent by the computing chip, the first switching chip can query the routing table or MAC address table based on the source address and destination address carried in the data packet to obtain the forwarding path of the data packet, determine the next-hop address, and then forward the data packet to the next-hop address.
[0039] If the first or second switching chip is a Layer 2 switch (link layer switch), a MAC address table can be generated using switch forwarding algorithms such as ARP. The source address and destination address can be MAC addresses. If the first or second switching chip is a Layer 3 switch (network layer switch), a routing table can be generated using routing algorithms such as RIP and BGP. The source address and destination address can be IP addresses. This application does not impose any specific limitations on this.
[0040] The forwarding path recorded in the routing table or MAC address table is determined based on the connection relationship between the first switching chip and the computing chip.
[0041] Optionally, in each fusion node, each first switching chip can establish a communication connection with all computing chips. Communication between computing chips within each fusion node can be achieved through any first switching chip 120 within the same fusion node. The forwarding path of data packets between computing chips within each fusion node can include any first switching chip 120. Each switching node establishes a communication connection with all fusion nodes through a connector, enabling data communication between each fusion node through any second switching chip. The forwarding path of data packets between each fusion node can include any second switching chip.
[0042] Optionally, each first switching chip can establish a communication connection with some computing chips. Data communication between multiple computing chips within the converged node can determine the forwarding path based on the first switching chip connected to it. If the sender and receiver are connected to the same first switching chip, the first switching chip can realize data communication between the first switching chip and some computing chips through a routing table. The forwarding path of the data packet may include the source address, the address of the same first switching chip, and the destination address.
[0043] If the sender and receiver are connected to different first switching chips and there is no direct connection between the two first switching chips, then their data communication can be achieved through a second switching chip connected to their respective first switching chips. The forwarding path of the message may include the source address, the address of the first switching chip 1, the address of the second switching chip 1, the address of the first switching chip 2, and the destination address. The first switching chip 1 is connected to the computing chip where the source address is located, the first switching chip 2 is connected to the computing chip where the destination address is located, and the second switching chip 1 is connected to both the first switching chip 1 and the first switching chip 2.
[0044] If the sender and receiver are connected to different first switching chips, and there is a direct connection between the two first switching chips, then data communication can be achieved through the two first switching chips connected to the sender and receiver. The forwarding path of the message may include the source address, first switching chip 1, first switching chip 2, and destination address.
[0045] In simple terms, if each computing chip is connected to only one first switching chip, the data packets generated by the computing chip can be directly sent to the first switching chip connected to it. If each computing chip is connected to multiple first switching chips, after generating a data packet, the computing chip can forward the data packet to the first switching chip with better network conditions based on the network conditions reported by at least one of the first switching chips connected to it. Alternatively, the computing system may also include a management node, which can determine the first switching chip to forward the data packet based on load balancing or other network management algorithms. This application does not make specific limitations. Alternatively, the computing chip can also select the first switching chip connected to it that can process the data packet based on the identifier carried in the data packet. For example, the identifier can be the identifier of the first forwarding plane, then only the first switching chip of the first forwarding plane can forward the data packet. This application does not make specific limitations.
[0046] The above implementation, without affecting data packet transmission, allows the system to increase the bandwidth available to each computing chip by reducing the number of computing chips connected to the first switching chip, and expand the scale of the computer cluster by increasing the number of the first switching chips, thereby solving the performance bottleneck problem of the computer cluster.
[0047] In one possible implementation, if the number of switching nodes is increased, and the converged nodes are placed horizontally and vertically, the number of switching nodes can be increased horizontally, and the length of the first housing containing the converged nodes can be adaptively increased. This allows the expanded first housing and a larger number of switching nodes to be connected via backplane-less orthogonal connectors. Similarly, if the number of converged nodes is increased vertically, the number of converged nodes can be increased vertically, and the length of the second housing containing the switching nodes can be adaptively increased. This allows the expanded second housing and a larger number of converged nodes to be connected via backplane-less orthogonal connectors.
[0048] Of course, if the merging nodes are placed vertically and the swapping nodes are placed horizontally, then the number of swapping nodes can be increased in the vertical direction and the number of merging nodes can be increased in the horizontal direction. This will not be elaborated on here.
[0049] In the above implementation, the fusion nodes and switching nodes are orthogonally connected. When the scale of the computing system expands, it is only necessary to expand the number of switching nodes or fusion nodes horizontally or vertically, which makes the computing system provided by this application highly scalable and feasible.
[0050] In one possible implementation, the computing system may include one or more symmetric multi-processing (SMP) systems, and each fusion node may also include one or more SMP systems. An SMP system refers to a set of processors aggregated on a server, which includes multiple CPUs. These processors share memory and other resources on the server, such as a shared bus architecture, allowing workloads to be evenly distributed across all available processors. That is, one SMP system corresponds to one OS domain and multiple computing chips. These computing chips can be within the same fusion node or across different fusion nodes; this application does not specifically limit this.
[0051] For example, a computing system may include 16 fusion nodes. These 16 fusion nodes could include 16 SMP systems, where one fusion node corresponds to one SMP system. Alternatively, the 16 fusion nodes could include 8 SMP systems, where two fusion nodes correspond to one SMP system. Or, the 16 fusion nodes could be fusion nodes 1 through fusion nodes 16, where fusion node 1 includes two SMP systems, fusion nodes 2 and 3 form one SMP system, and fusion nodes 4 through 16 form another SMP system. These examples are for illustrative purposes only and do not constitute a specific limitation in this application.
[0052] In the above implementation, since the computing chips are interconnected through the first and second switching nodes, the computing chips in each SMP system can be flexibly combined according to the business needs of the actual application scenario. This allows the computing system provided by this application to meet the user's needs for multiple SMP systems and improve the user experience.
[0053] Secondly, a communication method is provided, which is applied to a computing system. The computing system includes multiple fusion nodes and multiple switching nodes. A first switching node is coupled to a first fusion node via a connector. The first switching node is used to realize communication connections between the first fusion node and other fusion nodes among the multiple fusion nodes. The first fusion node includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to realize communication connections between the multiple computing chips. The method includes the following steps: a first computing chip of the first fusion node generates a data packet, wherein the destination address of the data packet is the address of a second computing chip; the first fusion node forwards the data packet according to the address of the fusion node where the second computing chip is located.
[0054] The first switching node in a multi-switching node is coupled to the first converged node via a connector. Multiple computing chips in the first converged node are connected to the first switching chip. In this way, even if the number of computing chips connected to the first switching chip decreases, the number of computing chips in the entire computing system can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth that can be allocated to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0055] In one possible implementation, the specific steps for the first fusion node to forward data packets based on the address of the fusion node where the second computing chip is located can be as follows: when the second computing chip is a computing chip within the first fusion node, the first fusion node forwards the data packets to the destination address through at least one first switching chip; when the second computing chip is a computing chip within the second fusion node, the first fusion node sends data packets to the first switching node among multiple switching nodes, and the first switching node sends data packets to the second fusion node.
[0056] In one possible implementation, the connection between multiple convergence nodes and multiple switching nodes is an orthogonal connection, and the connectors include backplane-less orthogonal connectors or optical blind mating connectors.
[0057] In one possible implementation, the connector includes a high-speed connector that enables orthogonal connections between multiple fusion nodes and multiple switching nodes by twisting a 90-degree angle.
[0058] In one possible implementation, the network topology between multiple converged nodes and multiple switching nodes includes a network topology without a central switching node and a network topology with a central switching node.
[0059] In one possible implementation, when the network topology is a network topology without a central switching node, the communication connection between multiple computing systems is achieved through the switching node; when the network topology is a network topology with a central switching node, the communication connection between multiple computing systems is achieved through the central switching node.
[0060] In one possible implementation, the number of converged nodes and switching nodes in the computing system is determined based on the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
[0061] In one possible implementation, at least one first switching chip includes a first switching chip of a first switching plane and a first switching chip of a second switching plane, the first switching node includes a plurality of second switching chips, the plurality of second switching chips include a second switching chip of a first switching plane and a second switching chip of a second switching plane, and the first switching plane and the second switching plane undertake different services.
[0062] In one possible implementation, the first and second switching planes are network-isolated switching planes, or the first and second switching planes use different communication protocols.
[0063] Thirdly, a converged node is provided, which can be applied to a computing system. The computing system includes multiple converged nodes and multiple switching nodes. A first switching node among the multiple switching nodes is coupled to a first converged node via a connector. The first switching node is used to realize communication connections between the first converged node and other converged nodes among the multiple converged nodes. The first converged node among the multiple converged nodes includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to realize communication connections between the multiple computing chips. The converged node includes a computing unit and a first switching unit. The computing unit is used to generate data packets. The source address of the data packet is the address of the first computing chip in the converged node, and the destination address is the address of the second computing chip. The first switching unit is used to forward the data packets according to the address of the converged node where the second computing chip is located.
[0064] The first switching node in a multi-switching node is coupled to the first converged node via a connector. Multiple computing chips in the first converged node are connected to the first switching chip. In this way, even if the number of computing chips connected to the first switching chip decreases, the number of computing chips in the entire computing system can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth that can be allocated to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0065] In one possible implementation, the first switching unit is used to forward data packets to a destination address through at least one first switching chip when the second computing chip is a computing chip within the first fusion node, and to send data packets to a first switching node among multiple switching nodes when the second computing chip is a computing chip within the second fusion node, and the first switching node sends data packets to the second fusion node.
[0066] In one possible implementation, the connection between multiple convergence nodes and multiple switching nodes is an orthogonal connection, and the connectors include backplane-less orthogonal connectors or optical blind mating connectors.
[0067] In one possible implementation, the connector includes a high-speed connector that enables orthogonal connections between multiple fusion nodes and multiple switching nodes by twisting a 90-degree angle.
[0068] In one possible implementation, the network topology between multiple converged nodes and multiple switching nodes includes a network topology without a central switching node and a network topology with a central switching node.
[0069] In one possible implementation, when the network topology is a network topology without a central switching node, the communication connection between multiple computing systems is achieved through the switching node; when the network topology is a network topology with a central switching node, the communication connection between multiple computing systems is achieved through the central switching node.
[0070] In one possible implementation, the number of converged nodes and switching nodes in the computing system is determined based on the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
[0071] In one possible implementation, at least one first switching chip includes a first switching chip of a first switching plane and a first switching chip of a second switching plane, the first switching node includes a plurality of second switching chips, the plurality of second switching chips include a second switching chip of a first switching plane and a second switching chip of a second switching plane, and the first switching plane and the second switching plane undertake different services.
[0072] In one possible implementation, the first and second switching planes are network-isolated switching planes, or the first and second switching planes use different communication protocols.
[0073] Fourthly, a switching node is provided, which can be applied to a computing system. The computing system includes multiple converged nodes and multiple switching nodes. A first switching node is coupled to a first converged node via a connector. The first switching node is used to realize communication connections between the first converged node and other converged nodes among the multiple converged nodes. The first converged node includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to realize communication connections between the multiple computing chips. The switching node includes a second switching unit, which is used to receive data packets sent by the first converged node. The source address of the data packet is the address of the first computing chip in the first converged node, and the destination address is the address of the second computing chip in the second converged node. The second switching unit is also used to forward the data packet to the second converged node according to the destination address.
[0074] The first switching node in a multi-switching node is coupled to the first converged node via a connector. Multiple computing chips in the first converged node are connected to the first switching chip. In this way, even if the number of computing chips connected to the first switching chip decreases, the number of computing chips in the entire computing system can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth that can be allocated to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0075] In one possible implementation, the connection between multiple convergence nodes and multiple switching nodes is an orthogonal connection, and the connectors include backplane-less orthogonal connectors or optical blind mating connectors.
[0076] In one possible implementation, the connector includes a high-speed connector that enables orthogonal connections between multiple fusion nodes and multiple switching nodes by twisting a 90-degree angle.
[0077] In one possible implementation, the network topology between multiple converged nodes and multiple switching nodes includes a network topology without a central switching node and a network topology with a central switching node.
[0078] In one possible implementation, when the network topology is a network topology without a central switching node, the communication connection between multiple computing systems is achieved through the switching node; when the network topology is a network topology with a central switching node, the communication connection between multiple computing systems is achieved through the central switching node.
[0079] In one possible implementation, the number of converged nodes and switching nodes in the computing system is determined based on the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
[0080] In one possible implementation, at least one first switching chip includes a first switching chip of a first switching plane and a first switching chip of a second switching plane, the first switching node includes a plurality of second switching chips, the plurality of second switching chips include a second switching chip of a first switching plane and a second switching chip of a second switching plane, and the first switching plane and the second switching plane undertake different services.
[0081] In one possible implementation, the first and second switching planes are network-isolated switching planes, or the first and second switching planes use different communication protocols.
[0082] Fifthly, a computing system is provided, the computing system comprising modules for performing communication methods in any possible implementation of the above aspects or aspects.
[0083] In a sixth aspect, a computing device is provided, the computing device including a processor and a memory, the memory for storing code, and the processor for executing the code to implement the functions of the operation steps performed by the fusion node as described in the second aspect.
[0084] In a seventh aspect, a computing device is provided, the computing device including a processor and a power supply circuit for supplying power to the processor, the processor being configured to perform the operational steps as described in the fusion section of the second aspect.
[0085] Eighthly, a communication device is provided, the communication device including a processor and a memory, the memory for storing code, and the processor for executing the code to implement the functions of the operation steps performed by the switching node as described in the second aspect.
[0086] In a ninth aspect, a computer cluster is provided, comprising multiple computing systems, which communicate with each other to collaboratively process tasks. Each computing system in the cluster can be a computing system described in the first aspect. The multiple computing systems can establish communication connections with a central switching node (such as the second rack mentioned above). Each central switching node is used to realize the communication connections between the computing systems. The network topology between the multiple computing systems and the multiple central switching nodes can be a network topology with a central switching node as described above, such as a fat tree network.
[0087] In a tenth aspect, a computer cluster is provided, comprising multiple computing systems that communicate with each other and collaboratively process tasks. Each computing system in the cluster can be the computing system described in the first aspect. The communication connection between each computing system is implemented through a switching node within each computing system. The network topology between the multiple computing systems can be a network topology without a central switching node as described above, such as a Dragonfly network, a Dragonfly+ network, or a Torus network, etc.
[0088] Eleventhly, a computer storage medium is provided, which stores instructions that, when executed on a computer, cause the computer to perform the methods described in the above aspects.
[0089] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0090] Figure 1 This is a schematic diagram of the structure of a computing system provided in this application;
[0091] Figure 2 This is a schematic diagram of the structure of a computing system provided in this application scenario;
[0092] Figure 3 This is a schematic diagram illustrating the connection relationship between a fusion node and a switching node provided in this application;
[0093] Figure 4 This is a schematic diagram of a four-layer fat tree network provided in this application;
[0094] Figure 5 This is a structural diagram of a cabinet provided in this application in another application scenario;
[0095] Figure 6 This is a structural diagram of a cabinet provided in this application in another application scenario;
[0096] Figure 7 This is a flowchart illustrating the steps of a communication method provided in this application scenario.
[0097] Figure 8 This is a flowchart illustrating the steps of a communication method provided in this application in another application scenario;
[0098] Figure 9 This is a schematic diagram of the structure of a computing system provided in this application;
[0099] Figure 10 This is a schematic diagram of the structure of a computing system provided in this application. Detailed Implementation
[0100] To address the performance bottleneck issue in computer clusters caused by the limited switching capacity of the aforementioned switching nodes, this application provides a computing system including multiple converged nodes and multiple switching nodes. A first converged node among the multiple converged nodes includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to establish communication connections between the multiple computing chips. The first switching node is used to establish communication connections between the first converged node and other converged nodes among the multiple converged nodes. The first switching node is coupled to the first converged node via a connector, and the multiple computing chips in the first converged node are connected to the first switching chip. The coupling indicates that the first switching node is directly connected to the first converged node via the connector, or that the coupling indicates that the first switching node is indirectly connected to the first converged node via the connector.
[0101] In this way, even if the number of computing chips connected to the first switching chip decreases, the number of computing chips in the entire computing system can be increased by increasing the number of the first switching chips. Conversely, by reducing the number of computing chips connected to each first switching chip, the bandwidth that can be allocated to each computing chip can be increased. This can increase the available bandwidth of each computing chip while expanding the scale of the computer cluster, thereby solving the performance bottleneck problem of the computer cluster.
[0102] It's important to note that a computing system can be a high-density rack-and-rack integrated architecture. The rack housing the computing system not only houses converged nodes and switching nodes but also facilitates communication connections between different converged nodes and switching nodes. Specifically, computing systems include rack-mount servers and cabinet servers. Rack-mount servers can include rack servers, blade servers, and high-density servers. Rack-mount servers typically have an open structure, allowing converged nodes and switching nodes within each computing system to reside within the same rack. Blade servers and high-density servers offer higher computing density and save more hardware space compared to rack servers, but have lower storage capacity. Therefore, the appropriate server type can be selected based on the specific application scenario. Cabinet servers can be entire racks of servers, featuring a fully enclosed or semi-enclosed structure. A single cabinet server can include one or more computing systems, including converged nodes and switching nodes.
[0103] Alternatively, chassis servers can also be housed in rack servers; in other words, multiple chassis servers can be placed in a single rack server.
[0104] To make this application easier to understand, the following description will use a rack-mounted server as an example, where the converged nodes and switching nodes in the computing system are placed in the same rack.
[0105] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of a computing system 400 provided in this application. This computing system is any server in a computer cluster, specifically a rack-mount server or a chassis server. Figure 1 In the example shown, the computing system is a rack-mount server, specifically a first rack 400. The first rack 400 may include a first enclosure 100 and a second enclosure 200, which can establish a communication connection via a link 300. The number of the first rack 400, the first enclosure 100, and the second enclosure 200 can be one or more; this application does not impose a specific limitation. Furthermore, the first rack 400 may also include a power module, a management module, fans, etc., which this application does not impose a specific limitation on.
[0106] The enclosure can refer to a chassis or other enclosure that houses and secures components, and supports and protects the various components within the chassis. The enclosure may include a shell, brackets, various switches and indicator lights on the panel, etc. The enclosure may be made of a combination of steel and plastic; this application does not impose specific limitations. In the embodiments of this application, each first enclosure 100 includes a fusion node 110, and each second enclosure 200 includes a switching node 210.
[0107] The fusion node 110 protected by the first housing 100 is a computing device with computing capabilities. This computing device can serve as a computing node in a computing system cluster, providing computing power to the cluster. The fusion node 110 includes at least a chip and an interface, and may also include more components depending on the actual application scenario, such as a motherboard, memory, hard drive, heatsink, graphics card, PCIe, etc. This application does not impose specific limitations.
[0108] In this embodiment, the chips of the fusion node 110 may include a computing chip 130 and a first switching chip 120, wherein the number of computing chips 130 and first switching chips 120 may be one or more. The computing chip 130 may be a chip in a computing system cluster that processes computing tasks, such as a central processing unit (CPU), graphics processing unit (GPU), neural network processing unit (NPU), data processing unit (DPU), etc., and this application does not impose specific limitations. The computing chip 130 may be a computing chip used for processing services within a cluster in application scenarios such as HPC and AI, or a storage chip used by storage nodes in a distributed storage scenario, and this application does not impose specific limitations.
[0109] The first switching chip 120 may have the function of transmitting electrical signals and / or optical signals and transmitting data based on rules. It can provide a dedicated electrical or optical signal path for any two computing chips 130 connected to the first switching chip 120. This application does not make any specific limitations.
[0110] The first switching chip 120 may store a routing table or a media access control (MAC) address table. The routing table or MAC address table may include multiple entries, each of which may represent a forwarding path. Each entry may include at least a source address, a destination address, and a corresponding next-hop address. After receiving a data packet sent by the computing chip 130, the first switching chip 120 may query the routing table or MAC address table based on the source address and destination address carried in the data packet to obtain the forwarding path of the data packet, determine the next-hop address, and then forward the data packet to the next-hop address.
[0111] In specific implementations, if the first switching chip 120 is a Layer 2 switch (link layer switch), a MAC address table can be generated using switch forwarding algorithms such as Address Resolution Protocol (ARP), and the source and destination addresses can be MAC addresses. If the first switching chip 120 is a Layer 3 switch (network layer switch), a routing table can be generated using routing algorithms such as Routing Information Protocol (RIP) and Border Gateway Protocol (BGP), and the aforementioned source and destination addresses can be IP addresses; this application does not impose specific limitations on this.
[0112] Optionally, a first switching chip 120 can be connected to each computing chip 130. For example, if the first housing includes computing chips 1 to 32, and also includes a first switching chip 1 and a first switching chip 2, then the first switching chip 1 can establish a communication connection with computing chips 1 to 32, and the first switching chip 2 can also establish a communication connection with computing chips 1 to 21. It should be understood that the above examples are for illustration only and this application does not impose specific limitations.
[0113] Optionally, a first switching chip 120 can also be connected to some computing chips 130. For example, if the first enclosure includes computing chips 1 to 32, and also includes first switching chip 1 and first switching chip 2, then first switching chip 1 can establish communication connections with computing chips 1 to 16, and first switching chip 2 can establish communication connections with computing chips 17 to 32. In specific implementation, the number of computing chips 130 connected to each first switching chip 120 can be determined based on the bandwidth requirements of the computing chips within the first enclosure 100, as well as the number of ports and switching capacity of the first switching chip. For example, if the bandwidth requirement of the computing chips is 100Gb, and the switching capacity of the first switching chip is 12.8Tbps, providing 32 downlink ports, then each port can provide a maximum bandwidth of 200Gb. When the number of computing chips connected to the first switching chip does not exceed 64, the 100Gb bandwidth requirement of the computing chips can be met. Furthermore, the number of computing chips connected to each first switching chip can be further determined by considering other conditions such as the cabinet capacity and the scale of computers required by the user. It should be understood that the above examples are for illustrative purposes only and are not intended to be specific limitations in this application.
[0114] Optionally, communication connections can also be established between the first switching chips 120, which is not specifically limited in this application.
[0115] It should be noted that the first switching chip 120 and the computing chip 130 are housed in the same enclosure. Optionally, the first switching chip 120 and the computing chip 130 can establish a communication connection through the bus inside the computing device, such as the Peripheral Component Interconnect Express (PCIe) bus, or the extended industry standard architecture (EISA) bus, unified bus (Ubus or UB), compute express link (CXL), cachecoherent interconnect for accelerators (CCIX), etc. This application does not make specific limitations.
[0116] Switching node 210 can be a network device capable of forwarding electrical or optical signals. This switching node includes at least a second switching chip 220 and a second housing. The number of second switching chips 220 can be one or more. In specific implementations, the second housing of switching node 210 is different from the first housing of converged node 110, but the first and second housings are deployed in the same first rack.
[0117] The second switching chip 220 can be used to transmit electrical and / or optical signals and transmit data based on forwarding rules. It can provide a dedicated electrical or optical signal path for any two computing chips 130 connected to the switching chip, which is not specifically limited in this application. It should be noted that the first switching chip 120 and the second switching chip 220 can be the same or different models, types, and specifications of switching chips. The appropriate first switching chip 120 and second switching chip 220 can be selected according to the actual application scenario. The second switching chip 220 can also store a routing table or a MAC address table. This routing table or MAC address table can include multiple entries, each representing a forwarding path. Each entry can at least include a source address, a destination address, and a corresponding next-hop address. After receiving a data packet, the second switching chip 220 can query the routing table or MAC address table based on the source and destination addresses carried in the data packet to obtain the forwarding path of the data packet, determine the next-hop address, and then forward the data packet to the next-hop address. For details, please refer to the description of the first switching chip 120 above; it will not be repeated here.
[0118] As one possible implementation, the computer cluster may include one or more first racks, each first rack may include at least one switching node 210 and at least one converged node 110, each converged node 110 may include at least one computing chip 130 and at least one first switching chip 120, and each switching node 210 may include at least one second switching chip 220.
[0119] In this embodiment, data exchange between computing chips within a converged node 110 can be achieved through a first switching chip 120 within the same housing, while data exchange between different converged nodes 110 can be achieved through a second switching chip 220 within a switching node 210 in the same first rack. Each first switching chip connects to multiple computing chips, and each first switching node connects to multiple first switching chips. Thus, even if the number of computing chips connected to a first switching chip decreases, the number of computing chips in the entire computer cluster can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth available to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0120] For example, continuing with the previous example, suppose a switching chip has a switching capacity of 12.8Tbps and can provide 64 200Gb ports and 32 downlink ports. When one switch connects 128 computing chips, each computing chip has an available bandwidth of only 50Gb. In the first rack 400 provided in this application, each first switching chip 120 is connected to 32 computing chips 130, and each second switching chip 220 is connected to 32 first switching chips 120. The first rack 400 can provide 1024 200Gb ports. The cluster size has increased from 128 computing chips 130 to 1024 computing chips, and the port bandwidth has increased from 50Gb to 200Gb. This not only expands the scale of the computing cluster but also increases the bandwidth of each computing chip, solving the problem of computing chip performance being limited by bandwidth and improving the performance of the entire computer cluster.
[0121] Link 300 refers to the physical line between the first box 100 where the convergence node 110 is located and the second box 200 where the switching node 210 is located. Link 300 can include electrical connections and optical connections. The two connection methods are explained below.
[0122] First, regardless of whether an electrical or optical connection is used, the connection between switching node 210 and converged node 110 is orthogonal. Orthogonal connection refers to signal connectivity between two mutually perpendicular enclosures. For example, the first enclosure 100 containing converged node 110 is placed horizontally, and the second enclosure 200 containing switching node 210 is placed vertically, establishing a communication connection between the first enclosure 100 and the second enclosure 200. Alternatively, the second enclosure 200 can be placed horizontally, and the first enclosure 100 vertically. This application does not impose specific limitations. It should be understood that orthogonal connection between switching node 210 and converged node 110 reduces the length of the link 300 between them, thereby mitigating cable connection error rates caused by excessive cables or optical fibers, and reducing cable or optical fiber loss.
[0123] Secondly, the electrical connection is mainly achieved through electrical connectors to realize the communication connection between the converged node 110 and the switching node 210. The electrical connectors here are backplane-less orthogonal connectors. The backplane-less orthogonal connectors allow the switching node 210 and the converged node 110 to be directly orthogonally connected. The two are not connected through cables or optical fibers, but are directly coupled, thereby avoiding the problems of high cable connection error rate and high cable or optical fiber loss caused by too many cables or optical fibers.
[0124] Alternatively, the electrical connection can also be achieved between the converged node 110 and the switching node 210 using a combination of electrical connectors and cables. The electrical connectors used here are high-speed connectors with a backplane. This connection method requires a small number of cables. First, a flexible connection is established between the switching node 210, the converged node 110, and the high-speed connector using cables. Then, the high-speed connector is rotated 90 degrees to achieve an orthogonal connection between the switching node 210 and the converged node 110, making them nearly directly orthogonal. Because the number of cables is reduced, the problems of high cable connection error rates and high cable or fiber loss caused by an excessive number of cables or optical fibers can be mitigated.
[0125] Finally, the optical connection is mainly achieved between the convergence node 110 and the switching node 210 through optical connectors. These optical connectors can be blind-mating connectors, which achieve orthogonal connection between the two nodes via optical signals. The optical signals can be further enhanced through dense wavelength division multiplexing (DWDM) and high-density fiber optic connectors to increase the communication bandwidth between the switching node 210 and the convergence node 110, thereby increasing the bandwidth available to each computing chip. This allows the computer cluster to scale up and improve its performance. However, blind-mating connectors are more expensive than backplane-less orthogonal connectors; therefore, the specific link and implementation method of link 300 should be selected based on the actual application scenario.
[0126] In a specific implementation, the switching node 210 can establish a communication connection with each fusion node 110 through the link 300, so that data communication between each fusion node can be achieved through any switching node 210, thereby enabling communication connections between all computing chips in the first cabinet 400.
[0127] In one embodiment, the number of switching nodes 210 and converged nodes 110 in the first rack 400 can be determined based on at least one of the bandwidth requirements of the computing chip 130, the number of ports and switching capacity of the first switching chip 120, and the number of ports and switching capacity of the first switching node 210. Then, considering various environmental conditions in the actual application scenario, such as the dimensions of the first and second enclosures, the number of enclosures that can be installed in the rack, and the parallel processing capability required by the user, the number of second switching chips 220 in each switching node 210 and the number of computing chips 130 and first switching chips 120 in each converged node 110 can be determined.
[0128] For example, suppose a server rack can accommodate 256 computing chips. The user requires at least 200Gb bandwidth for computing chips 130. The selected first switching chip 120 and second switching chip 220 each have a switching capacity of 12.8Tbps, providing 64 200Gb ports and 32 downstream ports. With 256 computing chips, the entire first server rack 400 requires 9600 200Gb ports. Therefore, each first switching chip 120 can connect to a maximum of 32 computing chips 130, and each second switching chip 220 is connected to at least 8 first switching chips 120, thus realizing a computer cluster of 256 computing chips, with each computing chip having an available bandwidth of 200Gb. Furthermore, considering practical application scenarios, the number of second switching chips in each switching node and the number of first switching chips and compute chips in each converged node are determined. For example, a converged node includes 2 first switching chips and 32 compute chips, with each first switching chip connected to one of the 32 compute chips, resulting in a total of 8 converged nodes and 16 first switching chips. Each switching node includes 4 second switching chips, resulting in a total of 4 switching nodes and 16 second switching chips. This network scale not only ensures the available bandwidth of each compute chip and the scale of the entire rack's computer cluster, but also provides multiple parallel paths between each compute chip, resulting in good network fault tolerance and allowing for reasonable traffic distribution within pods to avoid overload issues.
[0129] Figure 2 This is a schematic diagram of the structure of a computing system provided in this application scenario, as shown below. Figure 2 As shown, the computing system can be Figure 1 The first cabinet 400 in this embodiment may include 16 converged nodes 110 and 8 switching nodes 210. The 16 converged nodes 110 and 8 switching nodes 210 can establish a communication connection through a link 300, which may be the aforementioned backplane-less orthogonal connector. Each converged node 110 includes 16 computing chips 130 and 2 first switching chips 120, and each switching node 210 includes 4 second switching chips 220.
[0130] It needs to be explained that, Figure 2The number of converged nodes 110 and switching nodes 210 in the first cabinet 400 are for illustrative purposes only. The number of computing chips 130 and first switching chips 120 in the converged node 110 are for illustrative purposes only. The number of second switching chips 220 in the switching node 210 is for illustrative purposes only. This application does not impose specific limitations. The number of converged nodes 110 and switching nodes 210 can be determined based on at least one of the bandwidth requirements of computing chips 130, the number of ports and switching capacity of first switching chips 120, and the number of ports and switching capacity of first switching nodes 210. It can also be determined in conjunction with the network topology of converged nodes 110 and switching nodes 210.
[0131] For example, suppose the network topology between converged node 110 and switching node 210 in the first rack is a fat tree network. A fat tree network is a large-scale non-blocking network built with a large number of low-performance switches and multiple layers of network. Typically, a fat tree network can include multiple clusters (pods) and multiple core switches connecting the pods. Each pod includes multiple Layer 1 switches and multiple Layer 2 switches. Each Layer 1 switch connects to one or more computing chips and is used for data communication between computing chips. Each Layer 2 switch connects to one or more Layer 1 switches and is used for data communication between Layer 1 switches. Each core switch connects to one or more pods and is used for data communication between pods. Assuming there are K pods in the first rack, then the number of Layer 1 switches is K / 2, the number of Layer 2 switches is (K / 2)², and the number of computing chips connected to each Layer 1 switch is (K / 2)². Combining the bandwidth requirements of the computing chip 130, the number of ports and switching capacity of the first switching chip 120, and the number of ports and switching capacity of the first switching node 210, the number of each converged node 110 and switching node 210 can be determined. It should be understood that the above example is based on a fat tree network. In other network structures, the number of fusion nodes 110 and exchange nodes 210 can be determined according to the rules of other network structures. This application does not make any specific limitations.
[0132] refer to Figure 1 As described in the embodiments, the 16 computing chips 130 and 2 first switching chips 120 in each fusion node 110 are deployed in the same first housing 100, and the 4 second switching chips 220 in each switching node are deployed in the same second housing 200. Data communication is achieved between the first housing 100 and the second housing 200 through orthogonal connections.
[0133] It should be understood that the switching node 210 and the converged node 110 are connected using a backplane-less orthogonal connector without the use of cables or optical fibers. This reduces the number and length of cables, which not only reduces signal loss but also avoids the problem of high connection error rate caused by too many cables or the problem of photoelectric conversion delay overhead caused by too many optical fibers.
[0134] In the specific implementation, after the computing chip, the first switching chip, and the second switching chip establish a communication connection, the first switching chip and the second switching chip can generate a routing table or a MAC address table according to their respective connection relationships through a forwarding algorithm. The routing table or MAC address table may include multiple entries, each entry may represent a forwarding path, and each entry may include at least the source address, the destination address, and the corresponding next-hop address. After receiving the data packet sent by the computing chip, the first switching chip can query the routing table or MAC address table according to the source address and destination address carried in the data packet to obtain the forwarding path of the data packet, determine the next-hop address, and then forward the data packet to the next-hop address.
[0135] If the first or second switching chip is a Layer 2 switch (link layer switch), a MAC address table can be generated using switch forwarding algorithms such as ARP. The source address and destination address can be MAC addresses. If the first or second switching chip is a Layer 3 switch (network layer switch), a routing table can be generated using routing algorithms such as RIP and BGP. The source address and destination address can be IP addresses. This application does not impose any specific limitations on this.
[0136] The forwarding paths recorded in the routing table or MAC address table are determined based on the connection relationship between the first switching chip 120 and the computing chip 130. The possible connection relationships between the first switching chip 120 and the computing chip 130 and the corresponding forwarding paths in this application are explained below.
[0137] Optionally, in each fusion node 110, each first switching chip 120 can establish a communication connection with all computing chips 130. Communication between computing chips 130 within each fusion node 110 can be achieved through any first switching chip 120 within the same fusion node 110. The forwarding path of data packets between computing chips 130 within each fusion node 110 may include any first switching chip 120. Each switching node 210 establishes a communication connection with all fusion nodes 110 through link 300, enabling data communication between each fusion node 110 through any second switching chip 220. The forwarding path of data packets between each fusion node 110 may include any second switching chip 220.
[0138] Optionally, each first switching chip 120 can establish a communication connection with some computing chips 130. Data communication between multiple computing chips within the converged node can determine the forwarding path based on the first switching chip connected to it. If the sender and receiver are connected to the same first switching chip, the first switching chip can realize data communication between the first switching chip 120 and some computing chips 130 through a routing table. The forwarding path of the data packet may include the source address, the address of the same first switching chip, and the destination address.
[0139] If the sender and receiver are connected to different first switching chips and there is no direct connection between the two first switching chips, then their data communication can be achieved through a second switching chip connected to their respective first switching chips. The forwarding path of the message may include the source address, the address of the first switching chip 1, the address of the second switching chip 1, the address of the first switching chip 2, and the destination address. The first switching chip 1 is connected to the computing chip where the source address is located, the first switching chip 2 is connected to the computing chip where the destination address is located, and the second switching chip 1 is connected to both the first switching chip 1 and the first switching chip 2.
[0140] If the sender and receiver are connected to different first switching chips, and there is a direct connection between the two first switching chips, then data communication can be achieved through the two first switching chips connected to the sender and receiver. The forwarding path of the message may include the source address, first switching chip 1, first switching chip 2, and destination address.
[0141] In specific implementations, each computing chip can be connected to a first switching chip, so the data packets generated by the computing chip can be directly sent to the first switching chip connected to it. In some embodiments, each computing chip can also be connected to at least one first switching chip, so after the computing chip generates a data packet, it can send the data packet to the first switching chip with better network conditions based on the network conditions fed back by at least one first switching chip connected to it to achieve packet forwarding. Alternatively, the first cabinet 400 may also include a management node, which can determine the first switching chip for forwarding the data packet based on load balancing or other network management algorithms. This application does not make specific limitations. Alternatively, the computing chip can also select the first switching chip connected to it that can process the first switching chip carrying the identifier carried in the data packet to achieve data packet forwarding based on the identifier carried in the data packet. For example, the identifier can be the identifier of the first forwarding plane, so only the first switching chip of the first forwarding plane can forward the data packet. This application does not make specific limitations.
[0142] It needs to be explained that, Figure 2In the example shown, although the connections between chips and between chips and connectors within the switching node 210 and the fusion node 110 are not explicitly shown, the connections can be encapsulated within the housing according to the aforementioned description of the connections between chips within the nodes. Specifically, communication connections can be established between the computing chip and the first switching chip within the fusion node 110, and between the first switching chip and the connector, via a bus. Similarly, communication connections can be established between the second switching chip and the connector within the switching node 210 via a bus. This bus can be a PCIe bus, EISA bus, UB bus, CXL bus, CCIX bus, etc., and this application does not impose any specific limitations.
[0143] In this embodiment, the first switching chip 120 is responsible for data communication between computing chips 130 within the same fusion node 110, and the second switching chip 220 in the switching node 210 is responsible for data communication between computing chips 130 in different fusion nodes 110. The following is in conjunction with... Figure 3 The specific process of data communication between computing chips 130 within the same node and between different nodes is illustrated with an example of the first switching chip 120.
[0144] As one possible implementation method, Figure 3 This is a schematic diagram illustrating the connection relationship between a fusion node and a switching node provided in this application. Figure 3 Merge node 1 and merge node 16 in the middle are Figure 2 Two of the 16 fusion nodes in the process, Figure 3 The exchange node 1 in the middle is Figure 2 The eight exchange nodes in the middle.
[0145] like Figure 3 As shown, computing chips C1 and C2 in converged node 1 are connected to the first switching chip L12, which is connected to link 300. The second switching chip L23 in switching node 1 is also connected to link 300. The first switching chip L14 in converged node 16 is connected to link 300 and is also connected to computing chip C3. Based on the connections between the chips, the routing table or MAC address table in each first switching chip can record multiple forwarding paths. The rules for these forwarding paths can be as follows.
[0146] When computing chips C1 and C2 communicate with each other within fusion node 1, they can do so through the first switching chip L12 within fusion node 1. The forwarding path of the data packets sent from computing chip C1 to computing chip C2 is as follows: computing chip C1 sends the data packets to the first switching chip L12, and the first switching chip L12 forwards the data packets to computing chip C2, thereby realizing data communication between computing chips within the same fusion node.
[0147] When computing chips C1 and C4 communicate with each other inside fusion node 1, the two computing chips are connected to different first switching chips. The forwarding path of the data packets sent from computing chip C1 to computing chip C4 is as follows: computing chip C1 can first send the data packets to the first switching chip L12, the first switching chip L12 sends the data packets to the second switching chip L23 in switching node 1, the second switching chip L23 then forwards the data packets to the first switching chip L11 in fusion node 1, and the first switching chip L11 then forwards the data packets to computing chip C4.
[0148] Optionally, if a communication connection is established between the first switching chip L11 and the first switching chip L12, the forwarding path of the data packet sent by the computing chip C1 to the computing chip C2 is as follows: the computing chip C1 can first send the data packet to the first switching chip L12, the first switching chip L12 forwards the data packet to the first switching chip L11, and the first switching chip L11 then forwards the data packet to the computing chip C4.
[0149] It should be noted that if there are multiple first switching chips connected to the computing chip at the sending end, or multiple first switching chips connected to the computing chip at the receiving end, then there can be multiple paths with different overhead between the computing chip at the sending end and the computing chip at the receiving end. The routing table or MAC address table on the local computing chip at the sending end can determine the optimal forwarding path through algorithms such as load balancing. The computing chip at the sending end can query the local routing table to obtain the optimal forwarding path, thus avoiding network congestion.
[0150] When computing chip C1 in fusion node 1 and computing chip C3 in fusion node 16 communicate with each other, data communication can be achieved through the first switching chip L12, the second switching chip L23 in fusion node 1, and the first switching chip L14 in fusion node 16. Specifically, computing chip C1 can first send data packets to the first switching chip L12. The first switching chip L12 can then forward the data packets to the second switching chip L23 on fusion node 1. The second switching chip L23 can then forward the data packets to the first switching chip L14 on fusion node 16. The first switching chip L14 can then forward the data packets to computing chip C3 within its own node.
[0151] In practice, data communication between converged nodes can be achieved through any second switching chip on any switching node, such as any second switching chip in switching node 8. The appropriate second switching chip can be determined based on the workload of the second switching chip in each switching node. Alternatively, users can set some switching nodes to be responsible for the data communication of some converged nodes, and no specific restrictions are imposed on the application.
[0152] It should be understood that Figure 3 For ease of description, the complete connections between the computing chips and the first switching chips within the fusion node, the connections between all the first switching chips and link 300, and the connections between all the second switching chips and link 300 are not shown. In practical applications, all computing chips in fusion node 1 can be connected to some or all of the first switching chips within fusion node 1, all the first switching chips in fusion node 1 can be connected to link 300, and all the second switching chips in fusion node 1 can be connected to link 300. Similarly, all the first switching chips in fusion node 16 can be connected to the connector.
[0153] refer to Figure 3 From the description of the embodiments, it can be seen that Figure 2 The data communication process between any computing chip in the fusion nodes 1 to 16, both within and between nodes, will not be elaborated here.
[0154] In one embodiment, the first cabinet 400 includes multiple converged nodes 110 and multiple switching nodes 210. These converged nodes 110 and switching nodes 210 can form an interconnection network. The network topology of the interconnection network can include a network topology without a central switching node and a network topology with a central switching node. The network topology without a central switching node and the network topology with a central switching node will be explained below.
[0155] Specifically, in a network topology without a central switching node, the interconnection network does not include other switch cabinets, but only the aforementioned computing systems (such as the first cabinet 400). Data communication between multiple computing systems is achieved through switching node 210. In a network topology with a central node, the interconnection network also includes other switch cabinets, which are used to realize data communication between multiple computing systems. These switch cabinets are the aforementioned central switching node.
[0156] In specific implementations, network topologies without a central switching node may include, but are not limited to, dragonfly (or dragonfly+) networks, torus networks, etc., while network topologies with a central switching node may include, but are not limited to, fat-tree networks. This application does not impose any specific limitations.
[0157] The dragonfly network comprises multiple groups, each essentially a sub-network enabling intra-group communication. Each group can be one of the aforementioned first racks 400. These groups are connected by links, which are relatively narrow, resembling the wide body and narrow wings of a dragonfly, hence the name "dragonfly network." In other words, the first racks 400 are connected by links, and data communication between them is achieved through switching nodes within each rack 400. It should be understood that because the dragonfly network uses narrow links between its groups, the number of global links is significantly reduced, lowering networking costs.
[0158] Furthermore, the groups in a dragonfly network can be connected in an all-to-all manner. This type of network is also called a dragonfly+ network. All-to-all means that each group has at least one direct link to every other group in the dragonfly network. Dragonfly+ networks have larger group sizes and support more connected groups. Within a group, a connection method with as few hops as possible is used, usually an all-to-all or flat butterfly structure. This can further shorten the link length, thereby providing low-latency communication.
[0159] The Torus network consists of multiple nodes, each of which can be the first rack 400 mentioned above. There are loopback links (or toroidal boundaries) between each node. The existence of loopback links means that there is no longer a distinction between nodes located at the topology center or edge in the Torus network. Each node has two adjacent nodes in each dimension, which allows for multiple forwarding paths between any two nodes. The network has high reliability, switching capacity, and scalability.
[0160] In specific implementations, in the above-mentioned network topology without a central switching node, namely the dragonfly network, the dragonfly+ network, and the torus network, the forwarding of data packets between the first racks 400 needs to go through the switching node 210 in each first rack 400. The forwarding path of the data packets may pass through one or more switching nodes 210, which is not specifically limited in this application.
[0161] A fat tree network is a large-scale, non-blocking network constructed using numerous low-performance switches and multi-layered networks. Typically, a fat tree network includes multiple clusters (pods) and multiple core switches connecting the pods. Each pod includes multiple Layer 1 switches and multiple Layer 2 switches. Each Layer 1 switch connects to one or more compute chips, handling data communication between these chips. Each Layer 2 switch connects to one or more Layer 1 switches, handling data communication between Layer 1 switches. Each core switch connects to one or more pods, handling data communication between them. Furthermore, core switches can form multiple sub-networks, with data communication between each sub-network handled by the layer above it, and so on. This allows for multiple parallel paths between any two compute nodes in the fat tree network, resulting in good fault tolerance and efficient traffic distribution within pods to prevent overload. It should be understood that fat tree network is a non-blocking network technology that uses a large number of switching chips to build a large-scale non-blocking network. The bandwidth of this network does not converge from bottom to top, and there are many data communication paths. There is always a path that can make the communication bandwidth reach the required bandwidth, which further solves the problem that the performance of computing chips is limited by network bandwidth and improves the performance of computer clusters.
[0162] In the specific implementation, in the network topology structure with a central switching node, i.e., the fat tree network, each first cabinet 400 can establish a communication connection with one or more central switching nodes, and the forwarding of data packets between the first cabinets 400 needs to be achieved through one or more of the aforementioned central switching nodes.
[0163] It should be understood that the network topology composed of multiple converged nodes 110 and multiple switching nodes 210 can also include other commonly used large-scale network topologies that are sensitive to bandwidth and latency. These will not be listed here, and this application does not impose any specific limitations.
[0164] The following example illustrates the computing system provided in this application by taking the network topology between multiple computing systems (i.e., the first cabinet 400 mentioned above) as a fat tree network topology with a central switching node.
[0165] The interconnected network composed of multiple first cabinets 400 may also include multiple second cabinets, which may be switch cabinets. Each second cabinet may include multiple third switching chips. Each second cabinet is used to realize the communication connection between multiple first cabinets. Taking a fat tree network as an example, the first switching chip and computing chip in the first cabinet 400 can form the first layer network in the fat tree network, the second switching chip and the first switching chip can form the second layer network in the fat tree network, and the third switching chip and the second switching chip can form the third layer network in the fat tree network. Among them, the third switching chips can also form multi-layer networks, forming the third layer network, the fourth layer network, etc. of the fat tree network. This application does not limit the number of network layers formed by the third switching chips.
[0166] In practice, the second cabinet and the first cabinet 400 can establish a communication connection via cable or optical fiber, and this application does not make any specific limitations.
[0167] For example, Figure 4 This is a schematic diagram of a four-layer fat tree network provided in this application, as shown below. Figure 4 As shown, the first layer of this four-layer fat tree network is implemented by multiple first switching chips L1, the second layer by multiple second switching chips L2, the third layer by multiple third switching chips L3, and the fourth layer by multiple fourth switching chips L4. The first switching chips L1, second switching chips L2, third switching chips L3, and fourth switching chips L4 can be the same or different switching chips; this application does not impose specific limitations.
[0168] Specifically, the first switching chip L1 and the computing chip can be deployed in the first enclosure, the second switching chip L2 can be deployed in the second enclosure, the first enclosure and the second enclosure can be deployed in the first rack, and the third switching chip L3 and the fourth switching chip L4 can be deployed in the second rack. It should be understood that the third switching chip L3 and the fourth switching chip L4 can also be deployed in different racks, and this application does not make specific limitations.
[0169] In specific implementation, each computing chip can establish a communication connection with each first switching chip L1 in the same box, each first switching chip L1 can establish a communication connection with each second switching chip L2 in the same rack, each second switching chip L2 can establish a communication connection with each third switching chip L3, and each third switching chip L3 can establish a communication connection with each fourth switching chip L4.
[0170] It should be understood that fat-tree networks construct large-scale non-blocking networks by using a large number of switching chips. Although fat-tree networks have the advantage of theoretically non-converging bandwidth from bottom to top, the large number of switching chips and complex connections leads to an excessive number of cables or optical fibers in the network, which can easily cause problems such as difficult cable connections, connection errors, and high photoelectric conversion losses. Figure 4 In the four-layer fat-tree network shown, the first and second layers are encapsulated in the same cabinet. The computing chip and the first switching chip L1 are in the first cabinet, and one or more second switching chips L2 are in the second cabinet. The first and second cabinets are connected via a backplane-less orthogonal connector, which greatly reduces the number of cables or optical fibers in the first cabinet containing the first and second cabinets. Figure 4 The four-layer fat tree network shown significantly reduces the number of cables or optical fibers in the network, thus solving the problems of cable connection difficulties, connection errors, and large photoelectric conversion losses that often occur in fat tree networks due to the excessive number of cables or optical fibers.
[0171] It needs to be explained that, Figure 4 In the example shown, each converged node's first box contains *a* computing chips and one first switching chip L1; each second box contains one second switching chip L2; each first rack contains *b* first boxes and *c* second boxes; each second rack contains *d* third switching chips L3 and *e* fourth switching chips L4; and the computer cluster contains *m* first racks and *n* first racks. It should be understood that... Figure 4 For illustrative purposes only, this application does not limit the number of computing chips and the number of first switching chips L1 in the first housing, the number of second switching chips L2 in the second housing, the number of first and second housings in the first rack, the number of third switching chips L3 and fourth switching chips L4 in the second housing, or the number of first and second racks. Specifically, the number of layers in the fat-tree network and the number of switching chips in each layer can be determined based on the bandwidth required by each computing chip, the number of ports provided by each switching chip, and the switching capacity; this application does not impose specific limitations.
[0172] It needs to be explained that, Figure 4 In the example shown, each computing chip is connected to each first switching chip L1, each first switching chip L1 is connected to each second switching chip L2, each second switching chip L2 is connected to each third switching chip L3, and each third switching chip L3 is connected to each fourth switching chip L4. It should be understood that... Figure 4For illustrative purposes only, in a concrete implementation, each computing chip can also be connected to a portion of the first switching chip L1, each second switching chip can also be connected to a portion of the third switching chip L3, and each third switching chip L3 can also be connected to a portion of the fourth switching chip L4, ensuring that each computing chip can communicate with other computing chips within the computer cluster.
[0173] Optionally, if each first switching chip L1 is connected to each computing chip, then computing chips within the same housing can communicate with each other through the first switching chip L1 within the same node.
[0174] Optionally, if each computing chip within the same housing is connected to a portion of the first switching chip L1, then when the computing chips within the same housing interact with each other, if the sending and receiving computing chips are connected to the same first switching chip L1, then data communication can be achieved through the commonly connected first switching chip L1. For details, please refer to... Figure 3 The embodiment describes the steps involved in sending a data packet from computing chip C1 to computing chip C2. If the sender and receiver are not connected to the same first switching chip L1, data communication can be achieved through their respective connected first switching chips L1 and the second switching chip L2 within the same rack. For details, please refer to [reference needed]. Figure 3 The steps and flow of the computing chip C1 sending a data packet to the computing chip C4 in the embodiment are described, and will not be repeated here.
[0175] For example, if a user requires a computer cluster consisting of 60,000 computing chips, each requiring four 200Gb ports, the cluster would need to provide 240,000 200Gb ports. Using a 12.8Tbps switching capacity, providing 64 200Gb ports, and 32-port switching chips as an example, the following configuration could be established based on user requirements and the specifications of each computing chip: Figure 4 The illustrated four-layer fat-tree network, in which the number of computing chips is a = 8, the number of first boxes is b = 32, the number of second boxes is c = 32, the number of third switching chips is d = 64, the number of fourth switching chips (L4) is e = 32, the number of first racks is m, and the number of second racks is n, are approximately 200. This configuration can meet user requirements, ensuring that 60,000 computing chips can each have four 200Gb ports. It should be understood that the above example is for illustrative purposes only and this application does not impose specific limitations.
[0176] In one embodiment, the plurality of first switching chips within the converged node includes first switching chips of a first switching plane and first switching chips of a second switching plane. The plurality of second switching chips within the switching node includes second switching chips of a first switching plane and second switching chips of a second switching plane. Simply put, the network composed of the first and second switching chips within the first rack can include multiple switching planes. A portion of the first and second switching chips are responsible for data communication on the first switching plane, while another portion of the first and second switching chips are responsible for data communication on the second switching plane. It should be understood that the aforementioned first and second switching planes do not specify that the technical solution of this application only supports two switching planes, but rather express that the technical solution of this application supports different switching planes. In practice, it can support more than two switching planes, and the specific number of switching planes can be determined according to the actual application scenario.
[0177] Optionally, the first and second switching planes can be network-isolated switching planes. Specifically, for business security considerations, to achieve network isolation between user service 1 and service 2 and avoid network security issues, a layer 2 or higher fat-tree network can be formed using the cabinet provided in this application, and a corresponding switching plane can be assigned to each switching chip. This allows the user to use the first switching plane for data communication of service 1 and the second switching plane for data communication of service 2. No data communication occurs between the two switching planes; that is, the first and second switching chips of the first switching plane will not process data packets from the second switching plane, and vice versa, thereby achieving network isolation between the first and second switching planes. It should be understood that the above examples are for illustrative purposes only and this application does not impose specific limitations.
[0178] Optionally, the first and second switching planes can handle different services. For example, user service 1 uses the first switching plane, while service 2 uses the second switching plane, thus avoiding network latency caused by service conflicts. By using multiple switching planes, the computer cluster can provide users with greater bandwidth capabilities, improving the user experience.
[0179] Optionally, the first switching plane and the second switching plane use different communication protocols. For example, service 1 requires communication protocol 1, and service 2 requires communication protocol 2. Using the cabinet provided in this application, and allocating corresponding switching planes to each switching chip, the first switching plane uses communication protocol 1 to process service 1, and the second switching plane uses communication protocol 2 to process service 2. It should be understood that the above examples are for illustration only, and this application does not make any specific limitations.
[0180] For example, such as Figure 5 As shown, Figure 5 This is a structural diagram of a server rack provided in this application in another application scenario. Figure 2 In the application scenario shown, the cabinet includes only one switching plane. Figure 5 In the application scenario shown, the rack includes two switching planes. The rack comprises 16 converged nodes and 10 switching nodes. Each converged node includes 8 computing chips and 3 first switching chips. First switching chip L11-1 and first switching chip L12-1 belong to the first switching plane, while first switching chip L11-2 belongs to the second switching plane. Each switching node includes 4 second switching chips. Switching nodes 1-8 belong to the first switching plane, while switching nodes 9 and 10 belong to the second switching plane.
[0181] It should be understood that Figures 2-3 The described structure can be understood as a network structure with a single switching plane. Figure 5 Is Figure 2 and Figure 3 Based on the network structure of the first switching plane shown, a second switching plane is added. Specifically, each fusion node adds a first switching chip L11-2 as the first switching chip of the second switching plane, and two new switching nodes 9 and 10 are added as switching nodes of the second switching plane, with each switching node including four second switching chips. Thus, Figure 5 In the network structure shown, the first switching plane includes 32 first switching chips and 32 second switching chips, and the second switching plane includes 16 first switching chips and 8 second switching chips.
[0182] In practical implementation, when computing chips communicate data through the second switching plane, computing chips within the same fusion node can use the first switching chip of the second switching plane for time-based data communication, while computing chips between different fusion nodes can use the second switching chip within a switching node of the second switching plane to achieve data communication. Similarly, when using the first switching plane for data communication, the computing chips use both the first and second switching chips of the first switching plane to achieve data communication, which will not be elaborated further here.
[0183] Understandably, computing chips can connect to first switching chips on different switching planes according to business needs. There is no communication connection between the first switching chips on different switching planes, nor is there a communication connection between the second switching chips on different switching planes. This can achieve network isolation between different switching planes, allowing different switching planes to handle different services and use different communication protocols, thus better meeting the diverse needs of users and improving the user experience.
[0184] Optionally, depending on user needs, the first cabinet may include more switching planes, such as a third switching plane, a fourth switching plane, etc. Each switching plane may also include more first switching chips and second switching chips. This application does not make specific limitations.
[0185] In practical implementation, if the number of switching nodes is increased, it can be similar to... Figure 5 As shown, the number of switching nodes is increased horizontally, and the length of the first enclosure containing the converged nodes is adaptively increased, allowing the expanded first enclosure and a larger number of switching nodes to be connected via backplane-less orthogonal connectors. Similarly, if the number of converged nodes is increased, the number of converged nodes can be increased vertically, and the length of the second enclosure containing the switching nodes can be adaptively increased, allowing the expanded second enclosure and a larger number of converged nodes to be connected via backplane-less orthogonal connectors.
[0186] It needs to be explained that, Figure 3 and Figure 5 In the example, the merging node is placed horizontally and the exchanging node is placed vertically. In a specific implementation, the merging node can also be placed vertically and the exchanging node can be placed horizontally. This application does not make any specific limitation.
[0187] It should be understood that the first and second boxes can be directly connected by inserting backplane-less orthogonal connectors without the need for cables or optical fibers. Therefore, a cabinet including multiple first and second boxes can greatly reduce the number of cables or optical fibers, avoid cable connection errors, and reduce cable and optical transmission losses.
[0188] In one embodiment, the network structure of the first rack 400 provided in this application can be a two-layer or higher fat-tree network structure, and includes at least one switching plane. Specifically, a computing chip can simultaneously connect to the first switching chips of multiple switching planes. The first switching chip of the first switching plane establishes a communication connection with the second switching chip of the first switching plane, and the first switching chip of the second switching plane establishes a communication connection with the second switching chip of the second switching plane. First switching chips of different switching planes are not connected to each other, and second switching chips of different switching planes are also not connected to each other, thereby achieving the goal of having multiple switching planes in one rack. Similarly, if the network within the rack is a two-layer or higher fat-tree network, then the second switching chip of the first switching plane is connected to the third switching chip of the first switching plane, and the second switching chips and third switching chips of different switching planes are not connected to each other, and so on, thereby ensuring that the first rack 400 can include more layers of fat-tree networks and more switching planes.
[0189] For example, such as Figure 6 As shown, Figure 6This is a structural diagram of a server rack provided in this application under another application scenario. The server rack includes a first switching plane and a second switching plane, wherein the first switching plane is a four-layer fat-tree network, that is... Figure 4 The four-layer fat tree network shown will not be elaborated upon here. The second switching plane is a two-layer fat tree network. Each first cabinet includes b first switching chips L1 and c second switching chips L2 of the second switching plane. Each first box in the first cabinet includes one first switching chip L1 of the second switching plane, and each second box includes one second switching chip L2 of the second switching plane.
[0190] It needs to be explained that, Figure 1 The second switching chip of the first switching plane and the second switching chip of the second switching plane are packaged in different second boxes. In some embodiments, the second box may also include multiple second switching chips of the switching plane. In specific implementation, if the user requires network isolation of the switching plane, then different switching planes can be packaged in different second boxes. If the user does not need network isolation of the switching plane, but only needs different switching planes to carry out different services, or different switching planes to run different communication protocols, then different switching planes can be packaged in the same second box, saving more hardware resources.
[0191] exist Figure 6 In the cabinet shown, for data communication between computing chips within the same enclosure, if computing chip 1 needs to send a data packet to computing chip a within the first enclosure 1 via the first switching plane, computing chip 1 can send the data packet to the first switching chip L1-1 on the first switching plane within the same enclosure, so that L1-1 forwards the data packet to computing chip a, thus achieving data communication between computing chips within the same enclosure within the first switching plane. If computing chip 1 needs to send a data packet to computing chip a within the first enclosure 1 via the second switching plane, computing chip 1 can send the data packet to the first switching chip L1-1 on the second switching plane within the same enclosure, so that L1-1 forwards the data packet to computing chip a, thus achieving data communication between computing chips within the same enclosure within the second switching plane.
[0192] Similarly, for data communication between computing chips in different boxes, when implemented through the first switching plane, data communication is achieved through the first switching chip and the second switching chip of the first switching plane; when implemented through the second switching plane, data communication is achieved through the first switching chip and the second switching chip of the second switching plane. This will not be elaborated further here.
[0193] Furthermore, for networks at Layer 2 and above, the third switching chip can also be divided into the third switching chip of the first switching plane and the third switching chip of the second switching plane. When the user needs to implement the first switching plane, data communication is achieved through the first switching chip, the second switching chip and the third switching chip of the first switching plane. This will not be elaborated further here.
[0194] It should be noted that the computing system provided in this application can also be a chassis-type server, such as the aforementioned blade server, rack server, and high-density server, etc. This chassis-type server is similar to... Figure 1 The first rack shown has a similar structure. This rack-type server includes a rack, which houses multiple converged nodes 110 and switching nodes 210. It can also deploy other components required for a rack-type server, such as power supplies, fans, and management nodes, which will not be listed here. Similarly, the second rack mentioned above can also be a rack-type device. This rack-type device includes multiple third switching chips. For details, please refer to the descriptions of the first rack 400 and the second rack above, which will not be repeated here.
[0195] In one possible implementation, the computing system (such as the first rack 400 mentioned above) may include one or more symmetrical multi-processing (SMP) systems, and each fusion node may also include one or more SMP systems. An SMP system refers to a set of processors aggregated on a server, which includes multiple CPUs. The processors share memory and other resources on the server, such as a shared bus structure, allowing the workload to be evenly distributed across all available processors. That is, one SMP system corresponds to one OS domain and multiple computing chips. These computing chips can be within the same fusion node or within different fusion nodes; this application does not specifically limit this.
[0196] For example, a computing system may include 16 fusion nodes. These 16 fusion nodes could represent 16 SMP systems, with one fusion node corresponding to one SMP system. Alternatively, the 16 fusion nodes could represent 8 SMP systems, with 2 fusion nodes corresponding to one SMP system. Or, the 16 fusion nodes could be fusion node 1 to fusion node 16, where fusion node 1 represents 2 SMP systems, fusion node 2 and fusion node 3 represent one SMP system, and fusion nodes 4 to 16 represent one SMP system. These examples are for illustrative purposes only and do not constitute a specific limitation in this application.
[0197] It should be understood that since the computing chips are interconnected through the first and second switching nodes, the computing chips in each SMP system can be flexibly combined according to the business needs of the actual application scenario, and this application does not make any specific limitations.
[0198] In summary, this application provides a computing system comprising multiple converged nodes and multiple switching nodes. A first converged node among the multiple converged nodes includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to establish communication connections between the multiple computing chips. The first switching node is used to establish communication connections between the first converged node and other converged nodes among the multiple converged nodes. The first switching node is coupled to the first converged node via a connector, and the multiple computing chips in the first converged node are connected to the first switching chip. Thus, even if the number of computing chips connected to the first switching chip decreases, the total number of computing chips in the rack can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth available to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0199] The above text combined Figures 1 to 6 The architecture of the computing system provided in this application has been described in detail. Next, in conjunction with the appendix... Figure 7 and attached Figure 8 This application further explains the data communication process within the aforementioned computing system.
[0200] Figure 7 This is a flowchart illustrating the steps of a communication method provided in this application scenario. Figure 8 This is a flowchart illustrating the steps of a communication method provided in this application in another application scenario, wherein, Figure 7 and Figure 8 The difference is that, Figure 7 The communication method shown is applied in a scenario where the computing chip and the first switching chip are fully connected. Figure 8 The application scenario of the communication method shown is that the computing chips within the same fusion node are not fully connected to all the first switching chips.
[0201] Figure 7 and Figure 8 The communication method shown can be applied to a computing system 1000, which can be as follows: Figures 1-6 The computing system described in the embodiments, exemplarily, may be... Figures 1-6The first cabinet 400 shown can be configured with a network topology that includes a central switch, such as a fat tree network. The first cabinet 400 includes one or more converged nodes and one or more switching nodes. Each converged node includes one or more computing chips and at least one first switching chip, and each switching node includes one or more second switching chips. For details, please refer to the preceding description. Figures 1-6 The description in the document is not specifically limited in this application.
[0202] The method may include the following steps S710 to S760. It should be noted that steps S710 and S720 describe the steps of a communication method between computing chips within the same fusion node, while steps S730 to S760 describe the steps of a communication method between computing chips within different fusion nodes.
[0203] Step S710: The computing chip 1 sends a first data packet to the first switching chip 1, wherein the destination address of the first data packet is the address of the computing chip 2. The computing chip 1, computing chip 2 and the first switching chip 1 are chips within the fusion node 1. The fusion node 1 is encapsulated in a first housing, and the computing chip 1 and computing chip 2 are connected to the first switching chip 1.
[0204] It should be noted that if there are multiple first switching chips connected to the computing chip 1 at the transmitting end, then there can be multiple paths with different overhead between the transmitting computing chip 1 and the receiving computing chip 2. The local routing table or MAC address table of the transmitting computing chip can determine the optimal forwarding path through algorithms such as load balancing. The transmitting computing chip can query its local routing table or MAC address table to obtain the optimal forwarding path, thus avoiding network congestion. The routing algorithm or switch forwarding algorithm for determining the optimal forwarding path is not specifically limited in this application.
[0205] refer to Figure 7 It is known that computing chip 1 in fusion node 1 is connected to the first switching chip 1, and computing chip 2 is also connected to the first switching chip 1. Therefore, the first data packet sent by computing chip 1 to computing chip 2 can be forwarded through the first switching chip 1. Simply put, data communication between the computing chip acting as the sender and the computing chip acting as the receiver within the same fusion node can be achieved through the first switching chip, which has established connections with both the sender and the receiver.
[0206] Step S720: The first switching chip 1 forwards the first data packet to the computing chip 2.
[0207] In a specific implementation, after computing chip 1 and computing chip 2 establish a communication connection with the first switching chip 1, the first switching chip 1 can record the addresses of all computing chips connected to it and establish a routing table. This routing table records multiple transmission paths, so that after the first switching chip 1 receives a data packet, it can query the routing table according to the source address and destination address carried by the data packet. The algorithm used for querying can be a routing algorithm, such as the Routing Information Protocol (RIP), the Border Gateway Protocol (BGP), etc. This application does not make any specific limitations.
[0208] It should be understood that steps S710 and S720 describe the data communication process within a fusion node, while steps S730 to S760 describe the communication method between computing chips in different fusion nodes. If computing chip 1 does not need to communicate with computing chips in other fusion nodes, steps S730 to S760 can be omitted; if computing chip 1 does not need to communicate with computing chips in the same fusion node, steps S710 and S720 can be omitted, and steps S730 to S760 can be executed directly; of course, computing chip 1 can also execute steps S730 to S760 first and then execute steps S710 and S720, which is not specifically limited in this application.
[0209] In a specific implementation, the aforementioned computing chip 1 can be Figure 3 In the embodiment, computing chip C1 and computing chip 2 can be... Figure 3 In the embodiment, the computing chip C2 and the first switching chip 1 can be... Figure 3 For a detailed description of the first switching chip L12 in the embodiment, steps S710 to S720, please refer to [link / reference]. Figure 3 The steps and flow of data communication between computing chip C1 and computing chip C2 in the embodiment will not be repeated here.
[0210] Step S730: Computing chip 1 sends a second data packet to the first switching chip 1, wherein the destination address of the second data packet is the address of computing chip 3, computing chip 1 is the computing chip in fusion node 1, and computing chip 3 is the computing chip in fusion node 2.
[0211] Step S740: The first switching chip 1 forwards the second data packet to the second switching chip 1 within the switching node 1.
[0212] It should be understood and referenced. Figures 1-6As described in the embodiment, the fusion node and the switching node establish a communication connection through a connector. This connector is an orthogonal connector without a backplane. Therefore, the first switching chip 1 can establish a communication connection with the second switching chip 1. After the first switching chip 1 receives the second data packet carrying the address of the computing chip 3, it can determine the next hop address as the address of the switching node 1 based on the address of the computing chip 3 and the routing table, and then send the address of the second data packet to the switching node 1.
[0213] As mentioned above, each second switching chip within each converged node can establish communication connections with all first switching chips. This results in numerous data communication paths, ensuring that any computing chip can find a path that meets the required bandwidth. Therefore, the second switching chip 1 can be determined by switching node 1 based on the idle status of all second switching chips within the node.
[0214] Step S750: The second switching chip 1 forwards the second data packet to the first switching chip 2 of the converged node 2.
[0215] As can be seen from the foregoing, the first switching chip in each fusion node establishes a communication connection with each second switching chip. The second switching chip 1 is connected to the first switching chip 1 and the first switching chip 2. Therefore, the second switching chip 1 can forward the second data packet to the fusion node 2 where the computing chip 3 is located based on the address of the computing chip 3 carried in the second data packet.
[0216] Step S760: The first switching chip 2 forwards the second data packet to the computing chip 3.
[0217] It should be understood that the first switching chip inside the converged node can forward the second data packet to the computing chip 3 within the same node based on the address of the computing chip 3 carried in the received second data packet.
[0218] In a specific implementation, the aforementioned computing chip 1 can be Figure 3 In the embodiment, computing chip C1 and computing chip 3 can be... Figure 3 In the embodiment, the computing chip C3 and the first switching chip 1 can be... Figure 3 In the embodiment, the first switching chip L12 and the first switching chip 3 can be... Figure 3 In the embodiment, the first switching chip L14 and the second switching chip 1 can be... Figure 3 For a detailed description of the second switching chip L23 in the embodiment, steps S730 to S760, please refer to [link / reference]. Figure 3 The steps and flow of data communication between computing chip C1 and computing chip C3 in the embodiment will not be repeated here.
[0219] It should be understood that in steps S710 and S720 above, computing chip 1 and computing chip 2 are connected to the same first switching chip 1. Therefore, when computing chip 1 and computing chip 2 communicate data, data packets can be forwarded through the first switching chip 1. Referring to the foregoing, each first switching chip within a converged node can be connected to all computing chips within the same node, or each first switching chip can be connected to some computing chips. In this case, it is possible that computing chip 1, acting as the sender, and computing chip 4, acting as the receiver, are connected to different first switching chips. The following discussion, in conjunction with... Figure 8 The communication method between computing chips within the same fusion node in this case is described.
[0220] Figure 8 This is a flowchart illustrating the steps of a communication method provided in this application in another application scenario. Figure 8 In the application scenario shown, the computing chip at the transmitting end and the computing chip at the receiving end are connected to different first switching chips, such as... Figure 8 As shown, the method may include the following steps:
[0221] Step S810: Computing chip 1 sends a third data packet to the first switching chip 1, wherein the destination address of the third data packet is the address of computing chip 4, computing chip 4 and computing chip 1 are computing chips within the fusion node 1, and computing chip 1 is connected to the first switching chip 1, and computing chip 4 is connected to the first switching chip 2.
[0222] Step S820: The first switching chip 1 forwards the third data packet to the second switching chip 1 of the switching node 1.
[0223] It should be understood that there is no communication connection between the first switching chip 1 and the second switching chip 2. However, each second switching chip can establish a communication connection with each first switching chip. The first switching chip can forward the third data packet to the second switching chip 1, and then the second switching chip 1 forwards the third data packet to the second switching chip 2. In specific implementation, the first switching chip 1 can send the third data packet to the switching node 1. The switching node 1 determines which second switching chip 1 to use to forward the third data packet based on the idle status of each second switching chip.
[0224] Step S830: The second switching chip 1 forwards the third data packet to the first switching chip 2.
[0225] Step S840: The first switching chip 2 forwards the third data packet to the computing chip 4.
[0226] In a specific implementation, the aforementioned first switching chip 1 can be Figure 3 The first switching chip L12 and the first switching chip 2 can be Figure 3 The first switching chip L11 in the process, and the computing chip 1 can be Figure 3 The computing chip C1 and computing chip 4 in the middle can be Figure 3 The computing chip C4 and the second switching chip 1 can be... Figure 3 The second switching chip L23 in the above steps S810 to S840 can be found in the description. Figure 3 The steps and flow of data communication between computing chip C1 and computing chip C4 in the embodiment will not be repeated here.
[0227] In simple terms, for data communication between computing chips within the same fusion node, if the sending computing chip and the receiving computing chip are connected to the same first switching chip, then data communication between the sending and receiving ends can be achieved through this first switching chip. If the sending computing chip and the receiving computing chip are connected to different first switching chips, then the sending end can forward data packets to a second switching chip through the first switching chip connected to it, and the second switching chip will then forward the data packets to the first switching chip connected to the receiving end, and then forward the data packets to the receiving end through the first switching chip connected to the receiving end.
[0228] It should be understood that Figure 7 and Figure 8 The communication method described herein is for a computer cluster that is a two-layer fat-tree network. If the computer cluster forms a fat-tree network with more than two layers, for example... Figure 4 The four-layer fat tree network shown can also achieve data communication between different cabinets through a third switching chip.
[0229] Still with Figure 4 For example, in this embodiment, if the computing chip 1 in the first box 1 of the first rack 1 sends a data packet to the computing chip 1 in the first box 1 of the first rack m, the data packet needs to pass through the first switching chip L1-1 in the first box 1 of the first rack 1, then through the second switching chip L2-1 in the first rack 1 (or any other second switching chip in the same rack, which is not specifically limited in this application), then through the third switching chip L3-1 in the second rack 1 (or any third switching chip in any second rack), then through the second switching chip L2-1 in the second box 1 of the first rack m (or any second switching chip in the first rack m), and finally through the first switching chip L1-1 in the first box 1 of the first rack m, and finally be transmitted to the computing chip 1 in the first box 1 of the first rack m.
[0230] In summary, if there is a directly connected first switching chip between the transmitting and receiving computing chips, then they can communicate through this first switching chip. If there is no directly connected first switching chip between them, then they can communicate through a second switching chip directly connected to their respective first switching chips. If there is no directly connected second switching chip between the two first switching chips, then the two second switching chips connected to each of the two first switching chips can be identified, and they can communicate through a third switching chip connected to these two second switching chips. And so on. Examples are not provided here.
[0231] Furthermore, if the computer cluster comprises multiple switching planes, refer to Figure 5 and Figure 6 As illustrated in the embodiments, each switching plane has a corresponding first switching chip and a second switching chip. If the network layer of the computer cluster is two or more layers, each switching plane also has a corresponding third switching chip. When computing chips use different switching planes for data communication, they can achieve data communication through the first, second, and third switching chips corresponding to that switching plane. The communication method can be found in [reference needed]. Figure 7 and Figure 8 The relevant descriptions of the communication methods shown will not be repeated here.
[0232] In practical implementation, if a user requires multiple switching planes for network isolation, then the switching chips of different switching planes are not connected and do not share resources. If the user does not require network isolation but only wants different switching planes to handle different services or use different communication protocols, then the switching chips of different switching planes can be connected or shared. The data packets generated by the computing chip can carry the identifier of the switching plane. For example, when using the first switching plane for data communication, the data packets generated by the computing chip can carry the identifier of the first switching plane. In this way, when the first switching chip receives a data packet, it can determine whether to forward it based on the identifier. If the switching chip of the second switching plane receives a data packet carrying the identifier of the first switching plane, then the switching chip does not need to process the data packet.
[0233] It should be noted that when the computing system is a chassis-based server, the communication process within the computing system can be referenced. Figure 7 and Figure 8 The description of the embodiments will not be repeated here.
[0234] In summary, this application provides a communication method applied in a computer cluster. The computer cluster includes multiple first racks, each first rack including multiple converged nodes and multiple switching nodes. Each first converged node includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to establish communication connections between the multiple computing chips. The first switching node is used to establish communication connections between the first converged node and other converged nodes. The first switching node is coupled to the first converged node via a connector, and the multiple computing chips in the first converged node are connected to the first switching chip. Thus, even if the number of computing chips connected to the first switching chip decreases, the total number of computing chips in the rack can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth available to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0235] Figure 9 This is a schematic diagram of the structure of a computing system provided in this application. The computing system 1000 can be... Figure 1 and Figure 8 The computing system 1000 described in the embodiment may include switching nodes 210 and converged nodes 110. The switching node 210 includes a second switching unit 221, and the converged node 110 includes a computing unit 131 and a first switching unit 121. Multiple converged nodes 110 are coupled to multiple switching nodes 210 via connectors. Figure 9 Although the computing system only draws one exchange node 210 and one fusion node 110, in the actual implementation, the number of fusion nodes 110 and exchange nodes 210 can be multiple, and this application does not make a specific limitation.
[0236] The computing unit 131 of the fusion node 110 is used to process computing tasks and generate data packets. The source address of the data packet is the address of the first computing chip within the fusion node, and the destination address is the address of the second computing chip. The first switching unit 121 of the fusion node 110 forwards the data packet to the destination address through the first switching chip within the fusion node when the second computing chip is a computing chip within the fusion node. The first switching unit 121 is used to forward the data packet to the first switching node when the second computing chip is a computing chip outside the fusion node, so that the first switching node can forward the data packet to the destination address.
[0237] In a specific implementation, the computing unit 131 can... Figures 1 to 8 The computing chip 130 in the embodiment is implemented, and the computing unit 131 is executable. Figure 7Steps S710 and S730 in the embodiment and Figure 8 Step S810 in the embodiment. The first switching unit 121 can be accessed via... Figures 1 to 8 The first switching chip 120 in the embodiment is implemented. The first switching unit 121 is executable. Figure 7 Steps S720, S740, and S760 in the embodiment can also be executed. Figure 8 Steps S820 and S840 in the embodiment. The second switching unit 221 can be accessed via... Figures 1 to 8 The second switching chip 220 in the embodiment is implemented, and the second switching unit 221 is executable. Figure 7 Step S750 in the embodiment and Figure 8 Step S830 in the embodiment.
[0238] The second switching unit 221 of the switching node 210 is used to receive data packets sent by the first switching chip of the first fusion node among multiple fusion nodes 110. The source address of the data packet is the address of the first switching chip of the first fusion node, and the destination address is the address of the second computing chip in the second fusion node. The second switching unit 221 is used to forward the data packets to the first switching chip of the second fusion node 110, so that the first switching chip can forward the data packets to the second computing chip. The second fusion node 110 includes the first switching chip and the computing chip.
[0239] In one possible implementation, the connection between the fusion node 110 and the switching node 210 is an orthogonal connection, and the connector includes a backplane-less orthogonal connector or an optical blind mating connector.
[0240] In one possible implementation, the connector includes a high-speed connector that achieves an orthogonal connection between the fusion node and the first switching node by twisting a 90-degree angle.
[0241] In one possible implementation, the converged node 110 includes a first switching chip for a first switching plane and a first switching chip for a second switching plane; a first switching unit 121 is used to forward data packets through the first switching chip of the first switching plane when the data packet carries the identifier of the first switching plane; and a first switching unit 121 is used to forward data packets through the first switching chip of the second switching plane when the data packet carries the identifier of the second switching plane.
[0242] In one possible implementation, the first switching plane and the second switching plane are network-isolated switching planes; or, the first switching plane and the second switching plane carry out different services; or, the first switching plane and the second switching plane use different communication protocols.
[0243] In one possible implementation, the network topology between multiple fusion nodes 110 and multiple switching nodes 210 includes a network topology without a central switching node and a network topology with a central switching node. The network topology without a central switching node may include, but is not limited to, dragonfly networks, dragonfly+ networks, torus networks, etc., while the network topology with a central switching node may include, but is not limited to, fat tree networks. For details regarding the network topologies without and with central switching nodes, please refer to the preceding descriptions; they will not be repeated here.
[0244] In one possible implementation, when the network topology is a network topology without a central switching node, the switching node is used to realize the communication connection between multiple computing systems; when the network topology is a network topology with a central switching node, the communication connection between multiple computing systems is realized through the central switching node.
[0245] In one possible implementation, the number of converged nodes and switching nodes in the computing system is determined based on the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
[0246] In summary, this application provides a computing system comprising multiple converged nodes and multiple switching nodes. A first converged node among the multiple converged nodes includes multiple computing chips and at least one first switching chip. The at least one first switching chip is used to establish communication connections between the multiple computing chips. The first switching node is used to establish communication connections between the first converged node and other converged nodes among the multiple converged nodes. The first switching node is coupled to the first converged node via a connector, and the multiple computing chips in the first converged node are connected to the first switching chip. Thus, even if the number of computing chips connected to the first switching chip decreases, the total number of computing chips in the rack can be increased by increasing the number of first switching chips. Conversely, reducing the number of computing chips connected to each first switching chip increases the bandwidth available to each computing chip. This expands the scale of the computer cluster while increasing the available bandwidth of each computing chip, thereby solving the performance bottleneck problem of the computer cluster.
[0247] Figure 10 This is a schematic diagram of the structure of a computing system provided in this application. The computing system can be... Figures 1 to 9 The first cabinet 400 in the embodiment includes a plurality of computing devices 1000 and a plurality of communication devices 2000.
[0248] Furthermore, the computing device 1000 includes a computing chip 1001, a storage unit 1002, a storage medium 1003, a communication interface 1004, and a first switching chip 1007. The computing chip 1001, the storage unit 1002, the storage medium 1003, the communication interface 1004, and the first switching chip 1007 communicate via a bus 1005, and also via other means such as wireless transmission.
[0249] The computing chip 1001 comprises at least one general-purpose processor, such as a CPU, NPU, or a combination of a CPU and a hardware chip. The aforementioned hardware chip is an Application-Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), or a combination thereof. The aforementioned PLD is a Complex Programmable Logic Device (CPLD), a Field-Programmable Gate Array (FPGA), a Generic Array Logic (GAL), or any combination thereof. The computing chip 1001 executes various types of digital storage instructions, such as software or firmware programs stored in the storage unit 1002, enabling the computing device 1000 to provide a wide range of services.
[0250] In a specific implementation, as one example, the computing chip 1001 includes one or more CPUs, for example... Figure 10 CPU0 and CPU1 are shown in the diagram.
[0251] In a specific implementation, as one example, the computing device 1000 also includes multiple computing chips, for example... Figure 10 The computing chips 1001 and 1006 are shown. Each of these computing chips can be a single-core processor or a multi-core processor. Here, a processor refers to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0252] The first switching chip 1007 is the switching chip of the switch, and it consists of at least one general-purpose processor. This switching chip can be an ASIC, a PLD, or a combination thereof. The PLD can be a CPLD, an FPGA, a GAL, or any combination thereof. The first switching chip 1007 can execute software or firmware programs stored in the storage unit 1002 to realize data communication between computing chips within the computing device 1000.
[0253] In a specific implementation, as one example, the computing device 1000 also includes at least one first switching chip, for example... Figure 6 The first switching chip 1007 and the first switching chip 1008 are shown.
[0254] The storage unit 1002 is used to store program code, and its execution is controlled by the computing chip 1001 to perform the above-mentioned tasks. Figures 1-9 The processing steps of the computing chip are described in any embodiment. The program code includes one or more software units, which are... Figure 9 The computing unit in this embodiment is used to process computing tasks and generate data packets. The computing unit is used to execute... Figure 7 Steps S710 and S730 and their optional steps in the embodiments, and Figure 8 Step S810 and its optional steps in the embodiments.
[0255] The storage unit 1002 is also used to store program code, which is controlled by the first switching unit 1007 to execute the above-mentioned code. Figures 1-9 The processing steps of the first switching chip in any embodiment. The program code includes one or more software units, which are... Figure 9 The first switching unit in this embodiment is configured to forward a data communication request to the destination address when the destination address of the data packet is the address of a computing chip within the fusion node; and to forward the data packet to the switching node when the destination address of the data packet is the address of a computing chip outside the fusion node, so that the switching node can forward the data packet to the destination address. The first switching unit is configured to perform... Figure 6 Steps S720, S740, and S760, and their optional steps, in the embodiments, Figure 8 Steps S820, S840, and their optional steps in the embodiments will not be described again here.
[0256] Furthermore, the communication device 2000 includes a second switching chip 2001, a storage unit 2002, a storage medium 2003, and a communication interface 2004. The second switching chip 2002, the storage unit 2002, the storage medium 2003, and the communication interface 2004 communicate via a bus 2005, and also via other means such as wireless transmission.
[0257] The second switching chip 2002 is the switching chip of the switch, which consists of at least one general-purpose processor. This switching chip can be an ASIC, a PLD, or a combination thereof. The aforementioned PLD can be a CPLD, an FPGA, a GAL, or any combination thereof.
[0258] In a specific implementation, as one example, the communication device 2000 includes a plurality of second switching chips 2002, for example... Figure 10 The second switching chip 2002 and the second switching chip 2006 are shown in the figure.
[0259] The storage unit 2002 is used to store program code, and its execution is controlled by the second switching chip 2002 to perform the above-mentioned tasks. Figures 1-9 The processing steps of the second switching chip in any embodiment. The program code includes one or more software units, which are... Figure 9 The second switching unit in this embodiment is configured to receive data packets sent by the first switching chip of the first fusion node among multiple fusion nodes, and forward the data packets to the first switching chip of the second fusion node, so that the first switching chip of the second fusion node can forward the data packets to the second computing chip. The source address of the data packet is the address of the first computing chip of the first fusion node, and the destination address is the address of the second computing chip of the second fusion node. The second switching unit is configured to perform... Figure 7 Steps S740, S750, and their optional steps in the embodiments are also used to perform Figure 8 Step S830 and its optional steps in the embodiments will not be described again here.
[0260] Storage units 1002 and 2002 include read-only memory, random access memory, volatile memory, or non-volatile memory, or both. The non-volatile memory is read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory is random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM). Other examples include hard disks, USB flash drives, flash memory, SD cards, Memory Sticks, etc., where hard disks include hard disk drives (HDDs), solid-state drives (SSDs), and mechanical hard disks (HDDs), etc., which are not specifically limited in this application.
[0261] Storage media 1003 and storage media 2003 are carriers for storing data, such as hard disk, USB flash drive, flash memory, SD card, memory stick, etc. The hard disk can be a hard disk drive (HDD), solid state disk (SSD), mechanical hard disk (HDD), etc., and this application does not make specific limitations.
[0262] Communication interfaces 1004 and 2004 are wired interfaces (e.g., Ethernet interfaces), internal interfaces (e.g., Peripheral Component Interconnect Express (PCIe) bus interfaces), wired interfaces (e.g., Ethernet interfaces), or wireless interfaces (e.g., cellular network interfaces or wireless LAN interfaces) for communicating with other servers or units.
[0263] Bus 2005 and Bus 1005 are Peripheral Component Interconnect Express (PCIe) buses, or Extended Industry Standard Architecture (EISA) buses, Unified Bus (Ubus or UB), Compute Express Link (CXL), Cache Coherent Interconnect for Accelerators (CCIX), etc. Bus 2005 and Bus 1005 are divided into address bus, data bus, and control bus.
[0264] In addition to the data bus, buses 2005 and 1005 also include power buses, control buses, and status signal buses. However, for clarity, all buses are labeled as bus 2005 and bus 1005 in the diagram.
[0265] In this embodiment, the first switching chip 1007 in the computing device 1000 is coupled to the second switching chip 2001 in the communication device 2000 via a connector. Each second switching chip 2001 and each first switching chip 1007 establish a communication connection through an orthogonal connection, as detailed above. The connectors include, but are not limited to, backplane-less orthogonal connectors, optical blind-mating connectors, and high-speed connectors. The high-speed connector requires first establishing a flexible connection between the computing device 1000, the communication device 2000, and the high-speed connector using cables, and then rotating the high-speed connector 90 degrees to achieve an orthogonal connection between the switching node 210 and the fusion node 110.
[0266] It needs to be explained that, Figure 10 This is merely one possible implementation of an embodiment of this application. In actual applications, the computing device 1000 may include more or fewer components, which is not limited here. For content not shown or described in the embodiments of this application, please refer to the foregoing. Figures 1-9The relevant descriptions in the embodiments will not be repeated here.
[0267] This application provides a computer cluster, including multiple Figure 10 The computing system shown communicates with multiple computing systems to collaboratively process tasks. Each computing system can establish a communication connection with a central switching node (such as the second rack mentioned earlier). Each central switching node is used to implement communication connections between the computing systems. The network topology between the multiple computing systems and the multiple central switching nodes can be a network topology with a central switching node as described earlier, such as a fat-tree network. For any content not shown or described in the embodiments of this application, please refer to the foregoing... Figures 1-9 The relevant descriptions in the embodiments will not be repeated here.
[0268] This application provides another computer cluster, including multiple Figure 10 The computing system shown communicates with multiple computing systems to collaboratively process tasks. Communication between each computing system is achieved through a switching node within each system. The network topology between the multiple computing systems can be a centrally located switching node-less network topology as described above, such as the Dragonfly network, Dragonfly+ network, or Torus network, etc. For details regarding anything not shown or described in the embodiments of this application, please refer to the foregoing... Figures 1-9 The relevant descriptions in the embodiments will not be repeated here.
[0269] This application embodiment also provides a computing device, which includes a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to implement, for example... Figures 1-9 The embodiments describe the functions of the operation steps performed by the fusion node.
[0270] The above embodiments are implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments are implemented, in whole or in part, in the form of a computer program product. A computer program product includes at least one computer instruction. When the computer program instruction is loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer is a general-purpose computer, a special-purpose computer, a computer network, or other programming device. The computer instructions are stored in a computer read-only storage medium or transmitted from one computer read-only storage medium to another, for example, computer instructions are transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer read-only storage medium is any medium that a computer can access or a data storage node such as a server or data center that contains at least one set of media. The media is a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., high-density digital video disc (DVD), or a semiconductor medium. A semiconductor medium is an SSD.
[0271] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent repairs or substitutions within the technical scope disclosed in the present invention, and such repairs or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A rack-mount server, characterized in that, The rack-mount server includes at least one converged node and at least one switching node; The at least one fusion node includes at least one computing chip and at least one first switching chip, the at least one computing chip and the at least one first switching chip are in the same housing, and the at least one first switching chip is used to realize the communication connection between the at least one computing chip; The at least one switching node includes a first switching node, which is used to implement communication connections between the first fusion node and other fusion nodes. The at least one fusion node includes the first fusion node. The coupling method between the first fusion node and the first switching node is an orthogonal connection, and the first switching node is coupled to the first fusion node through a backplane-less connector.
2. The rack-mount server according to claim 1, characterized in that, The at least one first switching chip includes a third switching chip and a fourth switching chip, a portion of the at least one computing chip is connected to the third switching chip, and the other portion of the at least one computing chip is connected to the fourth switching chip.
3. The rack-mount server according to claim 1, characterized in that, The at least one first switching chip includes a third switching chip, and each of the at least one computing chip is connected to the third switching chip.
4. The rack-mount server according to claim 3, characterized in that, The connectors include backplate-less orthogonal connectors or optical blind mating connectors.
5. The rack-mount server according to claim 3, characterized in that the connector includes a high-speed connector, wherein the high-speed connector achieves an orthogonal connection between the first fusion node and the first switching node by twisting a 90-degree angle.
6. The rack-mount server according to any one of claims 1-5, characterized in that, The number of the at least one converged node and the at least one switching node is determined based on at least one of the bandwidth requirements of the computing chip, the number of ports and switching capacity of the first switching chip, and the number of ports and switching capacity of the first switching node.
7. The rack-mount server according to any one of claims 1-5, characterized in that, The at least one first switching chip includes a first switching chip of a first switching plane and a first switching chip of a second switching plane. The first switching node includes a plurality of second switching chips, which include the second switching chips of the first switching plane and the second switching plane. The first switching plane and the second switching plane undertake different services.
8. A computer cluster, characterized in that, The computer cluster includes a plurality of rack-mounted servers as described in any one of claims 1-7, and the plurality of rack-mounted servers communicate with each other through a switching node.
Citation Information
Patent Citations
Backplane connecting system and method for blade server
CN104951022A
Multidimensional switch network
US20050195808A1