Communication board card, topological structure and communication system thereof
By using high-performance electrical switching chips and optical modules without switching capabilities in communication boards, the backup lines and topological structures in L lines are designed, and the failure rate and cost of high bandwidth interconnection domains are solved, and the high availability of data center-scale cluster interconnection is achieved, reducing resource waste and cost.
Patent Information
- Application Number
- CN202510222489.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The existing high-bandwidth interconnection domain has high problems in terms of failure rate and cost, which leads to the communication performance of the entire supernode being affected when the computing card fails, and the entire supernode needs to be removed from the system for repair, resulting in wasted resources.
A communication board is designed, using a high-performance electrical switching chip and an optical module without switching capabilities. The expansion card is equipped with L lines, at least one of which is used to form a backup line, supporting dynamic loops within and between machines, reducing the risk of single point failure of network connections.
By improving the availability and flexibility of network connections, it reduces failure rates and resource waste, and supports highly available point-to-point data center scale cluster interconnection, effectively reducing costs.
Smart Images

Figure CN120075164A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular, to a communication board, a topology structure, and a communication system thereof. Background Art
[0002] As the scale and complexity of artificial intelligence models (hereinafter referred to as models) continue to increase, the computational amount of the models grows exponentially, posing higher requirements for hardware computing power. If the hardware computing power cannot keep up with the growth rate of the computational amount, the training time of the model will become extremely long, and the inference speed will also decrease, thus unable to meet the needs of practical applications. Based on this, AI accelerators have become a key technology to address large-scale computing requirements. An AI accelerator is a hardware device or system specifically designed to accelerate artificial intelligence and machine learning tasks. By optimizing the computing architecture and hardware design, it significantly improves the operating efficiency of AI algorithms. Among them, the GPU, as a typical representative of AI accelerators, is widely used in data centers and edge devices due to its powerful parallel computing capabilities and mature ecosystem.
[0003] However, in large-scale artificial intelligence computing scenarios, the computing power of a single GPU is already difficult to meet the requirements. For example, in the case of a traditional single-machine 8-card GPU configuration, both performance and scalability face bottlenecks when dealing with complex tasks. If one attempts to build a larger-scale GPU supernode (such as a single supernode containing 32 or 64 GPUs) to handle the computing requirements, due to the throughput limitations of PCIe and R-NIC, the data transfer bandwidth will become a bottleneck, resulting in severe communication latency, thereby significantly reducing the computing power utilization rate. To solve this problem, in related technologies, high-bandwidth technologies that have already achieved efficient data transfer within a single node are extended outside the machine, enabling high-speed data transfer channels to be established between multiple nodes. That is, through this GPU supernode interconnection topology, supernodes containing more intelligent computing chips can be constructed, thereby greatly enhancing the computing power and utilization rate of the cluster.
[0004] In this regard, the inventors found that there are at least the following technical problems in the related technologies:
[0005] Under the existing supernode design architecture, a computing card failure can lead to communication problems within the supernode, thereby affecting the performance of the entire supernode. Specifically, if any one of the computing cards within the supernode fails during task execution, the communication between the remaining computing cards will be disrupted and unable to work in the original efficient manner, so they will be forced to reduce their operating speed. Therefore, in order to ensure that the overall performance of the system is not greatly affected, measures need to be taken to remove the entire supernode from the system and then replace it with a backup supernode. After replacement, the faulty module can be repaired offline. For example, take the TP32 supernode which contains 32 computing cards. The failure of one computing card may cause serious communication problems in the entire supernode. To ensure the stable operation of the system, in addition to the faulty computing card, 31 normal computing cards need to be removed. These normal computing cards can actually still work properly, but due to the limitations of the system architecture, they cannot be effectively utilized in the faulty supernode environment, which will bring extremely high costs.
[0006] In summary, in the related art, the failure rate and cost of the high-bandwidth interconnection domain are relatively high. Summary of the Invention
[0007] An object of the present application is to provide a communication board card, a topology structure, and a communication system thereof, at least to solve the technical problems of relatively high failure rate and high cost in the high-bandwidth interconnection domain in the related art.
[0008] To achieve the above object, some embodiments of the present application provide the following aspects:
[0009] In a first aspect, some embodiments of the present application further provide a communication board card. The communication board card includes: a plurality of computing cards, each computing card is used to execute a computing task; a plurality of switching modules, the switching module includes a high-performance electrical switching chip and a first optical module; each switching module is connected to at least two computing cards to realize data exchange between the computing cards; the first optical module is an optical module without switching ability; the throughput of the switching module is 2 to L times the throughput of the computing card, where L is a positive integer; each switching module is configured with an expansion card, and each expansion card is provided with L lines; among the L lines, at least one line is used to form a backup line, and one line is used to form a topology structure.
[0010] In a second aspect, some embodiments of the present application further provide a topology structure, which is specifically constructed by using the high-speed communication board card and the server as described above; wherein, the topology structure includes at least one of a one-dimensional linear topology structure, a two-dimensional mesh topology structure, parallel lines, and a full-connection topology structure.
[0011] In a third aspect, some embodiments of the present application also provide a communication system, which includes a topology structure as described above, as well as a controller and a task scheduler; the controller is used to maintain topology information of the topology structure, and update the server status when the server is identified as a fault; the task scheduler is used to calculate the target topology based on the updated server status, and adjust the expansion card status according to the target topology to change the connection method of the server in the topology structure.
[0012] Compared with the related art, the solution provided in the embodiment of the present application cleverly improves the design of the traditional communication board card of the computing card. By using a high-performance electrical switching chip to replace the traditional double electrical switching chip, and combining the first optical module, the throughput of the high-performance electrical switching chip is increased to 2 to L times the throughput of the computing card, thereby expanding the throughput of the switching module; since each high-performance electrical switching chip is configured with an expansion card, each expansion card can be provided with L lines; among the L lines, at least one line is used to form a backup line, thereby improving the availability of network connections, and one line is used to form a topological structure, thereby supporting dynamic ring formation within and between machines, without relying on ring-shaped external interconnection, so it is conducive to flexible ring design and obtains innovative topological structures. In addition, the technical solution provided in this embodiment can support high-availability point-to-point (P2P) data center scale cluster interconnection, and can effectively support large-scale deployment, which can effectively reduce the failure rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0014] Figure 1 An exemplary schematic diagram of a communication board provided according to some embodiments of the present application;
[0015] Figure 2 is an exemplary schematic diagram of a traditional communication board in the related art;
[0016] Figure 3 An exemplary schematic diagram of an expansion card design based on the second optical module provided according to some embodiments of the present application;
[0017] Figure 4 An exemplary schematic diagram of a backup line in a communication board provided according to some embodiments of the present application;
[0018] Figure 5An exemplary schematic diagram of a physical loop design based on a communication board according to some embodiments of the present application;
[0019] Figure 6 Another exemplary schematic diagram of a physical loop design based on a communication board according to some embodiments of the present application;
[0020] Figure 7 An exemplary schematic diagram of a one-dimensional linear topology structure according to some embodiments of the present application;
[0021] Figure 8 An exemplary schematic diagram of a two-dimensional mesh topology structure according to some embodiments of the present application;
[0022] Figure 9 An exemplary schematic diagram of a parallel line and full-connection topology structure according to some embodiments of the present application;
[0023] Figure 10 Another exemplary schematic diagram of a parallel line and full-connection topology structure according to some embodiments of the present application;
[0024] Figure 11 Another exemplary schematic diagram of a parallel line and full-connection topology structure according to some embodiments of the present application. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0026] The following terms are used herein.
[0027] Computing power utilization rate: The full English name is Model FLOPs Utilization, abbreviated as MFU.
[0028] PCIE: The full English name is Peripheral Component Interconnect Express, which is a high-speed serial computer expansion bus standard.
[0029] R-NIC: The full English name is Remote Network Interface Controller, a remote network interface controller.
[0030] Intelligent computing chip: A dedicated chip used in an intelligent computing center (AIDC), different from the CPU chips in traditional general-purpose servers. Intelligent computing chips include types such as GPUs, FPGAs, and ASICs, which have excellent parallel computing capabilities and are particularly suitable for processing a large number of simple matrix operation tasks in AI algorithms.
[0031] Electric switch: The full English name is Electric Packet Switch, abbreviated as EPS.
[0032] Optical circuit switch: The full English name is Optical Circuit Switch, abbreviated as OCS.
[0033] High Bandwidth Domain: Abbreviated as "HBD" in English, is a group of intelligent computing systems interconnected by Ultra-Bandwidth (HB). The communication bandwidth between intelligent computing chips within an HBD is several times that between intelligent computing chips across different HBDs. By enhancing the communication rate between GPUs, HBD has become a key direction for improving computing power requirements. In model parallel processing, restricting the large and non-overlapping parts of data within one HBD can effectively improve computing efficiency. Through high-speed switching technology, HBD can be extended from within a single machine to between multiple machines, supporting the expansion of more GPUs and building an efficient computing power cluster. The core requirements of HBD include ultra-high bandwidth, high energy efficiency ratio, and low total cost of ownership.
[0034] UBB: The full English name is Open Compute Project Universal Baseboard, and 2.0 is a universal baseboard standard for open computing projects. Switch-Based design refers to a network architecture based on switches, which realizes the connection and data exchange between each node through switches. In this design, the switch plays a key role and can provide flexible network configuration and efficient data transmission.
[0035] Expansion Card: That is, an expansion card, usually used to expand the functions of a computer system. In the fields of servers and high-performance computing, expansion cards can be used to add functions such as network interfaces, storage interfaces, and computing cards.
[0036] Optical interconnection module: A module that uses optical signals for data transmission.
[0037] Reconfigurable optical interconnection module: Based on the optical interconnection module, it adds configurability and can flexibly adjust the optical signal transmission path and method according to different application requirements.
[0038] Ring-AllReduce: A commonly used communication mode in distributed training, used to efficiently exchange and update data between multiple computing nodes.
[0039] Time Division Multiplexing: Allows multiple signals to be transmitted on the same physical line in time slices.
[0040] First Embodiment
[0041] The first embodiment of the present application relates to a communication board. The communication board can be applied to, but not limited to, intelligent computing servers, so that the communication board can connect the server to a high-bandwidth interconnection domain. As Figure 1 shown, the communication board includes:
[0042] Multiple computing cards, each computing card is used to execute computing tasks;
[0043] Multiple switching modules, the switching module includes a high-performance electrical switching chip and a first optical module; each switching module is connected to at least two computing cards to implement data exchange between the computing cards; the first optical module is an optical module without switching ability; the throughput of the switching module is 2 to L times the throughput of the computing card, where L is a positive integer;
[0044] Each switching module is configured with an expansion card, and each expansion card is provided with L lines; among the L lines, at least one line is used to form a backup line, and one line is used to form a topology structure.
[0045] Specifically, the different computing cards are connected end to end, and the computing cards are used to execute computing tasks; the different switching modules are connected end to end, and the high-performance electrical switching chips in each switching module are respectively connected to at least two computing cards through the first optical module to implement data exchange between the computing cards; the first optical module is an optical module without switching ability; the throughput of the high-performance electrical switching chip is 2 to L times the throughput of the computing card, where L is a positive integer.
[0046] Specifically, each switching module is configured with an expansion card, and each expansion card is provided with L lines. Among the L lines, at least one line is used to form a backup line, and one line is used to form a topology structure.
[0047] It should be noted that for ease of understanding, this embodiment takes the intelligent computing server design standard widely used in the industry: OCP UBB 2.0 as an example to illustrate the communication board. However, the intelligent computing server design standard applicable to the communication board is not limited to OCP UBB 2.0, and this embodiment does not make specific limitations in this regard.
[0048] See Figure 2 shown, Figure 2In the related art, it is an exemplary schematic diagram of a traditional communication board under the OCP UBB 2.0 standard. In this example, the traditional communication board includes multiple computing cards, which are represented by OAM. Among them, there are a total of eight computing cards, which are respectively labeled as OAM0, OAM1, OAM2, OAM3, OAM4, OAM5, OAM6, and OAM7. Connections are established between different computing cards through colored lines. And, a single electrical switching chip is represented by Switch, and there are a total of eight electrical switching chips, which are respectively labeled from Switch0 to Switch7. Each electrical switching chip includes multiple ports; each electrical switching chip is connected to multiple computing cards through multiple lines. It can be seen that in the traditional technical solution as Figure 2 shown, an electrical switching chip equal to the number of computing cards is configured for each server, so that each computing card has a corresponding electrical switching chip to implement the communication function. And, a distributed switching device with a throughput twice that of the computing card can be configured on each electrical switching chip.
[0049] Combined with Figure 1 shown, the communication board provided in this embodiment may include: multiple computing cards, which are sequentially labeled as computing card 1, computing card 2, computing card 3... computing card N from left to right; these computing cards are used to execute specific computing tasks, such as matrix operations in deep learning, etc. Under the computing cards, there are multiple high-performance electrical switching chips of the switching module, and each high-performance electrical switching chip is configured with an expansion card through a first optical module; the first optical module is an optical module without switching ability. Multiple expansion cards are sequentially labeled as expansion card 1, expansion card 2... expansion card M from left to right. These expansion cards are used to connect the computing cards through the high-performance electrical switching chips and provide data exchange and transmission functions.
[0050] Furthermore, in this application, the throughput of the switching module is increased to 2 to L times the throughput of the computing card. For example, if the throughput of the computing card is 100 Gbps, then the throughput of the improved switching module can reach 300 Gbps to L×100 Gbps. By doing so, each expansion card can be externally connected to more lines (L lines). Specifically, each expansion card can be provided with L lines, which enables the communication board to connect more servers in scenarios such as data centers, thereby expanding the network connection ability. And, these lines can be time-division multiplexed, so as to further improve the utilization rate of the lines and achieve the purpose of transmitting more data without increasing physical lines.
[0051] Specifically, in this embodiment, the specific method of increasing the throughput of the switching module to 2 to L times the throughput of the computing card can be as follows: Replace the traditional single-rate electrical switching chip with a high-performance electrical switching chip to double the throughput of the high-performance electrical switching chip, and then double the number of the first optical modules connected to the high-performance electrical switching chip. Here, the first optical module is an optical module without switching capabilities. Here, doubling the number of the first optical modules means doubling based on the number of the first optical modules that can be connected by the traditional single-rate electrical switching chip.
[0052] Among them, one of the reasons for increasing the throughput of the switching module to 2 to L times the throughput of the computing card in this embodiment is that: One multiple of the throughput is used to form a topology structure to allow data to be transmitted on multiple paths and improve the reliability of the network. Another multiple of the throughput is used to form a backup line. In this way, when the main path fails, data can be transmitted through the backup line to ensure the continuity of communication. That is to say, more data transmission paths can be achieved through the increased throughput. For example, even if some computing cards (such as computing card 1 and computing card 2) fail, data can be transmitted through other paths (such as directly connected to computing card 3), thus bypassing the fault point. This design allows the system to remain operational when multiple components fail because there are multiple backup lines.
[0053] Exemplarily, the line can specifically be: a line formed by using optical fibers or cables. Here, since each line can support the complete throughput capacity of a single intelligent computing chip, the efficiency of data transmission can be ensured, the data transmission requirements of the intelligent computing chip during high-speed operation can be met, and data bottlenecks caused by insufficient line transmission capacity can be avoided. Combined Figure 1 with this, there are optical fibers labeled Fiber1, Fiber2... FiberL below the expansion card 1, and these optical fibers are used to transmit data between the expansion cards 1. In this way, the computing cards are connected and data is exchanged through the expansion cards. The expansion card 1 transmits data through optical fibers, which can ensure the data flow between each computing card. Through the collaborative work of the computing card and the expansion card 1 of the switching module, efficient data processing and operation are achieved.
[0054] In practical applications, the traditional single-rate electrical switching chip can be arbitrarily replaced with a high-performance electrical switching chip according to the specific requirements and design of the system. For example, the Figure 2 shown switch3 can also be replaced with the high-performance electrical switching chip. In this way, relevant personnel can decide which single-rate electrical switching chip to replace according to the location where the fault occurs, the topology structure of the system, and the convenience of maintenance and upgrade. However, it should be noted that in this embodiment, at least two single-rate electrical switching chips should be replaced with the high-performance electrical switching chip.
[0055] In short, the communication board provided in this embodiment improves the traditional communication board. Specifically, a high-performance electrical switching chip is used to replace the traditional single electrical switching chip, and combined with the improvement in the number of the first optical modules, the throughput of the high-performance electrical switching chip is increased to 2 to L times the throughput of the computing card, and each expansion card can externally connect more lines (L lines).
[0056] It can be understood that some technical solutions in the related art adopt a design based on a centralized electrical switch. A full-interconnection topology is built inside the supernode, so that each GPU is directly connected to other GPUs in the supernode, and the supernode is integrated into the existing data center architecture by using the traditional data center network for networking. However, since this design uses a large number of high-throughput switching chips with relatively high costs, the switching chips cannot be increased indefinitely to expand the bandwidth, resulting in a limited scale of the high-bandwidth domain. In addition, in this design, the failure of the server may affect the normal communication of the GPUs connected to it. Therefore, when the server fails, these GPUs that cannot work properly will have fragmentation problems, resulting in their inability to effectively participate in the computing tasks and reducing the resource utilization rate and computing efficiency of the entire system.
[0057] Some other technical solutions in the related art adopt a design based on a centralized optical path switch. This design connects 64 TPU computing cards through cables to form a 4x4x4 cube structure. This cube structure is a basic unit, and then 64 such cubes are connected to build a large-scale supernode containing 4096 TPUs. This supernode has powerful computing resources. However, this solution also has limitations. (1) The impact of centralized OCS failure is significant: Since the entire design is based on a centralized OCS, once this centralized OCS fails, it will affect all TPU cubes connected to it. Because the communication between all TPU cubes depends on this central switch, its failure will cause communication interruption, making a large number of TPUs unable to work properly, seriously affecting the operation of the entire supernode. (2) Performance degradation caused by a single chip failure: When a single TPUv4 chip fails, the system will try to reroute the traffic passing through this chip to other paths. However, this will cause bandwidth contention with the original traffic on the new path, thereby reducing the data transmission speed and resulting in performance degradation of the entire system; (3) It is impossible to completely avoid the internal fragmentation problem: Although this solution schedules in units of 64-card TPU cubes, it is still impossible to completely avoid the internal fragmentation problem. That is to say, during the resource allocation process, there will be some situations where TPU resources cannot be fully utilized, reducing the overall resource utilization rate. (4) Poor real-time performance: Currently, mainstream commercial optical path switches are usually based on piezoelectric and microelectromechanical principles. This technology takes several milliseconds to modify the path, resulting in poor real-time performance. This makes it impossible to adjust the optical path in a timely manner according to the actual situation during the task operation. Therefore, it is generally only configured during task initialization and used as a static path, limiting the adaptability and flexibility of the system in a dynamic environment.
[0058] The communication board card in the embodiment of the present application can make the switching function not concentrated on one device, allowing the internal and external throughput connection ratios to be adjusted through different external wiring methods. This enables the managers of the data center to flexibly allocate the bandwidth resources between the internal computing cards and between the server and external devices according to the actual business requirements and network conditions. For example, in some tasks that require a large amount of data to be transmitted externally, the external throughput connection ratio can be appropriately increased to meet the data output requirements; while in tasks mainly focused on local computing, the internal connection performance can be emphasized, and the performance of the entire system can be optimized through dynamic adjustment.
[0059] It is not difficult to find that, compared with the related art, in the solution provided by the embodiment of the present application, the design of the traditional communication board card of the computing card facing the outside is ingeniously improved. By using a high-performance electrical switching chip to replace the traditional single electrical switching chip and combining with the first optical module, the throughput of the high-performance electrical switching chip is increased to 2 to L times the throughput of the computing card, expanding the throughput of the switching module; since each high-performance electrical switching chip is configured with an expansion card, and each expansion card can be provided with L lines; among the L lines, at least one line is used to form a backup line, thereby improving the availability of the network connection, and one line is used to form a topology structure, thereby supporting in-machine and inter-machine dynamic ring formation without relying on a ring-shaped external interconnection, so it is beneficial to a flexible ring formation design and an innovative topology structure is obtained. In addition, the technical solution provided in this embodiment can support high-availability point-to-point (P2P) data center-scale cluster interconnection, and can effectively support large-scale deployment, and can effectively reduce the failure rate.
[0060] Second Embodiment
[0061] The second embodiment of the present application relates to a communication board card. The second embodiment is a technical solution parallel to the first embodiment. The difference is that: in the first embodiment, the switching module is specifically implemented by a high-performance electrical switching chip and a first optical module. In the second embodiment of the present application, the switching module is specifically implemented by a second optical module, and the second optical module is an optical module with switching capabilities. Compared with the first embodiment, considering that the high-performance electrical switching chip is not economical, the technical solution of the second embodiment of the present application can significantly reduce the hardware cost and volume.
[0062] Specifically, the improvement of this embodiment compared to the traditional communication board card as shown in Figure 2 is that: the second optical module is used to replace each single electrical switching chip (Switch0 to Switch7) in Figure 2 . Exemplarily, through the three interfaces characterized by OSFP0 to 3 in Figure 2 , the second optical module can be inserted to increase the throughput of each second optical module to twice the original.
[0063] Exemplarily, as shown in Figure 3 , it is a schematic diagram of the expansion card design based on the second optical module. It can be seen that the improved design can adapt to a smaller-grained switching module. Among them, the design diagram of the expansion card includes:
[0064] Optical fiber interface (206): The system includes two optical fiber interfaces, which are respectively used for inputting and outputting optical fiber signals.
[0065] On-board laser module (205): Connected to the fiber optic array (203), it provides a laser light source for optical signal transmission, and is connected to an external optical fiber through the fiber optic interface (206) to achieve the input and output of optical signals. There are two fiber optic interfaces on the expansion card, and each interface corresponds to a group of on-board laser modules and fiber optic arrays.
[0066] Fiber optic array (203): The on-board laser module (205) is connected to the reconfigurable on-board optical interconnection module (200) through the fiber optic array (203) for multiplexing and demultiplexing optical signals.
[0067] Reconfigurable on-board optical interconnection module (200): It includes a silicon photonics chip (201) and an electrical chip (202). The silicon photonics chip (201) and the electrical chip (202) are encapsulated on the packaging substrate (204). Moreover, the silicon photonics chip (201) and the electrical chip (202) are connected to the on-board laser module (205) through the fiber optic array (203) and are responsible for processing the interconnection and routing of optical signals.
[0068] Expansion card PCB board (207): It is the physical carrier of the entire expansion card and carries the above-mentioned modules and other components.
[0069] Voltage regulation modules (209) and (210): Used to provide appropriate voltages for different components on the expansion card to ensure the normal operation of each component.
[0070] Reconfigurable optical interconnection expansion card (100): The expansion card is connected to an external device or network through the high-speed interface (211) to achieve high-speed data transmission.
[0071] Retimers (208): There are four on the expansion card, used to shape and re-time signals to ensure the quality and stability of signal transmission.
[0072] High-speed interface (211): Used for high-speed data communication between the expansion card (100) and other devices.
[0073] The above design based on the expansion card can be applied to but not limited to: scenarios with extremely high requirements for data transmission speed and bandwidth such as high-performance computing data centers, supercomputers, and artificial intelligence training platforms. In these scenarios, traditional electrical signal transmission often cannot meet the data transmission requirements, while the expansion card in this embodiment can provide an efficient solution.
[0074] It is not difficult to find that in the embodiment of the present application, the switching module is specifically implemented by the second optical module, and the second optical module is an optical module with switching capabilities. Compared with the first embodiment, the technical solution of the second embodiment of the present application can significantly reduce the hardware cost and volume, so as to support smaller-grained switching modules and be compatible with general and new switching devices.
[0075] Third Embodiment
[0076] The third embodiment of the present application relates to a communication board. The third embodiment is an improvement based on the first embodiment. Specifically, in the third embodiment of the present application, some of the expansion cards in the expansion card, together with the switching module connected to the part of the expansion card, are replaced with a PCB board. By doing so, the cost can be further reduced.
[0077] Specifically, in some examples, some of the expansion cards in the expansion card are replaced with a PCB board, that is, a solution of mixing expansion cards with different configurations is used. In this embodiment, some of the expansion cards, such as M expansion cards (at least two) are configured with a switching module, and the rest use a board card only containing PCB connections. In this way, the cost can be further reduced. For those connections with low demand for switching functions, a board card only containing PCB connections can be used, and the cost of the PCB board is relatively low. For key connections that require switching functions, expansion cards configured with a switching module are used, which optimizes the cost while ensuring the system performance. For example, Figure 1 the expansion card 2 in
[0078] It should be noted that this embodiment can also be an improvement based on the second embodiment.
[0079] It is not difficult to find that in the embodiments of the present application, by replacing some of the expansion cards in the expansion card, together with the switching module connected to the part of the expansion card, with a PCB board, the purpose of further reducing the cost can be achieved.
[0080] Fourth Embodiment
[0081] The fourth embodiment of the present application relates to a communication board. This embodiment aims to elaborate in detail on the implementation principle, especially the physical loop design principle, around the beneficial effects of the communication board.
[0082] Specifically, the communication board provided in this embodiment has at least the following beneficial effects:
[0083] (1) In a traditional point-to-point (P2P) connection, if a certain server fails, the normal servers connected to it will have a reduced communication ability due to the change of the communication link. The present application cleverly utilizes the ability of the switching device to add a backup line. When the server fails, the switching device can automatically switch the communication traffic to the backup line. In this way, the communication of the connected normal servers is not affected, and the degradation of the communication ability is avoided. The design of this backup line greatly improves the fault tolerance and reliability of the system, ensuring that the data center can still operate stably and efficiently when some servers fail.
[0084] As Figure 4 shown, GPU0, GPU1, and GPU2 are connected through switching chips (dOCS0, dOCS1, dOCS2, dOCS3). The specific connections are as follows:
[0085] GPU0 is connected to dOCS1 through dOCS0.
[0086] GPU1 is connected to dOCS2 through dOCS1.
[0087] GPU2 is connected to dOCS3 through dOCS2.
[0088] In addition, a backup line is designed: there is a backup line between dOCS0 and dOCS3.
[0089] When GPU1 fails, communication can be ensured in the following way: the backup line between dOCS0 and dOCS3 comes into play, bypassing the faulty GPU1, thus maintaining communication between GPU0 and GPU2.
[0090] (2) Traditional high-speed domain designs usually rely on a central switch to achieve communication between servers, and this approach has limitations. The central switch may become a performance bottleneck, and once it fails, it will affect the communication of the entire data center. When building a large-scale data center, the addition and management of servers are not convenient enough. However, this solution can use the switching capabilities of communication boards to replace the central switch in traditional high-speed domain designs. A smart computing server cluster connecting the entire data center can be built only through point-to-point (P2P) connections. Specifically, the centralized switching capabilities in the traditional cluster are dispersed to each expansion card, realizing the transformation from centralized switching to distributed switching. Distributed switching allows each server to directly communicate with other servers without relying on a central switch, which is beneficial to reducing the impact caused by the failure of a single switching chip and reducing the failure rate.
[0091] Adopting point-to-point connections is more direct and flexible, improving communication efficiency, enhancing the reliability and scalability of the system, and effectively avoiding communication interruption problems caused by the failure of the central switch. When building a large-scale data center, it is more convenient to add and manage servers.
[0092] (3) In the traditional network architecture, the ring connection between racks requires additional wiring, which not only increases costs and complexity but also is prone to wiring failures. Moreover, in the traditional configuration, the connection between two GPUs is point-to-point. To achieve collective communication, it often relies on a centralized switch, and the cluster also needs to be connected in a ring structure. However, this application has made significant improvements by leveraging the functions of the switching module. In this solution, each optical module on each server, or rather, the cable connected to each GPU, has switching capabilities, allowing a group of servers to perform full-throughput Ring-AllReduce communication while being linearly connected. This enables any two interconnected servers to form a ring network covering all computing cards for cluster communication, eliminating the need for additional inter-rack ring wiring and no longer relying on a centralized switch.
[0093] The specific principle is to add an optical module with switching capabilities between the traditional point-to-point connections of two GPUs. Each switching module can be connected to multiple GPUs, thus forming a loop at the switching module. This ring network belongs to a common network topology structure, where all nodes (servers) are sequentially connected to form a closed loop. By constructing a ring network by the server itself, the network structure and wiring are simplified, the flexibility of the network, the reliability and efficiency of communication are improved, and the network construction and maintenance costs are reduced.
[0094] In some examples, the design scheme for constructing the in-machine physical ring is as follows:
[0095] a Computing card connection: At least two computing cards are connected to each switching module. These two computing cards can communicate with other external computing cards simultaneously or directly with each other.
[0096] b Formation of the in-machine physical ring: When all expansion cards of a server are configured for internal communication between two computing cards, a physical ring connecting all computing cards can be constructed within the machine. For example, Figure 4 shows the construction method of this in-machine physical ring. In this physical ring, the intelligent computing chips only need to communicate to perform communication operations such as Ring-based and All-Reduce. Due to different connection methods of different optical / electrical modules, there can be multiple rings with different sequences within the machine.
[0097] In some examples, such as Figure 5As shown, in the traditional direct connection method, GPU0 and GPU2 are directly connected to form a point-to-point connection. In this way, if we want to form a ring network between GPU0 and GPU1 and between GPU2 and GPU3, additional physical connections are required, such as the Backup Fiber shown in the figure, to achieve ring redundancy. However, after introducing modules with switching capabilities (such as dOCS0 and dOCS1), the situation has changed. These switching modules enable data to be flexibly transmitted between different GPUs without the need for direct physical connections. Specifically:
[0098] GPU0 is connected to GPU2 through dOCS0, and at the same time, dOCS0 is also connected to dOCS1.
[0099] GPU1 is connected to GPU3 through dOCS1, and at the same time, dOCS1 is also connected to dOCS0.
[0100] GPU2 is connected to GPU0 through dOCS2, and at the same time, dOCS2 is also connected to dOCS3.
[0101] GPU3 is connected to GPU1 through dOCS3, and at the same time, dOCS3 is also connected to dOCS2.
[0102] In this way, data can be transmitted from dOCS0 to dOCS1, allowing a ring connection to be formed between GPU0 and GPU1. Similarly, GPU2 and GPU3 can also form a ring connection through dOCS2 and dOCS3. This design no longer requires additional physical loop connections, such as Backup Fiber, to achieve ring redundancy. Therefore, this design based on switching modules provides higher flexibility and reliability because it allows dynamic connections to be established between GPUs without relying on fixed physical connections. In addition, this distributed switching ability also helps to reduce the impact caused by the failure of a single switching chip, thereby improving the fault tolerance of the entire system.
[0103] For another example, as Figure 6 shown, in the traditional connection method, the connections between GPUs are linear. For example, GPU0 is connected to GPU2, GPU2 is connected to GPU4, and GPU4 is connected to GPU6. If we want to form a ring network among these GPUs (such as GPU0, GPU2, GPU4, GPU6), additional physical connections are required, such as pulling a wire from GPU6 back to GPU0 to complete the ring connection.
[0104] In this embodiment, after introducing modules with switching capabilities (such as dOCS), data transmission becomes more flexible, and no additional physical connections are required to form a ring network. Specifically:
[0105] GPU0 is connected to GPU2 via dOCS0.
[0106] The left - hand side switching chip (dOCS2) of GPU2 is connected to the left - hand side switching chip (dOCS1) of GPU3, forming a path.
[0107] The right - hand side switching chip (dOCS4) of GPU4 is connected to the right - hand side switching chip (dOCS3) of GPU5, forming another path.
[0108] GPU6 is connected to GPU4 via dOCS6.
[0109] In this way, two ring paths can be formed: one is the path through dOCS0, dOCS2, dOCS4, dOCS6, and the other is the path through dOCS1, dOCS3. Thus, GPU0, GPU2, GPU4, GPU6 and GPU1, GPU3, GPU5, GPU7 can respectively form two rings without the need for additional physical connection loops, such as the line from GPU6 back to GPU0.
[0110] (4) This application supports a variety of switching devices, including switching devices from existing switching chips to new optical modules with embedded optical switching units, etc., indicating that the improved communication board has strong compatibility and adaptability. In the field of data communication, there are a wide variety of switching devices, and different switching devices vary in performance, cost, applicable scenarios, etc. Supporting a variety of switching devices means that the communication board can flexibly select the appropriate switching device according to the actual needs and the specific situation of the data center. For example, in some cost - sensitive scenarios, existing switching chips can be selected; while in scenarios with higher requirements for high - speed and large - capacity data transmission, new optical modules with embedded optical switching units can be adopted, so as to give full play to the advantages of different switching devices and meet diverse data communication needs.
[0111] In summary, the design improvement of the high - speed communication board of the computing card in this application brings significant advantages in terms of device compatibility, network architecture, communication reliability, etc., providing strong support for building an efficient and stable data center intelligent computing server cluster.
[0112] Furthermore, based on the design of high-performance communication boards, this application proposes three different network topologies, specifically: constructed by using the high-speed communication boards and servers provided in any one of the first to third embodiments; wherein, the topology includes at least one of a one-dimensional linear topology, a two-dimensional mesh topology, parallel lines, and a fully connected topology. Through these topologies, the connection can be restricted between adjacent servers, making the network structure relatively simple and easy to manage. When deployed on a large scale, this local connection method can reduce the wiring complexity and cost. Thus, as the business demand grows, the data center can easily add more servers to achieve gradual network expansion without difficulties in deployment and expansion due to the complexity of the topology.
[0113] Meanwhile, in a large data center environment with a large number of computing cards (ten-thousand-card cluster), it is inevitable that servers will fail. These three topologies mainly concentrate the connections between adjacent servers, so that the impact range of a single server failure is limited to a very small local area. When a certain server fails, since its connections with other servers are mainly locally adjacent, the normal operation of other relatively distant servers is minimally affected. Thus, the high reliability and stability of the entire ten-thousand-card cluster are ensured.
[0114] In summary, these three topologies are innovative achievements based on the design of high-performance communication boards, with good scalability, which can effectively improve the flexibility and applicability of the system, and adapt to the data center requirements of different scales, from the interconnection of internal servers in a small-scale computer room to a large-scale cross-computer room cluster system. They are very convenient for large-scale deployment and expansion, and also have high reliability in a ten-thousand-card cluster, providing strong technical support for building a large, efficient, and stable data center network. Next, these three topologies will be exemplarily described through the fifth to seventh embodiments.
[0115] The fifth embodiment
[0116] The fifth embodiment of this application relates to a topology. Specifically, it is a one-dimensional linear topology, as Figure 7 shown.
[0117] Specifically, in the one-dimensional linear topology, based on the communication board, at least two expansion cards each containing a switching module are configured for each server; using each one connection on the expansion card, the servers are strung into the one-dimensional linear topology. That is to say, the servers are sequentially connected through the connections on the expansion cards, forming a linear connection structure. Moreover, this linear connection can be infinitely extended, capable of connecting all the servers in the entire data center, and the connection method is a one-way sequential connection without the need to form a loop.
[0118] Specifically, according to the physical loop formation design principle mentioned in the fourth embodiment, any continuous X servers with an arbitrary length X can be intercepted from the formed one-dimensional linear topology structure. Here, the "length X" refers to the number of servers, that is, X consecutive servers can be selected.
[0119] Exemplarily, by reconfiguring the switching module, the N×X computing cards corresponding to the intercepted X servers can be formed into a physical connection loop. Wherein, the N represents the number of computing cards, and the X represents the number of servers. That is to say, the originally linearly arranged servers and their computing cards can have their connection structure changed into a ring shape by adjusting the configuration of the switching module to form a physical connection loop to meet specific communication requirements.
[0120] Exemplarily, a ring collective communication algorithm can be executed on the physical connection loop. Among them, the ring collective communication algorithm is a communication method applicable to the ring connection structure, which belongs to the prior art and will not be elaborated in this embodiment. Through this design, the communication rate can reach the throughput of the computing card, which means that the data transmission speed in this ring structure can match the data processing speed of the computing card, thereby improving the collaborative efficiency of data communication and computing, reducing the data transmission waiting time, and thus enhancing the performance of the entire system.
[0121] Optionally, in some embodiments, in the one-dimensional linear topology structure, connections between non-adjacent servers are established according to the remaining expansion cards in the communication board card and the remaining backup lines on each expansion card for the one-dimensional linear topology structure. That is to say, in addition to the wires used to string the servers into a one-dimensional linear topology, there are remaining expansion cards and other backup lines on each expansion card for one-dimensional wiring. These spare expansion cards and lines can be connected to non-adjacent servers to handle special situations. For example, Figure 7 taking server No. 0 as an example, it can use a set of backup wires of the left and right expansion cards and 2 sets of wires of the upper and lower expansion cards to connect to other different servers that are not adjacent to it. In this way, connections between non-adjacent servers can be established through these backup lines.
[0122] Furthermore, when a server fails, the servers on both sides of the failed server can use these additional wires. In this way, through these spare wires, the servers on both sides can bypass and isolate the failed server.
[0123] Furthermore, after isolating the faulty server, the servers on both sides and other relevant servers can use these backup connections to reorganize into a new ring connection. By doing so, even when some servers fail, the entire system can still maintain normal communication, preventing resources such as GPUs connected to the faulty servers from being wasted, and thus maintaining the overall operation and resource utilization efficiency of the system.
[0124] It should be noted that in a server system, the number of lines for handling fault scenarios is crucial. The more lines there are, the more likely the system is to remain non-degraded when facing more server failures. This is because more backup lines provide more communication path options. When some servers or lines fail, the system can reconstruct communication connections through these backup lines to maintain the overall function.
[0125] Optionally, in some embodiments, the backup lines are specifically connected to the left and right servers of the server in a half-by-half sequential manner. As shown above Figure 7 These backup lines can be connected to the left and right servers of the server in a half-by-half sequential manner. This connection method can optimize the usage efficiency of the backup lines to a certain extent, enabling the system to switch communication paths more flexibly when a failure occurs.
[0126] For example, when the backup lines can be connected to the left and right servers of the server in a half-by-half sequential manner, in a given data center, when the typical value of the overall failure of the computing cards is 0.9%, it can be calculated that:
[0127] For a 4-card server, when M×L > 6, under the given typical failure value, the system can effectively handle the failure. Where M represents the number of expansion cards and L represents the number of lines. For a task that requires 32 computing cards, compared to the ideal situation (i.e., no failure situation), less than 0.5% of the computing cards will be wasted. This means that in a 4-card server configuration, as long as the relevant parameters of the backup lines meet certain conditions, even in the case of typical failures, the impact on the waste of computing card resources during task execution is small.
[0128] For an 8-card server, when the condition M×L > 8 is met, similarly for a task that requires 32 computing cards, compared to the ideal situation, less than 0.5% of the computing cards will be wasted. This shows that in an 8-card server configuration, there are also corresponding requirements for the parameters of the backup lines to ensure the efficient utilization of computing card resources in a failure scenario and reduce waste.
[0129] It can be understood that in some embodiments, in the scenario where there are R N-card servers in a cabinet, the number of lines connected by one cabinet to one side is:
[0130]
[0131] The above is an expression for calculating the number of cables. Among them, M is the number of the expansion cards, L represents the number of the lines, and i is the cumulative parameter.
[0132] It can be seen that since the number of cables connected to one side of the cabinet calculated is small, and a small number of cables means a reduced connection complexity and possibly a reduced cost. Therefore, in actual server deployment, whether it is cross-cabinet group deployment (combining and connecting servers in different cabinets) or cross-data center deployment (connecting servers in different data centers), it is easier to operate.
[0133] It is not difficult to find that in this embodiment, a one-dimensional linear topology structure is provided. Through this structure, more computing cards are allowed to calculate together with higher efficiency, thereby enhancing the distributed computing ability of the AI accelerator. And, under the condition of the same failure rate, the fragmentation rate is reduced, which is beneficial to further reducing the cost.
[0134] Sixth Embodiment
[0135] The sixth embodiment of the present application relates to a topology structure. Specifically, the topology structure is: a two-dimensional mesh topology structure, as Figure 8 shown.
[0136] Specifically, in the two-dimensional mesh topology structure, based on the communication board card, more than two expansion board cards including switching modules are configured for each server, so that the servers form an intertwined connection manner on the plane.
[0137] Specifically, since each server is equipped with more than two expansion board cards including switching modules, there are more than 4 groups of connection lines. These connection lines are the physical basis for data transmission and connection between servers. Among them, 4 groups of connection lines are used to construct a two-dimensional mesh structure (2D Tours), that is, Figure 8 the black line part in. The two-dimensional mesh topology structure enables the servers to form an intertwined connection manner on the plane, which helps to realize multi-path transmission of data between different servers, improve the efficiency and reliability of data transmission, and enhance the overall performance of the system.
[0138] Optionally, in some embodiments, in the two-dimensional mesh topology, connections between non-adjacent servers are established based on the remaining expansion cards in the communication board and the remaining backup lines on each expansion card for the two-dimensional mesh topology. That is, the remaining connections serve as backup lines, and their main function is to bypass these faulty nodes and maintain the normal operation of the system when a server (node) fails. This improves the fault tolerance of the system and avoids a significant decline in the performance of the entire system or the inability to use some functions properly due to the failure of individual servers.
[0139] Exemplarily, the backup lines can be connected to any non-adjacent servers. Taking the above typical design as an example, this design has 8 groups of connections, and four groups of backup lines are connected to the positions of the diagonal lines. This connection method enables the backup lines to form a more flexible backup connection architecture in the system. When a server fails, the data transmission path can be quickly reconstructed through these backup lines to ensure the continuity of system communication.
[0140] Specifically, according to the physical loop formation design principle mentioned in the fourth embodiment, when a task requires X servers for distributed computing, in the initialization stage, only an uninterrupted line with a length of X needs to be found in the entire grid (i.e., the two-dimensional mesh topology formed by the connections between servers). Here, the line refers to the connection lines between servers, and backup lines can be used. The internal computing cards of the X servers determined in this way can form a physical loop. This means that before the distributed computing starts, by reasonably using the connection lines to select appropriate servers to construct a physical loop structure that meets the computing requirements, the distributed computing can be ensured to proceed.
[0141] It can be understood that in this topology, only when all the servers connected to a server fail, will this server be disconnected from the entire topology and thus unable to participate in the distributed computing. This indicates that this topology has high reliability because for a server to be completely disconnected from the topology, all the servers connected to it around need to fail simultaneously, and the probability of this situation occurring is extremely low.
[0142] In some examples, according to the simulation results, in a ten-thousand-card cluster (i.e., a server cluster containing a large number of computing cards), the possibility of wasting resources due to isolated nodes (i.e., servers disconnected from the entire topology) can be almost ignored. Therefore, in a large-scale cluster environment, this topology can effectively avoid resource waste caused by node failures, ensure the efficient utilization of cluster resources, and provide strong support for the stable operation of distributed computing.
[0143] It is not difficult to find that in this embodiment, a two-dimensional network topology structure is provided. Through this structure, more computing cards are allowed to calculate together with higher efficiency, thereby enhancing the distributed computing ability of the AI accelerator. Moreover, under the condition of the same failure rate, the fragmentation rate is reduced, which is beneficial to further reducing costs.
[0144] The seventh embodiment
[0145] The seventh embodiment of this application relates to a topology structure. Specifically, the topology structure is: parallel lines and a full-connection topology structure, as Figure 9 and Figure 10 shown.
[0146] Specifically, in the parallel lines and full-connection topology structure, the servers are divided into multiple parallel lines, and the servers with the same index on each line are connected using a full-connection topology.
[0147] Exemplarily, all servers can be first divided into multiple parallel lines. Then, a full-connection topology (i.e., Full-Mesh topology, where there is a direct connection line between any two nodes in the network) is used to connect the servers with the same index on each line. This design method combines the characteristics of parallel lines and full-connection, aiming to construct an efficient and flexible server connection architecture.
[0148] Optionally, in some embodiments, in the parallel lines and full-connection topology structure, at least two expansion cards containing switching modules are equipped for each server; among them, two groups of lines are used to connect adjacent servers on the same line segment, and the remaining connection lines are used to construct a full-connection plane in the parallel lines and full-connection topology structure.
[0149] Among them, two groups of lines are used to connect adjacent servers on the same line segment, which ensures the sequential connection of servers on their respective parallel lines, forming the basic structure of the parallel lines. And the remaining several groups of connection lines are used to construct the plane of Full-Mesh, that is, to realize the full-connection between the servers with the same index on different parallel lines. Through this distribution of the connection line functions, there is both a sequential connection relationship and a cross-line full-connection relationship between the servers, thus forming a complex and efficient topology structure.
[0150] See Figure 9 and Figure 10 , which respectively show the connection topology situations when there are 6 groups of lines and 8 groups of lines. Through these diagrams, it is possible to more intuitively see the specific connection methods and topology forms between the servers under different numbers of connection lines, which helps to understand the structural characteristics and data transmission paths of this new topology under different configurations. This new topology design has advantages in improving data transmission efficiency, enhancing the reliability and scalability of the system, etc.
[0151] Specifically, according to the physical loop formation design principle mentioned in the fourth embodiment, in the topology structure example shown as Figure 11 below, the computing device represents the node, and the situation where normal nodes cannot be interconnected only occurs in two specific cases. The first case is that all nodes in a plane fail; the second case is that the corresponding nodes on the upper and lower planes corresponding to the non-failed nodes all fail. Since the occurrence conditions of these two cases are relatively harsh, in general node failure scenarios, the topology structure can maintain the interconnection between normal nodes.
[0152] In practical applications, for the scenario of 6 groups of lines and 8 computing cards per server, through the quantitative research and analysis of the reliability of the topology structure under specific configurations, it can be known that the probability of non-interconnection will be as low as 10 -4 . Therefore, in this specific topology structure and server configuration, due to the extremely low probability that normal nodes cannot be used due to node failure, it shows that this topology structure has high reliability in this scenario, can effectively ensure that the system can still operate normally when nodes fail, and reduce the service interruption and resource waste problems caused by node failure.
[0153] It can be understood that in this topology structure, most of the connections are located in the mesh plane (fully connected plane, that is, a planar structure where any two nodes have direct connection lines). Since these connections are concentrated in the mesh plane, they can be deployed within a cabinet or between adjacent cabinets. Because the spatial distance inside the cabinet or between adjacent cabinets is relatively close, it is convenient for laying and connecting lines, which is conducive to reducing the wiring cost and difficulty and improving the efficiency of system construction.
[0154] Among them, only a few lines need to extend along the direction of the parallel lines. This feature is very beneficial to cluster expansion. Because there are fewer lines extending along the direction of the parallel lines, when performing large-scale cluster deployment, complex long-distance wiring is not required, making the expansion connection of the entire cluster easier. Whether it is connecting the servers in the entire computer room or realizing the deployment of servers across computer rooms, it can be completed more easily.
[0155] It is not difficult to find that this embodiment provides a parallel line and a fully connected topology structure. Through this structure, more computing cards are allowed to calculate together with higher efficiency, thereby enhancing the distributed computing ability of the AI accelerator, and, in the case of the same failure rate, reducing the fragmentation rate, which is conducive to further reducing costs.
[0156] Eighth Embodiment
[0157] The eighth embodiment of this application relates to a communication system, and the system includes the topology structure described in any one of the fifth to seventh embodiments, as well as a controller and a task scheduler;
[0158] The controller is used to maintain the topology information of the topology structure and update the server status when a server is identified as faulty.
[0159] The task scheduler is used to calculate a target topology based on the updated server status and adjust the expansion card status according to the target topology to change the connection mode of the server in the topology structure.
[0160] Specifically, in the GPU cluster construction stage, the servers are connected according to a pre-designed topology structure. Connecting accurately according to this topology lays the foundation for the efficient operation of the subsequent cluster and ensures that there are specific data transmission paths and connection relationships between the servers.
[0161] Specifically, the controller can maintain the topology information of the topology structure and update the server status when a server is identified as faulty.
[0162] Exemplarily, the controller is responsible for maintaining the complete topology information of the entire cluster. That is to say, the controller always knows the position of each server in the topology structure, the connection relationship with other servers and other key information.
[0163] Exemplarily, when a user's task is issued to the cluster, the task scheduler can obtain the detailed information of the current task, including but not limited to the size of the communication group. Task information such as the communication group size is crucial for determining how to efficiently execute the task on the existing server resources. Then, based on the existing server status (normal or faulty) and the connection situation between the servers, the optimal topology for this task, that is, the target topology, is calculated. The target topology is the server connection mode that can make the task run with the highest efficiency under the actual conditions of the current cluster. For example, paths passing through faulty servers can be avoided as much as possible, or according to the data interaction characteristics of the task, the most closely connected server combination can be selected.
[0164] Furthermore, when the controller detects a faulty server, the task scheduler can adjust the expansion card status to change the connection mode of the server in the topology structure to achieve rapid reconfiguration. After updating the server status, the task scheduler can recalculate the target topology based on the new server status. This is because the appearance of a faulty server changes the overall situation of the cluster, and the original target topology may no longer be applicable. The recalculated target topology can adapt to the new cluster environment and ensure that the task continues to run efficiently. The task scheduler sends the new optimal topology to each server. In this way, by adjusting the status of the expansion card, the connection mode of the server in the topology structure can be changed to achieve rapid reconfiguration.
[0165] Exemplarily, after changing the connection mode of the server in the topology, the task can be restarted. In this way, even if a server failure occurs during the task execution, the cluster can continue to complete the task by dynamically adjusting the topology, minimizing the impact of the failure on the task execution and ensuring the stability of the cluster and the continuity of task processing.
[0166] As described in Figure 1 above, in the normal operating state, each expansion card (such as expansion card 1) is connected to the expansion card of another server through an optical fiber (such as Fiber1). This is the standard connection mode when the server cluster is working properly, ensuring the normal transmission of data between servers. When the cluster monitoring system detects a failure of the peer server, the communication board based on this embodiment can immediately calculate a new topology (the specific calculation method can adopt the mature topology calculation method in the related technology, which will not be elaborated in this application). After the new topology is calculated, the cluster management software will receive a notification to handle the failure. Specifically, an instruction can be sent to the firmware that controls the expansion card. After the firmware of the expansion card receives the instruction, another optical fiber Fiber (as a backup line) can be quickly enabled and connected to a new normally operating server. This switching process is very fast and can be completed within milliseconds. Since the backup line is pre-planned, the system can immediately switch to the backup connection when a failure is detected, rather than waiting for reconfiguration or physical connection, thus minimizing the downtime of the cluster caused by server failure. Those skilled in the art can understand that the traditional fault recovery method generally requires removing the faulty server from the rack and reconnecting it, which usually takes a longer time.
[0167] It is not difficult to find that the communication system provided in this embodiment includes the topology described in any one of the fifth to seventh embodiments, as well as a controller and a task scheduler, and can immediately switch to the backup connection when a failure is detected, rather than waiting for reconfiguration or physical connection, thus minimizing the downtime of the cluster caused by server failure.
[0168] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0169] Based on the above embodiments, it is not difficult to find that the protected content of this application includes the communication board design and the topology structure based on the communication board design. The two complement each other and can jointly support the interconnection of AI computing chips in environments such as data centers. Compared with related technologies, the solutions provided by the embodiments of this application have at least the following beneficial effects:
[0170] (1) It supports the interconnection and interoperability between any AI computing chips, is not restricted by the chip type, and has strong versatility. That is to say, the solution provided by the embodiments of this application enables these different types of chips to be interconnected and work together, breaking down the barriers between chips, so that the data center can flexibly select and combine different chips according to actual needs, and fully utilize the advantages of various chips.
[0171] (2) The topology structure design provided by the embodiments of this application can support high-bandwidth domains of any size, can connect all computing cards in the data center, and can reasonably allocate the corresponding number of computing cards for high-bandwidth parallel tasks of any scale, thus significantly accelerating the process of large-scale model training.
[0172] (3) During the operation of the data center, it is inevitable that computing cards may fail. In traditional architectures, the failure of individual computing cards can lead to unreasonable resource allocation and resource fragmentation, resulting in some computing resources not being effectively utilized. However, the flexible fault scheduling mechanism of this solution can centrally manage all computing cards. When a failure occurs, it can timely adjust the resource allocation, avoid resource waste, and improve the resource utilization rate of the entire system. That is to say, through the flexible fault scheduling mechanism, the centralized management of all computing cards in the data center is realized, effectively alleviating the problem of resource fragmentation, thereby enhancing the utilization rate of system resources.
[0173] (4) This application can ensure high availability when the system fails, specifically manifested in being able to isolate the faulty single server and its related network, thus narrowing the scope of the impact of the failure; and by utilizing the hardware characteristics, it can avoid the bandwidth degradation of the remaining servers, significantly improving the reliability of the system.
[0174] (5) This application can be compatible with the existing UBB2.0 specification, enabling this application to be seamlessly connected with existing technologies and networks, without the need for large-scale transformation of the data center, reducing the application cost and implementation difficulty. This makes this solution widely applicable to different types of data centers, meeting the needs of various application scenarios, and having strong practicality and promotion value.
[0175] (6) Traditional supernode architectures require high hardware costs and complex design solutions, which limit their large-scale deployment. However, this application can implement a supernode architecture at a relatively low cost, making large-scale deployment more economically feasible.
[0176] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0177] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims may also be implemented by one unit or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.
[0178] As described above, these are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A communication board, characterized in that: The communication board comprises: Multiple computing cards, each computing card is used to perform computing tasks; A plurality of switching modules, wherein the switching modules include a high-performance electrical switching chip and a first optical module; each switching module is connected to at least two computing cards to realize data exchange between the computing cards; the first optical module is an optical module without switching capability; the throughput of the switching module is 2 to L times the throughput of the computing card, where L is a positive integer; Each switching module is configured with an expansion card, and each expansion card is provided with L lines; among the L lines, at least one line is used to form a backup line, and one line is used to form a topology structure.
2. The communication board according to claim 1, characterized in that: The switching module is specifically a second optical module, and the second optical module is an optical module with switching capability.
3. The communication board according to claim 1 or 2, characterized in that: Some of the expansion cards, together with the switching modules connected to the some of the expansion cards, are replaced with PCB boards.
4. A topological structure, characterized in that: The topological structure is specifically constructed using the communication board and server as described in any one of claims 1 to 3; wherein the topological structure includes at least one of a one-dimensional linear topological structure, a two-dimensional mesh topological structure, a parallel line and a fully connected topological structure.
5. The topological structure according to claim 4, characterized in that: In the one-dimensional linear topology, at least two expansion cards including switching modules are configured for each server based on the communication board; The servers are connected in series to form the one-dimensional linear topology using a copy of the connection on each expansion card.
6. The topological structure according to claim 5, characterized in that: In the one-dimensional linear topology structure, connections between non-adjacent servers are established based on the remaining expansion cards in the communication board and the remaining backup lines on each set of expansion cards used for the one-dimensional linear topology structure.
7. The topological structure according to claim 6, characterized in that: The backup line specifically adopts a method of connecting the left and right servers of the server in half.
8. The topological structure according to claim 4, characterized in that: In the two-dimensional mesh topology, based on the communication board, more than two expansion boards including switching modules are configured for each server, so that the servers form an interwoven connection on a plane.
9. The topological structure according to claim 8, characterized in that: In the two-dimensional mesh topology structure, connections between non-adjacent servers are established based on the remaining expansion cards in the communication board and the remaining backup lines on each set of expansion cards used for the two-dimensional mesh topology structure.
10. The topological structure according to claim 4, characterized in that: In the parallel line and fully connected topology structure, the server is divided into multiple parallel lines, and the fully connected topology is used to connect the servers with the same index on each line.
11. The topological structure according to claim 10, characterized in that: In the parallel line and fully connected topology, each server is equipped with at least two expansion cards including switching modules; wherein two groups of lines are used to connect adjacent servers on the same line segment, and the remaining lines are used to construct a fully connected plane.
12. A communication system, the system comprising the topology structure according to any one of claims 4 to 11, as well as a controller and a task scheduler; The controller is used to maintain the topology information of the topology structure, and update the server status when any of the servers is identified as a failure; The task scheduler is used to calculate the target topology according to the updated server status, and adjust the expansion card status according to the target topology to change the connection mode of the server in the topology structure.
Citation Information
Cited By
Management board, universal substrate and monitoring method
CN120407490A
Multi-card interconnection system
CN120929410A