Server system and data transmission method of server system

By optimizing the communication link design between the acceleration unit and switching component of the server system, the problems of complex expansion schemes and limited interconnection of computing components were solved, achieving efficient data transmission and improved computing capabilities.

CN121462413BActive Publication Date: 2026-03-31INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies have complex expansion schemes for computing components, limited interconnection capabilities, and difficulty in achieving high-bandwidth, low-latency supernode systems.

Method used

The design employs K acceleration units and L first switching components. Each acceleration unit includes M acceleration components and L second switching components. By optimizing the communication link design, the scale of high-speed signal lines is reduced, thereby lowering the system complexity.

Benefits of technology

It increases the scale of extended interconnects and reduces system complexity, enabling high-bandwidth, low-latency data transmission, improving computing power and accelerating the efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462413B_ABST
    Figure CN121462413B_ABST
Patent Text Reader

Abstract

The application discloses a server system and a data transmission method of the server system, and relates to the field of computer hardware. The server system comprises K acceleration units and L first switching components, each of the acceleration units comprises M acceleration components and L second switching components; in each of the acceleration units, each of the acceleration components is connected with the L second switching components through L first communication links respectively; each of the second switching components is connected with the M acceleration components through M first communication links respectively; each of the second switching components is connected with the corresponding first switching component through M second communication links, the transmission rate of the first communication link is different from the transmission rate of the second communication link; and each of the first switching components is connected with the corresponding second switching component through M second communication links respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer hardware technology, and in particular to a server system and a data transmission method for the server system. Background Technology

[0002] To meet the needs of computing power expansion and improvement, it is necessary to build a supernode system with tens or even hundreds of high bandwidth and low latency through extended interconnect technology. In the current extended interconnect scheme of supernodes, the interconnection of components is usually achieved through direct connection of high-speed lines or small onboard switching components. The switching components need to have a considerable number of ports. This architecture is highly complex and it is difficult to achieve full interconnection between each component.

[0003] This shows that the related technologies suffer from problems such as high complexity of computing component expansion schemes and limited interconnection. Summary of the Invention

[0004] This application provides a server system and a data transmission method for the server system, in order to at least solve the problems of high complexity of expansion schemes for computing components and limited interconnection in related technologies.

[0005] This application provides a server system, including: K acceleration units and L first switching components, each of the K acceleration units including M acceleration components and L second switching components; wherein, in each acceleration unit, each of the M acceleration components is connected to L second switching components through L first communication links; each of the L second switching components is connected to the M acceleration units through M first communication links; the L second switching components correspond one-to-one with the L first switching components, and each of the L second switching components is connected to a corresponding first switching component through M second communication links, wherein the transmission rate of one of the L first communication links is different from the transmission rate of one of the M second communication links; each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit through M second communication links; wherein K, L, M, N, and P are all positive integers greater than or equal to 2.

[0006] This application also provides a data transmission method for a server system. The server system includes: K acceleration units and L first switching components. Each of the K acceleration units includes M acceleration components and L second switching components. In each acceleration unit, each of the M acceleration components is connected to the L second switching components via L first communication links. Each of the L second switching components is connected to the M acceleration components via M first communication links. The L second switching components correspond one-to-one with the L first switching components, and each of the L second switching components is connected to a corresponding first switching component via M second communication links. The transmission rate of one of the L first communication links is equal to the transmission rate of one of the M second communication links. The rates are different; each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit through M second communication links; K, L, M, N and P are all positive integers greater than or equal to 2; the above method includes: obtaining data to be transmitted through the first acceleration component in the first acceleration unit of the K acceleration units, wherein the data to be transmitted is data to be transmitted to the second acceleration component in the second acceleration unit of the K acceleration units; in response to the data to be transmitted, transmitting the data to be transmitted to the second acceleration component via the data transmission path from the first acceleration component to the second acceleration component, wherein the data transmission path includes a second switching component in the first acceleration unit, a first switching component in the L first switching components and a second switching component in the second acceleration unit.

[0007] This application discloses a server system comprising: K acceleration units and L first switching components. Each of the K acceleration units includes M acceleration components and L second switching components. Within each acceleration unit, each of the M acceleration components is connected to the L second switching components via L first communication links. Each of the L second switching components is connected to the M acceleration units via M first communication links. The L second switching components correspond one-to-one with the L first switching components, and each of the L second switching components is connected to its corresponding first switching component via M second communication links. The transmission rate of one of the L first communication links differs from the transmission rate of one of the M second communication links. Each of the L first switching components is connected to its corresponding second switching component in each acceleration unit via M second communication links. Where K, L, M, N, and P are all positive integers greater than or equal to 2. This system optimizes the communication link design between components, reduces the scale of high-speed signal lines, and lowers system complexity. Therefore, it can solve the problems of high complexity of computing component expansion schemes and limited interconnection in related technologies, and achieve the technical effect of increasing the scale of expansion interconnection and reducing system complexity. Attached Figure Description

[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic diagram of the structure of an optional server system according to an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of the structure of an optional accelerator card server system according to an embodiment of this application.

[0011] Figure 3 This is a schematic diagram of another server system according to an embodiment of this application.

[0012] Figure 4 This is a schematic diagram of the structure of another server system according to an embodiment of this application.

[0013] Figure 5 This is a schematic diagram of the structure of another switching unit according to an embodiment of this application.

[0014] Figure 6 This is a schematic diagram of the structure of another switching unit according to an embodiment of this application.

[0015] Figure 7 This is a schematic diagram of the structure of a computing system according to an embodiment of this application.

[0016] Figure 8 This is a schematic diagram of the structure of another computing system according to an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0018] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0019] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] According to one aspect of an embodiment of this application, a switching unit is provided. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0021] Figure 1 This is a schematic diagram of an optional server system according to an embodiment of this application, such as... Figure 1As shown, the system includes: K acceleration units 101 and L first switching components 102. Each of the K acceleration units 101 includes M acceleration components 103 and L second switching components 104. In each acceleration unit 101, each of the M acceleration components 103 is connected to the L second switching components 104 via L first communication links 105. Each of the L second switching components 104 is connected to the M acceleration components 103 via the M first communication links 105. The L second switching components 104 correspond one-to-one with the L first switching components 102. Each of the L second switching components 104 is connected to the corresponding first switching component 102 through M second communication links 106. The transmission rate of one of the L first communication links 105 is different from the transmission rate of one of the M second communication links 106. Each of the L first switching components 102 is connected to the second switching component 104 corresponding to each first switching component 102 in each acceleration unit 101 through M second communication links 106. Wherein, K, L, M, N and P are all positive integers greater than or equal to 2.

[0022] With the rapid development of large models in deep learning, particularly in tasks such as natural language processing, visual understanding, and speech recognition, this trend has placed new demands on computing infrastructure. While traditional Graphics Processing Units (GPUs) have advantages in parallel computing, the memory capacity of a single GPU is insufficient to handle such massive model parameters. The growth rate of large model parameters far outpaces the increase in GPU memory capacity. This means that to run these large models on GPUs, more GPUs are needed to share storage requirements. Furthermore, as model complexity increases, model structures are gradually shifting towards sparse designs, which helps reduce computational resource requirements but also increases the need for cross-GPU data transfer.

[0023] Therefore, constructing extended high-speed interconnect domains with high bandwidth and low latency has become a key measure to improve the linearity of computing power expansion. Through extended interconnects, multiple GPUs can be connected into a high-performance supernode system, enabling north-south interconnects between different layers within the data center. This ensures efficient and real-time data exchange, thereby significantly improving overall computing power and accelerating the model training process. Here, north-south interconnects are like... Figure 2As shown, southbound interconnect mainly refers to the direct communication connection between GPU nodes, also known as card-to-card interconnect, while northbound interconnect mainly refers to the communication network between server nodes, involving data exchange and task scheduling between servers. Through network core switches and access switches, the GPU resource pool can be horizontally scaled out to achieve multi-machine interconnection, while through network interfaces and switches, the number of GPUs in a single node can be increased (for example, a server has multiple GPU cards), and the computing power of a single node can be vertically scaled up.

[0024] Based on the aforementioned extended interconnection scheme, related technologies include server systems, in order to... Figure 3 Taking the Optimized Accelerator Module (OAM) server as an example, the Central Processing Unit (CPU) is used to perform system management and task scheduling. The CPU can be connected to the Network Interface Card (NIC), Non-Volatile Memory Express (NVMe), and multiple OAM accelerator cards through a high-speed serial bus standard (Peripheral Component Interconnect Express, PCIe) switch. The GPU is located in the OAM accelerator card and is mainly responsible for large-scale parallel computing tasks. There are inter-card interconnect links between the eight OAMs, and the GPUs can be directly connected through high-speed cables or small onboard switching components. With the above architecture, a CPU can connect to more peripherals through multiple switches to meet the parallel needs of large-scale peripherals. However, as the complexity of the model increases, the bandwidth of the inter-card interconnect links may not be able to meet the rapidly increasing communication needs. Direct connections between GPUs or indirect connections through small switching components face insufficient bandwidth during high-speed data exchange, thus becoming a performance bottleneck and limiting the efficiency of GPU scaling up. Since the number of straight-line routes between each GPU is limited, when the number of GPUs increases to a certain extent, the complexity of the network topology becomes too high, which also leads to the strain on link resources and affects efficient communication between GPUs. This architecture limits the bandwidth and network scale of southbound GPU scaling up.

[0025] This shows that the related technologies suffer from problems such as high complexity of computing component expansion schemes and limited interconnection.

[0026] To at least partially solve the above-mentioned technical problems, this embodiment provides a server system comprising: K acceleration units and L first switching components. Each of the K acceleration units includes M acceleration components and L second switching components. Within each acceleration unit, each of the M acceleration components is connected to the L second switching components via L first communication links. Each of the L second switching components is connected to the M acceleration components via M first communication links. The L second switching components correspond one-to-one with the L first switching components, and each of the L second switching components is connected to its corresponding first switching component via M second communication links. The transmission rate of one of the L first communication links differs from the transmission rate of one of the M second communication links. Each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit via M second communication links. Here, K, L, M, N, and P are all positive integers greater than or equal to 2. This optimizes the design of communication links between components, reduces the scale of high-speed signal lines, and lowers system complexity. Therefore, it can solve the problems of high complexity of computing component expansion schemes and limited interconnection in related technologies, and achieve the technical effect of increasing the scale of expansion interconnection and reducing system complexity.

[0027] Optionally, such as Figure 4 As shown, the expansion of switching ports can be achieved based on multiple switching components. The switching components can be combined into a switching unit. The switching unit can include two layers of switching components formed by cascading first-layer switching components and second-layer switching components. The switching unit can include at least one first-layer switching component and at least two second-layer switching components. The first-layer switching component (i.e., L1 Switch) can be located in the upper layer of the cascaded structure and can be used to build non-blocking interconnects between second-layer switching components (L2 Switch). Based on 6 switching units and a GPU that supports 48 lanes PCIe Gen5 southbound interconnect, a 36-card scale-up system can be built.

[0028] Optionally, in this embodiment, the server may include K acceleration units and L first switching components. Each of the K acceleration units may include M acceleration components and L second switching components. Here, the acceleration units can be units within the server used for accelerated computing, and each acceleration unit can be used to independently execute parallel computing tasks. The acceleration units can be indirectly connected to other acceleration units through the first switching components to achieve larger-scale parallel computing.

[0029] Optionally, the acceleration unit may include acceleration components, and an acceleration unit may include M acceleration components for performing computations. These components may include, but are not limited to, GPUs, Tensor Processing Units (TPUs) or Data Processing Units (DPUs), which can be used to perform specific types of data processing and computation.

[0030] Optionally, each acceleration unit may also include L second switching components, which can be used to aggregate the communication links of the acceleration components, reduce the complexity of communication between acceleration components, and provide a connection channel between the acceleration components and the first switching components.

[0031] Optionally, the M acceleration components can be connected to the second switching component through L first communication links. Each acceleration component is connected to the second switching component through an independent link, which enables fast data exchange, reduces communication latency, and improves the efficiency of parallel processing.

[0032] Optionally, each of the L second switching components can be connected to the M acceleration components through M first communication links respectively; the L second switching components can correspond one-to-one with the L first switching components, that is, one second switching component can be connected to one first switching component.

[0033] Optionally, a first switching component can be connected to multiple second switching components to enable data transmission between the multiple second switching components, that is, data transmission between multiple acceleration units. Through the above connection strategy, a non-blocking data transmission path can be achieved from the acceleration unit to the first switching component and then to the second switching components, avoiding the delay caused by repeated data transmission between multiple switching components. Each of the L second switching components only handles communication with the first switching component directly connected to it, simplifying the network topology.

[0034] Optionally, the transmission rate of one of the L first communication links may differ from the transmission rate of one of the M second communication links. Here, due to process limitations, different components may use different communication protocols and achieve different transmission rates. For example, the transmission channel in the first communication link may only be a fifth-generation peripheral component interconnect bus channel (i.e., Gen5 channel), while the transmission channel in the second communication link may be a sixth-generation peripheral component interconnect bus channel (i.e., Gen6 channel). The transmission rate of the first transmission channel may be the Gen5 transmission rate, and the transmission rate of the second transmission channel may be the Gen6 transmission rate. In large-scale CPU interconnect scenarios, the mismatch in rates may affect the overall processing efficiency of the system.

[0035] Here, "Gen" (Generation) refers to different versions or generations of the Peripheral Component Interconnect Express (PCIe) protocol. PCIe is a high-speed serial bus standard widely used in computer hardware to connect processors, various devices on the motherboard, and other system components. Over time, it has undergone several versions, each bringing performance improvements. For example, PCIe 5.0 (the fifth generation of PCIe) offers significant improvements in data transfer rate and bandwidth compared to PCIe 4.0 (the fourth generation of PCIe). A single x16 PCIe 5.0 link can provide a higher data transfer rate than a single x16 PCIe 4.0 link.

[0036] Optionally, different transmission rates can be matched through a second switching component. By adapting the transmission channels and the number of switching components, a non-blocking network for processing corresponding data transmissions can be constructed.

[0037] Optionally, the uplink and downlink bandwidths of the second switching component can be matched, that is, the bandwidths of the first communication link and the second communication link can be matched.

[0038] For example, K=4, L=6, M=8, N=8, P=4, as... Figure 5 As shown, a 16-card server system can be constructed, including 4 acceleration units and 6 first switching components. Each of the 4 acceleration units includes 8 acceleration components and 6 second switching components. In each acceleration unit, each acceleration component is connected to 6 second switching components through 6 first communication links (i.e., 6 Gen5x8 groups). Each second switching component is connected to 8 acceleration components through 8 first communication links (i.e., 8 Gen5 x8 groups). The 6 second switching components correspond one-to-one with the 6 first switching components. Each second switching component is connected to the corresponding first switching component through 8 second communication links (i.e., 8 Gen6x4 groups). One first communication link includes 8 first transmission channels, and one second communication link includes 4 second transmission channels. The transmission rates of the first and second transmission channels are different.

[0039] Optionally, after the transmission rate matching and link aggregation are achieved through the second switching component, the signal between each acceleration unit and the first switching component can be a Gen6 PCIe signal, which greatly reduces the number of signals and the number of high-density connectors between boards used when building an actual server system, thereby reducing the system construction cost.

[0040] According to the embodiments provided in this application, a server system includes: K acceleration units and L first switching components. Each of the K acceleration units includes M acceleration components and L second switching components. In each acceleration unit, each of the M acceleration components is connected to the L second switching components via L first communication links. Each of the L second switching components is connected to the M acceleration components via M first communication links. The L second switching components correspond one-to-one with the L first switching components, and each of the L second switching components is connected to its corresponding first switching component via M second communication links. The transmission rate of one of the L first communication links is different from the transmission rate of one of the M second communication links. Each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit via M second communication links. Where K, L, M, N, and P are all positive integers greater than or equal to 2. This optimizes the communication link design between components, reduces the scale of high-speed signal lines, and lowers system complexity. Therefore, it can solve the problems of high complexity of computing component expansion schemes and limited interconnection in related technologies, and achieve the technical effect of increasing the scale of expansion interconnection and reducing system complexity.

[0041] In one exemplary embodiment, in each acceleration unit, M acceleration components are packaged into an accelerator module in groups of at least two acceleration components.

[0042] In related technologies, single-die packaging is usually used. However, in this embodiment, by packaging at least two acceleration components in one module, more computing resources can be integrated in the same or smaller physical space, significantly improving computing density. This is accomplished with less physical space and circuit wiring, solving the problem of large-scale interconnection signals in accelerator modules and the difficulty in implementation.

[0043] Optionally, the acceleration components within the accelerator module can share resources, such as cache and memory subsystem. Simultaneously, these components can work collaboratively, rapidly exchanging data via internal communication links (such as the first communication link) to achieve efficient execution of parallel computing tasks.

[0044] This embodiment demonstrates how integrating multiple GPUs or other types of accelerators into a single module enables more efficient resource utilization and communication paths, thereby improving computational density and performance.

[0045] In one exemplary embodiment, each acceleration unit is integrated onto a board, and different acceleration units among the K acceleration units are integrated onto different boards.

[0046] In this embodiment, each acceleration unit can be integrated onto a single board, and different acceleration units can be integrated onto different boards, making each acceleration unit an independent module.

[0047] Modular design allows for the dynamic addition or reduction of the number of acceleration units based on the needs of the computing task. In the field of high-performance computing, as the complexity of the model and the amount of data increase, computing requirements may change over time. Modular design enables the system to adapt to these changes. By simply adding or removing boards, computing power can be expanded or reduced, improving the system's flexibility and scalability. It can also reduce the complexity of inter-board wiring, thereby reducing signal latency and improving data transmission efficiency.

[0048] This embodiment simplifies system architecture design and improves system performance and reliability by integrating each acceleration unit onto a separate board.

[0049] In one exemplary embodiment, each first switching component is integrated onto a board, and different first switching components among the L first switching components are integrated onto different boards.

[0050] In this embodiment, each first switching component can be independently integrated onto its own board, realizing a modular design of the system. This not only optimizes the signal transmission path and reduces signal interference and attenuation, but also maximizes the use of limited space resources and improves the density and efficiency of the server.

[0051] Alternatively, the topology can be constructed with each acceleration unit as a minimum unit, and this minimum unit can be a board. Each second switching component can also be a board, making the connection simpler and greatly improving the feasibility of the topology.

[0052] By integrating different first switching components onto different boards, this embodiment can reduce system complexity, reduce signal interference and attenuation, and improve server density and efficiency.

[0053] In one exemplary embodiment, the server system further includes Q processing units, and each acceleration unit further includes R third switching units. Each of the Q processing units corresponds to a portion of the K acceleration units, and different processing units among the Q processing units correspond to different acceleration units among the K acceleration units. The R third switching units of each acceleration unit are all connected to the corresponding processing unit of each acceleration unit, where Q and R are both positive integers greater than or equal to 1.

[0054] In each acceleration unit, R third switching components are connected to the corresponding processing component of each acceleration unit. Each of the R third switching components is connected to some of the M acceleration components. Different of the R third switching components are connected to different of the M acceleration components.

[0055] In this embodiment, the server system may further include Q processing units, which may be CPUs, DPUs, or other types of processors. This embodiment does not limit the specific processing units.

[0056] Optionally, each processing unit may correspond to a portion of the K acceleration units, and can be used to control the connected acceleration units to achieve load balancing and avoid overloading of the processing unit.

[0057] The communication link between the processing unit and the acceleration unit may be limited by the number of ports, transmission rate and latency of the processing unit. For example, taking the CPU as the processing unit, the CPU has limited I / O bandwidth, and directly connecting multiple acceleration cards will quickly exhaust the CPU's PCIe resources.

[0058] Optionally, each acceleration unit may further include R third switching components for connection to the corresponding processing unit, enabling data transmission between the processing unit and the acceleration unit. In each acceleration unit, all R third switching components may be connected to the corresponding processing unit. Each of the R third switching components may be connected to some of the M acceleration units. Different R third switching components may be connected to different M acceleration units. These components can receive data from multiple acceleration units and perform link aggregation, transmitting data from multiple acceleration units to a single processing unit. The third switching components can receive and redistribute data, reducing the burden on the processing unit through a more efficient data routing and distribution mechanism.

[0059] For example, such as Figure 5As shown, the server system may include processing units CPU0 and CPU1. Each acceleration unit may also include two third switching units (SW0 and SW1). Among the two processing units, CPU0 can correspond to acceleration units 0 and 1 through the third switching units, and CPU1 can correspond to acceleration units 2 and 3 through the third switching units. Furthermore, each acceleration unit in each acceleration unit can be connected to the first switching unit through the second switching unit to achieve non-blocking interconnection between acceleration units. Specifically, in each acceleration unit, each of the eight acceleration units can be connected to the second switching unit through a set of Gen5 x8 links. Each of the six second switching units can receive data transmitted through the eight sets of Gen5 x8 links and achieve non-blocking communication between acceleration units through the six first switching units. Here, in order to achieve uplink and downlink bandwidth matching, each second switching unit and the first switching unit can transmit data through eight sets of Gen6 links. The downlink transmission links of a single first switching unit can be a total of 32 sets of Gen6 x4 transmission links, thereby constructing a 16-card 32-component supernode system interconnection topology.

[0060] For example, Q processing components can be integrated onto a single board to achieve a modular design of the system.

[0061] In this embodiment, receiving and transmitting data to the processing unit via the processing unit and the third switching unit helps to improve the overall performance, communication efficiency, and scalability of the server system.

[0062] In one exemplary embodiment, the server system further includes a plurality of control units, the plurality of control units including:

[0063] The first control unit is connected to L first switching components and is used to control the L first switching components;

[0064] There are K second control units, each corresponding to one of the K acceleration units. Each of the K second control units is connected to the corresponding acceleration unit. Each second control unit is used to control the acceleration unit corresponding to it.

[0065] In related technologies, to achieve remote control of servers integrating a large number of GPU nodes and high-performance switching nodes, an out-of-band control scheme is typically adopted. This involves remotely controlling the server through a dedicated interface. In this scheme, the out-of-band control of the server is primarily handled by the Baseboard Management Controller (BMC) located on the CPU node. For example, ... Figure 6As shown, the baseboard management controller can be used to achieve overall system control, and it also needs to perform computing unit control, acceleration module control, switching unit control, power supply control, heat dissipation control, etc.

[0066] In the aforementioned out-of-band control scheme, once the BMC (Block Controller) fails, the control function of the entire server node will be lost, making it impossible to remotely monitor or control the server's status. This severely impacts operational efficiency and system availability. Furthermore, in this architecture, there is often only one control path; if this path fails, the control function of the entire server will be unusable, easily leading to control paralysis. In addition, this control scheme typically only provides relatively coarse monitoring and cannot meet the needs for fine-grained real-time monitoring of the health status of acceleration components such as GPUs. Therefore, it cannot provide early warnings of potential equipment failures. When a failure is detected, manual intervention is often required for fault diagnosis and recovery, a process that is inefficient, prolongs the time it takes for the server to return to normal operation, and negatively impacts business continuity and user experience.

[0067] To at least partially solve the above-mentioned technical problems, in this embodiment, the server system further includes multiple control units, which include: a first control unit connected to L first switching components and used to control the L first switching components; and K second control units, each of the K second control units corresponding to one of the K acceleration units, with each second control unit connected to the acceleration unit corresponding to each second control unit, and each second control unit used to control the acceleration unit corresponding to each second control unit.

[0068] Optionally, the first control unit can act as the main control unit, communicating with L first switching components for centralized control of the first switching components. The first switching components can be used to realize communication between acceleration units and to acquire communication data within the acceleration units. By centrally controlling the first switching components, the first control unit can effectively configure and optimize communication. Furthermore, the first control unit can monitor the status of the first switching components in real time. Once a fault is detected, it can quickly take measures, such as reconfiguring communication paths, activating redundant links, or replacing the faulty component, thereby improving the reliability and availability of the system.

[0069] Optionally, each of the K second control units can correspond to one acceleration unit, and can be used to refine control and optimize the internal communication and state of the corresponding acceleration unit. Each second control unit is responsible for the control of one acceleration unit. By allocating the second control units to each acceleration unit, the control complexity of the main control unit can be reduced, the excessive load of single-point control can be avoided, and the optimization and balancing of the system can be achieved.

[0070] Alternatively, the first control unit can be located in a separate module or integrated on a single board.

[0071] In this embodiment, server control is achieved through a first control unit and multiple second control units, which can establish multiple management network paths, distribute control pressure, and improve control efficiency.

[0072] In one exemplary embodiment, the server system further includes:

[0073] The local area network (LAN) switching component includes an external network interface. The LAN switching component is connected to each of the multiple control units. The LAN switching component is used to aggregate network signals from each control unit and exchange network signals via the external network interface.

[0074] In this embodiment, the server system may also include a local area network switching component (i.e., a LANSwitch module) for aggregating network signals from each control unit and exchanging network signals via an external network interface.

[0075] Optionally, the LAN switching component can aggregate network signals from the first control unit and multiple second control units. Regardless of the location of the control unit on the server, the aggregation of network signals can be achieved through the LANSwitch module, thereby enabling communication with external networks.

[0076] Optionally, the local area network switching component can interact with external networks via an external network interface.

[0077] Optionally, in the event of a network connection failure in the control unit, its control signals can be rerouted via other available links in the local area network switching component to ensure the continued availability of control functions.

[0078] Optionally, the LAN switching component can intelligently control internal network traffic, optimize the use of network resources through load balancing mechanisms, avoid overload of a single network path, and ensure efficient and stable data transmission.

[0079] In this embodiment, by connecting multiple control units through a local area network switching component, a complete control network topology can be constructed, improving control efficiency and system reliability.

[0080] In one exemplary embodiment, the server system further includes a programmable logic device and a path switcher. The programmable logic device is connected to a first control unit and the path switcher, respectively. The path switcher is connected to the Universal Asynchronous Receiver / Transmitter (UART) interface of the first control unit and the UART interface of each of a plurality of controlled components. The plurality of controlled components includes K second control units and L first switching components.

[0081] A programmable logic device is used to send a drive signal to a path switcher in response to the control of a first control unit. The drive signal is used to drive the path switcher to switch the UART interface connected to the UART interface of the first control unit among multiple UART interfaces of multiple controlled components.

[0082] A path switcher is used to switch, in response to a drive signal, the UART interface among multiple UART interfaces of multiple controlled components that is connected to the UART interface of the first control unit.

[0083] In this embodiment, the server system may also include a Complex Programmable Logic Device (CPLD) and a path switcher for switching control paths.

[0084] Optionally, the programmability of the CPLD allows its internal logic to be updated to adapt to changes in server architecture or adjustments to management strategies. In other words, the CPLD can dynamically adjust its control logic according to the actual needs of the server, providing flexible hardware support.

[0085] Optionally, the path switcher can automatically switch to the Universal Asynchronous Receiver-Transmitter (UART) interface of the specified controlled component based on the drive signal issued by the CPLD, and establish a direct communication link between the first control unit and the target component.

[0086] For example, the path switcher can be one of the following: an analog switch group or a multiplexer. Here, the CPLD can send control commands to the path switcher through its output port, and the path switcher (which can be an analog switch group or a multiplexer) performs physical switching according to the received commands. An analog switch group is an on / off switch that can be controlled by digital signals to connect or disconnect analog (or digital) signal paths. It can directly interrupt the current connection and change the connection point in various ways (such as electromagnetic relays, transistor configurations, etc.) to connect the TX (transmit) and RX (receive) lines of the UART port of the first control unit to the corresponding UART channel of the target controlled component. The multiplexer has multiple input terminals and a common output terminal. It can route signals to the target controlled component by selecting the correct input channel, realizing a direct communication link from the main control unit to the controlled component.

[0087] Through intelligent switching by the path switcher, the first control unit can directly access the controlled components, including multiple second control units and the first switching component, without having to traverse the entire network.

[0088] Optionally, the path switcher can provide redundant paths between the first control unit and the UART links of each controlled component. When a link fails, the path switcher can quickly switch to another available link according to the instructions of the CPLD, thereby achieving automatic fault avoidance and improving the reliability and availability of the system.

[0089] In this embodiment, by switching the communication path between the first control unit and the controlled component using a CPLD and a path switcher, control efficiency can be improved.

[0090] In one exemplary embodiment, the programmable logic device is further configured to receive a switching command from a first control unit via a general purpose input / output (GPIO) interface of the programmable logic device, wherein the switching command carries a component identifier of a target controlled component among a plurality of controlled components; parse the switching command and generate a drive signal based on the parsing result of the switching command, wherein the drive signal is used to drive a path switcher to switch the UART interface connected to the UART interface of the first control unit to the UART interface of the target controlled component.

[0091] Optionally, the CPLD can receive control signals from the first control unit. These control signals may include operation instructions for the path switcher, such as specifying the universal asynchronous transceiver interface of the controlled component to which the switch is to. This can be a switching command sent via a general purpose input / output (GPIO) interface connected to the CPLD, which includes a component identifier of the target controlled component. The component identifier of the target controlled component may be a component number or a location code, which can be used to specify the controlled component or the UART interface of the controlled component.

[0092] Based on the received control signals, the CPLD can internally perform logical judgments, decode the control signals through predefined logic circuits, and generate corresponding drive signals to control the operation of the path switcher. These drive signals can include direct control information for the path switcher, such as activating specific analog switches or setting the selection bits of the multiplexer.

[0093] When the drive signal from the CPLD reaches the path switcher, it can trigger the operation of the path switcher, such as triggering the corresponding analog switch or selecting the correct output path in the multiplexer. Thus, the UART interface of the first control unit can be connected to the UART interface of the target controlled component, establishing a direct communication link between the two. The first control unit can then directly access and control the target controlled component without going through a complex intermediate network or bus.

[0094] In this embodiment, by switching the communication path between the first control unit and the controlled component using a CPLD and a path switcher, control efficiency and response speed can be improved.

[0095] In one exemplary embodiment, the path switcher includes multiple signal paths and multiple channel switches, each channel switch corresponding to one of the multiple signal paths, and each signal path corresponding to one of the multiple controlled components. Each signal path is a signal path between the UART interface of the first control unit and the UART interface of the corresponding controlled component; wherein...

[0096] A path switcher is used to switch the switch state of the target channel switch corresponding to the target controlled component to the closed state in response to a drive signal, and to switch the switch state of the other channel switches in the multiple channel switches to the open state.

[0097] Optionally, the path switcher may include multiple signal paths and multiple channel switches, and can dynamically establish and disconnect UART communication links with each controlled component in response to drive signals generated by the CPLD. The multiple signal paths correspond one-to-one with the multiple controlled components, and each signal path can be directly connected to the UART interface of the first control unit and the UART interface of a controlled component in the system.

[0098] Optionally, multiple channel switches correspond one-to-one with multiple signal paths. That is, each channel switch can correspond to a specific signal path, and thus to a specific controlled component in the system. Each signal path can be equipped with a channel switch to control the on / off state of the signal path.

[0099] Optionally, the channel switch can be non-electronic (such as a relay) or electronic (such as a transistor), and its state can determine whether the first control unit can communicate with a specific controlled component through the signal path.

[0100] Optionally, the path switcher can dynamically control the state of the channel switch based on the drive signal to establish or disconnect UART communication with different controlled components.

[0101] Optionally, the path switcher can receive a drive signal from the CPLD. The drive signal may include a component identifier of the target controlled component, indicating the target controlled component for which a communication link needs to be established. Based on the component identifier of the target controlled component in the drive signal, the path switcher can switch the channel switch state corresponding to the target controlled component to closed. When the channel switch is closed, the signal path is activated, and a communication link is established between the UART interface of the first control unit and the UART interface of the target controlled component.

[0102] Optionally, the path switcher can switch all other channel switches except the target channel switch to the on state to ensure that only one signal path is in operation, thus avoiding signal conflicts.

[0103] Optionally, by closing the target channel switch, the first control unit can directly access the target controlled component and perform control operations. After communication is complete, it can also control the target channel switch to open again via a drive signal, disconnecting the communication link and preparing for the next communication switchover.

[0104] In this embodiment, the communication path between the first control unit and the controlled component is switched by CPLD and path switcher. Each controlled component only needs a dedicated signal path with the first control unit, which can improve control efficiency and response speed, simplify system design, reduce system design complexity, and reduce hardware costs.

[0105] In one exemplary embodiment, the server system further includes an inter-device bus (I2C) extension component, which is connected to the I2C interface of the first control unit and the I2C interface of each first switching component.

[0106] Optionally, the first control unit can be connected to the first switching component via an inter-integrated circuit (I2C).

[0107] In this embodiment, the server system may include an I2C expansion component to enhance the functionality of the I2C bus. Here, the I2C expansion component is a dedicated hardware or software component in the server system used to enhance and expand the I2C communication capability. The I2C expansion component connects to the I2C interface of the first control unit, and then to the I2C interface of each first switching component, forming an extended I2C network. This overcomes the limitations of the I2C bus itself, such as signal attenuation, communication distance, and the number of connectable devices, thus meeting the monitoring and control needs of a large number of devices within the server system.

[0108] This embodiment demonstrates how communication can be achieved via the I2C bus and enhanced through I2C expansion components, thereby building a unified control network and improving system reliability.

[0109] In one exemplary embodiment, the server system further includes L level conversion components, each corresponding to one of the L first switching components, and the improved inter-device bus I3C interface of the first control unit is connected to the I3C interface of each first switching component via the level conversion component corresponding to each first switching component.

[0110] Here, the Improved Inter-Integrated Circuit (I3C) is a unified sensor / control bus upgraded from I2C. It can provide higher bandwidth, lower power consumption, and more advanced addressing and in-band interrupt functions, and has a higher communication rate than the traditional I2C.

[0111] Alternatively, a level conversion component can be used to accommodate potential voltage differences.

[0112] Optionally, the level conversion component can convert the signal from one voltage level to another through a built-in voltage conversion circuit, ensuring the correct interpretation and transmission of the signal under different voltage environments. Through the level conversion component, the I3C signal emitted by the first control unit can be adjusted from its native voltage level to the voltage level adapted to the first switching component, and vice versa.

[0113] Optionally, the improved inter-device bus I3C interface of the first control unit can be connected to the I3C interface of each first switching component via a level conversion component corresponding to each first switching component. This not only simplifies system design and wiring but also reduces system complexity.

[0114] In this embodiment, communication between the first control unit and the first calling unit is achieved through L level conversion components and the I3C interface, which can improve communication efficiency.

[0115] In one exemplary embodiment, the server system further includes a Peripheral Component Interconnect Bus (PCIe) switching component, which is connected to the PCIe root complex of the first control unit; wherein...

[0116] A PCIe switching component is used to expand the PCIe port connected to the PCIe root complex of the first control unit into L PCIe ports, wherein the L PCIe ports are connected one-to-one with the L PCIe ports of the first switching component.

[0117] Here, the PCIe switching component is a hardware device used to increase the connectivity of the PCIe bus, allowing multiple devices to connect to the same PCIe root complex simultaneously.

[0118] Optionally, the PCIe root complex can be located in the first control unit, and the PCIe switching component can be used to expand a single PCIe port of the root complex into multiple ports, thereby enabling the access of more peripheral devices.

[0119] Optionally, the PCIe switching component can be used to expand the PCIe ports connected to the PCIe root complex of the first control unit to L PCIe ports, wherein the L PCIe ports are connected one-to-one with the L PCIe ports of the first switching component. The PCIe switching component can be used to implement signal routing and control, ensuring that data is transmitted correctly between the first control unit and the first switching component. It can dynamically allocate bandwidth resources and handle packet forwarding according to the PCIe protocol, providing a flexible and efficient communication mechanism.

[0120] Optionally, by connecting the PCIe ports of multiple first switching components to the PCIe switching component, bandwidth aggregation can be achieved, and dispersed PCIe resources can be centrally managed, thereby improving overall communication efficiency.

[0121] Similar to the previous embodiments, the server system may also include I2C expansion components and I3C expansion components. Correspondingly, the first control unit can be connected to the first switching component through an I2C interface, an I3C interface, or a PCIe port, or directly achieve point-to-point communication with the first switching component through a UART interface. By integrating I2C, I3C, and PCIe links into a server system architecture, redundant communication paths are achieved by providing multiple communication links. In the event of a single port failure, the system automatically switches to other links. Even if one link fails, the system can still communicate and control through other links. Corresponding to different types of communication needs applicable to different links, the different advantages of each link can be utilized. For example, I3C is suitable for high-speed and device status monitoring, while PCIe is suitable for high-speed data transmission and control of high-performance computing devices.

[0122] For example, such as Figure 7 As shown, each acceleration component (which may be a GPU) has a corresponding second control unit. The first control unit can communicate with the acceleration component or the first switching component via an I2C interface, an I3C interface, or a PCIe interface. Communication can be achieved by expanding the PCIe interface using a PCIe switching component. The first control unit can interact with multiple second control units via network signals through the external network interface of the LAN switching component, or it can achieve direct communication path switching with multiple second control units through programmable logic devices and path switches.

[0123] In this embodiment, the first control unit directly controls L first switching components via PCIe and acts as a connection point to facilitate communication between the first switching components. This allows for the construction of a more complex and efficient network architecture, enabling high-speed data transmission between components within the server system.

[0124] In one exemplary embodiment, a first control unit is configured to control the server system when the first control unit is operating normally;

[0125] At least one of the K second control units is used to take over the control functions of the first control unit in the event of a failure of the first control unit.

[0126] In related technologies, in order to achieve redundant control of the system, two control units can be set for a single node, one as the main control unit and the other as a backup control unit, which is used to take over in the event of failure of the main control unit to achieve redundant control. This design is both complex and costly.

[0127] In this embodiment, only one second control unit is provided for each acceleration unit. The second control unit can be used to take over the control function of the first control unit in the event of a failure of the first control unit.

[0128] Optionally, when the first control unit is in normal operation, it is used to control the server system while the first control unit is operating normally. It can be used to coordinate all activities within the server and control the interaction between various hardware components and software services.

[0129] Optionally, the fault status of the first control unit can be checked periodically, or the first control unit can proactively report a fault. If a fault is determined to have occurred in the first control unit, at least one second control unit can be activated to take over the control functions of the first control unit.

[0130] Alternatively, the link switching of the control unit can be achieved through multiplexers and logic control, thereby reducing hardware redundancy and lowering design complexity and cost.

[0131] Optionally, if at least one of the second control units used to take over the control functions of the first control unit fails, other second control units among the K second control units can be activated to take over, and so on, until the last second control unit in the system.

[0132] In this embodiment, since each node only needs one control unit, resource utilization is more concentrated and efficient, which simplifies system design and fault recovery process, improves the system's self-healing ability, and enhances the overall reliability of the system.

[0133] In one exemplary embodiment, K second control units form a backup chain in a specified order;

[0134] The first second control unit on the backup chain is used to take over the control functions of the first control unit in the event of a failure of the first control unit;

[0135] The backup chain includes other second control units besides the first second control unit, which are used to take over the control functions of the first control unit and the second control units preceding the other second control units in the event of a failure of the first control unit and the other second control units preceding the other second control units.

[0136] Optionally, the server system may have a preset backup link management mechanism based on K second control units. The backup link refers to a redundant link formed by K second control units arranged in a preset order, which is used to back up the control functions of the first control unit.

[0137] Optionally, the backup chain can be formed based on a preset order, starting from the first second control unit in the second control unit and continuing until the Kth second control unit. This preset order can be based on the performance, location, or other system indicators of the second control units, and is not limited in this embodiment.

[0138] Optionally, in the event of a failure of the first control unit, the first second control unit on the backup chain can take over all control functions of the first control unit to ensure that the control of the server system is not interrupted.

[0139] Optionally, when the first control unit and the first second control unit on the backup chain fail successively, the second second control unit on the backup chain can take over the control functions of the first control unit, as well as the control functions of the first second control unit before the failure. Similarly, if the first control unit and all other second control units preceding it fail, the other second control units on the backup chain (excluding the first second control unit) can take over the control functions of the first control unit, as well as the control functions of the second control units preceding it. This process can proceed sequentially according to the order of the second control units on the backup chain, meaning that the later the second control unit is, the more control functions it can cumulatively take over.

[0140] Optionally, when each second control unit takes over, it not only takes over the functions of the first control unit, but also the functions of all the failed second control units before it, which can form a dynamic function transfer chain, ensuring the continuity and integrity of the control functions. Even in the case of multi-level faults, the system can still maintain basic operation through the subsequent second control units.

[0141] This embodiment simplifies the design complexity of the server system, enhances system reliability and fault tolerance, and improves resource utilization through the backup chain mechanism of K second control units.

[0142] In one exemplary embodiment, in each acceleration unit, a first communication link includes N first transmission channels, and a second communication link includes P second transmission channels. The transmission rate of one of the N first transmission channels and the transmission rate of one of the P second transmission channels are different.

[0143] Optionally, a first communication link may include N first transmission channels, and a second communication link may include P second transmission channels. Here, multiple transmission channels can form a first communication link, constituting a communication port. A transmission channel may include two differential pairs, and a transmission channel may be an independent bidirectional serial channel. Each transmission channel may have an independent data stream, but at the protocol layer, multiple transmission channels can operate simultaneously to transmit different parts of the same batch of data. Multiple transmission channels can be combined into a communication port or a communication link.

[0144] For example, a Gen5 x16 port can include 16 independent PCIe lanes, each operating at Gen5 speeds.

[0145] Alternatively, a communication link can be divided into multiple communication links with smaller bandwidths. For example, a Gen6x16 communication link can be divided into two Gen6x8 communication links.

[0146] In this embodiment, by combining the first transmission channel into a communication link, the number of communication ports can be reduced, and the system design can be simplified.

[0147] Embodiments of this application also provide a data transmission method for a server system. The server system includes: K acceleration units and L first switching components. Each of the K acceleration units includes M acceleration components and L second switching components. In each acceleration unit, each of the M acceleration components is connected to the L second switching components via L first communication links. Each of the L second switching components is connected to the M acceleration components via M first communication links. The L second switching components correspond one-to-one with the L first switching components. Each of the L second switching components is connected to its corresponding first switching component via M second communication links. The transmission rate of one of the L first communication links is different from the transmission rate of one of the M second communication links. Each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit via M second communication links. K, L, M, N, and P are all positive integers greater than or equal to 2. Figure 8 This is a flowchart illustrating a data transmission method for a server system according to an embodiment of this application, as shown below. Figure 8 As shown, the method includes:

[0148] Step S802: Obtain the data to be transmitted through the first acceleration component in the first acceleration unit among the K acceleration units, wherein the data to be transmitted is the data to be transmitted to the second acceleration component in the second acceleration unit among the K acceleration units;

[0149] Step S804: In response to the data to be transmitted, the data to be transmitted is transmitted to the second acceleration unit via the data transmission path from the first acceleration unit to the second acceleration unit, wherein the data transmission path includes a second switching unit in the first acceleration unit, a first switching unit in one of L first switching units, and a second switching unit in the second acceleration unit.

[0150] It should be noted that step S802 in this embodiment can be executed by the first acceleration component in the first acceleration unit among the K acceleration units 101 in the foregoing embodiments. For a description of the features corresponding to the acceleration unit 101 in the foregoing embodiments, please refer to the relevant descriptions of the embodiments corresponding to the acceleration unit 101, which will not be repeated here.

[0151] For example, the computing component can be an accelerator card, and the transmission channel supported by the accelerator card can be a high-speed peripheral component interconnect transmission channel.

[0152] Optionally, for a description of the features in the embodiment corresponding to the computing component, please refer to the relevant description of the embodiment corresponding to the computing component in the foregoing embodiments, which will not be repeated here.

[0153] The embodiments provided in this application obtain data to be transmitted through a first acceleration component in a first acceleration unit among K acceleration units, wherein the data to be transmitted is data to be transmitted to a second acceleration component in a second acceleration unit among K acceleration units; in response to the data to be transmitted, the data to be transmitted is transmitted to the second acceleration component via a data transmission path from the first acceleration component to the second acceleration component, wherein the data transmission path includes a second switching component in the first acceleration unit, a first switching component in one of L first switching components, and a second switching component in one of the second acceleration units. This can solve the problems of high complexity and limited interconnection in the related technologies, and achieve the technical effect of increasing the scale of the extended interconnection and reducing the system complexity.

[0154] In one exemplary embodiment, the server system further includes a plurality of control units, which include a first control unit and K second control units. The first control unit is connected to L first switching components, and the K second control units correspond one-to-one with K acceleration units. Each of the K second control units is connected to the acceleration unit corresponding to each second control unit. The method further includes: controlling the L first switching components through the first control unit; and controlling the acceleration unit corresponding to each second control unit through each second control unit.

[0155] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0156] The foregoing has provided a detailed description of a server system and a data transmission method for that server system. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server system, characterized by Comprise: K acceleration units and L first switching components, each of the K acceleration units comprises M acceleration components and L second switching components; wherein, In each of the acceleration units, each of the M acceleration components is connected to the L second switching components respectively through L first communication links; each of the L second switching components is connected to the M acceleration components respectively through M first communication links; the L second switching components correspond to the L first switching components one by one, each of the L second switching components is connected to the corresponding first switching component through M second communication links, the transmission rate of one of the L first communication links is different from the transmission rate of one of the M second communication links; Each of the L first switching components is connected to the second switching component corresponding to each first switching component in each acceleration unit through M second communication links respectively. Wherein, K, L and M are positive integers greater than or equal to 2.

2. The server system of claim 1, wherein, In each of the acceleration units, the M acceleration components are packaged into an accelerator module in groups of at least two acceleration components.

3. The server system of claim 1, wherein, Each of the acceleration units is integrated into a board card, and different acceleration units in the K acceleration units are integrated into different board cards.

4. The server system of claim 1, wherein, Each of the first switching components is integrated into a board card, and different first switching components in the L first switching components are integrated into different board cards.

5. The server system of claim 1, wherein, The server system further comprises Q processing components, each of the acceleration units further comprises R third switching components, each of the Q processing components corresponds to part of the K acceleration units, different processing components in the Q processing components correspond to different acceleration units in the K acceleration units, the R third switching components of each acceleration unit are connected to the processing component corresponding to each acceleration unit, Q and R are positive integers greater than or equal to 1; wherein, In each of the acceleration units, the R third switching components are connected to the processing component corresponding to each acceleration unit, each of the R third switching components is connected to part of the M acceleration components, and different third switching components in the R third switching components are connected to different acceleration components in the M acceleration components.

6. The server system of claim 5, wherein, The Q processing components are integrated into a board card.

7. The server system of claim 1, wherein, The server system further comprises a plurality of control units, the plurality of control units comprise: A first control unit connected to the L first switching components for controlling the L first switching components; K second control units, the K second control units correspond to the K acceleration units one by one, each of the K second control units is connected to the acceleration unit corresponding to each second control unit, and each second control unit is used for controlling the acceleration unit corresponding to each second control unit.

8. The server system of claim 7, wherein, The server system further comprises: A local area network switching component, wherein the local area network switching component comprises an external network interface, the local area network switching component is connected with each of the plurality of control units, and the local area network switching component is configured to aggregate network signals of the each of the plurality of control units and interact with the network signals via the external network interface.

9. The server system of claim 7, wherein, The server system further comprises a programmable logic device and a path switcher, the programmable logic device is connected with the first control unit and the path switcher respectively, the path switcher is connected with a universal asynchronous receiver transmitter (UART) interface of the first control unit and a UART interface of each of a plurality of controlled components, the plurality of controlled components comprising the K second control units and the L first switching components; wherein, The programmable logic device is configured to send a driving signal to the path switcher in response to control of the first control unit, wherein the driving signal is used to drive the path switcher to switch a UART interface connected with the UART interface of the first control unit from the plurality of UART interfaces of the plurality of controlled components. The path switcher is configured to switch the UART interface connected with the UART interface of the first control unit from the plurality of UART interfaces of the plurality of controlled components in response to the driving signal.

10. The server system of claim 9, wherein, The programmable logic device is further configured to receive a switching command of the first control unit through a general-purpose input / output (GPIO) interface of the programmable logic device, wherein the switching command carries a component identifier of a target controlled component from the plurality of controlled components; analyze the switching command, and generate the driving signal based on an analysis result of the switching command, wherein the driving signal is used to drive the path switcher to switch the UART interface connected with the UART interface of the first control unit to the UART interface of the target controlled component.

11. The server system of claim 10, wherein, The path switcher comprises a plurality of signal paths and a plurality of channel switches, the plurality of channel switches correspond to the plurality of signal paths one by one, the plurality of signal paths correspond to the plurality of controlled components one by one, and each of the plurality of signal paths is a signal path between the UART interface of the first control unit and a UART interface of the corresponding controlled component of the each of the plurality of signal paths; wherein, The path switcher is configured to switch a switch state of a target channel switch corresponding to the target controlled component from the plurality of channel switches to a closed state in response to the driving signal, and switch switch states of other channel switches from the plurality of channel switches except the target channel switch to an open state.

12. The server system of claim 9, wherein, The path switcher is one of an analog switch group and a multiplexer.

13. The server system of claim 7, wherein, The server system further comprises an inter-device bus (I2C) extension component, the I2C extension component is connected with an I2C interface of the first control unit and an I2C interface of the each of the first switching components respectively.

14. The server system of claim 7, wherein, The server system further comprises L level conversion components corresponding to the L first switching components one by one, and the I3C interface of the improved inter-device bus of the first control unit is connected with the I3C interface of each first switching component via the level conversion component corresponding to each first switching component.

15. The server system of claim 7, wherein, The server system further comprises a peripheral component interconnect bus (PCIe) switching component connected with the PCIe root complex of the first control unit; wherein, The PCIe switching component is configured to expand the PCIe port connected with the PCIe root complex of the first control unit into L PCIe ports, wherein the L PCIe ports are connected with the PCIe ports of the L first switching components one by one.

16. The server system of claim 7, wherein, The first control unit is configured to control the server system in a normal operation mode of the first control unit; At least one of the K second control units is configured to take over the control function of the first control unit in a failure mode of the first control unit.

17. The server system of claim 16, wherein, The K second control units form a backup chain in a specified order; The first second control unit in the backup chain is configured to take over the control function of the first control unit in a failure mode of the first control unit. The second control units other than the first second control unit in the backup chain are configured to take over the control function of the first control unit and the control function of the second control units before the other second control units in a failure mode of the first control unit and the second control units before the other second control units.

18. The server system of any one of claims 1 to 17, wherein, In each acceleration unit, a first communication link comprises N first transmission channels, and a second communication link comprises P second transmission channels, wherein the transmission rate of a first transmission channel in the N first transmission channels is different from the transmission rate of a second transmission channel in the P second transmission channels; and N and P are both positive integers greater than or equal to 2.

19. A data transmission method of a server system, characterized by, The server system comprises K acceleration units and L first switching components, each of the K acceleration units comprises M acceleration components and L second switching components; wherein in each of the acceleration units, each of the M acceleration components is connected to the L second switching components respectively through L first communication links; each of the L second switching components is connected to the M acceleration components respectively through M first communication links; the L second switching components correspond to the L first switching components one by one, each of the L second switching components is connected to the corresponding first switching component through M second communication links, one first communication link comprises N first transmission channels, one second communication link comprises P second transmission channels, the transmission rate of one first transmission channel in the N first transmission channels is different from the transmission rate of one second transmission channel in the P second transmission channels; each of the L first switching components is connected to the second switching component corresponding to the first switching component in each of the acceleration units through M second communication links respectively; K, L, M, N and P are all positive integers greater than or equal to 2; the method comprises: acquiring, by a first acceleration component in a first acceleration unit of the K acceleration units, to-be-transmitted data, wherein the to-be-transmitted data is data to be transmitted to a second acceleration component in a second acceleration unit of the K acceleration units; in response to the to-be-transmitted data, transmitting the to-be-transmitted data to the second acceleration component via a data transmission path from the first acceleration component to the second acceleration component, wherein the data transmission path comprises one second switching component in the first acceleration unit, one first switching component in the L first switching components and one second switching component in the second acceleration unit.

20. The method of claim 19, wherein, The server system further comprises a plurality of control units, the plurality of control units comprises a first control unit and K second control units, the first control unit is connected to the L first switching components, the K second control units correspond to the K acceleration units one by one, each of the K second control units is connected to the acceleration unit corresponding to the second control unit; The method further comprises: controlling the L first switching components by the first control unit; controlling the acceleration unit corresponding to each second control unit by each second control unit.

Citation Information

Patent Citations

  • Multi-accelerator card computing power test communication management method

    CN118860923A

  • Accelerator cluster super node device and computing acceleration device

    CN222814492U