Multifunctional configuration method suitable for orthogonal server cabinet and orthogonal server cabinet

By introducing external switching nodes into the orthogonal server rack and adjusting the data transmission protocol, the problem of limited computing power was solved, enabling more efficient data processing and computational task completion.

CN121056418BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511580999.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-01-27
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

The maximum data computing capacity of an orthogonal server rack is constrained by the number of computing nodes and switching nodes, as well as bandwidth, and cannot meet the needs of computing tasks with large computational loads.

Method used

By introducing external switching nodes into the orthogonal server rack and using different protocols for data transmission, full interconnection between computing nodes and internal switching nodes is achieved, and the data transmission protocol is dynamically adjusted to expand computing capabilities.

Benefits of technology

Without altering the rack structure, fully utilize internal and external network resources to accelerate data processing and computation tasks, meeting the demands of computationally intensive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056418B_ABST
    Figure CN121056418B_ABST
Patent Text Reader

Abstract

The application discloses a multifunctional configuration method suitable for an orthogonal server cabinet and the orthogonal server cabinet, relates to the technical field of servers, and generates a horizontal expansion function signal with a first value through a first controller of a computing node according to the calculation amount of a current computing task, and then sends the horizontal expansion function signal to a second controller to control the uplink port to be opened, so that the external switching node and the internal switching node are used to exchange and transmit data of the computing node to realize a service function. The maximum data calculation capacity of the orthogonal server cabinet is improved through external expansion. The problem that the maximum data calculation capacity of the orthogonal server cabinet is constrained by the number of computing nodes and switching nodes and bandwidth in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a multi-functional configuration method suitable for orthogonal server racks, an orthogonal server rack and a server system. Background Technology

[0002] Orthogonal architecture chassis with 90-degree interlocking front and rear connections are becoming an emerging trend in the industry. In orthogonal architecture, the combination of compute nodes and switching nodes is relatively fixed. Furthermore, the total number of compute node mounting slots and the total number of switching node mounting slots are also fixed. This means that the maximum data computing capacity of orthogonal server racks in related technologies is constrained by the number of compute nodes, switching nodes, and bandwidth. Summary of the Invention

[0003] This application provides a multi-functional configuration method for orthogonal server racks, an orthogonal server rack and a server system, to at least solve the problem in related technologies where the maximum data computing capacity of orthogonal server racks is constrained by the number of computing nodes and switching nodes and bandwidth.

[0004] This application provides a multi-functional configuration method applicable to orthogonal server racks. The orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. Each compute node includes a first controller, and each internal switching node includes a second controller and has an uplink port for connecting to an external switching node, which is a switching node other than the internal switching node. The method includes: the first controller acquiring business requirements, including the computational load of the current compute task; if the computational load of the current compute task is greater than a preset computational load, the first controller sends a horizontal expansion function signal with a first value to the second controller, so that the second controller controls the uplink port to open and sets the data transmission protocol of the uplink port to a first protocol, thereby using the external switching node and the internal switching node to exchange and transmit data between the compute nodes to realize business functions. The first protocol supports the transmission of data from the compute node to the external switching node, and the transmission protocol between the internal switching node and the compute node is a second protocol. The first protocol and the second protocol are different.

[0005] This application provides a multi-functional configuration method applicable to orthogonal server racks. The current orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. The compute nodes include a first controller, and the internal switching nodes include a second controller and have an uplink port. The uplink port is used to connect to external switching nodes, which are switching nodes other than the internal switching nodes of the current orthogonal server rack. The method includes: the second controller receiving a horizontal expansion function signal with a first value sent by the first controller. The horizontal expansion function signal with the first value is generated by the first controller when the business requirement indicates that the computational workload of the current computing task is greater than a preset computational workload; controlling the uplink port to open according to the horizontal expansion function signal with the first value and setting the data transmission protocol of the uplink port to a first protocol, thereby using the external switching nodes and internal switching nodes to exchange and transmit data of the compute nodes to realize business functions. The first protocol supports the transmission of data from the compute nodes to the external switching nodes, and the transmission protocol between the internal switching nodes and the compute nodes is a second protocol. The first protocol and the second protocol are different.

[0006] This application also provides an orthogonal server rack, comprising: a plurality of orthogonally connected compute nodes and a plurality of internal switching nodes, wherein the internal switching nodes have uplink ports for connecting to external switching nodes, the external switching nodes being switching nodes other than the internal switching nodes, each compute node including a first controller, and each switching node including a second controller; the first controller is used to execute the multi-functional configuration method of the first aspect applicable to the orthogonal server rack; the second controller is used to execute the multi-functional configuration method of the second aspect applicable to the orthogonal server rack.

[0007] This application also provides a server system, including: any type of orthogonal server rack; an external switching node, wherein the internal switching node in the orthogonal server rack is connected to the external switching node through an uplink port.

[0008] This application enables the use of a first protocol to transmit data from the compute node to the external switching node after the uplink port is opened, while a second protocol ensures data transmission between the compute node and the internal switching node, providing a fully interconnected connection. By employing the first protocol for horizontal scaling and combining it with the second protocol to maintain internal communication, the compute node can fully utilize internal and external network resources, accelerating data processing and computational task completion to achieve business objectives. For computationally intensive tasks, the requirements can be met by connecting to external switching nodes without altering the current orthogonal server rack structure. This solves the problem in related technologies where the maximum data computing capacity of orthogonal server racks is constrained by the number of compute nodes and switching nodes, as well as bandwidth. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating a multi-functional configuration method suitable for orthogonal server racks, provided in this application embodiment;

[0011] Figure 2 A schematic diagram illustrating bandwidth allocation between a first type of computing node and an internal switching node, provided for embodiments of this application;

[0012] Figure 3 A schematic diagram of a first orthogonal server rack switch installation scheme provided in this application embodiment;

[0013] Figure 4 A schematic diagram of a second orthogonal server rack switch installation scheme provided in this application embodiment;

[0014] Figure 5 A schematic diagram illustrating bandwidth allocation between a second type of computing node and an internal switching node, provided in an embodiment of this application;

[0015] Figure 6 A schematic diagram of a server system provided in an embodiment of this application;

[0016] Figure 7 This is a schematic diagram of an orthogonal server rack provided in an embodiment of this application;

[0017] Figure 8 A schematic diagram of the pin definition of the connector provided in the embodiments of this application;

[0018] Figure 9 A schematic diagram of the identification node slot provided in an embodiment of this application;

[0019] Figure 10 A flowchart illustrating another multi-functional configuration method applicable to orthogonal server racks, provided as an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] To address the technical problem mentioned in the background section: the maximum data computing capacity of an orthogonal server rack is constrained by the number of computing nodes and switching nodes, as well as bandwidth.

[0024] The first aspect of this application provides a multi-functional configuration method suitable for orthogonal server racks, which is executed by a first controller in a computing node in the current orthogonal server rack. Specifically, the first controller may be a first BMC (Baseboard Management Controller).

[0025] The current orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. The compute nodes include a first controller, and the internal switching nodes include a second controller and have an uplink port. The uplink port is used to connect to external switching nodes, which are switching nodes other than the internal switching nodes.

[0026] The difference between internal and external switching nodes in this solution lies in their installation location. Internal switching nodes are those located inside the current orthogonal server rack, while external switching nodes are those located outside the current orthogonal server rack. In other words, external switching nodes do not belong to the current orthogonal server rack but are located outside other server racks. For example, if there are two orthogonal server racks, designated as the first orthogonal server rack and the second orthogonal server rack, then the switching nodes in the first orthogonal server rack are considered external switching nodes relative to the second orthogonal server rack.

[0027] like Figure 1 As shown, the method includes:

[0028] Step S101: The first controller obtains the business requirements, which include the computational load of the current computing task;

[0029] The first controller monitors and analyzes the running status, task scheduling information, and GPU (Graphics Processing Unit) (also referred to as accelerator) usage on computing nodes to perceive changes in business needs in real time and obtain the computational load of the current computing task.

[0030] The computational load of the current computing task can be comprehensively evaluated by considering task characteristics (such as algorithm complexity), dataset size, expected processing time, and resource consumption metrics (such as CPU, GPU, and memory utilization).

[0031] Step S102: If the computational load of the current computing task exceeds the preset computational load, the first controller sends a horizontal expansion function signal with a first value to the second controller. This causes the second controller to open the uplink port and set the data transmission protocol of the uplink port to the first protocol. This allows external and internal switching nodes to exchange and transmit data between the computing nodes to achieve business functions. The first protocol supports data transmission from the computing node to the external switching node, while the transmission protocol between the internal switching node and the computing node is the second protocol. The first and second protocols are different. For example, the first value can be 1. Of course, the first value can also be set to other values ​​according to actual needs. Setting the first value is mainly to ensure that the second controller opens the uplink port when it receives the horizontal expansion function signal with the first value.

[0032] Specifically, the difference between installing an external switching node and other server racks within the current orthogonal server rack is that these racks also have compute nodes installed. Opening the uplink port enables a connection between the external and internal switching nodes. Data from the compute nodes in the current orthogonal server rack can then be transmitted to compute nodes in other server racks via the external switching node. This allows compute nodes in the current orthogonal server rack and other server racks to participate in computation, thus improving computing power and enabling business functions to be implemented even with large computational demands. This approach can be applied to the computation of large AI models.

[0033] A preset computational load is used as the trigger threshold for opening the uplink port. If the computational load of the current task exceeds the preset threshold, the first controller will detect this change and trigger the configuration adjustment process. The first controller sends a horizontal expansion function signal with a first value to the second controller. This signal includes the uplink port opening command and protocol switching requirements to activate the support of external switching nodes and enhance data transmission capabilities.

[0034] The first protocol is designed for horizontal scaling and can effectively support large data volume transmission between computing nodes and external switching nodes. It typically has higher transmission rates and lower latency to facilitate the handling of computationally intensive tasks.

[0035] The second protocol is a standard data transmission protocol used between internal switching nodes and compute nodes, which optimizes communication efficiency within the rack.

[0036] The preset computational load is derived from historical data analysis and statistical modeling. It is used to set the computational load standard under normal operating conditions and serves as a benchmark for dynamic system adjustment.

[0037] In some implementations, the design of the first and second protocols can be combined with different application scenarios. The former focuses on high bandwidth and scalability, while the latter focuses on the efficiency and stability of internal communication. The differences between them ensure the flexibility and responsiveness of the system under different computing requirements.

[0038] After the uplink port is opened, the first protocol is used to transmit data from the compute node to the external switch node, while the second protocol ensures data transmission between the compute node and the internal switch node, and the compute node and the internal switch node are fully interconnected. By using the first protocol for horizontal scaling and combining it with the second protocol to maintain internal communication, the compute node can fully utilize internal and external network resources, accelerating data processing and the completion of computing tasks, thereby achieving business goals. For computationally intensive tasks, the requirements can be met by connecting to the external switch node without changing the current orthogonal server rack structure. This solves the problem in related technologies where the maximum data computing capacity of orthogonal server racks is constrained by the number of compute nodes and switch nodes and bandwidth.

[0039] In some embodiments, the method further includes:

[0040] The first controller sends a slot support capability signal to the second controller, so that the second controller, under the condition of satisfying the first preset condition, determines the horizontal expansion capability of the switching node based on the slot support capability signal, the total number of uplink ports that the internal switching node can provide, and the bandwidth of each internal switching node. The horizontal expansion capability represents the maximum bandwidth that each switching node supports for external expansion.

[0041] Wherein, the slot support capability signal represents the maximum number of switching node slots supported by the current orthogonal server rack, the first preset condition represents that the bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth, the sum of the allocated downlink bandwidth of all internal switching nodes is equal to the total external bandwidth of the computing node, the port of the internal switching node includes the uplink port and the downlink port, and the downlink port supports the implementation of the allocated downlink bandwidth.

[0042] In some schemes, the allocated uplink bandwidth and allocated downlink bandwidth of internal switching nodes are set to be equal;

[0043] The sum of the allocated downlink bandwidth of all internal switching nodes equals the total external bandwidth of the compute node, signifying full interconnection between the compute node and the internal switching nodes. This can be understood as the total number of ports on the internal switching nodes being fixed; ensuring full interconnection between the compute node and the internal switching nodes must be prioritized, meaning downlink ports are configured to support full interconnection first. Only after achieving full interconnection are the remaining ports configured as uplink ports.

[0044] The downlink port supports the implementation of the allocated downlink bandwidth, as explained below:

[0045] For example, an internal switching node has a total of 32 ports and a total bandwidth of 25.6T. In order to achieve full interconnection between the compute node and the internal switching node, that is, to make the sum of the allocated downlink bandwidth of all the internal switching nodes equal to the total external bandwidth of the compute node, the allocated downlink bandwidth needs to be set to 12.8T. To support the transmission of data with an allocated downlink bandwidth of 12.8T, 16 ports are needed. Therefore, 16 of the 32 ports of the internal switching node are used as downlink ports, and the other 16 are used as uplink ports.

[0046] The total number of ports in an internal switching node is limited by the size of the internal switching node; the size here can be understood as the volume of the internal switching node.

[0047] See Figure 2 and Figure 3 With the addition of two extra switch slots (equivalent to the switch node slots mentioned in the text), a maximum of 6 internal switch nodes can be installed in this configuration. The slot support capability signal indicates that the maximum number of switch node slots supported by the current orthogonal server rack is 6. See also... Figure 3 Due to the limitations of the chassis size and the need to meet the first preset condition, each internal switching node can provide a maximum of 16 uplink ports.

[0048] See Figure 4 and Figure 5In this scenario, a maximum of four internal switching nodes can be installed. The slot support capability signal indicates that the maximum number of switching node slots supported by the current orthogonal server rack is four. See also... Figure 4 Due to the limitations of the chassis size and the need to meet the first preset condition, each internal switching node can provide a maximum of 32 uplink ports.

[0049] Each uplink port can expand the bandwidth by the same amount, so 32 uplink ports can expand the bandwidth by a maximum of twice that of 16 uplink ports.

[0050] The horizontal scaling capability of the switching node is determined based on the slot support capability signal, the total number of uplink ports that the internal switching node can provide, and the bandwidth of each internal switching node. Specifically, this is achieved as follows:

[0051] The bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth. Figure 2 and Figure 3 The maximum number of slots in the internal switching node is 6, but two slots are idle. The allocated uplink bandwidth and allocated downlink bandwidth of the internal switching node are both 12.8T. Each internal switching node provides 16 uplink ports, which can expand the 12.8T allocated downlink bandwidth to the outside, that is, achieve 100% expansion.

[0052] The bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth. If the maximum number of switching node slots is 6 and there are no free slots, the allocated uplink bandwidth and allocated downlink bandwidth of the internal switching node are both 12.8T. Each internal switching node provides 16 uplink ports, which can expand the 12.8T allocated downlink bandwidth to the outside, or achieve 100% expansion.

[0053] The bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth. Figure 4 and Figure 5 The number of internal switching node slots is 4. The allocated uplink bandwidth and allocated downlink bandwidth of the internal switching node are both 25.6T. Each internal switching node provides 32 uplink ports, which can expand the 25.6T allocated downlink bandwidth to the outside, that is, achieve 100% expansion.

[0054] The bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth. The maximum number of switching node slots is 6, with no free slots. The allocated uplink bandwidth and allocated downlink bandwidth of the internal switching node are both 25.6T. Due to the chassis size limitation, each internal switching node can only provide 16 uplink ports to expand half of the 25.6T allocated downlink bandwidth to the outside, i.e., to achieve a 50% expansion.

[0055] Therefore, the horizontal expansion capability of a switching node cannot be determined solely based on the slot support capability signal. It is also necessary to combine the total number of uplink ports that the internal switching nodes can provide and the bandwidth of each internal switching node to determine the horizontal expansion capability of the switching node. Of course, these judgments need to be based on the first preset condition.

[0056] In this embodiment of the application, each computing node is provided with a plurality of first connectors, and each switching node is provided with a plurality of second connectors. The first connectors and the second connectors are plugged into each other. The pin definitions of the first connectors and the second connectors are the same. Each first connector includes a slot support capability identification pin and an expansion identification pin. The slot support capability identification pin is used to transmit a slot support capability signal, and the expansion identification pin is used to transmit a horizontal expansion function signal.

[0057] If the computational load of the current computing task is greater than the preset computational load, the first controller sends a horizontal expansion function signal with a first value to the second controller, including: if the computational load of the current computing task is greater than the preset computational load, the first controller sends the horizontal expansion function signal with the first value to the second controller through the expansion identifier pin.

[0058] The first controller sends the slot support capability signal to the second controller, including: the first controller sends the slot support capability signal to the second controller through the slot support capability identification pin.

[0059] For orthogonal connections implemented via orthogonal connectors (including the first connector and the second connector), the pins of the orthogonal connectors can be defined to obtain slot support capability identification pins and expansion identification pins. The slot support capability identification pins are used to transmit slot support capability signals, and the expansion identification pins are used to transmit lateral expansion function signals. This ensures efficient and accurate data transmission between compute nodes and switching nodes.

[0060] In this embodiment of the application, before the first controller sends the slot support capability signal to the second controller, the method further includes:

[0061] The first controller obtains the total external bandwidth of all computing nodes;

[0062] Under the condition that the second preset condition is met, the first controller determines the number of internal switching nodes to be installed corresponding to all computing nodes based on the total external bandwidth, the width of the switching node slots, the total number of switching node slots under each width of the switching node slots, and the bandwidth of each internal switching node.

[0063] The first controller determines the maximum number of switching node slots supported by the current orthogonal server rack based on the number of internal switching nodes installed corresponding to all the computing nodes and the width of the current internal switching node's switching node slot.

[0064] The second preset condition indicates that the bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth, and the sum of the allocated downlink bandwidth of all the internal switching nodes is equal to the total external bandwidth of the computing node.

[0065] See Figure 3 With a narrower switch node slot width, a maximum of 6 internal switch nodes can be installed; see [link / reference] Figure 4 With a relatively wide switching node slot, up to four internal switching nodes can be installed.

[0066] Specifically, the second preset condition includes the following conditions: First condition: The bandwidth of the internal switching node includes allocated uplink bandwidth and allocated downlink bandwidth, and the sum of allocated uplink bandwidth and allocated downlink bandwidth is equal to the bandwidth of the internal switching node;

[0067] The second condition is that the total allocated downlink bandwidth of the internal switching nodes is equal to the total external bandwidth of the computing nodes to satisfy the full interconnection between the computing nodes and the internal switching nodes.

[0068] Provided that the first and second conditions are met, the number of internal switching nodes to be installed corresponding to all computing nodes is determined based on the total external bandwidth, the width of the switching node slot, the total number of switching node slots under the width of the switching node slot, and the bandwidth of each internal switching node.

[0069] The width of a switch node slot and the total number of switch node slots within that slot are considered as a third condition: the total width of all internal switch nodes does not exceed the width of the rack. In other words, the relationship between the width of a switch node slot and the total number of switch node slots within that slot is limited by the width of the rack. For example, Figure 3 A maximum of 6 internal switching nodes can be installed within the specified width;

[0070] Under the premise of satisfying the first, second, and third conditions, the number of internal switching nodes corresponding to all computing nodes is determined based on the total external bandwidth, the width of the switching node slot, the total number of switching node slots under that width, and the bandwidth of each internal switching node. In other words, the final number of internal switching nodes must satisfy both the requirements of hardware structure convenience and data processing requirements.

[0071] The first controller determines the maximum number of switching node slots supported by the current orthogonal server rack based on the number of internal switching nodes installed corresponding to all the computing nodes and the width of the current internal switching node's switching node slot. Specifically, this is implemented as follows:

[0072] See Figure 3 If the number of internal switching nodes installed is determined to be 4, and the switching node slots are relatively narrow, then the maximum number of switching node slots supported by the current orthogonal server rack is determined to be 6.

[0073] If the number of internal switching nodes installed is determined to be 6, and the switching node slots are relatively narrow, then the maximum number of switching node slots supported by the current orthogonal server rack is determined to be 6.

[0074] See Figure 4 If the number of internal switching nodes installed is determined to be 4, and the switching node slots are relatively wide, then the maximum number of switching node slots supported by the current orthogonal server rack is determined to be 4.

[0075] In other words, given a fixed width of the switching node slots, the maximum number of switching node slots supported by the current orthogonal server rack is also fixed. The only difference is the bandwidth matching between the ultimately installed switching nodes and compute nodes, meaning there can be empty slots.

[0076] In a more specific implementation, when the second preset condition is met, the first controller determines the number of internal switching nodes to be installed corresponding to all the computing nodes based on the total external bandwidth, the width of the switching node slots, the total number of switching node slots under each width of the switching node slot, and the bandwidth of each internal switching node, including:

[0077] Under the condition that the second preset condition and the third preset condition are met, the first controller determines the number of internal switching nodes that are compatible with all computing nodes based on the total external bandwidth, the width of the switching node slot, the total number of switching node slots under each width of the switching node slot, and the allocated uplink bandwidth and allocated downlink bandwidth of each internal switching node.

[0078] The third preset condition indicates that the allocated uplink bandwidth and the allocated downlink bandwidth of each of the internal switching nodes are equal.

[0079] The allocated uplink bandwidth and allocated downlink bandwidth of each internal switching node are set to be equal, and the sum of the allocated downlink bandwidth of all internal switching nodes is equal to the total external bandwidth of the computing node.

[0080] By setting the allocation of uplink bandwidth to be equal to the allocation of downlink bandwidth, the system achieves a balance in bidirectional communication, avoiding the upstream and downstream bandwidth imbalance problem that often exists in traditional networks, and ensuring the symmetry and efficiency of data transmission.

[0081] Set the total allocated downlink bandwidth of all internal switching nodes to equal the total outbound bandwidth of the compute nodes to satisfy full interconnection;

[0082] See Figure 2 Each compute node includes four GPUs, each GPU has a bandwidth of 1.6T, and there are a total of 8 compute nodes. The total external bandwidth of all compute nodes is: 1.6 * 4 * 8 = 51.2T.

[0083] See Figure 3 With a narrower switch node slot width, a maximum of 6 internal switch nodes can be installed; combined with Figure 2 and Figure 3 For internal switching nodes with a bandwidth of 25.6T, the bandwidth of the internal switching node includes the allocated uplink bandwidth and the allocated downlink bandwidth. Figure 2 The internal switching nodes are allocated 12.8T for both uplink and downlink bandwidth. 51.2 / 12.8=4, which means that if the total bandwidth of the allocated downlink bandwidth of the internal switching nodes is equal to the total external bandwidth of the computing nodes, then 4 internal switching nodes are required.

[0084] If the compute nodes are replaced with 2.4T medium-bandwidth switches, and the allocation method remains unchanged, each compute node will have an external bandwidth of 2.4 * 4 = 9.6T. The total external bandwidth of the 8 compute nodes will be 9.6 * 8 = 76.8T. 76.8 / 12.8 = 6, so 6 switches are needed (equivalent to internal switching nodes). Compared to the case of installing 4 switches, the two additional slots can be used at this time.

[0085] When installing a high-bandwidth 3.2T GPU, a corresponding 51.2T switch is needed due to the increased bandwidth. Eight compute nodes and four switches form a fully interconnected CLOS architecture. Each compute node has an external bandwidth of 3.2 * 4 = 12.8T, and the total external bandwidth for all eight compute nodes is 12.8 * 8 = 102.4T. Each switch has a bandwidth of 51.2T, with 25.6T allocated for both uplink and downlink, the uplink being used for scale-out expansion. Therefore, both uplink and downlink bandwidths are 25.6T. 102.4 / 25.6T = 4 switches. Therefore, four switches are needed. See also... Figure 4 With a relatively wide switching node slot, up to four internal switching nodes can be installed.

[0086] When supporting ultra-high bandwidth GPUs (4.0T), calculations show that 8 compute nodes would require 6 51.2T switches to meet the downlink bandwidth requirements.

[0087] Of course, the aforementioned 32 optical ports corresponding to 4 switches, 16 optical ports corresponding to 6 switches, and so on, are only for explaining the implementation principle of the solution in this application and do not limit the scope of protection of the solution in this application.

[0088] It is evident that the solution presented in this application can adapt GPUs with different bandwidths to different switch installation methods, and its external expansion capabilities vary depending on the chassis size. Compared to related technologies where a single orthogonal rack can only accommodate a specific type of GPU and its corresponding specific switch, and cannot flexibly switch between GPUs with different bandwidths and their corresponding switches, the flexible configuration of the solution presented in this application offers a significant technical advantage.

[0089] In some implementations, the method further includes: a first controller receiving a horizontal scaling capability signal from a second controller, and adjusting the data forwarding logic between the compute node and the internal switching node according to the horizontal scaling capability signal to meet business requirements. The data forwarding logic includes the forwarding logic for intermediate result data calculated by the compute node, wherein the horizontal scaling capability signal characterizes the horizontal scaling capability of the switching node.

[0090] Since the horizontal scaling capability signal indicates the maximum uplink bandwidth that the internal switching nodes can support, in the specific implementation of services, the first controller adjusts the data forwarding logic based on the maximum uplink bandwidth supported by the internal switching nodes and the allocated downlink bandwidth of the internal switching nodes to meet service requirements. Specifically, if the maximum uplink bandwidth supported by the internal switching nodes is larger, more data forwarding tasks can be allocated to the external switching nodes; conversely, if the maximum uplink bandwidth supported by the internal switching nodes is smaller, fewer data forwarding tasks can be allocated to the external switching nodes. That is, by adjusting the number of data forwarding tasks executed by the internal and external switching nodes, the most reasonable utilization of resources is achieved, which not only meets service requirements but also improves computing speed.

[0091] In massively parallel computing tasks, efficient transmission of intermediate result data becomes crucial, directly impacting the execution speed and overall efficiency of the computation. By adjusting the data forwarding logic, the transmission path of intermediate result data between computing nodes and internal switching nodes can be optimized, reducing data transmission latency and increasing data throughput.

[0092] See Figure 6Each compute node has multiple first connectors, and each switching node has multiple second connectors. The first and second connectors are plugged into each other (see [link]). Figure 7 The orthogonal connectors in the diagram have the same pin definitions for both the first and second connectors (see [link]). Figure 8 (Pin definitions in [reference]), the first connector includes a compute node identification pin (see [reference]). Figure 8 The pins numbered 1, 2, and 3 in the circuit are used to identify the mounting slot of the computing node. The slot information of the computing node output by different second connectors of each switching node is different. If no switching node is installed in the switching node slot, the computing node identification pin of the first connector corresponding to the switching node slot will have no information, and the internal state of the computing node will be characterized as high impedance.

[0093] The method also includes:

[0094] First reading step: The computing node reads the information of the computing node identification pin from the second connector from multiple first connectors in a first arrangement order until the information of the first non-high impedance computing node identification pin is identified. The first arrangement order is the arrangement order of multiple first connectors on a computing node.

[0095] First identification step: Determine the mounting slot information of the corresponding computing node based on the information of the first computing node identifier pin that is not in a high-impedance state;

[0096] The first binding step is to bind the installation slot information of the compute node with the expert information, so as to send the installation slot information of the compute node to the expert system for operation and maintenance management of the compute node.

[0097] Combination Figure 9 Understanding is that the internal switching node transmits the compute node's slot information (equivalent to the compute node identification pin information in this document) to the compute node via the connector. The compute node reads the compute node identification pin information from the second connector in a left-to-right order until it identifies the first non-high-resistance compute node identification pin, thus determining the compute node's mounting slot information. The meaning of high-resistance state is explained in [see...]. Figure 9 If the first and second first connectors from left to right are not connected to the switch, the information of the corresponding compute node identification pins cannot be read. At this time, the compute node will exhibit a high impedance state inside the compute node. When the third first connector is read, the information of the compute node identification pins can be read as 000, indicating that the compute node's mounting slot is the first mounting slot.

[0098] Figure 9The eight types of information represented by 000, 001, 010, 011, 100, 101, 110, and 111 correspond to the installation slots of the eight computing nodes.

[0099] After a compute node identifies its own installation slot information, it binds this information with expert information and sends the installation slot information to the expert system for operation and maintenance management. The technical advantages of this approach are as follows: The solution of identifying and binding the compute node's installation slot information to expert information, and then sending this information to the expert system for operation and maintenance management, offers the following significant technical advantages:

[0100] The compute nodes can automatically identify and report their installation slot information in the rack, which provides the expert system with accurate equipment location data, helping to quickly locate problem nodes and speed up troubleshooting and repair.

[0101] The binding of slot information with expert information means that the system can automatically call the most suitable expert resources for operation and maintenance based on the actual location of the equipment, which improves operation and maintenance efficiency and reduces the waste of response time caused by information delays or incorrect positioning.

[0102] The slot information and expert information of the computing nodes are fed back to the expert system, supporting remote monitoring and diagnosis. Operation and maintenance personnel can obtain equipment status and environmental information without being on-site, which greatly improves the level of automation of operation and maintenance management.

[0103] Based on accurate slot information, the expert system can intelligently schedule maintenance resources, optimize maintenance paths, avoid blind searching or improper scheduling, reduce maintenance costs, and improve resource utilization.

[0104] Through expert systems, the operation and maintenance process and results can be visualized, making it easy for operation and maintenance personnel and managers to view equipment status and historical records at any time, thus enhancing the transparency and traceability of operations.

[0105] By combining slot information with the data analysis capabilities of expert systems, the failure history of each computing node can be tracked, failure trends can be analyzed, a basis for preventive maintenance can be provided, and the possibility of future failures can be reduced.

[0106] When a computing node fails, the expert system can quickly locate the fault and call upon the appropriate expert to handle the situation, thus shortening business interruption time and ensuring business continuity and service quality.

[0107] Based on the real-time status and location information of computing nodes, the expert system can intelligently warn of potential fault risks and proactively perform maintenance in advance, thus avoiding business impact caused by sudden failures.

[0108] This technology, which binds compute nodes' installation slot information with expert information and sends it to the expert system, not only enables real-time monitoring of equipment status but also optimizes operation and maintenance management processes, improving fault response speed and resource scheduling efficiency. Through automated and intelligent operation and maintenance methods, the system can ensure business continuity, improve service quality, reduce operation and maintenance costs, and enhance the stability and competitiveness of data centers or computing platforms. This technology is a significant manifestation of intelligent operation and maintenance management in modern data centers and is of great importance for improving the reliability and efficiency of the overall computing and network environment.

[0109] In this embodiment of the application, the switching node includes an accelerator, and the method further includes:

[0110] Get the bandwidth of the accelerator in the compute node;

[0111] The type of the first connector is determined based on the bandwidth of the accelerator. The type of the first connector represents the total number of pins of the first connector. Specifically, the larger the bandwidth of the accelerator, the more pins the connector has, and vice versa.

[0112] A prompt message is generated based on the type of the first connector, and the prompt message is sent to the connector assembly equipment for the installation of the first connector.

[0113] By acquiring the accelerator bandwidth and automatically determining the type of the first connector accordingly, the system can achieve customized and automated hardware configuration. This means that whether it's a low-bandwidth GPU or a high-bandwidth GPU, the most suitable connector type can be obtained, ensuring data transmission efficiency and signal integrity.

[0114] Accelerator bandwidth is positively correlated with the total number of connector pins; that is, the greater the bandwidth, the more pins are required, and vice versa. This matching mechanism ensures that the connector can carry the required high-speed data transmission, avoiding performance bottlenecks caused by insufficient pins or cost waste caused by over-configuration.

[0115] Based on the type of the first connector, a prompt message is generated and sent to the connector assembly equipment, ensuring the accuracy of information during the assembly process and avoiding human error. Simultaneously, this information also promotes compatibility between GPUs and switch configurations with different bandwidths, enabling the same chassis to accommodate a wider variety of hardware combinations.

[0116] Sending prompts to the connector assembly equipment provides precise guidance for the production process. The assembly equipment can automatically select and install the correct first connector based on the received information, significantly improving assembly efficiency and accuracy.

[0117] By ensuring a precise match between the connector type and the accelerator bandwidth, the system can achieve optimal network performance, including key metrics such as data transmission rate, latency, and throughput, providing users with a smoother computing experience.

[0118] In a more specific embodiment, determining the type of the first connector based on the bandwidth of the accelerator includes:

[0119] If the bandwidth of the accelerator is within the first bandwidth range, then the first connector is determined to be a first type of connector;

[0120] If the accelerator's bandwidth is within the second bandwidth range, then the first connector is determined to be a second type connector. The minimum value of the second bandwidth range is greater than the maximum value of the first bandwidth range. The second type connector is formed by splicing together multiple first type connectors, and the positions of the first type connectors that are adapted to the first bandwidth range are the same among the different second type connectors.

[0121] See Figure 8 The first bandwidth range uses a 32DP connector, and the second bandwidth range uses a 64DP connector. The 64DP connector is made up of two 32DP connectors joined together, specifically, the two connectors are fixed together with a vertical separator. However, the orthogonal connector for the smaller bandwidth needs to be located to the left of the two connectors, in a fixed position, to facilitate the reading of the identification data.

[0122] This technical solution dynamically adjusts the connector type to adapt to accelerators with different bandwidths, with particular emphasis on the selection and design of the first and second type of connectors within the first and second bandwidth ranges. Its technical advantages are mainly reflected in the following aspects:

[0123] By distinguishing between the first bandwidth range and the second bandwidth range, the system can accurately select the connector type that matches the accelerator bandwidth, avoiding performance bottlenecks or resource waste caused by improper connector selection.

[0124] The second type of connector (such as 64DP) is formed by splicing multiple first type connectors (such as 32DP). It can not only meet the needs of high bandwidth accelerators, but also ensure compatibility and scalability between different bandwidth accelerators, and reserve space for future technology upgrades and bandwidth requirements.

[0125] Regardless of whether the accelerator is in the first bandwidth range or the second bandwidth range, the first type of connector adapted to the first bandwidth range remains in the same position as the second type of connector. This allows compute nodes and internal switching nodes to be interchanged under different bandwidth configurations, simplifying the standardized design and maintenance process of the equipment.

[0126] Two 32DP connectors are fixed together in a vertically separated structure to form a 64DP connector, which not only simplifies the connector installation process, but also ensures precise alignment and signal integrity between the connectors, and reduces the complexity of maintenance and debugging.

[0127] The small-bandwidth orthogonal connector is fixed on the left side, which facilitates the reading of identification data, simplifies the system identification and configuration process, and improves the efficiency of equipment installation and troubleshooting.

[0128] By using the first type of connector as the basic unit to build the second type of connector, the development cost of customized hardware can be reduced, while ensuring the quality of signal transmission under high bandwidth requirements.

[0129] Intelligent selection of connector type based on bandwidth requirements avoids over-configuration or under-configuration, achieves optimal resource utilization, and reduces the overall system operating cost.

[0130] In this embodiment, each switching node further includes a switching processor connected to the uplink port. If the computational load of the current computing task exceeds a preset computational load, the first controller sends a horizontal scaling function signal with a first value to the second controller, causing the second controller to control the uplink port to open, including:

[0131] If the computational load of the current computational task exceeds the preset computational load, the first controller sends a horizontal expansion function signal with a first value to the second controller, so that the second controller forwards the horizontal expansion function signal with the first value to the switching processor to enable it through the uplink port of the switching processor.

[0132] See Figure 6 The switching processor represents the core of the switching node. It sends a horizontal scaling signal with a first value to the second controller via the first controller, which then forwards it to the switching processor. This enables intelligent dynamic activation of the uplink ports on the switching node. This technical solution significantly improves the network flexibility, resource utilization efficiency, and overall performance of the computing platform, and is of great value for building highly resilient cloud computing and data center environments. It not only solves the problem of network resource allocation when computing load fluctuates, but also brings convenience to operation and maintenance management, saves costs, and provides users with consistent and high-quality computing services.

[0133] In this embodiment of the application, the method further includes: if the computational amount of the current computing task is less than or equal to the preset computational amount, the first controller sends a horizontal expansion function signal with a second value to the second controller, so that the second controller controls the uplink port to close.

[0134] By dynamically monitoring the actual computational load of computing tasks, unnecessary uplink ports are automatically shut down when the computational load falls below a preset threshold, reducing energy consumption. This technology can significantly save electricity costs, especially in large-scale data center or cloud platform environments, which aligns with the development trend of green computing.

[0135] Dynamically closing uplink ports avoids unnecessary network resource consumption, reduces the wear and tear on hardware devices, decreases long-term maintenance and replacement costs, and improves overall economic efficiency.

[0136] The automatic shutdown mechanism prevents the abuse of network resources under low load conditions, helps maintain stable system operation, and avoids system instability or security vulnerabilities caused by excessive resource consumption.

[0137] When computational load is low, closing the uplink port can free up network bandwidth, making room for other higher priority or more computationally demanding tasks, thereby improving the overall network throughput and the flexibility of task scheduling.

[0138] Dynamically closing uplink ports allows the system to react quickly to network topology changes, ensuring that network resources are always matched to current computing needs and avoiding network congestion and packet loss.

[0139] In this embodiment of the application, the method further includes:

[0140] Obtain the type of the current computing task and the computing objective to be achieved;

[0141] Determine the computational load of the current computational task based on the type of the current computational task and the computational goal to be achieved.

[0142] Examples of computational goals to be achieved are as follows:

[0143] The goal is to train a deep learning model to achieve an accuracy of over 98% on a specific dataset, while keeping the training time under 12 hours. In machine learning, computational objectives may include model accuracy, training time, and the number of iterations required to reach convergence. For example, training an image recognition model may require iterative optimization on a large amount of image data until the model can accurately classify new data, and the training process must be completed within a reasonable timeframe.

[0144] The task is to process a 100GB dataset, perform data cleaning, feature extraction, and data analysis, and output statistical reports and visualizations. The entire data processing workflow must be completed within 4 hours, with CPU and GPU utilization maintained above 70%. In data analysis scenarios, the computational goal can be to process datasets of a specific size, achieve a certain processing speed, and output analytical results in a specific format. For example, in financial risk analysis, it may be necessary to process a large amount of transaction data in a short period of time to identify potential risk patterns and generate reports for decision-makers.

[0145] Providing real-time rendering for virtual reality applications or online games ensures a frame rate of at least 60fps and network latency within 50 milliseconds. In the virtual reality or gaming field, computational goals typically involve the smoothness of real-time rendering (frame rate), network latency, and interactive response time. For example, a virtual reality game might require high-quality real-time rendering of images in complex 3D environments while maintaining low latency to provide an immersive gaming experience.

[0146] By analyzing the type of computing task and the computing goals to be achieved, the system can more accurately predict the computing resources required to execute the task, including CPU, GPU, memory, and network bandwidth, thereby providing a basis for the rational allocation of resources.

[0147] Based on the nature and objectives of the task, computational load prediction can more accurately estimate the task's execution time, data processing volume, and communication requirements, avoiding over- or under-allocation of resources.

[0148] During task execution, changes in task type and computational objectives are monitored in real time, and the allocation of computing resources is dynamically adjusted to ensure that the task can be completed in the shortest time with the optimal configuration.

[0149] Through real-time analysis, the system can quickly respond to changes in computing demands, make timely resource scheduling decisions, and effectively improve the response speed and processing efficiency of the computing platform.

[0150] Accurate computational load prediction and dynamic resource scheduling can maximize resource utilization efficiency, reduce resource waste, and save operating costs.

[0151] By reasonably predicting the computational load and making advance resource allocation and network bandwidth preparations, performance bottlenecks that may occur during the computation process can be effectively avoided, ensuring the smooth execution of computational tasks.

[0152] In this embodiment of the application, the method further includes: determining a preset computing volume based on the total external bandwidth of all computing nodes, the utilization rate of each computing node when executing computing tasks, and the storage space requirements required during task execution, wherein the storage space requirements include initial data loading, intermediate result storage, and final output result storage.

[0153] This approach not only focuses on computing resources (such as CPU and GPU utilization) but also covers the full range of network bandwidth and storage space requirements, providing a more comprehensive assessment of resource requirements and helping to plan and allocate resources more accurately.

[0154] It particularly emphasizes the three stages of storage space requirements—initial data loading, intermediate result storage, and final output result storage—which helps to analyze the entire data processing process in detail and ensure that there are sufficient storage resources to support each stage.

[0155] The preset computational load is determined based on real-time monitoring and analysis, and resource allocation can be dynamically adjusted according to the actual needs of the computational task, thereby improving the flexibility and efficiency of resource management.

[0156] By performing detailed calculations in advance, contention for computing, network, and storage resources can be avoided in advance, ensuring that each computing node and task can obtain the optimal resource allocation required during execution, thereby reducing resource conflicts and waiting time.

[0157] Accurate estimation of computational load enables the system to more rationally arrange the order of task execution and resource allocation, thereby accelerating task execution and improving the overall platform's computational efficiency.

[0158] Ensure sufficient resources for data transmission, computation, and storage operations to prevent performance bottlenecks during task execution and guarantee the smooth progress of the task.

[0159] This application also provides a multi-functional configuration method suitable for orthogonal server racks, specifically applied to a second controller. The current orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. Each compute node includes a first controller, and each internal switching node includes a second controller and has an uplink port for connecting to external switching nodes. The external switching nodes are switching nodes other than the internal switching nodes of the current orthogonal server rack. See [link to previous document]. Figure 10 The methods include:

[0160] Step S1001: Receive the horizontal expansion function signal with the first value sent by the first controller. The horizontal expansion function signal with the first value is generated by the first controller when the business requirement indicates that the computational amount of the current computing task is greater than the preset computational amount.

[0161] Step S1002: The uplink port is opened according to the horizontal expansion function signal with the first value and the data transmission protocol of the uplink port is set to the first protocol. The external switching node and the internal switching node are used to exchange and transmit data to the computing node to realize the business function. The first protocol supports the transmission of data from the computing node to the external switching node, and the transmission protocol between the internal switching node and the computing node is the second protocol. The first protocol and the second protocol are different.

[0162] After the uplink port is opened, the first protocol is used to transmit data from the compute node to the external switch node, while the second protocol ensures data transmission between the compute node and the internal switch node, and the compute node and the internal switch node are fully interconnected. By using the first protocol for horizontal scaling and combining it with the second protocol to maintain internal communication, the compute node can fully utilize internal and external network resources, accelerating data processing and the completion of computing tasks, thereby achieving business goals. For computationally intensive tasks, the requirements can be met by connecting to the external switch node without changing the current orthogonal server rack structure. This solves the problem in related technologies where the maximum data computing capacity of orthogonal server racks is constrained by the number of compute nodes and switch nodes and bandwidth.

[0163] See Figure 6 , Figure 8 and Figure 9 Each computing node is provided with multiple first connectors, and each switching node is provided with multiple second connectors. The first connectors and second connectors are plugged into each other. The pin definitions of the first connectors and second connectors are the same. The second connector includes a switching node identification pin, which is used to identify the mounting slot of the switching node. The slot information of the switching node output by different first connectors of each computing node is different. If no computing node is installed in the computing node slot, the switching node identification pin of the second connector corresponding to the computing node slot has no information, and the internal state of the switching node is characterized as high impedance.

[0164] The method also includes:

[0165] Second reading step: The switching node reads the information of the switching node identification pin from the first connector from multiple second connectors in a second arrangement order until the information of the first switching node identification pin that is not in a high impedance state is identified. The second arrangement order is the arrangement order of multiple second connectors on a switching node.

[0166] The second identification step: Determine the installation slot information of the corresponding switching node based on the information of the first switching node identification pin that is not in a high-impedance state.

[0167] The second binding step is to bind the installation slot information of the switching node with the fault information of the switching node in order to diagnose and restore faulty switching nodes.

[0168] Figure 9 The horizontal numbers 000, 001, 010, 011, 100, and 101 correspond to six different installation slots for switches.

[0169] The pin definitions of the first and second connectors are identical, ensuring universality and compatibility between computing nodes and switching nodes, and maintaining stable connection and data exchange even in scenarios with different bandwidth requirements.

[0170] The system automatically reads and identifies the installation slot information via the exchange node identification pin on the second connector, eliminating the need for manual configuration. This greatly simplifies system deployment and maintenance, and improves management efficiency.

[0171] By binding the installation slot information of the switching node with its fault information, the specific slot where the faulty node is located can be quickly located when a fault occurs, which facilitates rapid diagnosis and repair and reduces the time required for troubleshooting.

[0172] In slots where no computing node is installed, the exchange node identification pin of the second connector is in a high-impedance state. This mechanism clearly distinguishes between installed and uninstalled states, avoids identification errors, and enhances the robustness of the system.

[0173] In the first reading step, the switching node reads the identification information in a fixed second sequence until it finds the first valid slot information. This process can dynamically update the configuration based on the real-time installation status of the compute nodes, enabling the system to self-adjust and adapt to constantly changing hardware layouts.

[0174] Once the exact slot of the faulty switching node is identified, the system can automatically trigger the corresponding fault recovery process, including but not limited to reconfiguring network paths and adjusting load balancing, to ensure service continuity and reliability.

[0175] Automatic slot identification and intelligent fault binding simplify the hardware maintenance process, reduce reliance on professional technicians, and make daily operation and maintenance simpler and more efficient.

[0176] Intelligent fault diagnosis and recovery reduces downtime caused by hardware failures, avoids productivity losses and customer dissatisfaction caused by untimely fault diagnosis, and saves maintenance and operating costs in the long run.

[0177] In this embodiment of the application, different internal switching nodes are allocated the same uplink bandwidth, and all uplink ports of each internal switching node support the same allocated uplink bandwidth. The method further includes:

[0178] Obtain the allocated uplink bandwidth and the number of uplink ports for the internal switching nodes;

[0179] When the allocated uplink bandwidth of the internal switching node is less than or equal to the first allocated uplink bandwidth, and the number of uplink ports is not less than the first number, the maximum uplink bandwidth supported by the internal switching node is equal to the allocated uplink bandwidth of the internal switching node.

[0180] For example, if the allocated uplink bandwidth of an internal switching node is less than or equal to 12.8T and the number of uplink ports is greater than 16, it can support the full expansion of 12.8T.

[0181] When the allocated uplink bandwidth of the internal switching node is less than or equal to the second allocated uplink bandwidth and greater than the first allocated uplink bandwidth, and the number of uplink ports is not less than the second number, the maximum uplink bandwidth supported by the internal switching node is equal to the allocated uplink bandwidth of the internal switching node, the second allocated uplink bandwidth is greater than the first allocated uplink bandwidth, and the second number is greater than the first number.

[0182] For example, if the allocated uplink bandwidth of an internal switching node is less than or equal to 25.6T and the number of uplink ports is greater than 32, it can support the full expansion of 25.6T.

[0183] If the allocated uplink bandwidth of the internal switching node is greater than the second allocated uplink bandwidth, and the number of uplink ports is less than the second number, the maximum uplink bandwidth supported by the internal switching node is less than the allocated uplink bandwidth of the internal switching node.

[0184] For example, if the allocated uplink bandwidth of an internal switching node is greater than 25.6T and the number of uplink ports is less than 32, it does not support expanding the entire 25.6T; for example, if the number of uplink ports is 16, it only supports external expansion to 12.8T.

[0185] By determining the relationship between the allocated uplink bandwidth of the internal switching node and the preset first and second allocated uplink bandwidth, the external expansion requirements of the computing node can be accurately matched, ensuring full utilization of bandwidth resources.

[0186] When the allocated uplink bandwidth of the internal switching node meets the conditions, it can automatically support the full expansion of the maximum bandwidth without manual intervention, thus improving the automation level of bandwidth management.

[0187] The solution takes into account the number of uplink ports. Only when the number of ports meets or exceeds the preset first and second numbers can the internal switching nodes support the expansion of the maximum uplink bandwidth. This mechanism enables the network topology to be adaptively adjusted according to the hardware configuration, ensuring the stability and scalability of the network.

[0188] By linking the number of ports to bandwidth expansion capabilities, the solution strikes a balance between physical space constraints and bandwidth requirements, avoiding the problem of bandwidth expansion being limited by insufficient port numbers.

[0189] When the allocated uplink bandwidth of an internal switching node is greater than the preset value, but the number of uplink ports is insufficient to meet the requirements, the system will automatically adjust and use only the maximum uplink bandwidth that can be actually supported, thus avoiding the waste of bandwidth resources.

[0190] This application embodiment also provides an orthogonal server rack, including:

[0191] Multiple computing nodes and multiple internal switching nodes are orthogonally connected. The internal switching nodes have uplink ports for connecting to external switching nodes. The external switching nodes are switching nodes other than the internal switching nodes. Each computing node includes a first controller and each switching node includes a second controller.

[0192] The first controller is used to execute the multi-functional configuration method applicable to orthogonal server racks in the first aspect;

[0193] The second controller is used to execute the second aspect of the multi-functional configuration method applicable to orthogonal server racks.

[0194] The orthogonal server rack features the following functions: after the uplink port is opened, it uses a first protocol to transmit data from the compute nodes to external switching nodes, and a second protocol to ensure data transmission between the compute nodes and internal switching nodes, with a fully interconnected connection between them. By employing the first protocol for horizontal scaling and combining it with the second protocol to maintain internal communication, compute nodes can fully utilize internal and external network resources, accelerating data processing and computational task completion, thereby achieving business goals. For computationally intensive tasks, the requirements can be met by connecting to external switching nodes without changing the current orthogonal server rack structure. This solves the problem in related technologies where the maximum data computing capacity of orthogonal server racks is constrained by the number of compute nodes and switching nodes, as well as bandwidth.

[0195] See Figure 6 Each compute node is equipped with multiple first connectors, and each internal switching node is equipped with multiple second connectors. The first and second connectors are plugged into each other to ensure that the compute node and any switching node are orthogonal and pluggable. The first controller is connected to the first connector, and the second controller is connected to the second connector. The use of the first and second connectors satisfies the pluggable function, allowing a single chassis to accommodate multiple numbers of switches.

[0196] The first connector and the second connector have the same pin definitions;

[0197] Specifically, the first connector and the second connector are high-density connectors. High-density connectors play a crucial role in electronic devices and systems, especially in high-performance computing, servers, data centers, and other applications that require high bandwidth and high signal integrity.

[0198] The plurality of identification pins of each of the first connectors are located in the same column or multiple adjacent columns on the same side of all the signal transmission pins, and the plurality of identification pins of each of the second connectors are located in the same row or multiple adjacent rows on the same side of all the signal transmission pins. See also Figure 8 The connector has pins divided into high-speed differential signal pins and identification pins. The identification pins are located on one side of the high-speed differential signal pins, which facilitates wiring.

[0199] Each of the aforementioned switching nodes includes an accelerator (GPU), which is connected to the first connector. See also Figure 6 The first controller is connected to the identification pin in the connector, and the accelerator is connected to the high-speed differential signal pin in the connector. This enables the transmission of high-speed differential signals. The second controller is connected to the identification pin in the connector, and the switching processor is connected to the high-speed differential signal pin in the connector.

[0200] Figure 8 PIN4, representing slot support capability, is set on the compute node based on whether the chassis supports a total of 4 or 6 switch slots. It is strongly tied to the GPU bandwidth of the compute node. For example, small / medium / ultra-high bandwidth GPUs correspond to 6 switch slots, while high bandwidth GPUs correspond to 4 switch slots. Small bandwidth GPUs are 1.6T GPUs, medium bandwidth GPUs are 2.4T GPUs, ultra-high bandwidth GPUs are 4.0T GPUs, and high bandwidth GPUs are 3.2T GPUs. Based on the actual GPU bandwidth, this PIN information is transmitted from the compute node to the switch. 0 represents 4 switch slots, and 1 represents 6 switch slots.

[0201] Figure 8 PIN8 is used to determine the scale-out function. A value of 1 indicates that the uplink port is open, and a value of 0 indicates that the uplink port is closed. It is shared by compute nodes and switching nodes.

[0202] Each of the first connectors includes a slot support capability identification pin and an expansion identification pin. The slot support capability identification pin is used to transmit a slot support capability signal, and the expansion identification pin is used to transmit a lateral expansion function signal. The slot support capability signal represents the maximum number of switching node slots supported by the current orthogonal server rack.

[0203] The first connector and the second connector include differential signal transmission pins for transmitting differential data.

[0204] The specific functions of the slot support capability signal transmitted by the slot support capability identifier pin and the lateral extension function signal transmitted by the extension identifier pin are described in the embodiments above.

[0205] In addition, the first connector includes a compute node identification pin, which is used to identify the mounting slot of the compute node. The slot information of the compute node output by the different second connectors of each switch node is different. If no switch node is installed in the switch node slot, the compute node identification pin of the first connector corresponding to the switch node slot has no information, and the internal characteristics of the compute node are high impedance.

[0206] The second connector includes a switching node identification pin, which is used to identify the mounting slot of the switching node. The slot information of the switching node output by the first connector of each computing node is different. If a computing node is not installed in the computing node slot, the switching node identification pin of the second connector corresponding to the computing node slot has no information, and the internal state of the switching node is characterized as high impedance.

[0207] The configuration of the compute node identifier pin and the exchange node identifier pin in this embodiment is to facilitate the identification of the slot of the node and the slot of the exchange node. For the specific identification principle, please refer to the embodiment mentioned above.

[0208] High-density connectors can carry a large amount of data signals, providing high-speed data transmission capabilities. This means that within a smaller space, the connector can support more data channels, which is crucial for systems requiring high-bandwidth communication (such as massively parallel computing, high-speed network switching, etc.), enabling rapid data flow within the system.

[0209] Due to their high-density design, these connectors can achieve more connection points in a limited space, helping to reduce equipment size, server rack footprint, and improve data center space utilization. High-density connectors are an ideal choice for applications that demand compact design and high-density integration.

[0210] High density does not mean sacrificing signal integrity and electromagnetic compatibility. High-density connectors employ advanced signal isolation and shielding technologies to minimize crosstalk and electromagnetic interference between signals, ensuring excellent signal quality and stability even during high-bandwidth transmission. This is crucial for maintaining the reliability and accuracy of high-speed data transmission.

[0211] To cope with the frequent plugging and unplugging operations in servers and data centers, high-density connectors feature a robust and durable design with excellent contact stability and mechanical strength, capable of withstanding multiple plugging and unplugging operations without affecting performance. This ensures connection reliability and system stability in long-term operation and high-frequency use environments.

[0212] High-density connectors can not only transmit data signals, but also support power, control signals and multiple communication protocols, such as PCI Express and Ethernet. This achieves multi-functional integration, simplifies the complexity of cabling and device connections, and improves the integration of devices and the overall performance of the system.

[0213] See Figure 6 The switching node also includes a switching processor, which is connected to the second connector and the second controller. The switching processor is used to control the opening or closing of the uplink port. The switching processor represents the core of the switching node. It sends a horizontal expansion function signal with a first value to the second controller through the first controller, and then forwards it to the switching processor to realize the intelligent dynamic opening of the uplink port of the switching node.

[0214] Additionally, if the bandwidth of the accelerator in the computing node is within a first bandwidth range, then the first connector is a first type connector; if the bandwidth of the accelerator in the computing node is within a second bandwidth range, then the first connector is a second type connector, wherein the minimum value of the second bandwidth range is greater than the maximum value of the first bandwidth range.

[0215] The second type of connector is formed by splicing together multiple first type connectors. A separation structure is provided between any two adjacent first type connectors, and the positions of the first type connectors that are adapted to the first bandwidth range in different second type connectors are the same.

[0216] By subdividing the accelerator's bandwidth range into a first bandwidth range and a second bandwidth range, the connector type can be precisely selected based on the accelerator's actual bandwidth requirements. The first type of connector is suitable for lower bandwidth levels, while the second type of connector, composed of multiple first-type connectors, is suitable for higher bandwidth demands. This design ensures a precise match between the connector and the accelerator, optimizing signal transmission efficiency and quality.

[0217] The second type of connector design allows the chassis to support higher bandwidth accelerators, such as high-bandwidth and ultra-high-bandwidth GPUs, without altering the basic structure. At the same time, this separation structure ensures good compatibility and a unified interface standard even when different types of connectors are used interchangeably, facilitating future hardware upgrades and expansions.

[0218] Regardless of the connector type used, the position of the first-type connector, which is adapted to the first bandwidth range, is fixed within the second-type connector, greatly simplifying the system configuration process. There's no longer a need to remember complex connector layouts and pin definitions; simply focusing on the accelerator's bandwidth level and its correspondence with the connector type allows for rapid hardware configuration and signal path planning.

[0219] A separator structure is provided between two adjacent Type 1 connectors. This design facilitates rapid location and isolation of faults. When a Type 1 connector fails, the separator structure allows for quick identification of the fault location. Only the faulty connector needs to be replaced or repaired, without affecting other parts of the entire Type 2 connector, thus reducing maintenance difficulty and shortening recovery time.

[0220] By splicing the first type of connectors to form the second type of connector on demand, the connectivity requirements of high-bandwidth accelerators can be met while avoiding unnecessary hardware redundancy. This low-cost, scalable solution allows data centers to dynamically adjust resource allocation based on actual load, avoiding over-investment for short-term peak demand and thus maximizing cost-effectiveness.

[0221] The isolation structure design helps reduce crosstalk and electromagnetic interference (EMI) between signals, which is especially important for high-bandwidth connections. By physically isolating Type I connectors, signal integrity and system performance stability can be effectively maintained even in high-density cabling environments, ensuring the quality and speed of data transmission.

[0222] Furthermore, the switch node slots for installing the switching nodes are movable and removable to accommodate switch nodes with different numbers of uplink ports. The switch nodes with different numbers of uplink ports correspond to switch node slots of different widths, and the number of internal switch nodes installed is related to the total external bandwidth of all the computing nodes. The movable and removable slots, together with the first and second connectors, enable a single chassis to accommodate multiple numbers and widths of switches, achieving flexible configuration.

[0223] Specifically, the orthogonal server rack includes sliding rails mounted on the inner wall of the rack. The switching node slots are mounted on and movable along these sliding rails. The movable nature of the switching node slots is achieved through the rails installed inside the rack, thus allowing different numbers of internal switching nodes to be installed in the same rack.

[0224] This application also provides a server system, including: an orthogonal server rack; and an external switching node, wherein the internal switching node in the orthogonal server rack is connected to the external switching node through an uplink port.

[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0226] Embodiments of this application also provide a multi-functional configuration device suitable for orthogonal server racks, used to execute the multi-functional configuration method suitable for orthogonal server racks according to embodiments of this application.

[0227] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the multi-functional configuration method applicable to orthogonal server racks.

[0228] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the multi-functional configuration method applicable to orthogonal server racks when run.

[0229] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0230] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the multi-functional configuration method applicable to orthogonal server racks.

[0231] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the multi-functional configuration method applicable to orthogonal server racks.

[0232] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0233] The foregoing has provided a detailed description of a multi-functional configuration method for orthogonal server racks, as well as the orthogonal server rack and server system. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A multi-functional configuration method suitable for orthogonal server racks, characterized in that, The current orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. Each compute node includes a first controller, and each internal switching node includes a second controller and has an uplink port for connecting to an external switching node, which is a switching node other than the internal switching nodes. The method includes: The first controller acquires the business requirements, which include the computational load of the current computing task; If the computational load of the current computing task is greater than the preset computational load, the first controller sends a horizontal scaling function signal with a first value to the second controller, so that the second controller controls the uplink port to open and sets the data transmission protocol of the uplink port to a first protocol, thereby using the external switching node and the internal switching node to exchange and transmit data of the computing node to realize the business function. The first protocol supports the transmission of data of the computing node to the external switching node, and the transmission protocol between the internal switching node and the computing node is a second protocol. The first protocol is different from the second protocol.

2. The multi-functional configuration method according to claim 1, characterized in that, The method further includes: The first controller sends a slot support capability signal to the second controller, so that the second controller, under the condition of meeting a first preset condition, determines the horizontal scaling capability of the switching node based on the slot support capability signal, the total number of uplink ports that each internal switching node can provide, and the bandwidth of each internal switching node. The horizontal scaling capability represents the maximum bandwidth that each switching node supports for external expansion. Wherein, the slot support capability signal represents the maximum number of switching node slots supported by the current orthogonal server rack, the first preset condition represents that the bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth, the sum of the allocated downlink bandwidth of all internal switching nodes is equal to the total external bandwidth of the computing node, the ports of the internal switching nodes include the uplink port and the downlink port, and the downlink port supports the implementation of the allocated downlink bandwidth.

3. The multifunctional configuration method according to claim 2, characterized in that, Each computing node is provided with multiple first connectors, and each switching node is provided with multiple second connectors. The first connectors and second connectors are plugged into each other. The pin definitions of the first connectors and second connectors are the same. Each first connector includes a slot support capability identification pin and an expansion identification pin. The slot support capability identification pin is used to transmit the slot support capability signal, and the expansion identification pin is used to transmit the lateral expansion function signal. If the computational load of the current computing task is greater than the preset computational load, the first controller sends a horizontal expansion function signal with a first value to the second controller, including: if the computational load of the current computing task is greater than the preset computational load, the first controller sends the horizontal expansion function signal with the first value to the second controller through the expansion identifier pin. The first controller sends the slot support capability signal to the second controller, including: the first controller sends the slot support capability signal to the second controller through the slot support capability identification pin.

4. The multi-functional configuration method according to claim 1, characterized in that, The method further includes: The first controller obtains the total external bandwidth of all the computing nodes; When the second preset condition is met, the first controller determines the number of internal switching nodes to be installed corresponding to all the computing nodes based on the total external bandwidth, the width of the switching node slot, the total number of switching node slots under each width of the switching node slot, and the bandwidth of each internal switching node. The first controller determines the maximum number of switching node slots supported by the current orthogonal server rack based on the number of internal switching nodes installed corresponding to all the computing nodes and the width of the current internal switching node's switching node slot. The second preset condition indicates that the bandwidth of each internal switching node is equal to the sum of the allocated uplink bandwidth and the allocated downlink bandwidth, and the sum of the allocated downlink bandwidth of all the internal switching nodes is equal to the total external bandwidth of the computing node.

5. The multi-functional configuration method according to claim 4, characterized in that, Under the condition of satisfying the second preset condition, the first controller determines the number of internal switching nodes to be installed corresponding to all the computing nodes based on the total external bandwidth, the width of the switching node slots, the total number of switching node slots under each width of the switching node slots, and the bandwidth of each internal switching node, including: Under the condition that the second preset condition and the third preset condition are met, the first controller determines the number of internal switching nodes that are adapted to all the computing nodes based on the total external bandwidth, the width of the switching node slot, the total number of switching node slots under each width of the switching node slot, and the allocated uplink bandwidth and the allocated downlink bandwidth of each internal switching node. The third preset condition indicates that the allocated uplink bandwidth and the allocated downlink bandwidth of each of the internal switching nodes are equal.

6. The multi-functional configuration method according to claim 2, characterized in that, The method further includes: The first controller receives a horizontal scaling capability signal from the second controller and adjusts the data forwarding logic between the compute node and the internal switching node according to the horizontal scaling capability signal to meet service requirements. The data forwarding logic includes the forwarding logic for intermediate result data calculated by the compute node. The horizontal scaling capability signal characterizes the horizontal scaling capability of the switching node.

7. The multi-functional configuration method according to claim 1, characterized in that, Each computing node is provided with multiple first connectors, and each switching node is provided with multiple second connectors. The first connectors and second connectors are plugged into each other. The pin definitions of the first connectors and second connectors are the same. The first connector includes a computing node identification pin, which is used to identify the mounting slot of the computing node. The slot information of the computing node output by different second connectors of each switching node is different. If no switching node is installed in the switching node slot, the computing node identification pin of the first connector corresponding to the switching node slot has no information, and the internal state of the computing node is characterized as high impedance. The method further includes: First reading step: The computing node reads information from the computing node identification pins of the second connectors from the plurality of first connectors in a first arrangement order until the information of the first non-high impedance computing node identification pin is identified. The first arrangement order is the arrangement order of the plurality of first connectors on a computing node. First identification step: Determine the mounting slot information of the corresponding computing node based on the information of the first identification pin of the computing node that is not in a high-impedance state; The first binding step is to bind the installation slot information of the computing node with the expert information, so as to send the installation slot information of the computing node to the expert system for operation and maintenance management of the computing node.

8. The multi-functional configuration method according to claim 7, characterized in that, The switching node includes an accelerator, and the method further includes: Obtain the bandwidth of the accelerator in the computing node; The type of the first connector is determined based on the bandwidth of the accelerator, and the type of the first connector represents the total number of pins of the first connector; A prompt message is generated based on the type of the first connector, and the prompt message is sent to the connector assembly equipment to install the first connector.

9. The multi-functional configuration method according to claim 8, characterized in that, Determining the type of the first connector based on the bandwidth of the accelerator includes: If the bandwidth of the accelerator is within a first bandwidth range, then the first connector is determined to be a first type of connector; If the bandwidth of the accelerator is within the second bandwidth range, then the first connector is determined to be a second type connector, where the minimum value of the second bandwidth range is greater than the maximum value of the first bandwidth range.

10. The multi-functional configuration method according to claim 1, characterized in that, Each of the aforementioned switching nodes further includes a switching processor connected to the uplink port. If the computational load of the current computing task exceeds a preset computational load, the first controller sends a horizontal scaling function signal with a first value assigned to it to the second controller, causing the second controller to control the uplink port to open, including: If the computational load of the current computational task is greater than the preset computational load, the first controller sends the horizontal scaling function signal with the first value to the second controller, so that the second controller forwards the horizontal scaling function signal with the first value to the switching processor to enable the uplink port through the switching processor.

11. The multi-functional configuration method according to claim 1, characterized in that, The method further includes: If the computational load of the current computational task is less than or equal to the preset computational load, the first controller sends a horizontal scaling function signal with a second value to the second controller, so that the second controller controls the uplink port to close.

12. The multi-functional configuration method according to claim 1, characterized in that, The method further includes: Obtain the type of the current computing task and the computing objective to be achieved; The computational load of the current computing task is determined based on the type of the current computing task and the computing objective to be achieved.

13. The multi-functional configuration method according to claim 1, characterized in that, The method further includes: The preset computational load is determined based on the total external bandwidth of all the computing nodes, the utilization rate of each computing node when executing computing tasks, and the storage space requirements during task execution. The storage space requirements include initial data loading, intermediate result storage, and final output result storage.

14. A multi-functional configuration method suitable for orthogonal server racks, characterized in that, The current orthogonal server rack includes orthogonally connected compute nodes and internal switching nodes. Each compute node includes a first controller, and each internal switching node includes a second controller and has an uplink port for connecting to an external switching node. The external switching node is any switching node other than the internal switching node of the current orthogonal server rack. The method includes: The second controller receives a horizontal scaling function signal with a first value sent by the first controller. The horizontal scaling function signal with the first value is generated by the first controller when the business requirement indicates that the computational amount of the current computing task is greater than the preset computational amount. The uplink port is opened according to the horizontal expansion function signal with the first value assigned, and the data transmission protocol of the uplink port is set to the first protocol. The external switching node and the internal switching node are used to exchange and transmit data of the computing node to realize the business function. The first protocol supports the transmission of data of the computing node to the external switching node, and the transmission protocol between the internal switching node and the computing node is the second protocol. The first protocol is different from the second protocol.

15. The multifunctional configuration method according to claim 14, characterized in that, Each computing node is provided with multiple first connectors, and each switching node is provided with multiple second connectors. The first connectors and second connectors are plugged into each other. The pin definitions of the first connectors and second connectors are the same. The second connector includes a switching node identification pin, which is used to identify the mounting slot of the switching node. The slot information of the switching node output by different first connectors of each computing node is different. If no computing node is installed in the computing node slot, the switching node identification pin of the second connector corresponding to the computing node slot has no information, and the internal state of the switching node is characterized as a high-resistance state. The method further includes: Second reading step: The switching node reads information from the switching node identification pins of the first connector from a plurality of second connectors in a second arrangement order until the information of the first switching node identification pin that is not in a high impedance state is identified. The second arrangement order is the arrangement order of a plurality of second connectors on a switching node. The second identification step: Determine the mounting slot information of the corresponding switching node based on the information of the first identification pin of the switching node that is not in a high-impedance state; The second binding step is to bind the installation slot information of the switching node with the fault information of the switching node in order to diagnose and restore the faulty switching node.

16. The multi-functional configuration method according to claim 14, characterized in that, The different internal switching nodes have the same allocated uplink bandwidth, and all uplink ports of each internal switching node support the same allocated uplink bandwidth. The method further includes: Obtain the allocated uplink bandwidth of the internal switching node and the number of uplink ports; When the allocated uplink bandwidth of the internal switching node is less than or equal to the first allocated uplink bandwidth, and the number of uplink ports is not less than the first number, the maximum uplink bandwidth supported by the internal switching node is equal to the allocated uplink bandwidth of the internal switching node. When the allocated uplink bandwidth of the internal switching node is less than or equal to the second allocated uplink bandwidth and greater than the first allocated uplink bandwidth, and the number of uplink ports is not less than the second number, the maximum uplink bandwidth supported by the internal switching node is equal to the allocated uplink bandwidth of the internal switching node, the second allocated uplink bandwidth is greater than the first allocated uplink bandwidth, and the second number is greater than the first number. When the allocated uplink bandwidth of the internal switching node is greater than the second allocated uplink bandwidth, and the number of uplink ports is less than the second number, the maximum uplink bandwidth supported by the internal switching node is less than the allocated uplink bandwidth of the internal switching node.

17. An orthogonal server rack, characterized in that, include: Multiple compute nodes and multiple internal switching nodes are orthogonally connected. Each internal switching node has an uplink port for connecting to an external switching node, which is a switching node other than the internal switching nodes. Each compute node includes a first controller, and each switching node includes a second controller. The first controller is used to execute the multi-functional configuration method for orthogonal server racks as described in any one of claims 1 to 13; The second controller is used to execute the multi-functional configuration method for orthogonal server racks as described in any one of claims 14 to 16.

18. The orthogonal server rack according to claim 17, characterized in that, Each computing node is provided with multiple first connectors, and each internal switching node is provided with multiple second connectors. The first connectors and the second connectors are plugged into each other so that the computing node and any of the switching nodes are orthogonal to each other and pluggable. The first controller is connected to the first connector, and the second controller is connected to the second connector. The first connector and the second connector have the same pin definition. Each first connector includes a slot support capability identifier pin and an expansion identifier pin. The slot support capability identifier pin is used to transmit the slot support capability signal, and the expansion identifier pin is used to transmit the horizontal expansion function signal. The slot support capability signal represents the maximum number of switching node slots supported by the current orthogonal server rack. The first connector and the second connector include differential signal transmission pins for transmitting differential data.

19. The orthogonal server rack according to claim 18, characterized in that, The first connector includes a compute node identification pin, which is used to identify the mounting slot of the compute node. The slot information of the compute node output by different second connectors of each switch node is different. If no switch node is installed in the switch node slot, the compute node identification pin of the first connector corresponding to the switch node slot has no information, and the internal characteristics of the compute node are high impedance. The second connector includes a switching node identification pin, which is used to identify the mounting slot of the switching node. The slot information of the switching node output by the first connector of each computing node is different. If a computing node is not installed in the computing node slot, the switching node identification pin of the second connector corresponding to the computing node slot has no information, and the internal state of the switching node is characterized as high impedance.

20. The orthogonal server rack according to claim 18, characterized in that, The plurality of identification pins of each of the first connectors are located in the same column or multiple adjacent columns on the same side of all the signal transmission pins, and the plurality of identification pins of each of the second connectors are located in the same row or multiple adjacent rows on the same side of all the signal transmission pins.

21. The orthogonal server rack according to claim 18, characterized in that, The switching node also includes: A switching processor is connected to the second connector and the second controller respectively, and the switching processor is used to control the opening or closing of the uplink port.

22. The orthogonal server rack according to claim 18, characterized in that, Each of the switching nodes includes an accelerator, which is connected to the first connector.

23. The orthogonal server rack according to claim 22, wherein if the bandwidth of the accelerator in the computing node is within a first bandwidth range, the first connector is a first type connector; if the bandwidth of the accelerator in the computing node is within a second bandwidth range, the first connector is a second type connector, wherein the minimum value of the second bandwidth range is greater than the maximum value of the first bandwidth range. The second type of connector is formed by splicing together multiple first type connectors. A separation structure is provided between any two adjacent first type connectors, and the positions of the first type connectors that are adapted to the first bandwidth range in different second type connectors are the same.

24. The orthogonal server rack according to claim 17, characterized in that, The switching node slots for installing the switching nodes are movable and removable slots to accommodate switching nodes with different numbers of uplink ports. The switching nodes with different numbers of uplink ports correspond to switching node slots of different widths. The number of internal switching nodes installed is related to the total external bandwidth of all the computing nodes.

25. The orthogonal server rack according to claim 24, characterized in that, The orthogonal server rack includes a sliding rail installed on the inner wall of the orthogonal server rack, and the switching node slots are installed on the sliding rail and are movable on the sliding rail.

26. A server system, characterized in that, include: Orthogonal server rack as described in any one of claims 17 to 25; An external switching node is connected to the internal switching node in the orthogonal server rack via an uplink port.

Citation Information

Patent Citations

  • Multi-accelerator card heterogeneous server and resource link reconstruction method

    CN117687956A

  • Equipment interconnection system and method, equipment, medium and program product

    CN119052199A