Two-way artificial intelligence server

By introducing multiple AI topology switching mechanisms and a PCIe controller into the dual-socket AI server, the problem of the single topology of traditional servers is solved, enabling flexible AI resource allocation and efficient data transmission, improving server performance and stability, and making it suitable for AI application scenarios that require rapid response.

CN223598233UActive Publication Date: 2025-11-25HANGZHOU EBOYLAMP ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202423003294.1
Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-11-25
Estimated Expiration
2034-12-06

AI Technical Summary

Technical Problem

Traditional dual-socket AI server topologies are too simplistic and cannot meet the flexible allocation and scheduling needs of AI resources in different application scenarios, resulting in prolonged communication time and failing to maximize server performance.

Method used

It adopts a switching mechanism for multiple AI topologies (balanced topology, parallel topology, and cascaded topology), and optimizes the communication path between the CPU and high-speed peripheral devices by adjusting the cable plugging and unplugging of the PCIe bridge chip and MCIO interface. Combined with domestic CPUs and PCIe controllers, it supports hot-swapping and efficient data transmission.

Benefits of technology

It enables topology adjustments based on different AI training models and application requirements, reducing CPU resource consumption, lowering communication latency, improving server performance and efficiency, supporting domestic technologies, and ensuring system stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN223598233U_ABST
    Figure CN223598233U_ABST
Patent Text Reader

Abstract

A double-channel artificial intelligence server comprises a first CPU, a second CPU, a first MCIO interface connected with the first CPU and a second MCIO interface connected with the second CPU. The system also comprises a PCIe bridge sheet 1 and a PCIe bridge sheet 2. The PCIe bridge piece I comprises an uplink connection port I; the PCIe bridge sheet II comprises an uplink connection port II; the uplink connection port II is connected with the MCIO interface II; the first PCIe bridge piece and the second PCIe bridge piece are each provided with a plurality of PCIe signal channels used for being connected with high-speed peripheral equipment in an expanded mode. And one path of PCIe signal path expanded by the PCIe bridge piece I is connected with the MCIO interface III. According to the utility model, the server is allowed to be switched among different AI topologies, so that the differentiated requirements of clients on distribution and scheduling of AI resources in different application scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The utility model relates to the field of artificial intelligence especially relates to a double -way artificial intelligence server. BACKGROUND

[0002] Artificial intelligence server: also called AI server, is the high -performance computer system specially for running artificial intelligence application and big data calculation processing.

[0003] The traditional artificial intelligence server based on the domestic CPU is the double -way artificial intelligence server with two CPUs installed on the server mainboard.

[0004] The defect of prior art is that the topology form is single and cannot be changed. Utility model content

[0005] In view of the above defects of prior art, the utility model provides a double -way artificial intelligence server, allows the server to switch between different AI topologies, thereby satisfying the differentiated demand of customer to the allocation and scheduling of AI resource under different application scenarios.

[0006] To achieve the above purpose, the technical scheme adopted by the utility model is:

[0007] A dual-path artificial intelligence server, comprising a CPU one, a CPU two, an MCIO interface one connected with the CPU one and an MCIO interface two connected with the CPU two; further comprising a PCIe bridge one and a PCIe bridge two; the PCIe bridge one comprises an uplink port one; the PCIe bridge two comprises an uplink port two; the uplink port two is connected with the MCIO interface two; the PCIe bridge one and the PCIe bridge two are both extended with multiple PCIe signal paths for connecting high-speed peripheral devices; one of the PCIe signal paths extended by the PCIe bridge one is connected with an MCIO interface three.

[0008] The AI topology form comprises three types, namely balanced topology, parallel topology and cascade topology, and the three forms have advantages and disadvantages in performance and CPU usage rate of the server, and the corresponding topology form can be selected according to different AI training models. The PCIe signal path can be used to connect different high-speed peripheral devices. The MCIO interface is used to help flexible connection between network devices in the server.

[0009] The number of CPUs determines the computing power of the server; the communication efficiency between high-speed peripheral devices depends on whether the communication is through the CPU; and the number of PCIe buses affects the data transmission bandwidth between the CPU and the high-speed peripheral devices. The dual-path artificial intelligence server of the utility model adjusts the AI topology form through internal cable plugging and replacement. The adjustment conditions are as follows:

[0010] Balanced topology. The uplink port one of the PCIe bridge one is connected with the MCIO interface one. One PCIe bridge is connected under each CPU, and each PCIe bridge is extended with a group of PCIe signal paths. This topology form allows any two high-speed peripheral devices connected by PCIe signals in the group to communicate P2P through the PCIe bridge, and the communication efficiency of the PCIe signal path in the group is high. P2P communication, namely point-to-point communication, is a network communication mode in which participants in the network directly exchange data or information instead of transferring through the CPU. The advantage of balanced topology is that two groups of PCIe signal paths are mounted under two CPUs respectively, load balancing, and the CPU computing power is higher.

[0011] Parallel topology. The uplink port one of the PCIe bridge one is connected with the MCIO interface two. Two PCIe bridges are connected with the CPU two, and each PCIe bridge is connected with a group of PCIe signal paths. The advantage of this topology is that the CPU and the high-speed peripheral devices communicate through two PCIe buses, which can provide higher data transmission bandwidth; and the two high-speed peripheral devices outside the group can communicate P2P through one CPU and one PCIe bridge, which is more efficient than the balanced topology (the two high-speed peripheral devices outside the group need to communicate through two CPUs), and data can be quickly transmitted between the PCIe signal path and the CPU, improving the speed and efficiency of data transmission. But the disadvantage is that all PCIe signal paths are mounted under the same CPU, and the server computing power is lower than the balanced topology.

[0012] Cascade topology. The uplink port two of the PCIe bridge two is connected with the MCIO interface three. The CPU two is directly connected with the PCIe bridge two, and the PCIe bridge two is interconnected with another PCIe bridge one, and each PCIe bridge is connected with a group of PCIe signal paths below. The advantage of this topology is that the CPU two and all high-speed peripheral devices share one PCIe bus, which can save PCIe bus resources, and the disadvantage is that all PCIe signal paths are mounted under the same CPU, and the server computing power is lower than the balanced topology.

[0013] As a preferred, the dual-path artificial intelligence server comprises a mainboard and an expansion board; the CPU one and the CPU two are arranged on the mainboard; and the PCIe bridge one and the PCIe bridge two are arranged on the expansion board.

[0014] The CPU one and the CPU two are installed on the mainboard, and the PCIe bridge one and the PCIe bridge two are installed on the expansion board. By installing the CPU and the PCIe bridge on the mainboard and the expansion board respectively, the space layout inside the server can be optimized, and the integration of components can be improved. Separating the CPU and the PCIe bridge can improve the heat dissipation management, and the heat dissipation scheme such as heat sink or fan can be designed respectively to adapt to the respective heat load. At the same time, this layout supports modular design, which is convenient for maintenance and upgrading.

[0015] As a preferred, the MCIO interface one comprises a mainboard MCIO interface one and an expansion board MCIO interface one; and the MCIO interface two comprises a mainboard MCIO interface two and an expansion board MCIO interface two.

[0016] The mainboard MCIO interface one is connected with the CPU one, the mainboard MCIO interface two is connected with the CPU two, the extension board MCIO interface one is connected with the PCIe bridge one, and the extension board MCIO interface two is connected with the PCIe bridge two.

[0017] As preferred, the PCIe signal path is connected with a PCIe slot at the end.

[0018] The PCIe slot is a slot on the computer mainboard, which is used to connect and install various high-speed expansion cards, such as graphics processing units (GPU), network interface cards (NIC), sound cards, storage cards, etc.

[0019] As preferred, the PCIe slot is connected with an extension MCIO interface.

[0020] The extension MCIO interface is used to further enhance the network device expansion capability of the server. Through the extension MCIO interface connected by the PCIe slot, the server can connect more external network devices, such as high-speed network cards, fiber channel cards, storage expansion units, etc., thereby improving the network communication and data processing capability of the server. The extension MCIO interface provides additional connection points, so that the server can access more network resources, and is suitable for application scenarios that need to process a large number of network requests and data exchange.

[0021] As preferred, the CPU one and the CPU two are both connected with a PCIe controller.

[0022] The PCIe controller is a hardware component responsible for managing data transmission between the computer processor CPU and the PCIe bus. It implements the PCIe protocol to ensure efficient and reliable data transmission between the CPU and other devices connected through the PCIe bus. The PCIe controller supports high-speed serial data transmission and supports full-duplex communication mode, that is, data can be transmitted in both directions at the same time without waiting for the transmission in one direction to complete, improving the real-time performance and efficiency of communication. The PCIe controller of the utility model has advanced control and management functions, supports hot plug, that is, devices can be safely inserted or removed without shutting down the power supply.

[0023] As preferred, a heat dissipation module is further arranged on the mainboard, and the heat dissipation module comprises a refrigeration fan.

[0024] The heat dissipation module is installed on the mainboard, which can more effectively manage and dissipate the heat generated by the CPU, memory and other components, and keep the hardware running at a suitable working temperature. Through effective heat dissipation, the refrigeration fan helps to prevent system failure, restart or performance degradation caused by overheating, thereby improving the stability and reliability of the entire server system.

[0025] Compared with the prior art, the beneficial effects of the utility model lie in:

[0026] 1. The utility model discloses an internal cable plug-in replacement through uplink port one, uplink port two and MCIO interface three, realizes the adjustment of AI topological form, can adjust the AI topological form through simple cable plug-in operation according to different AI training models and application demand, to adapt to different work load and performance requirement.

[0027] 2. Through different topological forms, the server can optimize the communication path between PCIe signal paths, reduce the occupation of CPU resources, and reduce the communication delay, thereby improving the performance and efficiency of the server in various AI application scenarios.

[0028] 3. The utility model uses domestic CPU + AI topological switching technology, and the server system adopts domestic technology and equipment on key technologies and core equipment, realizes high self-controllability, provides safety guarantee for the current global information wave in the national key information technology field, prevents information leakage and malicious attack.

[0029] 4. The utility model uses the latest domestic CPU, and the PCIe rate of communication with the expansion board can reach GEN4, and the direct PCIe rate of GPU acceleration card can reach GEN5, and the bandwidth reaches x16, compared with other platforms of domestic server, its performance is higher, and the rate is faster.

[0030] 5. Since CPU one and CPU two are arranged on the mainboard of the server, and PCIe bridge one and PCIe bridge two are arranged on the dedicated expansion board, therefore, the space utilization in the server is optimized, the integration of components is improved, the overall system is more compact and efficient, and it is helpful to design the respective heat dissipation solutions.

[0031] 6. Since the stable connection between the mainboard MCIO interface one and the expansion board MCIO interface one, and the stable connection between the mainboard MCIO interface two and the expansion board MCIO interface two, the distribution of MCIO interface on the mainboard and the expansion board is more reasonable, the neatness and logicality of wiring are significantly improved, not only the high-speed data transmission between the mainboard and the expansion board is ensured, but also the replacement and maintenance work of the mainboard and the expansion board is greatly simplified, and the maintenance efficiency and operation reliability of the whole system are improved. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is a schematic diagram of the topology of the dual-path artificial intelligence server of embodiment 1;

[0033] Figure 2 is a schematic diagram of the structure of the dual-path artificial intelligence server in the form of a balanced topology;

[0034] Figure 3 is a schematic diagram of the structure of the dual-path artificial intelligence server in the form of a parallel topology;

[0035] Figure 4 is a schematic diagram of the structure of the dual-path artificial intelligence server in the form of a cascaded topology.

[0036] wherein:

[0037] 11, CPU one; 12, CPU two; 21, MCIO interface one; 211, mainboard MCIO interface one; 212, extension board MCIO interface one; 22, MCIO interface two; 221, mainboard MCIO interface two; 222, extension board MCIO interface two; 23, MCIO interface three; 31, PCIe bridge one; 32, PCIe bridge two. DETAILED DESCRIPTION

[0038] In order to make the technical means, creative features, purposes and effects achieved by the utility model easy to understand, the utility model will be further described in combination with specific diagrams. However, the utility model is not limited to the following embodiments.

[0039] It should be understood that the structures, proportions, sizes, etc. shown in the drawings attached to the present specification are only used to cooperate with the content disclosed in the specification for understanding and reading by those skilled in the art, and are not used to limit the implementation conditions of the utility model, so they do not have technical substantive significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effects and purposes that can be achieved by the utility model, should still fall within the scope of the technical content disclosed by the utility model.

[0040] Embodiment 1:

[0041] As Figure 1The double-path artificial intelligence server shown comprises CPU 11, CPU 12, MCIO interface 21 connected with CPU 11, and MCIO interface 22 connected with CPU 12; further comprising PCIe bridge 31 and PCIe bridge 32; PCIe bridge 31 comprises uplink port 1; PCIe bridge 32 comprises uplink port 2; uplink port 2 is connected with MCIO interface 22; PCIe bridge 31 and PCIe bridge 32 are both extended with multiple PCIe signal paths for connecting high-speed peripheral devices; one of the PCIe signal paths extended by PCIe bridge 31 is connected with MCIO interface 23.

[0042] The AI topology form comprises three types, namely balanced topology, parallel topology and cascade topology, and each of the three forms has advantages and disadvantages in terms of server performance and CPU usage rate, and the corresponding topology form can be selected according to different AI training models. The PCIe signal path can be used to connect different high-speed peripheral devices. The MCIO interface is used to help flexible connection between network devices in the server.

[0043] The end of the PCIe signal path is connected with a PCIe slot. The PCIe slot is a slot on a computer motherboard, which is used to connect and install various high-speed expansion cards such as graphic processing units (GPUs), network interface cards (NICs), sound cards, storage cards, etc. The PCIe slot supports full-duplex communication, and data can be transmitted and received simultaneously in both directions, improving data transmission efficiency.

[0044] The number of CPUs determines the computing power of the server; the communication efficiency between high-speed peripheral devices depends on whether the communication passes through the CPU; and the number of PCIe buses affects the data transmission bandwidth between the CPU and the high-speed peripheral devices. In the embodiment, each PCIe bridge is extended with a group of PCIe signal paths, and each group has five PCIe signal paths. The PCIe bridge 31 is extended with five PCIe signal paths, four of which are directly connected to the PCIe slot, and one of which is connected to the MCIO interface 23; the PCIe bridge 32 is extended with five PCIe signal paths, all of which are directly connected to the PCIe slot. The double-path artificial intelligence server of the present application is replaced by internal cables, and the AI topology form is adjusted. The adjustment conditions are as follows:

[0045] For example, if the CPU 11 is used for AI training, the CPU 12 is used for other tasks, and the CPU 11 is connected with the MCIO interface 21 and the MCIO interface 23, and the CPU 12 is connected with the MCIO interface 22 and the MCIO interface 24, then the balanced topology is formed, as shown in FIG. 2. Figure 2The double-path artificial intelligence server shown is a balanced topology. The uplink port 1 of the PCIe bridge 1 31 is connected to the MCIO interface 1 21. The PCIe slot of the PCIe bridge 1 31 is connected with four GPUs, and the MCIO interface 3 23 is connected with a NIC. The PCIe slot of the PCIe bridge 2 32 is connected with four GPUs and a NIC. This topology allows any two high-speed peripheral devices connected by PCIe signals in the group to communicate P2P through the PCIe bridge, and the communication efficiency of the PCIe signal path in the group is high. The advantage of the balanced topology is that the two groups of GPUs are mounted under two CPUs respectively, the load is balanced, the CPU computing power is higher, and the concurrent bandwidth between the GPU and the CPU is higher.

[0046] As shown in Figure 3 The double-path artificial intelligence server shown is a parallel topology. The uplink port 1 of the PCIe bridge 1 31 is connected to the MCIO interface 2 22. Two PCIe bridges are connected under the CPU 2 12, and each PCIe bridge is connected with a group of PCIe signal paths. The advantage of this topology is that the CPU and the high-speed peripheral device communicate through two PCIe x16 buses, which can provide higher data transmission bandwidth; and the two high-speed peripheral devices of the group outside can communicate P2P through a CPU and a PCIe bridge, which is more efficient than the balanced topology (the two high-speed peripheral devices outside the group need to communicate through two CPUs), and the data can be quickly transmitted between the PCIe signal path and the CPU, improving the speed and efficiency of data transmission. But the disadvantage is that all PCIe signal paths are mounted under the same CPU, and the server computing power is lower than the balanced topology.

[0047] As shown in Figure 4 The double-path artificial intelligence server shown is a cascade topology. The uplink port 2 of the PCIe bridge 2 32 is connected to the MCIO interface 3 23. The CPU 2 12 is directly connected to the PCIe bridge 2 32, and the PCIe bridge 2 32 is interconnected with another PCIe bridge 1 31, and each PCIe bridge is connected with a group of PCIe signal paths. The advantage of this topology is that the CPU 2 12 and all high-speed peripheral devices share a PCIe x16 bus, which can save PCIe bus resources, and the disadvantage is that all PCIe signal paths are mounted under the same CPU, and the server computing power is lower than the balanced topology.

[0048] The dual-path artificial intelligence server comprises a mainboard and an extension board; CPU one 11 and CPU two 12 are arranged on the mainboard; PCIe bridge one 31 and PCIe bridge two 32 are arranged on the extension board. CPU one 11 and CPU two 12 are installed on the mainboard, and PCIe bridge one 31 and PCIe bridge two 32 are installed on the extension board. By installing the CPU and the PCIe bridge on the mainboard and the extension board respectively, the spatial layout inside the server can be optimized, and the integration of components can be improved. Separating the CPU and the PCIe bridge can improve heat dissipation management and design separate heat dissipation schemes, such as heat sinks or fans, to adapt to their respective heat loads. At the same time, this layout supports modular design, facilitating maintenance and upgrading.

[0049] MCIO interface one 21 comprises mainboard MCIO interface one 211 and extension board MCIO interface one 212; MCIO interface two 22 comprises mainboard MCIO interface two 221 and extension board MCIO interface two 222. The mainboard MCIO interface one 211 is connected with CPU one 11, the mainboard MCIO interface two 221 is connected with CPU two 12, the extension board MCIO interface one 212 is connected with PCIe bridge one 31, and the extension board MCIO interface two 222 is connected with PCIe bridge two 32. The mainboard MCIO interface one 211 and the extension board MCIO interface one 212, and the mainboard MCIO interface two 221 and the extension board MCIO interface two 222 are connected through cables. The MCIO interfaces are distributed on the mainboard and the extension board, making the wiring layout clearer, not only ensuring high-speed data transmission between the mainboard and the extension board, but also making the replacement and maintenance of the mainboard and the extension board more convenient and efficient.

[0050] Specifically, the PCIe slot is connected with an extension MCIO interface. The extension MCIO interface is used to further enhance the network device expansion capability of the server. Through the extension MCIO interface connected by the PCIe slot, the server can connect more external network devices, such as high-speed network cards, fiber channel cards, storage expansion units, etc., thereby improving the network communication and data processing capability of the server. The extension MCIO interface provides additional connection points, so that the server can access more network resources, and is suitable for application scenarios that need to process a large amount of network requests and data exchange.

[0051] The CPU one 11 and the CPU two 12 are connected with PCIe controllers. The PCIe controllers are connected with the MCIO slot through a PCIe bus, mainly including clock signals, control signals, in-situ signals and data signals, so that high-speed data interaction between the CPU and the PCIe slot is ensured. The PCIe controller is a kind of hardware component, responsible for managing data transmission between the computer processor CPU and the PCIe bus. It implements the PCIe protocol, ensuring that data can be efficiently and reliably transmitted between the CPU and other devices connected through the PCIe bus. The PCIe controller supports high-speed serial data transmission and supports full-duplex communication mode, that is, data can be transmitted in two directions at the same time, without waiting for the transmission in one direction to be completed, improving the real-time performance and efficiency of communication. The PCIe controller of the utility model has advanced control and management functions, supports hot plugging, that is, devices can be safely inserted or removed without shutting down the power supply.

[0052] The dual-path artificial intelligence server also includes a heat dissipation module on the mainboard; the heat dissipation module includes a refrigeration fan. The heat dissipation module is installed on the mainboard, which can more effectively manage and dissipate the heat generated by the CPU, memory and other components, keeping the hardware running at a suitable working temperature. Through effective heat dissipation, the refrigeration fan helps to prevent system failures, restarts or performance degradation caused by overheating, thereby improving the stability and reliability of the entire server system.

Claims

1. A dual-processor artificial intelligence server, comprising a CPU one (11), a CPU two (12), an MCIO interface one (21) connected to the CPU one (11), and an MCIO interface two (22) connected to the CPU two (12); characterized in that, It also includes PCIe bridge chip one (31) and PCIe bridge chip two (32); PCIe bridge chip one (31) includes an uplink connection port one; PCIe bridge chip two (32) includes an uplink connection port two; the uplink connection port two is connected to the MCIO interface two (22); both PCIe bridge chip one (31) and PCIe bridge chip two (32) are extended with multiple PCIe signal paths for connecting high-speed peripheral devices; one of the PCIe signal paths extended by PCIe bridge chip one (31) is connected to MCIO interface three (23).

2. The dual-path artificial intelligence server according to claim 1, characterized in that, The dual-path AI server includes a motherboard and an expansion board; CPU 1 (11) and CPU 2 (12) are located on the motherboard; PCIe bridge 1 (31) and PCIe bridge 2 (32) are located on the expansion board.

3. The dual-path artificial intelligence server according to claim 2, characterized in that, The MCIO interface one (21) includes the motherboard MCIO interface one (211) and the expansion board MCIO interface one (212); the MCIO interface two (22) includes the motherboard MCIO interface two (221) and the expansion board MCIO interface two (222).

4. The dual-path artificial intelligence server according to claim 1, characterized in that, The PCIe signal path ends at a PCIe slot.

5. The dual-path artificial intelligence server according to claim 4, characterized in that, The PCIe slot is connected to an expansion MCIO interface.

6. The dual-path artificial intelligence server according to claim 1, characterized in that, Both CPU 1 (11) and CPU 2 (12) are connected to a PCIe controller.

7. The dual-path artificial intelligence server according to claim 2, characterized in that, It also includes a heat dissipation module disposed on the motherboard; the heat dissipation module includes a cooling fan.