GPU server and cabinet

By vertically mounting the backplane and GPU board inside the server and optimizing the layered layout of the GPU board using high-speed connectors and cable trays, the problems of limited GPU quantity and communication performance are solved, achieving high-efficiency computing density and improved communication performance.

CN223582412UActive Publication Date: 2025-11-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202422988198.0
Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-11-21
Estimated Expiration
2034-12-04

AI Technical Summary

Technical Problem

In existing technologies, the limited number of GPUs results in low computing power density, and the communication performance between GPUs is limited, which affects the training or inference efficiency of AI models.

Method used

By setting up a vertically connected backplane and GPU board inside the server, using high-speed connectors and cable trays to achieve a layered setup of the GPU board, and using connectors for point-to-point communication, combined with a switch for network communication, data transmission between GPUs is optimized.

Benefits of technology

It improves the computing power density and communication performance of GPU devices, thereby enhancing the training or inference efficiency of AI models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN223582412U_ABST
    Figure CN223582412U_ABST
Patent Text Reader

Abstract

The utility model provides a GPU server and a cabinet, and relates to the technical field of artificial intelligence, in particular to the technical fields of cloud computing, intelligent hardware and the like. The GPU server comprises a backboard; the plurality of GPU boards are vertically connected with the back board, and the height of the GPU boards is a preset height; a GPU and a connector are arranged on the GPU board, and the connector is used for connecting different GPUs. The computing power density and the communication performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of cloud computing, intelligent hardware and the like, and more particularly to a GPU server and a cabinet. BACKGROUND

[0002] A server is a basic device for building cloud computing, and the server is provided with a central processing unit (CPU), a memory, a hard disk, a power supply and the like. With the development of artificial intelligence (AI) technology, a graphics processing unit (GPU) has become one of the important processors. Therefore, an AI server is usually provided with multiple GPUs to meet the AI demand. SUMMARY

[0003] The present disclosure provides a GPU server and a cabinet.

[0004] According to an aspect of the present disclosure, a GPU server is provided, comprising: a backboard; a plurality of GPU boards, the GPU boards being connected perpendicularly to the backboard, and the height of the GPU boards being a preset height; a GPU and a connector being arranged on the GPU board, and the connector being used for connecting different GPUs.

[0005] According to another aspect of the present disclosure, a cabinet is provided, comprising: a plurality of GPU servers, each GPU server being as described in any one of the above aspects.

[0006] According to the technical solution of the present disclosure, the computing power density of the GPU device can be improved, and the communication performance between GPUs can be improved.

[0007] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0008] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0009] Figure 1 is a schematic view according to the first embodiment of the present disclosure;

[0010] Figure 2 is a schematic view according to the second embodiment of the present disclosure;

[0011] Figure 3 is a schematic view according to the third embodiment of the present disclosure;

[0012] Figure 4 is a schematic diagram of a communication mode between GPUs according to an embodiment of the present disclosure;

[0013] Figure 5 is a schematic diagram of another communication mode between GPUs according to an embodiment of the present disclosure;

[0014] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0015] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0016] Figure 8 is a schematic diagram of a cable bridge according to an embodiment of the present disclosure;

[0017] Figure 9 is a schematic diagram according to a sixth embodiment of the present disclosure;

[0018] Figure 10 is a schematic diagram according to a seventh embodiment of the present disclosure;

[0019] Figure 11 is a schematic diagram of connection of multiple GPUs and multiple switches according to an embodiment of the present disclosure;

[0020] Figure 12 is Figure 11 a schematic diagram of connection of data transmission ports of corresponding GPUs and switches;

[0021] Figure 13 is a schematic diagram of connection of multiple GPUs and a single switch according to an embodiment of the present disclosure;

[0022] Figure 14 is Figure 13 a schematic diagram of connection of data transmission ports of corresponding GPUs and switches. DETAILED DESCRIPTION

[0023] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in the context of being only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, descriptions of well-known functions and structures are omitted in the following description for the sake of clarity and conciseness.

[0024] As AI models become larger and larger, the number of parameters becomes larger and larger, in order to meet the demand for greater computing power, more GPUs are needed.

[0025] In the related art, GPUs are generally plugged on the mainboard through a Peripheral Component Interconnect Express (PCIe) slot inside the same server. PCIe is a high-speed serial computer expansion bus standard, which has the advantage of high data transmission rate.

[0026] The GPUs in different servers generally communicate through a network (e.g., Ethernet).

[0027] However, this layout has a limited number of GPUs inside each server, resulting in low computing power density. The GPUs in different servers communicate through a network, which affects the communication performance between GPUs and further affects the training or inference efficiency of AI models.

[0028] To improve computing power density and communication performance, the present disclosure provides the following embodiments.

[0029] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure. The present embodiment provides a GPU server.

[0030] As shown in Figure 1 The GPU server 100 includes a backboard 101 and a plurality of GPU boards 102a-102n. The backboard 101 is located inside the server case and is usually a Printed Circuit Board (PCB) that serves as a signal transmission function.

[0031] The GPU boards are connected perpendicularly to the backboard 101. Generally, the backboard is perpendicularly arranged inside the case, specifically perpendicularly to the mainboard inside the case. Since the GPU boards are perpendicularly connected to the backboard, the GPU boards are usually horizontally arranged inside the case.

[0032] The GPU boards are provided with GPUs and connectors for connecting different GPUs.

[0033] Specifically, the bottom layer of the GPU board is a PCB, which can be provided with a slot for plugging the GPU, or the GPU can be integrated on the PCB, so as to realize the setting of the GPU on the GPU board through plugging or integration. The PCB can also be provided with a connector for connecting different GPUs.

[0034] The height of the GPU board is a predetermined height, for example, the height of the GPU board is 1U. U is the height unit, 1U = 1.75 inches, about 4.445 cm.

[0035] The number of GPU boards is multiple, which means at least two. The specific number can be set according to the demand and the height of the GPU server in which it is located.

[0036] As shown in Figure 1 , assuming that the GPU boards in the GPU server are represented by 102a-102n respectively, these GPU boards are parallel to each other and are all connected to the backboard vertically to achieve the layered arrangement of different GPU boards.

[0037] In terms of specific connection mode, the GPU board can be plugged into the backboard or connected through a cable.

[0038] Taking the GPU plugging into the GPU board as an example, a preset number of slots, such as PCIe slots, are provided on each GPU board, and the GPU is plugged into the slot. For example, the preset number is 4, so a maximum of 4 GPUs can be plugged into each GPU board.

[0039] A connector is also provided on each GPU board, which is used to connect different GPUs, which can be located on the same GPU board or on different GPU boards.

[0040] In order to realize high-speed communication between GPUs, the connector can be a high-speed connector, which is used to transmit data signals at a high rate (such as greater than 20 Gbps), 20 Gbps means transmitting 20 Gbits of data per second.

[0041] In this embodiment, the multiple GPU boards are connected to the backboard vertically, which can realize the layered arrangement of multiple GPU boards, and multiple layers of GPU boards can be arranged as needed to improve the computing power density. By connecting different GPUs through the connector, communication between GPUs through the connector can be realized, which can realize point-to-point communication and improve communication performance compared to network communication.

[0042] The above embodiments show GPU-related components. In actual scenarios, the GPU server can also include other components such as CPUs, etc. Therefore, the present disclosure also provides other embodiments.

[0043] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure, and the present embodiment provides a GPU server.

[0044] Taking two GPU boards as an example, as shown in Figure 2 , the GPU server 200 includes a backboard 201 and multiple GPU boards 202a-202b, and further includes a CPU board 203, which is connected to the backboard vertically. The CPU board can be provided with elements such as CPUs.

[0045] In some embodiments, the GPU server further comprises a switch board 204, which is vertically connected to the backboard. The switch board is provided with network cards, and different GPU servers communicate with each other through the switch board.

[0046] In terms of specific connection mode, the CPU board, the GPU board and the switch board can be plugged into the backboard, or can also be connected to the backboard through a cable.

[0047] In this embodiment, by arranging the CPU board, the switch board and the like in the server, different scene requirements can be met, and the performance of the server can be improved.

[0048] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure, and the present embodiment provides a GPU server. In order to more clearly understand the internal structure of the GPU server, the present embodiment shows an exploded view of the internal components of the GPU server.

[0049] As shown in Figure 3 , the GPU server comprises a CPU board, a backboard and a GPU board. The CPU board is provided with a CPU, and the present embodiment takes two CPUs as an example. In installation, the backboard is vertically connected to the CPU board, and each GPU board is vertically connected to the backboard, so as to realize the layered arrangement of the GPU board.

[0050] The present embodiment takes two GPU boards as an example, and each GPU board is provided with a plurality of slots, so that a plurality of GPUs can be plugged in. In the present embodiment, four GPUs are plugged into each GPU board as an example.

[0051] On each GPU board, a plurality of GPUs are horizontally arranged on the GPU board, for example, referring to Figure 3 , four GPUs are horizontally plugged into the GPU board.

[0052] In the present embodiment, by horizontally arranging a plurality of GPUs on the GPU board, the number of GPUs can be increased, and the computing power density can be further improved. Among them, the horizontal arrangement means that the plane where the GPUs are located is horizontal relative to the GPU board, and the plane where the GPUs are located is vertical relative to the backboard. In addition, a plurality of slots can be arranged on the GPU board to plug in a plurality of GPUs.

[0053] Each GPU board is further provided with a connector, and the present embodiment takes a high-speed connector as an example, that is, a connector that can transmit a high-speed data signal (such as greater than 20 Gbps). Each GPU corresponds to at least one high-speed connector, and different GPUs are connected through the high-speed connector to transmit a data signal.

[0054] The high-speed connectors can be connected through a cable, which can be a cable for transmitting a high-speed data signal, such as a PCIe cable.

[0055] In order to facilitate the management of the cables, the cables can be placed in a cable device, such as Figure 3 As shown, the cable device can be a cable tray, which is a device for supporting and protecting cables and is commonly used in electrical engineering.

[0056] In this embodiment, by setting the cable device, the management of the cables can be facilitated, and the data transmission performance can be improved.

[0057] Further, by taking the cable tray as the cable device, the management of the cables can be efficiently and conveniently achieved.

[0058] Figure 4 is a schematic diagram of a communication mode between GPUs according to an embodiment of the present disclosure.

[0059] As shown in Figure 4 , assuming that two GPUs in mutual communication are a first GPU and a second GPU, and the high-speed connectors corresponding to the two GPUs are a first high-speed connector and a second high-speed connector, the first GPU and the second GPU communicate through the first high-speed connector and the second high-speed connector, specifically, the first high-speed connector and the second high-speed connector are connected through a cable, which can be a signal line for transmitting high-speed signals, such as a PCIe cable.

[0060] Taking the transmission of a data signal from the first GPU to the second GPU as an example, the data signal generated by the first GPU is first transmitted to the first high-speed connector, and the first high-speed connector transmits the data signal to the second high-speed connector through the cable placed in the cable tray, and the second high-speed connector transmits the data signal to the second GPU.

[0061] When two GPUs communicate, they can communicate directly point-to-point as shown in Figure 4 . In the case where the number of GPU ports is limited, in order to realize the communication between a single GPU and any other GPU, a switch can be set between the GPUs to realize the communication between the GPUs through the switch.

[0062] Figure 5 is a schematic diagram of another communication mode between GPUs according to an embodiment of the present disclosure.

[0063] As shown in Figure 5 , each GPU is connected to a switch through a high-speed connector and a cable, so that different GPUs can communicate through the switch, meeting the communication needs when the number of GPU ports is small.

[0064] The above embodiments describe a GPU server, which is usually placed in a cabinet.

[0065] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure, which provides a cabinet.

[0066] As shown in Figure 6 , the cabinet 600 includes a plurality of GPU servers 601a, 601b-601m.

[0067] Each GPU server is as shown in any of the above embodiments.

[0068] In this embodiment, by arranging the above GPU servers in the cabinet, the computing power density of the cabinet can be improved, and the communication performance between GPUs can be improved.

[0069] The GPUs in the same GPU server are connected through the connectors, and in the cabinet scenario, the GPUs in different GPU servers are also connected through the connectors. At this time, the connectors can also be connected through cables, and accordingly, the cables can also be managed by the cable device.

[0070] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure, which provides a cabinet.

[0071] As shown in Figure 7 , the cabinet includes a plurality of GPU servers, and a cable device such as a cable tray can also be installed on the cabinet, for example, fixed to the rear of the cabinet by screws.

[0072] In this way, the GPUs in different GPU servers can communicate through the connectors and the cables placed in the cable tray.

[0073] In this embodiment, by arranging the cable device, the cables can be managed conveniently, and the data transmission performance can be improved.

[0074] Further, by taking the cable tray as the cable device, the management of the cables can be efficiently and conveniently realized.

[0075] A plurality of GPUs are usually arranged on each GPU board, and these GPUs can be divided into a plurality of groups. Accordingly, a plurality of cable devices can be arranged, each of which corresponds to a group of GPUs, that is, the cables corresponding to each group of GPUs are placed in each cable device.

[0076] For example, four GPUs are arranged on each GPU board, and these four GPUs can be divided into two groups, that is, the two GPUs on the left side form one group, and the two GPUs on the right side form another group. At this time, as shown in Figure 7As shown, two cable trays can be arranged, the left cable tray is used to place the cables corresponding to the two GPUs on the left side, and the right cable tray is used to place the cables corresponding to the two GPUs on the right side.

[0077] In this embodiment, by arranging the cable device into multiple groups, different cable devices can be used to place the cables corresponding to different groups of GPUs, which can avoid the problem of signal transmission loss caused by excessively long cables and improve data transmission efficiency.

[0078] The above example of one cable device corresponding to each group of GPUs can also be that multiple GPUs correspond to one cable device, that is, the GPUs are multiple, the cable device is one, and the cables corresponding to the multiple GPUs are placed in the one cable device.

[0079] For example, only the GPUs in a single GPU server are needed, such as two layers of GPU boards in the single GPU server, 4 GPUs are arranged on each GPU board, and in the case of 8 GPUs, a cable device can be used to place the cables.

[0080] In this embodiment, by arranging multiple GPUs to correspond to one cable device, resources and space can be saved, and resource utilization can be improved.

[0081] In the case of needing more GPUs, multiple GPU servers can be arranged, each GPU server includes multiple GPU boards, multiple (such as 4) GPUs are arranged on each GPU board, and the GPUs in different GPU servers communicate through high-speed connectors and cables, and the cables are placed in the cable tray.

[0082] At this time, the multiple GPU servers are arranged in layers along the up-down direction (vertical direction) in the cabinet, and the multiple GPU boards are also arranged in layers along the up-down direction (vertical direction) in the GPU server. Therefore, as shown, Figure 7 a higher cable tray (cable tray) is needed.

[0083] In the case of needing fewer GPUs, such as only the GPUs in a single GPU server are needed, such as two layers of GPU boards in the single GPU server, 4 GPUs are arranged on each GPU board, and in the case of 8 GPUs, cable tray replacement can be performed, and a smaller cable tray can be used to replace a larger cable tray.

[0084] In this way, the cable tray can be replaced in different scenarios.

[0085] Figure 8 is a schematic diagram of replacing the cable tray according to an embodiment of the present disclosure.

[0086] AsFigure 8 As shown in the figure, when the number of required GPUs changes from more to less, the larger cable bridge (represented by the first cable bridge) corresponding to more GPUs can be removed and replaced by the smaller cable bridge (represented by the second cable bridge) corresponding to less GPUs, and the second cable bridge is used to place the cables corresponding to the two-layer GPU boards in the single-GPU server.

[0087] Conversely, if the number of required GPUs changes from less to more, the second cable bridge can be replaced by the first cable bridge to connect a larger number of GPUs.

[0088] In this way, through the replacement of the cable bridge, various scenarios can be flexibly adapted.

[0089] Different GPUs can use the direct communication mode as shown in Figure 4 , or can also use the mode of communication through the switch as shown in Figure 5 .

[0090] For this purpose, a switch can also be provided in the cabinet to realize communication between GPUs through the switch.

[0091] The switch can exchange data signals between GPUs in different GPU servers, or can also exchange data signals between GPUs in the same GPU server, and the setting position of the switch can be different based on different scenarios.

[0092] In this embodiment, by providing a switch, the switch can be used to exchange signals between different GPUs, which is suitable for scenarios where the number of GPUs is large but the number of ports is small, and can meet the large-scale GPU communication demand.

[0093] Figure 9 is a schematic diagram according to the sixth embodiment of the present disclosure, and the present embodiment provides a cabinet.

[0094] In this embodiment, the exchange of data signals between GPUs in different GPU servers by the switch is taken as an example. At this time, the switch is arranged at the intermediate position of the plurality of GPU servers.

[0095] As shown in Figure 9 , taking four GPU servers as an example, represented by the first GPU server to the fourth GPU server, the switch is placed at the intermediate position of these GPU servers, that is, two GPU servers (the first GPU server and the second GPU server) are placed above the switch, and two GPU servers (the third GPU server and the fourth GPU server) are placed below the switch.

[0096] In the embodiment, when the switch exchanges signals of GPUs in different GPU servers, the switch is arranged in the middle of the GPU servers, so that the different GPU servers are evenly distributed on both sides of the switch, and signal transmission loss is reduced.

[0097] Figure 10 is a schematic view according to the seventh embodiment of the present disclosure, and the embodiment provides a cabinet.

[0098] In the embodiment, the switch exchanges data signals of GPUs in the same GPU server, for example. At this time, the switch is arranged adjacent to the GPU server.

[0099] For example, referring to Figure 10 , it is assumed that the switch needs to exchange data signals of GPUs in the first GPU server, at this time, the positions of the first GPU server and the second GPU server can be exchanged, so that the first GPU server is adjacent to the switch. It can be understood that the position of the second GPU server and the switch can also be exchanged, so that the first GPU server is adjacent to the switch.

[0100] In the embodiment, when the switch exchanges signals of GPUs in the same GPU server, the switch is arranged adjacent to the same GPU server, so that the signal transmission path between the GPUs is shortened, and the signal transmission loss is reduced.

[0101] In order to realize redundancy backup, multiple data transmission ports can be arranged on a single GPU to improve reliability and transmission efficiency.

[0102] For example, for each GPU, there are 4 data transmission ports thereon, when a certain data transmission port is busy, other data transmission ports can be used for data signal transmission.

[0103] At this time, multiple switches can be arranged, for example, 4 switches are arranged.

[0104] For multiple data transmission ports and multiple switches, each data transmission port can be connected with a corresponding switch, or all data transmission ports can be connected with the same switch.

[0105] Figure 11 is a connection schematic view of multiple GPUs and multiple switches according to the embodiment of the present disclosure. Figure 12 is Figure 11 a connection schematic view of data transmission ports of the corresponding GPU and the switch.

[0106] For example, Figure 11As shown, assuming that 4 GPUs are arranged on each GPU board and 4 data transmission ports exist for each GPU, 4 switches are arranged, represented by a first switch to a fourth switch, each GPU is connected with all switches, and each switch is connected with all GPUs.

[0107] At this time, as shown in Figure 12 each data transmission port is connected with a corresponding switch, so that each GPU is connected with each switch, and each switch is connected with each GPU.

[0108] For example, the four data transmission ports are respectively referred to as a first port to a fourth port, the first port is connected with the first switch, the second port is connected with the second switch, and the rest are similar.

[0109] In specific signal transmission, taking that the first GPU sends data signals to the second GPU and data transmission is performed through the first port as an example, the data signal output by the first port of the first GPU is first transmitted to the first connector corresponding to the first GPU, then transmitted to the first switch corresponding to the first port through the first connector and a cable, and then transmitted to the second GPU through the second connector corresponding to the second GPU by the first switch through a cable, specifically to the first port of the second GPU, so that signal transmission between the first GPU and the second GPU is realized.

[0110] In the embodiment, data transmission is performed by using mutually corresponding data transmission ports and switches, which can realize redundant backup and improve data transmission efficiency and reliability.

[0111] The connection mode of each data transmission port and a corresponding switch described above can be applied to a large-scale GPU scenario. When the number of GPUs is small, for example, in the scenario of communication of 8 GPUs in the same GPU server, only a single switch can be used for signal exchange, and at this time, multiple data transmission ports on each GPU are connected with the single switch.

[0112] Figure 13 is a connection diagram of multiple GPUs and a single switch according to an embodiment of the disclosure. Figure 14 is Figure 13 a connection diagram of data transmission ports of a corresponding GPU and a switch.

[0113] As shown in Figure 13 the multiple GPUs are 8 and located in the same GPU server and arranged in two layers, that is, the GPU server includes two layers of GPU boards, and each layer of GPU board includes 4 GPUs, and the GPUs are all connected with the same switch.

[0114] Taking the four switches as an example, the same switch can be any one of the four switches, or each switch can be taken as the same switch respectively.

[0115] At this time, as shown in Figure 14 each GPU has a plurality of data transmission ports, and in this embodiment, four data transmission ports are taken as an example, and each data transmission port is connected with the same switch, that is, all data transmission ports are connected with the same switch.

[0116] For example, the four data transmission ports are called first port to fourth port, and the same switch is the first switch, and the first port to the fourth port are connected with the first switch.

[0117] In this embodiment, all data transmission ports of the GPU are connected with the same switch, which can be applied to a small-scale GPU scenario to improve data transmission efficiency.

[0118] It can be understood that the same or similar contents in different embodiments in the disclosure embodiments can be mutually referred to.

[0119] It can be understood that the "first", "second" and the like in the disclosure embodiments are only used for distinction, and do not represent the importance level, time sequence, and the like.

[0120] It can be understood that in the disclosure embodiments, the directions or position relationships indicated by "up", "down", "left", "right", "front", "back", "inside", "outside", "top", "bottom", "vertical", "horizontal" and the like are the directions or position relationships shown in the drawings, and are only used for describing the relative position relationship between the components or constituent parts, and do not particularly limit the specific installation direction of the components or constituent parts.

[0121] In addition, in addition to being used to indicate the direction or position relationship, the above-mentioned part of the terms can also be used to indicate other meanings, for example, the term "up" can also be used to indicate a certain dependent relationship or connection relationship in some cases. For those skilled in the art, the specific meaning of these terms in the disclosure can be understood according to the specific situation.

[0122] In addition, the terms "mount", "set", "provided with", "provided with", "connect", "connect" should be understood broadly. For example, it can be fixedly connected, detachably connected, or integrally configured; it can be mechanically connected, or electrically connected; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication between two devices, elements or constituent parts. For those skilled in the art, the specific meaning of the above-mentioned terms in the disclosure embodiments can be understood according to the specific situation.

[0123] In addition, the shapes, structures, proportions, sizes, etc. drawn in the drawings of the present disclosure are only used to match the disclosed content, for the technicians in the art to understand and read, and are not used to limit the defined conditions of the implementation of the present disclosure. Any modification of shape, structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that can be achieved by the present disclosure, should still fall within the scope of the disclosed technical content.

[0124] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A GPU server, comprising: a backboard; a plurality of GPU boards, the GPU boards being connected to the backboard perpendicularly, and the height of the GPU boards being a preset height; a GPU and a connector being arranged on the GPU board, the connector being used for connecting different GPUs. 2.The server according to claim 1, further comprising: a CPU board and / or a switching board; the CPU board and / or the switching board being connected to the backboard perpendicularly. 3.The server according to claim 1, wherein: a plurality of slots are arranged on the GPU board to plug a plurality of GPUs; the plurality of GPUs are arranged horizontally on the GPU board. 4.A cabinet, comprising: a plurality of GPU servers; each of the GPU servers being as claimed in any one of claims 1-3. 5.The cabinet according to claim 4, wherein: the connectors are connected through cables; the cabinet further comprises: a cable device for placing the cables.

6. The cabinet of claim 5, wherein, the cable device comprises: a cable bridge. 7.The cabinet according to claim 6, wherein: the GPUs are a plurality of GPUs, and the plurality of GPUs are divided into a plurality of groups, the cable devices are a plurality of cable devices, and the cables corresponding to each group of GPUs are placed in each cable device, respectively; alternatively, the GPUs are a plurality of GPUs, and the cable devices are one cable device, and the cables corresponding to the plurality of GPUs are placed in the one cable device. 8.The cabinet according to claim 4, further comprising: a switch, and the GPUs in the GPU servers are connected to the switch through the connectors. 9.The cabinet according to claim 8, wherein: the switch is used for signal switching of the GPUs in different GPU servers; the switch is located at a middle position of the different GPU servers. 10.The cabinet according to claim 9, wherein: the GPUs have a plurality of data transmission ports, and the switches are a plurality of switches; each data transmission port of the GPUs is connected to a corresponding switch. 11.The cabinet according to claim 8, wherein: the switch is used for signal switching of the GPUs in a same GPU server; the switch is located at an adjacent position of the same GPU server. 12.The cabinet according to claim 11, wherein: the GPUs have a plurality of data transmission ports, and the switches are one switch; the plurality of data transmission ports of the GPUs are connected to the one switch.