Server and server cluster
Through the data interaction expansion card and the OCS module, the direct data interconnection of the computing accelerator card is solved, and the problem of insufficient data communication rate between the computing core devices is improved, and the computing efficiency and resource utilization of artificial intelligence servers and server clusters are improved.
Patent Information
- Application Number
- CN202521257869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2035-06-19
AI Technical Summary
The increase in data communication rate between computing core devices is difficult to match the computing speed of computing core devices, resulting in waste of computing resources and the overall computing efficiency of artificial intelligence servers.
The data interaction expansion card and the OCS module are used to realize the direct data interconnection between the computing accelerator card, bypass switches and network cards and adjust the data transmission topology through the configuration of the OCS module.
It improves the data communication rate between the computing accelerator cards, reduces communication delay, improves the computing resource utilization rate and overall computing efficiency of artificial intelligence servers and server clusters, and enhances the adaptability to different artificial intelligence network models.
Smart Images

Figure CN223205857U_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of server communication technology, and in particular to a server and a server cluster. Background Art
[0002] An artificial intelligence server usually includes multiple computing core devices such as GPUs (Graphics Processing Units, graphics processors) (the computing core devices in this disclosure may also be referred to as computing accelerator cards). The computing core devices interact with each other through data links. For different artificial intelligence models (such as neural network models), different computing core devices need to complete different computing tasks and transfer data between computing core devices through data links, thereby completing related tasks during the operation of the artificial intelligence model (such as training tasks, inference tasks, etc.) through the coordinated cooperation between multiple computing core devices. Depending on the different artificial intelligence model structures, the data links between the computing core devices will also have different topological forms.
[0003] With the development of artificial intelligence technology, the computing speed of computing core devices has been continuously improved. In contrast, it is increasingly difficult to further increase the data communication rate of the data link between computing core devices. As a result, during the operation of the artificial intelligence model, the computing core device needs to spend time waiting for the reception of computing data, resulting in a large proportion of the data reception waiting time of the computing core device during the operation of the artificial intelligence model. On the one hand, it causes a waste of computing resources, and on the other hand, it makes it difficult for the overall computing efficiency of the artificial intelligence server to match the computing speed of the computing core device itself.
[0004] It can be seen that how to improve the data communication rate of the data link between the computing core devices to improve the overall computing resource utilization of the artificial intelligence server, and improve the matching degree between the overall computing efficiency of the artificial intelligence server and the computing speed of the computing core devices therein has become a direction that requires continuous development and innovation. Summary of the Invention
[0005] In view of this, the present disclosure provides a server and a server cluster to help improve the data communication rate of the data link between the computing core devices, thereby helping to improve the overall computing resource utilization of the artificial intelligence server, and helping to improve the overall computing efficiency of the artificial intelligence server and even the server cluster composed of artificial intelligence servers and the degree of matching with the computing speed of the computing core devices therein.
[0006] The technical solution of the present disclosure is achieved as follows:
[0007] According to one aspect of an embodiment of the present disclosure, a server is provided, including:
[0008] At least two computing acceleration cards, each of which includes at least two data interaction ports;
[0009] At least one data interaction expansion card, the data interaction expansion card comprising a server internal connection port, a server external connection port group, and a data interaction channel, the server internal connection port and the server external connection port group being coupled via the data interaction channel, the server external connection port group being used to connect to a device external to the server;
[0010] Among them, the data interaction ports with the same identification between different computing acceleration cards are coupled to the server internal connection ports of the same data interaction expansion card, wherein the number of the server internal connection ports is equal to the number of the connected data interaction ports.
[0011] In one possible implementation, each of the data interaction ports includes multiple data channels;
[0012] The server internal connection port includes a plurality of data channel connection ends connected one-to-one with the plurality of data channels respectively;
[0013] The server external connection port group includes multiple data channel connection ports, the number of the data channel connection ports is the same as the number of the data channels, wherein the identifiers of the data channels in any one of the data channel connection ports are the same, and the identifiers of the data channels between different data channel connection ports are different.
[0014] In one possible implementation, the server includes a UBB;
[0015] The computing acceleration card and the data interaction expansion card are plugged into the UBB;
[0016] The UBB is provided with a data transmission network connecting the data interaction port and the internal connection port of the server.
[0017] In one possible implementation, the server further includes:
[0018] processor; and,
[0019] A PCIe switching chip is coupled to the computing acceleration card and to the processor, forming a data interaction path between the computing acceleration card and the processor.
[0020] In one possible implementation, the server further includes:
[0021] A NIC card is coupled to the PCIe switch chip, and is used to establish a data interaction connection between the server and a network switch outside the server.
[0022] In one possible implementation, the number of the data interaction expansion cards does not exceed the number of the data interaction ports in a single computing acceleration card.
[0023] According to another aspect of an embodiment of the present disclosure, a server is provided, including:
[0024] At least two computing acceleration cards, each of which includes at least two data interaction ports;
[0025] at least one data interaction expansion card, the data interaction expansion card comprising a server internal connection port, a server external connection port group, and a data interaction channel, the server internal connection port and the server external connection port group being coupled via the data interaction channel; and
[0026] An OCS module, the OCS module being coupled between the server external connection port group of the data interaction expansion card and constituting a data interaction path between the at least two computing acceleration cards;
[0027] Among them, the data interaction ports with the same identification between different computing acceleration cards are coupled to the server internal connection ports of the same data interaction expansion card, wherein the number of the server internal connection ports is equal to the number of the connected data interaction ports.
[0028] According to another aspect of an embodiment of the present disclosure, a server cluster is provided, including:
[0029] The server according to any one of the above items, wherein the number of the servers is at least two;
[0030] An OCS module is coupled between the server external connection port groups of the data interaction expansion cards of at least two of the servers, forming a data interaction path between the computing acceleration cards in at least two of the servers.
[0031] In a possible implementation manner, the server external connection port group includes a data path configuration end, and the data path configuration end is used to transmit data interaction path configuration information of the OCS module.
[0032] In one possible implementation, the server cluster further includes:
[0033] The data interaction network configuration server is coupled to the server and is used to generate the data interaction path configuration information.
[0034] In a possible implementation manner, the number of the OCS modules is not less than the number of the servers, each of the servers is coupled to at least one OCS module, and different OCS modules are connected via optical fibers.
[0035] It can be seen from the above scheme that the server disclosed in the present invention can realize the data connection of the computing acceleration card in the server bypassing the switch through the data interaction expansion card therein, providing a hardware foundation for the one-to-one direct data interconnection of the computing acceleration card. The server disclosed in the present invention can realize the one-to-one direct data interconnection of each computing acceleration card in the same server or between different servers by changing the connection relationship between the server external connection port groups of the data interaction expansion card. The server disclosed in the present invention also realizes one-to-one full connection between computing acceleration cards in a single server. The server disclosed in the present invention can, on the one hand, eliminate the direct interconnection circuit between computing acceleration cards in the UBB, reducing the difficulty of UBB design. On the other hand, the server can flexibly adjust the data transmission network topology relationship of the direct interconnection between each computing acceleration card according to the characteristics of the artificial intelligence network model being run, and can make the data communication link between each computing acceleration card bypass the switch, NIC card in the server and the IB network card and switch outside the server, eliminating the data transmission delay caused by these nodes, which helps to improve the overall model training and reasoning performance of the server and server cluster.
[0036] In the server cluster disclosed in the present invention, the data interaction expansion card in each server is used to realize the direct connection between the OCS module and the data interaction port of the computing acceleration card in each server, so that the data transmission between different computing acceleration cards can bypass the network card, switch and other data forwarding devices, thereby eliminating the transmission delay caused by the network card and switch during the data transmission process. In addition, the delay of one-to-one data transmission based on the OCS module is significantly lower than the delay of data transmission when the servers are interconnected through the network card and switch. At the same time, through the configuration of the OCS module, flexible adjustment of the data link topology for different artificial intelligence network models is also realized, thereby increasing the configurable flexibility of the server cluster.
[0037] The servers and server clusters disclosed herein help to improve the data communication rate of the data links between computing acceleration cards such as OAM in artificial intelligence servers, reduce communication latency, help improve the overall computing resource utilization of artificial intelligence servers, help improve the overall computing efficiency of artificial intelligence servers and even server clusters composed of artificial intelligence servers and the degree of matching with the computing speed of the computing core devices therein, and help improve the flexibility of the data link topology structure of artificial intelligence servers and artificial intelligence server clusters in adaptive adjustment to different artificial intelligence network models. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a schematic diagram of a partial structure of an artificial intelligence server in an OAM form in related technology;
[0039] Figure 2 This is a schematic diagram of the structure of an artificial intelligence server cluster in related technology;
[0040] Figure 3 is a structural diagram of a server according to an exemplary embodiment;
[0041] Figure 4 is a schematic diagram of a link topology structure from a data interaction port to a server external connection port group according to an exemplary embodiment;
[0042] Figure 5 is another schematic diagram of a link topology structure from a data interaction port to a server external connection port group according to an exemplary embodiment;
[0043] Figure 6 is a schematic structural diagram of another server according to an exemplary embodiment;
[0044] Figure 7 is a structural diagram of another server according to an exemplary embodiment;
[0045] Figure 8 is a schematic diagram of a server cluster according to an exemplary embodiment;
[0046] Figure 9 This is a schematic diagram of the link configuration between OCS modules;
[0047] Figure 10 This is a schematic diagram of a topological structure of a fully interconnected OAM machine implemented by using the server and server cluster of the embodiment of the present disclosure;
[0048] Figure 11 This is a schematic diagram of a topological structure of an OAM part intra-machine interconnection and part inter-machine interconnection implemented by the server and server cluster of the embodiment of the present disclosure;
[0049] Figure 12 A schematic diagram of a server topology in which two machines are interconnected and implemented using the server and server cluster of the embodiment of the present disclosure;
[0050] Figure 13 The present invention is a schematic diagram of a server topology implemented by using the server and server cluster of the embodiment of the present invention, which realizes switching between two-machine interconnection and four-machine interconnection by configuring the OCS module. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and examples.
[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0053] Figure 1 This is a partial structural diagram of an artificial intelligence server in the OAM form in the related technology. Figure 1 As shown, OAM stands for Open Accelerator Module, which is a specification established by the OCP (Open Compute Project) community under the server project. This specification standardizes the accelerator module, simplifies the design of artificial intelligence infrastructure, and helps shorten the hardware design cycle. The OAM is equipped with an artificial intelligence processor, which can be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In this application, a computing accelerator card may refer to the OAM.
[0054] like Figure 1 As shown in the related art, OAM-based AI servers are usually located in a single server, with eight OAMs plugged into a UBB (Universal Baseboard). Through circuit connections in the UBB, the eight OAMs are directly interconnected within the UBB board. The circuit topology of the direct interconnections between the OAMs within the UBB is fixed, and cannot be flexibly adjusted based on the characteristics of the AI model actually used by the server, resulting in limited performance optimization of the server. Furthermore, as Figure 1As shown, the server also includes a PCIe switch chip and two CPUs (Central Processing Units), wherein the eight OAMs are each coupled to the PCIe switch chip, and the two CPUs are respectively coupled to different PCIe switch chips. The two CPUs are interconnected with the four OAMs through different PCIe switch chips, and the two CPUs are interconnected with each other.
[0055] On the other hand, in related technologies, since the artificial intelligence server realizes direct interconnection of 8 OAMs through circuit connections in the UBB, it also leads to the complexity of the wiring structure within the UBB and a high number of stacked circuit layers in the UBB, which increases the signal interference between different traces in the UBB, increases the difficulty of UBB development and the material cost of UBB production, resulting in a low overall cost-effectiveness of the UBB.
[0056] Figure 2 This is a schematic diagram of the structure of an artificial intelligence server cluster in related technologies. Figure 2 As shown, in related technologies, the horizontal expansion of AI servers within a cluster is achieved by interconnecting AI servers through data interconnection and forwarding devices based on the IB / RoCE protocol. Data interconnection and forwarding devices, such as IB (InfiniBand) network cards and RoCE (RDMA over Converged Ethernet) switches, are used. In this approach, data exchange between AI servers is subject to transmission delays caused by the data interconnection and forwarding devices. Furthermore, data transmission between OAMs within different AI servers is also subject to transmission delays caused by switches (such as PCIe switching chips) and NICs (Network Interface Cards) within the OAM servers. Therefore, in an AI server cluster, the delays caused by various data forwarding devices on the data links between OAMs within different AI servers further extend the waiting time for OAM signals within the AI server cluster, thereby affecting the overall model training and inference performance of the AI server cluster.
[0057] In light of this, the present disclosure provides a server and server cluster to help increase the data communication rate of the data links between computing accelerator cards such as OAM in AI servers, reduce communication latency, and thereby help improve the overall computing resource utilization of AI servers. It also helps improve the overall computing efficiency of AI servers, and even server clusters composed of AI servers, and the degree to which it matches the computing speed of the computing core devices therein. It also helps increase the flexibility of the data link topology of AI servers and AI server clusters in adaptively adjusting to different AI network models.
[0058] Figure 3 FIG. 1 is a schematic diagram showing the structure of a server according to an exemplary embodiment. Figure 3 As shown, in the exemplary embodiment, the server mainly includes at least two computing acceleration cards 1 and at least one data interaction expansion card 2. Each computing acceleration card 1 includes at least two data interaction ports 11. In the exemplary embodiment, the structures of the various computing acceleration cards 1 are the same, and the number and structure of the data interaction ports 11 between the various computing acceleration cards 1 are the same. The data interaction expansion card 2 includes a server internal connection port 21, a server external connection port group 22, and a data interaction channel 23. The server internal connection port 21 and the server external connection port group 22 are coupled through the data interaction channel 23 and perform data interaction through the data interaction channel 23. The server external connection port group 22 is used to connect to an external server device to perform data interaction with the external server device. The data interaction ports 11 with the same identifier between different computing acceleration cards 1 are coupled to the server internal connection port 21 of the same data interaction expansion card 2, and the number of the server internal connection ports 21 is equal to the number of the connected data interaction ports 11.
[0059] In an exemplary embodiment, the computing accelerator card 1 can be, for example, an OAM form factor, and include an artificial intelligence chip. The artificial intelligence chip can be any of a GPU, TPU, NPU, DPU, APU, and GPGPU. In conjunction with the use scenarios and forms of artificial intelligence servers using OAM forms in related technologies, in an exemplary embodiment, a specific embodiment of the at least two computing accelerator cards 1 in the server can be eight. Each computing accelerator card 1 includes eight data exchange ports 11, and the number of data exchange expansion cards 2 does not exceed the number of data exchange ports 11 in a single computing accelerator card 1, i.e., the number of data exchange expansion cards 2 does not exceed eight. Based on this, in an exemplary embodiment, in conjunction with the use scenarios and forms of artificial intelligence servers using OAM forms in related technologies, the server can include eight computing accelerator cards 1 and a maximum of eight data exchange expansion cards 2. Each computing accelerator card 1 includes eight data exchange ports 11. Each data exchange expansion card 2 has eight server-internal connection ports 21, and each data exchange expansion card 2 has one server-external connection port group 22. The eight data interaction expansion cards 2 may be named EXP0, EXP1, EXP2, EXP3, EXP4, EXP5, EXP6, and EXP7, respectively.
[0060] In an exemplary embodiment, a server includes eight different computing accelerator cards 1, and each computing accelerator card 1 can perform one-to-one direct data exchange with other modules (or devices, or equipment, such as other computing accelerator cards 1) via multiple data exchange ports 11. To distinguish these computing accelerator cards 1 and data exchange ports 11, these computing accelerator cards 1 and data exchange ports 11 are identified by name plus number. For example, the eight different computing accelerator cards 1 in a server can be identified as OAM0, OAM1, OAM2, OAM3, OAM4, OAM5, OAM6, and OAM7, respectively, and the eight data exchange ports 11 in each computing accelerator card 1 can be identified as PORT0, PORT1, PORT2, PORT3, PORT4, PORT5, PORT6, and PORT7, respectively. Based on this, in the illustrative embodiment, the data interaction ports 11 with the same identification between different computing acceleration cards 1 are coupled to the server connection port 21 of the same data interaction expansion card 2, which means: the respective PORT0 of OAM0 to OAM7 is coupled to the server connection port 21 of EXP0, the respective PORT1 of OAM0 to OAM7 is coupled to the server connection port 21 of EXP1, and so on, the respective PORT7 of OAM0 to OAM7 is coupled to the server connection port 21 of EXP7.If a unique identifier is used to distinguish each data interaction port 11 in each computing accelerator card 1, in an exemplary embodiment, the eight data interaction ports 11 in OAM0 can be respectively identified as PORT00, PORT01, PORT02, PORT03, PORT04, PORT05, PORT06, and PORT07; the eight data interaction ports 11 in OAM1 can be respectively identified as PORT10, PORT11, PORT12, PORT13, PORT14, PORT15, PORT16, and PORT17; the eight data interaction ports 11 in OAM2 can be respectively identified as PORT20, PORT21, PORT22, PORT23, PORT24, PORT25, PORT26, and PORT27; and the eight data interaction ports 11 in OAM3 can be respectively identified as PORT30, PORT31, PORT32, PORT33, and PORT34. , PORT35, PORT36, PORT37, the 8 data interaction ports 11 in OAM4 can be respectively identified as PORT40, PORT41, PORT42, PORT43, PORT44, PORT45, PORT46, PORT47, the 8 data interaction ports 11 in OAM5 can be respectively identified as PORT50, PORT51, PORT52, PORT53, PORT54, PORT55, PORT56, PORT57, the 8 data interaction ports 11 in OAM6 can be respectively identified as PORT60, PORT61, PORT62, PORT63, PORT64, PORT65, PORT66, PORT67, the 8 data interaction ports 11 in OAM7 can be respectively identified as PORT70, PORT71, PORT72, PORT73, PORT74, PORT75, PORT76, PORT77. Based on this, in an illustrative embodiment, the data interaction ports 11 with the same identification between different computing accelerator cards 1 are coupled to the server connection port 21 of the same data interaction expansion card 2, which means: PORT00, PORT10, PORT20, PORT30, PORT40, PORT50, PORT60, PORT70 are coupled to the server connection port 21 of EXP0, PORT01, PORT11, PORT21, PORT31, PORT41, PORT51, PORT61, PORT71 are coupled to the server connection port 21 of EXP1, and so on, PORT07, PORT17, PORT27, PORT37, PORT47, PORT57, PORT67, PORT77 are coupled to the server connection port 21 of EXP7.
[0061] It can be seen that the number of server connection ports 21 of each data interaction expansion card 2 and the number of connected data interaction ports 11 are both 8, that is, the number of server connection ports 21 and the number of connected data interaction ports 11 are equal.
[0062] In the illustrative embodiment, there is a situation where it is not necessary to connect all the data interaction ports 11 of the computing accelerator card 1 to the outside through the data interaction expansion card 2. Therefore, the number of data interaction expansion cards 2 can be determined according to the design and application scenario requirements. For example, when only four data interaction ports 11 of the computing accelerator card 1 need to be connected to the outside through the data interaction expansion card 2, four data interaction expansion cards 2 can be used. For example, when PORT0, PORT2, PORT4 and PORT6 of the computing accelerator card 1 need to be connected to the outside through the data interaction expansion card 2, the server of the embodiment of the present disclosure may include four data interaction expansion cards 2, namely EXP0, EXP2, EXP4 and EXP6.
[0063] In an illustrative embodiment, each data interaction port 11 includes multiple data channels; the server internal connection port 21 includes multiple data channel connection ends that are respectively connected one-to-one with the multiple data channels; the server external connection port group 22 includes multiple data channel connection ports, and the number of data channel connection ports is the same as the number of data channels, wherein each data channel connection port includes data channels with the same identification. In other words, the identifications of the data channels in any data channel connection port are the same, and the identifications of the data channels between different data channel connection ports are different.
[0064] Based on OAM specifications in related art, each data exchange port 11 includes eight data channels. Therefore, in the exemplary embodiment, the server internal connection port 21 includes eight data channel connection terminals that are connected one-to-one with the eight data channels, and the server external connection port group 22 includes eight data channel connection ports. The number of data channel connection ports is eight, and the number of data channels is also eight. Therefore, the number of data channel connection ports is the same as the number of data channels. If the eight data channels in the data interaction port 11 are named lane0, lane1, lane2, lane3, lane4, lane5, lane6, and lane7 respectively, and the eight data channel connection ports are named connector0, connector1, connector2, connector3, connector4, connector5, connector6, and connector7 respectively, then all the data channels in connector0 are lane0, all the data channels in connector1 are lane1, all the data channels in connector2 are lane2, all the data channels in connector3 are lane3, all the data channels in connector4 are lane4, all the data channels in connector5 are lane5, all the data channels in connector6 are lane6, and all the data channels in connector7 are lane7.
[0065] Figure 4 This is a schematic diagram of a link topology structure from a data interaction port to a server external connection port group according to an exemplary embodiment. Figure 4 The embodiment shown only takes PORT2 of the data interaction ports of each OAM as an example, wherein the related link topology of OAM2 to OAM6 is omitted. The omitted part can be referred to Figure 4 The related link topology of OAM0, OAM1 and OAM7 is realized. Figure 4 As shown, PORT2 of each of OAM0 to OAM7 is coupled to the corresponding server-internal connection port of EXP2. Taking the server-internal connection ports of EXP2 as an example, PORT2 of OAM0 is coupled to EXP2_P0 of EXP2, PORT2 of OAM1 is coupled to EXP2_P1 of EXP2, ..., PORT2 of OAM7 is coupled to EXP2_P7 of EXP2. Figure 4It can be seen that the data interaction ports (for example, PORT2) with the same identifier between different computing accelerator cards (OAM0 to OAM7) are coupled to the server-internal connection ports (for example, EXP2_P0 to EXP2_P7) of the same data interaction expansion card (for example, EXP2). The number (8) of the server-internal connection ports (for example, EXP2_P0 to EXP2_P7) of each data interaction expansion card (for example, EXP2) is equal to the number (8) of the connected data interaction ports (PORT2 of OAM0 to PORT2 of OAM7).
[0066] like Figure 4 As shown, in the exemplary embodiment, PORT2 of OAM0 to PORT2 of OAM7 each include 8 data channels of lanes 0 to lanes 7, based on which EXP2_P0 of EXP2 includes 8 data channel connection ends connected one-to-one with lanes 0 to lanes 7 of PORT2 of OAM0, EXP2_P1 of EXP2 includes 8 data channel connection ends connected one-to-one with lanes 0 to lanes 7 of PORT2 of OAM1, ..., EXP2_P7 of EXP2 includes 8 data channel connection ends connected one-to-one with lanes 0 to lanes 7 of PORT2 of OAM7, EXP The server external connection port group of 2 includes 8 data channel connection ports of connector0, connector1, ..., connector7. The data channels in connector0 are all lane0 (each lane0 corresponds to lane0 in PORT2 of different OAM), the data channels in connector1 are all lane1 (each lane1 corresponds to lane1 in PORT2 of different OAM), ..., and the data channels in connector7 are all lane7 (each lane7 corresponds to lane7 in PORT2 of different OAM).
[0067] Figure 5 is another schematic diagram of a link topology structure from a data interaction port to a server external connection port group according to an exemplary embodiment. Figure 5 The embodiment shown is illustrated by using PORT4 of each OAM data exchange port, and each PORT4 of OAM0 to OAM7 is coupled to EXP4_P0, EXP4_P1, ..., EXP2_P7 of EXP4 respectively. Figure 4 In comparison, except for the data interaction port of the computing accelerator card and the data interaction expansion card, Figure 5 The topology and Figure 4 same.
[0068] Figure 4、 Figure 5 The embodiment shown only takes EXP2 and EXP4 as examples for illustration. The related link topologies of other data interaction expansion cards and computing acceleration cards are the same as those in the embodiment shown in FIG. Figure 4 、 Figure 5 The same, no further details will be given here. Figure 4 、 Figure 5 It can be seen from the embodiment shown that for any data channel connection port in any data interaction expansion card, the data channel is the data channel with the same identifier as all the data interaction ports connected to the any data interaction expansion card, for example Figure 4 As shown, the data channels in connector 1 in EXP2 are lane 1 in PORT2 of OAM0, lane 1 in PORT2 of OAM1, lane 1 in PORT2 of OAM2, lane 1 in PORT2 of OAM3, lane 1 in PORT2 of OAM4, lane 1 in PORT2 of OAM5, lane 1 in PORT2 of OAM6, and lane 1 in PORT2 of OAM7 connected to EXP2.
[0069] It can be seen that the above Figure 4 、 Figure 5 In the topological structure design, the data channel routing topology is concentrated in the data interaction expansion card, realizing the modular design of the data channel routing. Figure 4 、 Figure 5 It can be seen that although different data interaction expansion cards are connected to different data interaction ports, the topological structure of the data channel routing within the data interaction expansion card is completely consistent. Therefore, the reusability of the data interaction expansion card is enhanced, so that different data interaction expansion cards of the same structure can be coupled to different data interaction ports respectively. At the same time, after adopting the topological structure of the data channel routing within the data interaction expansion card, it helps to simplify the corresponding data channel routing topological structure in the UBB within the server, which helps to reduce the wiring difficulty of the UBB. In addition, the same data channel connection port of the data interaction expansion card in the embodiment of the present disclosure is a data channel with the same identifier, which helps to meet personalized needs such as centralized management and data interaction control of data channels with the same identifier in different data interaction ports. The server of the embodiment of the present disclosure can be applied to artificial intelligence servers in the OAM form in related technologies. Figure 6 is a schematic diagram showing the structure of another server according to an exemplary embodiment, such as Figure 6As shown, in addition to the related structures in the above embodiments, in the exemplary embodiment, the server of the embodiment of the present disclosure also includes UBB3, the computing acceleration card 1 and the data interaction expansion card 2 are plugged into the UBB3, and the UBB3 is provided with a data transmission network connecting the data interaction port and the connection port in the server. Figure 4 As shown, the data transmission network between PORT2 of OAM0 to OAM7 and EXP2_P0 to EXP2_P7 of EXP2 is arranged in UBB3. Based on this, UBB3 is provided with slots for the computing acceleration card 1 and the data interaction expansion card 2 to be plugged in.
[0070] like Figure 6 As shown, in the exemplary embodiment, the server of the present disclosure further includes a processor 4 and a PCIe switch chip 5. The PCIe switch chip 5 is coupled to the computing accelerator card 1 and to the processor 4, forming a data interaction path between the computing accelerator card 1 and the processor 4. Based on the artificial intelligence server structure of the OAM form in the related art, in the exemplary embodiment, the server of the present disclosure may include eight computing accelerator cards 1, two processors 4 and two PCIe switch chips 5. Figure 6 In the illustrated embodiment, the four computing accelerator cards 1 on the left can be coupled to the left PCIe switch chip 5, and the four computing accelerator cards 1 on the right can be coupled to the right PCIe switch chip 5. The left PCIe switch chip 5 can be coupled to the left processor 4, and the right PCIe switch chip 5 can be coupled to the right processor 4. The left and right processors 4 are coupled to each other. To distinguish by name, the four computing accelerator cards 1 on the left can correspond to OAM0 to OAM3, respectively. The left PCIe switch chip 5 can be referred to as the first PCIe switch chip. The four computing accelerator cards 1 on the right can correspond to OAM4 to OAM7, respectively. The right PCIe switch chip 5 can be referred to as the second PCIe switch chip. The left processor 4 can be referred to as the first processor, and the right processor 4 can be referred to as the second processor. Based on this, OAM0 to OAM3 can be coupled to the first PCIe switch chip, OAM4 to OAM7 can be coupled to the second PCIe switch chip, the first PCIe switch chip can be coupled to the first processor, the second PCIe switch chip can be coupled to the second processor, and the first and second processors are coupled to each other.
[0071] like Figure 6As shown, in the exemplary embodiment, the server of the present disclosure further includes a NIC card 6. The NIC card 6 is coupled to the PCIe switch chip 5 and is used to establish a data exchange connection between the server and a network switch outside the server. In the exemplary embodiment, the number of NIC cards 6 can be one or two. When one NIC card 6 is used, the NIC card 6 is simultaneously coupled to two PCIe switch chips 5 to achieve network connectivity for the server. When two NIC cards 6 are used, the two NIC cards 6 are coupled one-to-one to two PCIe switch chips 5 to achieve network connectivity for the server.
[0072] In an illustrative embodiment, the server external connection port group 22 can be opened on the end panel of the server. Depending on the environment and deployment requirements of the server, the server external connection port group 22 can be opened on the front panel, rear panel, etc. of the server. In order to improve data transmission efficiency, the data channel connection port of the server external connection port group 22 can adopt the QSFP-DD interface specification and realize direct interconnection between the computing acceleration cards 1 through the inserted optical module.
[0073] Figure 7 FIG. 1 is a structural diagram of another server according to an exemplary embodiment. Figure 7 As shown, the server mainly includes at least two computing acceleration cards 1, at least one data interaction expansion card 2 and an OCS module 200. Each computing acceleration card 1 includes at least two data interaction ports 11. In the exemplary embodiment, the structures of the various computing acceleration cards 1 are the same, and the number and structure of the data interaction ports 11 between the various computing acceleration cards 1 are the same. The data interaction expansion card 2 includes a server internal connection port 21, a server external connection port group 22 and a data interaction channel 23. The server internal connection port 21 and the server external connection port group 22 are coupled through the data interaction channel 23 and perform data interaction through the data interaction channel 23. The OCS module 200 is coupled between the server external connection port group 22 of the data interaction expansion card 2, forming a data interaction path between at least two computing acceleration cards 1. The data interaction ports 11 with the same identifier between different computing acceleration cards 1 are coupled to the server internal connection port 21 of the same data interaction expansion card 2, and the number of the server internal connection ports 21 is equal to the number of the connected data interaction ports 11.
[0074] In the server of this embodiment, a data interaction path between the computing acceleration cards 1 in the server is established using the data interaction expansion card 2 and the OCS module 200. This data interaction path can replace the direct interconnection path between the computing acceleration cards 1 in the UBB board of the server. Therefore, it is not necessary to design a corresponding fully connected direct interconnection path wiring network between the computing acceleration cards 1 in the UBB, which helps to simplify the wiring structure in the UBB, and further helps to reduce signal interference between different traces in the UBB, helps to reduce the difficulty of UBB development and the material cost of UBB production, and helps to improve the overall cost-effectiveness of the UBB.
[0075] In addition to implementing the functions of an OAM-type artificial intelligence server in the related art, the server of the disclosed embodiment can also realize data connections of computing accelerator cards in the server bypassing switches (such as PCIe switch chips) through the data interaction expansion card therein, providing a hardware foundation for one-to-one direct data interconnection of computing accelerator cards. The server using the disclosed embodiment can realize on-demand one-to-one direct data interconnection of various computing accelerator cards within the same server or between different servers by changing the connection relationship between the server external connection port groups of the data interaction expansion card. The disclosed server also realizes one-to-one full connection between computing accelerator cards within a single server. The server using the disclosed embodiment can, on the one hand, eliminate the circuit for direct interconnection between computing accelerator cards in the UBB, reducing the difficulty of UBB design. On the other hand, it can enable the server to flexibly adjust the data transmission network topology relationship of direct interconnection between various computing accelerator cards according to the characteristics of the running artificial intelligence network model. The data communication links between various computing accelerator cards can bypass the server internal switch, NIC card, and external IB network card and switch, thereby eliminating the data transmission delay caused by these nodes, which helps to improve the overall model training and inference performance of the server and server cluster.
[0076] Figure 8 is a schematic diagram of a server cluster according to an exemplary embodiment, such as Figure 8 As shown, the server cluster includes the server 100 and the OCS module 200 according to any of the above embodiments. There are at least two servers 100, and the OCS module 200 is coupled between the server external connection port groups 22 of the data interaction expansion cards 2 of at least two servers 100, forming a data interaction path between the computing acceleration cards in the at least two servers 100.
[0077] The data transmission network topology of the direct interconnection between each computing accelerator card can be flexibly adjusted according to the characteristics of the running artificial intelligence network model. In the exemplary embodiment, the server external connection port group 22 also includes a data path configuration terminal, which is used to transmit the data interaction path configuration information of the OCS module 200. That is, the data interaction path configuration information can be sent to the OCS module 200 through the data path configuration terminal, so that the OCS module 200 establishes the data interaction path between the computing accelerator cards based on the data interaction path configuration information. In this way, the data interaction path configuration information can be directly sent to the OCS module 200 through the server 100, wherein the corresponding data interaction path configuration information transmission line for the OCS module 200 can be set in the UBB of the server 100 to realize the transmission of the data interaction path configuration information from the server 100 to the OCS module 200.
[0078] In an illustrative embodiment, the data interaction path configuration information of the OCS module 200 includes configuration instructions for connecting or closing corresponding data channels in the OCS module 200. Under the control of the configuration instructions, the OCS module 200 realizes the connection of relevant data channels of relevant data interaction ports between relevant computing acceleration cards, thereby realizing one-to-one direct interconnection of computing acceleration cards.
[0079] To manage and configure the data transmission network topology for direct interconnection between computing accelerator cards, in an exemplary embodiment, the server cluster may further include a data interaction network configuration server. The data interaction network configuration server is coupled to the server 100 and configured to generate data interaction path configuration information. In an exemplary embodiment, the data interaction network configuration server may generate data interaction path configuration information based on different artificial intelligence network model tasks and send the data interaction path configuration information to the OCS module 200 via the server 100.
[0080] In an exemplary embodiment, any one of the multiple servers 100 may also serve as a data interaction network configuration server.
[0081] In an exemplary embodiment, the number of OCS modules 200 is not less than the number of servers 100 , each server 100 is coupled to at least one OCS module 200 , and different OCS modules 200 are connected via optical fibers.
[0082] Figure 9 This is a schematic diagram of the link configuration between OCS modules, such as Figure 9As shown, the OCS module 200 on the left can be configured to be connected to at least one of the two OCS modules 200 on the right through the optical fiber between the OCS modules 200. The OCS module 200 on the left can also be configured to loop back from one data channel of the OCS module 200 to another data channel of the OCS module 200, so that direct interconnection between different computing accelerator cards in the server can be achieved. Figure 4 and Figure 5 The link topology shown in the figure can realize direct interconnection between one computing accelerator card and another computing accelerator card in the same server. Figure 4 As shown, as required, the loopback configuration of the OCS module 200 connected to connector 0 can achieve direct interconnection between lane 0 of PORT 2 of OAM0 and lane 0 of PORT 2 of OAM1.
[0083] In addition, in the exemplary embodiment, based on the artificial intelligence server cluster structure in the related art, the server cluster of the embodiment of the present disclosure may also include interconnection devices such as IB network cards and switches. Each server 100 may also realize the cluster interconnection structure in the related art through its own NIC card and the IB network card and switch in the server cluster.
[0084] In the server cluster of the embodiment of the present disclosure, the data interaction expansion card 2 in each server 100 is used to realize a direct connection between the OCS module 200 and the data interaction port of the computing acceleration card in each server 100, so that the data transmission between different computing acceleration cards can bypass data forwarding devices such as network cards and switches, thereby eliminating the transmission delay caused by the network cards and switches during the data transmission process, and the delay of one-to-one direct data transmission based on the OCS module 200 is significantly lower than the delay of data transmission when the servers 100 are interconnected through network cards and switches. At the same time, through the configuration of the OCS module 200, flexible adjustment of the data link topology for different artificial intelligence network models is also realized, thereby increasing the configurable flexibility of the server cluster.
[0085] The servers and server clusters of the disclosed embodiments help to improve the data communication rate of the data links between computing acceleration cards such as OAM in artificial intelligence servers, reduce communication latency, help improve the overall computing resource utilization of artificial intelligence servers, help improve the overall computing efficiency of artificial intelligence servers and even server clusters composed of artificial intelligence servers and the degree of matching with the computing speed of the computing core devices therein, and help improve the flexibility of the data link topology structure of artificial intelligence servers and artificial intelligence server clusters in adaptive adjustment to different artificial intelligence network models.
[0086] Based on the OAM specifications in the relevant technology, a single server contains 8 computing accelerator cards 1, and each computing accelerator card 1 includes 8 data interaction ports 11. Based on this, the maximum number of data interaction expansion cards 2 can be 8. When 8 data interaction expansion cards 2 are used and all 8 data interaction expansion cards 2 are coupled to the OCS module 200, through the configuration of the OCS module 200, it is possible to use part of the data interaction ports 11 of the computing accelerator card 1 to establish full intra-machine interconnection between the 8 computing accelerator cards 1 in a single server, and to use another part of the data interaction ports 11 of the computing accelerator card 1 to establish cross-server interconnection of computing accelerator cards 1 between servers.
[0087] Figure 10 FIG. 1 is a schematic diagram of a topological structure of a fully interconnected OAM machine implemented by using the server and server cluster of the embodiment of the present disclosure, such as Figure 10 As shown in the figure, in the fully interconnected topology of the OAM machine, OAM0 to OAM7 are 8 OAMs in the same server. In order to clearly present the fully interconnected topology of the OAM machine, the structure of other parts of the server is omitted in the figure. Figure 10 As shown, each OAM from OAM0 to OAM7 has 8 data exchange ports, wherein OAM0 is directly interconnected with OAM1, OAM2, OAM3, and OAM7 through its 8 data exchange ports, OAM1 is directly interconnected with OAM0, OAM2, OAM3, and OAM6 through its 8 data exchange ports, OAM2 is directly interconnected with OAM0, OAM1, OAM3, and OAM5 through its 8 data exchange ports, and OAM3 is directly interconnected with OAM0, OAM1, OAM3, and OAM5 through its 8 data exchange ports. M1, OAM2, and OAM4 are directly interconnected. OAM4 is directly interconnected with OAM3, OAM5, OAM6, and OAM7 through its eight data interaction ports. OAM5 is directly interconnected with OAM2, OAM4, OAM6, and OAM7 through its eight data interaction ports. OAM6 is directly interconnected with OAM1, OAM4, OAM5, and OAM7 through its eight data interaction ports. OAM7 is directly interconnected with OAM0, OAM4, OAM5, and OAM6 through its eight data interaction ports.
[0088] Figure 11 FIG. 1 is a schematic diagram of a topological structure of an OAM part interconnection and part inter-machine interconnection implemented by the server and server cluster of the embodiment of the present disclosure, such as Figure 11 As shown in the figure, in this topology, OAM0 to OAM7 are 8 OAMs in the same server. In order to clearly present the topology, the structure of other parts in the server is omitted in the figure. Figure 11As shown, each OAM from OAM0 to OAM7 has 8 data exchange ports, among which OAM0 is directly interconnected with OAM1, OAM2, OAM3, OAM5, OAM6, and OAM7 through 6 of its 8 data exchange ports. The other 2 of the 8 data exchange ports of OAM0 are used as data exchange ports for interconnection between machines. OAM1 is directly interconnected with OAM0, OAM2, OAM3, OAM4, OAM6, and OAM7 through 6 of its 8 data exchange ports. The other two of the interaction ports are used as data interaction ports for interconnection between machines. OAM2 is directly interconnected with OAM0, OAM1, OAM3, OAM4, OAM5, and OAM7 through 6 of its 8 data interaction ports. The other two of the 8 data interaction ports of OAM2 are used as data interaction ports for interconnection between machines. OAM3 is directly interconnected with OAM0, OAM1, OAM2, OAM4, OAM5, and OAM6 through 6 of its 8 data interaction ports. The other two of the 8 data interaction ports of OAM3 are used as data interaction ports for interconnection between machines. OAM4 is directly interconnected with OAM1, OAM2, OAM3, OAM5, OAM6, and OAM7 through 6 of its 8 data interaction ports. The other 2 of OAM4's 8 data interaction ports are used as data interaction ports for inter-machine interconnection. OAM5 is directly interconnected with OAM0, OAM2, OAM3, OAM4, OAM6, and OAM7 through 6 of its 8 data interaction ports. The other 2 of OAM5's 8 data interaction ports are used as data interaction ports for inter-machine interconnection. OAM6 directly interconnects with OAM0, OAM1, OAM3, OAM4, OAM5, and OAM7 through six of its eight data interaction ports. The other two of OAM6's eight data interaction ports are used as data interaction ports for inter-machine interconnection. OAM7 directly interconnects with OAM0, OAM1, OAM2, OAM4, OAM5, and OAM6 through six of its eight data interaction ports. The other two of OAM7's eight data interaction ports are used as data interaction ports for inter-machine interconnection.
[0089] Figure 12The present invention is a schematic diagram of a server topology of two interconnected machines implemented by the server and server cluster of the embodiment of the present invention. In combination with the aforementioned embodiment, the first server 101 includes OAM1 to OAM7, and the second server 102 includes another OAM1 to OAM7, totaling 16 OAMs. Each OAM is equivalent to an expert, totaling 16 experts. The two-machine interconnection switches the topology structure of each OAM interconnection to partial intra-machine interconnection and partial inter-machine interconnection through the OCS module. The inter-machine interconnection method can follow the principle of interconnection between OAMs with the same identifier, for example, OAM1 in the first server 101 is interconnected with OAM1 in the second server 102. Figure 12 The server cluster can realize parallel computing of 16 experts. In the exemplary embodiment, Figure 12 The solid line between the first server 101 and the second server 102 in the figure represents a DAC (Direct Attach Copper) direct connection between the servers in the related art. The DAC direct connection method can adopt the QSFP-DD interface specification. Figure 12 The dotted line between the first server 101 and the second server 102 in FIG. 1 represents a method of connecting the first server 101 and the second server 102 through an OCS module in an embodiment of the present disclosure. Figure 12 In the illustrated embodiment, the method of connecting the first server 101 and the second server 102 through the OCS module can be coordinated with the DAC direct connection form between servers in the related technology to realize the interconnection topology of multiple computing acceleration cards, so as to help realize the hardware adaptation and flexible adjustment of servers and server clusters for various artificial intelligence models.
[0090] Figure 13 This is a schematic diagram of a server topology implemented by using the server and server cluster of the embodiment of the present disclosure to switch between two-machine interconnection and four-machine interconnection by configuring the OCS module, as shown in FIG. Figure 13 As shown, the solid line represents a two-machine interconnection topology connection, and the dotted line represents a four-machine interconnection topology connection. By configuring the various OCS modules connected to the first server 101, the second server 102, the third server 103, and the fourth server 104, two-machine interconnection (including the two-machine interconnection between the first server 101 and the second server 102 and the two-machine interconnection between the third server 103 and the fourth server 104) and the four-machine interconnection from the first server 101 to the fourth server 104 can be achieved. It can be seen that the servers and server clusters using the embodiments of the present disclosure can achieve flexible changes in the direct interconnection configuration between the various OAMs in the servers, thereby facilitating the adaptive adjustment of the server cluster data link topology to various artificial intelligence network models.
[0091] The dual-machine interconnection implemented based on the present disclosure has obvious advantages over the RDMA network interconnection in the related art. In the RDMA network interconnection scenario based on the related art, a single server only supports a maximum of 8 computing accelerator cards, so the dimension of tensor parallelism can only reach 8, that is, TP8. Because the OAMs can communicate directly through a direct link, the dimension of tensor parallelism can reach 16, that is, TP16. Because the dual-machine interconnection implemented based on the present disclosure adopts the OCS module, the communication delay between OAMs is 2 to 5us (microseconds), while the communication delay between OAMs based on the RDMA network interconnection method in the related art is 50 to 200us. Therefore, compared with the RDMA network interconnection method in the related art, the dual-machine interconnection implemented based on the present disclosure has a smaller communication delay between OAMs.
[0092] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A server, characterized in that: include: At least two computing acceleration cards, each of which includes at least two data interaction ports; At least one data interaction expansion card, the data interaction expansion card comprising a server internal connection port, a server external connection port group, and a data interaction channel, the server internal connection port and the server external connection port group being coupled via the data interaction channel, the server external connection port group being used to connect to a device external to the server; Among them, the data interaction ports with the same identification between different computing acceleration cards are coupled to the server internal connection ports of the same data interaction expansion card, wherein the number of the server internal connection ports is equal to the number of the connected data interaction ports.
2. The server according to claim 1, wherein: Each of the data interaction ports includes multiple data channels; The server internal connection port includes a plurality of data channel connection ends connected one-to-one with the plurality of data channels respectively; The server external connection port group includes multiple data channel connection ports, the number of the data channel connection ports is the same as the number of the data channels, wherein the identifiers of the data channels in any one of the data channel connection ports are the same, and the identifiers of the data channels between different data channel connection ports are different.
3. The server according to claim 1, wherein: The server includes UBB; The computing acceleration card and the data interaction expansion card are plugged into the UBB; The UBB is provided with a data transmission network connecting the data interaction port and the internal connection port of the server.
4. The server according to claim 1, wherein: The server further includes: processor; and, A PCIe switching chip is coupled to the computing acceleration card and to the processor, forming a data interaction path between the computing acceleration card and the processor.
5. The server according to claim 4, wherein: The server further includes: A NIC card is coupled to the PCIe switch chip, and is used to establish a data interaction connection between the server and a network switch outside the server.
6. The server according to claim 1, wherein: The number of the data interaction expansion cards does not exceed the number of the data interaction ports in a single computing acceleration card.
7. A server, characterized in that: include: At least two computing acceleration cards, each of which includes at least two data interaction ports; at least one data interaction expansion card, the data interaction expansion card comprising a server internal connection port, a server external connection port group, and a data interaction channel, the server internal connection port and the server external connection port group being coupled via the data interaction channel; and An OCS module, the OCS module being coupled between the server external connection port group of the data interaction expansion card and constituting a data interaction path between the at least two computing acceleration cards; Among them, the data interaction ports with the same identification between different computing acceleration cards are coupled to the server internal connection ports of the same data interaction expansion card, wherein the number of the server internal connection ports is equal to the number of the connected data interaction ports.
8. A server cluster, characterized in that: include: The server according to any one of claims 1 to 6, wherein the number of the servers is at least two; An OCS module is coupled between the server external connection port groups of the data interaction expansion cards of at least two of the servers, forming a data interaction path between the computing acceleration cards in at least two of the servers.
9. The server cluster according to claim 8, wherein: The server external connection port group further includes a data path configuration end, and the data path configuration end is used to transmit data interaction path configuration information of the OCS module.
10. The server cluster according to claim 9, wherein: The server cluster further includes: The data interaction network configuration server is coupled to the server and is used to generate the data interaction path configuration information.
11. The server cluster according to claim 8, wherein: The number of the OCS modules is not less than the number of the servers. Each server is coupled to at least one OCS module, and different OCS modules are connected via optical fibers.
Citation Information
Cited By
Computer system, server optimization method, electronic device and storage medium
CN121217598A