A data switching chip and server

By introducing an AI Switch chip into the server, direct data transmission is achieved, solving the problem of data transmission latency in multi-server parallel processing, improving data transmission efficiency and bandwidth, and enhancing the efficiency of multi-server parallel processing.

CN112148663BActive Publication Date: 2026-03-10HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

When multiple servers process data in parallel, the data transmission latency between servers is relatively long, which affects the efficiency of parallel data processing.

Method used

The system employs an AI Switch chip, which includes a controller, an AI interface, a PCIe interface, and a network interface. The controller connects directly to the AI ​​chip and the processor, enabling direct data transmission through the AI ​​interface and the network interface, thus avoiding the need for other chips or modules and improving data transmission efficiency.

Benefits of technology

It reduces data transmission latency between servers, improves data transmission efficiency and bandwidth, and enhances the efficiency of multi-server parallel processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112148663B_ABST
    Figure CN112148663B_ABST
Patent Text Reader

Abstract

This application provides an artificial intelligence (AI) switching chip and server. The AI ​​switching chip includes a first AI interface, a first network interface, and a controller. The first AI interface is used for the AI ​​switching chip to connect to a first AI chip in a first server, where the first AI chip is any one of multiple AI chips in the first server. The first network interface is used for the AI ​​switching chip to connect to a second server. The controller receives data sent by the first AI chip through the first AI interface and then sends the data to the second server through the first network interface. Through the AI ​​switching chip, when a server sends data from its AI chip to another server, it can directly receive the data sent by the AI ​​chip through the AI ​​interface and then send it to the other server through one or more network interfaces connected to the controller. This reduces the latency and increases the efficiency of data transmission from the AI ​​chip in the server to other servers.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of servers, and in particular to a data exchange chip and a server. BACKGROUND

[0002] With the development of computer technology, artificial intelligence (AI) and big data are widely used in more and more fields such as machine vision and natural language processing. Artificial intelligence is the simulation of human consciousness and thinking process, and is usually implemented through a neural network model. The training of the neural network model often uses an AI chip. The AI chip is a module used to process a large amount of computing tasks in artificial intelligence applications. There can be one or more AI chips in a server. With the increasing demand of application scenarios, the size of the neural network and the size of the data set grow rapidly. Training a neural network model through a single server requires a long time. In order to adapt to the growth of the size of the neural network and the data, and shorten the time of training the neural network model, it is usually necessary to use AI chips in multiple servers to perform parallel processing of data. For example, different servers are responsible for training different network layers of the neural network model, or different data of the same layer of the network are distributed to AI chips in different servers for training, and then the calculation results of all servers are combined in a certain way. However, when multiple servers perform parallel processing of data, a large amount of data needs to be transmitted between different servers. How to reduce the time delay when transmitting data between servers is a technical problem to be solved to improve the efficiency of data parallel processing. SUMMARY

[0003] The present application provides a chip for data exchange and a server for improving the efficiency of data transmission between servers.

[0004] In a first aspect, the present application provides an artificial intelligence exchange AI Switch chip, comprising: a first AI interface, a first network interface, and a controller, wherein,

[0005] The first AI interface is used for the AI Switch chip to connect a first AI chip in a first server through the first AI interface. The first server includes the AI Switch chip and a plurality of AI chips. The first AI chip is any one of the plurality of AI chips.

[0006] The first network interface is used for the AI Switch chip to connect a second server through the first network interface.

[0007] The controller is connected to the first AI interface and the first network interface, respectively, and is used to receive first data sent by the first AI chip through the first AI interface, and then send the first data to the second server through the first network interface.

[0008] Through the AI Switch chip, when the server needs to send data in the AI chip to another server, the controller in the AI Switch chip does not need to pass through other chips or modules, and can directly receive the data sent by the AI chip through the AI interface connected with the controller, and then send the data to another server through the network interface connected with the controller, so that the controller in the server has smaller time delay when receiving data from the AI chip and has higher efficiency.

[0009] In a possible implementation, the AI Switch chip further includes a peripheral bus interface standard (PCIe) interface, and the PCIe interface is connected with the controller and the processor in the first server respectively.

[0010] Before receiving the first data sent by the first AI chip through the first AI interface, the controller is further configured to receive control information sent by the processor through the PCIe interface, and the control information carries an identifier of the first AI chip; when receiving the first data sent by the first AI chip through the first AI interface, the controller is specifically configured to receive the first data sent by the first AI chip through the first AI interface according to the identifier of the first AI chip.

[0011] In a possible implementation, the AI Switch chip further includes a second network interface connected with the second server, and the controller is further configured to receive second data sent by the first AI chip through the first AI interface, and send the second data to the second server through the second network interface. The AI Switch chip includes a plurality of network interfaces, and the controller in the AI Switch chip can send data in the AI chip to other servers through the plurality of network interfaces after receiving the data, so that the bandwidth of data transmission between servers can be improved, and the time delay of data transmission between servers can be reduced.

[0012] In a possible implementation, the first AI interface in the AI Switch chip is connected with the processor through the first AI chip, and the controller is further configured to receive third data sent by the processor through the first AI interface and the first AI chip, and send the third data to the second server through the first network interface. The AI Switch chip can receive data sent by the processor through the PCIe interface or the first AI interface and the first AI chip, so that the AI Switch chip can be connected with the processor through two paths and receive data in the processor through the two paths.

[0013] In a possible implementation, the controller is further configured to: receive fourth data sent by the first AI chip through the first AI interface, and then send the fourth data to a second AI chip in the plurality of AI chips in the server through the second AI interface. The controller receives data sent by one AI chip through an AI interface, and then sends the data to another AI chip through another AI interface, so as to realize data interaction between the AI chips in the server.

[0014] In a possible implementation, the first AI interface is an HSSI interface.

[0015] In a possible implementation, the first AI interface is an HSSI interface.

[0015] In a second aspect, the present application provides a server, comprising a processor, a plurality of artificial intelligence (AI) chips, and an AI switch chip. The AI switch chip is connected to the processor through a first peripheral component interconnect express (PCIe) interface in the AI switch chip, and is connected to the plurality of AI chips through a plurality of AI interfaces in the AI switch chip. The AI switch chip is configured to: receive control information sent by the processor through the PCIe interface, the control information comprising an identifier of a first AI chip, the first AI chip being any one of the plurality of AI chips; receive first data sent by the first AI chip through a first AI interface in the AI switch chip, and send the first data to another server through a first network interface in the AI switch chip.

[0016] In a possible implementation, the AI switch chip further comprises a second network interface, and the AI switch chip is connected to another server through the second network interface. The AI switch chip is further configured to: receive second data sent by the first AI chip through the first AI interface; and then send the second data to the another server through the second network interface.

[0017] In a possible implementation, the first AI interface is connected to the processor through the first AI chip. The AI switch chip is further configured to: receive third data sent by the processor through the first AI interface and the first AI chip; and then send the third data to the another server through the first network interface.

[0018] In a possible implementation, the AI Switch chip is connected to the processor through a PCIe interface, and before the AI Switch chip receives the third data sent by the processor through the first AI interface and the first AI chip, the processor needs to determine whether a path connecting the processor through the PCIe interface has data congestion. The controller can be connected to the processor through the PCIe interface, or be connected to the processor through the first AI interface and the first AI chip. When the AI Switch chip needs to receive data in the processor, the AI Switch chip can receive the data sent by the processor through any one of the paths, or receive the data sent by the processor through both of the paths at the same time. When the controller receives the data sent by the processor through any one of the paths, the processor can obtain the congestion conditions of the two data paths, and select the data path with lighter congestion to receive the data sent by the processor, so as to reduce the time delay of receiving data.

[0019] In a possible implementation, the AI Switch chip can also receive fourth data in the first AI chip through the second AI interface, and send the fourth data to a second AI chip in the plurality of AI chips. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0021] Figure 1A is a topological structure schematic diagram of a plurality of servers connected to each other provided by an embodiment of the present application;

[0022] Figure 1B is another topological structure schematic diagram of a plurality of servers connected to each other provided by an embodiment of the present application;

[0023] Figure 2 is a structure schematic diagram of a server with a plurality of AI chips provided by an embodiment of the present application;

[0024] Figure 3 is a structure schematic diagram of a server provided by an embodiment of the present application;

[0025] Figure 4 is a specific structure schematic diagram of a server provided by an embodiment of the present application,

[0026] Figure 5 is a schematic diagram of an AI chip connected to an AI Switch chip provided by an embodiment of the present application;

[0027] Figure 6 is a structural schematic diagram of an AI switch chip provided by an embodiment of the present application;

[0028] Figure 7 is a connection schematic diagram of an AI chip and an AI switch chip provided by an embodiment of the present application;

[0029] Figure 8 is a working schematic diagram of a server provided by an embodiment of the present application;

[0030] Figure 9 is a data transmission method provided by an embodiment of the present application;

[0031] Figure 10 is a structural schematic diagram of another multi-AI chip server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0032] The present application will be specifically described below with reference to the accompanying drawings. First, special terms involved in the present application are introduced:

[0033] Artificial intelligence (AI) chip: a module used for processing a large amount of computing tasks in artificial intelligence applications, and one or more AI chips can be included in a server.

[0034] Network interface controller (NIC): also known as a network card, the NIC is computer hardware designed to allow computers to communicate on a network, and the NIC of a server is used to connect one server to another server, or to establish a connection between the server and a network device such as a switch.

[0035] Peripheral component interface express (PCIe) interface: a high-speed serial computer expansion bus standard interface, the PCIe interface is used for high-speed serial point-to-point double-channel high-bandwidth transmission, the devices connected by the PCIe interface are allocated exclusive channel bandwidth and do not share bus bandwidth, and mainly support end-to-end reliable transmission.

[0036] Peripheral Component Interface Express Switch (PCIe Switch) chip: The PCIe Switch chip is a module for expanding the PCIe link, and the PCIe link uses an end-to-end connection mode, so that only one device or component can be connected at each end of the PCIe link. Therefore, the PCIe link must be expanded using the PCIe Switch chip, so that multiple devices or components can be connected at one end of the PCIe link. The PCIe Switch chip is connected to other devices or components through a PCIe bus.

[0037] High-speed serial interface (HSSI) interface: an expansion interface using serial communication mode, including universal serial bus (USB), high-definition multimedia interface (HDMI), mobile industry processor interface (MIPI), etc.

[0038] In the field of artificial intelligence, with the rapid growth of the scale of neural networks and the scale of data sets, it is difficult for the computing power of one or more AI chips within a single server to train a large-scale neural network using a large-scale training set. Therefore, multiple servers (containing more AI chips) are needed to perform parallel processing of data. For example, a model parallel training method is used to distribute different network layers of a neural network model to different servers for training. After a single server processes the data, it needs to send the processed data to other servers for use by other servers for training.

[0039] When multiple servers are used for parallel processing of data, the servers can be directly connected (i.e., data transmission between two servers does not pass through other devices), or they can be connected through a router or switch (i.e., data transmission between two servers is forwarded through a router or switch). When multiple servers are directly connected, any two servers can be directly connected through one or more network interface controllers (e.g., the NICs of any two servers are connected by a network cable). When data needs to be transmitted between any two servers, the bandwidth available for data transmission between the two servers is the sum of the bandwidths of the NICs connected between the two servers. When multiple servers are connected through a router or switch, the maximum bandwidth available for data transmission between any two servers can be the sum of the bandwidths of the NICs possessed by each server.

[0040] For example, if each server has 8 NICs, and three servers are directly connected to each other through full interconnection, the network topology is as shown in Figure 1A , any two servers can be connected to each other through 4 NICs, and the maximum bandwidth for data transmission between any two servers is the sum of the bandwidths of the 4 NICs. When multiple servers are connected to each other through routers or switches, the maximum bandwidth for data transmission between any two servers can be the sum of the bandwidths of the NICs of each server, as shown in Figure 1B , Figure 1B is a schematic diagram of three servers connected through routers or switches. When server A and server B need to transmit data, and neither server A nor server B has data transmission with server C, server A and server B can transmit data through 8 NICs, i.e., the maximum bandwidth for data transmission between any two servers is the sum of the bandwidths of the 8 NICs. During data transmission between a server and another server, the key factor affecting data transmission efficiency is the path of internal data transmission of the server and the bandwidth for data transmission between the two servers.

[0041] As shown in Figure 2 , Figure 2 is a schematic diagram of a server containing multiple AI chips. This type of server includes 8 NICs and multiple artificial intelligence switching AI Switch chips. Each NIC is connected to two AI chips through a PCIe Switch chip and is connected to a processor through a PCIe Switch chip. The AI Switch chip is used to connect multiple AI chips within the server, enabling data exchange between any two AI chips within the server through the AI Switch chip. When a server needs to send data in an AI chip to another server, the NIC receives the data sent by the AI chip through the PCIe Switch chip and then sends it to another server connected to the NIC. For example, Figure 2When the server shown in the figure is used to cooperatively train a neural network model with other servers, if a server needs to send data in an AI chip to other servers, since one NIC connects two AI chips through one PCIe Switch chip, one NIC needs to receive data sent by one or two AI chips through one PCIe Switch chip. The data in the two AI chips needs to be transmitted to the same NIC through the same PCIe Switch chip, so the data paths of the two AI chips have a coincident part (i.e. the PCIe bus connected between the PCIe Switch chip and the NIC), wherein the data path of the AI chip refers to the path for transmitting data between the AI chip and the NIC. Since the bandwidth of the PCIe bus is small, when the amount of data to be transmitted is large, congestion may occur in the path from the PCIe Switch chip to the NIC, which increases the delay in transmitting data between servers. Further, the NIC receives data of one AI chip through the PCIe Switch chip or the NIC sends data to one AI chip through the PCIe Switch chip, which also increases the delay in transmitting data between servers. Figure 2 When the amount of data to be transmitted is large in the server shown in the figure, congestion may occur in the path from the PCIe Switch chip to the NIC, which increases the delay in transmitting data between servers. Further, the NIC receives data of one AI chip through the PCIe Switch chip or the NIC sends data to one AI chip through the PCIe Switch chip, which also increases the delay in transmitting data between servers.

[0042] Under the trend of rapid growth of the scale of neural networks and the scale of data sets, when multiple servers are used for parallel processing of data, data transmission between servers will become frequent. If the server shown in the figure is used, one NIC needs to receive data in one or more AI chips through the PCIe Switch chip and transmit the data to another server, which will cause the delay in transmitting data between servers to be long. If the delay in transmitting data between servers is long, the efficiency of parallel processing of multiple servers will be reduced, and the advantage of parallel processing cannot be fully utilized. Figure 2

[0043] To solve the problem that the delay in transmitting data between servers is long due to slow transmission speed when multiple servers are used for parallel processing of data, an embodiment of the present application provides a server, which is described below in combination with the drawings. As shown in Figure 3 Figure 3 is a structural schematic diagram of a server provided by an embodiment of the present application, which includes a processor, an AI chip, an AI Switch chip and a PCIe Switch chip. The processor and the AI chip and the processor and the AI Switch chip are connected through the PCIe Switch chip. The AI chip and the AI Switch chip are connected through the AI interface in the AI chip and the AI Switch chip. Exemplarily, Figure 4 ​​is a specific structural schematic diagram of the above-mentioned server provided by the present application, which comprises a processor, a plurality of AI chips, one or more AI Switch chips and a plurality of PCIe Switch chips. The processor can be connected with the plurality of AI chips through one PCIe Switch chip and connected with the plurality of AI Switch chips through the PCIe Switch chip, the plurality of AI chips and the plurality of AI Switch chips are connected with each other through AI interfaces, and the connection mode of the plurality of AI chips and the plurality of AI Switch chips is as shown in Figure 5 each AI chip and each AI Switch chip comprises a plurality of AI interfaces, each AI chip can be connected with one AI interface in one AI Switch chip through one AI interface, or can be connected with a plurality of AI interfaces in one AI Switch chip through a plurality of AI interfaces, so that each two AI chips in the plurality of AI chips in the server can be connected with each other through the AI Switch chip. Exemplarily, Figure 5 is a connection schematic diagram between any two AI chips in the server and the AI interfaces in any three AI Switch chips in the plurality of AI Switch chips, and each AI chip in the diagram is connected with one AI interface in one AI Switch chip through one AI interface. Alternatively, the above-mentioned AI interface can be an HSSI interface, or other types of interfaces for connecting chips, which are not limited in the present application.

[0044] Figure 6Fig. 1 is a schematic diagram of an AI Switch chip according to an embodiment of the present application. The AI Switch chip includes a controller, one or more AI interfaces, one or more PCIe interfaces, and one or more network interfaces. The controller is connected to the AI interfaces, the PCIe interfaces, and the network interfaces, respectively, receives data or control information through the AI interfaces, the PCIe interfaces, or the network interfaces, and sends the received data to an AI chip or another server through a designated interface. The AI interfaces are used for the AI Switch chip to connect to multiple AI chips in a server. Each AI interface on the AI Switch chip is connected to an AI interface on an AI chip. The controller is connected to the AI interfaces, receives data in the AI chip through the AI interfaces, and sends the received data to an AI chip connected to the other AI interfaces, thereby achieving data transmission between AI chips in the server, or sends the received data to another server through the network interfaces, thereby achieving data transmission between servers. The PCIe interfaces are used for the AI Switch chip to connect to a processor through a PCIe Switch chip connected to the PCIe interfaces. The network interfaces are used for the server to connect to another server, or to connect to a switch or a router, thereby achieving data transmission between servers. In the case shown in Fig. 1, the connection mode of the AI chip and the AI Switch chip is shown in Fig. 2, i.e., multiple AI Switch chips are connected to multiple AI chips through different AI interfaces, so that one AI Switch chip can be connected to multiple AI chips, and multiple AI chips can exchange data within a server or between servers; multiple AI Switch chips can also be connected to one AI chip, so that multiple AI Switch chips can simultaneously exchange data within a server or between servers for the same AI chip. Figure 6 Figure 7

[0045] Figure 8 ​​A working schematic diagram of a server is provided in the embodiments of the present application. The AI Switch chip can be connected to the processor through the PCIe Switch chip, or connected to the processor through an AI chip and a PCIe Switch chip. For the AI chip, the controller is connected to the processor through the PCIe interface and the path (i.e., path 1 in the figure) of the PCIe Switch chip, which is referred to as a control path. The processor sends control information to the controller in the AI Switch chip through the control path, and the controller receives the data sent by the AI chip according to the control information. The controller is connected to the AI chip through the AI interface (i.e., path 2 in the figure), which is referred to as a data path. Each AI chip can be connected to the controller in an AI Switch chip through an AI interface to form a data path. Therefore, each AI chip can be connected to the controllers in multiple AI Switch chips through multiple AI interfaces to establish multiple data paths. One or more controllers that receive the control information can receive the data sent by the AI chip through the data path. It can be understood that the AI Switch chip includes multiple AI interfaces and multiple PCIe interfaces, Figure 8 only one AI interface in the AI Switch chip is connected to one AI chip, and only one PCIe interface in the AI Switch chip is connected to one PCIe Switch chip.

[0046] Based on the above server, the embodiments of the present application provide a data transmission method. The method is applied to a server system that uses multiple servers for parallel processing of data, as shown in Figure 9 The method includes the following steps.

[0047] In S102, the AI Switch chip receives the control information sent by the processor.

[0048] The above control information is used to instruct the AI Switch chip to receive the data in the target AI chip through the controller. The control information includes the identification of the target AI chip, such as the ID of the target AI chip and the interface number of the target AI chip. The ID of the target AI chip is the AI chip that the AI Switch chip needs to receive data from, and the interface number of the target AI chip is the AI interface connected by the target AI chip. For example, when training a neural network model using a model parallel method, one server is responsible for training one network layer. When a server receives data that needs to be processed, the processor of the server allocates the received data to multiple AI chips in the server for processing. When an AI chip finishes processing the allocated data, the processor sends control information to the controller in one or more AI Switch chips through the control path to instruct the controller that receives the control information to receive the data sent by the target AI chip through the AI interface.

[0049] S104, the AI Switch chip receives the data sent by the target AI chip according to the control information.

[0050] Since the controller is located in the internal of the AI Switch chip, the controller is connected with the AI chip through the AI interface in the AI Switch chip, so the controller can receive the data sent by the target AI chip through the AI interface connected with the target AI chip. After receiving the above control information, the controller needs to obtain the ID of the target AI chip and the port number of the target AI chip in the control information, determine the target AI chip according to the ID of the target AI chip in the control information, and determine the AI interface connected with the target AI chip according to the port number of the target AI chip, and then receive the data sent by the target AI chip.

[0051] S106, the AI Switch chip sends the received data to the target server.

[0052] After the AI Switch chip receives the corresponding data in the target AI chip according to the above control information, the AI Switch chip sends the received data to the target server through the network interface connected with the network interface.

[0053] It is worth noting that in the above data transmission method, the controller can directly receive the data in the AI chip through the AI interface, as shown in Figure 7 Since the controller is connected with the AI chip through the AI interface, the controller can directly receive the data in the AI chip through the AI interface. That is, in the above S104, the controller does not need to pass through other chips or modules, and can directly receive the data sent by the AI chip through the AI interface. And the controller is connected with multiple network interfaces, the data received by the controller from the AI chip can be transmitted to the target server through the multiple network interfaces, so the controller in the AI Switch chip provided by the present application has smaller time delay when receiving data from the AI chip and transmitting data, and has higher efficiency.

[0054] It is worth noting that in the server provided by the present application, when the data amount to be transmitted by one of the AI chips is large, the controllers in multiple AI Switch chips can be used to simultaneously receive the data sent by the AI chip, and the data can be sent to another server through the network interfaces in the multiple AI Switch chips, so as to reduce the time delay of data transmission. As shown in Figure 8As shown, each AI chip is connected with each AI Switch chip through an AI interface, and each AI chip can be connected with the controller in each AI Switch chip through the AI interface in the AI Switch chip, thereby forming a data path with the controller in each AI Switch chip. When data sent by a certain AI chip needs to be received, the data in the AI chip can be received by the controllers in multiple AI Switch chips connected with the AI chip at the same time, that is, the processor can send control information to the controllers in multiple AI Switch chips to receive the data in the target AI chip through the multiple controllers, and then send the received data to another server through the network interfaces in the multiple AI Switch chips. For example, in the above Figure 8 AI chip is connected with each of the three AI Switch chips through an AI interface, and when data in one of the AI chips needs to be sent to another server, a data path can be formed between the AI chip and each of the three AI Switch chips connected with the AI chip, and then the data in the AI chip is received through the three data paths, and then the received data is sent to another server through the network interfaces in the three AI Switch chips. Therefore, the server provided in the present application can simultaneously receive data in the same AI chip through multiple controllers, and simultaneously send data to another server through the network interfaces in multiple AI Switch chips. The server receiving the data can also receive the data through multiple network interfaces. When the same amount of data in the AI chip needs to be sent, the server provided in the present application can provide greater bandwidth, faster speed, higher efficiency and smaller time delay for receiving and sending data in the AI chip.

[0055] Further, in the server provided in the present application, when the controller needs to send data in all AI chips to another server, the data in each AI chip can be transmitted to another server through any one or more network interfaces in the AI Switch chip. Figure 2 In the server shown, one NIC needs to receive data in two AI chips, and the data in the two AI chips needs to be transmitted to the NIC through the same PCIe bus. Since the bandwidth of the PCIe bus is limited, the data in the two AI chips cannot be transmitted to the NIC through the same PCIe bus at the same time. Therefore, the data in the two AI chips needs to be transmitted to the NIC through the same PCIe bus at different times, and the data in the two AI chips needs to be transmitted to the NIC through the same PCIe bus at different times. Figure 2In the server shown, when the amount of data to be transmitted is large, congestion may occur on the path from the PCIe Switch chip to the NIC, leading to increased latency in data transmission between servers. In the server provided in this application, each AI Switch chip includes one or more network interfaces, and the number of network interfaces can exceed the number of AI chips. Therefore, when data from all AI chips needs to be sent to another server, the controller in each AI Switch chip can receive data from one or more AI chips through the AI ​​interface, and then send the received data out through different network interfaces. Each AI chip has at least one corresponding network interface to send its data, thus avoiding congestion. Figure 2 In servers, because a single NIC needs to transmit data from multiple AI chips, when the data volume is too large, channel congestion occurs, leading to high latency in data transmission between servers.

[0056] For example, if the server provided in this application includes 2 AI chips and 3 AI Switch chips, and each AI Switch chip includes 2 network interfaces. According to... Figure 8 The diagram illustrates the connection relationships between the AI ​​chips and the AI ​​interfaces in the AI ​​Switch chip, as well as the internal connections of the AI ​​Switch chip. All six network interfaces in the server can connect to each AI chip and transmit data from those AI chips. When each AI chip connects to each AI Switch chip via only one AI interface, and it's necessary to receive data from both AI chips simultaneously, the controller in each AI Switch chip can receive data from both AI chips through the AI ​​interface connected to each AI chip. Then, one network interface is used to transmit data from one AI chip, and another network interface transmits data from the other AI chip. That is, when it's necessary to send data from all AI chips to another server, each AI chip can have at least one corresponding network interface for transmitting the data received from that AI chip by the controller.

[0057] It is worth noting that the server provided in this application allows for the configuration of the number of network interfaces according to actual needs, adapting to the bandwidth required for transmitting data from the AI ​​chip. Since the controller is located inside the AI ​​Switch chip, it can directly access the data in the AI ​​chip through the AI ​​interface on the AI ​​Switch chip. When configuring the network interfaces in this server, only one network interface needs to be configured in each AI Switch chip; no additional PCIe Switch chip connected to the AI ​​chip is required. Furthermore, the number of interfaces connecting each AI chip to the AI ​​interface in the AI ​​Switch chip can remain unchanged; only the newly added network interfaces and the interface connection cables between the network interfaces and the controller need to be added. It is understood that the above... Figure 2 In the server shown, each AI chip can connect to multiple NICs by increasing the number of NICs. However, adding one NIC requires one AI chip's interface, and each AI chip has a limited number of interfaces. Furthermore, both the processor and AI chips connect to the NICs via PCIe switch chips, and each PCIe switch chip has a limited number of interfaces. Therefore, increasing the number of NICs also requires increasing the number of PCIe switch chips. Figure 2 Adding a NIC to the server structure shown would limit the number of NICs that can be added due to the limited number of AI chip interfaces, and would also result in an excessive number of internal chips, leading to structural complexity.

[0058] The aforementioned controller can also be used to receive data sent by the processor. When receiving data from the processor, it can simultaneously receive data through two data paths or select the less congested data path from the two data paths. For example... Figure 7 As shown, each controller can connect to the processor solely through a PCIe switch chip, while each controller connects to an AI chip via an AI interface. The AI ​​chip connects to the processor via a PCIe switch chip. Therefore, the controller can also connect to the processor via both the AI ​​chip and the PCIe switch chip. Thus, the controller can connect via... Figure 7 Paths 1 and 3 are connected to the processor. Path 1 can serve as a control path, through which the processor sends control information to the controller, as described above, the controller sends control information to the processor via path 1. Path 1 can also serve as a data path, in which case the data path between the controller and the processor includes... Figure 7The controller can receive the data sent by the processor through any one of the above-mentioned data paths 1 or 3, or simultaneously receive the data sent by the processor through the above-mentioned two data paths. When the controller receives the data sent by the processor through any one of the above-mentioned data paths, the processor can obtain the congestion conditions of the two data paths and select the data path with a lighter congestion degree to send data to the controller. The processor can receive the data sent to the controller through the two data paths or select the data path with a lighter congestion degree from the two data paths, thereby reducing the time delay of the processor sending data.

[0059] When the processor needs to send data in the processor memory or the AI chip to another server, the network interface is integrated in the AI Switch chip. After receiving the control information, the controller can directly receive the data in the AI chip through the interface between the AI Switch chip and the AI chip, thereby reducing the time for the AI Switch chip to receive data in the processor or the AI chip and reducing the time for data transmission between servers. Further, when the server needs to send data in an AI chip, the server can receive data in the same AI chip through the controllers in multiple AI Switch chips and send the data in the AI chip through the network interfaces in the multiple controllers. When all the data in the AI chips needs to be sent, at least one network interface can be provided for each AI chip to receive and send data in an AI chip, thereby reducing the time for data transmission between devices when transmitting the same amount of data and improving the efficiency of data reading and transmission. In addition, when data in the processor needs to be sent, the processor can send data to the AI Switch chip through two data paths or select a data path with a lighter congestion degree from the two data paths to send data to the AI Switch chip, thereby reducing the time for the processor to send data and reducing the time delay.

[0060] The following takes a server including 16 AI chips and 12 AI Switch chips as an example to analyze the structure of the server and the bandwidth or time delay during data transmission. As shown in FIG. 1, Figure 10 Figure 10 ​is a schematic diagram of a server provided by an embodiment of the present application, which includes two identical substrates, each of which includes 8 AI chips and 6 AI Switch chips, if each AI Switch chip includes 4 network interfaces, then the server provided in the embodiment of the present application includes a total of 48 network interfaces. Each AI chip has 6 AI interfaces, each AI Switch chip has 18 AI interfaces, each AI interface on each AI chip is connected to one AI Switch chip, and the two substrates are connected through the AI interfaces between the AI Switch chips, so each AI chip can be connected to 12 AI Switch chips. Since each AI chip is connected to each AI Switch chip through an AI interface, when data in an AI chip needs to be sent, it can be received by the controllers in any 6 AI Switch chips, wherein each AI Switch chip provides 1 network interface to send the received data in the AI chip. When data in 16 AI chips needs to be sent, each AI chip can receive data in the AI chip through the controllers in 3 AI Switches, and send data through one network interface in each AI Switch chip. If three servers described above are connected in a full interconnection manner as shown in Figure 1A , any two servers can be connected to each other through 24 network interfaces, and the maximum bandwidth when data is transmitted between any two servers is the sum of the bandwidths of the 24 network interfaces. When three servers are connected to each other through a router or a switch network as shown in Figure 1B , if server A and server B need to transmit data, and server A and server B do not have data transmission with server C, then server A and server B can transmit data through 48 network interfaces, that is, the maximum bandwidth when data is transmitted between any two servers can be the sum of the bandwidths of the 48 network interfaces.

[0061] It should be understood that the structure of the chip or server described in the above embodiments is only exemplary, and the functions of the chip or server described above can also be implemented in other ways. For example, the structure of the AI Switch chip described above is only a logical division of functions, and actual implementation can also have another division method, for example, receiving data in the AI chip through a network card and sending the data to other servers, implementing data exchange between AI chips in the server through a crossbar matrix, etc. In addition, the functional chips in the embodiments of the present application can be inherited on one module, or each chip can exist separately.

[0062] The above description of the embodiments of the present application is a description of the principles of the chip and the server provided in the present application using specific examples. The above embodiments are only used to help understand the core idea of the present application. For those skilled in the art, the specific implementation and application range will be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An artificial intelligence switch (AI Switch) chip, comprising: Comprising: a first AI interface connected to a first AI chip in a first server, the first server comprising the AI Switch chip and a plurality of AI chips, the first AI chip being any one of the plurality of AI chips; a first network interface connected to a second server; a second network interface connected to the second server; a controller connected to the first AI interface, the first network interface, and the second network interface respectively, and configured to: receive first data sent by the first AI chip through the first AI interface; send the first data to the second server through the first network interface; receive second data sent by the first AI chip through the first AI interface; send the second data to the second server through the second network interface.

2. The Al Switch chip of claim 1, wherein, The AI Switch chip further comprises: a peripheral component interconnect express (PCIe) interface connected to the controller and a processor in the first server respectively; The controller is further configured to receive control information sent by the processor through the PCIe interface, the control information carrying an identifier of the first AI chip; The controller is configured to receive the first data sent by the first AI chip through the first AI interface according to the identifier of the first AI chip.

3. The Al Switch chip of claim 1, wherein, The AI Switch chip is further connected to a processor through the first AI chip, and the controller is further configured to: receive third data sent by the processor through the first AI interface and the first AI chip; send the third data to the second server through the first network interface.

4. The Al Switch chip of any of claims 1-3, wherein, The controller is further configured to: receive fourth data sent by the first AI chip through the first AI interface; send the fourth data to a second AI chip in the plurality of AI chips through a second AI interface in the AI Switch chip.

5. The Al Switch chip of claim 1, wherein, The first AI interface is an HSSI interface.

6. A server, characterized by Comprising: a processor configured to send control information; a plurality of artificial intelligence (AI) chips; an AI Switch chip connected to the processor through a peripheral component interconnect express (PCIe) interface in the AI Switch chip, and connected to a plurality of AI chips through a plurality of AI interfaces in the AI Switch chip; The AI Switch chip is configured to: receive the control information sent by the processor through the PCIe interface, the control information comprising an identifier of a first AI chip, the first AI chip being any one of the plurality of AI chips; receive first data sent by the first AI chip through a first AI interface in the AI Switch chip; send the first data to another server through a first network interface in the AI Switch chip; The AI Switch chip further comprises a second network interface connected to the another server; The AI Switch chip is further configured to: receive second data sent by the first AI chip through the first AI interface; send the second data to the other server through the second network interface.

7. The server of claim 6, wherein, The AI Switch chip is further connected with the processor through the first AI chip, and the AI Switch chip is further configured to: receive third data sent by the processor through the first AI chip through the first AI interface; send the third data to the other server through the first network interface.

8. The server of claim 7, wherein, The processor is further configured to: determine that a path connected with the processor through the PCIe interface exists data congestion; and send the third data to the AI Switch chip through the first AI chip.

9. The server of any one of claims 6-8, wherein: The AI Switch chip is further configured to receive fourth data in the first AI chip through a second AI interface and send the fourth data to a second AI chip in the plurality of AI chips.

Citation Information

Patent Citations

  • A computing cluster and a computing cluster configuration method

    CN109739802A