Communication method, device, storage medium, and program product

By installing a universal interface in a heterogeneous GPU cluster, direct communication between heterogeneous GPUs is achieved, solving the problem of low communication efficiency in heterogeneous GPU clusters and improving communication efficiency and system performance.

CN119226193BActive Publication Date: 2025-11-18SHANGHAI BIREN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411719315.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-11-18
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

In heterogeneous GPU clusters, the communication efficiency between heterogeneous GPUs is low, which leads to increased communication latency and affects training efficiency.

Method used

By installing a universal interface on the artificial intelligence chip, data transmission interfaces compatible with different manufacturers and models are integrated to achieve direct communication, avoiding multiple data copies. Memory registration and copy operations are used to optimize the communication path.

Benefits of technology

It improves communication efficiency between heterogeneous GPUs, reduces communication latency, enhances the overall performance and compatibility of computing systems, and supports future technological development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226193B_ABST
    Figure CN119226193B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and provides a communication method, equipment, a storage medium and a program product, wherein the method comprises the following steps: based on a general interface installed at a local artificial intelligence chip, a data transmission interface matched with the local artificial intelligence chip is called, and the general interface is integrated with data transmission interfaces matched with various types of artificial intelligence chips; based on the data transmission interface matched with the local artificial intelligence chip, communication is carried out with a peer artificial intelligence chip through a direct connection transmission path, the peer artificial intelligence chip is installed with the general interface, and a data transmission interface matched with the peer artificial intelligence chip is called through the general interface. The method, equipment, storage medium and program product provided by the application realize direct connection communication between various types of artificial intelligence chips, thereby reducing communication delay between heterogeneous artificial intelligence chips and improving communication efficiency between the heterogeneous artificial intelligence chips.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a communication method, device, storage medium and program product. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, especially the rise of large language models (LLM), the demand for computing resources has shown explosive growth. The size of these models, the size of the training data, and the number of GPU (Graphics Processing Unit) resources required all grow at an exponential rate.

[0003] In some cases, in order to meet the training needs, even thousands or even tens of thousands of GPUs are needed. However, in the current GPU cloud service and resource usage environment, these thousands of GPUs may come from different manufacturers or belong to different product models of the same manufacturer, have different hardware architectures, that is, are heterogeneous.

[0004] Under such circumstances, different models and different manufacturers of GPUs are mixed and used to support large-scale model training, that is, heterogeneous mixed training (also known as "heterogeneous training"). Heterogeneous mixed training can well solve the "heterogeneous island" problem of current data center GPUs.

[0005] The communication efficiency between GPUs of different manufacturers or different models directly affects the overall efficiency of heterogeneous mixed training. How to improve the communication efficiency between heterogeneous GPUs to achieve efficient interconnection and intercommunication between heterogeneous GPUs has become a hot research topic in the industry. SUMMARY

[0006] The present application provides a communication method, device, storage medium and program product to solve the low communication efficiency between GPUs in related technologies.

[0007] The present application provides a communication method, comprising:

[0008] Based on the general interface installed at the local artificial intelligence chip, a data transmission interface adapted to the local artificial intelligence chip is called, and the general interface is integrated with data transmission interfaces adapted to various types of artificial intelligence chips;

[0009] Based on the data transmission interface adapted to the local artificial intelligence chip, communication is performed with the opposite end artificial intelligence chip through a direct connection transmission path, the opposite end artificial intelligence chip is installed with the general interface and calls a data transmission interface adapted to the opposite end artificial intelligence chip through the general interface;

[0010] The direct connection transmission path is composed of the local artificial intelligence chip, the local network card connected with the local artificial intelligence chip, the opposite end network card connected with the opposite end artificial intelligence chip, and the opposite end artificial intelligence chip.

[0011] According to the communication method provided by the application, the data transmission interface comprises a memory copy interface.

[0012] The local artificial intelligence chip communicates with the opposite end artificial intelligence chip through a direct connection transmission path based on the data transmission interface matched with the local artificial intelligence chip.

[0013] Memory copy is performed in the target memory area of the local artificial intelligence chip based on the memory copy interface matched with the local artificial intelligence chip, and data in the target memory area is transmitted to the opposite end artificial intelligence chip through the direct connection transmission path after the copy is completed.

[0014] According to the communication method provided by the application, the data transmission interface further comprises a memory registration interface.

[0015] The memory copy interface matched with the local artificial intelligence chip is further used for performing memory copy in the target memory area of the local artificial intelligence chip before the memory copy.

[0016] The memory registration interface matched with the local artificial intelligence chip is used for performing memory registration in the local artificial intelligence chip to obtain the target memory area in the memory of the local artificial intelligence chip.

[0017] According to the communication method provided by the application, the general interface installed at the local artificial intelligence chip is used for calling the data transmission interface matched with the local artificial intelligence chip.

[0018] The general interface installed at the local artificial intelligence chip is used for calling the data transmission interface matched with at least one of the manufacturer and the product model of the local artificial intelligence chip.

[0019] According to the communication method provided by the application, the data transmission interface matched with the local artificial intelligence chip is used for communicating with the opposite end artificial intelligence chip through a direct connection transmission path.

[0020] The data transmission interface matched with the local artificial intelligence chip is used for transmitting a data block to the opposite end artificial intelligence chip through a direct connection transmission path, and the size of the data block is a target size.

[0021] The target size is one of a plurality of candidate sizes.

[0022] According to the communication method provided by the application, the data transmission interface matched with the local artificial intelligence chip is called based on the general interface installed at the local artificial intelligence chip, and the data transmission interface matched with the local artificial intelligence chip is called.

[0023] The test data is divided into a plurality of data blocks in the candidate size.

[0024] The data transmission interface matched with the local artificial intelligence chip is sent to the opposite end artificial intelligence chip through the direct transmission path based on the data transmission interface matched with the local artificial intelligence chip, and the data transmission time in the candidate size is obtained.

[0025] The target size is determined based on the data transmission time in the plurality of candidate sizes.

[0026] According to the communication method provided by the application, the data transmission interface matched with the local artificial intelligence chip is called based on the general interface installed at the local artificial intelligence chip, and the data transmission interface matched with the local artificial intelligence chip is called.

[0027] In the case that it is determined that the local artificial intelligence chip supports direct memory access between chips and the general interface is installed at the local artificial intelligence chip, the general interface is registered.

[0028] The application further provides a communication device, comprising:

[0029] The calling unit is used for calling the data transmission interface matched with the local artificial intelligence chip based on the general interface installed at the local artificial intelligence chip, and the general interface is integrated with the data transmission interface matched with each type of artificial intelligence chip.

[0030] The communication unit is used for communicating with the opposite end artificial intelligence chip through the direct transmission path based on the data transmission interface matched with the local artificial intelligence chip, and the opposite end artificial intelligence chip is installed with the general interface and calls the data transmission interface matched with the opposite end artificial intelligence chip through the general interface.

[0031] The direct transmission path is composed of the local artificial intelligence chip, the local network card connected with the local artificial intelligence chip, the opposite end network card connected with the opposite end artificial intelligence chip and the opposite end artificial intelligence chip.

[0032] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the communication method as described above when executing the program.

[0033] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the communication method.

[0034] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the communication method.

[0035] The communication method, device, storage medium and program product provided by the application realize unified management of data transmission operations for various artificial intelligence chips through the general interface integrated with the data transmission interfaces respectively matched with the various artificial intelligence chips, and based on this, the data transmission interfaces matched with the various artificial intelligence chips can be called to realize direct connection communication between the various artificial intelligence chips, so as to reduce communication delay between the heterogeneous artificial intelligence chips and improve communication efficiency between the heterogeneous artificial intelligence chips. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0037] Figure 1 is a structural schematic diagram of a computing system based on a heterogeneous GPU cluster in the related art.

[0038] Figure 2 is a communication flow schematic diagram between GPUs in the related art.

[0039] Figure 3 is one of the flow schematic diagrams of the communication method provided by the application.

[0040] Figure 4 is the second flow schematic diagram of the communication method provided by the application.

[0041] Figure 5 is the third flow schematic diagram of the communication method provided by the application.

[0042] Figure 6 is the fourth flow schematic diagram of the communication method provided by the application.

[0043] Figure 7 is a structural schematic diagram of the communication device provided by the application.

[0044] Figure 8 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0045] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application and not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0046] With the increase of the parameter size of large language models and the amount of training data, the demand for storage resources and hardware computing power has exceeded the upper limit that a single computing device can provide. In order to break through this limitation and improve the throughput of training, a 3D parallel training strategy is often used, that is, distributed collaborative training of the model is implemented through data parallelism, tensor parallelism and pipeline parallelism. Here, distributed collaborative training refers to the use of cooperative work between multiple computing devices to jointly participate in the training process of the model to achieve training acceleration and performance improvement of the model. It should be noted that the computing device here is an artificial intelligence chip, which can be, for example, a GPU, a TPU (Tensor Processing Unit), an NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit), etc. The present application does not make specific limitations on this.

[0047] In distributed collaborative training, thousands or even tens of thousands of GPUs are needed to meet the training requirements. Traditional distributed training requires completely consistent GPU devices (same manufacturer and same model), and resource shortage has become a bottleneck for large-scale model training, especially for large language model training.

[0048] In the current GPU cloud service and resource usage environment, these thousands of GPUs may come from different manufacturers or belong to different product models of the same manufacturer, have different hardware architectures, that is, are heterogeneous. In the embodiments of the present application, a cluster composed of GPUs with different hardware architectures is referred to as a heterogeneous GPU cluster.

[0049] For example, Figure 1 is a structural schematic diagram of a computing system based on a heterogeneous GPU cluster in the related art. As Figure 1As shown, the system can include a host 110 and a device end 120, where the device end 120 can include a plurality of first-type devices, a plurality of second-type devices, and other types of devices, which can all be regarded as GPUs or other types of artificial intelligence chips. The host 110 is connected to the plurality of devices of the device end 120, for controlling the devices to perform various computing tasks and cooperatively complete training of a model. At the device end 120, the plurality of first-type devices are all devices of the same manufacturer and same model, which form a homogeneous device cluster, while the first-type devices and the second-type devices can be devices of different manufacturers or different models, which form a heterogeneous device cluster. The heterogeneous mixed training of the model is that, in the process of training the model, the different types of devices (including homogeneous and heterogeneous devices) are jointly used to participate in the training task of the model. These devices can have different computing capabilities, memory sizes, data formats, and other characteristics, but they can work cooperatively to complete the training of the model under the control and coordination of the host.

[0050] In a heterogeneous GPU cluster, different manufacturers and different models of GPUs are involved, and due to the differences in hardware characteristics among the GPUs, communication efficiency problems often occur among the GPUs.

[0051] In related technologies, the communication among GPUs needs to first move data from the memory of one GPU to the system memory, and then transmit the data to the system memory of another computer through the RDMA (Remote Direct Memory Access) technology, and then move the data from the system memory of the other computer to the memory of another GPU.

[0052] For example, Figure 2 is a schematic diagram of the communication process among GPUs in related technologies. In Figure 2 , GPU1 and CPU (Central Processing Unit) 1 belong to one computer, and GPU2 and CPU2 belong to another computer. GPU1 and GPU2 can be GPUs produced by different manufacturers, or GPUs produced by the same manufacturer but of different product models. To realize the communication from GPU1 to GPU2, steps ① to ④ can be completed.

[0053] In the step 1, the data to be transmitted is copied from the HBM (High Bandwidth Memory) of the GPU 1 to the memory of the CPU 1. In the step 2, the data is transmitted from the memory of the CPU 1 to the memory of the CPU 2 through the NIC (Network Interface Controller) by the TCP (Transmission Control Protocol). In the step 3, the data is copied from the memory of the CPU 2 to the HBM of the GPU 2. In the step 4, the layout conversion is performed on the data in the GPU 2, so that the layout of the data obtained by the transmission can conform to the layout of the GPU 2.

[0054] In the steps 1 and 3, the data copying between the HBM of the GPU and the memory of the CPU can be realized by the PCI (Peripheral Component Interconnect) SW (Switch).

[0055] It can be understood that in this process, multiple data copying is involved, including from the GPU to the CPU and from the CPU to the GPU, and the multiple data copying will inevitably increase the communication delay and affect the communication efficiency between the GPUs.

[0056] In order to improve the communication efficiency between the GPUs, the GPU Direct RDMA technology emerges as the times require. The GPU Direct RDMA is a network communication method based on the DMA (Direct Memory Access) technology. In the DMA, the hardware controller can directly exchange data with the memory without the intervention of the CPU, which directly reduces the burden of the CPU and improves the communication efficiency between the hardware controller and the memory. The RDMA is a high-performance remote direct memory access realized by the zero-copy network technology and the kernel memory bypass technology.

[0057] In the zero-copy network technology, the NIC can directly transmit data with the application memory without copying the data to the kernel and then copying the data from the kernel to the application memory. This eliminates the copying step between the application memory and the kernel, thereby reducing the transmission delay. In the kernel memory bypass technology, the application program can directly send commands through the NIC without calling through the kernel. This reduces the number of environment switching between the memory space of the kernel and the space of the application program, further improving the communication efficiency.

[0058] The application of GPU Direct RDMA technology allows a GPU of one computer to directly access the memory of a GPU of another computer, thereby reducing the number of data copying and effectively reducing the communication delay.

[0059] However, the implementation of the RDMA technology relies on specific software programs that need to be closely coordinated with the hardware and software design of the GPU to enable direct access to the remote memory. Due to the differences in hardware and software design of GPUs from different manufacturers or different models of the same manufacturer, the software programs relied on by RDMA have not been adapted to GPUs from different manufacturers or different models of the same manufacturer. This makes the RDMA technology usually only support GPUs produced by a specific manufacturer.

[0060] With the wide application of GPUs from various manufacturers and GPUs of various product models in high-performance computing and data centers, how to implement GPU Direct RDMA communication across the network between GPUs from different manufacturers or different models, that is, how to implement GPU Direct RDMA communication between heterogeneous GPUs, is still a problem to be solved.

[0061] Figure 3 is one of the flowcharts of the communication method provided by the present application, as shown in Figure 3 The communication method comprises the following steps:

[0062] In step 310, based on a general interface installed at the local artificial intelligence chip, a data transmission interface adapted to the local artificial intelligence chip is called, and the general interface is integrated with data transmission interfaces adapted to various types of artificial intelligence chips.

[0063] Specifically, the communication method is applied to the communication between two artificial intelligence chips, which may be, for example, GPU, TPU, NPU, APU, GPGPU, etc., and the embodiments of the present application do not make specific limitation thereto.

[0064] For convenience of distinction, one of the two artificial intelligence chips, which is the execution subject of the communication method, is referred to as the local artificial intelligence chip, and the other is referred to as the opposite artificial intelligence chip. Both of the two artificial intelligence chips can regard themselves as the local artificial intelligence chip, and both of the two artificial intelligence chips are the opposite artificial intelligence chip of each other. It can be understood that the two artificial intelligence chips herein can come from different manufacturers or belong to different product models of the same manufacturer, that is, the two artificial intelligence chips herein can have different hardware architectures, and the two artificial intelligence chips can be heterogeneous.

[0065] For the local artificial intelligence chip, a general interface can be pre-installed at the local artificial intelligence chip. The general interface here is an interface obtained by abstracting data transmission operations respectively adapted to artificial intelligence chips of various manufacturers and various product models produced by the various manufacturers, that is, the general interface integrates data transmission interfaces respectively adapted to artificial intelligence chips of various manufacturers and various product models produced by the various manufacturers, and thus can support communication between the local artificial intelligence chip and artificial intelligence chips of any type.

[0066] Here, the data transmission interface is an interface for calling data transmission operations, and the data transmission operations can include memory registration, memory copy, etc. It can be understood that the memory registration, memory copy, etc. operations adapted to different types of artificial intelligence chips can be different, but the functions of the memory registration, memory copy, etc. operations adapted to different types of artificial intelligence chips are consistent, so the functions of specific operations can be taken as an abstraction basis to integrate various data transmission operations adapted to different types of artificial intelligence chips, and to hide the underlying hardware implementation details of different data transmission operations, thereby obtaining a general interface that can call data transmission interfaces adapted to various types of artificial intelligence chips.

[0067] For the case that the local artificial intelligence chip pre-installs the general interface, the general interface can be used to determine and call the data transmission interface adapted to the local artificial intelligence chip from the data transmission interfaces integrated in the general interface and adapted to various types of artificial intelligence chips. It can be understood that the data transmission interface adapted to the local artificial intelligence chip here can be a data transmission interface corresponding to the manufacturer of the local artificial intelligence chip, or a data transmission interface corresponding to the product model of the local artificial intelligence chip, etc., which is not limited in the embodiments of the present application.

[0068] Step 320, based on the data transmission interface adapted to the local artificial intelligence chip, communicate with the opposite end artificial intelligence chip through a direct transmission path, the opposite end artificial intelligence chip installs the general interface and calls the data transmission interface adapted to the opposite end artificial intelligence chip through the general interface; the direct transmission path is composed of the local artificial intelligence chip, the local network card connected to the local artificial intelligence chip, the opposite end network card connected to the opposite end artificial intelligence chip, and the opposite end artificial intelligence chip.

[0069] Specifically, the opposite end artificial intelligence chip also installs the general interface as the local artificial intelligence chip does, and can call the data transmission interface adapted to the opposite end artificial intelligence chip through the general interface, so that the data transmission interface called by the opposite end artificial intelligence chip can also support communication between the opposite end artificial intelligence chip and artificial intelligence chips of any type.

[0070] Thus, the communication with the opposite end artificial intelligence chip can be realized based on the data transmission interface matched with the opposite end artificial intelligence chip and the general interface call.

[0071] Here, the communication between the local end artificial intelligence chip and the opposite end artificial intelligence chip is realized through the direct connection transmission path.

[0072] Figure 4 is the second flowchart of the communication method provided by the application, the direct connection transmission path can be in the form shown in Figure 4 , that is, the data at the GPU1 of the local end can be transmitted to the network card NIC of the opposite end through the network card NIC of the local end, and then transmitted to the GPU2 of the opposite end, in the direct connection transmission path, the data direct connection transmission from the GPU of the local end to the GPU of the opposite end is realized without passing through the CPU of the local end and the CPU of the opposite end. Specifically in Figure 4 , the communication from the GPU1 to the GPU2 can be completed through step 1 and step 2. Wherein, step 1 is to copy the data to be transmitted from the HBM of the GPU1 to the HBM of the GPU2 through the network card NIC1 and the network card NIC2. Step 2 is to convert the data layout in the GPU2, so that the data layout obtained by the transmission can conform to the layout of the GPU2.

[0073] It can be understood that, compared with the way of realizing the communication between two GPUs through steps 1 to 4 in the related art shown in Figure 2 , in the embodiment of the application, the communication between two GPUs is realized through the direct connection transmission path, which can avoid multiple data copying including from GPU to CPU and from CPU to GPU, thereby greatly reducing the communication delay and improving the communication efficiency.

[0074] And the premise of building a direct transmission path between the heterogeneous local artificial intelligence chip and the opposite artificial intelligence chip to realize the direct communication between the local artificial intelligence chip and the opposite artificial intelligence chip is that the data transmission interface suitable for itself is called at the local artificial intelligence chip and the opposite artificial intelligence chip respectively. It can be understood that, compared with the RDMA technology in the related art which only supports specific manufacturers or specific models, resulting in a large number of artificial intelligence chips that cannot adapt to the software program of the RDMA technology, since the universal interface in the present application can integrate the data transmission interfaces suitable for each type of artificial intelligence chip respectively, therefore, whether the local artificial intelligence chip or the opposite artificial intelligence chip, the data transmission interface suitable for itself can be called through the universal interface, thereby ensuring the compatibility of various types of artificial intelligence chips in the communication method provided in the embodiment of the application. Based on the data transmission interface suitable for itself, the local artificial intelligence chip and the opposite artificial intelligence chip can perform data transmission operations including memory registration, memory copy and the like suitable for itself, thereby breaking the communication barrier between the local artificial intelligence chip and the opposite artificial intelligence chip due to heterogeneity, to realize the direct communication between the local artificial intelligence chip and the opposite artificial intelligence chip, thereby realizing zero-copy transmission (GPU Direct RDMA, GDR) between heterogeneous artificial intelligence chips without detouring the CPU.

[0075] In the method provided in the embodiment of the application, the universal interface integrated with the data transmission interfaces suitable for each type of artificial intelligence chip realizes unified management of data transmission operations for each type of artificial intelligence chip, based on which the data transmission interface suitable for each type of artificial intelligence chip can be called to realize direct communication between each type of artificial intelligence chip, thereby reducing the communication delay between heterogeneous artificial intelligence chips and improving the communication efficiency between heterogeneous artificial intelligence chips.

[0076] And since the universal interface integrates the data transmission interfaces suitable for each type of artificial intelligence chip respectively, it can support various artificial intelligence chips at present and in the future, the communication method executed based on the universal interface has excellent compatibility and expandability, and is easy to integrate into the RDMA communication framework, providing support for future technology development.

[0077] In addition, compared with the communication between chips realized by the CPU, the direct communication between each type of artificial intelligence chip in the embodiment of the application can improve the utilization rate of the PCIE (Peripheral Component Interconnect Express, Peripheral Component Interconnect Express bus) bandwidth, and further optimize the overall performance of the computing system.

[0078] Based on the above embodiment, the data transmission interface includes a memory copy interface;

[0079] In step 320, the data transmission interface matched with the local artificial intelligence chip is used to communicate with the opposite artificial intelligence chip through the direct transmission path, including:

[0080] Based on the memory copy interface matched with the local artificial intelligence chip, memory copy is performed in the target memory area of the local artificial intelligence chip, and after the copy is completed, the data in the target memory area is transmitted to the opposite artificial intelligence chip through the direct transmission path.

[0081] Specifically, in the general interface, the memory copy interface matched with various artificial intelligence chips is integrated. Here, the memory copy interface is a type of interface in the data transmission interface for implementing memory copy operation. It can be understood that different artificial intelligence chips can correspond to different memory copy interfaces, and the memory copy interfaces corresponding to different artificial intelligence chips are consistent in functionality, i.e., they are all used to implement memory copy operation.

[0082] It can be understood that memory copy refers to the process of copying the data of one memory area to another memory area. In the embodiments of the present application, it specifically refers to the process of copying data from a certain memory area of the local artificial intelligence chip to the target memory area of the local artificial intelligence chip. Here, the target memory area is the memory area where data used for communication with the opposite artificial intelligence chip is stored. The target memory area can be any memory area of the local artificial intelligence chip, or a memory area pre-divided by the local artificial intelligence chip for communication with the opposite artificial intelligence chip. The embodiments of the present application do not make specific limitation on this.

[0083] Correspondingly, in the communication process, memory copy can be performed on the data to be transmitted in the local artificial intelligence chip based on the memory copy interface matched with the local artificial intelligence chip, so that the data to be transmitted can be stored in the target memory area of the local artificial intelligence chip. After the memory copy is completed, the local artificial intelligence chip can transmit the data in the target memory area to the opposite artificial intelligence chip through the direct transmission path, thereby realizing communication between the local artificial intelligence chip and the opposite artificial intelligence chip.

[0084] Based on any of the above embodiments, the data transmission interface further includes a memory registration interface.

[0085] In step 320, before the memory copy is performed in the target memory area of the local artificial intelligence chip based on the memory copy interface matched with the local artificial intelligence chip, it further includes:

[0086] Based on the memory registration interface adapted to the local AI chip, memory registration is performed on the local AI chip to obtain the target memory region in the local AI chip's memory.

[0087] Specifically, the general interface also integrates memory registration interfaces adapted to various artificial intelligence chips. Here, the memory registration interface is a type of interface used in the data transmission interface to implement memory registration operations. It can be understood that different artificial intelligence chips can correspond to different memory registration interfaces, and the functionality of the memory registration interfaces corresponding to different artificial intelligence chips is consistent, that is, they are all used to implement memory registration operations.

[0088] It is understood that memory registration refers to the operation of marking a memory region for a specific purpose or state. In this embodiment of the invention, it specifically refers to the operation of marking a memory region in the local AI chip as a memory region used for communication with the remote AI chip.

[0089] Accordingly, during communication, a memory registration operation can be performed on the local AI chip's memory region via a direct transmission path, based on a memory registration interface compatible with the local AI chip. This registers a memory region on the local AI chip for communication with the remote AI chip, thus obtaining the target memory region within the local AI chip. After memory registration, subsequent memory copy operations based on the memory copy interface can copy the data to be transmitted to the target memory region of the local AI chip. Once the memory copy is complete, the local AI chip can then transmit the data within the target memory region to the remote AI chip via the direct transmission path.

[0090] In the method provided in this embodiment of the invention, memory registration and memory copy are abstracted to realize data transmission. Thus, when communicating between heterogeneous artificial intelligence chips, memory registration and memory copy can be realized for the local artificial intelligence chip by selecting the memory registration interface and memory copy interface adapted to the local artificial intelligence chip, thereby ensuring efficient communication between heterogeneous artificial intelligence chips.

[0091] Furthermore, the abstraction of the memory registration interface and the memory copy interface enables unified management of memory registration and copying for different types of artificial intelligence chips, thereby improving the communication efficiency and reliability between chips.

[0092] For example, Figure 5 This is the third flowchart illustrating the communication method provided by the present invention, as shown below. Figure 5 As shown, a universal interface can be installed on each of the two artificial intelligence chips.

[0093] At each AI chip, a memory registration interface adapted to its own can be called through a general interface to allocate device memory, thereby achieving memory registration. This is understandable. Figure 5 The result of memory registration is the target memory region at each smart chip. The target memory region is then registered with the network card connected to it, so that the network card can send the data in the target memory region to the network card connected to the peer AI chip during communication, and write the data received by the network card connected to the peer AI chip into the target memory region.

[0094] Furthermore, each AI chip can adjust its memory copy interface via a universal interface to copy data to be transmitted from the user buffer to the target memory area, thus achieving device memory copying. After the memory copy is completed, it can trigger the data transmission operation of the network card connected to it, thereby realizing communication between AI chips.

[0095] Furthermore, before any two AI chips can communicate, each chip must perform the following steps to initialize inter-chip communication:

[0096] First, initialize the RDMA device, which can be an artificial intelligence chip.

[0097] Then, create a PD (Protection Domain), a CQ (Completion Queue), and a QP (Queue Pair). After that, a QP connection can be established, and QP information can be exchanged between the two communicating parties.

[0098] Following this, each AI chip can call its own compatible memory registration interface to register the target memory region, thereby creating the memory. Additionally, during data transmission and reception, each AI chip can call its own compatible memory copy interface to copy the data to be transmitted to the target memory region, thus performing a memory copy.

[0099] Based on any of the above embodiments, in step 310, the step of calling a data transmission interface adapted to the local AI chip based on the general interface installed on the local AI chip includes:

[0100] Based on the universal interface installed on the local AI chip, a data transmission interface compatible with at least one of the manufacturers and product models of the local AI chip is invoked.

[0101] Specifically, to enable communication between AI chips, for the local AI chip, a data transmission interface compatible with the local AI chip can be determined and called from the data transmission interfaces integrated in the general interface that are compatible with various AI chips.

[0102] Since AI chips from different manufacturers may use different data transmission operations to achieve data transmission, or AI chips of different product models from the same manufacturer may also use different data transmission operations to achieve data transmission, when determining the data transmission interface that is compatible with the local AI chip, at least one of the manufacturer and product model of the local AI chip can be used as the identifier of the data transmission interface that the local AI chip is compatible with.

[0103] Therefore, the data transmission interface adapted to the local AI chip can be a data transmission interface adapted to the manufacturer of the local AI chip, a data transmission interface adapted to the product model of the local AI chip, or a data transmission interface adapted to both the manufacturer and the product model of the local AI chip.

[0104] Based on any of the above embodiments, step 320, which involves communicating with the peer AI chip via a direct transmission path using a data transmission interface adapted to the local AI chip, includes:

[0105] Based on the data transmission interface adapted to the local AI chip, data blocks are transmitted to the remote AI chip via a direct transmission path, and the size of the data blocks is the target size;

[0106] The target size is one of several candidate sizes.

[0107] Specifically, during the communication process between the local AI chip and the remote AI chip, the data transmitted based on the data transmission interface adapted to the local AI chip can be transmitted in the form of data blocks. That is, the data to be transmitted can be divided into multiple data blocks, and multiple data blocks can be transmitted sequentially to form a data stream.

[0108] The data block here refers to a data unit with a fixed structure and format. The size of the data block can be adjusted according to the actual transmission scenario. Here, the size of the data block is denoted as the target size, which can be one of several candidate sizes.

[0109] That is, multiple candidate sizes can be preset. Each candidate size can be selected as the target size for the data block. In actual transmission scenarios, one of the multiple candidate sizes can be selected as the target size, and data transmission can be performed accordingly. It is understood that the candidate size selected as the target size can be the candidate size with the highest transmission efficiency among multiple candidate sizes, or the default candidate size of the local AI chip, or the candidate size used in the previous communication with the peer AI chip. This embodiment of the invention does not specifically limit this.

[0110] Based on any of the above embodiments, before step 320, the method further includes:

[0111] The test data is divided into multiple data blocks under the candidate size;

[0112] Based on the data transmission interface adapted to the local AI chip, multiple data blocks of the candidate size are sent to the remote AI chip through a direct transmission path to obtain the data transmission time of the candidate size;

[0113] The target size is determined based on the data transmission time under multiple candidate sizes.

[0114] Specifically, before implementing communication between the local AI chip and the remote AI chip based on the data transmission interface, the transmission efficiency of each candidate size between the local AI chip and the remote AI chip can be tested first, and the candidate size with the highest transmission efficiency can be selected as the target size.

[0115] In this process, the test data can be determined first. The test data here is the data used to test the transmission efficiency of each candidate size. The test data can be pre-constructed data or any other data that can be used for communication between artificial intelligence chips.

[0116] After obtaining the test data, it can be divided into multiple data blocks under each candidate size. Understandably, when multiple candidate sizes exist, the test data needs to be divided for each candidate size. Therefore, under each candidate size, there are multiple data blocks with the same size as that candidate size.

[0117] For each candidate size, multiple data blocks of that size can be transmitted to the peer AI chip via a direct transmission path using a data transmission interface compatible with the local AI chip. The time taken to transmit these data blocks is recorded, thus obtaining the data transmission time for that candidate size. It is understandable that since data blocks are transmitted for each candidate size, a corresponding data transmission time can be obtained for each size. This data transmission time reflects the communication efficiency between the local and peer AI chips for the same amount of data transmitted. Generally, a longer data transmission time indicates lower communication efficiency, and a shorter data transmission time indicates higher communication efficiency.

[0118] Therefore, after obtaining the data transmission time for each candidate size, a target size can be selected from the candidate sizes based on the data transmission time for each candidate size. For example, the candidate size with the shortest data transmission time can be selected from multiple candidate sizes as the target size.

[0119] In the method provided in this embodiment of the invention, the target size is determined based on the data transmission time under multiple candidate sizes, which helps to improve the communication efficiency between artificial intelligence chips.

[0120] Based on any of the above embodiments, before step 310, the method further includes:

[0121] If it is determined that the local AI chip supports direct memory access between chips and that a general interface is installed on the local AI chip, then general interface registration is performed.

[0122] Specifically, before calling the data transmission interface based on the general interface of the local AI chip, functional judgment, installation detection and interface registration can be performed on the local AI chip.

[0123] The functional determination refers to determining whether the local AI chip itself supports direct memory access between chips, i.e., whether the AI ​​chip itself has GDR (GPU Direct RDMA) functionality. It is understood that the communication method provided in this embodiment of the invention is executed only if the local AI chip supports direct memory access between chips. If the local AI chip does not support direct memory access between chips, then the local AI chip needs to communicate with the peer AI chip through a host device such as a CPU.

[0124] The installation detection is used to detect whether a universal interface is installed on the local AI chip. It is understood that the communication method provided in this embodiment of the invention is only executed if the local AI chip has a universal interface installed. If the local AI chip does not have a universal interface installed, it cannot call the data transmission interface adapted to itself through the universal interface and needs to communicate with the peer AI chip through a host device such as a CPU.

[0125] Interface registration refers to registering a general-purpose interface as a plug-in at the local AI chip. This enables the local AI chip to manage and use the general-purpose interface, allowing it to communicate by calling a compatible data transmission interface. It is understood that the communication method provided in this embodiment requires interface registration before calling the data transmission interface based on the general-purpose interface. If interface registration fails, the data transmission interface cannot be called through the general-purpose interface. In this case, the local AI chip needs to communicate with the peer AI chip through a host device such as a CPU.

[0126] For example, Figure 6 This is the fourth flowchart illustrating the communication method provided by the present invention, as shown below. Figure 6 As shown, for the local AI chip, we can first perform GDR function detection, that is, detect whether the local AI chip supports GDR function, which means determining whether the local AI chip itself supports direct memory access between chips. If the local AI chip is found to support GDR function, then proceed to the next step; otherwise, communication between chips is performed through RDMA communication between CPUs.

[0127] If the local AI chip is found to support GDR functionality, GDR plug-in detection can be performed. Here, GDR plug-in refers to the general interface as described in this embodiment of the invention. GDR plug-in detection checks whether the local AI chip has a general interface installed. If a general interface is detected, proceed to the next step; otherwise, communication between the chips is achieved through RDMA communication between CPUs.

[0128] The local AI chip has been found to have a universal interface, allowing for GDR plugin registration. This GDR plugin registration refers to the universal interface registration as described in this embodiment. If the interface registration is successful, proceed to the next step; otherwise, inter-chip communication is achieved via RDMA communication between CPUs.

[0129] Once the interface registration is successful, GDR (Generative Data Redirect) can be performed. Specifically, based on the general interface installed on the local AI chip, a data transmission interface compatible with the local AI chip can be called to achieve communication with the remote AI chip. The communication referred to here can include multiple rounds of data transmission and / or data reception operations.

[0130] The communication device provided by the present invention is described below. The communication device described below and the communication method described above can be referred to in correspondence.

[0131] Figure 7 This is a schematic diagram of the communication device provided by the present invention, as shown below. Figure 7 As shown, the device includes:

[0132] The calling unit 710 is used to call a data transmission interface adapted to the local artificial intelligence chip based on a general interface installed on the local artificial intelligence chip. The general interface integrates data transmission interfaces adapted to various types of artificial intelligence chips.

[0133] The communication unit 720 is used to communicate with the peer AI chip through a direct transmission path based on a data transmission interface adapted to the local AI chip. The peer AI chip is equipped with the general interface and calls the data transmission interface adapted to the peer AI chip through the general interface.

[0134] The direct transmission path consists of the local AI chip, the local network interface card connected to the local AI chip, the peer network interface card connected to the peer AI chip, and the peer AI chip.

[0135] In the device provided in the embodiments of the present invention, a universal interface with data transmission interfaces adapted to various types of artificial intelligence chips is integrated to realize unified management of data transmission operations for various types of artificial intelligence chips. Based on this, the data transmission interfaces adapted to various types of artificial intelligence chips can be called to realize direct communication with various types of artificial intelligence chips, thereby reducing the communication latency between heterogeneous artificial intelligence chips and improving the communication efficiency between heterogeneous artificial intelligence chips.

[0136] Based on the above embodiments, the data transmission interface includes a memory copy interface;

[0137] The communication unit is specifically used for:

[0138] Based on the memory copy interface adapted to the local AI chip, memory copying is performed in the target memory area of ​​the local AI chip, and after the copy is completed, the data in the target memory area is transmitted to the peer AI chip through the direct transmission path.

[0139] Based on any of the above embodiments, the data transmission interface further includes a memory registration interface;

[0140] The communication unit is also used for:

[0141] Based on the memory registration interface adapted to the local AI chip, memory registration is performed on the local AI chip to obtain the target memory region in the local AI chip's memory.

[0142] Based on any of the above embodiments, the calling unit is specifically used for:

[0143] Based on the universal interface installed on the local AI chip, a data transmission interface compatible with at least one of the manufacturers and product models of the local AI chip is invoked.

[0144] Based on any of the above embodiments, the communication unit is specifically used for:

[0145] Based on the data transmission interface adapted to the local AI chip, data blocks are transmitted to the remote AI chip via a direct transmission path, and the size of the data blocks is the target size;

[0146] The target size is one of several candidate sizes.

[0147] Based on any of the above embodiments, the communication unit is further configured to:

[0148] The test data is divided into multiple data blocks under the candidate size;

[0149] Based on the data transmission interface adapted to the local AI chip, multiple data blocks of the candidate size are sent to the remote AI chip through a direct transmission path to obtain the data transmission time of the candidate size;

[0150] The target size is determined based on the data transmission time under multiple candidate sizes.

[0151] Based on any of the above embodiments, the calling unit is further configured to:

[0152] If it is determined that the local AI chip supports direct memory access between chips and that a general interface is installed on the local AI chip, then general interface registration is performed.

[0153] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a communication method, which includes:

[0154] Based on the general interface installed on the local AI chip, a data transmission interface adapted to the local AI chip is called. The general interface integrates data transmission interfaces adapted to various types of AI chips.

[0155] Based on the data transmission interface adapted to the local AI chip, communication is established with the remote AI chip via a direct transmission path. The remote AI chip is equipped with the general interface and calls the data transmission interface adapted to the remote AI chip through the general interface.

[0156] The direct transmission path consists of the local AI chip, the local network interface card connected to the local AI chip, the peer network interface card connected to the peer AI chip, and the peer AI chip.

[0157] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0158] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to execute the communication methods provided by the above methods, the method comprising:

[0159] Based on the general interface installed on the local AI chip, a data transmission interface adapted to the local AI chip is called. The general interface integrates data transmission interfaces adapted to various types of AI chips.

[0160] Based on the data transmission interface adapted to the local AI chip, communication is established with the remote AI chip via a direct transmission path. The remote AI chip is equipped with the general interface and calls the data transmission interface adapted to the remote AI chip through the general interface.

[0161] The direct transmission path consists of the local AI chip, the local network interface card connected to the local AI chip, the peer network interface card connected to the peer AI chip, and the peer AI chip.

[0162] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the communication methods provided by the methods described above, the method comprising:

[0163] Based on the general interface installed on the local AI chip, a data transmission interface adapted to the local AI chip is called. The general interface integrates data transmission interfaces adapted to various types of AI chips.

[0164] Based on the data transmission interface adapted to the local AI chip, communication is established with the remote AI chip via a direct transmission path. The remote AI chip is equipped with the general interface and calls the data transmission interface adapted to the remote AI chip through the general interface.

[0165] The direct transmission path consists of the local AI chip, the local network interface card connected to the local AI chip, the peer network interface card connected to the peer AI chip, and the peer AI chip.

[0166] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0167] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A communication method, characterized in that, include: Based on the universal interface installed on the local AI chip, a data transmission interface adapted to the local AI chip is called. The universal interface integrates data transmission interfaces adapted to various types of AI chips. The universal interface supports communication between the local AI chip and any type of AI chip. Based on the data transmission interface adapted to the local AI chip, communication is established with the remote AI chip via a direct transmission path. The remote AI chip is equipped with the general interface and calls the data transmission interface adapted to the remote AI chip through the general interface. The direct transmission path consists of the local AI chip, the local network card connected to the local AI chip, the peer network card connected to the peer AI chip, and the peer AI chip. The local AI chip and the peer AI chip are from different manufacturers or different product models of the same manufacturer. The communication with the peer AI chip via a direct connection path, based on a data transmission interface adapted to the local AI chip, includes: Determine multiple candidate sizes; Based on the data transmission interface adapted to the local AI chip, data blocks are transmitted to the remote AI chip via a direct transmission path, and the size of the data blocks is the target size; The target size is one of the plurality of candidate sizes; The data transmission interface adapted to the local AI chip communicates with the remote AI chip via a direct transmission path, and prior to this, it also includes: The test data is divided into multiple data blocks under the candidate size; Based on the data transmission interface adapted to the local AI chip, multiple data blocks of the candidate size are sent to the remote AI chip through a direct transmission path to obtain the data transmission time of the candidate size; The target size is determined based on the data transmission time under multiple candidate sizes; The data transmission interface includes a memory copy interface; The communication with the peer AI chip via a direct connection transmission path, based on a data transmission interface adapted to the local AI chip, includes: Based on the memory copy interface adapted to the local AI chip, memory copying is performed in the target memory area of ​​the local AI chip, and after the copying is completed, the data in the target memory area is transmitted to the peer AI chip through the direct transmission path. The data transmission interface also includes a memory registration interface; The process of copying memory in the target memory region of the local AI chip using a memory copy interface adapted to the local AI chip includes, prior to: Based on the memory registration interface adapted to the local AI chip, memory registration is performed on the local AI chip to obtain the target memory region in the local AI chip's memory.

2. The communication method according to claim 1, characterized in that, The method of calling a data transmission interface adapted to the local AI chip based on a general interface installed on the local AI chip includes: Based on the universal interface installed on the local AI chip, a data transmission interface compatible with at least one of the manufacturers and product models of the local AI chip is invoked.

3. The communication method according to any one of claims 1 to 2, characterized in that, The method of calling a data transmission interface adapted to the local AI chip based on a universal interface installed on the local AI chip also includes: If it is determined that the local AI chip supports direct memory access between chips and that a general interface is installed on the local AI chip, then general interface registration is performed.

4. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the communication method as described in any one of claims 1 to 3.

5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the communication method as described in any one of claims 1 to 3.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the communication method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Artificial intelligence chip, data transmission method and data transmission system

    CN116614433A