Cross-core trunking communication system, cross-core trunking communication method, electronic equipment and storage medium

By designing a cross-chip cluster communication system, the limitations of communication libraries for clusters of chips from the same source are solved, enabling efficient communication between chips from different sources, improving the flexibility and resource utilization of the cluster, supporting cluster communication for various hardware architectures, and meeting the needs of large-scale computing.

CN121750658APending Publication Date: 2026-03-27BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, communication libraries for chip clusters based on the same source rely on the hardware characteristics of specific manufacturers, resulting in poor scalability and flexibility of the cluster, failing to meet diverse computing needs, limiting resource utilization and performance, and lacking the possibility of cross-manufacturer technology integration.

Method used

A cross-chip cluster communication system is designed, including a chip interface module, a separate communication module for chips of the same origin, and a unified communication module for chips of different origins. The communication type is determined by acquiring chip information, and efficient same-origin or heterogeneous cluster communication is achieved using adapters and the unified communication module. It supports multiple AI frameworks and hardware architectures and provides a unified communication programming model and topology components.

Benefits of technology

It enables efficient communication between chips from different manufacturers, improves communication efficiency across chip clusters, supports flexible selection of multiple hardware architectures, meets the computing power requirements of large-scale heterogeneous training, simplifies the user experience, and promotes the transparency and diversified utilization of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750658A_ABST
    Figure CN121750658A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-core trunking communication system, a cross-core trunking communication method, electronic equipment and a storage medium. The cross-core trunking communication system comprises a chip interface module used for obtaining information of a plurality of chips, and the information of the plurality of chips is used for determining whether the plurality of chips belong to homologous chips or heterologous chips; the homologous chip independent communication module comprises adapters matched with various chips and is used for homologous chip cluster communication; and the heterogeneous chip unified communication module is used for heterogeneous chip cluster communication. The cross-core trunking communication system provided by the invention not only can be suitable for a cluster composed of chips of the same manufacturer, but also can be suitable for a cluster composed of chips of different manufacturers. Therefore, the user can select the computing power resource required for executing the service more flexibly.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of chip communication, in particular to a cross-chip cluster communication system, a cross-chip cluster communication method, an electronic device and a storage medium. BACKGROUND

[0002] With the rapid expansion of artificial intelligence (AI) models, computing power resources have become a core problem that AI vendors must solve. In the current computing environment, cluster communication libraries play a crucial role, especially in high-performance computing (HPC), big data processing, and cloud computing. These libraries provide a mechanism that allows processes distributed across different computing nodes to exchange data efficiently.

[0003] The cluster communication library of the same source chip usually depends on the hardware characteristics and optimization of a specific manufacturer, thereby realizing efficient data transmission and processing capability. For example, some libraries may be optimized for a specific manufacturer's processor architecture or network technology (such as Intel's Omni-Path or NVIDIA's NVLink). This close binding relationship ensures the best performance in the same source chip environment.

[0004] However, this design that relies on the same manufacturer's chip also has significant limitations. First, it limits the scalability and flexibility of the cluster, as all nodes must use chips from the same manufacturer. This not only increases costs but also limits the user's ability to choose the most suitable hardware according to specific application needs. Second, as computing tasks become more diverse, a single manufacturer's chip may not be able to meet all computing needs, leading to low resource utilization and performance bottlenecks. Finally, this design is not conducive to technological innovation and diversity, as it tends to lock in specific hardware and technology ecosystems, limiting the possibility of cross-manufacturer technology integration.

[0005] Currently, there is no heterogeneous communication library that supports cross-chip high-speed interconnection in the industry. Therefore, in order to alleviate the shortage of computing power, promote the compatibility and interconnection of different computing power, realize large-scale pool training of different chips, and maximize the advantages of various devices, a heterogeneous communication library has become an urgent link to be solved. SUMMARY

[0006] The present application provides a cross-chip cluster communication system, a cross-chip cluster communication method, an electronic device and a storage medium.

[0007] In a first aspect, the present application provides a cross-chip cluster communication system. The system comprises: a chip interface module configured to obtain information of a plurality of chips, the information of the plurality of chips being used to determine whether the plurality of chips belong to a same-source chip or a different-source chip; a same-source chip individual communication module comprising an adapter matched with the plurality of chips, and configured to implement cluster communication for same-source chips; and a different-source chip unified communication module configured to implement cluster communication for different-source chips.

[0008] In a second aspect, the present application provides a cross-chip cluster communication method. The method comprises: obtaining information of a plurality of chips; determining a type of current cross-chip cluster communication according to the information of the plurality of chips; when the current cross-chip cluster communication is a same-source chip cluster communication, invoking a same-source chip individual communication module to implement cluster communication; and when the current cross-chip cluster communication is a different-source chip cluster communication, invoking a different-source chip unified communication module to implement cluster communication.

[0009] In a third aspect, an electronic device is provided. The electronic device comprises a processor and a memory, the memory being configured to store program instructions, the program instructions being executed by the processor to implement the cross-chip cluster communication method.

[0010] In a fourth aspect, a readable storage medium is provided, and is configured to store program instructions. The program instructions are executed by the processor to implement the cross-chip cluster communication method.

[0011] The cross-chip cluster communication system provided by the present application can be applied to a cluster composed of chips of the same manufacturer, and can also be applied to a cluster composed of chips of different manufacturers. Therefore, a user can more flexibly select the computing power resources required to perform a service. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0013] Figure 1 According to an embodiment of the present application, a cross-chip cluster communication system is provided.

[0014] Figure 2 According to another embodiment of the present application, a cross-chip cluster communication system is provided.

[0015] Figure 3 A schematic diagram of implementing inter-cluster kernel inter-cluster communication by a cross-chip cluster communication system is shown.

[0016] Figure 4is a schematic diagram of a communication architecture related to heterogeneous chip cluster communication.

[0017] Figure 5 is a schematic diagram of an implementation manner of inter-node RDMA communication provided by the present application.

[0018] Figure 6 is a cross-chip cluster communication method according to an embodiment of the present application.

[0019] Figure 7 is a topology construction method of cross-chip cluster communication according to an embodiment of the present application.

[0020] Figure 8 is an inter-cluster communication method according to an embodiment of the present application.

[0021] Figure 9 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] The technical solutions in the present application will be described clearly and completely in the present application in combination with the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0023] As described above, there is no homogenous communication library supporting cross-chip high-speed interconnection in the industry at present, and the CPU forwarding technology is mainly used to realize cross-chip heterogeneous training, which leads to low efficiency of heterogeneous training in large-scale scenarios. In addition, the collection communication libraries (NCCL, RCCL, HCCL, CNCL) matched by mainstream chip manufacturers (Nvidia, AMD, Huawei, Cambrian etc.) lack unified interfaces and protocols, which greatly increases the development barriers of cross-chip high-speed interconnection.

[0024] In view of this situation, we hope to design a heterogeneous unified communication library based on standard network protocol (IB / RoCE) supporting cross-node communication of homogenous / heterogeneous chips and multiple deep learning frameworks, realizing multi-chip computing power compatible interconnection and assisting the real meaning of computing power transparency. The cross-chip cluster communication library provided by the present application can realize one or more of the following functions:

[0025] The upper layer of the cross-chip cluster communication library provided by the present application interfaces with multiple AI frameworks (such as PyTorch, Tensorflow, Paddlepaddle, etc.), and connects multiple frameworks-unified communication library-multi-chip link, greatly simplifying the use difficulty of users in different frameworks and chip architectures.

[0026] The cross-chip cluster communication library provided in the application creates a unified interface and standard for the industry mainstream chip manufacturers' set communication library, helps to build a production environment and research and development ecology based on a multi-chip computing power pool, and effectively meets the growing computing power demand brought by large model training.

[0027] The cross-chip cluster communication library provided in the application breaks through the cross-node chip RDMA high-speed interconnection by building a unified communication programming model, and natively supports homogenous / heterogeneous communication operations (Send / Recv and AllReduce, etc.). Compared with the existing cross-chip communication through CPU forwarding in the industry, the efficiency of cross-chip communication can be greatly improved, and the large-scale heterogeneous mixed training is enabled.

[0028] The cross-chip cluster communication library provided in the application supports multiple AI frameworks (such as PyTorch, Tensorflow, Paddlepaddle, etc.) through framework plugins (Framwork Plugin) and unified application program interfaces (API) upward, and supports multiple hardware architectures downward, with high compatibility.

[0029] The cross-chip cluster communication library provided in the application calls the vendor's self-developed communication library (such as NCCL, RCCL, HCCL, etc.) through the homogenous adapter (Adaptor) in the homogenous scenario, and realizes efficient intra-node (Intra-node) and inter-node (Inter-node) point-to-point (P2P) and collective (Collective) communication through the heterogeneous chip unified communication module in the heterogeneous scenario.

[0030] The cross-chip cluster communication library provided in the application can perform advanced work such as collective communication algorithm optimization, remote direct memory access (RDMA) transmission optimization, fault tolerance and elasticity optimization, flow control optimization, and adaptive adjustment optimization based on the unified communication programming model, unified topology component and unified communication component provided by the heterogeneous chip unified communication module.

[0031] Figure 1 According to an embodiment of the application, a cross-chip cluster communication system is shown. As shown in Figure 1 The system includes a chip interface module 11, a homogenous chip separate communication module 12 and a heterogeneous chip unified communication module.

[0032] The chip interface module 11 is used to obtain information of a plurality of chips, and the information of the plurality of chips is used to determine whether the plurality of chips belong to homogenous chips or heterogeneous chips. Wherein, the homogenous chips refer to a plurality of chips from the same manufacturer, and the heterogeneous chips refer to a plurality of chips from different manufacturers.

[0033] The chip interface module 11 is responsible for obtaining information of multiple chips. These information can include but not limited to the manufacturer, model, architecture, version, performance parameters (such as processor speed, memory size, power consumption, etc.) of the chips, as well as supported communication protocols and interfaces, etc. These information can be obtained by querying the hardware abstraction layer (HAL) of the chips, or by executing specific system commands or API calls. In some cases, the information of the chips can also be read from the configuration file or firmware of the system.

[0034] Once the information of the chips is obtained, the chip interface module 11 can determine whether these chips are homogenous or heterogeneous. If the manufacturer, model and architecture of all chips are the same, then these chips can be considered as homogenous chips. This means that they can have the same performance characteristics and communication interfaces, and therefore can use the same communication protocol and algorithm for cluster communication. If the chips come from different manufacturers, then these chips can be considered as heterogeneous chips. This means that they can have different performance characteristics and communication interfaces, and therefore may need to use different communication protocols and algorithms for cluster communication. It should be noted that if the chips come from the same manufacturer but have different models or architectures, they can theoretically use the corresponding homogenous communication library of the manufacturer for cluster communication, and therefore can be considered as homogenous chips; however, they can also be split according to the model and architecture, and therefore can also be considered as heterogeneous chips in some embodiments.

[0035] The homogenous chip individual communication module 12 includes adapters 121 (such as adapter A, adapter B and adapter C) that match multiple chips, for homogenous chip cluster communication. Figure 1 The adapters A, B and C shown are used for homogenous chip cluster communication.

[0036] The homogenous chip individual communication module 12 includes adapters 121 that match multiple chips. These adapters are specially designed and can match specific models of chips to implement communication interface docking with the chips. For example, it can include an NCCL adapter corresponding to Nvidia chips, an HCCL adapter corresponding to Huawei chips, an RCCL adapter corresponding to AMD chips, etc. When cluster communication is needed, the homogenous chip individual communication module 12 will select the adapter that matches the target chip for communication.

[0037] The heterogeneous chip unified communication module 13 is used for heterogeneous chip cluster communication.

[0038] The heterogeneous chip unified communication module 13 is used to process communication problems between chips from different manufacturers with different models and architectures. Its function is to provide a general communication framework so that different types of chips can seamlessly exchange data and cooperate. In order to achieve this function, the heterogeneous chip unified communication module 13 can use interface standardization, communication protocol conversion, data format adaptation and other technologies. Its specific explanation will be introduced later.

[0039] The cross-chip cluster communication system provided by the present application can be applied to a cluster composed of chips of the same manufacturer and a cluster composed of chips of different manufacturers. Therefore, users can more flexibly select the computing power resources needed to perform services.

[0040] Figure 2 According to another embodiment of the present application, a cross-chip cluster communication system is shown. As shown in the figure, the cross-chip cluster communication system can include one or more of the following modules or components: a unified communication component 21, a homogenous chip individual communication module 22, a heterogeneous chip unified communication module 23, a unified topology component 24, a framework docking module 25, a user interface 26, etc. It should be understood that in actual application, the cross-chip cluster communication system can only include part of these modules / components to realize its corresponding functions, and should not be understood as necessarily including all the aforementioned modules / components. Figure 2

[0041] The framework docking module 25 is used to dock with the service framework. For example, the service framework can be any AI-related service framework. The framework docking module 25 and the user interface 26 constitute the upper interface. It supports various AI frameworks (including PyTorch plug-ins, Tensorflow plug-ins, PaddlePaddle plug-ins, etc.) upwards, integrates various communication library backends (including initialization / termination functions, communication operation functions, memory operation functions, etc.) downwards, and supports various hardware architectures. The functions of the upper interface are implemented as follows (taking docking with PyTorch as an example for illustration):

[0042] • Define and register the cross-chip cluster communication system

[0043] ​Define the cross-core cluster communication library: For convenience of description, the cross-core cluster communication system is referred to as FlagCX. When interfacing with the upper layer framework through the interface, a communication backend class BackendFlagCX is first defined, which inherits from the base backend class Backend, used to implement specific communication operations. This class includes a constructor and a series of communication operation methods, such as broadcast, allreduce, reduce, allgather, etc. Among them, the base backend class refers to an abstract class or interface that defines the communication operations and interfaces that all communication backend classes must implement. It provides a unified programming model for different communication technologies, allowing the upper layer application or library to call the communication operations of different backends in a consistent manner. The base backend class usually does not contain specific implementation logic, but declares a series of methods and properties, which are then implemented in the derived communication backend class. The communication backend class refers to the class responsible for implementing a specific communication protocol or mechanism. In a distributed system or cluster communication library, the backend usually refers to the underlying technology or service that supports data transmission and message passing. In the BackendFlagCX class, specific methods are implemented for various cluster communication operations. These methods accept tensors as input, perform the corresponding communication operations, and return a work object representing the status and result of the asynchronous communication operation. Further, a work class workFlagCX class is created, which inherits from the work base class. This class is used to encapsulate and manage the execution status of communication operations, including checking whether the operation is complete, whether it is successful, and waiting for the operation to complete. In the workFlagCX class, methods are provided to allow the asynchronous results of the operation to be obtained, which is implemented through the interaction with the future object. Among them, the Future object is a programming pattern used to handle operations that may take a long time, especially in parallel and asynchronous programming. When an asynchronous operation is started, a Future object is immediately obtained.

[0044] Registering the cross-core cluster communication library: define a static method createBackendFlagCX that creates and initializes a BackendFlagCX object. This method receives the configuration information of the cluster (such as storage objects, ranks, size, timeout settings) as parameters. By defining a special constructor BackendFlagCXConstructor and marking it with the __attribute__((constructor)) attribute, this constructor is automatically executed when the module is loaded. In this constructor, use the pybind11 library to register the createBackendFlagCX method in the Python module, allowing Python code to create and use the BackendFlagCX communication backend through the registered backend name "flagcx", thereby realizing the registration of the backend constructor. In the definition of the BackendFlagCX class, the implementation of the constructor is included. When the library containing this class is loaded, the constructor is automatically executed, completing the registration of the communication backend.

[0045] • Compile the dynamic link library using PyTorch's CppExtension tool

[0046] First, import the required modules and libraries. This includes os (for handling operating system-related operations), torch (PyTorch library for deep learning and tensor computation), setuptools (for packaging Python projects), and torch.utils.cpp_extension (for compiling and loading C++ extensions), among others. Next, define the source files and include directories. Source files are C++ code files, and include directories are directories where header files are located. In an example, the source file is "src / flagcx.cpp", and the include directory is the "include / " directory in the current file's directory. Then create a CppExtension object: create a C++ extension using torch.utils.cpp_extension.CppExtension. This object takes three arguments: the name (in this example, "torch_flagcx"), a list of source files, and a list of include directories. Then set up the project using setuptools' setup function: this function takes four arguments: the name, the version, a list of extension modules (which can be one or more modules), and a dictionary of command classes that maps the 'build_ext' command to torch.utils.cpp_extension.BuildExtension. This command class is used to compile and load C++ extensions. Finally, execute this script, which will compile the C++ extension and create a dynamic link library named "torch_flagcx". This library can be imported and used in Python.

[0047] • Runtime loading of FlagCX backend

[0048] When the upper layer framework needs to call the cross-chip cluster communication library, the registered class can be loaded and the compiled library can be imported to implement the relevant operations. Taking the implementation of global reduction as an example, first, the PyTorch library is imported, which is a widely used deep learning framework that provides rich tensor operations and automatic differentiation functions. At the same time, the distributed computing module torch.distributed of PyTorch is imported, which provides a set of APIs for communication across multiple processes or computing nodes. In addition, the custom C++ extension torch_flagcx that was compiled earlier is imported, which implements the FlagCX backend for optimizing distributed computing performance. Then the distributed process group is initialized: use the torch.distributed.init_process_group function to initialize a distributed process group. This function needs to specify the backend name (in this case, “flagcx”), indicating that the custom FlagCX backend will be used for communication. In addition, the rank (unique identifier) of the current process and the world_size (total number of processes in the process group) also need to be specified. This step is the basis of distributed computing, ensuring that each process can identify and communicate with each other through the FlagCX backend. Then perform the distributed tensor reduction operation: call the torch.distributed.all_reduce function to perform the reduction operation on the specified tensor. The op parameter of this function specifies the type of reduction operation, which is dist.ReduceOp.SUM in this case, indicating a sum operation on the tensor. This operation will be executed cooperatively in all participating processes, with each process contributing its local tensor, and all processes will eventually get the result of the reduction operation.

[0049] The homologous chip individual communication module 22 is for a homologous scenario, and can call a communication library (such as NCCL, RCCL, HCCL, etc.) of a corresponding manufacturer according to different hardware.

[0050] The heterogeneous chip unified communication module 23 is for a heterogeneous scenario, and includes a unified communication programming model, a unified topology component 24, and a unified communication component 21. The unified communication programming model is used for communication algorithm development, and mainly includes two technical paths: 1) generating a set communication algorithm implementation under a specific topology architecture through a communication algorithm synthesizer (unified communication algorithm synthesis module 232); and 2) developing a communication algorithm through a programming primitive provided by a communication algorithm abstraction interface (unified communication algorithm abstraction module 231), supporting cross-node RDMA cluster communication of different chips.

[0051] The unified communication algorithm abstraction module 231 can define a set of communication primitives to realize the portable programming of the collective communication algorithm, and is configured to uniformly abstract the communication primitives used in the heterogeneous chip cluster communication. In the unified communication programming model provided in the present application, a two-layer design can be used, i.e., a Cluster to Cluster (C2C) mode and a Primitive mode. In some embodiments, the communication primitives include basic operations, communication negotiation operations, and communication calculation operations.

[0052] In the C2C mode, by referring to the implementation of cross-chip heterogeneous P2P communication, the cards in the same Cluster use the vendor self-developed collective communication algorithm (flagcxAllReduce, etc.), the cross-Cluster heterogeneous communication uses flagcxSend / Recv, and the calculation operation provides a unified interface (flagcxSum, flagcxProd, flagcxMinMax, etc.). Different implementations can be selected according to the type of the underlying chip, or a portable programming model already available in the industry, such as SCCL, Kokkos, RAJA, etc., can be used to realize the calculation operation. Figure 3 A schematic diagram of realizing the inter-cluster communication of the cluster kernel through the cross-chip cluster communication system is shown.

[0053] The following is an example of realizing the global reduction (allreduce) through the C2C mode.

[0054] First, initialize the homogeneous subgroup communication group: use the flagcxCommInitRank function of the flagcx API to initialize the communication group of the homogeneous subgroup. Then initialize the heterogeneous peer communication group: also use the flagcxCommInitRank function of the flagcx API to initialize the communication group of the heterogeneous peer. Perform the reduce-scatter operation of the subgroup: use the flagcxReduceScatter function to perform the reduce-scatter operation in the subgroup. Perform the sending / receiving operation of the heterogeneous peer: use the flagcxSend and flagcxRecv functions to perform the data sending and receiving between the heterogeneous peers. Perform the local reduce operation of the heterogeneous peer: here, two options are provided. One is to use the reduction operation function provided by flagcx, such as flagcxSum. The other is to use the kernel of a third party, such as SYCL. For SYCL, we need to submit a handler of the processor and execute the reduction operation in parallel through the parallel_for function. Perform the all-gather operation of the subgroup: finally, use the flagcxAllGather function to perform the all-gather operation of the subgroup. Through the above steps, the global reduction operation in the C2C mode can be successfully implemented.

[0055] In the primitive mode, the following several primitive operations are abstracted for the communication kernel: basic operations, such as barrier, set_buffer, mod_buffer, etc.; communication negotiation operations, such as wait, post, etc.; and communication calculation operations, such as reduce_copy, etc. The implementation of the primitives on different hardware is developed by the manufacturer or the user, and the user can develop different collective communication algorithms according to the free combination of the above primitives.

[0056] The following is an example of implementing global reduction (allreduce) through the primitive mode.

[0057] First, data buffer preparation is performed: the set_buffer primitive is used to prepare the data buffer to store data for subsequent communication and computation operations. Then, a synchronization barrier is performed to ensure that all participating nodes have reached a synchronized state before starting the subsequent operations through the barrier primitive. Wait for peer nodes: the wait primitive is used to wait for a signal from one or more peer nodes, which is part of the communication negotiation, ensuring that all nodes are ready for data exchange. Perform the second synchronization barrier: the barrier primitive is used again to ensure that all nodes have reached a synchronized state again before proceeding to the next step. Perform the reduce copy operation: the reduce copy operation is performed through the reduce_copy primitive, which is part of the communication computation. In this step, the data of each node will be subjected to a specified reduction operation (such as sum, maximum, etc.) and the result will be copied to the specified buffer. Perform the third synchronization barrier: the barrier primitive is used to ensure that all nodes are synchronized again after the reduce copy operation is completed. Finally, post peer nodes: the post primitive is used to send a signal to one or more peer nodes that the current node's operation has been completed or is ready to receive data. Through the above steps, the global reduction operation in the primitive mode can be successfully implemented.

[0058] The design purpose of the unified communication algorithm synthesis module 232 is to automatically generate an optimal communication algorithm according to the cluster topology structure, and to generate an optimal communication algorithm according to the topology structure of the heterogeneous chip cluster communication. This module can use a specific topology structure to use CAS (communication algorithm synthesizer) to automatically construct a homogeneous / heterogeneous set communication algorithm.

[0059] The way to generate an optimal communication algorithm based on a specific topology structure can refer to existing technologies, such as the MSCCL technology described in Guiding Collective Algorithm Synthesis using Communication Sketches. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) (pp. 593-612). In this application, it will not be repeated.

[0060] The unified topology component 24 defines a set of cluster topology standard definitions, mainly including logical devices (physical device abstraction component 241) and high-speed interconnection links (interconnection link abstraction component 242). Among them, the physical device abstraction component 241 is used to provide a physical device abstraction tool. The interconnection link abstraction component 242 is used to provide an interconnection link abstraction tool.

[0061] The physical device abstraction component 241 can abstract a chip into a joint structure, which can represent different physical devices, including an accelerated processing unit (APU), a PCT switch, a CPU (NUMA network node), a NET (network interface card, NIC), and the like. The APU can represent different acceleration devices, such as Nvidia GPU, AMD GPU, Intel GPU, Google TPU, Ascend NPU, Kunlun XPU, FPGA, ASIC, and the like.

[0062] In one example, the chip can be physically abstracted in the following way: first, detect the NUMA node of the CPU, read the architecture, manufacturer, model, and the like of the CPU; read the identifier of the NUMA node; retrieve the affinity information of the CPU to determine the available CPU cores and the NUMA nodes to which they belong. Second, rely on the information and tools provided by the device manufacturer to identify various accelerator devices in the system, such as GPU, TPU, NPU, and the like. Then, identify the PCI devices in the system, for example, in one example, the device with code "0x060400" can be identified as a PCI switch; the device with code "0x068000" is classified as NVS; the device with code "0x068001" is classified as CPU; the device with code "0x03" is classified as GPU; and the device with code "0x02" is classified as NIC. Finally, identify the network interface card connected to the PCI, for example, for the network interface card directly connected to the device, the manufacturer-provided API can be used for identification. Through the above steps, the physical device abstraction component can effectively detect and abstract various physical devices in the system.

[0063] The interconnect link abstraction component 242 can divide the interconnect links designed in the system architecture into five types, including local (LOC), cache-coherent interconnect (CCI), peripheral component interconnect (PCI), symmetric multiprocessor interconnect (SYS), and network interconnect (NET). The CCI is, for example, NVlink, GenZ, CXL, CCIX, CAPI, and the like. The SYS is, for example, quick path interconnect (QPI), uniform path interconnect (UPI), and the like.

[0064] In one example, the multiple chip related links can be probed in the following way: first, use the API provided by the device manufacturer to obtain the CCI link information. Then, detect and read the PCT link information, including the current link speed, the current link width, the maximum link speed and / or the maximum link width, etc. Next, detect the SYS link information, which depends on the CPU architecture, manufacturer and model, and the following are some reference values: POWER architecture: 32.0 GB / s; ARM architecture: 6.0 GB / s; Intel architecture: SKL 10.0 GB / s, other models: 6.0 GB / s; AMD architecture: 16.0 GB / s; ZHAOXIN architecture: YONGFENG 9.0 GB / s, other models: 6.0 GB / s. Finally, detect the NET link information, including: reading the rate of the InfiniBand device, reading the rate of the Ethernet device, etc. For some specific devices, the API provided by the device manufacturer can also be used to obtain the link information. Through the above steps, effective probing of the interconnection link can be achieved.

[0065] After completing the physical device abstraction and interconnection link abstraction, the unified topology component 24 can construct the cluster topology. It follows the following steps: traverse and add PCI-based unified topology nodes; and traverse and add CCI-based unified topology nodes.

[0066] In the process of adding PCI-based unified topology devices, traverse the CPU nodes, add CPU devices for each node; and traverse the child nodes, if the child node is a PCI device, add the device and its child devices, if the child node is a NET device, add the connected NIC device and NET endpoint, and establish the PCI type link between CPU-NIC; traverse all CPUs to establish SYS type links between them.

[0067] In the process of adding PCI devices and their child devices, add PCT devices. Then traverse the child devices, if the child device is an APU device, add the device; if the child device is a NET device, add the connected NIC device; if the child device is a PCI device, add the device; finally, establish the PCI type link between the parent node and the child node.

[0068] In the process of adding NET devices, add NIC devices, add NET devices as endpoints, and establish NET type links between NIC and NET.

[0069] In the process of adding CCI-based unified topology devices, traverse the CCI devices, and establish CCI type links between AFU and its remote devices (such as AFU or CPU) for each device.

[0070] In this way, the cluster topology can be constructed.

[0071] The unified communication component 21 mainly abstractly designs communication primitives, supports remote direct memory access (RDMA), transmission control protocol (TCP), shared memory (SHM), point-to-point (P2P), and the like. Through the component, different chip cross-node RDMA P2P communication based on IB / RoCE can be realized, and cross-chip cluster communication systems are allowed to fully utilize various hardware resources to realize interconnection communication. The unified communication component 21 includes at least one of the following: a unified communication transmission component, configured to provide an intra-node and inter-node communication abstraction tool; a unified communication protocol component, configured to provide an abstraction tool of different communication protocols; a unified communication memory component, configured to provide an abstraction and encapsulation tool of memory management; or a unified communication service component, configured to provide a communication service interface. The unified communication component 21 is used for remote direct memory access (RDMA) between multiple nodes.

[0072] The unified communication transmission component 211 abstracts three types of transmissions for all intra-node and inter-node high-speed interconnection communication, which are point-to-point (P2P), shared memory (SHM), and network (NET). All transmissions define different implementations of the following transmission primitives:

[0073] Initialization (setup): used for initializing related resources required by communication transmission.

[0074] Connection (connect): used for establishing a connection of a communication opposite end.

[0075] Release (free): used for releasing related resources used by communication transmission.

[0076] Shared initialization (sharedInit): used for initializing shared resources (if any) for proxy operation.

[0077] Progress (progress): used for following the status of proxy operation.

[0078] Registration (register): used for registering a buffer required by communication.

[0079] Deregistration (deregister): used for canceling a registered communication buffer.

[0080] Figure 4 is a schematic diagram of a communication architecture related to a heterogeneous chip cluster communication. In Figure 4In the figure, the transmission with label 1 can be an intra-node interconnection transmission complying with a high-speed interconnection interface or protocol such as NVlink, XPU link, HCCS, MetaLink, etc. The transmission with label 2 can be an intra-node interconnection transmission complying with the PCIe high-speed serial computer expansion bus standard. The transmission with label 3 can be an intra-CPU node or inter-CPU node interconnection transmission complying with the QPI / UPI interconnection interface. The transmission with label 4 can be an inter-node interconnection transmission complying with the IB or Ethernet network communication protocol. The point-to-point communication can be implemented through CUDA, CANN, ROCm communication library, chip architecture or computing platform, etc. The shared memory can be implemented through the POSIX standard, SysV mechanism, XPMEM mechanism, etc. The network communication (NET) can be implemented through the RDMA technology, TCP / IP mechanism, etc.

[0081] The focus of the unified communication transmission can be on intra-node P2P and inter-node RDMA. Based on the unified abstract design, the following can be implemented: 1) point-to-point communication of homologous chips in a node; 2) point-to-point communication of homologous chips across nodes; and 3) point-to-point communication of heterologous chips across nodes.

[0082] The point-to-point communication of homologous chips in a node can directly call the point-to-point (P2P) device-to-device (D2D) interface provided by the manufacturer to be implemented. The point-to-point communication of homologous / heterologous chips across nodes is implemented by directly passing through inter-node RDMA in a kernel bypass manner. As shown in the figure, the implementation of the inter-node RDMA will be described in detail below. Figure 5

[0083] First, device initialization is performed. The InfiniBand (IB) interface and devices available in the system are detected and identified. A list of all InfiniBand devices in the system is obtained. The obtained device list is traversed, and each device is processed. The number of processed devices does not exceed a predefined maximum number of devices. For each device, the context of the device is opened, and the attributes of the device are queried and obtained. Each physical port of the device is traversed, and for each port, the attributes of the port are queried and obtained. According to the queried device and port attributes, the configuration of the device is set, and if necessary, a merging operation of the device is performed. After the processing of the device is completed, the context of the device is closed. After the processing of all devices is completed, the device list is released.

[0084] ​Then connection setup is performed, including: 1) listening for connection requests: initializing a listening port using socket technology, preparing to receive connection requests, and causing the system to listen for inbound connection requests on a specific port. 2) establishing a connection: initializing a send communication, changing the current state from START to CONNECT; establishing a connection with the remote end using socket technology; allocating a protection domain for each consolidated device and creating a completion queue to manage completion events for RDMA operations, initializing the protection domain (PD) and completion queue (CQ); creating a set of queue pairs for the communication session and querying the ECE (Explicit Congestion Notification) configuration; registering memory regions for data structures used in the communication to support RDMA access. This includes registering FIFO queues and configuring support for RoCE; modifying the state of the queue pairs, first setting to IBV_QPS_RTR (ready to receive) and then to IBV_QPS_RTS (ready to send) to start the RDMA communication; updating the current communication state according to different stages of the RDMA communication, including SEND (sending data), CONNECTING (connecting), and CONNECTED (connected). 3) accepting a connection: initializing a receive communication and changing the current state from START to ACCEPT; accepting connection requests from the remote end using socket technology; allocating a protection domain for each device on the receiving end and creating a completion queue; creating a queue pair for the communication session and setting its state to IBV_QPS_RTR and IBV_QPS_RTS to support RDMA data reception and transmission; for devices that support GPU direct memory access, registering FIFO queues and Flush dummy buffers to enable GDR (GPU Direct RDMA) support; updating the current communication state according to different stages of the RDMA communication, including RECV (receiving data), SEND (sending data), and PENDINGREADY (waiting for a remote end ready signal).

[0085] Memory registration is then performed, providing an external interface that internally calls the corresponding process to register the memory.

[0086] Non-blocking data transfer (send / receive) operations are then performed and the operation status is checked. Specifically, a non-blocking send operation sends data from local memory to remote memory, but does not wait for the operation to complete. A send request can be placed in the send queue, ready to be sent to the remote node, by calling a specific function. This function needs to specify the context of the send communication, the data to be sent, the size of the data, the tag, and the memory handle. A non-blocking receive operation receives data from remote memory to local memory, but does not wait for the operation to complete. A receive request can be placed in the receive queue, ready to receive data from the remote node, by calling a specific function. This function needs to specify the context of the receive communication, the number and location of the receive buffers, the size of the buffers, the tag, and the memory handle. In addition, a non-blocking flush can also be performed, which sends a flush request to the remote memory to ensure that all previous non-blocking send operations have completed. A flush request can be placed in the send queue, ready to be sent to the remote node, by calling a specific function. This function needs to specify the context of the receive communication, the number and location of the flush requests, the size of the requests, the tag, and the memory handle. In addition, an operation status check can also be performed, which checks the status of non-blocking send, receive, or flush operations to determine whether they have completed. The completion queue is checked by calling a specific function to obtain the number and size of completed operations. This function needs to specify the requests to be checked, the completion status, and the size of the operations.

[0087] After the transfer is completed, clean-up and close operations after the completion of the communication can be performed, such as closing the send communication, closing the receive communication, and closing the listen communication. The close send communication operation releases resources related to the send communication after the send operation is completed, including destroying the queue pair (QP) related to the send communication, unregistering the memory region previously registered for the send operation, destroying the completion queue used to manage the completion events of the send operation, and releasing the protection domain previously allocated for the send operation. The close receive communication operation releases resources related to the receive communication after the receive operation is completed, with steps similar to closing the send communication. The close listen communication operation releases resources related to the listen communication after the listen operation is completed, including closing the Socket connection previously used to listen for connection requests. In this way, after the completion of the RDMA communication, various types of resources related to the communication can be gradually released, including the queue pair, the memory region, the completion queue, and the protection domain, to ensure efficient use and management of system resources.

[0088] In some embodiments, the unified communication component 21 can further include a unified communication protocol module 212. This module provides a set of high-level abstraction over different communication protocols, enabling different communication technology support. For example, it includes: Sync / Async Communication, Blocking / Non-Blocking Operation, Eager / Rendezvous transport protocol, Tag-Matching communication mode, RMA / Atomics communication mode / operation, Active Messages communication mode, Stream communication mode, etc.

[0089] In some embodiments, the unified communication component 21 can further include a unified communication memory module 213. This module is an abstraction and encapsulation of memory management on different architectures, including CDA, CANN, RoCM, etc. (mainly related to GDR enablement), as well as memory operations for host memory (e.g. mmap, munmap, mremap, shmat, etc. functions).

[0090] In some embodiments, the unified communication component 21 can further include a unified communication service module 214. This module provides a set of unified communication service interfaces, including data types, network interfaces (IB / RoCE, TCP / IP Socket), debugging interfaces, timing interfaces, analysis interfaces, statistical interfaces, adaptive adjustment interfaces, fault-tolerant resilience interfaces, etc. The implementation of the network transmission interface is described as follows:

[0091] 1) Transmission establishment is performed. First, transmission resources are set and allocated, and proxy settings, request buffers, send flags, request sizes, response buffers, response sizes, and completion states, etc. are specified. Send branch: if the send flag is true, enter the send branch; if the proxy function is enabled, enter the proxy branch. In the proxy branch, a specific function is used to start listening for connection requests, which requires specifying the listening port and address, as well as other related parameters; if the proxy function is not enabled, perform other related settings and initialization operations. Receive branch: if the send flag is false, enter the receive branch and perform related settings and initialization operations.

[0092] 2) Transmission association. First, establish a transmission connection, call a specific function, specify proxy settings, request buffer, send flag, request size, response buffer, response size, and completion status. Send branch: if the send flag is true, enter the send branch. If the proxy function is enabled, enter the proxy branch. In the proxy branch, call a specific function to start connecting the opposite end. This function needs to specify the target port and address, as well as other related parameters; if the proxy function is not enabled, call the same function to start connecting the opposite end. Receive branch: if the send flag is false, enter the receive branch. If the proxy function is enabled, enter the proxy branch. In the proxy branch, call a specific function to start receiving the connection. This function needs to specify the listening port and address, as well as other related parameters. If the proxy function is not enabled, call the same function to start receiving the connection.

[0093] 3) Transmission execution. First, specify proxy settings and send flag. Send branch: if the send flag is true, enter the send branch. If D2D is enabled, copy data from the source address to the buffer; asynchronously send data; and test the send state. Receive branch: if the send flag is false, enter the receive branch. Asynchronously receive data; test the receive state; asynchronously flush data; and, if D2D is enabled, copy data from the buffer to the source address.

[0094] 4) Transmission release. First, specify proxy settings and send flag. Send branch: if the send flag is true, enter the send branch. If the proxy function is enabled, call a specific function to close the send channel, which is responsible for releasing all resources used in the send process, including the buffer and the connection handle. If the proxy function is not enabled, other resource release operations may need to be performed to ensure that all send-related resources are properly cleaned up. Receive branch: if the send flag is false, enter the receive branch. Call a specific function to close the receive channel. This function is also responsible for releasing all resources used in the receive process.

[0095] The application also provides a cross-core cluster communication method. As shown in the figure, the method comprises operations S301 to S304. Figure 6

[0096] In S301, information of a plurality of chips is obtained.

[0097] In S302, the type of the current cross-core cluster communication is determined according to the information of the plurality of chips.

[0098] In S303, when the cross-core cluster communication is a same-source chip cluster communication, a same-source chip individual communication module is called to implement cluster communication.

[0099] ​Specifically, the communication with each chip can be initialized by calling the adapter matched with each chip; and the collective communication can be performed using the native collective communication algorithm of each chip or the uniform communication algorithm of each chip.

[0100] In S304, when the cross-chip cluster communication is a heterogeneous chip cluster communication, a heterogeneous chip uniform communication module is called to implement the cluster communication.

[0101] Specifically, a uniform abstract topology structure among the chips can be constructed; an optimal communication algorithm can be generated according to the uniform abstract topology structure; and the cluster communication can be performed according to the optimal communication algorithm under the uniform abstract topology structure.

[0102] The present application also provides a topology construction method for cross-chip cluster communication. As shown in Figure 7 the method includes operations S401 to S404.

[0103] In S401, the information of a plurality of chips involved in a service requirement is obtained.

[0104] In S402, the plurality of chips are physically device abstracted to obtain the node information involved in the communication task.

[0105] In S403, the interconnection link information between each node is detected and determined.

[0106] In S404, the cluster communication topology is constructed based on the node information and the interconnection link information.

[0107] The present application also provides a cluster-to-cluster communication method. As shown in Figure 8 the method includes operations S501 to S505.

[0108] In S501, the same-source nodes are divided to obtain a plurality of sub-clusters, and the plurality of sub-clusters include heterogeneous nodes.

[0109] In S502, a reduction dispersion operation of the sub-cluster is performed.

[0110] In S503, data exchange between the sub-clusters is performed, and the reduction dispersion operation result of the sub-cluster is sent to other sub-clusters.

[0111] In S504, each node in the sub-cluster performs a reduction operation to reduce the received data and the local data.

[0112] In S505, a global collection operation is performed in the sub-cluster, and the reduction result on each node in the sub-cluster is collected on all nodes in the sub-cluster.

[0113] Figure 9 is a schematic block diagram of an electronic device 600 provided by an embodiment of the present application. As shown inFigure 9 As shown, the electronic device 600 includes a processor 601 and a memory 602, and the processor 601 is communicatively connected with the memory 602. The electronic device 600 can be, for example but not limited to, various high-performance computing devices, servers, supercomputers, cloud computing platforms, big data processing devices, etc. In some embodiments, the electronic device 600 can further include a transceiver module for transmitting / receiving data, or only include a transmitting module for transmitting data, or only include a receiving module for receiving data. The memory 602 of the electronic device 600 is used to store program instructions, which can be executed by the processor 601 to implement the cross-core cluster communication method described in any of the foregoing embodiments.

[0114] It should be understood that the processor of the embodiments of the present application can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the method embodiments described above can be completed by an integrated logic circuit or an instruction in the form of software in the processor.

[0115] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. It should be noted that the memory of the system and method described herein is intended to include, but not limited to, these and any other suitable types of memory. The embodiments of the present application also provide a computer readable storage medium for storing a computer program.

[0116] Optionally, the computer readable storage medium can be applied to the electronic device in the embodiments of the present application, and the computer program causes the computer to execute the corresponding processes realized by the electronic device in each method of the embodiments of the present application. For the sake of brevity, they will not be described here.

[0117] The embodiments of the present application also provide a computer program product, which includes computer program instructions.

[0118] Optionally, the computer program product can be applied to the electronic device in the embodiments of the present application, and the computer program instructions cause the computer to execute the corresponding processes realized by the electronic device in each method of the embodiments of the present application. For the sake of brevity, they will not be described here.

[0119] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0120] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A cross-chip trunking communication system, characterized in that, include: A chip interface module is used to acquire information from multiple chips, and the information from the multiple chips is used to determine whether the multiple chips are of the same origin or different origin. Individual communication modules for homogeneous chips, including adapters compatible with various chips, for cluster communication of homogeneous chips; as well as A unified communication module for heterogeneous chips, used for cluster communication of heterogeneous chips.

2. The system as described in claim 1, characterized in that, The heterogeneous chip unified communication module includes a unified communication algorithm abstraction module, which is used to uniformly abstract the communication primitives used in the heterogeneous chip cluster communication.

3. The system as described in claim 2, characterized in that, The communication primitives include basic operations, communication negotiation operations, and communication computation operations.

4. The system as described in claim 1, characterized in that, The heterogeneous chip unified communication module includes a unified communication algorithm synthesis module, which is used to generate the optimal communication algorithm according to the topology of the heterogeneous chip cluster communication.

5. The system as claimed in claim 1, wherein, It also includes a unified topology component, which includes: Physical device abstraction components are used to provide physical device abstraction tools; and The interconnection abstraction component provides interconnection abstraction tools.

6. The system of claim 1, wherein, It also includes a unified communications component, which includes at least one of the following: Unified communication transport components are used to provide abstract tools for communication within and between nodes; Unified communication protocol components provide abstraction tools for different communication protocols; Unified communications memory component, used to provide abstraction and encapsulation tools for memory management; The Unified Communications Service component is used to provide communication service interfaces.

7. The system as described in claim 6, characterized in that: The unified communication component is used for interconnection between multiple nodes via Remote Direct Memory Access (RDMA).

8. The system as described in claim 1, characterized in that, Also includes: The framework integration module is used to interface with the business framework.

9. A cross-chip trunking communication method, comprising: Obtain information from multiple chips; Based on the information from the multiple chips, determine the type of current cross-chip cluster communication; When the current cross-chip cluster communication is same-source chip cluster communication, the individual communication module of the same-source chip is called to realize cluster communication. as well as When the current cross-chip cluster communication is heterogeneous chip cluster communication, the unified communication module for heterogeneous chips and the individual communication module for homogeneous chips are invoked to realize cluster communication.

10. The method of claim 9, wherein, The method of calling the individual communication module of the same chip to achieve cluster communication includes: Invoke the adapters matched to each chip and initialize communication with each chip; and Collective communication is performed using the native collective communication algorithm of each chip.

11. The method of claim 9, wherein, The method of calling the heterogeneous chip unified communication module to achieve cluster communication includes: Construct a unified abstract topology among the chips; Generate the optimal communication algorithm based on the unified abstract topology; and Cluster communication is performed according to the optimal communication algorithm under the unified abstract topology.

12. An electronic device comprising a processor and a memory, the memory being used to store program instructions which, when executed by the processor, are used to implement the method described in any one of claims 9 to 11.

13. A readable storage medium for storing program instructions, wherein, When the program instructions are executed by the processor, they are used to implement the method described in any one of claims 9 to 11.