RDMA (Remote Direct Memory Access) data transmission method, network equipment, system and electronic equipment

By introducing hardware engines to the network equipment, assembling and transmitting working elements, the problem of low RDMA data transmission efficiency in the MoE model is solved, efficient and low-latency data transmission is achieved, and dependence on CPU/GPU resources is avoided.

CN120196573AActive Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202510338839.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the generative large language model, the MoE model has low RDMA data transmission efficiency due to the large number of scattered small data block transmission requirements, and relies on CPU/GPU resources, resulting in performance bottlenecks.

Method used

By introducing a first engine and a second engine in a network device, the first engine assembles the working elements based on the hardware offloading of the asynchronous copy instruction set and transmits them to the second engine. The second engine stores the working elements in the WQE Buffer and transmits them to the RDMA engine to realize data processing and feedback.

Benefits of technology

It reduces the number of data handling times and delays during RDMA data transmission, realizes the decoupling of RDMA data transmission tasks and CPU/GPU resources, and provides an efficient and low-latency RDMA data transmission solution without relying on CPU/GPU resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196573A_ABST
    Figure CN120196573A_ABST
Patent Text Reader

Abstract

The invention provides an RDMA data transmission method, network equipment and a system, and relates to the technical field of communication, in particular to the technical fields of chip processors, computing power clusters, ensemble communication operation, large models and the like. The method is applied to a network device comprising an xPU and an RNIC, the xPU comprises a first engine, and the RNIC comprises a second engine communicating with the first engine, a WQE Buffer and an RDMA engine communicating with the second engine. The method comprises the steps that a first engine assembles a WE based on a hardware unloading asynchronous copy instruction set and transmits the WE to a second engine; the second engine stores the WE in a WQE Buffer and transmits the WE to the RDMA engine; the RDMA engine executes data processing based on the WE, wherein the data processing comprises data transmission, memory access and queue management; the first engine receives feedback of data processing transmitted via the second engine. The RDMA data transmission task is directly unloaded to the special engine through the hardware asynchronous instruction set, decoupling of the RDMA data transmission task and computing resources of the system is achieved, and an efficient and low-delay RDMA data transmission solution which does not depend on CPU / GPU resources is provided for data transmission intensive scenes such as an MoE model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technologies, specifically to technical fields such as chip processors, computing power clusters, collective communication operations, large models, etc., and particularly relates to RDMA data transmission methods, network devices, systems, electronic devices, computer-readable storage media, and computer program products. Background Art

[0002] In generative large language models, the MoE (Mixture of Experts) model can process different tasks through multiple "expert" networks. In the MoE scenario, there are a large number of scattered small data block transmission requirements.

[0003] The RDMA (Remote Direct Memory Access) protocol is a network protocol that allows a computer to directly access the memory of a remote computer and is commonly used in scenarios such as high-performance computing (HPC), data centers, storage networks, etc. Summary of the Invention

[0004] Embodiments of the present disclosure propose an RDMA data transmission method, network device, system, electronic device, computer-readable storage medium, and computer program product.

[0005] In a first aspect, embodiments of the present disclosure propose an RDMA data transmission method, which is applied to a network device including a processor xPU and an RDMA network interface controller RNIC. The xPU includes a first engine, and the RNIC includes a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine. The method includes: the first engine assembles a work element WE based on a hardware offload asynchronous copy instruction set and transmits the WE to the second engine; the second engine stores the WE in the WQE Buffer and transmits it to the RDMA engine; the RDMA engine performs data processing based on the WE, and the data processing includes data transmission, memory access, and queue management; the first engine receives feedback on the data processing transmitted via the second engine.

[0006] Second aspect, an embodiment of the present disclosure proposes another RDMA data transmission method, which is applied to a network system including multiple network devices. The network device includes a processor xPU and an RDMA network interface controller RNIC. The xPU includes a first engine, and the RNIC includes a second engine communicating with the first engine, a work queue buffer WQEBuffer, and an RDMA engine communicating with the second engine. The method includes: the first engine of the first network device assembles a work element WE based on a hardware offloading asynchronous copy instruction set and transmits the WE to the second engine of the first network device; the second engine of the first network device stores the WE in the WQE Buffer and transmits it to the RDMA engine of the first network device; the RDMA engine of the first network device initiates a request to the second network device based on the WE; the second network device completes the request and feeds back to the first engine of the first network device, where the first network device and the second network device are different network devices among the multiple network devices.

[0007] Third aspect, an embodiment of the present disclosure provides an RDMA data transmission network device, which includes: a processor xPU, including a first engine; an RDMA network interface controller RNIC, including a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine, where the first engine is configured to assemble a work element WE based on a hardware offloading asynchronous copy instruction set and transmit the WE to the second engine; the second engine is configured to store the WE in the WQE Buffer and transmit it to the RDMA engine; the RDMA engine is configured to perform data processing based on the WE and transmit feedback of the data processing to the first engine via the second engine, and the data processing includes data transmission, memory access, and queue management.

[0008] Fourth aspect, an embodiment of the present disclosure provides an RDMA data transmission network system, which includes: a plurality of network devices, where the network device includes a processor xPU and an RDMA network interface controller RNIC. The xPU includes a first engine, and the RNIC includes a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine. Among them, the first engine of the first network device is configured to assemble a work element WE based on a hardware offloading asynchronous copy instruction set and transmit the WE to the second engine of the first network device; the second engine of the first network device is configured to store the WE in the WQE Buffer and transmit it to the RDMA engine of the first network device; the RDMA engine of the first network device is configured to initiate a request to the second network device based on the WE; the second network device is configured to complete the request and feedback it to the first engine of the first network device, where the first network device and the second network device are different network devices among the plurality of network devices.

[0009] Fifth aspect, an embodiment of the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can implement the RDMA data transmission method described in the first aspect and the second aspect.

[0010] Sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and when the computer instructions are used to make a computer execute, the computer can implement the RDMA data transmission method described in the first aspect and the second aspect.

[0011] Seventh aspect, an embodiment of the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the computer program can implement the RDMA data transmission method described in the first aspect and the second aspect.

[0012] In the RDMA data transmission solution provided by the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE (Device Work Queue Element) at hardware speed and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, by means of the hardware asynchronous instruction set, the RDMA data transmission task is directly offloaded to the dedicated hardware engine, realizing the decoupling of the RDMA data transmission task from the CPU / GPU resources, thus providing an RDMA data transmission solution that is efficient and low-latency and does not rely on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent: Figure 1 An exemplary system architecture to which the present disclosure can be applied; Figure 2 A block diagram of an RDMA data transmission network system architecture provided for an embodiment of the present disclosure; Figure 3 A flowchart of an RDMA data transmission method provided for an embodiment of the present disclosure; Figure 4 A block diagram of another RDMA data transmission network system architecture provided for an embodiment of the present disclosure; Figure 5 A flowchart of another RDMA data transmission method provided for an embodiment of the present disclosure; Figure 6 A block diagram of yet another RDMA data transmission network system architecture provided for an embodiment of the present disclosure; Figure 7 A schematic structural diagram of an electronic device suitable for executing the RDMA data transmission method provided for an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0016] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0017] Figure 1 An exemplary system architecture 100 is shown that can apply the embodiments of the RDMA data transmission method, network device, system, electronic device, and computer-readable storage medium of the present disclosure.

[0018] As Figure 1 shown, the system architecture 100 may include a server 110, an RNIC 120, and an xPU 130. The server 110 may include a CPU (Central Processing Unit), a host memory, etc., and can provide various services through various built-in applications. For example, the server 110 can provide a running environment and resource management for the xPU 130, allocate memory, and manage its communication with the RNIC (Remote Direct Memory Access Network Interface Controller) 120.

[0019] In the present disclosure, "x" in xPU (x Processing Unit, which can be regarded as a general term for various processors) is a wildcard, and xPU can be regarded as a general term for various types of processors. For example, xPU 130 may include at least one of a CPU, a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Processing Unit), a DPU (Data Processing Unit), a VPU (Vision Processing Unit), a QPU (Quantum Processing Unit), and an APU (Accelerated Processing Unit). In addition, xPU 130 generally appears as hardware, and in special scenarios (such as simulation scenarios), it may also appear as software or a software running product, which is not specifically limited in the present disclosure.

[0020] In addition, the system architecture 100 may include multiple xPUs, such as a first xPU 1311, a second xPU 1312, a third xPU 1313, a fourth xPU 1314, and a fifth xPU 1315, etc., which exist in the form of a graphics card. It should be noted that Figure 1 the form of the RNIC 120 is simplified, and only a small number of xPUs 130 are shown as examples. However, those skilled in the art can set the number, type, and connection relationship of the server 110, RNIC 120, and xPU 130 in the system architecture 100 according to actual needs, which is not limited in the present disclosure.

[0021] The RNIC 120 can be used to provide communication between the server 110 and the xPU 130. The RNIC 120 may include various connection types, and the present disclosure does not limit the hardware form, interface type, and protocol support adopted by the RNIC 120.

[0022] In addition, the user can use a terminal device to interact with the server 110 through, for example, the RNIC 120 to receive or send messages, etc. Various applications for realizing information communication between the two can be installed on the server 110, the xPU 130, and the terminal device, such as training task distribution applications, training strategy optimization applications, instant messaging applications, etc.

[0023] The terminal device and the server 110 can be either hardware or software. When the terminal device is hardware, it can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop computers, desktop computers, etc.; when the terminal device is software, it can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, which is not specifically limited herein. When the server 110 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can also be implemented as a single server; when the server 110 is software, it can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, which is not specifically limited herein. The graphics processor that constitutes the xPU 130 usually appears as hardware, but of course it can also appear as software or a software running product in special scenarios (such as simulation scenarios), which is not specifically limited herein.

[0024] The server 110 can provide various services through various built-in applications. For example, the server 110 can include at least one of a file server, a database server, a mail server, a Web server, an application server, a game server, IaaS (Infrastructure as a Service), PaaS (Platform as a Service), SaaS (Software as a Service), and an AI training server. Among them, the file server can centrally store and share files (such as an enterprise internal document library); the database server can run database management systems such as MySQL and PostgreSQL; the mail server can handle the sending, receiving, and storage of emails; the Web server can host websites and Web applications (such as Apache and Nginx); the application server can run middleware; the game server can handle the real-time interaction logic of multiplayer online games; IaaS can provide virtual machines, storage, and network resources; PaaS can provide a development platform and tools; SaaS can deliver software through the network; the AI training server can be equipped with GPU / TPU to accelerate the training of deep learning models, etc.

[0025] Figure 2 A block diagram of an RDMA data transmission network system architecture 100 provided for the embodiments of the present disclosure.

[0026] As Figure 1 and Figure 2As shown, the system architecture 100 may include network devices, where the network devices may include an RNIC 120 and an xPU 130. The xPU 130 may include a first engine 131. The RNIC 120 may include a second engine 121 that communicates with the first engine 131, an RDMA engine 122 that communicates with the second engine 121, and a WQE Buffer (Work Queue Element Buffer) 123.

[0027] In addition, the xPU 130 may further include a core computing unit 132 and a memory management unit 133, where the types of the core computing unit 132 and the memory management unit 133 of different types of xPU 130 may be different. For example, when the xPU 130 includes a GPU, its core computing unit 132 may include an SM (Streaming Multiprocessor); when the xPU 130 includes a CPU, its core computing unit 132 may include an Arithmetic Logic Unit (ALU), a Control Unit (CU), etc.

[0028] In addition, when the xPU 130 executes a program, the core computing unit 132 generates a virtual address, and the memory management unit 133 may be responsible for converting the virtual address into a physical address. Optionally, when the xPU 130 includes a GPU or a CPU, its memory management unit 133 may include an MMU (Memory Management Unit).

[0029] The RDMA engine 122 of the RNIC 120 may perform various data processes, where the data processes may include data transmission, memory access, queue management, etc.

[0030] It should be understood that Figure 1 and Figure 2 the numbers of servers, networks, network devices, etc. in

[0031] Figure 3 is only illustrative. According to the implementation requirements, there may be any number of servers, networks, and network devices. Figure 3 For a flowchart of an RDMA data transmission method provided by an embodiment of the present disclosure, with reference to Step 201: The first engine assembles a work element WE based on a hardware offloading asynchronous copy instruction set and transmits the WE to the second engine; This step aims to trigger a communication protocol including a hardware offload instruction set by a hardware engine (the first engine), where the hardware engine can be directly integrated inside the xPU and can assemble WE (Work Element) by itself. Through, for example, PCIe BAR space mapping, the assembled WE can be directly output to the second engine located in the RNIC. This can reduce the dependence of RDMA data transmission on CPU performance or GPU performance. In addition, it can also reduce the number of data transfers during RDMA data transmission and reduce the data latency during the above transmission process.

[0032] Specifically, the hardware offload asynchronous copy instruction set may include multi-level synchronous control data transmission instructions. Among them, the multi-level synchronous control data transmission instructions may include thread-level synchronous control data transmission instructions, thread group-level synchronous control data transmission instructions, storage block-level synchronous control data transmission instructions, and global-level synchronous control data transmission instructions.

[0033] For example, the hardware offload asynchronous copy instruction set may include: ibcp.async(dst.addr, src.addr, length); ibcp.async.wait_thread; ibcp.async.wait_wrap; ibcp.async.wait_block; ibcp.async.wait_all.

[0034] Among them, ibcp.async(dst.addr, src.addr, length) is an instruction for initiating asynchronous data transmission, which is used to initiate asynchronous data transmission from the source address src.addr to the destination address dst.addr with a length of length; ibcp.async.wait_thread is a thread-level wait synchronization instruction; ibcp.async.wait_wrap is a thread group-level wait synchronization instruction; ibcp.async.wait_block is a storage block-level wait synchronization instruction; ibcp.async.wait_all is a global-level wait synchronization instruction.

[0035] Taking ibcp.async(dst.addr, src.addr, length) as an example, this initiates an asynchronous one-sided RDMA data transmission, which can directly write data from the local, such as xPU memory (src.addr), to the memory of a remote device (dst.addr) without waiting for the transmission to complete. This instruction can be executed by a dedicated hardware engine, avoiding occupying the resources of the xPU.

[0036] Take ibcp.async.wait_thread as an example. It waits for all the operations of RDMA data transfers initiated by ibcp.async(dst.addr, src.addr, length) within the current xPU thread to complete, and its synchronization granularity reaches the thread level. After initiating multiple RDMA data transfers within the same thread, it can ensure that subsequent operations are not carried out until these transfers are completed.

[0037] Take ibcp.async.wait_wrap as an example. It waits for all the operations of RDMA data transfers initiated by ibcp.async(dst.addr, src.addr, length) within the current xPU thread group to complete, and its synchronization granularity reaches the thread group level. After initiating multiple RDMA data transfers within the same thread group, it can ensure that subsequent operations are not carried out until these transfers are completed.

[0038] Take ibcp.async.wait_block as an example. It waits for all the operations of RDMA data transfers initiated by ibcp.async(dst.addr, src.addr, length) within the current xPU storage block to complete, and its synchronization granularity reaches the storage block level. After initiating multiple RDMA data transfers within the same storage block, it can ensure that subsequent operations are not carried out until these transfers are completed.

[0039] Take ibcp.async.wait_all as an example. It waits for all the operations of RDMA data transfers initiated by ibcp.async(dst.addr,src.addr, length) before within the xPU to complete, and its synchronization granularity reaches the global level. After all the initiated RDMA data transfers, it can ensure that subsequent operations are not carried out until all the transfers are completed.

[0040] Therefore, in some embodiments of the present disclosure, through the hardware offloading asynchronous copy instruction set including multi-level synchronous control data transfer instructions, the RDMA data transfer task is directly offloaded to a dedicated hardware engine, completely releasing the computing resources of the xPU. It returns immediately after initiating the RDMA data transfer, and the xPU can continue to execute computing tasks and can be synchronized as needed through multi-level synchronous control data transfer instructions, maximizing the utilization of computing resources. In addition, multi-level synchronization such as thread, thread group, storage block, and global is provided to adapt to different scenario requirements.

[0041] Optionally, in some embodiments of the present disclosure, the first engine assembling the work element WE based on the hardware offload asynchronous copy instruction set may include: receiving an RDMA operation request; assembling the RDMA operation request into a WE based on the hardware offload asynchronous copy instruction set; and writing the WE into a second register, where the second register is located in the RNIC and can communicate with the second engine.

[0042] Specifically, the first engine may receive an RDMA operation request sent by any xPU in the system architecture through the NOC (Network on Chip), such as an RDMA Copy Async message. The RDMA operation request is assembled into a WE based on the hardware offload asynchronous copy instruction set. The NOC is a communication architecture inside the integrated circuit, used to connect multiple components within the chip (such as processors, memories, etc.), and can achieve efficient data transmission and task scheduling.

[0043] The WE may include at least one of Remote info, Local info, Opcode, QPN (Queue Pair Number), and TAG, where TAG is used to mark the thread identifier ID of the RDMA operation request or the block ID of the source.

[0044] For example, Remote info may be represented as (Remote VA, Rkey); Local info may be represented as (localVA, lkey, length), where Remote VA is the virtual address of the target memory on the remote device, Rkey is the remote memory access key; local VA is the virtual address of the memory on the local device; lkey is the local memory access key; and length is the length of the data transmission in bytes. The WE information together constitutes the basic parameters of the RDMA data transmission operation, which can ensure the efficient and secure execution of the RDMA data transmission.

[0045] Optionally, the method of writing the WE into the second register may include the device direct connection PCIe P2P method. The PCIe P2P method (Peer to Peer) means direct communication between two devices or components without passing through the processor or memory for transfer. This transmission method significantly improves the processing speed of the RDMA data transmission task and the overall system performance. For example, the PCIe P2P method is direct hardware-level communication, which bypasses the processor and memory; realizes asynchronous operation optimization. After writing the WE into the second register, the hardware engine (the second engine) communicating with the second register can automatically process the task (for example, storing the WE into the WQE Buffer) without waiting for software instructions; in addition, the direct hardware path can reduce the data transmission delay.

[0046] Step 202: The second engine stores the WE in the WQE Buffer and transfers it to the RDMA engine; Based on Step 201, this step aims to process DWQE at hardware speed using a hardware engine (the second engine) and release the main processing unit resources of the CPU and RNIC. On the basis of reducing the single - operation latency and improving the throughput of network devices, the RDMA data transfer task is directly offloaded to a dedicated hardware engine through a hardware asynchronous instruction set, realizing the decoupling of the RDMA data transfer task from the resources of the CPU / GPU. Therefore, it provides an RDMA data transfer solution that is efficient and low - latency without relying on CPU / GPU resources for data - transfer - intensive scenarios such as the MoE model.

[0047] Specifically, the RDMA protocol is a network protocol that allows a computer to directly access the memory of a remote computer and is commonly used in scenarios such as high - performance computing (HPC), data centers, and storage networks. RDMA can achieve high - speed data transfer between hosts without involving the intervention of the operating system kernel, so it has extremely low latency and high bandwidth. When using the RDMA protocol, the sender (i.e., the communication initiator) directly writes data into the memory of the receiver (usually called the communication recipient) without the intervention of the operating system. This direct memory access reduces the intermediate layer in the data transfer process, thus improving the data transfer efficiency.

[0048] However, taking data - transfer - intensive scenarios such as the MoE model as an example, each sample in the model can activate multiple experts, which leads to a large number of scattered small - data - block transfer requirements. In the RDMA data transfer process, each task needs to generate a WQE (Work Queue Element), trigger data processing through a DB (Doorbell), and actively check for completion events in the CQ (Completion Queue) through a thread.

[0049] In some embodiments, a CPU thread can be used to encapsulate a task unit to be sent (e.g., a network request) as a WQE and submit it to the send queue for processing. However, in the RDMA data transfer process, the number of QPs is relatively large, or in technologies such as RDMA, the number of basic units for managing send / receive queues is large. To meet the large number of scattered small - data - block transfer requirements and implement tasks such as generating WQE, triggering DB, and checking CQ, a large amount of CPU resources will be consumed. Therefore, when the data transfer overhead grows exponentially, system performance becomes the bottleneck for model scale - up.

[0050] In some other embodiments, the processing of multiple QP tasks can be completed through the multi-threaded parallel capabilities of the GPU. For example, the tasks of generating WQE, triggering the DB, and checking the CQ are completed through the multi-threads of the GPU. This decouples the dependence of the RDMA data transfer task on the CPU performance. However, in the case where the overhead of data transfer grows exponentially, it will consume a large amount of GPU SM computing resources.

[0051] Therefore, in the embodiments of the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed, and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transfer process and reduces data latency; in addition, through the hardware asynchronous instruction set, the RDMA data transfer task is directly offloaded to the dedicated hardware engine, realizing the decoupling of the RDMA data transfer task from the CPU / GPU resources. Therefore, it provides an RDMA data transfer solution that is efficient and low-latency without relying on CPU / GPU resources for data transfer-intensive scenarios such as the MoE model.

[0052] Specifically, in some embodiments of the present disclosure, the second engine storing the WE in the WQE Buffer and sending it to the RDMA engine may include: parsing the WE and storing the WE in the WQE Buffer; based on the occupancy rate of the WQE Buffer being greater than or equal to a preset threshold, transmitting multiple WE stored in the WQE Buffer to the host memory of the network device.

[0053] The second engine can, for example, receive the WE from the PCIe bus and verify the integrity of the WE information after receiving the WE. Check whether the received WE information is complete and whether there is data loss or corruption. For example, verify whether the length in the WE information meets the expectation and whether the checksum in the WE information is correct, etc. If the WE information is incomplete or there is an error, the WE can be selected to be discarded and the error information can be recorded.

[0054] In addition, the second engine can also parse the WE and parse each field in the WE to clarify the specific requirements of this operation. For example, parse fields such as Remote info (Remote VA, Rkey), local Info (local VA, lkey, length), Opcode, QPN, TAG, etc.

[0055] After parsing the WE, the second engine can extract the key information in the WE. Extract the key information from the parsed fields to provide a basis for subsequent processing and decision-making. For example, determine the operation type according to the Opcode. Optionally, the operation type can include read, write, atomic operation, etc.; and determine the queue peer to which the operation belongs according to the QPN.

[0056] According to the information obtained by parsing, the second engine can also allocate cache space in the WQE Buffer for this WE. Based on the occupancy rate of the WQE Buffer being greater than or equal to a preset threshold, transfer multiple WEs stored in the WQE Buffer to the host memory of the network device. In other words, if the on-chip WQE Buffer is full, the WQE can be pushed to the host memory of the network device (such as the WQE external buffer 111 of the host memory shown). This hierarchical storage can ensure performance through on-chip caching (e.g., WQE Buffer), provide capacity redundancy through external memory (e.g., WQE external buffer 111), and ultimately achieve an efficient and reliable WQE processing pipeline. Figure 2 As shown in the WQE external buffer 111 of the host memory). This hierarchical storage can ensure performance through on-chip caching (e.g., WQEBuffer), provide capacity redundancy through external memory (e.g., WQE external buffer 111), and ultimately achieve an efficient and reliable WQE processing pipeline.

[0057] In addition, after parsing the WE, the second engine can schedule the processing order. For example, determine the processing order of this WE according to factors such as the priority of the operation and the queue status. High-priority WEs are processed first. This hardware arbitration mechanism automatically assigns task priorities and can avoid queue blocking.

[0058] Optionally, the second engine stores the WE in the WQE Buffer and sending it to the RDMA engine may further include: generating a DB, where the DB may include at least one of the queue pair information to be processed, the starting address of multiple WEs, and the number of multiple WEs; transmitting the DB to the RDMA engine can trigger the RDMA engine to perform data processing, where the RDMA engine can read the WE from the WQE Buffer or the WQE external buffer in the above host memory.

[0059] The generation and transmission of the DB itself is the sign of the start of the task. When the DB is written to a specific memory area (such as a predefined location in the on-chip cache or host memory), after the RDMA engine detects the existence of the DB through hardware logic, it automatically starts the processing flow without software or additional interrupt signal intervention. In addition, the DB can be located in the shared area of the on-chip cache or host memory, and the RDM engine can directly access it, which avoids the data copy overhead. After the RDM engine finishes processing the current data processing, it can immediately obtain the next data processing through the doorbell mechanism of the DB without waiting for an interrupt response.

[0060] Step 203: The RDM engine performs data processing based on the WE, and the data processing can include data transmission, memory access, and queue management; This step aims to describe that the RDMA engine realizes RDMA data transmission processing with low latency, high bandwidth, and low xPU resource occupancy by performing data processing tasks including data transmission, memory access, and queue management. Its core advantage lies in completely offloading complex network operations (such as transmission control, memory management) to hardware, enabling the CPU to focus on business logic, thereby significantly improving the performance of distributed systems.

[0061] Optionally, the RDMA engine can achieve high-speed data transmission between hosts through direct memory access technology without involving the intervention of the operating system kernel, so it has extremely low latency and high bandwidth. For example, the send WE contains the source address and the destination address. Through direct memory access technology, the RDMA engine can directly transfer the data from the source address to the send buffer of the RNIC without copying.

[0062] In addition, the RDMA engine can support multiple types of transmissions, such as Send / Receive, Read / Write, etc. Send / Receive can be understood as two-way data transfer between local or remote nodes, and Read / Write can be understood as directly accessing the memory of a remote node. This can eliminate the context switching and interrupt overhead involved in the xPU and reduce the data transmission latency to the microsecond level.

[0063] Through the Memory Window technology, the RDMA engine can map the virtual address of a remote node to a physical address, thus supporting direct memory access across nodes.

[0064] In addition, each queue pair (QP) can contain an independent send queue (SQ) and receive queue (RQ), so the RDMA engine supports concurrent processing of multiple task flows.

[0065] Step 204: The first engine receives the feedback on the data processing transmitted via the second engine.

[0066] This step aims to feedback the data processing completion message based on the RDMA operation request to the xPU that issues the RDMA operation request. Optionally, the first engine can transmit the feedback on this data processing to the xPU that issues the RDMA operation request through the NOC. The NOC is a communication architecture inside an integrated circuit used to connect multiple components within the chip (such as processors, memories, etc.) and can achieve efficient data transmission and task scheduling.

[0067] Specifically, after the RDMA engine completes the above data processing, the RDMA engine can generate a Completion Event (CE) message; the second engine receives the CE message and writes the received CE message into the first register, where the first register is located in the xPU and can communicate with the first engine. This can ensure the reliability of RDMA data transmission.

[0068] Both the first register and the second register described above can be understood as part of the Base Address Register (BAR) register. The first engine, the second engine, and the RDMA engine in the system architecture can be located in different hardware modules or processing units (e.g., xPU), and can communicate with each other through NOC or shared registers (e.g., BAR registers). This avoids the need for complex hardware connections that may be required for direct communication between the first engine, the second engine, and the RDMA engine, reduces the design complexity, and can ensure that different hardware modules or processing units work independently and reduce interference with each other.

[0069] In addition, the BAR register is a register reported by PCIe devices (e.g., RNIC, xPU, etc.) to the server during the enumeration phase, and is used to declare the memory space required by the device (e.g., configuration space, I / O space, memory-mapped space, etc.). The server can assign a physical address to the BAR register. Thus, as part of memory-mapped I / O, the BAR register provides a standardized interface, and data interaction can be completed through simple read and write operations, reducing the design difficulty and power consumption.

[0070] Optionally, the method of writing the received CE message into the first register can include the PCIe P2P method. This transmission method significantly improves the processing speed of RDMA data transmission tasks and the overall system performance. For example, the PCIe P2P method is direct hardware-level communication that bypasses the processor and memory; it achieves asynchronous operation optimization. After writing the CE into the first register, the hardware engine (the first engine) that communicates with the first register can automatically process tasks (e.g., parse the CE) without waiting for software instructions; in addition, the direct hardware path can reduce data transmission latency.

[0071] As an option, the CE message can include at least one of the following: an operation code (Opcode), a queue pair number (QPN), and a tag (TAG), where the TAG is used to mark the thread identifier (ID) of the RDMA operation request or the block ID of the source. The CE information together constitutes the basic parameters of the RDMA data transmission operation, and can ensure the efficient and secure execution of the RDMA data transmission.

[0072] After receiving the CE message, the first engine can generate a task completion message based on the CE message, pass it to the NOC, and transmit it to the xPU that issued the RDMA operation request through the NOC.

[0073] Therefore, in the RDMA data transmission scheme provided by the embodiments of the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed, and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, by directly offloading the RDMA data transmission task to the dedicated hardware engine through the hardware asynchronous instruction set, the decoupling of the RDMA data transmission task from the CPU / GPU resources is achieved, thus providing an RDMA data transmission solution that is efficient and low-latency without relying on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0074] Figure 4 It is a block diagram of another RDMA data transmission network system architecture 100 provided by the embodiments of the present disclosure. Figure 5 It is a flowchart of another RDMA data transmission method provided by the embodiments of the present disclosure.

[0075] As Figure 2 and Figure 4 shown, the system architecture 100 may include multiple network devices, such as a first network device and a second network device, where the first network device may include a server A and an xPUA, etc., and the second network device may include a server B and an xPUB, etc.

[0076] In addition, it should be noted that the system architecture 100 may also include other network devices in addition to the first network device and the second network device. Figure 4 Only a small number of network devices in the system architecture 100 are shown as examples. Those skilled in the art can set the number, type, and connection relationship of the network devices in the system architecture 100 according to actual needs, and the present disclosure does not limit this.

[0077] Each network device among the multiple network devices may include an RNIC and an xPU. The xPU may include a first engine, and the RNIC may include a second engine communicating with the first engine, an RDMA engine communicating with the second engine, and a WQE Buffer.

[0078] In addition, in each of the multiple network devices, the xPU may further include a core computing unit and a memory management unit, where the types of the core computing unit and the memory management unit of different types of xPUs may be different. For example, in the case where the xPU includes a GPU, its core computing unit may include SM; in the case where the xPU includes a CPU, its core computing unit may include an arithmetic logic unit, a control unit, etc.

[0079] In addition, when the xPU executes a program, the core computing unit generates a virtual address, and the memory management unit may be responsible for converting the virtual address into a physical address. Optionally, in the case where the xPU includes a GPU or a CPU, its memory management unit may include an MMU.

[0080] In each of the multiple network devices, the RDMA engine of the RNIC may perform various data processing, where the data processing may include data transmission, memory access, queue management, etc.

[0081] It should be understood that Figure 4 the numbers of servers, networks, network devices, etc. in

[0082] are merely illustrative. According to the implementation requirements, there may be any number of servers, networks, and network devices.

[0083] Reference Figure 5 , process 300 may include the following steps: Step 301: The first engine of the first network device assembles a work element WE based on the hardware offload asynchronous copy instruction set and transmits the WE to the second engine; The purpose of this step is to trigger the communication protocol including the hardware offload instruction set by the hardware engine (the first engine), where the hardware engine may be directly integrated inside the xPU and can assemble the WE (Work Element) by itself. Through, for example, PCIe BAR space mapping, etc., the assembled WE can be directly output to the second engine located in the RNIC. This can reduce the dependence of RDMA data transmission on CPU performance or GPU performance. In addition, the number of data transfers during the RDMA data transmission process can be reduced, and the data delay during the above transmission process can be reduced.

[0084] In addition, in some embodiments of the present disclosure, before performing step 301, the RDMA data transmission method may further include: creating an interface corresponding to at least a part of the MR (Memory Region) of RDMA, where the interface information includes a unique identifier UniqueID; based on this interface, allocating the MR to multiple network devices, and establishing an MR page table for the network devices, where the memory access credentials mKey of the multiple MR page tables are all UniqueID; multiple network devices perform memory handle exchange; binding the memory handle obtained after the exchange to the information of the WE, and establishing a virtual address mapping relationship, so as to facilitate remote access initiated through the memory handle and UniqueID.

[0085] Specifically, when supported by the RNIC hardware, a new RDMA memory region can be provided to create the above interface. For example, the instruction rdma_register_mr_unique(va, length, uniqueID) is used to create the interface. UniqueID is the ID used to identify collective communication, and all nodes in the cluster will complete synchronization during communication initialization. The nodes described here can be understood as network devices in the system architecture.

[0086] The memory region can be understood as an abstraction of a continuous memory segment in RDMA. The interface allows users to create an MR by specifying a virtual address (va), a length (length), and a unique identifier (uniqueID). UniqueID can ensure that in cluster communication, each node can identify and synchronize specific communication tasks.

[0087] Taking the first network device as an example, the xPUA can apply for HBM (High Bandwidth Memory) through the cudaMalloc function, create a memory region corresponding to the first network device using the above interface; establish an MR page table for the first network device; and make the memory access credentials mKey of the MR page table all be UniqueID. Among them, after creating a memory region for the first network device and establishing an MR page table, the operation of RDMA data transmission can be located in the memory region of the first network device, and mKey can be used as the key credential for this memory access.

[0088] Similarly, a memory region can also be created and an MR page table can be established in the second network device and the remaining network devices. In this way, the operation of RDMA data transmission can also be located in the memory region of the second network device and the memory regions of the remaining network devices, and mKey can be used as the key credential for this memory access. In addition, multiple network devices in the system architecture can also access each other, and using the same UniqueID can ensure the communication association between multiple network devices.

[0089] Optionally, the application program of the second network device can assemble a memory handle (xPUB Memory Handle), interact with the remaining network devices (e.g., the first network device) through Socket or other means, and pass the memory handle to the remaining network devices. The memory handle may include key information for accessing remote memory, such as the virtual address (xPUB va) and length (length) of the second network device.

[0090] Similarly, the application program of the first network device can assemble a memory handle (xPUA Memory Handle), interact with the remaining network devices (e.g., the second network device) through Socket or other means, and pass the memory handle to the remaining network devices. The memory handle may include key information for accessing remote memory, such as the virtual address (xPUAva) and length (length) of the first network device.

[0091] The network device can bind the memory handle obtained after the exchange to the information of the WE and establish a virtual address mapping relationship to facilitate remote access through the memory handle and the UniqueID.

[0092] For example, the network device can bind the memory handle obtained after the exchange to the QP and specify which QP is to complete the access to the network device corresponding to the memory handle. In addition, the QP is edited in the information of the WE. In this way, the WE can specify the communication path for the network device to access, ensuring that data can be accurately transmitted between network devices.

[0093] Optionally, taking the first network device as an example, the first network device can apply for a local reserved virtual address (e.g., Local Reserved VA) space locally through the CudaMallocReserved function, establish mapping information for it, and make it point to the memory handle of the remaining network devices (e.g., the xPUB Memory Handle of the second network device). The mapping information can be managed by the MMU of the first network device or in the form of table entries. In other words, the network device can reserve a section of virtual address space locally as the local reserved virtual address. After establishing the mapping, when the network device accesses the local reserved virtual address, the local reserved virtual address can be converted into the actual address and relevant communication information of the network device to be accessed through the mapping information, thus realizing remote memory access.

[0094] On this basis, taking the first network device as an example, the xPU of the first network device can start a kernel program Kernel Launch, which is executed in parallel by multiple threads and initiates remote memory access through a hardware offload asynchronous copy instruction set. The UniqueID is injected into the process context and can be transmitted to the MMU and RNIC through the hardware offload asynchronous copy instruction set later for associating the metadata of remote memory access.

[0095] Taking ibcp.async(xPUB Reserved VA, Local VA, length) as an example, the hardware offload asynchronous copy instruction set is transmitted to the MMU, and the MMU can complete address translation, converting the reserved virtual address xPUB ReservedVA of the second network device into the corresponding virtual address xPUB VA and QPN information of the second network device. Then, the MMU transmits the virtual address xPUA VA, xPUB VA, QPN, length, and UniqueID information of the first network device to the first engine.

[0096] Optionally, the hardware offload asynchronous copy instruction set may include ibcp.async(dst.addr, src.addr, length). When Dst.addr is xPUB Reseverd VA, it means writing the memory pointed to by xPUA local va of the local first network device to the memory of the remote second network device; when Src.addr is xPUB Reseverd VA, it means reading the xPUB Memory of the remote second network device to the memory pointed to by xPUA local va of the local first network device.

[0097] UniqueID is used to identify the ID of the collective communication, and the synchronization of all nodes in the cluster has been completed during communication initialization. In other words, at the starting point of the RDMA data operation, by starting the kernel program and carrying the UniqueID, it can be ensured that the operation can be accurately associated with a specific communication task, avoiding confusion between different communication tasks.

[0098] The role of the MMU is to convert the virtual address into a physical address or the corresponding remote address information. Through address translation, the actual memory address of the network device can be determined, and the relevant operation information can be transmitted to the first engine and the second engine for subsequent preparation work of RDMA data transmission.

[0099] The first engine assembles the work element WE based on the above information and the hardware offload asynchronous copy instruction set and transmits the WE to the second engine.

[0100] Step 302: The second engine of the first network device stores the WE in the WQE Buffer and transfers it to the RDMA engine of the first network device; Based on step 301, this step aims to use a hardware engine (the second engine) to process the DWQE at hardware speed and release the CPU and the main processing unit resources of the RNIC. On the basis of reducing the latency of a single operation and improving the throughput of the network device, the RDMA data transfer task is directly offloaded to a dedicated hardware engine through a hardware asynchronous instruction set, realizing the decoupling of the RDMA data transfer task from the CPU / GPU resources. Therefore, it provides an RDMA data transfer solution that is efficient and low-latency and does not rely on CPU / GPU resources for data transfer-intensive scenarios such as the MoE model.

[0101] Step 303: The RDMA engine of the first network device initiates a request to the second network device based on the WE; This step aims to describe that the RDMA engine realizes RDMA data transfer processing with low latency, high bandwidth, and low xPU resource occupancy by executing data processing tasks including data transfer, memory access, and queue management. Its core advantage is that it completely offloads complex network operations (such as transmission control and memory management) to hardware, enabling the CPU to focus on business logic, thereby significantly improving the performance of the distributed system.

[0102] Optionally, in some embodiments of the present disclosure, the RDMA engine of the first network device may obtain the QPC (Queue Pair Context) corresponding to the queue label QPN of the WE based on the WE and determine the operation type according to the QPC; the RDMA engine of the first network device may also obtain the memory handle and UniqueID of the second network device based on the WE; the RDMA of the first network device completes packet encapsulation based on the WE and sends a request to the second network device.

[0103] Step 304: The second network device completes the above request and feeds it back to the first engine of the first network device.

[0104] This step aims to feed back the data processing completion message based on the RDMA operation request to the xPU that issued the RDMA operation request. Optionally, the first engine may transmit the feedback of this data processing to the xPU that issued the RDMA operation request through the NOC. The NOC is an internal communication architecture of an integrated circuit used to connect multiple components within the chip (such as processors, memories, etc.) and can achieve efficient data transfer and task scheduling.

[0105] Thus, in the RDMA data transmission solution provided by the embodiments of the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed, and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, by using the hardware asynchronous instruction set, the RDMA data transmission task is directly offloaded to the dedicated hardware engine, realizing the decoupling of the RDMA data transmission task from the CPU / GPU resources, thus providing an RDMA data transmission solution that is efficient and low-latency without relying on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0106] Figure 6 FIG. 100 is a block diagram of another RDMA data transmission network system architecture provided by the embodiments of the present disclosure.

[0107] Combined Figure 5 and Figure 6 , the following will describe the RDMA data transmission method by taking the RDMA data transmission task as a write operation as an example. Specifically, the first network device may initiate an ibcp.async write operation request to write data from the local memory of the first network device to the remote memory of the second network device. It should be noted that since the content involved in the RDMA data transmission method described above can be fully or partially applied to the RDMA data transmission method described below, the related or similar content will not be repeated. However, those skilled in the art can understand that the structural features, method processes, implementation principles, and technical effects achieved by the RDMA data transmission methods described herein are similar.

[0108] Referring Figure 5 , the process 300 may include the following steps: Step 301: The first engine of the first network device assembles a work element WE based on the hardware offloading asynchronous copy instruction set and transmits the WE to the second engine; The purpose of this step is to trigger the communication protocol including the hardware offloading instruction set by the hardware engine (the first engine), where the hardware engine can be directly integrated inside the xPU and can assemble the WE (Work Element) by itself. Through, for example, PCIe BAR space mapping, the assembled WE can be directly output to the second engine located in the RNIC. This can reduce the dependence of RDMA data transmission on CPU performance or GPU performance. In addition, it can also reduce the number of data transfers during the RDMA data transmission process and reduce the data latency during the above transmission process.

[0109] In addition, in some embodiments of the present disclosure, before performing step 301, the RDMA data transmission method may further include: creating an interface corresponding to at least a part of the memory region MR of RDMA, where the interface information includes a unique identifier QPC; based on this interface, allocating the MR to multiple network devices, and establishing an MR page table for the network devices, where the memory access credentials mKey of the multiple MR page tables are all UniqueID; implementing memory handle exchange among the multiple network devices; binding the obtained memory handle after the exchange to the information of the WE, and establishing a virtual address mapping relationship to facilitate initiating remote access through the memory handle and UniqueID.

[0110] During the execution of step 301, the xPU of the first network device may start a kernel program KernelLaunch, which is executed in parallel by multiple threads, and initiate remote memory access through a hardware offload asynchronous copy instruction set. UniqueID is injected into the process context and can be subsequently transmitted to the MMU and RNIC through the hardware offload asynchronous copy instruction set for associating the metadata of remote memory access.

[0111] UniqueID is used to identify the ID of the collective communication, and the synchronization of all nodes in the cluster has been completed during communication initialization. In other words, at the starting point of the RDMA data operation, by starting the kernel program and carrying UniqueID, it can be ensured that the operation can be accurately associated with a specific communication task, avoiding confusion between different communication tasks.

[0112] The role of the MMU is to convert the virtual address into a physical address or the corresponding remote address information. Through address translation, the actual memory address of the network device can be determined, and the relevant operation information can be passed to the first engine and the second engine for subsequent preparation work of RDMA data transmission.

[0113] Taking ibcp.async(xPUB Reserved VA, Local VA, length) as an example, the hardware offload asynchronous copy instruction set is transmitted to the MMU, and the MMU can complete the address translation, converting the reserved virtual address xPUB ReservedVA of the second network device into the corresponding virtual address xPUB VA and QPN information of the second network device. Then, the MMU transmits the virtual address xPUA VA, xPUB VA, QPN, length, and UniqueID information of the first network device to the first engine.

[0114] The first engine assembles the work element WE based on the above information and the hardware offload asynchronous copy instruction set, and transmits the WE to the second engine.

[0115] Step 302: The second engine of the first network device stores the WE in the WQE Buffer and transfers it to the RDMA engine of the first network device; Based on Step 301, this step aims to process DWQE at hardware speed using a hardware engine (the second engine) and release the CPU and the main processing unit resources of the RNIC. On the basis of reducing the latency of a single operation and improving the throughput of the network device, the RDMA data transfer task is directly offloaded to a dedicated hardware engine through a hardware asynchronous instruction set, achieving the decoupling of the RDMA data transfer task from the CPU / GPU resources. Therefore, it provides an RDMA data transfer solution that is efficient and low-latency without relying on CPU / GPU resources for data transfer-intensive scenarios such as the MoE model.

[0116] Step 303: The RDMA engine of the first network device initiates a request to the second network device based on the WE; This step aims to describe that the RDMA engine realizes RDMA data transfer processing with low latency, high bandwidth, and low xPU resource occupancy by performing data processing tasks including data transfer, memory access, and queue management. Its core advantage is to completely offload complex network operations (such as transmission control and memory management) to hardware, enabling the CPU to focus on business logic, thus significantly improving the performance of the distributed system.

[0117] Specifically, the RDMA engine of the first network device can obtain the corresponding QPC based on the QPN, initiate an RDMA Write operation, and complete the encapsulation of the RDMA Write Request message according to the hardware offload asynchronous copy instruction set and the WE, and send it to a remote network device (such as the second network device). For example, the data in the high-bandwidth memory of the second network device can be obtained according to the xPUB va, UniqueID (as the mKey), and length information.

[0118] Step 304: The second network device completes the above request and feeds back to the first engine of the first network device.

[0119] This step aims to feed back the data processing completion message based on the RDMA operation request to the xPU that issued the RDMA operation request. Optionally, the first engine can transmit the feedback of this data processing to the xPU that issued the RDMA operation request through the NOC. The NOC is an internal communication architecture of an integrated circuit used to connect multiple components within the chip (such as processors, memories, etc.), and can achieve efficient data transfer and task scheduling.

[0120] For example, after the RDMA engine of the second network device receives the above request message (e.g., data write packet), it can write the payload data in the request message into the memory of the second network device according to the remote endpoint transport header, UniqueID, and length information, where the remote endpoint transport header may include the virtual address of the second network device (xPUB VA, xPUB Virtual Address). By offloading the asynchronous copy instruction set and WE information in hardware, the second network device can be accurately located, and the xPU of the second network device can be bypassed, reducing the intervention and overhead of the xPU, thus significantly improving the efficiency of RDMA data transmission and avoiding the xPU from becoming the bottleneck of data transmission.

[0121] The RDMA engine of the second network device can also return an ACK (Acknowledgment, confirmation response packet) to the RDMA engine of the first network device to inform the first network device that the RDMA write request message it sent has been successfully received and processed. This can ensure the reliability of RDMA data transmission.

[0122] After the RDMA engine of the first network device receives the ACK, it can generate a CE message according to the QP and WQE and deliver the CE message to the second engine. The second engine can assemble the received CE message, for example, making the format and content of the message more standardized and unified. Then, it is passed to the first engine in a P2P manner. By generating and delivering the CE message, the RDMA engine of the first network device can timely inform the hardware engine that the operation has been completed. This enables the hardware engine to perform subsequent processing in a timely manner, such as updating status information, releasing resources, or triggering new tasks. In addition, the P2P method avoids the latency and bandwidth contention problems brought by message transmission through intermediate nodes or shared buses. Through the P2P method, the CE message can be quickly transferred from the second engine to the first engine, improving the response speed and overall performance of the network device.

[0123] Therefore, in the RDMA data transmission solution provided by the embodiments of the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed, and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, by directly offloading the RDMA data transmission task to the dedicated hardware engine through the hardware asynchronous instruction set, the decoupling of the RDMA data transmission task from the CPU / GPU resources is achieved. Therefore, it provides an RDMA data transmission solution that is efficient and low-latency without relying on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0124] Referring again to Figure 2 , as an implementation of the above RDMA data transmission method, the embodiments of the present disclosure provide an RDMA data transmission network device, which can correspond to Figure 3 the method shown, and this network device can be specifically applied to various electronic devices.

[0125] As Figure 2 shown, the RDMA data transmission network device provided by the embodiments of the present disclosure may include an RNIC 120 and an xPU 130. The xPU 130 may include a first engine 131, and the RNIC 120 may include a second engine 121 communicating with the first engine 131, an RDMA engine 122 communicating with the second engine 121, and a WQE Buffer 123. The first engine 131 is configured to assemble a work element WE based on the hardware offloading asynchronous copy instruction set and transmit the WE to the second engine; the second engine 121 is configured to store the WE in the WQE Buffer 123 and transmit it to the RDMA engine 122; the RDMA engine 122 is configured to perform data processing based on the WE and transmit the feedback of the data processing to the first engine 131 via the second engine 121, and the data processing includes data transmission, memory access, and queue management.

[0126] In addition, the xPU 130 may further include a core computing unit 132 and a memory management unit 133, where the types of the core computing unit 132 and the memory management unit 133 of different types of xPU 130 may be different. For example, when the xPU 130 includes a GPU, its core computing unit 132 may include an SM (Streaming Multiprocessor); when the xPU 130 includes a CPU, its core computing unit 132 may include an Arithmetic Logic Unit (ALU), a Control Unit (CU), etc.

[0127] In addition, when the xPU 130 executes a program, the virtual address generated by the core computing unit 132, and the memory management unit 133 may be responsible for converting the virtual address into a physical address. Optionally, when the xPU 130 includes a GPU or a CPU, its memory management unit 133 may include an MMU (Memory Management Unit).

[0128] Optionally, the xPU 130 may include at least one of a CPU, a GPU, a TPU, an NPU, a DPU, a VPU, a QPU, and an APU. In addition, the xPU 130 generally appears as hardware, and may also appear as software or a software running product in special scenarios (for example, an emulation scenario), which is not specifically limited in this disclosure.

[0129] Optionally, the RDMA data transmission network device may include multiple xPU 130. In other words, those skilled in the art may set the number, type, and connection relationship of the RNIC 120 and the xPU 130 in the RDMA data transmission network device according to actual needs, which is not limited in this disclosure.

[0130] In some embodiments of the present disclosure, the hardware offloading asynchronous copy instruction set may include multi-level synchronous control data transmission instructions, where the multi-level synchronous control data transmission instructions may include thread-level synchronous control data transmission instructions, thread group-level synchronous control data transmission instructions, storage block-level synchronous control data transmission instructions, and global-level synchronous control data transmission instructions. By using the hardware offloading asynchronous copy instruction set including multi-level synchronous control data transmission instructions, the RDMA data transmission task is directly offloaded to a dedicated hardware engine, completely releasing the computing resources of the xPU. Immediately return after initiating the RDMA data transmission, and the xPU can continue to execute the computing task and can synchronize as needed through the multi-level synchronous control data transmission instructions, maximizing the utilization of computing resources. In addition, multi-level synchronization such as thread, thread group, storage block, and global is provided to adapt to different scenario requirements.

[0131] Optionally, the first engine 131 may be configured to assemble a work element WE based on a hardware offload asynchronous copy instruction set. For example, the first engine 131 may receive an RDMA operation request; assemble the RDMA operation request into a WE based on the hardware offload asynchronous copy instruction set; and write the WE into a second register, where the second register is located in the RNIC 120 and can communicate with the second engine 121.

[0132] The WE may include at least one of Remote info, Local info, Opcode, QPN (Queue Pair Number), and TAG, where the TAG is used to mark the thread identifier ID of the RDMA operation request or the block ID of the source. The WE information together constitutes the basic parameters for the RDMA data transfer operation, which can ensure the efficient and secure execution of the RDMA data transfer.

[0133] Optionally, the first engine 131 may be configured to write the WE into the second register in a device direct PCIe P2P manner. The PCIe P2P manner is a direct hardware-level communication that bypasses the processor and memory; it achieves asynchronous operation optimization. After writing the WE into the second register, the hardware engine (the second engine) communicating with the second register can automatically process tasks (such as storing the WE in the WQE Buffer) without waiting for software instructions; in addition, the direct hardware path can reduce data transfer latency.

[0134] In some embodiments of the present disclosure, the second engine 121 may be configured to parse the WE and store the WE in the WQE Buffer 123; based on the occupancy rate of the WQE Buffer 123 being greater than or equal to a preset threshold, transfer multiple WEs stored in the WQE Buffer 123 to the host memory of the network device.

[0135] Specifically, the second engine 121 can receive the WE from the PCIe bus, for example, and verify the integrity of the WE information after receiving it. Check whether the received WE information is complete, without data loss or corruption. For example, verify whether the length in the WE information meets the expectation, whether the checksum in the WE information is correct, etc. If the WE information is incomplete or there is an error, the WE can be optionally discarded and the error information can be recorded. In addition, the second engine 121 can also parse the WE, parse each field in the WE to clarify the specific requirements of the operation. According to the parsed information, the second engine 121 can also allocate buffer space in the WQE Buffer 123 for the WE. If the on-chip WQE Buffer 123 is full, the WQE can be pushed to the host memory of the network device. This hierarchical storage can ensure performance through on-chip caching (e.g., WQE Buffer 123), provide capacity redundancy through external memory, and ultimately achieve an efficient and reliable WQE processing pipeline.

[0136] Optionally, after parsing the WE, the second engine 121 can also schedule the processing order. For example, according to factors such as the priority of the operation and the queue status, determine the processing order of the WE. High-priority WEs are processed first. This hardware arbitration mechanism automatically assigns task priorities and can avoid queue blocking.

[0137] In addition, the second engine 121 can also generate a DB, where the DB can include at least one of the queue pair information to be processed, the starting addresses of multiple WEs, and the number of multiple WEs; transmitting the DB to the RDMA engine can trigger the RDMA engine 122 to perform data processing, where the RDMA engine 122 can read the WE from the WQE Buffer 123 or the external buffer of the WQE in the above host memory.

[0138] The generation and transmission of the DB itself is a sign of the start of the task. When the DB is written to a specific memory area (such as a predefined location in on-chip caching or host memory), the RDMA engine 122 can automatically start the processing flow after detecting the existence of the DB through hardware logic, without the intervention of software or additional interrupt signals. Additionally, the DB can be located in a shared area of on-chip caching or host memory, and the RDM engine 122 can directly access it, which avoids the data copy overhead. After the RDM engine 122 finishes processing the current data processing, it can immediately obtain the next data processing through the doorbell mechanism of the DB without waiting for an interrupt response.

[0139] Optionally, the RDMA engine 122 can achieve high-speed data transmission between hosts through direct memory access technology without involving the intervention of the operating system kernel, thus having extremely low latency and high bandwidth. For example, the sending WE contains the source address and the destination address. Through direct memory access technology, the RDMA engine 122 can directly transfer the data from the source address to the sending buffer of the RNIC 120 without copying.

[0140] In addition, the RDMA engine 122 can support multiple types of transmissions, such as Send / Receive, Read / Write, etc. Among them, Send / Receive can be understood as two-way data transfer between local or remote nodes, and Read / Write can be understood as directly accessing the memory of a remote node. This can eliminate the context switching and interrupt overhead involved by the xPU 130 and reduce the data transmission latency to the microsecond level. Through the memory window technology, the RDMA engine 122 can map the virtual address of a remote node to a physical address, thus supporting direct memory access across nodes. In addition, each QP can contain an independent sending queue and receiving queue, so the RDMA engine 122 supports concurrent processing of multiple task flows.

[0141] After the RDMA engine 122 finishes the above data processing, the RDMA engine 122 can generate a CE message; the second engine 121 receives the CE message and writes the received CE message into the first register, where the first register is located in the xPU 130 and can communicate with the first engine 131.

[0142] Both the first register and the second register described above can be understood as part of the BAR (Base Address Register) register. The first engine 131, the second engine 121, and the RDMA engine 122 can be located in different hardware modules or processing units in the RDMA data transmission network device and can communicate with each other through the NOC or shared registers (such as the BAR register). This avoids the complex hardware connections that may be required for the direct transfer between the first engine 131, the second engine 121, and the RDMA engine 122, reduces the design complexity, and can ensure the independent operation of different hardware modules or processing units and reduce mutual interference.

[0143] In addition, the BAR register is a register reported by PCIe devices (such as RNIC, xPU, etc.) to the server during the enumeration phase and is used to declare the memory space required by the device (such as configuration space, I / O space, memory mapped space, etc.). The server can assign a physical address to the BAR register. Thus, as part of the memory mapped I / O, the BAR register provides a standardized interface and can complete data interaction through simple read and write operations, reducing the design difficulty and power consumption.

[0144] Optionally, the second engine 121 can write the received CE message into the first register in a PCIe P2P manner. This transmission method significantly improves the processing speed of RDMA data transmission tasks and the overall system performance. For example, the PCIe P2P method is a hardware-level direct communication that bypasses the processor and memory; it implements asynchronous operation optimization. After writing CE into the first register, the hardware engine (first engine 131) communicating with the first register can automatically process the task (for example, parse CE) without waiting for software instructions; in addition, the direct hardware path can reduce data transmission delays.

[0145] As an option, the CE message may include: at least one of an operation code Opcode, a queue pair number QPN, and a tag TAG, wherein TAG is used to mark the thread identifier ID of the RDMA operation request or the source block ID. The CE information together constitutes the basic parameters of the RDMA data transmission operation, which can ensure the efficient and secure execution of RDMA data transmission.

[0146] After receiving the CE message, the first engine 131 may generate a task completion message based on the CE message, and pass the message to the NOC, which is then transmitted to the xPU that issued the RDMA operation request through the NOC.

[0147] Therefore, in the RDMA data transmission solution provided by the embodiment of the present disclosure, the first engine is directly integrated into the xPU, and the communication protocol including the hardware offload instruction set can be triggered by the hardware engine; the second engine is a dedicated pre-processing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed and release the main processing unit resources of the CPU and RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, the RDMA data transmission task is directly offloaded to the dedicated hardware engine through the hardware asynchronous instruction set, which realizes the decoupling of the RDMA data transmission task from the CPU / GPU resources, thereby providing a high-efficiency and low-latency RDMA data transmission solution that does not rely on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0148] Reference again Figure 1 , Figure 4 and Figure 6 As an implementation of the above RDMA data transmission method, the embodiment of the present disclosure provides an RDMA data transmission network system, and the system architecture 100 can be used with Figure 5 Corresponding to the method shown, the network system can be specifically applied to various electronic devices.

[0149] The RDMA data transmission network system architecture 100 provided by the embodiments of the present disclosure may include multiple network devices, such as a first network device and a second network device. The first network device may include an xPUA and an RNIC, etc., and the second network device may include a server B and an xPUB, etc. Both the xPUA and the xPUB include a first engine. The RNIC includes a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine.

[0150] The system architecture 100 may further include the remaining network devices other than the first network device and the second network device. Figure 4 and Figure 6 Only a small number of network devices in the system architecture 100 are shown as examples. Those skilled in the art can set the number, type, and connection relationship of the network devices in the system architecture 100 according to actual needs, and the present disclosure does not limit this.

[0151] In addition, in each of the multiple network devices, the xPU may further include a core computing unit and a memory management unit, where the types of the core computing unit and the memory management unit of different types of xPUs may be different. For example, in the case where the xPU includes a GPU, its core computing unit may include an SM; in the case where the xPU includes a CPU, its core computing unit may include an arithmetic logic unit, a control unit, etc.

[0152] In addition, when the xPU executes a program, the core computing unit generates a virtual address, and the memory management unit may be responsible for converting the virtual address into a physical address. Optionally, in the case where the xPU includes a GPU or a CPU, its memory management unit may include an MMU.

[0153] In each of the multiple network devices, the RDMA engine of the RNIC may perform various data processes, where the data processes may include data transmission, memory access, and queue management, etc.

[0154] It should be noted that since the content described in the RDMA data transmission network above may be fully or partially applicable to the RDMA data transmission system described below, the related or similar content will not be repeated. However, those skilled in the art can understand that the structural features, method processes, implementation principles, and technical effects achieved by the RDMA data transmission schemes described in both are similar.

[0155] The first engine of the first network device can be configured to assemble a work element WE based on a hardware offload asynchronous copy instruction set and transmit the WE to the second engine of the first network device; the second engine of the first network device can be configured to store the WE in the WQE Buffer and transmit it to the RDMA engine of the first network device; the RDMA engine of the first network device can be configured to initiate a request to the second network device based on the WE; the second network device is configured to complete the request and feedback it to the first engine of the first network device, where the first network device and the second network device are different network devices among multiple network devices in the system architecture 100.

[0156] As Figure 1 shown, the system architecture 100 may further include a server 110. The server 110 can be configured to create an interface corresponding to at least a part of the MR of the RDMA, and the interface information includes a unique identifier UniqueID; based on this interface, allocate the MR to multiple network devices and establish an MR page table of the network devices, where the memory access credentials mKey of multiple MR page tables are all UniqueID; multiple network devices implement memory handle exchange; bind the memory handle obtained after the exchange to the information of the WE and establish a virtual address mapping relationship to facilitate remote access through the memory handle and UniqueID.

[0157] UniqueID is used to identify the ID of the collective communication, and the synchronization of all nodes in the cluster has been completed during communication initialization. In other words, at the starting point of the RDMA data operation, by starting the kernel program and carrying UniqueID, it can ensure that the operation can be accurately associated with a specific communication task, avoiding confusion between different communication tasks.

[0158] Optionally, the hardware offload asynchronous copy instruction set may include ibcp.async(dst.addr, src.addr, length). When Dst.addr is the xPUB Reseverd VA, it means writing the memory pointed to by the xPUA local va of the local first network device to the memory of the remote second network device; when Src.addr is the xPUB Reseverd VA, it means reading the xPUB Memory of the remote second network device to the memory pointed to by the xPUA localva of the local first network device.

[0159] In some embodiments of the present disclosure, the RDMA engine of the first network device is configured to obtain a QPC corresponding to the queue label QPN of the WE based on the WE, and determine the operation type according to the QPC; the RDMA engine of the first network device is further configured to obtain the memory handle and UniqueID of the second network device based on the WE; the RDMA of the first network device completes packet encapsulation based on the WE and sends a request to the second network device.

[0160] After receiving the above request packet (e.g., data write packet), the RDMA engine of the second network device can write the payload data in the request packet into the memory of the second network device according to the remote endpoint transport header, UniqueID, and length information, where the remote endpoint transport header may include the virtual address (xPUB VA, xPUB Virtual Address) of the second network device. By offloading the asynchronous copy instruction set and WE information through hardware, the second network device can be accurately located, and the xPU of the second network device can be bypassed, reducing the intervention and overhead of the xPU, thereby significantly improving the efficiency of RDMA data transmission and avoiding the xPU from becoming the bottleneck of data transmission.

[0161] The RDMA engine of the second network device can also return an ACK (Acknowledgment, confirmation response packet) to the RDMA engine of the first network device to inform the first network device that the RDMA write request packet it sent has been successfully received and processed. This can ensure the reliability of RDMA data transmission.

[0162] After receiving the ACK, the RDMA engine of the first network device can generate a CE message according to the QP and WQE and deliver the CE message to the second engine. The second engine can assemble the received CE message, for example, making the format and content of the message more standardized and unified. Then, it is passed to the first engine in a P2P manner. By generating and delivering the CE message, the RDMA engine of the first network device can timely inform the hardware engine that the operation has been completed. This enables the hardware engine to perform subsequent processing in a timely manner, such as updating status information, releasing resources, or triggering new tasks. In addition, the P2P method avoids the latency and bandwidth contention problems brought by message transmission through intermediate nodes or shared buses. Through the P2P method, the CE message can be quickly transferred from the second engine to the first engine, improving the response speed and overall performance of the network device.

[0163] Therefore, in the RDMA data transmission solution provided by the embodiments of the present disclosure, the first engine is directly integrated inside the xPU, and the communication protocol including the hardware offloading instruction set can be triggered by this hardware engine; the second engine is a dedicated preprocessing hardware engine, which is integrated in the RNIC and can process DWQE at hardware speed, and release the main processing unit resources of the CPU and the RNIC. This reduces the number of data transfers during the RDMA data transmission process and reduces data latency; in addition, by using the hardware asynchronous instruction set, the RDMA data transmission task is directly offloaded to the dedicated hardware engine, realizing the decoupling of the RDMA data transmission task from the CPU / GPU resources. Therefore, it provides an RDMA data transmission solution that is efficient and low-latency without relying on CPU / GPU resources for data transmission-intensive scenarios such as the MoE model.

[0164] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the RDMA data transmission method described in any of the above embodiments.

[0165] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions for enabling a computer to implement the RDMA data transmission method described in any of the above embodiments when executed.

[0166] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which can implement the RDMA data transmission method described in any of the above embodiments when executed by a processor.

[0167] Figure 7 A schematic block diagram of an example electronic device 400 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0168] As Figure 7As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0169] Multiple components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0170] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the RDMA data transfer method. For example, in some embodiments, the RDMA data transfer method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the RDMA data transfer method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the RDMA data transfer method by any other appropriate means (e.g., by means of firmware).

[0171] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0172] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0173] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0174] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0175] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0176] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the deficiencies of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0177] According to the technical solution of the embodiment of the present disclosure, after receiving a training task indicating to train a generative large language model, first determine the computing power cluster composed of a single type of chip for training and the target type of chip specifically used by the computing power cluster, then determine the performance evaluation of the target type of chip under the preset performance categories, and then combine the weight distribution of each preset performance category under different alternative training strategies and use this weight distribution to perform weighted processing on the performance evaluations of the corresponding performance categories, and calculate the comprehensive training evaluation corresponding to each alternative training strategy according to the weighted performance evaluations. Finally, select the appropriate target training strategy according to the evaluation parameters of each comprehensive training evaluation. That is, by applying the technical solution provided in this embodiment, it is possible to determine the appropriate target training strategy without the need to perform trial training for each alternative training strategy one by one, reducing the resource overhead and time overhead consumed by performing trial training, and can quickly determine the target training strategy suitable for training the generative large language model by the computing power cluster composed of the target type of chip at a low cost, so as to complete the training task with only less time overhead and resource overhead, indirectly improving the model training efficiency.

[0178] It should be understood that the various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is made herein.

[0179] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A RDMA data transmission method, the method being applied to a network device including a processor xPU and an RDMA network interface controller RNIC, the xPU including a first engine, the RNIC including a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine, the method comprising: The first engine assembles a work element WE based on a hardware offload asynchronous copy instruction set, and transmits the WE to the second engine; The second engine stores the WE in the WQE Buffer and transmits it to the RDMA engine; The RDMA engine performs data processing based on the WE, wherein the data processing includes data transmission, memory access, and queue management; The first engine receives feedback of the data processing transmitted via the second engine.

2. The method according to claim 1, wherein: The hardware offload asynchronous copy instruction set includes multi-level synchronous control data transmission instructions, The multi-level synchronous control data transmission instructions include thread-level synchronous control data transmission instructions, thread group-level synchronous control data transmission instructions, storage block-level synchronous control data transmission instructions and global-level synchronous control data transmission instructions.

3. The method according to claim 1, wherein: The xPU further includes a first register in communication with the first engine, and the first engine receiving feedback of the data processing transmitted via the second engine includes: After completing the data processing, the RDMA engine generates a completion event CE message; The second engine receives the CE message, and writes the received CE message into the first register; Among them, the method of writing the received CE message into the first register includes a method of directly connecting the device to PCIe P2P, and the CE message includes at least one of an operation code Opcode, a queue pair number QPN and a tag TAG, and the TAG is used to mark the thread identifier ID of the RDMA operation request or the block ID of the source.

4. The method according to claim 3, wherein: Also includes: The first engine transmits the feedback of the data processing to the xPU that issued the RDMA operation request through the on-chip network NOC.

5. The method according to claim 1, wherein: The RNIC also includes a second register communicating with the second engine, and the first engine assembling a work element WE based on a hardware offload asynchronous copy instruction set includes: Receive RDMA operation request; Based on the hardware offload asynchronous copy instruction set, assembling the RDMA operation request into the WE; Writing the WE into the second register; Among them, the method of writing the WE into the second register includes a method of directly connecting the device to PCIe P2P, and the WE includes at least one of remote information Remote info, local information Local info, operation code Opcode, queue pair number QPN and tag TAG, and the TAG is used to mark the thread identifier ID of the RDMA operation request or the block ID of the source.

6. The method according to claim 1, wherein: The second engine stores the WE in the WQE Buffer and sends the WE to the RDMA engine, including: Parsing the WE and storing the WE in the WQE Buffer; Based on the occupancy rate of the WQE Buffer being greater than or equal to a preset threshold, the plurality of WEs stored in the WQE Buffer are transmitted to the host memory of the network device.

7. The method according to claim 6, wherein: The second engine stores the WE in the WQE Buffer and sends the WE to the RDMA engine, and further includes: Generate a doorbell signal DB, wherein the DB includes at least one of queue pair information to be processed, start addresses of the plurality of WEs, and the number of the plurality of WEs; The DB is transmitted to the RDMA engine to trigger the RDMA engine to perform the data processing. The RDMA engine reads the WE from the WQE Buffer or the host memory.

8. A RDMA data transmission method, the method being applied to a network system including a plurality of network devices, the network device including a processor xPU and an RDMA network interface controller RNIC, the xPU including a first engine, the RNIC including a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine, the method comprising: The first engine of the first network device assembles a work element WE based on a hardware offload asynchronous copy instruction set, and transmits the WE to the second engine of the first network device; The second engine of the first network device stores the WE in the WQE Buffer, and transmits the WE to the RDMA engine of the first network device; The RDMA engine of the first network device initiates a request to the second network device based on the WE; The second network device completes the request and feeds back to the first engine of the first network device, The first network device and the second network device are different network devices among the multiple network devices.

9. The method according to claim 8, further comprising: Creating an interface corresponding to at least a portion of the memory region MR of the RDMA, wherein the interface information includes a unique identifier UniqueID; Based on the interface, the MR is allocated to the plurality of network devices, and a MR page table of the network device is established, wherein the credentials mKey for accessing memory of the plurality of MR page tables are all the UniqueID; A plurality of the network devices implement memory handle exchange; The memory handle obtained after the exchange is bound to the information of the WE, and a virtual address mapping relationship is established to facilitate initiating remote access through the memory handle and the UniqueID.

10. The method according to claim 9, wherein: The RDMA engine of the first network device initiates a request to the second network device based on the WE, including: The RDMA engine of the first network device obtains a queue pair context QPC corresponding to a queue pair number QPN of the WE based on the WE, and determines an operation type according to the QPC; The RDMA engine of the first network device obtains the memory handle and the UniqueID of the second network device based on the WE; The RDMA of the first network device completes message encapsulation based on the WE, and sends a request to the second network device.

11. An RDMA data transmission network device, comprising: A processor xPU, including a first engine; An RDMA network interface controller RNIC includes a second engine communicating with the first engine, a work queue buffer WQEBuffer, and an RDMA engine communicating with the second engine, Wherein, the first engine is configured to assemble a work element WE based on a hardware offload asynchronous copy instruction set, and transmit the WE to the second engine; The second engine is configured to store the WE in the WQE Buffer and transmit it to the RDMA engine; The RDMA engine is configured to perform data processing based on the WE and transmit feedback of the data processing to the first engine via the second engine, the data processing including data transmission, memory access and queue management.

12. An RDMA data transmission network system, comprising a plurality of network devices, wherein the network device comprises a processor xPU and a network device of an RDMA network interface controller RNIC, wherein the xPU comprises a first engine, the RNIC comprises a second engine communicating with the first engine, a work queue buffer WQE Buffer, and an RDMA engine communicating with the second engine, in, The first engine of the first network device is configured to assemble a work element WE based on a hardware offload asynchronous copy instruction set, and transmit the WE to the second engine of the first network device; The second engine of the first network device is configured to store the WE in the WQE Buffer and transmit the WE to the RDMA engine of the first network device; The RDMA engine of the first network device is configured to initiate a request to the second network device based on the WE; The second network device is configured to complete the request and feed back to the first engine of the first network device, The first network device and the second network device are different network devices among the multiple network devices.

13. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the RDMA data transmission method according to any one of claims 1 to 10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the RDMA data transmission method according to any one of claims 1 to 10.

15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the RDMA data transmission method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and system for accelerating set communication through RDMA (Remote Direct Memory Access) communication

    CN113553279A

  • Data processing method, device and system and computer readable storage medium

    CN113849293A

  • RDMA communication method and device based on coroutine, and storage medium

    CN114090483A

  • Computing system, PCI device manager and initialization method thereof

    CN114138702A

  • Multi-engine multi-stream parallel operation method and system

    CN117492955A

Cited By

  • Computing chip and device, data processing method, storage medium and program product

    CN120448326A

  • Computing chip and device, data processing method, storage medium and program product

    CN120448326B

  • Data transmission method and device, chip and electronic equipment

    CN121547430A

  • Data processing method and equipment of heterogeneous accelerator, medium and program product

    CN122331998A