Data transmission method and system, and medium

By using the DPDK framework and multiple send/receive threads in user space for data transmission, the problems of frequent interruptions and performance degradation in kernel-space data transmission are solved, achieving efficient data transmission and improved system performance.

CN120825530APending Publication Date: 2025-10-21BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510895032.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In existing technologies, data transmission is frequently interrupted and suffers significant performance loss when it occurs in kernel mode, leading to increased system overhead.

Method used

By using the DPDK framework in user space, multiple sets of send and receive threads poll the cache library and perform direct memory access, data transmission between the host side and the card side is achieved, avoiding kernel mode switching and interruption.

Benefits of technology

It reduces system interruptions and performance consumption, and improves data transmission efficiency and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825530A_ABST
    Figure CN120825530A_ABST
Patent Text Reader

Abstract

The invention provides a data transmission method and system and a medium, and relates to the field of artificial intelligence, and the method comprises the steps that a host side carries out user state polling on a first cache library of the host side and a second cache library of a card side through multiple groups of receiving and transmitting threads, whether to-be-transmitted data is to be transmitted or not is determined, and the to-be-transmitted data is artificial intelligence AI model parameter data; and under the condition that the data to be transmitted exists in the first cache library and / or the second cache library, performing data transmission processing on the data to be transmitted through the multiple groups of receiving and transmitting threads, packaging threads and / or packaging threads. According to the method provided by the invention, the host side can directly read the data in the user mode through the DPDK framework, and the AI model parameter data is transmitted between the host side and the card side, so that the system interruption is reduced, the overhead and the performance consumption are reduced, and the system performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a data transmission method, system, and medium. Background Art

[0002] The implementation of computing power acceleration server systems is usually built based on the Linux system, and usually uses PCIE drivers to implement high-speed data communication and signaling transmission in the kernel state. Summary of the Invention

[0003] The present disclosure provides a data transmission method, system, and medium to solve the problems in related technologies of frequent interruptions, high overhead, and performance loss caused by the need to perform data transmission in kernel mode.

[0004] A first aspect embodiment of the present disclosure proposes a data transmission method, including: performing user-mode polling on a first cache library on a host side and a second cache library on a card side through multiple groups of transceiver threads to determine whether there is data to be transmitted, where the data to be transmitted is artificial intelligence (AI) model parameter data; when there is data to be transmitted in the first cache library and / or the second cache library, performing data transmission processing on the data to be transmitted through multiple groups of transceiver threads, packet assembly threads and / or packet receiving threads.

[0005] In some embodiments of the present disclosure, data polling is performed on the first cache library on the host side and the second cache library on the card side respectively through multiple groups of transceiver threads to determine whether there is data to be transmitted, including: performing user-state polling on the first cache library through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache library, and performing user-state polling on the second cache library through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache library; when there is training parameter data to be transmitted in the first cache library and / or training result data to be transmitted in the second cache library, it is determined that there is data to be transmitted.

[0006] In some embodiments of the present disclosure, when there is data to be transmitted in the first cache library and / or the second cache library, data transmission processing is performed on the data to be transmitted through multiple groups of sending and receiving threads, packet assembly threads and / or packet receiving threads, including at least one of the following: through direct memory access, the training parameter data to be transmitted is sent to the card side using the packet assembly thread and the first sending thread; through direct memory access, the training result data to be transmitted is received from the card side using the packet receiving thread and the first receiving thread.

[0007] In some embodiments of the present disclosure, direct memory access is used to send the training parameter data to be transmitted to the card side using a packet assembly thread and a first sending thread, including: packetizing the first training parameter data to be transmitted in the first cache library through direct memory access and a packet assembly thread to obtain a first data packet; using the first sending thread to send the first data packet to the card side, so that the card side uses the first training parameter data to be transmitted in the first data packet to perform model training.

[0008] In some embodiments of the present disclosure, direct memory access is used to receive training result data to be transmitted from the card side using a packet receiving thread and a first receiving thread, including: receiving the first training result data to be transmitted in the second cache through direct memory access and the packet receiving thread to obtain a second data packet; using the first receiving thread to receive the second data packet from the card side; notifying the upper-layer artificial intelligence AI application corresponding to the first training result data to be transmitted, so that the upper-layer AI application uses the second data packet to perform model data processing.

[0009] In some embodiments of the present disclosure, the method also includes: initializing the data plane development kit DPDK framework and registering multiple groups of transceiver threads; determining the corresponding board on the card side based on the vendor identifier and the device identifier; abstracting the hardware capabilities of the board corresponding to the card side and starting multiple groups of transceiver threads; and binding the multiple groups of transceiver threads to the CPU core.

[0010] In the above embodiment, the data transmission method directly reads data in user mode through the DPDK framework on the host side, and transmits AI model parameter data between the host side and the card side, thereby reducing system interruptions, reducing overhead and performance consumption, and improving system performance.

[0011] The second aspect of the present disclosure proposes a data transmission system, including: a host side and a card side, the host side including: a first cache library, a polling module, and a transceiver module; the card side including: a second cache library; the polling module including multiple groups of transceiver threads, the polling module is used to perform user-mode polling on the first cache library on the host side and the second cache library on the card side through the multiple groups of transceiver threads respectively to determine whether there is data to be transmitted, and the data to be transmitted is artificial intelligence AI model parameter data; the transceiver module includes a packet assembly thread and a packet receiving thread, and the transceiver module is used to perform data transmission processing on the data to be transmitted through the multiple groups of transceiver threads, the packet assembly thread and / or the packet receiving thread when there is data to be transmitted in the first cache library and / or the second cache library.

[0012] In some embodiments of the present disclosure, the polling module is further used to: perform user-state polling on the first cache library through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache library, and perform user-state polling on the second cache library through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache library; when there is training parameter data to be transmitted in the first cache library and / or training result data to be transmitted in the second cache library, determine that there is data to be transmitted.

[0013] In some embodiments of the present disclosure, the transceiver module is also used to: package the first training parameter data to be transmitted in the first cache library through direct memory access and a package package thread to obtain a first data packet; use the first sending thread to send the first data packet to the card side, so that the card side uses the first training parameter data to be transmitted in the first data packet to perform model training.

[0014] In some embodiments of the present disclosure, the transceiver module is also used to: receive and process the first training result data to be transmitted in the second cache library through direct memory access and a packet receiving thread to obtain a second data packet; use the first receiving thread to receive the second data packet from the card side; the host side also includes a calling module for notifying the upper-layer artificial intelligence AI application corresponding to the first training result data to be transmitted, so that the upper-layer AI application can use the second data packet to perform model data processing.

[0015] In some embodiments of the present disclosure, the polling module is also used to initialize the DPDK framework and register multiple groups of transceiver threads; determine the corresponding board on the card side based on the vendor identifier and device identifier; abstract the hardware capabilities of the board corresponding to the card side, and start multiple groups of transceiver threads, and bind the multiple groups of transceiver threads to the CPU core for processing.

[0016] In the above embodiment, the data transmission system uses the DPDK framework to directly read data in user mode and transmit AI model parameter data, thereby reducing system interruptions, reducing overhead and performance consumption, and improving system performance.

[0017] An embodiment of a third aspect of the present disclosure provides a data transmission device, which is configured to execute any one of the methods described in the first aspect of the present disclosure.

[0018] An embodiment of the fourth aspect of the present disclosure provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor, when used to run the computer program, executes any one of the methods described in the first aspect of the present disclosure.

[0019] The fifth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute any one of the methods described in the first aspect of the present disclosure.

[0020] In summary, according to the data transmission method proposed in the present invention, the host side directly accesses the memory and performs user-state polling on the first cache library and the second cache library, so that when there is data to be transmitted, multiple groups of sending and receiving threads and packet assembly threads and / or packet receiving threads are used to transmit the AI ​​model parameter data between the host side and the card side, thereby avoiding system interruption, reducing system overhead and performance consumption, and improving system performance.

[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0023] Figure 1 This is a schematic diagram of the kernel driver framework;

[0024] Figure 2 Processing logic diagram for DPDK;

[0025] Figure 3 A flowchart of a data transmission method provided in an embodiment of the present disclosure;

[0026] Figure 4 A schematic diagram of a data transmission process according to an embodiment of the present disclosure;

[0027] Figure 5 A schematic diagram of a data transmission process according to an embodiment of the present disclosure;

[0028] Figure 6 An architectural diagram of a data transmission system provided in an embodiment of the present disclosure;

[0029] Figure 7 This is a schematic diagram of a resident thread;

[0030] Figure 8 A schematic structural diagram of a data transmission device provided in an embodiment of the present disclosure;

[0031] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0033] like Figure 1 As shown in the kernel driver framework diagram, the general kernel driver framework has disadvantages such as high complexity, lack of high-level abstraction layer support and high difficulty in development; in system application, it has disadvantages such as long call chain, high system call overhead and frequent interruptions; in maintenance and updating, it has disadvantages such as kernel version compatibility issues and high maintenance costs.

[0034] The Data Plane Development Kit (DPDK) software is a set of userspace libraries and drivers that accelerate packet processing workloads running on all major CPU architectures. Figure 2 As shown in the DPDK processing logic diagram, DPDK bypasses the Linux kernel and performs packet processing in the userland to maximize I / O performance. DPDK achieves this by using a polling mode driver (PMD) running in the userland, constantly checking the incoming packet queue to see if new data has arrived, thereby reducing the latency of the system call link and improving the throughput of the data transmission path.

[0035] However, general-purpose AI accelerator cards typically use standard PCIE drivers on the host side to implement upper-layer AI model data transmission, which has the following disadvantages:

[0036] 1. The training data of the AI ​​model needs to be transmitted in kernel state. Each transmission takes about 100 microseconds, not including cache misses caused by context switching.

[0037] 2. Data must be copied between kernel mode and user mode, which results in a lot of CPU consumption and global lock contention.

[0038] 3. There is a system call overhead for sending and receiving data.

[0039] 4. Currently, when most kernel-mode drivers work in multi-core server scenarios, in order to ensure multi-core consistency, they will cause performance loss due to bus locking and memory barriers.

[0040] 5. The communication link from AI applications to the computing core of the AI ​​accelerator card is too long, and the software design and maintenance costs are high.

[0041] In order to solve the problems existing in the related art, the present disclosure proposes a data transmission method and system. Aiming at the drawbacks of the kernel driver solution, a data transmission system is proposed to achieve efficient data packet forwarding.

[0042] The data transmission system and data transmission method proposed in this disclosure can be applied to AI acceleration scenarios in data centers, and can be added to large server hosts in the form of PCIE boards to provide acceleration for AI model inference training.

[0043] The data transmission system and data transmission method proposed in the present disclosure can provide the host-side underlying software stack capabilities for the AI ​​accelerator card, realizing high-speed unloading of AI model parameter data.

[0044] The data transmission system and data transmission method proposed in the present disclosure can be applied to large-scale model training scenarios in the cloud, as well as to real-time edge data processing scenarios, etc.

[0045] The data transmission method and data transmission system proposed in the present disclosure are applied to multi-core server scenarios.

[0046] The data transmission method, system, and medium provided in this application are described in detail below with reference to the accompanying drawings.

[0047] Figure 3 The data transmission method of the present disclosure is executed by the host side, such as Figure 3 As shown, the following steps are included:

[0048] Step 301 : Perform user-mode polling on a first cache bank on the host side and a second cache bank on the card side through multiple groups of transceiver threads to determine whether there is data to be transmitted.

[0049] In some embodiments, the data to be transmitted is artificial intelligence (AI) model parameter data.

[0050] In some embodiments, multiple groups of transceiver threads, i.e., one sending thread and one receiving thread, constitute a group of transceiver threads, and the multiple groups of transceiver threads include at least two groups of transceiver threads. The specific number of groups of transceiver threads can be preset, which is not limited by the present disclosure.

[0051] In some embodiments, user-mode polling can be a PMD (Poll Mode Driver) in DPDK (Data Plane Development Kit). Since DPDK bypasses the kernel, it avoids the overhead of copying data from kernel mode to user mode and the overhead of interrupt response context switching, thereby achieving higher efficiency.

[0052] In some embodiments, user-mode polling is used on the host side to avoid interruptions caused by kernel-mode related processing and reduce switching overhead caused by interruptions.

[0053] In some embodiments, the sending thread and receiving thread in the multiple groups of sending and receiving threads can be used to perform user-mode polling on the host-side cache and the card-side cache respectively to determine whether there is data in the cache that needs to be processed for data transmission.

[0054] In some embodiments, the DPDK framework applied on the host side adopts large page memory technology, that is, the first cache library and the second cache library adopt large page memory technology to reduce the problem of cache failure.

[0055] In the above embodiment, the large page memory technology can reduce TLB Miss (Translation Lookaside Buffer Miss) and Cache Miss.

[0056] In some embodiments, the data to be transmitted is AI model parameter data, that is, data including AI model training data and AI model training results, so as to achieve high-speed unloading of AI model parameter data.

[0057] In some embodiments, user-state polling is performed on the first cache library on the host side and the second cache library on the card side through multiple groups of transceiver threads to determine whether there is data to be transmitted, including: performing user-state polling on the first cache library through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache library, and performing user-state polling on the second cache library through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache library; when there is training parameter data to be transmitted in the first cache library and / or training result data to be transmitted in the second cache library, it is determined that there is data to be transmitted.

[0058] In some embodiments, multiple sending threads are used to perform user-state polling on the first cache library on the host side to determine whether there is training parameter data to be transmitted that is unloaded from the host side to the card side, and multiple receiving threads are used to perform user-state polling on the second cache library on the card side to determine whether there is training result data to be transmitted that is returned from the card side to the host side.

[0059] In some embodiments, user-state polling of the first cache by multiple sending threads and user-state polling of the second cache by multiple receiving threads are performed synchronously, that is, the sending threads and the receiving threads perform user-state polling at the same time to determine whether there is data to be transmitted in the cache.

[0060] For example, when the PMD driver is enabled, there will be two groups of threads that respectively detect the data (training parameters) unloaded from the host to the EP (card side) and the data (training results) returned by the EP to the host. Therefore, the PMD driver will poll as a user-mode resident thread to see whether there is any training data that needs to be unloaded to the card side or whether there are any training results that need to be returned to the host side for calculation.

[0061] In some embodiments, when there is training parameter data to be transmitted in the first cache, it is determined that there is training parameter data to be transmitted.

[0062] In some embodiments, when there is training result data to be transmitted in the second cache, it is determined that there is training result data to be transmitted.

[0063] In some embodiments, when there is training parameter data to be transmitted in the first cache and there is training result data to be transmitted in the second cache, it is determined that there is training parameter data to be transmitted and training result data to be transmitted.

[0064] In some embodiments, when there is no training parameter data to be transmitted in the first cache and no training result data to be transmitted in the second cache, it is determined that there is no data to be transmitted, and data transmission processing is not required.

[0065] In some embodiments, the method also includes: initializing the data plane development kit DPDK framework and registering multiple groups of transceiver threads; determining the corresponding board on the card side based on the vendor identifier and device identifier; abstracting the hardware capabilities of the board corresponding to the card side and starting multiple groups of transceiver threads; binding the multiple groups of transceiver threads to the CPU core.

[0066] In some embodiments, a DPDK framework is deployed on the host side, and multiple groups of sending and receiving threads are registered in the DPDK framework.

[0067] For example, after the DPDK framework is initialized, the PMD driver will query the board of the corresponding manufacturer based on the Vendor ID and Device ID.

[0068] In some embodiments, hardware capability abstraction for the corresponding card side board can be achieved by hiding the underlying complex implementation and exposing a simplified interface. Hardware capability abstraction can be achieved through a hardware abstraction layer, instruction set abstraction, or functional model. The specific method of hardware capability abstraction is not limited in this disclosure.

[0069] In some embodiments, the hardware capability abstraction of the board is, for example, the hardware capability abstraction of DMA (Direct Memory Access), NPU (Neural Network Processing Unit), C2C (chip-to-chip communication), etc.

[0070] In some embodiments, binding multiple groups of transceiver threads to the CPU core may be performed by affinity binding, so that the multiple groups of transceiver threads have a binding relationship with the CPU core, that is, each group of transceiver threads has a one-to-one binding relationship with the CPU core.

[0071] In some embodiments, when the present disclosure is applied to a multi-core server scenario, the binding relationship between multiple groups of sending and receiving threads on the host side and the CPU core can reduce the scheduling switching overhead generated by data transmission between multiple CPU cores.

[0072] For example, the hardware capabilities of the board are abstracted, such as DMA, NPU, C2C, etc., and then the sending thread and the receiving thread are started, and affinity binding is performed, and then polling is started to see if there are any incoming data packets.

[0073] In the above embodiment, by using the DPDK framework and binding multiple groups of sending and receiving threads to the CPU core, the performance consumption caused by thread scheduling can be reduced, and the program execution efficiency can be further improved.

[0074] Step 302 : When there is data to be transmitted in the first cache and / or the second cache, data transmission processing is performed on the data to be transmitted by using multiple groups of sending and receiving threads, packet assembly threads and / or packet receiving threads.

[0075] In some embodiments, multiple groups of sending and receiving threads, packet forming threads, and packet receiving threads may be resident threads.

[0076] In some embodiments, when there is data to be transmitted in the first cache, data transmission processing is performed on the data to be transmitted through multiple groups of transceiver threads and packet assembly threads.

[0077] In some embodiments, when there is data to be transmitted in the second cache, data transmission processing is performed on the data to be transmitted through multiple groups of sending and receiving threads and packet receiving threads.

[0078] In some embodiments, when there is data to be transmitted in the first cache and the second cache, data transmission processing is performed on the data to be transmitted through multiple groups of sending and receiving threads, packet assembly threads, and packet receiving threads.

[0079] In some embodiments, data transmission processing is performed on the data to be transmitted through multiple groups of sending and receiving threads, packet assembly threads and / or packet receiving threads, including at least one of the following: through direct memory access, using the packet assembly thread and the first sending thread to send the training parameter data to be transmitted to the card side; through direct memory access, using the packet receiving thread and the first receiving thread to receive the training result data to be transmitted from the card side.

[0080] In some embodiments, when there is training parameter data to be transmitted in the first cache, the training parameter data to be transmitted is sent to the card side through direct memory access using a package assembly thread and a first sending thread, wherein the first sending thread is any one of a plurality of sending threads, that is, the package assembly thread and the first sending thread are used to unload the training parameter data to be transmitted from the host side to the card side.

[0081] In some embodiments, when there is training result data to be transmitted in the second cache, the training result data to be transmitted is received from the card side using a packet receiving thread and a first receiving thread through direct memory access, wherein the first receiving thread is any one of a plurality of receiving threads, that is, the packet receiving thread and the first receiving thread return the training result data to be transmitted from the card side to the host side.

[0082] In some embodiments, when there is training parameter data to be transmitted in the first cache library and training result data to be transmitted in the second cache library, the training parameter data to be transmitted is unloaded from the host side by direct memory access and the packet assembly thread and the first sending thread are used respectively, and the training result data to be transmitted is returned from the card side to the host side by using the packet receiving thread and the first receiving thread.

[0083] In the above embodiment, user-mode polling is used on the host side to poll the first cache library and the second cache library to determine whether there is training parameter data to be transmitted and training result data to be transmitted, so as to use direct memory access to perform data transmission processing of the data to be transmitted between the host side and the card side through a packet assembly thread and / or a packet receiving thread and multiple groups of sending and receiving threads, so as to avoid the complexity and risk of kernel-mode driving, reduce the overhead of interrupt processing and memory copying, and improve system performance and data transmission efficiency.

[0084] Figure 4 A flow chart of a data transmission method provided in an embodiment of the present disclosure, based on Figure 3 The embodiment shown, as Figure 4 As shown, Figure 4 right Figure 3 Step 302 in the embodiment is further described, and includes the following steps:

[0085] Step 401 : Packaging the first training parameter data to be transmitted in the first cache through direct memory access and a packet assembly thread to obtain a first data packet.

[0086] In some embodiments, the packet assembly process may be a process in which the packet assembly thread retrieves the first training parameter data to be transmitted from the first cache through direct memory access, and encapsulates the first training parameter data to be transmitted into a complete data packet that meets protocol requirements according to preset rules.

[0087] In some embodiments, the preset rules can be customized according to different needs or scenarios, which is not limited by this disclosure.

[0088] In some embodiments, the first training parameter data to be transmitted may be any one of the training parameter data to be transmitted.

[0089] In some embodiments, direct memory access can bypass the kernel to directly interact with the first cache library for data, thereby achieving direct data transmission and reducing memory copy overhead.

[0090] For example, applications based on DPDK can directly access the system cache for data transmission. User-space applications can directly access the socket buffer through user-mode drivers and enter the receive / send queue through direct memory access (DMA), thereby avoiding interaction with the kernel, reducing memory copy overhead, and achieving direct data transmission.

[0091] Step 402: Use a first sending thread to send a first data packet to the card side, so that the card side uses the first to-be-transmitted training parameter data in the first data packet to perform model training.

[0092] In some embodiments, the first sending thread is any one of a plurality of sending threads, and is used for data transmission processing of a piece of training parameter data to be transmitted, that is, data transmission processing of the first piece of training parameter data to be transmitted.

[0093] In some embodiments, the first training parameter data to be transmitted is the AI ​​model parameter data unloaded from the host side to the card side, and the card side uses the first training parameter data to be transmitted to perform model training of the AI ​​model.

[0094] For example, the direction from host to EP is the direction of loading model data, and the upper-layer model training data needs to be abstracted into data of custom protocol packages for unloading.

[0095] In the above embodiment, through direct memory access, using the package assembly thread and the first sending thread, the training parameter data to be transmitted on the host side can be unloaded to the card side, realizing user-mode driver to perform direct hardware register access and data transmission, reducing the design difficulty of the underlying software of the AI ​​accelerator card, and increasing the unloading rate of AI model training data, thereby improving system performance.

[0096] Figure 5A schematic diagram of a data transmission process provided by an embodiment of the present disclosure, based on Figure 3-Figure 4 The embodiment shown, Figure 5 right Figure 3 Step 302 in the embodiment is further described as follows: Figure 5 As shown, the following steps are included:

[0097] Step 501 : Through direct memory access and a packet receiving thread, the first training result data to be transmitted in the second cache is processed to obtain a second data packet.

[0098] In some embodiments, the packet receiving process may be a process in which the packet receiving thread obtains the training result data to be transmitted from the cache library on the card side, and parses and verifies the data according to preset rules.

[0099] In some embodiments, the preset rules can be customized according to different needs or scenarios, which is not limited by this disclosure.

[0100] In some embodiments, the first training result data to be transmitted may be any one of the training result data to be transmitted.

[0101] In some embodiments, direct memory access can bypass the kernel to directly interact with the second cache library for data, thereby achieving direct data transmission and reducing memory copy overhead.

[0102] For example, in this solution, the application can directly access the system cache based on DPDK for data transmission; while in the kernel-mode driver solution, the application needs to first access the system call in the kernel space to access the driver through the software stack, and the driver accesses the buffer to copy and transmit data packets. Compared with the kernel-mode driver, this solution uses the user-mode driver to reduce the overhead of interrupt processing and kernel access, improve system performance, and improve data transmission efficiency.

[0103] For example, user-space applications can directly access the socket buffer through user-mode drivers and enter the receive / transmit queue through direct memory access (DMA). Compared to kernel-mode driver solutions, this can avoid the overhead of accessing kernel space, improve system performance, and increase data transmission efficiency.

[0104] Step 502: Use the first receiving thread to receive a second data packet from the card side.

[0105] In some embodiments, the first receiving thread is any one of a plurality of receiving threads, and is used for data transmission processing of a piece of training result data to be transmitted, that is, data transmission processing of the first piece of training result data to be transmitted.

[0106] In some embodiments, the first training result data to be transmitted is the AI ​​model result data returned by the card side to the host side, and the host side uses the training result data for subsequent processing.

[0107] For example, the EP side is in the direction of receiving model training results from the host, and is responsible for returning the parameter data of the completed AI model training to the host.

[0108] Step 503: Notify the upper-layer artificial intelligence (AI) application corresponding to the first training result data to be transmitted, so that the upper-layer AI application uses the second data packet to perform model data processing.

[0109] In some embodiments, after the host side receives the second data packet of the first training result data to be transmitted, it needs to notify the corresponding AI application so that the AI ​​application can use the second data packet to perform model data processing.

[0110] In some embodiments, for different AI applications or second data packets, the notification methods may be different or the same, and the model data processing methods may be different or the same, which is not limited by the present disclosure.

[0111] In some embodiments, the method of notifying the upper-layer AI application can be to use an API call to notify the upper-layer AI application, which can be a synchronous API call or an asynchronous API call. Among them, the synchronous API call is, for example, after the training is completed, the host side directly uses HTTP POST to request the interface of the upper-layer application, carrying the second data packet storage path or metadata.

[0112] In some embodiments, the method for notifying the upper-layer AI application may be to use a message queue to notify the upper-layer AI application; or the upper-layer AI application may monitor a parameter file, trigger a callback, and notify the upper-layer AI application through the callback function; or the training process may write parameters to a shared memory area and wake up the application process through a semaphore.

[0113] In some embodiments, model data processing may be performed differently for different AI applications, such as using a second data packet to update the model, or starting model evaluation, restarting the inference service process, or hot loading the model server.

[0114] In the above embodiment, through direct memory access, using the packet receiving thread and the first receiving thread, the training result data to be transmitted on the card side can be returned to the host side, realizing user-mode driver to perform direct hardware register access and data transmission, reducing the design difficulty of the underlying software of the AI ​​accelerator card, and improving the transmission rate of AI model training data, thereby improving system performance.

[0115] In summary, according to the data transmission method proposed in the present invention, the host side directly accesses the memory and performs user-mode polling on the first cache library and the second cache library, so that when there is data to be transmitted, multiple groups of sending and receiving threads and packet assembly threads and / or packet receiving threads are used to perform data transmission processing on the host side and the card side, thereby avoiding system interruption, reducing system overhead and performance consumption, and improving system performance.

[0116] Figure 6 This is an architecture diagram of a data transmission system proposed in an embodiment of the present disclosure, such as Figure 6 As shown, the data transmission system 600 includes a host side 610 and a card side 620 .

[0117] The host side includes: a first buffer library 611, a polling module 612, and a transceiver module 613;

[0118] The card side includes: a second cache library 621;

[0119] The polling module 612 includes multiple groups of transceiver threads, which are bound to the CPU core. The polling module is used to perform user-mode polling on the first cache library on the host side and the second cache library on the card side through the multiple groups of transceiver threads to determine whether there is data to be transmitted.

[0120] The transceiver module 613 includes a packet assembly thread and a packet receiving thread. When there is data to be transmitted in the first cache library and / or the second cache library, the transceiver module is used to perform data transmission processing on the data to be transmitted through multiple groups of transceiver threads, packet assembly threads and / or packet receiving threads.

[0121] In some embodiments, the polling module can be an NPU PMD driver module, the transceiver module can be an NPU LIB module, the transceiver module can provide a calling API for the upper-level AI application, and the training data of the upper-level AI model can directly call the PMD driver through the LIB interface, using the hardware DMA capability to achieve efficient data transmission from the host to the board.

[0122] For example, in the computing power acceleration card system of this solution, the NPUPMD driver and NPU LIB are mainly implemented on the host side based on the DPDK framework. The PMD driver is responsible for abstracting the AI ​​accelerator card hardware and implementing polling sending and receiving of data packets. The NPU LIB provides a calling API for the upper-level AI framework. The training data of the upper-level AI model can be directly called to the underlying PMD driver through the LIB interface, and efficient data transmission from the host to the board is achieved with the help of the hardware DMA capability. This process achieves zero copy data transmission and greatly reduces the number of hardware interrupts. At the same time, the data sending and receiving threads are bound to the CPU affinity to reduce the cache miss of the process, which is suitable for AI data center acceleration scenarios.

[0123] In some embodiments, the polling module is also used to: perform user-state polling on the first cache library through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache library, and perform user-state polling on the second cache library through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache library; when there is training parameter data to be transmitted in the first cache library and / or training result data to be transmitted in the second cache library, determine that there is data to be transmitted.

[0124] For example, the PMD driver polls as a user-mode resident thread to determine whether any training data needs to be offloaded to the card for calculation, and performs software encapsulation of the hardware DMA to implement channels for hardware multiplexing. When the PMD driver is enabled, two groups of threads will respectively detect the data offloaded from the host to the EP (training parameters) and the data returned by the EP to the host (training results).

[0125] In some embodiments, the transceiver module is also used to: package the first training parameter data to be transmitted in the first cache library through direct memory access and a package package thread to obtain a first data packet; use the first sending thread to send the first data packet to the card side, so that the card side uses the first training parameter data to be transmitted in the first data packet to perform model training.

[0126] In some embodiments, the transceiver module is further used to: receive and process the first training result data to be transmitted in the second cache through direct memory access and a packet receiving thread to obtain a second data packet; and use the first receiving thread to receive the second data packet from the card side.

[0127] For example, a resident thread such as Figure 7As shown, the direction from host to EP is the direction of loading model data, abstracting the upper-layer model training data into data of custom protocol packets for unloading, wherein data is loaded through npu_enqueue_buf, npu_gen_txll is the package assembly thread, and npu_xmit_txll is the sending thread; the direction from EP to host is the direction of receiving model training results, which is responsible for returning the parameter data of the AI ​​model training completed to the host side and notifying the upper-layer AI application for corresponding processing, wherein npu_gen_rxll is the package assembly thread, npu_xmit_rxll is the packet receiving thread, and data is obtained through npu_dequeue_buf.

[0128] In some embodiments, the host side also includes a calling module 614, which is used to notify the upper-layer artificial intelligence AI application corresponding to the first training result data to be transmitted, so that the upper-layer AI application can use the second data packet to perform model data processing.

[0129] For example, the EP is responsible for receiving model training results from the host, returning the parameter data of the AI ​​model training to the host side, and notifying the upper-layer AI application to perform corresponding processing.

[0130] In some embodiments, the polling module is further used to initialize the DPDK framework and register multiple groups of sending and receiving threads; and determine the corresponding board on the card side according to the vendor identifier and the device identifier.

[0131] In some embodiments, the polling module is further used to abstract the hardware capabilities of the corresponding board on the card side, and start multiple groups of transceiver threads; and bind the multiple groups of transceiver threads to the CPU core for processing.

[0132] For example, after the DPDK framework is initialized, the PMD driver probes the corresponding manufacturer's card based on the Vendor ID and Device ID. It abstracts the card's hardware capabilities, such as DMA, NPU, and C2C, starts the send and receive threads, performs affinity binding, and then begins polling for incoming data packets.

[0133] In the above embodiment, the data transmission system is used to perform Figure 3-Figure 5 The data transmission method shown is used to realize data transmission between the host side and the card side, and to realize high-speed unloading of the training data of the AI ​​model to the card side and return of the result data of the AI ​​model to the host side. It will not be repeated here.

[0134] The following is a specific implementation of a data transmission method driven by user mode based on DPDK (Data Plane Development Kit):

[0135] This solution directly operates hardware registers in upper-level applications, reads underlying data packets, and reduces system calls. At the same time, it uses CPU affinity binding and huge page (huge page memory technology) to reduce TLB Miss (Transfer Look-aside Table, high-speed address index cache miss) and Cache Miss, thereby achieving efficient data forwarding.

[0136] This solution is used in the computing power acceleration card system. On the host side, the NPU PMD driver and NPU LIB are mainly implemented based on the DPDK framework. The PMD driver is responsible for abstracting the AI ​​accelerator card hardware and implementing polling and receiving of data packets. The NPU LIB provides a calling API for the upper-level AI framework. The training data of the upper-level AI model can be directly called to the underlying PMD driver through the LIB interface, and efficient data transmission from the host to the board is achieved with the help of the hardware DMA capability. This process achieves zero copy data transmission and greatly reduces the number of hardware interrupts. At the same time, the data sending and receiving threads are bound to the CPU affinity to reduce the cache miss of the process, which is suitable for AI data center acceleration scenarios.

[0137] System implementation comparison: In the original kernel-mode driven system, the application must first access the kernel space system call to access the driver through the software stack, and then the driver accesses the buffer to copy and transmit data packets. In the system of this solution, to address the shortcomings of the general kernel device driver, this solution introduces the DPDK tool library, implements user-mode operation based on the DPDK tool library, and optimizes the application's data processing overhead.

[0138] Data link comparison: In the original kernel-mode driver solution, user-space applications must first access the kernel-mode socket buffer, driver, and ring buffer to enter the receive / send queue and send data to the driver. In this solution, user-space applications can directly access the socket buffer through the user-mode driver and enter the receive / send queue through direct memory access (DMA).

[0139] In the AI ​​model data offload scenario, the PMD driver polls as a user-mode resident thread to determine whether there is training data that needs to be offloaded to the card for calculation. It also encapsulates the hardware DMA in software to implement channels for hardware multiplexing. When the PMD driver is enabled, two groups of threads will respectively detect the data offloaded from the host to the EP (training parameters) and the data returned by the EP to the host (training results).

[0140] First, after the DPDK framework is initialized, the PMD driver will query the corresponding manufacturer's board based on the Vendor ID and Device ID.

[0141] Abstract the hardware capabilities of the board, such as DMA, NPU, C2C, etc., start the sending thread and receiving thread, and perform affinity binding, then start polling for incoming data packets, resident threads such as Figure 7 As shown, the direction from host to EP is the direction of loading model data, abstracting the upper-layer model training data into data of custom protocol packets for unloading, wherein data is loaded through npu_enqueue_buf, npu_gen_txll is the package assembly thread, and npu_xmit_txll is the sending thread; the direction from EP to host is the direction of receiving model training results, which is responsible for returning the parameter data of the AI ​​model training completed to the host side and notifying the upper-layer AI application for corresponding processing, wherein npu_gen_rxll is the package assembly thread, npu_xmit_rxll is the packet receiving thread, and data is obtained through npu_dequeue_buf.

[0142] In summary, this solution has the following beneficial effects:

[0143] 1. Implemented user-mode driver, bypassing the operating system kernel and processing data packets directly in user space, while avoiding the complexity and risks of kernel driver;

[0144] 2. User-mode drivers and polling mode reduce interrupt processing and memory copy overhead, thereby improving performance. At the same time, large page memory technology reduces TLB misses and improves memory access efficiency.

[0145] 3. Threads can be bound to specific CPU cores to reduce the performance consumption caused by thread scheduling and further improve program execution efficiency.

[0146] Figure 8 Schematic diagram of a data transmission device 800 provided in an embodiment of the present disclosure. Figure 8 As shown, the device includes:

[0147] The polling module 810 is used to perform user-mode polling on the first cache library on the host side and the second cache library on the card side through multiple groups of transceiver threads to determine whether there is data to be transmitted, and the data to be transmitted is artificial intelligence AI model parameter data.

[0148] The transmission module 820 is configured to, when there is data to be transmitted in the first cache and / or the second cache, perform data transmission processing on the data to be transmitted through multiple groups of sending and receiving threads, packet assembly threads and / or packet receiving threads.

[0149] In some embodiments, the polling module is also used to perform user-state polling on the first cache library through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache library, and to perform user-state polling on the second cache library through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache library; when there is training parameter data to be transmitted in the first cache library and / or training result data to be transmitted in the second cache library, determine that there is data to be transmitted.

[0150] In some embodiments, the transmission module is also used to send the training parameter data to be transmitted to the card side through direct memory access, using the package assembly thread and the first sending thread; and to receive the training result data to be transmitted from the card side through direct memory access, using the package receiving thread and the first receiving thread.

[0151] In some embodiments, the transmission module is also used to package the first training parameter data to be transmitted in the first cache library through direct memory access and a package package thread to obtain a first data packet; and use the first sending thread to send the first data packet to the card side, so that the card side uses the first training parameter data to be transmitted in the first data packet to perform model training.

[0152] In some embodiments, the transmission module is also used to receive and process the first training result data to be transmitted in the second cache through direct memory access and a packet receiving thread to obtain a second data packet; use the first receiving thread to receive the second data packet from the card side; and notify the upper-layer artificial intelligence AI application corresponding to the first training result data to be transmitted, so that the upper-layer AI application can use the second data packet to perform model data processing.

[0153] In some embodiments, the polling module is also used to initialize the data plane development kit DPDK framework and register multiple groups of transceiver threads; determine the corresponding board on the card side based on the vendor identifier and device identifier; abstract the hardware capabilities of the board corresponding to the card side and start multiple groups of transceiver threads; bind the multiple groups of transceiver threads to the CPU core for processing.

[0154] In summary, the data transmission device proposed in the present disclosure directly reads data in user mode by using the DPDK framework, and transmits AI model parameter data, thereby reducing system interruptions, reducing overhead and performance consumption, and improving system performance.

[0155] In the embodiments provided above, the methods and devices provided in the embodiments of the present application are introduced. In order to implement the various functions of the methods provided in the embodiments of the present application, the electronic device may include a hardware structure and a software module, and implement the aforementioned functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. One of the aforementioned functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.

[0156] Figure 9 FIG1 is a block diagram of an electronic device 900 for implementing the above-mentioned data transmission method according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0157] Reference Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902 , a memory 904 , a power component 906 , a multimedia component 908 , an audio component 910 , an input / output (I / O) interface 912 , a sensor component 914 , and a communication component 916 .

[0158] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 902 may include one or more modules to facilitate interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate interaction between the multimedia component 908 and the processing component 902.

[0159] The memory 904 is configured to store various types of data to support operations on the electronic device 900. Examples of such data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0160] The power supply component 906 provides power to the various components of the electronic device 900. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 900.

[0161] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0162] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), and when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 also includes a speaker for outputting audio signals.

[0163] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0164] The sensor assembly 914 includes one or more sensors for providing various aspects of status assessment for the electronic device 900. For example, the sensor assembly 914 can detect the open / closed state of the electronic device 900, the relative positioning of components, such as the display and keypad of the electronic device 900. The sensor assembly 914 can also detect changes in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and temperature changes of the electronic device 900. The sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0165] The communication component 916 is configured to facilitate wired or wireless communication between the electronic device 900 and other devices. The electronic device 900 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0166] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.

[0167] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the instructions can be executed by the processor 920 of the electronic device 900 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0168] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the data transmission method described in the above embodiments of the present disclosure.

[0169] An embodiment of the present disclosure further provides a computer program product, including a computer program. The computer program is used by a processor to execute the data transmission method described in the above embodiment of the present disclosure.

[0170] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0171] Throughout this specification, references to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" indicate that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0172] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0173] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0174] It should be understood that various parts of the embodiments of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0175] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0176] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in either hardware or software functional modules. If the integrated modules are implemented as software functional modules and sold or used as standalone products, they may also be stored in a computer-readable storage medium. The aforementioned storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0177] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A data transmission method, characterized in that: The method is executed by the host side, and the method includes: Performing user-mode polling on the first cache bank on the host side and the second cache bank on the card side through multiple groups of transceiver threads to determine whether there is data to be transmitted, where the data to be transmitted is artificial intelligence (AI) model parameter data; When there is data to be transmitted in the first cache and / or the second cache, data transmission processing is performed on the data to be transmitted through the multiple groups of sending and receiving threads, packet assembly threads and / or packet receiving threads.

2. The method according to claim 1, characterized in that The method of performing user-mode polling on the first cache bank on the host side and the second cache bank on the card side through multiple groups of transceiver threads to determine whether there is data to be transmitted includes: Performing the user-state polling on the first cache through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache, and performing the user-state polling on the second cache through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache; When there is training parameter data to be transmitted in the first cache and / or training result data to be transmitted in the second cache, it is determined that there is the data to be transmitted.

3. The method according to claim 2, characterized in that When there is data to be transmitted in the first cache and / or the second cache, performing data transmission processing on the data to be transmitted by the multiple groups of sending and receiving threads, packet assembly threads, and / or packet receiving threads includes at least one of the following: Sending the to-be-transmitted training parameter data to the card side using the package assembly thread and the first sending thread through direct memory access; The training result data to be transmitted is received from the card side by using the packet receiving thread and the first receiving thread through the direct memory access.

4. The method according to claim 3, characterized in that The method of sending the training parameter data to be transmitted to the card side by using the package assembly thread and the first sending thread through direct memory access includes: By using the direct memory access and the packet assembly thread, the first training parameter data to be transmitted in the first cache is packaged to obtain a first data packet; The first sending thread is used to send the first data packet to the card side, so that the card side uses the first to-be-transmitted training parameter data in the first data packet to perform model training.

5. The method according to claim 3, characterized in that The receiving the to-be-transmitted training result data from the card side by using the packet receiving thread and the first receiving thread through the direct memory access includes: Performing packet processing on the first training result data to be transmitted in the second cache through the direct memory access and the packet receiving thread to obtain a second data packet; Using the first receiving thread to receive the second data packet from the card side; Notify the upper-layer artificial intelligence (AI) application corresponding to the first training result data to be transmitted, so that the upper-layer AI application uses the second data packet to perform model data processing.

6. The method according to claim 1, characterized in that The method further comprises: Initialize the data plane development kit DPDK framework and register the multiple groups of sending and receiving threads; Determine the board corresponding to the card side according to the supplier identification and the device identification; Abstracting the hardware capabilities of the board corresponding to the card side, and starting the multiple groups of sending and receiving threads; Binding the multiple groups of sending and receiving threads to the CPU core.

7. A data transmission system, characterized in that: Including host side and card side, The host side includes: a first cache library, a polling module, and a transceiver module; The card side includes: a second cache library; The polling module includes multiple groups of transceiver threads, and the polling module is used to perform user-mode polling on the first cache library on the host side and the second cache library on the card side through the multiple groups of transceiver threads to determine whether there is data to be transmitted, and the data to be transmitted is artificial intelligence AI model parameter data; The transceiver module includes a packet assembly thread and a packet receiving thread. The transceiver module is used to perform data transmission processing on the data to be transmitted through the multiple groups of transceiver threads, the packet assembly thread and / or the packet receiving thread when there is data to be transmitted in the first cache library and / or the second cache library.

8. The data transmission system according to claim 7, characterized in that: The polling module is also used to: Performing the user-state polling on the first cache through multiple sending threads to determine whether there is training parameter data to be transmitted in the first cache, and performing the user-state polling on the second cache through multiple receiving threads to determine whether there is training result data to be transmitted in the second cache; When there is training parameter data to be transmitted in the first cache and / or training result data to be transmitted in the second cache, it is determined that there is the data to be transmitted.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

10. A program product, characterized in that The method comprises computer instructions for causing a computer to execute the interactive method according to any one of claims 1 to 6.