A data processing system, method, device, medium and program product

By creating queue pairs in the host's extended memory module and utilizing a cache consistency protocol, the problem of insufficient number of acceleration device queue pairs is solved, dynamic allocation of queue pairs on demand and fast data transmission are achieved, and data transmission efficiency is improved.

CN120492369BActive Publication Date: 2025-10-21LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510984602.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-21
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

The onboard registers of the acceleration device are limited, resulting in an insufficient number of queue pairs provided by the network port, which makes it easy for congestion and data loss to occur during concurrent connections during the communication process.

Method used

By creating multiple queue pairs in the host's extended memory module and synchronizing the target queue pairs to the acceleration device using the cache consistency protocol, queue pairs can be dynamically allocated on demand and remote direct data access can be achieved, thereby improving data transmission efficiency.

Benefits of technology

The number of queue pairs has been expanded, solving the problem of insufficient queue pair resources and improving the data transmission efficiency and the speed of data sending and receiving operations between the host and the acceleration device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492369B_ABST
    Figure CN120492369B_ABST
Patent Text Reader

Abstract

The application discloses a data processing system, method, device, medium and program product in the computer technical field. In the application, the acceleration device provides a remote direct data access module and an extension memory module of a host computer, so that the host computer creates multiple queue pairs in the extension memory module to solve the problem of insufficient queue pairs and expand the number of queue pairs; when the host computer confirms that the queue pairs corresponding to a target port for communication with a third-party device are insufficient, the target queue pairs are selected from the extension memory module and configured to the target port, so that the on-demand dynamic allocation of the queue pairs is realized; after the target queue pairs are updated, the host computer synchronizes the target queue pairs to the acceleration device by using a cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pairs to improve the data transmission efficiency between the host computer and the acceleration device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing system, method, device, medium and program product. Background Art

[0002] Currently, accelerators have limited onboard registers, and based on these registers, they can provide a limited number of queue pairs for network ports. When there are too many concurrent connections during communication, insufficient available queue pairs may lead to communication congestion, which can easily cause data loss.

[0003] Therefore, how to provide more queue pairs for network ports is a problem that those skilled in the art need to solve. Summary of the Invention

[0004] In view of this, an object of the present application is to provide a data processing system, method, device, medium, and program product to provide more queue pairs for a network port.

[0005] In a first aspect, the present application provides a data processing system, comprising: a host and an acceleration device, the host and the acceleration device being connected via a cache consistency protocol; the acceleration device comprising: a remote direct data access module and an extended memory module of the host; the host being configured to: create multiple queue pairs in the extended memory module; when there are insufficient queue pairs corresponding to a target port for communicating with a third-party device, selecting an idle target queue pair in the extended memory module and configuring it at the target port; after updating the target queue pair, synchronizing the target queue pair to the acceleration device using the cache consistency protocol; the acceleration device being configured to: utilize the remote direct data access module and the target queue pair to implement data transmission and reception operations with the third-party device.

[0006] In a second aspect, the present application provides a data processing method, which is applied to a host, and the host is connected to an acceleration device through a cache consistency protocol; the acceleration device includes: a remote direct data access module and an extended memory module of the host; the data processing method includes: creating multiple queue pairs in the extended memory module; when the queue pairs corresponding to the target port for communicating with the third-party device are insufficient, selecting an idle target queue pair in the extended memory module and configuring it at the target port; after updating the target queue pair, synchronizing the target queue pair to the acceleration device using the cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to realize data transmission and reception operations between the third-party device.

[0007] In a third aspect, the present application provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the aforementioned disclosed data processing method.

[0008] In a fourth aspect, the present application provides a non-volatile storage medium for storing a computer program, wherein the computer program implements the aforementioned disclosed data processing method when executed by a processor.

[0009] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which implements the steps of the aforementioned disclosed data processing method when executed by a processor.

[0010] From the above scheme, it can be seen that the present application provides a data processing system, including: a host and an acceleration device, the host and the acceleration device are connected via a cache consistency protocol; the acceleration device includes: a remote direct data access module and an extended memory module of the host; the host is used to: create multiple queue pairs in the extended memory module; when the queue pairs corresponding to the target port for communicating with the third-party device are insufficient, select an idle target queue pair in the extended memory module and configure it at the target port; after updating the target queue pair, synchronize the target queue pair to the acceleration device using the cache consistency protocol; the acceleration device is used to: use the remote direct data access module and the target queue pair to realize data sending and receiving operations between the third-party device.

[0011] It can be seen that the beneficial effects of the present application are as follows: the acceleration device provides a remote direct data access module and an extended memory module of the host, so that the host creates multiple queue pairs in the extended memory module to solve the problem of insufficient number of queue pairs and expand the number of queue pairs; when the host confirms that the queue pairs corresponding to the target port for communication with the third-party device are insufficient, an idle target queue pair is selected in the extended memory module and configured at the target port, thereby realizing on-demand dynamic allocation of queue pairs; after updating the target queue pair, the host synchronizes the target queue pair to the acceleration device using the cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to improve the data transmission efficiency between the host and the acceleration device, and can also quickly realize data transmission and reception operations with third-party devices. It can be seen that this solution uses remote direct data access technology and cache consistency protocol to realize rapid data transmission and queue pair synchronization between the host and the acceleration device, and provides more queue pairs for communication, which can solve the problem of insufficient queue pair resources.

[0012] Correspondingly, the data processing method, device, medium and program product provided by this application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0014] Figure 1 A schematic diagram of a data processing system disclosed in this application;

[0015] Figure 2 A schematic diagram of another data processing system disclosed in this application;

[0016] Figure 3 A flow chart of a data processing method disclosed in this application;

[0017] Figure 4 A schematic diagram of an electronic device disclosed in this application;

[0018] Figure 5 A server structure diagram provided for this application;

[0019] Figure 6 This is a terminal structure diagram provided for this application. DETAILED DESCRIPTION

[0020] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0021] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0023] Currently, accelerators have limited onboard registers, and based on these onboard registers, the number of queue pairs that can be provided for network ports is limited. When there are too many concurrent connections during the communication process, insufficient available queue pairs may cause communication congestion, which can easily lead to data loss. To this end, this application provides a data processing solution that can provide more queue pairs for port communication, realize on-demand dynamic allocation of queue pairs, and improve the data transmission efficiency between the host and the accelerator.

[0024] See also Figure 1 As shown, an embodiment of the present application discloses a data processing system comprising: a host and an acceleration device, the host and the acceleration device being connected via a cache coherence protocol; the acceleration device comprising: a remote direct data access module and an extended memory module of the host. The extended memory module not only increases the host memory but also provides more queue pairs for the port.

[0025] The host is configured to: create multiple queue pairs in an extended memory module; when insufficient queue pairs correspond to a target port for communication with a third-party device, select an idle target queue pair in the extended memory module and configure it at the target port; after updating the target queue pair, synchronize the target queue pair to the acceleration device using a cache coherence protocol, so that the acceleration device is configured to: utilize the remote direct data access module and the target queue pair to perform data transmission and reception operations with the third-party device. The acceleration device has multiple ports for communication with the third-party device, and the target port is at least one of them.

[0026] In this embodiment, the host and the acceleration device belong to the same target end, and data is sent and received between the target end and the third-party device. That is, both sending and receiving operations are performed by the target end. Specifically, when the host sends data to the third-party device, it uses the acceleration device to send data; when the host receives data from the third-party device, it uses the acceleration device to receive data.

[0027] It should be noted that the accelerator device also includes multiple registers, which can also create queue pairs for communication between the accelerator device and the host. After creating multiple queue pairs, the host also configures attribute information for each of the multiple queue pairs, including source address, destination address, port number, etc.

[0028] In one embodiment, the host determines whether there are sufficient queue pairs corresponding to a port, including: evaluating the number of queue pairs required based on a received data transmission / reception request; if the number of queue pairs exceeds the number of queue pairs configured for the target port, determining that the number of queue pairs corresponding to the target port is insufficient. The data transmission / reception request may include key parameters affecting the number of queue pairs, such as the data volume. If the number of queue pairs does not exceed the number of queue pairs configured for the target port, determining that the number of queue pairs corresponding to the target port is sufficient, and current communication can be directly carried out through the target port without increasing the number of queue pairs for the target port.

[0029] In one embodiment, the host also obtains device information of the remote direct data access module, allocates a protection domain for the remote direct data access module, registers its own memory area used for sending and receiving data with the remote direct data access module, and determines corresponding memory handles and other operations to achieve management and configuration of the remote direct data access module.

[0030] In order to realize data transmission between the host and the acceleration device, the host also binds the target queue pair to the memory area corresponding to the received and sent data, and updates the target queue pair accordingly. Specifically, the host determines the memory area and data information corresponding to the received and sent data, and constructs a data receiving and sending request; and adds the data receiving and sending request to the target queue pair. In another example, the host determines the memory area, data length and target address of the data to be sent, and constructs a data sending request; adds the data sending request to the sending queue in the target queue pair, so that the acceleration device is used to: use the remote direct data access module and the sending queue to realize the corresponding data sending operation; after the data sending operation is completed, the sending completion queue in the target queue pair is updated accordingly; after the sending completion queue is updated, the sending completion queue is synchronized to the host using the cache consistency protocol. The host can subsequently learn whether the data sending operation is completed by querying the sending completion queue. The sending completion queue can record information such as the success or failure result of each data sending operation.

[0031] In another example, a host is configured to: determine a memory area for data to be received and construct a data receive request; add the data receive request to a receive queue in a target queue pair, causing the acceleration device to: implement a corresponding data receive operation using a remote direct data access module and the receive queue; upon completion of the data receive operation, update a receive completion queue in the target queue pair; and, after updating the receive completion queue, synchronize the receive completion queue to the host using a cache coherence protocol. The host can subsequently determine whether the data receive operation is complete by querying the receive completion queue, which can record information such as the success or failure of each data receive operation.

[0032] In one embodiment, the host further includes a cache coherence management module connected to both the remote direct data access module and the extended memory module. The cache coherence management module is configured to manage cache coherence protocol communication between the host and the acceleration device. Specifically, data is transmitted between the cache coherence management module and the extended memory module using the cxl.cache protocol.

[0033] In this embodiment, the acceleration device provides a remote direct data access module and an extended memory module of the host, so that the host creates multiple queue pairs in the extended memory module to solve the problem of insufficient number of queue pairs and expand the number of queue pairs; when the host confirms that the queue pairs corresponding to the target port for communication with the third-party device are insufficient, an idle target queue pair is selected in the extended memory module and configured at the target port, thereby realizing on-demand dynamic allocation of queue pairs; after updating the target queue pair, the host synchronizes the target queue pair to the acceleration device using the cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to improve the data transmission efficiency between the host and the acceleration device, and can also quickly realize data transmission and reception operations with the third-party device. It can be seen that this solution uses remote direct data access technology and cache consistency protocol to realize fast data transmission and queue pair synchronization between the host and the acceleration device, and provides more queue pairs for communication, which can solve the problem of insufficient queue pair resources.

[0034] It should be noted that queue pairs include the Send Queue (SQ), Receive Queue (RQ), Send Completion Queue, and Receive Completion Queue. The SQ stores pending data and parameters related to send operations. Each request includes information such as the memory address and length of the data to be sent. The RQ stores pending data and parameters related to receive operations. The Send Completion Queue and Receive Completion Queue store information about sent and received results, such as whether the transmission or reception was successful. The cache coherence protocol is the CXL protocol. By leveraging the PCIe physical layer, the CXL protocol ensures compatibility and interoperability with PCIe, allowing devices to efficiently process shared data and supporting the expansion of memory bandwidth and capacity. The CXL protocol includes: CXL.io: responsible for traditional I / O communication, compatible with existing PCIe devices; CXL.cache: provides cache support for acceleration devices and memory; and CXL.memory: provides low-latency direct communication between the host and memory.

[0035] See Figure 2Another data processing system includes a host computer and an FPGA board (i.e., an accelerator). The FPGA board includes a MEM memory module, an RDMA (Remote Direct Memory Access) module, two network modules, and a CXL management module. Multiple groups of QPs (or queue pairs) are configured within the FPGA board's internal memory module: 1QP0-QPN to NQP0-QPN. The first group of QPs (1QP0-QPN) controls data transmission and reception on Net Port 0, while the second group of QPs (1QP0-QPN) controls data transmission and reception on Net Port 1. A Queue Pair (QP) is the most fundamental logical unit in RDMA communication, managing bidirectional data transmission between nodes.

[0036] exist Figure 2 In the FPGA board, RDMA functionality is implemented, using local memory to expand QP resources based on the cxl.cache protocol. When the QP group resources corresponding to a port are insufficient, the expanded QP resources are dynamically allocated to the corresponding port. Due to the cache coherence of CXL.cache, when the host program configures a QP, cxl.cache notifies the FPGA of the updated QP content, allowing the FPGA to send and receive data.

[0037] like Figure 2 As shown in the figure, host application APP0 uses the PCIE bar to manipulate the QPs of port 0 for data transmission and reception, while host application APP1 uses the PCIE bar to manipulate the QPs of port 1 for data transmission and reception. The host application uses the CXL.cache function to configure the mem memory module as host extended memory with cache coherence. N additional groups of QPs are configured in the mem memory module to expand QPs. These QPs are dynamically allocated based on network port usage. When new applications (APP2-N) require QPs to control network ports, some QPs are allocated from the extended memory mem, allowing more applications to access the network ports for data transmission and reception. When configuring QPs in the extended memory mem, due to cache coherence, once the host application completes the QP configuration, the cxl module, using the cxl.cache protocol, notifies the RDMA module that the QPs have been configured and data transmission and reception is possible.

[0038] Reference Figure 2The process of system initialization of the RDMA environment includes: the host program obtains relevant information about the RDMA device, such as the device handle, protection domain, etc. Usually, the RDMA device can be opened using the ibv_open_device function, and the protection domain is allocated through the ibv_alloc_pd function. Then, rdma_cm_id and qp_init_attr are prepared. Specifically, if multiple QPs are to be created, the rdma_create_qp function needs to be called multiple times, using the same p_ibv_pd but different rdma_cm_id and qp_attr. Each rdma_cm_id identifies an RDMA communication endpoint and can be created by calling the rdma_create_id function. The qp_init_attr structure is used to specify the initialization properties of the QP, including the size of the send and receive queues, the completion queue (CQ), etc. When multiple APP applications (APP2-N) need to run, QPs are dynamically allocated from the extended memory mem area for configuration. An APP may need to allocate multiple QPs. After preparing the rdma_cm_id and qp_init_attr, call the rdma_create_qp function multiple times to create multiple QPs. Each QP is created independently and needs to be configured and managed separately. Each created QP needs to be connected and configured to be in the correct state for data transfer. This typically involves setting QP attributes such as source and destination addresses, port numbers, and establishing a connection using the rdma_connect function.

[0039] In the RDMA technology architecture, P_ibv_pd is a pointer type to the ibv_pd structure (the P_ prefix typically indicates a pointer). Its core function is to manage memory protection and access permissions. ibv_pd (infiniBand verbs protectiondomain) is a key data structure in RDMA programming interfaces (such as verbs), used to isolate memory resources between different applications and ensure the security and legality of data access. rdma_cm_id is a core data structure in Connection Management (CM), used to represent and manage the context of an RDMA connection. It serves as an abstract carrier for key components in RDMA communication, such as devices, queue pairs (QPs), and address resolution, similar to socket descriptors in traditional network programming.

[0040] Before data transmission, the host program must register the memory regions used for sending and receiving data with the RDMA device. This is accomplished by calling the ibv_reg_mr function. The registered memory regions are assigned a memory handle (MR). Furthermore, send and receive buffers must be allocated for each QP. When data needs to be sent, it is sent based on the specific QP. The specific process is as follows: First, a send work request (WR) is constructed, specifying the data buffer, length, and destination address. Then, the ibv_post_send function is called to add the WR to the send work queue (SQ) of the specified QP. The RDMA hardware then performs data transmission based on the WR in the SQ. When receiving data, it is also based on the specific QP. First, some receive work requests (WR) must be sent in advance. The ibv_post_recv function is called to add the WR to the QP's receive work queue (RQ), notifying the hardware of the ready receive buffers. When data arrives, the hardware stores the data in the specified receive buffer based on the WR in the RQ and notifies the host program through the completion queue (CQ).

[0041] When using the QP configuration of the extended area mem for data transmission and reception, the CXL module notifies RDMA to trigger the event.

[0042] After the RDMA hardware completes a send or receive operation, it adds a completion event to the corresponding completion queue (CQ). The host program needs to poll the CQ periodically or asynchronously to check the data transfer completion status. The host program retrieves the completion event from the CQ by calling the ibv_get_cq_event function and determines the success of the data transfer based on the information in the event (including the error code). If the send or receive operation fails, appropriate error handling (partial or full retransmission) is required.

[0043] This implementation enables high-speed interconnection and memory sharing between the host and the acceleration device. High-speed, reliable data transmission between devices is achieved through the CXL protocol based on PCIe. Through the CXL.memory protocol, the host can directly access the memory on the FPGA board, enabling memory resource sharing and expanding the number of RDMAQPs. The memory sharing feature of CXL can be used to store some RDMA connection status data originally stored in the host memory in the FPGA board's memory, reducing the burden on the host memory and PCIe bus. The CXL protocol maintains consistency between the CPU memory space and the memory of the connected devices. When expanding the number of RDMAQPs, multiple devices may simultaneously access and modify QP-related status information. CXL's cache consistency mechanism ensures the accuracy and consistency of this data, avoiding data conflicts and errors, allowing devices such as the FPGA and CPU to work more closely together to handle a large number of RDMA connections.

[0044] The following introduces a data processing method provided in an embodiment of the present application. The data processing method described below can be referenced with other embodiments described in this document.

[0045] An embodiment of the present application discloses a data processing method, which is applied to a host, wherein the host is connected to an acceleration device via a cache consistency protocol; the acceleration device includes: a remote direct data access module and an extended memory module of the host.

[0046] See also Figure 3 As shown, the data processing method disclosed in the embodiment of the present application includes:

[0047] S301: Create multiple queue pairs in an extended memory module.

[0048] S302: When the number of queue pairs corresponding to the target port for communicating with the third-party device is insufficient, select an idle target queue pair from the extended memory module and configure it at the target port.

[0049] S303: After updating the target queue pair, synchronize the target queue pair to the acceleration device using a cache coherence protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to implement data transmission and reception operations with the third-party device.

[0050] In one embodiment, the host is configured to: after creating the multiple queue pairs, configure attribute information for each of the multiple queue pairs.

[0051] In one embodiment, the host is configured to: evaluate the required number of queue pairs according to the received data sending and receiving request; if the number of queue pairs exceeds the number of queue pairs configured for the target port, determine that the queue pairs corresponding to the target port are insufficient.

[0052] In one embodiment, the host is configured to: if the number of queue pairs does not exceed the number of configured queue pairs of the target port, confirm that the queue pairs corresponding to the target port are sufficient.

[0053] In one embodiment, the host is configured to obtain device information of the remote direct data access module and allocate a protection domain to the remote direct data access module.

[0054] In one embodiment, the host is configured to register a memory area of ​​the host for sending and receiving data in the remote direct data access module and determine a corresponding memory handle.

[0055] In one embodiment, the host is configured to bind the target queue pair to a memory area for sending and receiving data, and update the target queue pair accordingly.

[0056] In one embodiment, the host is configured to: determine a memory area and data information corresponding to data transmission and reception, and construct a data transmission and reception request; and add the data transmission and reception request to a target queue pair.

[0057] In one embodiment, the host is used to: determine the memory area, data length and target address of the data to be sent, and construct a data sending request; add the data sending request to the sending queue in the target queue pair; the acceleration device is used to: use the remote direct data access module and the sending queue to implement the corresponding data sending operation; after the data sending operation is completed, update the sending completion queue in the target queue pair accordingly; after updating the sending completion queue, synchronize the sending completion queue to the host using the cache consistency protocol.

[0058] In one embodiment, the host is used to: determine the memory area of ​​the data to be received and construct a data reception request; add the data reception request to the reception queue in the target queue pair; the acceleration device is used to: use the remote direct data access module and the reception queue to implement the corresponding data reception operation; after the data reception operation is completed, update the reception completion queue in the target queue pair accordingly; after updating the reception completion queue, synchronize the reception completion queue to the host using the cache consistency protocol.

[0059] In one embodiment, the host further includes: a cache consistency management module connected to both the remote direct data access module and the extended memory module; the cache consistency management module is used to manage the communication connection of the cache consistency protocol between the host and the acceleration device.

[0060] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0061] It can be seen that this embodiment provides a data processing method that uses remote direct data access technology and cache consistency protocol to achieve fast data transmission and queue pair synchronization between the host and the acceleration device, and provides more queue pairs for network port communication, which can solve the problem of insufficient queue pair resources.

[0062] An electronic device provided in an embodiment of the present application is introduced below. The electronic device described below can be referenced with other embodiments described herein.

[0063] See also Figure 4 As shown, an embodiment of the present application discloses an electronic device, including: a memory 401 for storing a computer program; a processor 402 for executing the computer program to implement the method disclosed in any of the above embodiments.

[0064] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: create multiple queue pairs in the extended memory module; when the queue pairs corresponding to the target port for communicating with the third-party device are insufficient, select an idle target queue pair in the extended memory module and configure it at the target port; after updating the target queue pair, synchronize the target queue pair to the acceleration device using the cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to implement data transmission and reception operations between the third-party device.

[0065] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: after creating multiple queue pairs, configuring attribute information for each of the multiple queue pairs.

[0066] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: evaluate the required number of queue pairs based on the received data sending and receiving request; if the number of queue pairs exceeds the number of configured queue pairs of the target port, confirm that the queue pairs corresponding to the target port are insufficient.

[0067] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: if the number of queue pairs does not exceed the number of configured queue pairs of the target port, confirming that the queue pairs corresponding to the target port are sufficient.

[0068] In this embodiment, when the processor executes the computer program stored in the memory, the processor may specifically implement the following steps: obtaining device information of the remote direct data access module, and allocating a protection domain for the remote direct data access module.

[0069] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: registering its own memory area for sending and receiving data in the remote direct data access module, and determining a corresponding memory handle.

[0070] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: binding the target queue pair with the memory area for sending and receiving data, and updating the target queue pair accordingly.

[0071] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: determine the memory area and data information corresponding to the sent and received data, and construct a data sending and receiving request; add the data sending and receiving request to the target queue pair.

[0072] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: determine the memory area, data length and target address of the data to be sent, and construct a data sending request; add the data sending request to the sending queue in the target queue pair.

[0073] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: use the remote direct data access module and the sending queue to implement the corresponding data sending operation; after the data sending operation is completed, update the sending completion queue in the target queue pair accordingly; after updating the sending completion queue, use the cache consistency protocol to synchronize the sending completion queue to the host.

[0074] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: determine the memory area of ​​the data to be received and construct a data reception request; add the data reception request to the reception queue in the target queue pair.

[0075] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: use the remote direct data access module and the receiving queue to implement the corresponding data receiving operation; after the data receiving operation is completed, update the receiving completion queue in the target queue pair accordingly; after updating the receiving completion queue, use the cache consistency protocol to synchronize the receiving completion queue to the host.

[0076] Furthermore, the embodiment of the present application also provides an electronic device. The electronic device can be Figure 5 The server shown can also be Figure 6 The terminal shown. Figure 5 and Figure 6Each of the diagrams is a structural diagram of an electronic device according to an exemplary embodiment, and the contents in the diagrams cannot be considered as any limitation on the scope of use of the present application.

[0077] Figure 5 This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, which is loaded and executed by the processor to implement the relevant steps of the data processing disclosed in any of the aforementioned embodiments.

[0078] In this embodiment, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world. The specific interface type can be selected according to specific application needs and is not specifically limited here.

[0079] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0080] The operating system is used to manage and control the hardware devices and computer programs on the server, enabling the processor to operate and process data in the memory. It can be Windows Server, NetWare, Unix, Linux, etc. In addition to computer programs capable of performing the data processing methods disclosed in any of the aforementioned embodiments, computer programs can also include computer programs capable of performing other specific tasks. Data can include data such as application update information and other data such as application developer information.

[0081] Figure 6 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may specifically include but is not limited to a smartphone, tablet computer, laptop computer or desktop computer.

[0082] Generally, the terminal in this embodiment includes: a processor and a memory.

[0083] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content required to be displayed on the display. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0084] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is used to store at least the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the data processing method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to update information of the application.

[0085] In some embodiments, the terminal may further include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0086] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation to the terminal, and may include more or fewer components than shown in the figure.

[0087] A non-volatile storage medium provided in an embodiment of the present application is introduced below. The non-volatile storage medium described below can be referenced with other embodiments described herein.

[0088] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed in the aforementioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium that, as a carrier for resource storage, may be a read-only memory, random access memory, a magnetic disk, or an optical disk. The resources stored thereon include an operating system, a computer program, and data, and the storage method may be either temporary or permanent.

[0089] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, the following steps can be specifically implemented: creating multiple queue pairs in the extended memory module; when the queue pairs corresponding to the target port for communicating with the third-party device are insufficient, selecting an idle target queue pair in the extended memory module and configuring it at the target port; after updating the target queue pair, synchronizing the target queue pair to the acceleration device using the cache consistency protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to implement data transmission and reception operations between the third-party device.

[0090] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, the following steps may be specifically implemented: after creating multiple queue pairs, configuring attribute information for each of the multiple queue pairs.

[0091] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: evaluate the required number of queue pairs based on the received data sending and receiving request; if the number of queue pairs exceeds the number of configured queue pairs of the target port, confirm that the queue pairs corresponding to the target port are insufficient.

[0092] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, the following steps may be specifically implemented: if the number of queue pairs does not exceed the number of configured queue pairs of the target port, confirming that the queue pairs corresponding to the target port are sufficient.

[0093] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it may specifically implement the following steps: obtaining device information of the remote direct data access module, and allocating a protection domain for the remote direct data access module.

[0094] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: registering its own memory area for sending and receiving data in the remote direct data access module, and determining the corresponding memory handle.

[0095] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, the following steps may be specifically implemented: binding the target queue pair to the memory area for sending and receiving data, and updating the target queue pair accordingly.

[0096] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: determine the memory area and data information corresponding to the sent and received data, and construct a data sending and receiving request; add the data sending and receiving request to the target queue pair.

[0097] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: determine the memory area, data length and target address of the data to be sent, and construct a data sending request; add the data sending request to the sending queue in the target queue pair.

[0098] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: use the remote direct data access module and the sending queue to implement the corresponding data sending operation; after the data sending operation is completed, update the sending completion queue in the target queue pair accordingly; after updating the sending completion queue, use the cache consistency protocol to synchronize the sending completion queue to the host.

[0099] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: determine the memory area of ​​the data to be received and construct a data reception request; add the data reception request to the reception queue in the target queue pair.

[0100] In this embodiment, when the processor executes the computer program stored in the non-volatile storage medium, it can specifically implement the following steps: use the remote direct data access module and the receiving queue to implement the corresponding data receiving operation; after the data receiving operation is completed, update the receiving completion queue in the target queue pair accordingly; after updating the receiving completion queue, use the cache consistency protocol to synchronize the receiving completion queue to the host.

[0101] A computer program product provided in an embodiment of the present application is introduced below. The computer program product described below can be referenced with other embodiments described herein.

[0102] A computer program product comprises a computer program / instruction, which implements the steps of the aforementioned data processing method when executed by a processor.

[0103] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments are implemented.

[0104] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0105] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0106] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A data processing system, characterized in that: include: A host and an acceleration device, wherein the host and the acceleration device are connected via a cache coherence protocol; The acceleration device includes: a remote direct data access module and an extended memory module of the host; The host is configured to: create a plurality of queue pairs in the extended memory module; when a queue pair corresponding to a target port for communicating with a third-party device is insufficient, select an idle target queue pair in the extended memory module and configure it on the target port; the target port being at least one of the plurality of ports in the acceleration device; and synchronize the queue pairs between the host and the acceleration device based on a cache coherence protocol; The acceleration device is used to implement data transmission and reception operations with a third-party device by utilizing the remote direct data access module and the target queue pair.

2. The system according to claim 1, wherein: The host is configured to: after creating the multiple queue pairs, respectively configure attribute information for each of the multiple queue pairs.

3. The system according to claim 1, wherein: The host is configured to: evaluate the required number of queue pairs according to the received data sending and receiving request; if the number of queue pairs exceeds the number of configured queue pairs of the target port, confirm that the queue pairs corresponding to the target port are insufficient.

4. The system according to claim 3, characterized in that The host is configured to: if the number of queue pairs does not exceed the number of configured queue pairs of the target port, confirm that the queue pairs corresponding to the target port are sufficient.

5. The system according to claim 1, wherein: The host is used to obtain device information of the remote direct data access module and allocate a protection domain for the remote direct data access module.

6. The system according to claim 1, wherein: The host is used to register its own memory area for sending and receiving data in the remote direct data access module and determine the corresponding memory handle.

7. The system according to claim 6, characterized in that The host is used to bind the target queue pair to the memory area for sending and receiving data.

8. The system according to claim 6, wherein: The host is used to: determine the memory area and data information corresponding to the data to be sent and received, and construct a data sending and receiving request; and add the data sending and receiving request to the target queue pair.

9. The system according to claim 8, characterized in that The host is used to: determine the memory area, data length and target address of the data to be sent, and construct a data sending request; add the data sending request to the sending queue of the target queue pair; The acceleration device is used to: use the remote direct data access module and the sending queue to implement a corresponding data sending operation; after the data sending operation is completed, update the sending completion queue in the target queue pair accordingly; after updating the sending completion queue, use the cache consistency protocol to synchronize the sending completion queue to the host.

10. The system according to claim 8, wherein: The host is used to: determine the memory area of ​​the data to be received and construct a data receiving request; add the data receiving request to the receiving queue in the target queue pair; The acceleration device is used to: use the remote direct data access module and the receiving queue to implement a corresponding data receiving operation; after the data receiving operation is completed, update the receiving completion queue in the target queue pair accordingly; after updating the receiving completion queue, use the cache consistency protocol to synchronize the receiving completion queue to the host.

11. The system according to any one of claims 1 to 10, characterized in that The acceleration device further includes: a cache consistency management module connected to both the remote direct data access module and the extended memory module; The cache consistency management module is used to manage the communication connection of the cache consistency protocol between the host and the acceleration device.

12. A data processing method, characterized in that: Applied to a host, the host is connected to an acceleration device via a cache consistency protocol; the acceleration device comprises: a remote direct data access module and an extended memory module of the host; The data processing method includes: creating a plurality of queue pairs in the extended memory module; When there are insufficient queue pairs corresponding to a target port for communicating with a third-party device, selecting an idle target queue pair in the extended memory module and configuring it at the target port; the target port is at least one of the multiple ports in the acceleration device; The queue pair synchronization between the host and the acceleration device is achieved based on a cache coherence protocol, so that the acceleration device uses the remote direct data access module and the target queue pair to implement data transmission and reception operations with a third-party device.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to claim 12.

14. A non-volatile storage medium, characterized in that: Used for storing a computer program, wherein the computer program implements the method according to claim 12 when executed by a processor.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method of claim 12 is implemented.

Citation Information

Patent Citations

  • Method and system for realizing RDMA network card request queue

    CN119271618A

  • Dynamic network adapter queue pair allocation

    US20110252419A1