Data Processing Method and Related Products for Remote Direct Memory Access
By allocating host memory resources for each transceiver work queue in the local terminal device, the data processing exceptions caused by limited chip memory resources are solved, and more efficient and stable data storage and processing are achieved.
Patent Information
- Application Number
- CN202310067975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-01-30
AI Technical Summary
When the remote computer sends the requested memory access data back to the local computer, the limited memory resources of the chip lead to data processing exceptions, especially when the amount of memory access data is too large, the memory access data is stored in the chip memory in the prior art, which is prone to abnormalities.
By allocating free host memory resources in the local terminal device, a data cache work queue is exclusively for each sending and receiving work queue based on the queue depth of the sending and receiving work queue, storing the access data obtained by remote direct memory access operations.
It improves the stability and efficiency of terminal devices to store and process data access, avoids exceptions caused by insufficient chip memory resources, simplifies memory management and improves performance.
Smart Images

Figure CN115964319B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data transmission, and particularly relates to a data processing method for remote direct memory access and related products. Background Art
[0002] RDMA (Remote Direct Memory Access) has characteristics such as high bandwidth and low latency. Using RDMA communication technology can improve system throughput, reduce the network communication latency of the system, and save precious CPU (Central Processing Unit) resources in a computer. Therefore, it has been widely used in the field of data center storage and computing networks currently.
[0003] However, in the prior art, after the remote computer transmits the memory access data of the request back to the local computer, the local computer will store the memory access data in the chip memory. Therefore, when the quantity of memory access data is excessive and the chip memory resources are limited, it will cause abnormalities in the data processing process of the local computer. Summary of the Invention
[0004] This application provides a data processing method for remote direct memory access and related products, aiming to improve the efficiency and stability of a terminal device in processing access data obtained from a remote direct memory access operation.
[0005] In a first aspect, an embodiment of this application provides a data processing method for remote direct memory access, which is applied to a first terminal device in a remote memory access system. The remote memory access system includes the first terminal device and a second terminal device. The method includes:
[0006] Upon detecting an execution instruction for a remote direct memory access operation for the second terminal device, execute the remote direct memory access operation, and obtain a plurality of memory access work queues corresponding to the current remote direct memory access operation. The memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues include at least one transceiver work queue. A single transceiver work queue includes a single memory receiving queue and a single memory sending queue. The single memory receiving queue or memory sending queue is used to store work queue elements, and a single work queue element is used to represent the request information of a single work request;
[0007] Perform a memory resource allocation operation according to the queue depth of the at least one transceiver work queue, obtain at least one memory resource information, and store each memory resource information in the work queue information of the corresponding transceiver work queue. The at least one memory resource information corresponds one-to-one with the at least one transceiver work queue. The queue depth is used to represent the number of work queue elements stored in the corresponding transceiver work queue. The memory resource allocation operation refers to the operation of allocating the idle host memory in the first terminal device to store the access data obtained by performing the remote direct memory access operation. A single memory resource information is used to indicate the corresponding single data cache work queue. The data cache work queue is used to cache the access data associated with the corresponding transceiver work queue. The data cache work queue is a work queue arranged with multiple cache work queue elements. The work queue information is used to store the information associated with the corresponding transceiver work queue;
[0008] Receive at least one group of access data sent by the second terminal device in response to the remote direct memory access operation. The at least one group of access data corresponds one-to-one with the at least one transceiver work queue;
[0009] For each group of access data in the at least one group of access data, perform the following operations:
[0010] Determine the target memory resource information associated with the currently processed access data;
[0011] Determine at least one target cache work queue element in the target data cache work queue according to the target memory resource information;
[0012] Store the currently processed access data into the at least one target cache work queue element.
[0013] In a second aspect, an embodiment of the present application provides a data processing device for remote direct memory access, which is applied to a first terminal device in a remote memory access system. The remote memory access system includes the first terminal device and a second terminal device. The device includes:
[0014] A first execution unit, configured to detect an execution instruction for a remote direct memory access operation for the second terminal device, execute the remote direct memory access operation, and obtain a plurality of memory access work queues corresponding to the current remote direct memory access operation. The memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues includes at least one transceiver work queue. A single transceiver work queue includes a single memory receive queue and a single memory send queue. The single memory receive queue or memory send queue is used to store work queue elements, and a single work queue element is used to represent the request information of a single work request.
[0015] A resource allocation unit, configured to perform a memory resource allocation operation according to the queue depth of the at least one transceiver work queue, obtain at least one memory resource information, and store each memory resource information in the work queue information of the corresponding transceiver work queue.
[0016] A receiving unit, configured to receive at least one set of access data sent by the second terminal device in response to the remote direct memory access operation. The at least one set of access data corresponds to the at least one transceiver work queue one by one.
[0017] A second execution unit, configured to perform the following operations for each set of access data in the at least one set of access data: determine the target memory resource information associated with the currently processed access data; determine at least one target cache work queue element in the target data cache work queue according to the target memory resource information; store the currently processed access data in the at least one target cache work queue element.
[0018] In a third aspect, an embodiment of the present application provides a terminal device, including a processor, a memory, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the processor. The programs include instructions for performing the steps in the first aspect of the embodiments of the present application.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps in the first aspect of the embodiments of the present application are implemented.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, some or all of the steps described in the first aspect of the embodiments of the present application are implemented.
[0021] It can be seen that by obtaining at least one transceiver work queue corresponding to the current remote direct memory access operation executed by the first terminal device, and then performing a memory resource allocation operation according to the queue depths of the at least one transceiver work queue to obtain at least one memory resource information, the free memory resources of the host are allocated to it; finally, after receiving at least one set of access data sent by the second terminal device in response to the remote direct memory access operation, the at least one set of access data is stored in at least one target cache work queue element determined by the associated target memory resource information. In this way, compared with the prior art of storing memory access data in the internal resources of the chip, the present application proposes to allocate the free host memory resources to the transceiver work queues in advance, so that the access data corresponding to each transceiver work queue has a dedicated data cache work queue for data storage, improving the stability and efficiency of the terminal device for storing and processing access data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 is a structural block diagram of a remote memory access system provided by an embodiment of the present application;
[0024] Figure 2 is a schematic flowchart of a data processing method for remote direct memory access provided by an embodiment of the present application;
[0025] Figure 3 is a schematic structural diagram of a memory access work queue provided by an embodiment of the present application;
[0026] Figure 4 is a schematic structural diagram of a transceiver work queue provided by an embodiment of the present application;
[0027] Figure 5 is an interaction schematic diagram of a first terminal device and a second terminal device provided by an embodiment of the present application;
[0028] Figure 6 is a schematic diagram of a first resource information addressing provided by an embodiment of the present application;
[0029] Figure 7 is a schematic diagram of a second resource information addressing provided by an embodiment of the present application;
[0030] Figure 8It is a block diagram of the functional units of a data processing device for remote direct memory access provided by an embodiment of the present application;
[0031] Figure 9 It is a block diagram of the functional units of another data processing device for remote direct memory access provided by an embodiment of the present application;
[0032] Figure 10 It is a block diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0035] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0036] Please refer to Figure 1 , Figure 1 It is a block diagram of a remote memory access system provided by an embodiment of the present application. As Figure 1As shown, the remote memory access system 100 includes a first terminal device 110 and a second terminal device 120, and the first terminal device 110 and the second terminal device 120 are communicatively connected. The first terminal device obtains the memory access data of the second terminal device by performing a remote direct memory access operation, and when the first terminal device performs the remote direct memory access operation, the first terminal device performs a memory resource allocation operation according to at least one transceiver work queue corresponding to the remote direct memory access operation. After the first terminal device receives at least one set of access data sent by the second terminal device in response to the remote direct memory access operation, the at least one set of access data is stored in at least one target cache work queue that has been allocated. Among them, the first terminal device 110 and the second terminal device 120 can be computer devices such as tablet computers and notebook computers. One first terminal device 110 can simultaneously correspond to multiple second terminal devices 120, or the remote memory access system 100 includes multiple first terminal devices 110, and each first terminal device 110 corresponds to one or more second terminal devices 120.
[0037] Based on this, an embodiment of the present application provides a data processing method for remote direct memory access. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0038] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a data processing method for remote direct memory access provided by an embodiment of the present application. The method is applied to the first terminal device 110 in the remote memory access system 100 as Figure 1 described above. The remote memory access system 100 includes a first terminal device 110 and a second terminal device 120; the method includes:
[0039] Step 201, detecting an execution instruction for a remote direct memory access operation for the second terminal device, performing the remote direct memory access operation, and obtaining a plurality of memory access work queues corresponding to the current remote direct memory access operation.
[0040] Among them, the memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues include at least one transceiver work queue. A single transceiver work queue includes a single memory receive queue and a single memory send queue. The single memory receive queue or the memory send queue is used to store work queue elements, and a single work queue element is used to represent the request information of a single work request.
[0041] Among them, the remote direct memory access operation refers to the RDMA operation, that is, the direct access from the memory of one computer to the memory of another computer. During this process, the operating system of either party is not involved. At the same time, there is no need to copy data between the application memory of the terminal device and the data buffer of the operating system. It should be noted that when the local computer, that is, the first terminal, executes the RDMA operation, one of the most important concepts is the memory access work queue (Work Queue, WQ). Please refer to Figure 3 , Figure 3 is a schematic structural diagram of a memory access work queue provided by an embodiment of the present application. As shown in the figure, the memory access work queue is a queue for storing work requests. This work request is the "work" sent by the software in the first terminal device to the hardware, and the hardware will complete the task according to the work request sent by the software. The work queue elements sent by the software to the hardware are stored in the memory access work queue. The work queue elements can be understood as a kind of "task description". This description contains the tasks that the software hopes the hardware to do and the detailed information about this task. For example, a certain task is like this: "I want to send the data with a length of 10 bytes located at address 0x12345678 to the opposite node". After receiving the task, the hardware will fetch the data from the memory, assemble the data packet, and then send it. A memory access work queue can contain many work queue elements, or it can have no work queue elements. Based on the concept of this memory access work queue, the concepts of the send-receive work queue (Queue Pair, QP), the memory send queue (SendQueue, SQ), and the memory receive queue (Receive Queue, RQ) will be introduced next. Please refer to Figure 4 , Figure 4 is a schematic structural diagram of a send-receive work queue provided by an embodiment of the present application. As Figure 4 shown, there are sending and receiving ends in any communication process, among which Figure 4On the left is the transceiver work queue of the sending end, and on the right is the transceiver work queue of the receiving end. Each transceiver work queue includes a memory receive queue and a memory send queue. The memory send queue in the transceiver work queue on the left contains two work queue elements, which means the software of the sending end has sent two send work requests to the hardware. The content of one send work request can be "Please send the 10-byte data with the address 0x111111111 in the memory to node B". In this work request, the hardware of the sending end will retrieve this 10-byte data from the memory and send it for communication to node B. At this time, we default this receiving end to be node B. After this receiving end receives this data, its software sends a work request in the memory receive queue, and its content can be "Please prepare to receive data and store the received data in the memory area with the address 0x222222222". Then, the hardware of the receiving end will store the 10-byte data sent by the sending end in a specific memory area. Simply put, the memory receive queue is specifically used to store send tasks, and the memory send queue is specifically used to store receive tasks. In one send-receive process, the sending end needs to put the work queue element representing one send task into the memory receive queue. Similarly, the receiving-end software needs to send a work queue element representing a receive task to the hardware so that the hardware knows where to store the received data in the memory.
[0042] It should be noted that in one remote direct memory access operation, there can be multiple transceiver work queues. They can be processed by the hardware simultaneously or not simultaneously. However, in the same memory access work queue, the work queue elements are consumed by the hardware in sequence. This is also the concept of a queue, that is, a first-in-first-out data structure. The work queue element first sent by the software will also be the first to be processed by the hardware.
[0043] Step 202: Perform a memory resource allocation operation according to the queue depth of at least one transceiver work queue to obtain at least one memory resource information, and store each memory resource information in the work queue information (QPC) of the corresponding transceiver work queue.
[0044] Among them, the at least one memory resource information corresponds to the at least one transceiver work queue one by one. The queue depth is used to represent the number of work queue elements stored in the corresponding transceiver work queue. The memory resource allocation operation refers to the operation of allocating the idle host memory in the first terminal device to store the access data obtained by performing the remote direct memory access operation. A single memory resource information is used to indicate the corresponding single data cache work queue. The data cache work queue is used to cache the access data associated with the corresponding transceiver work queue. The data cache work queue is a work queue arranged with multiple cache work queue elements. The work queue information is used to store the information associated with the corresponding transceiver work queue.
[0045] Among them, the data cache work queue (Read / Atomic Response&Ack Queue, RAQ) is a work queue maintained by the chip itself inside the chip in the first terminal device. It is used to cache the access data sent by the second terminal device in the remote memory access system. This access data is obtained by the first terminal device performing a remote direct memory access operation on the second terminal device. The chip can be a common network card for communication or other chip devices, which is not limited in this case. After the chip caches the data, the subsequent module will read the data in the data cache work queue according to their needs for consumption. It should be noted that in the prior art, the chip implements the processing of caching data in the data cache work queue in the form of a free list. Its feature is that all the send and receive work queues in a single remote direct memory access operation share a data cache work queue. Briefly speaking, the free list organizes all the memory blocks in the internal memory of the chip in the form of a linked list. When allocating memory, scan each free memory block in the free list to find a memory block with a size that meets the requirements, and remove the memory block from the free list. When releasing memory, the released memory block is reinserted into the free list. However, there are certain deficiencies in using the free list method for the internal free memory of the chip. For example, the form of the free list occupies internal resources of the chip, and when the internal resources of the chip are limited, the effect is not good. When the number and queue depth of the send and receive work queues reach a certain threshold, it is easy to cause the data cache work queue to overflow, resulting in abnormal sending / receiving of the send and receive work queues. Moreover, the implementation and management of the free list form are complex, affecting the sending and receiving performance. When multiple send and receive work queues operate on the free list simultaneously, the performance deteriorates. Therefore, in the data processing method for remote direct memory access provided in this application, the memory resource allocation operation performed according to the queue depth of the send and receive work queue allocates the free host memory in the first terminal device, so that each send and receive work queue can exclusively own a corresponding data cache work queue. And by storing the memory resource information allocated to each send and receive work queue in the work queue information of the corresponding send and receive work queue, during the subsequent storage process, the access data corresponding to each send and receive work queue can find the corresponding allocated free memory for storage. And because the resources are allocated in the host memory, there will be no memory bottleneck in the storage space, saving the internal resources of the chip and not affecting the operation of the chip itself. Compared with the prior art, it is simpler, more stable and reliable, and has higher efficiency.
[0046] In a possible example, the first terminal device includes a driver, a host memory, a chip, and a computer system bus, where the driver, the host memory, and the chip are respectively connected to the computer system bus. Based on the data processing method of remote direct memory access provided in this application, the driver software in the driver is used to apply for the memory of the transceiver work queue, that is, perform a memory resource allocation operation according to the queue depth of the transceiver work queue, and the driver software will select any one of continuous memory allocation and discontinuous memory allocation for memory allocation during allocation, and generate corresponding memory resource information and store it in the work queue information. After the allocation, after the receiving end of the remote memory access engine in the chip receives the data sent back by the second terminal device, the chip caches these access data into the host memory of the first terminal device according to the corresponding memory resource information, that is, the chip plays a role in managing the data cache work queue and performing data storage operations.
[0047] Step 203: Receive at least one set of access data sent by the second terminal device in response to the remote direct memory access operation.
[0048] Wherein, the at least one set of access data corresponds to the at least one transceiver work queue one by one.
[0049] Wherein, as mentioned above, a transceiver work queue contains one or more work requests, and the hardware in the first terminal device performs an operation to access the memory of the second terminal device according to the work requests. The data that the first terminal device wants to access and obtains from the second terminal device is the access data, which is the data sent by the second terminal device in response to the remote direct memory access operation of the first terminal device.
[0050] Wherein, please refer to Figure 5 , Figure 5 is an interaction schematic diagram of a first terminal device and a second terminal device provided in an embodiment of this application. As shown in Figure 5As shown in the figure, on the left is the structure of the second terminal device, and on the right is the structure of the first terminal device. Each terminal device includes a central processing unit (i.e., CPU), a host memory, and a chip (such as a network card). Among them, the central processing unit, the host memory, and the chip are all connected to the bus of the computer system of the terminal device. The first terminal device and the second terminal device communicate with each other through the network card. Based on the remote direct memory access operation provided in this application, please note the dotted arrow in the figure. After the first terminal device on the right executes the remote direct memory access operation, a segment of data in the host memory of the second terminal device is copied into the host memory of the first terminal device, and the CPUs at both ends hardly participate in the data transmission process (only participate in the control plane). The network card of the second terminal device directly copies the data into the internal storage space, and then the hardware assembles the packets at each layer and sends them to the network card of the peer through the physical link. After the network card of the peer receives the data, it strips the packet headers and check codes at each layer and directly stores the accessed data into the pre-allocated host memory, that is, the pre-allocated data cache work queue.
[0051] In a possible example, the accessed data includes at least one set of data, and the at least one set of data corresponds one-to-one with the at least one target cache work queue element; the type of the at least one set of data includes at least one of the following: memory data, operation success response instruction, operation failure response instruction; and, the type of the work request includes at least one of the following: memory read operation, memory rewrite operation, atomic operation.
[0052] Among them, since a transceiver work queue corresponding to a set of accessed data may contain multiple work queue elements, that is, multiple work requests, then the accessed data may include more than one set of data. The type of data is divided into memory data, operation success response instruction, and operation failure response instruction according to the type of work situation; the type of work request then includes memory read operation, memory rewrite operation, and atomic operation.
[0053] Among them, the work request corresponding to the memory data is a memory read operation (Read Response), that is, the hardware of the first terminal device wants to obtain the memory data at a specific address of the second terminal device, so this memory read operation will be executed to read the memory of the second terminal device to obtain the memory data. The work requests corresponding to the operation success response instruction and the operation failure response instruction are the memory rewrite operation (Write Response) and the atomic operation (Atomic Response), that is, when the first terminal device executes these operations on the memory of the second terminal device, it does not need to transmit the memory data, and only needs the second terminal device to give feedback on whether the operation is successful.
[0054] It can be seen that in this example, by determining the data type in the access data and the type of the work request, the accuracy and stability of the data processing of the remote direct memory access can be improved.
[0055] Step 204, for each group of access data in at least one group of access data, perform the following operations.
[0056] Among them, the following operations include: determining the target memory resource information associated with the currently processed access data; determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information; storing the currently processed access data into the at least one target cache work queue element.
[0057] Among them, since the previous memory resource allocation operation allocates free host memory space for each transceiver work queue to cache the corresponding access data, when the access data arrives, the target memory resource information of the associated transceiver work queue is determined for subsequent caching operations. This can ensure that each host memory space matches the size of the corresponding access data, and there will be no problem of not being able to store it. And since the cache work queue element for finally storing the data, after determining the target data cache work queue according to the target memory resource information, it is also necessary to determine the cache work queue element for caching the data in it, because there may be a situation where some work queue elements in the target data cache work queue are lost or have already occupied memory.
[0058] In a possible example, when the memory resource information is the first resource information; the determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information includes: determining the target cache work queue according to the base address information; determining the head information and the tail information of the target cache work queue according to the target cache work queue; determining the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
[0059] Among them, the first resource information includes the base address information of the data cache work queue and at least one queue element position, where the queue element position is used to represent the arrangement position of the corresponding target cache work queue element in the target data cache work queue; the head information is used to represent the number of cache work queue elements whose memory in the target data cache work queue has been occupied, and the tail information is used to represent the total number of cache work queue elements in the target data cache work queue.
[0060] Among them, since there are two allocation situations in the memory resource allocation operation, one is continuous allocation and the other is non - continuous allocation. When allocating resources to each transceiver work queue, one of these situations appears randomly. In order to accurately find the address of the corresponding cache work queue element in the host memory, therefore, for different allocation situations, this solution has different delivery of resource information and matching of addressing modes. Please refer to Figure 6 , Figure 6 which is a schematic diagram of a first resource information addressing provided by an embodiment of the present application. As Figure 6 shown, Figure 6 what is shown is that in the first allocation situation, that is, in the case of continuous memory allocation, the first resource allocation information includes the base address information of the data cache work queue and at least one queue element position. In this case, the addressing mode is to first find the address of the corresponding data cache work queue in the host memory of the first terminal device through the base address information. After determining the target data cache work queue, obtain the head information and tail information of this work queue, and then determine the corresponding target cache work queue element according to the queue element position. Figure 6 The example given in
[0061] is that a target data cache work queue with a cacheable data memory size of 64 bits is determined through the base address information, and then the cache work queue element 2 with a memory size of 15 bits in the target data cache work queue is determined as the target cache work queue element according to the queue element position, so as to facilitate subsequent storage operations.
[0062] It can be seen that in this example, when the memory resource information is the second resource information; determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information includes: determining a target memory address according to the memory address information; determining the target data cache work queue in the target memory address according to the target memory address and the target work queue index; determining the head information and the tail information of the target cache work queue according to the target data cache work queue; and determining the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
[0063] Wherein, the second resource information includes memory address information, a target work queue index, and at least one queue element position. The memory address information is used to indicate a memory address storing address information of one or more data cache work queues.
[0064] Wherein, in this example, please refer to Figure 7 , Figure 7 which is a schematic diagram of addressing the second resource information provided by an embodiment of this application. As Figure 7 shown, Figure 7 what is shown is that in the second allocation case, that is, in the case of non - contiguous allocation, the second resource allocation information includes memory address information, a target work queue index, and at least one queue element position. The addressing mode corresponding to the second resource allocation information is a two - level lookup mode. That is, first, according to the memory address information in the second resource information, find the target memory address storing the address information of data cache work queue 1 and data cache work queue 2. Then, according to the target work queue index, determine that data cache work queue 1 is the target data cache work queue we are looking for. Then, locate data cache work queue 1 according to the address information of data cache work queue 1 in the target memory address. Finally, according to the queue element position, determine that cache work queue element 2 in data cache work queue 1 is the target cache work queue element.
[0065] Wherein, the meaning of the head information is that in a data cache work queue, since the number of its work queue elements may be ten, eight, or just one, the sequence numbers of the work queue elements starting from this work queue may be uncertain. There may be missing ones or the work queue elements have been occupied. For example, normally, there are 10 cache work queue elements in a data cache work queue, and the numbers will be arranged from cache work queue element 0 to cache work queue element 9. However, due to the missing of two work queue elements and the occupation of two work queue elements in this data cache work queue, the actual numbers of the work queue elements used for storage are cache work queue element 4. Then, the head information is used to represent this information. And the tail information is used to represent the number of the last cache work queue element in this data cache work queue. Combining with the head information, it is possible to know how many work queue elements are used for storage and which work queue elements are free. Then, combining with the queue element position, the final target cache work queue element can be determined.
[0066] It can be seen that in this example, when the memory resource information is the second resource information, the addressing mode for determining the target cache work queue element in the target data cache work queue becomes a two - level lookup mode, which can accurately find the memory address in the case of non - contiguous allocation of memory resources, improving the flexibility, stability, and efficiency of data processing for remote direct memory access.
[0067] In a possible example, determining the head information and tail information of the target cache work queue according to the target data cache work queue includes: determining a plurality of reference work queue elements according to the target data cache work queue; sequentially numbering the plurality of reference work queue elements from small to large starting from the dequeue end of the target data cache work queue to obtain a plurality of queue element numbers; traversing the memory occupancy of the plurality of reference work queue elements, determining that the queue element number with the smallest value corresponding to the reference work queue element with the memory occupancy being unoccupied is the head number, and determining that the queue element number with the largest value corresponding to the reference work queue element with the memory occupancy being unoccupied is the tail number; generating head information according to the head number and generating tail information according to the tail number.
[0068] Wherein, the plurality of queue element numbers correspond to the plurality of reference work queue elements one by one.
[0069] Wherein, in this example, by traversing the memory occupancy and numbering, the positions and quantities of the work queue elements in the target data cache work queue that do not occupy memory are determined, so as to generate the head information and tail information, which is convenient for subsequently determining the target cache work queue elements that are free for storage in combination with the queue element positions.
[0070] It can be seen that in this example, a plurality of reference work queue elements are determined according to the target data cache work queue, and then the head information and tail information are generated by numbering and traversing the memory occupancy, ensuring that the free work queue elements in the target data cache work queue that can be used for storage are determined, and improving the stability and efficiency of data storage.
[0071] In a possible example, the access data carries a serial number; determining the target memory resource information associated with the currently processed access data includes: determining a target transceiver work queue according to the target serial number carried by the currently processed access data; determining the corresponding target work queue information according to the target transceiver work queue; and obtaining the target memory resource information according to the target work queue information.
[0072] Wherein, the serial number is used to indicate the transceiver work queue associated with the access data.
[0073] Among them, the sequence number is synchronously generated when generating the transceiver work queue during a remote direct memory access operation, and is associated with the generation time of the transceiver work queue. The smaller the sequence number corresponding to the transceiver work queue generated earlier, and the larger the sequence number corresponding to the transceiver work queue generated later. After the sequence number is generated, it will flow along with the entire operation in the form of data. Therefore, when obtaining access data from the second terminal device, the access data transmitted back by the second terminal device to the first terminal device carries this sequence number, which can help the hardware of the first terminal device identify which transceiver work queue is associated, so as to determine its corresponding work request and target memory resource information.
[0074] It can be seen that in this example, by setting the access data to carry the sequence number, it is convenient for the access data to accurately determine the target transceiver work queue associated with it, so as to determine the target memory resource information, improving the accuracy and efficiency of data processing.
[0075] In a possible example, the storing the currently processed access data into the at least one target cache work queue element includes: sequentially matching at least one set of data in the currently processed access data and the at least one target cache work queue element according to a preset matching rule and a preset matching order to obtain at least one matching information;
[0076] According to the at least one matching information, storing each set of data in the currently processed access data into the corresponding target cache work queue element.
[0077] Among them, the matching information is used to represent the corresponding relationship between a single set of data and a single target cache work queue element, and the data and target cache work queue elements corresponding to any two matching information are different. Among them, the preset matching order means that the data of the type of memory data is the first matching order, and the data of the type of operation success reply instruction or operation failure reply instruction is the second matching order; and, the preset matching rule means traversing the memory size of the at least one target cache work queue element, and matching the target cache work queue element with the memory size closest to the memory size applied for by the currently processed data for data caching.
[0078] Among them, it is obvious that the matching operation is to match the most suitable target cache work queue element for each group of data to store it, and preferably, it can ensure that the free memory size of the target cache work queue element storing it is equivalent to the memory size that this group of data wants to apply for, which can greatly save memory resources. Therefore, during the matching, the design of the matching order is that the data of the type of memory data that requires significantly more memory comes to match first, and then the instruction data that is only used to represent whether the operation is successful is matched. In terms of the matching rule, by obtaining the memory size of each target cache work queue element, the closest work queue element is determined for data caching. And of course, one more point that this closest one needs to meet is that the free memory size of the target cache work queue element used for storage is greater than the memory size that the corresponding data wants to apply for to ensure that it can be stored.
[0079] It can be seen that in this example, by presetting the matching rule and the matching order, each group of data in the currently processed access data can match the most suitable target cache work queue element for data storage, improving the efficiency of data storage and the stability of the operation of the terminal device, and improving the utilization rate of memory resources.
[0080] It can be seen that Figure 2 is a schematic flowchart of a data processing method for remote direct memory access provided by an embodiment of the present application. As shown in the figure, by obtaining at least one transceiver work queue corresponding to the current remote direct memory access operation executed by the first terminal device, and then performing a memory resource allocation operation according to the queue depth of the at least one transceiver work queue to obtain at least one memory resource information and allocate the free memory resources of the host for it; finally, after receiving at least one group of access data sent by the second terminal device in response to the remote direct memory access operation, store the at least one group of access data into at least one target cache work queue element determined by the associated target memory resource information. In this way, compared with the prior art of storing memory access data in the internal resources of the chip, the present application proposes to allocate the free host memory resources to the transceiver work queue in advance, so that the access data corresponding to each transceiver work queue exclusively enjoys a data cache work queue for data storage, improving the stability and efficiency of the terminal device for storing and processing access data.
[0081] The following is an embodiment of the device of the present application. The embodiment of the device of the present application and the embodiment of the method of the present application belong to the same concept and are used to execute the method described in the embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the device of the present application are shown in the embodiment of the device of the present application. For the specific technical details not disclosed, please refer to the description of the embodiment of the method of the present application, and details will not be repeated here.
[0082] An embodiment of the present application provides a data processing device for remote direct memory access. The data processing device is applied to a first terminal device in a remote memory access system, and the remote memory access system includes the first terminal device and a second terminal device. Specifically, the data processing device is configured to execute the steps performed by the first terminal device in the above data processing method for remote direct memory access. The data processing device for remote direct memory access provided by the embodiment of the present application may include modules corresponding to the respective steps.
[0083] In the embodiment of the present application, the data processing device may be divided into functional modules according to the above method example. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. The division of modules in the embodiment of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods.
[0084] In the case of dividing each functional module corresponding to each function, Figure 8 is a block diagram of the functional unit composition of a data processing device for remote direct memory access provided by an embodiment of the present application. The device is applied to the Figure 1 first terminal device 110 as described above, such as Figure 8As shown in the figure, the remote direct memory access data processing device 80 includes: a first execution unit 801, which is configured to detect an execution instruction for a remote direct memory access operation for the second terminal device, execute the remote direct memory access operation, and obtain a plurality of memory access work queues corresponding to the current remote direct memory access operation. The memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues includes at least one transceiver work queue. A single transceiver work queue includes a single memory receive queue and a single memory send queue. The single memory receive queue or memory send queue is used to store work queue elements, and a single work queue element is used to represent the request information of a single work request; a resource allocation unit 802, which is configured to perform a memory resource allocation operation according to the queue depth of the at least one transceiver work queue, obtain at least one memory resource information, and store each memory resource information in the work queue information of the corresponding transceiver work queue; a receiving unit 803, which is configured to receive at least one set of access data sent by the second terminal device in response to the remote direct memory access operation. The at least one set of access data corresponds to the at least one transceiver work queue one by one; a second execution unit 804, which is configured to perform the following operations for each set of access data in the at least one set of access data: determine the target memory resource information associated with the currently processed access data; determine at least one target cache work queue element in the target data cache work queue according to the target memory resource information; and store the currently processed access data in the at least one target cache work queue element.
[0085] In a possible example, when the memory resource information is the first resource information, the first resource information includes the base address information of the data cache work queue and at least one queue element position. The queue element position is used to represent the arrangement position of the corresponding target cache work queue element in the target data cache work queue. In terms of determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information, the second execution unit 804 is specifically configured to: determine the target cache work queue according to the base address information; determine the head information and tail information of the target cache work queue according to the target cache work queue. The head information is used to represent the number of cache work queue elements in the target data cache work queue whose memory has been occupied, and the tail information is used to represent the total number of cache work queue elements in the target data cache work queue; and determine the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
[0086] In a possible example, when the memory resource information is the second resource information, the second resource information includes memory address information, a target work queue index, and at least one queue element position, where the memory address information is used to indicate a memory address storing the address information of one or more data cache work queues; in terms of determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information, the second execution unit 804 is further specifically configured to: determine a target memory address according to the memory address information; determine the target data cache work queue in the target memory address according to the target memory address and the target work queue index; determine the head information and the tail information of the target cache work queue according to the target data cache work queue; and determine the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
[0087] In a possible example, in terms of determining the head information and the tail information of the target cache work queue according to the target data cache work queue, the second execution unit 804 is further specifically configured to: determine a plurality of reference work queue elements according to the target data cache work queue; sequentially number the plurality of reference work queue elements from small to large starting from the dequeue end of the target data cache work queue to obtain a plurality of queue element numbers, and the plurality of queue element numbers correspond to the plurality of reference work queue elements one by one; traverse the memory occupancy of the plurality of reference work queue elements, determine that the queue element number with the smallest value corresponding to the reference work queue element with an unoccupied memory occupancy is the head number, and determine that the queue element number with the largest value corresponding to the reference work queue element with an unoccupied memory occupancy is the tail number; generate head information according to the head number, and generate tail information according to the tail number.
[0088] In a possible example, the access data carries a sequence number, and the sequence number is used to indicate a transceiver work queue associated with the access data; in terms of determining the target memory resource information associated with the access data currently being processed, the second execution unit 804 is further specifically configured to: determine a target transceiver work queue according to the target sequence number carried by the access data currently being processed; determine the corresponding target work queue information according to the target transceiver work queue; and obtain the target memory resource information according to the target work queue information.
[0089] In a possible example, the access data includes at least one set of data, and the at least one set of data corresponds one-to-one with the at least one target cache work queue element; the type of the at least one set of data includes at least one of the following: memory data, operation success reply instruction, operation failure reply instruction; and, the type of the work request includes at least one of the following: memory read operation, memory rewrite operation, atomic operation.
[0090] In a possible example, in terms of storing the currently processed access data into the at least one target cache work queue element, the second execution unit 804 is specifically further configured to: sequentially match at least one set of data in the currently processed access data and the at least one target cache work queue element according to a preset matching rule and a preset matching order to obtain at least one matching information, where the matching information is used to represent the correspondence between a single set of data and a single target cache work queue element, and the data and target cache work queue elements corresponding to any two matching information are different from each other. Among them, the preset matching order means that the data of the type of memory data is the first matching order, and the data of the type of operation success reply instruction or operation failure reply instruction is the second matching order; and, the preset matching rule means traversing the memory size of the at least one target cache work queue element and matching the target cache work queue element with the memory size closest to the memory size applied for by the currently processed data for data caching; according to the at least one matching information, store each set of data in the currently processed access data into the corresponding target cache work queue element.
[0091] In the case of adopting an integrated unit, as Figure 9 shown, Figure 9 is a functional unit composition block diagram of another remote direct memory access data processing device provided by an embodiment of the present application. In Figure 9 it, the remote direct memory access data processing device 90 includes: a processing module 902 and a communication module 901. The processing module 902 is used to control and manage the actions of the remote direct memory access data processing device. For example, the steps of the first execution unit 801, the resource allocation unit 802, the receiving unit 803, and the second execution unit 804, and / or used to execute other processes of the technologies described herein. The communication module 901 is used to support the interaction between the remote direct memory access data processing device and other devices. As Figure 9 shown, the remote direct memory access data processing device may further include a storage module 903, and the storage module 903 is used to store the program code and data of the remote direct memory access data processing device.
[0092] Among them, the processing module 902 may be a processor or a controller. For example, it may be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 901 may be a transceiver, an RF circuit, a communication interface, etc. The storage module 903 may be a memory.
[0093] Among them, all relevant contents of each scenario involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here. The above remote direct memory access data processing device 90 can all execute the above Figure 2 The remote direct memory access data processing device shown.
[0094] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more collections of available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state drive.
[0095] Figure 10 is a structural block diagram of a terminal device provided by an embodiment of this application. As Figure 10As shown, the terminal device 1000 may include one or more of the following components: a processor 1001, and a memory 1002 coupled to the processor 1001. The memory 1002 may store one or more computer programs, and the one or more computer programs may be configured to implement the methods described in the above embodiments when executed by the one or more processors 1001. The terminal device 1000 may be the first terminal device 110 and the second terminal device 120 in the above embodiments.
[0096] The processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire terminal device 1000 using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1002, and by calling data stored in the memory 1002, it performs various functions of the terminal device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 1001 and may be implemented separately through a communication chip.
[0097] The memory 1002 may include random access memory (RAM) and may also include read-only memory (ROM). The memory 1002 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1002 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), and instructions for implementing the above method embodiments. The data storage area may also store data created during the use of the terminal device 1000.
[0098] It can be understood that the terminal device 1000 may include more or fewer structural elements than those in the above block diagram, which is not limited here.
[0099] An embodiment of this application also provides a computer storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, part or all of the steps of any one of the methods described in the above method embodiments are implemented.
[0100] An embodiment of this application also provides a computer program product. The above computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the above computer program is operable to cause a computer to execute part or all of the steps of any one of the methods described in the above method embodiments.
[0101] It should be understood that in various embodiments of this application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0102] In several embodiments provided by this application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0103] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0105] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, magnetic disks, optical disks, volatile memories, or non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM), etc., and various media that can store program code.
[0106] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions without departing from the spirit and scope of the present invention, and can make various changes and modifications, including combinations of the above different functions and implementation steps, including software and hardware implementation manners, all within the protection scope of the present invention.
Claims
1. A data processing method for remote direct memory access, characterized in that, A first terminal device applied to a remote memory access system, the remote memory access system including the first terminal device and a second terminal device, the method comprising: Detect an execution instruction for a remote direct memory access operation for the second terminal device, execute the remote direct memory access operation, and obtain a plurality of memory access work queues corresponding to the current remote direct memory access operation. The memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues includes at least one transceiver work queue. A single transceiver work queue includes a single memory receive queue and a single memory send queue. The single memory receive queue or memory send queue is used to store work queue elements, and a single work queue element is used to represent request information of a single work request; Perform a memory resource allocation operation according to the queue depth of the at least one transceiver work queue to obtain at least one memory resource information, and store each memory resource information in the work queue information of the corresponding transceiver work queue. The at least one memory resource information and the at least one transceiver work queue are in one-to-one correspondence. The queue depth is used to represent the number of work queue elements stored in the corresponding transceiver work queue. The memory resource allocation operation refers to an operation of allocating free host memory in the first terminal device to store access data obtained by executing the remote direct memory access operation. A single memory resource information is used to indicate a corresponding single data cache work queue. The data cache work queue is used to cache access data associated with the corresponding transceiver work queue. The data cache work queue is a work queue arranged with a plurality of cache work queue elements. The work queue information is used to store information associated with the corresponding transceiver work queue; Receive at least one set of access data sent by the second terminal device in response to the remote direct memory access operation. The at least one set of access data and the at least one transceiver work queue are in one-to-one correspondence; For each set of access data in the at least one set of access data, perform the following operations: Determine target memory resource information associated with the currently processed access data; Determine at least one target cache work queue element in a target data cache work queue according to the target memory resource information; Store the currently processed access data into the at least one target cache work queue element.
2. The method according to claim 1, wherein When the memory resource information is first resource information, the first resource information includes base address information of the data cache work queue and at least one queue element position, where the queue element position is used to represent the arrangement position of the corresponding target cache work queue element in the target data cache work queue. The determining at least one target cache work queue element in the target data cache work queue according to the target memory resource information includes: Determine the target cache work queue according to the base address information; Determine the head information and tail information of the target cache work queue according to the target cache work queue. The head information is used to represent the number of cache work queue elements in the target data cache work queue whose memory has been occupied, and the tail information is used to represent the total number of cache work queue elements in the target data cache work queue; Determine the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
3. The method according to claim 1, wherein When the memory resource information is the second resource information, the second resource information includes memory address information, a target work queue index, and at least one queue element position. The memory address information is used to indicate the memory address storing the address information of one or more data cache work queues. The determining the at least one target cache work queue element in the target data cache work queue according to the target memory resource information includes: Determine the target memory address according to the memory address information; Determine the target data cache work queue in the target memory address according to the target memory address and the target work queue index; Determine the head information and tail information of the target cache work queue according to the target data cache work queue; Determine the at least one target cache work queue element according to the head information, the tail information, and the at least one queue element position.
4. The method according to claim 2 or 3, characterized in that, The determining the head information and tail information of the target cache work queue according to the target data cache work queue includes: Determine a plurality of reference work queue elements according to the target data cache work queue; Number the plurality of reference work queue elements in ascending order starting from the dequeue end of the target data cache work queue to obtain a plurality of queue element numbers, and the plurality of queue element numbers correspond to the plurality of reference work queue elements one by one; Traverse the memory occupancy of the plurality of reference work queue elements, determine that the queue element number with the smallest value corresponding to the reference work queue element with the memory occupancy being unoccupied is the head number, and determine that the queue element number with the largest value corresponding to the reference work queue element with the memory occupancy being unoccupied is the tail number; Generate head information according to the head number and generate tail information according to the tail number.
5. The method according to any one of claims 1 to 3, characterized in that, The access data carries a sequence number (QPN), and the sequence number is used to indicate the send-receive work queue associated with the access data. The determining the target memory resource information associated with the currently processed access data includes: Determine the target send-receive work queue according to the sequence number carried by the currently processed access data; Determine the corresponding target work queue information according to the target send-receive work queue; Obtain the target memory resource information according to the target work queue information.
6. The method according to any one of claims 1-3, characterized in that, The access data includes at least one set of data, and the at least one set of data corresponds to the at least one target cache work queue element one by one. The type of the at least one set of data includes at least one of the following: Memory data, operation success reply instruction, operation failure reply instruction; and, The types of the work requests include at least one of the following: Memory read operation, memory rewrite operation, atomic operation.
7. The method according to claim 6, wherein storing the currently processed access data into the at least one target cache work queue element includes: Sequentially match at least one set of data in the currently processed access data and the at least one target cache work queue element according to a preset matching rule and a preset matching order to obtain at least one matching information, where the matching information is used to represent the corresponding relationship between a single set of data and a single target cache work queue element, and the data and target cache work queue elements corresponding to any two matching information are different from each other, where The preset matching order means that the data of the type of memory data is the first matching priority, and the data of the type of operation success reply instruction or operation failure reply instruction is the second matching priority; and, the preset matching rule means traversing the memory sizes of the at least one target cache work queue element, and matching the target cache work queue element with the memory size closest to the memory size requested by the currently processed data for data caching; According to the at least one piece of matching information, store each group of data in the currently processed access data into the corresponding target cache work queue element.
8. A data processing device for remote direct memory access, characterized in that Applied to a first terminal device in a remote memory access system, the remote memory access system includes the first terminal device and a second terminal device, and the device includes: A first execution unit, configured to detect an execution instruction for a remote direct memory access operation for the second terminal device, execute the remote direct memory access operation, and obtain a plurality of memory access work queues corresponding to the current remote direct memory access operation. The memory access work queue refers to a queue for storing work requests sent from software in the first terminal device to hardware. The plurality of memory access work queues include at least one transceiver work queue. A single transceiver work queue includes a single memory receiving queue and a single memory sending queue. The single memory receiving queue or memory sending queue is used to store work queue elements, and a single work queue element is used to represent the request information of a single work request; A resource allocation unit, configured to perform a memory resource allocation operation according to the queue depth of the at least one transceiver work queue, obtain at least one memory resource information, and store each memory resource information in the work queue information of the corresponding transceiver work queue; A receiving unit, configured to receive at least one group of access data sent by the second terminal device in response to the remote direct memory access operation, and the at least one group of access data corresponds to the at least one transceiver work queue one by one; A second execution unit, configured to perform the following operations for each group of access data in the at least one group of access data: determine the target memory resource information associated with the currently processed access data; determine at least one target cache work queue element in the target data cache work queue according to the target memory resource information; store the currently processed access data into the at least one target cache work queue element.
9. A terminal device, characterized in that, Including a processor, a memory, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the processor. The programs include instructions for performing the steps in the method according to any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Simple Flow Control Protocol Over RDMA
US20090319701A1
RDMA-based communication method, node, system, and medium
WO2022142562A1