Memory semantic tunnel communication method and device based on RDMA engine

By performing packet reorganization and adaptation in RDMA technology, the inefficiency and delay problems of RDMA technology in small packet memory semantic communication are solved, and more efficient and reliable data transmission is achieved.

CN120111005AActive Publication Date: 2025-06-06SHANGHAI SUIYUAN TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510592934.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing RDMA technology has problems such as low efficiency, long transmission delay and poor reliability in high-frequency scenarios in memory semantic communication scenarios that process small data packets.

Method used

By receiving locally sent memory semantic data, packet reorganization is performed to obtain merged data blocks, and the merged data blocks are adapted to RDMA payload information through the RDMA engine, backed up and encapsulated and sent to the remote device.

Benefits of technology

This method improves the payload transmission efficiency by reducing the additional overhead of the packet header during the packaging process, reduces the overall access delay, and ensures reliable data transmission through data backup.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111005A_ABST
    Figure CN120111005A_ABST
Patent Text Reader

Abstract

The invention discloses a memory semantic tunnel communication method and device based on an RDMA engine, and the method comprises the steps: receiving locally transmitted memory semantic data, and carrying out the data package recombination of the memory semantic data, and obtaining a merged data block; adapting the merged data block through a direct memory access (RDMA) engine interface adapter to obtain RDMA effective load information; and backing up the RDMA payload information in a local memory through an RDMA engine, packaging the RDMA payload information after backup is completed, and sending the packaged RDMA payload information to remote equipment. Small memory semantic data are merged into data blocks through data packet recombination, and the merged data blocks are sent after being packaged once, so that the extra overhead of packet headers in the packaging process is reduced, and the transmission efficiency of effective loads is improved; according to the method, data stored in a memory is directly pushed locally instead of being actively captured, so that low delay during first transmission of a data packet is ensured, the overall access delay can be remarkably reduced, and reliable transmission of the data is ensured through data backup.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a memory semantic communication method and device based on an RDMA engine. Background Art

[0002] With the rapid development of computer system architecture and high-performance computing, the efficiency and stability of network communication have become key factors affecting overall performance. In the current Remote Direct Memory Access (RDMA) technology, although its main advantage is efficient and low-latency direct memory access, there are still some problems in processing memory semantic communication scenarios of small data packets.

[0003] These problems mainly manifest themselves in low efficiency, long transmission delay and poor reliability in high-frequency scenarios. This is because in artificial intelligence scenarios, each small packet needs to have additional header information. When transmitting high-frequency small data packets, the header occupancy rate is too high, which will significantly reduce the transmission efficiency of the effective load. The traditional RDMA transmission implementation process has a complex software data transmission process, involving multiple direct memory access (DMA) operations, which increases the complexity of data sending operations and the difficulty of communication system scheduling. Multiple links will lead to a significant increase in transmission delay. In addition, data loss or damage may occur during data transmission, resulting in data transmission failure. Therefore, the overall effect of the existing RDMA data transmission method is not ideal. Summary of the invention

[0004] The present invention provides a memory semantic tunnel communication method based on an RDMA engine to improve the overall performance of RDMA memory semantic communication.

[0005] According to a first aspect of the present invention, a memory semantic tunnel communication method based on an RDMA engine is provided, comprising:

[0006] Receiving memory semantic data sent locally, and reorganizing data packets of the memory semantic data to obtain a merged data block;

[0007] Adapting the merged data block to obtain RDMA payload information through a direct memory access (RDMA) engine interface adapter;

[0008] The RDMA payload information is backed up in a local memory through an RDMA engine, and the backed-up RDMA payload information is packaged and sent to a remote device.

[0009] According to another aspect of the present invention, a memory semantic tunnel communication device based on an RDMA engine is provided, comprising:

[0010] A data reorganization module, used for receiving memory semantic data sent locally, and reorganizing data packets of the memory semantic data to obtain a merged data block;

[0011] An RDMA adapter module is used to adapt the merged data block to obtain RDMA payload information through a direct memory access RDMA engine interface adapter;

[0012] The data encapsulation module is used to back up the RDMA payload information in a local memory through an RDMA engine, and to encapsulate the backed-up RDMA payload information and send it to a remote device.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method described in any embodiment of the present invention when executed.

[0018] The technical solution of the embodiment of the present invention combines small memory semantic data into data blocks through data packet reorganization, and encapsulates the combined data blocks once before sending them, thereby reducing the additional overhead of the packet header during the encapsulation process, thereby improving the transmission efficiency of the effective load; using local direct push rather than actively grabbing data in the memory to ensure low latency when the data packet is transmitted for the first time, which can significantly reduce the overall access latency, and ensure reliable data transmission through data backup.

[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 It is a flowchart of a memory semantic tunnel communication method based on an RDMA engine provided according to Embodiment 1 of the present invention;

[0022] Figure 2 2 is a schematic diagram of an application framework for memory semantic encapsulation and decapsulation provided according to Embodiment 1 of the present invention;

[0023] Figure 3 It is a schematic diagram of the principle of memory semantic tunnel communication based on an RDMA engine provided according to the first embodiment of the present invention;

[0024] Figure 4 is a schematic diagram of a principle of data interaction provided according to the first embodiment of the present invention;

[0025] Figure 5 is a schematic diagram of a main memory path of three data paths provided according to Embodiment 1 of the present invention;

[0026] Figure 6 It is a flowchart of a memory semantic tunnel communication method based on an RDMA engine provided according to Embodiment 2 of the present invention;

[0027] Figure 7 This is a schematic diagram of the structure of an RDMA engine interface adapter provided according to Embodiment 2 of the present invention;

[0028] Figure 8 It is a structural diagram of a memory semantic tunnel communication device based on an RDMA engine provided according to Embodiment 3 of the present invention;

[0029] Fig. 9 It is a schematic diagram of the structure of an electronic device provided by Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 A flowchart of a memory semantic tunnel communication method based on an RDMA engine is provided for the first embodiment of the present invention. This embodiment is applicable to the case of communicating internal semantic data. The method can be executed by a memory semantic tunnel communication device based on an RDMA engine, which can be implemented in the form of hardware and / or software. Figure 1 As shown, the method includes:

[0034] Step S101, receiving memory semantic data sent locally, and reorganizing the memory semantic data into data packets to obtain a merged data block.

[0035] Among them, Figure 2 The figure shows the application framework diagram of memory semantic encapsulation and decapsulation in this embodiment. In the sending direction, it mainly involves key information extraction, data packet reassembly, RDMA engine interface adapter adaptation, and RDMA engine encapsulation and sending, while in the receiving direction, it mainly involves RDMA engine decapsulation, RDMA engine interface adapter processing, data packet splitting, and key information reassembly. Figure 3 The figure shows the principle diagram of memory semantic tunnel communication based on RDMA engine. Figure 3 right Figure 2 The relevant structures involved in the process are specifically described, and the relevant paths for sending and receiving data are specifically described. And when the local near-end device and the remote device use Figure 3 When communicating according to the principle shown in the figure, the specific principle diagram of data interaction is as follows Figure 4 As shown, a near-end device can communicate with multiple far-end devices at the same time. Figure 4 The example of communicating with one remote device is used for explanation. The specific number of remote devices communicating with the near-end device is not limited in this implementation.

[0036] Optionally, memory semantic data sent locally is received, and data packets of the memory semantic data are reassembled to obtain a merged data block, including: receiving memory semantic data sent by a local processor through an on-chip network bus, and sending the memory semantic data to a bus interface module; extracting parameters of each received memory semantic data through the bus interface module to obtain high-performance on-chip bus protocol AXI control information, memory semantic valid data and queue identifier, wherein the AXI control information includes the type of memory access; reassembling memory semantic valid data with the same queue identifier and belonging to the same type through a data packet reassembly component to obtain a merged data block, wherein the type includes a write request, a read request or a read response.

[0037] Specifically, in this embodiment, different types of memory semantic data sent by the local processor through the on-chip network bus will be received. In order to identify the specific type of memory semantics, the memory semantic data will be sent to the bus interface module, and the bus interface module will extract parameters from each received memory semantic data to obtain high-performance on-chip bus protocol AXI control information: AXI Meta, memory semantic valid data: AXI Payload and queue identifier. The AXI control information: AXI Meta specifically includes the type of memory access, such as local write request, local read request or remote read response. Of course, this embodiment is only an example and does not limit the specific type of memory access. In addition, different types of memory semantic valid data, such as write requests, read requests, write responses and read responses, respectively adopt standard memory semantic access message structures to reduce additional overhead in the RDMA protocol encapsulation. For example, the message structure corresponding to the write request includes data attributes, target address and data payload; the message structure corresponding to the read request includes data attributes and target address; the message structure corresponding to the write response includes response status; the message structure corresponding to the read response includes data attributes, response status and data payload. Of course, this implementation is only an example and does not limit the specific form of the message structure corresponding to different types of memory semantic valid data.

[0038] It should be noted that the queue identifier in this embodiment can be extracted indirectly, that is, first obtain the target address contained in the memory semantic data, and query the address queue relationship table according to the target address to obtain the queue identifier that matches the target address, wherein the address queue relationship table contains the correspondence between the target address and the queue identifier; the queried queue identifier is used as the queue identifier extracted from the memory semantic data. Of course, this embodiment is only an example, and does not limit the specific method of extracting the queue identifier. It can be seen from this that the memory semantic data with the same queue identifier needs to be sent to the same remote device.

[0039] Optionally, memory semantic valid data with the same queue identifier and belonging to the same type are reassembled through a data packet reassembly component to obtain a merged data block, including: grouping the memory semantic valid data of the same type according to the queue identifier in input order through the data packet reassembly component to obtain memory semantic data groups, wherein each memory semantic valid data group contains memory semantic valid data with the same type and queue identifier; obtaining a merge setting condition, and merging the memory semantic data groups according to the merge setting condition to obtain original merged data, wherein the merge setting condition includes the maximum allowable merged packet data and the maximum allowable merge duration; obtaining a data header based on the original merged data, and combining the original merged data with the data header to obtain a merged data block, wherein the number of merged data blocks corresponding to each memory semantic data group is at least one.

[0040] Specifically, in this embodiment, after obtaining the type and queue identifier of each memory semantic valid data, the data packet reassembly component can group the memory semantic valid data according to the queue identifier and type, and generate a larger data block by merging the data in the same queue, thereby improving the effective load ratio, and in this embodiment, independent write channels and read channels are used to merge local memory semantic data of different types, respectively, so as to achieve efficient management of data. For example, when the received memory semantic valid data are: M, N, K, M+1, K+1, N+1, N+2, N+3, K+2, M+2, and the types of the above memory semantic valid data are the same, for example, they can be local write requests, so in this case, when grouping, there is no need to consider the difference in type, and only the above memory semantic valid data need to be grouped according to the queue identifier in the input order to obtain the memory semantic data group, for example, memory semantic data group 1 = {N N+1 N+2 N+3}, memory semantic data group 2 = {K K+1 K+2 K+3}, memory semantic data group 3 = {M M+1 M+2 M+3}. In addition, in order to further improve the transmission efficiency, the present embodiment will pre-configure the merge setting conditions according to user needs. The merge setting conditions include the maximum allowed number of merged packets and the maximum allowed merge duration. The merge timeout mechanism is determined based on the maximum allowed merge duration. The timeout mechanism ensures that the data that has not been merged within a certain period of time can be directly transmitted to avoid excessive system response delay, thereby meeting the real-time requirements while transmitting efficiently. And in the present embodiment, after weighing the transmission delay and the RDMA transmission efficiency corresponding to different merge numbers, the merge number selected by the user will be used as the maximum allowed number of merged packets according to the actual transmission requirements, so that when the memory semantic data groups are subsequently merged, the number of each merged data block must not exceed the set maximum allowed number of merged packets. In the present embodiment, after obtaining the memory semantic data groups, each memory semantic data group will be merged according to the above-mentioned merge setting conditions to obtain the original merged data. If the maximum allowed merge time is reached before the data in the memory semantic data group is merged, the merge process is terminated, and the data merge result obtained at the termination time is used as the original merged data; if the maximum allowed merge time is not reached before the data in the memory semantic data group is merged, the memory semantic data group is merged according to the maximum allowed merge package number to obtain the original merged data. When merging, both the merge time and the merge number must be considered, so as to improve the data transmission efficiency while ensuring the real-time transmission. In addition, in this embodiment, the number of memory semantic valid data contained in the original merged data, such as 4; the type, such as a local read request; the target address, such as address 1, are also obtained, and the above-obtained number, type and target address are combined to obtain a data header, and the original merged data is combined with the data header to obtain a merged data block.

[0041] It should be noted that in this embodiment, different types of memory semantic data are merged separately using different merging components. For example, the local write request merging component only merges the valid memory semantic data of the local write request, the local read request merging component only merges the valid memory semantic data of the local read request, and the remote read response merging component only merges the valid memory semantic data of the remote read response, and each merging component can be executed independently and in parallel without affecting each other, thereby further improving the efficiency of data merging. Among them, local mainly refers to the sender of the request, and remote mainly refers to the receiver of the request, and for a local proximal device, it can correspond to multiple remote devices, that is, a proximal device can interact with multiple remote devices at the same time, which is not repeated in this embodiment.

[0042] Optionally, the method also includes: when it is determined that the type includes a write request, constructing a first write task that matches the memory semantic data, issuing the first write task, and updating the write status that matches the first write task in real time, wherein the write status includes issued but not completed or issued and completed; in response to a synchronization mark write request for the memory semantic data, constructing a synchronization mark update task that matches the memory semantic data; querying the target write status that matches the memory semantic data, and determining whether the target write status is issued and completed, and if so, issuing the synchronization mark update task, otherwise, blocking the synchronization mark update task.

[0043] Among them, a data transmission and synchronization is decomposed into Data+Flag mode. Among them, Data is the data to be transmitted, and Flag is the synchronization mark. When the receiving end queries the Flag update, the data transmission is considered to be completed. Due to the complex topology of the AI ​​communication network and the internal bus of the chip, it is necessary to strictly ensure that the Flag update is later than the Data write-through memory to ensure efficient data order preservation. However, as the chip bus topology becomes more and more complex, Data updates and Flag updates may go through different data paths. At this time, the existing hardware structure cannot guarantee that Data must be written earlier than Flag, which causes data disorder and significantly affects system reliability. Therefore, when it is determined that the type of memory semantic data is a write request, that is, when the memory semantic data on the local proximal device needs to be written to the remote chip, in order to ensure that the memory semantic data is written thoroughly before the subsequent data is written, the write status of the Data is updated in real time to provide a reliable update basis for the subsequent update of the corresponding Flag of the Data, so as to accurately determine whether the update of the corresponding Flag needs to be blocked based on the write status of the Data, and realize the final data barrier, that is, in complex AI communication scenarios, ensure that the Data is written through the memory before the Flag, solve the problem of disordered data transmission in complex AI communication scenarios, realize data order preservation in complex AI communication scenarios, and effectively improve system reliability. In addition, the data transmission order preservation method can realize data transmission order preservation in the AI ​​interconnected network, that is, when the chip of the proximal device writes data to the memory of another remote device chip, and the data corresponding synchronization mark is updated in the memory of the remote chip, the aforementioned data transmission order preservation method can strictly ensure that the synchronization mark is updated later than the data is written through the memory, so as to realize efficient data barrier.

[0044] For example, Figure 5 The figure shows a schematic diagram of a main memory path with three data paths. Figure 5As shown, the dotted line indicates the sending path corresponding to the normal write request (the path for CPU1 to access the main memory a2), the dotted line indicates the sending path corresponding to the synchronization mark write request (the path for CPU1 to access b2), and the solid line indicates the path for reading the synchronization mark (the path for CPU2 to access the main memory b2). After the CPU1 in the near-end device receives the response of all write data returned by the network port a, it executes the step of updating the Flag, thereby ensuring that the Flag must arrive at the network port later than the Data. When CPU1 sends a network data packet to update the Flag, the network data packet carries the attribute of whether a fence (data barrier) is required. The fence attribute passes through SOC bus 1 and then through the communication network to the remote device, and is used at the network port b of the remote device. On the network path between the near-end device and the remote device, the Rocv2 protocol is used to ensure that all data packets will not be out of order. The data packet across the AI ​​communication network carries the attribute of whether the current data packet needs to be fenced. At the remote device network port b, check whether the data packet carries the fence attribute. If it does, after the remote device network port receives the response of all write Data returned by the main memory a2, the Flag data packet carrying the Fence attribute is sent to SOC bus 2. This can ensure that the Flag arrives at the target device's network port b later than the Data.

[0045] Step S102: adapting the merged data block through a direct memory access (RDMA) engine interface adapter to obtain RDMA payload information.

[0046] Optionally, the merged data block is adapted to obtain RDMA payload information through direct memory access to an RDMA engine interface adapter, including: performing RDMA adaptation on the merged data block through the RDMA engine interface adapter to obtain adaptation information, and combining the adaptation information with the merged data block to obtain RDMA payload information, wherein the adaptation information includes an information header and an information tail; monitoring the RDMA payload information through the RDMA engine interface adapter, and generating an automatic response to feed back to the local processor when it is determined that a failure timeout occurs in the RDMA payload information.

[0047] Specifically, in this embodiment, the sequential issuance of memory semantic data is realized through the above-mentioned data order preservation method, and the memory semantic data is reorganized through the data preservation and reorganization component to obtain the merged data block, and the merged data block is adapted through the RDMA engine interface adapter to obtain the adaptation information, wherein the adaptation information can be the information header RDMA Header and the information tail, and the information header includes the adapter adaptation parameters for the merged blocks, such as the identifier of the RDMA tunnel used, the occupied RDMA transmission resources, etc., and the information tail is used to indicate the end position of the data block. Of course, this embodiment is only an example, and does not limit the specific parameters of the adaptation information. Among them, the RDMA engine interface adapter in this embodiment is connected to the RDMA engine, and notifies the RDMA engine of the RDMA payload information, so that the RDMA engine pulls the RDMA payload information and performs RDMA encapsulation.

[0048] Step S103: back up the RDMA payload information in the local memory through the RDMA engine, and encapsulate the backed-up RDMA payload information and send it to the remote device.

[0049] Optionally, the RDMA payload information is backed up in a local memory through the RDMA engine, and the backed-up RDMA payload information is encapsulated and sent to a remote device, including: generating queue data to be backed up according to the RDMA payload information through a data sending component of the RDMA engine, and backing up the queue data to be backed up in a local memory; obtaining queue data not to be backed up, and arbitrating current data to be encapsulated from the queue data to be backed up and the queue data not to be backed up, wherein the queue data not to be backed up includes a data reception response and / or target retransmission data; adding a packet header to the data to be encapsulated to obtain encapsulated data, and sending the encapsulated data to the remote device through a switch.

[0050] Specifically, the RDMA payload information in this embodiment may include a read request type, a write request type, and a read response type, so the queue data to be backed up may be the data to be transmitted sent by the local processor of the near-end device, and the data to be transmitted composed of the data requested to be read by the remote device, and exists in the form of a queue. In this embodiment, the queue data to be backed up may be stored in the local memory to realize the backup of the queue data to be backed up, and while backing up the queue data to be backed up to the local memory, the queue data to be backed up in the RDMA packet structure (the structure of the RDMA data packet) is sent to the target data receiving end. In addition, in this embodiment, non-queue data to be backed up may be obtained, and the non-queue data to be backed up may be other data to be sent to the remote device in addition to the queue data to be backed up, and exists in the form of a queue, for example, data reception response and / or target retransmission data, and the target retransmission data may be data that the remote device fails to receive and is backed up in the local memory. The current data to be encapsulated may be the data that is currently required to be sent to the remote device and has not been encapsulated as arbitrated. The target encapsulated data may be a data packet encapsulated by the current data to be encapsulated according to the RDMA packet structure, and the encapsulated data may be a data packet encapsulated by the current data to be encapsulated according to the RDMA packet structure. Among them, when sending the queue data to be backed up to the remote device, if there is also queue data not to be backed up to be sent to the remote device at the same time, the queue data to be backed up and the queue data not to be backed up can be sorted according to the preset data sorting rules, and the data at the top of the sorting list is used as the current data to be encapsulated, and then the current data to be encapsulated is encapsulated according to the RDMA packet structure to obtain the target encapsulated data, and then the target encapsulated data is sent to the target data receiving end via Ethernet.

[0051] It should be noted that, in this embodiment, before sending non-backup queue data to the remote device, it also includes: parsing the RDMA data to be received sent by the remote device, and generating a data reception response when the RDMA data to be received passes the verification; and / or, when there is a data reception anomaly in the remote device, determining the target retransmission data in the local memory. Among them, the RDMA data to be received can be data in the RDMA packet structure sent by the remote device to the local processor. If the remote device sends the RDMA data to be received to the local processor side, the RDMA data to be received sent by the remote device is decapsulated based on the RDMA packet structure, and the data legitimacy of the decapsulated data is checked. When the decapsulated data passes the data legitimacy check, it indicates that the RDMA data to be received has passed the verification, and a data reception response for the RDMA data to be received is generated. If there is a data reception anomaly at the target data receiving end, indicating that the historical backup queue data (the queue data backed up in the local memory before the aforementioned queue data to be backed up, and the data composition of the historical backup queue data can be specifically referred to the queue data to be backed up) sent by the local processor to the remote device has not been successfully received, then the target retransmission data that has not been successfully received by the target data receiving end is queried from the local memory.

[0052] It is worth mentioning that, in this embodiment, after parsing the RDMA data to be received sent by the remote device, it also includes determining the packet header information of the target queue data successfully received by the remote device when the RDMA data to be received passes the verification; and deleting data from the local memory according to the packet header information of the target queue data. Among them, the target queue data can be data sent by the local processor to the remote device and successfully received by the remote device, existing in the form of a queue. If the RDMA data to be received passes the verification, it can be further determined that the remote device has successfully received the target queue data sent by the local processor, and then the packet header information of the target queue data is obtained, so as to locate the space for storing the target queue data from the network card memory according to the packet header information of the target queue data, and clear the space.

[0053] Optionally, adding a packet header to the data to be encapsulated to obtain the encapsulated data includes: adding an RDMA packet header to the data to be encapsulated based on the RDMA packet component in the RDMA tunnel to obtain a first packet result; adding UDP and IP packet headers to the first packet result based on the UDP and IP packet components in the RDMA tunnel to obtain a second packet result; adding an Ethernet packet header to the second packet result based on the Ethernet component in the RDMA tunnel to obtain the encapsulated data.

[0054] Among them, after merging the memory semantic data to obtain the merged data block, and arbitrating the merged data block to determine the data to be encapsulated, in order to realize data transmission between different protocols, it is necessary to use the RDMA tunnel to add a header to the data to be encapsulated, and when encapsulating, different packet components are used to add headers with different contents to the data to be encapsulated. For example, an RDMA header is added through an RDMA packet component, a UDP or IP header is added through a UDP or IP packet component, and an Ethernet header is added through an Ethernet component. Therefore, the header in the obtained encapsulated data specifically includes three items, namely, an RDMA header, a UDP and IP header, and an Ethernet header. In addition, in this embodiment, different RDMA headers are used for different types of data, as shown in Table 1 below, which are the RDMA headers corresponding to the memory write request type:

[0055] Table 1

[0056]

[0057] Table 2 below shows the RDMA packet header corresponding to the memory read request type:

[0058] Table 2

[0059]

[0060] Table 3 below shows the RDMA packet header corresponding to the memory read response type:

[0061] Table 3

[0062]

[0063] In addition, in order to distinguish the difference between the RDMA tunnel and the ordinary RDMA message, the present embodiment adopts the custom Opcode programming to realize the memory semantic operation, and the following table 4 shows the type example of realizing the memory semantic operation:

[0064] Table 4

[0065]

[0066] In this implementation, by merging multiple small memory semantic data and then encapsulating them, it is only necessary to add a packet header once, thereby significantly reducing the proportion of encapsulated data occupied by the packet header and improving the transmission efficiency of the effective load; and by logically merging multiple discrete local memory semantic data, the number of packets transmitted per unit time is reduced, and the hardware requirements for packet processing performance are reduced; under the premise of ensuring that the delay is controlled, a reasonable merging strategy is used to reduce the hardware resource consumption of frequent transmission of small packets in the RDMA tunnel.

[0067] Optionally, the method also includes: receiving remote encapsulated data sent by a remote device through a data receiving component of the RDMA engine, and decapsulating the remote encapsulated data to obtain remote RDMA payload information; sending the remote RDMA payload information to the RDMA engine interface adapter, and obtaining a remote merged data block according to the RDMA payload information through the RDMA engine interface adapter; splitting the remote merged data block through a data packet splitting component to obtain remote memory semantic valid data and remote AXI control information; merging the remote memory semantic valid data and the remote AXI control information through a bus interface module to obtain remote memory semantic data, and transmitting the remote memory semantic data to the local through an on-chip network bus.

[0068] Optionally, the remote memory semantic data is transmitted to the local via the on-chip network bus, including: when the type of the remote memory semantic data is a remote write request or a remote read request, the remote memory semantic data is sent to the local memory; when the type of the remote memory semantic data is a local read response, the remote memory semantic data is sent to the local processor.

[0069] It should be noted that, in this embodiment, on the one hand, the memory semantic data can be merged, and the encapsulated data obtained by the merge encapsulation can be sent to the remote device, and on the other hand, the encapsulated data sent by the remote device can be unpacked, and the different types of remote merged data blocks sent after the RDMA component is unpacked can be split, wherein the remote merged data block contains remote memory semantic data of the same type and queue identifier, and also contains a data header. When splitting the remote merged data block according to the data header to obtain the remote memory semantic data, different components are used for different types of merged data blocks, for example, the remote write request splitting component only splits the remote write request merged data block, the remote read request splitting component only splits the remote read request merged data block, and the local read response splitting component only splits the local read response merged data block, and each splitting component can be executed independently and in parallel without affecting each other, thereby further improving the efficiency of data splitting. And when data splitting is performed, the data in the receiving direction passes through the RDMA engine, the RDMA engine interface adapter, the data packet splitting component, the bus interface module and the NOC bus respectively, and the processing process of each component in the receiving direction is the reverse process of the above-mentioned sending direction. Since the specific principles of each component have been specifically explained above, they will not be repeated in this embodiment.

[0070] In the implementation manner of the present application, small memory semantic data are merged into data blocks through data packet reorganization, and the merged data blocks are encapsulated once and then sent, thereby reducing the additional overhead of the packet header during the encapsulation process, thereby improving the transmission efficiency of the effective load; using local direct push instead of actively grabbing data in the memory to ensure low latency during the first transmission of the data packet, which can significantly reduce the overall access latency, and ensure reliable data transmission through data backup.

[0071] Embodiment 2

[0072] Figure 6 A flowchart of a memory semantic tunnel communication method based on an RDMA engine is provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiment and specifically describes the process of monitoring RDMA payload information through an RDMA engine interface adapter and generating an automatic response feedback to a local processor when it is determined that a failure timeout occurs in the RDMA payload information. Figure 6 As shown, the method includes:

[0073] Step S201: monitor incoming RDMA payload information through the RDMA engine interface adapter, and record the acquired queue status information into a queue pair status record table.

[0074] Specifically, this application is mainly aimed at the situation where chips in the same system use queue pairs (QP) to perform RDMA communication in the RDMA scenario, such as Figure 7 As shown, it is a schematic diagram of the structure of the RDMA engine interface adapter in this embodiment, which mainly includes a queue pair state record table, an automatic response state machine, a multi-queue first-in first-out component, a timeout queue record table and a response filter. Of course, the RDMA engine interface adapter in this embodiment also includes other structures, and in this embodiment, only the structure related to the automatic response to the fault is displayed. When the structure is located in the near-end device, when the near-end device sends the RDMA payload information in the form of a queue to the remote device to send a fault, the RDMA engine interface adapter can be used to perform an automatic response inside the near-end device to avoid the near-end device from failing to obtain the response of the remote device due to information failure.

[0075] In this embodiment, when the near-end device and the remote device are performing data transmission, the near-end device sends the RDMA payload information to the remote device in the form of a queue. Figure 7The RDMA engine interface adapter shown will monitor the queue pairs flowing through the transmission data path in real time. The RDMA payload information in this embodiment includes multiple queue pairs, and is sent as a whole when sent to the remote device, but can be transmitted in the form of queue pairs in the transmission data path inside the near-end chip. A queue identifier is marked in each queue pair, so the queue pairs with the same queue identifier belong to the same RDMA payload information. In this embodiment, the specific number of queues included in the RDMA payload information sent by the near-end device to the remote device is not limited. The RDMA engine interface adapter obtains the queue identifier and detection timestamp of each queue pair flowing through the sending data path through monitoring, determines the target queue to which the queue pair belongs according to the queue identifier, and updates the queue timestamp of the target queue recorded in the queue pair status record table according to the detection timestamp. For example, when it is determined that the queue identifier marked by the queue pair a that has most recently flown through the sending data path is 1, and the detection time is T1, the queue timestamp TimeStamp of queue 1 is updated to T1 in the queue pair status record table, and the outstanding counter of queue 1 is increased by 1, wherein the outstanding counter refers to the number of unfinished items for which the request direction of the proximal device has sent a request to the remote device but has not received a response from the remote device. Therefore, the queue timestamp and flight quantity of each queue are specifically recorded in the queue pair status record table, and the queue timestamp and flight quantity of each queue are constantly changing during the monitoring period.

[0076] After recording the acquired queue status information in the queue pair status record table, the method further includes: acquiring key information and queue identifiers of the queue pair; and saving the corresponding relationship between the key information and the queue identifiers in the multi-queue first-in-first-out component.

[0077] Step S202: poll the queue pair status record table through the automatic response state machine in the RDMA engine interface adapter, and determine the running status of each queue according to the queue status information.

[0078] Among them, the queue pair status record table is polled through the automatic response state machine in the RDMA engine interface adapter, including: the queue pair status record table is polled through the automatic response state machine in the RDMA engine interface adapter to obtain the queue status information of each queue; for each queue, it is determined whether the number of flights in the queue status information is 0, if so, the operation status of the queue currently polled is directly determined to be normal, otherwise, the operation status of the queue currently polled is determined according to the current polling time and the queue timestamp in the queue status information. And the operation status of the queue currently polled is determined according to the current polling time and the queue timestamp in the queue status information, which can be specifically adopted by obtaining a pre-configured time threshold and the time difference between the current polling time and the queue timestamp; judging whether the time difference is greater than the time threshold, if so, the operation status of the queue currently polled is determined to be a fault timeout, otherwise, the operation status of the queue currently polled is determined to be normal. Therefore, the operation status in this embodiment includes two states: fault timeout and normal.

[0079] Step S203: record the queue identifier of the operation state of fault timeout into the timeout queue record table, and obtain the target queue identifier corresponding to the RDMA payload information.

[0080] Specifically, in this implementation, the queue identifier whose operating status is determined to be a fault timeout is recorded in the timeout queue record table. Since the queue status record table is updated in real time during the process of the near-end device sending information to the far-end device, the corresponding timeout queue record table is also updated in real time.

[0081] Step S204: when it is determined that the target queue identifier is located in the timeout queue record table, an automatic response is generated for the RDMA payload information and fed back to the local processor.

[0082] Among them, in this embodiment, the target queue identifier corresponding to the RDMA payload information is obtained, and when it is determined that the target queue identifier is located in the timeout queue record table, an automatic response is generated for the RDMA payload information, and the fault queue in the automatic response generated by the RDMA payload information is filtered and discarded, thereby achieving isolation between the proximal chip and the remote device. And while generating an automatic response for the faulty queue, it does not affect the acquisition of the remote device's response for the normal queue, that is, the normal queue can still be accessed normally, so that the proximal device can obtain the execution result based on the response of all the queues obtained, that is, the proximal device is always in a keep-alive state, and the operation of the entire device or even the entire system will not be interrupted due to the failure of a single queue inside the proximal device, thereby achieving isolation between the internal chip fault and the external network.

[0083] In the implementation mode of the present application, small memory semantic data is merged into data blocks through data packet reorganization, and the merged data blocks are encapsulated once before being sent, which reduces the extra overhead of the packet header during the encapsulation process, thereby improving the transmission efficiency of the effective load; using local direct push instead of actively grabbing the data in the memory to ensure low latency when the data packet is first transmitted, which can significantly reduce the overall access latency, and ensure reliable data transmission through data backup. And when the near-end chip requester causes a queue failure in the request sent to the remote chip due to a port failure, an automatic response will be generated for the fault queue and fed back to the near-end chip requester, and the fault queue does not interfere with other normal queues, so the network failure does not affect the normal operation of the near-end chip, and fault convergence and isolation are achieved. And the near-end chip is kept alive through automatic response, which can ensure the normal operation of other chips connected to the near-end chip to achieve the normal operation of the entire system, thereby improving the reliability and availability of the system.

[0084] Embodiment 3

[0085] Figure 8 A schematic diagram of the structure of a memory semantic tunnel communication device based on an RDMA engine is provided in Embodiment 4 of the present invention. Figure 8 As shown, the device includes: a data reorganization module 310, an RDMA adaptation module 320 and a data encapsulation module 330.

[0086] The data reorganization module 310 is used to receive the memory semantic data sent locally, and reorganize the data packets of the memory semantic data to obtain the merged data blocks;

[0087] The RDMA adapter module 320 is used to adapt the merged data block to obtain RDMA payload information through a direct memory access RDMA engine interface adapter;

[0088] The data encapsulation module 330 is used to back up the RDMA payload information in the local memory through the RDMA engine, and to encapsulate the backed-up RDMA payload information and send it to the remote device.

[0089] Optionally, a data reorganization module is used to receive memory semantic data sent by the local processor through the on-chip network bus, and send the memory semantic data to the bus interface module;

[0090] Extract parameters of each received memory semantic data through the bus interface module to obtain high-performance on-chip bus protocol AXI control information, memory semantic valid data and queue identifier, wherein the AXI control information includes the type of memory access;

[0091] The memory semantic valid data with the same queue identifier and belonging to the same type are reassembled through a data packet reassembly component to obtain a merged data block, wherein the type includes a write request, a read request or a read response.

[0092] Optionally, the data reorganization module is further used to group the same type of memory semantic valid data according to the queue identifier in the input order through the data packet reorganization component to obtain memory semantic data groups, wherein each memory semantic valid data group contains memory semantic valid data with the same type and queue identifier;

[0093] Acquire a merge setting condition, and merge the memory semantic data group according to the merge setting condition to obtain original merged data, wherein the merge setting condition includes a maximum allowable merged package data and a maximum allowable merge duration;

[0094] A data header is obtained according to the original merged data, and the original merged data and the data header are combined to obtain a merged data block, wherein the number of merged data blocks corresponding to each memory semantic data group is at least one.

[0095] Optionally, the device further includes an order-preserving module, which is used to construct a first write task matching the memory semantic data when the determined type includes a write request, issue a task for the first write task, and update a write status matching the first write task in real time, wherein the write status includes issued but not completed or issued and completed;

[0096] In response to a synchronization tag write request for memory semantic data, construct a synchronization tag update task matching the memory semantic data;

[0097] Query the target write status that matches the memory semantic data, and determine whether the target write status has been issued and completed. If so, issue the synchronization mark update task; otherwise, block the synchronization mark update task.

[0098] Optionally, the RDMA adaptation module includes an RDMA configuration unit, which is used to perform RDMA adaptation on the merged data block through an RDMA engine interface adapter to obtain adaptation information, and combine the adaptation information with the merged data block to obtain RDMA payload information;

[0099] The fault automatic response unit is used to monitor the RDMA payload information through the RDMA engine interface adapter, and when it is determined that the RDMA payload information has a fault timeout, an automatic response is generated and fed back to the local processor.

[0100] Optionally, a fault automatic response unit is used to monitor incoming RDMA payload information through an RDMA engine interface adapter, and record the acquired queue status information into a queue pair status record table, wherein the queue status information includes a queue timestamp and a flight quantity;

[0101] The queue pair status record table is polled through the automatic response state machine in the RDMA engine interface adapter, and the operation status of each queue is determined according to the queue status information, wherein the operation status includes fault timeout or normal;

[0102] Record the queue ID of the running state as fault timeout into the timeout queue record table, and obtain the target queue ID corresponding to the RDMA payload information;

[0103] When it is determined that the target queue identifier is located in the timeout queue record table, an automatic response is generated for the RDMA payload information and fed back to the local processor.

[0104] Optionally, a data encapsulation module, used to generate the queue data to be backed up according to the RDMA payload information through the data sending component of the RDMA engine, and back up the queue data to be backed up in the local storage;

[0105] Acquire the non-to-be-backed up queue data, and arbitrate the current to-be-encapsulated data from the to-be-backed up queue data and the non-to-be-backed up queue data, wherein the non-to-be-backed up queue data includes a data receiving response and / or target retransmission data;

[0106] A packet header is added to the data to be encapsulated to obtain the encapsulated data, and the encapsulated data is sent to the remote device through the switch.

[0107] Optionally, a data encapsulation module, configured to add an RDMA packet header to the data to be encapsulated based on an RDMA packet component in the RDMA tunnel to obtain a first packet result;

[0108] Adding UDP and IP headers to the first packet result based on the UDP and IP packet components in the RDMA tunnel to obtain a second packet result;

[0109] An Ethernet header is added to the second packet result based on the Ethernet component in the RDMA tunnel to obtain encapsulated data.

[0110] Optionally, the apparatus further comprises a remote data processing module, which is used to receive remote encapsulated data sent by a remote device through a data receiving component of the RDMA engine, and decapsulate the remote encapsulated data to obtain remote RDMA payload information;

[0111] Sending remote RDMA payload information to the RDMA engine interface adapter, and obtaining remote merged data blocks according to the RDMA payload information through the RDMA engine interface adapter;

[0112] The remote merged data block is split by the data packet splitting component to obtain the remote memory semantic valid data and remote AXI control information;

[0113] The remote memory semantic valid data and the remote AXI control information are merged through the bus interface module to obtain the remote memory semantic data, and the remote memory semantic data is transmitted to the local through the on-chip network bus.

[0114] Optionally, the remote data processing module is further used to send the remote memory semantic data to the local memory when the type of the remote memory semantic data is a remote write request or a remote read request;

[0115] When the type of the remote memory semantic data is a local read response, the remote memory semantic data is sent to the local processor.

[0116] The memory semantic tunnel communication device based on the RDMA engine provided in the embodiment of the present invention can execute the memory semantic tunnel communication method based on the RDMA engine provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0117] Embodiment 5

[0118] Fig. 9 The present invention is a block diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0119] like Fig. 9As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11 in communication, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0120] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0121] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a memory semantic tunnel communication method based on an RDMA engine.

[0122] In some embodiments, the memory semantic tunnel communication method based on the RDMA engine can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the memory semantic tunnel communication method based on the RDMA engine described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the memory semantic tunnel communication method based on the RDMA engine in any other appropriate manner (for example, by means of firmware).

[0123] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0124] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0125] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0127] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0128] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0129] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0130] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A memory semantic tunnel communication method based on RDMA engine, characterized in that: include: Receiving memory semantic data sent locally, and reorganizing data packets of the memory semantic data to obtain a merged data block; Adapting the merged data block to obtain RDMA payload information through a direct memory access (RDMA) engine interface adapter; The RDMA payload information is backed up in a local memory through an RDMA engine, and the backed-up RDMA payload information is packaged and sent to a remote device.

2. The method according to claim 1, characterized in that The receiving of the memory semantic data sent locally and reorganizing the data packets of the memory semantic data to obtain the merged data block includes: Receiving memory semantic data sent by the local processor via the on-chip network bus, and sending the memory semantic data to the bus interface module; Extracting parameters of each of the received memory semantic data through the bus interface module to obtain high-performance on-chip bus protocol AXI control information, memory semantic valid data and queue identifier, wherein the AXI control information includes a memory access type; The memory semantic valid data with the same queue identifier and belonging to the same type are reassembled by a data packet reassembly component to obtain the merged data block, wherein the type includes a write request, a read request or a read response.

3. The method according to claim 2, characterized in that The step of reorganizing the memory semantic valid data having the same queue identifier and belonging to the same type by a data packet reorganization component to obtain the merged data block includes: The data packet reassembly component groups the same type of memory semantic valid data in input order according to the queue identifier to obtain memory semantic data groups, wherein each of the memory semantic valid data groups contains memory semantic valid data of the same type and queue identifier; Acquire a merge setting condition, and perform data merging on the memory semantic data group according to the merge setting condition to acquire original merged data, wherein the merge setting condition includes a maximum allowable merged package data and a maximum allowable merge duration; A data header is obtained according to the original merged data, and the original merged data and the data header are combined to obtain the merged data block, wherein the number of the merged data blocks corresponding to each of the memory semantic data groups is at least one.

4. The method according to claim 2, characterized in that: The method further comprises: When it is determined that the type includes a write request, a first write task matching the memory semantic data is constructed, the first write task is issued, and a write status matching the first write task is updated in real time, wherein the write status includes issued but not completed or issued and completed; In response to a synchronization tag write request for the memory semantic data, construct a synchronization tag update task matching the memory semantic data; Query the target write status that matches the memory semantic data, and determine whether the target write status has been issued and completed. If so, issue the synchronization mark update task; otherwise, block the synchronization mark update task.

5. The method according to claim 1, characterized in that: The step of adapting the merged data block to obtain RDMA payload information through a direct memory access (RDMA) engine interface adapter includes: Performing RDMA adaptation on the merged data block through the RDMA engine interface adapter to obtain adaptation information, and combining the adaptation information with the merged data block to obtain the RDMA payload information; The RDMA payload information is monitored by the RDMA engine interface adapter, and when it is determined that the RDMA payload information has a fault timeout, an automatic response is generated and fed back to the local processor.

6. The method according to claim 5, characterized in that The RDMA payload information is monitored by the RDMA engine interface adapter, and when it is determined that the RDMA payload information has a fault timeout, an automatic response is generated and fed back to the local processor, including: Monitoring the incoming RDMA payload information through the RDMA engine interface adapter, and recording the acquired queue status information into a queue pair status record table, wherein the queue status information includes a queue timestamp and a flight quantity; The queue pair status record table is polled through an automatic response state machine in the RDMA engine interface adapter, and the operation status of each queue is determined according to the queue status information, wherein the operation status includes fault timeout or normal; Record the queue identifier of the running state as fault timeout into the timeout queue record table, and obtain the target queue identifier corresponding to the RDMA payload information; When it is determined that the target queue identifier is located in the timeout queue record table, an automatic response is generated for the RDMA payload information and fed back to the local processor.

7. The method according to claim 1, characterized in that The RDMA payload information is backed up in a local memory by the RDMA engine, and the backed-up RDMA payload information is packaged and sent to a remote device, including: Generate the to-be-backed-up queue data according to the RDMA payload information through the data sending component of the RDMA engine, and back up the to-be-backed-up queue data in the local memory; Acquire non-to-be-backed-up queue data, and arbitrate current to-be-encapsulated data from the to-be-backed-up queue data and the non-to-be-backed-up queue data, wherein the non-to-be-backed-up queue data includes data reception response and / or target retransmission data; A packet header is added to the data to be encapsulated to obtain encapsulated data, and the encapsulated data is sent to the remote device through a switch.

8. The method according to claim 7, characterized in that The step of adding a header to the data to be encapsulated to obtain encapsulated data includes: Adding an RDMA packet header to the data to be encapsulated based on the RDMA packet component in the RDMA tunnel to obtain a first packet result; Adding UDP and IP packet headers to the first packet result based on the UDP and IP packet components in the RDMA tunnel to obtain a second packet result; An Ethernet header is added to the second packet result based on the Ethernet component in the RDMA tunnel to obtain the encapsulated data.

9. The method according to claim 1, characterized in that: The method further comprises: Receiving the remote encapsulated data sent by the remote device through the data receiving component of the RDMA engine, and decapsulating the remote encapsulated data to obtain remote RDMA payload information; Sending the remote RDMA payload information to the RDMA engine interface adapter, and acquiring the remote merged data block according to the RDMA payload information through the RDMA engine interface adapter; Splitting the remote merged data block by a data packet splitting component to obtain remote memory semantic valid data and remote AXI control information; The remote memory semantic valid data and the remote AXI control information are merged through the bus interface module to obtain the remote memory semantic data, and the remote memory semantic data is transmitted to the local through the on-chip network bus.

10. The method according to claim 9, characterized in that The step of transmitting the remote memory semantic data to the local memory via the on-chip network bus includes: When the type of the remote memory semantic data is a remote write request or a remote read request, the remote memory semantic data is sent to the local memory; When the type of the remote memory semantic data is a local read response, the remote memory semantic data is sent to the local processor.

11. A memory semantic tunnel communication device based on RDMA engine, characterized in that: include: A data reorganization module, used for receiving memory semantic data sent locally, and reorganizing data packets of the memory semantic data to obtain a merged data block; An RDMA adapter module is used to adapt the merged data block to obtain RDMA payload information through a direct memory access RDMA engine interface adapter; The data encapsulation module is used to back up the RDMA payload information in a local memory through an RDMA engine, and to encapsulate the backed-up RDMA payload information and send it to a remote device.

12. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so as to enable the at least one processor to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 10 when executed.

Citation Information

Patent Citations

  • Query engine system for distributed memory database based on RDMA

    CN107329814A

  • A method and a device for accessing NAND FLASH by an AXI bus

    CN109726149A

  • DDR arbitration and scheduling method and system based on AXI protocol

    CN113641603A

  • Data storage method and system, storage access configuration method and related equipment

    CN116414735A

  • NVMe write data processing method, terminal and storage medium

    CN118860290A