Remote data transmission method based on RDMA protocol

Through the long-distance data transmission method based on the RDMA protocol, the problem of insufficient performance of the traditional TCP protocol stack at high link speed is solved, and efficient and low-latency data transmission is achieved, which is suitable for cloud computing and distributed architectures.

CN119996404AActive Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510154730.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The traditional TCP network protocol stack cannot meet the high performance and low latency requirements of cloud computing and distributed architectures at high link rates, resulting in high CPU resource consumption and increased processing latency.

Method used

Using the long-distance data transmission method based on the RDMA protocol, reliable connection is established through RDMACM, efficient address buffer pools and sliding windows are designed, necessary RDMA resources are created, and performance is optimized using a batch processing mechanism.

Benefits of technology

It reduces the impact of the network environment on transmission efficiency, significantly improves the speed and efficiency of RDMA data transmission, reduces CPU utilization, and realizes low latency and high throughput data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996404A_ABST
    Figure CN119996404A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote data transmission, in particular to a remote data transmission method based on an RDMA (Remote Direct Memory Access) protocol, which comprises the following steps: a sending end and a receiving end establish reliable connection, transmit related information of a file and exchange key parameters required by subsequent communication; a receiving end designs an adaptive address buffer pool and a proper sliding window size; the receiving end and the sending end create necessary resources before RDMA transmission, and then start file transmission; a batch processing mechanism of completion queue elements is designed, multiple pieces of completion information are combined, and the performance is further optimized; the sending end reads and blocks file content, requests an available address from an address pool of the receiving end, writes the file blocks into a specified memory of the receiving end, writes the memory into a specified position of a hard disk after writing is completed, and reuses the address; and after file transmission is completed, the created related resources are released, so that the speed and efficiency of RDMA data transmission are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of long-distance data transmission, and in particular to a long-distance data transmission method based on the RDMA protocol. Background Art

[0002] With the continuous advancement of cloud computing technology, distributed architecture has been widely used in high-performance computing and storage due to its high performance, high reliability and easy scalability. As an important infrastructure supporting distributed business, data centers have developed rapidly in recent years.

[0003] In order to meet the transmission requirements of upper-layer services, data center networks are constantly upgrading. The link rate has been increased from the early 10Gbps to the current 100Gbps, and a new generation of hardware devices supporting 400Gbps rates have been launched. For example, NVIDIA's fourth-generation in-machine GPU communication link NV-Link has a total bandwidth of up to 900Gbps for single-GPU in-machine communication. At the same time, the business requirements for basic transmission latency are becoming increasingly stringent, with the goal of microseconds. For example, cloud storage needs to provide performance close to local storage. Taking Alibaba Cloud's enhanced SSD as an example, its 4KB small message latency commitment is controlled within 200 microseconds. For distributed machine learning tasks, in order to efficiently train models and process data, the infrastructure network needs to compress the tail latency of short messages to less than 50 microseconds.

[0004] However, the traditional TCP network protocol stack can no longer meet the above requirements. In terms of link rate, the TCP protocol stack needs to occupy about 1Hz CPU frequency to process 1bit message. For a 100Gbps link, the TCP protocol stack will consume about 100GHz CPU resources, which will bring a huge burden to service providers. In addition, when the TCP protocol sends and receives messages, the kernel needs to perform multiple context switches and data copies. Each context switch takes about 5 to 10 microseconds, resulting in a processing delay of tens of microseconds overall, which becomes a performance bottleneck.

[0005] In this context, Remote Direct Memory Access (RDMA) technology has been introduced into data centers. RDMA can directly read and write memory without the intervention of the operating systems of the two communicating parties. Its characteristics of bypassing the kernel, zero load, and zero copy give it the advantages of high throughput and low latency, making it a key technology for improving data center network performance. RDMA technology can reduce the processing latency of the protocol stack to 1 millisecond and reduce the CPU occupancy rate to 5% in 40Gbps high-speed transmission, effectively overcoming the performance defects of the TCP protocol stack at high link rates.

[0006] There are currently three main hardware implementations of RDMA, namely Infiniband, iWARP and RoCE. Among them, RoCE can perform RDMA communication in traditional Ethernet, and only requires the use of special network cards, which greatly reduces the deployment cost. Currently, the RoCE protocol includes two implementation versions, RoCEv1 and RoCEv2. In terms of flow control, a priority-based flow control mechanism is adopted, and the Go-Back-N retransmission mechanism is integrated inside the network card. RoCEv1 is implemented based on the Ethernet link layer and only allows communication between two hosts in the same broadcast domain. RoCEv2 is implemented based on the UDP layer of Ethernet. The main difference from RoCEv1 is that it can use UDP and IP network layers, which enables RoCEv2 devices to support IP routing and support communication between hosts in different broadcast domains.

[0007] SoftRoCE is a software implementation of RoCEv2, which is compatible with common devices without RDMA network cards. It enables RDMA services to be deployed on devices that lack dedicated network cards, greatly reducing the deployment cost of RDMA networks. SoftRoCE is suitable for any Ethernet environment and does not rely on NIC, Switch, L2QoS and other support. It is mainly used for testing, application development, and environments that lack RDMA hardware but need to deploy RDMA applications. Its implementation can be divided into two parts: coupling with the API provided by libibverbs in user space and coupling with the TCP / IP protocol stack through the driver in kernel space. Users transmit RoCE data for virtual RDMA devices through the UDP tunnel of an Ethernet card. Summary of the invention

[0008] The purpose of the present invention is to provide a long-distance data transmission method based on the RDMA protocol to solve the problems raised in the above background technology.

[0009] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0010] A long-distance data transmission method based on the RDMA protocol, the method comprising:

[0011] S100, when the file transmission starts, the sending end establishes a reliable connection with the receiving end through the RDMACM mode, transmits the relevant information of the file, and exchanges the key parameters required for the subsequent communication;

[0012] S200, the receiving end designs an adaptive address buffer pool according to the file size and available memory capacity, and designs an appropriate sliding window size according to the network bandwidth, packet loss rate and delay to optimize transmission performance;

[0013] S300, the receiving end and the sending end create necessary resources before RDMA transmission, including Queue Pair, Protection Domain, Memory Region and Completion Queue, and then start the file transfer;

[0014] S400, design a batch processing mechanism for completion queue elements to merge multiple completion information and further optimize performance;

[0015] S500, the sending end reads the file content and divides it into blocks, requests an available address from the address pool of the receiving end, writes the file blocks into the designated memory of the receiving end through RDMAWRITE, and after the writing is completed, writes the memory into the designated location of the hard disk through RDMACM communication, and reuses the address;

[0016] S600: After the file transfer is completed, the created related resources are released.

[0017] Preferably, in S200, the address buffer pool adapted to the file size and available memory capacity is designed, including:

[0018] S201-1, determining the capacity of available memory, and obtaining the available capacity of the memory through memory specifications or performance tests;

[0019] S201-2, when the file size is less than or equal to the available memory capacity / 4, allocate an address buffer pool equivalent to the file size;

[0020] S201-3. When the file size is greater than the available memory capacity / 4, an address buffer pool of the available memory capacity / 4 is used.

[0021] Preferably, in S200, a suitable sliding window size is designed according to the bandwidth, packet loss rate and delay of the network, including:

[0022] S202-1. Determine the bandwidth, packet loss rate, and delay of the network, and obtain the bandwidth, packet loss rate, and delay through performance testing;

[0023] S202-2. Since the network bandwidth determines the maximum throughput, the delay determines the network response speed and user experience, and the packet loss rate determines the network reliability, set the appropriate sliding window size according to the application scenario and system resources:

[0024] Set the delay to RTT (unit: ms), the packet loss rate to P, and the bandwidth to B (unit: bps);

[0025] According to the delay and packet loss rate, the window size is set as:

[0026] If bandwidth B is considered, the window size should meet the network throughput requirement as follows:

[0027] Based on this, combined with bandwidth, packet loss rate and delay, the sliding window size is obtained as

[0028] S202-3. Formulate strategies based on different scenarios adjusted , real-time monitoring of packet loss rate and delay, dynamic adjustment of sliding window size, balance of throughput and stability: W adjusted =W×(1-P);

[0029] When P>0.01 is in a high packet loss rate scenario, reduce W to avoid congestion;

[0030] In a high-latency scenario where RTT>100ms, balance the relationship between bandwidth and latency according to the formula and increase W.

[0031] Preferably, the batch processing mechanism for completing the queue elements is designed in S400, including:

[0032] S401, design a fixed value CQ_MAX_SIZE, and when the number of entries in the CQ reaches the set value, process it, otherwise return the information that the CQ is empty;

[0033] S402. Design a timeout value CQ_TIME_THRESHOLD. When no CQE is generated for a period of time, it will actively obtain CQEs less than CQ_MAX_SIZE in the CQ.

[0034] S403, the processed CQE is used to notify the server-side memory to write to the hard disk, so CQ_MAX_SIZE is determined by the block size and the address buffer pool size, and CQ_TIME_THRESHOLD is determined by the network bandwidth and the file size. The specific formula is as follows:

[0035] Assume the block size is P, the address buffer pool size is M, then CQ_MAX_SIZE = M / P*10;

[0036] Assume that the network bandwidth is B and the file size is F, then CQ_TIME_THRESHOLD = F / B*5.

[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the above-mentioned long-distance data transmission method based on the RDMA protocol.

[0038] A computer device comprises a memory, a processor and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps in the above-mentioned long-distance data transmission method based on the RDMA protocol are implemented.

[0039] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0040] The present invention designs an efficient address buffer pool at the receiving end, fully reuses memory resources, and flexibly adjusts the sliding window according to the network packet loss rate and delay conditions to reduce the impact of the network environment on transmission efficiency; in addition, the design of batch completion queue elements reduces CPU utilization, thereby significantly improving the speed and efficiency of RDMA data transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0042] Figure 1 It is a flow chart of a long-distance data transmission method based on the RDMA protocol of the present invention;

[0043] Figure 2 It is a RDMACM link building process of a long-distance data transmission method based on the RDMA protocol of the present invention;

[0044] Figure 3 The invention discloses an RDMA resource creation process of a long-distance data transmission method based on an RDMA protocol;

[0045] Figure 4 The invention discloses an RDMA resource release process of a long-distance data transmission method based on an RDMA protocol. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] See also Figure 1-Figure 4 , the present invention provides a technical solution:

[0048] Embodiment 1:

[0049] S100, when the file transmission starts, the sending end establishes a reliable connection with the receiving end through the RDMACM mode, transmits the relevant information of the file, and exchanges the key parameters required for the subsequent communication;

[0050] The specific steps of using the RDMACM method to establish a reliable connection with the receiving end in step S100 are as follows:

[0051] like Figure 2 As shown, the sender process is as follows:

[0052] 1. Create an event channel

[0053] Open a channel for reporting communication events. Asynchronous events will be reported to users through the event channel. The corresponding method is

[0054] struct rdma_event_channel*rdma_create_event_channel(void)

[0055] 2. Create a connection ID

[0056] Create an identifier for tracking communication information. The corresponding method is

[0057] int rdma_create_id(struct rdma_event_channel*channel,struct rdma_cm_id**id,void*context,enum rdma_port_space ps)

[0058] 3. Binding address

[0059] Associate the source address with the rdma_cm_id. The corresponding method is

[0060] int rdma_bind_addr(struct rdma_cm_id*id,struct sockaddr*addr)

[0061] 4. Create QP

[0062] Allocate the QP associated with the specified rdma_cm_id and convert it for transmission and reception. The corresponding method is

[0063] int rdma_create_qp(struct rdma_cm_id*id,struct ibv_pd*pd,struct ibv_qp_init_attr*qp_init_attr)

[0064] 5. Create a Listener and return the port / address

[0065] Initialize the Listener for incoming connection requests or packet service lookups. The corresponding method is

[0066] int rdma_listen(struct rdma_cm_id*id,int backlog)

[0067] 6. Waiting for connection request

[0068] Retrieve communication events until the RDMA_CM_EVENT_CONNECT_REQUEST event is received. The corresponding method is

[0069] int rdma_get_cm_event(struct rdma_event_channel*channel,struct rdma_cm_event**event)

[0070] 7. Accept the connection request

[0071] Wait for the receiver to establish a connection. The corresponding method is

[0072] int rdma_accept(struct rdma_cm_id*id,struct rdma_conn_param*conn_param)

[0073] 8. Wait for the connection to be established

[0074] Retrieve communication events until the RDMA_CM_EVENT_CONNECT_REQUEST event is received and the connection establishment is complete;

[0075] The receiving process is as follows:

[0076] 1. Create an event channel

[0077] Same as sender

[0078] 2. Create a connection ID

[0079] Same as sender

[0080] 3. Binding address

[0081] Resolve the destination address and optional source address from the IP address to an RDMA address. If successful, the specified rdma_cm_id will be bound to the local device. The corresponding method is

[0082] int rdma_resolve_addr(struct rdma_cm_id*id,struct sockaddr*src_addr,struct sockaddr*dst_addr,int timeout_ms)

[0083] 4. Wait for address resolution

[0084] Retrieve communication events until the RDMA_CM_EVENT_ADDR_RESOLVED event is received, the address resolution is completed, and the RDMA route pointing to the target address is resolved to establish a connection. The corresponding method is

[0085] int rdma_resolve_route(struct rdma_cm_id*id,int timeout_ms)

[0086] 5. Wait for routing resolution

[0087] Retrieve communication events until the RDMA_CM_EVENT_ROUTE_RESOLVED event is received and the route resolution is completed

[0088] 6. Create QP

[0089] Same as the sender

[0090] 7. Establish a connection

[0091] Request to establish a connection with the sender. The corresponding method is

[0092] int rdma_connect(struct rdma_cm_id*id,struct rdma_conn_param*conn_param)

[0093] 8. Wait for the connection to be established

[0094] Same as the sender.

[0095] S200, the receiving end designs an adaptive address buffer pool according to the file size and available memory capacity, and designs an appropriate sliding window size according to the network bandwidth, packet loss rate and delay to optimize transmission performance;

[0096] The specific steps for designing the address buffer pool size in step S200 are as follows:

[0097] S201-1. It is necessary to determine the capacity of the available memory, and obtain the available capacity of the memory through memory specifications or performance tests;

[0098] S201-2, when the file size is less than or equal to the available memory capacity / 4, an address buffer pool equivalent to the file size is allocated;

[0099] S201-3. When the file size is greater than the available memory capacity / 4, an address buffer pool of the available memory capacity / 4 is used.

[0100] Among them, in step S2, the specific steps of designing a suitable sliding window size according to the packet loss rate and delay of the network are as follows:

[0101] S202-1. It is necessary to determine the bandwidth, packet loss rate and delay of the network, and obtain the bandwidth, packet loss rate and delay through performance testing;

[0102] S202-2. Since the network bandwidth determines the maximum throughput, the delay determines the network response speed and user experience, and the packet loss rate determines the network reliability, the appropriate sliding window size can be set according to the application scenario and system resource conditions;

[0103] S202-3, let the delay be R (unit: ms), the packet loss rate be P, and the bandwidth be B (unit: bps)

[0104] According to the delay and packet loss rate, the window size is set to:

[0105] If bandwidth B is considered, the window size should meet the network throughput requirement as follows:

[0106] Based on this, combined with bandwidth, packet loss rate and delay, the sliding window size is obtained as

[0107] S202-3. Formulate strategies based on different scenarios adjusted , real-time monitoring of packet loss rate and delay, dynamic adjustment of sliding window size, balance of throughput and stability: W adjusted =W×(1-P);

[0108] When P>0.01 is in a high packet loss rate scenario, reduce W to avoid congestion;

[0109] In a high-latency scenario where RTT>100ms, balance the relationship between bandwidth and latency according to the formula and increase W.

[0110] S300, the receiving end and the sending end create necessary resources before RDMA transmission, including Queue Pair, Protection Domain, Memory Region and Completion Queue, and then start the file transfer;

[0111] like Figure 3As shown, in step S3, the receiving end and the sending end create necessary resources before RDMA transmission. The specific steps are as follows:

[0112] S301. Allocate PD:

[0113] Protection domain (PD) is used to group and isolate resources (such as queue pairs, memory areas, etc.) to ensure that operations between different queue pairs and memory areas do not interfere with each other. A protection domain is allocated to the device by calling the ibv_alloc_pd() method.

[0114] S302. Create CQ

[0115] The Completion Queue (CQ) is a core component that is mainly used to manage the completion notification of RDMA operations. The Completion Queue records the completion status of RDMA operations (such as sending, receiving, writing, etc.). When an RDMA operation is completed, a Completion Queue Entry (CQE) is written to the associated CQ. The program queries the CQE to confirm whether the operation is successfully completed and creates a completion queue by calling the ibv_create_cq() method.

[0116] S303, create QP

[0117] Queue Pair (QP) is one of the core components of RDMA communication, responsible for the management and execution of data sending, receiving and related operations. RDMA communication can be abstracted as communication between QPs, and queue pairs are created by calling the ibv_create_qp() method.

[0118] S304, Register MR

[0119] Memory Region (MR) is a key concept in RDMA operation. MR is a bridge between applications and RDMA network hardware for direct memory access. It provides safe, efficient and direct access to memory. Registered memory RDMA stipulates that the memory where the data is located must be registered before transmitting data. The memory region is registered by calling the ibv_reg_mr() method.

[0120] S305. Create threads and related resources

[0121] The sending end creates a thread for reading files, a sending thread for RDMA WRITE, and a completion queue thread. The receiving end creates a thread for writing files and an address pool.

[0122] S400, design a batch processing mechanism for completion queue elements to merge multiple completion information and further optimize performance;

[0123] The design scheme of the batch processing mechanism based on the completion queue elements in step S400 is as follows:

[0124] After each packet is sent, the network card sends a work completion message to the kernel. This is relatively inefficient and the system needs to spend a lot of time polling for work completion information. If the completion information can be sent to the kernel at one time after a fixed number of packets are sent, the system's work efficiency will be greatly improved.

[0125] S401, a fixed value CQ_MAX_SIZE is designed. When the number of entries in the CQ reaches the set value, processing is performed, otherwise the information that the CQ is empty is returned;

[0126] S402. A timeout value CQ_TIME_THRESHOLD is designed. When no CQE is generated for a period of time, the system will actively obtain CQEs less than CQ_MAX_SIZE in the CQ.

[0127] S403, the processed CQE is mainly used to notify the server-side memory to write to the hard disk, so CQ_MAX_SIZE is determined by the block size and the address buffer pool size, and CQ_TIME_THRESHOLD is determined by the network bandwidth and the file size. The specific formula is as follows:

[0128] Assume the block size is P, and the address buffer pool size is MCQ_MAX_SIZE = M / P*10

[0129] Assume the network bandwidth is B, and the file size is FCQ_TIME_THRESHOLD = F / B*5

[0130]

[0131] S500, the sending end reads the file content and divides it into blocks, requests an available address from the address pool of the receiving end, writes the file blocks into the designated memory of the receiving end through RDMAWRITE, and after the writing is completed, writes the memory into the designated location of the hard disk through RDMACM communication, and reuses the address;

[0132] The transmission design scheme in step S500 is as follows:

[0133] S501, the reading thread reads file information in blocks;

[0134] S502, requesting address information from the receiving end address pool;

[0135] S503, writing the file information into the receiving end memory via RDMA WRITE;

[0136] S504, the completion queue receives the completion information of the receiving end writing into the memory, and notifies the receiving end to write into the hard disk;

[0137] S505, write to the specified location on the hard disk, and add the address to the address pool again.

[0138] S600: After the file transfer is completed, the created related resources are released.

[0139] like Figure 4 As shown, the steps of releasing resources in S600 are as follows:

[0140] S601, clean up work requests and buffers, complete or cancel all unfinished requests in the send queue and receive queue, and release the data buffer allocated to the RDMA operation;

[0141] S602, cancel memory registration, use ibv_dereg_mr() to release the memory region (Memory Region, MR) registered by ibv_reg_mr();

[0142] S603, release the completion queue (CQ), call ibv_destroy_cq() to release the completion queue resources;

[0143] S604, destroy the queue pair (QP), and use ibv_destroy_qp() to release the queue pair (Queue Pair, QP);

[0144] S605, release the protection domain (PD), use ibv_dealloc_pd() to release the protection domain;

[0145] S606, release the completion event channel, and use ibv_destroy_comp_channel() to release the completion event channel;

[0146] S607: Close the device context. Use ibv_close_device() to close the RDMA device context.

[0147] Embodiment 2:

[0148] The computer-readable storage medium of this embodiment stores a computer program, which, when executed by a processor, implements the steps of a long-distance data transmission method based on the RDMA protocol in Embodiment 1.

[0149] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.

[0150] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0151] Embodiment 3:

[0152] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of a long-distance data transmission method based on the RDMA protocol in Embodiment 1 are implemented.

[0153] In this embodiment, the processor may be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, readily available programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0154] Those skilled in the art will appreciate that the disclosed content of the embodiments may be provided as methods, systems, or computer program products. Therefore, the present solution may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present solution may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.

[0155] The present solution is described with reference to the method according to the embodiment of the present solution and the flowchart and / or block diagram of the computer program product. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions; these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 one or more processes and / or methods Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0156] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 one or more processes and / or methods Figure 1 A function specified in one or more boxes.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 one or more processes and / or methods Figure 1 The steps for the functions specified in one or more boxes.

[0158] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0159] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A long-distance data transmission method based on the RDMA protocol, characterized in that: The method comprises: S100, when the file transmission starts, the sending end establishes a reliable connection with the receiving end through the RDMACM mode, transmits the relevant information of the file, and exchanges the key parameters required for the subsequent communication; S200, the receiving end designs an adaptive address buffer pool according to the file size and available memory capacity, and designs an appropriate sliding window size according to the network bandwidth, packet loss rate and delay to optimize transmission performance; S300, the receiving end and the sending end create necessary resources before RDMA transmission, including Queue Pair, ProtectionDomain, Memory Region and Completion Queue, and then start the file transfer; S400, design a batch processing mechanism for completion queue elements to merge multiple completion information and further optimize performance; S500, the sending end reads the file content and divides it into blocks, requests an available address from the address pool of the receiving end, writes the file blocks into the designated memory of the receiving end through RDMA WRITE, and after the writing is completed, writes the memory into the designated location of the hard disk through RDMACM communication, and reuses the address; S600: After the file transfer is completed, the created related resources are released.

2. A long-distance data transmission method based on the RDMA protocol as claimed in claim 1, characterized in that: The address buffer pool adapted to the file size and available memory capacity is designed in S200, including: S201-1, determining the capacity of available memory, and obtaining the available capacity of the memory through memory specifications or performance testing; S201-2, when the file size is less than or equal to the available memory capacity / 4, allocate an address buffer pool equivalent to the file size; S201-3. When the file size is greater than the available memory capacity / 4, an address buffer pool of the available memory capacity / 4 is used.

3. A long-distance data transmission method based on the RDMA protocol as claimed in claim 1, characterized in that: In S200, a suitable sliding window size is designed according to the bandwidth, packet loss rate and delay of the network, including: S202-1. Determine the bandwidth, packet loss rate, and delay of the network, and obtain the bandwidth, packet loss rate, and delay through performance testing; S202-2. Set an appropriate sliding window size according to the application scenario and system resources: Set the delay to RTT, the packet loss rate to P, and the bandwidth to B; According to the delay and packet loss rate, the window size is set as: If bandwidth B is considered, the window size should meet the network throughput requirement as follows: Based on this, combined with bandwidth, packet loss rate and delay, the sliding window size is obtained as S202-3. Formulate strategies based on different scenarios adjusted , real-time monitoring of packet loss rate and delay, dynamic adjustment of sliding window size, balance of throughput and stability: W adjusted =W×(1-P); When P>0.01 is in a high packet loss rate scenario, reduce W to avoid congestion; In a high-latency scenario where RTT>100ms, balance the relationship between bandwidth and latency according to the formula and increase W.

4. A long-distance data transmission method based on the RDMA protocol as claimed in claim 1, characterized in that: The batch processing mechanism for completing the queue elements in S400 includes: S401, design a fixed value CQ_MAX_SIZE, and when the number of entries in the CQ reaches the set value, process it, otherwise return the information that the CQ is empty; S402. Design a timeout value CQ_TIME_THRESHOLD. When no CQE is generated for a period of time, it will actively obtain CQEs less than CQ_MAX_SIZE in the CQ. S403, the processed CQE is used to notify the server-side memory to write to the hard disk, so CQ_MAX_SIZE is determined by the block size and the address buffer pool size, and CQ_TIME_THRESHOLD is determined by the network bandwidth and the file size. The specific formula is as follows: Assume the block size is P, the address buffer pool size is M, then CQ_MAX_SIZE = M / P*10; Assume that the network bandwidth is B and the file size is F, then CQ_TIME_THRESHOLD = F / B*5.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a long-distance data transmission method based on the RDMA protocol as described in any one of claims 1 to 4 are implemented.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps in the long-distance data transmission method based on the RDMA protocol as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • RDMA network data transmission method based on erasure code

    CN108631947A

  • A lightweight asynchronous message implementation method based on multi-sliding window concurrency

    CN109067506A

  • Data migration method and data node

    CN109144972A

  • Remote memory direct access method and related device

    CN113608686A

  • TTE network communication method based on RDMA

    CN116089331A