A lightweight method for the key I / O path of a distributed storage system

By extending the method of registering memory attributes and generating parallel subrequests in a distributed storage system, the problems of RDMA memory management redundancy and data copying are solved, and the effect of reducing request processing delay and improving system performance is achieved.

CN115509452BActive Publication Date: 2025-05-30SHANDONG HAILIANG INFORMATION TECH RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211194147.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-05-30
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

In distributed storage systems, the prior art is unable to effectively manage RDMA memory, resulting in repeated organization management and data transfer on system key I/O paths, increasing request processing delays and system performance degradation.

Method used

By extending RDMA registration memory attributes in the standard protocol network layer, increasing remote read and write permissions, and parsing and reconstructing private protocol communication requests at the virtual block device layer, parallel sub-requests are generated, and redundant management and data copying are reduced.

Benefits of technology

Eliminates redundant RDMA memory management on key I/O paths, reduces the number of requests and releases of memory resources, reduces the latency of request processing, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115509452B_ABST
    Figure CN115509452B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight method for the key I / O path of a distributed storage system. Compared with the existing distributed storage system, in the present invention, the standard protocol network layer is no longer directly connected to the physical storage device, but is connected to the physical storage devices located at multiple different remote nodes through the virtual block device layer. At the same time, based on this architectural change, the present invention extends the RDMA registered memory attributes in the standard protocol network layer to give it remote read and write permissions. Since the memory management module in the NVMe-oF standard protocol Target is reused, the redundant management of the RDMA memory on the key I / O path is eliminated, the number of applications and releases of memory resources is reduced, and the processing delay of requests is reduced. It solves the problem that the communication queue in the private protocol communication connection and the RDMA memory in the NVMe-oF standard protocol are in different protection domains and thus cannot access the memory, eliminates the extra data copy in the key path, and improves the overall performance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer distributed storage, and more specifically, relates to a method for lightweighting the key I / O path of a distributed storage system. Background Art

[0002] RDMA technology is a technology developed to solve problems such as insufficient bandwidth and high latency during network transmission. Compared with traditional TCP / IP networks, RDMA technology bypasses the operating system kernel and directly uses the system bus to complete the data exchange between memory and the network card. At this time, the RDMA operation only requires the direct participation of the CPU when sending control commands, and the direct memory access controller completes the data transmission. Therefore, its CPU utilization rate is very low. In the fields of distributed storage and high-performance computing, RDMA technology has been widely applied to the network communication module of the system.

[0003] The proposal of NVMe-oF is an extension of the NVMe protocol for Ethernet and Fibre Channel, etc. It uses a message-based model to transmit requests and process responses between the host side and the target storage device through the network. The original purpose of its design is to replace PCIe to extend the communication distance between the NVMe host and the NVMe storage subsystem at the cost of very little performance loss, and to realize cross-network block storage services with high performance, high resource utilization, high scalability, and fault isolation.

[0004] NVMe-oF usually uses the RDMA network as the transport layer protocol to give full play to its high-speed interconnection advantages. In this type of distributed storage system, it is necessary to pre-register the RDMA memory to reduce system latency. Therefore, the system needs to manage this part of the RDMA memory. In addition, when data is sent from the user layer to the network layer, it also needs to be copied into the RDMA memory before it can be sent.

[0005] Since the control command flow is unidirectional, even if the private protocol communication connection of the backend distributed storage system uses an RDMA network, it cannot interact with the network layer of the NVMe-oF protocol. The virtual block device layer can only passively receive information from the network layer. After packaging and attribute removal, the information from the network layer seen by the virtual block device layer has lost the key RDMA attributes, and only ordinary memory information is exposed to the service interface of the distributed storage system. The above working method is because at the beginning of the NVMe-oF protocol design, the data in this memory was directly operated by the DMA controller, and the controller only needed the local read and write permission lkey. Therefore, when the request is passed to the virtual block device layer, only the existing permission information will be hidden. In addition, even if this part of the permission information can be obtained, remote access cannot be completed. This will cause the private protocol communication connection part to reorganize and manage the necessary RDMA memory again on the critical I / O path of the system and perform data migration. These redundant operations will lead to an increase in request processing latency and a decrease in system performance. Summary of the Invention

[0006] Aiming at the defects of the prior art, the purpose of the present invention is to provide a lightweight method for the critical I / O path of a distributed storage system, aiming to reduce the system request processing delay and improve the system performance.

[0007] To achieve the above purpose, in a first aspect, the present invention provides a lightweight method for the critical I / O path of a distributed storage system. The distributed storage system includes a standard protocol network layer, a virtual block device layer, and a distributed storage node connected in sequence. The virtual block device layer interacts with the distributed storage node through a private protocol communication connection. The method includes the following steps:

[0008] S1, before the original request is passed from the standard protocol network layer to the virtual block device layer, expand and store the check information magic_number of the original request, the remote access permission rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory are located in the standard protocol network layer, and the context information device_ctx of the RDMA network device;

[0009] S2. After the original request is passed to the virtual block device layer, the virtual block device layer parses the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the target data storage address offset offset, the target data length len, magic_number, rkeys, pd, and device_ctx carried in the original request. When it is determined that the original request is a normal read / write request through magic_number, a private protocol communication request is reconstructed based on the parsed iovs, iov_cnt, offset, len, magic_number, rkeys, pd, and device_ctx. Then, multiple parallel sub-requests are generated in combination with the offset in the original request;

[0010] S3. Send each sub-request to the distributed storage node through the private protocol communication connection;

[0011] S4. The distributed storage node receives and parses each sub-request, and initiates remote read / write through the private protocol communication connection.

[0012] Further, in S2, after the original request is passed to the virtual block device layer, only the starting address of the original request is encapsulated in the private protocol communication request;

[0013] When generating multiple parallel sub-requests, first parse the starting address of the original request from the encapsulated private protocol communication request, and then directly obtain the associated field information by combining the starting address with the offset.

[0014] Further, before S3, it also includes:

[0015] S3′. Determine whether the private protocol communication connection has been initialized. If so, execute S3; if not, initialize the private protocol communication connection using pd and device_ctx, and then execute S3.

[0016] Further, when initializing the private protocol communication connection, two-phase initialization is adopted with the second-phase delayed initialization. The first-phase initialization is completed at system startup, and the second-phase initialization is completed by parsing pd and device_ctx in the private protocol communication request.

[0017] Further, in S4, if the sub-request is a read request, read data from the local SSD into the local memory, and use the RDMA write operation in combination with rkeys to send the data to the memory area indicated by the sub-request; if the sub-request is a write request, use the RDMA read operation in combination with rkeys to read the data into the local memory, and then write it to the local SSD.

[0018] In a second aspect, the present invention provides a distributed storage system, which includes a standard protocol network layer, a virtual block device layer, and distributed storage nodes that are connected in sequence. The virtual block device layer communicates and interacts with the distributed storage nodes through a private protocol connection;

[0019] The standard protocol network layer is used to expand and store the check information magic_number of the original request, the remote access permissions rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory in the standard protocol network layer are located, and the context information device_ctx of the RDMA network device before the original request is transmitted from the standard protocol network layer to the virtual block device layer;

[0020] The virtual block device layer is used to parse the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the target data storage address offset offset, the target data length len, magic_number, rkeys, pd, and device_ctx carried in the original request after the original request is transmitted to the virtual block device layer. When it is determined that the original request is a normal read / write request through magic_number, a private protocol communication request is reconstructed based on the parsed iovs, iov_cnt, offset, len, magic_number, rkeys, pd, and device_ctx; and multiple parallel sub-requests are generated in combination with the offset in the original request;

[0021] The virtual block device layer is further used to send each sub-request to the distributed storage node through the private protocol communication connection;

[0022] The distributed storage node is used to receive and parse each sub-request, and initiate remote read / write through the private protocol communication connection.

[0023] In a third aspect, the present invention provides a computer-readable storage medium, which includes a stored computer program. When the computer program is run by a processor, it controls the device where the computer-readable storage medium is located to execute the method for lightweighting the key I / O path of the distributed storage system as described in the first aspect.

[0024] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0025] Compared with existing distributed storage systems, the standard protocol network layer often directly connects to physical storage devices. In the present invention, the standard protocol network layer no longer directly connects to physical storage devices, but instead connects to physical storage devices located at multiple different remote nodes through a virtual block device layer. At the same time, based on this architectural change, the present invention extends the RDMA registered memory attributes in the standard protocol network layer to give it remote read and write permissions. Since the memory management module in the NVMe-oF standard protocol Target is reused, redundant management of RDMA memory on the critical I / O path is eliminated, the number of applications and releases of memory resources is reduced, and the processing latency of requests is lowered. The problem that the communication queue in the private protocol communication connection and the RDMA memory in the NVMe-oF standard protocol are in different protection domains and thus cannot access memory is solved, and the extra data copy in the critical path is eliminated, improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of a method for lightweighting the critical I / O path of a distributed storage system provided by an embodiment of the present invention;

[0027] Figure 2 is a data structure diagram of requests at different stages of the critical I / O path provided by an embodiment of the present invention;

[0028] Figure 3 is a flowchart of a read operation of a method for lightweighting the critical I / O path of a distributed storage system provided by an embodiment of the present invention;

[0029] Figure 4 is a flowchart of a write operation of a method for lightweighting the critical I / O path of a distributed storage system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0031] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0032] Refer to Figure 1 and in combination with Figures 2 to 4 , the present invention provides a method for lightweighting the critical I / O path of a distributed storage system, and the method includes operations S1 - S4.

[0033] Operation S1: Before the original request is transferred from the standard protocol network layer to the virtual block device layer, expand and store the check information magic_number of the original request, the remote access permissions rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory are located in the standard protocol network layer, and the context information device_ctx of the RDMA network device.

[0034] In this embodiment, the RDMA registered memory attributes in the standard protocol network layer are expanded. Based on the local read and write permission lkey, the remote read and write permission rkey is added. In the original request from the NVMe-oF standard protocol network layer, four additional memory spaces are opened up to store the check information magic_number of the original request, the remote access permissions rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory are located in the standard protocol network layer, and the context information device_ctx of the RDMA network device. Among them, magic_number is used to indicate the type of the original request; rkeys are reconstructed in each of the mapped sub-requests, and each RDMA memory segment has an associated permission information rkey. At most 16 rkeys can be stored in rkeys; pd and device_ctx are used for the lazy initialization of the private protocol communication connection.

[0035] In addition, two global variables are maintained, including the status variable InitFlag and the lock variable InitLock. The former indicates whether the two-phase initialization of the private protocol communication connection has been completed, and the latter is used as a synchronization lock during the initialization of the private protocol communication connection. A predefined field access semantics is provided to allow local reverse control flow to access the above information. Specifically, InitFlag being False indicates that the second phase of the private protocol communication connection MSG_M has not been initialized, and InitFlag being True indicates that the private protocol communication connection has been initialized. When InitLock is in the lock state, it means that the second phase of the private protocol communication connection is being initialized. When InitLock is in the unlock state, it means that the state of the private protocol communication connection is indicated by the InitFlag variable.

[0036] Operation S2: After the original request is passed to the virtual block device layer, the virtual block device layer parses the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the target data storage address offset offset, the target data length len, magic_number, rkeys, pd, and device_ctx carried in the original request. When it is determined that the original request is a normal read / write request through magic_number, a private protocol communication request is reconstructed based on the parsed iovs, iov_cnt, offset, len, magic_number, rkeys, pd, and device_ctx; then multiple parallel sub-requests are generated in combination with the offset in the original request.

[0037] As a preferred method, after the original request is passed to the virtual block device layer, only the starting address of the original request is encapsulated in the private protocol communication request; when generating multiple parallel sub-requests, first parse the starting address of the original request from the encapsulated private protocol communication request, and then directly obtain the associated field information by combining the starting address with the offset.

[0038] Operation S3: Send each sub-request to the distributed storage node through the private protocol communication connection.

[0039] In this embodiment, before S3, it also includes:

[0040] S3′: Determine whether the private protocol communication connection has been initialized. If it has, execute S3; if not, initialize the private protocol communication connection using pd and device_ctx, and then execute S3.

[0041] Furthermore, when initializing the private protocol communication connection, two-stage initialization is adopted with the second stage being delayed initialization. The first stage of initialization is completed when the system starts, and the second stage of initialization is completed by parsing pd and device_ctx in the private protocol communication request.

[0042] Specifically, S3′ includes:

[0043] S31′: Check the status variable to determine whether the private protocol communication connection has completed two-stage initialization using pd and device_ctx. If not, execute S32′; otherwise, execute S33′.

[0044] S32': Try to obtain the lock variable. After successfully locking, check the status variable again. If the check passes again, use pd and device_ctx in the private protocol communication request to complete the establishment of the communication connection, application and configuration of RDMA resources, etc. Set the status variable to avoid initialization replay operations. If the second check fails, execute S33'.

[0045] S33': Traverse the RDMA memory segments in the private protocol communication request level by level. Combine the target data address and the actual data length to map each RDMA memory segment to the actual storage node. Reconstruct the mapped sub-request by combining the remote access permissions rkeys of the associated RDMA memory, and send the sub-request to the actual storage node using the private protocol communication connection.

[0046] Operation S4, the distributed storage node receives and parses each sub-request, and initiates remote read and write through the private protocol communication connection.

[0047] In this embodiment, if the sub-request is a read request, read data from the local SSD into the local memory, and use the RDMA write operation combined with rkeys to send the data to the memory area indicated by the sub-request; if the sub-request is a write request, use the RDMA read operation combined with rkeys to read the data into the local memory, and then write it to the local SSD.

[0048] Next, the present invention will be further described in detail in combination with specific read and write operations.

[0049] As Figure 2 shown. In addition to having the check information magic_number, the remote access permissions rkeys of the associated RDMA memory, the communication queue pair and the protection domain pd where the registered memory is located in the standard protocol network layer, and the context information device_ctx of the RDMA network device, the reconstructed private protocol communication request flame_request also has the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the offset of the target data storage address offset, the target data length len, and the data length cpl_len that has been read and written for the current request. The private protocol communication request sub-request chunk_request has information such as the target allocation identifier id, the offset offset_inchunk of the target data in the target slice, the length len_inchunk in the target slice, the remote buffer address remote_addr, and the remote access permission rkey corresponding to the buffer.

[0050] The verification information magic_number, protection domain pd, and RDMA network device context information each occupy 8 bytes. The remote access permissions rkeys can accommodate a maximum of 16 memory segment permission identifiers, and each remote permission identifier occupies 4 bytes, totaling 64 bytes. Additionally, the status variable InitFlag and the lock variable InitLock are maintained. InitFlag also has a related variable ChunkSize, which represents the shard size when the virtual volume is mapped to the actual storage node. If magic_number is equal to the globally predefined value MAGIC_NUMBER (0xABCD1234), it indicates that the request is a normal read / write request from the standard protocol network layer; otherwise, it indicates that the request is a detection request during the construction of the backend virtual device.

[0051] In this embodiment, ChunkSize is 1GB, and the NVMe-oF protocol data transfer block is set to a maximum of 4KB. In the original request, iovs[0].base is NULL, iovs[0].iov_len is 0, iovs[1].base is NULL, iovs[1].iov_len is 0, and iovcnt is 0. pd is NULL, and device_ctx is NULL. offset is 0x3FFFF000, and len is 0x2000. magic_number is 0xABCD1234, rkeys[0] is 0, and rkeys[1] is 0. As Figure 3 shown, the process of a data read operation is as follows:

[0052] (1) When registering RDMA memory, add the IBV_ACCESS_REMOTE_WRITE and IBV_ACCESS_REMOTE_READ flags to extend it to remotely access RDMA memory. When constructing the network layer of the original request NVMe-oF standard protocol, the target data length is 0x2000. Therefore, two RDMA memory segments with a length of 0x1000 are constructed. After applying for buffer space at this time, iovs[0].base is 0x7FFD9E9C8000, iovs[0].iov_len is 0x1000, iovs[1].base is 0x7FFD9E9D4000, iovs[0].iov_len is 0x1000, and iovcnt is 2. pd is 0xBC63B0, device_ctx is 0xBDB7B0, rkeys[0] is 0xC3389, and rkeys[1] is 0xC3389. Reconstruct the RDMA memory segment information iovs, the remote access permission information rkeys, the RDMA device context information device_ctx, and the protection domain information pd where the RDMA memory segment is located with the target data address offset and the target data length len into a private protocol communication request flame_request, and send it to the virtual block device layer for processing.

[0053] (2) After receiving the private protocol communication request flame_request, the virtual block device layer obtains the magic_number field information based on the private protocol communication request flame_request. At this time, if magic_number is equal to MAGIC_NUMBER (0xABCD1234), it means that this request is a normal read / write request, and set the currently read data length cpl_len to 0.

[0054] (3) If the status variable InitFlag is False during the first check, attempt to lock the lock variable InitLock. If the lock variable successfully changes to the lock state and the status variable InitFlag is still False during the second check, use the pd and device_ctx of the private protocol communication request to perform the second-phase initialization of the private protocol communication connection, that is, construct a communication queue pair based on the protection domain pd0xBC63B0 and the RDMA network device context device_ctx 0xBDB7B0. After the initialization operation is completed, set the status variable InitFlag to True; if the status variable InitFlag becomes True during the second check, skip the initialization operation and release the lock variable to the unlock state. If the lock variable fails to be locked, the current thread is suspended until the lock is successful. At this time, the check of the status variable InitFlag must be True, so the initialization operation is skipped. If the status variable InitFlag is True during the first check, directly skip the initialization operation.

[0055] (4) After double-checking, process the RDMA memory segment information one by one. The buffer pointed to by iovs[0].base has a size of 0x1000, the target address is 0x3FFFF000, which crosses the boundary of the Chunk shard. Generate sub-request 1, its id is 0, offset_inchunk is 0x3FFFF000, len_inchunk is 0x1000, the remote address remote_addr is 0x7FFD9E9C8000, and the remote access permission rkey is 0xC3389. Generate sub-request 2, its id is 1, offset_inchunk is 0, len_inchunk is 0x1000, the remote address remote_addr is 0x7FFD9E9D4000, and the remote access permission rkey is 0xC3389. Send the sub-requests to the target storage node in parallel for processing.

[0056] (5) The storage node that receives sub-request 1 parses the Chunk shard id information in the request, converts it to the SSD address, and reads the data into the local RDMA memory. The communication queue pair writes the data to the remote_addr 0x7FFD9E9C8000 using the remote access permission rkey 0xC3389. Since the receiver of this communication queue pair and the remote_addr registered memory are in the same protection domain pd0xBC63B0, this write operation is successful. The processing of sub-request 2 is the same as above.

[0057] (6) After the storage node writes data to remote_addr 0x7FFD9E9C8000, it triggers the callback of the private protocol communication request flame_request and updates the cpl_len field in it to 0x1000. After the storage node writes data to remote_addr 7FFD9E9D4000, it triggers the second callback of the private protocol communication request flame_request and updates the cpl_len field in it to 0x2000.

[0058] (7) When the cpl_len in the private protocol communication request flame_request is updated to 0x2000, which is equal to the target data length len, it indicates that the request is completed. At this time, a hierarchical feedback is made to the virtual block device layer, and thus a read operation is completed.

[0059] In this embodiment, the ChunkSize is 1GB, and the maximum data block transmitted by the NVMe-oF protocol is set to 128KB. In the original request, iovs[0].base is 0x7FFDA0E91000, iovs[0].iov_len is 0x2000, and iovcnt is 1. pd is NULL, and device_ctx is NULL. offset is 0x3FFFF000, and len is 0x4000. The magic_number is 0xABCD1234, and rkeys[0] is 0. As Figure 4 shown, the process of a data write operation is as follows:

[0060] (1) When registering the RDMA memory, the IBV_ACCESS_REMOTE_WRITE and IBV_ACCESS_REMOTE_READ flags are added to expand it to be remotely accessible to the RDMA memory. When constructing the original request NVMe-oF standard protocol network layer, pd is 0xBD72B0, device_ctx is 0xBEC6B0, and rkeys[0] is 0x84E83. The RDMA memory segment information iovs, the remote access permission information rkeys, the RDMA device context information device_ctx, and the protection domain information pd where the RDMA memory segment is located are reconstituted into the private protocol communication request flame_request and sent to the virtual block device layer for processing.

[0061] (2) After the virtual block device layer receives the private protocol communication request flame_request, it obtains the magic_number field information based on the private protocol communication request flame_request. At this time, if magic_number is equal to MAGIC_NUMBER (0xABCD1234), it indicates that this request is a normal read / write request, and the currently read data length cpl_len is set to 0.

[0062] (3) If the status variable InitFlag is False during the first check, then it attempts to lock the lock variable InitLock. If the lock variable successfully changes to the lock state, and the status variable InitFlag is still False during the second check, then the second-phase initialization of the private protocol communication connection is performed using the pd and device_ctx of the private protocol communication request, that is, a communication queue pair is constructed based on the protection domain pd 0xBD72B0 and the RDMA network device context device_ctx 0xBEC6B0. After the initialization operation is completed, the status variable InitFlag is set to True. If the status variable InitFlag changes to True during the second check, then the initialization operation is skipped and the lock variable is released to the unlock state. If the lock variable fails to be locked, the current thread is suspended until the lock is successful. At this time, the check of the status variable InitFlag must be True, so the initialization operation is skipped. If the status variable InitFlag is True during the first check, then the initialization operation is directly skipped.

[0063] (4) After the double-check, the RDMA memory segment information is processed one by one. The buffer size pointed to by iovs[0].base is 0x4000, the target address is 0x3FFFF000, which crosses the boundary of the Chunk shard. Sub-request 1 is generated, with its id being 0, offset_inchunk being 0x3FFFF000, len_inchunk being 0x1000, remote address remote_addr being 0x7FFDA0E91000, and remote access permission rkey being 0x84E83. Sub-request 2 is generated, with its id being 1, offset_inchunk being 0, len_inchunk being 0x3000, remote address remote_addr being 0x7FFDA0E92000, and remote access permission rkey being 0x84E83. The sub-requests are sent to the target storage node in parallel for processing.

[0064] (5) The storage node that receives Sub-request 1. The communication queue pair reads data from remote_addr 0x7FFDA0E91000 to the local RDMA memory using the remote access right rkey0xC3389. Since the receiver of this communication queue pair and the remote_addr registered memory are in the same protection domain pd0xBD72B0, this remote read operation is successful. Parse the Chunk shard id information in the request, convert it to the SSD address and write the data to the SSD. The processing of Sub-request 2 is the same as above.

[0065] (6) After the storage node writes the data corresponding to 0x7FFDA0E91000 to the SSD, it triggers the callback of the private protocol communication request flame_request and updates the cpl_len in it to 0x1000. After the storage node writes the data corresponding to 0x7FFDA0E92000 to the SSD, it triggers the second callback of the private protocol communication request flame_request and updates the cpl_len in it to 0x4000.

[0066] (7) When the cpl_len in the private protocol communication request flame_request is updated to 0x4000, which is equal to the length len of the data to be written, it means that the request is completed. At this time, a hierarchical feedback is made to the virtual block device layer, and thus a write operation is completed.

[0067] In summary, compared with the prior art, the beneficial effects of the present invention are mainly as follows: In the prior art, there is an overlap in the RDMA memory management function in the bridging of the NVMe-oF standard protocol and the private protocol communication connection. In addition, due to the fine-grained protection domain restrictions, data migration occurs on the critical I / O path. Through the incremental extension of NVMe-oF, the present invention opens up the metadata and data access tunnels for the private protocol communication connection without affecting the original working semantics. The metadata access tunnel eliminates the redundant mechanism for the private protocol communication connection to re-maintain the RDMA memory with remote access rights by adding and maintaining the remote access right rkey of the RDMA memory segment and directly mapping the discrete RDMA memory segments to discrete actual storage nodes, reduces the number of applications and releases of RDMA memory resources, and reduces the processing latency of the system. The data access tunnel solves the problem that the communication queue pair in the private protocol and the RDMA memory in the standard protocol are in different protection domains and thus cannot access the memory by maintaining the communication queue pair and the protection domain and RDMA network device context information where the registered memory is located, and by customizing the two-phase initialization of the private protocol and the second-phase double-check lazy initialization strategy, eliminates the data migration in the critical I / O path, reduces the processing latency of the system requests, and improves the system performance.

[0068] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A lightweight method for the critical I / O path of a distributed storage system, characterized in that, the distributed storage system includes a standard protocol network layer, a virtual block device layer, and distributed storage nodes connected in sequence. The virtual block device layer communicates with the distributed storage nodes through a private protocol. The method includes the following steps: S1. Before the original request is transmitted from the standard protocol network layer to the virtual block device layer, expand and store the check information magic_number of the original request, the remote access permission rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory are located in the standard protocol network layer, and the context information device_ctx of the RDMA network device; S2. After the original request is transmitted to the virtual block device layer, the virtual block device layer parses the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the target data storage address offset offset, the target data length len, magic_number, rkeys, pd, and device_ctx carried in the original request. When it is determined that the original request is a normal read / write request through magic_number, reconstruct a private protocol communication request based on the parsed iovs, iov_cnt, offset, len, magic_number, rkeys, pd, and device_ctx; then generate multiple parallel sub-requests in combination with the offset in the original request; S3. Send each sub-request to the distributed storage node through the private protocol communication connection; S4. The distributed storage node receives and parses each sub-request, and initiates remote read / write through the private protocol communication connection.

2. The lightweight method for the critical I / O path of a distributed storage system according to claim 1, characterized in that, in S2, after the original request is transmitted to the virtual block device layer, only the first address of the original request is encapsulated in the private protocol communication request; when generating multiple parallel sub-requests, first parse the first address of the original request from the encapsulated private protocol communication request, and then directly obtain the associated field information by combining the first address with the offset.

3. The lightweight method for the critical I / O path of a distributed storage system according to claim 1 or 2, characterized in that, before S3, it further includes: S3′. Judge whether the private protocol communication connection has been initialized. If so, execute S3; if not, initialize the private protocol communication connection using pd and device_ctx, and then execute S3.

4. The lightweight method for the critical I / O path of a distributed storage system according to claim 3, characterized in that, when initializing the private protocol communication connection, two-stage initialization is adopted and the second stage is delayed initialization. The first stage of initialization is completed when the system starts, and the second stage of initialization is completed by parsing pd and device_ctx in the private protocol communication request.

5. The key I / O path lightweight method for a distributed storage system according to claim 1, wherein, in S4, if the sub-request is a read request, read data from the local SSD to the local memory, and use the RDMA write operation combined with rkeys to send the data to the memory area indicated by the sub-request; if the sub-request is a write request, use the RDMA read operation combined with rkeys to read the data to the local memory, and then write it to the local SSD.

6. A distributed storage system, wherein, it includes a standard protocol network layer, a virtual block device layer, and distributed storage nodes connected in sequence. The virtual block device layer communicates with the distributed storage nodes through a private protocol; the standard protocol network layer is used to expand and store the check information magic_number of the original request, the remote access permission rkeys associated with the RDMA memory, the protection domain pd where the communication queue pair and the registered memory are located in the standard protocol network layer, and the context information device_ctx of the RDMA network device before the original request is transmitted from the standard protocol network layer to the virtual block device layer; the virtual block device layer is used to parse the RDMA memory segment information iovs, the number of RDMA memory segments iov_cnt, the target data storage address offset, the target data length len, magic_number, rkeys, pd, and device_ctx carried in the original request after the original request is transmitted to the virtual block device layer, and when it is determined that the original request is a normal read / write request through magic_number, reconstruct a private protocol communication request based on the parsed iovs, iov_cnt, offset, len, magic_number, rkeys, pd, and device_ctx; then generate multiple parallel sub-requests in combination with the offset in the original request; the virtual block device layer is further used to send each sub-request to the distributed storage node through the private protocol communication connection; the distributed storage node is used to receive and parse each sub-request, and initiate remote read / write through the private protocol communication connection.

7. A computer-readable storage medium, wherein, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the computer-readable storage medium is located to execute the key I / O path lightweight method for a distributed storage system according to any one of claims 1 to 5.