Local data read-write method and read-write device
By obtaining the virtual address of the NVMe target process and performing data read and write operations based on the address, the problem of local storage access in the NVMe-oF environment needs to be transmitted over the network, achieving the effect of reducing bandwidth consumption and delay.
Patent Information
- Application Number
- CN202510122706.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
In a dual-machine hyperconverged deployment environment based on NVMe-oF, although high-performance remote storage access is achieved, there is unnecessary bandwidth consumption and latency because data communication must be completed over the network, and even local storage access follows a complete network transmission process.
By determining the local data read or write request initiated by the NVMe-oF client, obtain the virtual address of the NVMe target process and perform data read or write operations based on the virtual address to avoid transmission over the network.
It reduces unnecessary bandwidth consumption and latency, improves the efficiency of local storage access, and ensures extremely low data read and write latency and overhead in dual-machine high availability scenarios.
Smart Images

Figure CN119938563A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a local data reading and writing method and a reading and writing device. Background Art
[0002] In a dual-machine hyper-converged deployment environment based on NVMe-oF (Non-Volatile Memory express over Fabrics), although high-performance remote storage access is achieved and the system's fault tolerance is improved, there is a major flaw: even if the application only needs to interact with local storage, data communication must still be completed through the network. This architectural feature leads to unnecessary bandwidth consumption and additional latency, because NVMe-oF was originally designed to connect remote storage resources across the network, but failed to optimize the local storage access path. This defect means that even in a high-availability configuration, when the client only needs to access the server on the local machine, it must follow the complete network transmission process, which not only wastes network resources, but also introduces unnecessary time delays. Summary of the invention
[0003] The present application embodiment provides a local data reading and writing method, including:
[0004] Determine the read or write request for local data initiated by the NVMe-oF client;
[0005] Obtain a virtual address of the NVMe target process; wherein the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated according to the physical address;
[0006] The NVMe-oF server performs data read or write operations based on the virtual address of the NVMe target process.
[0007] Optionally, the request includes a physical address for a local data read or write operation; or,
[0008] After determining the local data read or write request initiated by the NVMe-oF client, it further includes: obtaining a physical address for the local data read or write operation.
[0009] Optionally, before determining the local data read or write request initiated by the NVMe-oF client, the method further includes:
[0010] Receives a request to read or write data.
[0011] Optionally, the NVMe-oF client initiates a local data read or write request to the NVme driver, and the buffer corresponding to the local data read or write request is registered with the RDMA network card, and requests allocation of a memory area; the RDMA network card allocates a corresponding memory area for the buffer.
[0012] As an option, the process in which the memory mapping module in the NVMe driver obtains the virtual address of the NVMetarget process according to the physical address mapping is:
[0013] The NVMe-oF server sends an address query request to the NVMe-oF client through a communication module;
[0014] The NVMe-oF client queries the corresponding physical address according to the memory area based on the address query request;
[0015] The memory mapping module maps the physical address to obtain the virtual address of the NVMe target process.
[0016] Optionally, the physical address is allocated according to the physical address, including: the physical address is allocated by malloc or mmap;
[0017] After obtaining the virtual address of the NVMe target process, the method further includes: binding the physical address and the allocated virtual address of the NVMe target process.
[0018] Optionally, the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, including:
[0019] The NVMe-oF server reads or writes the lower storage device using asynchronous input and output, asynchronous input and output ring, or user-mode NVMe input and output virtual address mode according to the virtual address of the NVMe target process.
[0020] Optionally, a memory area mapping table is provided in the NVMe-oF client, and the memory area mapping table is used to store memory areas and their corresponding virtual addresses, and physical addresses converted from the virtual addresses.
[0021] Optionally, before the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, the NVMe-oF server further includes:
[0022] A memory area mapping table cache is set on the NVMe-oF server, and the memory area mapping table cache is used to cache the correspondence between the buffer area and the memory area of the NVMe-oF client.
[0023] The embodiment of the present application also provides a local data reading and writing device, including:
[0024] Memory, used to provide data reading or writing services;
[0025] An NVMe driver is used to determine a request for reading or writing local data, obtain a virtual address of an NVMe target process, and perform a data read or write operation on the buffer based on the virtual address of the NVMe target process; wherein the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated according to the physical address. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A flowchart of a local data reading and writing method according to an embodiment of the present application;
[0027] Figure 2 A flowchart of a memory mapping module in an NVMe driver in a local data reading and writing method according to an embodiment of the present application obtaining a virtual address of an NVMe target process according to a physical address obtained by conversion;
[0028] Figure 3 A schematic diagram of the structure of a local data reading and writing device according to an embodiment of the present application;
[0029] Figure 4 This is a schematic diagram of the structure of the local data reading and writing system according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] Various aspects and features of the present application are described herein with reference to the accompanying drawings.
[0031] It should be understood that various modifications may be made to the embodiments of the present application. Therefore, the above description should not be considered as limiting, but only as an example of an embodiment. Other modifications within the scope and spirit of the present application will occur to those skilled in the art.
[0032] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0033] These and other characteristics of the present application will become apparent from the following description of a preferred form of embodiment given as a non-limiting example with reference to the accompanying drawings.
[0034] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art will be able to readily implement many other equivalent forms of the present application.
[0035] The above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description when taken in conjunction with the accompanying drawings.
[0036] Specific embodiments of the present application are described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments applied for are merely examples of the present application, which may be implemented in a variety of ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that obscure the present application. Therefore, the specific structural and functional details applied for herein are not intended to be limiting, but merely serve as a basis and representative basis for the claims to teach those skilled in the art to use the present application in a variety of ways with substantially any suitable detailed structure.
[0037] This specification may use the phrases "in one embodiment," "in another embodiment," "in yet another embodiment," or "in other embodiments," all of which may refer to one or more of the same or different embodiments according to the present application.
[0038] Since the introduction of the NVMe-oF industry standard in 2016, it has extended the capabilities of high-performance storage array controllers to remote structures through network transmission protocols such as Ethernet, Fibre Channel, RoCE or InfiniBand, replacing the traditional PCIe bus connection method. NVMe-oF is based on message passing rather than shared memory architecture, allowing NVMe commands and responses to be efficiently transmitted between endpoints. All-flash arrays widely use NVMe-oF for remote storage access, separating the server from the client, and achieving data exchange through network interconnection.
[0039] In order to reduce TCO (Total Cost of Ownership), dual-machine hyper-convergence deployment has become a popular solution, that is, two servers each provide storage and computing services, and back up each other to ensure that a single failure will not affect business continuity. This solution usually integrates multiplexing high availability functions, and each virtual volume establishes at least two connection paths: one pointing to the local storage and the other pointing to the peer server. When the primary path fails, the traffic can be seamlessly switched to the backup path to ensure that the application continues to run unaffected, thereby improving the system's fault tolerance, reducing network latency, and improving transmission efficiency.
[0040] However, existing solutions have shortcomings. For solutions that use NVMe-oF clients to connect to the local and remote targets respectively, when only local storage needs to be interacted with, communication must still go through the network, resulting in unnecessary bandwidth consumption and increased latency. In a dual-machine high-availability configuration, if a server fails, the RDMA function of the network card of the other server may fail due to the lack of power supply of the RDMA direct line, and the expected performance cannot be achieved. In addition, although there are attempts to optimize the IO path by introducing a shared memory channel, this requires the client to explicitly specify the use of the local shared memory interface, and because it involves memory replication, it increases latency and resource usage.
[0041] Although current NVMe-oF applications have achieved efficient access to remote storage, in dual-machine high-availability scenarios, how to effectively reduce network dependence, reduce bandwidth consumption, and avoid performance degradation caused by hardware failures still needs further exploration and improvement. Therefore, this application provides a local data reading and writing method and reading and writing device based on the NVMe over RDMA (Remote Direct Memory Access) protocol.
[0042] The local data reading and writing method provided in the embodiment of the present application can be used in a dual-machine high-availability scenario. In this scenario, when the local machine is not faulty, data is read locally, and the local storage service directly accesses the data of the user process through the NVMe driver, with extremely low latency and overhead, without occupying the network card bandwidth; in addition, when data is written, since there is no memory copy, a lot of latency will be reduced.
[0043] In addition, in the NVMe-oF architecture, users do not need to care about the complex technical details of the underlying layer when operating storage resources. They only need to complete storage operations through the NVMe-oF client, which plays a key role in the entire process. The NVMe-oF client uses the local IP address and the remote IP address to establish a connection based on the user's needs. Specifically, the NVMe-oF client will first bind the local network interface based on the local IP to ensure the normal operation of the local end of the network communication; at the same time, it will establish a communication link with the server where the remote storage device is located based on the remote IP, so that data can be transmitted between the local and remote. In this process, when performing storage operations, users do not need to know details such as how NVMe-oF uses technologies such as RDMA to achieve efficient storage access, how data is transmitted in the network, and how to interact with memory between different nodes. The operating experience is the same as using local storage devices. Users can use familiar file system operation commands to perform storage operations, such as reading and writing files, just like operating local storage, because the NVMe-oF client will encapsulate, transmit, process, and receive and transmit responses to user requests and operations, making the entire storage access process transparent and convenient for users.
[0044] The local data reading and writing method of the present application is described in detail below with reference to the accompanying drawings. Figure 1 Flow chart of the local data reading and writing method according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:
[0045] S100: Determine a request for reading or writing local data initiated by an NVMe-oF client.
[0046] In an embodiment of the present application, the NVMe-oF server and the NVMe-oF client are located in the same host, and the NVMe-oF server runs in user mode and is connected to the NVMe-oF client.
[0047] The NVMe-oF server is configured to provide NVMe-oF storage services and manage actual storage resources, for example, connecting to local NVMe SSD (Solid State Drive) devices. The NVMe-oF server is also responsible for listening to connection requests from NVMe-oF clients. Once the connection is established, the NVMe-oF server can perform corresponding operations on the local storage device, such as reading or writing data, based on the request of the NVMe-oF client. At the same time, the NVMe-oF server is also used to handle tasks such as data cache management, storage device scheduling, and connection management with NVMe-oF clients to ensure efficient data transmission and stable operation of storage devices.
[0048] The NVMe-oF client is configured to request access to storage resources through the NVMe-oF protocol. The NVMe-oF client is similar to an initiator, making an operation request for stored data to the NVMe-oF server. The NVMe-oF client is used to construct data read or write request messages that comply with the NVMe-oF protocol specification and send these requests to the local NVMe-oF server. Before sending a request, the NVMe-oF client needs to establish a connection with the NVMe-oF server and perform necessary handshakes and negotiations to determine the functions and parameters supported by both parties. When the NVMe-oF client receives a response message returned by the NVMe-oF server, it processes the data according to the response content, for example, confirming whether the data is written successfully, or providing the read data to the upper-layer application for use.
[0049] When the NVMe-oF client needs to obtain data from the storage device, it sends a data read request to the NVMe-oF server. The read request contains key information such as the logical address of the data to be read in the storage device and the length of the data to be read. After receiving the read request, the NVMe-oF server locates the corresponding data block on the local storage device according to the logical address in the read request, then reads the data and returns it to the NVMe-oF client.
[0050] When the NVMe-oF client has data to be stored on the storage device, it sends a data write request to the NVMe-oF server. The write request contains information such as the logical address of the data to be written in the storage device and the actual data content to be written. After receiving the write request, the NVMe-oF server writes the data to the logical address location corresponding to the local storage device and returns a response message to the NVMe-oF client to inform the NVMe-oF client whether the data write operation is successful.
[0051] It should be noted that before determining the request for reading or writing local data initiated by the NVMe-oF client, it also includes: receiving a request for reading or writing data. That is, after receiving the request from the NVMe-oF client, a determination is first made to determine whether the request is a request for a read or write operation on local data, thereby determining the request for reading or writing local data initiated by the NVMe-oF client.
[0052] S200, obtain the virtual address of the NVMe target process.
[0053] Among them, the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is obtained by allocating the NVMe-oF server according to the physical address.
[0054] In modern operating systems, the memory used by applications when running is usually virtual memory. Virtual memory is a logical memory concept that provides a unified, continuous and independent memory space for each NVMe-oF client application, which makes it easier to write and run programs without considering the actual distribution of physical memory and memory conflicts between different applications. However, operations such as reading and writing data must ultimately be implemented at the actual physical address.
[0055] In the embodiment of the present application, the NVMe driver can be set in the host, in the NVMe-oF server, or in the NVMe-oF client. The NVMe driver is provided with a memory mapping module, which converts the virtual address used by the application of the NVMe-oF client into the corresponding physical address according to the memory management mechanism of the operating system, so as to ensure that the application can actually find the actual physical storage location where the data is stored when accessing the memory.
[0056] The local data read or write request initiated by the NVMe-oF client includes a physical address for the local data read or write operation. Alternatively, after determining the local data read or write request initiated by the NVMe-oF client, further comprising: obtaining a physical address for the local data read or write operation.
[0057] In a multi-process system environment, different processes usually need to share data or interact with each other. The memory mapping module in the NVMe driver will convert the physical address and remap it to a virtual address that can be used by the NVMe target process from the perspective of the NVMe target process. In other words, the memory corresponding to the original physical address has a corresponding virtual address representation from the perspective of other processes. Now, through the operation of the memory mapping module, the NVMe target process can also access this memory with a virtual address that conforms to its own memory management logic, thereby realizing the sharing or reallocation of memory between different processes, and thus facilitating operations such as data interaction between processes.
[0058] In addition, the virtual address of the NVMe target process can also be allocated by the NVMe-oF server according to the physical address. Specifically, the NVMe-oF server can allocate the virtual address of the NVMe target process through the malloc function or the mmap function. The memory mapping module binds the physical address of the NVMe-oF client and the allocated virtual address of the NVMe target process to achieve efficient data transmission, and ensures data consistency by supporting the zero-copy mechanism and reducing memory access overhead, avoiding data inconsistency problems and unifying memory status, optimizing system resource utilization, reducing memory fragmentation and CPU resource usage, and supporting remote direct memory access and other features, adapting to RDMA technology, and promoting efficient collaboration between different nodes in a distributed storage environment.
[0059] Among them, the malloc function is a standard library function for dynamically allocating memory. It can allocate a memory space of a specified size on the heap (a memory area of the program used for dynamic memory allocation) and return a pointer to the starting address of the memory. When the NVMe-oF server uses malloc to allocate a virtual address for the NVMe target process, it means that the server can request the operating system to dynamically allocate a continuous virtual address space for the NVMe target process as needed. It should be noted that when using malloc, it is initially allocated only in the virtual address space, and the actual allocation of physical memory is usually delayed. Only when the program accesses the memory area will the operating system allocate the corresponding physical memory for it through the page fault interrupt mechanism and establish a mapping relationship from the virtual address to the physical address.
[0060] mmap is another memory mapping function that provides more powerful functions than malloc. With mmap, you can map files or devices to the address space of a process, or create anonymous mappings. In the NVMe-oF server scenario, mmap can map part or all of the storage area of the NVMe storage device to the virtual address space of the NVMe target process, so that the NVMe target process can access data on the storage device like accessing memory. Similar to malloc, the virtual address allocated by mmap only exists in the virtual space at the beginning, and physical memory is allocated as needed during subsequent accesses.
[0061] In the NVMe-oF server, allocating virtual addresses to the NVMe target process helps manage storage resources, improve memory usage efficiency, and more flexibly handle requests from NVMe-oF clients. For example, when the NVMe-oF server receives a read or write request from an NVMe-oF client, it can map the request to the corresponding physical storage location on the storage device through the virtual address, facilitating data read and write operations.
[0062] In actual applications, the NVMe-oF server can select the appropriate memory allocation method (malloc or mmap) according to specific circumstances, such as the frequency of requests, the amount of data, the characteristics of the storage device, etc., to ensure that the NVMe-oF server can run efficiently and stably and meet the storage requirements of different NVMe-oF clients.
[0063] S300, the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process.
[0064] In an embodiment of the present application, when the NVMe-oF server receives the physical address information of the NVMe-oF client, it can directly map the physical address to the corresponding storage unit on the storage device, which can avoid additional address conversion processes and locate the data location on the storage device more quickly.
[0065] Specifically, the NVMe-oF server can access the user-mode NVMe through the physical address. Traditional storage device drivers and operations are usually executed in the kernel mode of the operating system, while user-mode NVMe allows NVMe devices to be directly operated in user space. This can reduce the context switching overhead between kernel mode and user mode, improve performance, and also give developers more flexibility.
[0066] When the NVMe-oF server receives the virtual address of the NVMe target process, the NVMe-oF server first converts the virtual address to the physical address of the system physical memory. This physical address conversion is from the virtual address space of the program to the physical address space of the system physical memory. It is mainly to achieve memory management and protection, and ensure that the program can access the memory normally under the management of the operating system. Then, the NVMe-oF server further maps the physical address of the system physical memory to the physical address of the storage device. This mapping process can direct the operation of the program to the storage device, so that the storage device can correctly receive the read and write requests. Finally, the NVMe-oF server can send the read and write requests to the corresponding physical location on the storage device, so that the NVMe-oF server can implement read and write operations based on the virtual address of the NVMe target process.
[0067] In one embodiment, in the above step S100, the physical address of the NVMe-oF client is obtained by converting the virtual address corresponding to the memory area of the NVMe-oF client by the memory mapping module in the NVMe driver.
[0068] In an embodiment of the present application, the memory mapping module set in the NVMe driver converts the virtual address corresponding to the client memory area into a physical address. This conversion process is very important because in actual data transmission and storage operations, the hardware device ultimately needs to accurately access and process data through the physical address. Without this conversion process, the hardware device cannot correctly locate the data stored in the NVMe-oF client memory, resulting in errors in data transmission and processing.
[0069] Furthermore, the NVMe-oF client initiates a local data read or write request to the NVme driver, and the buffer corresponding to the local data read or write request is registered with the RDMA network card, and requests to allocate a memory area; the RDMA network card allocates a corresponding memory area for the buffer.
[0070] In the embodiment of the present application, according to the RDMA communication principle, RDMA allows the application of one computer to directly access the memory of another computer without the intervention of the CPU, cache or operating system of the target host. This direct access can greatly improve the efficiency of data transmission and reduce the overhead caused by data copying and context switching in traditional network communication.
[0071] Ordinary memory buffers are not directly accessible to RDMA network cards. This is because the memory management mechanism of the operating system protects memory layout and access. When an application uses RDMA communication, it is necessary to allow the RDMA network card to recognize and directly access a specific buffer (i.e., buf). By registering a buffer with the RDMA network card, the RDMA network card is actually informed of the location, size, and other information of this memory, so that the RDMA network card knows that this memory can be directly operated by it.
[0072] The RDMA network card's access rights to memory are managed through the registration mechanism to ensure that the network card can only access the registered buffer, prevent illegal access to other memory, and ensure the security of the system.
[0073] When the buffer is successfully registered with the RDMA network card, a MR (Memory Region) will be obtained. MR is equivalent to a handle that is used to uniquely identify the registered buffer and contains the permission information required by the RDMA network card to access the MR. For example, MR can specify whether the memory area is readable, writable, or executable, as well as the remote node's access rights to the memory area. In actual communication, the RDMA network card uses MR to locate and operate the registered buffer.
[0074] In a specific embodiment, in the above step S300, the NVMe-oF server can also perform data read or write operations based on the physical address of the NVMe-oF client. When performing data read or write operations, the NVMe-oF server directly accesses the user-mode NVMe using the input and output physical address (IOPA, I / O Physical Address) according to the physical address.
[0075] Among them, the input and output physical address method is a method used to directly access physical memory addresses in computer systems. By using this method, the NVMe-oF server can directly interact with the NVMe-related memory area in the user state through the physical address provided by the NVMe-oF client, bypassing some intermediate links (for example, some memory management links in the kernel state).
[0076] Traditional storage access needs to be buffered and processed by the kernel, which incurs additional overhead. The NVMe-oF server uses this method to directly read and write data to NVMe-related memory areas in user mode. This not only reduces the overhead of data copying and context switching, improves data transmission efficiency, but also makes better use of the high-performance characteristics of NVMe devices. For example, in some database applications with extremely high real-time requirements, the NVMe-oF server can quickly respond to user I / O requests and improve the performance of the entire system by directly accessing user-mode NVMe.
[0077] After receiving the request containing the physical address, the NVMe-oF server performs the corresponding data read or write operation according to the physical address. For example, if it is a read request, the NVMe-oF server reads the data from the connected storage device and then sends the data directly to the memory location corresponding to the physical address specified by the NVMe-oF client; if it is a write request, the NVMe-oF server reads the data from the physical address specified by the NVMe-oF client and then writes it to the locally connected storage device.
[0078] In one embodiment, Figure 2As shown, the process by which the memory mapping module in the NVMe driver obtains the virtual address of the NVMe target process according to the converted physical address mapping is:
[0079] S210. The NVMe-oF server sends an address query request to the NVMe-oF client through a communication module.
[0080] The communication module can communicate via system calls, ioctl (I / O Control) or shared memory.
[0081] S220, the NVMe-oF client queries the corresponding physical address according to the memory area based on the address query request.
[0082] Further, a memory region mapping table (ie, MR mapping table) is provided in the NVMe-oF client, and the MR mapping table can be used to store the memory region and its corresponding virtual address, and the converted physical address of the NVMe-oF client. The corresponding physical address can be obtained by querying the MR mapping table.
[0083] S230, the memory mapping module maps the physical address to obtain the virtual address of the NVMe target process.
[0084] In a specific embodiment, in the above step S300, the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, specifically including:
[0085] The NVMe-oF server uses AIO (Asynchronous Input / Output), IOURing (Asynchronous I / ORing) or user-mode NVMe's IOVA (Input / Output Virtual Address) to read or write to the underlying storage device according to the virtual address of the NVMe target process.
[0086] In a specific embodiment, before the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, it also includes:
[0087] A memory region mapping table cache is set on the NVMe-oF server. The memory region mapping table cache is used to cache the correspondence between the buffer and the memory region of the NVMe-oF client.
[0088] Every time a read or write operation is performed, if the NVMe-oF client is queried for the correspondence between the MR and the buffer, it means that a certain search and confirmation process must be performed. The NVMe-oF client needs to traverse the relevant data structure or execute some search algorithms to determine the MR corresponding to the current read or write operation, which will undoubtedly bring additional time overhead. Especially in the scenario of frequent read and write operations, this repeated query behavior will accumulate and have an adverse effect on the overall storage performance, resulting in slower read and write speeds and increased response delays.
[0089] However, in actual application, the correspondence between the final buffer of the NVMe-oF client and the MR is relatively stable and does not change frequently. For example, an application continuously performs read and write operations on a fixed range of data. Based on this feature, a memory area mapping table cache (cache mechanism) is introduced to store the previously queried and determined correspondence between the MR and the buffer in this cache. The next time a read or write operation is performed, the corresponding correspondence is first searched in the cache. If it can be found, subsequent operations can be performed directly based on the information in the cache without querying the NVMe-oF client. This skips the time-consuming query process and greatly speeds up the entire read and write operation, thereby improving the overall performance of the system.
[0090] In one embodiment, the NVMe-oF client is also provided with a Multipath module, which can be located in the NVMe driver. When there are multiple data transmission paths between the NVMe-oF client and the NVMe-oF server, the Multipath module can monitor the status of these paths in real time. When a path fails, it can quickly switch data transmission to other available paths to ensure the continuity of data access and avoid data transmission interruption due to a single path failure, thereby improving the reliability of the entire system.
[0091] In one embodiment, a management configuration module can be set on both the NVMe-oF client and the NVMe-oF server to ensure system adaptation and optimization. For example, on the NVMe-oF client, the management configuration module can adjust the buffer size of data transmission according to the performance of the network card used to make full use of the network bandwidth; on the NVMe-oF server, according to the characteristics of the NVMe storage device, set a suitable I / O queue depth to optimize the read and write performance of the storage device.
[0092] In a specific embodiment, the data reading and writing interaction process between the NVMe-oF client and the NVMe-oF server includes a data reading process and a data writing process.
[0093] The data reading process is as follows:
[0094] On the NVMe-oF client, the user application initiates a data read request to the NVMe driver, where the buffer corresponding to the data read request is located in the user application. According to the RDMA communication principle, the buffer corresponding to the data read request needs to be registered with the RDMA network card and obtain an MR.
[0095] The NVMe driver stores the MR and its corresponding virtual address in the MR mapping table, converts the virtual address corresponding to the buffer into a physical address, and stores the physical address in the MR mapping table. It should be noted that the NVMe driver can perform specific operations on this buffer, one of which is to prevent the memory where the buffer is located from being paged, so as to ensure its stability in the memory, avoid performance degradation and data errors caused by memory paging operations during the interaction with the NVMe device, and ensure the efficient and stable operation of the NVMe device.
[0096] The NVMe-oF client communicates with the NVMe-oF server through the RDMA protocol. Specifically, the NVMe-oF client sends an rdma_send instruction to the NVMe-oF server, where the rdma_send instruction includes MR, buffer address, and read parameters.
[0097] After receiving the rdma_send instruction, the NVMe-oF server initiates an address query request to the NVMe-oF client through the communication module. Based on the address query request, the NVMe driver in the NVMe-oF client queries the corresponding physical address according to the memory area, and remaps the physical address through the memory mapping module to obtain the virtual address of the NVMe target process. The NVMe-oF client returns the virtual address and physical address of the NVMe target process to the NVMe-oF server.
[0098] The NVMe-oF server can read the underlying storage device using AIO (Asynchronous Input / Output), IOURing (Asynchronous I / ORing) or user-mode NVMe's IOVA (Input / Output Virtual Address) according to the virtual address of the NVMe target process. The NVMe-oF server can also directly access the user-mode NVMe using the input / output physical address (IOPA, I / O Physical Address) according to the physical address.
[0099] After the NVMe-oF server has finished reading the data, the data has been stored in the physical page corresponding to the virtual address of the NVMe-oF client, that is, the buffer corresponding to the read request of the NVMe-oF client has obtained the required data. The NVMe-oF server notifies the NVMe-oF client that the data reading is completed.
[0100] The data writing process is:
[0101] On the NVMe-oF client, the user application initiates a data write request to the NVMe driver, where the buffer corresponding to the data write request is located in the user application. According to the RDMA communication principle, the buffer corresponding to the data write request needs to be registered with the RDMA network card and obtain an MR.
[0102] The NVMe driver stores the MR and its corresponding virtual address in the MR mapping table, converts the virtual address corresponding to the buffer into a physical address, and stores the physical address in the MR mapping table. It should be noted that the NVMe driver can perform specific operations on this buffer, one of which is to prevent the memory where the buffer is located from being paged, so as to ensure its stability in the memory, avoid performance degradation and data errors caused by memory paging operations during the interaction with the NVMe device, and ensure the efficient and stable operation of the NVMe device.
[0103] The NVMe-oF client communicates with the NVMe-oF server through the RDMA protocol. Specifically, the NVMe-oF client sends an rdma_send instruction to the NVMe-oF server, where the rdma_send instruction includes MR, buffer address, and write parameters.
[0104] After receiving the rdma_send instruction, the NVMe-oF server initiates an address query request to the NVMe-oF client through the communication module. Based on the address query request, the NVMe driver in the NVMe-oF client queries the corresponding physical address according to the memory area, and remaps the physical address through the memory mapping module to obtain the virtual address of the NVMe target process. The NVMe-oF client returns the virtual address and physical address of the NVMe target process to the NVMe-oF server.
[0105] The NVMe-oF server can use AIO (Asynchronous Input / Output), IOURing (Asynchronous I / ORing) or user-mode NVMe's IOVA (Input / Output Virtual Address) to write to the underlying storage device based on the virtual address of the NVMe target process. The NVMe-oF server can also use the input / output physical address (IOPA, I / O Physical Address) to directly access the user-mode NVMe based on the physical address.
[0106] After the NVMe-oF server finishes writing the data, the data has been written from the NVMe-oF client virtual address to the backend storage device. The NVMe-oF server notifies the NVMe-oF client that the data writing is complete.
[0107] Based on the same inventive concept, Figure 3 As shown, the embodiment of the present application also provides a local data reading and writing device, including:
[0108] The receiving module is configured to receive a data read or write request initiated by the NVMe-oF client at the NVMe-oF server.
[0109] An acquisition module is configured to obtain the address of the NVMe-oF client, wherein the address is the physical address of the NVMe-oF client or the virtual address of the NVMe target process; the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated by the NVMe-oF server according to the physical address.
[0110] The execution module is configured to execute a data read or write operation based on the physical address of the NVMe-oF client or the virtual address of the NVMetarget process by the NVMe-oF server.
[0111] Optionally, the execution module is further configured to:
[0112] The NVMe-oF server directly accesses the user-state NVMe using the input and output physical address method according to the physical address. Alternatively, the NVMe-oF server reads or writes the lower storage device using asynchronous input and output, asynchronous input and output ring or user-state NVMe input and output virtual address method according to the virtual address of the NVMe target process.
[0113] Optionally, the physical address of the NVMe-oF client is obtained by converting a virtual address corresponding to a memory area of the NVMe-oF client by a memory mapping module in the NVMe driver.
[0114] Furthermore, the memory area of the NVMe-oF client is obtained by registering with the RDMA network card according to the buffer corresponding to the data read or write request initiated by the NVMe-oF client.
[0115] Optionally, the process by which the memory mapping module in the NVMe driver obtains the virtual address of the NVMetarget process according to the physical address mapping is:
[0116] The NVMe-oF client receives the address query request of the NVMe-oF server through the communication module. The NVMe-oF client queries the corresponding physical address according to the memory area based on the address query request. The memory mapping module maps the physical address to obtain the virtual address of the NVMe target process.
[0117] Optionally, the NVMe-oF server allocates the virtual address of the NVMe target process using malloc or mmap according to the physical address. After obtaining the virtual address of the NVMe target process, the memory mapping module binds the physical address of the NVMe-oF client and the allocated virtual address of the NVMe target process.
[0118] Optionally, a memory area mapping table is provided in the NVMe-oF client, and the memory area mapping table is used to store memory areas and their corresponding virtual addresses, and physical addresses converted from the virtual addresses.
[0119] Optionally, the NVMe-oF server is provided with a memory area mapping table cache, and the memory area mapping table cache is used to cache the correspondence between the buffer area and the memory area of the NVMe-oF client.
[0120] Based on the same inventive concept, Figure 4 As shown, the embodiment of the present application also provides a local data reading and writing device, including:
[0121] Memory is used to provide data reading or writing services.
[0122] An NVMe driver is used to determine a request for reading or writing local data, obtain a virtual address of an NVMe target process, and perform a data read or write operation on the buffer based on the virtual address of the NVMe target process; wherein the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated according to the physical address.
[0123] In the scenario where the NVMe-oF server and NVMe-oF client are set up in the same host, the NVMe-oF server runs in user mode and connects to the NVMe-oF client. The NVMe-oF server communicates with the NVMe-oF target through the NVMe-oF protocol.
[0124] Among them, Buffer is set in both NVMe-oF server and NVMe-oF client. NVMe-oF client sends the read and write requests of upper-layer applications to NVMe-oF server reliably and efficiently in the form of NVMe commands through NVMe Driver, Nic Driver and HCA (Host Channel Adapter), thus realizing remote access and control of storage data.
[0125] The NVMe-oF client passes its address information to the NVMe driver by initiating a system call to Char Dev (character device). The NVMe driver encapsulates the address information according to the rules of the NVMe-oF protocol stack and transmits it to the NVMe-oF server through the communication module set in the NVMe-oF client. The address of the NVMe-oF client is the physical address of the NVMe-oF client or the virtual address of the NVMe target process.
[0126] Specifically, the virtual address of the NVMe target process is obtained by MemReMap (memory mapping module) in the NVMe-oF client according to the physical address mapping, or is allocated by the NVMe-oF server according to the physical address.
[0127] The NVMe-oF server performs data read or write operations based on the physical address of the NVMe-oF client or the virtual address of the NVMe target process.
[0128] When the NVMe-oF server receives a data read request initiated by the NVMe-oF client, the Blk Driver (block device driver) reads data from the NVMe or HDD (hard disk drive). The read data is passed to AIO, Uring or UIO (user space I / O) for processing through the Blk Driver. The processed data is passed to the Buffer of the NVMe-oFTarget (ie, the NVMe-oF server) through SDS (software defined storage).
[0129] When the NVMe-oF server receives a data write request initiated by the NVMe-oF client, the data first enters the NVMe-oF Target's Buffer. The data in the Buffer is passed to one or more I / O processing mechanisms in AIO, Uring, or UIO through SDS. These mechanisms process the data according to the type of request and system configuration, such as asynchronous processing and the use of high-performance I / O interfaces. The processed I / O request is passed to the Blk Driver by the corresponding I / O processing mechanism. The Blk Driver converts the I / O request into instructions suitable for NVMe or HDD and writes the data to the corresponding storage device.
[0130] It should be noted that Mem (shared memory) is set in the host, and the Buffer of the NVMe-oF client and the Buffer of the NVMe-oF server are both connected to the same shared memory space.
[0131] By using shared memory and RDMA technology, multiple data copy operations can be avoided in the traditional data transmission process. Data can be directly transferred between the buffers of the NVMe-oF client and the NVMe-oF server, which can significantly reduce CPU load and latency. RDMA allows data to be transferred directly between memories, bypassing the intervention of the CPU, which can improve the efficiency of data transmission.
[0132] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and protection scope of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present application.
Claims
1. A local data reading and writing method, comprising: Determine the read or write request for local data initiated by the NVMe-oF client; Obtain a virtual address of the NVMe target process; wherein the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated according to the physical address; The NVMe-oF server performs data read or write operations based on the virtual address of the NVMe target process.
2. The method according to claim 1, wherein the request comprises a physical address for a local data read or write operation; or After determining the local data read or write request initiated by the NVMe-oF client, further comprising: Gets the physical address for a local data read or write operation.
3. The method according to claim 1, before determining the request for reading or writing local data initiated by the NVMe-oF client, further comprising: Receives a request to read or write data.
4. According to the method described in claim 1, the NVMe-oF client initiates a local data read or write request to the NVme driver, the buffer corresponding to the local data read or write request is registered with the RDMA network card, and requests to allocate a memory area; the RDMA network card allocates a corresponding memory area for the buffer.
5. According to the method of claim 1, the process of the memory mapping module obtaining the virtual address of the NVMe target process according to the physical address mapping is: The NVMe-oF server sends an address query request to the NVMe-oF client through a communication module; The NVMe-oF client queries the corresponding physical address according to the memory area based on the address query request; The memory mapping module maps the physical address to obtain the virtual address of the NVMe target process.
6. The method according to claim 1, wherein the physical address allocation comprises: The physical address is allocated by malloc or mmap; After obtaining the virtual address of the NVMe target process, the method further includes: binding the physical address and the allocated virtual address of the NVMe target process.
7. According to the method of claim 1, the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, comprising: The NVMe-oF server reads or writes the lower storage device using asynchronous input and output, asynchronous input and output ring, or user-mode NVMe input and output virtual address mode according to the virtual address of the NVMe target process.
8. According to the method described in claim 1, a memory area mapping table is provided in the NVMe-oF client, and the memory area mapping table is used to store memory areas and their corresponding virtual addresses, and physical addresses obtained by converting the virtual addresses.
9. The method according to claim 1, before the NVMe-oF server performs a data read or write operation based on the virtual address of the NVMe target process, further comprising: A memory area mapping table cache is set on the NVMe-oF server, and the memory area mapping table cache is used to cache the correspondence between the buffer area and the memory area of the NVMe-oF client.
10. A local data reading and writing device, comprising: Memory, used to provide data reading or writing services; NVMe driver, used to determine the read or write request for local data and obtain the virtual address of the NVMe target process; And performing a data read or write operation on the buffer based on the virtual address of the NVMe target process; wherein the virtual address of the NVMe target process is obtained by mapping the memory mapping module in the NVMe driver according to the physical address, or is allocated according to the physical address.