Request division data storage system based on equipment perception

By introducing low-latency auxiliary storage devices and request division modules into the SSD file system, user requests are divided into sub-requests suitable for SSD and auxiliary storage devices, the problem that existing SSD file systems cannot fully utilize the performance of high bandwidth SSDs, and efficient data storage performance and scalability are achieved.

CN120179160APending Publication Date: 2025-06-20HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189974.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing SSD file systems are difficult to fully utilize the performance potential of high-bandwidth SSDs under multiple workloads, resulting in low write throughput and high write latency.

Method used

A secondary storage device that supports byte addressing is introduced with low latency, and user requests are divided into sub-requests for data blocks, data pages and data non-aligned pages through the request division module, and sub-requests are converted into IO operations through the IO processing module, optimizing the IO processing mechanism to make full use of the performance of SSDs and auxiliary storage devices.

Benefits of technology

It effectively improves data storage performance, avoids the expensive alignment overhead and page cache overhead in traditional SSD storage systems, maximizes the performance potential of SSD, and supports high scalability and fast persistence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179160A_ABST
    Figure CN120179160A_ABST
Patent Text Reader

Abstract

The invention discloses a request division data storage system based on equipment perception, which belongs to the technical field of computer storage and comprises an SSD (Solid State Disk), auxiliary storage equipment and a file system, the file system comprises: a request division module, which is used for dividing a user request into sub-requests for each storage unit and sending the sub-requests to an IO processing module; the unified file mapping module is used for managing and positioning each storage unit; wherein the storage unit comprises a data block, a data page and a data non-aligned page; the data block is located in the SSD and used for placing large-scale aligned request data, and the data page and the data non-aligned page are located in the auxiliary storage device and used for placing small-scale aligned request data and small-scale non-aligned request data respectively; and the IO processing module is used for processing the data block sub-requests sent to the SSD and other sub-requests sent to the auxiliary storage device in parallel according to the parallel capability and the IO capability of each storage device at the bottom layer. The performance potential of the SSD can be fully exerted, and the overall performance of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer storage, and more specifically, relates to a data storage system for request partitioning based on device awareness. Background Art

[0002] With the continuous development of computer storage technology, solid-state drives (SSDs) are replacing slow mechanical hard drives (HDDs) as the mainstream storage devices. Currently, SSDs have been widely deployed, such as from the Internet of Things (IoT), mobile devices, servers, to large data centers. Due to the high internal parallelism of SSDs, the bandwidth of SSDs continues to grow. For example, the read / write bandwidth of PCIe4.0 SSDs is approximately 5GB / s - 7GB / s and 3GB / s - 6GB / s, while the latest PCIe5.0 SSDs have nearly doubled this bandwidth. At the same time, the size of the basic read / write unit of SSDs, the SSD-page, has also increased from 4KB or 8KB in the past to the currently common 16KB.

[0003] The file system is used to manage the underlying storage device and provide a logical file / directory view and corresponding file operations to upper-layer applications. Using the standard POSIX interface, applications call file read / write operations with byte-addressable offsets and lengths. The file system then converts these arbitrarily patterned user requests into IO operations allowed by the device.

[0004] While considering the characteristics of SSDs, mainstream SSD file systems still inherit the traditional block storage stack and use a page cache to convert file requests into one or more 512B or 4KB bios aligned according to the underlying physical page addresses. Each bio corresponds to one or more consecutive physical pages. For unaligned and page-cache-miss write operations, the file system uses read-modify-write (RMW) processing to force the write request to be aligned to the page boundary and then submit the bio. RMW processing first reads the corresponding page from the SSD, then updates the page, and finally writes the aligned page to the SSD. In addition, when the host's write operation is unaligned on the minimum access unit in the SSD (i.e., the SSD-page) and misses the internal cache of the SSD, the SSD must perform a similar RMW operation to internally align the SSD-page, which results in high alignment overhead and expensive page-cache overhead.

[0005] With the emergence of data-intensive applications such as machine learning, cloud computing, graph processing, and databases, file systems should make full use of the growing I / O capabilities of modern SSDs and provide comprehensive support for various hybrid workload patterns in a cost-effective manner (e.g., mixed small and large, unaligned and aligned, read and write), and support high scalability and fast persistence. However, due to high alignment overhead, expensive page cache overhead, and insufficient I / O parallelism, existing SSD file systems are difficult to fully exploit the high-bandwidth SSD performance potential under various workloads.

[0006] Generally speaking, in servers deployed with high-speed SSDs that require high performance and cost-effectiveness, existing file systems have low write throughput and high write latency, and cannot fully exploit the I / O performance of high-performance storage devices. Summary of the Invention

[0007] In view of the deficiencies of the prior art and the improvement requirements, the present invention provides a data storage system based on device-aware request partitioning, aiming to fully exploit the performance potential of high-bandwidth SSDs and improve the overall performance of the system.

[0008] To achieve the above object, according to one aspect of the present invention, there is provided a data storage system based on device-aware request partitioning, including: an SSD, an auxiliary storage device, and a file system; the latency of the auxiliary storage device is lower than that of the SSD, and it supports byte addressing; the file system includes: a request partitioning module, a unified file mapping module, and an I / O processing module;

[0009] The request partitioning module is configured to perform device-aware read / write request partitioning operations when receiving user requests from upper-layer applications, so as to partition the user requests into sub-requests for each storage unit, and send the sub-requests to the I / O processing module; the storage units include: data blocks, data pages, and data unaligned pages;

[0010] The size of the data block is an integer multiple of the minimum access unit size of the SSD; the size of the data page is an integer multiple of the minimum access unit size of the auxiliary storage device; the data unaligned page includes a header and a data part, the data part is used to store data segments smaller than the data page size, and the header is used to store the in-page offsets and sizes of each data segment; the read / write request partitioning operation preferentially partitions the user requests into sub-requests for data blocks, and for the part that cannot be partitioned into sub-requests for data blocks, it preferentially partitions it into sub-requests for data pages, and for the part that cannot be partitioned into sub-requests for data pages, it partitions it into sub-requests for data unaligned pages;

[0011] The unified file mapping module is configured to manage and locate each storage unit, map file offsets to the addresses of the storage units, and identify the types of storage units;

[0012] The IO processing module is used to convert the received sub-requests into IO operations and execute them. Among them, the IO operations obtained by converting the sub-requests of data blocks are sent to the SSD, and the IO operations obtained by converting the sub-requests of data pages and data misaligned pages are sent to the auxiliary storage device.

[0013] Furthermore, the read / write request partitioning operation includes: determining the type of the user request. If it is an append write request, then execute the append write request partitioning operation; if it is an overwrite write request, then execute the overwrite write request partitioning operation; if it is a mixed write request that simultaneously includes append write and overwrite write, then divide the user request into an append write request and an overwrite write request and respectively execute the append write request partitioning operation and the overwrite write request partitioning operation; if it is a read request, then execute the read request partitioning operation.

[0014] The append write request partitioning operation includes: partitioning the intermediate data aligned with the logical block from the data to be written, using the data before the intermediate data to fill the previous unfilled logical block, generating a write sub-request for the data block for the intermediate data, and generating a write sub-request for the data page for the data after the intermediate data.

[0015] The overwrite write request partitioning operation includes: partitioning the intermediate data aligned with the logical block from the data to be written, dividing the header data before the intermediate data into a data part aligned with the data page and a data segment misaligned with the data page, dividing the tail data after the intermediate data into a data part aligned with the data page and a data segment misaligned with the data page, generating a write sub-request for the data block for the intermediate data, generating a write sub-request for the data page for the data part aligned with the data page, and generating a write sub-request for the data misaligned page for the data segment misaligned with the data page.

[0016] The read request partitioning operation includes: locating all storage units corresponding to the read request and dividing the read request into read sub-requests for each storage unit.

[0017] Among them, the logical block size is the same as the data block size.

[0018] Furthermore, the IO processing module executes IO operations based on an adaptive parallel IO processing mechanism.

[0019] Among them, the adaptive parallel IO mechanism sets M1 dedicated IO threads for each SSD, and sets the thread interleaving binding granularity of each SSD to an integer multiple of the data block; the adaptive parallel IO mechanism sets M2 dedicated IO threads for each auxiliary storage device, and sets the thread interleaving binding granularity of each auxiliary storage device to an integer multiple of the data page; both M1 and M2 are positive integers greater than or equal to 1, and M1 ≥ M2.

[0020] Moreover, after receiving the partitioned sub-requests, the IO processing module determines the size of each IO according to the thread interleaving binding granularity, and uses corresponding dedicated IO threads to process these IOs.

[0021] Furthermore, the unified file mapping module maintains two tree-like index structures for each file, which are respectively used to locate the storage units in the SSD and the auxiliary storage device; when locating the storage units, a parallel search strategy is adopted, and two threads are used to search these two tree-like index structures simultaneously;

[0022] Alternatively, the unified file mapping module maintains a single tree-like index structure for each file, which is used to uniformly locate the storage units in the SSD and the auxiliary storage device; when locating data, only one search of the tree-like index structure is required to locate all the storage units corresponding to the request, and there is no need to search multiple times.

[0023] Furthermore, read-write range locks are set on the leaf nodes of the tree-like index structure to support concurrent writing to non-overlapping regions of the file and concurrent reading of all regions of the file.

[0024] Furthermore, the IO processing module uses the Direct IO mode to execute the write sub-requests to the SSD, uses the Buffered IO mode to execute the read sub-requests to the SSD, and uses the memory semantics access mode to execute the write and read sub-requests to the auxiliary storage device.

[0025] Furthermore, the file system also supports the asynchronous write mode;

[0026] Moreover, in the asynchronous write mode, the IO processing module uses the Buffered IO mode to execute the write and read sub-requests to the SSD, and uses the memory semantics access mode to execute the write and read sub-requests to the auxiliary storage device.

[0027] Furthermore, when the IO processing module processes the write sub-requests of data misaligned pages, it first merges the data segments to be written with the existing data segments intersecting within the data misaligned pages, then writes the merged data segments, and modifies the information in the header of the data misaligned pages.

[0028] Furthermore, the file system also supports an additional access mode; the additional access mode includes: only using the auxiliary storage mode and only using the SSD mode;

[0029] In the only using the auxiliary storage mode, the processing methods of the file system for the write requests from the upper-layer applications include: directly writing the write requests into the data pages of the auxiliary storage device;

[0030] In the case of only using the auxiliary storage mode, the file system processes read requests from upper-layer applications by sending the read requests to the request partitioning module, so that the request partitioning module partitions the read requests into sub-requests for each storage unit and sends the sub-requests to the IO processing module;

[0031] In the case of only using the SSD mode, the file system processes write requests from upper-layer applications as follows: if the write request is aligned to the boundary of the data block, the write request is directly written into the data block of the SSD; if the write request is not aligned to the boundary of the data block, the corresponding data is first read from the auxiliary storage device to align the write request to the boundary of the data block, and then the aligned write request is sent to the SSD;

[0032] In the case of only using the SSD mode, the file system processes read requests from upper-layer applications by sending the read requests to the request partitioning module, so that the request partitioning module partitions the read requests into sub-requests for each storage unit and sends the sub-requests to the IO processing module.

[0033] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0034] (1) The present invention introduces an auxiliary storage device that supports byte addressing with low latency in addition to the SSD. Based on the access characteristics of the SSD and the auxiliary storage device, three types of storage units, namely data blocks, data pages, and data misaligned pages, are proposed, and read / write requests from upper-layer applications are partitioned into sub-requests for the storage units. The sub-requests sent by the present invention to the corresponding underlying storage devices are always friendly to the device access characteristics. Therefore, the storage devices can work in their optimal mode for a long time and process these access requests more efficiently. In the process of processing write requests, the present invention does not need to rely on the page cache to perform operations similar to RMW, but aligns the requests to the minimum access unit of the SSD through request partitioning within the system. Therefore, the present invention can fully exert the performance potential of the SSD, effectively improve the data storage performance, and avoid the expensive alignment overhead in traditional SSD storage systems.

[0035] (2) In the preferred solution of the present invention, an adaptive parallel IO processing mechanism is introduced during the execution of IO operations, enabling the file system to automatically process the partitioned sub-requests in parallel and fully considering the parallel capabilities and access characteristics of the devices. Therefore, the performance potential of the devices can be maximally utilized. In addition, the present invention can be well applied to multi-device scenarios. Only by allocating corresponding dedicated IO threads for each device, the system has high scalability.

[0036] (3) In a preferred embodiment of the present invention, with the tree - shaped index structure as the core, a stable and fast data location function can be provided through a parallel search mechanism, thereby implementing an efficient unified file mapping mechanism, which can manage storage units uniformly and efficiently, and reduce the complexity of the file system introduced by the request partitioning mechanism. In some alternative embodiments, two tree - shaped index structures are maintained for each file, respectively used to locate storage units in the SSD and the auxiliary storage device. By searching the two index structures in parallel, the upper limit of the performance of request data location can be improved; in other alternative embodiments, a single tree - shaped index structure is maintained for each file, used to uniformly locate storage units in the SSD and the auxiliary storage device, thereby reducing the storage and computing resources required to maintain the tree - shaped index structure.

[0037] (4) In a preferred embodiment of the present invention, read - write range locks are set on the leaf nodes of the tree - shaped index structure, which can support concurrent write access to non - overlapping regions of the file and concurrent read access to all regions, further improving the concurrency performance of the file system.

[0038] (5) The present invention adopts multiple access data paths for the SSD and the auxiliary storage device respectively, which can further improve the performance of the file system and reduce the software stack overhead. Specifically, the present invention uses the Direct IO mode to execute write sub - requests for the SSD, which can avoid page cache overhead and significantly reduce the cost of fsync(); uses the Buffered IO mode to execute read sub - requests for the SSD, which can utilize a large number of optimized designs of page cache for read operations; uses the memory semantics access mode to execute write sub - requests and read sub - requests for the auxiliary storage device, which can be compatible with a variety of different storage devices that support memory semantics.

[0039] (6) The present invention further provides a mode of only using the auxiliary storage, a mode of only using the SSD, and an asynchronous write mode, enabling users to flexibly choose to use the underlying storage devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is an architecture diagram of a storage system based on device - aware request partitioning provided by an embodiment of the present invention;

[0041] Figure 2 It is a flowchart of read / write request partitioning operations provided by an embodiment of the present invention;

[0042] Figure 3 It is an example diagram of read / write request partitioning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0044] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0045] In order to solve the technical problem that the existing file system cannot exert the performance potential of high-performance SSDs, resulting in low system performance, the present invention provides a data storage system based on device-aware request partitioning. The overall concept is that, on the basis of introducing an auxiliary storage device with low latency and supporting byte addressing, actively utilize the rich semantics and information of the file system to identify and partition user requests from upper-layer applications, so that the resulting sub-requests match the access characteristics of the underlying storage devices. The resulting sub-requests are transparently and quickly processed in a manner that conforms to the parallel and I / O capabilities of the devices, changing the access mode to the storage devices at the root, enabling the devices to operate in their optimal mode, and thus making full use of the performance potential of the underlying storage devices.

[0046] The present invention does not require the auxiliary storage device to provide high bandwidth, high scalability, and large capacity. It only utilizes its access characteristics of low latency and byte addressing and can be applicable to various auxiliary storage devices, including non-volatile memory, CXL-SSD, etc.

[0047] Based on the above concept, in an embodiment of the present invention, a data storage system based on device-aware request partitioning is provided, as Figure 1 shown, including: an SSD, an auxiliary storage device, and a file system; the latency of the auxiliary storage device is lower than that of the SSD and it supports byte addressing; the file system includes: a request partitioning module, a unified file mapping module, and an I / O processing module;

[0048] The request partitioning module is used to perform device-aware read / write request partitioning operations when receiving user requests from upper-layer applications, so as to partition the user requests into sub-requests for each storage unit and send the sub-requests to the I / O processing module;

[0049] The unified file mapping module is used to manage and locate each storage unit, map the file offset to the address of the storage unit, and identify the type of storage unit;

[0050] The IO processing module is used to convert the received sub-requests into IO operations and execute them. Among them, the IO operations obtained by converting the sub-requests of data blocks are sent to the SSD, and the IO operations obtained by converting the sub-requests of data pages and data misaligned pages are sent to the auxiliary storage device.

[0051] In this embodiment, three types of storage units are set according to the characteristics of the underlying storage device, namely data blocks, data pages, and data misaligned pages. The size of the data block is an integer multiple of the minimum access unit size of the SSD. Optionally, in this embodiment, the data block size is set to 32KB. The size of the data page is an integer multiple of the minimum access unit size of the auxiliary storage device. Optionally, in this embodiment, the data page size is 4KB. The data misaligned page includes a header and a data part. The data part is used to store data segments smaller than the data page size, and the header is used to store the internal offset and size of each data segment. Optionally, in this embodiment, the header size of the data misaligned page is 64B, and the data part size is 4KB.

[0052] The solid-state drive is only used to place data blocks, and the auxiliary storage device is used to place data pages and data misaligned pages. The auxiliary storage device can be any memory with low latency and byte-addressing characteristics, such as various types of non-volatile memory, solid-state drives based on the CXL protocol, and so on.

[0053] When performing the read / write request partitioning operation, the alignment priority principle and the fragmentation minimization principle are followed. Among them, the alignment priority principle is to preferentially divide the large and logically block-aligned parts in the request into a whole, and the fragmentation minimization principle is to place the data into as few storage units as possible. Correspondingly, in this embodiment, the read / write request partitioning operation preferentially divides the user request into sub-requests of data blocks. For the part of the sub-request that cannot be divided into data blocks, it is preferentially divided into sub-requests of data pages. For the part of the sub-request that cannot be divided into data pages, it is divided into sub-requests of data misaligned pages. The specific process of the read / write request partitioning operation is as Figure 2 shown, including:

[0054] Judge the type of the user request. If it is an append write request, execute the append write request partitioning operation. If it is an overwrite write request, execute the overwrite write request partitioning operation. If it is a mixed write request that contains both append write and overwrite write, divide the user request into an append write request and an overwrite write request and then execute the append write request partitioning operation and the overwrite write request partitioning operation respectively. If it is a read request, execute the read request partitioning operation;

[0055] The additional write request partitioning operation includes: partitioning intermediate data aligned with logical blocks from the data to be written, filling the previous incomplete logical block with the data before the intermediate data, generating a write sub-request for a data block for the intermediate data, and generating a write sub-request for a data page for the data after the intermediate data;

[0056] The overwrite write request partitioning operation includes: partitioning intermediate data aligned with logical blocks from the data to be written, dividing the header data before the intermediate data into a data part aligned with data pages and a data segment not aligned with data pages, dividing the tail data after the intermediate data into a data part aligned with data pages and a data segment not aligned with data pages, generating a write sub-request for a data block for the intermediate data, generating a write sub-request for a data page for the data part aligned with data pages, and generating a write sub-request for a data non-aligned page for the data segment not aligned with data pages;

[0057] The read request partitioning operation includes: locating all storage units corresponding to the read request and partitioning the read request into read sub-requests for each storage unit;

[0058] Among them, the logical block size is the same as the data block size.

[0059] In the above read / write request partitioning operations, the determination of the user request type can be completed according to the file size, write request offset, and size, and this information all belongs to the rich semantics and information of the file system. In addition, when the file system places the partitioned data segment into a data non-aligned page, it first merges the data segment with the existing data segment intersecting within the data non-aligned page, and then updates the header of the data non-aligned page according to the merge result to reduce file data fragmentation during runtime.

[0060] Figure 3 As shown, it is an example when processing the write request partitioning using the above read / write request partitioning operations. This example is an example of a write operation starting from an offset of 89KB with a length of 107KB to a file with a size of 152KB. As Figure 3As shown, the data in the 0 - 32KB area, 64KB - 89KB area, 92KB - 109KB area, and 111KB - 128KB area of the file are stored in data blocks, the data in the 32KB - 64KB area and 128KB - 152KB area are stored in data pages, and the data in the 89KB - 92KB area and 109KB - 111KB area are stored in data misaligned pages. Through the above read / write request partitioning, the write request is first divided into an overwrite write request for the 89KB - 152KB area and an append write request for the 152KB - 196KB area. According to the data distribution of the file, the two logically block - aligned areas of 96KB - 128KB in the overwrite write request and 160KB - 192KB in the append write request are divided into data block sub - write requests. The 128KB - 152KB area and 92KB - 96KB area of the overwrite write request are divided into data page sub - write requests, and 89KB - 92KB is divided into data misaligned page sub - write requests. The 152KB - 160KB area and 192KB - 196KB area of the append write request are divided into data page sub - write requests.

[0061] Based on the above read / write request partitioning operation, the user request is divided into sub - requests that are adapted to the access characteristics of the underlying storage device (SSD or auxiliary storage device). After the IO processing module sends the divided sub - requests to the corresponding storage device, these sub - requests can be transparently and quickly processed in a manner that conforms to the device's parallel capabilities and IO capabilities, thereby fully utilizing the performance potential of the underlying storage device. To further improve system performance, as a preferred implementation, in this embodiment, the IO processing module performs IO operations based on an adaptive parallel IO processing mechanism;

[0062] Among them, the adaptive parallel IO mechanism sets M1 dedicated IO threads for each SSD, and sets the thread interleaving binding granularity of each SSD to an integer multiple of the data block; the adaptive parallel IO mechanism sets M2 dedicated IO threads for each auxiliary storage device, and sets the thread interleaving binding granularity of each auxiliary storage device to an integer multiple of the data page; both M1 and M2 are positive integers greater than or equal to 1, and M1≥M2. Optionally, in this embodiment, M1 = 32, M2 = 4, the thread interleaving binding granularity of the SSD is the same as the data block size, and the thread interleaving binding granularity of the auxiliary storage device is the same as the data page size;

[0063] Moreover, after receiving the divided sub - requests, the IO processing module determines the size of each IO according to the thread interleaving binding granularity and uses the corresponding dedicated IO thread to process these IOs.

[0064] After setting the IO dedicated threads and thread interleaving binding granularity of each storage device, each dedicated IO thread contains a dedicated request queue. All dedicated IO threads are interleaved and bound to the address space of the storage device at the selected IO size granularity. The divided sub-requests are further divided into the request queues of these dedicated IO threads according to their access addresses and are adaptively executed in parallel.

[0065] This embodiment uses an adaptive parallel IO mechanism to automatically parallelize and process the sub-requests after request division, fully considering the parallel capabilities and access characteristics of the devices, so as to maximize the utilization of the device performance potential. This embodiment can also be well used in the scenario of multiple devices, only by allocating corresponding dedicated IO threads for each device, which ensures the high scalability of the file system.

[0066] To further improve the performance of the file system and reduce the software stack overhead, this embodiment adopts multiple access data paths for SSDs and auxiliary storage devices respectively. Specifically, in this embodiment, the IO processing module uses the Direct IO mode to execute the write sub-requests to the SSDs to avoid page cache overhead and significantly reduce the fsync() cost; uses the Buffered IO mode to execute the read sub-requests to the SSDs, so that a large number of optimization designs for read operations can be made using the page cache; uses the memory semantics access mode to execute the write sub-requests and read sub-requests to the auxiliary storage devices to be compatible with various storage devices that support memory semantics.

[0067] In practical applications, the unified file mapping module can selectively adopt an explicit parallel search mechanism or an implicit parallel search mechanism;

[0068] In the explicit parallel search mechanism, the unified file mapping module maintains two tree-like index structures for each file, which are respectively used to locate the storage units in the SSDs and auxiliary storage devices. During the search, the two index structures are searched in parallel. This mechanism further uses a parallel optimization mechanism to identify the situation where all storage units corresponding to the user request are managed by a single index structure, and terminate the search for the other index structure at this time, so as to reduce unnecessary CPU occupancy and search overhead.

[0069] In the implicit parallel search mechanism, the unified file mapping module maintains a tree-like index structure for each file, which is used to uniformly locate the storage units in the SSDs and auxiliary storage devices. This index structure is a merged form of the above two index structures. Searching this index structure is logically equivalent to searching the two index structures in parallel.

[0070] The specific implementation of the tree - shaped index structure is the same as that in the key - value storage system, which will not be elaborated here. The tree - shaped index structure is naturally suitable for read - write range locks. Therefore, in this embodiment, regardless of which parallel search mechanism is adopted, read - write range locks are set on the leaf nodes of the tree - shaped index structure to support concurrent writing to non - overlapping regions of the file and concurrent reading of all regions of the file.

[0071] The unified file mapping mechanism provided in this embodiment takes the tree - shaped index structure as the core and utilizes the parallel search mechanism, which can provide stable and fast data location functions. By setting range locks, the parallelism of the file system can be further improved.

[0072] In this embodiment, the file system waits for all sub - requests to be completed before considering the request processing completed and returning the corresponding response to the user.

[0073] Generally speaking, in view of the problems of high alignment overhead and page cache overhead in the existing SSD file system, as well as insufficient IO parallelism, this embodiment uses a device - aware read / write request partitioning module to avoid alignment overhead (this overhead accounts for 17% - 76% of the write overhead of traditional SSD file systems), uses an efficient data path to avoid expensive page cache overhead (this overhead accounts for 6% - 39% of the write overhead of traditional SSD file systems), and uses a parallel IO processing mechanism module to overcome the problem of insufficient IO parallelism (improving the write performance by 29% - 93%). It effectively improves the performance of the file system and maximizes the performance potential of the SSD.

[0074] To further optimize the performance of the file system, in this embodiment, the file system will also perform data migration between the SSD and the auxiliary storage device when the preset migration conditions are met. The migration conditions include: the remaining space in the auxiliary storage device is lower than the preset threshold, there are continuous cold data pages in the auxiliary storage device, etc. The specific method of data migration is: select continuous (cold) data pages in the auxiliary storage device, merge them into data blocks, generate write sub - requests for the data blocks and send them to the IO processing module to write the corresponding data to the SSD, and then delete the migrated data in the auxiliary storage device.

[0075] To enable users to more flexibly select the underlying storage device, in this embodiment, the file system also supports the asynchronous write mode;

[0076] Moreover, in the asynchronous write mode, the IO processing module uses the Buffered IO mode to execute write sub - requests and read sub - requests to the SSD, and uses the memory - semantic access mode to execute write sub - requests and read sub - requests to the auxiliary storage device.

[0077] Furthermore, in this embodiment, the file system also supports additional access modes; the additional access modes include: only using the auxiliary storage mode and only using the SSD mode.

[0078] In the only using auxiliary storage mode, all write requests are not split but directly sent to the auxiliary storage device. However, for read requests, the file system still divides them according to the above-mentioned device-aware read / write request division method and then hands them over to the IO processing module for processing.

[0079] In the only using SSD mode, all write requests are not split but directly sent to the SSD. Since the SSD does not support unaligned writes, the file system will first read the corresponding data from the auxiliary storage device to align the write request to the data block boundary. If the corresponding data does not exist on the auxiliary storage device (such as unaligned append writes), the missing part of the request is filled with 0 to align it to the data block boundary. It is easy to understand that if the write request is originally aligned to the data block boundary, it can be directly written to the SSD. For read requests, the file system still divides them according to the above-mentioned device-aware read / write request division method and then hands them over to the IO processing module for processing.

[0080] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A data storage system based on device-aware request partitioning, characterized in that: include: SSD, secondary storage devices and file systems; The auxiliary storage device has a lower latency than the SSD and supports byte addressing; The file system comprises: a request division module, a unified file mapping module and an IO processing module; The request division module is used to perform a device-aware read / write request division operation upon receiving a user request from an upper-layer application, so as to divide the user request into sub-requests for each storage unit, and send the sub-requests to the IO processing module; the storage unit includes: a data block, a data page, and a data non-aligned page; the size of the data block is an integer multiple of the minimum access unit size of the SSD; the size of the data page is an integer multiple of the minimum access unit size of the auxiliary storage device; the data non-aligned page includes a header and a data part, the data part is used to store data segments that are less than the data page size, and the header is used to store the page offset and size of each data segment; the read / write request division operation preferentially divides the user request into sub-requests for data blocks, and preferentially divides the part that cannot be divided into sub-requests for data blocks into sub-requests for data pages, and divides the part that cannot be divided into sub-requests for data pages into sub-requests for data non-aligned pages; The unified file mapping module is used to manage and locate each storage unit, map the file offset to the address of the storage unit, and identify the type of storage unit; The IO processing module is used to convert the received sub-requests into IO operations and execute them; wherein the IO operations obtained by converting the sub-requests of data blocks are sent to the SSD, and the IO operations obtained by converting the sub-requests of data pages and data non-aligned pages are sent to the auxiliary storage device.

2. The data storage system based on device-aware request partitioning as claimed in claim 1, characterized in that: The read / write request division operation includes: judging the type of the user request, if it is an append write request, performing an append write request division operation; if it is an overwrite write request, performing an overwrite write request division operation; if it is a mixed write request including both append write and overwrite write, dividing the user request into an append write request and an overwrite write request, and then performing an append write request division operation and an overwrite write request division operation respectively; if it is a read request, performing a read request division operation; The additional write request division operation includes: dividing the intermediate data aligned with the logic blocks from the data to be written, filling the previous unsatisfied logic block with the data before the intermediate data, generating a write sub-request for the data block for the intermediate data, and generating a write sub-request for the data page for the data after the intermediate data; The overwrite write request division operation includes: dividing the intermediate data aligned with the logic blocks from the data to be written, dividing the head data before the intermediate data into the data part aligned with the data page and the data segment not aligned with the data page, dividing the tail data after the intermediate data into the data part aligned with the data page and the data segment not aligned with the data page, generating a write sub-request for the data block for the intermediate data, generating a write sub-request for the data page for the data part aligned with the data page, and generating a write sub-request for the data non-aligned page for the data segment not aligned with the data page; The read request division operation includes: locating all storage units corresponding to the read request, and dividing the read request into read sub-requests for each storage unit; The logical block size is consistent with the data block size.

3. The data storage system based on device-aware request partitioning according to claim 1 or 2, characterized in that: The IO processing module performs IO operations based on an adaptive parallel IO processing mechanism; The adaptive parallel IO mechanism sets M1 dedicated IO threads for each SSD, and sets the thread interleaving binding granularity of each SSD to an integer multiple of a data block; the adaptive parallel IO mechanism sets M2 dedicated IO threads for each auxiliary storage device, and sets the thread interleaving binding granularity of each auxiliary storage device to an integer multiple of a data page; M1 and M2 are both positive integers greater than 1, and M1≥M2; Furthermore, after receiving the divided sub-requests, the IO processing module determines the size of each IO according to the thread interleaving binding granularity, and uses corresponding dedicated IO threads to process these IOs.

4. The data storage system based on device-aware request partitioning as claimed in claim 1, characterized in that: The unified file mapping module maintains two tree index structures for each file, which are used to locate the storage units in the SSD and the auxiliary storage device respectively; when locating the storage unit, a parallel search strategy is adopted, and two threads are used to search the two tree index structures simultaneously; Alternatively, the unified file mapping module maintains a single tree index structure for each file, which is used for uniformly locating storage units in the SSD and the auxiliary storage device.

5. The data storage system based on device-aware request partitioning as claimed in claim 4, characterized in that: Read-write range locks are set on the leaf nodes of the tree index structure to support concurrent writing of disjoint areas of the file and concurrent reading of all areas of the file.

6. The data storage system based on device-aware request partitioning as claimed in claim 3, characterized in that: The IO processing module uses the Direct IO mode to execute the write sub-request to the SSD, uses the Buffered IO mode to execute the read sub-request to the SSD, and uses the memory semantic access mode to execute the write sub-request and the read sub-request to the auxiliary storage device.

7. The data storage system based on device-aware request partitioning as claimed in claim 3, characterized in that: The file system also supports asynchronous write mode; Furthermore, in the asynchronous write mode, the IO processing module uses the Buffered IO mode to execute the write sub-request and the read sub-request for the SSD, and uses the memory semantic access mode to execute the write sub-request and the read sub-request for the auxiliary storage device.

8. The data storage system based on device-aware request partitioning according to claim 1 or 2, characterized in that: When processing a write sub-request of a data non-aligned page, the IO processing module first merges the data segment to be written with the existing data segment intersecting in the data non-aligned page, then writes the merged data segment, and modifies the information of the data non-aligned page header.

9. The data storage system based on device-aware request partitioning according to claim 1 or 2, characterized in that: The file system also supports additional access modes; the additional access modes include: auxiliary storage only mode and SSD only mode; In the auxiliary storage only mode, the file system processes the write request from the upper layer application in the following manner: directly writing the write request into the data page of the auxiliary storage device; In the auxiliary storage only mode, the file system processes a read request from an upper layer application by sending the read request to the request division module, so that the request division module divides the read request into sub-requests for each storage unit and sends the sub-requests to the IO processing module; In the SSD-only mode, the file system processes the write request from the upper layer application in the following manner: if the write request is aligned to the boundary of the data block, the write request is directly written to the data block of the SSD; if the write request is not aligned to the boundary of the data block, the corresponding data is first read from the auxiliary storage device to align the write request to the boundary of the data block, and then the aligned write request is sent to the SSD; In the SSD-only mode, the file system processes the read request from the upper-layer application by sending the read request to the request partitioning module, so that the request partitioning module partitions the read request into sub-requests for each storage unit, and sends the sub-requests to the IO processing module.