Scalable File System, Method, Medium and Device Based on NVM Hybrid Memory

Through the file locking strategy, log pre-allocation and page recycling mechanism based on data blocks, the lock granularity and concurrency of the NVM hybrid memory file system are optimized, and the scalability problem in high concurrency scenarios is solved, and higher throughput and multi-threading performance is achieved.

CN116126806BActive Publication Date: 2025-07-22SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211684982.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-07-22
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

The existing file system based on NVM hybrid memory is insufficient in high concurrency scenarios, especially when multiple threads read and write the same file, lock conflicts lead to serialized execution, affecting system performance.

Method used

Using file locking strategies based on data blocks, log pre-allocation strategies and page-based garbage collection mechanisms, optimize lock granularity and concurrency, reduce lock holding time, and improve pipeline concurrent execution.

Benefits of technology

It significantly improves the throughput and scalability of the file system in high concurrency scenarios, reduces garbage collection waiting time, and improves the performance of multi-thread access to the same file.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116126806B_ABST
    Figure CN116126806B_ABST
Patent Text Reader

Abstract

The present invention provides a scalable file system, method, medium and device based on NVM hybrid memory, including: a file lock policy based on data blocks: defining the granularity of locking during file system reading and writing as data blocks of a fixed size, and when an application sends a read / write request to the file system, the file system will perform a locking operation on the data blocks containing the requested data range; a pre-allocation policy for logs: performing data persistence operations on write logs and write data in the file system to maximize concurrency during I / O execution; a page-based garbage collection mechanism: tracking the usage of each page through a page-based counter, and using a separate lock for each log page to control whether it can be garbage collected. The present invention optimizes the scalability of the file system based on NVM hybrid memory, and improves the throughput of the file system in the scenario of multi-threaded reading and writing of the same file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of NVM hybrid memory, and in particular, to a scalable file system, method, medium and device based on NVM hybrid memory. Background Art

[0002] As a computer subsystem that provides user data persistence, the file system plays a very important role in the operating system. Through the abstraction and organization of data in the persistent device, user data is stored in the persistent device connected to the system to complete the data persistence operation, allowing the recovery of data after power failure. The abstraction and organization of file data, the consistency protection of data under multi-user access, and the speed of file reading and writing are all important considerations in the design and implementation of the file system.

[0003] The rapid development of computer systems has continuously improved the processing power of computing devices. The computing power of the central processing unit (Central Computing Process, abbreviated as CPU) is no longer limited by the number of transistors stacked on the chip. To meet the high-speed processing performance requirements of the central processing unit, the solution of multi-core central processing units provides a new direction for the development of the CPU, greatly improving the computing power and computing speed of the CPU, and enabling the possibility of a single machine to process high-performance computing tasks. Currently, modern commercial servers are often configured with hundreds of central processing unit cores, allowing multi-core processors to perform multi-task calculations simultaneously, or to cooperate to complete a single complex task, or to complete different types of computing tasks.

[0004] While the processor performance is improving, the demand for I / O operations by application programs is also constantly expanding. Complex and intensive computing tasks are highly dependent on data. On the one hand, the data required for computing tasks needs to be obtained from the persistent device; on the other hand, the calculated data often has the need for persistence and needs to be finally stored in the persistent device through the file system. As the scale of application programs increases, the volume of persistent data is also constantly expanding, and the performance bottleneck of the system falls on the data I / O process. The massive concurrent applications and data reads pose new challenges to the I / O capabilities of the system. In addition to requiring low latency and high throughput support, new requirements are also put forward for the scalability of the file system and high-speed storage devices.

[0005] Regarding the scalability problem of the I / O process, on the one hand, hardware manufacturers increase the concurrent support at the hardware level by optimizing the arrangement of storage media and constructing multiple I / O queues; on the other hand, there are also many related studies on the scalability of the file system at the software level, hoping to slow down the performance degradation problem in high-concurrency scenarios through software.

[0006] The emergence of new non-volatile memory (Non-Non-Volatile Memory, abbreviated as NVM) provides an opportunity for high-performance file systems. New non-volatile memories include ReRam, PCM, and 3D-XPoint, etc., which have characteristics such as low latency and byte access, making it possible for the development of high-performance file systems. In addition, with the release of Intel Optane non-volatile memory, the commercialization of non-volatile memory provides convenience for commercial software written based on non-volatile memory. Compared with traditional file systems that use mechanical hard disks and solid-state drives as external storage, such as ext4 and F2FS, people now hope to utilize the characteristics of low read / write latency and high throughput of NVM to develop a high-performance file system based on NVM hybrid memory.

[0007] With the increase in the number of cores of general-purpose processors, the metrics for measuring the performance of file systems have added requirements for high scalability on the basis of latency and throughput. Although there are currently many studies on the scalability of file systems for traditional storage media, the research on the high scalability of file systems for NVM hybrid memory is still in its infancy. Conducting research on high-scalability file systems for NVM hybrid memory is of great research significance and application value for the high-quality development of file systems.

[0008] Patent document CN107862064A (application number: CN201711133827.X) discloses a high-performance, scalable lightweight file system based on NVM, including: superblock, inode table, hash table, segment table, metadata log, data log, bitmap, and file data space; inodes are stored in the inode table, and each node stores necessary metadata information; segments are stored in the segment table, and each segment stores information about a continuous area organized in bytes; the file system naming layer is organized by the global hash table, and each hash bucket is a linked list that links nodes with the same hash value; the data of each file is managed by a segment-based file B+ tree, and each segment represents a corresponding file data fragment as a leaf node of the file tree; both the metadata log and the data log contain multiple log files; the bitmap represents the usage of each data block in the file system; the file data space stores file data and is managed in units of 4KB-sized blocks.

[0009] In addition, existing research on improving the scalability of file systems is based on the assumption that users will not read and write the same file. The file system protects the consistency of read and write data for files through file-based locks. When any two requests need to write to the same file system, they can only be executed serially due to lock conflicts. Therefore, it is necessary to optimize for the situation where concurrent requests read and write different regions of the same file.

[0010] Based on the above requirements, the present invention conducts research on the scalability of a file system based on NVM hybrid memory, optimizes a typical log-based file system based on NVM hybrid memory, optimizes the granularity of locks through block-level locks, allows a higher degree of concurrency, and proposes a pre-allocation strategy for logs and a page-based garbage collection scheme, which better optimizes the execution of the read-write pipeline, thereby improving the scalability of the file system based on NVM hybrid memory. Summary of the Invention

[0011] Aiming at the deficiencies in the prior art, the object of the present invention is to provide a scalable file system, method, medium, and device based on NVM hybrid memory.

[0012] The scalable file system based on NVM hybrid memory provided by the present invention includes:

[0013] A file lock policy based on data blocks: Define the granularity of locking during file system reading and writing as data blocks of a fixed size. When an application sends a read / write request to the file system, the file system will lock the data blocks containing the requested data range.

[0014] A pre-allocation strategy for logs: Perform data persistence operations on write logs and write data in the file system to maximize concurrency during I / O execution.

[0015] A page-based garbage collection mechanism: Track the usage of each page through page-based counters, and use a separate lock for each log page to control whether it can be garbage collected.

[0016] Preferably, the file lock policy based on data blocks includes: When the size of a read / write request exceeds the size of a single data block, lock multiple consecutive data blocks at the beginning of processing the request; Determine the granularity of the file lock when the file system is mounted. After the size of the lock is fixed, make the index structure of the file lock into an array structure for quick positioning through offsets.

[0017] Preferably, the pre-allocation strategy for logs includes: Recombine the pipeline of write requests, integrate multiple writes of logs into a single write, and decompose the two strongly correlated operations of recording logs and updating the tail pointer. Update the file log tail pointer during log space allocation and complete the persistence operation of the log without locks.

[0018] Preferably, the page-based garbage collection mechanism includes: During the garbage collection process, only process some pages that allow garbage collection; Immediately release the relevant locks after the page-based garbage collection ends, allowing other applications to access the data in a timely manner, and reducing the time the system is suspended due to garbage collection.

[0019] The method for implementing a scalable file system based on NVM hybrid memory provided by the present invention includes:

[0020] Step 1: Define the granule of locking during read / write operations of the file system as a data block of a fixed size. When an application sends a read / write request to the file system, the file system will lock the data blocks containing the requested data range.

[0021] Step 2: Perform data persistence operations on write logs and write data in the file system to maximize concurrency during I / O execution.

[0022] Step 3: Track the usage of each page through a page-based counter, and use a separate lock for each log page to control whether it can be garbage collected.

[0023] Preferably, Step 1 includes: when the size of a single read / write request exceeds the size of a single data block, lock multiple consecutive data blocks at the beginning of processing the request; determine the granule of the file lock when the file system is mounted. After the size of the lock is fixed, make the index structure of the file lock into an array structure and perform fast positioning through offsets.

[0024] Preferably, Step 2 includes: recombine the pipeline of write requests, integrate multiple writes of the log into a single write, and decompose the two strongly related operations of recording the log and updating the tail pointer. Update the tail pointer of the file log during log space allocation and complete the persistence operation of the log without locks.

[0025] Preferably, Step 3 includes: during the garbage collection process, only process some pages that allow garbage collection; immediately release the relevant locks after the page-based garbage collection ends, allowing other applications to access the data in a timely manner and reducing the time for the system to be suspended due to garbage collection.

[0026] A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the method for implementing a scalable file system based on NVM hybrid memory as described above are implemented.

[0027] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the method for implementing a scalable file system based on NVM hybrid memory as described above are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] (1) The present invention designs a data block-based file lock strategy for the NVM hybrid memory file system. By reducing the granularity of the data lock, it supports concurrent requests for reading and writing operations on the same file, thereby improving the scalability of the file system in the scenario of high-concurrency reading and writing of the same file;

[0030] (2) The present invention proposes a log pre-allocation strategy, constructs a new file system read and write operation pipeline, converts high-latency operations into a lock-free design, allows more steps in the pipeline to be executed concurrently, and improves the concurrent execution degree of the pipeline;

[0031] (3) The present invention constructs a page-based garbage collection mechanism, maintains the read and write counters of each log page, and reduces the waiting time of the garbage collection process and the time the system is suspended by the garbage collection thread through a page-based locking scheme. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0033] Figure 1 This is a diagram of the scalable file system architecture based on NVM hybrid memory;

[0034] Figure 2 It is the execution pipeline diagram of a single write request;

[0035] Figure 3 This is a performance comparison chart of multi-threaded reading of private data blocks of the same shared file in a specific embodiment;

[0036] Figure 4 This is a performance comparison chart of multiple threads writing private data blocks of the same shared file in a specific embodiment. DETAILED DESCRIPTION

[0037] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0038] Example:

[0039] like Figure 1, the present invention provides a scalable file system based on NVM hybrid memory, including a file lock policy based on data blocks, a pre-allocation policy for logs, and a page-based garbage collection mechanism. Among them, the file lock based on data blocks is the core of the present invention and also a specific implementation for refining the granularity of file locks. By refining the file locks, it allows the same file to be held by multiple write threads, thus ensuring the scalability of the file system. In addition, in order to minimize the holding time of locks, a pre-allocation policy for logs is proposed. The page-based garbage collection mechanism effectively reduces the waiting time of mutex locks by sacrificing the requirement for the integrity of garbage collection. The pre-allocation policy for logs and the page-based garbage collection mechanism also provide consistency protection for the file system.

[0040] 1. File lock policy based on data blocks

[0041] The present invention defines the granularity of locking during file system reading and writing as data blocks of a fixed size. When an application sends a read / write request to the file system, the file system will lock the data blocks containing the requested data range. When the size of a single request exceeds the size of a single data block, multiple consecutive data blocks will be locked at the beginning of processing the request. In the case of a fixed data block size, it may result in a larger locked range than the actual read / write data range. However, for larger files, compared with file-level locks, the locked range of the fine-grained locks based on data blocks is much smaller than the lock of the entire file, so it can also bring more concurrent support. In addition, the granularity of file locks is determined when the file system is mounted. After the size of the lock is fixed, the index structure of the file lock can be made into an array structure and quickly located through offsets, thereby reducing the exclusive requirements for locking / unlocking operations on the file system, which provides very friendly support for scalability.

[0042] 2. Pre-allocation policy for logs

[0043] The concurrency of file reading and writing is determined by the design of the file system. For high scalability support, it is required that there are not too many metadata operations in a single read / write operation, otherwise high-concurrency requests cannot be supported due to the exclusivity of metadata modification. In order to maximize the concurrency during I / O execution, a pre-allocation policy for logs is designed to allow more steps in the read / write pipeline to execute concurrently.

[0044] For read / write operations, the largest latency comes from the I / O latency of the persistent medium, that is, the latency of loading or storing data from / to the persistent device. In order to increase the concurrency support of the file system, it is necessary to avoid holding exclusive locks for high-latency operations as much as possible.

[0045] In the file system, the two steps that require data persistence operations are writing logs and writing data. Compared with the independence of file data, the writing of log data is closely related to metadata. To reduce the use of exclusive locks during data writing, the present invention designs a pre-allocation strategy for logs. At the same time, considering the consistency problem of the file system, the pipeline of write requests is reorganized. The multiple writes of logs are integrated into a single write, and the two strongly related operations of recording logs and updating the tail pointer are decomposed, as Figure 2 . Update the file log tail pointer during log space allocation, and then complete the persistence operation of the log without locks.

[0046] 3. Page-based garbage collection mechanism

[0047] The page-based garbage collection mechanism is a brand-new garbage collection mechanism designed to minimize the impact of the garbage collection thread on other ongoing threads. It can perform fast garbage collection on some data to reduce the waiting time of the garbage collection thread and other threads.

[0048] By designing a page-based counter to track the usage of each page, a separate lock is used for each log page to control whether it can be garbage collected. On the one hand, during the garbage collection process, the system only needs to process some pages that allow garbage collection, effectively reducing the waiting time of the garbage collection thread. On the other hand, after the page-based garbage collection ends, the relevant locks are immediately released, allowing other applications to access the data in a timely manner, reducing the time the system is suspended due to garbage collection.

[0049] The implementation of the Scalable File System for NVM Hybrid Memory (S-NOVA) adheres to the design principle of enhancing the scalability of the file system. The core of the implementation of this file system is to optimize the granularity of file locks, allowing different threads to access and modify the same file concurrently. And through the division of the request process, a pipeline is constructed, and the processing power of the CPU is fully utilized in the form of a pipeline to improve the scalability of the file system in high-concurrency scenarios.

[0050] In this embodiment, the configuration of the running platform is determined as follows. In terms of hardware, the model of the system hardware is:

[0051] (1) CPU: Intel Xeon Gold 6230 processor with 20 cores

[0052] (2) Memory RAM: 250GB

[0053] (3) Storage: Intel Optane DCPMM 4×512GB

[0054] And the settings of the software system are:

[0055] (1) Host platform: Ubuntu 16.04.1 LTS

[0056] (2) Kernel: Linux 4.18.0

[0057] Here, the fxmark tool is used to test the scalability of the file system, and the improvement of the file system scalability is reflected by recording the throughput of the file system in a specific scenario.

[0058] Figure 3 Shows the throughput of the file system S-NOVA designed by the present invention when multi-threadedly reading the private data blocks of the same shared file. Figure 3 As can be seen from the data, the throughput of S-NOVA nearly linearly increases in the scenario of multi-threadedly reading the private data blocks of the same shared file. When the number of threads reaches 40, 119M requests can be processed per second, which is 22.4 times that of the traditional NVM-based hybrid memory file system NOVA.

[0059] Figure 4 Shows the throughput of the file system S-NOVA designed by the present invention when multi-threadedly writing the private data blocks of the same shared file. Figure 4 As can be seen from the data, the throughput of S-NOVA in the scenario of multi-threadedly reading the private data blocks of the same shared file is significantly higher than that of the traditional NVM-based hybrid memory file system NOVA, with a maximum performance improvement of 120%.

[0060] Through this embodiment, it can be found that the scalable file system based on NVM hybrid memory designed by the present invention has a significant increase in throughput in a high-concurrency scenario and has good scalability.

[0061] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware component.

[0062] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A scalable file system based on NVM hybrid memory, characterized in that, Comprising: File lock policy based on data blocks: Defines the granularity of locking during file system read / write operations as data blocks of a fixed size. When an application sends a read / write request to the file system, the file system locks the data blocks containing the requested data range. Log pre-allocation policy: Performs data persistence operations on write logs and write data in the file system to maximize concurrency during I / O execution. Page-based garbage collection mechanism: Tracks the usage of each page through page-based counters and uses a separate lock for each log page to control whether it can be garbage collected. The file lock policy based on data blocks includes: When the size of a single read / write request exceeds the size of a single data block, lock multiple consecutive data blocks at the beginning of request processing; Determine the granularity of the file lock during file system mounting. After the lock size is fixed, make the index structure of the file lock into an array structure for fast positioning through offsets. The log pre-allocation policy includes: Recombine the pipeline of write requests, integrate multiple writes of the log into a single write, and decompose the two strongly related operations of logging and updating the tail pointer. Update the file log tail pointer during log space allocation and complete the log persistence operation without locks.

2. The scalable file system based on NVM hybrid memory according to claim 1, wherein The page-based garbage collection mechanism includes: During garbage collection, only process some pages that allow garbage collection; Immediately release the relevant locks after page-based garbage collection ends, allowing other applications to access data in a timely manner and reducing the system suspension time due to garbage collection.

3. A method for implementing an extensible file system based on NVM hybrid memory, characterized in that Comprising: Step 1: Define the granularity of locking during file system read / write operations as data blocks of a fixed size. When an application sends a read / write request to the file system, the file system locks the data blocks containing the requested data range. Step 2: Perform data persistence operations on write logs and write data in the file system to maximize concurrency during I / O execution. Step 3: Track the usage of each page through page-based counters and use a separate lock for each log page to control whether it can be garbage collected. The said Step 1 includes: When the size of a single read / write request exceeds the size of a single data block, lock multiple consecutive data blocks at the beginning of request processing; Determine the granularity of the file lock during file system mounting. After the lock size is fixed, make the index structure of the file lock into an array structure for fast positioning through offsets. The said Step 2 includes: Recombine the pipeline of write requests, integrate multiple writes of the log into a single write, and decompose the two strongly related operations of logging and updating the tail pointer. Update the file log tail pointer during log space allocation and complete the log persistence operation without locks.

4. The method for implementing an extensible file system based on NVM hybrid memory according to claim 3, wherein The said Step 3 includes: During garbage collection, only process some pages that allow garbage collection; Immediately release the relevant locks after page-based garbage collection ends, allowing other applications to access data in a timely manner and reducing the system suspension time due to garbage collection.

5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for implementing a scalable file system based on NVM hybrid memory as described in claim 3 or 4.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for implementing an NVM hybrid memory-based scalable file system described in claim 3 or 4.

Citation Information

Patent Citations

  • High-performance extensible lightweight file system based on NVM

    CN107862064A

  • Decentralized storage method and device for data de-duplication

    CN113449065A