Multi-stream management method for application data and flash memory device

By obtaining the compression rate and life cycle of application data, clustering and mapping, the write amplification problem of solid-state drives is solved, transparent multi-stream management of the host is realized, and the performance and life of flash memory devices are improved.

WO2025138539A1PCT designated stage expired Publication Date: 2025-07-03DAPUSTOR CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/093410
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-05-15
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing technology has write amplification problems in solid-state drives, resulting in reduced bandwidth resource consumption and service life in storage. The commonly used multi-stream management mechanism requires modifying the operating system kernel and relying on the host, and the problem of writing amplification cannot be transparently optimized.

Method used

By obtaining the compression rate of application data, clustering and generating software streams, estimating the life cycle of each software stream, establishing a mapping relationship between the software stream and the hardware stream, and storing data with similar life cycles in the same hardware stream, optimizing the write amplification problem.

Benefits of technology

Reduces the number of times data is moved on the flash memory device, optimizes the write amplification problem, improves the performance and life of the flash memory device, and at the same time realizes transparent multi-stream management of the host.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024093410_03072025_PF_FP_ABST
    Figure CN2024093410_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of storage management, and discloses a multi-stream management method for application data and a flash memory device. The method comprises: clustering data having similar compression ratios to generate a plurality of software streams; estimating the life cycle of each software stream, so as to map the software streams having similar life cycles into a same hardware stream; and storing, in storage positions corresponding to the hardware streams, data corresponding to the software streams. According to the present application, on the basis of compression ratios, application data having similar life cycles is further stored in storage positions corresponding to hardware streams, so as to mitigate the write amplification problem of the flash memory device.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-stream management method for application data and flash memory device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311863315.4, and application name “Multi-stream management method of application data and flash memory device”, all contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of storage management technology, and in particular to a multi-stream management method for application data and a flash memory device. Background Art

[0004] Solid-state drives (SSDs) based on non-volatile storage media (NAND FLASH) can support "write-in-place". "Write-in-place" means writing data directly to the storage medium. Before performing a write operation, NAND FLASH needs to erase the old data and recycle the invalid pages through the garbage collection mechanism. However, when garbage collection is performed on data, a large number of valid pages still remain in the block to be recycled. These valid pages must be written to an idle location first before the block can be erased and recycled for subsequent use.

[0005] However, the write-back operation of valid pages causes additional write operations, leading to the write amplification problem. Write amplification not only consumes bandwidth resources in the storage and reduces read and write performance, but also reduces the service life of the solid-state drive. Therefore, solid-state drives that can support multi-stream features need to support data isolation at the hardware level, allowing application software to divert data with different update modes (for example: append updates, in-place updates) and different life cycles, so that data with similar life cycles are placed in the same physical block, and data with large differences in life cycles are placed in different physical blocks. When garbage collection is performed, the life cycle of the data in the recycled physical blocks is basically invalid. Therefore, reducing the amount of data that needs to be written back and thus reducing the write amplification problem of the solid-state drive, for solid-state drives that support multi-stream features, how to perform accurate, efficient and scalable multi-stream management is crucial.

[0006] Currently, the most commonly used multi-stream management mechanism is an automatic traffic diversion management mechanism based on statistical information and host opacity. This automatic traffic diversion management mechanism estimates the data lifecycle based on the characteristic information of the data corresponding to the statistical write request, and further maps data with different lifecycles to different flow identifiers through an algorithm. The disadvantages of this method are as follows:

[0007] (1) It is necessary to collect statistics on the characteristic information of a large amount of input or output data, which consumes the host's memory resources.

[0008] (2) The kernel of the operating system still needs to be modified. For example, the kernel of the operating system includes a virtual file system, a file system, a device driver, etc., and cannot be transparent to the host. The above method has a high dependence on application software and is not transparent to the host, and cannot effectively improve the write amplification problem of the solid-state drive.

[0009] Application Contents

[0010] The embodiments of the present application provide a multi-stream management method for application data and a flash memory device, which solves the technical problem of data write amplification and can store application data with similar life cycles in the storage location corresponding to the hardware stream to optimize the write amplification problem of the flash memory device.

[0011] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0012] In a first aspect, an embodiment of the present application provides a method for managing multiple streams of application data, the method comprising:

[0013] Get the compression ratio of application data;

[0014] Clustering application data according to the compression rate of application data to generate multiple software flows;

[0015] Estimate the lifecycle of each software flow;

[0016] Clustering the software flows according to the life cycle of each software flow to establish a first mapping relationship between each software flow and a hardware flow;

[0017] According to the first mapping relationship, the data corresponding to the software stream is stored in the storage location corresponding to the hardware stream.

[0018] In a second aspect, an embodiment of the present application provides a flash memory device, comprising:

[0019] A compression module, configured to determine a compression algorithm corresponding to the application data according to a data type corresponding to the application data, and compress the application data according to the compression algorithm to obtain a compression ratio of the application data;

[0020] a first clustering module, configured to cluster the application data according to a compression rate of the application data to generate a plurality of software flows;

[0021] Lifecycle estimation module, used to estimate the lifecycle of each software flow;

[0022] a second clustering module, configured to cluster the software flows according to the life cycle of each software flow, so as to establish a first mapping relationship between each software flow and a hardware flow;

[0023] A flash memory conversion module, configured to store the data corresponding to the software stream to a storage location corresponding to the hardware stream according to the first mapping relationship;

[0024] A remapping module, configured to recycle valid data in a storage location corresponding to a hardware stream, so as to move the valid data in the storage location to a preset storage location;

[0025] Flash media, used to store application data.

[0026] In a third aspect, an embodiment of the present application provides a flash memory device, comprising:

[0027] at least one processor; and,

[0028] a memory communicatively connected to at least one processor; wherein,

[0029] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to perform the multi-stream management method of application data as described in the first aspect.

[0030] The beneficial effects of the embodiments of the present application are: different from the related art, the embodiments of the present application provide a multi-stream management method for application data, the method including: obtaining the compression rate of the application data; clustering the application data according to the compression rate of the application data to generate multiple software streams; estimating the life cycle of each software stream; clustering the software streams according to the life cycle of each software stream to establish a first mapping relationship between each software stream and the hardware stream; according to the first mapping relationship, storing the data corresponding to the software stream to the storage location corresponding to the hardware stream.

[0031] By clustering data with similar compression rates, generating multiple software streams, estimating the life cycle of each software stream, mapping software streams with similar life cycles to the same hardware stream, and storing the data corresponding to the software stream in the storage location corresponding to the hardware stream, this application can store application data with similar life cycles in the storage location corresponding to the hardware stream to optimize the write amplification problem of flash memory devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] One or more embodiments are exemplarily described by corresponding drawings, which do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, and unless otherwise stated, the dimensions in the drawings do not constitute proportional limitations.

[0033] FIG1 is a schematic diagram of a connection relationship between a host and a flash memory device provided in an embodiment of the present application;

[0034] FIG2 is a flow chart of a manual diversion mechanism provided in an embodiment of the present application;

[0035] FIG3 is a flow chart of an automatic traffic diversion mechanism based on statistical information provided in an embodiment of the present application;

[0036] FIG4A is a schematic diagram of a process for clustering and dividing application data based on compression ratio according to an embodiment of the present application;

[0037] FIG4B is a schematic diagram of a process for performing two-stage clustering and diversion of application data according to an embodiment of the present application;

[0038] FIG4C is a schematic diagram of a process of a two-stage multi-stream mapping method provided in an embodiment of the present application;

[0039] FIG5 is a flow chart of a multi-stream management method for application data provided in an embodiment of the present application;

[0040] FIG6 is a schematic diagram of a detailed flow chart of step S501 of FIG5 ;

[0041] FIG7 is a schematic diagram showing an example of compression ratios corresponding to different types of data provided by an embodiment of the present application;

[0042] FIG8 is a schematic diagram of a detailed flow chart of step S503 of FIG5 ;

[0043] FIG9 is a schematic diagram of a detailed flow chart of step S5031 of FIG8 ;

[0044] FIG10 is a schematic diagram of a detailed flow chart of step S504 of FIG5 ;

[0045] FIG11 is a schematic diagram of a process for recovering valid data in a physical page according to an embodiment of the present application;

[0046] 12 is a schematic diagram of a process for establishing a second mapping relationship between a hardware flow and a preset hardware flow according to an embodiment of the present application;

[0047] FIG13 is a schematic diagram of a detailed flow chart of step S1101 of FIG11 ;

[0048] FIG14 is a schematic diagram of a detailed flow chart of step S1112 of FIG13 ;

[0049] FIG15 is a schematic diagram of a detailed flow chart of step S1121 of FIG14 ;

[0050] FIG16 is a schematic diagram of the overall structure of a flash memory device provided in an embodiment of the present application;

[0051] FIG17 is a schematic structural diagram of another flash memory device provided in an embodiment of the present application.

[0052] Description of reference numerals: DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0054] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments.

[0055] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.

[0056] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0057] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.

[0058] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0059] Before introducing the embodiments of the present application, a brief introduction to the relevant technologies known to the inventors of the present application is first given to facilitate subsequent understanding of the embodiments of the present application.

[0060] Multi-stream technology is an SSD interface that allows the host to communicate data semantics. Based on the data semantics provided by the host, the multi-stream SSD can optimally place written data. When the data semantics reflect its expected lifespan, the SSD can store data with similar lifespans in the same erase block, minimizing the likelihood that data in the same erase block will expire together. This reduces the number of valid pages migrated during garbage collection and reduces write amplification, ultimately improving SSD performance and lifespan.

[0061] The technical solution of this application is described in detail below with reference to the accompanying drawings:

[0062] Please refer to FIG1 , which is a schematic diagram of a connection relationship between a host and a flash memory device provided in an embodiment of the present application;

[0063] As shown in Figure 1, the flash memory device 100 is used to process the application data of the host 200, classify and compress the application data of the host 200 according to the data type, cluster the compressed application data to generate multiple software streams, and estimate the life cycle of each software stream. Furthermore, a mapping relationship is established between the software stream corresponding to the application data of the host 200 and the hardware stream of the flash memory device 100. According to the mapping relationship, the application data corresponding to the software streams with different life cycles are stored in different hardware streams respectively, so as to realize the separate storage of application data with large differences in life cycles, which can optimize the write amplification problem of data in the flash memory device 100.

[0064] The flash memory device 100 includes a solid-state drive (SSD), a USB flash drive, a memory card, a flash drive (a storage device that integrates a USB flash drive and a memory card) and other storage devices for storing application data.

[0065] Among them, the solid-state drive includes flash memory media. The flash memory media serves as the storage medium of the solid-state drive and is also called flash memory, Flash, NAND Flash memory or Flash particles. The flash memory media serves as a storage unit for application data, system data, etc.

[0066] NAND Flash memory is a non-volatile memory that can permanently store data even without a current supply. NAND Flash memory uses a single transistor as a binary signal storage unit. Its structure is very similar to that of ordinary semiconductor transistors, but the difference is that the single transistor in NAND Flash memory incorporates a floating gate and a control gate. The floating gate is used to store electrons. Its surface is covered by a layer of silicon oxide insulator and is coupled to the control gate via a capacitor. When negative electrons are injected into the floating gate under the action of the control gate, the storage state of the single crystal of NAND Flash memory 12 changes from "1" to "0". When the negative electrons are removed from the floating gate, the storage state changes from "0" to "1". The insulator covering the floating gate surface is used to trap the negative electrons in the floating gate, enabling data storage. That is, the storage unit of NAND Flash memory 12 is a floating gate transistor, which uses the floating gate transistor to store data in the form of charge. The amount of stored charge is related to the voltage applied to the floating gate transistor. Specifically, whether data is stored in the storage unit depends on whether the voltage of the stored charge is greater than a preset voltage threshold Vth.

[0067] Based on the different levels of storage cell voltage, NAND Flash memory can be divided into SLC, MLC, TLC, and QLC. When the storage cell of a NAND Flash memory is of SLC type, a single storage cell only stores one bit of data, 1 or 0. When the voltage of the charge stored in a single storage cell is greater than the preset voltage threshold Vth, it indicates that the data of the storage cell is 1; when the voltage of the charge stored in a single storage cell is less than the preset voltage threshold Vth, it indicates that the data of the storage cell is 0. Since the default data of a storage cell is 1, the data of a storage cell being 0 indicates that the charge stored in the single storage cell is released. When the charge is released to a certain extent, causing the voltage of the floating gate transistor to be less than the preset voltage threshold Vth, the operation of writing data 0 is completed. When the storage cell of the NAND Flash memory 12 is of MLC type, TLC type or QLC type, a single storage cell can store multi-bit data. Taking 2-bit data as an example, multiple preset voltage thresholds Vth are defined by controlling the amount of charge inside the storage cell. For writing data, the amount of charge inside the storage cell is controlled through the charging process so that it falls within different intervals of two adjacent preset voltage thresholds Vth, corresponding to different data 00, 01, 10, 11. For reading data, the current inside the corresponding storage cell is obtained, and then the reading is completed through a series of decoding circuits to parse the data 00, 01, 10, 11 stored in the storage cell.

[0068] A NAND Flash memory consists of at least one chip. Each chip consists of several physical blocks, and each physical block consists of several pages. A physical block is the smallest unit for performing erase operations on a NAND Flash memory, and a page is the smallest unit for performing read and write operations on a NAND Flash memory. The capacity of a NAND Flash memory is equal to the number of physical blocks, the number of pages in a physical block, and the capacity of a page.

[0069] Since the storage unit of NAND Flash memory is implemented by floating gate technology, the charge on the floating gate layer can change the distribution state of electrons, thereby changing the state of the storage unit. However, the charge on the floating gate layer will leak over time. Therefore, it is necessary to regularly erase and refresh the charge to maintain data persistence.

[0070] Please refer to FIG2 , which is a flow chart of a manual diversion mechanism provided in an embodiment of the present application;

[0071] The manual diversion mechanism is mainly based on the developer's experience to judge the corresponding life cycles of different types of application data, that is, the life cycle of the application data is determined before the application data is input into the storage location corresponding to different hardware streams, and the data with different life cycles are mapped to the corresponding hardware streams according to the storage locations corresponding to different hardware stream identifiers. Among them, the life cycle of the application data stored in the storage location corresponding to each hardware stream is different.

[0072] It is understandable that due to the physical characteristics of flash memory devices and the settings of flash memory devices, the data transmission speed and concurrent stream support capabilities of flash memory devices are subject to certain limitations. For example, the data transmission of flash memory devices needs to follow certain protocols and standards, such as SATA, PCIe, etc. These protocols and standards will limit the number of streams supported by flash memory devices. Therefore, the number of hardware streams supported by different flash memory devices is determined by hardware or other technologies.

[0073] As shown in Figure 2, the manual diversion mechanism is further explained by taking the RocksDB key-value database as an example. The RocksDB key-value database runs on the host. Before the data files in the RocksDB key-value database are diverted and stored to the multi-stream solid-state disk, the life cycle of different data files in the RocksDB key-value database is determined based on the developer's experience. By modifying the source code corresponding to the RocksDB key-value database, different "write hints" are set for the software flows corresponding to different files, that is, the hardware flow corresponding to the software flow is manually modified. Among them, the data files in the RocksDB key-value database include the write-ahead log (WAL) and the data in the Memtable. The Memtable (Memory Table) is a data structure in the host memory used to store memory data. Before the data in the Memtable is flushed to the multi-stream solid-state disk, it needs to be sorted through the Sorted Strings Table. The multi-stream solid-state disk reads memory data from the Memtable using a file called a Write Table (SSTable), and stores files with different life cycles in the memory data into different SSTable files respectively. In the process of writing the memory data and the write-ahead log from the application layer of the host to the multi-stream solid-state disk, the write hint will also be sent down layer by layer along with the memory data and the write-ahead log. When the multi-stream solid-state disk receives the write hint, it converts the write hint into the corresponding hardware flow identifier to store the data in the software stream corresponding to the write hint corresponding to the hardware flow identifier into the storage location corresponding to the hardware flow identifier.

[0074] For example: when the write hint is a long life cycle, the write hint is converted into a hardware flow identifier 3, and the storage location corresponding to the hardware flow identifier 3 is used to store data files with a long life cycle; when the write hint is a medium life cycle, the write hint is converted into a hardware flow identifier 2, and the storage location corresponding to the hardware flow identifier 2 is used to store data files with a medium life cycle; when the write hint is a short life cycle, the write hint is converted into a hardware flow identifier 1, and the storage location corresponding to the hardware flow identifier 1 is used to store data files with a short life cycle, for example: log files often need to be refreshed, that is, the life cycle of the log files is short, then the log files are stored in the storage location corresponding to the hardware flow identifier 1; when the write hint is default, the write hint is converted into a hardware flow identifier, and the hardware flow identifier is 0, and the storage location corresponding to the hardware flow identifier 0 is used to store the default data of the host, such as: system configuration information, application configuration information, default files or data, etc.

[0075] The manual diversion mechanism used in FIG2 to store host data in a multi-stream solid-state disk has the following disadvantages: (1) it requires developers to have a deep understanding of host data; (2) it requires the source code of the application to be modified, which is highly invasive to the application; (3) when storing data generated by different applications in a multi-stream solid-state disk, each application needs to be adapted to the multi-stream solid-state disk, which has poor scalability; (4) when there are too many applications, the complexity of multi-stream management of application data increases. Based on this, the embodiment of the present application provides a method for diverting and managing application data with an automatic diversion mechanism to overcome the above disadvantages.

[0076] Specifically, please refer to FIG3 , which is a flow chart of an automatic traffic diversion mechanism based on statistical information provided by an embodiment of the present application;

[0077] As shown in Figure 3, the operating system kernel includes a statistical information collection module, a life cycle estimation module, and a multi-stream mapping module. The statistical information collection module is used to collect statistical characteristic information of application data. For example, the characteristic information includes the access frequency, rewrite interval, and program counter summary of the logical block address corresponding to the application data. The program counter summary includes the values ​​recorded by the program counter, such as the number of times the program is executed, the average execution time of the program, the longest execution time of the program, etc.

[0078] The life cycle estimation module estimates the life cycle of the data of the write request based on the above-mentioned characteristic information, and then automatically maps the data to hardware streams of different priorities according to the estimated life cycle. For example: in the driver layer of the multi-stream solid-state disk, by counting the logical block addresses of the write requests, characteristic information such as the access heat (frequency), recency (recency) and sequentiality (sequentiality) of the data corresponding to the logical block address can be obtained. This characteristic information can reflect the data life cycle to a certain extent. For example, the higher the access heat, the shorter the life cycle of the data, the longer the life cycle of the data that has not been accessed for a long time, and the closer the life cycle of data with adjacent addresses.

[0079] The multi-stream mapping module is used to map data with similar life cycles to the same hardware flow identifier through multi-queue, clustering and other algorithms, so as to store data with similar life cycles in the storage location corresponding to the same hardware flow identifier, thereby realizing automatic data diversion.

[0080] The above-mentioned automatic diversion mechanism can realize automatic data diversion, but it still has the following disadvantages: (1) It needs to be modified in the operating system kernel, such as the virtual file system, file system, block layer, device driver and other modules, which is relatively complex. (2) It needs to count a large amount of I / O feature information, which consumes a large amount of host memory resources. (3) It does not consider the write amplification problem of multi-stream solid-state drives. (4) The diversion strategy is not evaluated.

[0081] Based on this, the present application provides a multi-stream management method for application data, which does not require statistics of a large amount of I / O feature information and saves the host's memory resources.

[0082] Please refer to FIG4A , which is a schematic diagram of a process for clustering and dividing application data based on compression ratio according to an embodiment of the present application;

[0083] As shown in FIG. 4A , the flash memory device 100 includes a compression module 101 , a first clustering module 102 , a flash memory conversion module 105 , and a flash memory medium 106 .

[0084] The compression module 101 receives application data sent by the host and compresses the application data. The application data includes multiple types of data files, and different types of data files have different corresponding compression rates.

[0085] The first clustering module 102 clusters the application data according to the compression ratios corresponding to different types of data files, and maps data files with similar compression ratios to the same hardware flow to establish a mapping relationship between the data files and the hardware flow.

[0086] The flash conversion module 105 maps the data file to the hardware stream and stores the data file in the flash medium 106, wherein the flash medium 106 includes multiple super blocks, each hardware stream corresponds to a super block one by one, and each hardware stream corresponds to a hardware stream identifier. The data file is stored in the flash medium 106, that is, according to the hardware stream identifier corresponding to the data file, the data file is stored in the super block corresponding to the hardware stream identifier corresponding to the data file.

[0087] In an embodiment of the present application, application data with different compression rates are clustered by a first clustering module to achieve multi-stream mapping of a solid-state flash memory device. However, clustering and splitting application data based solely on compression rate cannot optimize the write amplification problem of the flash memory device.

[0088] In view of this, an embodiment of the present application provides a two-stage clustering and diversion method for application data, which can further cluster application data with similar life cycles in the same hardware flow based on the compression rate to optimize the write amplification problem of the flash memory device.

[0089] Specifically, please refer to FIG4B , which is a schematic diagram of a process for performing two-stage clustering and diversion of application data according to an embodiment of the present application;

[0090] As shown in FIG4B , the flash memory device 100 includes a compression module 101 , a first clustering module 102 , a lifecycle estimation module 103 , a second clustering module 104 , a flash memory conversion module 105 , and a flash memory medium 106 .

[0091] The compression module 101 receives application data sent by the host and compresses the application data. The application data includes multiple types of data files, and different types of data files have different corresponding compression rates.

[0092] The first clustering module 102 clusters the application data according to the compression ratios corresponding to different types of data files, and generates multiple software streams, wherein data files with similar compression ratios are mapped into the same software stream.

[0093] The lifecycle estimation module 103 estimates the lifecycle of each software flow based on the compression ratio of the data file.

[0094] The second clustering module 104 clusters the software flows based on their lifecycles, mapping software flows with similar lifecycles to the same hardware flow. This establishes a mapping relationship between software flows and hardware flows, where each hardware flow corresponds to multiple software flows. It is understood that the number of hardware flows depends on the flash memory device type and is typically smaller than the number of software flows. Therefore, mapping software flows to hardware flows ensures that the software flows match the hardware flows.

[0095] The flash memory conversion module 105 stores the data in the software stream into the super block corresponding to the hardware stream according to the mapping relationship between the software stream and the hardware stream.

[0096] In an embodiment of the present application, data with similar compression rates are clustered through a first clustering module to map application data with the same compression rate to the same software stream. Furthermore, software streams are clustered through a second clustering module to map software streams with similar life cycles to the same hardware stream. Clustering application data in two stages can place application data with similar life cycles together, avoid occupying multi-stream resources, and solve the problem of mismatch between software streams and hardware streams.

[0097] The above-mentioned multi-stream management method places application data with similar life cycles in the same hardware stream, which can effectively improve the write amplification problem of flash memory devices. However, the same hardware stream may still contain data with similar compression rates but longer life cycles, resulting in multiple rewrites during subsequent garbage collection, causing write amplification problems. To further improve this problem, an embodiment of the present application provides a two-stage multi-stream mapping method, which can effectively improve the write amplification problem during the garbage collection process.

[0098] Please refer to FIG4C , which is a flowchart of a two-stage multi-stream mapping method provided in an embodiment of the present application;

[0099] As shown in Figure 4C, Figure 4C is improved on the basis of Figure 4B. In order to solve the problem of write amplification caused by the garbage collection process mentioned in Figure 4B, a remapping module 107 is added to the flash memory device 100. The remapping module 107 is used to map all hardware streams to a new hardware stream. During the garbage collection process, the physical pages in the super block corresponding to each hardware stream with a longer life cycle and valid data are written back to the new hardware stream, which can effectively optimize the write amplification problem.

[0100] In an embodiment of the present application, data with similar compression rates are clustered through a first clustering module to map application data with the same compression rate to the same software stream. Further, software streams are clustered through a second clustering module to map software streams with similar life cycles to the same hardware stream. Clustering application data in two stages can place application data with similar life cycles together, avoid occupying multi-stream resources, and solve the problem of mismatch between software streams and hardware streams. At the same time, data with longer life cycles and validity in the data blocks corresponding to the hardware streams are remapped through a remapping module, which can effectively improve the write amplification problem of flash memory devices.

[0101] Please refer to FIG5 , which is a flowchart of a multi-stream management method for application data provided by an embodiment of the present application;

[0102] As shown in FIG5 , the multi-stream management method for application data includes:

[0103] Step S501: Obtaining the compression ratio of application data;

[0104] In an embodiment of the present application, there is an intrinsic correlation between the compression rate of application data and the life cycle of application data. For example, in the RocksDB and Cassandran application scenarios, the compression rate of WAL files is low, but the life cycle of WAL files is short, and the compression rate of SSTable files is high, but the life cycle of SSTable files is long. That is, in general, files with low compression rates have short life cycles, and files with high compression rates have long life cycles. Therefore, the compression rate can be used to indirectly characterize the life cycle, thereby firstly diverting application data based on the compression rate, and preliminarily mapping application data with similar life cycles into one software stream.

[0105] Specifically, the application data is compressed using a compression algorithm to obtain a compression ratio of the application data, wherein the compression algorithm includes LZ4, ZSTD, snappy, DEFLATE and other compression algorithms.

[0106] In the embodiment of the present application, by compressing the application data, the total amount of data written to the flash memory device can be reduced.

[0107] Please refer to FIG6 again, which is a detailed flowchart of step S501 in FIG5 ;

[0108] As shown in FIG6 , step S501 includes:

[0109] Step S5011: Determine the compression algorithm corresponding to the application data according to the data type corresponding to the application data;

[0110] In an embodiment of the present application, application data has multiple data types. In order to improve the compression ratio and compression rate of application data and improve the write performance of the flash memory device, different compression algorithms are selected for application data of different data types. For example, log files and temporary data are data that are transmitted and stored quickly. The LZ4 algorithm can be used to process log files and temporary data faster. For large data sets, such as pictures and videos, the ZSTD algorithm can be used to process large data sets in batches to improve compression efficiency.

[0111] Specifically, the compression algorithm corresponding to the application data is determined based on the data type corresponding to the application data, where the data type includes structured data, unstructured data, real-time data, text data, log data, etc. It is understandable that different compression algorithms have different speeds and efficiencies when processing different types of data. For example, the LZ4 algorithm is more suitable for processing data that requires fast transmission and storage, such as log files and temporary data; the ZSTD algorithm is more suitable for processing large data sets, such as video, audio, and images; and the snappy algorithm is suitable for processing file-type data, such as text files and binary files.

[0112] Please refer to FIG7 again, which is a schematic diagram illustrating an example of compression ratios corresponding to different types of data provided by an embodiment of the present application;

[0113] As shown in Figure 7, the compression ratio of each data block in the first database file is basically distributed around 0.17. The first database file includes the nci_database database file. The data compression ratio of the web page file is basically distributed around 0.4. The web page file includes the webster web page. The data compression ratio of the text file is basically distributed around 0.5. The text file includes the dickens_txt text file. The data compression ratio of each data block in the second database file is basically distributed around 0.74. The second database file includes the osdb_database database file. The data compression ratio of the image file is basically distributed around 0.82. The image file includes the xray_image picture. The data compression ratio of the binary file is basically distributed around 0.85. The binary file includes the sao_bin text file.

[0114] From the above data, we can see that the compression rates of data of different data types are different. The compression rates of some different data types vary greatly, such as the data compression rate difference between the nci_database database file and the xray_image picture. However, the compression rates of some different data types vary little, such as the data compression rate difference between the osdb_database database file and the sao_bin text file.

[0115] In an embodiment of the present application, in the process of compressing data, the selection of a suitable compression algorithm needs to consider the compression ratio of the data, the performance requirements of the flash memory device, and the life cycle of the data. For example: whether over-compression of data will affect I / O performance. For example: for data that is frequently read (data with a short life cycle), lightweight compression can decompress the data faster when reading. For example, in the Cassandra scenario, the write-ahead log is frequently used, so there is no need to compress the write-ahead log. SSTable is a data structure in Cassandra. SSTable is divided into multiple layers. The heat of the data stored in each layer of SSTable is different. For example, the data in the 0th layer SSTable has the highest heat, and the data in the nth layer SSTable (the last layer) has the lowest heat. A lightweight compression algorithm is used for the data in the 0th layer SSTable, and a heavyweight compression algorithm is used for the data in the nth layer SSTable. The higher the heat of the data, the more times the data is accessed and the shorter the life cycle of the data. The lower the heat of the data, the fewer times the data is accessed and the longer the life cycle of the data.

[0116] In an embodiment of the present application, a lightweight compression algorithm is used for data with high popularity, and the compression rate of the data with high popularity is low. A heavyweight compression algorithm is used for data with low popularity, and the compression rate of the data with low popularity is high. That is, the life cycle corresponding to the data with a low compression rate is short, and the life cycle corresponding to the data with a high compression rate is long. The life cycle of the data is indirectly represented by the compression rate, so as to further divert and manage the application data according to the compression rate of the data.

[0117] Step S5012: compressing the application data according to a compression algorithm to obtain a compression ratio of the application data;

[0118] Specifically, the application data is compressed according to a compression algorithm determined by the data type of the application data to obtain a compression ratio of the compressed application data.

[0119] In an embodiment of the present application, different compression algorithms are used for data of different data types, which can avoid the additional delay caused by the compression algorithm on the write critical path, thereby avoiding the impact of the compression algorithm on the write performance of the flash memory device.

[0120] Step S502: clustering the application data according to the compression rate of the application data to generate multiple software flows;

[0121] Specifically, the application data is clustered based on its compression rate. That is, based on the compression rate of the application data, a clustering algorithm is used to classify application data with similar compression rates into the same category of data to generate multiple software flows. The clustering algorithm includes a K-means algorithm, a DBSCAN algorithm, a hierarchical clustering algorithm, a spectral clustering algorithm, a collaborative filtering algorithm, a random forest algorithm, a Gaussian process regression clustering algorithm, and the like. The number of software flows output by the clustering algorithm (the number of categories) can be set according to specific needs. It should be noted that whether the compression rates are similar can be determined by whether the compression rates of the application data are within the same preset compression rate range. For example, the preset compression rate range is [0.1, 0.2]. Assuming that the compression rate of a text file is 0.17 and the compression rate of an image is 0.1, the compression rates of the text file and the image are both within the preset compression rate range, and the compression rates of the text file and the image are similar.

[0122] Please refer to FIG7 again, which is a schematic diagram illustrating an example of compression ratios corresponding to different types of data provided by an embodiment of the present application;

[0123] As shown in Figure 7, the compression ratio of each data block in the first database file (such as the nci_database database file) is basically distributed around 0.17, the data compression ratio of web page files (such as the webster web page) is basically distributed around 0.4, the data compression ratio of text files (such as the dickens_txt text file) is basically distributed around 0.5, the data compression ratio of each data block in the second database file (such as the osdb_database database file) is basically distributed around 0.74, the data compression ratio of image files (such as the xray_image picture) is basically distributed around 0.82, and the data compression ratio of binary files (such as the sao_bin binary file) is basically distributed around 0.85.

[0124] It can be seen from the above data that the compression rates of data of files of the same type are basically the same, so the data of files of the same type are clustered in the same software stream. The data compression rates of the image file in Figure 7, the data compression rates of each data block in the second database file, and the data compression rates of the binary file are all high and close, that is, the corresponding life cycles of the image file, the second database file, and the binary file are relatively close, so the image file, the second database file, and the binary file are clustered in the same software stream.

[0125] In an embodiment of the present application, data with different compression rates are clustered through a clustering algorithm, so that data of the same file type are clustered in the same software stream, and data of different file types are clustered in different software streams, and application data with similar life cycles are clustered in the same software stream, and application data with distant life cycles are clustered in different software streams, so as to realize multi-stream identification of flash memory devices.

[0126] Step S503: estimating the life cycle of each software flow;

[0127] In the embodiments of the present application, the compression rate is only an indirect representation of the lifecycle of the application data. The compression rates of the application data corresponding to different software streams vary significantly. Software streams with these significantly different compression rates may have similar lifecycles. In this case, distributing software streams with similar lifecycles into different hardware streams would waste multi-stream resources. Therefore, it is necessary to estimate the lifecycle of each software stream so that software streams with potentially similar lifecycles can be mapped into the same hardware stream to conserve multi-stream resources. It should be noted that the number of hardware streams is limited by the flash memory device and will not be excessive.

[0128] In the embodiment of the present application, each software stream corresponds to at least one logical block address, the logical block address includes a plurality of logical pages, and the data corresponding to the software stream is stored in the logical pages.

[0129] Specifically, the time interval between two writes to each logical page is recorded, the life cycle of each logical page is estimated based on the time interval corresponding to each logical page, and the average life cycle of all logical pages corresponding to each software flow is calculated to estimate the life cycle corresponding to each software flow.

[0130] Please refer to FIG8 again, which is a detailed flowchart of step S503 in FIG5 ;

[0131] As shown in FIG8 , step S503 includes:

[0132] Step S5031: Estimate the lifecycle of all logical pages corresponding to each software flow;

[0133] Specifically, the time interval between the two most recent adjacent write operations on each logical page is recorded, and the life cycle of all logical pages corresponding to each software flow is estimated based on the time interval between the two most recent adjacent write operations on each logical page.

[0134] Please refer to FIG9 again, which is a detailed flowchart of step S5031 in FIG8 ;

[0135] As shown in FIG9 , step S5031 includes:

[0136] Step S5311: Counting the timestamps of each logical page access, wherein the timestamps include a first timestamp and a second timestamp, and the first timestamp and the second timestamp are adjacent timestamps;

[0137] Specifically, the timestamp of each logical page access is counted, that is, the timestamp of each logical page write operation is recorded, and the timestamps of the two most recent adjacent write operations are taken. The timestamps include a first timestamp and a second timestamp. The first timestamp and the second timestamp are adjacent timestamps, and the first timestamp is smaller than the second timestamp.

[0138] Step S5312: Calculate the difference between the first timestamp and the second timestamp to determine the life cycle of each logical page;

[0139] Specifically, the difference between the first timestamp and the second timestamp is calculated, that is, the value obtained by subtracting the first timestamp from the second timestamp is the life cycle of the logical page. The life cycle of each logical page can be determined using the above calculation method.

[0140] Step S5032: Calculate the average value of the lifecycles of all logical pages corresponding to each software flow to determine the lifecycle of each software flow;

[0141] Specifically, calculate the average life cycle of all logical pages corresponding to each software flow, add up the life cycles of all logical pages corresponding to each software flow to obtain the total life cycle, divide the total life cycle by the number of logical pages corresponding to the software flow, and obtain the average life cycle of all logical pages. This average value is the life cycle of the software flow.

[0142] It should be noted that the life cycle of the software flow can also be determined by other methods, which are not limited here.

[0143] Step S504: clustering the software flows according to the life cycle of each software flow to establish a first mapping relationship between each software flow and the hardware flow;

[0144] In an embodiment of the present application, each hardware stream corresponds to a preset life cycle. For example, assuming there are three hardware streams, the storage location corresponding to the first hardware stream is used to store data with a life cycle of one day, the storage location corresponding to the second hardware stream is used to store data with a life cycle of seven days, and the storage location corresponding to the third hardware stream is used to store data with a life cycle of one month. Before establishing the first mapping relationship between each software stream and the hardware stream, it is necessary to determine whether the life cycle of the current software stream is within the preset period corresponding to the current hardware stream. If the life cycle of the current software stream is five days, the current software stream is mapped to the second hardware stream.

[0145] In the embodiment of the present application, each hardware flow corresponds to a hardware flow number one by one. For example, there are three hardware flows, and the three hardware flow numbers are 0, 1, and 2. It should be noted that the preset life cycle corresponding to the hardware flow number is not affected by the order of the hardware flow numbers. For example, the preset life cycle of the hardware flow number 1 can be greater than the preset life cycle of the hardware flow number 0, and the preset life cycle of the hardware flow number 1 can also be less than the preset life cycle of the hardware flow number 0. The specific setting is based on actual needs and is not limited here.

[0146] Specifically, according to the life cycle of each software stream, the software streams are clustered using a clustering algorithm to determine whether the life cycle of each software stream is within the preset period corresponding to the hardware stream. If it is within the preset period corresponding to the hardware stream, a first mapping relationship between each software stream and the hardware stream is established, that is, each software stream is mapped to the hardware stream number corresponding to a hardware stream. For example, assuming that there are currently 10 software streams and 4 hardware streams, and the life cycles of software streams 1-4 are within the preset life cycle of the hardware stream with hardware stream number 0, software streams 1-4 are mapped to the hardware stream with hardware stream number 0. The number of clusters of the clustering algorithm is limited according to the number of hardware streams. For example, if there are currently 8 hardware streams, the number of clusters of the clustering algorithm is set to 8.

[0147] In the embodiment of the present application, by establishing a first mapping relationship between each software flow and a hardware flow, software flows with close life cycles are grouped into the same hardware flow, and one hardware flow corresponds to multiple software flows.

[0148] In an embodiment of the present application, by estimating the life cycle of each software stream, software streams with large differences in compression rates and close life cycles are mapped to the same hardware stream, thereby avoiding the occupation of multi-stream resources. At the same time, multiple software streams are mapped to one hardware stream, solving the problem of mismatch between the number of software streams and the number of hardware streams.

[0149] Please refer to FIG10 again, which is a detailed flowchart of step S504 in FIG5 ;

[0150] As shown in FIG10 , the step S504 includes:

[0151] Step S5041: Determine whether the life cycle corresponding to the current software flow is within the preset life cycle corresponding to the current hardware flow;

[0152] Specifically, determine whether the life cycle corresponding to the current software flow is within the preset life cycle corresponding to the current hardware flow. If it is within the preset life cycle corresponding to the current hardware flow, jump to step S5042; otherwise, jump to step S5044.

[0153] Step S5042: Mapping the software flow to the hardware flow corresponding to the preset lifecycle to determine a first mapping relationship between the software flow and the hardware flow;

[0154] Specifically, if the life cycle corresponding to the current software flow is within the preset life cycle corresponding to the current hardware flow, the current software flow is mapped to the current hardware flow through a clustering algorithm to determine the mapping relationship between the current software flow and the current hardware flow, and jump to step S5043.

[0155] Step S5043: Move to the next software flow;

[0156] Specifically, if the mapping relationship between the current software flow and the current hardware flow has been determined, the process moves to the next software flow and continues to execute step S5041 until all software flows are clustered into corresponding hardware flows.

[0157] Step S5044: Do not map the current software stream;

[0158] Specifically, if the life cycle corresponding to the current software flow is not within the preset life cycle corresponding to the current hardware flow, jump to step S5045.

[0159] Step S5045: moving to the next hardware flow;

[0160] Specifically, if it is determined that the life cycle corresponding to the current software flow is not within the preset life cycle corresponding to the current hardware flow, the process moves to the next hardware flow and continues to execute step S5041 until all software flows are clustered into corresponding hardware flows.

[0161] In the embodiment of the present application, a first mapping relationship between software streams and hardware streams is established to implement multi-stream management of application data.

[0162] Step S505: storing the data corresponding to the software stream to the storage location corresponding to the hardware stream according to the first mapping relationship;

[0163] Specifically, each hardware stream corresponds to a super block one by one. According to the first mapping relationship between the software stream and the hardware stream, the data corresponding to the software stream is stored in the storage location corresponding to the hardware stream, that is, the data corresponding to the software stream is stored in the super block corresponding to the hardware stream.

[0164] In an embodiment of the present application, data with similar compression rates are clustered to generate multiple software streams, and the life cycle of each software stream is estimated to map software streams with similar life cycles to the same hardware stream, and the data corresponding to the software stream is stored in the storage location corresponding to the hardware stream. This allows application data with similar life cycles to be stored in the storage location corresponding to the hardware stream, and can reduce the number of times data is moved on the flash memory device, thereby optimizing the write amplification problem of the flash memory device.

[0165] In an embodiment of the present application, the diversion management of the host's application data is directly implemented inside the flash memory device without relying on the host's application program. The host is only responsible for transmitting data, thereby achieving a host transparent effect.

[0166] Please refer to FIG11 again, which is a schematic diagram of a process for recovering valid data in a physical page according to an embodiment of the present application;

[0167] As shown in FIG11 , the valid data in the physical page is recycled, including:

[0168] Step S1101: Reclaiming valid data in a physical page corresponding to a storage location corresponding to a hardware stream;

[0169] In an embodiment of the present application, when the garbage collection mechanism in the flash memory device is started, the storage space that is no longer in use will be automatically cleaned up to ensure the continuity and availability of the storage space.

[0170] In an embodiment of the present application, during the garbage collection process, a small amount of valid data with a long life cycle may still exist in the storage location corresponding to the hardware stream. In this case, the valid data with a long life cycle in the storage location corresponding to the hardware stream needs to be moved to another storage location before it can be erased. Since the valid data with a long life cycle needs to be rewritten each time garbage collection is performed, a write amplification problem is caused. Therefore, it is necessary to remap this type of data to a hardware stream with a longer life cycle to reduce the number of times this type of data is moved.

[0171] Specifically, the storage location corresponding to the hardware stream is a super block, which is composed of multiple physical blocks. Each physical block includes multiple physical pages. The physical pages are used to store data. The valid data in the storage location corresponding to the hardware stream is recycled, that is, the valid data in the physical pages corresponding to the storage location corresponding to the hardware stream is recycled.

[0172] Please refer to FIG. 12 again, which is a schematic diagram of a process for establishing a second mapping relationship between a hardware flow and a preset hardware flow according to an embodiment of the present application;

[0173] As shown in FIG12 , establishing a second mapping relationship between the hardware flow and the preset hardware flow includes:

[0174] Step S1201: establishing a second mapping relationship between each hardware flow and a preset hardware flow, wherein the preset hardware flow corresponds to a preset storage location;

[0175] Specifically, before recycling the valid data in the physical page corresponding to the storage location corresponding to the hardware stream, a second mapping relationship between each hardware stream and the preset hardware stream is established in advance. The preset hardware stream is used to store the valid data with a longer life cycle in the hardware stream to the preset storage location corresponding to the preset hardware stream.

[0176] For example, for the physical pages of a normal hardware stream, when the write-back frequency exceeds a threshold, the data will be written to hardware stream 0, which is numbered as hardware stream 0. Hardware stream 0 is used to store "colder" data. If the physical page of hardware stream 0 is written back again during garbage collection, the data in the physical page will be written to hardware stream 1, which is numbered as hardware stream 1. Hardware stream 1 is used to store "extremely cold" data. If the physical page of hardware stream 1 is written back again during garbage collection, the data will still be written back to hardware stream 1. If the physical page of hardware stream 1 is overwritten, the new data will be written to the dedicated hardware stream 0. If the physical page of hardware stream 0 is overwritten, the new data will be written to the original hardware stream, that is, the hardware stream given by the second clustering module.

[0177] By storing cold data in a preset storage location corresponding to a preset hardware flow, where cold data is data with a long life cycle or data that is not frequently used, the embodiment of the present application can effectively improve the write amplification problem caused by the mixing of cold and hot data.

[0178] Please refer to FIG13 again, which is a detailed flowchart of step S1001 in FIG10 ;

[0179] As shown in FIG13 , step S1101 includes:

[0180] Step S1111: Counting feedback information of the physical page corresponding to the storage location corresponding to each hardware flow;

[0181] Specifically, feedback information for the physical page corresponding to the storage location corresponding to each hardware stream is collected. This feedback information includes the write-back frequency of the physical page and the hardware stream number to which the physical page belongs. The write-back frequency of the physical page indicates how often the physical page is used in write and erase operations. Write-back refers to the fact that valid data on the physical page still exists during garbage collection and needs to be written to a new location.

[0182] In an embodiment of the present application, the improvement effect of the write amplification of each hardware flow can be evaluated based on the feedback information of the physical page. For example: the write-back frequency determined by the write-back frequency of a hardware flow is greater than or equal to the frequency threshold, indicating that there is valid data with a long life cycle in the hardware flow, then the valid data with a long life cycle in the hardware flow needs to be moved to the preset hardware flow to improve the write amplification problem of the hardware flow.

[0183] Step S1112: moving the valid data in the storage location to a preset storage location according to the feedback information;

[0184] Specifically, the feedback information includes the write-back frequency of the physical page and the hardware stream number where the physical page is located. Based on the write-back frequency of the physical page, it can be determined whether to move the valid data in the physical page to the preset storage location, or based on the hardware stream number where the physical page is located, it can be determined whether to move the valid data in the physical page to the preset storage location according to the state machine.

[0185] Please refer to FIG. 14 again, which is a detailed flowchart of step S1112 of FIG. 13 ;

[0186] As shown in FIG14 , step S1112 includes:

[0187] Step S1121: determining the write-back frequency of the physical page within a period of time according to the write-back frequency of the physical page, and moving valid data in the physical page corresponding to the storage location to a preset storage location according to the write-back frequency;

[0188] Specifically, based on the write-back frequency of the physical page, the write-back frequency of the physical page within a period of time is determined, that is, the number of times the physical page is referenced within a period of time. Based on the write-back frequency, that is, the number of times the physical page is referenced within a period of time, it is determined whether the number of times the physical page is referenced is greater than or equal to the frequency threshold. If the number of times the physical page is referenced is greater than or equal to the frequency threshold, it means that the physical page is frequently used. When garbage collection is performed on the valid data in the physical page, the valid data in the physical page is remapped to the same hardware stream. If the number of times the physical page is referenced is less than the frequency threshold, the valid data of the physical page is re-moved to the preset storage location corresponding to the preset hardware stream. The preset storage location includes a physical block. The physical block includes multiple physical pages. The valid data of the physical page is moved to the physical page in the physical block. Wherein, a period of time refers to a preset time threshold. The preset time threshold is set according to specific needs. For example, it is set to a value such as 0.5s, 1s, etc., which is not limited here.

[0189] In the embodiment of the present application, by counting the feedback information of the physical pages, the write amplification improvement effect of the flash memory device can be effectively evaluated.

[0190] Step S1122: according to the hardware stream number of the physical page, move the valid data in the storage location to the preset storage location according to the state machine;

[0191] In an embodiment of the present application, the state machine is used to record the state of a physical page, and the state of a physical page includes an idle state, a used state, and a recycled state.

[0192] Specifically, according to the hardware stream number where the physical page is located, the state of the physical page in the super block corresponding to the hardware stream number is searched in the state machine. If the state of the physical page is a recycled state, the valid data in the physical page corresponding to the recycled state in the storage location is moved to the preset storage location.

[0193] Please refer to FIG15 again, which is a schematic diagram of a detailed flow chart of step S1121 in FIG14 ;

[0194] As shown in FIG15 , step S1121 includes:

[0195] Step S1123: determining whether the write-back frequency of the physical page is greater than or equal to the frequency threshold;

[0196] Specifically, it is determined whether the write-back frequency of the physical page is greater than or equal to the frequency threshold. If the write-back frequency of the physical page is greater than or equal to the frequency threshold, the process jumps to step S1124; otherwise, the process jumps to step S1125.

[0197] Step S1124: moving valid data in the physical page corresponding to the storage location to a preset storage location according to the second mapping relationship;

[0198] Specifically, for the physical pages of a normal hardware stream, when the number of write-backs exceeds a threshold, the data will be written to dedicated hardware stream 0, numbered hardware stream 0, which is used to store "relatively cold" data. If the physical page of hardware stream 0 is written back again during garbage collection, the data will be written to hardware stream 1, numbered hardware stream 1, which is used to store "extremely cold" data. If the physical page of dedicated hardware stream 1 is written back again during garbage collection, the data will still be written back to dedicated hardware stream 1. If the physical page of hardware stream 1 is overwritten, the new data will be written to dedicated hardware stream 0. If the physical page of hardware stream 0 is overwritten, the new data will be written to the original hardware stream, i.e., the hardware stream provided by the second clustering module.

[0199] It is understood that "relatively cold" data refers to data that is infrequently used but occasionally used, while "extremely cold" data refers to data that is rarely used. For example, data can be determined to be first-cold data or second-cold data based on the time interval between the two most recent write operations on each logical page, where first-cold data is "relatively cold" data and second-cold data is "extremely cold" data. For example, if the time interval exceeds a first time threshold, the data is determined to be first-cold data. Further, if the time interval exceeds a second time threshold, the data is determined to be second-cold data, where the first time threshold is less than the second time threshold.

[0200] Step S1125: Do not move the data of the current physical page;

[0201] Specifically, if the write-back frequency of a physical page is less than the frequency threshold, it means that the physical page is frequently referenced, and there is no need to move the data of the current physical page.

[0202] In an embodiment of the present application, by remapping the hardware stream to the preset hardware stream, the data with a long life cycle and valid in the hardware stream is moved to the preset storage location corresponding to the preset hardware stream, so as to reduce the number of times the data with a long life cycle is moved, thereby improving the write amplification problem of the flash memory device.

[0203] In an embodiment of the present application, data with similar compression rates are clustered to generate multiple software streams, and the life cycle of each software stream is estimated to map software streams with similar life cycles to the same hardware stream, and the data corresponding to the software stream is stored in the storage location corresponding to the hardware stream. The present application can store application data with similar life cycles in the storage location corresponding to the hardware stream to optimize the write amplification problem of the flash memory device.

[0204] Please refer to FIG16 , which is a schematic diagram of the overall structure of a flash memory device provided in an embodiment of the present application;

[0205] As shown in FIG16 , the flash memory device 100 includes a compression module 101 , a first clustering module 102 , a lifecycle estimation module 103 , a second clustering module 104 , a flash memory conversion module 105 , a flash memory medium 106 , and a remapping module 107 .

[0206] The compression module 101 is connected to the first clustering module 102 and is configured to obtain application data, determine a compression algorithm corresponding to the application data based on the data type of the application data, and compress the application data based on the compression algorithm to obtain a compression ratio of the application data. Compression algorithms include LZ4, ZSTD, snappy, DEFLATE, and other compression algorithms.

[0207] The first clustering module 102 is connected to the compression module 101 and the life cycle estimation module 103, and is used to cluster the application data according to the compression rate of the application data to generate multiple software flows. The first clustering module 102 is used to run a clustering algorithm, which includes a K-means algorithm, a DBSCAN algorithm, a hierarchical clustering algorithm, a spectral clustering algorithm, a collaborative filtering algorithm, a random forest algorithm, a Gaussian process regression clustering algorithm, and other clustering algorithms.

[0208] Preferably, the first clustering module 102 includes a K-means clustering module, where K is the number of software flows. The K-means clustering module is configured to run a K-means algorithm, setting the number of clusters to the number of software flows, to cluster the application data and generate K software flows. When initializing the K-means clustering algorithm, the compression ratio (center point) of the software flow is preset to an initial value. However, during the operation of the K-means clustering algorithm, the compression ratio (center point) of the software flow itself will adaptively change or update according to the compression ratio of the logical page data of the software flow.

[0209] It is understandable that the number of software streams is determined based on the developer's experience and is usually a large value, such as 32. Based on the data compression rate, user data, that is, write data, is divided into 32 clusters, that is, 32 software streams.

[0210] The lifecycle estimation module 103 is connected to the first clustering module 102 and the second clustering module 104 and is used to estimate the lifecycle of each software flow.

[0211] The second clustering module 104 is connected to the lifecycle estimation module 103 and the flash conversion module 105 and is configured to cluster the software flows according to the lifecycle of each software flow to establish a first mapping relationship between each software flow and a hardware flow.

[0212] Preferably, the second clustering module 104 includes a K-means clustering module, where K is the number of hardware flows. The K-means clustering module is used to run the K-means algorithm, setting the number of clusters to the number of hardware flows to cluster the application data and generate K hardware flows. When initializing the K-means clustering algorithm, the life cycle (center point) of the hardware flow is preset as an initial value. However, during the operation of the K-means clustering algorithm, the life cycle (center point) of the hardware flow itself will adaptively change or update with the life cycle of the software flow.

[0213] It's understandable that the number of hardware flows is smaller than the number of software flows. For example, if the hardware only supports eight flows, the 32 software flows must be clustered into the eight hardware flows. The second clustering still uses the K-means algorithm, with cluster distances calculated based on the estimated flow lifetimes. The second clustering output is eight flows, which matches the actual hardware.

[0214] Flash translation module 105 is connected to second clustering module 104 and flash memory medium 106 and is configured to store data corresponding to the software stream in a storage location corresponding to the hardware stream based on the first mapping relationship. Flash translation module 105 includes a flash translation layer (FTL), and various functions of flash translation module 105 are implemented by the flash translation layer (FTL).

[0215] The flash memory medium 106 is connected to the flash memory conversion module 105 and the remapping module 107 and is used to store application data.

[0216] The remapping module 107 is connected to the flash conversion module 105 and the flash medium 106 and is used to recycle the valid data in the storage location corresponding to the hardware flow to move the valid data in the storage location to a preset storage location.

[0217] In an embodiment of the present application, the compression rate of the application data is obtained; the application data is clustered according to the compression rate of the application data to generate multiple software flows; the life cycle of each software flow is estimated; the software flows are clustered according to the life cycle of each software flow to establish a first mapping relationship between each software flow and the hardware flow; according to the first mapping relationship, the data corresponding to the software flow is stored in the storage location corresponding to the hardware flow.

[0218] By clustering data with similar compression rates, generating multiple software streams, estimating the life cycle of each software stream, mapping software streams with similar life cycles to the same hardware stream, and storing the data corresponding to the software stream in the storage location corresponding to the hardware stream, this application can store application data with similar life cycles in the storage location corresponding to the hardware stream to optimize the write amplification problem of flash memory devices.

[0219] Please refer to FIG17 , which is a schematic diagram of the structure of another flash memory device provided in an embodiment of the present application;

[0220] As shown in Figure 17, the flash memory device 100 includes one or more processors 108 and a memory 109. Figure 17 takes one processor 101 as an example.

[0221] The processor 108 and the memory 109 may be connected via a bus or other means. FIG17 takes the bus connection as an example.

[0222] The processor 108 is used to provide computing and control capabilities to control the flash memory device 100 to perform corresponding tasks, for example, to control the flash memory device 100 to perform the multi-stream management method for application data in any of the above-mentioned method embodiments, the multi-stream management method for application data comprising: obtaining a compression ratio of the application data; clustering the application data according to the compression ratio of the application data to generate multiple software streams; estimating the life cycle of each software stream; clustering the software streams according to the life cycle of each software stream to establish a first mapping relationship between each software stream and a hardware stream; and storing the data corresponding to the software stream in a storage location corresponding to the hardware stream according to the first mapping relationship.

[0223] By clustering data with similar compression rates, generating multiple software streams, estimating the life cycle of each software stream, mapping software streams with similar life cycles to the same hardware stream, and storing the data corresponding to the software stream in the storage location corresponding to the hardware stream, this application can store application data with similar life cycles in the storage location corresponding to the hardware stream to optimize the write amplification problem of flash memory devices.

[0224] The processor 108 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or any combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0225] The memory 109, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the multi-stream management method for application data in the embodiment of the present application. The processor 108 can implement the multi-stream management method for application data in any of the above method embodiments by running the non-transitory software programs, instructions and modules stored in the memory 109. Specifically, the memory 109 may include a volatile memory (VM), such as a random access memory (RAM); the memory 109 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD) or other non-transitory solid-state storage device; the memory 109 may also include a combination of the above types of memories.

[0226] The memory 109 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 109 may optionally include a memory remotely located relative to the processor 108, and such remote memory may be connected to the processor 108 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0227] One or more modules are stored in the memory 109 and, when executed by one or more processors 108 , perform the multi-stream management method for application data in any of the above method embodiments, for example, perform the steps shown in FIG. 5 described above.

[0228] The present application also provides a non-volatile computer-readable storage medium, such as a memory including program code, which can be executed by a processor to implement the multi-stream management method of application data in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0229] The present application also provides a computer program product comprising one or more program codes stored in a non-volatile computer-readable storage medium. A processor of a flash memory device reads the program codes from the non-volatile computer-readable storage medium and executes the program codes to perform the steps of the multi-stream management method for application data provided in the above-described embodiment.

[0230] Those skilled in the art will understand that all or part of the steps for implementing the above embodiments may be accomplished by hardware, or by hardware related to program code, and the program may be stored in a non-volatile computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0231] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the relevant technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.

[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Based on the idea of ​​the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present application as above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-stream management method for application data, characterized in that The method includes: Obtaining the compression ratio of the application data; Clustering the application data according to the compression ratio of the application data to generate multiple software flows; Estimating the life cycle of each software flow; Clustering the software flows according to the life cycle of each software flow to establish a first mapping relationship between each software flow and a hardware flow; Storing the data corresponding to the software flow at the storage location corresponding to the hardware flow according to the first mapping relationship.

2. The method according to claim 1, characterized in that, The obtaining the compression ratio of the application data includes: Determining the compression algorithm corresponding to the application data according to the data type corresponding to the application data; Compressing the application data according to the compression algorithm to obtain the compression ratio of the application data.

3. The method according to claim 2, wherein Each software flow corresponds to at least one logical block address, the logical block address includes multiple logical pages, and the data corresponding to the software flow is stored in the logical pages; The estimating the life cycle of each software flow includes: Estimating the life cycle of all logical pages corresponding to each software flow; Calculating the average value of the life cycles of all logical pages corresponding to each software flow to determine the life cycle of each software flow.

4. The method according to claim 3, wherein The estimating the life cycle of all logical pages corresponding to each software flow includes: Counting the timestamps of access to each logical page, where the timestamps include a first timestamp and a second timestamp, and the first timestamp and the second timestamp are adjacent timestamps; Calculating the difference between the first timestamp and the second timestamp to determine the life cycle of each logical page.

5. The method according to claim 4, characterized in that, Each hardware flow corresponds to a preset life cycle one by one. The clustering the software flows according to the life cycle of each software flow to establish the first mapping relationship between the software flow and the hardware flow includes: If the life cycle of the software flow is within the preset life cycle, mapping the software flow to the hardware flow corresponding to the preset life cycle to determine the first mapping relationship between the software flow and the hardware flow, where each hardware flow corresponds to multiple software flows.

6. The method according to claim 5, wherein The storage location includes multiple physical pages, and the method further includes: Recycling the valid data in the physical pages corresponding to the storage location corresponding to the hardware flow, including: Counting the feedback information of the physical pages corresponding to the storage location corresponding to each hardware flow; Moving the valid data in the storage location to a preset storage location according to the feedback information.

7. The method according to claim 6, characterized in that, The feedback information includes the write-back frequency of the physical page and the hardware flow number where the physical page is located; The moving the valid data in the storage location to a preset storage location according to the feedback information includes: Determining the write-back frequency of the physical page within a period of time according to the write-back frequency of the physical page, and moving the valid data in the physical page corresponding to the storage location to a preset storage location according to the write-back frequency; And / or Moving the valid data in the storage location to a preset storage location according to the state machine according to the hardware flow number where the physical page is located; Before moving the valid data in the storage location to a preset storage location, the method further includes: Establish a second mapping relationship between each of the hardware streams and the preset hardware stream, where the preset hardware stream corresponds to the preset storage location.

8. The method according to claim 7, wherein The moving the valid data in the physical page corresponding to the storage location to the preset storage location according to the write-back frequency includes: If the write-back frequency is greater than or equal to the frequency threshold, move the valid data in the physical page corresponding to the storage location to the preset storage location according to the second mapping relationship.

9. A flash memory device, characterized in that, Includes: A compression module, configured to determine a compression algorithm corresponding to the application data according to the data type of the application data, and compress the application data according to the compression algorithm to obtain the compression rate of the application data; A first clustering module, configured to cluster the application data according to the compression rate of the application data to generate multiple software streams; A life cycle estimation module, configured to estimate the life cycle of each of the software streams; A second clustering module, configured to cluster the software streams according to the life cycle of each of the software streams to establish a first mapping relationship between each of the software streams and the hardware stream; A flash conversion module, configured to store the data corresponding to the software stream to the storage location corresponding to the hardware stream according to the first mapping relationship; A remapping module, configured to recycle the valid data in the storage location corresponding to the hardware stream to move the valid data in the storage location to the preset storage location; A flash medium, configured to store the application data.

10. A flash memory device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-stream management method for application data according to any one of claims 1-8.

Citation Information

Patent Citations

  • File compression storage method, device and equipment and storage medium

    CN112100143A

  • Data storage device for managing memory resources using flash translation layer with condensed mapping information

    CN112114743A

  • Data storage method, data storage device, storage medium and product

    CN113568573A

  • Memory management method and electronic equipment

    CN116243850A

  • Multi-stream management method of application data and flash memory device

    CN117806986A