Data writing method, flash memory device and computer readable storage medium

By using machine learning models in flash devices to identify file block types and allocate file blocks, the problem of naked-level access conflicts is solved, and the read performance of flash devices is improved.

CN120066402AActive Publication Date: 2025-05-30DAPUSTOR CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411958399.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

When existing flash devices write multiple files concurrently, they cannot effectively avoid die-level access conflicts, resulting in degradation in read performance.

Method used

The data type of each file block is identified through the machine learning model, and the same type of file blocks are distributed successively on each die, reducing the die-level conflict.

Benefits of technology

Effectively reduces die-level conflicts and improves the read performance of flash memory devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066402A_ABST
    Figure CN120066402A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data writing method, flash memory equipment and a computer readable storage medium, the method comprises the following steps: obtaining file block data refreshed by a host, one file block data corresponding to one file block; based on a machine learning model, the file block data is identified to determine the data type of each file block, and the file blocks have a plurality of data types; determining a plurality of file blocks of the same data type according to the data type of each file block; a plurality of file blocks of the same data type are sequentially written into a plurality of bare chips, the file blocks of the same data type correspond to the bare chips in sequence, the data type of each file block is recognized through a machine learning model, and the file blocks of the same type are sequentially and circularly distributed to the bare chips. According to the flash memory device, bare chip level conflicts can be reduced, and the reading performance of the flash memory device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of data processing, and particularly to a data writing method, a flash memory device, and a computer-readable storage medium. Background Art

[0002] A flash memory device refers to a storage device manufactured based on flash memory technology (Flash Memory). For example, a solid-state drive (SSD) generally uses multiple flash memory chips (NAND Flash) as its backend storage medium. A flash memory chip usually contains multiple dies, and each die contains multiple storage units for storing bit data. For a single die, only one request can be processed at the same time, while multiple dies allow the solid-state drive to process multiple input / output requests in parallel. The higher the degree of parallel data processing of these dies, the greater the bandwidth and throughput of the solid-state drive. On the contrary, if all the physical pages of the data to be read are located on the same die, these pages can only be read one by one and cannot be read simultaneously, ultimately resulting in a sharp decline in the read performance of the solid-state drive. This phenomenon is also known as die-level access conflict.

[0003] To avoid the performance degradation caused by die-level access conflict, the flash translation layer in most traditional solid-state drive firmware distributes die schemes in the order of writing data, that is, selects dies in a cyclic manner and allocates the physical pages on the die to the write request, so as to distribute the physical pages corresponding to the consecutive logical pages of the file on different dies as much as possible. However, when writing multiple files concurrently, the flash translation layer cannot guarantee continuous die allocation for the same file. At the same time, as the file system ages, the data of a file is likely to be only aggregated on some dies, resulting in die-level conflicts on these dies, thereby causing a decline in the read performance of the flash memory device.

[0004] Currently, to solve die-level conflicts, it is usually through the way of eliminating file fragmentation. For example, copy the fragmented file to a new continuous space and then delete the fragmented area. However, for a large-capacity hard disk, the sorting time of this method is relatively long, and the sorting process requires a large amount of system resources, thereby reducing the read performance of the flash memory device. Summary of the Invention

[0005] The embodiments of the present application provide a data writing method, a flash memory device, and a computer-readable storage medium. By using a machine learning model to identify the data types of each file block and sequentially and cyclically allocate the file blocks of the same type on each die, the present application can reduce die-level conflicts and improve the read performance of the flash memory device.

[0006] The embodiments of the present application provide the following technical solutions:

[0007] In a first aspect, the embodiments of the present application provide a data writing method applied to a flash memory device. The flash memory device includes flash memory chips, and the flash memory chips include multiple dies. The method includes:

[0008] Obtain the file block data brushed by the host, where one file block data corresponds to one file block;

[0009] Based on a machine learning model, identify the file block data to determine the data type of each file block, where there are several data types of file blocks;

[0010] According to the data type of each file block, determine several file blocks of the same data type;

[0011] Write several file blocks of the same data type into several dies in sequence, where the file blocks of the same data type correspond to the dies in sequence.

[0012] In some embodiments, the method further includes:

[0013] Train the machine learning model to obtain a trained machine learning model, including:

[0014] Obtain an original data set, where the original data set includes a training data set and a validation data set, and the original data set includes multiple file block data;

[0015] Train the machine learning model with the training data set to optimize the parameters of the machine learning model and obtain an optimized machine learning model;

[0016] According to the performance of the validation data set on the machine learning model during the optimization process, select the machine learning model at the best performance moment as the final model to obtain a trained machine learning model.

[0017] In some embodiments, according to the data type of each file block, determining several file blocks of the same data type includes:

[0018] Group the file block data according to the data type of each file block to obtain several file block groups, where the file block groups correspond to the data types one by one;

[0019] According to the file block groups, determine several file blocks of the same data type, where each file block group includes several file blocks.

[0020] In some embodiments, the method further includes:

[0021] Sort several file block groups based on the number of file blocks in each file block group to obtain the sorted several file block groups;

[0022] Write several file blocks of the same data type into several dies in sequence, including:

[0023] Write all the file blocks in each file block group into several dies in sequence according to the sequence of the several file block groups.

[0024] In some embodiments, one down-brush batch contains multiple file block groups;

[0025] The method further includes:

[0026] Construct a die position record table, where the die position record table is used to record the die position information of the die where the file block of each data type is written last time, and the file block of each data type includes the file blocks of the same type of file block groups in different down-brush batches;

[0027] Write several file blocks in each file block group into several dies in sequence, including:

[0028] Obtain the first die position information, where the first die position information is the die position information of the first die, and the first die is the die where the same type of file blocks in the down-brush batches corresponding to the file block data type included in the current down-brush batch are written last time;

[0029] Determine the second die position information according to the first die position information, where the second die position information is the die position information of the second die, and the second die is the next die of the first die;

[0030] Write the first file block of each file block group in the current down-brush batch into the second die, and write the remaining file blocks of each file block group into the dies after the second die in sequence.

[0031] In some embodiments, the second die includes multiple physical pages, and the physical pages correspond to serial numbers one by one;

[0032] Write the first file block of each file block group in the current down-brush batch into the second die, including:

[0033] Obtain the first free physical page of the second die, where the first free physical page is the free physical page with the earliest serial number in the second die;

[0034] Write the first file block of each file block group in the current down-brush batch into the first free physical page.

[0035] In some embodiments, the flash memory device includes a write cache space for caching the file block data flushed by the host;

[0036] The method further includes:

[0037] Determine whether the occupied space of the file block data in the write cache space is greater than a preset space threshold;

[0038] If so, write all the file blocks in the write cache space to the die;

[0039] If not, control the write cache space to continue caching the file block data flushed by the host.

[0040] In some embodiments, before writing the file block to the die, the method further includes:

[0041] Determine the write mode of the file block, where the write mode is an append write mode or an overwrite write mode;

[0042] Determining the write mode of the file block includes:

[0043] Obtain a flash memory mapping table, where the flash memory mapping table is used to record the mapping relationship between the logical block address and the physical page address of each file block;

[0044] Query the flash memory mapping table to determine whether the logical block address corresponding to the current file block exists in the flash memory mapping table;

[0045] If so, determine that the write mode of the current file block is the overwrite write mode, where the overwrite write mode is used to write the current file block to a third die, where the third die is the die corresponding to the logical block address of the current file block;

[0046] If not, determine that the write mode of the current file block is the append write mode, where the append write mode is used to write the current file block to a second die, where the second die is the next die after the first die, and the first die is the die to which the last file block of each file type in the already flushed batches of the current flush batch is written.

[0047] In a second aspect, an embodiment of the present application provides a flash memory device, including:

[0048] A processor and a memory, where the processor is used to execute the executable program code in the memory, and when the executable program code is executed, the processor executes the instructions of the data writing method as described in the first aspect.

[0049] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed, the data writing method as described in the first aspect is implemented.

[0050] The beneficial effects of the embodiments of this application are as follows: Different from the prior art, the embodiments of this application provide a data writing method applied to a flash memory device. The flash memory device includes flash memory chips, and each flash memory chip includes multiple dies. The method includes: obtaining file block data brushed down by a host, where one file block data corresponds to one file block; based on a machine learning model, identifying the file block data to determine the data type of each file block, where there are several types of file block data types; according to the data type of each file block, determining several file blocks of the same data type; writing several file blocks of the same data type into several dies in sequence, where the file blocks of the same data type correspond to the dies in sequence. This application can identify the data type of each file block through a machine learning model and sequentially and cyclically allocate file blocks of the same type to each die, thereby reducing die-level conflicts and improving the reading performance of the flash memory device. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.

[0052] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this application;

[0053] Figure 2 is a schematic flowchart of a data writing method provided by an embodiment of this application;

[0054] Figure 3 is a schematic diagram of a file block classification technology based on machine learning provided by an embodiment of this application;

[0055] Figure 4 is a schematic flowchart of a process for obtaining a trained machine learning model provided by an embodiment of this application;

[0056] Figure 5 is Figure 2 a detailed flowchart of step S203 in

[0057] Figure 6 is a schematic flowchart of a process for obtaining several sorted file block groups provided by an embodiment of this application;

[0058] Figure 7 is a schematic flowchart of a process for constructing a die position record table provided by an embodiment of this application;

[0059] Figure 8 is Figure 2 a detailed flowchart of step S204 in

[0060] Figure 9 Is Figure 8 The detailed process schematic diagram of step S241 in;

[0061] Figure 10 Is Figure 9 The detailed process schematic diagram of step S2413 in;

[0062] Figure 11 It is a schematic diagram of a file block placement technology based on a write cache provided by an embodiment of the present application;

[0063] Figure 12 It is a schematic diagram of a process for determining whether the occupied space of file block data in the write cache space is greater than a preset space threshold provided by an embodiment of the present application;

[0064] Figure 13 It is a schematic diagram of a process for determining the write mode of a file block provided by an embodiment of the present application;

[0065] Figure 14 Is Figure 13 The detailed process schematic diagram of step S1301 in;

[0066] Figure 15 It is a schematic diagram of processing an overwrite write request by a file block placement technology based on a write cache provided by an embodiment of the present application;

[0067] Figure 16 It is a schematic diagram of processing an overwrite write request by a file block placement technology based on a mapping table provided by an embodiment of the present application;

[0068] Figure 17 It is a schematic diagram of the structure of a flash memory device provided by an embodiment of the present application.

[0069] Explanation of the reference numerals in the drawings:

[0070] Label Name Label Name 100 Application Scenario 10 Host 20 Flash Device 210 Processor 220 Memory Detailed implementation manners

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0072] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0073] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in this specification in the description of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.

[0074] The technical solution of this application will be specifically described below in conjunction with the accompanying drawings of the specification:

[0075] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application scenario provided by an embodiment of this application.

[0076] As Figure 1 shown, the application scenario 100 includes a host 10 and a flash device 20, where the host 10 is communicatively connected to the flash device 20.

[0077] In an embodiment of this application, the host 10 includes an application layer, a file system layer, and an operating system layer. Among them, the application layer includes multiple application programs. The host 10 is used to generate data in the application layer through application programs. For example, data is generated through user input, calculation, or file operations; by calling the file system interface, the data is written to the file system layer; the file system layer organizes the data into file blocks and sends a write request to the operating system; after the operating system receives the write request sent by the file system, the write request is processed through the I / O subsystem, where the write request corresponds to multiple file block data; according to the write request, the file block data is flushed to the flash device 20 through a host interface (such as SATA, NVMe, etc.).

[0078] In an embodiment of this application, the flash device 20 includes flash chips, and the flash chips include multiple dies. The flash device 20 is used to obtain the file block data flushed by the host 10, where one file block data corresponds to one file block; based on a machine learning model, the file block data is identified to determine the data type of each file block, where there are several data types of file blocks; according to the data type of each file block, several file blocks of the same data type are determined; and several file blocks of the same data type are sequentially written to several dies, where the file blocks of the same data type correspond to the dies in sequence.

[0079] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a data writing method provided by an embodiment of this application.

[0080] Among them, the data writing method is applied to a flash memory device, which includes flash memory chips, and each flash memory chip includes a plurality of dies.

[0081] As Figure 2 shown, the data writing method includes:

[0082] Step S201: Obtain the file block data downloaded by the host.

[0083] Specifically, the flash memory device is communicatively connected to the host, and the flash memory device obtains the file block data downloaded by the host. Among them, one file block data corresponds to one file block, and the file block refers to the data storage unit in the file system, that is, the smallest allocation unit of data in the file system. The size of the file block is usually determined by the file system and can be fixed or variable. Common block sizes include but are not limited to 4KB, 8KB, or 16KB.

[0084] The existing file type recognition tasks need to process a large amount of data with complex types. The existing technical solutions using statistical recognition methods have problems such as poor recognition rate and few recognized types. For example, the method of directly determining the file type according to features such as the longest common substring and subsequence, and byte frequency distribution. Since the file type may be affected by multiple factors (such as encoding, compression, encryption, etc.), simply relying on the longest common substring or subsequence of the text content may not be able to accurately distinguish different types of files. Therefore, the above existing technical solutions have low accuracy in file type recognition.

[0085] To solve the problems of the above existing technical solutions, this application proposes a file block classification technology based on machine learning. This technology uses a large-scale training data set to train the current advanced machine learning model. In the solid-state drive, the trained model can be used to automatically predict the data type of the file block, and determine the placement location of the data on the flash memory according to the data type of the file block.

[0086] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a file block classification technology based on machine learning provided by an embodiment of this application.

[0087] As Figure 3As shown in the figure, the present application combines the following advantages of the machine learning model: First, the machine learning algorithm can automatically learn and capture the main patterns of the input data, improving the accuracy of recognition. Among them, the data input into the model can be a combination of statistical features of the original file blocks according to different algorithm models, or the original file can be directly used as the input. For example, the Sceadan algorithm uses unigrams and bigrams as the features of the file block bytes and uses a support vector machine model to predict the file block type. The FiFTy algorithm, on the other hand, directly uses the byte values of the original file as the input and uses a CNN combination network to predict the file block type. Second, with the intelligence of solid-state drives, some solid-state drives have added a computable module that allows the recognition of file blocks by loading and using a trained machine learning module in the drive. For example, the space occupied by the FiFTy model that can recognize 75 file types is only 3.5MB, and the lightweight version only requires 1.7MB. This machine learning algorithm can be fully loaded into the computable module of the solid-state drive, and the present application proposes a file block classification module based on machine learning.

[0088] The machine learning model is placed on the computable module of the solid-state drive, and this machine learning model can be used to implement the function of classifying the data written from the host side to the solid-state drive side. The unclassified data blocks can be labeled with different types of file types, such as CSV, GIF, and ARW, etc., after being calculated and predicted by the model. The selection of the machine learning model can be based on the requirements of the actual task, and the available models include Sceadan, FiFTy, ByteRCNN, etc. Here, taking the FiFTy model with a large number of recognized types (75 types), high accuracy (77.5%), and great influence as an example, the FiFTy model can predict the file type of the 512-byte or 4KB original file byte blocks after segmentation. The embedding layer in the model can automatically learn the semantic information of the input bytes, and the CNN layer can learn the relationship and main patterns between the input and output data. Its training process includes the following 3 steps:

[0089] Step (1): Prepare the original data set. The source of the data can be selected according to the requirements of the file classification types in different task scenarios. If it is a general model, an open-source data set that has been marked with file types can be selected. If it is a specific scenario, such as the RocksDB scenario, it is necessary to manually mark the target file types (such as WAL, SST files, etc.) and then input them into the machine learning model for training. After obtaining the data, it is also necessary to divide the data into a training set, a validation set, and a data set in a non-overlapping manner, and the division ratio is usually 8:1:1.

[0090] Step (2): Input the input part and the target value part of the training set data into the machine learning model for learning. Through learning, that is, the machine learning model continuously fits the error between the predicted value calculated according to the input data and the specified target value. The fitting method generally used is to continuously correct the parameters of each layer in the model according to the error between the prediction result and the target result of the training set by using the gradient descent algorithm, so as to obtain an optimized machine learning model.

[0091] Step (3): Control the end of the training of the machine learning model during the optimization process through the validation data set to obtain the trained machine learning model. When the prediction result of the validation data set starts to deteriorate, the training process needs to be stopped to end the training. At this time, the training is completed, and thus the trained machine learning model is obtained. The trained model can be directly added to the computable module of the flash memory device to predict the data type of file blocks of 512 bytes or 4KB in size.

[0092] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for obtaining a trained machine learning model provided by an embodiment of the present application.

[0093] As Figure 4 shown, the process for obtaining a trained machine learning model includes:

[0094] Step S401: Obtain the original data set.

[0095] Specifically, obtain the original data. The source of the original data can be selected according to the requirements of file classification types in different task scenarios. For example, if it is a general model, an open-source data set that has been marked with file types can be selected. If it is a specific scenario, such as the RocksDB scenario, it is necessary to manually mark the target file types (such as WAL, SST files, etc.) in advance and then input them into the machine learning model for training. After obtaining the original data, it is also necessary to divide the original data into a training set, a validation set, and a data set in a non-overlapping manner. The division ratio is usually 8:1:1, so as to obtain the original data set.

[0096] In some embodiments, the input data can also be selected according to different algorithm models. For example, the original data set can be a combination of statistical features of original file blocks or an original file. For example, the Sceadan algorithm uses unigrams and bigrams as the features of file block bytes and uses a support vector machine model to predict the file block type. The FiFTy algorithm directly uses the byte values of the original file as the input and uses a CNN combination network to predict the file block type.

[0097] Step S402: Train the machine learning model with the training data set to optimize the parameters of the machine learning model and obtain an optimized machine learning model.

[0098] Specifically, after obtaining the original data set, the input part and the target value part of the training data set are input into the machine learning model for learning. The so-called learning means that the machine learning model continuously fits the error between the predicted value calculated according to the input data and the specified target value. The fitting method generally used is to continuously correct the parameters of each layer in the model according to the error between the prediction result and the target result of the training set by using the gradient descent algorithm to obtain an optimized machine learning model.

[0099] Step S403: Select the machine learning model at the moment with the best performance as the final model based on the performance of the machine learning model on the validation data set during the optimization process to obtain a trained machine learning model.

[0100] Specifically, during the training process, in order to ensure that the optimized machine learning model is generalized rather than overfitted, it is necessary to monitor its performance on the validation data set during the training process, and verify the optimized machine learning model through the validation data set to obtain a trained machine learning model. The performance of the validation data set on the machine learning model during the optimization process includes the accuracy of the prediction result (Accuracy). Among them, the accuracy of the prediction result is characterized by the proportion of correctly predicted samples in the total samples. The higher the accuracy, the better the performance of the machine learning model. When the prediction result of the validation data set starts to deteriorate, the training process needs to be stopped and the training ends. That is, based on the performance of the validation data set on the machine learning model during the optimization process, by observing the accuracy of the prediction result of the validation data set, select the machine learning model at the moment with the best performance as the final model, that is, the machine learning model with the highest accuracy of the prediction result as the final model. At this time, the training is completed, and then a trained machine learning model is obtained. The trained model can be directly added to the computable module of the flash device to predict the data type of file blocks of 512 bytes or 4KB in size.

[0101] Step S202: Identify the file block data based on the machine learning model to determine the data type of each file block.

[0102] Specifically, based on the above-obtained trained machine learning model, identify multiple file block data, where one file block data corresponds to one file block, obtain multiple file blocks, and then predict the data types of multiple file blocks to output the data type of each file block through the machine learning model. Among them, there are several data types of file blocks, and the machine learning model includes but is not limited to models such as Sceadan, FiFTy, and ByteRCNN.

[0103] In the embodiments of the present application, if the trained machine learning model is the FiFTy model, the data types of the file blocks that the FiFTy model can predict include, but are not limited to, data types such as CSV, EXE, GIF, ARW, etc. In a specified scenario, a more targeted machine learning model can be manually trained to more accurately classify files. For example, in the RocksDB scenario, the data types predicted by the machine learning model include, but are not limited to, SST, WAL, etc.; in the MySQL scenario, the data types predicted by the machine learning model include, but are not limited to, ibd, binlog, and redo, etc.

[0104] Step S203: Determine several file blocks of the same data type according to the data type of each file block.

[0105] Specifically, please refer to Figure 5 , Figure 5 Yes Figure 2 is the detailed process schematic diagram of step S203 in

[0106] As Figure 5 shown, step S203: Determine several file blocks of the same data type according to the data type of each file block, including:

[0107] Step S231: Group the file block data according to the data type of each file block to obtain several file block groups.

[0108] Specifically, group the file block data according to the data type of each file block to divide the file blocks of the same data type into the same file block group, thereby obtaining several file block groups, where the file block groups correspond one-to-one with the data types. For example, assuming there are 5 file blocks of different data types in the file block data, 5 file block groups can be determined.

[0109] Step S232: Determine several file blocks of the same data type according to the file block groups;

[0110] Specifically, after determining the file block groups according to the data type, determine several file blocks in each file block group according to the file block groups, that is, determine several file blocks of the same data type, where each file block group includes several file blocks.

[0111] In the embodiments of the present application, the present application proposes a file block placement technology based on a write cache. According to the file type recognition result of the machine learning model, the file blocks of the same type in the write cache of the flash memory device are concentrated and written into the flash memory, ensuring that the data of the same file type can make more full use of each die.

[0112] Please refer to againFigure 6 , Figure 6 is a schematic flowchart of a process for obtaining a plurality of sorted file block groups provided by an embodiment of the present application.

[0113] As Figure 6 shown, the process for obtaining a plurality of sorted file block groups includes:

[0114] Step S601: Sort a plurality of file block groups based on the number of file blocks in each file block group to obtain a plurality of sorted file block groups.

[0115] Specifically, when the write cache threshold is reached and data needs to be flushed, sort a plurality of file block groups based on the number of file blocks in each file block group to obtain a plurality of sorted file block groups, that is, sort the file block groups according to the number of file blocks of each file type from more to less. Among them, the write cache threshold can be set according to actual needs. Preferably, the write cache threshold is set to 20% of the memory size of dynamic random access memory (DRAM). For example, when the size of the SSD write cache reaches 20% of the DRAM, it is forced to flush once. Then, according to the plurality of sorted file block groups, the file blocks of the same type are sequentially flushed together. Finally, different file type data in a single flush can be effectively isolated to a certain extent, thereby reducing die-level conflicts. At the same time, when flushing, by cyclically and sequentially placing data on each die, it can ensure that the same type of data can better utilize the parallelism of the die when being read.

[0116] In the embodiment of the present application, the write cache data flushed in different batches will still cause relatively significant die-level conflicts with each other. Specifically, when the time intervals between writing multiple file blocks of the same file on the host side are very different, these file blocks will exist in different batches of write cache data, and then these file blocks are flushed from the write cache to the backend flash in different batches. At this time, it cannot be guaranteed that the data in the next batch can be placed immediately after the data in the previous batch in the die order, which will cause a certain degree of die-level conflict. Therefore, it is necessary to record the die positions (i.e., which die) where the previous batch of each file was written, and the data in the next batch is placed according to this record. At this time, it is necessary to construct a die position record table for the die position information where the previous batch of file blocks of each data type was written.

[0117] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a process for constructing a die position record table provided by an embodiment of the present application.

[0118] As Figure 7 shown, the process for constructing a die position record table includes:

[0119] Step S701: Construct a die position record table.

[0120] Specifically, construct a die position record table, where the die position record table is used to record the die position information of the die into which the file blocks of each data type are written. Among them, the file blocks of each data type include the file blocks of the file block groups of different download batches. It can be understood that a download batch is all the data in the memory downloaded at one time, and these data can be classified into multiple file block groups after classification. Therefore, a download batch contains multiple file block groups. The die position record table is used to record the die position information of the die into which the file blocks of each data type are written for the last time. Among them, the file blocks of each data type include the file blocks of the same type of file block groups of different download batches. In subsequent steps, the die position information of the die into which the file blocks of the file block groups of different download batches are written can be obtained by querying the die position record table.

[0121] Step S204: Write several file blocks of the same data type into several dies in sequence.

[0122] Among them, the file blocks of the same data type and the dies correspond in sequence, which means that several file blocks of the same data type are written into several dies in sequence according to the order.

[0123] Specifically, please refer to Figure 8 , Figure 8 is Figure 2 the detailed process schematic diagram of step S204 in

[0124] As Figure 8 shown, step S204: Write several file blocks of the same data type into several dies in sequence, including:

[0125] Step S241: Write all the file blocks in each file block group into several dies in sequence according to the order of several file block groups.

[0126] Specifically, please refer to Figure 9 , Figure 9 is Figure 8 the detailed process schematic diagram of step S241 in

[0127] As Figure 9 shown, step S241: Write all the file blocks in each file block group into several dies in sequence according to the order of several file block groups, including:

[0128] Step S2411: Obtain the first die position information.

[0129] Specifically, first obtain the current download batch, and determine the first die where the last write of the same type of file block in the download batches corresponding to the file block data types included in the current download batch is located as the first die. By querying the die position record table, obtain the first die position information, where the first die position information is the die position information of the first die, and the first die is the die where the last file block of each file block group in the previous download batch of the current download batch is written. The first die is the die where the last write of the same type of file block in the download batches corresponding to the file block data types included in the current download batch is located.

[0130] Step S2412: Determine the second die position information according to the first die position information.

[0131] Specifically, the first die position information includes the position of the die where the last write of the same type of file block in the download batches corresponding to the file block data types included in the current download batch is located. Determine this position as the first die position. Furthermore, according to the first die position information, the next position of the first die position can be obtained to determine the second die position information, where the second die position information is the die position information of the second die, and the second die is the next die of the first die.

[0132] Step S2413: Write the first file block of each file block group in the current download batch to the second die, and sequentially write the remaining file blocks of each file block group to the dies after the second die.

[0133] Specifically, please refer to Figure 10 , Figure 10 Yes Figure 9 is the detailed process schematic diagram of step S2413 in

[0134] As Figure 10 shown, step S2413: Write the first file block of each file block group in the current download batch to the second die, and sequentially write the remaining file blocks of each file block group to the dies after the second die, including:

[0135] Step S4131: Obtain the first free physical page of the second die.

[0136] Specifically, each die includes multiple physical pages. Sort the multiple physical pages in the die to determine the serial number of each physical page, and determine the unwritten full physical pages as free physical pages. Among them, the second die includes multiple physical pages, and the physical pages and serial numbers correspond one by one, that is, each physical page corresponds to a serial number one by one. Obtain the first free physical page of the second die, where the first free physical page is the free physical page with the earliest serial number in the second die.

[0137] In the embodiment of the present application, sorting multiple physical pages in the die includes determining the serial number of each physical page according to the address of the physical page. The address of the physical page is its unique identifier on the die, similar to the address in a computer's memory, and is used to locate and access a specific physical page. During the sorting process, first, the address information of each physical page is obtained, and these address information are generally represented in binary, hexadecimal, or other formats; then, according to the address information of each physical page, a sorting algorithm (such as bubble sort, quick sort, merge sort, etc.) is used to sort the physical pages. The purpose of sorting is to determine the relative position of each physical page on the die, so as to assign a unique serial number to each physical page, and this serial number can be used for subsequent access, management, and maintenance operations.

[0138] Step S4132: Write the first file block of each file block group in the current downflush batch to the first free physical page.

[0139] Specifically, when flushing file blocks of the same data type in the next batch, according to the die position record table, determine the second die position information, and then according to the second die position information, write the first file block of each file block group in the current downflush batch to the first free physical page in the second die, that is, for file blocks of the same data type, write the data of the current file block to the free physical page with the earliest serial number in the next die after the die where the data was last written.

[0140] In the embodiment of the present application, this method can ensure that when flushing file data of the same data type in multiple batches, the data of files of the same data type can still be scattered to each die in an orderly manner. At the same time, since the data of each file block will be written to the free physical page with the earliest serial number of each die, it can also ensure that there will not be a large number of free physical pages between different files, making full use of the space of the flash memory chip.

[0141] Please refer to Figure 11 , Figure 11 which is a schematic diagram of a file block placement technology based on a write cache provided by the embodiment of the present application.

[0142] As Figure 11As shown, assume that there are 16 dies in the current flash memory, namely die 1 to die 16. The data types of file blocks include CSV type, GIF type, and ARW type. Among them, the CSV type corresponds to CSV file blocks, the GIF type corresponds to GIF file blocks, and the ARW type corresponds to ARW file blocks. During the first write process of file blocks, if there are 23 CSV file blocks (serial numbers 1 to 23), 19 GIF file blocks (serial numbers 1 to 19), and 10 ARW file blocks (serial numbers 1 to 10) that need to be written into the flash memory, first, according to the number of file blocks of each data type, sort the file block groups. The file blocks sorted in the front are written into the flash memory first. Since the number of CSV file blocks is greater than the number of GIF file blocks, and the number of GIF file blocks is greater than the number of ARW file blocks, first write the CSV file blocks into the flash memory, then write 19 GIF file blocks into the flash memory, and finally write 10 ARW file blocks into the flash memory. Through the die position record table, record that the last die written for the CSV file blocks is die 7, the last die written for the GIF file blocks is die 10, and the last die written for the ARW file blocks is die 4, that is, complete the data download of this file block.

[0143] During the second write process of file blocks, the circles in the flash memory represent the data downloaded last time, and the squares represent the data to be downloaded this time. Assume that there are 15 CSV file blocks to be written this time, and their serial numbers are 24 to 38; there are 3 ARW file blocks to be written this time, and their serial numbers are 11 to 13. Then, since the number of CSV file blocks to be written this time is greater than the number of ARW file blocks to be written this time, first write the CSV file blocks into the flash memory. Through the die position record table, it is obtained that the last die written for the CSV file blocks last time is die 7, so it is determined that the position where the CSV file block with serial number 24 is written this time is die 8, write the CSV file block with serial number 25 into die 9, write the CSV file block with serial number 26 into die 10, and so on, and obtain that the last die written for the CSV file blocks this time is die 6. Then write the ARW file blocks into the flash memory. Through the die position record table, it is obtained that the last die written for the ARW file blocks last time is die 4, so it is determined that the position where the ARW file block with serial number 11 is written this time is die 5, write the ARW file block with serial number 12 into die 6, write the ARW file block with serial number 13 into die 7, and further obtain that the last die written for the ARW file blocks this time is die 7.

[0144] During the writing process of the third file block, assume that there are 11 GIF file blocks to be written this time, and their serial numbers are from 20 to 30; through the die position record table, it is obtained that the last die written for the GIF file block last time is die 10, so it is determined that the position for writing the GIF file block with serial number 20 this time is die 11, write the GIF file block with serial number 20 to die 11, write the CSV file block with serial number 21 to die 12, and so on, and it is obtained that the last die written for the CSV file block this time is die 5.

[0145] During the writing process of the fourth file block, assume that there are 17 ARW file blocks to be written this time, and their serial numbers are from 14 to 30; through the die position record table, it is obtained that the last die written for the ARW file block last time is die 7, so it is determined that the position for writing the ARW file block with serial number 14 this time is die 8, write the ARW file block with serial number 15 to die 9, write the ARW file block with serial number 16 to die 10, and so on, and it is obtained that the last die written for the ARW file block this time is die 8. It should be noted that Figure 10 each die includes 8 physical pages. Each time a file block is written to a die, the file block needs to be written to the free physical page with the earliest serial number in the die. For example, before writing the ARW file block with serial number 29, the physical pages with serial numbers 5 to 8 in die 7 are free physical pages. At this time, it is determined that the free physical page with the earliest serial number in the die is the physical page with serial number 5, so write the ARW file block with serial number 29 to the physical page with serial number 5.

[0146] In the embodiment of the present application, taking the CSV file block (green) as an example, if the rule of writing the first file block of each file block group in the current download batch to the second die is not followed, the 24th CSV file block should have been placed at the position of the blue 11th file block in the flash memory (located on die 5), but since the die record table shows that the last file written for the CSV file was the green 23rd file (located on die 7), therefore, this write needs to find a free block at the position of die 8 for writing. At this time, it is ensured that the CSV file data can be placed sequentially and dispersedly on the die. In addition, if the rule of writing the first file block of each file block group in the current download batch to the first free physical page is not followed, there is a section of free physical pages between the 38th CSV (green) file block and the 20th GIF (yellow) file block before writing the ARW (blue) data. As the data is continuously written, this section of free physical pages will cause a large number of data holes in the flash memory. However, if each die preferentially selects the free physical page with the earliest serial number for writing, this section of free physical pages will finally be effectively used, thereby avoiding the generation of a large number of data holes in the flash memory.

[0147] Please refer to Figure 12 ,Figure 12 It is a schematic flow chart for determining whether the occupied space of the file block data in the write cache space is greater than a preset space threshold provided by an embodiment of the present application.

[0148] In an embodiment of the present application, the flash memory device includes a write cache space for caching file block data flushed by the host.

[0149] Such as Figure 12 As shown, the process of determining whether the occupied space of the file block data in the write cache space is greater than a preset space threshold includes:

[0150] Step S1201: Obtain the occupied space of the file block data in the write cache space.

[0151] Specifically, each time the host flushes file block data, obtain the occupied space of the file block data in the current write cache space.

[0152] Step S1202: Determine whether the occupied space of the file block data in the write cache space is greater than a preset space threshold.

[0153] Specifically, the flash memory device will set a write cache in the dynamic random access memory in the disk to buffer the data written by the host into the disk. After reaching the preset space threshold, the data in the write cache will be flushed to the backend flash memory medium. That is, determine whether the occupied space of the file block data in the write cache space is greater than the preset space threshold. If the occupied space of the file block data in the write cache space is greater than the preset space threshold, go to step S1203; if the occupied space of the file block data in the write cache space is less than or equal to the preset space threshold, go to step S1204. It should be noted that the preset space threshold can be set according to actual needs. For example, set the preset space threshold to 20% of the write cache space. If the occupied space of the file block data in the current write cache space reaches 20% of the write cache space, then step S1203.

[0154] Step S1203: Write all file blocks in the write cache space to the die.

[0155] Specifically, if the occupied space of the file block data in the write cache space is greater than the preset space threshold, it means that the occupied space of the file block data in the current write cache space meets the condition for writing the file block to the flash memory, then write all file blocks in the write cache space to the die.

[0156] Step S1204: Control the write cache space to continue caching the file block data flushed by the host.

[0157] Specifically, if the occupied space of the file block data in the write cache space is less than or equal to the preset space threshold, it indicates that the occupied space of the file block data in the current write cache space does not meet the condition for writing the file block to the flash memory. Then, control the write cache space to continue caching the file block data flushed by the host, and cache the file block data flushed by the host in the write cache space.

[0158] In the embodiments of the present application, flushing file blocks of the same data type in the write cache can be applied to the append write mode of files, but it cannot be applied to the die-level conflict in the overwrite write mode. Specifically, since in-place update is not possible inside the flash memory, the data written in a file overwrite write operation and the data to be overwritten are located at different positions in the flash memory. This will destroy the sequential placement order of this same file data on the die, resulting in a significant increase in the probability of flash access conflict when this part of the data is read. Based on this, the present application proposes a method for writing file blocks to the flash memory according to different write modes of file blocks, where the write mode includes the overwrite write mode and the append write mode.

[0159] Please refer to Figure 13 , Figure 13 which is a schematic flowchart of a process for determining the write mode of a file block provided by an embodiment of the present application.

[0160] As Figure 13 shown, the process for determining the write mode of a file block includes:

[0161] Step S1301: Determine the write mode of the file block.

[0162] Specifically, please refer to Figure 14 , Figure 14 which is Figure 13 a refined flowchart of step S1301 in

[0163] As Figure 14 shown, step S1301: Determine the write mode of the file block, includes:

[0164] Step S1311: Obtain the flash memory mapping table.

[0165] Specifically, obtain the flash memory mapping table, where the flash memory mapping table refers to the conversion mapping table from the logical block address to the physical page address in the flash memory device. The flash memory mapping table records the physical page address information corresponding to the logical block address of the written data, and the write mode of the file block can be identified and judged through the flash memory mapping table in the subsequent steps.

[0166] Step S1312: Query the flash memory mapping table to determine whether the logical block address corresponding to the current file block exists in the flash memory mapping table.

[0167] Specifically, query the flash mapping table to determine whether the logical block address corresponding to the current file block exists in the flash mapping table. If the logical block address corresponding to the current file block already exists in the flash mapping table, go to step S1313; if the logical block address corresponding to the current file block does not exist in the flash mapping table, go to step S1314.

[0168] Step S1313: Determine that the write mode of the current file block is the overwrite write mode.

[0169] Specifically, if the logical block address corresponding to the current file block already exists in the flash mapping table, it indicates that the current write mode is the overwrite write mode. Among them, the overwrite write mode is used to write the current file block to the third die, where the third die is the die corresponding to the logical block address of the current file block; for the overwrite write mode, determine the die corresponding to the logical block address of the current file block as the third die, and determine the file block to be overwritten on the third die according to the logical block address of the current file block. The flash translation layer of the flash device will write the current file block on the same die as the file block to be overwritten according to the logical block address of the file block to be overwritten.

[0170] Step S1314: Determine that the write mode of the current file block is the append write mode.

[0171] Specifically, if the logical block address corresponding to the current file block does not exist in the flash mapping table, it indicates that the current write mode is the append write mode. Among them, the append write mode is used to write the current file block to the second die, where the second die is the next die of the first die, and the first die is the die where the last file block of each file type in the current download batch has been downloaded.

[0172] Please refer to Figure 15 , Figure 15 It is a schematic diagram of a technology for processing an overwrite write request based on a write cache for file block placement provided by an embodiment of the present application.

[0173] Such as Figure 15As shown, when determining the write mode of the current file block without using the flash mapping table, when the file block overwrites and writes to the same logical block address, the flash translation layer of the solid-state drive will default to allocate physical page addresses to the file block in the append write mode. Specifically, in the flash mapping table, it cannot be ensured that the physical page addresses Y2 and Y3 allocated to the overwritten logical block addresses X2 and X3 are in the same die position as before. For example, assume that the logical block address corresponding to file block No. 65 is the same as the physical page address of file block No. 3, but the overwritten file block No. 65 of CSV (green) is not placed on die 3 corresponding to the overwritten file block No. 3, but is placed on die 1 in the append write mode based on the data placement technology of the write cache, resulting in a pile-up of valid CSV files on die 1 and sparse CSV data on die 3, ultimately leading to serious die-level conflicts when the CSV file is read.

[0174] Please refer to Figure 16 , Figure 16 which is a schematic diagram of a file block placement technology based on a mapping table for processing overwrite write requests provided by an embodiment of the present application.

[0175] As Figure 16 shown, assume that the currently to-be-written file block is the CSV file block numbered 66. By querying the flash mapping table, it is determined that the logical block address corresponding to the CSV file block numbered 66 exists in the flash mapping table. If the logical block address corresponding to the current file block already exists in the flash mapping table, it indicates that the write mode corresponding to the current file block is the overwrite write mode, where the overwrite write mode is used to write the current file block to the third die, where the third die is the die corresponding to the logical block address of the current file block, that is, die 6; for the overwrite write mode, the flash translation layer of the flash device writes the CSV file block numbered 66 on the same die 6 as the overwritten file block according to the logical block address of the overwritten file block.

[0176] For another example, when using the flash mapping table, in the flash mapping table, the physical page addresses allocated to the overwritten logical block addresses X2 and X3 are in the same die position as the previous physical page addresses, that is, the data is still placed on die 2 and die 3; in the flash, taking the CSV file block No. 75 as an example, the file block data is placed on die 3; the data placement method for the overwrite writes of the remaining file blocks is the same as the above method, and will not be elaborated here.

[0177] In the embodiment of the present application, the CSV file blocks can be very evenly distributed on each die. The die with the most CSV file blocks is die 15, with 4 CSV file blocks, and the die with the fewest CSV file blocks contains 4 CSV file blocks, and the difference is only 1. And compared with Figure 14A solution for placing overwritten data without using flash mapping table information, the difference between the die with the most CSV file blocks and the die with the least CSV file blocks is 4. When a large amount of data is written subsequently, this phenomenon will be more serious, leading to die-level conflicts.

[0178] In an embodiment of the present application, the present application provides a data writing method, which is applied to a flash memory device. The flash memory device includes flash memory chips, and the flash memory chips include multiple dies. The method includes: obtaining file block data brushed by the host, where one file block data corresponds to one file block; based on a machine learning model, identifying the file block data to determine the data type of each file block, where there are several types of file block data types; according to the data type of each file block, determining several file blocks of the same data type; writing several file blocks of the same data type into several dies in sequence, where the file blocks of the same data type and the dies correspond in sequence. By using the machine learning model to identify the data type of each file block and cyclically distributing the file blocks of the same type to each die in sequence, the present application can reduce die-level conflicts and improve the reading performance of the flash memory device.

[0179] Please refer to again Figure 17 , Figure 17 is a schematic structural diagram of a flash memory device provided by an embodiment of the present application.

[0180] As Figure 17 shown, the flash memory device 20 includes one or more processors 210 and a memory 220. Among them, Figure 17 one processor 210 is taken as an example in

[0181] The processor 210 and the memory 220 can be connected through a bus or other means, Figure 17 Taking connection through a bus as an example in

[0182] The processor 210 is used to provide computing and control capabilities to control the flash memory device 20 to execute corresponding tasks. For example, controlling the flash memory device 20 to execute the data writing method in any one of the above method embodiments. The data writing method is applied to a flash memory device. The flash memory device includes flash memory chips, and the flash memory chips include multiple dies. The method includes: obtaining file block data brushed by the host, where one file block data corresponds to one file block; based on a machine learning model, identifying the file block data to determine the data type of each file block, where there are several types of file block data types; according to the data type of each file block, determining several file blocks of the same data type; writing several file blocks of the same data type into several dies in sequence, where the file blocks of the same data type and the dies correspond in sequence.

[0183] By using a machine learning model to identify the data types of each file block and sequentially and cyclically allocate file blocks of the same type on each die, the present application can reduce die-level conflicts and improve the read performance of the flash device.

[0184] The processor 210 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0185] The memory 220, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the data writing method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory 220, the processor 210 can implement the data writing method in any one of the above method embodiments. Specifically, the memory 220 can include a volatile memory (VM), such as a random access memory (RAM); the memory 220 can also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or other non-transitory solid-state storage devices; the memory 220 can also include a combination of the above types of memories.

[0186] The memory 220 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 220 may optionally include a memory remotely located with respect to the processor 210, and these remote memories may be connected to the processor 210 through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] One or more modules are stored in the memory 220 and, when executed by one or more processors 210, perform the data writing method in any of the above method embodiments. For example, perform the Figure 2 respective steps shown above.

[0188] In the embodiments of the present application, the flash device 20 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The flash device 20 may also include other components for implementing the functions of the device, which will not be elaborated here.

[0189] The embodiments of the present application also provide a non-volatile computer-readable storage medium, such as a memory including program code, and the above program code can be executed by a processor to complete the data writing method in the above embodiments. For example, the non-volatile computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0190] The embodiments of the present application also provide a non-volatile computer-readable storage medium, such as a memory including program code, and the above program code can be executed by a processor to complete the data writing method in the above embodiments. For example, the non-volatile computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0191] The embodiments of the present application also provide a computer program product, which includes one or more program codes stored in a non-volatile computer-readable storage medium. The processor of the flash device reads the program code from the non-volatile computer-readable storage medium, and the processor executes the program code to complete the method steps of the data writing method provided in the above embodiments.

[0192] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or by hardware related to program code. The program can be stored in a non-volatile computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.

[0193] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the method of the above embodiments can be completed by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data writing method, characterized in that: Applied to a flash memory device, the flash memory device includes a flash memory chip, the flash memory chip includes a plurality of bare chips, and the method includes: Obtain the file block data flushed from the host, where one file block data corresponds to one file block; Based on the machine learning model, the file block data is identified to determine the data type of each file block, wherein the data type of the file block has several types; According to the data type of each of the file blocks, determining a number of file blocks of the same data type; A plurality of file blocks of the same data type are sequentially written into a plurality of the bare chips, wherein the file blocks of the same data type correspond to the bare chips sequentially.

2. The method according to claim 1, characterized in that The method further comprises: Training the machine learning model to obtain a trained machine learning model includes: Acquire an original data set, wherein the original data set includes a training data set and a verification data set, wherein the original data set includes a plurality of file block data; Training the machine learning model using the training data set to optimize the parameters of the machine learning model to obtain an optimized machine learning model; According to the performance of the verification data set on the machine learning model during the optimization process, the machine learning model with the best performance is selected as the final model to obtain the trained machine learning model.

3. The method according to claim 1, characterized in that The step of determining a plurality of file blocks of the same data type according to the data type of each file block includes: According to the data type of each file block, the file block data is grouped to obtain a plurality of file block groups, wherein the file block groups correspond to the data types one by one; According to the file block group, a plurality of file blocks of the same data type are determined, wherein each of the file block groups includes a plurality of the file blocks.

4. The method according to claim 3, characterized in that The method further comprises: sorting the plurality of file block groups based on the number of file blocks in each file block group to obtain sorted plurality of file block groups; The step of sequentially writing a plurality of file blocks of the same data type into a plurality of the bare chips comprises: All the file blocks in each of the file block groups are written into the plurality of bare chips in sequence according to the sequence of the plurality of file block groups.

5. The method according to claim 4, characterized in that A refresh batch includes a plurality of the file block groups; The method further comprises: Constructing a die location record table, wherein the die location record table is used to record the die location information of the die to which the file blocks of each data type were last written, wherein the file blocks of each data type include file blocks of the same type of file block groups in different batches of flashing; The step of sequentially writing the plurality of file blocks in each of the file block groups into the plurality of bare chips comprises: Obtaining first die location information, wherein the first die location information is die location information of a first die, and the first die is a die to which a file block of the same type of a downloaded batch corresponding to a file block data type included in a current downloaded batch is last written; Determine second die position information according to the first die position information, wherein the second die position information is die position information of a second die, and the second die is a next die of the first die; The first file block of each file block group of the current batch is written to the second die, and the remaining file blocks of each file block group are sequentially written to the die after the second die.

6. The method according to claim 5, characterized in that The second die includes a plurality of physical pages, and the physical pages correspond to the serial numbers one by one; Writing the first file block of each file block group of the current batch to the second bare chip includes: Acquire a first idle physical page of the second die, wherein the first idle physical page is an idle physical page with the first sequence number in the second die; The first file block of each file block group of the current flush batch is written to the first free physical page.

7. The method according to claim 1, characterized in that The flash memory device includes a write cache space, and the write cache space is used to cache file block data flushed from the host; The method further comprises: Determine whether the occupied space of the file block data in the write cache space is greater than a preset space threshold; If so, writing all file blocks of the write cache space to the die; If not, the write cache space is controlled to continue caching the file block data flushed from the host.

8. The method according to claim 1, characterized in that Before writing the file block to the die, the method further includes: Determine a write mode of the file block, wherein the write mode is an append write mode or an overwrite write mode; The step of determining the write mode of the file block includes: Obtaining a flash memory mapping table, wherein the flash memory mapping table is used to record a mapping relationship between a logical block address and a physical page address of each file block; Query the flash memory mapping table to determine whether the logical block address corresponding to the current file block exists in the flash memory mapping table; If yes, determine that the write mode of the current file block is an overwrite write mode, wherein the overwrite write mode is used to write the current file block to a third die, wherein the third die is a die corresponding to the logical block address corresponding to the current file block; If not, determine that the write mode of the current file block is an append write mode, wherein the append write mode is used to write the current file block to a second bare chip, wherein the second bare chip is the next bare chip of the first bare chip, and the first bare chip is the bare chip to which the last file block of each file type of the current refresh batch is written.

9. A flash memory device, characterized in that: include: A processor and a memory, wherein the processor is used to execute an executable program code in the memory, and when the executable program code is executed, the processor executes instructions of the data writing method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the data writing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Solid state disk performance optimization method and device, computer equipment and storage medium

    CN114153398A

  • Memory system and control method

    CN115576873A

  • Enabling faster and regulated device initialization times

    US10877900B1

  • Identifying memory block write endurance using machine learning

    US20180357535A1

  • Logical-to-physical mapping of data groups with data locality

    US20210191850A1