A data storage method and a data storage system based on a CPU
By optimizing the data writing strategy based on the data type and cache module status of the CPU, the problems of slow NAND flash memory writing speed and low resource utilization efficiency are solved, and more efficient data storage is achieved.
Patent Information
- Application Number
- CN202411616419.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-13
AI Technical Summary
When using NAND flash memory for data storage, the prior art does not distinguish between continuous data and random data, resulting in slow writing speed, large write amplification, and low resource utilization efficiency.
The CPU performs targeted data processing procedures based on the type of data to be stored, the type of cache module and the remaining capacity, and writes data to NAND flash memory, including data integration and migration, and optimizes the data writing strategy.
It improves the data writing speed, significantly reduces the write amplification coefficient, improves resource utilization efficiency, and extends the service life of NAND flash memory.
Smart Images

Figure CN119473156B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of storage control, and in particular, to a data storage method and a data storage system based on a CPU. Background Art
[0002] NAND flash is a non-volatile memory that can still store data after power-off. With the gradual replacement of mechanical hard disks by solid-state drives in the industrial application market, NAND flash has also been widely used in various electronic products due to its excellent read and write performance, large storage capacity, and low power consumption.
[0003] In the prior art, when using NAND flash for data storage, data is usually directly written into NAND flash without processing. This method has the following defects: when dealing with random data, it is directly written into Nand flash without sorting. When writing and garbage collecting, on the one hand, it will lead to a slow writing speed, and on the other hand, during garbage collection, the write amplification is relatively large. Moreover, some technologies set up a temporary storage area, but do not distinguish between continuous data and random data, resulting in low resource utilization efficiency. Summary of the Invention
[0004] To solve one or more of the above technical problems, the present application provides a data storage method and a data storage system based on a CPU.
[0005] In a first aspect, the present application provides a data storage method based on a CPU, adopting the following technical solution:
[0006] The CPU is respectively connected to a cache module and a NAND flash; the method includes:
[0007] The cache module continuously receives data to be stored;
[0008] The CPU reads the data to be stored in the cache module and determines the type of the data to be stored;
[0009] The CPU reads the data to be stored in the cache module and determines the type of the data to be stored;
[0010] The CPU obtains the type and remaining capacity of the cache module, and based on the type of the data to be stored, the type and remaining capacity of the cache module, and a pre-constructed data processing model, processes the data to be stored and writes it into the NAND flash;
[0011] Among them, the type of the data to be stored includes continuous data and random data; the type of the cache module includes a non-volatile cache module and a volatile cache module.
[0012] By adopting the above technical solution, the CPU executes corresponding data processing procedures according to factors such as the type of data to be stored, the type and remaining capacity of the cache module, and its own system operation status, performs targeted and refined processing on the data to be stored, and writes the processed data to be stored into the NAND flash memory, improving the speed of writing data into the NAND flash memory, significantly reducing the write amplification factor of the NAND flash memory, improving the resource utilization efficiency, enhancing the performance of the NAND flash memory, and extending the service life of the NAND flash memory.
[0013] In a specific feasible implementation, the data processing model includes a first data processing sub-model, a second data processing sub-model, a third data processing sub-model, and a fourth data processing sub-model;
[0014] When the type of the cache module is a non-volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored based on the first data processing sub-model and writes it into the NAND flash memory;
[0015] When the type of the cache module is a non-volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored based on the second data processing sub-model and writes it into the NAND flash memory;
[0016] When the type of the cache module is a volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored based on the third data processing sub-model and writes it into the NAND flash memory;
[0017] When the type of the cache module is a volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored based on the fourth data processing sub-model and writes it into the NAND flash memory.
[0018] In a specific feasible implementation, the CPU processes the data to be stored based on the first data processing sub-model and writes it into the NAND flash memory, which specifically includes:
[0019] The CPU determines whether the system is busy based on its own system operation status;
[0020] If so, the CPU controls the cache module to temporarily store the received data to be stored and continuously monitors its own system operation status. When the system is idle, the CPU moves the data to be stored in the cache module to the NAND flash memory;
[0021] If not, the CPU directly writes the data to be stored in the cache module into the NAND flash memory.
[0022] By adopting the above technical solution, when the type of the cache module is a non-volatile cache module, since the non-volatile cache module has the characteristic of long-term data storage, therefore, for continuous data, when the system is idle, the continuous data is directly moved and written into the NAND flash as much as possible, which can reduce the occupation of the cache module, reduce the occupation of the CPU system, avoid the problem of a large flash write amplification factor caused by the CPU system being busy, avoid system congestion, and ensure the reliability of data storage.
[0023] In a specific feasible implementation, the CPU processes the data to be stored and writes it into the NAND flash based on the second data processing sub-model, specifically including:
[0024] The CPU determines whether the system is busy based on its own system running state;
[0025] If so, the CPU controls the cache module to temporarily store the received data to be stored, and continuously monitors its own system running state. Until the system is idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash;
[0026] If not, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash.
[0027] By adopting the above technical solution, when the type of the cache module is a non-volatile cache module and the data to be stored is random data, first, it is judged whether the remaining capacity of the cache module is sufficient. If it is not sufficient, the CPU quickly performs data integration and moving operations on the data to be stored; if it is sufficient, the CPU controls the cache module to continuously temporarily store the received data to be stored. At the same time, the CPU performs data integration and moving operations when the system is not busy according to its own system running state.
[0028] In a specific feasible implementation, the CPU processes the data to be stored and writes it into the NAND flash based on the third data processing sub-model, specifically including:
[0029] The CPU directly moves the data to be stored in the cache module to the NAND flash.
[0030] By adopting the above technical solution, when the type of the cache module is a volatile cache module, since the data stored in the volatile cache module will be cleared in case of accidental conditions such as system power-off or restart, therefore, in order to improve the security of data transmission and storage, if the data to be stored is continuous data at this time, the continuous data is directly moved and written into the NAND flash.
[0031] In a specific feasible implementation, the CPU processes the data to be stored based on the fourth data processing sub-model and writes it into the NAND flash memory, which specifically includes:
[0032] The CPU determines whether the remaining capacity of the cache module is lower than a first preset capacity threshold based on the remaining capacity information of the cache module;
[0033] If so, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory;
[0034] If not, the CPU controls the cache module to temporarily store the received data to be stored and determines whether the system is busy based on its own system operation status;
[0035] If idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory;
[0036] If busy, the CPU continuously monitors its own system operation status until the system is idle, then the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory.
[0037] By adopting the above technical solution, if the data to be stored is random data, when the remaining capacity of the volatile cache module is relatively small, the CPU quickly integrates the random data into large files and then moves them to the NAND flash memory, avoiding the problem of data loss caused by the inability to continue storing data when the remaining capacity of the cache module is insufficient. When the remaining capacity of the volatile cache module is relatively large, it means that the volatile cache module can continue to store data at this time, so the CPU performs data integration and transfer operations when its own system is idle.
[0038] In a specific feasible implementation, the data processing model further includes a power failure protection processing sub-model; the CPU is used to detect the power failure risk of the cache module based on the power failure protection processing sub-model and execute corresponding protection measures, which specifically include:
[0039] If the type of the cache module is a volatile cache module, the CPU continuously obtains the power supply voltage of the cache module and determines whether there is a power failure risk for the cache module based on the fluctuation condition of the power supply voltage;
[0040] If so, the CPU moves the data to be stored in the cache module to the NAND flash memory.
[0041] By adopting the above technical solution, by considering the waveform condition of the power supply voltage of the volatile cache module, when the power supply voltage is unstable, there may be a risk of power failure in the volatile cache module. To avoid data loss, at this time, the CPU quickly moves the data to be stored in the cache module to the NAND flash for storage, making the data transmission process more secure and reliable.
[0042] In a specific feasible implementation, when the type of the cache module is a non-volatile cache module, the cache module is an area divided inside the NAND flash or a device external to the NAND flash;
[0043] When the type of the cache module is a volatile cache module, the cache module is a device external to the NAND flash.
[0044] By adopting the above technical solution, when the cache module is of the non-volatile cache type, those skilled in the art can divide a part of the area inside the NAND flash as the cache module, or choose a device external to the NAND flash as the cache module; when the cache module is of the volatile cache type, those skilled in the art can choose a device external to the NAND flash as the cache module.
[0045] In a specific feasible implementation, the process by which the CPU integrates the data to be stored in the cache module specifically includes:
[0046] The CPU divides each of the data to be stored into corresponding data sets based on the data size of the data to be stored;
[0047] The CPU sorts the data to be stored in each of the data sets based on the sorting rules corresponding to each data set to generate corresponding data sequences;
[0048] The CPU performs matching and combination between the data to be stored in each of the data sequences to form the integrated data to be stored.
[0049] In a second aspect, the present application provides a data storage system, adopting the following technical solution: The system includes a CPU, a cache module, and a NAND flash. The data storage system applies the data storage method based on the CPU described in the first aspect or any feasible implementation of the first aspect above.
[0050] In summary, the technical solution of the present application at least includes the following beneficial technical effects:
[0051] 1. The CPU executes corresponding data processing procedures based on factors such as the type of data to be stored, the type and remaining capacity of the cache module, and its own system operation status, performs targeted and refined processing on the data to be stored, and writes the processed data to be stored into the NAND flash memory, which improves the speed of writing data into the NAND flash memory, significantly reduces the write amplification factor of the NAND flash memory, improves the resource utilization efficiency, enhances the performance of the NAND flash memory, and extends the service life of the NAND flash memory.
[0052] 2. When the type of the cache module is a non-volatile cache module, since the non-volatile cache module has the characteristic of long-term data preservation, for continuous data at this time, when the system is idle, the continuous data is directly moved and written into the NAND flash memory as much as possible, which can reduce the occupation of the cache module, reduce the occupation of the CPU system, avoid the problem of a large write amplification factor of the NAND flash memory caused by the CPU system being busy, avoid system congestion, and ensure the reliability of data storage; for random data, when the remaining capacity of the cache module is insufficient, the CPU quickly performs data integration and transfer operations. If it is sufficient, the CPU controls the cache module to continuously cache the received data to be stored, and at the same time, when the system is not busy, the CPU performs data integration and transfer operations.
[0053] 3. When the type of the cache module is a volatile cache module, since the data stored in the volatile cache module will be cleared in the event of system power-off or restart, etc., therefore, in order to improve the security of data transmission and storage, for continuous data, the continuous data is directly moved and written into the NAND flash memory; for random data, when the remaining capacity of the cache module is relatively small, the CPU quickly performs integration and transfer operations to avoid the problem of data loss caused by the inability to continue storing data when the remaining capacity of the cache module is insufficient. When the remaining capacity of the cache module is relatively large, it means that the cache module can continue to store data at this time, and the CPU performs data integration and transfer operations when its own system is idle. Description of the Drawings
[0054] Figure 1 is a schematic diagram of the connection relationship of the CPU in the embodiment of the present application;
[0055] Figure 2 is the main flowchart of the data storage method in the embodiment of the present application;
[0056] Figure 3 is the processing flowchart of the data to be stored when the cache module is a non-volatile cache module in the embodiment of the present application;
[0057] Figure 4 is the processing flowchart of the data to be stored when the cache module is a volatile cache module in the embodiment of the present application;
[0058] Figure 5 It is a schematic diagram of the data integration process in the embodiments of the present application. Detailed implementation manners
[0059] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0060] The embodiments of the present application provide a data storage method based on a CPU. Refer to Figure 1 , the CPU is respectively connected to a cache module and a NAND flash memory; specifically, the cache module is a high-speed cache module.
[0061] Refer to Figure 2 , the data storage method includes:
[0062] The cache module continuously receives data to be stored;
[0063] The CPU reads the data to be stored in the cache module and determines the type of the data to be stored;
[0064] The CPU obtains the type and remaining capacity of the cache module, and based on the type of the data to be stored, the type and remaining capacity of the cache module, and a pre-constructed data processing model, processes the data to be stored and writes it into the NAND flash memory;
[0065] Among them, the types of the data to be stored include continuous data and random data; the types of the cache module include a non-volatile cache module and a volatile cache module.
[0066] In a possible implementation manner, the CPU may include two data input interfaces. One data input interface is connected to the cache module, and the other data input interface is used to directly receive the data to be stored. That is to say, the CPU can not only determine the type of the data to be stored by reading the data to be stored in the cache module, but also determine the type of the data to be stored by directly receiving and reading the data to be stored.
[0067] In a possible implementation manner, the data processing model includes a first data processing sub-model, a second data processing sub-model, a third data processing sub-model, and a fourth data processing sub-model;
[0068] When the type of the cache module is a non-volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored and writes it into the NAND flash memory based on the first data processing sub-model;
[0069] When the type of the cache module is a non-volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored based on the second data processing sub-model and writes it into the NAND flash memory;
[0070] When the type of the cache module is a volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored based on the third data processing sub-model and writes it into the NAND flash memory;
[0071] When the type of the cache module is a volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored based on the fourth data processing sub-model and writes it into the NAND flash memory.
[0072] In a possible implementation manner, referring to Figure 3 , the CPU processes the data to be stored based on the first data processing sub-model and writes it into the NAND flash memory, which specifically includes:
[0073] The CPU determines whether the system is busy based on its own system running state;
[0074] If so, the CPU controls the cache module to temporarily store the received data to be stored and continuously monitors its own system running state. Until the system is idle, the CPU moves the data to be stored in the cache module to the NAND flash memory;
[0075] If not, the CPU directly writes the data to be stored in the cache module into the NAND flash memory.
[0076] Through the above process, when the type of the cache module is a non-volatile cache module, since the non-volatile cache module has the characteristic of long-term data preservation, therefore, for continuous data, when the system is idle, the continuous data is directly moved and written into the NAND flash memory as much as possible, which can reduce the occupation of the cache module, reduce the occupation of the CPU system, avoid the problem of a large NAND flash write amplification factor caused by the CPU system being busy, avoid system congestion, and ensure the reliability of data storage.
[0077] In a possible implementation manner, referring to Figure 3 , the CPU processes the data to be stored based on the second data processing sub-model and writes it into the NAND flash memory, which specifically includes:
[0078] The CPU determines whether the system is busy based on its own system running state;
[0079] If so, the CPU controls the cache module to temporarily store the received data to be stored, and continuously monitors the running state of its own system. Until the system is idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory;
[0080] If not, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory.
[0081] In the above process, when the type of the cache module is a non-volatile cache module and the data to be stored is random data, first judge whether the remaining capacity of the cache module is sufficient. If it is not sufficient, the CPU quickly performs data integration and transfer operations on the data to be stored; if it is sufficient, the CPU controls the cache module to continuously temporarily store the received data to be stored. At the same time, the CPU performs data integration and transfer operations according to the running state of its own system when the system is not busy.
[0082] Since each piece of random data is a relatively small file, the CPU integrates the data to be stored in the cache module into large files and then moves them to the NAND flash memory. This integration and transfer process can not only improve the writing speed of random data, but also improve the speed by more than 10 times compared with the method of directly writing random data into the NAND flash memory, significantly improving the writing speed of random data and significantly reducing the write amplification factor and extending the service life of the NAND flash memory.
[0083] In a possible implementation manner, referring to Figure 4 , the CPU processes the data to be stored based on the third data processing sub-model and writes it into the NAND flash memory, specifically including:
[0084] The CPU directly moves the data to be stored in the cache module to the NAND flash memory.
[0085] When the type of the cache module is a volatile cache module, since the data stored in the volatile cache module will be cleared in the event of an accident such as a system power-off or restart, in order to improve the security of data transmission and storage, if the data to be stored is continuous data at this time, the continuous data is directly moved and written into the NAND flash memory.
[0086] In a possible implementation manner, referring to Figure 4 , the CPU processes the data to be stored based on the fourth data processing sub-model and writes it into the NAND flash memory, specifically including:
[0087] The CPU judges whether the remaining capacity of the cache module is lower than a first preset capacity threshold based on the remaining capacity information of the cache module;
[0088] If so, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory;
[0089] If not, the CPU controls the cache module to temporarily store the received data to be stored, and based on its own system operation status, determines whether the system is busy;
[0090] If idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory;
[0091] If busy, the CPU continuously monitors its own system operation status until the system is idle, then the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory.
[0092] In the above process, if the data to be stored is random data, when the remaining capacity of the cache module is relatively small, the CPU quickly integrates the random data into large files and then moves them to the NAND flash memory, avoiding the problem of data loss caused by the inability to continue storing data when the remaining capacity of the cache module is insufficient. When the remaining capacity of the cache module is relatively large, it means that the cache module can continue to store data at this time, so the CPU performs data integration and transfer operations when its own system is idle.
[0093] In a possible implementation manner, for the determination of the first preset capacity threshold, those skilled in the art can adjust it according to the performance and / or actual usage status of the cache module; exemplarily, for a cache module with relatively low performance, the first preset capacity threshold can be set to 20% of the total capacity of the cache module, that is, when the current remaining capacity of the cache module is lower than 20% of the total capacity, it can be considered that the remaining capacity of the cache module has reached a relatively low level and cannot continue to receive data to be stored; for a volatile cache module with relatively high performance, the first preset capacity threshold can be set to 10% of the total capacity of the cache module, that is, when the current remaining capacity of the cache module is lower than 10% of the total capacity, it can be considered that the remaining capacity of the cache module has reached a relatively low level and cannot continue to receive data to be stored.
[0094] In a possible implementation manner, the data processing model further includes a power-off protection processing sub-model; the CPU is used to detect the power-off risk of the cache module based on the power-off protection processing sub-model and perform corresponding protection measures, specifically including:
[0095] If the type of the cache module is a volatile cache module, the CPU obtains the power supply voltage of the cache module in real time and determines whether the cache module has a power-off risk based on the fluctuation of the power supply voltage;
[0096] If so, the CPU moves the data to be stored in the cache module to the NAND flash memory.
[0097] In the above process, by considering the waveform condition of the power supply voltage of the volatile cache module, when the power supply voltage is unstable, there may be a risk of power loss in the volatile cache module. To avoid data loss, at this time, the CPU quickly moves the data to be stored in the cache module to the NAND flash memory for storage, making the data transmission process more secure and reliable.
[0098] In a possible implementation manner, the process of the CPU integrating the data to be stored in the cache module specifically includes the following steps S1 - S3:
[0099] S1, the CPU, based on the data size D of the data to be stored size , divides each piece of the data to be stored into corresponding data sets.
[0100] Specifically, in this step, the division of the data to be stored can be performed based on a preset data set and the corresponding data range of each data set; for example, the preset data set includes a first data set, a second data set, and a third data set, and the first data set, the second data set, and the third data set respectively correspond to different data ranges, and the data ranges increase in sequence; for example, the data range corresponding to the first data set is (0, A], the data range corresponding to the second data set is (A, B], and the data range corresponding to the third data set is (B, C); and A < B ≤ S < C, where S is the page size of the NAND flash memory; C is the file size threshold, which can be considered as the value for dividing continuous data and random data, and A and B are two set boundary values.
[0101] S2, the CPU sorts the data to be stored in each data set based on the sorting rule corresponding to each data set to generate a corresponding data sequence.
[0102] In this step, the sorting rules corresponding to each data set are different; continuing with the above example, for the first data set, its sorting rule is: according to the data size D of each piece of the data to be stored in the first data set size , sorts each piece of the data to be stored to generate a first data sequence; preferably, the sorting can be performed in ascending order.
[0103] For the second data set, its sorting rule is: first, based on the data size D of each piece of the data to be stored in the second data set size , the page size S of the NAND flash memory, and the multiple calculation equation, calculates the multiple n corresponding to each piece of the data to be stored; where the multiple calculation equation is: (n - 1)*S < Dsize ≤n*S, where n≥1 and is an integer. Then, based on the page size S of the NAND flash memory and the data size D of each data to be stored size and the corresponding multiple n, calculate the data difference △D corresponding to each data to be stored; where △D = n*S - D size . Finally, based on the data difference △D corresponding to each data to be stored, sort each data to be stored to generate a second data sequence; preferably, based on the data difference △D corresponding to each data to be stored, sort the corresponding data to be stored in ascending order.
[0104] For the third data set, its sorting rule is the same as that of the second data set. It also calculates the multiple n first, then calculates the data difference △D, and finally sorts the data to be stored according to the data difference △D. Details are not described here again.
[0105] S3. Match and combine the data to be stored in each data sequence to form the integrated data to be stored.
[0106] In this step, matching and combining the data to be stored in each data sequence includes steps S31 - S32:
[0107] S31. Mark the data to be stored in the second data sequence and the third data sequence as data to be matched.
[0108] S32. For each data to be stored in the first data sequence, sequentially find a matching item from the data to be matched according to the data matching rule, and integrate the data to be stored and the corresponding matching item into one data to form the integrated data to be stored.
[0109] Specifically, the data matching rule includes: for the data to be stored in the first data sequence, sequentially find a matching item from each data to be matched. If the sum of the data size of the data to be stored in the first data sequence and the data size of the data to be matched D size-sum satisfies: T < D size-sum ≤n*S; then determine the data to be matched as the matching item of the data to be stored in the first data sequence. Where T is a preset ratio threshold, related to S, for example, T = 90%S; n is the multiple n corresponding to the data to be matched calculated as described above. It should be noted that when sequentially finding a matching item from each data to be matched in this step, it is possible to preferentially find in the second data sequence in order. If not found, then find in the third data sequence in order.
[0110] Through the above process, the functions of integrating small data into big data and integrating random addresses into continuous addresses are realized, thereby ensuring the improvement of random data writing and reading performance.
[0111] It can be understood by those skilled in the art that for the data to be stored that cannot be matched, that is, the data to be stored for which no matching items are found in the first data sequence, and the data in the second data sequence and the third data sequence that are not used as matching items, they are marked as data to be integrated again, and after waiting for the next batch of data to be stored to be divided, they are merged into the corresponding data set, and the above integration process is performed again. The next batch of data to be stored still uses the same integration method. That is, for the data to be stored for which no matching items are found in the first data sequence, they are merged into the first data set of the next batch; for the data in the second data sequence that are not used as matching items, they are merged into the second data set of the next batch; for the data in the third data sequence that are not used as matching items, they are merged into the third data set of the next batch, which will not be repeated here. With the continuous receipt of data, the fusion of new data and unmatched data is also continuously realized.
[0112] Combine the following Figure 5 Explanation of the above integration process:
[0113] The data range of the first data set is preset to (0, 200 Byte], the data range of the second data set is preset to (200 Byte, 600 Byte], and the data range of the third data set is preset to (600 Byte, 1200 Byte); and the page size S of the NAND flash memory is 512 Byte.
[0114] First, based on the data size D of the data to be stored size , and divide each data to be stored into a corresponding data set respectively, and the data to be stored in the first data set includes P1, P2, P3, and P4, and the data size of P1 is 20Byte, the data size of P2 is 45Byte, the data size of P3 is 50Byte, and the data size of P4 is 30Byte; the data to be stored in the second data set includes Q1, Q2, Q3, and Q4, and the data size of Q1 is 300Byte, the data size of Q2 is 400Byte, the data size of Q3 is 350Byte, and the data size of Q4 is 550Byte; the data to be stored in the third data set includes R1, R2, R3, and R4, and the data size of R1 is 700Byte, the data size of R2 is 650Byte, the data size of R3 is 1100Byte, and the data size of R4 is 710Byte.
[0115] If sorted in ascending order, after sorting the data to be stored in the first data set, the first data sequence {P1, P4, P2, P3} can be obtained; after sorting the data to be stored in the second data set, the second data sequence {Q2, Q3, Q1, Q4} can be obtained; after sorting the data to be stored in the third data set, the third data sequence {R4, R1, R2, R3} can be obtained.
[0116] After obtaining the three data sequences, first, for P1, search for matching items in the second and third data sequences in turn to form the integrated data to be stored; the matching process will not be elaborated here.
[0117] In a possible implementation manner, after integrating the data to be stored with the corresponding matching items into one data to form the integrated data to be stored in step S32, the data to be stored and the corresponding matching items can also be marked with a partition identifier, and a mapping relationship can be established based on the integrated data to be stored; it is convenient to first read out the integrated data to be stored based on the mapping table during subsequent reading, and then extract the corresponding data according to the partition identifier. As Figure 5 shown, if Q2 is the matching item of P1, then a partition identifier is used between Q2 and P1.
[0118] In a possible implementation manner, to avoid the risk of power failure, the CPU integrates the data to be stored in the cache module, and every first preset period T1, moves the integrated data to be stored to the NAND flash memory, and can also establish a mapping table to mark the integrated data to be stored; further, the CPU can also release the integrated data from its own cache every second preset period T2; further, the CPU can also move the data to be stored that cannot be matched to the NAND flash memory every third preset period T3 and temporarily store it in its own cache, waiting for a new matching opportunity for the next batch of data.
[0119] An ideal situation achieved by applying such an integration method is that the data size of the data to be stored in the first data sequence is below 0.5S, and the data size of the data to be matched in the second data sequence is between 0.5S and S. After integrating these two data, they can be combined into data with a size of S and then moved.
[0120] In a possible implementation manner, before the cache module continuously receives the data to be stored, it further includes:
[0121] The CPU determines a file size threshold C based on the first data processing capacity of the CPU; the file size threshold C is used to distinguish between continuous data and random data; the first data processing capacity represents the capacity allocated by the CPU for processing the data to be stored.
[0122] Specifically, if the file size of the data to be stored exceeds the file size threshold C, the CPU determines that the data to be stored is continuous data; if the file size of the data to be stored does not exceed the file size threshold C, the CPU determines that the data to be stored is random data.
[0123] Therefore, after the staff allocates the first data processing capacity in the CPU according to actual needs, the CPU automatically determines the file size threshold C, which facilitates subsequent determination of the type of data to be stored.
[0124] Therefore, in the data storage method of the present application, the CPU performs corresponding data processing processes according to factors such as the type of data to be stored, the type and remaining capacity of the cache module, and its own system operating status, performs targeted and refined processing on the data to be stored, and writes the processed data to be stored into the NAND flash memory, improving the speed of writing data into the NAND flash memory, significantly reducing the write amplification factor of the NAND flash memory, improving the resource utilization efficiency, improving the performance of the NAND flash memory, and extending the service life of the NAND flash memory.
[0125] In a possible implementation manner, when the type of the cache module is a non-volatile cache module, the cache module is an area divided inside the NAND flash memory or a device external to the NAND flash memory;
[0126] When the type of the cache module is a volatile cache module, the cache module is a device external to the NAND flash memory.
[0127] The device external to the above NAND flash memory is other devices connected to the NAND flash memory. Therefore, when the cache module is of the non-volatile cache type, those skilled in the art can divide a part of the area inside the NAND flash memory as the cache module, or can choose a device external to the NAND flash memory as the cache module; when the cache module is of the volatile cache type, those skilled in the art can choose a device external to the NAND flash memory as the cache module.
[0128] The embodiment of the present application provides a data storage system, which includes a CPU, a cache module, and a NAND flash memory, and the data storage system applies the above data storage method based on the CPU.
[0129] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application shall be covered within the protection scope of the present application.
Claims
1. A data storage method based on CPU, characterized in that, The CPU is respectively connected to a cache module and a NAND flash memory; the data storage method includes: The cache module continuously receives data to be stored; The CPU reads the data to be stored in the cache module and determines the type of the data to be stored; The CPU obtains the type and remaining capacity of the cache module, and based on the type of the data to be stored, the type and remaining capacity of the cache module, and a pre-constructed data processing model, processes the data to be stored and writes it into the NAND flash memory; Wherein, the type of the data to be stored includes continuous data and random data; the type of the cache module includes a non-volatile cache module and a volatile cache module; The data processing model includes a second data processing sub-model; when the type of the cache module is a non-volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored and writes it into the NAND flash memory based on the second data processing sub-model, specifically including: the CPU determines whether the system is busy based on its own system operation state; if so, the CPU controls the cache module to temporarily store the received data to be stored, and continuously monitors its own system operation state until the system is idle, then the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory; if not, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory; The process of the CPU integrating the data to be stored in the cache module specifically includes: the CPU divides each data to be stored into corresponding data sets based on the data size of the data to be stored; the CPU sorts the data to be stored in each data set based on the corresponding sorting rule of each data set to generate corresponding data sequences; the CPU matches and combines the data to be stored between each data sequence to form the integrated data to be stored.
2. The data storage method based on CPU according to claim 1, characterized in that: The data processing model further includes a first data processing sub-model, a third data processing sub-model, and a fourth data processing sub-model; When the type of the cache module is a non-volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored and writes it into the NAND flash memory based on the first data processing sub-model; When the type of the cache module is a volatile cache module and the type of the data to be stored is continuous data, the CPU processes the data to be stored and writes it into the NAND flash memory based on the third data processing sub-model; When the type of the cache module is a volatile cache module and the type of the data to be stored is random data, the CPU processes the data to be stored and writes it into the NAND flash memory based on the fourth data processing sub-model.
3. The data storage method based on CPU according to claim 2, characterized in that, The CPU processes the data to be stored and writes it into the NAND flash memory based on the first data processing sub-model, specifically including: The CPU determines whether the system is busy based on its own system operation state; If so, the CPU controls the cache module to temporarily store the data to be stored received, and continuously monitors the running state of its own system. Until the system is idle, the CPU moves the data to be stored in the cache module to the NAND flash memory; If not, the CPU directly writes the data to be stored in the cache module into the NAND flash memory.
4. The data storage method based on CPU according to claim 2, characterized in that, The CPU processes the data to be stored based on the third data processing sub-model and writes it into the NAND flash memory, specifically including: The CPU directly moves the data to be stored in the cache module to the NAND flash memory.
5. The data storage method based on CPU according to claim 2, characterized in that The CPU processes the data to be stored based on the fourth data processing sub-model and writes it into the NAND flash memory, specifically including: The CPU determines whether the remaining capacity of the cache module is lower than the first preset capacity threshold based on the remaining capacity information of the cache module; If so, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory; if not, the CPU controls the cache module to temporarily store the data to be stored received and determines whether the system is busy based on the running state of its own system; If idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory; if busy, the CPU continuously monitors the running state of its own system. Until the system is idle, the CPU integrates the data to be stored in the cache module and moves the integrated data to be stored to the NAND flash memory.
6. The data storage method based on CPU according to claim 2, wherein: The data processing model further includes a power-off protection processing sub-model; The CPU is used to detect the power-off risk of the cache module based on the power-off protection processing sub-model and execute corresponding protection measures, specifically including: If the type of the cache module is a volatile cache module, the CPU continuously obtains the power supply voltage of the cache module and determines whether there is a power-off risk for the cache module based on the fluctuation condition of the power supply voltage; If so, the CPU moves the data to be stored in the cache module to the NAND flash memory.
7. The data storage method based on CPU according to claim 1, wherein: When the type of the cache module is a non-volatile cache module, the cache module is an area divided inside the NAND flash memory or a device external to the NAND flash memory; When the type of the cache module is a volatile cache module, the cache module is a device external to the NAND flash memory.
8. A data storage system, characterized in that, Comprising a CPU, a cache module and a NAND flash memory, the data storage system applies the CPU-based data storage method according to any one of claims 1-7.
Citation Information
Patent Citations
Data storage method and device and data query method and device
CN115344201A
Random write instruction processing method, SMR hard disk and computer equipment
CN116048430A
User device including flash and random write cache and method writing data
US20100174853A1