Configuration method, usage method and configuration system of storage device compression algorithm template
By analyzing data patterns in the storage device and generating a compression algorithm template, matching the optimal compression algorithm for data compression, the problem of increasing the number of flash erases is solved, extending the device life and improving performance.
Patent Information
- Application Number
- CN202510559066.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-30
AI Technical Summary
When existing storage devices write large amounts of repetitive or regular data, they cause the number of erases of flash memory to increase, which loses the service life of the flash memory and reduces the read and write performance.
By receiving host data, the data mode is analyzed according to the preset carry counting system, the preset number of bits and the preset numerical sequence, and the compression algorithm is called to generate a compression algorithm template, matching the optimal compression algorithm for data compression, reducing the number of erasing times of flash memory.
Reduce the number of erasing times of storage devices, extend service life and improve read and write performance, and enhance device stability and data processing capabilities.
Smart Images

Figure CN120085811B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of static storage technology, and in particular to a configuration method, a use method and a configuration system for a storage device compression algorithm template. Background Art
[0002] When writing data to existing storage devices, data can be written directly to the flash memory or written to the cache first and then flushed to the flash memory. Both of these methods do not process the data. When writing large amounts of repetitive data or data with a regular pattern, the number of times the flash memory needs to be erased and written increases significantly, shortening the flash memory's lifespan and causing a decrease in its read and write performance. Therefore, there is room for improvement. Summary of the Invention
[0003] In view of the shortcomings of the prior art described above, the purpose of the present invention is to provide a configuration method, usage method and configuration system for a storage device compression algorithm template, which is used to improve the problem in the prior art that the host data is not reasonably compressed, resulting in a large increase in the number of erase and write times of the flash memory.
[0004] To achieve the above-mentioned and other related objectives, the present invention provides a method for configuring a storage device compression algorithm template, comprising:
[0005] receiving host data, parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and splitting the host data into a plurality of host sub-data according to the data pattern;
[0006] For each data mode of the host sub-data, call all the preset compression algorithms to compress it to obtain the corresponding compressed data, and calculate the corresponding compression ratio;
[0007] For each data pattern, the compression algorithm with the lowest compression ratio is recorded as the optimal compression algorithm, and each data pattern is matched with its corresponding optimal compression algorithm to generate a compression algorithm template;
[0008] The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data.
[0009] In one embodiment of the present invention, the step of parsing the host data according to a preset numbering system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern includes:
[0010] Divide the host data into a plurality of host sub-data according to the writing order of the host data, the preset carry counting system, and the preset number of bits;
[0011] The plurality of host sub-data are compared with a preset numerical sequence to obtain data patterns corresponding to the plurality of host sub-data.
[0012] In one embodiment of the present invention, the step of calling all preset compression algorithms to compress the host sub-data of each data mode to obtain corresponding compressed data includes:
[0013] Determine whether the data patterns of two adjacent host sub-data are the same, and merge the two adjacent host sub-data if they are the same, and separate the two adjacent host sub-data if they are different;
[0014] All preset compression algorithms are called, and according to the arrangement order of the plurality of host sub-data in the host data, the merged host sub-data and the unmerged host sub-data are compressed respectively to obtain corresponding compressed data.
[0015] In one embodiment of the present invention, before the step of receiving host data and parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, the step further includes:
[0016] According to a preset radix counting system, a preset number of bits and a preset numerical sequence, a data mode of host data is configured, and the host data is written into a storage device.
[0017] In one embodiment of the present invention, the data pattern includes at least an all-0 data pattern, an all-1 data pattern, an alternating data pattern, an increasing data pattern, a decreasing data pattern, a pseudo-random data pattern, a checkerboard data pattern, an all-A data pattern, a custom data pattern, and a special character data pattern.
[0018] In one embodiment of the present invention, after the steps of obtaining the compression algorithm with the lowest compression ratio for each data pattern and recording it as the optimal compression algorithm, and matching each data pattern with its corresponding optimal compression algorithm to generate a compression algorithm template, the following steps are further included:
[0019] A data mapping table is created, each data pattern is mapped to a compression algorithm that matches it, an algorithm mapping relationship is generated, and the algorithm mapping relationship is written into the data mapping table.
[0020] In one embodiment of the present invention, after the steps of creating a data mapping table, mapping each data pattern to its matching compression algorithm, generating an algorithm mapping relationship, and writing the algorithm mapping relationship into the data mapping table, the method further includes:
[0021] The host data is compressed to generate compressed data, the logical address of the host data is mapped to the physical address of the corresponding compressed data to generate a data mapping relationship, and the data mapping relationship is written into the data mapping table.
[0022] In one embodiment of the present invention, after the steps of compressing the host data to generate compressed data, mapping the logical addresses of the host data to the physical addresses of the corresponding compressed data to generate a data mapping relationship, and writing the data mapping relationship into the data mapping table, the method further includes:
[0023] After compressing the host data to generate compressed data, the compression algorithm corresponding to the data pattern in the host data is obtained based on the algorithm mapping relationship, and the compression algorithm of the host data is associated with its corresponding data mapping relationship.
[0024] The present invention also provides a method for using a storage device compression algorithm template, comprising:
[0025] Receiving and parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and splitting the host data into a plurality of host sub-data according to the data pattern;
[0026] For the host sub-data of each data pattern, a matching optimal compression algorithm is selected from the compression algorithm template, and the host sub-data of each data pattern is compressed by the matching optimal compression algorithm to obtain corresponding compressed data.
[0027] The present invention also provides a configuration system for a storage device compression algorithm template, comprising:
[0028] a parsing unit, configured to receive and parse the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and split the host data into a plurality of host sub-data according to the data pattern;
[0029] A compression unit is used to compress the host sub-data of each data mode by calling all preset compression algorithms to obtain corresponding compressed data and calculate the corresponding compression ratio;
[0030] A generating unit is configured to obtain, for each data pattern, a compression algorithm with the lowest compression ratio as an optimal compression algorithm, and match each data pattern with its corresponding optimal compression algorithm to generate a compression algorithm template;
[0031] The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data.
[0032] As described above, the configuration method, usage method and configuration system of the storage device compression algorithm template of the present invention reduce the number of erase and write times of the storage device by reasonably compressing the host data, thereby reducing the service life of the storage device and improving the read and write performance of the storage device. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 A schematic flow chart of a method for configuring a storage device compression algorithm template provided by an embodiment of the present invention.
[0035] Figure 2 A schematic diagram illustrating the correspondence between data modes and compression algorithms provided by an embodiment of the present invention.
[0036] Figure 3 A schematic diagram of a flow chart of configuring a compression algorithm for a compression module provided in one embodiment of the present invention.
[0037] Figure 4 A schematic diagram of the structure of a storage device provided in one embodiment of the present invention.
[0038] Figure 5 A schematic diagram of the structure of a compression module in a storage device provided by an embodiment of the present invention.
[0039] Figure 6 This is a structural block diagram of a system for configuring a storage device compression algorithm template provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0041] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0042] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0043] See also Figures 1 to 6 This invention proposes a configuration method, usage method, and configuration system for a storage device compression algorithm template. These methods can be applied to storage devices such as embedded multi-media cards (EMMCs), solid-state drives (SSDs), and universal flash storage (UFS). This invention can reduce the time it takes to read and write large amounts of duplicate data on storage devices, shorten the lifespan of storage devices, and improve read and write performance. A detailed description of this method is provided below using specific embodiments.
[0044] See also Figure 1 In one embodiment of the present invention, a configuration method for a storage device compression algorithm template is proposed, which may include the following steps.
[0045] Step S10: Receive host data, parse the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and split the host data into multiple host sub-data according to the data pattern.
[0046] Specifically, host data refers to application data written by the host. The host can write host data and send commands to the main controller 10, or read host data and receive commands from the main controller 10. The host can be a communication device such as a personal computer (PC), a tablet computer (Pad), or a mobile phone.
[0047] For example, the data pattern includes at least one of an all-0 data pattern, an all-1 data pattern, an alternating data pattern, an increasing data pattern, a decreasing data pattern, a pseudo-random data pattern, a checkerboard data pattern, an all-A data pattern, a custom data pattern, and a special character data pattern.
[0048] All-0 pattern (00000000)
[0049] The binary representation is 0000000000000000000000000000000000, and its uses are as follows.
[0050] Initialization test: Used to verify the behavior of the storage device in the initialized state. Error detection: The all-0 pattern can be used to test whether the storage device can correctly process the lowest level of data. Power consumption test: The all-0 pattern is often used to test the performance of the storage device in low power mode.
[0051] All 1s mode (FFFFFFFF)
[0052] Binary representation: 1111111111111111111111111111111, and its uses are as follows.
[0053] Maximum level test: Used to verify the behavior of the storage device under the highest level data. Error detection: The all-1 pattern can be used to test whether the storage device can correctly handle the highest level data. Endurance test: The all-1 pattern is often used to test the endurance of the storage device, especially during high-frequency write and erase operations.
[0054] Alternating pattern (AAAAAAAA or 55555555)
[0055] The binary representation of AAAAAAAA is: 1010101010101010101010101010101010.
[0056] The binary representation of 55555555 is: 0101010101010101010101010101010101, and its uses are as follows.
[0057] Signal integrity testing: Alternating patterns can be used to test the signal integrity of storage devices, especially during high-speed data transmission. Bit flip testing: Alternating patterns can be used to detect whether a storage device has bit flip issues. Interference testing: Alternating patterns can be used to test the performance of a storage device when subjected to interference.
[0058] Incremental mode (01020304...)
[0059] Example: 0102030405060708... is used as follows.
[0060] Sequential Write Test: Incremental mode can be used to test the performance of a storage device during sequential write operations. Data Integrity Test: Incremental mode can be used to verify that a storage device can correctly store and read sequential data. Error Location: If a data read error occurs, Incremental mode can help quickly locate the error.
[0061] Decreasing mode (FFFEFDFC...)
[0062] Example: FFFEFDCFBFAF9F8... is used as follows.
[0063] Reverse Sequence Test: Descending mode can be used to test the storage device's performance in reverse sequential write operations. Data Integrity Test: Descending mode can be used to verify whether the storage device can correctly store and read reverse sequential data. Error Location: Descending mode can help quickly locate the location of errors.
[0064] Pseudo-random data pattern
[0065] Example: A data pattern generated using a pseudo-random number generator (PRNG) for the following purposes.
[0066] Random Write Test: Pseudo-random data patterns can be used to test the performance of storage devices in random write operations. Stress Test: Pseudo-random data patterns can be used to stress test storage devices, simulating the data distribution in actual use. Error Detection: Pseudo-random data patterns can be used to detect the error rate of storage devices under complex data patterns.
[0067] CheckerboardPattern
[0068] Example: AA55AA55AA55AA55
[0069] Binary representation: 10101010010101011010101001010101... The uses are as follows.
[0070] Interference test: Checkerboard pattern can be used to test the performance of storage devices when subjected to interference. Bit flip test: Checkerboard pattern can be used to detect whether the storage device has bit flip issues. Signal integrity test: Checkerboard pattern can be used to test the signal integrity of the storage device.
[0071] All A mode (A5A5A5A5)
[0072] Binary representation: 10100101101001011010010110100101, and its uses are as follows.
[0073] Specific data pattern test: The all-A pattern can be used to test the performance of a storage device under a specific data pattern. Error detection: The all-A pattern can be used to test whether a storage device can correctly handle a specific data pattern. Durability test: The all-A pattern can be used to test the durability of a storage device.
[0074] User-defined mode
[0075] Example: Users can define specific data patterns as needed, such as DEADBEEF (commonly used for debugging), which is used as follows.
[0076] Debugging tools: User-defined modes can be used to debug the firmware and hardware of storage devices. Specific scenario testing: Users can define specific data patterns based on actual application scenarios to test the performance of storage devices in specific scenarios.
[0077] Blending Mode
[0078] Example: Combining multiple data patterns, such as AA5500FF0102A55A... is used as follows.
[0079] Complex scenario testing: Hybrid mode can be used to test the performance of storage devices under complex data patterns. Stress testing: Hybrid mode can be used to stress test storage devices, simulating the data distribution in actual use. Error detection: Hybrid mode can be used to detect the error rate of storage devices under complex data patterns.
[0080] Special character mode
[0081] Example: Use special characters in the ASCII character set, such as !@#$%^&*(), for the following purposes.
[0082] Compatibility testing: Special character patterns can be used to test the compatibility of storage devices when processing non-standard data. Error detection: Special character patterns can be used to detect whether storage devices can correctly process non-standard data.
[0083] As can be seen above, different data patterns (such as all 0s, all 1s, alternating patterns, increasing patterns, and pseudo-random patterns) are used in different testing and debugging scenarios. These data patterns can help developers and testers verify the performance, reliability, and endurance of storage devices. By using these data patterns, storage device performance can be more comprehensively evaluated and its design and algorithms can be optimized.
[0084] In step S10 , the step of parsing the host data according to a preset radix counting system, a preset number of bits and a preset numerical sequence to obtain a corresponding data pattern may include steps S110 and S120 .
[0085] Step S110 : dividing the host data into a plurality of host sub-data according to the writing order of the host data, the preset carry counting system, and the preset number of bits.
[0086] Specifically, the writing order of the host data can be used to represent the arrangement order of different data in the host data. The value of the corresponding digit in the radix number system depends on its position in the sequence, and can be binary, octal, decimal, or hexadecimal. In this embodiment, binary is used.
[0087] In this embodiment, the preset bit length can be 32 bits, meaning that one host subdata is 32 bits. 32 bits is compatible with processor word lengths (e.g., 32-bit / 64-bit systems) and storage interface bit widths (e.g., eMMC's 8-bit bus). A 32-bit pattern is sufficiently long to expose cross-byte bit errors (e.g., adjacent bit interference). Furthermore, the preset bit width can be adjusted based on actual test objectives. For example, the preset bit width can be 64 bits / 128 bits for more complex, high-precision signal integrity testing.
[0088] Step S120 : Compare the plurality of host sub-data with a preset numerical sequence to obtain data patterns corresponding to the plurality of host sub-data.
[0089] Specifically, the multiple host sub-data can be compared with sequences of different data patterns (such as all 0s, all 1s, alternating patterns, increasing patterns, pseudo-random patterns, etc.) to confirm the data patterns corresponding to the multiple host sub-data.
[0090] In addition, the received host data may correspond to multiple sets of host data, which are sent to the storage device sequentially, each set of host data having the same data pattern, and different sets of host data having different data patterns. Alternatively, the received host data may correspond to multiple sets of host data, which are sent to the storage device simultaneously, each set of host data having the same data pattern, and different sets of host data having different data patterns. Alternatively, the received host data may correspond to a single set of host data, and the single set of host data may have multiple data patterns.
[0091] Step S20: For each data mode of the host sub-data, all preset compression algorithms are called to perform compression processing to obtain corresponding compressed data, and the corresponding compression ratio is calculated.
[0092] In one embodiment of the present invention, step S20 may include step S210 and step S220.
[0093] Step S210 : determining whether the data patterns of two adjacent host sub-data are the same, and merging the two adjacent host sub-data if they are the same, and separating the two adjacent host sub-data if they are different.
[0094] Specifically, since the host data is a batch of data and has multiple different data patterns, there may be cases where two adjacent host sub-data have the same data pattern, and there may also be cases where two adjacent host sub-data have different data patterns. When compressing the host data, the same compression algorithm needs to be used for host data with the same data pattern, and different compression algorithms need to be used for host data with different data patterns.
[0095] Therefore, it is determined whether the data patterns of two adjacent host sub-data are the same. If they are the same, the two adjacent host sub-data are merged; if they are different, the two adjacent host sub-data are separated.
[0096] Step S220 , calling all preset compression algorithms, and compressing the merged host sub-data and the unmerged host sub-data respectively according to the arrangement order of the plurality of host sub-data in the host data, to obtain corresponding compressed data.
[0097] Specifically, in Figure 2 In the data pattern diagram, Data Pattern 1 to Data Pattern 32 represent 32 data patterns abstracted by the host, with 512 bytes or 4 KB as the basic unit. The compression algorithm can be divided into 32 models, each corresponding to a specific Data Pattern.
[0098] Step S30: For each data pattern, the compression algorithm with the lowest compression ratio is obtained and recorded as the optimal compression algorithm. Each data pattern is matched with its corresponding optimal compression algorithm to generate a compression algorithm template. The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data.
[0099] Specifically, the compression ratio is expressed as the ratio of the compressed data size to the corresponding host data size. The compression ratio is calculated as the ratio of compressed data to uncompressed data, or, more accurately, compressed data to host data. A smaller compression ratio indicates better compression, resulting in smaller compressed data and less storage space occupied by the flash memory 50. This reduces the time it takes to write to the flash memory 50 and increases transmission speeds.
[0100] like Figure 1 、 Figure 2 、 Figure 3 As shown, when the host data has only a single data pattern, all preset compression algorithms can be called to compress and obtain compressed data, and the compression algorithm with the lowest compression rate is recorded as the optimal compression algorithm, and the data pattern is matched with its corresponding optimal compression algorithm.
[0101] like Figure 3 As shown, when the host data has only multiple data modes or all data modes, the host data can be parsed to obtain the corresponding different data modes. Then, for each data mode, all preset compression algorithms can be used to compress the compressed data. The compression algorithm with the lowest compression ratio is recorded as the optimal compression algorithm, and each data mode is matched with its corresponding optimal compression algorithm.
[0102] Table 1. Compression algorithm compression ratio correspondence table
[0103]
[0104] As shown in Table 1 and Figure 2 、 Figure 3 As shown, Data1Pattrrn represents the first data pattern, and CompressionAlgorithm1 represents the first compression algorithm. Under each data pattern, the compression algorithm with the lowest compression ratio is obtained and recorded as the optimal compression algorithm, and each data pattern is matched with its corresponding optimal compression algorithm to generate a compression algorithm template.
[0105] Among them, the compression rate of the first data pattern Data1Pattrrn in the first compression algorithm Compression Algorithm1 can reach 90%, and the compression rate of the first data pattern Data1Pattrrn in other compression algorithms Compression Algorithm2 to Compression Algorithm32 is greater than 90%, indicating that the first data pattern Data1Pattrrn can be the most suitable for the first compression algorithm Compression Algorithm1.
[0106] See also Figure 1 In one embodiment of the present invention, before the step of receiving host data in step S30 and parsing the host data according to a preset radix counting system, a preset number of bits and a preset numerical sequence to obtain a corresponding data pattern, step S100 is also included.
[0107] Step S100: configuring a data mode of host data according to a preset radix counting system, a preset number of bits and a preset numerical sequence, and writing the host data into a storage device.
[0108] Specifically, for a storage device, the host data written thereto is arbitrary. In different application scenarios, the host data written to the storage device may have different data volumes and corresponding data modes.
[0109] Therefore, considering the need to match each data pattern with its corresponding optimal compression algorithm, the data pattern of the host data can be configured according to a preset radix counting system, a preset number of bits and a preset numerical sequence, and the host data can be written to the storage device.
[0110] See also Figure 1In one embodiment of the present invention, after step S30, that is, for each data pattern, the compression algorithm with the lowest compression ratio is obtained and recorded as the optimal compression algorithm, and each data pattern is matched with its corresponding optimal compression algorithm. After generating a compression algorithm template, the storage device testing method also includes step S410.
[0111] Step S410: Create a data mapping table, map each data pattern to its matching compression algorithm, generate an algorithm mapping relationship, and write the algorithm mapping relationship into the data mapping table.
[0112] Specifically, by establishing an algorithm mapping relationship between each data pattern and its matching compression algorithm, when the host data is subsequently compressed, the algorithm mapping relationship can be directly searched from the data mapping table, thereby selecting a matching compression algorithm based on the data pattern of the host data for compression processing.
[0113] See also Figure 1 In one embodiment of the present invention, after step S410, step S420 and step S430 are also included.
[0114] Step S420: compress the host data to generate compressed data, map the logical address of the host data to the physical address of the corresponding compressed data to generate a data mapping relationship, and write the data mapping relationship into a data mapping table.
[0115] Step S430: After compressing the host data to generate compressed data, based on the algorithm mapping relationship, obtain the compression algorithm corresponding to the data pattern in the host data, and associate the compression algorithm of the host data with its corresponding data mapping relationship.
[0116] In addition, after the host data is compressed to generate compressed data, the data mapping relationship between the logical address of the host data and the physical address of the corresponding compressed data is written into the data mapping table.
[0117] Specifically, the data mapping table stores mapping information between logical and physical addresses of host data. Logical addresses are relative addresses used in user programs, also known as virtual addresses. Logical addresses are generated by the host controller 10 and are used to access data in the flash memory 50. Physical addresses are actual addresses in the flash memory 50, also known as real addresses. By querying the mapping information in the data mapping table using a logical address, the corresponding physical address can be found, allowing the host data at the corresponding physical address in the flash memory 50 to be read, written, or moved.
[0118] In one embodiment of the present invention, after compressing the host data to generate compressed data, the compression algorithm corresponding to the host data is obtained based on the algorithm mapping relationship, and the compression algorithm of the host data is associated with the data mapping relationship.
[0119] See also Figure 4 The present invention provides a storage device that may include a main controller 10, a compression controller 20, a compression module 40, and a flash memory 50. The main controller 10 is configured to receive host data. A batch of host data may have a single data pattern or multiple data patterns, and the data patterns are used to test, debug, and verify the performance and reliability of the storage device.
[0120] Specifically, the compression controller 20 is used to parse the host data to obtain a corresponding data pattern. The compression module 40 is used to call a matching compression algorithm according to the data pattern of the host data, compress the host data to generate compressed data, and transmit the compressed data to the flash memory 50.
[0121] Thus, by using compression algorithms tailored to the different data patterns of the host data and compressing the host data, the amount of data written to the storage device's flash memory 50 can be reduced, thereby improving the storage device's data processing capabilities. By reducing the number of erase and write cycles required for the storage device's flash memory 50, the lifespan of the flash memory 50 is reduced. By compressing the host data, the time it takes to write the host data to the flash memory 50 is increased, thereby reducing data loss in the event of an abnormal power outage and improving the storage device's operational stability.
[0122] Table 2. Data compression description table
[0123]
[0124] Specifically, for example, if 262,144 KB of host data is stored in 4-KB units, the total number of units required is 262,144 / 4 = 65,536. The compression module 40's firmware internal algorithm detects the data format within each of the 65,536 basic compression units, along with the corresponding compression algorithm. The host data is then compressed based on the corresponding compression algorithm to produce compressed data. The total amount of compressed data is 19,239.5 KB, significantly reducing the actual storage space of the flash memory 50.
[0125] See also Figure 4 and Figure 5In one embodiment of the present invention, the compression module 40 may include a data cache 410 , a first cache 420 , a compression algorithm module 430 , a data check module 440 , a tag module 450 and a second cache 460 .
[0126] Specifically, the data cache area 410 stores host data received from the write cache area 70 .
[0127] The compression algorithm module 430 can check the data pattern of the host data received by the data buffer area 410, call the optimal compression algorithm according to the data pattern, and compress the host data to obtain compressed data.
[0128] The first buffer area 420 is used to store the compressed data cached and compressed from the data buffer area 410 .
[0129] The data verification module 440 is used to verify the compressed data.
[0130] The second buffer area 460 is the buffer where the compressed data is finally output. The second buffer area 460 is used to transmit the compressed data to the flash memory 50 .
[0131] Tag module 450 is used to establish a data mapping relationship between the logical block address (LBA) of host data and the physical block address (PBA) of compressed data formed by compressing the host data, as well as to establish the compression algorithm type corresponding to the compressed data formed by compressing the host data. The purpose of tag module 450 is to infer whether the data at the physical block address of flash memory 50 has been compressed. If the data at the physical block address of flash memory 50 has a corresponding compression algorithm type, it indicates that the data has been compressed.
[0132] See also Figure 4 In one embodiment of the present invention, the storage device further includes a table cache area 80, which stores a data mapping table, maps each data pattern with its corresponding optimal compression algorithm, generates an algorithm mapping relationship, and writes the algorithm mapping relationship into the data mapping table.
[0133] In addition, the compression algorithm may be mapped to the data pattern that matches it to obtain an algorithm mapping relationship, and the algorithm mapping relationship may be written into the data mapping table.
[0134] In addition, when the compression module 40 compresses the host data in the main controller 10 to generate compressed data, the main controller 10 writes the data mapping relationship between the logical address of the host data and the physical address of the corresponding compressed data into the data mapping table.
[0135] Specifically, the data mapping table stores mapping information between logical and physical addresses of host data. Logical addresses are relative addresses used in user programs, also known as virtual addresses. Logical addresses are generated by the host controller 10 and are used to access data in the flash memory 50. Physical addresses are actual addresses in the flash memory 50, also known as real addresses. By querying the mapping information in the data mapping table using a logical address, the corresponding physical address can be found, allowing the host data at the corresponding physical address in the flash memory 50 to be read, written, or moved.
[0136] See also Figure 4 In one embodiment of the present invention, when the compression module 40 compresses the host data to generate compressed data, the main controller 10 is used to obtain the compression algorithm corresponding to the host data based on the algorithm mapping relationship, and associate the compression algorithm of the host data with the data mapping relationship.
[0137] See also Figure 4 In one embodiment of the present invention, the storage device further includes a decompression module 30. The decompression module 30 selects a corresponding decompression algorithm based on the data mapping relationship and compression algorithm corresponding to the compressed data to decompress the compressed data and generate host data. The decompression algorithm is preconfigured, and one decompression algorithm corresponds to one compression algorithm.
[0138] Specifically, by writing the mapping relationship between the logical address of the host data and the physical address of the corresponding compressed data into the data mapping table, and also associating the type of the configured compression algorithm corresponding to the host data with the mapping relationship of the host data, when the decompression module 30 reads data from the physical block address of the flash memory 50, it can check whether the data has the corresponding compression algorithm type.
[0139] When the data has a corresponding compression algorithm type, it indicates that the data is compressed into compressed data. The decompression module 30 can select a corresponding decompression algorithm based on the type of compression algorithm to decompress the compressed data, thereby obtaining corresponding host data.
[0140] When the data does not have a corresponding compression algorithm type, it indicates that the data has not been compressed into host data. The decompression module 30 can directly read the host data and transmit the host data to the main controller 10 .
[0141] See also Figure 4In one embodiment of the present invention, the storage device further includes a write cache area 70 and a read cache area 60. The write cache area 70 is used to cache host data read by the main controller 10, and the host data is compressed by the compression module 40. The read cache area 60 is used to cache host data decompressed by the decompression module 30, and the decompressed host data is transmitted to the host through the main controller 10.
[0142] See also Figure 4 and Figure 5 In one embodiment of the present invention, after compressing the host data in the main controller 10 to generate compressed data, the compression module 40 verifies the compressed data and writes the verification code into the data mapping table. In the data mapping table, the verification code corresponding to the compressed data is associated with the mapping relationship between the compressed data.
[0143] Specifically, the application of checksums in solid-state storage primarily aims to ensure data integrity and reliability. With the advancement of storage technology, checksums play a crucial role, particularly in flash memory-based storage devices such as solid-state drives (SSDs), eMMC, and UFS. Setting a checksum allows for error detection and correction, data recovery, improved durability, long-term storage stability, and enhanced security.
[0144] See also Figure 4 In one embodiment of the present invention, the storage device further includes a power supply module 101 , a serial port module 102 , an interface module 103 , a read-only memory 104 and a static random access memory 105 .
[0145] Specifically, the power supply module 101 can provide power to the main controller 10. The serial port module (UART) 102 is a widely used interface for serial communications, allowing two devices to exchange data through a serial port without using a common clock signal. The read-only memory 104 can be used to store firmware for the storage device.
[0146] The present invention also proposes a method for using a storage device compression algorithm template, which may include the following steps.
[0147] Step S510: Receive and parse the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and split the host data into multiple host sub-data according to the data pattern.
[0148] Step S520: For the host sub-data of each data pattern, select a matching optimal compression algorithm from the compression algorithm template, and compress the host data of each data pattern using the matching optimal compression algorithm to obtain corresponding compressed data.
[0149] It can be seen that according to different data modes of the host data, using corresponding compression algorithms to compress the host data can reduce the amount of data written to the flash memory 50 of the storage device and improve the data processing capability of the storage device.
[0150] See also Figure 6 The present invention further provides a storage device testing system 100 , which may include a parsing unit 110 , a compression unit 120 , and a generation unit 130 .
[0151] The parsing unit 110 is used to receive host data, parse the host data according to a preset radix counting system, a preset number of bits and a preset numerical sequence to obtain a corresponding data pattern, and split the host data into multiple host sub-data according to the data pattern.
[0152] The compression unit 120 is used to call all preset compression algorithms to perform compression processing on the host sub-data of each data mode to obtain corresponding compressed data, and calculate the corresponding compression ratio.
[0153] The generating unit 130 is configured to obtain, for each data pattern, a compression algorithm with the lowest compression ratio as an optimal compression algorithm, and match each data pattern with its corresponding optimal compression algorithm to generate a compression algorithm template.
[0154] The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data.
[0155] Thus, by using compression algorithms tailored to the different data patterns of the host data and compressing the host data, the amount of data written to the storage device's flash memory 50 can be reduced, thereby improving the storage device's data processing capabilities. By reducing the number of erase and write cycles required for the storage device's flash memory 50, the lifespan of the flash memory 50 is reduced. By compressing the host data, the time it takes to write the host data to the flash memory 50 is increased, thereby reducing data loss in the event of an abnormal power outage and improving the storage device's operational stability.
[0156] In summary, the configuration method, usage method, and configuration system for a storage device compression algorithm template disclosed in this invention reduce the time it takes to read and write large amounts of duplicate data on a storage device, shorten the lifespan of the storage device, and improve read and write performance. Therefore, this invention effectively overcomes the shortcomings of the prior art and possesses high industrial value.
[0157] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. A method for configuring a storage device compression algorithm template, characterized in that: include: receiving host data, parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and splitting the host data into a plurality of host sub-data according to the data pattern; For each data mode of the host sub-data, call all the preset compression algorithms to compress it to obtain the corresponding compressed data, and calculate the corresponding compression ratio; For each data pattern, the compression algorithm with the lowest compression ratio is recorded as the optimal compression algorithm, and each data pattern is matched with its corresponding optimal compression algorithm to generate a compression algorithm template; The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data. The step of parsing the host data according to the preset radix counting system, the preset number of bits and the preset numerical sequence to obtain the corresponding data pattern includes: Divide the host data into a plurality of host sub-data according to the writing order of the host data, the preset carry counting system, and the preset number of bits; Comparing the plurality of host sub-data with a preset numerical sequence to obtain data patterns corresponding to the plurality of host sub-data; The step of calling all preset compression algorithms to compress the host sub-data of each data mode to obtain corresponding compressed data includes: Determine whether the data patterns of two adjacent host sub-data are the same, and merge the two adjacent host sub-data if they are the same, and separate the two adjacent host sub-data if they are different; All preset compression algorithms are called, and according to the arrangement order of the plurality of host sub-data in the host data, the merged host sub-data and the unmerged host sub-data are compressed respectively to obtain corresponding compressed data.
2. The method for configuring a storage device compression algorithm template according to claim 1, wherein: Before the step of receiving host data and parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, the method further includes: According to a preset radix counting system, a preset number of bits and a preset numerical sequence, a data mode of host data is configured, and the host data is written into a storage device.
3. The method for configuring a storage device compression algorithm template according to claim 1, wherein: The data pattern includes at least an all-0 data pattern, an all-1 data pattern, an alternating data pattern, an increasing data pattern, a decreasing data pattern, a pseudo-random data pattern, a checkerboard data pattern, an all-A data pattern, a custom data pattern, and a special character data pattern.
4. The method for configuring a storage device compression algorithm template according to claim 1, wherein: After the steps of obtaining the compression algorithm with the lowest compression ratio for each data pattern and recording it as the optimal compression algorithm, and matching each data pattern with its corresponding optimal compression algorithm to generate a compression algorithm template, the method further includes: A data mapping table is created, each data pattern is mapped to a compression algorithm that matches it, an algorithm mapping relationship is generated, and the algorithm mapping relationship is written into the data mapping table.
5. The method for configuring a storage device compression algorithm template according to claim 4, wherein: After the steps of creating a data mapping table, mapping each data pattern to its matching compression algorithm, generating an algorithm mapping relationship, and writing the algorithm mapping relationship into the data mapping table, the method further includes: The host data is compressed to generate compressed data, the logical address of the host data is mapped to the physical address of the corresponding compressed data to generate a data mapping relationship, and the data mapping relationship is written into the data mapping table.
6. The method for configuring a storage device compression algorithm template according to claim 5, characterized in that: After the steps of compressing the host data to generate compressed data, mapping the logical addresses of the host data to the physical addresses of the corresponding compressed data to generate a data mapping relationship, and writing the data mapping relationship into the data mapping table, the method further includes: After compressing the host data to generate compressed data, the compression algorithm corresponding to the data pattern in the host data is obtained based on the algorithm mapping relationship, and the compression algorithm of the host data is associated with its corresponding data mapping relationship.
7. A method for using a storage device compression algorithm template, applying the configuration method of the storage device compression algorithm template according to any one of claims 1 to 6, characterized in that: include: Receiving and parsing the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and splitting the host data into a plurality of host sub-data according to the data pattern; For the host sub-data of each data pattern, a matching optimal compression algorithm is selected from the compression algorithm template, and the host sub-data of each data pattern is compressed by the matching optimal compression algorithm to obtain corresponding compressed data.
8. A configuration system for a storage device compression algorithm template, characterized in that: include: a parsing unit, configured to receive host data, parse the host data according to a preset radix counting system, a preset number of bits, and a preset numerical sequence to obtain a corresponding data pattern, and split the host data into a plurality of host sub-data according to the data pattern; A compression unit is used to compress the host sub-data of each data mode by calling all preset compression algorithms to obtain corresponding compressed data and calculate the corresponding compression ratio; A generating unit is configured to obtain, for each data pattern, a compression algorithm with the lowest compression ratio as an optimal compression algorithm, and match each data pattern with its corresponding optimal compression algorithm to generate a compression algorithm template; The compression ratio is expressed as the ratio of the amount of compressed data to the amount of its corresponding host data. The step of parsing the host data according to the preset radix counting system, the preset number of bits and the preset numerical sequence to obtain the corresponding data pattern includes: Divide the host data into a plurality of host sub-data according to the writing order of the host data, the preset carry counting system, and the preset number of bits; Comparing the plurality of host sub-data with a preset numerical sequence to obtain data patterns corresponding to the plurality of host sub-data; The step of calling all preset compression algorithms to compress the host sub-data of each data mode to obtain corresponding compressed data includes: Determine whether the data patterns of two adjacent host sub-data are the same, and merge the two adjacent host sub-data if they are the same, and separate the two adjacent host sub-data if they are different; All preset compression algorithms are called, and according to the arrangement order of the plurality of host sub-data in the host data, the merged host sub-data and the unmerged host sub-data are compressed respectively to obtain corresponding compressed data.
Citation Information
Patent Citations
Data compression method and device
CN117220685A