Storage system optimization method based on data compression and dynamic caching

Through dynamic base difference compression and decompression methods and dynamic cache line design, the independent problems of cache management and data compression are solved, efficient data storage and processing are realized, system performance is improved, and efficient and real-time requirements of modern data-intensive applications are met.

CN120336207APending Publication Date: 2025-07-18SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353215.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, cache management and data compression technologies are independent of each other and cannot flexibly adapt to the diverse data access modes and data characteristics, resulting in low cache utilization, many memory access times, and long data access delays, making it difficult to find a balance between compression ratio, compression/decompression delays and hardware complexity.

Method used

Dynamic base value difference compression and decompression methods are adopted to flexibly select base value and difference for data compression, and adjust cache line parameters according to the compression format, combined with dynamic cache line design, optimize cache resource allocation, and improve cache hit rate and system performance.

Benefits of technology

It significantly improves the utilization rate of cache resources, reduces memory access overhead, improves the data processing and storage performance of computer systems, and meets the efficiency and real-time requirements of modern data-intensive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336207A_ABST
    Figure CN120336207A_ABST
Patent Text Reader

Abstract

The invention provides a storage system optimization method based on data compression and dynamic caching, and belongs to the technical field of computer storage and data optimization. Dynamic base value difference compression and decompression are carried out, and data compression is realized by flexibly selecting a base value and calculating a difference value; according to the dynamic cache line compressed and decompressed based on the dynamic base value difference value, related parameters of the cache line are flexibly and intelligently adjusted according to the compression format selected after compression, so that the utilization efficiency of cache resources is improved, and meanwhile, the overall performance of a system is enhanced. Through the dynamic adaptation mechanism, the allocation of cache resources is optimized, the cache hit rate is improved, and the memory access overhead is reduced, so that the data processing and storage performance of a computer system is remarkably improved, the strict requirements of modern data intensive application on the high efficiency and the real-time performance of the system are met, and the defects of the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer storage and data optimization, and particularly to an optimization method for a storage system based on data compression and dynamic caching. Background Art

[0002] With the wide application of emerging technologies such as artificial intelligence, big data analysis, and cloud computing, the amount of data has shown an exponential growth trend. To address this challenge, data storage and compression technologies have been continuously evolving. As a key component for improving the performance of computer systems, on-chip caches are also being continuously optimized to increase data access speed and reduce processor waiting time. At the same time, data compression technologies have been deeply studied and widely applied in various application scenarios, and various algorithms have emerged in an endless stream, aiming to reduce data redundancy, improve data storage efficiency and transmission speed, thereby alleviating storage pressure and reducing the occupancy of memory bandwidth.

[0003] However, the current data storage and compression technologies still face many dilemmas. In terms of cache management, the traditional fixed cache line size design lacks flexibility and cannot adapt to diverse data access patterns and data characteristics. The data block sizes generated by different applications vary greatly. The fixed cache line size either results in wasted space when storing small data blocks, reducing cache utilization; or frequently causes cache misses when processing large data blocks, increasing the number of memory accesses, prolonging data access latency, and thus affecting the performance of the entire system. In the field of data compression, although existing algorithms can achieve data compression to a certain extent, it is often difficult for them to find an ideal balance among compression ratio, compression / decompression latency, and hardware complexity. Some high-compression ratio algorithms require complex calculation processes during decompression, consuming a large amount of time and computing resources, and are not suitable for scenarios with high real-time requirements; while simple compression algorithms have low computational overhead, but the compression effect is limited and they cannot fully exploit the redundant information in the data. More critically, the existing data compression and cache management technologies are independent of each other and lack an effective coordination mechanism, resulting in the inability to fully utilize the advantages of both in practical applications. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides an optimization method for a storage system based on data compression and dynamic caching, which ensures that under the premise of low hardware complexity and short decompression latency, a reasonable data compression method is selected to achieve a high compression ratio, effectively utilize the redundant information in the data, and improve the storage efficiency of the cache. At the same time, it innovatively enables the cache line to autonomously match an appropriate cache line size according to the format selected in the compression system. Through this dynamic adaptation mechanism, the allocation of cache resources is optimized, the cache hit rate is increased, the memory access overhead is reduced, thereby significantly improving the data processing and storage performance of the computer system, meeting the strict requirements of modern data-intensive applications for system efficiency and real-time performance, and making up for the deficiencies of the prior art.

[0005] The technical solution of the present invention is as follows:

[0006] An optimization method for a storage system based on data compression and dynamic caching, comprising:

[0007] Dynamic base value difference compression and decompression, which realizes data compression by flexibly selecting the base value and calculating the difference;

[0008] The dynamic cache line based on dynamic base value difference compression and decompression flexibly and intelligently adjusts the relevant parameters of the cache line according to the selected compression format after compression, so as to improve the utilization efficiency of cache resources and enhance the overall performance of the system.

[0009] Furthermore,

[0010] (1) Dynamic base value difference compression and decompression

[0011] Dynamic base value difference compression and decompression realizes data compression by flexibly selecting the base value and calculating the difference, mainly including a data reading module, a base value selection module, a difference calculation module, and a format selection module.

[0012] The data reading module accurately reads 32-byte storage row data from the storage. For this 32-byte data, it is divided into several data units according to four formats: 8-byte base value, 4-byte base value, 2-byte base value, and no base value division, and is sent in parallel to the difference calculation modules corresponding to the base values.

[0013] The data is sent to multiple difference calculation modules with different difference representation capabilities. Each difference calculation module has a specific difference representation ability, such as 4-byte difference, 2-byte difference, 1-byte difference, etc. Different difference calculation modules perform subtraction operations on each data unit and the first data unit (referred to as the base value unit). In addition, the difference calculation unit also strictly executes a special value compression strategy. For the all-0 or all-1 data that appears in the 32-byte storage row data unit, a special efficient compression method is adopted without performing the conventional base value difference calculation.

[0014] After obtaining the accurate difference, the format judgment module will determine whether the difference is within the range it can represent. For example, the difference calculation module for 2-byte differences will check whether each difference can be completely represented within 2 bytes. If there is a difference greater than 2 bytes, then this difference calculation module will determine that the current data unit cannot be compressed by it, will not output a result, and is regarded as unable to be compressed. And when the differences of the data units processed by a certain difference calculation module are all within its representable range, this module will output these compressed data units as valid results. Therefore, in some situations, the format judgment module will concurrently send multiple groups of compressed data for the format selection module to judge and select. Here, a 4-bit compressed format information code is set, where bits 0-2 are used to represent 8 compression formats, and the 3rd bit is used as the valid bit for compression to represent compression information such as the selected difference, base value, and whether it is compressed.

[0015] The format selection module receives the stored row data after being compressed with different base values, difference combinations, and special values, as well as the accompanying compressed format information code. When multiple compression processing units output results, the format selection module will use the pre-established compression format table as the basis for priority judgment, as shown in the figure below. The compression format table specifies the byte size of the compressed data under different compression formats, and the smaller the size, the higher the priority. For example, in the compression format with a base value of 8 bytes and a difference of 1 byte, a 32-byte stored row is compressed into the form of "8-byte base value + 3 * 1-byte difference", that is, a compressed row of 11 bytes. If none of the units in the upper-layer compression processing module successfully outputs compressed data, that is, it is considered that the stored row cannot pass through any compression processing process, the format selection module will directly output the original 32-byte stored row data without compression. This not only ensures that when compression is possible, the most optimized compression format can be selected to minimize the data storage space occupied, but also ensures that when effective compression cannot be performed, the data will not be lost or misprocessed, guaranteeing the integrity and availability of the data and providing an accurate data basis for subsequent data storage and use.

[0016] The decompression module is as follows:

[0017] The decompression module first receives the data after compression processing and the compressed format information code. The decompression module performs operations on the base value and each difference. Specifically, according to the compressed format information code, it adds the difference to the base value. For example, adding difference 1 to the base value gets another value 1, adding difference 2 to the base value gets another value 2, adding difference 3 to the base value gets another value 3, and so on, gradually restoring each value in the original data unit. During the entire decompression process, for the data part that was previously processed using a special value compression strategy (such as marking or encoding a large number of occurrences of 0 or 1), the decompression module can accurately identify it and perform reverse processing according to the corresponding rules to restore it to the original all-0 or all-1 data form.

[0018] (2) Dynamic cache line:

[0019] The dynamic cache line design based on dynamic base value difference compression and decompression aims to flexibly and intelligently adjust the relevant parameters of the cache line according to the selected compression format after compression, thereby significantly improving the utilization efficiency of cache resources and enhancing the overall performance of the system. This part includes a configuration generation module, a dynamic storage module, and a dynamic replacement management module.

[0020] During the dynamic base value difference compression process, the format selection module will select the optimal compression format as the output, and at the same time output the compression format information code containing compression information such as the base value and the number of difference bytes to the dynamic cache module. These information can accurately reflect the storage requirements of the compressed data, providing a solid and reliable basis for the adjustment of the dynamic cache line. For example, if the finally selected compression format indicates that the amount of compressed data is significantly reduced, it means that the granularity of the cache line can be correspondingly reduced to avoid waste of cache space; on the contrary, if the amount of compressed data changes little, the size of the cache line can be maintained or appropriately adjusted to ensure the efficiency of data access. The following is a detailed description of each module:

[0021] The configuration generation module is the core control part of the dynamic cache line, mainly used to read the compression format information code and parse its information to send to the subsequent modules, providing key parameter support for the subsequent memory configuration and decompression operations. When the cache line is loaded, the hardware decoding circuit will automatically parse the compression format information code and quickly generate the memory array configuration signal.

[0022] The dynamic storage module includes a programmable memory array, which is composed of multiple basic storage units. The fixed 32-byte physical storage space is divided into 4 non-splittable 8-byte physical blocks, which are the basic units of storage. The size and combination mode of the logical basic unit vary according to different compression formats. For example, in the compression format of 8-byte base value and 1-byte difference, the logical unit size is 11 bytes, which will occupy 2 physical blocks (a total of 16 bytes), of which 11 bytes are used to store valid data, and the remaining 5 bytes are used to store various auxiliary information; in the compression format of 4-byte base value plus 2-byte difference, the logical unit size is 18 bytes, occupying 3 physical blocks (a total of 24 bytes), the valid data is 18 bytes, and 6 bytes are used to store other data; in the non-compressed mode, the logical unit is 32 bytes, occupying all 4 physical blocks. These basic storage units are flexibly connected through a crossbar switch matrix, supporting multiple combination modes.

[0023] The crossbar switch matrix is used to implement address remapping, accurately converting the logical unit address into the physical block address. For example, when writing data, when accessing the 11B data of logical unit 0, the crossbar switch matrix can map it to the first 11B space in the 16B storage block composed of physical blocks 0 and 1; when reading data, the crossbar switch matrix will simultaneously read the 16-byte data of physical blocks 0 and 1, and extract the first 11B valid data, masking the last 5B data. Through this method, it is possible to effectively prevent unaligned access across physical blocks, ensuring the accuracy and stability of data access.

[0024] The dynamic storage module integrates an adaptive alignment engine, which integrates a shifter and a mask generator. When dealing with unaligned access, it can automatically perform data bit-width conversion, flexibly switch between different data bit-widths such as 8 bits, 16 bits, 32 bits, etc. according to actual needs; perform byte boundary alignment compensation to ensure the boundary accuracy of data during storage and reading; and perform cross-access protection on compressed blocks to prevent data errors or chaos during access.

[0025] The dynamic replacement management module intelligently determines the cache line eviction priority according to the compression characteristics and access patterns to balance space efficiency and access performance. A value evaluation model (where w1, w2, and w3 are three weights, with different values in different states. For example, in the high-load state, w1:w2:w3 = 5:3:2, emphasizing space optimization; in the normal state, w1:w2:w3 = 3:4:3, as a balanced mode). Through this model, the value scores of cache lines are calculated by integrating multiple dimensions of indicators, and the cache lines with higher scores have higher priorities. At the same time, there are three levels of eviction queues: the main queue is for highly compressed data with high-frequency access, using the LRU (Least Recently Used) replacement mechanism; the secondary queue is for highly compressed data with infrequent access, using the LFU (Least Frequently Used) replacement mechanism; the emergency queue is for uncompressed data and data urgently requiring memory access, using the FIFO (First In First Out) replacement mechanism.

[0026] The beneficial effects of the present invention are

[0027] This technical solution integrates dynamic base value difference compression and decompression and dynamic cache line design, with significant advantages. Through dynamic base value difference compression, the base value and difference can be flexibly selected according to data characteristics, greatly improving the compression ratio, reducing the occupied space of data storage, and at the same time ensuring the accuracy and efficiency of decompression. The dynamic cache line intelligently adjusts parameters according to the compression format, improves the utilization rate of cache resources, reduces access latency, balances space efficiency and access performance, provides all-round optimization for data storage, processing and access, and significantly enhances the overall performance of the system. Brief Description of the Drawings

[0028] Appendix Figure 1It is a schematic diagram of dynamic base value difference compression of the present invention;

[0029] Appendix Figure 2 It is a compression format table of the present invention;

[0030] Appendix Figure 3 It is a decompression schematic diagram of the present invention;

[0031] Appendix Figure 4 It is a schematic diagram of the dynamic cache line architecture of the present invention;

[0032] Appendix Figure 5 It is a schematic diagram of the working process of the present invention. Detailed implementation manners

[0033] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] The present invention provides an optimization method for a storage system based on data compression and dynamic caching. Through a dynamically selectable data compression algorithm, it ensures that under the premise of low hardware complexity and short decompression delay, a reasonable data compression method is selected to achieve a high compression ratio, effectively utilize redundant information in the data, and improve the storage efficiency of the cache. At the same time, it innovatively enables the cache line to autonomously match a suitable cache line size according to the format selected in the compression system. Through this dynamic adaptation mechanism, the allocation of cache resources is optimized, the cache hit rate is increased, the memory access overhead is reduced, thereby significantly improving the data processing and storage performance of the computer system, meeting the strict requirements of modern data-intensive applications for system efficiency and real-time performance, and making up for the deficiencies of the existing technology.

[0035] Suppose a file storage system that needs to process data stored on the hard disk. There is a current storage line of data with a content of 32 bytes. After the data reading module accurately reads this 32-byte storage line of data from the hard disk, it starts to divide it into several data units according to four formats: 8-byte base value, 4-byte base value, 2-byte base value, and no base value division. For the 8-byte base value format, 4 data units are divided and these data units are sent in parallel to the difference calculation module responsible for the 8-byte base value. Similarly, for the 4-byte base value format, the 32-byte data is divided into 8 data units, each unit being 4 bytes, and it is sent to the difference calculation module responsible for the 4-byte base value; the 2-byte base value format and the no base value division format also divide the data units according to their respective rules and send them to the corresponding difference calculation modules.

[0036] The difference calculation module starts to work. Taking the difference calculation module with an 8-byte base value as an example, it performs subtraction operations on the subsequent three data units and the base value unit. And during the calculation process, the difference calculation unit will check if there are data with all 0s or 1s. If so, it will adopt a special value compression strategy. However, there is no such situation in this example.

[0037] After each difference calculation module obtains the differences, the format judgment module starts to work. For example, the difference calculation module responsible for 2-byte differences checks if each difference can be represented within the 2-byte range. Suppose a certain difference it calculates exceeds the 2-byte range, then this module determines that the current data unit cannot be compressed by it and does not output the result; if the differences are all within the range, it outputs the compressed data unit as a valid result. Multiple difference calculation modules may concurrently send multiple sets of compressed data to the format selection module.

[0038] The format selection module receives this compressed data and the 4-bit compressed format information code. Suppose in the compressed format of an 8-byte base value and 1-byte differences, the obtained compressed data is the 8-byte base value plus 3 1-byte differences, totaling 11 bytes, and the number of bytes of the compressed data obtained by other compressed formats is greater than 11 bytes. According to the rule in the compressed format table that "the smaller the size, the higher the priority", the format selection module selects the compressed format of an 8-byte base value and 1-byte differences as the final output, and at the same time outputs the compressed format information code containing compression information such as the base value and the number of difference bytes to the dynamic cache module.

[0039] In the dynamic cache line part, the configuration generation module reads the compressed format information code and generates a storage array configuration signal after parsing. In the programmable storage array in the dynamic storage module, its fixed 32-byte physical storage space is divided into 4 8-byte physical blocks. Since the compressed format of an 8-byte base value and 1-byte differences is selected, when physically storing, physical block 0 stores the 8B base value, and physical block 1 stores Δ1, Δ2, Δ3, and 5B padding data. When the upper layer requests to read logical unit 0 during logical unit access, the crossbar switch matrix will concurrently read 16B data from physical block 0 and block 1. The data splicer extracts the first 11B valid data from it and filters out the last 5B through a mask, and outputs an 11B data stream that conforms to the logical unit definition. From the perspective of space utilization, if there are two such logical units, a total of 4 physical blocks (32B) are occupied, the total effective data is 11B×2 = 22B, and the metadata space is 5B×2 = 10B, strictly ensuring that the total physical space of 32B does not exceed the limit. The adaptive alignment engine in the dynamic storage module can automatically perform data bit-width conversion when processing unaligned access, such as switching between 8 bits, 16 bits, and 32 bits, to perform byte boundary alignment compensation and protect cross-access of compressed blocks.

[0040] The dynamic replacement management module calculates the value score of the cache line according to the compression characteristics and access patterns, using the value evaluation model (assuming a normal state). If this set of data has a high access frequency and good temporal locality, its value score is calculated to be high and it is placed in the main queue, adopting the LRU replacement mechanism.

[0041] When data needs to be read for decompression, the decompression module receives the data after compression processing and the compression format information code. According to the compression format information code, the difference is added to the base value to gradually restore each value in the original data unit. For example, the previously obtained difference is added to the base value to restore the initial 32-byte storage line data, completing the data decompression process and restoring it to the original data state for subsequent use.

[0042] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. An optimization method for a storage system based on data compression and dynamic caching, characterized in that it includes: Dynamic base value difference compression and decompression, which realizes data compression by flexibly selecting a base value and calculating the difference; A dynamic cache line based on dynamic base value difference compression and decompression, which flexibly and intelligently adjusts the relevant parameters of the cache line according to the selected compression format after compression, so as to improve the utilization efficiency of cache resources and enhance the overall performance of the system.

2. The method according to claim 1, characterized in that Dynamic base value difference compression and decompression includes a data reading module, a format judgment module, a difference calculation module, a format selection module, and a decompression module; The data reading module accurately reads 32-byte storage line data from the storage; for this 32-byte data, it is divided into several data units according to four formats: 8-byte base value, 4-byte base value, 2-byte base value, and no base value division, and is sent in parallel to the difference calculation modules corresponding to the base values; The data is sent to several difference calculation modules with different difference representation capabilities. Each difference calculation module has a specific difference representation ability. Different difference calculation modules perform subtraction operations on each data unit and the first data unit, that is, the base value unit; After obtaining the accurate difference, the format judgment module will judge whether the difference is within the range it can represent. If so, it will output the result; The format selection module receives the storage line data processed by compression with different base values, differences, and special values, as well as the accompanying compression format information code; when there are output results from several compression processing units, the format selection module will use the pre-established compression format table as the priority judgment basis; The decompression module is as follows: The decompression module first receives the data processed by compression and the compression format information code; the decompression module performs operations on the base value and each difference. Specifically, according to the compression format information code, the difference is added to the base value.

3. The method according to claim 2, characterized in that In addition, the difference calculation unit will also strictly implement the special value compression strategy. For the data all 0 or 1 that appears in the 32-byte storage line data unit, a compression method is adopted, and there is no need to perform conventional base value difference calculation.

4. The method according to claim 2, characterized in that In some situations, the format judgment module will concurrently send array compressed data for the format selection module to judge and select; here, a 4-bit compression format information code is set, where bits 0-2 are used to represent 8 compression formats, and the 3rd bit is used as the valid bit of compression to represent the selected difference, base value, and whether it is compressed.

5. The method according to claim 2, characterized in that The dynamic cache line includes a configuration generation module, a dynamic storage module, and a dynamic replacement management module; The configuration generation module is used to read the compression format information code, parse its information and send it to the subsequent modules, providing parameter support for the subsequent memory configuration and decompression operations; when the cache line is loaded, the hardware decoding circuit will automatically parse the compression format information code to generate a memory array configuration signal; The dynamic storage module includes a programmable storage array composed of several basic storage units; a fixed 32-byte physical storage space is divided into 4 non-splittable 8-byte physical blocks, which are the basic units of storage; the size and combination mode of the logical basic units vary according to different compression formats, and these basic storage units are flexibly connected through a crossbar switch matrix, supporting several combination modes; An adaptive alignment engine is integrated in the dynamic storage module, and this engine integrates a shifter and a mask generator; when dealing with unaligned accesses, it can automatically perform data bit-width conversion and can flexibly switch between different data bit-widths such as 8 bits, 16 bits, 32 bits, etc. according to actual needs; Perform byte-boundary alignment compensation to ensure the boundary accuracy of data during storage and reading; and perform cross-access protection on compressed blocks; The dynamic replacement management module intelligently determines the eviction priority of cache lines according to compression characteristics and access patterns to balance space efficiency and access performance.

6. The method according to claim 2, wherein, The crossbar switch matrix is used to implement address remapping, accurately converting the logical unit address into a physical block address.

7. The method according to claim 6, wherein, The dynamic replacement management module introduces a value evaluation model Among them, w1, w2, and w3 are three weights with different numerical values in different states.

8. The method according to claim 7, wherein, Through this model, the value scores of cache lines are calculated by synthesizing indicators in several dimensions, and the cache lines with higher scores have higher priorities; at the same time, there are three levels of eviction queues: the main queue is for highly compressed data with high-frequency access, and the LRU replacement mechanism is adopted; the secondary queue is for highly compressed data with infrequent access, and the LFU replacement mechanism is adopted; the emergency queue is for uncompressed data and data urgently requiring memory access, and the FIFO replacement mechanism is adopted.