Data management method and device, equipment and medium
By determining whether the data meets the online compression conditions in the storage device and performing online or background compression processing, the dependence problem of storage devices on hardware compression cards is solved, and dynamic balance of performance and efficiency and cost reduction are achieved.
Patent Information
- Application Number
- CN202510715327.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, storage devices have strong dependence on hardware compression cards, resulting in performance bottlenecks and increased costs under high load conditions, and the compression algorithm is not able to adapt to different data types.
By determining whether the data to be written meets the online compression conditions, perform online compression or directly write to the disk, and perform compression processing in the background, combining online compression and background compression, the compression algorithm is dynamically adjusted to adapt to different data types.
The dynamic balance between real-time performance and storage efficiency of storage devices is achieved, reducing dependence on hardware compression cards, avoiding additional delays, and reducing the cost of the whole machine.
Smart Images

Figure CN120234306A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data compression technology, and particularly to data management methods, devices, equipment, and media. Background Art
[0002] Currently, the cost of all-flash storage is relatively high, and there is a problem of limited number of erase-write cycles. Through compression technology, the occupied space of data can be reduced, storage costs can be saved, storage efficiency can be improved, and the data written to disk is less after compression, which can effectively reduce IO (Input / Output) operations, reduce write amplification, improve the overall throughput, and greatly reduce the number of erase-write cycles when the same volume of data is written to disk, thereby reducing flash wear and extending the service life of SSD (Solid State Drive).
[0003] In related technologies, the compression function is implemented by externally inserting a QAT (Quick Assist Technology) compression card. This type of compression with the help of a compression card usually adds a specific module on the host IO path of the storage system for interacting with the compression card. For example, after the storage system receives a write IO issued by the host, it goes through a series of processes to reach the module that interacts with the compression card, and this module submits it to the compression card. After waiting for the IO to be processed by the compression card, it returns to the host IO process of the storage system for subsequent processing until the data is written to disk. However, this method is highly dependent on the compression card and belongs to real-time compression. Once the compression card fails, data compression cannot be performed.
[0004] Therefore, how to reduce the dependence of storage devices on hardware compression cards during the data compression process is an urgent problem to be solved currently. Summary of the Invention
[0005] This application provides a data management method, device, equipment, and medium to at least solve the problem of reducing the dependence of storage devices on hardware compression cards during the data compression process in related technologies.
[0006] This application provides a data management method, and the method includes: In response to a data write request sent by a host, determining whether the data to be written in the data write request meets the online compression condition; If the data to be written meets the online compression condition, performing online compression processing on the data to be written, obtaining first compressed data, and writing the first compressed data to a disk; If the data to be written does not meet the online compression condition, write the data to be written to the disk, and when background data compression is triggered, perform background compression processing on the data to be written to obtain second compressed data, and update the data to be written to the second compressed data.
[0007] The present application also provides a data management device, which includes: A condition judgment module, configured to, in response to a data write request sent by a host, judge whether the data to be written in the data write request meets the online compression condition; An online compression module, configured to, if the data to be written meets the online compression condition, perform online compression processing on the data to be written to obtain first compressed data, and write the first compressed data to the disk; A background compression module, configured to, if the data to be written does not meet the online compression condition, write the data to be written to the disk, and when background data compression is triggered, perform background compression processing on the data to be written to obtain second compressed data, and update the data to be written to the second compressed data.
[0008] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above data management methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above data management methods are implemented.
[0010] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above data management methods are implemented.
[0011] Through the present application, by judging whether the data to be written meets the online compression condition, for the data that meets the compression condition, online compression processing is performed, and for the data that does not meet the compression condition, the data is directly written to the disk and compressed in the background, which can avoid the additional delay caused by online compression. Through the combination of online compression and background compression, the storage device can achieve a dynamic balance between real-time performance and storage efficiency, and, through the cooperation of online compression and background compression, the coupling relationship between the compression function and the host IO process is weakened, eliminating the dependence on the hardware compression card. Description of the Drawings
[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a schematic flowchart of a data management method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a data management device provided by an embodiment of the present application. Detailed implementation manners
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0016] To enable those skilled in the art of this technology to better understand the solutions of the present application, the following will further elaborate on the present application in detail with reference to the drawings and specific implementation manners.
[0017] Term explanation: QTA compression card: It is a hardware acceleration card used to accelerate data compression and decompression, mainly used to improve the efficiency of storage and network transmission.
[0018] Host IO path: It refers to the path of a series of hardware components and software modules through which data passes when the host performs input / output operations with external storage devices (such as hard disks, solid-state drives, etc.).
[0019] In the related art, the compression function is implemented by an external QAT (Quick Assist Technology) compression card. This compression method relying on the compression card usually adds a specific module to the host IO path of the storage system for interacting with the compression card. For example, after the storage system receives a write IO sent by the host, it goes through a series of processes to reach the module that interacts with the compression card, and then this module submits it to the compression card. After waiting for the IO to be processed by the compression card, it returns to the host IO process of the storage system for subsequent processing until the data is written to the disk. However, this method is highly dependent on the compression card and belongs to real-time compression. Once the compression card fails, data compression cannot be performed. Moreover, when the business pressure is low, the compression card may not become a bottleneck. However, once the business pressure reaches a certain level, the processing flow of the compression card will inevitably affect the IO response time and performance. When the business pressure reaches the maximum concurrency that the compression card can handle, the compression card will become a performance bottleneck, and the cost of solving this hardware bottleneck is relatively high. Furthermore, the compression algorithm in the compression card is fixed, and the same compression algorithm can only be used regardless of the type of data (text, picture, video), with high limitations. Finally, due to the dependence of this compression technology on the hardware compression card, when enabling the compression function, at least one compression card must be configured for each device, which increases the shipping cost of the storage device.
[0020] Based on the above problems, embodiments of the present application provide a data management method, which is applied to an all-flash storage device / system. In combination with the execution process of the data management method, the method will be described in detail.
[0021] Refer to Figure 1 As shown, the data management method provided by the embodiments of the present invention includes the following steps: S11. In response to a data write request sent by the host, determine whether the data to be written in the data write request meets the online compression condition.
[0022] Specifically, the data write request carries the data to be written. The all-flash storage system / device responds to the data write request sent by the host and determines whether the data to be written meets the online compression condition.
[0023] Exemplarily, for a write IO sent by the host, in the all-flash stack, it is generally split into one data block after another for processing. In response to the data write request sent by the host, the write IO is split into multiple data blocks, that is, multiple data to be written. By determining whether each data block meets the online compression condition, it is realized to determine whether the data to be written meets the online compression condition.
[0024] In some embodiments, the above step S11 (determine whether the data to be written in the data write request meets the online compression condition) can be implemented in the following manner: Determine whether the proportion of preset data in the data to be written is greater than a preset proportion, and whether the front-end load is less than a preset load; If the proportion of preset data in the data to be written is greater than the preset proportion, and the front-end load is less than the preset load, then determine that the data to be written meets the online compression condition; If the proportion of preset data in the data to be written is less than or equal to the preset proportion, and / or the front-end load is greater than or equal to the preset load, then determine that the data to be written does not meet the online compression condition.
[0025] Among them, the preset data can be non-zero data. The preset proportion can be determined according to the actual application scenario. For example, the preset proportion can be selected as 20%, 30%, etc., or other reasonable values, and no specific limitation is made here. When the proportion of non-zero data in the data to be written is relatively high, it indicates that the randomness of the data to be written is relatively large, and the repetition pattern and redundant information are relatively small. In this case, even if the data to be written is compressed, the space-saving effect obtained may be limited and may not be able to make up for the system resources and time costs consumed during the compression process.
[0026] The preset load refers to the maximum input / output request volume, data throughput, concurrent connection number, etc. that the front end of the all-flash storage system / device can handle. In the embodiments of the present disclosure, the preset load can represent the maximum input / output request volume that the front end of the all-flash storage system / device can handle. The preset load can be set according to the actual application scenario. For example, the preset load can be selected as 30, 50, etc., or other reasonable values, and no specific limitation is made here.
[0027] Specifically, by determining whether the front-end load is less than the preset load, and determining whether the proportion of preset data in the data to be written is greater than the preset proportion, it can be determined whether the data to be written meets the online compression condition. If the front-end load is less than the preset load, and the proportion of non-zero data in the data to be written is less than or equal to the preset proportion, then determine that the target data block meets the online compression condition; if the front-end load is greater than or equal to the preset load, and / or the proportion of non-zero data in the data block to be written is greater than the preset proportion, then determine that the target data block does not meet the online compression condition, skip the online compression process, and directly write to disk.
[0028] Under high load conditions, that is, when the front-end load is greater than or equal to the preset load, the all-flash storage system actively adjusts the concurrency of online compression. By dynamically adjusting the concurrency, the system can reasonably allocate resources such as CPU and memory according to the actual load situation, avoiding over-occupation or waste of resources. Among a unit quantity of data (such as a cumulative 100 data blocks), a part of them (such as 50 data blocks) is allowed not to be processed by online compression and is directly written to the disk. Under high load conditions, reducing the amount of data for online compression can effectively relieve the pressure on the online compression module and avoid performance bottlenecks caused by insufficient compression processing capabilities.
[0029] Exemplarily, assume that the preset quantity is 50. Under high load conditions, it is judged whether the number of data blocks in the data to be written is greater than 50. If the number of data blocks is greater than 50, 50 data blocks in the data to be written are processed by online compression, and the remaining 50 data blocks are not processed by online compression and are directly written to the disk. By reducing the overhead of online compression, the all-flash storage system can write data to the flash memory faster, giving full play to the high-performance advantages of the flash memory. This can effectively improve the IO processing ability of the all-flash storage system under high load conditions, reduce the overall latency, and prevent online compression from becoming a performance bottleneck.
[0030] By judging whether the proportion of preset data (such as compressible data) in the data to be written is greater than the preset proportion, and whether the front-end load is less than the preset load, the system can dynamically decide whether to perform online compression, and can accurately judge whether to perform online compression according to the actual compressibility of the data and the current load situation of the system, avoiding unnecessary compression operations. For example, when the data is not compressible or the compression effect is not good (such as a low proportion of preset data), or the system load is too high, the online compression will be skipped and the data will be directly written to the disk, avoiding the additional resource overhead caused by the compression process.
[0031] Since frequent online compression processing may increase the latency of data writing. By performing online compression under appropriate circumstances, the system can avoid the increase in latency caused by the compression process, thereby improving the overall IO performance. That is, performing online compression when the front-end load is low can make full use of system resources and improve the throughput of data processing. Through online compression, the system can reduce the amount of data before writing to the disk, thereby saving storage space.
[0032] In some embodiments, the above step S11 (judging whether the data to be written in the data write request meets the online compression condition) can also be implemented in the following manner: Pre-compress the first preset number of bytes of the data to be written, and calculate the compression ratio according to the compressed data and the data before compression; Determine whether the compression ratio is greater than a preset compression ratio and whether the front - end load is less than a preset load; If the compression ratio is greater than the preset compression ratio and the front - end load is less than the preset load, it is determined that the data to be written meets the online compression condition; If the compression ratio is less than or equal to the preset compression ratio and / or the front - end load is greater than or equal to the preset load, it is determined that the data to be written does not meet the online compression condition.
[0033] Among them, the preset value can be selected according to the actual situation. For example, the preset value can be selected as 256, or other reasonable values, and no specific limitation is made here. The preset compression ratio can be set according to the actual situation. For example, the preset compression ratio can be set to 40%, 50%, etc., or other reasonable values, and no specific limitation is made here.
[0034] Specifically, pre - compress the first preset number of bytes of the data to be written, calculate the compression ratio of the first preset number of bytes of the data to be written according to the compressed data and the uncompressed data. If the compression ratio is greater than the preset compression ratio and the front - end load is less than the preset load, it is determined that the data to be written meets the online compression condition, and the data to be written can be processed by online compression. If the compression ratio is less than or equal to the preset compression ratio and / or the front - end load is greater than or equal to the preset load, it is determined that the data to be written does not meet the online compression condition, and the data to be written is directly written to the disk, and the compression processing task is completed through background compression subsequently.
[0035] Exemplarily, taking the first 256 bytes of the data to be written as an example, although the first 256 bytes of the data to be written only account for a small part of the entire data to be written, they can contain some key information of the data to be written, such as file headers, metadata, etc. These information have certain regularity and repeatability, and have important reference value for judging the compressibility of the data. Selecting the first 256 bytes for analysis can greatly reduce the calculation amount and improve the processing efficiency of the all - flash storage system on the premise of ensuring a certain judgment accuracy.
[0036] Suppose the size of the data to be written is 8K, the preset compression ratio is 30%, pre - compress the first 256 bytes of the data to be written, and after compression, it is 200 bytes. According to the compression ratio calculation formula, the compression ratio can be calculated as (256 bytes - 200 bytes) / 256 bytes×100%≈21.9%. Since 21.9%<30%, the data to be written does not meet the online compression condition, and the data to be written is directly written to the disk, and the compression processing task for the data to be written is completed through background compression subsequently.
[0037] S12. If the data to be written meets the online compression condition, perform online compression processing on the data to be written, obtain the first compressed data, and write the first compressed data to the disk.
[0038] Among them, the first data is the data that meets the online compression condition. In the embodiments of the present disclosure, the first compression algorithm is a compression algorithm applicable to the online compression scenario. For example, the first compression algorithm can be the LZ77 algorithm, LZ78, LZMA algorithm, etc., and specific limitations are not made here.
[0039] Specifically, if the data to be written meets the online compression condition, perform online compression processing on the data to be written through the online compression algorithm, obtain the first compressed data, and write the first compressed data to the disk.
[0040] S13. If the data to be written does not meet the online compression condition, write the data to be written to the disk, and when the background data compression is triggered, perform background compression processing on the data to be written, obtain the second compressed data, and update the data to be written to the second compressed data.
[0041] Specifically, if the data to be written does not meet the online compression condition, write the data to be written to the disk, and when the background data compression is triggered, perform background compression processing on the data to be written through the background compression algorithm, obtain the second compressed data, and update the data to be written to the second compressed data. Among them, the time consumption of the online compression algorithm is lower than that of the background compression algorithm. There are many choices for the compression algorithm, such as the LZ77 algorithm, LZ78 algorithm, LZMA algorithm, etc. An adaptive compression process with a compression ratio from low to high, or an adaptive compression process with a time consumption from low to high, or multiple different compression algorithms can be implemented to handle different file types and business scenarios. In the embodiments of the present disclosure, two compression algorithms can be implemented, a simple algorithm with low time consumption and low compression ratio, and a complex algorithm with high time consumption and high compression ratio. It can be understood that generally, for an algorithm with a low compression ratio, its time consumption is less; for an algorithm with a high compression ratio, the corresponding time consumption is more. Therefore, a compression algorithm with less time consumption can be selected for online compression, and a compression algorithm with a high compression ratio can be selected for background compression.
[0042] For the background compression task, the appropriate concurrency can be selected by monitoring the front-end load. When the front-end pressure is high, actively reduce or even stop the background compression task. When the front-end pressure is low, increase the concurrency of the background compression task.
[0043] The online compression algorithm ensures fast response, while the background compression algorithm further optimizes the storage space in the background, enabling the system to exhibit good performance and efficiency in different scenarios. At the same time, the compression function is implemented through different compression algorithms, that is, through software, which decouples the compression process from the hardware compression card and eliminates the storage device's dependence on the hardware compression card.
[0044] In the embodiments of the present disclosure, the compression function is implemented through a software pre-compression algorithm, which is different from the way of using a hardware compression card to implement the compression function in the related art. This way not only achieves complete decoupling from the hardware and improves robustness, but also prevents the hardware from becoming a performance bottleneck and reduces the impact of the compression process on the front-end performance. Additionally, it can implement the function of selecting different compression algorithms for different types of data through the background compression process, improving the compression ratio. Finally, reducing the use of hardware compression cards can also reduce the overall cost of the machine.
[0045] In some embodiments, before writing the first compressed data to the disk, the preset identifier of the data to be written is updated to a first identifier.
[0046] Among them, the first identifier is used to identify the data obtained through online compression processing.
[0047] The preset identifier is used to indicate whether the data to be written has been compressed, and further, the preset identifier can also be used to indicate which algorithm the data to be written is compressed by. For example, two bits can be used to represent the preset identifier. "00" indicates that the data to be written has not been compressed, "01" indicates that the data to be written has been compressed, and "02" indicates that the data to be written has been compressed by the background compression algorithm. Another example is that "00" indicates that the data to be written has not been compressed, "01" indicates that the data to be written has been compressed by compression algorithm 1, "02" indicates that the data to be written has been compressed by compression algorithm 2, "03" indicates that the data to be written has been compressed by compression algorithm 3, etc.
[0048] Specifically, before writing the first compressed data to the disk, the preset identifier of the data to be written is updated to the first identifier, that is, the first identifier is used to indicate that the first compressed data has been processed by the online compression algorithm.
[0049] Before writing the data to be written to the disk, the preset identifier is updated to a second identifier.
[0050] Among them, the second identifier is used to identify uncompressed data.
[0051] Specifically, before writing the data to be written to the disk, update the preset identifier to the second identifier, that is, the second identifier is used to indicate that the data to be written has not been compressed.
[0052] Before updating the data to be written to the second compressed data, update the preset identifier to the third identifier.
[0053] Among them, the third identifier is used to identify the data obtained through background compression processing.
[0054] Specifically, after obtaining the second compressed data and before updating the data to be written to the second compressed data, the preset identifier of the second compressed data can be updated to the third identifier first, that is, the third identifier is used to indicate that the second compressed data has been compressed by the background compression algorithm.
[0055] By updating the identifier of the data before writing it to the disk, it is ensured that the compression state of the data is always consistent with the identifier. At the same time, the system can distinguish online-compressed, uncompressed, and background-compressed data according to the identifier, thus supporting multiple compression strategies. In this way, in subsequent operations such as reading and backup, the all-flash storage system can quickly identify the compression state of the data according to the identifier, and then adopt the corresponding decompression algorithm for processing to adapt to different application scenarios. For example, the online compression method is used in scenarios with high real-time requirements, and the background compression method is used in scenarios with high storage efficiency requirements.
[0056] Optionally, the above step S13 (when triggering background data compression, perform background compression processing on the data to be written) can be implemented in the following way: Obtain the data in the disk with the preset identifier being the second identifier; Perform background compression processing on the data with the preset identifier being the second identifier.
[0057] Specifically, when triggering background data compression and performing background compression processing on the data to be written, the data with the preset identifier being the second identifier in the disk can be obtained, that is, the uncompressed data, and background compression processing is performed on the uncompressed data.
[0058] In some embodiments, the above step 13 (perform background compression processing on the data to be written to obtain the second compressed data) can be implemented in the following way: Periodically read the data to be written stored in the disk and identify the data type of the data to be written; Determine the background compression algorithm corresponding to the data to be written according to the data type; Perform compression processing on the data to be written according to the background compression algorithm corresponding to the data to be written to obtain the second compressed data.
[0059] Specifically, the data to be written stored on the disk is read periodically, and the data type of the data to be written is identified. For example, the data types include types such as pictures, audio, video, text, etc. According to different data types, the background compression algorithm corresponding to the data to be written is determined. For example, picture compression algorithm, audio compression algorithm, video compression algorithm, text compression algorithm, etc. The data to be written is compressed according to the background compression algorithm corresponding to the data to be written to obtain the second compressed data.
[0060] As data is continuously written and updated, the system can periodically check and process new data to ensure that all data undergoes optimized compression processing. And the storage system may contain multiple types of data, and the periodic check and processing mechanism can ensure that each type of data is properly processed. By identifying the data type and selecting the appropriate compression algorithm, the system can apply the most effective compression strategy for different types of data. Since different types of data (such as text, pictures, video, logs, etc.) have different compression characteristics, by selecting the most suitable compression algorithm, the system can achieve a higher compression ratio, thereby reducing the occupation of storage space and saving storage costs. In addition, the background compression processing does not occupy the resources of real-time tasks, so it does not affect the real-time performance of the system, enabling the system to still maintain good performance under high load conditions.
[0061] In some embodiments, in response to a data reading request sent by the host, it is determined whether the data to be read of the data reading request has been compressed; If the data to be read has been compressed, the data to be read is decompressed and the decompressed data to be read is sent to the host; If the data to be read has not been compressed, the data to be read is read from the disk and the data to be read is sent to the host.
[0062] Specifically, in response to a data reading request sent by the host, it is determined whether the data to be read of the data reading request has been compressed. If the data to be read has been compressed, the data to be read is decompressed and the decompressed data to be read is sent to the host, If the data to be read has not been compressed, the data to be read is read from the disk and the data to be read is sent to the host.
[0063] In some embodiments, the above step (determining whether the data to be read of the data reading request has been compressed) can be implemented in the following manner: According to the preset identifier of the data to be read, it is determined whether the data to be read has been compressed.
[0064] Among them, the preset identifier is used to indicate whether the data to be read has been compressed. Further, the preset identifier can also be used to indicate which algorithm the data to be read is compressed by. For example, two bits can be used to represent the preset identifier. "00" indicates that the data to be read has not been compressed, and "01" indicates that the data to be read has been compressed. Another example is that "00" indicates that the data to be read has not been compressed, "01" indicates that the data to be read has been compressed by an online compression algorithm, and "02" indicates that the data to be read has been compressed by a background compression algorithm. Another example is that "00" indicates that the data to be read has not been compressed, "01" indicates that the data to be read has been compressed by compression algorithm 1, "02" indicates that the data to be read has been compressed by compression algorithm 2, "03" indicates that the data to be read has been compressed by compression algorithm 3, and so on.
[0065] Specifically, according to the preset identifier of the data to be read, it is determined whether the data to be read has been compressed.
[0066] Exemplarily, assume that "00" indicates that the data to be read has not been compressed, "01" indicates that the data to be read has been compressed by an online compression algorithm, and "02" indicates that the data to be read has been compressed by a background compression algorithm. When the preset identifier of the data to be read is "00", it is determined that the data to be read has not been compressed.
[0067] Optionally, determining whether the data to be read has been compressed according to the preset identifier of the data to be read includes: If the preset identifier of the data to be read is the first identifier or the third identifier, it is determined that the data to be read has been compressed; If the preset identifier of the data to be read is the second identifier, it is determined that the data to be read has not been compressed.
[0068] If the preset identifier of the data to be read is the second identifier, it is determined that the data to be read has not been compressed.
[0069] Among them, the first identifier is used to identify the data obtained through online compression processing. The second identifier is used to identify the uncompressed data. The third identifier is used to identify the data obtained through background compression processing.
[0070] Specifically, if the preset identifier of the data to be read is the first identifier or the third identifier, it is determined that the data to be read has been compressed; if the preset identifier of the data to be read is the second identifier, it is determined that the data to be read has not been compressed.
[0071] In some embodiments, the above step (decompressing the data to be read) can be implemented in the following manner: Obtain the compression algorithm for the data to be read; Determine the decompression algorithm corresponding to the compression algorithm according to the compression algorithm for the data to be read; Perform decompression processing on the data to be read according to the decompression algorithm.
[0072] Specifically, by obtaining the compression algorithm for the data to be read, determining the decompression algorithm corresponding to the compression algorithm according to the compression algorithm for the data to be read; and performing decompression processing on the data to be read according to the decompression algorithm.
[0073] Optionally, the above step (obtaining the compression algorithm for the data to be read) can be implemented in the following manner: When the preset identifier of the data to be read is the first identifier, obtain the online compression algorithm for the data to be read; When the preset identifier of the data to be read is the third identifier, obtain the background compression algorithm for the data to be read.
[0074] Specifically, when the preset identifier of the data to be read is the first identifier, obtain the online compression algorithm for the data to be read; when the preset identifier of the data to be read is the third identifier, obtain the background compression algorithm for the data to be read. It can be understood that both the online compression algorithm and the background compression algorithm can include various different types of compression algorithms, which can be selected according to the actual application scenario and are not specifically limited here.
[0075] Exemplarily, assume that "00" indicates that the data to be read has not been compressed, "01" indicates that the data to be read has been compressed by compression algorithm 1, "02" indicates that the data to be read has been compressed by compression algorithm 2, "03" indicates that the data to be read has been compressed by compression algorithm 3, etc. If the compression identifier is "01", it means that the data to be read has been compressed by compression algorithm 1, then perform decompression processing on the data to be read based on the decompression algorithm corresponding to compression algorithm 1, and send the decompressed data to be read to the host. If the compression identifier is "00", it means that the data to be read has not been compressed, then directly read the data to be read from the disk and send the data to be read to the host.
[0076] When reading data, the all-flash storage system can quickly determine whether the data needs decompression processing according to the identifier. For uncompressed data (the second identifier), the system can directly read it without performing decompression operations, thereby reducing the latency caused by decompression. For compressed data (the first identifier or the third identifier), the system can select an appropriate decompression algorithm according to the identifier to quickly restore the data and improve the reading efficiency.
[0077] The data management method provided by the embodiments of the present disclosure determines whether the data to be written meets the online compression condition. For the data that meets the compression condition, online compression processing is performed. For the data that does not meet the compression condition, the data is directly written to the disk and compressed in the background, which can avoid the additional delay caused by online compression. By combining online compression and background compression, the storage device can achieve a dynamic balance between real-time performance and storage efficiency. Moreover, by cooperating with online compression and background compression, the coupling relationship between the compression function and the host I / O process is weakened, eliminating the dependence on the hardware compression card.
[0078] Figure 2 FIG. is a schematic structural diagram of a data management device 200 provided by the present disclosure, as Figure 2 shown, including: A condition judgment module 210, configured to respond to a data write request sent by a host, and judge whether the data to be written in the data write request meets the online compression condition; An online compression module 220, configured to perform online compression processing on the data to be written if the data to be written meets the online compression condition, obtain first compressed data, and write the first compressed data to the disk; A background compression module 230, configured to write the data to be written to the disk if the data to be written does not meet the online compression condition, and perform background compression processing on the data to be written when background data compression is triggered, obtain second compressed data, and update the data to be written to the second compressed data.
[0079] As an optional implementation manner of the embodiments of the present disclosure, the device further includes an identifier update module, configured to: Before writing the first compressed data to the disk, update a preset identifier of the data to be written to a first identifier; the first identifier is used to identify the data obtained through online compression processing; Before writing the data to be written to the disk, update the preset identifier to a second identifier; the second identifier is used to identify the uncompressed data; Before updating the data to be written to the second compressed data, update the preset identifier to a third identifier; the third identifier is used to identify the data obtained through background compression processing.
[0080] As an optional implementation manner of the embodiments of the present disclosure, the background compression module 230 is specifically configured to: Obtain the data with the preset identifier being the second identifier in the disk; Perform background compression processing on the data with the preset identifier being the second identifier.
[0081] As an alternative implementation manner of an embodiment of the present disclosure, the background compression module 230 is further specifically configured to: Periodically read the data to be written stored in the disk, and identify the data type of the data to be written; Determine the background compression algorithm corresponding to the data to be written according to the data type; Perform compression processing on the data to be written according to the background compression algorithm corresponding to the data to be written to obtain second compressed data.
[0082] As an alternative implementation manner of an embodiment of the present disclosure, the condition judgment module 210 is specifically configured to: Judge whether the proportion of preset data in the data to be written is greater than a preset proportion, and whether the front-end load is less than a preset load; If the proportion of preset data in the data to be written is greater than the preset proportion and the front-end load is less than the preset load, it is determined that the data to be written meets the online compression condition; If the proportion of preset data in the data to be written is less than or equal to the preset proportion, and / or the front-end load is greater than or equal to the preset load, it is determined that the data to be written does not meet the online compression condition.
[0083] As an alternative implementation manner of an embodiment of the present disclosure, the device further includes: A judgment module, configured to determine whether the data to be read in the data reading request has been compressed in response to a data reading request sent by the host; A decompression module, configured to, if the data to be read has been compressed, perform decompression processing on the data to be read and send the decompressed data to be read to the host; A reading module, configured to, if the data to be read has not been compressed, read the data to be read from the disk and send the data to be read to the host.
[0084] As an alternative implementation manner of an embodiment of the present disclosure, the decompression module is specifically configured to: Obtain the compression algorithm of the data to be read; Determine the decompression algorithm corresponding to the compression algorithm according to the compression algorithm of the data to be read; Perform decompression processing on the data to be read according to the decompression algorithm.
[0085] For the description of the features in the corresponding embodiment of the data management device 200, reference may be made to the relevant description in the corresponding embodiment of the data management method, which will not be elaborated here one by one.
[0086] The data management device provided by the embodiments of the present disclosure determines whether the data to be written meets the online compression condition. For the data that meets the compression condition, online compression processing is performed. For the data that does not meet the compression condition, the data is directly written to the disk and compressed in the background, which can avoid the additional delay caused by online compression. Through the combination of online compression and background compression, the storage device can achieve a dynamic balance between real-time performance and storage efficiency. Moreover, by means of the cooperation of online compression and background compression, the coupling relationship between the compression function and the host IO process is weakened, eliminating the dependence on the hardware compression card.
[0087] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the data management method.
[0088] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the data management method when running.
[0089] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, etc., various media that can store computer programs.
[0090] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the data management method are implemented.
[0091] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the data management method are implemented.
[0092] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0093] The above has introduced in detail a data management method provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A data management method, characterized in that, The method includes: In response to a data writing request sent by a host, determining whether the data to be written in the data writing request meets the online compression condition; If the data to be written meets the online compression condition, performing online compression processing on the data to be written, obtaining first compressed data, and writing the first compressed data to a disk; If the data to be written does not meet the online compression condition, writing the data to be written to the disk, and when background data compression is triggered, performing background compression processing on the data to be written, obtaining second compressed data, and updating the data to be written to the second compressed data.
2. The data management method according to claim 1, wherein The method further includes: Before writing the first compressed data to the disk, updating a preset identifier of the data to be written to a first identifier; the first identifier is used to identify data obtained through online compression processing; Before writing the data to be written to the disk, updating the preset identifier to a second identifier; the second identifier is used to identify uncompressed data; Before updating the data to be written to the second compressed data, updating the preset identifier to a third identifier; the third identifier is used to identify data obtained through background compression processing.
3. The data management method according to claim 2, wherein When background data compression is triggered, performing background compression processing on the data to be written includes: Obtaining data in the disk whose preset identifier is the second identifier; Performing background compression processing on the data whose preset identifier is the second identifier.
4. The data management method according to claim 1, characterized in that Performing background compression processing on the data to be written to obtain second compressed data includes: Periodically reading the data to be written stored in the disk and identifying the data type of the data to be written; Determining a background compression algorithm corresponding to the data to be written according to the data type; Performing compression processing on the data to be written according to the background compression algorithm corresponding to the data to be written to obtain second compressed data.
5. The data management method according to claim 1, characterized in that Determining whether the data to be written in the data writing request meets the online compression condition includes: Determining whether the proportion of preset data in the data to be written is greater than a preset proportion, and whether the front-end load is less than a preset load; If the proportion of preset data in the data to be written is greater than the preset proportion, and the front-end load is less than the preset load, determining that the data to be written meets the online compression condition; If the proportion of preset data in the data to be written is less than or equal to the preset proportion, and / or the front-end load is greater than or equal to the preset load, determining that the data to be written does not meet the online compression condition.
6. The data management method according to claim 1, wherein The method further includes: In response to a data reading request sent by a host, determining whether the data to be read in the data reading request has been compressed; If the data to be read has been compressed, performing decompression processing on the data to be read and sending the decompressed data to be read to the host; If the data to be read has not been compressed, reading the data to be read from the disk and sending the data to be read to the host.
7. The data management method according to claim 6, wherein Performing decompression processing on the data to be read includes: Obtaining the compression algorithm of the data to be read; Determine the decompression algorithm corresponding to the compression algorithm according to the compression algorithm of the data to be read; Perform decompression processing on the data to be read according to the decompression algorithm.
8. A data management device, characterized in that, The device includes: A condition judgment module, configured to respond to a data writing request sent by a host, and judge whether the data to be written in the data writing request meets the online compression condition; An online compression module, configured to perform online compression processing on the data to be written if the data to be written meets the online compression condition, obtain first compressed data, and write the first compressed data into a disk; A background compression module, configured to write the data to be written into a disk if the data to be written does not meet the online compression condition, and perform background compression processing on the data to be written when background data compression is triggered, obtain second compressed data, and update the data to be written to the second compressed data.
9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to implement the steps of the data management method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data management method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Data writing method, storing device of memory and control circuit unit of memory
CN104881240A
Database optimization method and apparatus
CN106528896A
Data compression method and device, computer equipment and storage medium
CN111459404A
System and method for an improved real-time adaptive data compression
US20180300087A1