Data compression method, electronic device and computer program product

By determining the number of data to be compressed in real time and adjusting the compression level dynamically, the problem of low data compression efficiency in traditional storage systems is solved, and more efficient data compression and resource utilization are achieved.

CN114816222BActive Publication Date: 2025-06-06EMC IP HLDG CO LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110088559.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-22
Publication Date
2025-06-06
Estimated Expiration
2041-01-22

AI Technical Summary

Technical Problem

Data compression efficiency in traditional storage systems is low, resulting in waste of storage space and computing resources.

Method used

By determining the number of data to be compressed in real time, dynamically adjusting the compression level to achieve adaptive compression of the data.

Benefits of technology

The overall compression rate of data is improved and the efficiency of the storage space and computing resources utilization of the storage system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816222B_ABST
    Figure CN114816222B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a data compression method, an electronic device, and a computer program product, the method comprising: determining the amount of data to be compressed in a storage system; determining a target compression level for compressing the data to be compressed based on the amount of data to be compressed; and compressing the data to be compressed according to the target compression level. In this way, the data to be compressed can be compressed using a compression level corresponding to the amount of data to be compressed, thereby improving the efficiency of data compression in the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to a data compression method, an electronic device, and a computer program product. Background Art

[0002] In the era of big data, the contradiction between the existence of massive data and the limited storage space of the storage system has raised the demand for data compression. It is understandable that the lower the compression rate of the data, the smaller the storage data occupied by the compressed data, and the better the compression effect. A better compression rate means a larger logical capacity, more data that can be cached in the SSD cache, and better restoration performance. However, as the compression rate of the data decreases, the calculation required by the storage system will increase accordingly, and the amount of data that can be compressed in the same time will also decrease accordingly. That is, in traditional storage systems, the efficiency of data compression is low. Summary of the invention

[0003] Embodiments of the present disclosure provide a scheme for data compression.

[0004] In a first aspect of the present disclosure, a data compression method is provided, the method comprising determining an amount of data to be compressed in a storage system. The method further comprises determining a target compression level for compressing the data to be compressed based on the amount of data to be compressed. The method further comprises compressing the data to be compressed according to the target compression level.

[0005] In a second aspect of the present disclosure, an electronic device is provided, comprising a processor; and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform actions, the actions comprising: determining the amount of data to be compressed in a storage system; determining a target compression level for compressing the data to be compressed based on the amount of data to be compressed; and compressing the data to be compressed according to the target compression level.

[0006] In a third aspect of the present disclosure, there is provided a computer program product tangibly stored on a computer readable medium and comprising machine executable instructions which, when executed, cause a machine to perform any of the steps of the method according to the first aspect.

[0007] This Summary is provided to introduce a selection of concepts in a simplified form that are further described in the Detailed Description below. This Summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other purposes, features and advantages of the present disclosure will become more apparent by describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present disclosure. In the accompanying drawings:

[0009] Figure 1 A schematic diagram showing an exemplary environment according to an embodiment of the present disclosure is shown;

[0010] Figure 2 a schematic diagram showing various graphs of various metrics related to compression levels;

[0011] Figure 3 A flowchart showing a process of data compression according to an embodiment of the present disclosure is shown;

[0012] Figure 4 A flowchart showing a process of determining a target compression level according to an embodiment of the present disclosure

[0013] Figure 5 A schematic diagram of a process for data compression in a data storage task according to an embodiment of the present disclosure;

[0014] Figure 6 A schematic diagram showing a process for data compression in a data recycling task according to an embodiment of the present disclosure; and

[0015] Figure 7 A block diagram of an example device that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0016] The principles of the present disclosure will be described below with reference to several example embodiments shown in the accompanying drawings.

[0017] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "a set of example embodiments". The term "another embodiment" means "a set of additional embodiments". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] As discussed above, in a storage system, the compression ratio refers to the ratio of the size of the compressed data to the size of the original data. The smaller the compression ratio of the data, the less space the compressed data occupies, but the more computing resources consumed by the compression. Therefore, in a unit time, for example, a storage system with limited computing resources due to hardware conditions can compress less data, and vice versa. In order to balance the amount of data that the storage system can compress in a unit time (sometimes also referred to as data throughput, or throughput, in this article) and the compression ratio, in a storage system, a fixed compression level is usually selected to compress the data to meet the requirements of the above two. (At this compression level, the compressed data obtained by compressing data with the same data segment size using the same hardware device will have a roughly fixed compression ratio. However, since the data throughput can change in real time over time, for a storage system with a possibly large peak value of the data throughput, the compression level cannot usually be selected too high, which will result in insufficient compression of the data when the real-time data throughput is small, resulting in a waste of storage space and / or computing resources.

[0019] In order to at least partially solve the above shortcomings, an embodiment of the present disclosure provides a data compression scheme for adaptively determining a compression level for data compression according to the amount of data to be compressed. The scheme can determine the amount of data to be compressed in real time and select a compression level that matches the amount of data to be compressed for compression.

[0020] Based on such a data compression scheme, since the compression level adopted by the storage system is adapted to the amount of data to be compressed, the overall compression rate of the data can be improved, thereby improving the utilization efficiency of the storage space and / or computing resources of the storage system.

[0021] Figure 1 FIG. 1 is a schematic diagram of an exemplary environment 100 according to an embodiment of the present disclosure, in which the device and / or method according to an embodiment of the present disclosure may be implemented. Figure 1 As shown, Figure 1 As shown, the exemplary environment may include a storage system 150. The storage system 150 may include a computing device 105 for processing various operations for data storage, including but not limited to data compression and decompression, data deduplication, data storage, data backup and recovery.

[0022] The storage system 150 may include (multiple) storage disks for storing data, not shown. The storage disk may be various types of devices with storage functions, including but not limited to hard disks (HDDs), solid state disks (SSDs), removable disks, any other magnetic storage devices, and any other optical storage devices, or any combination thereof.

[0023] The computing device 105 may be configured to compress the data to be compressed 110 to obtain compressed data 130. The compressed data 130 may be stored in a storage disk to save storage space of the storage disk.

[0024] In some embodiments, the storage system 150 may be a storage system for data backup, which is configured with a deduplication device (sometimes also referred to as deduplication or data deduplication in this article) to remove duplicate parts in the data and only store non-duplicate parts, thereby achieving efficient use of storage space. The storage system 150 may further compress the deduplicated data. In some embodiments, the storage system 150 may compress the data using a corresponding coprocessor using various compression technologies such as QuickAssist accelerated compression technology (QAT). In some embodiments, the storage system 150 may be configured to compress the data using various compression levels provided by various compression technologies.

[0025] It is understood that in some cases, the amount of data to be compressed 110 may vary over time. For example, in the case of storing data that is not the first backup, the amount of deduplicated data (in other words, the amount of data to be compressed) is much less than the amount of original data. Therefore, if the computing resources that the storage system can provide are the same, a better compression level can be used to compress the data in such a case to achieve various benefits of using a higher compression ratio, such as logical capacity, more data that can be cached in the SSD cache, and better restore performance.

[0026] Therefore, the storage system 150 (for example, the computing device 150 of the storage system) can compress the data to be compressed 110 according to the target compression level 130 that matches the amount of data to be compressed 110, so that the compressed data 130 is as small as possible while ensuring that the latency requirements of the data processing operation are met.

[0027] The following will combine Figures 2 to 6 The process according to the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0028] Figure 2 A schematic diagram 200 showing multiple tables of various indicators related to compression level. It should be noted that Figure 2Only the charts of various indicators corresponding to various compression levels provided by the QAT compression technology under the same hardware configuration are shown. It is understandable that similar charts can be obtained by those skilled in the art through testing of the storage system when different hardware configurations and / or other compression technologies are adopted.

[0029] Table 210 shows the amount of data that the storage system can compress at various compression levels per unit time (seconds, s) according to an embodiment of the present disclosure, i.e., throughput (GB / s). Compression and / or decompression can be divided into two types: dynamic type and static type, which can refer to dynamic Huffman data compression and / or decompression, and static Huffman data compression and / or decompression, respectively.

[0030] For example, the QAT compression technology can provide dynamic compression level 1 to dynamic compression level 4 (sometimes referred to as dynamic level in this article), and static compression level 1 to static compression level 4 (sometimes referred to as static level in this article). At different compression levels, the throughput is different. In addition, the data segment size, such as 1K, 4K, 8K, 16K, 64K (unit: bit B), may also affect the throughput.

[0031] Table 220 shows the ratio of data compressed by the storage system to the data size, i.e., the compression ratio, at various compression levels according to an embodiment of the present disclosure. It can be seen from Table 220 that for the same data, a better compression level can provide a lower compression ratio, in other words, the compressed data requires less storage space.

[0032] It can be seen from Table 210 and Table 220 that the throughput corresponding to the better compression level is lower, in other words, the amount of data that the storage system can compress per unit time is smaller. If the better compression level is still used when the amount of data to be compressed is large, it may cause an overall delay in data compression of the storage system. Therefore, the compression delay may vary significantly depending on the compression level.

[0033] Table 230 shows the amount of data compressed according to various compression levels that the storage system can decompress per unit time (seconds, s) according to an embodiment of the present disclosure, i.e., throughput (GB / s). It can be seen from Table 230 that at different compression levels, the amount of data that the storage system can decompress is not much different, so the decompression delay will not change significantly due to the compression level. In the case where the decompression delay is basically the same, more data can be obtained by decompressing data compressed at a higher compression level per unit time.

[0034] It is understandable that the exact values ​​of various indicators similar to those shown in the above charts may vary when adopting different other hardware configurations and / or other compression technologies, but the relationships between them are similar to those described above with reference to the above charts.

[0035] Figure 3 FIG. 3 is a flowchart of a data compression process 300 according to an embodiment of the present disclosure. The process 300 may be performed in Figure 1 The system is implemented at the computing device 105 shown in FIG.

[0036] At 302 , the computing device 105 may determine an amount of data 110 to be compressed in the storage system 150 .

[0037] Specifically, the data to be compressed is data that is expected to be compressed using various compression techniques or algorithms. In some embodiments, the data to be compressed may be data to be stored (e.g., to be backed up) in a storage system. In some embodiments, the data to be compressed may also be data obtained after the data to be stored is deduplicated. After the compression process, the data may be stored in a storage disk for subsequent retrieval.

[0038] In other embodiments, the stored data already stored in the storage disk may also be compressed using various compression techniques or algorithms when performing data recycling processes such as garbage collection. The data to be compressed may also be the data to be recycled.

[0039] The amount of data can be obtained in various ways. In some embodiments, this can be achieved by using various monitors of the memory to monitor the parameters of the amount of data in real time, and additionally or alternatively, using such parameters to calculate the amount of data to be compressed. For example, for the data to be stored, this can be achieved by detecting the flow rate or network bandwidth of the data input into the storage system.

[0040] At 304 , the computing device 105 may determine a target compression level 120 for compressing the data to be compressed 110 based on the amount of the data to be compressed 110 .

[0041] It is understandable that the better the compression level, the higher the compression rate, and thus the less storage space required for the compressed data. However, the storage system usually needs to meet certain delay requirements when processing data. In some cases, for a large amount of data to be compressed per unit time, using, for example, the best compression level is likely to cause the storage system to process the data too long, and thus fail to meet the predetermined delay requirements. In other cases, for a small amount of data to be compressed per unit time, using, for example, the worst compression level can meet the predetermined delay requirements, but is likely to cause unnecessary occupation of storage space.

[0042] Therefore, the computing device can select the optimal compression level that is suitable for the amount of data to be compressed according to the amount of data to be compressed, so that the predetermined delay requirement can be met and the compression rate of the compressed data can be maximized.

[0043] At 306, the computing device 105 may compress the data to be compressed according to the target compression level 120. The compressed data 130 may be further stored in a storage disk.

[0044] In this way, the computing device can determine and compress the data to be compressed using a compression level corresponding to the amount of data to be compressed, thereby improving the compression efficiency of data in the storage system.

[0045] It is understandable that the computing device 105 can dynamically adjust the target compression level as the amount of data to be compressed changes, so that the compression level used can compress the data to be compressed in a timely manner without causing excessive delay. In some embodiments, the computing device 105 can respond to changes in the computing resources available for compressing the data to be compressed and / or the amount of data to be compressed in the storage system, for example, by performing the above steps 302-204 again to update the target compression level.

[0046] In the case where, for example, the amount of data to be compressed per unit time decreases and / or the computing resources that the storage system can use to compress data increases, the computing device can determine whether to use a higher compression level to compress the data, so as to obtain a higher compression rate for the data and save storage space; and in the case where, for example, the amount of data to be compressed per unit time increases and / or the computing resources that the storage system can use to compress data decreases, the computing device can determine whether to use a lower compression level to compress the data, so as to increase the amount of data that the storage system can process with limited computing resources by reducing the compression rate of the data. In this way, adaptive adjustment of the compression rate can be achieved, the utilization rate of the computing resources of the storage system is maximized, the efficiency of the storage system in compressing data is improved, and storage space is saved.

[0047] Figure 4FIG. 4 is a flowchart showing a process 400 for determining a target compression level according to an embodiment of the present disclosure. The process 400 may be performed at Figure 1 The system is implemented at the computing device 105 shown in FIG.

[0048] At 402 , the computing device 105 may determine computing resources in the storage system 150 that are available for compressing the data to be compressed 110 .

[0049] Specifically, if the hardware configuration of the storage system remains unchanged, the total computing resources that it can provide will be fixed, and the total computing resources are, for example, associated with the hardware configuration and / or operating parameters of the coprocessor dedicated to the compression operation. However, there may be a situation where not all computing resources are used to compress data. Therefore, the amount of data that the storage system can compress according to each compression level may also change over time.

[0050] In some embodiments, the computing device may detect the ratio of the available portion of computing resources of the storage system to the total computing resources of the storage system (sometimes referred to herein as utilization). The computing device may determine the computing resources available for compressing the data to be compressed based on the total computing resources and the ratio. The ratio may be obtained, for example, by a utilization monitor.

[0051] At 404 , the computing device 105 may determine, based on the computing resources, a first amount of data that the storage system 150 is capable of compressing according to each of a plurality of candidate compression levels.

[0052] It is understandable that since the computing resources consumed for compressing the same amount of data according to various compression levels are different, the throughput corresponding to each candidate compression level can be determined based on the allocated computing resources for determining the target compression level.

[0053] In some embodiments, the computing device may obtain a compression level mapping table associated with the total computing resources of the storage system. The compression level mapping table may include a correspondence between multiple candidate compression levels and multiple second quantities, each second quantity being the amount of data that the storage system can compress according to the corresponding candidate compression level (i.e., throughput). For example, Table 1 below shows an example of a compression level mapping table for 128KB data segment compression.

[0054] Table 1

[0055]

[0056]

[0057] The data in the compression level mapping table can be obtained based on the test of the storage system. In the example of Table 1 above, the data is obtained from Figure 2 It is understood that the candidate compression levels included in the compression level mapping table can be selected according to the compression effects provided by the various compression levels of various compression techniques, so that the compression ratio can be optimized for each range of throughput and additionally or alternatively for each data segment size.

[0058] The computing device may determine a first quantity corresponding to each candidate compression level based on the compression level mapping table and the determined computing resources.

[0059] For example, if the determined computing resources account for 50% of the total computing resources, the amount of data that can be compressed according to static level 1, dynamic level 1, dynamic level 2, dynamic level 3, and dynamic level 4 can be determined to be 52.5, 43, 24.5, 18, and 12.5, respectively, based on a compression level mapping table such as Table 1 above.

[0060] At 406 , computing device 105 may select a target compression level 120 from the plurality of candidate compression levels such that the amount of to-be-compressed data 110 matches a first amount corresponding to target compression level 120 .

[0061] The selected target compression level may be a compression level that can meet the storage system latency requirements while providing an optimal compression ratio. More specifically, at the selected target compression level, the amount of data to be compressed that the storage system can compress is just greater than or equal to the amount of data to be compressed.

[0062] For example, if it is determined that all computing resources are available for compression, and the amount of data to be compressed is 40 GB / s, then dynamic level 2 matching the amount may be selected to compress the data to be compressed.

[0063] If it is determined that all computing resources are available for compression, and the amount of data to be compressed changes to 20 GB / s, for example, due to a change in the deduplication rate, dynamic level 4 matching the amount may be selected for compressing the data to be compressed.

[0064] If it is determined that the ratio of available computing resources to total computing resources changes to 50%, and the amount of data to be compressed is 20 GB / s, dynamic level 2 matching the amount may be selected for compressing the data to be compressed.

[0065] In some embodiments, the computing device can select a compression level mapping table including multiple candidate compression levels that matches the size of the data segment to be compressed based on the size of the data segment to be compressed, and select a target compression level that matches the amount of data to be compressed and the size of the data segment therein.

[0066] In this way, the computing device can select the optimal compression level to compress the data to be compressed, so that the storage space occupied by the compressed data is saved.

[0067] Figure 5 A schematic diagram of a process 500 for data compression in a data storage task according to an embodiment of the present disclosure. The process 500 may be performed at Figure 1 The system is implemented at the computing device 105 shown in FIG.

[0068] The computing device 105 may determine the amount of data to be stored to the storage system. For example, the computing device 105 may be configured with a data volume monitor 512 for counting the amount of data to be stored 502 to determine the amount of data to be stored 514.

[0069] The computing device 105 may determine a deduplication rate for the data to be stored. For the data to be stored 502, the computing device may be configured with a deduplication device 504, for example, to delete portions of the data to be stored 502 that are repeated with the stored data. The deduplicated data to be stored may be used as the data to be compressed 510.

[0070] The computing device may be configured with a deduplication rate monitor 506 to monitor the operation of the deduplication unit 504 , for example in real time, to determine the deduplication rate 506 for the data to be stored 502 .

[0071] The computing device 105 may determine the amount of data to be compressed 510 based on the amount of data to be stored and the deduplication rate. For example, if the amount of data to be stored is 100 GB / s and the deduplication rate is 60%, the amount of data to be compressed may be determined to be equal to 100×(1-60%)=40 GB / s. The above determination process of the amount of data to be compressed may be configured to be executed at the controller 525 of the computing device 105, for example.

[0072] In this way, when selecting a compression level for data compression, the influence of the deduplication rate can be taken into account to more accurately select the target compression level and improve the efficiency of data compression in the storage system.

[0073] The computing device may further include a utilization monitor 522 configured to monitor in real time the ratio of computing resources available for compression to total computing resources, and determine it as a utilization 524. In some embodiments, the utilization monitor 522 may determine the utilization 524 for each of the different tasks being performed simultaneously. For example, the utilization monitor 522 may determine the utilization of a compression task related to data to be stored, the utilization of a compression task related to data to be recycled (described in more detail below), and the utilization of other tasks (e.g., a decompression task).

[0074] The controller 525 may also be configured to determine the amount of data to be compressed and select an appropriate target compression level 520 based on a predetermined compression level mapping table 526, the amount of data to be compressed, the size information of the data segment to be compressed, the utilization 524 and / or the total computing resources of the storage system. The selection process of the target compression level 520 is described above with reference to Figure 3 and Figure 4 It has been described and will not be repeated here.

[0075] Using the selected target compression level 520, the compressor 515 of the computing device 105 may execute a compression algorithm on the data to be compressed 510 to obtain compressed data 530. In some embodiments, the compressor 515 may be implemented by a coprocessor using QAT technology.

[0076] Figure 6 FIG. 6 is a schematic diagram showing a process 600 for data compression in a data recovery task according to an embodiment of the present disclosure. The process 600 may be performed in Figure 1 The system is implemented at the computing device 105 shown in FIG.

[0077] Data recycling is sometimes also referred to as garbage collection (GC) processing. It is understandable that in a storage system, data recycling is not always performed, and is usually performed when the storage system is idle (ie, the computing resources of the storage system are not all utilized).

[0078] The data recovery process involves decompression and recompression to reorganize the data on the storage disk. It is desirable that in the data recovery process, the data originally compressed at a poorer compression level is recompressed with a better compression level so that the storage space on the storage disk is further saved. It is understandable that during the data recovery process, the throughput requirements of the system also need to be met.

[0079] The computing device 105 may determine the amount of data to be recycled 602 in the storage system. This may be achieved, for example, by a GC monitor 612 in the computing device. The GC monitor is configured to count the amount of data to be recycled 602 in real time to determine the amount of data to be recycled 614.

[0080] The computing device 105 may determine the amount of data to be compressed based on the amount of data to be recycled 614. This may be configured to be performed at the controller 625 of the computing device 105, for example.

[0081] The computing device 105 may include a utilization monitor 622 for determining utilization 624 (eg, utilization of a compression task associated with data to be recycled). The utilization monitor 622 may be associated with a reference Figure 5 The utilization monitor 522 described above is similar and will not be described again here.

[0082] The controller 625 may also be configured to determine the amount of data to be compressed and select an appropriate target compression level 620 based on the predetermined compression level mapping table 626, the amount of data to be compressed, the size information of the data segment to be compressed, the utilization rate 624 and / or the total computing resources of the storage system. The selection process of the target compression level 620 is described in detail above. Figure 3 and Figure 4 It has been described and will not be repeated here.

[0083] After decompression, the data to be recycled 602 can be used as the data to be compressed 610 for recompression to have a better compression rate.

[0084] Using the selected target compression level 620, the compressor 615 of the computing device 105 may execute a compression algorithm on the data to be compressed 610 to obtain compressed data 630. In some embodiments, the compressor 615 may be implemented by a coprocessor using QAT technology.

[0085] In this way, the optimal compression level can be selected for the data to be recycled, so that the compression rate of the recycled data is higher, storage space is saved, and the storage efficiency of the storage system is improved.

[0086] Figure 7 Schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. For example, the electronic device 700 can be used to implement Figure 1. As shown in the figure, the device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 to a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0087] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0088] The processing unit 701 performs the various methods and processes described above, such as any one of the processes 300 to 600. For example, in some embodiments, any one of the processes 300 to 600 may be implemented as a computer software program or a computer program product, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU 701, one or more steps in any of the processes 300 to 600 described above may be performed. Alternatively, in other embodiments, the CPU 701 may be configured to perform any one of the processes 300 to 600 by any other appropriate means (e.g., by means of firmware).

[0089] The present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0090] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, any non-temporary storage device, or any suitable combination of the above. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the above. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0091] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0092] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0093] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0094] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0095] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0096] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0097] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A data compression method, include: Determining the amount of data to be compressed in the storage system; Determining a target compression level for compressing the data to be compressed based on the amount of the data to be compressed; as well as Compressing the data to be compressed according to the target compression level, wherein determining the target compression level comprises: Determining a first amount of data that the storage system is capable of compressing according to each of a plurality of candidate compression levels, wherein determining the first amount of data that the storage system is capable of compressing according to each of the candidate compression levels comprises: Obtaining a compression level mapping table associated with total computing resources of the storage system, wherein the compression level mapping table includes a correspondence between the plurality of candidate compression levels and a plurality of second quantities, each second quantity being a quantity of data that can be compressed by the storage system according to a corresponding candidate compression level; and The first number corresponding to each candidate compression level is determined based on the compression level mapping table and the determined computing resources in the storage system that can be used to compress the data to be compressed.

2. The method according to claim 1, wherein determining the amount of the data to be compressed include: Determining the amount of data to be stored in the storage system; Determining a deduplication rate for the data to be stored; as well as Based on the amount of the data to be stored and the deduplication rate, the amount of the data to be compressed is determined.

3. The method according to claim 1, wherein determining the amount of the data to be compressed include: Determining the amount of data to be recycled in the storage system; as well as The amount of the data to be compressed is determined based on the amount of the data to be recycled.

4. The method of claim 1, wherein determining the target compression level include: Determining the computing resources in the storage system that can be used to compress the data to be compressed; as well as The target compression level is selected from the plurality of candidate compression levels so that the amount of the to-be-compressed data matches the first amount corresponding to the target compression level.

5. The method of claim 4, wherein determining the computing resource include: Detecting a ratio of the available partial computing resources of the storage system to the total computing resources of the storage system; as well as The computing resource is determined based on the total computing resource and the ratio.

6. The method of claim 1, wherein determining the target compression level include: The target compression level is updated in response to a change in at least one of the following: The computing resources in the storage system that can be used to compress the data to be compressed; as well as The amount of data to be compressed.

7. An electronic device, include: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs actions, the actions comprising: Determining the amount of data to be compressed in the storage system; Determining a target compression level for compressing the data to be compressed based on the amount of the data to be compressed; and Compressing the data to be compressed according to the target compression level, wherein determining the target compression level comprises: Determining a first amount of data that the storage system is capable of compressing according to each of a plurality of candidate compression levels, wherein determining the first amount of data that the storage system is capable of compressing according to each of the candidate compression levels comprises: Obtaining a compression level mapping table associated with total computing resources of the storage system, wherein the compression level mapping table includes a correspondence between the plurality of candidate compression levels and a plurality of second quantities, each second quantity being a quantity of data that can be compressed by the storage system according to a corresponding candidate compression level; and The first number corresponding to each candidate compression level is determined based on the compression level mapping table and the determined computing resources in the storage system that can be used to compress the data to be compressed.

8. The electronic device according to claim 7, wherein determining the amount of the data to be compressed include: Determining the amount of data to be stored in the storage system; Determining a deduplication rate for the data to be stored; as well as Based on the amount of the data to be stored and the deduplication rate, the amount of the data to be compressed is determined.

9. The electronic device according to claim 7, wherein determining the amount of the data to be compressed include: Determining the amount of data to be recycled in the storage system; as well as The amount of the data to be compressed is determined based on the amount of the data to be recycled.

10. The electronic device of claim 7, wherein determining the target compression level include: Determining computing resources in the storage system that can be used to compress the data to be compressed; as well as The target compression level is selected from the plurality of candidate compression levels so that the amount of the to-be-compressed data matches the first amount corresponding to the target compression level.

11. The electronic device according to claim 10, wherein determining the computing resource include: Detecting a ratio of the available partial computing resources of the storage system to the total computing resources of the storage system; as well as The computing resource is determined based on the total computing resource and the ratio.

12. The electronic device of claim 7, wherein determining the target compression level include: The target compression level is updated in response to a change in at least one of the following: The computing resources in the storage system that can be used to compress the data to be compressed; as well as The amount of data to be compressed.

13. A computer program product tangibly stored on a computer readable medium and comprising machine executable instructions which, when executed, cause a machine to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Dynamic data compression method for solid-state disc storage system

    CN105094709A

  • A data storage method and device

    CN109445719A