Block storage device and method for data compression

By detecting whether the data block can be deduplicated during the first operation stage of the block storage device and using large-block compression when it cannot be carried out, the problem of low efficiency of traditional storage devices when compressing large-sized data blocks is solved, and a more efficient data reduction effect is achieved.

CN114258558BActive Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080015357.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-23
Publication Date
2025-06-10
Estimated Expiration
2040-07-23

AI Technical Summary

Technical Problem

Traditional storage devices are inefficient when compressing large-sized data blocks, and do not support multiple data reduction methods and granular configurations, which limits the compression capabilities of data blocks.

Method used

A block storage device and method are provided, by detecting whether a data block can be deduplicated in the first operation stage, and if it cannot be performed, a large block compressed storage data block is employed. The device also supports deduplication and compression at the sub-block level of data blocks and similarity deduplication based on the similarity hash table.

Benefits of technology

It improves the compression efficiency of data blocks, supports a variety of data reduction methods and granular configurations, and can select the most effective data reduction technology in different storage areas, thereby improving the effect of data reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114258558B_ABST
    Figure CN114258558B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of storage devices, and more particularly to block storage devices. More specifically, a device and a method are provided that are capable of applying large-block compression to data blocks if deduplication cannot be performed. Accordingly, the present invention provides a block storage device (100) for data compression. The block storage device (100) is configured to: in a first operation phase, if the block storage device (100) determines that a data block (101) is to be written to a large-block storage area (102) of the block storage device (100), determine whether the data block (101) can be deduplicated. If the data block (101) cannot be deduplicated, the block storage device (100) stores the data block (101) using large-block compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage devices, and more particularly to block storage devices. More specifically, an apparatus and a method are provided that are capable of applying large-block compression to data blocks if deduplication cannot be performed. Background Art

[0002] Conventional storage devices that implement two-stage data reduction are designed to minimize CPU occupancy during an initial online stage and maximize data reduction during subsequent background processing.

[0003] During the online stage, a conventional storage device generates a hash fingerprint of an input data block and compares it with existing fingerprints. If a match is found, the storage device performs deduplication. That is, instead of storing the data block, the storage device stores a pointer to an existing identical data block. If deduplication cannot be performed, the system compresses and stores the data. During the background process, the conventional storage device attempts to further increase the size of the stored data blocks.

[0004] Conventional storage devices can use several data reduction methods during the online stage: as previously mentioned, the conventional method is fixed-size deduplication. In this method, the input data block is divided into aligned blocks of a fixed size (e.g., 4KB, 8KB, 16KB, etc.). For each block, a strong hash fingerprint is generated. If the block to be written has the same signature as a written block, they are considered identical. Thus, the data is not stored again, but a pointer to the same block is retained.

[0005] Another conventional method is similarity compression. In this method, a similarity hash function (e.g., a minhash function) is generated for each data block and stored in an opportunity table. If data blocks have the same similarity hash, the conventional storage device considers them to be similar. This means that a portion of the data in one data block is the same as a portion of the data in another block. Thus, a portion of one data block is the same as a portion of another data block, and overall, the data blocks are similar but not identical. This allows only the non-identical portions of the data blocks to be stored and differential compression to be performed using a pointer to the portion of the similar data block that is the same as the portion of the current data block, or using a similar block as a reference.

[0006] However, the disadvantage of conventional storage devices is that they only support compression of data blocks within a very small size range, which results in low compression efficiency. In addition, conventional storage devices do not support multiple data reduction methods, specifically, they do not support a granularity configuration method. Moreover, the two-stage method is limited to static configuration. Summary of the Invention

[0007] In view of the above problems, an object of an embodiment of the present invention is to improve a conventional storage device.

[0008] This object or other objects can be achieved by the embodiments of the present invention described in the appended independent claims. Advantageous implementations of the embodiments of the present invention are further defined in the dependent claims.

[0009] A first aspect of the present invention provides a block storage device for data compression, configured to: in a first operation phase, if the block storage device determines that a data block is to be written to a large block storage area of the block storage device, determine whether the data block can be deduplicated; if the data block cannot be deduplicated, store the data block using large block compression.

[0010] This ensures that if the area where the data block should be written is a large block storage area (i.e., usually a large block read and / or write area), the block storage device can apply large block compression to the data block. Due to the improved effectiveness of large block compression, this improves the effectiveness of data reduction.

[0011] Specifically, large block compression means that the entire data block is compressed at once. Specifically, large block compression means that the entire data block is not divided into sub-blocks before applying compression. Specifically, large block compression means that compression is applied to data blocks larger than a predefined threshold. Specifically, the predefined threshold can be one of the following file sizes: 128 kilobytes, 256 kilobytes, 512 kilobytes, 1 megabyte, 2 megabytes, 4 megabytes, 8 megabytes, and so on.

[0012] Specifically, determining that a data block is to be written to a large block storage area, determining whether the data block can be deduplicated, and large block compression of the data block are performed during the first operation phase.

[0013] Specifically, the large block storage area includes at least a part of a virtual disk or a physical disk; and / or at least a part of a virtual partition or a physical partition.

[0014] Specifically, deduplication includes: if a data block includes at least two identical sub-blocks, only store one of the identical sub-blocks, and for each remaining sub-block, retain a pointer to the stored sub-block. That is, instead of requiring N identical sub-blocks, only one sub-block and (N–1) pointers to that sub-block are required.

[0015] In one implementation of the first aspect, the large block storage area is an area where data blocks with an average size greater than a predefined threshold are read and / or written.

[0016] This ensures that the block storage device can detect a storage area where large block compression can be effectively applied.

[0017] Specifically, the predefined threshold can be one of the following file sizes: 128 kilobytes, 256 kilobytes, 512 kilobytes, 1 megabyte, 2 megabytes, 4 megabytes, 8 megabytes, and so on.

[0018] In another implementation of the first aspect, the block storage device is further configured to determine the average size according to read statistics and / or write statistics.

[0019] This provides a suitable metric for determining the average input / output (IO) size.

[0020] In yet another implementation of the first aspect, the first operation phase is the online phase.

[0021] This ensures that the block storage device can be used in the online phase.

[0022] Specifically, the online phase is executed in association with the write operation of the data block.

[0023] In yet another implementation of the first aspect, the block storage device is further configured to: if the data block can be deduplicated, perform deduplication and compression on the data block, and store the resulting data block.

[0024] This ensures that multiple data reduction techniques can be selected from the data block and / or multiple data reduction techniques can be applied to the data block to maximize the data reduction effect. In addition, the best reduction technique can be selected.

[0025] Specifically, in this case, the data block is divided into small sub-blocks (e.g., small sub-blocks of 8KB), deduplicated at a small granularity of 8KB, and compressed at a small granularity.

[0026] Specifically, the compression applied in this step is different from large-block compression.

[0027] In yet another implementation of the first aspect, the block storage device is further configured to: if the data block can be deduplicated, compare the size obtained by deduplicating and compressing the data block with the size obtained by performing large-block compression on the data block.

[0028] This ensures that several data reduction techniques can be evaluated.

[0029] In yet another implementation of the first aspect, the block storage device is further configured to: if the size obtained by deduplication and compression is greater than or equal to the size obtained by large-block compression, store the data block using large-block compression.

[0030] This ensures that the most effective data reduction technique is selected.

[0031] Specifically, the block storage device is further configured to: if the size obtained by deduplication and compression is smaller than the size obtained by large block compression, store the data block using deduplication and compression.

[0032] In another implementation of the first aspect, the block storage device is further configured to perform deduplication on a first sub-block portion of the data block, compress a second sub-block portion of the data block, and store the obtained data block, so as to perform deduplication and compression on the data block.

[0033] This provides a detailed implementation of simultaneously applying deduplication and compression to maximize the data reduction effect.

[0034] Specifically, the first sub-block portion and the second sub-block portion together form all sub-blocks of the data block. Specifically, the first portion and the second portion do not overlap.

[0035] Specifically, the compression applied in this step is different from large block compression because it is not applied to the entire data block, but to at least one sub-block of the data block.

[0036] In yet another implementation of the first aspect, the block storage device is further configured to: for each sub-block of the data block, write a similarity hash into a similarity hash table.

[0037] This ensures that in the first operation stage, the prerequisite for performing similarity deduplication in the second operation stage can be completed, that is, when the similarity hashes of the sub-blocks are most efficiently obtained and written.

[0038] Specifically, one similarity hash corresponds to one sub-block.

[0039] In another implementation of the first aspect, the block storage device is further configured to: in the second operation stage, determine whether the size of the data block stored in the first operation stage can be further reduced using similarity deduplication according to the similarity hash.

[0040] This ensures that in the second operation stage, the size of the data block can be further reduced, specifically the size reduction technique most suitable for the second operation stage.

[0041] Specifically, the similarity hash is the similarity hash stored in the similarity hash table.

[0042] Specifically, similarity deduplication includes that deduplication is also applied to all sub-blocks of the data block.

[0043] In another implementation of the first aspect, the block storage device is further configured to determine the space required to store the data block using similarity deduplication and the space required to store the data block in its current format.

[0044] This ensures that the effectiveness of similarity deduplication can be evaluated.

[0045] Specifically, the current format is either large data block compression or a combination of deduplication and compression.

[0046] In another implementation of the first aspect, the block storage device is further configured to: if the space required to store the data block using similarity deduplication is less than the space required to store the data block in its current format, then store the data block in the data memory using similarity deduplication only.

[0047] This ensures that in the second operation phase, the most effective data reduction method is selected. Specifically, if similarity deduplication is more effective than the current format used to store the data block, then only similarity deduplication is selected.

[0048] In yet another implementation of the first aspect, the second operation phase is an offline phase.

[0049] This ensures that the block storage device can be used in the offline phase.

[0050] Specifically, the offline phase is executed periodically. Specifically, periodically means according to a predefined time interval. Specifically, periodically means that this operation phase is executed independently of write operations and / or read operations.

[0051] A second aspect of the present invention provides a method for data compression, the method comprising the following steps: in a first operation phase, if a block storage device determines that a data block is to be written to a large block storage area of the block storage device, then the block storage device determines whether the data block can be deduplicated; if the data block cannot be deduplicated, then the block storage device stores the data block using large block compression.

[0052] In one implementation of the second aspect, the large block storage area is a data block read and / or write area with an average size greater than a predefined threshold.

[0053] In yet another implementation of the second aspect, the method further comprises the block storage device determining the average size according to read statistics and / or write statistics.

[0054] In yet another implementation of the second aspect, the first operation phase is an online phase.

[0055] In another implementation of the second aspect, the method further includes: if the data block can be deduplicated, the block storage device performs deduplication and compression on the data block, and the block storage device stores the obtained data block.

[0056] In yet another implementation of the second aspect, the method further includes: if the data block can be deduplicated, the block storage device compares the size obtained by performing deduplication and compression on the data block with the size obtained by performing large-block compression on the data block.

[0057] In yet another implementation of the second aspect, the method further includes: if the size obtained by deduplication and compression is greater than or equal to the size obtained by large-block compression, the block storage device stores the data block using large-block compression.

[0058] In another implementation of the second aspect, the method further includes the block storage device performing deduplication on the first sub-block portion of the data block, the block storage device performing compression on the second sub-block portion of the data block, and the block storage device storing the obtained data block, performing deduplication and compression on the data block.

[0059] In another implementation of the second aspect, the method further includes: for each sub-block of the data block, the block storage device writes a similarity hash into a similarity hash table.

[0060] In another implementation of the second aspect, the method further includes: in a second operation stage, the block storage device determines whether the size of the data block stored in the first operation stage can be further reduced using similarity deduplication based on the similarity hash.

[0061] In another implementation of the second aspect, the method further includes: the block storage device determines the space required to store the data block using similarity deduplication, and the block storage device determines the space required to store the data block in its current format.

[0062] In yet another implementation of the second aspect, the method further includes: if the space required to store the data block using similarity deduplication is less than the space required to store the data block in its current format, the block storage device stores the data block in the data memory using similarity deduplication only.

[0063] In yet another implementation of the second aspect, the second operation stage is an offline stage.

[0064] The second aspect and its implementation manners include the same advantages as those of the first aspect and its respective implementation manners.

[0065] A third aspect of the present invention provides a computer program product including instructions that, when executed by a computer, cause the computer to perform the steps of the method of the second aspect or any of its implementation manners.

[0066] The third aspect and its implementation manners include the same advantages as those of the second aspect and its respective implementation manners.

[0067] A fourth aspect of the present invention provides a non-transitory computer-readable storage medium including instructions that, when executed by a computer, cause the computer to perform the steps of the method of the second aspect or any of its implementation manners.

[0068] The fourth aspect and its implementation manners include the same advantages as those of the second aspect and its respective implementation manners.

[0069] In other words, the present invention provides a solution that determines which data reduction method will produce the best results. For example, compared with another available method, a data reduction method can produce better CPU performance and / or can use disk space more efficiently. In addition, a two-stage method can be fully utilized. For example, according to the design considerations and available resources of each stage, adopting different data reduction methods in each stage can reduce the required space more efficiently and effectively. In other words, the present invention adopts a solution that analyzes a series of parameters, such as the size of data blocks, the type of read / write operations usually performed at the current address, or the compressibility of data, and selects the most suitable data reduction method according to the results of the analysis. In addition, the analysis and decision-making process uses a two-stage method, in which differential logic is applied to online and background processing. This enables the block storage device to prioritize performance during online processing when the CPU occupancy rate is most critical, and to prioritize disk space during background processing when there is more available space on the CPU. Specifically, the block storage device can utilize at least one of these data reduction methods: deduplication of identical blocks, compression of large blocks, compression of small blocks, similarity compression.

[0070] It should be noted that all devices, elements, units, and modules described in this application can be implemented in software or hardware elements or any type of combination thereof. All steps performed by various entities described in this application and the functions described as being to be performed by various entities are intended to indicate that the corresponding entities are suitable or used to perform the corresponding steps and functions. Although in the description of the following specific embodiments, the specific functions or steps performed by external entities are not reflected in the description of the specific elements of the entity performing the specific steps or functions, those skilled in the art should clearly understand that these methods and functions can be implemented in the corresponding hardware elements or software elements or any type of combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In conjunction with the accompanying drawings, the following description of specific embodiments will set forth the aspects of the present invention and their implementation methods as described above.

[0072] Figure 1 A schematic diagram of a device provided by an embodiment of the present invention is shown;

[0073] Figure 2 A more detailed schematic diagram of a device provided by an embodiment of the present invention is shown;

[0074] Figure 3 A schematic diagram of an operation mode provided by the present invention is shown;

[0075] Figure 4 A schematic diagram of an operation mode provided by the present invention is shown;

[0076] Figure 5 Another schematic diagram of a method provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] Figure 1 A schematic diagram of a block storage device 100 provided by an embodiment of the present invention is shown. The block storage device 100 is used for data compression and thus is used to: determine whether data block 101 can be deduplicated if the block storage device 100 determines that the data block 101 is to be written to the large block storage area 102 of the block storage device 100.

[0078] In addition, the block storage device 100 is used to: store the data block 101 using large block compression if the data block 101 cannot be deduplicated. The data block 101 is specifically stored in the large block storage area 102.

[0079] All these steps are performed during a first operation phase. The first operation phase can specifically be an online phase.

[0080] Optionally, the large block storage area 102 can be a data block read and / or write area with an average size greater than a predefined threshold. The large block storage area can specifically be a part of a physical disk, a physical volume, a virtual disk, or a virtual volume. The average size of the data block to be written or read can optionally be based on read statistics and / or write statistics.

[0081] Figure 2 A more detailed schematic diagram of a block storage device 100 provided by an embodiment of the present invention is shown. Figure 2 The device 100 shown includes Figure 1 all the features and functions of the device 100, as well as the following optional features:

[0082] Such as Figure 2As shown, the block storage device 100 may optionally be used to: if the data block 101 can be deduplicated, deduplicate and compress the data block 101, and store the resulting data block 201. That is, perform deduplication and compression on the data block 101, rather than perform large-block compression. The difference between this compression and large-block compression is specifically that: it is applied to data blocks smaller than large blocks.

[0083] Optionally, if the data block 101 can be deduplicated, the block storage device 100 may compare the size obtained by deduplicating and compressing the data block 101 with the size obtained by performing large-block compression on the data block 101. This evaluation helps to determine a more effective data reduction method. If the size obtained by deduplication and compression is greater than or equal to the size obtained by large-block compression, the block storage device 100 may use large-block compression to store the data block 101.

[0084] As Figure 2 further shown, the block storage device 100 may deduplicate the first sub-block portion 202 of the data block 101. The block storage device 100 may also compress the second sub-block portion 203 of the data block 101. Then, the data block 201 obtained by deduplication and compression may be stored by the block storage device 100. This may be done, for example, in the large-block storage area 102 or in the traditional storage area of the block storage device 100. For example, if the size of the data block 101 is 1 MB, the first sub-block portion 202 may include, for example, two sub-blocks, each with a size of 256 KB, and the second sub-block portion 203 may include, for example, four sub-blocks, each with a size of 128 KB. However, any type of sub-block distribution following this principle is possible.

[0085] In a specific embodiment, two-level block sizes are used. In this embodiment, the size of the data block 101 is 512 KB and the size of the sub-block is 8 KB. Thus, the 512 KB data block 101 can be compressed into a single block or compressed into 64 sub-blocks of 8 KB size, where each 8 KB sub-block can be deduplicated or compressed individually.

[0086] As Figure 2 further shown, to prepare for similarity deduplication (which may be performed, for example, in the second operation stage of the block storage device 100), for each sub-block 202, 203 of the data block 101, the block storage device 100 may write a similarity hash 204 to the similarity hash table 205. This step is specifically performed during the first operation stage of the block storage device 100.

[0087] Optionally, in the second operation phase, the block storage device 100 may determine whether the size of the data block 101 stored in the first operation phase can be further reduced using similarity deduplication based on the similarity hash 204. The similarity hash 204 may be, for example, the similarity hash stored in the similarity hash table 205 during the first operation phase.

[0088] Optionally, the block storage device 100 may determine the space required to store the data block 101 using similarity deduplication. Optionally, the block storage device 100 may also determine the space required to store the data block (101) in its current format. This enables the use of similarity deduplication to store only the data block 101 in the data memory if the space required to store the data block 101 using similarity deduplication is less than the space required to store the data block 101 in its current format.

[0089] Specifically, the second operation phase may be an offline phase.

[0090] Figure 3 The operation scenario performed during the first operation phase is described in more detail.

[0091] In step 301, the block storage device 100 checks whether the data block 101 can be deduplicated (it can be determined that the same data block has been written to the disk based on the hash fingerprint), and if so, performs deduplication (see step 302). In any case, in step 301, the block storage device 100 generates a minimum hash (i.e., the similarity hash 204), which will be added to the opportunity table (i.e., the similarity hash table 205) so that similarity deduplication can be performed subsequently in the second operation phase.

[0092] If the data block 101 cannot be deduplicated, the block storage device 100 checks (in step 303) the prediction table to determine the available physical address range that is typically used for writing and reading small data blocks or writing and reading large data blocks (in other words, to determine whether the data block 101 will be written to the large block storage area 102). If the current range is typically used for writing and reading small data blocks, the system proceeds as normal (i.e., it stores the small data block), as shown in step 304.

[0093] If it is recognized that data block 101 is to be written to an area where most reads and writes are done in large chunks, the block storage device 100 analyzes (in step 305) whether large chunk compression will result in better compression than small chunk compression. If large chunk compression does not result in better compression, the block storage device proceeds as normal (i.e., it stores the compressed data block 101), see step 306. If large chunk compression will result in better compression, the block storage device 100 stores the compressed data block 101 using large chunk compression settings (see step 307). For example, this can be achieved in one of the following ways: Large chunk compression can be used to mark large chunk storage areas 102 as containing the same data. The block storage device 100 can, for example, write the same granular (physical) address at the same offset in each relevant large chunk storage area 102, or the block storage device 100 can set bits indicating the start and end of the range of the same large chunk storage area 102.

[0094] Furthermore, if data block 101 is eligible for deduplication, the block storage device 100 can also determine whether large chunk compression of data block 101 will result in better results, or whether a combination of deduplication and compression of data block 101 will result in better results, and select the better option.

[0095] Thus, at the end of the first operation phase, the block storage device 100 has generated and saved the minimum hash, and has either deduplicated data block 101 and / or has stored data block 101 compressed in small chunks; or has stored data block 101 using large chunk compression.

[0096] Figure 4 The operation scenarios performed during the first operation phase are described in more detail.

[0097] For example, as described above, online processing (i.e., the first operation phase) prioritizes reducing CPU occupancy rather than minimizing disk utilization. Then, during background processing (i.e., the second operation phase), when CPU occupancy is less critical, the block storage device 100 attempts to further optimize disk utilization, as described below. Figure 4 Step 401 specifically is performed Figure 3 after step 307.

[0098] As shown in step 402, during the second operation phase, for data blocks 401 compressed in large data chunks, the block storage device 100 checks the opportunity table (i.e., the similarity hash table 205) to see if there is a similarity deduplication option (i.e., if there is another data block 101 with the same minimum hash).

[0099] If similarity deduplication cannot be performed, the block storage device 100 does nothing, and the data block 401 remains large-block compressed, see step 403.

[0100] If similarity deduplication can be performed, the block storage device 100 analyzes whether similarity deduplication will produce better results than large-block compression, see step 404.

[0101] If similarity deduplication will produce lower data reduction than large-block compression, the system does nothing, and the data block 101 remains large-block compressed, see step 405.

[0102] If similarity deduplication will produce better data reduction than large-block compression, the system performs similarity deduplication, and the data block 101 is stored, for example, in 8k data blocks, which, if necessary, have similarity deduplication pointers.

[0103] Figure 5 A schematic diagram of a method 500 provided by an embodiment of the present invention is shown. The method 500 is used for data compression. In a first operation stage, the method includes the following steps: If the block storage device 100 determines that the data block 101 is to be written to the large-block storage area 102 of the block storage device 100, the block storage device 100 determines 501 whether the data block 101 can be deduplicated. The method includes (still in the first operation stage) the following another step: If the data block 101 cannot be deduplicated, the block storage device 100 stores 502 the data block 101 using large-block compression.

[0104] The present invention has been described in connection with different embodiments and implementations taken as examples. However, upon study of the drawings, the present invention, and the independent claims, those skilled in the art will be able to understand and implement other variations when practicing the claimed invention. In the claims as well as in the description, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. Stating certain measures in mutually different dependent claims does not indicate that a combination of these measures cannot be used in a beneficial implementation.

Claims

1. A block storage device (100) for data compression, characterized in that, for: in a first operation stage, - if the block storage device (100) determines that a data block (101) is to be written to a large block storage area (102) of the block storage device (100), determine whether the data block (101) can be deduplicated; the large block storage area (102) is a data block read and / or write area with an average size greater than a predefined threshold; - if the data block (101) cannot be deduplicated, store the data block (101) using large block compression; - if the data block (101) can be deduplicated, compare the size obtained by deduplicating and compressing the data block (101) with the size obtained by performing large block compression on the data block (101); the compression is different from the large block compression, and the large block compression represents performing a single overall compression on a data block; - if the size obtained by deduplication and compression is greater than or equal to the size obtained by large block compression, store the data block (101) using large block compression; - if the size obtained by deduplication and compression is less than the size obtained by large block compression, store the data block (101) using deduplication and compression.

2. The block storage device (100) according to claim 1, characterized in that, it is also used to determine the average size according to read statistics and / or write statistics.

3. The block storage device (100) according to claim 1, characterized in that, the first operation stage is an online stage.

4. The block storage device (100) according to claim 1, characterized in that, the block storage device (100) is also used to: if the data block (101) can be deduplicated, perform deduplication and compression on the data block (101) and store the resulting data block (201).

5. The block storage device (100) according to claim 1 or 4, characterized in that, the block storage device (100) is also used to perform deduplication on a first sub-block portion (202) of the data block (101), and perform compression on a second sub-block portion (203) of the data block (101) and store the resulting data block (201) to perform deduplication and compression on the data block (101).

6. The block storage device (100) according to claim 1, characterized in that, the block storage device (100) is also used to: for each sub-block (202, 203) of the data block (101), write a similarity hash (204) to a similarity hash table (205).

7. The block storage device (100) according to claim 1, characterized in that, it is also used to: in a second operation stage, determine whether the size of the data block (101) stored in the first operation stage can be further reduced using similarity deduplication according to the similarity hash (204).

8. The block storage device (100) according to claim 7, characterized in that, It is also used to determine the space required to store the data block (101) using similarity deduplication and the space required to store the data block (101) in its current format.

9. The block storage device (100) according to claim 8, wherein, it is also used to: if the space required to store the data block (101) using similarity deduplication is less than the space required to store the data block (101) in its current format, then use similarity deduplication to store only the data block (101) in the data memory.

10. The block storage device (100) according to any one of claims 7-9, wherein, the second operation phase is an offline phase.

11. A method (500) for data compression, wherein, the method (500) includes the following steps: in a first operation phase, - if the block storage device (100) determines that a data block (101) is to be written to the large block storage area (102) of the block storage device (100), then the block storage device (100) determines (501) whether the data block (101) can be deduplicated; the large block storage area (102) is a data block read and / or write area with an average size greater than a predefined threshold; - if the data block (101) cannot be deduplicated, then the block storage device (100) stores the data block (101) using large block compression (502); - if the data block (101) can be deduplicated, then compare the size obtained by deduplicating and compressing the data block (101) with the size obtained by performing large block compression on the data block (101); the compression is different from the large block compression, and the large block compression represents performing an overall one-time compression on a data block; - if the size obtained by deduplicating and compressing is greater than or equal to the size obtained by large block compression, then store the data block (101) using large block compression; - if the size obtained by deduplicating and compressing is less than the size obtained by large block compression, then store the data block (101) using deduplication and compression.

12. A computer program product, wherein, it includes instructions that, when the program is executed by a computer, cause the computer to perform the steps of the method (500) according to claim 11.

Citation Information

Patent Citations

  • Efficiently estimating data compression ratio of ad-hoc set of files in protection storage filesystem with stream segmentation and data deduplication

    US10430383B1

  • Deduplication and compression of data segments in a data storage system

    US20180329631A1

  • Storage system and data processing method for the same

    WO2010097960A1