Backup data compression method and device, equipment, storage medium and program product
By optimizing incremental compression based on the similarity and lifecycle concepts between data blocks, the problem of high overhead of incremental compression algorithms in existing technologies is solved, data recovery efficiency and storage logic are improved, and more efficient data management is achieved.
Patent Information
- Application Number
- CN202510781394.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Existing incremental compression algorithms have high overhead in the data compression and recovery process, especially when random read operations are frequent, resulting in unstable performance.
By obtaining the similarity between data blocks, determining the number of records of the basic data block, and selecting the target data block for incremental compression based on the preset threshold, the storage and recovery process is optimized by combining the data block life cycle concept.
It improves data compression effect, reduces random read overhead, improves data recovery efficiency and storage logic, and achieves more efficient data management.
Smart Images

Figure CN120704944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a backup data compression method, device, equipment, storage medium and program product. Background Art
[0002] Data deduplication has attracted increasing attention in data-intensive storage systems. As one of the most effective data reduction methods in recent years, fingerprint-based data deduplication technology eliminates duplicate data blocks by checking the security fingerprint of data blocks (i.e., SHA-1 / SHA-256 signatures), which has been widely used in commercial backup and archive storage systems.
[0003] The data compression method based on basic block and security fingerprint usually uses a fixed-length block method and a hash algorithm (such as SHA series, MD5, etc.) to quickly verify whether there are duplications between data blocks. However, this method also has a more obvious shortcoming, that is, it cannot effectively process similar data. Only two data blocks that are 100% similar will be judged as duplicates. The performance in terms of duplication speed and data compression effect is unsatisfactory.
[0004] Incremental compression is another data reduction method that is more scalable in large-scale storage systems. Similarity detection-based incremental compression aims to improve the performance of the basic strategy. However, the incremental compression algorithm incurs a certain degree of additional overhead in both compression and recovery, resulting in high overhead. Summary of the Invention
[0005] The present invention provides a backup data compression method, device, equipment, storage medium and program product, which are used to solve the problem of high overhead of existing incremental compression algorithms.
[0006] In order to solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a backup data compression method, comprising:
[0008] Obtaining similarity between any two data blocks in the backup data;
[0009] Determining the number of records of each data block including a basic data block according to the similarity;
[0010] Determining a target data block to be compressed in the data block according to the number of records and a preset number threshold;
[0011] Incremental compression is performed on the target data block to obtain compressed backup data.
[0012] Optionally, obtaining the similarity between any two data blocks of the backup data includes:
[0013] Performing a block operation on the backup data to obtain at least two initial data blocks;
[0014] Performing fingerprint deduplication processing on each of the initial data blocks to obtain at least two deduplication-processed data blocks;
[0015] Performing similarity calculation on the first data block and the second data block to obtain the similarity between any two of the data blocks;
[0016] The first data block is any one of the at least two deduplication-processed data blocks, and the second data block is a data block other than the first data block among the at least two deduplication-processed data blocks.
[0017] Optionally, determining the number of records of the basic data block included in each data block according to the similarity includes:
[0018] Acquire a first data block group whose similarity between any two data blocks is greater than a preset threshold, wherein the first data block group includes a third data block and a fourth data block corresponding to the third data block;
[0019] According to the first rule, the number of times that the basic data block is included in each of the data blocks is determined, wherein the first rule includes: determining the basic data blocks with the same data in the third data block and the fourth data block, and recording the basic data blocks included in the third data block and the fourth data block once accordingly.
[0020] Optionally, determining a target data block to be compressed in the data blocks according to the number of records and a preset number threshold includes:
[0021] Determining the data block whose recording times is greater than the preset times threshold as an initial compressed data block;
[0022] The initial compressed data block is subjected to noise removal processing to obtain the target data block to be compressed.
[0023] Optionally, the method further includes:
[0024] In a case where the compressed backup data consists of at least two copies of backup data, obtaining a target data block after incremental compression in each of the compressed backup data;
[0025] Performing data recovery on the target data blocks after incremental compression according to the first logical sequence corresponding to the target data blocks after incremental compression to obtain recovered backup data;
[0026] The first logical order is determined according to whether the target data block after incremental compression is compressed in each compressed backup data.
[0027] Optionally, the method further includes:
[0028] Obtaining a first compression sequence of the compressed backup data;
[0029] The first logical order is determined according to the first compression order and whether the target data block after the incremental compression is compressed in each compressed backup data.
[0030] Optionally, performing data recovery on the incrementally compressed target data blocks according to the first logical order corresponding to the incrementally compressed target data blocks to obtain recovered backup data includes:
[0031] determining, according to the first logical order, in the compressed target data blocks, a fifth data block that needs to be restored corresponding to each compressed backup data and a restoration order corresponding to the fifth data blocks;
[0032] According to the recovery sequence, data recovery is performed on the fifth data block in sequence to obtain the restored backup data.
[0033] In a second aspect, an embodiment of the present invention further provides a backup data compression device, comprising:
[0034] A first acquisition module is used to acquire the similarity between any two data blocks in the backup data;
[0035] A first determining module, configured to determine the number of records containing a basic data block in each of the data blocks according to the similarity;
[0036] A second determining module is configured to determine a target data block to be compressed among the data blocks according to the number of records and a preset number threshold;
[0037] The first processing module is configured to perform incremental compression on the target data block to obtain compressed backup data.
[0038] In a third aspect, an embodiment of the present invention further provides a backup data compression device, comprising: a processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the backup data compression method described in any one of the first aspects are implemented.
[0039] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium having a program stored thereon, and when the program is executed by a processor, the steps in the backup data compression method as described in any one of the first aspects are implemented.
[0040] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps in the backup data compression method as described in any one of the first aspects.
[0041] The beneficial effects of the present invention are:
[0042] The backup data compression method provided by the solution of the present invention obtains the similarity between any two data blocks in the data blocks of the backup data, determines the number of records including the basic data block in each data block based on the similarity, determines the target data blocks that need to be compressed in the data blocks based on the number of records and a preset number threshold, accelerates the acquisition process of the data blocks that need to be compressed, performs incremental compression on the target data blocks, and obtains compressed backup data. That is, the target data blocks selected by the present invention when performing incremental compression are data blocks with greater similarity, thereby improving the compression effect and reducing random read overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram showing the distribution of incrementally compressed data provided by the present invention;
[0044] Figure 2 A flowchart showing a backup data compression method provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram showing a similarity matrix provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram showing data blocks included in backup data provided by an embodiment of the present invention;
[0047] Figure 5 A schematic diagram showing the relationship between backup data and blocks provided by an embodiment of the present invention;
[0048] Figure 6 A schematic diagram showing a divided life cycle according to an embodiment of the present invention;
[0049] Figure 7 A schematic diagram showing a logic block provided by an embodiment of the present invention;
[0050] Figure 8 A schematic diagram showing the structure of a backup data compression device provided by an embodiment of the present invention;
[0051] Figure 9 A schematic diagram showing the structure of a backup data compression device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] In order to make the technical problems, technical solutions and advantages to be solved by the present application clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. In the following description, specific details such as specific configurations and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures has been omitted.
[0053] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0054] In the various embodiments of the present application, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0055] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0056] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.
[0057] Before describing the specific embodiments of the present invention, the following description is first given:
[0058] Data backup and recovery: Data backup refers to the process of copying data and databases to other storage media or locations in order to protect data security.
[0059] Incremental compression: Incremental compression is a data compression technology that reduces the size of backup data by storing only the data blocks that have changed since the last backup, thereby improving storage efficiency and reducing storage costs.
[0060] Lifecycle: The lifecycle refers to the complete time span of an object from its creation, use to its final deletion or archiving. In the field of data management, it specifically refers to the entire process of storage, use, and maintenance of data or data blocks in the system.
[0061] In the prior art, when performing incremental compression, similarity detection between data blocks requires comparative calculations between non-continuous data blocks. This operation will bring about a large number of random read operations, and when memory is limited, a large amount of memory swapping will occur, which will bring about input (Input) / output (Output) overhead. Random read means that when the disk reads data, the access to data blocks is disordered, that is, the position of the data blocks read each time on the disk is randomly distributed. In this reading mode, the disk head needs to move frequently to different positions to read data, so the performance is usually low. This uncontrollable random read will make the compression operation performance extremely unstable. During recovery, the incremental compression method will also produce performance bottlenecks. Suppose data blocks A and A' are a group of similar data blocks, which become the incremental deviation dA between data blocks A and A' and A after incremental compression processing. When recovering data block A', it is necessary to obtain data blocks A and dA at the same time, such as Figure 1 As shown in the figure, it is not difficult to find that although incremental compression of backup data reduces the overall space usage, it destroys the spatial locality between data blocks. This may also cause a large number of random reads during data recovery, resulting in uncontrollable additional performance overhead at the I / O level.
[0062] To solve the problem of high overhead in existing incremental compression algorithms, embodiments of the present invention provide a backup data compression method, apparatus, device, storage medium, and program product.
[0063] like Figure 2 As shown, an embodiment of the present invention provides a backup data compression method, including:
[0064] Step 201: Obtain the similarity between any two data blocks in the backup data.
[0065] It should be noted that the data blocks of the backup data are formed by dividing the backup data into data of a preset size (preset KB).
[0066] In this step, after the backup data is divided into at least two data blocks, the similarity between any two different data blocks is counted to obtain the similarity between the two data blocks.
[0067] Step 202: Determine the number of times each data block includes a record of a basic data block according to the similarity.
[0068] It should be noted that the basic data block refers to the portion of data that is identical between two data blocks during the process of performing incremental compression on the two data blocks.
[0069] In this step, the number of times the basic data block is recorded in each data block can be obtained according to the determined similarity between any two data blocks. The number of times the basic data block is recorded is used to subsequently determine whether the data block is compressed.
[0070] Step 203: Determine the target data block that needs to be compressed in the data blocks according to the number of records and a preset number threshold.
[0071] The target data blocks that need to be compressed are determined by using the number of times each data block records the basic data block and the preset number threshold, that is, reasonable target data blocks are selected for incremental compression, and selective incremental compression is performed between data blocks.
[0072] Step 204: perform incremental compression on the target data block to obtain compressed backup data.
[0073] Through the above steps, the target data blocks selected by the present invention during incremental compression are data blocks with relatively high similarity, thereby improving the compression effect and reducing random read overhead.
[0074] In an optional embodiment, obtaining the similarity between any two data blocks in the backup data includes:
[0075] The backup data is divided into blocks to obtain at least two initial data blocks.
[0076] Specifically, the backup data is divided into data of a preset size (preset KB) to obtain at least two initial data blocks after division.
[0077] Fingerprint deduplication processing is performed on each of the initial data blocks to obtain at least two deduplication-processed data blocks.
[0078] For example, for a data backup data, data blocks A, B, C, D, and E are obtained after block segmentation and fingerprint deduplication processing.
[0079] Performing similarity calculation on the first data block and the second data block to obtain the similarity between any two of the data blocks;
[0080] The first data block is any one of the at least two deduplication-processed data blocks, and the second data block is a data block other than the first data block among the at least two deduplication-processed data blocks.
[0081] Specifically, similarity calculation is performed between any two data blocks to obtain the similarity between the any two data blocks. The similarity between the any two data blocks can be represented in the form of a similarity matrix.
[0082] For example, Figure 3 As shown, Figure 3 The similarity matrix obtained by performing similarity statistics between any two data blocks in data blocks A, B, C, D, and E is shown in FIG. Figure 3 As shown, the similarity between data block A and data block B is 0.7, the similarity between data block A and data block C is 0.2, the similarity between data block A and data block D is 0.5, the similarity between data block A and data block E is 0.9, the similarity between data block B and data block C is 0.5, the similarity between data block B and data block D is 0.8, the similarity between data block B and data block E is 0.4, the similarity between data block C and data block D is 0.3, the similarity between data block C and data block E is 0.5, and the similarity between data block D and data block E is 0.3.
[0083] In an optional embodiment, determining the number of records of the basic data block included in each data block according to the similarity includes:
[0084] Acquire a first data block group whose similarity between any two data blocks is greater than a preset threshold, wherein the first data block group includes a third data block and a fourth data block corresponding to the third data block;
[0085] According to the first rule, the number of times that the basic data block is included in each of the data blocks is determined, wherein the first rule includes: determining the basic data blocks with the same data in the third data block and the fourth data block, and recording the basic data blocks included in the third data block and the fourth data block once.
[0086] It should be noted that by analyzing the backup data set (including multiple backup data), a threshold t equal to the total number of data blocks is set for the data blocks in the backup data set. Incremental compression is performed on the data blocks of the backup data set. During this process, the number of times each data block is selected as a base data block in the incremental compression is counted. A base data block refers to the portion of data that is identical between two data blocks during the incremental compression process. Ultimately, the number of data blocks selected as base data blocks more than the threshold t is counted, and it is found that the number of data blocks that meet the criteria accounts for approximately 30% to 60% of the total data set. During the incremental compression operation, the statistical distribution of each data block selected as a base data block exhibits a long-tail characteristic. When a reasonable threshold t is determined, data blocks selected more than the threshold t can produce an effective compression effect during the compression process. Data blocks that do not exceed the threshold have limited compression effect, while also incurring additional random read overhead.
[0087] In this optional embodiment, according to the similarity matrix, a first data block group with a similarity greater than a preset threshold can be obtained. The first data block group includes two data blocks. Figure 3 Taking the similarity matrix shown as an example, the preset threshold is set to 0.6, and the first data group selected includes: data block A and data block B, data block A and data block E, data block B and data block D. The number of records that include the basic data block in data block A is 2, the number of records that include the basic data block in data block B is 2, the number of records that include the basic data block in data block B is 2, the number of records that include the basic data block in data block D is 1, and the number of records that include the basic data block in data block E is 1.
[0088] In an optional embodiment, determining a target data block to be compressed in the data blocks according to the number of records and a preset number threshold includes:
[0089] Determining the data block whose recording times is greater than the preset times threshold as an initial compressed data block;
[0090] The initial compressed data block is subjected to noise removal processing to obtain the target data block to be compressed.
[0091] Continuing with the above example, the preset number threshold can be set to 1, and the data block A and the data block B having a record number greater than 1 are used as the initial compressed data blocks to be compressed.
[0092] Afterwards, the selected initial compressed data block can be filtered to remove noise blocks to obtain the target data block that needs to be compressed.
[0093] In this optional embodiment, the similarity matrix is used as an index to accelerate the acquisition of target data blocks in the compression process.
[0094] It should be noted that basic incremental compression results in a large number of random reads during data recovery, which reduces the spatial locality of the compressed data and incurs additional I / O overhead. To address this issue, a lifecycle-based incremental data compression data distribution method is needed to repartition and sort the data, ensuring a certain degree of continuity in the data set. This ensures that data blocks are accessed consistently when the disk reads data, mitigating the random read problem.
[0095] Assume that there are backup data 1, backup data 2, and backup data 3, and the data blocks they contain are as follows: Figure 4 As shown, backup data 1 includes data block A, data block B, data block C, data block D, data block E, data block F, and data block G; backup data 2 includes data block A, data block B', data block E, data block F', data block H, and data block I; and backup data 3 includes data block A, data block B', data block C, data block D', data block E, and data block J. Under the existing incremental compression method, the compressed data blocks are stored in each storage block in sequence, as shown in FIG. Figure 5 As shown in the figure, if you want to restore each backup data, backup data 2 and backup data 3 will require different numbers of random reads. That is, the restored data cannot be read in the order of ABCDEFG. Instead, various random jumps are required to restore different data blocks. The number of random reads is related to the number and distribution of compressed data blocks.
[0096] When managing and manipulating data, different minimum data units are usually defined in memory and on disk to facilitate reading, writing, and operations. For example, the minimum unit of in-memory data in a database is a tuple, while the minimum unit of disk data is a page. A page can contain multiple tuples. When performing searches, reads, or writes, the entire page containing the tuple is read from disk in order to retrieve the target tuple. In existing incremental compression algorithms, data blocks A and B can be understood as the minimum memory unit tuple in the above example, while Block represents the minimum disk unit page in the above example. Therefore, a Block can contain multiple pages. At the same time, when reconstructing data, when reading a Block, the data in the Block may not be completely relevant to the data to be reconstructed.
[0097] by Figure 5 The data recovery solution provided by the embodiment of the present invention is illustrated by taking backup data 1, backup data 2, backup data 3 and Block division in the target data block as an example. In this case, when data recovery is required, data backup 1 will depend on block1, block2, and block3 (data block A and data block B in the target data block depend on block1, data block C and data block D in the target data block depend on block2, and data block F and data block G in the target data block depend on block3). Therefore, data backup 2 will depend on block1, block3, block4, block5, and block6 (data block A in the target data block depends on block1, data block B' in the target data block depends on block1 and block4, data block E in the target data block depends on block4, and data block Data block F' in the target data block depends on block3 and block5, data block H in the target data block depends on block5, and data block I in the target data block depends on block6). Data backup 3 depends on block1, block2, block4, block6, and block7 (data block A in the target data block depends on block1, data block B' in the target data block depends on block1 and block4, data block C in the target data block depends on block2, data block D' in the target data block depends on block2 and block6, data block E in the target data block depends on block4, and data block J in the target data block depends on block7). It can be seen that during data recovery, random reads will occur in the compressed data of data backup 2 and data backup 3, affecting overall performance.
[0098] Furthermore, the method further comprises:
[0099] In a case where the compressed backup data consists of at least two copies of backup data, obtaining a target data block after incremental compression in each of the compressed backup data;
[0100] Performing data recovery on the target data blocks after incremental compression according to the first logical sequence corresponding to the target data blocks after incremental compression to obtain recovered backup data;
[0101] The first logical order is determined according to whether the target data block after incremental compression is compressed in each compressed backup data.
[0102] Therefore, in an embodiment of the present invention, the concept of life cycle is introduced, and a logical order (i.e., a first logical order) is given for at least two backup data. The first logical order is determined based on whether the target data block after incremental compression is compressed in each of the compressed backup data.
[0103] The start and end points of the life cycle of a data block may be understood as the order between the backup data of the data block that is first incrementally compressed and the backup data that is last incrementally compressed, as indicated by the first logical order.
[0104] Optionally, the method further includes:
[0105] Obtaining a first compression sequence of the compressed backup data;
[0106] For example, taking the three backup data as backup data 1, backup data 2 and backup data 3, the first compression order of the compressed backup data is determined according to the order in which the backup data is compressed, or according to the identification order, and the first compression order is backup data 1, backup data 2 and backup data 3.
[0107] The first logical order is determined according to the first compression order and whether the target data block after the incremental compression is compressed in each compressed backup data.
[0108] In such Figure 4 and Figure 5 In the example data, only data block F and data block G are used when restoring backup data 1, so their life cycle is 1-1. Data block E is used when restoring backup data 2 and backup data 3, so its life cycle is 2-3. Based on this definition, a layer of logical blocks can be constructed on the data blocks, and Block(x,y) can be used to represent the data blocks with a life cycle of x-y. Block(x,y) can be understood as the first logical order. The life cycle after division is as follows Figure 6 shown.
[0109] Optionally, performing data recovery on the incrementally compressed target data blocks according to the first logical order corresponding to the incrementally compressed target data blocks to obtain recovered backup data includes:
[0110] determining, according to the first logical order, in the compressed target data blocks, a fifth data block that needs to be restored corresponding to each compressed backup data and a restoration order corresponding to the fifth data blocks;
[0111] According to the recovery sequence, data recovery is performed on the fifth data block in sequence to obtain the restored backup data.
[0112] For example Figure 4 and Figure 5 The example data can be divided into the following blocks according to the above life cycle logic: Figure 7 A set of logic blocks is shown, which is used to indicate the recovery order.
[0113] When data blocks are stored sequentially in descending order (x, y), each backup data can be retrieved from the disk through approximately sequential reads, effectively reducing the I / O overhead caused by random reads during data recovery.
[0114] The embodiments of the present invention aim to address performance issues caused by random reads in incremental compression algorithms during both the data compression and recovery phases. They creatively introduce the concept of a data block lifecycle and utilize this concept to optimize the data storage and recovery process. This approach not only addresses the issues of frequent random reads and high I / O overhead in the prior art, but also significantly improves data recovery efficiency through logical block partitioning and sequential storage. Furthermore, the present invention improves the logic and organization of data storage, making data management more efficient and organized. This approach can bring substantial advancements to the field of data backup and recovery. The innovative introduction of the concept of a data block lifecycle and utilizes this concept to optimize the data storage and recovery process. This approach not only addresses the issues of frequent random reads and high I / O overhead in the prior art, but also significantly improves data recovery efficiency through logical block partitioning and sequential storage. Furthermore, the present invention improves the logic and organization of data storage, making data management more efficient and organized. Therefore, the present invention is significantly innovative in technology and can bring substantial advancements to the field of data backup and recovery.
[0115] The present invention addresses the problem of excessive random read overhead caused by random selection of basic data blocks in existing incremental compression methods. It creatively proposes a method for determining basic data blocks during the incremental compression process. Specifically, a similarity matrix between data blocks is introduced. The number of times each data block is selected as a basic data block is obtained based on the similarity matrix. Noise blocks are filtered using a threshold to remove noise blocks, thereby obtaining the basic data blocks for compression. Finally, the similarity matrix is used as an index to accelerate data block acquisition during the compression process, ensuring that the basic data blocks selected during incremental compression are those with a high degree of similarity to the data blocks to be compressed, thereby improving compression efficiency and reducing random read overhead. After determining the basic data blocks for the compression process based on the similarity between data blocks, the present invention creatively introduces the concept of a data block lifecycle. The start and end points of a data block's lifecycle are sorted by the backup data that first and last uses the data block in logical order, respectively. A layer of logical partitioning is then constructed on the data blocks. Consequently, when data is stored sequentially, from the backup data at the start point of its lifecycle to the backup data at the end point of its lifecycle, each backup data can be retrieved from the disk via a near-sequential read, alleviating the random read problem.
[0116] The data is selectively compressed based on the data compression ratio, which effectively reduces the compression cost.
[0117] The compressed data is divided reasonably to make it spatially local to ensure the performance of the data recovery process.
[0118] like Figure 8 As shown, an embodiment of the present invention further provides a backup data compression device, comprising:
[0119] A first acquisition module 801 is configured to acquire the similarity between any two data blocks in the backup data;
[0120] A first determining module 802 is configured to determine the number of records containing a basic data block in each data block according to the similarity;
[0121] A second determining module 803 is configured to determine a target data block to be compressed among the data blocks according to the number of records and a preset number threshold;
[0122] The second processing module 804 is configured to perform incremental compression on the target data block to obtain compressed backup data.
[0123] Optionally, the first acquisition module 801 includes:
[0124] A first processing unit is configured to perform a block operation on the backup data to obtain at least two initial data blocks;
[0125] a second processing unit, configured to perform fingerprint deduplication processing on each of the initial data blocks to obtain at least two deduplication-processed data blocks;
[0126] a third processing unit, configured to perform similarity calculation on the first data block and the second data block to obtain a similarity between any two of the data blocks;
[0127] The first data block is any one of the at least two deduplication-processed data blocks, and the second data block is a data block other than the first data block among the at least two deduplication-processed data blocks.
[0128] Optionally, the first determining module 802 includes:
[0129] a first acquiring unit, configured to acquire a first data block group whose similarity between any two data blocks is greater than a preset threshold, wherein the first data block group includes a third data block and a fourth data block corresponding to the third data block;
[0130] A first determination unit is used to determine the number of times that each data block includes a basic data block according to a first rule, wherein the first rule includes: determining the basic data blocks with the same data in the third data block and the fourth data block, and recording the basic data blocks in the third data block and the fourth data block once accordingly.
[0131] Optionally, the second determining module 803 includes:
[0132] A second determining unit, configured to determine the data block having the number of records greater than the preset number threshold as an initial compressed data block;
[0133] The fourth processing unit is configured to perform noise removal processing on the initial compressed data block to obtain the target data block to be compressed.
[0134] Optionally, the device further comprises:
[0135] A second acquisition module is configured to acquire, when the compressed backup data consists of at least two copies of backup data, a target data block after incremental compression in each of the compressed backup data;
[0136] a third processing module, configured to perform data recovery on the target data blocks after the incremental compression according to the first logical order corresponding to the target data blocks after the incremental compression, to obtain recovered backup data;
[0137] The first logical order is determined according to whether the target data block after incremental compression is compressed in each compressed backup data.
[0138] Optionally, the device further comprises:
[0139] A third acquisition module, configured to acquire a first compression sequence of the compressed backup data;
[0140] The fourth processing module is configured to determine the first logical order according to the first compression order and whether the target data block after the incremental compression is compressed in each compressed backup data.
[0141] Optionally, the third processing module includes:
[0142] a fifth processing unit, configured to determine, according to the first logical order, in the compressed target data blocks, a fifth data block that needs to be restored, corresponding to each compressed backup data, and a restoration order corresponding to the fifth data blocks;
[0143] The sixth processing unit is configured to restore the fifth data block in sequence according to the restoration sequence to obtain the restored backup data.
[0144] It should be noted that the backup data compression device provided in the embodiment of the present invention is a device capable of executing the above-mentioned backup data compression method. All embodiments of the above-mentioned backup data compression method are applicable to the device and can achieve the same or similar technical effects.
[0145] like Figure 9 As shown, an embodiment of the present invention also provides a backup data compression device, including: a processor 901; and a memory 903 connected to the processor 901 through a bus interface 902, the memory 903 is used to store programs and data used by the processor 901 when performing operations, and the processor 901 calls and executes the programs and data stored in the memory 903.
[0146] The transceiver 904 is connected to the bus interface 902 and is configured to receive and send data under the control of the processor 901. Specifically, the processor 901 is configured to read the program in the memory 903 and to perform the following process:
[0147] Obtaining similarity between any two data blocks in the backup data;
[0148] Determining the number of records of each data block including a basic data block according to the similarity;
[0149] Determining a target data block to be compressed in the data block according to the number of records and a preset number threshold;
[0150] Incremental compression is performed on the target data block to obtain compressed backup data.
[0151] Optionally, the processor 901 is configured to:
[0152] Performing a block operation on the backup data to obtain at least two initial data blocks;
[0153] Performing fingerprint deduplication processing on each of the initial data blocks to obtain at least two deduplication-processed data blocks;
[0154] Performing similarity calculation on the first data block and the second data block to obtain the similarity between any two of the data blocks;
[0155] The first data block is any one of the at least two deduplication-processed data blocks, and the second data block is a data block other than the first data block among the at least two deduplication-processed data blocks.
[0156] Optionally, the processor 901 is configured to:
[0157] Acquire a first data block group whose similarity between any two data blocks is greater than a preset threshold, wherein the first data block group includes a third data block and a fourth data block corresponding to the third data block;
[0158] According to the first rule, the number of times that the basic data block is included in each of the data blocks is determined, wherein the first rule includes: determining the basic data blocks with the same data in the third data block and the fourth data block, and recording the basic data blocks included in the third data block and the fourth data block once accordingly.
[0159] Optionally, the processor 901 is configured to:
[0160] Determining the data block whose recording times is greater than the preset times threshold as an initial compressed data block;
[0161] The initial compressed data block is subjected to noise removal processing to obtain the target data block to be compressed.
[0162] Optionally, the processor 901 is further configured to:
[0163] In a case where the compressed backup data consists of at least two copies of backup data, obtaining a target data block after incremental compression in each of the compressed backup data;
[0164] Performing data recovery on the target data blocks after incremental compression according to the first logical sequence corresponding to the target data blocks after incremental compression to obtain recovered backup data;
[0165] The first logical order is determined according to whether the target data block after incremental compression is compressed in each compressed backup data.
[0166] Optionally, the processor 901 is further configured to:
[0167] Obtaining a first compression sequence of the compressed backup data;
[0168] The first logical order is determined according to the first compression order and whether the target data block after the incremental compression is compressed in each compressed backup data.
[0169] Optionally, the processor 901 is specifically configured to:
[0170] determining, according to the first logical order, in the compressed target data blocks, a fifth data block that needs to be restored corresponding to each compressed backup data and a restoration order corresponding to the fifth data blocks;
[0171] According to the recovery sequence, data recovery is performed on the fifth data block in sequence to obtain the restored backup data.
[0172] Among them, Figure 9 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 901 and memory represented by memory 903. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and therefore will not be described further herein. The bus interface provides a user interface 905. The transceiver 904 may be a plurality of components, i.e., including a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium. The processor 901 is responsible for managing the bus architecture and general processing, and the memory 903 may store data used by the processor 901 when performing operations.
[0173] In addition, a specific embodiment of the present invention further provides a readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps in any one of the backup data compression methods described above.
[0174] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection of some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0175] In addition, the functional units in various embodiments of the present invention may be integrated into a single processing unit, each unit may be physically included separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional units.
[0176] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to perform some steps of the resource selection method described in various embodiments of the present invention, or to perform some steps of the information sending method described in various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, and other media that can store program code.
[0177] A specific embodiment of the present invention further provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 2 The various processes of the method embodiment shown can achieve the same technical effect, and to avoid repetition, they will not be described here.
[0178] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary personnel in this technical field, several improvements and modifications can be made without departing from the principles described in the present invention. These improvements and modifications are also within the scope of protection of the present invention.
Claims
1. A backup data compression method, characterized in that: include: Obtaining similarity between any two data blocks in the backup data; Determining the number of records of each data block including a basic data block according to the similarity; Determining a target data block to be compressed in the data block according to the number of records and a preset number threshold; Incremental compression is performed on the target data block to obtain compressed backup data.
2. The method according to claim 1, characterized in that Obtaining the similarity between any two data blocks in the backup data includes: Performing a block operation on the backup data to obtain at least two initial data blocks; Performing fingerprint deduplication processing on each of the initial data blocks to obtain at least two deduplication-processed data blocks; Performing similarity calculation on the first data block and the second data block to obtain the similarity between any two of the data blocks; The first data block is any one of the at least two deduplication-processed data blocks, and the second data block is a data block other than the first data block among the at least two deduplication-processed data blocks.
3. The method according to claim 1, characterized in that Determining the number of records of the basic data block included in each of the data blocks according to the similarity includes: Acquire a first data block group whose similarity between any two data blocks is greater than a preset threshold, wherein the first data block group includes a third data block and a fourth data block corresponding to the third data block; According to the first rule, the number of times that the basic data block is included in each of the data blocks is determined, wherein the first rule includes: determining the basic data blocks with the same data in the third data block and the fourth data block, and recording the basic data blocks included in the third data block and the fourth data block once accordingly.
4. The method according to claim 1, wherein Determining a target data block to be compressed in the data blocks according to the number of records and a preset number threshold includes: Determining the data block whose recording times is greater than the preset times threshold as an initial compressed data block; The initial compressed data block is subjected to noise removal processing to obtain the target data block to be compressed.
5. The method according to claim 1, wherein The method further comprises: In a case where the compressed backup data consists of at least two copies of backup data, obtaining a target data block after incremental compression in each of the compressed backup data; Performing data recovery on the target data blocks after incremental compression according to the first logical sequence corresponding to the target data blocks after incremental compression to obtain recovered backup data; The first logical order is determined according to whether the target data block after incremental compression is compressed in each compressed backup data.
6. The method according to claim 5, characterized in that The method further comprises: Obtaining a first compression sequence of the compressed backup data; The first logical order is determined according to the first compression order and whether the target data block after the incremental compression is compressed in each compressed backup data.
7. The method according to claim 5, characterized in that Performing data recovery on the incrementally compressed target data block according to the first logical order corresponding to the incrementally compressed target data block to obtain recovered backup data includes: determining, according to the first logical order, in the compressed target data blocks, a fifth data block that needs to be restored corresponding to each compressed backup data and a restoration order corresponding to the fifth data blocks; According to the recovery sequence, data recovery is performed on the fifth data block in sequence to obtain the restored backup data.
8. A backup data compression device, characterized in that: include: A first acquisition module is used to acquire the similarity between any two data blocks in the backup data; A first determining module, configured to determine the number of records containing a basic data block in each of the data blocks according to the similarity; A second determining module is configured to determine a target data block to be compressed among the data blocks according to the number of records and a preset number threshold; The first processing module is configured to perform incremental compression on the target data block to obtain compressed backup data.
9. A backup data compression device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the backup data compression method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program, and when the program is executed by a processor, the steps of the backup data compression method according to any one of claims 1 to 7 are implemented.
11. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps in the backup data compression method according to any one of claims 1 to 7.