Similar data detection method and device

Through the multi-level super feature method and false positive test mechanism, the dilemma between coverage and accuracy of N-Transform super feature method is solved, and the high coverage and high accuracy of similar data matching is achieved, which improves the compression rate and system efficiency of differential compression.

CN120277450APending Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410029222.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing N-Transform super feature method cannot take into account both coverage and accuracy in similar data detection, which makes it difficult to further improve the compression rate of differential compression.

Method used

The multi-level super feature method is used to utilize different levels of super features with different similarity detection thresholds. High-level super features are used to match high-similar data blocks, and low-level super features are used to match potential similar blocks. Combined with the false positive test mechanism, we ensure high coverage and high accuracy of the matching results.

Benefits of technology

The compression rate of differential compression is improved, high coverage and high accuracy of similar data matching are achieved, and the efficiency of data storage, transmission and distributed storage systems is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277450A_ABST
    Figure CN120277450A_ABST
Patent Text Reader

Abstract

The invention provides a similar data detection method and device, and the method comprises the steps: extracting a multi-level hyper-feature of a target data block, the multi-level hyper-feature comprises a plurality of levels of hyper-features, and the similarity detection threshold value of the high-level hyper-feature is higher than the similarity detection threshold value of the low-level hyper-feature; based on the multi-level hyper-features, similarity matching is carried out from a hyper-feature data set to obtain a matching result, and the hyper-feature data set comprises the multi-level hyper-features corresponding to the data blocks in the data block set; and based on a matching result, determining a similar data block corresponding to the target data block. According to the similar data detection method provided by the invention, similar matching is carried out through the multi-level hyper-features with different similar matching thresholds, the coverage rate and the accuracy rate of similar data matching are considered, and high coverage rate and high accuracy rate of similar data matching are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a similar data detection method and device. Background Art

[0002] Most of the current mainstream differential compression technologies use the N-Transform super-feature method derived from Broder's theorem to find similar blocks. Super-features are essentially a single-threshold filter that can only determine whether two data blocks are similar, and cannot compare which of multiple candidate reference blocks is more similar to the current data block to be compressed. Therefore, if the screening threshold is set high, potential similar blocks will be missed and the coverage rate will be reduced; if the screening threshold is set low, reference blocks with lower similarity may be matched, the accuracy rate will be reduced, and the data reduction rate will be reduced. No matter how the threshold is set, some coverage or accuracy will be lost, which will cause the compression rate of differential compression to encounter a bottleneck and be difficult to further improve. Therefore, the N-Transform super-feature method faces the problem of not being able to balance coverage and accuracy. Summary of the invention

[0003] The embodiments of the present application provide a similar data detection method, which performs similarity matching by utilizing multi-level super features with different similarity matching thresholds, takes into account both the coverage and accuracy of similar data matching, and achieves high coverage and high accuracy of similar data matching.

[0004] In a first aspect, the present application provides a similar data detection method, including extracting a multi-level super feature of a target data block, the multi-level super feature including multiple levels of super features, wherein a similarity detection threshold of a high-level super feature is higher than a similarity detection threshold of a low-level super feature; based on the multi-level super feature, performing similarity matching from a super feature data set to obtain a matching result, wherein the super feature data set includes the multi-level super features corresponding to each data block in a data block set; based on the matching result, determining a similar data block corresponding to the target data block.

[0005] The similar data detection method provided by the present application performs similar matching by using multi-level super features with different similar matching thresholds. The high-level super features in the multi-level super features have a higher similar matching threshold to ensure the high accuracy of the matching data, and the low-level super features have a lower similar matching threshold to ensure the coverage of the matching data, thereby achieving both the coverage and accuracy of similar data matching, and then achieving high coverage and high accuracy of similar data matching. Coverage refers to how many data blocks can find a highly similar reference block. Accuracy refers to whether the reference blocks found are similar enough to provide a higher data reduction rate benefit.

[0006] In a possible implementation, based on multi-level super features, a specific implementation of obtaining a matching result by matching from a super feature dataset is as follows: in the order from the highest to the lowest super feature level of the multi-level super features, match the multi-level super features of the target data block and the multi-level super features of the data block to be matched level by level, where the data block to be matched is any data block in the data block set; if the super features at any level are successfully matched, it is determined that the target data block and the data block to be matched are successfully matched.

[0007] In other words, when the target data block performs similarity matching from the data block set, it starts matching from the highest-level super features. If the highest-level super features are successfully matched, it is determined that the super features of the two data blocks (i.e., the target data block and the data block to be matched) are successfully matched, and then it is determined that the two data blocks are successfully matched. If no matching data block is found from the data block set through the highest-level super features, continue to match through the second-highest-level super features, and so on, until a data block is matched or the lowest-level super features are matched.

[0008] In this application, by matching in the order from the highest to the lowest of the multi-level super features, it is ensured that data blocks with high similarity detection thresholds are matched first, achieving high accuracy of the matching data; when no data block is matched from the data block set by the high-level super features, then use the low-level super features for matching. The low-level super features have lower similarity thresholds, ensuring that data blocks with low similarity thresholds are covered by the matching and no potential similar blocks are missed, thereby ensuring high coverage of the matching of similar data.

[0009] In this possible implementation, each level of the multi-level super features includes multiple super feature values; if the target data block and the data block to be matched have at least one same super feature value in the super features at any level, it is determined that the super features at any level are successfully matched. In the specific matching process, according to the first-fit principle, for the super features at the same level, select the first data block that has at least one common super feature value with the target data block as the successfully matched data block.

[0010] In another possible implementation, a specific implementation of determining the similar data block corresponding to the target data block based on the matching result is as follows: determine the data block obtained by matching (i.e., the data block to be matched that is successfully matched with the target data block) as the similar data block corresponding to the target data block.

[0011] In another possible implementation, another specific implementation of determining the similar data block corresponding to the target data block based on the matching result is as follows: Based on the target compression ratio and the compression ratio threshold, the matched data block is inspected, where the target compression ratio indicates the compression ratio corresponding to performing differential compression on the target data block with the matched data block as the reference block; if the inspection passes, the matched data block is determined as the similar data block corresponding to the target data block.

[0012] Since this application uses multi-level super features for similar data matching, the matched data block may be obtained by matching lower-level super features in the multi-level super features. Because the similarity detection threshold of the lower-level super features is relatively low, there may be a situation where the similarity between the matched data block and the target data block does not meet the standard. Therefore, after the data block is matched, this application will further inspect the matched data block to check whether it meets the standard. If the inspection passes, the matched data block is determined as the similar data block corresponding to the target data block. If the inspection fails, the matched data block is discarded and the matching continues. Ensure that the finally obtained similar data blocks meet the requirements.

[0013] In one example, a specific implementation of inspecting the matched data block based on the target compression ratio and the compression ratio threshold is to check whether the target compression ratio is greater than the compression ratio threshold. If so, it is determined that the inspection passes; if not, it is determined that the inspection fails.

[0014] Optionally, the compression ratio threshold can be a preset fixed compression ratio threshold. For example, the developer determines the compression ratio threshold based on experience. If it is judged according to experience that the compression ratio of differential compression is generally above 0.3, the compression ratio threshold is set to 0.3. If the target compression ratio is less than or equal to 0.3, it is determined that the matched data block is a false positive, the similarity does not meet the requirements, and the inspection fails. If the target compression ratio is greater than 0.3, it is determined that the matched data block is a true positive, the similarity meets the requirements, and the inspection passes.

[0015] Optionally, the compression ratio threshold can also be a dynamic compression ratio threshold, which is determined based on the compression ratio of historical lossless compression. For example, the compression ratio threshold is determined based on the average compression ratio of the compression ratios corresponding to the previous L lossless compression operations (i.e., the L lossless compression operations performed before the current time), and L is a positive integer greater than 1.

[0016] It can be understood that most storage systems or data transmission systems will have lossless compression (for example, common differential compression coding algorithms include a local lossless compression step to compress incremental blocks). For those with lossless compression records, the above inspection mechanism can be adopted. If there is extremely rare situation where a storage system or data transmission does not have lossless compression, lossless compression can be performed on data blocks at fixed periodic intervals to ensure that the system has lossless compression records.

[0017] By comparing the compression ratio corresponding to the differential compression of the matched data block and the target data block with the average compression ratio corresponding to the first L lossless compressions, it is determined whether the overall compression ratio of the data is improved, and further whether the similarity between the matched data block and the target data block meets the requirements, so as to realize the inspection of the matched data block with relatively small computational effort.

[0018] In another possible implementation, a specific implementation of extracting the multi-level super features corresponding to the target data block is to extract N eigenvalue of the target data block, where N is a positive integer greater than 1; based on multiple different super feature extraction parameters, perform super feature extraction on the N eigenvalue to obtain multi-level super features; wherein, the super feature extraction parameters indicate the number of groups of the N features and the number of eigenvalue in each group.

[0019] In another possible implementation, the specific implementation of extracting N eigenvalue of the target data block is to sample the target data block to obtain a sampled data block; perform feature extraction on the sampled data block through N different sliding hash functions to obtain N different eigenvalue sets; based on the smallest eigenvalue in each eigenvalue set among the N eigenvalue sets, determine the N eigenvalue of the target data block.

[0020] In another possible implementation, the similar data detection method provided by this application further includes when the target data block does not match a similar data block from the data block set, saving the multi-level super features corresponding to the target data block to the super feature data set to expand the super feature data set, which is beneficial to the subsequent similar detection of data blocks.

[0021] However, it should be noted that for the similar data block detection method provided by this application, the super features of incremental blocks are not stored in the super feature database, nor are such incremental blocks selected as reference blocks. The reason is that the resulting chain references will lead to serious data fragmentation and bring a large number of I / O operations when backing up / restoring incremental blocks.

[0022] The similar data detection method provided by this application can be applied to any scenario that requires similar data detection, such as similar data detection in data compression scenarios, similar data detection in data transmission scenarios, and similar data detection in data distributed storage scenarios, etc. This application does not make specific limitations on the application scenarios of similar data detection.

[0023] In a second aspect, the present application further provides a data compression method, including performing similar data detection on a stored data block set based on the similar data detection method described in the first aspect to obtain a similar data block corresponding to the data block to be compressed; using the similar data block as a reference block to perform differential compression on the data block to be compressed.

[0024] By performing similar data detection using the similar data detection method provided in the first aspect of the present application, taking into account both the coverage rate and the accuracy rate, high coverage rate and high accuracy rate of reference block matching are achieved, thereby improving the compression rate of differential compression.

[0025] In a third aspect, the present application further provides a data transmission method, including performing similar data detection on multiple data blocks to be transmitted based on the similar data detection method described in the first aspect to obtain a set of similar data in the multiple data blocks; performing differential compression on the set of similar data to obtain a compressed data set; and transmitting the compressed data set.

[0026] By performing similar data detection using the similar data detection method provided in the first aspect of the present application, the coverage rate and accuracy rate of reference block matching are improved, thereby providing the compression rate of differential compression. Then, data transmission is performed using the highly compressed data, improving the transmission efficiency of data transmission and saving bandwidth.

[0027] In a fourth aspect, the present application further provides a data storage method, applied to a distributed storage system. The distributed storage system includes multiple storage nodes. The data storage method includes performing similar data detection on the data in each storage node among the multiple storage nodes based on the similar data detection method described in the first aspect; determining a target storage node from the multiple storage nodes based on the similar data detection result, where the target storage node stores a data block similar to the data block to be stored; and storing the data block to be stored in the target storage node.

[0028] By performing similar data detection using the similar data detection method provided in the first aspect of the present application, the coverage rate and accuracy rate of similar data matching are improved, thereby improving the storage efficiency of distributed storage.

[0029] Fifth aspect, the present application further provides a similar data detection device, which includes a feature extraction module, a matching module, and a determination module. The feature extraction module is used to extract multi-level super features of the target data block. The multi-level super features include super features of multiple levels, and the similarity detection threshold of the super features of the higher level is higher than that of the super features of the lower level. The matching module is used to perform similarity matching from the super feature dataset based on the multi-level super features to obtain a matching result. The super feature dataset includes the multi-level super features corresponding to each data block in the data block set. The determination module is used to determine the similar data block corresponding to the target data block based on the matching result.

[0030] In another possible implementation, the matching module is specifically configured to match the multi-level super features of the target data block and the multi-level super features of the data block to be matched level by level in the order from the highest to the lowest super feature level of the multi-level super features. The data block to be matched is any data block in the data block set. If the super features of any level are successfully matched, it is determined that the target data block and the data block to be matched are successfully matched.

[0031] In another possible implementation, each level of super features in the multi-level super features includes multiple super feature values. If the target data block and the data block to be matched have at least one same super feature value in the super features of any level, it is determined that the super features of any level are successfully matched.

[0032] In another possible implementation, the determination module is specifically configured to determine the data block obtained by matching as the similar data block corresponding to the target data block.

[0033] In another possible implementation, the determination module is specifically configured to check the data block obtained by matching based on the target compression ratio and the compression ratio threshold. The target compression ratio indicates the compression ratio corresponding to performing differential compression on the reference data block and the target data block. The reference data block is the data block to be matched that is successfully matched with the target data block. If the check passes, the data block obtained by matching is determined as the similar data block corresponding to the target data block.

[0034] In another possible implementation, a specific implementation of checking the data block obtained by matching based on the target compression ratio and the compression ratio threshold is: checking whether the target compression ratio is greater than the compression ratio threshold. If so, it is determined that the check passes; if not, it is determined that the check fails.

[0035] Optionally, the target data block will be subjected to differential compression or lossless compression operations before being transmitted or stored. The compression ratio threshold is determined based on the compression ratio of historical lossless compression operations. Exemplarily, the compression ratio of historical lossless compression operations is determined based on the average compression ratio of the compression ratios corresponding to the previous L lossless compression operations, where L is a positive integer greater than 1.

[0036] In another possible implementation, the feature extraction module is specifically configured to extract N eigenvalue of the target data block, where N is a positive integer greater than 1; perform hyper-feature extraction on the N eigenvalues based on multiple different hyper-feature extraction parameters to obtain hyper-features at multiple levels; where the hyper-feature extraction parameters indicate the number of groups of the N features and the number of eigenvalues in each group.

[0037] In another possible implementation, a specific implementation of extracting N eigenvalues of the target data block is to sample the target data block to obtain a sampled data block; perform feature extraction on the sampled data block through N different sliding hash functions to obtain N different eigenvalue sets; determine the N eigenvalues of the target data block based on the smallest eigenvalue in each eigenvalue set of the N eigenvalue sets.

[0038] In another possible implementation, the similar data detection device provided in this application further includes a storage module, which is configured to save the multi-level hyper-features corresponding to the target data block to the hyper-feature dataset when no similar data block is matched from the data block set for the target data block.

[0039] In a sixth aspect, this application provides a data compression device, which includes a similar data detection module and a compression module. Among them, the similar data detection module is configured to perform similar data detection from the stored data block set based on the similar data detection method described in the first aspect to obtain a similar data block corresponding to the data block to be compressed; the compression module is configured to perform differential compression on the data block to be compressed with the similar data block as the reference block.

[0040] In a seventh aspect, this application provides a data transmission device, which includes a similar data detection module, a compression module, and a transmission module. Among them, the similar data detection module is configured to perform similar data detection from multiple data blocks to be transmitted based on the similar data detection method described in the first aspect to obtain a set of similar data in the multiple data blocks; the compression module is configured to perform differential compression on the set of similar data to obtain a compressed data set; the transmission module is configured to perform data transmission on the compressed data set.

[0041] In an eighth aspect, this application provides a data storage device, which can be applied to a distributed storage system. The distributed storage system includes multiple storage nodes. The data storage device provided in this application includes a similar data detection module, a determination module, and a storage module. Among them, the similar data detection module is configured to perform similar data detection on the data in each storage node among the multiple storage nodes based on the similar data detection method described in the first aspect; the determination module is configured to determine a target storage node from the multiple storage nodes based on the similar data detection result, and the target storage node stores a data block similar to the data block to be stored; the storage module is configured to store the data block to be stored to the target storage node.

[0042] In a ninth aspect, an embodiment of the present application provides a computing device, including a memory and a processor. Instructions are stored in the memory. When the instructions are executed by the processor, the methods described in the first aspect and / or the second aspect and / or the third aspect and / or the fourth aspect are implemented.

[0043] In a tenth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods described in the first aspect and / or the second aspect and / or the third aspect and / or the fourth aspect are implemented.

[0044] In an eleventh aspect, an embodiment of the present application further provides a computer program or a computer program product. The computer program or the computer program product includes instructions. When the instructions are executed, the computer is made to execute the methods described in the first aspect and / or the second aspect and / or the third aspect and / or the fourth aspect.

[0045] In a twelfth aspect, an embodiment of the present application further provides a chip, including at least one processor and a communication interface. The processor is configured to execute the methods described in the first aspect and / or the second aspect and / or the third aspect and / or the fourth aspect. Description of the Drawings

[0046] Figure 1 A schematic diagram showing the implementation process of an N-Transform super feature method is shown;

[0047] Figure 2 A schematic diagram showing the storage process of a storage system to which the similar data detection method provided by the embodiment of the present application can be applied is shown;

[0048] Figure 3 A flowchart showing a method for detecting similar data provided by an embodiment of the present application;

[0049] Figure 4 A specific implementation flowchart showing the matching of similar data blocks provided by an embodiment of the present application is shown;

[0050] Figure 5 A schematic diagram showing the structure of a similar data detection device provided by an embodiment of the present application;

[0051] Figure 6 A flowchart showing a data compression method is shown;

[0052] Figure 7 A schematic diagram showing the structure of a data compression device provided by an embodiment of the present application;

[0053] Figure 8 A flowchart showing a data transmission method provided by an embodiment of the present application is shown;

[0054] Figure 9 It is a schematic structural diagram of a data transmission device provided by an embodiment of the present application;

[0055] Figure 10 It shows a schematic structural diagram of a distributed storage system;

[0056] Figure 11 It shows a schematic flowchart of a data storage method;

[0057] Figure 12 It is a schematic structural diagram of a data storage device provided by an embodiment of the present application;

[0058] Figure 13 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0059] The term "and / or" mentioned in this article is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0060] The terms "first", "second", etc. in the description and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects. For example, the first-level super feature and the second-level super feature are used to distinguish different super features, rather than to describe a specific order of super features.

[0061] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0062] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements.

[0063] To facilitate understanding of the solutions of the embodiments of the present application, the technical terms involved in this article are first explained below.

[0064] Hash: A mapping algorithm that maps a larger input data to a smaller hash value of a unified specification. Theoretically, the probability that different input data generate the same hash value is very low.

[0065] Delta compression: A compression technique that, for the current data block, first attempts to find a highly similar data block among the existing data blocks through similar fingerprints, and then only saves the difference part of the current data block relative to this highly similar data block.

[0066] Similar fingerprint (sketch, SK): A special mark generated for each data block. The more similar the data blocks are, the greater the probability of generating the same similar fingerprint.

[0067] Reference block: An existing data block that is highly similar to the current data block and is used in delta compression.

[0068] Delta block: The difference part that needs to be saved for the current data block after performing delta compression.

[0069] Sliding window: For a data sequence, the interval from the i-th byte to the (i + w)-th byte is called a "window" with a width of w. The process of moving the left endpoint of the window from the first byte to the end of the data sequence is called "sliding", and all the windows experienced during this period are called a set of "sliding windows".

[0070] Feature: A value of a specific specification calculated for the original data block. Similar data blocks have a relatively high probability of obtaining the same feature.

[0071] Super feature (SF): The hash value calculated for a set of several features. If the data blocks are not highly similar, the probability of generating the same super feature is extremely low.

[0072] Hash table: An efficient data structure that can quickly query whether there is an object with the same hash value as the current object.

[0073] The similar data detection method and device provided by the embodiments of the present application can be applied to any scenario or solution that requires similar data detection, and can achieve both coverage and accuracy when performing similar data detection, and at the same time achieve high coverage and high accuracy of similar data matching.

[0074] For example, the similar data detection method and device provided by the embodiments of the present application can be applied to data compression scenarios, data transmission scenarios, data distributed storage scenarios, etc. The present application does not specifically limit the application scenarios of similar data detection.

[0075] Taking the similar data detection method and device provided by the embodiments of the present application applied to the data compression scenario as an example, the detailed implementation of the similar data detection method and device provided by the embodiments of the present application will be introduced in detail below.

[0076] In a storage system, how to efficiently reduce duplicate and redundant data is a key issue determining the storage efficiency of the entire system. Currently, there are mainly three types of data reduction techniques, including data deduplication, lossless compression, and differential compression.

[0077] Data deduplication: Calculate a fingerprint (FP) for each data block through a strong hashing method. If two data blocks have the same fingerprint, it means they are duplicate data blocks and do not need to be written to the storage device. Using a hash table can achieve efficient query and comparison of fingerprints, thereby efficiently implementing data deduplication. The disadvantage is that it can only handle exactly the same data blocks. If there are minor data modifications, the data deduplication technique cannot be applied.

[0078] Lossless compression: Lossless compression is a general term for various common compression methods, such as compression algorithms like RAR, ZIP, 7Z, ZSTD, etc. By analyzing the repeated data in the data string to build an index dictionary, and then compressing and encoding the data. It can be applied to data in any scenario. The disadvantage is that it can only compress and encode local data, cannot find and reduce duplicate data globally, and the compression ratio is usually low.

[0079] Differential compression: A technique for compressing similar data blocks. For the current data block, first try to find a highly similar data block among the existing data blocks, and then only save the difference part of the current data block relative to this highly similar data block, thus greatly reducing the amount of data to be stored. The key challenge in designing the differential compression technique lies in how to efficiently and accurately find similar data blocks. The more similar blocks are identified, the more data can be differentially compressed; the higher the similarity of the identified similar blocks, the higher the data reduction amount of the differential compression operation.

[0080] In differential compression, the process of finding similar blocks is called similarity detection. Most of the existing similarity feature calculation methods are based on the classic Broder theory. Broder described the similarity problem as the intersection and union problem in a set and proposed MinHash as the theoretical basis for estimating the similarity of sets. To quickly evaluate the similarity of two sets, Broder proposed the "Broder theorem": Suppose A and B are two sets, and F(A) and F(B) are the corresponding sets obtained by calculating the elements in A and B with the MinHash function F. If min(S) represents the smallest element in the set S, it can be expressed by the formula:

[0081]

[0082] The Broder theory describes that the probability that two sets have the same minimum hash elements is equal to the Jaccard coefficient of the two sets. If the elements in sets A and B are composed of all the shingles (dividing a file into overlapping fixed-length consecutive strings, called shingles) in two data blocks C1 and C2, according to the Broder theory, the probability that F(A) and F(B) have the same minimum element is equal to the similarity of the two data blocks.

[0083] The currently commonly used feature detection method based on the Broder theory is the N-Transform super-feature method. First, N-Transform calculates multiple hash functions and extracts multiple feature values. Then, super-feature combines multiple feature values into a super feature value for similarity matching. As long as there is a match of one same super feature value between two data blocks, it means they are similar. Finally, differential encoding is performed on this pair of similar data blocks.

[0084] Differential encoding is a compression technique derived from dictionary encoding. Different from dictionary encoding, differential compression does not encode within the compression window of a single object, but encodes between two objects. The compression process of the LZ77 compression algorithm can be simply understood as: if the subsequent string is the same as the previously processed string, find the longest match between them and replace it with the COPY command. The currently most commonly used differential compression tools, such as vcdiff, Xdelta, and Zdelta, etc., can be simply considered as performing LZ77 compression on two similar files. Suppose there are two similar data blocks A and B, and B is the "target block" to be compressed, then A is called the "reference block" of B. Differential compression will find the content that exists in B but not in A and write it into a file, and the process of generating this file is called differential encoding. Specifically, differential encoding will use the sliding window technique to simultaneously detect the content in data blocks A and B. For the content in B, if the same content exists in A, the COPY command is used for encoding, otherwise the INSERT command is used for encoding. The COPY command contains the offset and length of the repeated content in data block A, while the INSERT command contains the length and the content itself of the non-repeated content in data block B. The file generated containing the COPY and INSERT commands is called the differential block. The size of the differential block is generally smaller than data block B. When storing, the differential block will be saved instead of data block B, thus reducing the storage overhead. The rate of differential encoding and the size of the differential block are both related to the similarity of data blocks A and B. The higher the similarity, the faster the differential encoding speed and the smaller the differential block. Differential decoding is to regenerate data block B according to the differential block and data block A according to the commands.

[0085] In the differential compression schemes in the related art, most of them use the N-Transform super feature method derived from Broder's theorem to find similar blocks. First, a sliding window hash is performed on a given data block using a hash function, and the minimum value among all the hash values obtained is called a feature. Broder's theorem states that for two data blocks A and B, the probability that the features obtained by the same hash function are the same is equal to the set similarity of the two data blocks. Therefore, by calculating N features for two data blocks using N hash functions (usually obtained by performing N linear transformations on an original hash function), the similarity between the two data blocks can be efficiently estimated by comparing the number of identical features. The formula for feature calculation is as follows, where m i and a i are a set of random numbers:

[0086] h i (fp) = (m i *fp + a i ) mod 2^32

[0087] In a large-scale storage system, calculating the similarity between a certain data block and all other data blocks and selecting the data block with the highest similarity incurs a large computational overhead. In most cases, finding relatively similar data blocks can meet the requirements. Therefore, Broder further proposed the N-Transform super feature method to optimize the process of finding reference blocks. For the existing N features, Super-feature combines multiple of these feature values into a single super feature value for similarity matching. As long as two data blocks have a match for one of the same super feature values, it means they are similar. The strategy of combining multiple feature values into a super feature value avoids the calculation of specific similarities and greatly reduces the computational overhead in the matching process.

[0088] Figure 1 shows a schematic diagram of the implementation process of an N-Transform super feature method.

[0089] The original N-Transform super feature method has the problem that the computational overhead required to calculate N features is large, which slows down the overall speed of the system. In one solution, a sampling method can be used to accelerate the counting of feature calculations. First, a fixed sampling function is used for each data block to generate a smaller sample, and then the N-Transform super feature method is run on the small sample, greatly improving the execution efficiency of this method and making feature calculation no longer a performance bottleneck.

[0090] In the related art, the main problem faced in detecting similar data using the N-Transform superfeature method is that it is caught in a dilemma between coverage rate and accuracy rate, thus limiting the further improvement of the compression rate. The coverage rate refers to how many data blocks can find a highly similar reference block. The accuracy rate refers to the fact that the found reference block is similar enough to provide a high data reduction rate benefit. Essentially, a superfeature is a single-threshold filter that can only judge whether two data blocks are "similar" and cannot compare which one of multiple alternative reference blocks is more similar to the current data block to be compressed. Therefore, if the screening threshold is set high, potential similar blocks will be missed and the coverage rate will decrease (for example, 12 features form a superfeature, and all feature values must be exactly the same to be judged as similar); if the screening threshold is set low, reference blocks with relatively low similarity may be matched, the accuracy rate will decrease, and the data reduction rate will decline (for example, 12 feature values form 4 or 6 superfeatures, and as long as one feature is the same, it can be judged as similar). No matter how the threshold is set, a certain part of the coverage rate or accuracy rate will definitely be lost, which will lead to a bottleneck in the compression rate of differential compression and it is difficult to further improve it.

[0091] In view of the problems existing in similar data detection in the related art, the present application provides a method for detecting similar data. By extracting multi-level superfeatures for data blocks, different levels of superfeatures have different similarity detection thresholds, and multi-level superfeatures are used for similar data matching. The high-level superfeatures have a higher similarity detection threshold and are used to match data blocks with high similarity to ensure the accuracy of the matched data. The low-level superfeatures have a lower similarity detection threshold and are used to match as many potential similar data blocks as possible, simultaneously achieving high coverage rate and high accuracy rate of reference block matching.

[0092] The following details the specific implementation of the method for detecting similar data provided by the embodiments of the present application through the accompanying drawings.

[0093] Figure 2 The storage process schematic diagram of a storage system to which the method for detecting similar data provided by the embodiments of the present application can be applied is shown. As Figure 2 shown, when the storage system stores data, for example, the data to be stored includes Figure 2For the data blocks A, B, A and B' in it, first perform data deduplication on the data to be stored, and delete the duplicate data in the data to be stored. For example, after fingerprint comparison, it is found that data block A is a duplicate data block, then data block A is deleted, and only data blocks A, B and B' are stored. Then, perform differential compression on the data to be stored after the data deduplication operation. A necessary step in differential compression is similar data detection. Similar data blocks are detected through similar data detection. For example, after similar data detection, it is detected that data blocks B and B' in the data to be stored after the deduplication operation are similar data. Then, when storing, differential compression is performed on data blocks B and B', and the differential block △B' of data blocks B and B' is obtained. When storing, only the differential block △B' is stored, and data block B' is not stored. Finally, data blocks A, B and differential block △B' are stored in the storage medium, such as a disk, after lossless compression.

[0094] The similar data detection method provided by the embodiments of the present application can be applied to the similar data matching link in the differential compression stage. In the related art, a single-threshold similar fingerprint is relied on to query similar data. The similar data detection method provided by the embodiments of the present application uses a multi-level super feature query to replace the traditional query scheme, and improves the quality of the similar data query result without affecting the work of other components of the system, thereby improving the compression effect of differential compression and improving the data storage efficiency.

[0095] Figure 3 It is a schematic flowchart of a similar data detection method provided by the embodiments of the present application. The similar data detection method provided by the embodiments of the present application can be executed by any computing device, device, platform or device cluster that needs to perform similar data detection. For example, this method can be executed by a storage server. When storing data, the storage server detects the corresponding similar data blocks of the data blocks to be stored through the similar data detection method provided by the embodiments of the present application, realizes high coverage and high accuracy of similar data detection, and thereby improves the compression ratio of differential compression. As Figure 3 shown, the similar data detection method provided by the embodiments of the present application at least includes steps S301 to S303.

[0096] In step S301, extract the multi-level super features corresponding to the target data block.

[0097] Taking the data compression scenario as an example, the target data block is the data block to be compressed. Perform multi-level super feature extraction on the data block to be compressed to obtain the multi-level super features corresponding to the data block to be compressed. The following takes the data block to be compressed as data block A as an example to introduce the specific implementation of multi-level super feature extraction for the data block to be compressed.

[0098] First, use N different hash functions (for example, N hash functions obtained by performing N linear transformations on a single hash function) to extract features from data block A respectively, obtaining N feature values of data block A (which can be called original feature values). For example, the feature extraction of data block A by each hash function is as follows: the hash value set of data block A is extracted by means of sliding hash, and then the minimum hash value in the hash value set is selected as the feature value of data block A.

[0099] Then, use different hyper - feature extraction parameters to perform hyper - feature extraction on the N feature values of data block A respectively, obtaining multi - level hyper - features of data block A. The hyper - feature extraction parameters include k and s, where k represents dividing the N feature values into k groups, and s represents the number of feature values in each group.

[0100] Exemplarily, by default, set N = 12, and (k; s) are (3; 4), (4; 3), and (6; 2) respectively to extract three - level hyper - features of data block A. The extraction parameter of the first - level hyper - features (k; s)=(3; 4). According to the hyper - feature extraction parameter (k; s)=(3; 4), the 12 feature values of data block A are divided into 3 groups, with 4 feature values in each group. Then, hash calculation is performed on the feature values in each group, and the hyper - feature values of each group of feature values are extracted, obtaining 3 hyper - feature values, which are the first - level hyper - features. Similarly, the second - level hyper - features and the third - level hyper - features of data block A are extracted by using the hyper - feature extraction parameters (k; s)=(4; 3) and (k; s)=(6; 2) respectively.

[0101] Different super feature extraction parameters determine the similarity detection thresholds of different levels of super features. The similarity detection threshold of higher-level super features is higher than that of lower-level super features. That is, the similarity detection threshold of the first-level super features is greater than that of the second-level super features, and the similarity detection threshold of the second-level super features is greater than that of the third-level super features. In other words, the similarity of the data blocks obtained by matching higher-level super features is greater than that of the data blocks obtained by matching lower-level super features. For example, the super feature extraction parameters of the first-level super features are (k; s) = (3; 4), and the extracted super features include three super feature values, each of which is obtained by extracting four feature values. When matching through the first-level super features, it is necessary that at least one super feature value of the first-level super features of the data block to be matched and the first-level super features of data block A is the same for successful matching. That is to say, it is required that at least four consecutive feature values are the same to determine that the data block to be matched and data block A are successfully matched. The super feature extraction parameters of the second-level super features are (k; s) = (4; 3), and the extracted super features include four super feature values, each of which is obtained by extracting three feature values. When matching through the second-level super features, it is necessary that at least one super feature value of the second-level super features of the data block to be matched and the second-level super features of data block A is the same for successful matching. That is to say, it is required that at least three consecutive feature values are the same to determine that the data block to be matched and data block A are successfully matched. It can be seen that the data block matched by the first-level super features and data block A have at least 4 consecutive identical feature values, while the data block matched by the second-level super features and data block A have at least 3 consecutive identical feature values. Therefore, the similarity between the data block matched by the first-level super features and data block A must be greater than the similarity between the data block matched by the second-level super features and data block A.

[0102] From the above description, it can be concluded that the similarity detection threshold depends on the super feature extraction parameters k and s. Assuming that the similarity of two data blocks is p, the probability that at least one super feature value of the two data blocks is the same is:

[0103]

[0104] Of course, the number of feature values N of data block A and the super feature extraction parameters (k; s) can also be set to other values according to actual needs. For example, N = 12, and (k; s) are (2; 6), (3; 4), (4; 3), and (6; 2) respectively, to extract four levels of super features of data block A. Another example is that N = 8, and (k; s) are (2; 4) and (4; 2) respectively, to extract two levels of super features of data block A. There is no specific limitation for comparison with this application.

[0105] In another example, to reduce the computational overhead of calculating N eigenvalues, before calculating the eigenvalues of data block A, data block A can be sampled to generate a smaller sample, and then the N eigenvalues can be calculated on the small sample (i.e., the sampled data block A), and then multi-level super features can be extracted through different super feature extraction parameters. In this way, the computational overhead is reduced and the calculation speed of multi-level super features is increased.

[0106] Any sampling method can be used to sample data block A. For example, sampling can be performed at fixed intervals, such as sampling once every 10 bytes; sampling can also be performed when the sliding hash value meets certain mathematical conditions, such as when the sliding hash value can be divided evenly by a certain integer, or when the remainder of dividing by a certain integer is a specific value. The embodiments of the present application do not specifically limit the sampling method, and a suitable sampling method can be selected according to needs to sample data block A.

[0107] In step S302, based on the multi-level super features, similarity matching is performed on the super feature dataset corresponding to the data block set to obtain a matching result.

[0108] After the multi-level super features of data block A are extracted, similarity matching is performed on the super feature dataset using the multi-level super features of data block A to obtain a matching result.

[0109] It can be understood that the super feature dataset includes the multi-level super features of each data block in the data block set, and the data block set includes multiple stored data blocks. The stored data blocks can be understood as data belonging to the same program as data block A. For example, data block A and the stored data blocks are all data in the same database.

[0110] The matching result includes successful matching and the output data block (i.e., the matched data block), as well as failed matching, that is, no data block similar to data block A is matched.

[0111] In one example, matching can be performed step by step from the super feature dataset in the order of decreasing super feature levels of the multi-level super features. When the super features of a certain level are successfully matched, it is determined that the super features of the data block match the super features of data block A, and the data block is output. For example, first match the first-level super features with the first-level super features in the super feature dataset. When at least one super feature value of the first-level super features of the two data blocks is the same, it is determined that the first-level super features of the two data blocks are successfully matched, and then it is determined that the two data blocks are successfully matched, and the data block is output; if the first-level super features of the two data blocks are not successfully matched, then match the first-level super features of the next data block. If no data block is matched through the first-level super features in the entire dataset, continue to match through the second-level super features. If the second-level super features are successfully matched, output the successfully matched data block. If no data block is matched through the second-level super features in the dataset, continue to match through the third-level super features. If the third-level super features are successfully matched, output the successfully matched data block. If no data block is matched through the third-level super features in the dataset, output that the matching result is a failure, that is, no data block similar to data block A is matched in the dataset.

[0112] Of course, since the similarity detection threshold of the low-level super features is low, compared with the single-threshold super feature matching in the related art, the matching coverage rate of this application is much greater than that of the single-threshold super feature matching scheme in the related art, and potential similar blocks can be matched out, and the probability of matching failure is low.

[0113] Optionally, when using the super features of any level for matching, the first-fit principle is adopted for matching, and the first data block having at least one same super feature value as data block A is selected as the successfully matched data block, that is, the first data block successfully matched with data block A is output. After a data block is matched, the matching is stopped to save the calculation overhead.

[0114] By means of step-by-step matching in the order of decreasing super feature levels of the multi-level super features, on the one hand, it is ensured that the high-level super features will be preferentially matched, so that data blocks with higher similarity will be preferentially matched at the high-threshold level, ensuring the high quality of the matching result; on the other hand, the low-level super features have a lower similarity detection threshold, which can ensure that when the high-level super features fail to match, more potential similar blocks are covered, improving the coverage rate of similarity matching, avoiding missing potential similar blocks, and reducing the occurrence of matching failures.

[0115] In another example, it is also possible to first perform matching from the lowest-level super-features, use the lowest-level super-features to screen out data blocks with a certain similarity, and then perform matching from the data set level by level in descending order from the multi-level super-features other than the lowest-level super-features, increasing the matching efficiency.

[0116] In step S303, based on the matching result, determine the similar data block corresponding to the target data block.

[0117] Due to the existence of the lowest-level super-features in the multi-level super-feature similarity matching, some false positive similar blocks may be output in the final matching result. That is, when the higher-level super-features (for example, when the multi-level super-feature is a three-level super-feature, the first-level super-feature and the second-level super-feature) both fail to match, the lowest-level super-feature (for example, the third-level super-feature) will be used for matching. However, the lowest-level super-feature sacrifices some accuracy for high coverage, so the data blocks matched may have a similarity that does not meet the standard, that is, the similarity of the output similar block to data block A is too low to meet the requirements. For example, using the matched data block as the reference block to perform differential compression on data block A, the compression rate is too low to meet the requirements.

[0118] To solve the problem that the matched data block may have false positives, after the data block is matched, a false positive test is performed on the matched data block. If the test passes, it proves that the similarity of the data block meets the standard and is a "true positive", and the matched data block is used as the similar data block of data block A. If the test fails, it proves that the similarity of the data block does not meet the standard and is a "false positive", and the data block is discarded.

[0119] The difficulty of false positive detection comes from the ambiguity of false positive determination. To design a suitable mechanism to solve this problem, the criteria for false positive determination should be determined first. The goal of differential compression is to improve the overall compression ratio. The advantage of differential compression is that it performs differential compression by detecting the base blocks (i.e., similar data blocks) on the complete data set. In contrast, lossless compression can only remove redundancy within a very limited data range. Therefore, if differential compression of the base block and the new data block (i.e., the data block to be compressed) can improve the overall compression ratio, it can be determined as a true positive; if not, it can be determined as a false positive. However, such a determination criterion is difficult to directly apply to practical systems. Modern lossless compression algorithms usually require an input buffer of 128 KB or larger to achieve a good compression ratio, and the computational workload is very large. This input buffer is much larger than the length of the data blocks in the system (usually 4 KB or 8 KB), and each input buffer can contain dozens of data blocks. Testing one by one whether differential compression of each data block can bring a higher compression ratio will be very inefficient, because it requires calculating the original lossless compression ratio of a complete input buffer, and then rolling back and checking whether each block in the buffer can improve the compression ratio through differential compression.

[0120] An embodiment of the present application provides an adaptive false positive detection mechanism. This detection design is based on two observations. The first observation is that the compressibility of a data set usually does not change drastically. And, since the input buffer used by current lossless compression algorithms usually contains many data blocks and all the data in the buffer is compressed together, small changes in the compression ratios of different data blocks usually do not affect the compression ratio of the overall data in the buffer. The second observation is that common differential compression coding algorithms (such as xDelta) include a local lossless compression step to compress the delta blocks. Therefore, the encoded delta blocks are difficult to be compressed in subsequent lossless compression steps. The key idea of the adaptive false positive detection mechanism in the embodiment of the present application is to compare the differential compression ratio of the current block with the compression ratios of the previous L lossless compression records. Assume that the differential compression ratio is higher than the average compression ratio of the previous L lossless compression records. In this case, we mark the delta block as a true positive, otherwise as a false positive.

[0121] In other words, by using the matched data block as the base block for differential compression of data block A, and then calculating the compression ratio of the differential compression as the first compression ratio; then calculating the average compression ratio of the previous L lossless compression records as the second compression ratio; comparing the magnitudes of the two compression ratios, if the compression ratio of the differential compression is greater than the compression ratio of the lossless compression, it is determined that the matched data block is a true positive and the detection passes, that is, the similarity between the matched output data block and data block A meets the standard; otherwise, it is determined that the matched data block is a false positive and the detection fails, that is, the similarity between the matched output data block and data block A does not meet the standard.

[0122] In another example, the second compression ratio can also be the median of the L compression ratios corresponding to the first L lossless compression records, and then by comparing the magnitudes of the first compression ratio and the second compression ratio, it is determined whether the verification passes.

[0123] In this way, by designing a false positive verification mechanism for the data blocks output by the matching, it is possible to adapt to data sets with different characteristics, avoid false similar data with low similarity that is not sufficient to improve the overall compression ratio by differential compression, and avoid performing differential compression on the data blocks to be compressed using false positive data, thereby ensuring the compression ratio stability of the storage system in different environments.

[0124] It should be explained that most storage systems have lossless compression and there are lossless compression records, and the above verification mechanism can be adopted. If there are extremely few storage systems without lossless compression, the embodiments of the present application perform lossless compression on the data blocks at fixed periodic intervals to ensure that the storage system has lossless compression records.

[0125] In another example, the data blocks matched can also be directly used as the similar data blocks of data block A, increasing the similarity detection efficiency of the similar data blocks.

[0126] In another example, the similar data detection method provided by the embodiments of the present application further includes that when data block A does not match a similar data block from the data block set, for example, no data block is matched from the data block set through multi-level super features, that is, the matching result is a matching failure or the matched data block fails the verification; at this time, the multi-level super features corresponding to data block A are saved to the super feature data set to expand the super feature data set, which is beneficial to the subsequent similarity detection of data blocks. When data block A matches a similar data block (that is, the matched data block passes the verification), the multi-level super features of data block A are not saved.

[0127] That is to say, in the similar data block detection method provided by the present application, the super features of the incremental blocks are not stored in the super feature database, nor are such incremental blocks selected as the reference blocks. The reason is that the resulting chain references will cause serious data fragmentation and bring a large number of I / O operations when backing up / restoring the incremental blocks.

[0128] Figure 4 Shows a specific implementation flowchart of similar data block matching provided by the embodiments of the present application. As Figure 4As shown in the figure, first, feature extraction is performed on the target data block through sliding hashing and sampling, and a specific number (e.g., 12) of original feature values are extracted. Then, the 12 original feature values are grouped by different grouping methods. For example, the first grouping method is to divide the 12 original feature values into 4 groups in a way that each group has 4 feature values; the second grouping method is to divide the 12 original feature values into 4 groups in a way that each group has 3 feature values; the third grouping method is to divide the 12 original feature values into 6 groups in a way that each group has 2 feature values. Then, super-feature extraction is performed on each group. After super-feature extraction in the first grouping method, 3 super-feature values are obtained, and these 3 super-feature values constitute the first-level super-feature. After super-feature extraction in the second grouping method, 4 super-feature values are obtained, and these 4 super-feature values constitute the second-level super-feature. After super-feature extraction in the third grouping method, 6 super-feature values are obtained, and these 6 super-feature values constitute the third-level super-feature. Finally, similarity matching is performed in the order from the highest to the lowest super-feature level. The super-features at a higher level have a higher similarity matching threshold, ensuring the accuracy of the matched data blocks (i.e., it is highly probable that the similarity of the matched data blocks meets the standard), and the super-features at a lower level have a lower similarity matching threshold, ensuring the coverage rate of the matched data, achieving a balance between the coverage rate and accuracy of similar data block matching.

[0129] The similar data detection method provided by the embodiment of the present application extracts multi-level super-features of the data block to be compressed and performs matching in the order from the highest to the lowest super-feature level, ensuring that similar blocks with high expected similarity are preferentially matched, thereby improving the compression rate of differential compression. The super-features at a lower level ensure that as many potential similar blocks as possible can be matched, and at the same time, high coverage rate and high accuracy of similar block matching are achieved.

[0130] Based on the same concept as the embodiment of the foregoing similar data detection method, the embodiment of the present application also provides a similar data detection device 500. The similar data detection device 500 can be deployed on any device, equipment, platform, or device cluster with computing capabilities to execute, implement the similar data detection method provided by the embodiment of the present application, and achieve high coverage rate and high accuracy of similar data detection. The similar data detection device 500 includes units or modules for implementing Figure 3 and 4 each step in the similar data detection method shown in the figure.

[0131] Figure 5 This is a schematic structural diagram of a similar data detection device provided by the embodiment of the present application. As shown in Figure 5As shown in the figure, a similarity data detection device 500 includes a feature extraction module 501, a matching module 502, and a determination module 503. The feature extraction module 501 is used to extract multi-level super features corresponding to the target data block. The multi-level super features include super features at multiple levels, and the similarity detection threshold of the super features at a higher level is higher than that of the super features at a lower level. The matching module 502 is used to perform similarity matching from the super feature dataset corresponding to the data block set based on the multi-level super features to obtain a matching result. The super feature dataset includes the multi-level super features corresponding to each data block in the data block set. The determination module 503 is used to determine the similar data block corresponding to the target data block based on the matching result.

[0132] In another possible implementation, the matching module 502 is specifically configured to perform matching from the super feature dataset level by level in the order of the super feature levels of the multi-level super features from high to low. If the super features at any level are successfully matched, it is determined that the target data block and the data block to be matched are successfully matched.

[0133] In another possible implementation, each level of the multi-level super features includes multiple super feature values. For the matching of the super features at any level, if the target data block and the data block to be matched have at least one same super feature value in the super features at the same level, it is determined that the super features at any level are successfully matched.

[0134] In another possible implementation, the determination module 503 is specifically configured to determine the data block obtained by matching as the similar data block corresponding to the target data block.

[0135] In another possible implementation, the determination module 503 is specifically configured to perform an inspection on the data block obtained by matching based on a first compression ratio and a second compression ratio. The first compression ratio indicates the compression ratio corresponding to performing differential compression on the target data block with the data block obtained by matching as the reference block. The second compression ratio indicates the average compression ratio of the first L lossless compressions, where L is a positive integer greater than 1. If the inspection passes, the data block obtained by matching is determined as the similar data block corresponding to the target data block.

[0136] In another possible implementation, a specific implementation of performing an inspection on the data block obtained by matching based on the first compression ratio and the second compression ratio is: checking whether the first compression ratio is greater than the second compression ratio. If so, it is determined that the inspection passes; if not, it is determined that the inspection fails.

[0137] In another possible implementation, the feature extraction module 501 is specifically configured to extract N eigenvalue of the target data block, where N is a positive integer greater than 1; perform super feature extraction on the N eigenvalue based on multiple different super feature extraction parameters to obtain multiple levels of super features; where the super feature extraction parameters indicate the number of groups of the N features and the number of eigenvalue in each group.

[0138] In another possible implementation, a specific implementation of extracting N eigenvalue of the target data block is to sample the target data block to obtain a sampled data block; perform feature extraction on the sampled data block through N different sliding hash functions to obtain N different eigenvalue sets; determine the N eigenvalue of the target data block based on the smallest eigenvalue in each eigenvalue set of the N eigenvalue sets.

[0139] In another possible implementation, the similarity data detection device 500 provided in this application further includes a storage module 504, and the storage module is configured to store the multi-level super features corresponding to the target data block in the super feature dataset when the similar data block is not matched from the data block set for the target data block.

[0140] The similarity data detection device 500 according to the embodiment of this application can correspondingly execute the method described in the embodiment of this application, and the above and other operations and / or functions of each module in the similarity data detection device 500 are respectively for implementing Figure 3 and 4 the corresponding processes of each method in, for the sake of brevity, will not be described in detail here.

[0141] The similarity data detection device provided in the embodiment of this application can be deployed in the similarity data detection module of the storage system to complete the similarity data detection with high quality, so as to facilitate differential compression of the data block to be stored based on the detected similar data block.

[0142] Figure 6 The flowchart of a data compression method is shown. This data compression method can be applied to any storage system, such as Figure 6 shown, this data compression method at least includes step S601 and step S602.

[0143] In step S601, perform similarity data detection on the stored data block set to obtain the similar data block corresponding to the data block to be compressed.

[0144] According to the similarity data detection method described above, perform similarity data detection on the stored data block set to obtain the similar data block corresponding to the data block to be compressed. For the specific implementation of the similarity data detection method, refer to the above description, and for the sake of brevity, it will not be described in detail here.

[0145] In step S602, using the similar data block as the reference block, perform differential compression on the data block to be compressed.

[0146] After detecting the similar data block corresponding to the data block to be compressed from the stored data blocks, use this similar data block as the reference block to perform a differential compression operation on the data block to be compressed. For example, use the sliding window technique to simultaneously detect the content in the data block to be compressed and the reference block. For the content in the reference block, if the same content exists in the data block to be compressed, use the COPY command for encoding; otherwise, use the INSERT command for encoding. The COPY command contains the offset and length of the repeated content in data block A, while the INSERT command contains the length of the non-repeated content in data block B and the content itself. The resulting file containing the COPY and INSERT commands is called the differential block, which is also the result after performing the differential compression operation.

[0147] Based on the same concept as the embodiment of the foregoing data compression method, an embodiment of the present application also provides a data compression device 700. This data compression device 700 can be deployed in any storage system to implement the data compression method provided by the embodiment of the present application to achieve a high-quality compression effect. The data compression device 700 includes units or modules for implementing Figure 6 each step in the data compression method shown.

[0148] Figure 7 It is a schematic structural diagram of a data compression device provided by an embodiment of the present application. As Figure 7 shown, this data compression device 700 includes a similar data detection module 701 and a compression module 702. Among them, the similar data detection module 701 is used to perform similar data detection from the stored data block set based on the similar data detection method provided by the embodiment of the present application to obtain the similar data block corresponding to the data block to be compressed; the compression module 702 is used to use the similar data block as the reference block to perform differential compression on the data block to be compressed.

[0149] The similar data detection device provided by the embodiment of the present application can be deployed in the data transmission end device in the data transmission system to perform high-quality similar data detection, so as to facilitate differential compression of the data to be transmitted based on the detected similar data block, and then transmit the compressed data, saving bandwidth and increasing transmission efficiency.

[0150] Figure 8 It shows a schematic flow diagram of a data transmission method provided by an embodiment of the present application. This data transmission system can be deployed in any data transmission end device to save data transmission bandwidth and improve data transmission speed. As Figure 8 shown, this data transmission method includes at least step S801 to step S803.

[0151] In step S801, similarity data detection is performed on multiple data blocks to be transmitted, and a set of similar data in the multiple data blocks is obtained.

[0152] Exemplarily, the data to be transmitted includes data block A, data block B, data block C, and data block D. The similarity detection method provided in the embodiments of the present application is used to perform similarity detection on data block A, data block B, data block C, and data block D, and it is obtained that data block A and data block B are similar data blocks, and data block C and data block D are similar data blocks. For the specific implementation of the similarity data detection method, refer to the above description. For the sake of brevity, it will not be elaborated here.

[0153] In step S802, differential compression is performed on the set of similar data to obtain a compressed data set.

[0154] A differential compression operation is performed on the similar data blocks detected in the previous step. For example, a differential compression operation is performed on data block A and data block B to obtain data block A and differential block △(AB), and a differential compression operation is performed on data block C and data block D to obtain data block C and differential block △(CD). That is, the compressed data set is data block A and differential block △(AB), data block C and differential block △(CD).

[0155] In step S803, the compressed data set is transmitted.

[0156] The compressed data set, which is data block A and differential block △(AB), data block C and differential block △(CD), is transmitted to the destination device, saving transmission bandwidth and accelerating transmission efficiency.

[0157] Based on the same concept as the foregoing embodiment of a data transmission method, an embodiment of the present application also provides a data transmission device 900. The data transmission device 900 can be deployed at any data transmission end to implement the data transmission method provided in the embodiments of the present application, so as to save data transmission bandwidth and increase data transmission efficiency. The data transmission device 900 includes units or modules for implementing Figure 8 each step in the shown data transmission method.

[0158] Figure 9 This is a schematic structural diagram of a data transmission device provided in an embodiment of the present application. As Figure 9As shown in the figure, a data transmission device 900 includes a similar data detection module 901, a compression module 902, and a transmission module 903. The similar data detection module 901 is used to perform similar data detection on multiple data blocks to be transmitted based on the similar data detection method provided in the embodiments of the present application, and obtain a set of similar data in the multiple data blocks. The compression module 902 is used to perform differential compression on the set of similar data to obtain a compressed data set. The transmission module 903 is used to perform data transmission on the compressed data set.

[0159] The similar data detection device provided in the embodiments of the present application can be deployed on the management node of the distributed storage system, and is used to complete similar data detection with high quality, so as to store the data to be stored in the storage node with similar data.

[0160] Figure 10 The structure diagram of a distributed storage system is shown. As Figure 10 shown, the distributed storage system 100 includes a management node 101 and multiple storage nodes (such as storage node 102, storage node 103, storage node 104, and storage node 105).

[0161] The management node receives an IO request from the upper-layer application. The IO request carries the data to be stored. The management node decides which storage node among the multiple storage nodes to store the data to be stored. For example, the management node uses the similar data detection method provided in the embodiments of the present application to detect which storage node among the multiple storage nodes has a similar data block to the data to be stored, and then determines it as the target storage node, and stores the data to be stored in the target storage node. The storage node is provided with a storage medium (such as a disk) for storing data.

[0162] Figure 11 The flowchart of a data storage method is shown. This data storage method can be executed and implemented by Figure 10 the management node 101 in the distributed storage system shown, so as to achieve efficient data distributed storage. As Figure 11 shown, this data storage method at least includes step S1101 to step S1103.

[0163] In step S1101, similar data detection is performed on the data in each storage node among the multiple storage nodes.

[0164] According to the similar data detection method described above, similar data detection is performed on the data in each storage node among the multiple storage nodes to determine the storage node storing the similar data block to the data to be stored. For the specific implementation of the similar data detection method, refer to the above description. For the sake of brevity, it will not be elaborated here.

[0165] In step S1102 , based on the similar data detection result, a target storage node is determined from multiple storage nodes.

[0166] Exemplarily, by detecting similar data, it is determined that a data block similar to the data to be stored is stored in the storage node 102, and then the storage node 102 is determined to be the target storage node.

[0167] In step S1103, the data block to be stored is stored in the target storage node.

[0168] After determining the target storage node, the management node 101 sends a write request to the target storage node, which carries the data to be stored. The target storage node then uses the similar data block as a reference block, performs a differential compression operation on the data to be stored, obtains a differential block, and then stores the differential block to the disk to achieve distributed data storage.

[0169] Based on the same concept as the above-mentioned data storage method embodiment, the present application embodiment also provides a data storage device 1200, which can be deployed in the management node of any distributed storage system to implement the data storage method provided in the present application embodiment to achieve efficient distributed storage of node data. The data storage device 1200 includes a Figure 11 The units or modules of each step in the data storage method shown.

[0170] Figure 12 A schematic diagram of a data storage device provided in an embodiment of the present application. Figure 12 As shown, the data storage device 1200 includes a similar data detection module 1201, a determination module 1202 and a storage module 1203, wherein the similar data detection module 1201 is used to perform similar data detection on data in each storage node among multiple storage nodes based on the similar data detection method provided in the embodiment of the present application; the determination module 1202 is used to determine a target storage node from multiple storage nodes based on the similar data detection result, and the target storage node stores a data block similar to the data block to be stored; the storage module 1203 is used to store the data block to be stored in the target storage node.

[0171] The present application also provides a computing device, including at least one processor, a memory and a communication interface, wherein the processor is used to execute Figure 3 , 4 , the methods described in 6, 8, and 11.

[0172] Figure 13 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.

[0173] like Figure 13As shown, the computing device 1300 includes at least one processor 1301, a memory 1302, and a communication interface 1303. Among them, the processor 1301, the memory 1302, and the communication interface 1303 are communicatively connected, and the communication connection can be achieved by wire (such as a bus) or by wireless means. The communication interface 1303 is used to send and / or receive data sent by other devices; the memory 1302 stores computer instructions, and the processor 1301 executes the computer instructions to at least execute a similar data detection method in the foregoing method embodiments to achieve high coverage and high accuracy of similar data detection.

[0174] It should be understood that in the embodiments of the present application, the processor 1301 may be a central processing unit CPU, and the processor 1301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0175] The memory 1302 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1301. The memory 1302 may also include a non-volatile random access memory.

[0176] The memory 1302 can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0177] It should be understood that the computing device 1300 according to the embodiments of the present application can execute the methods shown in Figure 3 , 4 , 6, 8, 11. For a detailed description of the implementation of this method, please refer to the above. For the sake of brevity, it will not be repeated here.

[0178] Embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer instructions are executed by a processor, the methods mentioned above are implemented.

[0179] Embodiments of the present application provide a chip, which includes at least one processor and an interface. The at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the methods mentioned above.

[0180] Embodiments of the present application provide a computer program or a computer program product, which includes instructions. When the instructions are executed, the computer is made to execute the methods mentioned above.

[0181] Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0182] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0183] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above is only the specific implementation manner of this application and is not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included in the protection scope of this application.

Claims

1. A method for detecting similar data, characterized in that Including: Extracting multi-level super features of a target data block, where the multi-level super features include super features of multiple levels, and the similarity detection threshold of the super features at a higher level is higher than that of the super features at a lower level; Based on the multi-level super features, performing similarity matching from a super feature dataset to obtain a matching result, where the super feature dataset includes the multi-level super features of each data block in a data block set; Based on the matching result, determining a similar data block corresponding to the target data block.

2. The method according to claim 1, wherein The performing matching from the super feature dataset based on the multi-level super features to obtain a matching result includes: Sequentially matching the multi-level super features of the target data block and the multi-level super features of a data block to be matched in descending order of the super feature levels of the multi-level super features, where the data block to be matched is any data block in the data block set; If the super features at any level are successfully matched, it is determined that the target data block and the data block to be matched are successfully matched.

3. The method according to claim 2, characterized in that, Each level of super features in the multi-level super features includes multiple super feature values; If the target data block and the data block to be matched have at least one same super feature value in the super features at any level, it is determined that the super features at that level are successfully matched.

4. The method according to claim 2 or 3, characterized in that, The determining a similar data block corresponding to the target data block based on the matching result includes: Determining the data block to be matched that is successfully matched with the target data block as the similar data block corresponding to the target data block.

5. The method according to claim 2 or 3, characterized in that, The determining a similar data block corresponding to the target data block based on the matching result includes: Based on a target compression ratio and a compression ratio threshold, verifying the data block to be matched that is successfully matched with the target data block, where the target compression ratio indicates the compression ratio corresponding to a differential compression operation performed on a reference data block and the target data block, and the reference data block is the data block to be matched that is successfully matched with the target data block; If the verification passes, determining the data block to be matched that is successfully matched with the target data block as the similar data block corresponding to the target data block.

6. The method according to claim 5, characterized in that, The verifying the data block obtained by matching based on the target compression ratio and the compression ratio threshold includes: Verifying whether the target compression ratio is greater than the compression ratio threshold. If so, it is determined that the verification passes; if not, it is determined that the verification fails.

7. The method according to any one of claims 1-6, characterized in that, The extracting multi-level super features corresponding to the target data block includes: Extracting N feature values of the target data block, where N is a positive integer greater than 1; Based on multiple different super feature extraction parameters, performing super feature extraction on the N feature values to obtain the super features of multiple levels; Wherein, the super feature extraction parameters indicate the number of groups of the N features and the number of feature values in each group.

8. The method according to claim 7, characterized in that, The extracting N feature values of the target data block includes: Sampling the target data block to obtain a sampled data block; Performing feature extraction on the sampled data block through N different sliding hash functions to obtain N different sets of feature values; Based on the smallest feature value in each set of feature values in the N sets of feature values, determining the N feature values of the target data block.

9. The method according to any one of claims 1-8, characterized in that, Also including: When the target data block fails to match a similar data block from the data block set, save the multi-level super features corresponding to the target data block to the super feature data set.

10. A data compression method, characterized in that, Including: Based on the similar data detection method according to any one of claims 1-9, perform similar data detection from the stored data block set to obtain a similar data block corresponding to the data block to be compressed; Use the similar data block as a reference block to perform differential compression on the data block to be compressed.

11. A data transmission method, characterized in that, Including: Based on the similar data detection method according to any one of claims 1-9, perform similar data detection on multiple data blocks to be transmitted to obtain a set of similar data in the multiple data blocks; Perform differential compression on the set of similar data to obtain a compressed data set; Transmit the compressed data set.

12. A data storage method, applied to a distributed storage system, the distributed storage system comprising a plurality of storage nodes, characterized in that, Including: Based on the similar data detection method according to any one of claims 1-9, perform similar data detection on the data in each of the multiple storage nodes; Based on the similar data detection result, determine a target storage node from the multiple storage nodes, where the target storage node stores a data block similar to the data block to be stored; Store the data block to be stored in the target storage node.

13. A similar data detection device, characterized in that, Including: A feature extraction module for extracting multi-level super features of a target data block, where the multi-level super features include super features of multiple levels, and the similarity detection threshold of the super features of a higher level is higher than that of the super features of a lower level; A matching module for performing similarity matching from a super feature data set based on the multi-level super features to obtain a matching result, where the super feature data set includes the multi-level super features corresponding to each data block in the data block set; A determination module for determining a similar data block corresponding to the target data block based on the matching result.

14. The device according to claim 13, wherein The matching module is specifically used for: Match the multi-level super features of the target data block and the multi-level super features of the data block to be matched level by level in descending order of the super feature levels of the multi-level super features, where the data block to be matched is any data block in the data block set; If the super features of any level match successfully, it is determined that the target data block and the data block to be matched match successfully.

15. The device according to claim 14, characterized in that, Each level of super features in the multi-level super features includes multiple super feature values; If the target data block and the data block to be matched have at least one same super feature value in the super features of any level, it is determined that the super features of that level match successfully.

16. The device according to claim 14 or 15, characterized in that, The determination module is specifically used for: Determine the data block to be matched that matches the target data block as the similar data block corresponding to the target data block.

17. The device according to claim 14 or 15, characterized in that, The determination module is specifically used for: Based on a target compression ratio and a compression ratio threshold, verify the data block to be matched that matches the target data block, where the target compression ratio indicates the compression ratio corresponding to performing differential compression on a reference data block and the target data block, and the reference data block is the data block to be matched that matches the target data block; If the verification passes, determine the to-be-matched data block that successfully matches the target data block as the similar data block corresponding to the target data block.

18. The device according to claim 17, characterized in that, The verifying the data block obtained by matching based on the target compression ratio and the compression ratio threshold includes: Verify whether the target compression ratio is greater than the compression ratio threshold. If so, determine that the verification passes; if not, determine that the verification fails.

19. The device according to any one of claims 13-18, characterized in that, The feature extraction module is specifically configured to: Extract N eigenvalue of the target data block, where N is a positive integer greater than 1; Perform super-feature extraction on the N eigenvalue based on multiple different super-feature extraction parameters to obtain multiple levels of super-features; Among them, the super-feature extraction parameter indicates the number of groups of the N features and the number of eigenvalue in each group.

20. The device according to claim 19, characterized in that, The extracting N eigenvalue of the target data block includes: Sample the target data block to obtain a sampled data block; Extract features from the sampled data block through N different sliding hash functions to obtain N different eigenvalue sets; Determine the N eigenvalue of the target data block based on the smallest eigenvalue in each eigenvalue set of the N eigenvalue sets.

21. The device according to any one of claims 13-20, characterized in that, It further includes: A saving module, configured to save the multi-level super-features corresponding to the target data block to the super-feature dataset when the similar data block is not matched from the data block set for the target data block.

22. A data compression device, characterized in that, It includes: A similar data detection module, configured to perform similar data detection on the stored data block set based on the similar data detection method according to any one of claims 1-9 to obtain the similar data block corresponding to the to-be-compressed data block; A compression module, configured to perform differential compression on the to-be-compressed data block with the similar data block as the reference block.

23. A data transmission device, characterized in that, It includes: A similar data detection module, configured to perform similar data detection on multiple data blocks to be transmitted based on the similar data detection method according to any one of claims 1-9 to obtain a set of similar data in the multiple data blocks; A compression module, configured to perform differential compression on the set of similar data to obtain a compressed data set; A transmission module, configured to transmit the compressed data set.

24. A data storage device is applied to a distributed storage system, and the distributed storage system includes multiple storage nodes. It is characterized in that It includes: A similar data detection module, configured to perform similar data detection on the data in each of the multiple storage nodes based on the similar data detection method according to any one of claims 1-9; A determination module, configured to determine a target storage node from the multiple storage nodes based on the similar data detection result, where the target storage node stores a data block similar to the to-be-stored data block; A storage module, configured to store the to-be-stored data block in the target storage node.

25. A computing device, comprising a memory and a processor, characterized in that, Instructions are stored in the memory, and when the instructions are executed by a processor, the method according to any one of claims 1-12 is implemented.

26. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-12 is implemented.