A Cloud Storage Redundant Data Prediction Method and Device Based on Similar Data Detection

By using similar data detection methods in cloud storage backup, the hash fingerprint and similar feature groups of data blocks are extracted, and combined with the sampling method, a three-stage framework is built, which solves the problem of incomplete detection of similar data in the existing technology, and improves the redundant data deduplication performance and storage space utilization.

CN114579362BActive Publication Date: 2025-05-30NANHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210182503.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-05-30
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and estimate similar data in cloud storage backup, resulting in incomplete redundant data deletion and wasted storage space.

Method used

A cloud storage redundant data prediction method based on similar data detection is adopted. By extracting the hash fingerprint and similar feature groups of data blocks, combined with the sampling method of Bernoulli binomial distribution, a three-stage framework (feature extraction, sampling and scanning prediction) is constructed to estimate data redundancy.

Benefits of technology

Improves the deduplication performance of redundant data in cloud storage, reduces storage space waste, and maintains high estimation accuracy in the case of large data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114579362B_ABST
    Figure CN114579362B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for predicting redundant data in cloud storage based on similar data detection. The method includes: partitioning cloud storage data to obtain data blocks; traversing all data blocks and calculating the hash fingerprints corresponding to the data blocks by using a hash algorithm; calculating the similar feature groups of the data blocks by using the N-transform method; selecting m data blocks according to the size of the dataset to be predicted; traversing the set of data blocks composed of all the extracted data blocks, and circularly selecting m initial samples by using the Bernoulli binomial distribution; traversing the initial sample set composed of the initial samples, making judgments based on the hash fingerprints and the similar feature groups, and adding the duplicate data blocks that do not meet the conditions of the hash fingerprints and the similar feature groups to the base samples to obtain a base sample set; traversing the dataset to be predicted, and determining duplicate data and similar data based on the base sample set, so as to calculate an estimated value of data redundancy. The present invention can effectively improve the deduplication performance of redundant data in cloud storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer information storage, and particularly to a method and device for predicting redundant data in cloud storage based on similar data detection. Background Art

[0002] In the scenario of cloud storage backup, the funds that users need to pay are often proportional to the data to be stored. By detecting duplicate data and similar data in the stored data before backup, a large amount of storage space can be effectively saved, thereby reducing unnecessary expenses for users to purchase storage space.

[0003] Previously, Danny Harnik et al. proposed a two-stage framework in the field of deduplication estimation. This framework divides data into blocks and uses the hash value of block-level data as its unique identifier. After sampling the data set and scanning the complete data, the duplication rate of the data set is given. Although the above framework can accurately estimate duplicate data, it does not consider the similar data existing in the backup data. And researchers have proposed methods for block-level data similarity detection such as N-transform and Finesse, but these two methods only judge whether two data blocks are similar, do not give an estimated value of similarity, and also lack an overall framework to apply them to the deduplication estimation scenario. Summary of the Invention

[0004] The present invention aims to solve at least one of the above-mentioned related problems to a certain extent.

[0005] According to one aspect of the present invention, there is provided a method for predicting redundant data in cloud storage based on similar data detection, and the method for predicting redundant data in cloud storage includes:

[0006] Feature extraction stage of data blocks:

[0007] Divide the cloud storage data into blocks to obtain data blocks;

[0008] Traverse all data blocks, and use the hash algorithm to calculate the hash fingerprint corresponding to the data blocks;

[0009] Use the N-transform method to calculate the similar feature group of the data blocks;

[0010] Sample set acquisition stage:

[0011] According to the size of the data set to be predicted, determine that the number of required data blocks is m;

[0012] Traverse the set composed of all the extracted data blocks, and use the Bernoulli binomial distribution to cyclically select m initial samples;

[0013] Traverse the initial sample set composed of the initial samples, make judgments based on the hash fingerprint and the similar feature group, and add the duplicate data blocks that do not meet the conditions of the hash fingerprint and the similar feature group to the base samples to obtain the base sample set;

[0014] Prediction stage:

[0015] Traverse the data set to be predicted, and determine duplicate data and similar data based on the base sample set, so as to calculate the estimated value of data redundancy.

[0016] Furthermore, traversing the initial sample set composed of the initial samples, making judgments based on the hash fingerprint and the similar feature group, and adding the duplicate data blocks that do not meet the conditions of the hash fingerprint and the similar feature group to the base samples to obtain the base sample set, including:

[0017] Initialize the base samples to be empty, and record the attributes of each data block in the base samples:

[0018] Record ρ i as the compression ratio of data block i. If it is not compressed, ρ i = 1;

[0019] Record base i as the frequency of redundancy of data block i in the initial samples, and initialize it to 1;

[0020] Record count i as the frequency of redundancy of data block i in the entire data set, and initialize it to 0;

[0021] Traverse the initial sample set and make the following judgments:

[0022] If there is a data block in the base samples that is the same as the current data block in the initial sample set and the hash fingerprints of the same data blocks are also the same, then increase the base i of this data block in the current base samples by 1;

[0023] Otherwise, traverse the base samples. If the similar feature group of the current data block in the initial sample set has the same dimension as that of a certain data block in the base samples, record the number of similar features, calculate the similarity, and if the calculated maximum similarity is greater than the set similarity threshold, then increase the base i of the data block in the base samples by the similarity;

[0024] Otherwise, add the current data block in the initial sample set to the base samples to generate the base sample set.

[0025] Furthermore, traverse the set composed of all the extracted data blocks, and use the Bernoulli binomial distribution to cyclically select m initial samples, including:

[0026] Generate a random number according to the Bernoulli binomial distribution:

[0027]

[0028] where l is the number of data blocks included in the current data block set, and n is the total number of data blocks in the data set;

[0029] If k ≥ 1, select k random data blocks and add them to the initial sample. If k = 0, ignore;

[0030] The size of the initial sample set composed of the initial samples is m'. If m' is greater than m, randomly select m of the initial samples to form an initial sample set;

[0031] If m' is less than m, return to the step of selecting the initial sample, reselect the initial sample until an initial sample set composed of m initial samples is obtained.

[0032] Further, traverse the data set to be predicted, and determine duplicate data and similar data based on the base sample set, thereby calculating an estimated value of data redundancy, including:

[0033] Traverse the data set to be predicted and make the following judgments:

[0034] If there is a data block in the base sample that is the same as the data block in the initial sample set and the hash fingerprints of the same data blocks are also the same, then the attribute count of the corresponding data block in the current base sample i +1;

[0035] Otherwise, traverse the base sample. If the dimension of the similarity feature group of the current data block in the initial sample set is the same as that of a data block in the base sample, record the number of similar features, calculate the similarity, and if the calculated maximum similarity is greater than the set similarity threshold, then the count of the data block of the base sample i + similarity;

[0036] Otherwise, repeat the foregoing scanning steps again.

[0037] Further, the calculation formula for calculating the estimated value of data redundancy is:

[0038]

[0039] where B is the base sample obtained in the sampling stage; base i is the frequency of redundancy of data block i in the initial sample during the sampling stage; count i is the frequency of redundancy of data block i in the entire data set during the scanning stage; ρ i is the compression ratio of data block i.

[0040] Further, the cloud storage data is chunked to obtain data chunks, including:

[0041] Chunk the original data in cloud storage according to a fixed length.

[0042] According to another invention of the present invention, a cloud storage redundant data prediction device based on similar data detection is also disclosed. The cloud storage redundant data prediction device includes a memory and a processor;

[0043] A cloud storage redundant data prediction program that can run on the processor is stored on the memory. When the cloud storage redundant data prediction program is executed by the processor, it implements the steps of the cloud storage redundant data prediction method described in any of the previous items.

[0044] The cloud storage redundant data prediction algorithm based on similar data detection disclosed in the present invention is divided into three stages: a feature extraction stage, a sampling stage, and a scanning prediction stage. Compared with the existing technology, the advantages of the present invention are: 1. The original two-stage framework is improved to a three-stage framework, and the improved framework can adapt to different feature extraction methods and is easy to optimize; 2. The feature extraction stage extracts the overall features (overall hash values of data chunks) and similar features (SFs extracted based on N-transform), and the similarity can be obtained according to the comparison of feature values between different chunks; 3. Through sampling and scanning, it is allowed to estimate without reading most of the actual data in the case of a large data set, and a certain estimation accuracy is guaranteed, effectively improving the deduplication performance of cloud storage redundant data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings that form a part of this specification depict embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0046] Figure 1 It is a schematic diagram of the overall flow of the algorithm in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0047] The following will describe various exemplary embodiments of the present invention in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0048] At the same time, it should be understood that, for the sake of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0049] The description of at least one exemplary embodiment below is merely illustrative in nature and in no way serves as a limitation to the present invention or its application or use.

[0050] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0051] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods and devices should be regarded as part of the authorization specification.

[0052] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0053] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it will not be discussed further in subsequent drawings.

[0054] Combined Figure 1 To illustrate Embodiment 1, the cloud storage redundant data prediction algorithm based on similar data detection is divided into three stages: the feature extraction stage, the sampling stage, and the scanning stage.

[0055] In this embodiment, the hash algorithm and the N-transform method are applied to the feature extraction stage, and the specific process is as follows:

[0056] The data in the dataset is divided into blocks according to a fixed length;

[0057] Traverse all the data blocks and perform the following operations:

[0058] According to the selected hash algorithm, calculate the hash fingerprint of the data block;

[0059] Calculate the similar feature group SFs of the data block according to the N-transform method;

[0060] Through the hash fingerprint and the similar feature group SFs obtained by the extraction in the above steps, the features corresponding to each data block are generated.

[0061] Specifically, the principles of the two extraction means mainly adopted in the feature extraction stage are as follows:

[0062] (1) Extract the hash fingerprint of the entire data block through the hash algorithm

[0063] A hash algorithm can map a binary value string of any length to a binary value string of a fixed length. The resulting fixed-length binary value string is the hash value of the original binary value string under this hash algorithm. The hash value formed by the hash algorithm is related to each byte of the original binary value string. When the original binary value string changes, the generated hash value also changes. That is, the hash value can be used as a unique identifier for the original data.

[0064] The cloud storage data is divided into blocks to obtain data blocks. The divided data blocks are hashed, and the hash values obtained using hash algorithms (such as MD5, SHA1, etc.) can uniquely identify the current data block. The generated hash values are used as the hash fingerprints of these data blocks.

[0065] (2) Extract the similar feature groups of the data blocks through the N-transform method

[0066] N-transform is a currently popular block-level similarity detection method proposed by Broder, which can extract a fixed number of features from a data block. This method extracts features through Rabin fingerprints (rolling hash algorithm), and then groups and compares these features to detect feature similarity. Through this method, the similar feature groups SFs of the data blocks can be finally obtained. The specific algorithm process is as follows:

[0067] a. Initialize the N-dimensional feature group features to 0;

[0068] b. Traverse the current data block bit by bit and perform the following operations:

[0069] a) Record FP, which is the Rabin fingerprint of the current bit of this data block;

[0070] b. Traverse the N-dimensional feature values features and record the linear mapping value transform of this FP in this dimension i , if this mapping value is greater than the current dimension feature i , then assign feature i the value of transform i .

[0071]

[0072] Among them, a i and b i are randomly predefined data (linear transformation), and L is the length of the data block.

[0073] The final N-dimensional feature values are sequentially divided into x groups, each group contains N / x features (N / x is usually an integer), and Rabin hashing is performed on each group again to obtain the final similar feature group SFs.

[0074] SF x = Rabin(feature x·i ,...,feature x·i+i-1 ) (2)

[0075] The N-transform method can detect highly similar data blocks to the greatest extent for two reasons: ① The matching of a similar feature SF means that almost all the features grouped in the SF are the same, which can effectively reduce the possibility of false detection. ② Calculating multiple SFs can increase the probability of detecting highly similar data blocks.

[0076] The specific steps in the sampling stage are as follows:

[0077] (1) Determine that the number of data blocks required for the initial sample set is m according to the size of the data set to be predicted;

[0078] (2) Traverse the set composed of all the extracted data blocks, and use the Bernoulli binomial distribution to loop and select m initial samples. The selection steps are as follows:

[0079] a. Let l be the number of data blocks contained in the current data block set, and n be the total number of data blocks in the data set.

[0080] Generate a random number k according to the Bernoulli binomial distribution. If k ≥ 1, select k random data blocks and add them to the initial sample; if k = 0, ignore this data block set.

[0081] (3) The size of the initial sample selected in the loop in (2) is m'. If m' is greater than m, randomly select m samples from this sample; if m' is less than m, go back to (2) and use to reselect the samples until the size of the initial sample set composed of the initial samples is finally m.

[0082] (4) Initialize the base sample to be empty. For each data block in the base sample, the following attributes will be recorded:

[0083] Define ρ i : The compression ratio of this data block i. If it is not compressed, then ρ i = 1.

[0084] Define base i : The frequency of occurrence of the data block that is redundant (identical or similar) to this data block, initialized to 1. If it is the same as this data block, the frequency is incremented by 1; if it is similar, the similarity is added.

[0085] Define count i : In the scanning stage, the frequency of occurrence of data block i being redundant in the entire data set, initialized to 0.

[0086] (5) Traverse the initial sample set generated in (3) and perform the following operations:

[0087] a. If there is a data block in the base sample that is the same as the current data block in the initial sample set, and there is a data block in the base sample with the same hash fingerprint as the current data block, then increment the attribute base of the corresponding data block in the current base sample i by 1;

[0088] b. If the situation in a does not exist, then consider the following:

[0089] a) Traverse the base sample. If there is the same SF in the feature values SFs of the current data block and those of a data block in the base sample, record the number of the same SF features. Then, (the number of the same SF features / the size of SFs) is the similarity value of this data block.

[0090] b) Record the information of the data block with the highest similarity in the loop. If the maximum similarity is greater than the set similarity threshold (e.g., 50%), then increment the base of the data block in this base sample i by the similarity;

[0091] c. If none of the above situations match, then add this data block to the base sample, and finally obtain the base sample set.

[0092] The specific steps of the scanning prediction stage are as follows:

[0093] (1) Traverse the entire data set to be predicted and perform the following operations:

[0094] a. If the current data block exists in the base sample set, and there is a data block in the base sample set with the same hash fingerprint as the current data block, then increment the attribute count of the corresponding data block in the current base sample set i by 1;

[0095] b. If the situation in a does not exist, then consider the following:

[0096] a) Traverse the base sample. If there is the same SF in the feature values SFs of the current data block and those of a data block in the base sample, record the number of the same SF features.

[0097] The number of the same SF features / the size of SFs is the similarity of this data block;

[0098] b) If the maximum similarity is greater than the set similarity threshold (e.g., 50%, the same as the similarity threshold set in the previous stage), then increment the count of the data block in this base sample set i by the similarity;

[0099] If none of the above conditions are met, restart the scanning loop.

[0100] (2) After the scanning phase ends, the estimated value of data redundancy can be calculated according to the following formula. The data redundancy sought here includes the duplicate data calculated in a) of step b and the similar data calculated in b) of step b:

[0101]

[0102] where B is the base sample finally obtained in the sampling phase; base i is the number of times the data block i appears redundantly in the initial sample during the sampling phase; count i is the number of times the data block i appears redundantly in the entire data set during the scanning phase; ρ i is the compression ratio of the data block i.

[0103] The technical key points of the method of the present invention are as follows:

[0104] 1. Cut the data set into block-level sizes for fine-grained feature extraction;

[0105] 2. Adopt the N-transform method to extract similar features of block-level data;

[0106] 3. Use the base sample set obtained by sampling to replace the larger actual data for prediction.

[0107] According to the second aspect of the present application, a cloud storage redundant data prediction device based on similar data detection is also disclosed. The cloud storage redundant data prediction device includes a memory and a processor;

[0108] A cloud storage redundant data prediction program that can run on the processor is stored on the memory. When the cloud storage redundant data prediction program is executed by the processor, it implements the steps of the cloud storage redundant data prediction method described in any one of the preceding items.

[0109] The cloud storage redundant data prediction device based on similar data detection can run on computing devices such as desktop computers, notebooks, palmtop computers, and cloud servers. The device that can run the cloud storage redundant data prediction device based on similar data detection may include, but is not limited to, a processor and a memory.

[0110] Those skilled in the art can understand that the above examples are merely examples of a cloud storage redundant data prediction device based on similar data detection, and do not constitute a limitation on a cloud storage redundant data prediction device based on similar data detection. It may include more or fewer components than the examples, or combine certain components, or different components. For example, a cloud storage redundant data prediction device based on similar data detection may also include input / output devices, network access devices, buses, etc. The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of a cloud storage redundant data prediction device based on similar data detection, and connects various parts of the entire cloud storage redundant data prediction device through various interfaces and lines. The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the cloud storage redundant data prediction device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0111] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle scope of the present invention shall be included in the protection scope of the present invention.

[0112] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

Claims

1. A method for predicting redundant data in cloud storage based on similar data detection, characterized in that, the method for predicting redundant data in cloud storage includes: The stage of extracting the features of data blocks: Chunk the cloud storage data to obtain data blocks; Traverse all data blocks and calculate the hash fingerprints corresponding to the data blocks using the hash algorithm; Calculate the similar feature groups of the data blocks using the N-transform method; The stage of collecting the sample set: According to the size of the data set to be predicted, determine that the number of required data blocks is m; Traverse the set composed of all the extracted data blocks, and use the Bernoulli binomial distribution to cyclically select m initial samples; Traverse the initial sample set composed of the initial samples, make judgments based on the hash fingerprints and the similar feature groups, and add the duplicate data blocks that do not meet the conditions of the hash fingerprints and the similar feature groups to the base samples to obtain the base sample set; The scanning and prediction stage: Traverse the data set to be predicted, and determine the duplicate data and similar data based on the base sample set, so as to calculate the estimated value of data redundancy; Traverse the initial sample set composed of the initial samples, make judgments based on the hash fingerprints and the similar feature groups, and add the duplicate data blocks that do not meet the conditions of the hash fingerprints and the similar feature groups to the base samples to obtain the base sample set, including: Initialize the base sample to be empty and record the attributes of each data block in the base sample: Record ρ i is the compression ratio of data block i. If not compressed, ρ i = 1; Record base i It is the frequency of redundancy of data block i in the initial sample, initialized to 1; Record count i Is the frequency of redundancy of data block i in the entire data set, initialized to 0; Traverse the initial sample set and make the following judgments: If there is a data block in the base sample that is the same as the current data block in the initial sample set and the hash fingerprints of the same data blocks are also the same, then the attribute base of this data block in the current base sample is i +1; Otherwise, traverse the base samples. If the dimension of the similar feature group of the current data block in the initial sample set is the same as that of a data block in the base sample, record the number of similar features and calculate the similarity. If the calculated maximum similarity is greater than the set similarity threshold, then add the similarity to the base of the data block of the base sample i + similarity; Otherwise, add the current data block in the initial sample set to the base sample to generate the base sample set; Traverse the set composed of all the extracted data blocks, and use the Bernoulli binomial distribution to cyclically select m initial samples, including: Generate a random number according to the Bernoulli binomial distribution: where l is the number of data blocks included in the current data block set, n is the total number of data blocks in the data set, and B is the base sample; If k≥1, then select k random data blocks and add them to the initial samples. If k = 0, then ignore; The size of the initial sample set composed of the initial samples is m'. If m' is greater than m, then randomly select m of the initial samples to form the initial sample set; If m' is less than m, then return to the step of selecting the initial samples and re-select the initial samples until an initial sample set composed of m initial samples is obtained; Traverse the data set to be predicted, and determine the duplicate data and similar data based on the base sample set, so as to calculate the estimated value of data redundancy, including: Traverse the data set to be predicted and make the following judgments: If there is a data block in the base sample that is the same as the data block in the initial sample set and the hash fingerprints of the same data blocks are also the same, then the attribute count of the corresponding data block in the current base sample i is incremented by 1; Otherwise, traverse the base samples. If the dimension of the similar feature group of the current data block in the initial sample set is the same as that of a data block in the base sample, record the number of similar features and calculate the similarity. If the calculated maximum similarity is greater than the set similarity threshold, then add the similarity to the count of the data block of the base sample i + similarity; Otherwise, repeat the previous scanning step again; Among them, the specific process of the N-transform method is as follows: a. Initialize the N-dimensional feature group features to 0; b. Traverse the current data block bit by bit and perform the following operations: a) Record FP, which is the Rabin fingerprint under the current bit of this data block; b) Traverse the N-dimensional eigenvalue features and record the linear mapping value transform of this FP in this dimension i , if this mapping value is greater than the current dimension feature i , then assign feature i the value of transform i ; Among them, a i and b i are linearly varied with randomly predefined data, and L is the data block length; Divide the final N-dimensional feature values into x groups in sequence, each group contains N / x features, N / x is an integer, and perform Rabin hashing on each group again to obtain the final similar feature group SFs; SF x = Rabin(feature x·i ,...,feature x·i+i-1 )。 2. A method for predicting redundant data in cloud storage based on similar data detection according to claim 1, characterized in that, The calculation formula for estimating the data redundancy is as follows: Among them, B is the base sample obtained in the sampling stage; base i In the sampling stage, count is the frequency of redundancy of data block i in the initial sample; count i In the scanning stage, ρ is the frequency of redundancy of data block i in the entire data set; ρ i is the compression ratio of data block i.

3. A cloud storage redundant data prediction method based on similar data detection according to claim 1, characterized in that the cloud storage data is partitioned to obtain data blocks, including: partitioning the original data of the cloud storage according to a fixed length.

4. A cloud storage redundant data prediction device based on similar data detection, characterized in that the cloud storage redundant data prediction device includes a memory and a processor; a cloud storage redundant data prediction program is stored on the memory and runs on the processor, and when the cloud storage redundant data prediction program is executed by the processor, the steps of the cloud storage redundant data prediction method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Protocol-independent network redundant flow eliminating method

    CN103888317A

  • Big data redundancy detection method

    CN107562794A