System and method for sketch calculation
By calculating unique feature sketches for data segments and using them for in-line deduplication, the method addresses the challenges of large-scale data reduction in conventional techniques, achieving efficient data reduction and improved performance.
Patent Information
- Application Number
- JP2022537814
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-18
- Filing Date
- 2020-12-17
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2040-12-17
AI Technical Summary
Conventional deduplication techniques struggle to efficiently handle large-scale data reduction, especially for data generated by enterprise applications, due to exponential data growth and the computational overhead of comparing large index tables.
The method involves receiving an input data stream, generating segments from it, calculating a sketch for each segment that represents unique features, and using these sketches for in-line deduplication without generating a full index or comparing to a full index, thereby optimizing data reduction at high throughput.
This approach enables efficient data reduction at terabyte and petabyte scales, improving computational performance and reducing operational costs by minimizing the need for extensive indexing and comparison processes.
Smart Images

Figure 0007700128000001 
Figure 0007700128000002 
Figure 0007700128000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications
[0002] This application claims priority to U.S. Patent Application No. 16 / 718,686, filed on December 18, 2019; U.S. Patent Application No. 16 / 718,703, filed on December 18, 2019; and U.S. Patent Application No. 16 / 718,714, filed on December 18, 2019, each of which is hereby incorporated by reference in its entirety.
Background Art
[0003] A cloud storage system can store large amounts of data from client applications, such as enterprise applications. In many cases, a significant portion of the input data may be duplicated. Storing and processing the data may require large amounts of memory, storage space, and processing power. In some cases, data can be reduced before storage, for example, using deduplication or compression techniques. However, large - scale data reduction using small chunks and 1:1 chunk comparisons has been shown to be technically difficult or impractical because it consumes significant memory space (by adding a computational load to both the write and read processes) and degrades performance due to the large index tables generated. As a result, conventional deduplication techniques generally cannot handle data reduction for large amounts of data on the order of hundreds of terabytes or petabytes, especially data generated by enterprise applications.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Due to exponential scaling in data generation, there is a recognized need for methods and systems that can efficiently process large-scale data reduction while maintaining or improving computational performance. Data reduction can lead to a reduction in operational costs related to computing resources and storage.
[0005] The present disclosure provides systems and methods configured to optimally reduce data at high throughput (e.g., at least terabyte level per day, or 100 terabytes per day, or several terabytes per day, etc.) and on a large scale (e.g., at least petabyte scale). The systems and methods herein can be applied to data generated by various client applications, such as enterprise applications. As used herein, the term "data" can refer to any type of data, such as structured data, unstructured data, time-series data, relational data, etc. The term "enterprise application" can refer to a large-scale software system platform developed using enterprise architecture and designed to operate in an enterprise environment such as business or government. Although some embodiments of the present disclosure have been described with respect to enterprise applications, it should be understood that some embodiments herein can be made applicable to or adapted for non-enterprise applications or other small-scale applications.
Means for Solving the Problem
[0006] In one aspect, the present disclosure provides a method for sketch calculation, the method comprising: (a) receiving an input data stream from one or more client applications; (b) generating at least one segment from the input data stream, the at least one segment including a plurality of chunks; (c) calculating a sketch of the at least one segment, the sketch representing the at least one segment or including a set of features unique to the at least one segment such that the set of features corresponds to the at least one segment, the sketch being usable for in-line deduplication of at least one other input data stream received from one or more client applications without (i) generating a full index of the plurality of chunks or (ii) comparing the at least one other input data stream to a full index.
[0007] In another aspect, the present disclosure provides a method for sketch calculation, the method comprising: (a) receiving an input data stream from one or more client applications; (b) generating at least one segment from the input data stream, the at least one segment including a plurality of chunks; (c) calculating a sketch of the at least one segment, the sketch representing the at least one segment or including a set of features unique to the at least one segment such that the set of features corresponds to the at least one segment, the sketch being usable for in-line deduplication of at least one other input data stream received from one or more client applications without (i) generating a full index of the plurality of chunks or (ii) comparing the at least one other input data stream to a full index.
[0008] In another aspect, the present disclosure provides a method for sketch calculation, the method comprising: (a) receiving an input data stream from one or more client applications; (b) generating at least one segment from the input data stream, the at least one segment including a plurality of chunks; and (c) calculating a sketch of the at least one segment, the sketch representing the at least one segment or including a set of features that are specific to the at least one segment such that the set of features corresponds to the at least one segment, the sketch being usable for in-line deduplication of at least one other input data stream received from one or more client applications without (i) generating a full index of the plurality of chunks or (ii) comparing the at least one other input data stream to a full index.
[0009] In another aspect, the present disclosure provides a computer-implemented method for inline data deduplication using sketch computations. The method includes receiving, via a computer network, a first input data stream from one or more client applications; using at least one computer processor to generate a first segment from the first input data stream, the first segment being generated by identifying one or more breaks within the first input data stream for applying a hash function to the first input data stream to generate a plurality of chunks and assembling the plurality of chunks based on a target segment size range to form the first segment; calculating a sketch of the first segment, the sketch representing the first segment or including a set of features unique to the first segment; and using the first sketch for inline data deduplication of a second input data stream received from one or more client applications by determining at least a similarity between the first sketch and a second sketch of a second segment generated from the second input data stream.
[0010] In another aspect, the present disclosure provides a computer-implemented method for inline data deduplication using sketch computations, the method comprising receiving, via a computer network, a first input data stream from one or more client applications; using at least one computer processor to generate a first segment from the first input data stream, the first segment being generated by assembling a plurality of chunks based at least in part on one or more breaks within the first input data stream; calculating a sketch of the first segment, the sketch representing the first segment or including a set of features unique to the first segment; and using the first sketch for inline data deduplication of a second input data stream received from one or more client applications by determining at least a degree of similarity between the first sketch and a second sketch of a second segment generated from the second input data stream.
[0011] In another aspect, the present disclosure provides a method for data processing. The method includes receiving one or more input data streams from one or more client applications; generating at least a first segment and a second segment from the one or more input data streams, wherein the first segment includes a first set of chunks and the second segment includes a second set of chunks; calculating (i) a first set of fingerprints of the first set of chunks and (ii) a second set of fingerprints of the second set of chunks, wherein the first set of fingerprints or the second set of fingerprints includes a plurality of hashes generated using one or more hashing algorithms; comparing the first set of fingerprints with the second set of fingerprints to generate a similarity score; and when the similarity score is greater than or equal to a similarity threshold, processing the first set of chunks and the second set of chunks by performing a difference operation, wherein the difference operation includes at least (i) generating a reference hash set based on the first set of chunks and generating a second set of hashes based on the second set of chunks, (ii) comparing the second set of hashes with the reference hash set, and (iii) generating and storing a single pointer that references a subset of the second set of chunks having hashes that match in the reference hash set.
[0012] In another aspect, the present disclosure provides a method for data reduction, the method comprising: (a) receiving one or more input data streams from one or more client applications; (b) generating at least a first segment and a second segment from the one or more input data streams, wherein the first segment includes a first plurality of chunks and the second segment includes a second plurality of chunks; (c) calculating (i) a first sketch of the first segment and (ii) a second sketch of the second segment, wherein the first sketch represents the first segment or includes a set of first features unique to the first segment, the second sketch represents the second segment or includes a set of second features unique to the second segment, the set of first features corresponds to the first segment, and the set of second features corresponds to the second segment; (d) processing the first sketch and the second sketch to generate a similarity metric indicating whether the second segment is similar to the first segment; (e) selecting a hash function from a plurality of hash functions having different hashing strengths based at least in part on the similarity metric; and (f) subsequent to (e), (1) performing a difference operation on the second segment with respect to the first segment using the hash function selected in (e) when the similarity metric is greater than or equal to a similarity threshold, or (2) storing the first segment and the second segment in a database without performing the difference operation when the similarity metric is less than the similarity threshold.
[0013] In some embodiments, the set of features may include the minimum number of features that can be used to uniquely identify or distinguish at least one segment from another segment. In some embodiments, the minimum number of features may range from about 3 features to about 15 features. In some embodiments, the minimum number of features may include 15 or fewer features.
[0014] In some embodiments, at least one segment can have a size of at least about 1 megabyte (MB). In some embodiments, at least one segment can have a size in the range of about 1 megabyte (MB) to about 4 MB.
[0015] In some embodiments, the plurality of chunks can include at least about 100 chunks. In some embodiments, the plurality of chunks can include at least about 1000 chunks.
[0016] In some embodiments, the plurality of chunks can be of variable length. In some embodiments of the method, step (b) can further include generating a plurality of segments from the input data stream, the plurality of segments including at least one segment. In some embodiments, the segments among the plurality of segments have different sizes in the range of about 1 megabyte (MB) to about 4 MB. In some embodiments, the segments among the plurality of segments can have substantially the same size within the range of about 1 megabyte (MB) to about 4 MB.
[0017] In some embodiments of the method, step (b) can further include generating a fingerprint for each chunk of the plurality of chunks. In some embodiments, the fingerprint can be generated using one or more hashing algorithms. In some embodiments, the fingerprint can be generated using one or more non-hashing algorithms. In some embodiments, the set of features can be associated with a subset of chunks selected from the plurality of chunks.
[0018] In some embodiments, the set of features may include a set of fingerprints for a subset of the chunks. In some embodiments, the set of fingerprints may include a plurality of chunk hashes for a subset of the chunks. In some embodiments, the subset of the chunks may be less than about 10% of the plurality of chunks. In some embodiments, the subset of the chunks may be less than about 1% of the plurality of chunks. In some embodiments, the subset of the chunks may include from about 3 chunks to about 15 chunks.
[0019] In some embodiments, the subset of the chunks may be selected from the plurality of chunks using one or more fitting algorithms for a plurality of hashes generated for the plurality of chunks. In some embodiments, one or more fitting algorithms may be used to determine a minimum hash for each hash function of two or more different hash functions. In some embodiments, the plurality of hashes may be generated using two or more different hash functions. In some embodiments, the two or more different hash functions may be selected from the group consisting of Secure Hash Algorithm 0 (SHA-0), Secure Hash Algorithm 1 (SHA-1), Secure Hash Algorithm 2 (SHA-2), and Secure Hash Algorithm 3 (SHA-3).
[0020] In some embodiments, each feature of the set of features may include a minimum hash for each hash function of two or more different hash functions. In some embodiments, the set of features may include a vector of minimum hashes of two or more different hash functions. In some embodiments, the set of features may be provided as a linear combination of features including the vector.
[0021] In another aspect, the present disclosure provides a method for data processing, the method comprising: (a) receiving one or more input data streams from one or more client applications; (b) generating at least a first segment and a second segment from the one or more input data streams, the first segment including a first set of chunks and the second segment including a second set of chunks; (c) calculating (i) a first set of fingerprints of the first plurality of chunks and (ii) a second set of fingerprints of the second plurality of chunks; (d) processing the first set of fingerprints and the second set of fingerprints to determine that the first set of chunks and the second set of chunks meet a similarity threshold; and (e) processing the first set of chunks and the second set of chunks to determine one or more differences between the first segment and the second segment.
[0022] In some embodiments, the first segment and the second segment may be determined to be similar based at least on a similarity threshold.
[0023] In some embodiments, the similarity threshold may be at least about 50%. In some embodiments, the similarity threshold may indicate the degree of overlap between the first set of chunks and the second set of chunks.
[0024] In some embodiments, the second segment may be approximately the same size as the first segment. In some embodiments, the second segment may be substantially different in size from the first segment.
[0025] In some embodiments, the first segment and the second segment may each have a size in the range of about 1 megabyte (MB) to about 4 MB.
[0026] In some embodiments, the first set of chunks and the second set of chunks may have different numbers of chunks.
[0027] In other embodiments, the first set of chunks and the second set of chunks may have the same number of chunks.
[0028] In some embodiments, the first set of chunks and the second set of chunks may each include at least about 100 chunks. In some embodiments, the first set of chunks and the second set of chunks may each include at least about 1000 chunks.
[0029] In some embodiments, the first set of chunks and the second set of chunks may have variable lengths.
[0030] In some embodiments, the first set of fingerprints may be associated with a first subset of chunks selected from the first set of chunks, and the second set of fingerprints may be associated with a second subset of chunks selected from the second set of chunks. In some embodiments, the first set of fingerprints may include a first plurality of chunk hashes for the first subset of chunks, and the second set of fingerprints may include a second plurality of chunk hashes for the second subset of chunks. In some embodiments, the first subset of chunks may be less than about 10% of the first set of chunks. In some other embodiments, the first subset of chunks may be less than about 1% of the first set of chunks. In some further embodiments, the second subset of chunks may be less than about 10% of the second set of chunks. In some embodiments, the second subset of chunks may be less than about 1% of the second set of chunks.
[0031] In some embodiments, the subset of the first chunks and the subset of the second chunks may have the same number of chunks. In other embodiments, the subset of the first chunks and the subset of the second chunks may have different numbers of chunks. In some embodiments, the subset of the first chunks and the subset of the second chunks may each include from about 3 to about 15 chunks.
[0032] In some embodiments, the subsets of the first and second chunks may be selected from the sets of the first and second chunks using one or more fitting algorithms for a plurality of hashes generated for the sets of the first and second chunks. In some embodiments, the one or more fitting algorithms may include a minimum hash function.
[0033] In some embodiments, the set of the first fingerprints and the set of the second fingerprints may be generated using one or more hashing algorithms. In some embodiments, the one or more hashing algorithms may be selected from the group consisting of Secure Hash Algorithm 0 (SHA-0), Secure Hash Algorithm 1 (SHA-1), Secure Hash Algorithm 2 (SHA-2), and Secure Hash Algorithm 3 (SHA-3). In some embodiments, the set of the first fingerprints and the set of the second fingerprints may be generated using two or more different hashing algorithms selected from the group.
[0034] In some other embodiments, the set of the first fingerprints and the set of the second fingerprints may be generated using one or more non-hashing algorithms.
[0035] In a further aspect, the present disclosure provides a method for data reduction, the method comprising: (a) receiving one or more input data streams from one or more client applications; (b) generating at least a first segment and a second segment from the one or more input data streams, the first segment including a first plurality of chunks and the second segment including a second plurality of chunks; (c) calculating (i) a first sketch of the first segment and (ii) a second sketch of the second segment, the first sketch representing the first segment or including a set of first features unique to the first segment, the second sketch representing the second segment or including a set of second features unique to the second segment, the set of first features corresponding to the first segment and the set of second features corresponding to the second segment; (d) processing the first sketch and the second sketch to generate a similarity metric indicating whether the second segment is similar to the first segment; and (e) following (d), (1) performing a difference operation on the second segment with respect to the first segment if the similarity metric is greater than or equal to a similarity threshold, or (2) storing the first segment and the second segment in a database without performing the difference operation if the similarity metric is less than the similarity threshold.
[0036] In some embodiments, the difference operation of (e) may include: (i) generating a reference hash set for a first plurality of chunks of a first segment; and (ii) storing the reference hash set in a memory table. In some embodiments, the reference hash set may include weak hashes. In some embodiments, the reference hash set may be generated using a hashing function having a throughput of at least 1 gigabyte (GB). In some embodiments, the difference operation may further include: (iii) generating hashes for chunks among a second plurality of chunks of a second segment in a sequential rolling basis; and (iv) comparing the hashes with the reference hash set to determine whether there is a match.
[0037] In some embodiments, the difference operation may further include: (v) continuing to generate one or more other hashes for one or more subsequent chunks among the second plurality of chunks as long as the hash and one or more other hashes find a match from the reference hash set.
[0038] In some embodiments, the difference operation may further include: (vi) generating and storing a single pointer that references the chunk and one or more subsequent chunks when it is detected that the hash of the subsequent chunk cannot find a match from the reference hash set.
[0039] In some embodiments, the hash may be a weak hash.
[0040] In some embodiments, one or more other hashes may include weak hashes. In some embodiments, the hash and one or more other hashes may include weak hashes.
[0041] In some embodiments, for the next chunk of the second plurality, another hash may be generated and compared with the reference hash set to determine if there is a match before determining if the hash matches the reference hash set for the next chunk of the second plurality.
[0042] In some embodiments, when one or more input data streams are received from one or more client applications, a differential operation may be performed inline.
[0043] In some embodiments, the differential operation may reduce the first segment and the second segment into a plurality of homogeneous fragments. In some embodiments, the method may further include storing the plurality of homogeneous fragments in one or more cloud object data stores. In some embodiments, the method may further include generating an index that maps the plurality of uniform fragments to the first segment and the second segment. In some embodiments, the method further includes receiving a read request transmitted from one or more client applications, the read request being for an object that may include at least one of the first segment or the second segment, and reconstructing the first segment or the second segment by at least partially using (1) the plurality of homogeneous fragments stored in one or more cloud object data stores, and (2) the index, to generate an object in response to the read request. In some embodiments, the method may further include providing the generated object to the one or more client applications that transmitted the read request.
[0044] In some embodiments of the method, the processing of step (d) may include comparing a second set of features with a first set of features to determine if one or more features are common to both the first set and the second set.
[0045] In some embodiments, the second segment may be determined to be similar to the first segment if (i) the similarity metric is greater than or equal to a similarity threshold, or (ii) the second segment may be determined to not be similar to the first segment if the similarity metric is less than the similarity threshold. In some embodiments, the similarity threshold may be at least about 50%.
[0046] In some embodiments, the similarity metric may indicate the degree of overlap between the first segment and the second segment. In some embodiments, one or more features may be similar or identical in the first set and the second set.
[0047] In some embodiments, the second segment may be approximately the same size as the first segment.
[0048] In some other embodiments, the second segment may be substantially different in size from the first segment.
[0049] In some embodiments, the first segment and the second segment may each have a size in the range of about 1 megabyte (MB) to about 4 MB.
[0050] In some embodiments, the first set and the second set may each include from about 3 features to about 15 features.
[0051] In some embodiments, the first plurality of chunks and the second plurality of chunks may each include at least about 100 chunks.
[0052] In some embodiments, the first plurality of chunks and the second plurality of chunks may be of variable length.
[0053] In some embodiments of the present method, step (e) may further include storing the first sketch and the second sketch in a database when the similarity metric is less than the similarity threshold.
[0054] In some embodiments, the similarity metric may be a similarity score.
[0055] In some embodiments, the first segment is generated by determining that the sum of the lengths of a plurality of chunks is within a target segment size range, and the target segment size range is from 1 megabyte (MB) to 16 MB.
[0056] In some embodiments, the set of features is associated with a subset of a plurality of chunks selected from the plurality of chunks.
[0057] In some embodiments, the set of features includes a set of fingerprints for a subset of the plurality of chunks.
[0058] In some embodiments, the set of fingerprints includes a plurality of chunk hashes for a subset of the plurality of chunks.
[0059] In some embodiments, the subset of the plurality of chunks includes less than 10% of the plurality of chunks.
[0060] In some embodiments, the subset of the plurality of chunks includes from 3 to 15 chunks.
[0061] In some embodiments, the second segment is formed by assembling a set of chunks based on the target segment size range, and the set of chunks is generated from a second input data stream.
[0062] In some embodiments, the size of the second segment is different from the size of the first segment.
[0063] In some embodiments, the size of the second segment is the same as the size of the first segment.
[0064] In some embodiments, the first segment or the second segment has a size of at least 1 megabyte (MB).
[0065] In some embodiments, the size of the first segment depends on the sum of the lengths of a plurality of chunks.
[0066] In some embodiments, the first segment is generated by determining that the sum of the lengths of a plurality of chunks is within a target segment size range.
[0067] In some embodiments, the plurality of hashes are generated using two or more different hash functions selected from the group consisting of Secure Hash Algorithm 0 (SHA-0), Secure Hash Algorithm 1 (SHA-1), Secure Hash Algorithm 2 (SHA-2), and Secure Hash Algorithm 3 (SHA-3).
[0068] In some embodiments, assembling a plurality of chunks includes generating a plurality of chunks from a first input data stream by applying a hash function to the first input data stream to identify one or more breaks within the first input data stream, and summing the lengths of a subset of the plurality of chunks based on a target segment size range.
[0069] In some embodiments, the second segment is generated by assembling a set of chunks generated from a second input data stream.
[0070] In some embodiments, the first set of fingerprints is calculated based on a plurality of hashes associated with a subset of first chunks selected from the first set of chunks, and the second set of fingerprints is calculated based on a plurality of hashes associated with a subset of second chunks selected from the second set of chunks.
[0071] In some embodiments, the subset of chunks is sequential chunks.
[0072] In some embodiments, when it is determined that a subsequent chunk to the subset of chunks has a hash that does not find a match within the reference hash set, a single pointer is generated and stored.
[0073] In some embodiments, the second hash set is a weak hash.
[0074] In some embodiments, when the similarity metric has a first score, the first hash function is selected, and when the similarity metric has a second score lower than the first score, the second hash function is selected.
[0075] In some embodiments, the second hash function has a higher hashing strength than the first hash function.
[0076] In some embodiments, the strength of the selected hash function varies inversely with the score of the similarity metric.
[0077] Another aspect of the present disclosure provides a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere in this specification.
[0078] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere in this specification.
[0079] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which illustrates only exemplary embodiments of the present disclosure. As will be understood, the present disclosure is capable of other and different embodiments and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive. Incorporation by reference
[0080] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference into this specification to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the incorporated publications and patents or patent applications conflict with the disclosure contained herein, this specification is intended to supersede and / or prevail over such conflicting material.
[0081] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description that illustrates exemplary embodiments in which the principles of the invention are utilized, and from the appended drawings (also referred to herein as “figure” and “FIG.”). BRIEF DESCRIPTION OF THE DRAWINGS
[0082]
Figure 1
[0083]
Figure 2
[0084]
Figure 3
[0085]
Figure 4
[0086]
Figure 5
[0087]
Figure 6A
Figure 6B
[0088]
Figure 7A
Figure 7B
[0089]
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 8E
[0090]
Figure 9
[0091]
Figure 10
[0092]
Figure 11
[0093]
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0094] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art can conceive of numerous variations, modifications, and substitutions without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be used.
[0095] When the terms "at least", "greater than", or "above" are in front of the first numerical value of a series of two or more numerical values, the terms "at least", "greater than", or "above" always apply to each numerical value of that series of numerical values. For example, 1, 2, or 3 or more is equivalent to 1 or more, 2 or more, or 3 or more.
[0096] When terms such as "slightly", "less than", or "below" are in front of the first of a series of two or more numerical values, the terms "slightly", "less than", and "below" always apply to each numerical value in the series. For example, 3, 2, or 1 or below is equivalent to 3 or below, 2 or below, or 1 or below.
[0097] As used herein, the term "real time" generally refers to the simultaneous or substantially simultaneous occurrence of a first event or action relative to the occurrence of a second event or action. A real-time action or event can be executed within a response time of less than 1 second, less than one-tenth of a second, less than one-hundredth of a second, less than a millisecond, or less than that, at least relative to another event or action. Real-time actions can be executed by one or more computer processors.
[0098] As used herein, the term "sketch" generally refers to the fingerprint of at least one data unit, such as at least one data segment. A sketch may be used to describe or characterize the (one or more) data segments of a file or object. A sketch may include a set of features that can be used to describe or characterize the (one or more) data segments.
[0099] Terms such as "weak hashing algorithm" as used herein generally refer to a hashing algorithm that maximizes the number of data chunks hashed per unit time at the expense of reducing the total number of collisions of the hashed data chunks. A collision can occur when a hashing algorithm generates the same hash value for different data chunks.
[0100] As used herein, terms such as "strong hashing algorithm" generally refer to a hashing algorithm that sacrifices maximizing the number of hashed data chunks hashed per unit time to minimize the total number of collisions of hashed data chunks. A collision can occur when a hashing algorithm generates the same hash value for different data chunks. Summary
[0101] Data reduction can be a process of reducing the capacity required to store data. The data reduction system described herein can, among other advantages, increase storage efficiency, improve processing / calculation speed performance, and reduce costs.
[0102] Conventional methods of handling data reduction in a data storage system generally rely on segmenting data into chunks, generating a fingerprint (e.g., a hash) for the chunks of data, and storing the fingerprints in an in-memory table. After the fingerprints are calculated, a lookup can be performed in memory to compare a chunk with a new chunk. If the fingerprints match in memory and the new chunk is considered unique, the new chunk can be stored. If the new chunk is considered the same, a pointer to the first stored chunk can be stored. This process requires a significant amount of storage space to store each fingerprint. Further, this process requires a significant amount of processing and calculation time to generate each fingerprint for each chunk and then compare each fingerprint corresponding to each chunk with another fingerprint of a different chunk. The calculation and processing of the fingerprint for each chunk is time-consuming and can be computationally expensive.
[0103] The data reduction systems and methods described herein can at least address the drawbacks of conventional data deduplication techniques. For example, instead of performing direct comparisons of fingerprints (hashes) between all of the individual data chunks, the data reduction systems and methods provided herein use sketches to describe or characterize large segments of data within a file or object, and compare the sketches to determine whether two or more segments are of the same kind (e.g., similar). A sketch can include a set of features that can be used to describe or characterize large segment data. If the sketches of two or more segments are determined to be substantially similar, then the sketches can then be differentiated at a finer level (e.g., feature level, chunk level, etc.), at which point fingerprint comparison can be performed. Instead of pointers to individual chunks, a single pointer may be generated for a group of chunks whose fingerprints match. If the sketches of two or more segments are determined to be substantially different, the segments and their sets of features can be stored in a database, eliminating the need for chunk-level differentiation and saving computational resources for deduplication of other similar segments.
[0104] A data reduction system 1040 according to some embodiments herein may exist within an ecosystem, for example, as shown in FIG. 10. The ecosystem may include one or more client applications 1010, and one or more storage modules 1020 and 1030.
[0105] The data reduction system 1040 described in this specification may comprise one or more modules. As shown in FIG. 11, the modules may include, for example, a sketch calculation module 600, a sketch comparison module 700, a difference operation module 800, a data reconstruction module 900, a data chunking module 100, a data segmentation module 300, a variable segment sizing module 500, or various combinations thereof. The functions of each module can generally be described as follows.
[0106] The sketch calculation module may be configured to calculate one or more sketches for one or more data segments generated from one or more input data streams. After a sketch is generated for a segment, the sketch comparison module can determine whether the sketch of the new segment is substantially similar to the sketch of the previous segment. If the two sketches are determined not to be substantially similar, the new segment may be stored in the database. If the two sketches are determined to be substantially similar, the difference operation module is utilized to compare the chunks between the segments and determine whether the segment has one or more overlapping chunks common to both segments. The difference module may be configured to generate a sparse index array and store pointers to blocks of overlapping chunks. The blocks of overlapping chunks may be stored in the database as homogeneous fragments. When a read request is received from a client application, the data reconstruction module can reconstruct the requested object (requested by the client application) using the sparse index array and the homogeneous fragments generated by the difference module. Additional aspects regarding the partitioning of data using the data chunking module, the data segmentation module, or the variable segment sizing module are described in more detail elsewhere in this specification. I. Sketch Calculation
[0107] In one aspect, a method for sketch calculation is provided. A sketch may be a data structure that supports a set of pre-specified queries and updates to a database. A sketch may consume less memory space compared to storing all information of the entire segment. A sketch may be a fingerprint of the segment. Sketch calculation can be used to reduce memory requirements and speed up the data writing and reading processes. Sketch calculation may include generating a set of features from at least one segment. The set of features can be used as an approximate identifier of the segment (e.g., as a fingerprint) so that similar segments can be identified using those features (or a subset of those features). As described elsewhere in this specification, a sketch can be calculated by determining a set of features using a hashing algorithm (e.g., a hashing function) and / or other algorithms (e.g., non-hashing algorithms). A sketch may represent one or more data segments. In some embodiments, a sketch may be a metadata value. A sketch may be utilized to find matching or similar sketches associated with other segments from one or more input data streams. A sketch may be utilized to find matching or similar sketches associated with previously processed segments.
[0108] Sketches can be calculated using the sketch calculation module 600. An example of the sketch calculation module and sketch calculation is shown with reference to FIGS. 6A and 6B. As shown in FIG. 6B, an input data stream (610) can be used for sketch calculation. The method may include receiving an input data stream from one or more client applications (step 601). The input data stream may include a series of data made available over time. The input data stream may be a series of digitally encoded coherent signals (e.g., data packets, data packets, network packets, etc.) used to transmit or receive information in the process being transmitted. The input data stream may include data, data packets, files, objects, etc. The input data stream may include a set of extracted information. The input data stream may include raw data (e.g., unprocessed data, unstructured data, etc.). The input data stream may include structured data. The input data stream may be, for example, network traffic, a graph stream, a client application data stream, or a multimedia stream. The input data stream may include at least one segment.
[0109] The client application may be an application configured to run on a workstation or a personal computer. The workstation or personal computer may be within a network. The client application may include an enterprise application. In some embodiments, the enterprise application may be a large-scale software system platform designed to operate in a corporate environment. The enterprise application may be designed to interface or integrate with other applications used within an organization or without other applications. The enterprise application may be computer software used to meet the needs of an organization rather than individual users. Such organizations may include, for example, businesses, governments, and the like. The enterprise application may be an essential part of a (computer-based) information system. The enterprise application may support, for example, data management, business intelligence, business process management, knowledge management, customer relationship management, databases, enterprise resource planning, enterprise asset management, low-code development platforms, supply chain management, product data management, product life cycle management, networking and information security, online shopping, online payment processing, interactive product catalogs, automated billing systems, security, business process management, enterprise content management, IT service management, customer relationship management, enterprise resource planning, business intelligence, project management, collaboration, human resource management, manufacturing, occupational health and safety, enterprise application integration, information storage or enterprise form automation, etc. The complexity of the enterprise application may require special functions and specific knowledge.
[0110] As shown in FIG. 6B, the method may include generating at least one segment from an input data stream (step 602). The method may further include calculating a sketch of the at least one segment, as described elsewhere herein. As shown in FIG. 6B, one or more segments 620-622 within the input data stream 610 can be generated (step 602). In some embodiments, multiple segments can be generated from the input data stream. The multiple segments can include at least about 1, 5, 10, 15, 25, 100, 1000, 10000 or more segments. The multiple segments can have a size of at least about 1 kilobyte (KB), 10 KB, 100 KB, 500 KB, 1 megabyte (MB), 2 MB, 3 MB, 4 MB, 5 MB, 6 MB, 7 MB, 8 MB, 9 MB, 10 MB, or more. The multiple segments can have a size of up to about 10 MB, 9 MB, 8 MB, 7 MB, 6 MB, 4 MB, 3 MB, 2 MB, 1 MB, 500 KB, 100 KB, 10 KB, or less. The multiple segments can have a size of about 100 KB to 10 MB, 500 KB to 5 MB, or 1 MB to 4 MB. In some embodiments, each of the multiple segments can have a size in the range of about 1 MB to about 4 MB. In some embodiments, the multiple segments can have different sizes in the range of about 1 MB to about 4 MB.
[0111] As described elsewhere in this specification, segments can be generated from an input data stream. Each segment can include a plurality of chunks. As shown in FIG. 6B, the segment may be converted into a plurality of chunks 630. The segment may be converted into, for example, 1000 chunks. A chunk can contain data. A chunk may be a fragment of information. A chunk may be a unit of information. A chunk can contain a header. The header can indicate the parameters of the chunk. The parameters can include, for example, the type, comment, size, etc. of the chunk. The process of obtaining a segment and generating one or more chunks can be referred to as chunking. Chunking can include dividing the data within a data segment into several sections (e.g., chunks) of continuous data (e.g., from an input data stream). For example, chunking can be used to reduce the overhead of a central processing unit (CPU) or to reduce latency. In some embodiments, a segment can include at least about 5, 10, 15, 25, 100, 1000, 10000 or more chunks. A segment can include from about 2 to 10000 chunks, from 10 to 1000 chunks, or from 25 to 100 chunks. In some embodiments, a segment can include at least about 1000 chunks. The plurality of chunks may be of the same length or of variable length. The plurality of chunks may be of the same data size. The plurality of chunks may be of different data sizes. The plurality of chunks can have a size of at least about 0.1 kilobyte (KB), 0.5 KB, 1 KB, 2 KB, 3 KB, 4 KB, 5 KB, 6 KB, 7 KB, 8 KB, 9 KB, 10 KB, 50 KB or more. The size of the plurality of chunks may be in the range of 0.1 KB to 10 KB, 0.5 KB to 7 KB, or 1 KB to 4 KB. In some embodiments, the plurality of chunks can have a size in the range of 4 KB to 16 KB.
[0112] The method may further include generating a fingerprint for each of a plurality of chunks using step 604, as further shown in FIG. 6B. The fingerprint may be utilized to identify a particular data chunk. In some embodiments, the fingerprint may be generated using one or more hashing algorithms. The fingerprint may include one or more hash values generated by one or more hashing algorithms. As shown in FIG. 6B, the hashing algorithm may be executed on the chunks (e.g., 1000 chunks) to generate a plurality of hash values 640 (e.g., 1000 hash values). Each hash value may be a fingerprint associated with each respective chunk. In some embodiments, two or more hash values may be calculated for a particular chunk.
[0113] The hashing algorithms (e.g., hash (hashing) functions) described in this specification may include any method that can be used to map data of any size to a fixed-size value. The value returned by a hash function may be called a hash value, hash code, digest, or hash. These values may be used to index into a fixed-size table called a hash table. In some cases, encryption-grade hash functions may be used to generate fingerprints. Encryption-grade hash functions may be keyed, unkeyed, or use a combination thereof.The hash function may be selected from the group consisting of SHA-0, SHA-1, SHA-2, SHA-3, SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, SHA-512 / 256, SHA3-224, SHA3-256, SHA3-384, SHA3-512, SHAKE128, SHAKE256, BLAKE-256, BLAKE-512, BLAKE2s, BLAKE2b, BLAKE2X, ECOH, FSB, GOST, Grostl, HAS-160, HAVAL, JH, LSH, MD2, MD4, MD5, MD6, RadioGatun, RIPEMD, RIPEMD-128, RIPEMD-160, RIPEMD-320, Skein, Snefru, Spectral Hash, Streebog, SWIFFT, Tiger, Whirlpool, HMAC, KMCA, One-key MAC, PMAC, Poly1305-AES, SipHash, UMAC, VMAC, Pearson hashing, Paul Hsieh’s SuperFastHash, Buzhash, Fowler-Noll-Vo hash function, Jenkins hash function, Bernstein hash djb2, PJW hash, MurmurHash, Fast-Hash, SpookyHash, CityHash, FarmHash, MetroHash, number hash, xxHash, t1ha, cksum (Unix), CRC-16, CRC-32, Rabin fingerprint, aggregating hashing, general one-way hash functions, and Zobrist hashing. Additionally or alternatively, the fingerprint may also be generated using one or more non-hashing algorithms.
[0114] The method may further include steps of generating a plurality of features of the segment. The sketch of the segment may represent the segment or include a set of features (e.g., characteristics) unique to the segment. Feature generation or extraction can reduce the amount of resources required to describe the segment. The features may describe the most relevant information from the segment. The features of the segment may not change even when small deformations are introduced into the chunks. The features may describe relevant information from the segment so that the desired task (e.g., chunk comparison) can be performed using a reduced representation (e.g., sketch comparison of features) instead of using the entire set of chunks. The features may include, for example, specific items associated with the chunks within the segment. The items may include, for example, hash values generated by one or more hashing algorithms. The items may include, for example, integers (e.g., ID numbers, hash values), data types, file extensions, etc.
[0115] The set of features in the sketch of the segment may include the minimum number of features that can be used to uniquely identify or distinguish the segment. The set of features may include at least about 1, 2, 3, 4, 5, 10, 15, 25, 100, 100 or more features. The set of features may include at most about 100, 100, 25, 15, 10, 5, 4, 3, 2, or fewer features. The set of features may include about 1 to 100, 2 to 25, 3 to 15, or 5 to 10 features. In some embodiments, the set of features in the sketch of the segment may be in the range of about 3 to about 15 features. The set of features may include a linear combination of features. The set of features may be used to approximate the segment.
[0116] In some embodiments, the set of features may be associated with a subset of chunks selected from a plurality of chunks. The set of features may include a set of fingerprints for the subset of chunks. The set of fingerprints may include chunk hashes for the subset of chunks. The subset of chunks may be less than about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50% of the plurality of chunks. The subset of chunks may be about 1% - 50%, 5% - 40%, or 10% - 25% of the plurality of chunks. The subset of chunks may include at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, or more chunks. The subset of chunks may be selected from the plurality of chunks using one or more fitting algorithms for a plurality of hashes generated for the plurality of chunks. In some embodiments, the plurality of hashes may be generated using two or more different hash functions. The two or more different hash functions may include any of the hash functions as described elsewhere in this specification.
[0117] In some embodiments, one or more fitting algorithms may be used to determine the minimum hash of the entire set of chunks or a subset of chunks (step 605). The feature may correspond to the minimum hash value of the entire set of chunks or a subset of chunks. As shown in FIG. 6B, the fitting algorithm may be used to obtain the minimum hash value of 1000 hash values (e.g., fingerprints; 650) corresponding to a particular chunk. One or more fitting algorithms may be used to calculate the minimum hash value when the hash value of each chunk is generated. One or more fitting algorithms may be used to calculate the minimum hash value after all the hash values of the entire set of chunks have been generated. The minimum hash value of the set of chunks may be a feature (F0, 650) that can be used to represent or describe 1000 chunks. In some embodiments, one or more fitting algorithms may be used to determine the minimum hash for each of two or more different hash functions. In some cases, each feature of the set of features may include the minimum hash for each of two or more different hash functions. The set of features may include a vector of minimum hashes of two or more different hash functions.
[0118] In some embodiments, one or more fitting algorithms may be used to determine the maximum hash of the entire set of chunks or a subset of chunks. The feature may correspond to the maximum hash value of the entire set of chunks or a subset of chunks. In some cases, the maximum hash value may be used, additionally or alternatively, together with the minimum hash value for determining the set of features. In some cases, the feature vector may include one or more features generated from the minimum hash value and / or one or more features generated from the maximum hash value. In some cases, the feature may be a linear combination of one or more hashes generated by one or more hashing algorithms for one or more chunks.
[0119] For example, as shown in FIG. 6B, one or more hashing algorithms can be used to generate one or more features (e.g., F0, F1, F2, ..., F i ). For example, the SHA-2 hashing algorithm can be used to generate 1000 SHA-2 hash values. The minimum hash value in the set of SHA-2 hash values can be used to generate the first feature (F0). Next, the MD2 hashing algorithm can be used to generate 1000 MD2 hash function values. The minimum hash value in the set of MD2 hash values can be used to generate the second feature (F1). The features can be generated simultaneously. The features can be stored in a database or as described elsewhere in this specification. By storing features instead of the individual fingerprints (e.g., all hash values) of the chunks, the size of the data storage and memory requirements can be reduced. For example, in the case of a segment containing 1000 data chunks, 10 features can represent the entire segment (or the entire set of chunks). Instead of storing 1000 hash values for 1000 individual data chunks, the system described herein only needs to store 10 features, and thus, the memory storage can be reduced by three digits. In some cases, one or more features may be associated with one or more specific chunks within the set of chunks. For example, a specific chunk (e.g., chunk 1,1) can have a hash value that is the minimum hash value in the set.
[0120] As shown in FIG. 6B, a feature vector 670 can be generated by combining features. Using the features, a sketch 680 can be generated (step 606). The sketch 680 can include a set of features. The sketch can include one or more feature vectors. The sketch can be compared with one or more other sketches as described elsewhere in this specification. In some cases, the sketch may be calculated using, for example, a spatio-temporal sketching algorithm, a Count Sketch, a Count-min Sketch, a conservative update sketch, a Count-Min-Log Sketch, a Slim-Fat Sketch, or a Weight-Median Sketch. In some embodiments, the sketch may be generated using a similarity hashing algorithm or a similar function.
[0121] The sketch may be usable for in-line deduplication of at least one other segment from an input data stream received from one or more client applications. By using the sketch for in-line deduplication, reduction of large amounts of data (e.g., on the petabyte scale) can be achieved. The sketch may be usable for in-line deduplication without requiring a full index of multiple chunks. The sketch may be usable for in-line deduplication without having to look up all chunks in at least one other input data stream in a full index. II. Sketch Comparison
[0122] Sketch comparison can be performed using a sketch comparison module 700, as shown, for example, in FIG. 7A. The method may further include generating at least one other segment from at least one other input data stream (step 701). FIG. 7B shows a first input data stream 710 and a second input data stream 715 that can be used to generate a first segment 720 and a second segment 725. In some cases, the first segment 720 and the second segment 725 may be generated from the same data input stream. The second segment 725 can be generated using the methods described elsewhere in this specification. The method may further include calculating a sketch of the second segment (step 702). FIG. 7B shows a comparison of the sketch of the first segment 730 and the sketch of the second segment 745 (step 703). The sketch of the second segment can be calculated as described elsewhere in this specification. The sketch of the second segment may represent the second segment or include another set of features (e.g., characteristics) that are unique to the second segment. FIG. 7B shows that the sketch 730 of the first segment may include a set of features 740 and the sketch 735 of the second segment may include a set of features 745. The above features can be generated as described elsewhere in this specification.
[0123] The method may further include processing the first sketch and the second sketch, at least partially based on a similarity score, to determine whether the second segment is probabilistically similar to the first segment (step 704). Processing may include comparing a first set of features to a second set of features to determine whether one or more features are common to both sets. As shown in FIG. 7B, a sketch 730 of the first segment and a sketch 735 of the second segment may be compared (750) to determine features that are in both sets 764 and / or features that may not be shared by both sets (762, 766). Each sketch may include a different number of features. For example, sketch 730 may include 10 features while sketch 735 may include 6 features. Each sketch may include the same number of features. For example, sketch 730 may include 10 features while sketch 735 may include 10 features.
[0124] If the similarity score is greater than or equal to the similarity threshold, the second segment may be determined to be probabilistically similar to the first segment. If the similarity score is less than the similarity threshold, the second segment may be determined to not be probabilistically similar to the first segment. As shown in FIG. 7B, when the similarity score is greater than or equal to the threshold (770), the first segment and the second segment may be considered similar. In some cases, when the first segment and the second segment are considered to be similar, the method may further include performing a difference operation. The difference operation may be as described elsewhere in this specification. When the similarity score is less than or equal to the threshold (775), the first segment and the second segment may be considered not similar. In some cases, when the first segment and the second segment may be considered not similar, one or more chunks of the second segment may be stored in a database (step 705). The similarity threshold may be at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more. The similarity score threshold may be about 5% - 99%, 10% - 90%, 20% - 80%, 30% - 70%, or 40% - 50%. In some embodiments, the similarity threshold may be at least about 50%.
[0125] The similarity score may indicate the degree of overlap between the first segment and the second segment. In some cases, for example, if the first segment has 10 features, the second segment has 8 features, and it is found that 6 features match in both sets, the similarity score may be 50% (e.g., 6 matching features / 12 unique features). In some cases, the similarity score may be calculated for a specific segment. For example, if the first segment has 10 features, the second segment has 8 features, and it is found that 6 features match in both sets, the similarity score may be 6 / 10 (i.e., 60%) or 6 / 8 (i.e., 75%) respectively. In some cases, the number of features in the sketch of the first segment and the number of features in the sketch of the second segment may be the same. In some cases, the number of features in the sketch of the first segment and the number of features in the sketch of the second segment may be different.
[0126] The similarity score may indicate the number of matching features between the first segment and the second segment. As shown in FIG. 7B, the matching feature 764 can be seen in both sketches. One or more features may be similar or identical in both sets. One or more features may not be shared by both sets, or may not be common to both sets. One or more features may be a combination of matching and non-matching features between the first segment and the second segment. The second segment may be the same size as the first segment. The second segment may be a different size from the first segment. Each of the first segment and the second segment may have a size in the range of about 1 megabyte (MB) to about 4 MB. Each of the first segment and the second segment may have a size in the range as described elsewhere in this specification. III. Differential Calculation
[0127] The data difference operation may be performed using the difference operation module 800, as shown, for example, in FIG. 8A. The method may further include storing the features of the first segment and the second segment in a database. As shown in FIG. 8B, a sketch of a segment (e.g., a set of features of the segment, 810) and a chunk corresponding to the sketch may be stored in a database (820, 840). The method may further include storing the second segment and its set of features (e.g., sketch, 830) in the database if the similarity score is less than the similarity threshold. For example, if the similarity score between a sketch and another sketch is 15% and the threshold is set at 40%, the second segment and its set of features can be stored in the database. As shown in FIG. 8B, when two sketches are compared (e.g., 810 vs. 830) and the similarity score is below the threshold, the features of the sketch and chunk corresponding to that segment may be stored in database 840. In some cases, the databases may be the same. In some cases, the databases may be different.
[0128] The method may further include performing a difference operation between the first segment and the second segment when the similarity score is greater than or equal to the similarity threshold. For example, if the similarity score between a sketch and another sketch is 65% and the threshold is set at 40%, the chunks of the first segment and the chunks of the second segment may be differenced. As shown in FIG. 8C, when the similarity scores of the first segment and the second segment are greater than or equal to the threshold, the individual chunks of both segments can be compared (e.g., 850 vs. 860). The difference operation may include generating a reference hash set (870) for a plurality of chunks of the first segment (step 801). The hashes of the plurality of chunks of the first segment may be generated using one or more hashing algorithms as described elsewhere in this specification (step 802). The hashes of the plurality of chunks may be pre-generated hashes. The method may include storing the reference hash set in a memory table.
[0129] A reference hash set may include weak hashes and / or strong hashes. The strength of a hash may depend on the hashing algorithm. A weak hashing algorithm may generate one or more weak hashes. A strong hashing algorithm may generate one or more strong hashes. Weak and / or strong hashes may be generated using one or more hashing algorithms described elsewhere in this specification. A weak hashing algorithm may be a hashing algorithm with weak collision resistance. Weak collision resistance may indicate that the probability of not being able to find a collision is non-negligible. A strong hashing algorithm may be a hashing algorithm with strong collision resistance. Strong collision resistance may indicate that the probability of not being able to find a collision is negligible. A strong hashing algorithm may make it difficult to find inputs that map to the same hash value. A weak hashing algorithm may make it easier to find inputs that map to the same hash value than a strong hashing algorithm. A weak hashing algorithm may be more likely to cluster hash values (e.g., the mapping of keys to the same hash value) than a strong hash function. A strong hash function may have a uniform distribution of hash values.
[0130] The strength of a hashing algorithm (e.g., from weak to strong) may be on a gradient scale. The strength of a hashing algorithm may depend on the time scale of using the hashing algorithm, the complexity of the hashing algorithm, the implementation of the hashing algorithm, the benchmark of the central processing unit, or the cycles per byte, etc. The strength of a hashing algorithm may be determined using one or more statistical tests. One or more statistical tests may measure, for example, whether a hash function can be easily distinguished from a random function. The test may be, for example, determining whether a hash function exhibits an avalanche effect. The avalanche effect may be an effect in which any single-bit change in the input key affects half of the bits on average in the output.
[0131] A weak hashing algorithm may maximize the number of data chunks hashed per unit time. A weak hashing algorithm may maximize the number of data chunks hashed per unit time at the expense of reducing the total number of collisions of the hashed data chunks. A collision may occur when the hashing algorithm generates the same hash value for different data chunks. A strong hashing algorithm may minimize the total number of collisions of the hashed data chunks. A strong hashing algorithm may minimize the total number of collisions of the hashed data chunks at the expense of maximizing the number of hashed data chunks hashed per unit time. A collision may occur when the hashing algorithm generates the same hash value for different data chunks.
[0132] The reference hash set may be generated using a high-throughput hashing function having a throughput of at least 1 gigabyte. In some cases, the high-throughput hashing function may be a weak hashing algorithm. The hashing algorithm may be a hashing algorithm as described elsewhere in this specification. The similarity between two sketches / segments may determine the strength of the hashing algorithm used by the high-throughput hashing function. For example, if two sketches / segments have a particular similarity score, a particular hashing algorithm may be selected over another hashing algorithm. For example, if a high similarity score is considered likely between two sketches of two segments, a weak hashing function may be used. Since sketch comparison may be a first approximation in quantifying similarity, a weak hashing function may be used (e.g., the sketch may assist in determining that two segments are similar, and as a result, a weak hash may be used). Conversely, if the similarity score of the set of features between two segments is low, a strong hashing function may be used. In some cases, if two sketches have a low similarity score, a hashing function may not be used.
[0133] In some embodiments, when the similarity score is in the range of, for example, 70% to 90%, the first hashing algorithm may be used. In some cases, when the similarity score is greater than, for example, 90%, the method may use a second hashing function different from the first hash function. When the similarity score is less than 70% but greater than 50%, the method may use a third hashing function. When the similarity is less than 50%, for example, the method may not use a hashing algorithm and instead store two sketches (e.g., features) and two segments in the database. In some cases, after comparing the sketches of the two segments, if the two segments are considered not to be substantially similar, the advantage of distinguishing the two segments may be slight.
[0134] In some cases, various one or more parameters may be changed to assist the hashing function to maximize the hashing throughput capacity. For example, the parameter may be used to reduce the number of clock cycles required to generate the hash value and to adjust the hash value memory footprint, or the data word size, etc. The hash may be calculated iteratively. The hash may be calculated iteratively by adjusting the byte size given to the hashing algorithm. The byte size may be at least about 1 byte, 2 bytes, 4 bytes, 8 bytes, 16 bytes, 32 bytes, 64 bytes, or more. The byte size may be at most about 64 bytes, 32 bytes, 16 bytes, 8 bytes, 4 bytes, 2 bytes, 1 byte, or less. The byte size may be from about 1 byte to 64 bytes, from 1 byte to 16 bytes, or from 1 byte to 4 bytes.
[0135] In some embodiments, the performance of a hashing algorithm for high-throughput hash value generation may depend on the throughput data size (e.g., gigabytes). The performance of the hashing algorithm may depend on the throughput speed of the data for hash value generation (e.g., gigabytes per second). The performance of the hashing algorithm may depend on the strength of the hashing algorithm. For example, since a weak hashing function generally generates hash values faster than a strong hashing function, a weak hashing algorithm may provide a performance improvement if faster hash value generation is desired.
[0136] In some embodiments, the differential operation may further include generating a hash for a chunk among a plurality of chunks of the second segment in a sequential rolling basis. For example, the hash values of 1000 chunks may be generated by generating a first hash value of a first chunk until 1000 hash values are generated or a subset of 1000 hash values are generated, and then generating a second hash value of a second chunk. As shown in FIG. 8D, the hash value (HC1,861) of the chunk in the second segment may be calculated at an initial time (e.g., t1). The second hash value (HC2,862) of the chunk within the second segment may be calculated at a time after t1 (e.g., t2). The chunks may be adjacent to each other. Alternatively, the chunks need not be consecutive chunks (e.g., 861 vs. 863). In some cases, before generating another hash for the next chunk among the plurality of chunks of the second segment, the hash may be compared with a reference hash set to determine whether there is a match. In some cases, all the hashes of the plurality of chunks of the second segment may be generated simultaneously.
[0137] The differential operation may further include a step (step 803) of comparing the hash with a reference hash set to determine whether there is a match. The differential operation may further include a step (steps 802-804) of continuously generating one or more other hashes for one or more subsequent chunks of the plurality of chunks as long as the hash and one or more other hashes find a match from the reference hash set. The differential operation may further include a step (step 805) of generating and storing a single pointer that references the chunk and one or more subsequent chunks when it detects that the hash of the subsequent chunk cannot find a match from the reference hash set. The hash may be a weak hash as described elsewhere in this specification.
[0138] As shown in FIG. 8E, the second segment 860 may include a plurality of chunks that may be in a sequential order. The first hash value 871 may be compared with a reference hash value 880. The reference hash may be a hash generated from any segment prior to hashing of a subsequent segment. The reference hash may be a hash from the first segment. The reference hash may be a hash stored in a database. The reference hash value may be equivalent to the first hash value. In some cases, instead of generating a pointer at this point, sequential chunks may be inspected. If a sequential chunk (e.g., 872) has the same hash value as the reference hash value, the method may include a step of continuously checking the hash values (e.g., 871 to 874) of each chunk until a mismatch occurs (e.g., the hash value does not match the reference hash value, 874). At this time, a pointer may be stored with reference to each subsequent chunk (871 to 873). Storing a pointer following sequential chunk analysis reduces the number of pointers that need to be stored, which in addition to reducing memory usage, reduces the number of pointers that need to be accessed and may lead to an improvement in computational speed.
[0139] When one or more input data streams are received from one or more client applications, a difference operation may be executed inline. In some alternative embodiments, the difference operation may be executed offline. For example, the difference operation may be executed offline after one or more segments are stored in a database. The difference operation may be used to reduce the first segment and the second segment into a plurality of homogeneous fragments. The plurality of homogeneous fragments may be stored in one or more cloud object stores. The difference operation may be used to generate a sparse index including a set of reduced pointers. The set of reduced pointers may include a single pointer that refers to a series of sequential chunks. Using homogeneous fragments can reduce the number of chunks that need to be stored after the difference operation, thereby reducing the memory storage requirements. IV. Data Reconstruction
[0140] The reconstruction of data from homogeneous fragments may be performed using a data reconstruction module 900, as shown in FIG. 9A for example. The method may further include the step of receiving a read request from one or more client applications (step 910). The read request may be for an object including the first segment and / or the second segment. The method may further include the step of reconstructing the first and / or second segments using at least partially the plurality of homogeneous fragments and the sparse index to generate an object in response to the read request (step 920). The homogeneous fragments may include one or more data chunks, as described elsewhere in this specification.
[0141] The method may further include providing the reconstructed object to one or more client applications (step 930). The read request may utilize a sparse array index to quickly reconstruct or reorganize the object. The sparse index array may indicate each homogeneous fragment for reconstructing the object requested by the client application. Since the data reconstruction module can reconstruct the object using the sparse index and homogeneous fragments (e.g., a set of chunks) as opposed to all of the individual chunks, processing time and computing power can be saved. V. DATA CHUNKING
[0142] The input data stream described herein may be segmented into variable-sized segments. The segments of the data stream may be determined by pre-chunking the data stream into a set of chunks that can be assembled into one of the segments. Each segment may be deduplicated without wasting extra space. A deduplication chunk algorithm may be used to generate segments that include integer chunks. For example, by using a sliding window analysis of the data stream, chunks can be identified by finding natural breaks in the data stream to support chunks of 4 kilobytes (kB) to 16 kB. In this example, a natural break can be generated by calculating the hash of a 16-byte region and determining whether the hash has a pattern with the last 13 bits in the pattern being zero. The chunks may be further assembled into segments within a target range (e.g., 1 megabyte (MB) to 8 MB, 2 MB to 16 MB, or some other range).
[0143] FIG. 1 is a block diagram of a data chunking module 100 that can be used to deduplicate data before storing the data. In FIG. 1, data chunking module 100 can include a data storage 110 that can be used to store deduplicated chunks (e.g., chunks 104A, 104C, and 104D). Data storage 110 can be any type of data storage system that can deduplicate and / or store data (e.g., a storage system including a hard disk drive, a solid state drive, a memory, an optical drive, a tape drive, and / or another type of system capable of storing data; a distributed storage system; a cloud storage system; and / or another type of storage system). The data storage system can be a physical or virtual data storage system.
[0144] To deduplicate data stream 108, data chunking module 100 can split data stream 108 into a set of data blocks 102A-102C. For example, three data blocks 102A-102B are shown, and in alternative embodiments, more or fewer data blocks 102A-102C can exist. The size of the data blocks can be in the range of 1 MB to 16 MB (e.g., in the range of 1 MB to 8 MB, in the range of 2 MB to 16 MB, or some other range), and the size of the data blocks can be larger or smaller. In some cases, the data blocks may be evenly split, and each data block 102A-102C may have the same fixed size.
[0145] The deduplication component 106 deduplicates the data blocks 102A-102C by dividing each data block into smaller chunks 104A-104E, and can determine whether each of the chunks 104A-104E can currently be stored in the data storage 110. For example, for each of the chunks 104A-104E, the system 110 can calculate the fingerprint of that chunk 104A-104E. In this embodiment, the fingerprint may be a mechanism used to uniquely identify each chunk 104A-104E. The fingerprint can be an encrypted hash function, such as one of the Secure Hashing Algorithm (SHA) (e.g., SHA-1, SHA-256, etc., and / or another type of cryptographic hash function). The fingerprint of each of the chunks 104A-104E can uniquely identify the chunks 104A-104E (assuming no data collisions in fingerprint calculation). The fingerprint can be used to determine whether one of the chunks 104A-104E is currently stored in the data storage 110. The system 110 can store the chunk fingerprints in a database. For each of the chunks 104A-104E that can be stored, the data chunking module 100 can calculate the fingerprint for the chunk (e.g., chunk 104A) and determine whether that fingerprint exists in the fingerprint database. If the newly calculated fingerprint is not in the database, the system 100 can store the corresponding chunk. If the chunk fingerprint matches one of the fingerprints in the database, a copy of this chunk can currently be stored in the data storage 110. In this case, the system 100 may not need to store the chunk. Instead, the system can increment the count of the number of references to this chunk in the data storage and store the reference to that chunk. The reference count can be used to determine when that chunk can be deleted from the data storage 110.As shown in FIG. 1, since chunks 104A, 104C, and 104D are currently stored in data storage 110, system 100 may store chunks 104B and 104E for data block 102A. This is because data blocks 104A, 104C, and 104D may already be stored in data storage system 110. As a result, system 100 may not need to store those chunks. In some cases, the data storage system may exist outside of data chunking module 100.
[0146] As described with respect to FIG. 1, data chunking module 100 may divide each data block into smaller chunks and perform deduplication analysis at the chunk level. Data chunking module 100 may divide data blocks into chunks of equal size. However, this may lead to an insufficient determination of duplicate data because variable-size objects within data stream 102 may be randomly split into random chunks. Alternatively, data chunking module 100 may divide data blocks into variable-size chunks at more natural breaks in order to find different objects within the data stream. This can increase the likelihood of finding duplicate chunks within the data stream. In some cases, however, problems may occur because the system may have extra data chunks when determining variable-size chunks from fixed-size data blocks.
[0147] FIG. 2 is a block diagram of pre-chunking a data stream 200 into a set of data blocks and chunks. In FIG. 2, the data stream 200 can be divided into fixed-size data blocks 202A-202C (e.g., 1 MB). As described above, there may be more or fewer than three data blocks 202A-202C for the data stream 200. In the case of data block 202A, the system can divide data block 202A into chunks 204A-204E. Since the chunks can be of variable size (e.g., 4 kB to 16 kB), there may be a partition of extra data that does not conform to the chunk definition. The system can chunk data block 2020A into chunks 204A-204E, and there may be a chunk 206 of extra data that does not conform to the partitioning algorithm used by the system. For example, the system can use a sliding window to find chunks 204A-204E of 4 to 16 kB size by examining a 16-byte sliding window within data blocks 202A-202C and can search for a pattern having the last 13 bits in the pattern as zero. However, this may leave an extra chunk 206 that does not conform to the above pattern. As a result, since it is unlikely that another chunk will have the same fingerprint as the fingerprint of the extra chunk 206, this data may be wasted. The extra chunk 206 is stored as a separate chunk with a low likelihood of duplication. This may not be much of a problem when the chunk size is small, but as the chunk size (and possibly the data block size) increases, the potential for data waste may increase.
[0148] In some embodiments, a mitigation for this can be to examine the start of the next data block for chunks that include the extra chunk 206. For example, the extra chunk 206 may be analyzed in conjunction with the start of data block 202B, which is the next data block. By examining the next data block, the pre-chunking process can be serialized, thereby suppressing the parallelization of the entire deduplication process.
[0149] In some embodiments, instead of having fixed-size data blocks, the system may pre-chunk the data stream into variable-size segments using the same or similar criteria as those used to chunk the data blocks into multiple chunks. The system may analyze the data stream for chunks using the same or similar criteria as those used to chunk the data for the deduplication operation. If the system may have a sufficient number of chunks to include the amount of data within the segment's range (e.g., 1 MB - 8 MB, 2 MB - 16 MB, or some other range), the system may duplicate this segment. By performing this pre-chunking, the system may generate segments that can be chunked without having extra chunks, as described in FIG. 2 above. This can reduce waste and potentially increase parallelization. VI. Data Segmentation
[0150] FIG. 3 is a flowchart of a segmentation module 300 configured to segment a data stream. In FIG. 3, the segmentation module 300 may start by receiving a data stream at block 302. The data stream may be a file or another type of object that can be deduplicated. At block 302, the segmentation module 300 may pre-chunk the data stream to create segments of chunks. A segment may include a plurality of chunks without extra chunks. In some cases, the segmentation module 300 may pre-chunk the data stream using the same or similar criteria as the deduplication process for chunking data blocks. Pre-chunking may be further described in FIGS. 4 and 5 below. The segmentation module 300 may deduplicate the data stream using the segments at block 306. Deduplication may be performed sequentially or in parallel because the segments do not have extra chunks for deduplication. The segmentation module 300 may chunk each of the segments and perform deduplication on these chunks. For example, for each chunk, the segmentation module 300 may calculate a fingerprint for each chunk and use this fingerprint to determine whether this chunk is currently stored in the data storage. The segmentation module 300 may store the chunk fingerprint in a database. For each chunk to be stored, the segmentation module 300 may calculate the fingerprint of the chunk and determine whether its fingerprint exists in the fingerprint database. If the newly calculated fingerprint is not in the database, the segmentation module 300 may store the corresponding chunk. If the chunk fingerprint matches one of the fingerprints in the database, a copy of this chunk may be stored in the data storage. At block 308, this process may store the deduplicated data stream.The deduplicated data stream may include unique chunks that are not currently stored in the data storage. The segmentation module 300 may store the deduplicated data stream when the data stream is being written, or it can be done after initial storage (e.g., deduplication in the background).
[0151] As described above, the segmentation module 300 may pre-chunk the data stream into a set of segments that are ready for the deduplication process. FIG. 4 is a block diagram of pre-chunking the data stream 400 into a set of variable-size data blocks and chunks. In FIG. 4, the data stream 400 may be pre-chunked into variable-size segments. Each segment may be the total number of chunks (e.g., as shown in FIG. 2 above, without extra chunks). For example, segment 402A, which is smaller than segment 402B or 402C, may include chunks 404A - 404E. In some cases, chunks 404A - 404E may be of variable size and there are no extra chunks that are part of segment 402A. Segment 402A may be shown as a segment smaller than segment 402B or 402C, and in some cases, segment 402A does not necessarily have to be smaller than other data segments (e.g., it may be larger than one, some, or all of the segments, or the same size as another segment). VII. Variable Segment Sizing
[0152] FIG. 5 is a flowchart of a variable segment sizing module 500 that can determine variable-sized segments for duplicate elimination. In FIG. 5, the variable segment sizing module 500 may begin by receiving target segment information at block 502. The target segment information may have a range of bytes that can be used to determine a variable-sized segment. For example, the target segment range may be 1MB - 8MB, 2MB - 16MB, or some other range. At block 504, step 500 may receive a data stream. The data stream may be a file or another object that can be stored in data storage.
[0153] The variable segment sizing module 500 can calculate an offset from the start of the data stream at block 506. The variable segment sizing module 500 can calculate an offset that can be within the range of 4 kB to 16 kB that can be used to find a chunk. For example, the variable segment sizing module 500 can randomly calculate an offset that can be within the range of 4 kB to 16 kB from the start of the data stream. At block 508, an area for analysis can be selected. The variable segment sizing module 500 can select a 16-byte area to determine if there is a natural break in the data stream. Step 500 can calculate an area hash at block 510. The variable segment sizing module 500 can calculate the area hash using a rolling hash (e.g., Rabin-Karp, Rabin fingerprint, Cyclic fingerprint, Addler rolling hash, and / or any other type of rolling hash). The variable segment sizing module 500 can use a hash function algorithm as described elsewhere in this specification. The variable segment sizing module 500 can calculate this hash as a way to determine if there is a natural break in the data stream. At block 512, the variable segment sizing module 500 can determine if a chunk has been found. The variable segment sizing module 500 can determine if a chunk exists by determining that the hash calculated for the 16-byte area has at least 13 bits of the last bits of the hash being zero. The variable segment sizing module 500 can use different criteria to determine if a chunk has been found (e.g., a different number of zeros, a different pattern, etc.). If a chunk is found, execution can proceed to block 514. If a chunk is not found, execution can proceed to block 508, where a new area can be selected by advancing the window within the data stream for analysis.
[0154] In block 514, the variable segment sizing module 500 can determine whether a segment has been found. The variable segment sizing module 500 can determine that a segment has been found by summing the lengths of chunks that can be determined by the variable segment sizing module 500 for those that are not in the current portion of the identified segment. If the sum of these lengths is within the target segment size range, the variable segment sizing module 500 can determine that a new segment has been found and the execution can proceed to block 516. If no segment is found, the execution can proceed to block 508 and a new region can be selected by advancing the window in the data stream for the analysis of a new chunk. In block 518, the variable segment sizing module 500 can mark the segment for deduplication. The variable segment sizing module 500 can mark this segment for deduplication and the segment can be deduplicated later. Computer system
[0155] The present disclosure provides a computer system programmed to implement the methods of the present disclosure. FIG. 12 shows a computer system 1201 programmed or otherwise configured to capture an input data stream, generate one or more segments from the input data stream, generate hash values for one or more chunks of the one or more segments, generate features from the one or more hash values, calculate sketches for the one or more segments, compare the one or more sketches of the one or more segments from the one or more input data streams, differentiate the one or more segments, store the one or more chunks in a database, reduce data duplication, and reconstruct data from one or more read requests. The computer system 1201 can be adjusted in various aspects of sketch calculation, sketch comparison, segment differentiation, and data reconstruction of the present disclosure, such as, for example, a hashing algorithm for generating hash values for one or more chunks can be adjusted to obtain different features for sketch calculation and sketch comparison. The computer system 1201 can be a user's electronic device or a computer system located remotely with respect to the electronic device. The electronic device can be a mobile electronic device.
[0156] Computer system 1201 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 1205, which can be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 1201 also includes a memory or memory location 1210 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 1215 (e.g., hard disk), a communication interface 1220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1225 such as a cache, other memory, data storage, and / or an electronic display adapter. Memory 1210, storage unit 1215, interface 1220, and peripheral devices 1225 communicate with CPU 1205 via a communication bus (solid lines), such as a motherboard. Storage unit 1215 can be a data storage unit (or data repository) for storing data. Computer system 1201 can be operably coupled to a computer network ("network") 1230 with the aid of communication interface 1220. Network 1230 can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet communicating with the Internet. Network 1230 can be, in some cases, a telecommunications and / or data network. Network 1230 can include one or more computer servers that can enable distributed computing such as cloud computing. Network 1230 can, in some cases, implement a peer-to-peer network with the aid of computer system 1201, thereby enabling devices coupled to computer system 1201 to operate as clients or servers.
[0157] The CPU 1205 can execute a series of machine-readable instructions that can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 1210. The instructions can be targeted at the CPU 1205, which can then be programmed or otherwise configured in the CPU 1205 to implement the method of the present disclosure. Examples of operations performed by the CPU 1205 can include fetch, decode, execute, and write-back.
[0158] The CPU 1205 can be part of a circuit such as an integrated circuit. One or more other components of the system 1201 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0159] The storage unit 1215 can store files such as drivers, libraries, and stored programs. The storage unit 1215 can store user data, such as user preferences and user programs. The computer system 1201 can include one or more additional data storage units external to the computer system 1201, such as being located on a remote server that communicates with the computer system 1201 via an intranet or the Internet in some cases.
[0160] The computer system 1201 can communicate with one or more remote computer systems via the network 1230. For example, the computer system 1201 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. A user can access the computer system 1201 via the network 1230.
[0161] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored in an electronic memory location of the computer system 1201, such as, for example, the memory 1210 or the electronic storage unit 1215. The machine executable code or machine readable code can be provided in the form of software. In use, the code can be executed by the processor 1205. In some cases, the code can be retrieved from the storage unit 1215 and stored in the memory 1210 for easy access by the processor 1205. In some situations, the electronic storage unit 1215 can be excluded and the machine executable instructions can be stored in the memory 1210.
[0162] The code can be configured to be used in a machine having a processor adapted to execute the code either by pre-compiling the code or by compiling the code at runtime. The code can be provided in a selectable programming language and can execute the code either in a pre-compiled manner or in a runtime-compiled manner.
[0163] Aspects of the systems and methods provided herein, such as computer system 1201, may be embodied in programming. Various aspects of the technology can generally be considered a "product" or "manufacture" in the form of machine (or processor) executable code and / or associated data that is typically executed or embodied in a type of machine-readable medium. The machine executable code can be stored in an electronic memory unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "memory" type of medium can include any or all of the tangible memories of a computer, processor, etc., or their associated modules such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage for software programming at any time. All or part of the software may be communicated over the Internet or other various electrical communication networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from an administrative server or host computer to an application server's computer platform. Thus, another type of medium that can hold software elements includes physical interfaces between local devices, wired and optical landline networks, and light, electrical, and electromagnetic waves such as those used in various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that hold software. As used herein, unless limited to non-transitory and tangible "memory" media, terms such as "readable medium" of a computer or machine refer to any medium involved in providing instructions to a processor for execution.
[0164] Accordingly, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks such as any (one or more) storage device such as any computer that can be used to implement, for example, a database shown in the drawings. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables having wires that form a bus within a computer system; copper wire and optical fiber. Carrier wave transmission media can take the form of electrical signals or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Accordingly, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROM, DVD or DVD-ROM, other optical media, punch card paper tape, other physical storage media having patterns of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, other memory chips or cartridges, carrier waves that carry data or instructions, cables or links that carry such carrier waves, or other media that a computer can read programming code and data from. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0165] Computer system 1201 can include, or be communicable with, an electronic display 1235 that includes a user interface (UI) 1240 for providing, for example, a hashing algorithm for determination of features of a plurality of chunks for sketching. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0166] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented via software when executed by a central processing unit 1205. The algorithms can, for example, generate a minimum hash value from a set of hash values of multiple chunks.
[0167] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided within the specification. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Those skilled in the art may envision numerous variations, changes, and substitutions without departing from the present invention. Further, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative ratios described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in practicing the present invention. Accordingly, the present invention is intended to cover such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within these claims and their equivalents be covered thereby.
Claims
**Claim 1**: A method for data reduction executed by a system, comprising: (a) receiving one or more input data streams from one or more client applications; (b) generating at least a first segment and a second segment from the one or more input data streams, wherein the first segment includes a first plurality of chunks and the second segment includes a second plurality of chunks; (c) (i) calculating a first sketch of the first segment and (ii) calculating a second sketch of the second segment, wherein the first sketch represents the first segment or includes a set of first features unique to the first segment, the set of first features being calculated using a first subset of chunks selected from the first plurality of chunks, the second sketch represents the second segment or includes a set of second features unique to the second segment, the set of second features being calculated using a second subset of chunks selected from the second plurality of chunks, the set of first features corresponds to the first segment, and the set of second features corresponds to the second segment; (d) processing the first sketch and the second sketch to generate a similarity metric indicating whether the second segment is similar to the first segment; (e) subsequent to (d), (1) when the similarity metric is greater than or equal to a similarity threshold, performing a difference operation on the second segment with respect to the first segment, determining a group of chunks having a matching fingerprint between the first segment and the second segment, and generating and storing a single pointer to the group of chunks, (2) when the similarity metric is less than the similarity threshold, storing the first segment and the second segment in a database without performing the difference operation A method comprising the above steps. **Claim 2** The method according to claim 1, wherein the difference operation in step (e) includes (i) generating a reference hash set of the first plurality of chunks of the first segment and (ii) storing the reference hash set in a memory table. **Claim 3** The method according to claim 2, wherein the reference hash set includes weak hashes.
4. The method according to claim 2, wherein the difference operation further includes: (iii) generating a second hash set for the second plurality of chunks of the second segment; and (iv) comparing the second hash set with the reference hash set in lexicographical order to determine whether there is a match.
5. The method according to claim 4, wherein the difference operation further includes generating and storing a single pointer that collectively references the series of sequential chunks from a subset of the second plurality of chunks when it is determined that the series of sequential chunks has a hash that finds a match from the reference hash set and a subsequent chunk following the series of sequential chunks cannot be found to match from the reference hash set.
6. The method according to claim 5, wherein the single pointer is partially used to generate a sparse index including a reduced set of pointers.
7. The method according to claim 1, wherein the first subset of the chunks is less than 10% of the first plurality of chunks.
8. The method according to claim 1, wherein the first subset of the chunks is less than 1% of the first plurality of chunks.
9. The method according to claim 1, wherein the similarity threshold is at least about 50%.
10. The method according to claim 1, wherein the second subset of the chunks is less than 10% of the second plurality of chunks.
11. The method according to claim 1, wherein the second subset of the chunks is less than 1% of the second plurality of chunks.
12. The method according to claim 1, wherein the first plurality of chunks and the second plurality of chunks are of variable length.
13. The method according to claim 1, wherein the similarity metric is a similarity score.
Citation Information
Patent Citations
Storage controller, storage device, data processing method, and program
JP2017142683A
Use of Similarity Hash to Route Data for Improved Deduplication in a Storage Server Cluster
US20110099351A1
Sparse index bidding and auction based storage
US20120143715A1
Adaptive Index for Data Deduplication
US20120166448A1
Sampling based data de-duplication
US20120233135A1