Repeated data deletion method based on fuzzy clustering in distributed storage system

Through the fuzzy clustering and Bloom filter optimization methods, the problem of high fingerprint index expansion and calculation costs in distributed storage systems is solved, efficient deduplication is achieved, system performance and resource utilization are improved, and dynamic changes in data types and distribution are adapted to.

CN120560583AInactive Publication Date: 2025-08-29XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510663726.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In distributed storage systems, the scale of fingerprint indexes has expanded rapidly, and the I/O cost of fingerprint calculation and diversion check has continued to increase, becoming a bottleneck in system performance. Traditional algorithms lack the ability to perceive and adjust the system load, and cannot dynamically adjust the chunking granularity or clustering strategy. The resource utilization rate is low, making it difficult to cope with dynamic changes in data types and distributions, which affects the deduplication effect and system scalability.

Method used

The fuzzy clustering method is adopted to construct file mapping tables through file types and division strategies, use super-blocking ideas for blocking, combine with keams++ algorithm to improve clustering accuracy, and optimize the Bloom filter through Bloom filter, and use dichotomy and machine addressing length to optimize the Bloom filter, dynamically adjust the search range, reduce the calculation amount, improve the search efficiency, and reduce metadata management overhead.

Benefits of technology

Maintain system performance in a high-load environment, improve duplicate data recognition rate, reduce metadata management overhead, ensure data throughput and system fault tolerance, significantly improve system response efficiency and processing throughput, and adapt to multi-task conflicts and frequent fingerprint query scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560583A_ABST
    Figure CN120560583A_ABST
Patent Text Reader

Abstract

The invention discloses a method for deleting repeated data based on fuzzy clustering in a distributed storage system. The method comprises a partitioning module, a superblock module, a clustering center module, a to-be-stored file S, a similarity Ri module, a threshold value delta module, a fingerprint module, an index table module and a fingerprint index updating module. According to the fuzzy clustering-based duplicated data deletion method in the distributed storage system, a file mapping table is constructed by applying file types and a division strategy to similar clustering, files are classified according to duplicated contents, blocking is carried out according to the file types, a similar tree structure is created, and the duplicated data deletion efficiency is improved. According to the method, by layering and gradually reducing the search range, the search efficiency is improved, and the metadata management overhead is reduced while the repeated data recognition rate is improved, so that the system performance is kept, and the data throughput and the system fault-tolerant capability can also be ensured without sacrificing the partial duplicate removal efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer networks, graphic image processing, computer performance improvement, encryption and decryption, Chinese character input, and more specifically to a method for deleting duplicate data in a distributed storage system based on fuzzy clustering. Background Art

[0002] Driven by the diversification of mobile internet application scenarios, cloud computing platforms (including public and private cloud deployment models) have become the mainstream solution for enterprises to meet massive data storage and computing needs. This transformation not only meets business needs but also creates significant economic benefits. It is worth noting that a joint research report by Microsoft and EMC indicates that in actual workloads, primary memory and secondary storage devices have 50% and 85% data redundancy, respectively. This redundant data not only causes a serious waste of storage resources but also significantly increases the complexity of system operations and maintenance. Therefore, how to effectively eliminate data redundancy has become a key technical challenge in improving the performance of distributed storage systems.

[0003] As an efficient data compression technology, deduplication eliminates redundant data copies and retains only unique data instances, demonstrating significant advantages in optimizing storage space and saving transmission bandwidth. However, facing the exponential growth of data in data centers, traditional single-machine storage architectures are no longer able to meet demand, prompting academic research to shift its focus to deduplication system design in distributed environments.

[0004] Distributed deduplication systems have a core contradiction. When implementing data deduplication in a distributed architecture, system design faces a fundamental contradiction: how to maintain an ideal deduplication rate while ensuring processing efficiency (high throughput). This is specifically manifested in the following ways:

[0005] The negative correlation between chunk granularity and deduplication efficiency: Smaller chunk sizes can improve duplicate data identification rates, but they also lead to exponential increases in metadata management overhead, significant increases in I / O operation frequency, and decreased overall system performance.

[0006] Special requirements of distributed systems: Existing deduplication cluster systems (such as Hydrastor) usually use a larger fixed block size (such as 64KB) to ensure throughput and fault tolerance, but this will sacrifice some deduplication efficiency.

[0007] As data volumes continue to expand, the scale of fingerprint indexes rapidly increases, and the I / O costs of fingerprint calculation and duplicate checking continue to rise, becoming a key bottleneck limiting system performance. Traditional algorithms generally lack the ability to perceive and adjust system load, and are unable to dynamically adjust block granularity or clustering strategies, resulting in low resource utilization. Furthermore, existing fingerprint similarity calculation methods are inefficient, and static clustering mechanisms lack adaptability, making it difficult to cope with dynamic changes in data types and distributions, further impacting deduplication effectiveness and system scalability.

[0008] Therefore, there is an urgent need for an efficient and flexible data deduplication mechanism oriented to file type characteristics to improve the deduplication rate, reduce resource overhead, and enhance the system's processing capabilities and adaptability in complex application scenarios. Summary of the Invention

[0009] (1) Technical problems solved

[0010] The purpose of the present invention is to provide a method for deduplicating data based on fuzzy clustering in a distributed storage system, so as to solve the problems proposed in the above background technology, such as the rapid expansion of fingerprint index scale, the continuous increase in I / O costs of fingerprint calculation and duplication checking, which become the key bottleneck limiting system performance, and the general lack of perception and adjustment capabilities of traditional algorithms for system load, and the inability to dynamically adjust the block granularity or clustering strategy, resulting in low resource utilization. At the same time, the existing fingerprint similarity calculation method is inefficient, the static clustering mechanism lacks adaptability, and it is difficult to cope with the dynamic changes in data types and distributions, further affecting the deduplication effect and system scalability.

[0011] (2) Technical solution

[0012] To achieve the above object, the present invention provides the following technical solution: a method for deduplicating data in a distributed storage system based on fuzzy clustering, comprising blocks, super blocks, cluster centers, files to be stored S, similarity Ri, threshold δ, fingerprints, index tables, and a fingerprint index update module; characterized in that the method comprises the following steps:

[0013] Step 1: File segmentation and ultra-fast division;

[0014] Step 2: Initial cluster center construction module;

[0015] Step 3: Similarity calculation and clustering module;

[0016] Step 4: Fingerprint matching and deduplication module;

[0017] Step 5: Fingerprint index update module. This method of deduplication based on fuzzy clustering in distributed storage system builds a file mapping table by applying file type and division strategy to similar clustering, classifies files according to duplicate content, and divides them into blocks according to file type. It adopts the idea of ​​super block and applies it to similar clustering algorithm. It selects multiple consecutive data blocks to form large-grained super block, and calculates fingerprint for each super block. It uses different clustering algorithms for clustering according to file type, and improves clustering accuracy by using keams++ algorithm. It adopts coarse-grained and fine-grained methods in fingerprint comparison and deduplication. In the combined form, Bloom filters are used to improve speed and accuracy, and Bloom filters are optimized and improved. The binary search method and machine addressing length are used to optimize the Bloom filter. By recursively dividing the bit array of the Bloom filter, a tree-like structure is created, and the target data is quickly found by binary search. This method reduces unnecessary calculations and improves search efficiency by layering and gradually narrowing the search range, and improves its false positive problem. While improving the recognition rate of duplicate data, it reduces metadata management overhead, thereby maintaining system performance. It does not sacrifice some deduplication efficiency and can also ensure data throughput and system fault tolerance.

[0018] Preferably, the block segmentation is to process the files in the data stream by block segmentation through the CDC algorithm, to form a large-grained super block with multiple continuous data blocks, and to dynamically adjust the super block size according to the current system load. By setting the CPU usage and available memory thresholds, adaptive judgment of the system load status is achieved. The table is stored in DRAM to improve the efficiency of subsequent clustering and comparison.

[0019] Preferably, the super block is selected to contain fewer data blocks when the load is high, and more data blocks are selected when the load is low. Each block generates a unique fingerprint through a hash algorithm, and the super block calculates the overall fingerprint for subsequent clustering.

[0020] Preferably, the fingerprint is selected from each super block as the initial cluster center during the fingerprint extraction stage. The preferred strategy is to select the fingerprint with the smallest hash value, namely minhash. All cluster centers and their corresponding storage locations constitute a cluster table.

[0021] Preferably, in the initial clustering phase, the cluster centers are temporary cluster centers using representative fingerprints. The similarity Ri between the file to be stored S and each cluster center is calculated. If the similarity Ri is greater than a dynamically adjusted threshold δ, the k-means++ algorithm is used to cluster similar superblocks and update the cluster centers. If the similarity Ri is less than or equal to the threshold δ, the file to be stored is considered a new type, added as a new cluster center, and its address is recorded in the cluster table. Compared to traditional sampling methods, a two-stage optimization is performed: first, l characteristic fingerprints are intelligently selected for each cluster v to construct a cluster center vi; simultaneously, i typical fingerprints are extracted from the set to be stored u to form a representative set ui. Based on these two optimized representative sets, the calculated rw indicator can more accurately reflect the actual similarity between cluster v and set u.

[0022] Preferably, when comparing and deleting duplicate data, the cluster center first reads the cluster center fingerprint set with the highest similarity, preliminarily compares the fingerprint to be stored with its similarity Ri, and if it is lower than the threshold δ, skips the cluster. Different from the traditional Bloom filter that uses a fixed addressing length and a fixed number of hash functions, this method dynamically optimizes the Bloom filter parameters according to the current fingerprint similarity, that is, if the similarity exceeds the threshold, the binary search method is used to dynamically adjust the machine addressing length, and the appropriate Bloom filter bit array length m is quickly searched according to the fingerprint similarity.

[0023] Preferably, the fingerprint index updating module is responsible for maintaining and updating the fingerprint index in the system to ensure the real-time and accuracy of fingerprint comparison and deduplication operations.

[0024] Preferably, the fingerprint index update module will dynamically update the representative fingerprints in each cluster to the index structure after completing the similarity clustering of the super block. When new fingerprint data is generated, the system will insert it into the most matching existing cluster according to its similarity, and update the representative fingerprint of the cluster when necessary, so that the cluster center can always reflect the overall characteristics of the current data. When new fingerprint data is generated, the system will insert it into the most matching existing cluster according to its similarity, and update the representative fingerprint of the cluster when necessary, so that the cluster center can always reflect the overall characteristics of the current data. This process ensures that subsequent query operations are mainly performed on these representative fingerprints, reducing index access overhead and improving matching efficiency.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows: the deduplication method based on fuzzy clustering in the distributed storage system builds a file mapping table by applying the file type and division strategy to similar clustering, classifies files according to duplicate content, and divides them into blocks according to file type. The idea of ​​super block is adopted and applied to the similar clustering algorithm, multiple consecutive data blocks are selected to form a large-grained super block, and a fingerprint is calculated for each super block. Different clustering algorithms are used for clustering according to file type, and the accuracy of clustering is improved by adopting the keams++ algorithm. Coarse-grained and fine-grained methods are used in fingerprint comparison and deduplication. In the form of combining granularity, Bloom filters are used to improve speed and accuracy, and Bloom filters are optimized and improved. The binary search method and machine addressing length are used to optimize the Bloom filter. By recursively dividing the bit array of the Bloom filter, a tree-like structure is created, and the target data is quickly found by binary search. This method reduces unnecessary calculations and improves search efficiency by layering and gradually narrowing the search range, and improves its false positive problem. While improving the recognition rate of duplicate data, it reduces metadata management overhead, thereby maintaining system performance and ensuring data throughput and system fault tolerance without sacrificing some deduplication efficiency.

[0026] The deduplication method based on fuzzy clustering in a distributed storage system proposed in the present invention is particularly suitable for high-load concurrent environments. It can still maintain high system response efficiency and processing throughput in scenarios such as multi-task conflicts, frequent fingerprint queries and duplicate detection, effectively alleviating the problem of performance degradation of traditional methods when resource competition is fierce. An efficient fingerprint indexing method that integrates super-block and similarity clustering mechanisms is proposed, and a binary method is used to dynamically adjust the machine addressing length. The appropriate Bloom filter bit array length is quickly searched according to fingerprint similarity. Comparative experiments with existing approximate deduplication algorithms show that under the same resource configuration conditions, the method of the present invention can achieve better duplicate data identification effect with smaller memory overhead, especially in large-scale, multi-type file processing environments, showing stronger stability and scalability. Compared with traditional precise duplicate detection algorithms, the present invention has the advantages of high performance, high deduplication rate and low resource consumption, and has significant engineering application value and promotion potential.

[0027] This fuzzy clustering-based deduplication method for distributed storage systems integrates multiple core technologies, including a dynamic load regulation mechanism based on access popularity and change frequency to balance system load in high-concurrency environments; a super-block partitioning strategy to effectively reduce the number of fingerprints and improve the granularity of redundant identification; and an adaptive cluster center update strategy to maintain the accuracy and timeliness of fingerprint clustering. Furthermore, combined with an optimized fingerprint deduplication mechanism, it significantly improves the performance and resource utilization of the deduplication system. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a schematic diagram of the overall process structure of the present invention;

[0029] Figure 2 This is a schematic diagram of the cluster table structure of the present invention;

[0030] Figure 3 This is a schematic diagram of the structure of the distributed data deduplication system of the present invention;

[0031] Figure 4 This is a schematic diagram of the structure of the Bloom filter query operation process of the present invention;

[0032] Figure 5 This is the rw index calculation formula diagram of the present invention;

[0033] Figure 6 Graph showing the calculation formula for the Bloom filter bit array length m of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] See also Figure 1-6 The present invention provides a technical solution: a method for deduplicating data based on fuzzy clustering in a distributed storage system, comprising blocks, super blocks, cluster centers, files to be stored S, similarity Ri, threshold δ, fingerprints, index tables, and a fingerprint index update module. The method is characterized in that the method comprises the following steps:

[0036] Step 1: File segmentation and ultra-fast division;

[0037] Step 2: Initial cluster center construction module;

[0038] Step 3: Similarity calculation and clustering module;

[0039] Step 4: Fingerprint matching and deduplication module;

[0040] Step 5: Fingerprint index update module. This method of deduplication based on fuzzy clustering in distributed storage system builds a file mapping table by applying file type and division strategy to similar clustering, classifies files according to duplicate content, and divides them into blocks according to file type. It adopts the idea of ​​super block and applies it to similar clustering algorithm. It selects multiple consecutive data blocks to form large-grained super block, and calculates fingerprint for each super block. It uses different clustering algorithms for clustering according to file type, and improves clustering accuracy by using keams++ algorithm. It adopts coarse-grained and fine-grained methods in fingerprint comparison and deduplication. In the combined form, Bloom filters are used to improve speed and accuracy, and Bloom filters are optimized and improved. The binary search method and machine addressing length are used to optimize the Bloom filter. By recursively dividing the bit array of the Bloom filter, a tree-like structure is created, and the target data is quickly found by binary search. This method reduces unnecessary calculations and improves search efficiency by layering and gradually narrowing the search range, and improves its false positive problem. While improving the recognition rate of duplicate data, it reduces metadata management overhead, thereby maintaining system performance. It does not sacrifice some deduplication efficiency and can also ensure data throughput and system fault tolerance.

[0041] In the present invention, Figure 1 、 Figure 2 、 Figure 3 and Figure 4 As shown, the block segmentation is to block the files in the data stream through the CDC algorithm, form multiple continuous data blocks into large-grained super blocks, and dynamically adjust the super block size according to the current system load. By setting the CPU usage and available memory thresholds, adaptive judgment of the system load status is achieved. The table is stored in DRAM to improve the efficiency of subsequent clustering and comparison.

[0042] In the present invention, Figure 1 、 Figure 2 、 Figure 3 and Figure 4 As shown, the super block is selected to contain fewer data blocks when the load is high, and more data blocks are selected when the load is low. Each block generates a unique fingerprint through a hash algorithm, and the super block calculates the overall fingerprint for subsequent clustering.

[0043] In the present invention, Figure 1 、 Figure 2 、 Figure 3 and Figure 4 As shown, the fingerprint is selected from each super block in the fingerprint extraction stage. The representative fingerprint is used as the initial cluster center. The optimal strategy is to select the fingerprint with the smallest hash value, namely minhash. All cluster centers and their corresponding storage locations constitute a cluster table.

[0044] In the present invention, Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 and Figure 5 As shown, in the initial clustering phase, the cluster centers are temporarily selected using representative fingerprints. The similarity Ri between the file to be stored S and each cluster center is calculated. If the similarity Ri is greater than a dynamically adjusted threshold δ, the k-means++ algorithm is used to cluster similar superblocks and update the cluster centers. If the similarity Ri is less than or equal to the threshold δ, the file to be stored is considered a new type, added as a new cluster center, and its address is recorded in the cluster table. Compared to traditional sampling methods, this method performs a two-stage optimization. First, for each cluster v, l characteristic fingerprints are intelligently selected to construct a cluster center vi. Simultaneously, i representative fingerprints are extracted from the set to be stored u to form a representative set ui. Based on these two optimized representative sets, the calculated rw metric more accurately reflects the actual similarity between cluster v and set u.

[0045] In the present invention, Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 and Figure 6 As shown, the cluster center first reads the cluster center fingerprint set with the highest similarity during comparison and deduplication, and preliminarily compares the fingerprint to be stored with its similarity Ri. If it is lower than the threshold δ, the cluster is skipped. Unlike the traditional Bloom filter that uses a fixed addressing length and a fixed number of hash functions, this method dynamically optimizes the Bloom filter parameters according to the current fingerprint similarity. That is, if the similarity exceeds the threshold, the binary search method is used to dynamically adjust the machine addressing length, and the appropriate Bloom filter bit array length m is quickly searched according to the fingerprint similarity.

[0046] In the present invention, Figure 1 、 Figure 2 、 Figure 3 and Figure 4 As shown, the fingerprint index update module is responsible for maintaining and updating the fingerprint index in the system to ensure the real-time and accuracy of fingerprint comparison and duplicate data deduplication operations.

[0047] In the present invention, Figure 1 、 Figure 2 、 Figure 3 and Figure 4As shown, the fingerprint index update module will dynamically update the representative fingerprints in each cluster to the index structure after completing the similarity clustering of the super block. When new fingerprint data is generated, the system will insert it into the most matching existing cluster according to its similarity, and update the representative fingerprint of the cluster when necessary, so that the cluster center can always reflect the overall characteristics of the current data. When new fingerprint data is generated, the system will insert it into the most matching existing cluster according to its similarity, and update the representative fingerprint of the cluster when necessary, so that the cluster center can always reflect the overall characteristics of the current data. This process ensures that subsequent query operations are mainly performed on these representative fingerprints, reducing index access overhead and improving matching efficiency.

[0048] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions of the technical solution of the present invention by ordinary technicians in this field do not deviate from the essence and scope of the technical solution of the present invention.

Claims

1. A method for deduplicating data in a distributed storage system based on fuzzy clustering, comprising blocks, super blocks, cluster centers, files to be stored S, similarity Ri, threshold δ, fingerprints, index tables, and a fingerprint index update module. The method is characterized in that: The steps include: Step 1: File segmentation and ultra-fast division; Step 2: Initial cluster center construction module; Step 3: Similarity calculation and clustering module; Step 4: Fingerprint matching and deduplication module; Step 5: Fingerprint index update module.

2. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 1, characterized in that: The block segmentation is to block the files in the data stream through the CDC algorithm, form multiple continuous data blocks into large-grained super blocks, and dynamically adjust the super block size according to the current system load. By setting the CPU usage and available memory thresholds, adaptive judgment of the system load status is achieved.

3. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 2, characterized in that: The super block is selected to contain fewer data blocks when the load is high, and more data blocks are selected when the load is low. Each block generates a unique fingerprint through a hash algorithm, and the super block calculates the overall fingerprint for subsequent clustering.

4. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 3, characterized in that: The fingerprint is selected from each super block in the fingerprint extraction stage. The representative fingerprint is used as the initial cluster center. The optimal strategy is to select the fingerprint with the smallest hash value, namely minhash. All cluster centers and their corresponding storage locations constitute a cluster table.

5. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 4, characterized in that: The cluster center is a temporary cluster center using a representative fingerprint in the initial stage of clustering. The similarity Ri between the file to be stored S and each cluster center is calculated. If the similarity Ri is greater than the dynamically adjusted threshold δ, the k-means++ algorithm is used to cluster similar super blocks and update the cluster center. If the similarity Ri is less than or equal to the threshold δ, the file to be stored is regarded as a new type, added as a new cluster center and its address is recorded in the cluster table.

6. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 5, characterized in that: When comparing and deleting duplicate data, the cluster center fingerprint set with the highest similarity is first read, and the fingerprint to be stored is preliminarily compared with its similarity Ri. If it is lower than the threshold δ, the cluster is skipped.

7. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 6, characterized in that: The fingerprint index update module is responsible for maintaining and updating the fingerprint index in the system to ensure the real-time and accuracy of fingerprint comparison and duplicate data deduplication operations.

8. A method for deduplicating data in a distributed storage system based on fuzzy clustering according to claim 7, characterized in that: After completing the similarity clustering of the super block, the fingerprint index update module will dynamically update the representative fingerprint in each cluster to the index structure. When new fingerprint data is generated, the system will insert it into the most matching existing cluster based on its similarity, and update the representative fingerprint of the cluster when necessary, so that the cluster center can always reflect the overall characteristics of the current data.