Cloud backup platform data deduplication method based on distributed storage
By using technical means such as dynamic chunking algorithms, deep learning algorithms and reinforcement learning algorithms on the cloud backup platform, data redundancy problems in the cloud backup platform are solved, efficient and accurate data deduplication and storage optimization are achieved, and the overall performance and user experience of the cloud backup platform are improved.
Patent Information
- Application Number
- CN202510020907.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
There are data redundancy problems in cloud backup platforms, resulting in excessive storage space usage, increased storage costs, excessive network bandwidth and computing resources consumption, and the existing technology is difficult to effectively ensure data consistency and deduplication performance.
The data deduplication method based on distributed storage is adopted. The data is adaptively divided and multi-dimensional feature extraction through dynamic chunking algorithms and intelligent feature extraction algorithms, distributed hash tables and global indexes are established, and the deep learning algorithms are used for redundancy analysis and similarity matching, and dynamic optimization is combined with the hierarchical graph division algorithm and reinforcement learning algorithm.
Effectively identify and remove duplicate data blocks, improve the accuracy and efficiency of data deduplication, optimize data storage distribution, reduce uneven load of storage nodes, and improve the data storage and management efficiency of cloud backup platforms.
Smart Images

Figure CN119938406A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a data deduplication method for a cloud backup platform based on distributed storage. Background Art
[0002] With the rapid development of cloud computing and big data technology, more and more enterprises and individuals have begun to migrate data storage and backup to cloud platforms to ensure data security and improve data management efficiency. As an important cloud storage application, cloud backup platforms can provide users with remote data backup and recovery services. However, with the rapid growth of data volume, cloud backup platforms face many challenges, among which data redundancy is particularly prominent.
[0003] In cloud backup scenarios, user data backups often contain a large amount of duplicate data. For example, for different versions of the same file, only part of the content changes, while a large amount of content remains the same; or for files uploaded by multiple users, duplicate data may also exist due to the same data source. The existence of these duplicate data will take up a lot of storage space, increase storage costs, and also consume more network bandwidth and computing resources, affecting the performance and efficiency of the cloud backup platform. At present, there are still the following problems: when the existing technology establishes and maintains the global index, it is often difficult to effectively ensure data consistency, which easily leads to index failure or storage location mismatch, affecting data positioning and deduplication performance; traditional data deduplication methods mainly rely on simple hash matching, which cannot perform in-depth analysis of the characteristics of data blocks, and it is difficult to identify similar data blocks. It fails to fully tap the potential of data redundancy, resulting in low storage space utilization; when the existing technology stores data after deduplication, it lacks an effective storage optimization strategy, resulting in uneven data distribution and excessive load on some storage nodes, affecting system performance and stability. Summary of the invention
[0004] To solve the above problems, the present invention provides a data deduplication method for a cloud backup platform based on distributed storage, which solves the problem of how to effectively solve the data redundancy problem in the cloud backup platform, ensures data consistency and deduplication performance, and improves the accuracy of data deduplication by identifying similar data blocks through in-depth analysis. At the same time, it optimizes data storage distribution and reduces the problem of uneven load on storage nodes, thereby achieving high efficiency, accuracy and dynamic optimization of data deduplication, and improving the data storage and management efficiency of the cloud backup platform.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] The method for deduplication of data on a cloud backup platform based on distributed storage includes the following steps:
[0007] S1: Adaptively partition the backup data using a dynamic partitioning algorithm, generate a unique identification code for each partition, and extract multi-dimensional feature information based on an intelligent feature extraction algorithm, including content distribution features, redundancy pattern features, and compression potential assessment features;
[0008] S2: Establish a global index through a distributed hash table, map the unique identifier of each block to its storage location, and use a distributed consistency protocol to dynamically maintain the consistency of the index;
[0009] S3: Based on the global index, a deep learning algorithm is used to perform redundancy analysis and similarity matching on the block features, to identify duplicate or similar data blocks in the global scope, and to execute a deduplication strategy based on the detection results;
[0010] S4: Based on the access frequency and load conditions, a layered graph partitioning algorithm is used to map the deduplicated data blocks into a cross-node storage graph model, and the deduplicated data blocks are distributed and stored in multiple nodes;
[0011] S5: Based on the reinforcement learning algorithm, it monitors the performance of the deduplication process in real time, dynamically optimizes the block strategy, index frequency and storage distribution, and achieves the global optimization of performance and resource utilization.
[0012] Furthermore, the construction process of the dynamic block algorithm includes the following steps:
[0013] Based on the content characteristics of the data to be backed up, the sliding window technology is combined with the sensitivity adjustment mechanism to dynamically cut the data and generate an initial block set with boundary adaptability;
[0014] The initial block set is divided through a multi-level block optimization strategy, combined with the data content change rate and block granularity adjustment factor;
[0015] Generate a unique identification code for each block, wherein the identification code is generated based on the joint calculation of the encrypted hash algorithm and the block characteristics;
[0016] After the chunking is completed, the characteristics of each chunk are evaluated, and the chunking strategy is dynamically adjusted based on the chunk's stability index and specific content distribution rules.
[0017] Furthermore, the step S3 includes the following steps:
[0018] Obtain multi-dimensional feature information of blocks based on global index, and generate high-dimensional feature vectors using feature embedding algorithm;
[0019] Adaptive variational autoencoder is used to reduce the dimensionality of high-dimensional feature vectors to generate latent feature space representation, and density clustering analysis is used to evaluate the similarity and repeatability between blocks to construct a similarity matrix of block features.
[0020] Based on the similarity matrix, a distributed graph neural network is used to dynamically construct a block feature map, perform redundancy detection and repeated clustering analysis on the blocks, and improve the efficiency and accuracy of repeated detection through a cross-node distributed collaboration mechanism;
[0021] According to the duplicate detection results, the edge entropy optimal deduplication algorithm is used to perform deduplication operations on duplicate or similar data blocks, delete completely duplicate blocks, and implement a merge storage strategy for similar blocks.
[0022] Furthermore, the formula of the distributed graph neural network is as follows:
[0023]
[0024] in, Represents the feature vector of block v after being updated by the graph neural network; and Represents the multi-dimensional feature data of data blocks; F u and F v represents the additional feature vectors of nodes u and v, including relevant metadata obtained from the global index; N(v) represents the set of neighbor nodes of node v; d u and d v Indicates the number of other similar blocks it is connected to; represents the attention weight, that is, the similarity between node u and node v; W (l) Represents the trainable weight matrix of the graph neural network at layer l; represents the adaptive weight coefficient; σ represents the activation function.
[0025] Further, the step S4 includes the following steps:
[0026] Based on the deduplicated data blocks, the importance of the data blocks is dynamically evaluated using access frequency and load balancing strategies to calculate the storage priority of each data block.
[0027] According to the characteristics of the deduplicated data blocks and the node load status, a cross-node storage graph model is constructed using a layered graph partitioning algorithm. The nodes in the graph represent storage nodes, and the edges represent the transmission path weights between nodes.
[0028] Based on the constructed cross-node storage graph model, the graph optimization algorithm is used to generate the optimal block storage distribution solution;
[0029] The node storage status and the distribution status of access blocks are monitored in real time through the reinforcement learning algorithm, and the distribution storage strategy of deduplicated data blocks is dynamically adjusted.
[0030] Furthermore, the formula of the layered graph partitioning algorithm is as follows:
[0031]
[0032] Wherein, min F represents the objective function of hierarchical graph partitioning, which is used to describe the optimization goal of data block storage distribution; N represents the total number of data blocks; M represents the total number of storage nodes; represents the set of neighboring data blocks of data block i; w ij represents the transmission weight between data block i and data block j; d ij represents the network delay or transmission path weight between different nodes; f ij represents the similarity characteristics of data blocks i and j; c represents the weight coefficient; L k represents the current load of storage node k; C k Represents the capacity of storage node k.
[0033] Furthermore, step S4 also includes storing the data blocks determined to be duplicate only once, and managing other duplicate references through logical pointers; for newly added non-duplicate data blocks, efficient storage is performed by combining compression algorithms and hot and cold data tiering strategies, and an incremental data flow model is introduced to dynamically optimize the storage layout.
[0034] Furthermore, the establishment of the global index supports efficient associative mapping between high-dimensional features and physical storage locations through a multi-dimensional index mapping mechanism combined with a content-aware hash algorithm based on block features.
[0035] Furthermore, the reinforcement learning algorithm is based on a two-layer hierarchical optimization framework. The top layer uses a policy gradient method to optimize the global performance objectives of the deduplication process, and the bottom layer combines a multi-agent collaboration mechanism to achieve coordinated dynamic adjustment of the block strategy, index frequency, and storage mapping.
[0036] The beneficial effects of the present invention are:
[0037] The present invention uses a dynamic block algorithm to adaptively block the backup data, and extracts multidimensional feature information (such as content distribution features, redundant pattern features and compression potential assessment features) in combination with an intelligent feature extraction algorithm, which can accurately locate the redundancy and similarity of data, and effectively improve the accuracy and efficiency of data deduplication. A global index is established through a distributed hash table, and a distributed consistency protocol is used to dynamically maintain the consistency of the index, ensuring that the unique identification of the data block and the storage location mapping are accurate, while ensuring the high availability and reliability of the index in the distributed system. The deep learning algorithm is used to perform redundancy analysis and similarity matching on the characteristics of the data block, which can efficiently identify duplicate or similar data blocks in a global range, realize intelligent data deduplication, and effectively reduce storage space occupancy. The deduplicated data blocks are mapped to a cross-node storage graph model through a hierarchical graph partitioning algorithm, and distributed storage is performed in combination with data access frequency and node load conditions, further realizing data load balancing and improving the utilization of storage resources. Based on the reinforcement learning algorithm, the performance of the deduplication process is monitored in real time, and the block strategy, index update frequency and storage distribution are dynamically optimized to ensure that the data deduplication process reaches the global optimum in terms of performance and resource utilization, effectively improving the response speed and operation efficiency of the system. Through precise data deduplication and distributed storage mechanisms, the amount of duplicate data storage is significantly reduced, saving storage costs, while increasing the speed of data backup and recovery, and improving the overall performance and user experience of the cloud backup platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flow chart of a data deduplication method for a cloud backup platform based on distributed storage according to the present invention.
[0039] Figure 2 It is a flowchart of step S3 provided in one embodiment of the present invention.
[0040] Figure 3 It is a flowchart of step S4 provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0041] See also Figure 1-3 As shown, the present invention relates to a data deduplication method for a cloud backup platform based on distributed storage.
[0042] Example
[0043] The method for deduplication of data on a cloud backup platform based on distributed storage includes the following steps:
[0044] S1: Adaptively partition the backup data using a dynamic partitioning algorithm, generate a unique identification code for each partition, and extract multi-dimensional feature information based on an intelligent feature extraction algorithm, including content distribution features, redundancy pattern features, and compression potential assessment features;
[0045] The construction process of the dynamic block algorithm includes the following steps:
[0046] Based on the content characteristics of the data to be backed up, the sliding window technology is combined with the sensitivity adjustment mechanism to dynamically cut the data and generate an initial block set with boundary adaptability;
[0047] The initial block set is divided through a multi-level block optimization strategy, combined with the data content change rate and block granularity adjustment factor;
[0048] Specifically, the data content change rate of the initial block set is evaluated, and the degree of content change within the block is calculated (for example, through indicators such as data entropy value and repetition rate).
[0049] Dynamically adjust the block size based on the change rate results:
[0050] High rate of change: refine the block granularity and divide the data into smaller blocks;
[0051] Low change rate: Merge adjacent blocks to generate larger blocks, reducing the number of blocks and index burden.
[0052] Set the block size range (e.g. minimum 4KB, maximum 64KB), introduce a block size adjustment factor, and dynamically optimize blocks based on data characteristics. For example: in a network transmission environment, set a smaller block size for bandwidth-sensitive data streams to facilitate subsequent deduplication processing; in large file storage scenarios, increase the block size appropriately to improve storage efficiency. Through a multi-level block optimization strategy, the initial block set is re-divided into a block set with stronger content adaptability and better granularity.
[0053] Generate a unique identification code for each block, wherein the identification code is generated based on the joint calculation of the encrypted hash algorithm and the block characteristics;
[0054] Specifically, for each finally generated block, a cryptographic hash algorithm (such as SHA-256 or MD5) is used to calculate the hash value of the block. In order to improve the accuracy of the unique identification code and the distinguishability of the blocks, the hash result is jointly calculated in combination with the block characteristics (such as content distribution and change rate) to ensure uniqueness. The hash value of the block is combined with the content feature vector to generate a composite identification code. The generated unique identification code and the metadata corresponding to the block (such as block size, starting position, etc.) are recorded in the index table to provide a basis for subsequent data deduplication and storage mapping.
[0055] After the chunking is completed, the characteristics of each chunk are evaluated, and the chunking strategy is dynamically adjusted based on the chunk's stability index and specific content distribution rules.
[0056] It should be noted that the stability of each block is evaluated mainly based on the following indicators:
[0057] If a block remains stable in different versions, it is considered a high-stability block;
[0058] The degree of repetitiveness of chunk boundaries: Frequent changes in boundaries may indicate that the chunking strategy needs to be optimized.
[0059] Analyze the data content distribution of the blocks and identify specific types of data (such as duplicate data blocks and data blocks with high compression potential).
[0060] Dynamically adjust the block strategy based on content distribution rules: further refine the redundant data blocks; mark the compressible data blocks in advance to facilitate subsequent compression processing. Combining stability indicators and content rule analysis, the system will dynamically optimize block sensitivity, granularity adjustment factors and other parameters to gradually improve block quality and efficiency.
[0061] S2: A global index is established through a distributed hash table, the unique identifier of each block is mapped to its storage location, and a distributed consistency protocol is used to dynamically maintain the consistency of the index; the establishment of the global index supports efficient association mapping between high-dimensional features and physical storage locations through a multi-dimensional index mapping mechanism combined with a content-aware hash algorithm based on block features.
[0062] Specifically, a DHT ring structure is established based on a consistent hashing algorithm (such as Consistent Hashing), and the storage nodes are arranged according to the hash value space. For example, a unique node ID is assigned to each storage node, and mapped to a certain position in the logical ring through a hash function (such as SHA-256).
[0063] The unique identification code of the data block (such as SHA-256) is also calculated by the same hash function and stored on the corresponding node according to its hash value. The unique identification code of each data block is used as the key, and the storage location (node address, offset, etc.) is used as the value to form a key-value pair of <unique identification code, storage location>. If the hash value of a data block conflicts with multiple nodes, the virtual node mechanism is used to evenly distribute the data block to different nodes. In order to prevent uneven load on storage nodes, virtual node technology is used to divide a physical node into multiple logical virtual nodes to improve the balance of data block indexes.
[0064] The extracted data block features (such as content distribution features, redundant pattern features, and compression potential features) are compressed into low-dimensional space through dimensionality reduction algorithms (such as principal component analysis PCA and t-SNE). Content-aware hashing algorithms (such as SimHash and MinHash) are used to generate feature-related hash codes to ensure that similar data blocks are mapped to similar locations, accelerating similarity retrieval.
[0065] Design a two-layer index structure:
[0066] Primary index: The unique identifier and storage node address of the data block stored through DHT.
[0067] Auxiliary index: stores the feature hash code of the data block to support feature retrieval and similarity matching.
[0068] The distributed hash search algorithm is used for the primary index, and the query complexity is O(log N). For the auxiliary index, locality sensitive hashing (LSH) is used to achieve fast approximate search for similar features and reduce search overhead.
[0069] Use the Raft protocol or Paxos protocol to maintain the consistency of index data to ensure that the index can be correctly synchronized when a node fails or a new node is added.
[0070] Working principle of Raft protocol: All nodes elect a master node (Leader); the master node is responsible for receiving index update requests and synchronizing the updates to other nodes (Followers); if the master node fails, the system automatically re-elects the master node to ensure uninterrupted consistency maintenance.
[0071] When a data block is added, deleted, or modified, the node will send an index update request to the master node, which will be responsible for executing the update and synchronizing the update to all replica nodes through the consistency protocol. An index version control mechanism is implemented to avoid index conflicts between different nodes.
[0072] In order to improve system reliability, the index data is copied to multiple nodes to form a replication mechanism (such as three-copy storage). If a node fails, the index data can be restored from the replica node to ensure that the index service is not interrupted.
[0073] When a new node joins the system, the system automatically migrates the index of some data blocks from the existing node to the new node, redistributes the index data, and ensures the balanced index load. When a node leaves the system, the index of the data block is automatically migrated to other available nodes to maintain the integrity of the index data. According to the storage node load status and data access frequency, the system periodically executes the data redistribution algorithm to dynamically rebalance and optimize the global index.
[0074] S3: Based on the global index, a deep learning algorithm is used to perform redundancy analysis and similarity matching on the block features, to identify duplicate or similar data blocks in the global scope, and to execute a deduplication strategy based on the detection results;
[0075] Wherein, the step S3 comprises the following steps:
[0076] Obtain multi-dimensional feature information of blocks based on global index, and generate high-dimensional feature vectors using feature embedding algorithm;
[0077] Adaptive variational autoencoder is used to reduce the dimensionality of high-dimensional feature vectors to generate latent feature space representation, and density clustering analysis is used to evaluate the similarity and repeatability between blocks to construct a similarity matrix of block features.
[0078] Specifically, the high-dimensional feature vector is input into an adaptive variational autoencoder (VAE) to perform dimensionality reduction on the feature space.
[0079] The working process of the VAE model includes: mapping the high-dimensional input feature vector to the low-dimensional latent feature space to extract the core features of the data block; dynamically adjusting the dimension of the encoder according to the complexity of the feature so that the dimensionality reduction result can accurately retain the representativeness of the feature. By reconstructing the low-dimensional feature vector, the information retention in the dimensionality reduction process is ensured. The output result is the representation vector of each block in the low-dimensional latent space, which can more efficiently represent the redundancy and similarity of the block.
[0080] Use density clustering algorithms (such as DBSCAN or HDBSCAN) to cluster potential feature vectors and evaluate the similarity and repetition between blocks: high-density areas represent data blocks with high similarity or repeated content; outliers represent non-repetitive blocks with strong uniqueness. The output result is a similarity matrix between blocks, which is used for subsequent redundancy detection and cluster analysis.
[0081] Based on the similarity matrix, a distributed graph neural network is used to dynamically construct a block feature map, perform redundancy detection and repeated clustering analysis on the blocks, and improve the efficiency and accuracy of repeated detection through a cross-node distributed collaboration mechanism;
[0082] Specifically, based on the similarity matrix, a distributed graph neural network (such as GraphSAGE, GAT) is used to dynamically construct a block feature graph:
[0083] Node: feature vector of each data block;
[0084] Edge: Similarity between nodes (edge weight).
[0085] Graph neural networks learn through node representation and further extract the relational features between nodes.
[0086] Complete duplicate block: nodes with higher edge weights in the graph are marked as complete duplicates through clustering;
[0087] Similar blocks: Nodes with high edge weights but not completely identical are identified as similar blocks through similarity matching.
[0088] In a distributed storage environment, the computing tasks of the graph neural network are distributed and collaboratively executed through a cross-node communication mechanism to improve the efficiency and accuracy of duplicate detection: each node is responsible for calculating the feature map of the local block; the intermediate results are shared with other nodes through a decentralized algorithm (such as the Gossip protocol) to build a global block feature map.
[0089] According to the duplicate detection results, the edge entropy optimal deduplication algorithm is used to perform deduplication operations on duplicate or similar data blocks, delete completely duplicate blocks, and implement a merge storage strategy for similar blocks.
[0090] Specifically, the edge entropy optimal deduplication algorithm is used to perform deduplication operations on the detected duplicate or similar blocks:
[0091] Completely duplicate chunks: directly delete redundant chunks and keep only one physical copy; mark the storage locations of other duplicate chunks through logical references (such as pointers or index mappings).
[0092] Similar blocks: Use content difference merging technology (such as incremental storage method) to store only the differences between similar blocks, minimizing storage space usage.
[0093] Based on the deduplication results, update the storage information after deleting redundant data to the global index to ensure that the index is consistent with the storage layout. Update the cross-node storage distribution to ensure load balancing and efficient access to data storage.
[0094] Furthermore, the formula of the distributed graph neural network is as follows:
[0095]
[0096] in, Represents the feature vector of block v after being updated by the graph neural network; and Represents the multi-dimensional feature data of data blocks; F u and F v represents the additional feature vectors of nodes u and v, including relevant metadata obtained from the global index; N(v) represents the set of neighbor nodes of node v; d u and d v Indicates the number of other similar blocks it is connected to; represents the attention weight, that is, the similarity between node u and node v; W (l) Represents the trainable weight matrix of the graph neural network at layer l; represents the adaptive weight coefficient; σ represents the activation function.
[0097] S4: According to the access frequency and load conditions, a layered graph partitioning algorithm is used to map the deduplicated data blocks into a cross-node storage graph model, and the deduplicated data blocks are distributed and stored in multiple nodes; the step S4 also includes storing the data blocks determined to be duplicated only once, and managing other duplicate references through logical pointers; for newly added non-duplicate data blocks, the compression algorithm and the hot and cold data tiering strategy are combined for efficient storage, and the incremental data flow model is introduced to dynamically optimize the storage layout.
[0098] It should be noted that a compression algorithm (such as LZ4, ZSTD) is executed on the newly added non-duplicate data blocks to reduce the storage occupancy of the data blocks.
[0099] Hot and cold tiered storage strategy:
[0100] Hot data: Highly accessed data blocks are stored in high-performance tiers (such as SSDs or NVMe storage nodes). This ensures the lowest access latency and improves data access efficiency.
[0101] Cold data: Data blocks that are accessed infrequently are stored in large capacity layers (such as HDD storage nodes). Data blocks are compressed during storage to save storage space to the greatest extent possible.
[0102] During the data backup and storage process, new or changed data blocks are continuously monitored to build incremental data streams.
[0103] The incremental data is compared with the stored data in real time to identify the newly added data blocks and dynamically adjust the storage layout.
[0104] When the storage node load is unbalanced or the storage space is insufficient, the system dynamically migrates data blocks:
[0105] Migrate low-priority data blocks to nodes with lower load or more suitable for storage; keep high-priority data blocks on high-performance nodes. Keep data accessible during the migration process to avoid affecting system performance.
[0106] Wherein, the step S4 comprises the following steps:
[0107] Based on the deduplicated data blocks, the importance of the data blocks is dynamically evaluated using access frequency and load balancing strategies to calculate the storage priority of each data block.
[0108] Specifically, the access frequency of the deduplicated data blocks is monitored in real time, and the data block access logs within a period of time are collected. The number of accesses to each data block is counted using the sliding time window method:
[0109] High-frequency data blocks: These blocks are accessed frequently recently and are considered hot data.
[0110] Low-frequency data blocks: These blocks are accessed less frequently recently and are considered cold data.
[0111] The system monitors the storage occupancy, bandwidth usage and load status of each storage node in real time. It conducts a comprehensive evaluation of the node load and generates a node load score table: nodes with low load have higher storage availability; nodes with high load are avoided first when allocating storage.
[0112] Taking into account the access frequency of data blocks and node load conditions, the storage priority is dynamically calculated for each data block: high-frequency data blocks are preferentially stored in low-latency, high-bandwidth storage nodes; low-frequency data blocks are compressed and stored in cold data storage nodes with larger capacity.
[0113] According to the characteristics of the deduplicated data blocks and the node load status, a cross-node storage graph model is constructed using a layered graph partitioning algorithm. The nodes in the graph represent storage nodes, and the edges represent the transmission path weights between nodes.
[0114] Specifically, in the definition of the graph model: the nodes in the graph represent storage nodes; the edges represent the transmission paths between nodes, and the weights are set according to indicators such as transmission bandwidth, node load, and access delay.
[0115] Each storage node reports its own status information in real time, including: storage space usage; current network bandwidth load; I / O processing capacity and access latency.
[0116] Through the hierarchical graph partitioning algorithm (such as METIS, Kernighan-Lin, etc.): the storage nodes are divided into different levels: high-performance layer: suitable for storing high-priority, frequently accessed data blocks; large-capacity layer: suitable for storing compressed low-frequency data blocks; spare layer: as a failover and disaster recovery backup node. The principles of hierarchical division include: fully considering the load balancing of nodes; minimizing the transmission path overhead across nodes and optimizing the storage efficiency of data blocks.
[0117] Based on the constructed cross-node storage graph model, the graph optimization algorithm is used to generate the optimal block storage distribution solution;
[0118] Specifically, based on the constructed cross-node storage graph model, a graph optimization algorithm (such as the shortest path algorithm Dijkstra or the minimum spanning tree algorithm) is used to generate an optimal data block storage distribution solution.
[0119] High-priority data blocks: have the lowest path weight, the smallest transmission delay, and are preferentially allocated to high-performance nodes;
[0120] Low-priority data blocks: have high path overhead but low storage cost, and are allocated to large-capacity cold storage nodes.
[0121] For detected duplicate data blocks, the system only stores one physical copy; in other storage nodes, the reference relationship of duplicate data blocks is managed through logical pointers (such as soft links or metadata mapping tables) to reduce storage redundancy.
[0122] The node storage status and the distribution status of access blocks are monitored in real time through the reinforcement learning algorithm, and the distribution storage strategy of deduplicated data blocks is dynamically adjusted.
[0123] Furthermore, the formula of the layered graph partitioning algorithm is as follows:
[0124]
[0125] Wherein, min F represents the objective function of hierarchical graph partitioning, which is used to describe the optimization goal of data block storage distribution; N represents the total number of data blocks; M represents the total number of storage nodes; represents the set of neighboring data blocks of data block i; w ij represents the transmission weight between data block i and data block j; d ij represents the network delay or transmission path weight between different nodes; f ij represents the similarity characteristics of data blocks i and j; c represents the weight coefficient, which is used to balance the optimization priority of transmission cost and node load balancing; L k represents the current load of storage node k; C k Represents the capacity of storage node k.
[0126] S5: Based on the reinforcement learning algorithm, the performance of the deduplication process is monitored in real time, and the block strategy, index frequency and storage distribution are dynamically optimized to achieve the global optimization of performance and resource utilization; the reinforcement learning algorithm is based on a two-layer hierarchical optimization framework. The top layer uses the policy gradient method to optimize the global performance goals of the deduplication process, and the bottom layer combines the multi-agent collaboration mechanism to achieve coordinated dynamic adjustment of the block strategy, index frequency and storage mapping.
[0127] Specifically, the intelligent agent through reinforcement learning learns how to choose the best action to optimize the overall performance under different system states. A two-layer reinforcement learning framework is designed: Top-level optimization: responsible for global strategy decision-making and performance target optimization. Bottom-level optimization: responsible for dynamic adjustment of block, index and storage layout parameters, and realizing multi-agent collaboration.
[0128] The top-level optimization mainly makes decisions based on the global goals. The specific implementation steps include: setting overall performance goals according to the needs of data deduplication, such as maximizing the deduplication rate, improving storage space utilization efficiency, and reducing data processing delays. The top-level reinforcement learning agent regularly analyzes the current system status and dynamically optimizes global parameters, such as the granularity range of the block strategy and the index update cycle. For example: when the system load is high, it is preferred to reduce the block size to speed up data processing; when the system bandwidth pressure is high, the index update frequency is reduced to save resources.
[0129] The underlying layer uses multi-agent reinforcement learning, and each agent is responsible for block strategy, index maintenance and storage layout optimization. The specific implementation is as follows: monitor the characteristic distribution of data blocks and the current system performance, and dynamically adjust the granularity of blocks; dynamically adjust the frequency and strategy of index updates according to the frequency of data changes and system resource status; monitor the load of storage nodes and the frequency of data access, and dynamically adjust the storage layout of data blocks; share system status information among multiple agents, and through collaborative decision-making, ensure that block strategy, index update and data storage layout cooperate with each other to achieve the global optimum.
[0130] In summary, the present invention effectively identifies the redundant patterns and compression potential of data through dynamic block algorithm and multi-dimensional feature extraction, and performs deduplication operation using edge entropy optimal deduplication algorithm, thereby significantly reducing data redundancy and improving storage space utilization. The dynamic block algorithm is adopted, combined with the data content change rate and block granularity adjustment factor, and granularity adaptive optimization is performed for different data types (such as network data streams and large files), balancing storage efficiency and deduplication accuracy.
[0131] The present invention maintains a global index through a distributed hash table (DHT) and a consistency protocol to achieve efficient mapping between unique identifiers of data blocks and storage locations. A two-layer index structure (primary index and auxiliary index) is adopted to support efficient unique data location and similarity retrieval at the same time. Deep learning algorithms (such as variational autoencoders, density clustering and graph neural networks) are used to achieve feature dimensionality reduction, similarity clustering and redundancy analysis of data blocks, effectively improving the recognition accuracy of duplicate data, especially achieving efficient deduplication on a global scale.
[0132] The present invention realizes the distributed collaborative processing of data block deduplication tasks through distributed graph neural network and cross-node communication mechanism, avoids centralized bottlenecks, and improves the processing efficiency and scalability of the system. The layered graph partitioning algorithm is used to distribute the deduplicated data blocks to different nodes, and the hot and cold layered storage strategy and node load monitoring are combined to ensure efficient data storage and access, reduce network transmission overhead, and achieve storage load balancing.
[0133] The present invention uses a reinforcement learning algorithm to monitor the system status in real time and optimize the block strategy, index frequency and storage layout to ensure the global optimization of performance and resource utilization, and adapt to system load changes and data access patterns. Through a consistency protocol (such as the Raft protocol) and an index copy mechanism, it is ensured that when a node fails or a new node is added, the index and data can be correctly synchronized and restored to ensure the high reliability and disaster recovery capability of the system. Using an incremental data flow model, newly added data blocks are identified in real time, and non-duplicate data is efficiently stored in combination with a compression algorithm, further reducing storage space occupancy and improving data backup efficiency.
[0134] The above implementation modes are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering and technical personnel in the field shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A data deduplication method for a cloud backup platform based on distributed storage, characterized in that: The following steps are involved: S1: Adaptively partition the backup data using a dynamic partitioning algorithm, generate a unique identification code for each partition, and extract multi-dimensional feature information based on an intelligent feature extraction algorithm, including content distribution features, redundancy pattern features, and compression potential assessment features; S2: Establish a global index through a distributed hash table, map the unique identifier of each block to its storage location, and use a distributed consistency protocol to dynamically maintain the consistency of the index; S3: Based on the global index, a deep learning algorithm is used to perform redundancy analysis and similarity matching on the block features, to identify duplicate or similar data blocks in the global scope, and to execute a deduplication strategy based on the detection results; S4: Based on the access frequency and load conditions, a layered graph partitioning algorithm is used to map the deduplicated data blocks into a cross-node storage graph model, and the deduplicated data blocks are distributed and stored in multiple nodes; S5: Based on the reinforcement learning algorithm, it monitors the performance of the deduplication process in real time, dynamically optimizes the block strategy, index frequency and storage distribution, and achieves the global optimization of performance and resource utilization.
2. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 1, characterized in that: The construction process of the dynamic block algorithm includes the following steps: Based on the content characteristics of the data to be backed up, the sliding window technology is combined with the sensitivity adjustment mechanism to dynamically cut the data and generate an initial block set with boundary adaptability; The initial block set is divided through a multi-level block optimization strategy, combined with the data content change rate and block granularity adjustment factor; Generate a unique identification code for each block, wherein the identification code is generated based on the joint calculation of the encrypted hash algorithm and the block characteristics; After the chunking is completed, the characteristics of each chunk are evaluated, and the chunking strategy is dynamically adjusted based on the chunk's stability index and specific content distribution rules.
3. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 1, characterized in that: The step S3 comprises the following steps: Obtain multi-dimensional feature information of blocks based on global index, and generate high-dimensional feature vectors using feature embedding algorithm; Adaptive variational autoencoder is used to reduce the dimensionality of high-dimensional feature vectors to generate latent feature space representation, and density clustering analysis is used to evaluate the similarity and repeatability between blocks, and the similarity moment matrix of block features is constructed; Based on the similarity matrix, a distributed graph neural network is used to dynamically construct a block feature map, perform redundancy detection and repeated clustering analysis on the blocks, and improve the efficiency and accuracy of repeated detection through a cross-node distributed collaboration mechanism; According to the duplicate detection results, the edge entropy optimal deduplication algorithm is used to perform deduplication operations on duplicate or similar data blocks, delete completely duplicate blocks, and implement a merge storage strategy for similar blocks.
4. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 3 is characterized in that: The formula of the distributed graph neural network is as follows: in, Represents the feature vector of block v after being updated by the graph neural network; and Represents the multi-dimensional feature data of data blocks; F u and F v represents the additional feature vectors of nodes u and v, including relevant metadata obtained from the global index; N(v) represents the set of neighbor nodes of node v; d u and d v Indicates the number of other similar blocks it is connected to; represents the attention weight, that is, the similarity between node u and node v; W (l) Represents the trainable weight matrix of the graph neural network at layer I; represents the adaptive weight coefficient; σ represents the activation function.
5. The method for deduplication of cloud backup platform data based on distributed storage according to claim 1, characterized in that: The step S4 comprises the following steps: Based on the deduplicated data blocks, the importance of the data blocks is dynamically evaluated using access frequency and load balancing strategies to calculate the storage priority of each data block. According to the characteristics of the deduplicated data blocks and the node load status, a cross-node storage graph model is constructed using a layered graph partitioning algorithm. The nodes in the graph represent storage nodes, and the edges represent the transmission path weights between nodes. Based on the constructed cross-node storage graph model, the graph optimization algorithm is used to generate the optimal block storage distribution solution; The node storage status and the distribution status of access blocks are monitored in real time through the reinforcement learning algorithm, and the distribution storage strategy of deduplicated data blocks is dynamically adjusted.
6. The method for deduplication of cloud backup platform data based on distributed storage according to claim 5, characterized in that: The formula of the layered graph partitioning algorithm is as follows: Wherein, min F represents the objective function of hierarchical graph partitioning, which is used to describe the optimization goal of data block storage distribution; N represents the total number of data blocks; M represents the total number of storage nodes; represents the set of neighboring data blocks of data block i; w ij represents the transmission weight between data block i and data block j; d ij represents the network delay or transmission path weight between different nodes; f ij represents the similarity characteristics of data blocks i and j; c represents the weight coefficient; L k represents the current load of storage node k; C k Represents the capacity of storage node k.
7. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 1, characterized in that: The step S4 also includes storing the data blocks determined to be duplicate only once, and managing other duplicate references through logical pointers; for newly added non-duplicate data blocks, efficient storage is performed by combining compression algorithms and hot and cold data tiering strategies, and an incremental data flow model is introduced to dynamically optimize the storage layout.
8. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 1, characterized in that: The establishment of the global index supports efficient associative mapping between high-dimensional features and physical storage locations through a multi-dimensional index mapping mechanism combined with a content-aware hash algorithm based on block features.
9. The method for deduplicating data on a cloud backup platform based on distributed storage according to claim 1, characterized in that: The reinforcement learning algorithm is based on a two-layer hierarchical optimization framework. The top layer uses a policy gradient method to optimize the global performance objectives of the deduplication process, and the bottom layer combines a multi-agent collaboration mechanism to achieve coordinated dynamic adjustment of the block strategy, index frequency, and storage mapping.
Citation Information
Patent Citations
Big-data-oriented cloud disaster tolerant backup method
CN104932956A
A block-level data de-restorage system
CN109445702A
Large-scale data storage deduplication optimization method based on machine learning model
CN119003504A
Performing deduplication in a distributed filesystem
US9679040B1
Cited By
Data platform file migration method, computer program product and data platform
CN120277046A
Data platform file migration method, computer program product, and data platform
CN120277046B
Data preservation method for server and application system
CN120315946A
A data security method and application system for a server
CN120315946B
Dynamic geological mass data storage and high-speed retrieval method
CN120448395A