Cloud storage data compression deduplication system and use method thereof
By implementing dynamic adjustment measures for data preprocessing, compression, deduplication, layered storage and resource scheduling in the cloud storage system, the problem of insufficient dynamic adjustment capabilities of data access mode in the existing technology is solved, efficient data compression and deduplication, accurate layered storage of hot and cold data and optimized resource management, and the performance and adaptability of the system are improved.
Patent Information
- Application Number
- CN202411892594.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology lacks the ability to dynamically adjust data access mode, and cannot flexibly select appropriate compression or deduplication strategies based on data type, size or complexity, resulting in insufficient compression rate or deduplication accuracy. In terms of layered storage of hot and cold data, it is difficult to accurately identify hot and cold data, resulting in poor dynamic adaptability.
Through steps such as data preprocessing, data compression, data deduplication, data hierarchical storage and resource scheduling, data are dynamically identified and classified, streaming compression technology and sliding window hashing algorithms, combined with dynamic blocking algorithms and access frequency calculation models, accurate hierarchical storage and resource optimization of data are achieved.
It realizes precise layered storage of hot and cold data, improves the efficiency of data compression and deduplication, optimizes resource scheduling, improves the overall performance and adaptability of the system, and reduces storage space occupation and maintenance costs.
Smart Images

Figure CN120029988A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud storage data compression, and in particular to a cloud storage data compression and deduplication system and a use method thereof. Background Art
[0002] The cloud storage data compression and deduplication system is an efficient solution to the challenge of massive data storage in the big data era. The system significantly improves the efficiency of cloud storage and reduces costs through the two core technologies of data compression and deduplication.
[0003] Data compression technology uses algorithms to reduce data redundancy and saves original information in a smaller storage space. It supports both lossless and lossy compression modes to meet different types of data storage requirements. Data deduplication technology reduces redundant storage in the storage system by identifying and eliminating duplicate data copies. This technology is especially important in cloud environments to control the surge in data volume.
[0004] In addition, the system also includes optimization measures such as tiered storage, data pre-fetching and caching, network transmission optimization and load balancing, which further improve data access speed and overall performance. The application of these comprehensive technologies enables cloud storage services to process and store large-scale data sets more economically and efficiently, meeting the growing demand for data storage by enterprises and individuals.
[0005] The existing technology lacks the ability to dynamically adjust data access patterns and cannot flexibly select appropriate compression or deduplication strategies according to data type, size or complexity, resulting in insufficient compression rate or deduplication accuracy. In terms of tiered storage of hot and cold data, the existing solutions have a relatively simple calculation of data access frequency, making it difficult to accurately identify hot and cold data, resulting in poor dynamic adaptability. Therefore, a cloud storage data compression and deduplication system and a method for using the same are proposed to address the above problems. Summary of the invention
[0006] The purpose of the present invention is to provide a cloud storage data compression and deduplication system and a method of using the same, so as to solve the problem that the prior art lacks the ability to dynamically adjust data access patterns and cannot flexibly select appropriate compression or deduplication strategies according to data type, size or complexity, resulting in insufficient compression rate or deduplication accuracy. In terms of tiered storage of hot and cold data, the existing solutions have a relatively simple calculation of data access frequency, making it difficult to accurately identify hot data and cold data, which in turn leads to poor dynamic adaptability.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A cloud storage data compression and deduplication system and a method for using the same, comprising:
[0009] S1: Data preprocessing: Obtain the format, size and type of the uploaded data, identify and classify the format, size and type of the data, perform block-level deduplication on small data files, and implement block compression strategy on large files;
[0010] S2: Data compression: Perform compression operations on non-duplicate data, use streaming compression technology to process real-time data, and compress large files in blocks with a fixed block size B;
[0011] S3: Data deduplication:
[0012] S31: For each data block b i Calculate the hash value H(b i ), and query the global fingerprint library;
[0013] S32: There is H(b i ), marked as duplicate data;
[0014] S33: If H(b i ), then store the data block and update the fingerprint library;
[0015] S4: Data distribution storage: The compressed data is stored in layers according to the access frequency F(d), satisfying the following conditions:
[0016]
[0017] In the formula, F threshold is the threshold of access frequency;
[0018] Hot data is stored in high-performance storage media, and cold data is stored in low-cost storage media;
[0019] S5: Resource scheduling and optimization: Dynamically adjust the priority of compression and deduplication tasks according to the system resource status:
[0020] Use the task queue management mechanism to assign high-priority tasks to nodes with sufficient resources; adopt parallel processing for real-time tasks.
[0021] As further optimization content of the present invention, among others: in the data preprocessing step, the system will dynamically analyze the traffic pattern of the uploaded real-time streaming data, and according to the characteristics of the traffic peak and trough periods, give priority to the compression and deduplication operations during the off-peak period.
[0022] As a further optimization of the present invention, in the data deduplication step, data block b i The hash value H(b i ) is calculated using a sliding window algorithm based on a variable block size, specifically:
[0023] H(bi )=Hash(b i [k:k+w]),k=0,1,...,|b i |-w
[0024] Where w is the sliding window size, b i [k:k+w] represents data block b i The subsegment at position k.
[0025] As a further optimization of the present invention, the block compression operation in the data compression step adopts a dynamic block algorithm, and the block size B is dynamically adjusted in the following manner:
[0026] B=B min +α·(|b i |-B min )
[0027] Among them, B min is the minimum block size, |b i | is the data block size, α is the dynamic weight factor, and the block size is adjusted according to the complexity of the data block to optimize the compression ratio and speed.
[0028] As a further optimization of the present invention, in the data hierarchical storage step, the calculation formula of the access frequency F(d) is:
[0029]
[0030] In the formula, R represents the number of data reads, W represents the number of data writes, and W r and W w is the weight factor for reading and writing, T is the observation time window size;
[0031] Determine hot data and cold data based on access frequency and dynamically adjust tiering strategies.
[0032] As a further optimized content of the present invention, in the source scheduling and optimization step, the task priority calculation formula is:
[0033]
[0034] Where P(t) is the priority of task t, L(t) is the amount of remaining data of the task, C(t) is the estimated computing resource requirement of the task, and β is the regularization parameter to prevent the denominator from being zero.
[0035] Prioritize scheduling tasks with higher priority P(t) to computing nodes.
[0036] As further optimized content of the present invention, it includes: a data preprocessing module: used to identify the format, size classification and duplicate detection of the uploaded data, and select appropriate compression and deduplication strategies according to the data file type and size;
[0037] Compression module: adopts block compression strategy to divide large files into blocks of fixed size, compresses each block independently, and supports streaming compression technology for real-time data processing;
[0038] Deduplication module: Based on the global fingerprint library, duplicate data detection is achieved by calculating the hash value of files or data blocks, and fingerprint calculation is only performed on newly added or modified data blocks;
[0039] Tiered storage module: used to perform hot and cold tiered storage of data according to data access frequency: hot data is stored using a low compression algorithm, and cold data is stored using a high compression algorithm;
[0040] Resource scheduling module: monitors system resource status in real time and dynamically adjusts the priority of compression and deduplication tasks;
[0041] Metadata management module: records data compression, deduplication and storage location information, and supports rapid location and management of duplicate reference data.
[0042] As a further optimization of the present invention, the compression module supports a block compression strategy, which divides a large file into multiple data blocks b according to a fixed size B. 1 ,b 2 ,...,b n , each data block is compressed and stored independently to reduce memory usage and improve compression efficiency.
[0043] As a further optimized content of the present invention, the deduplication module implements deduplication by the following steps:
[0044] Step 1: Calculate data block b i The hash value H(b i );
[0045] Step 2: Query the global fingerprint database. If H(b i ) exists, it is marked as duplicate data and the reference count is updated;
[0046] Step 3: If H(b i ) does not exist, then b i Stored as a unique data block and record its fingerprint information.
[0047] As a further optimized content of the present invention, the hierarchical storage module determines hot data and cold data by access frequency threshold, and regularly cleans up the long-term unaccessed parts of the cold data.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] 1. In the present invention, accurate hierarchical storage of hot and cold data is realized through a multi-dimensional calculation model of access frequency: the access frequency is calculated according to the weight factors of read and write operations, and the classification strategy of hot data and cold data is dynamically adjusted; hot data is stored in high-performance media, and a low compression rate algorithm is preferentially used to ensure access speed; cold data is stored in low-cost media, and a high compression rate algorithm is used to save storage space. In addition, through traffic pattern analysis, compression and deduplication tasks are prioritized during off-peak periods to reduce resource pressure during peak periods and improve system adaptability;
[0050] 2. In the present invention, the efficiency of data compression and deduplication is improved by combining data segmentation, streaming compression technology and sliding window hash algorithm: for small files, block-level deduplication technology is adopted to effectively reduce redundant storage space; for large files, a dynamic segmentation algorithm is adopted to dynamically adjust the block size according to the complexity of the data, further optimizing the compression ratio and processing speed; the sliding window hash algorithm is introduced to significantly improve the accuracy of duplicate data detection, especially in scenarios with various data alignment methods, the efficiency of duplicate removal is significantly improved;
[0051] 3. In the present invention, the remaining data volume of the task is combined with the estimated computing resource demand through the task priority calculation formula, and the execution order of the tasks is dynamically adjusted: high-priority tasks are preferentially allocated to nodes with sufficient resources, which improves the overall processing efficiency of the system, and adopts parallel processing technology for real-time tasks, which reduces the delay in data uploading and processing;
[0052] 4. In the present invention, through modular design and optimization of metadata management, centralized management of compression, deduplication, and storage location information is achieved, which is convenient for rapid positioning and updating of data: the incremental deduplication mechanism only calculates fingerprints for newly added or modified data blocks, reducing unnecessary repeated operations, and regularly performs cleaning optimization on cold data to release storage space and further reduce maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A flowchart of a cloud storage data compression and deduplication system and a method for using the same according to the present invention;
[0054] Figure 2 The present invention is a system block diagram of a cloud storage data compression and deduplication system and a method for using the same. DETAILED DESCRIPTION
[0055] See also Figure 1-2 , the present invention provides a technical solution:
[0056] A cloud storage data compression and deduplication system and a method for using the same, comprising:
[0057] S1: Data preprocessing: Obtain the format, size and type of the uploaded data, identify and classify the format, size and type of the data, perform block-level deduplication on small data files, and implement block compression strategy on large files;
[0058] S2: Data compression: Perform compression operations on non-duplicate data, use streaming compression technology to process real-time data, and compress large files in blocks with a fixed block size B;
[0059] S3: Data deduplication:
[0060] S31: For each data block b i Calculate the hash value H(b i ), and query the global fingerprint library;
[0061] S32: There is H(b i ), marked as duplicate data;
[0062] S33: If H(b) does not exist in the fingerprint database i ), then store the data block and update the fingerprint library;
[0063] S4: Data distribution storage: The compressed data is stored in layers according to the access frequency F(d), satisfying the following conditions:
[0064]
[0065] In the formula, F threshold is the threshold of access frequency;
[0066] Hot data is stored in high-performance storage media, and cold data is stored in low-cost storage media;
[0067] S5: Resource scheduling and optimization: Dynamically adjust the priority of compression and deduplication tasks according to the system resource status:
[0068] Use the task queue management mechanism to assign high-priority tasks to nodes with sufficient resources; use parallel processing for real-time tasks. Through the comprehensive design of five steps, including data preprocessing, compression, deduplication, tiered storage, and resource scheduling, the system can efficiently manage uploaded data, reduce storage space usage, optimize data access and resource allocation, and improve the overall performance and adaptability of the cloud storage system.
[0069] As a technical solution for further implementation of this solution, in the data preprocessing step, the system will dynamically analyze the traffic pattern of the uploaded real-time streaming data, and give priority to compression and deduplication operations during the off-peak period according to the characteristics of the traffic peak and trough periods. Dynamic analysis is performed on the traffic pattern of the real-time streaming data, and compression and deduplication operations are given priority during the off-peak period to reduce the system pressure during the peak period, while optimizing storage performance and improving the processing efficiency of the system in a dynamic traffic environment;
[0070] As a technical solution for further implementing this solution, in the data deduplication step, data block b i The hash value H(b i ) is calculated using a sliding window algorithm based on a variable block size, specifically:
[0071] H(b i )=Hash(b i [k:k+w]),k=0,1,...,|b i |-w
[0072] Where w is the sliding window size, b i [k:k+w] represents data block b i In the sub-segment at position k, a sliding window algorithm is used to calculate the hash value, which improves the efficiency and flexibility of deduplication. In particular, in scenarios with various duplicate data alignment methods, duplicate data blocks can be detected and removed more accurately, further reducing storage overhead.
[0073] As a technical solution for further implementing this solution, the block compression operation in the data compression step adopts a dynamic block algorithm, and the block size B is dynamically adjusted in the following manner:
[0074] B=B min +α·(|b i |-B min )
[0075] Among them, B min is the minimum block size, |b i | is the data block size, α is the dynamic weight factor, and the block size is adjusted according to the complexity of the data block to optimize the compression ratio and speed. By dynamically adjusting the block size, the compression ratio and speed are optimized according to the data complexity, which not only meets the needs of large file processing, but also significantly reduces the compression time and resource occupation while ensuring the compression effect;
[0076] As a technical solution for further implementing this solution, in the data tiered storage step, the calculation formula for the access frequency F(d) is:
[0077]
[0078] In the formula, R represents the number of data reads, W represents the number of data writes, and W r and W w is the weight factor for reading and writing, T is the observation time window size;
[0079] Hot data and cold data are determined based on access frequency, and the tiering strategy is adjusted dynamically. The access frequency calculation introduces read and write weight factors, so that the tiered storage strategy can dynamically adapt to different data usage scenarios, effectively improving the access efficiency of high-frequency access data while reducing the storage cost of low-frequency data;
[0080] As a technical solution for further implementation of this solution, in the source scheduling and optimization step, the task priority calculation formula is:
[0081]
[0082] Where P(t) is the priority of task t, L(t) is the amount of remaining data of the task, C(t) is the estimated computing resource requirement of the task, and β is the regularization parameter to prevent the denominator from being zero.
[0083] Prioritize the scheduling of tasks with higher priority P(t) to computing nodes. By using the task priority calculation formula, the remaining data volume of the task is combined with the computing resource requirements to reasonably allocate computing resources, improve the intelligence of task scheduling, ensure that high-priority tasks are processed in a timely manner, and optimize the overall system performance.
[0084] As a technical solution for further implementation of this solution, it includes: data preprocessing module: used to identify the format, size classification and duplicate detection of uploaded data, and select appropriate compression and deduplication strategies according to the data file type and size;
[0085] Compression module: adopts block compression strategy to divide large files into blocks of fixed size, compresses each block independently, and supports streaming compression technology for real-time data processing;
[0086] Deduplication module: Based on the global fingerprint library, duplicate data detection is achieved by calculating the hash value of files or data blocks, and fingerprint calculation is only performed on newly added or modified data blocks;
[0087] Tiered storage module: used to perform hot and cold tiered storage of data according to data access frequency: hot data is stored using a low compression algorithm, and cold data is stored using a high compression algorithm;
[0088] Resource scheduling module: monitors system resource status in real time and dynamically adjusts the priority of compression and deduplication tasks;
[0089] Metadata management module: records data compression, deduplication, and storage location information, supports rapid location and management of duplicated data, and has clear system functions through modular design. It covers data preprocessing, compression, deduplication, hierarchical storage, resource scheduling, and metadata management. It can efficiently perform data compression and deduplication tasks and adapt to diverse cloud storage needs.
[0090] As a technical solution for further implementation of this solution, the compression module supports a block compression strategy, which divides a large file into multiple data blocks b according to a fixed size B. 1 ,b 2 ,...,b n , each data block is compressed and stored independently to reduce memory usage and improve compression efficiency. It supports block compression strategy and dynamic selection of compression algorithm, which greatly reduces memory usage and significantly improves compression efficiency and flexibility. It is particularly suitable for cloud storage scenarios that process large files;
[0091] As a technical solution for further implementing this solution, the deduplication module implements deduplication through the following steps:
[0092] Step 1: Calculate data block b i The hash value H(b i );
[0093] Step 2: Query the global fingerprint database. If H(b i ) exists, it is marked as duplicate data and the reference count is updated;
[0094] Step 3: If H(b i ) does not exist, then b i Store as a unique data block and record its fingerprint information, use the global fingerprint library to quickly mark duplicate data and update reference relationships, ensure the accuracy and efficiency of data deduplication, and further reduce redundant storage and computing resource waste;
[0095] As a technical solution for further implementation of this solution, the tiered storage module determines hot data and cold data through the access frequency threshold and regularly cleans up the long-term unaccessed parts of the cold data. The tiered storage module regularly cleans up the long-term unaccessed parts of the cold data to effectively release storage space. At the same time, combined with the priority storage strategy for hot data, it optimizes the storage structure of high-frequency and low-frequency data and improves the storage utilization of the system.
[0096] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the method and core ideas of the present invention. The above is only a preferred implementation of the present invention. It should be pointed out that due to the limitations of textual expression and the objective existence of infinite specific structures, ordinary technicians in this technical field can make several improvements, modifications or changes without departing from the principles of the present invention, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the protection scope of the present invention.
Claims
1. A method for using a cloud storage data compression and deduplication system, characterized in that: include: S1: Data preprocessing: Obtain the format, size and type of the uploaded data, identify and classify the format, size and type of the data, perform block-level deduplication on small data files, and implement block compression strategy on large files; S2: Data compression: Perform compression operations on non-duplicate data, use streaming compression technology to process real-time data, and compress large files in blocks with a fixed block size B; S3: Data deduplication: S31: For each data block b i Calculate the hash value H(b i ), and query the global fingerprint library; S32: There is H(b i ), marked as duplicate data; S33: If H(b i ), then store the data block and update the fingerprint library; S4: Data distribution storage: The compressed data is stored in layers according to the access frequency F(d), satisfying the following conditions: In the formula, F threshold is the threshold of access frequency; Hot data is stored in high-performance storage media, and cold data is stored in low-cost storage media; S5: Resource scheduling and optimization: Dynamically adjust the priority of compression and deduplication tasks according to the system resource status: Use the task queue management mechanism to assign high-priority tasks to nodes with sufficient resources; adopt parallel processing for real-time tasks.
2. A method for using a cloud storage data compression and deduplication system according to claim 1, characterized in that: In the data preprocessing step, the system dynamically analyzes the traffic pattern of the uploaded real-time streaming data, and prioritizes compression and deduplication operations during the low-peak period based on the characteristics of the traffic peak and low-peak periods.
3. The method for using the cloud storage data compression and deduplication system according to claim 1, characterized in that: In the data deduplication step, data block b i The hash value H(b i ) is calculated using a sliding window algorithm based on a variable block size, specifically: H(b i )=Hash(b i [k:k+w]),k=0,1,...,|b i |-w Where w is the sliding window size, b i [k:k+w] represents data block b i The subsegment at position k.
4. The method for using the cloud storage data compression and deduplication system according to claim 1, characterized in that: The block compression operation in the data compression step adopts a dynamic block algorithm, and the block size B is dynamically adjusted in the following manner: B=B min +a·(|b i |-B min ) Among them, B min is the minimum block size, |b i | is the data block size, α is the dynamic weight factor, and the block size is adjusted according to the complexity of the data block to optimize the compression ratio and speed.
5. The method for using the cloud storage data compression and deduplication system according to claim 1, characterized in that: In the data hierarchical storage step, the calculation formula of the access frequency F(d) is: In the formula, R represents the number of data reads, W represents the number of data writes, and W r and W w is the weight factor for reading and writing, T is the observation time window size; Determine hot data and cold data based on access frequency and dynamically adjust tiering strategies.
6. The method for using the cloud storage data compression and deduplication system according to claim 1, characterized in that: In the source scheduling and optimization step, the task priority calculation formula is: Where P(t) is the priority of task t, L(t) is the amount of remaining data of the task, C(t) is the estimated computing resource requirement of the task, and β is the regularization parameter to prevent the denominator from being zero. Prioritize scheduling tasks with higher priority P(t) to computing nodes.
7. A cloud storage data compression and deduplication system according to any one of claims 1 to 6, characterized in that: include: Data preprocessing module: used to identify the format, size classification and duplicate detection of uploaded data, and select appropriate compression and deduplication strategies according to the data file type and size; Compression module: adopts block compression strategy to divide large files into blocks of fixed size, compresses each block independently, and supports streaming compression technology for real-time data processing; Deduplication module: Based on the global fingerprint library, duplicate data detection is achieved by calculating the hash value of files or data blocks, and fingerprint calculation is only performed on newly added or modified data blocks; Tiered storage module: used to perform hot and cold tiered storage of data according to data access frequency: hot data is stored using a low compression algorithm, and cold data is stored using a high compression algorithm; Resource scheduling module: monitors system resource status in real time and dynamically adjusts the priority of compression and deduplication tasks; Metadata management module: records data compression, deduplication and storage location information, and supports rapid location and management of duplicate reference data.
8. The cloud storage data compression and deduplication system according to claim 1, characterized in that: The compression module supports a block compression strategy, which divides a large file into multiple data blocks b1, b2, ..., b according to a fixed size B. n , each data block is compressed and stored independently to reduce memory usage and improve compression efficiency.
9. The cloud storage data compression and deduplication system according to claim 1, characterized in that: The deduplication module implements deduplication through the following steps: Step 1: Calculate data block b i The hash value H(b i ); Step 2: Query the global fingerprint database. If H(b i ) exists, it is marked as duplicate data and the reference count is updated; Step 3: If H(b i ) does not exist, then b i Stored as a unique data block and record its fingerprint information.
10. The cloud storage data compression and deduplication system according to claim 1, characterized in that: The hierarchical storage module determines hot data and cold data by using an access frequency threshold, and regularly clears the long-term unaccessed portion of the cold data.
Citation Information
Cited By
Variable-length block compression and encryption supporting source end deduplication method and variable-length block compression and encryption supporting source end deduplication system
CN120523779A
Data management method and system of intelligent handheld terminal
CN121116934A
Mass data deduplication method and system
CN121411712A
A method and system for deduplication of massive amounts of data
CN121411712B
Industrial unstructured data processing method and system
CN121743278A