Data deduplication method and device, equipment and medium

Through monitoring and adaptive selection of deletion strategies, the problem that deletion technology cannot dynamically adapt to I/O load changes is solved, and the storage performance optimization and resource utilization are optimized, and the elastic expansion capabilities of the storage system are enhanced.

CN120540600APending Publication Date: 2025-08-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510682406.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing re-deletion technology solutions cannot dynamically adapt to I/O load changes, resulting in a degradation of storage performance at high loads, insufficient resource utilization at low loads, and lack of elastic expansion capabilities.

Method used

By monitoring the utilization rate of hardware resources, including the utilization rate of the central processor, memory rate, interconnection bandwidth rate between nodes and back-end disk rate, the optimal deletion strategy is adaptively selected in combination with the preset threshold value, and data deletion is performed, including the timing of deletion, deletion range, blocking strategy and hashing algorithm, etc., to dynamically adjust the deletion process.

Benefits of technology

It realizes dynamically adapting to I/O load changes, optimizing storage performance, improving resource utilization, and enhancing the elastic expansion capabilities of the storage system without sacrificing storage performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540600A_ABST
    Figure CN120540600A_ABST
Patent Text Reader

Abstract

The invention discloses a data deduplication method and device, equipment and a medium, and relates to the technical field of storage. According to the scheme, when the I / O request issued by the host side is received, the optimal deduplication strategy can be adaptively selected according to the hardware resource utilization rate under the current I / O load in combination with the corresponding hardware resource utilization rate threshold value, data deduplication is executed, dynamic adaptation to I / O load changes is achieved, the deduplication rate is guaranteed on the premise that the storage performance is not sacrificed, and the data deduplication efficiency is improved. The storage space is saved, the storage performance is optimized, and the storage system has higher elastic expansion capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and in particular to a data deduplication method, device, equipment and medium. Background Art

[0002] Deduplication is a core technology that optimizes storage space by identifying and eliminating redundant data. Its fundamental principle is to replace duplicate copies of data blocks with unique ones, thereby reducing storage resource redundancy. This technology is widely used in backup systems, cloud storage, virtualization, and cold data archiving. Deduplication efficiency is closely related to system resources such as the central processing unit (CPU), memory, storage input / output (I / O), and network bandwidth.

[0003] However, current deduplication solutions typically rely on static parameter configuration or simple on-the-ground adjustments, making them difficult to dynamically adapt to changes in I / O loads and failing to fully consider storage system performance bottlenecks. This results in storage performance degradation under high loads due to deduplication computational overhead, while under low loads, hardware resource utilization is insufficient and elastic scalability is lacking.

[0004] In view of the above, how to solve the problem that the current deduplication solution cannot dynamically adapt to changes in I / O load, resulting in decreased storage performance under high load, insufficient resource utilization under low load, and lack of elastic scalability, is an urgent problem that needs to be solved by technical personnel in this field. Summary of the Invention

[0005] The present invention provides a data deduplication method, device, equipment and medium to at least solve the problem that the current deduplication solution cannot dynamically adapt to I / O load changes, resulting in decreased storage performance under high load, insufficient resource utilization under low load, and lack of elastic expansion capability.

[0006] The present invention provides a data deduplication method, which is applied to a storage system; the method comprises:

[0007] When receiving an input / output request from the host side, monitoring hardware resource utilization according to a first preset period; wherein the hardware resource utilization includes at least central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization and backend disk utilization;

[0008] Obtain the hardware resource utilization threshold corresponding to each hardware resource utilization;

[0009] Determine a deduplication policy based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold, and perform data deduplication based on the deduplication policy and input / output requests.

[0010] The present invention also provides a data deduplication device, which is applied to a storage system; the device comprises:

[0011] A monitoring module configured to monitor hardware resource utilization according to a first preset period upon receiving an input / output request from the host side; wherein the hardware resource utilization includes at least central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization, and backend disk utilization;

[0012] An acquisition module is used to obtain a hardware resource utilization threshold corresponding to each hardware resource utilization;

[0013] The deduplication module is used to determine a deduplication strategy based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold, and perform data deduplication based on the deduplication strategy and input / output requests.

[0014] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data deduplication methods when executing the computer program.

[0015] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data deduplication methods are implemented.

[0016] The beneficial effect of the present invention is that when an I / O request sent by the host side is received, it can adaptively select the optimal deduplication strategy and execute data deduplication based on the hardware resource utilization under the current I / O load and the corresponding hardware resource utilization threshold, thereby achieving dynamic adaptation to I / O load changes, ensuring the deduplication rate without sacrificing storage performance, saving storage space, optimizing storage performance, and enabling the storage system to have higher elastic expansion capabilities.

[0017] In addition, the present invention also provides a data deduplication device, equipment and medium, which have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A schematic diagram of a data deduplication architecture provided by an embodiment of the present invention;

[0020] Figure 2 A flowchart of a data deduplication method provided by an embodiment of the present invention;

[0021] Figure 3 A schematic diagram of a data deduplication device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0023] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0025] Current deduplication solutions typically rely on static parameter configuration or simple on-the-ground adjustments, making them difficult to dynamically adapt to changes in I / O load and failing to fully consider storage system performance bottlenecks. This results in decreased storage performance under high load conditions due to deduplication computational overhead, while under low load conditions, hardware resource utilization is insufficient and elastic scalability is lacking. To address these issues, the present invention provides a data deduplication method.

[0026] It should be noted that the method provided by the present invention is applied to a storage system. In this system, a host uses a switch to connect to a storage system via protocols such as the Internet Small Computer System Interface (ISCSI), Fibre Channel-Small Computer System Interface (FC-SCSI), and Fibre Channel-Non-Volatile Memory Express (FC-NVME), forming a storage area network (SAN) system. The host issues I / O to the storage system, which is responsible for processing and scheduling deduplication I / O and storing the deduplicated data.

[0027] Figure 1 This is a schematic diagram of a data deduplication architecture provided by an embodiment of the present invention. Figure 1 As shown, the storage system has four modules that jointly implement the data deduplication method: a configuration module, a real-time resource monitoring module, a multi-dimensional deduplication feature selection module, and a deduplication execution module. The following describes the data deduplication method in detail, combining each module.

[0028] Figure 2 A flow chart of a data deduplication method provided by an embodiment of the present invention. Figure 2 As shown, the method includes:

[0029] S10: When receiving an input / output request sent by the host side, monitoring hardware resource utilization according to a first preset period.

[0030] Among them, hardware resource utilization includes at least central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization and backend disk utilization.

[0031] Specifically, after the host sends an I / O request to the storage system, the real-time resource monitoring module monitors hardware resource utilization according to a first preset period. It should be noted that in this embodiment, there is no restriction on the first preset period; for example, it can be 1 second. Furthermore, hardware resource utilization includes at least CPU utilization, memory utilization, inter-node interconnect bandwidth utilization, and backend disk utilization, and may also include other types of hardware resource utilization, depending on the specific implementation.

[0032] It should be noted that CPU utilization can be obtained using the top command. The ratio of the percentage of CPU occupied by the main processing logic process (plmain) to the number of CPU cores is the CPU utilization. Memory utilization is the ratio of the memory occupied by the current plmain process to the total memory size; inter-node interconnection bandwidth utilization is the ratio of the current node interconnection bandwidth to the total node interconnection bandwidth; backend disk utilization is the percentage of disk busy time. For example, in a striped RAID 5 system, the time a single disk spends processing I / O within 1 second is the backend disk utilization.

[0033] S11: Obtaining a hardware resource utilization threshold corresponding to each hardware resource utilization.

[0034] Furthermore, the real-time resource monitoring module obtains hardware resource utilization thresholds corresponding to each hardware resource utilization, pre-configured by the configuration module. For example, hardware resource utilization includes at least CPU utilization, memory utilization, inter-node interconnection bandwidth utilization, and backend disk utilization. Therefore, the corresponding hardware resource utilization thresholds should include at least the CPU utilization threshold, the memory utilization threshold, the inter-node interconnection bandwidth utilization threshold, and the backend disk utilization threshold. In this embodiment, the specific values ​​of the hardware resource utilization thresholds are not limited and are determined based on specific implementation circumstances.

[0035] S12: Determine a deduplication policy based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold, and perform data deduplication based on the deduplication policy and the input / output request.

[0036] After obtaining the hardware resource utilization and the corresponding threshold, the real-time resource monitoring module reports the hardware resource utilization and the corresponding threshold to the multi-dimensional deduplication feature selection module when the hardware resource utilization crosses the corresponding threshold. The multi-dimensional deduplication feature selection module has a built-in multi-dimensional deduplication feature selection function, which can determine the deduplication strategy based on the hardware resource utilization and the corresponding hardware resource utilization threshold, and report the deduplication strategy to the deduplication execution module. It should be noted that this embodiment does not limit the specific process of determining the deduplication strategy based on the hardware resource utilization and the corresponding hardware resource utilization threshold.

[0037] Finally, the deduplication execution module incorporates a built-in deduplication domain, a hash algorithm, and a hash fingerprint index. After obtaining the deduplication policy selected by the multi-dimensional deduplication feature selection module, it executes data deduplication based on the deduplication policy and I / O request. The deduplication process, when processing an I / O request, calculates the hash fingerprint value of a data block to check whether a data block with the same fingerprint already exists in the storage system. If the fingerprint already exists, the data block does not need to be re-stored. Instead, a pointer to the new metadata is added to the logical block address (LBA) of the stored data block, thereby preventing duplicate data from being written to disk and improving disk utilization.

[0038] In this embodiment, when an I / O request is received from the host side, the optimal deduplication strategy can be adaptively selected and data deduplication can be performed based on the hardware resource utilization under the current I / O load and the corresponding hardware resource utilization threshold, thereby achieving dynamic adaptation to I / O load changes, ensuring the deduplication rate without sacrificing storage performance, saving storage space, optimizing storage performance, and enabling the storage system to have higher elastic expansion capabilities.

[0039] In order to improve the performance of data deduplication, based on the above embodiment, in some embodiments, before monitoring the hardware resource utilization in the system according to the first preset period, the method further includes:

[0040] S13: Configure the initial hardware resource utilization threshold corresponding to each hardware resource utilization.

[0041] The initial hardware resource utilization thresholds include at least an initial CPU utilization threshold, an initial memory utilization threshold, an initial inter-node interconnection bandwidth utilization threshold, and an initial backend disk utilization threshold.

[0042] S14: Determine the overall status and hardware type of the system.

[0043] The overall status includes at least a single-controller cluster, a multi-controller cluster, a common cluster, and an active-active cluster; the hardware type includes at least the CPU model and the backend disk type.

[0044] S15: Adjust each initial hardware resource utilization threshold according to the overall state and hardware type to generate each hardware resource utilization threshold.

[0045] Specifically, when configuring the hardware resource utilization thresholds in the configuration module, the initial hardware resource utilization thresholds corresponding to the hardware resource utilizations are first configured. It is understood that the initial hardware resource utilization thresholds include at least the initial CPU utilization threshold, the initial memory utilization threshold, the initial inter-node interconnection bandwidth utilization threshold, and the initial backend disk utilization threshold.

[0046] Furthermore, determine the overall status and hardware type of the system. It should be noted that the overall status includes at least single-controller clusters, multi-controller clusters, standard clusters, and active-active clusters. A single-controller cluster is a cluster managed by a single controller, with a simple structure but the risk of single point failure. A multi-controller cluster contains multiple controllers, improving system availability and performance through load balancing and failover. A standard cluster generally refers to a cluster with basic redundancy and fault recovery capabilities, which may include a single-controller or multi-controller structure. An active-active cluster is a cluster in which two or more nodes are simultaneously active, jointly processing requests, providing high availability and load balancing. At the same time, the hardware type includes at least the CPU model and backend disk type, and may also include other types.

[0047] It's worth noting that the configuration module has a built-in deduplication feature offset function. After obtaining the overall status of the storage system, it pre-determines the sensitive characteristics of the storage and adjusts the initial hardware resource utilization thresholds based on the overall status and hardware type. Specifically, the initial hardware resource utilization thresholds are offset by a certain percentage to generate the final hardware resource utilization thresholds, thereby effectively improving data deduplication performance. The following is a detailed explanation:

[0048] In some embodiments, adjusting each initial hardware resource utilization threshold based on the overall status and hardware type includes:

[0049] S151: When the overall state is a single-controller cluster, increase the initial node interconnection bandwidth utilization threshold.

[0050] S152: When the overall state is an active-active cluster, reduce the initial CPU utilization threshold.

[0051] S152: When the CPU model is the target model, increase the initial CPU utilization threshold.

[0052] S153: When the backend disk type is a mechanical hard disk, reduce the initial backend disk utilization threshold.

[0053] Specifically, when the overall status of the storage system is determined to be a single-controller cluster, since a single-controller cluster has no inter-node communication, the inter-node communication bandwidth utilization threshold can be increased; while an active-active cluster focuses more on real-time synchronization between data and has high requirements for I / O processing speed, so the CPU utilization threshold is appropriately lowered to improve data deduplication efficiency.

[0054] Furthermore, the CPU model is obtained. If the CPU model is the target model, the initial CPU utilization threshold is increased. In this embodiment, there is no restriction on the target model. Finally, the backend disk type is obtained. If the backend disk type is a mechanical hard disk, fragmented data blocks will cause disk addressing difficulties, seriously affecting disk performance. In this case, the backend disk utilization threshold can be lowered to reduce the deduplication rate and avoid excessive data block fragmentation.

[0055] Based on the above embodiments, in some embodiments, a deduplication strategy is determined based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold, including:

[0056] S120: Determine the deduplication timing, deduplication range, block strategy, hash algorithm, and hot and cold data indexes based on the utilization rates of each hardware resource and the corresponding hardware resource utilization thresholds.

[0057] Among them, the deduplication timing includes online deduplication and background deduplication; the deduplication scope includes global deduplication and local deduplication; the blocking strategy includes fixed blocking and variable blocking; the hash algorithm includes a first hash algorithm and a second hash algorithm, and the collision rate and calculation speed of the first hash algorithm are respectively lower than the collision rate and calculation speed of the second hash algorithm.

[0058] In order to determine the deduplication strategy, in this embodiment, the deduplication timing, deduplication range, block strategy, hash algorithm and hot and cold data indexes are determined based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold.

[0059] It should be noted that deduplication timing includes online deduplication and background deduplication; online deduplication is to perform deduplication in real time when data is written to the storage system, reducing storage space usage, but may increase write latency; background deduplication is to perform deduplication through background tasks after data is written, which does not affect write performance but may cause temporary expansion of storage space. Deduplication scope includes global deduplication and local deduplication; global deduplication is to perform deduplication in the entire storage system, identify and eliminate all duplicate data blocks, and minimize storage space usage; local deduplication is to perform deduplication within a specific range (such as a single volume or directory), which is suitable for specific application scenarios and reduces computing overhead. Block strategies include fixed blocking and variable blocking; fixed blocking is to divide data into blocks of fixed size for deduplication, which is simple to implement but may not be able to effectively handle data changes; variable blocking is to dynamically adjust the block size according to the data content to improve deduplication efficiency, and is especially suitable for scenarios with frequent data changes; hash algorithms include a first hash algorithm and a second hash algorithm, and the collision rate and calculation speed of the first hash algorithm are respectively lower than the collision rate and calculation speed of the second hash algorithm. In this embodiment, there is no restriction on the specific types of the first hash algorithm and the second hash algorithm. For example, the first hash algorithm can be SHA-256, and the second hash algorithm can be xxHash64+xxHash3. The first hash algorithm provides high security and extremely low collision probability, making it suitable for security-sensitive applications, but it has a slower computational speed. The second hash algorithm provides extremely fast hash computation speed, making it suitable for performance-critical scenarios, but it has a relatively high collision rate. Finally, hot and cold data indexing refers to the optimization of hot and cold data tiered indexing. Specifically, hot and cold data tiering divides data into cold and hot data based on access frequency, storing them on different storage media to optimize storage performance and cost. At the same time, by optimizing the index structure, the retrieval efficiency of hash fingerprints is improved, reducing the computational overhead of deduplication operations.

[0060] Therefore, in this solution, the decision on which deduplication strategy to use is made based on the sensitivity of each deduplication technology dimension to different hardware resource utilization rates, so that it can be applied to working conditions under different I / O loads. The specific process of determining the deduplication strategy is described in detail below with reference to specific embodiments:

[0061] In some embodiments, the deduplication timing, deduplication range, block strategy, hash algorithm, and hot and cold data indexes are determined based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold, including:

[0062] S121: When the CPU utilization is less than the first CPU utilization threshold and the memory utilization is less than the first memory utilization threshold, determine that the deduplication opportunity is online deduplication.

[0063] S122: When the CPU utilization is not less than the first CPU utilization threshold, and / or the memory utilization is not less than the first memory utilization threshold, determine that the deduplication opportunity is background deduplication.

[0064] Specifically, deduplication occurs in two different scenarios: online and background. Online deduplication requires real-time hash calculation and hash fingerprint comparison, increasing CPU and memory overhead. Hash calculation and hash fingerprint comparison consume approximately 15%-30% of CPU power. When the fingerprint index resides in memory, approximately 1-3GB of memory space is required per TB of data. This is largely independent of inter-node interconnection bandwidth and the status of the backend disk.

[0065] Therefore, when the CPU utilization is less than the first CPU utilization threshold and the memory utilization is less than the first memory utilization threshold, online deduplication is determined to be the optimal deduplication time, prioritizing ensuring the deduplication rate without affecting the performance of the storage system. When the CPU utilization is not less than the first CPU utilization threshold and / or the memory utilization is not less than the first memory utilization threshold, background deduplication is determined to be the optimal deduplication time, prioritizing ensuring the performance of the storage system.

[0066] S123: When the CPU utilization is less than the third CPU utilization threshold, the memory utilization is less than the first memory utilization threshold, the inter-node interconnection bandwidth utilization is less than the inter-node interconnection bandwidth utilization threshold, and the backend disk utilization is less than the backend disk utilization threshold, the deduplication range is determined to be global deduplication.

[0067] S124: When the CPU utilization is not less than the third CPU utilization threshold, and / or the memory utilization is not less than the first memory utilization threshold, and / or the inter-node interconnection bandwidth utilization is not less than the inter-node interconnection bandwidth utilization threshold, and / or the back-end disk utilization is not less than the back-end disk utilization threshold, the deduplication range is determined to be local deduplication.

[0068] Secondly, deduplication is divided into global and local deduplication. For storage systems, global deduplication corresponds to a pool, while local deduplication corresponds to a volume (vdisk). Global deduplication achieves system-wide deduplication by synchronizing hash fingerprints across nodes. This increases CPU consumption and requires storing more hash fingerprints. Each hash fingerprint comparison or write of deduplicated data requires a large amount of cache data, cross-node data acquisition and transmission, and certain inter-node bandwidth requirements. Furthermore, more hash fingerprints need to be stored and read.

[0069] Therefore, when the CPU utilization is less than the third CPU utilization threshold, the memory utilization is less than the first memory utilization threshold, the inter-node interconnection bandwidth utilization is less than the inter-node interconnection bandwidth utilization threshold, and the back-end disk utilization is less than the back-end disk utilization threshold, the deduplication scope is determined to be global deduplication. Global deduplication will greatly improve the deduplication rate, and the measured deduplication rate can reach 65%-80%. When the CPU utilization is not less than the third CPU utilization threshold, and / or the memory utilization is not less than the first memory utilization threshold, and / or the inter-node interconnection bandwidth utilization is not less than the inter-node interconnection bandwidth utilization threshold, and / or the back-end disk utilization is not less than the back-end disk utilization threshold, the deduplication scope is determined to be local deduplication, thereby prioritizing storage performance.

[0070] S125: When the CPU utilization is less than the second CPU utilization threshold, and the memory utilization is less than the second memory utilization threshold, determining that the blocking strategy is variable blocking.

[0071] S126: When the CPU utilization is not less than the second CPU utilization threshold, and / or the memory utilization is not less than the second memory utilization threshold, determine that the blocking strategy is fixed blocking.

[0072] Furthermore, chunking strategies are categorized into fixed chunking and variable chunking. Variable chunking dynamically partitions deduplication blocks based on data content. It uses a sliding window algorithm (such as Rabin fingerprinting) to scan the data stream. Chunking is triggered when preset conditions (such as a hash modulus match) are met, generating data chunks of varying sizes (typically ranging from 4KB to 64KB). Rabin fingerprint calculation consumes 5%-10% of CPU resources. Memory consumption for variable chunking primarily comes from sliding window state maintenance, hash fingerprint index storage, and parallel computation cache consumption, and is unrelated to inter-node interconnect bandwidth and backend disk utilization.

[0073] Therefore, when the CPU utilization is less than the second CPU utilization threshold and the memory utilization is less than the second memory utilization threshold, the block strategy is determined to be variable block. When the CPU utilization is not less than the second CPU utilization threshold and / or the memory utilization is not less than the second memory utilization threshold, the block strategy is determined to be fixed block, and the block size is 64K.

[0074] S127: When the CPU utilization is less than the second CPU utilization threshold, the memory utilization is less than the third memory utilization threshold, and the backend disk utilization is less than the backend disk utilization threshold, determine that the hash algorithm is the first hash algorithm.

[0075] S128: When the CPU utilization is not less than the second CPU utilization threshold, and / or the memory utilization is not less than the third memory utilization threshold, and / or the backend disk utilization is not less than the backend disk utilization threshold, determine that the hash algorithm is the second hash algorithm.

[0076] The main hash algorithms used in storage systems are SHA-256 and xxHash64+xxHash3. SHA-256 is a computationally intensive algorithm, with advantages in low collision rate and tamper resistance. xxHash64 is a non-cryptographic hash algorithm designed for high speed and low collision rate; its core advantage lies in its extremely fast computation speed. SHA-256 requires calculating the 256-bit hash value of deduplicated data blocks, placing high demands on CPU performance while consuming very little additional memory. SHA-256 requires an additional 32 bytes of storage per data block, while xxHash64 requires 8 bytes per block. Therefore, SHA-256 has a certain impact on backend disk performance.

[0077] Therefore, when the CPU utilization is less than the second CPU utilization threshold, the memory utilization is less than the third memory utilization threshold, and the backend disk utilization is less than the backend disk utilization threshold, the hash algorithm is determined to be the first hash algorithm. When the CPU utilization is not less than the second CPU utilization threshold, and / or the memory utilization is not less than the third memory utilization threshold, and / or the backend disk utilization is not less than the backend disk utilization threshold, the hash algorithm is determined to be the second hash algorithm.

[0078] It should also be noted that since the collision rate of xxHash64 is (1 / 2) 32 When the data volume is large, deduplication errors are prone to occur, causing data inconsistency. Therefore, xxHash3 is used for secondary verification, and the actual processing performance is still higher than SHA-256. It should be noted that regardless of the hash algorithm used, a Bloom filter must be used before calculating the hash value to filter out data blocks that are not in the hash fingerprint table.

[0079] S129: When the memory utilization is less than the first memory utilization threshold, add a hot data index.

[0080] S130: When the memory utilization is not less than a third memory utilization threshold, reduce the hot data index.

[0081] Finally, the hot and cold data indexes are only related to the memory status. The cold data index is the hash fingerprint value stored on disk, and the hot data index is the hash fingerprint value stored in the cache. Therefore, when the memory utilization rate is less than the first memory utilization threshold, the hot data index is increased. When the memory utilization rate is not less than the third memory utilization threshold, the hot data index is decreased.

[0082] It is worth noting that the first CPU utilization threshold is less than the second CPU utilization threshold, and the second CPU utilization threshold is less than the third CPU utilization threshold; the first memory utilization threshold is less than the second memory utilization threshold, and the second memory utilization threshold is less than the third memory utilization threshold. In this embodiment, there is no restriction on the size of the first CPU utilization threshold, the second CPU utilization threshold, and the third CPU utilization threshold; there is no restriction on the size of the first memory utilization threshold, the second memory utilization threshold, and the third memory utilization threshold; there is no restriction on the size of the inter-node interconnection bandwidth utilization threshold and the backend disk utilization threshold. For example, the first CPU utilization threshold is set to 40%, the second CPU utilization threshold is set to 50%, and the third CPU utilization threshold is set to 80%; the first memory utilization threshold is set to 50%, the second memory utilization threshold is set to 60%, and the third memory utilization threshold is set to 90%; the inter-node interconnection bandwidth utilization threshold is set to 60%, and the backend disk utilization threshold is set to 90%.

[0083] In summary, this solution adaptively selects the optimal deduplication timing, deduplication range, deduplication block, hash algorithm, etc. based on the CPU, memory, node interconnection bandwidth, and backend disk utilization under different I / O load conditions. Without sacrificing storage performance, it ensures the deduplication rate, saves storage space, and optimizes storage performance.

[0084] Based on the above embodiment, in some embodiments, after performing data deduplication according to the deduplication policy and the input / output request, the method further includes:

[0085] S16: Monitor the hardware resource utilization according to the second preset period, and return to the step of obtaining the hardware resource utilization threshold corresponding to each hardware resource utilization.

[0086] The first preset period is shorter than the second preset period.

[0087] In practice, after data deduplication is executed based on the deduplication policy and input / output requests, adjustments to the deduplication policy inevitably lead to changes in hardware resource utilization. These changes are proactive, and the resource monitoring module will not immediately report them. Otherwise, the deduplication feature selection will be ineffective and hardware resource utilization will fluctuate.

[0088] Therefore, in order to avoid the above situation, in this embodiment, the hardware resource utilization is specifically monitored according to the second preset period, and the process returns to the step of obtaining the hardware resource utilization threshold corresponding to each hardware resource utilization. It should be noted that the first preset period is smaller than the second preset period. In this embodiment, there is no restriction on the second preset period. For example, when the first preset period is 1s, the second preset period can be set to 5min. That is, after executing the data deduplication interval of 5min, the hardware resources are re-monitored, and the process returns to the step of obtaining the hardware resource utilization threshold corresponding to each hardware resource utilization, and the data deduplication process is executed again to avoid causing fluctuations in the hardware resource utilization.

[0089] Based on the above embodiment, in some embodiments, after performing data deduplication according to the deduplication policy and the input / output request, the method further includes:

[0090] S17: Determine the deduplication rate of this data deduplication and system performance change information.

[0091] S18: Based on the deduplication rate and system performance change information, determine whether the current data deduplication meets the preset conditions; if not, proceed to step S19 today; if so, end.

[0092] S19: reconfigure the hardware resource utilization threshold corresponding to each hardware resource utilization, and return to the step of monitoring the hardware resource utilization according to the first preset period.

[0093] After deduplication is complete, you can evaluate the performance of the deduplication strategy to determine its effectiveness. Specifically, first determine the deduplication ratio and system performance change information for this deduplication operation. The deduplication ratio measures the ratio of the amount of duplicate data reduced by deduplication technology to the total amount of original data, directly reflecting the effectiveness of the deduplication strategy in eliminating redundant data. System performance change information assesses the impact of the deduplication operation on the overall performance of the storage system, including I / O throughput, latency, and CPU and memory utilization.

[0094] Furthermore, based on the deduplication rate and system performance change information, it is determined whether the data deduplication meets the preset conditions. It should be noted that the preset conditions are not restricted in this embodiment. For example, a higher deduplication rate indicates that the deduplication strategy is very effective and can significantly reduce the storage space occupied; a lower deduplication rate may mean that there is less duplicate data in the data set, or that the current deduplication strategy fails to fully identify and eliminate redundant data. Therefore, a deduplication rate threshold can be set as part of the preset conditions. At the same time, while ensuring a high deduplication rate, the impact on system performance should be as small as possible. An ideal deduplication strategy should reduce storage space usage while avoiding negative impacts on system performance as much as possible. In the case of significant performance degradation, it may be necessary to optimize the execution timing of the deduplication strategy or adjust the system resource configuration to balance space optimization and performance requirements. Therefore, a system performance change threshold (including but not limited to an I / O throughput threshold, a latency threshold, or a CPU and memory utilization threshold) can also be set as another part of the preset conditions.

[0095] Therefore, if it is determined that the current data deduplication does not meet the preset conditions, the hardware resource utilization thresholds corresponding to each hardware resource utilization are reconfigured, and the process returns to the step of monitoring hardware resource utilization according to the first preset period to perform the next deduplication. If it is confirmed that the current data deduplication meets the preset conditions, the process ends. This allows for the evaluation of the current data deduplication and accurately determines the effectiveness of the deduplication strategy.

[0096] In addition, to determine the effectiveness of the deduplication strategy, the amount of storage space saved can also be used to determine it. Specifically, storage space savings refers to the actual amount of storage space reduced by deduplication technology, which is a direct benefit brought by the deduplication strategy. Larger storage space savings indicate that the deduplication strategy can effectively optimize the utilization of storage resources and reduce storage costs, which is particularly important for large-scale data storage environments. Smaller storage space savings may require re-evaluation of the configuration of the deduplication strategy, such as adjusting the block size, hash algorithm, or deduplication range to improve the space optimization effect. Therefore, in the specific implementation, after completing data deduplication, it is also possible to determine whether the storage space savings meet a preset value. If so, it is confirmed that the data deduplication effect is good and there is no need to adjust the strategy; if not, it is considered that the data deduplication strategy needs to be adjusted. In this way, high-performance execution of data deduplication is guaranteed.

[0097] To help those skilled in the art better understand this solution, a specific implementation process of a data deduplication method is given below:

[0098] When the real-time resource monitoring module detects CPU utilization of 30%, memory utilization of 20%, inter-node interconnect bandwidth utilization of 30%, and backend disk utilization of 25%, the storage system is considered to be operating in a low-load state. Based on the deduplication feature selection function, the selected deduplication timing is online, the deduplication scope is global, the chunking strategy is variable, the hash algorithm is SHA-256, and the cold and hot data indexing method is hot data indexing.

[0099] After deduplication is completed based on the above deduplication policy, CPU utilization, memory utilization, and inter-node interconnection bandwidth utilization will inevitably increase. For example, if CPU utilization rises to 55%, memory utilization rises to 58%, and inter-node interconnection bandwidth utilization rises to 60%, these hardware resource utilizations have exceeded their corresponding thresholds. However, at this point, the real-time resource monitoring module will not report this to the multi-dimensional deduplication feature selection module. Otherwise, the deduplication feature selection will fail, and hardware resource utilization will fluctuate.

[0100] Based on this, after adjusting the multi-dimensional deduplication features, the real-time resource monitoring module will report hardware resource utilization again after 5 minutes. At this time, due to increased front-end I / O traffic, memory utilization reaches 65%, exceeding the corresponding second memory utilization threshold of 60%. The real-time resource monitoring module reports this to the multi-dimensional deduplication feature selection module. The multi-dimensional deduplication feature selection module detects that the memory change within the deduplication range has crossed the threshold and adjusts global deduplication to local deduplication. Continued increase in I / O traffic causes any one of CPU utilization, memory utilization, inter-node interconnect bandwidth utilization, or back-end disk utilization to reach 100%. If memory utilization reaches 100%, a comprehensive adjustment of the deduplication strategy will be triggered. The deduplication features are changed to background deduplication, local deduplication, fixed block, xxHash64+xxHash3, and cold data index, and the utilization of each hardware resource decreases. After 5 minutes, when any hardware resource utilization crosses the threshold, it is reported again to the multi-dimensional deduplication feature selection module for a new round of deduplication strategy adjustment.

[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0102] Figure 3 Schematic diagram of a data deduplication device provided by an embodiment of the present invention. The device is applied to a storage system; Figure 3 As shown, the device includes:

[0103] The monitoring module 10 is used to monitor the hardware resource utilization according to a first preset period when receiving an input / output request sent by the host side; wherein the hardware resource utilization includes at least the central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization and back-end disk utilization.

[0104] The acquisition module 11 is configured to acquire a hardware resource utilization threshold corresponding to each hardware resource utilization.

[0105] The deduplication module 12 is configured to determine a deduplication policy according to the utilization of each hardware resource and the corresponding hardware resource utilization threshold, and perform data deduplication according to the deduplication policy and the input / output request.

[0106] In some embodiments, it further includes:

[0107] A configuration submodule is used to configure the initial hardware resource utilization threshold corresponding to each hardware resource utilization; wherein the initial hardware resource utilization threshold includes at least an initial CPU utilization threshold, an initial memory utilization threshold, an initial inter-node interconnection bandwidth utilization threshold, and an initial backend disk utilization threshold;

[0108] The first determination submodule is used to determine the overall status and hardware type of the system; wherein the overall status includes at least a single-controller cluster, a multi-controller cluster, a common cluster, and an active-active cluster; the hardware type includes at least a central processing unit model and a backend disk type;

[0109] The adjustment module is used to adjust each initial hardware resource utilization threshold according to the overall state and hardware type to generate each hardware resource utilization threshold.

[0110] In some embodiments, the adjustment module includes:

[0111] The first adjustment submodule is used to increase the initial node interconnection bandwidth utilization threshold when the overall state is a single-control cluster;

[0112] The second adjustment submodule is configured to reduce the initial CPU utilization threshold when the overall state is an active-active cluster;

[0113] a third adjustment submodule, configured to increase the initial CPU utilization threshold when the CPU model is the target model;

[0114] The fourth adjustment submodule is configured to reduce the initial backend disk utilization threshold when the backend disk type is a mechanical hard disk.

[0115] In some embodiments, the deduplication module 12 includes:

[0116] The deduplication strategy determination module is used to determine the deduplication timing, deduplication range, block strategy, hash algorithm, and hot and cold data indexes based on the utilization of each hardware resource and the corresponding hardware resource utilization threshold;

[0117] Among them, the deduplication timing includes online deduplication and background deduplication; the deduplication scope includes global deduplication and local deduplication; the blocking strategy includes fixed blocking and variable blocking; the hash algorithm includes a first hash algorithm and a second hash algorithm, and the collision rate and calculation speed of the first hash algorithm are respectively lower than the collision rate and calculation speed of the second hash algorithm.

[0118] In some embodiments, the deduplication strategy determination module includes:

[0119] a second determining submodule, configured to determine that the deduplication opportunity is online deduplication when the CPU utilization is less than a first CPU utilization threshold and the memory utilization is less than a first memory utilization threshold;

[0120] a third determining submodule, configured to determine that the deduplication opportunity is background deduplication when the CPU utilization is not less than a first CPU utilization threshold and / or the memory utilization is not less than a first memory utilization threshold;

[0121] a fourth determination submodule, configured to determine that the deduplication range is global deduplication when the CPU utilization is less than a third CPU utilization threshold, the memory utilization is less than a first memory utilization threshold, the inter-node interconnection bandwidth utilization is less than the inter-node interconnection bandwidth utilization threshold, and the backend disk utilization is less than the backend disk utilization threshold;

[0122] a fifth determining submodule, configured to determine that the deduplication range is local deduplication when the CPU utilization is not less than a third CPU utilization threshold, and / or the memory utilization is not less than a first memory utilization threshold, and / or the inter-node interconnection bandwidth utilization is not less than the inter-node interconnection bandwidth utilization threshold, and / or the backend disk utilization is not less than the backend disk utilization threshold;

[0123] a sixth determining submodule, configured to determine that the partitioning strategy is variable partitioning when the CPU utilization is less than the second CPU utilization threshold and the memory utilization is less than the second memory utilization threshold;

[0124] a seventh determining submodule, configured to determine that the partitioning strategy is fixed partitioning when the CPU utilization is not less than a second CPU utilization threshold and / or the memory utilization is not less than a second memory utilization threshold;

[0125] an eighth determining submodule, configured to determine that the hash algorithm is the first hash algorithm when the CPU utilization is less than the second CPU utilization threshold, the memory utilization is less than the third memory utilization threshold, and the backend disk utilization is less than the backend disk utilization threshold;

[0126] a ninth determination submodule, configured to determine that the hash algorithm is the second hash algorithm when the CPU utilization is not less than the second CPU utilization threshold, and / or the memory utilization is not less than the third memory utilization threshold, and / or the backend disk utilization is not less than the backend disk utilization threshold;

[0127] a tenth determining submodule, configured to add a hot data index when the memory utilization is less than a first memory utilization threshold;

[0128] an eleventh determining submodule, configured to reduce a hot data index when the memory utilization is not less than a third memory utilization threshold;

[0129] Among them, the first CPU utilization threshold is less than the second CPU utilization threshold, and the second CPU utilization threshold is less than the third CPU utilization threshold; the first memory utilization threshold is less than the second memory utilization threshold, and the second memory utilization threshold is less than the third memory utilization threshold.

[0130] In some embodiments, it further includes:

[0131] A first monitoring submodule is configured to monitor the hardware resource utilization according to a second preset period and return to the step of obtaining a hardware resource utilization threshold corresponding to each hardware resource utilization;

[0132] The first preset period is shorter than the second preset period.

[0133] In some embodiments, it further includes:

[0134] An information determination module is used to determine the deduplication rate and system performance change information of this data deduplication;

[0135] The judgment module is used to judge whether the data deduplication meets the preset conditions based on the deduplication rate and system performance change information; if not, reconfigure the hardware resource utilization threshold corresponding to each hardware resource utilization, and return to the step of monitoring the hardware resource utilization according to the first preset period; if so, end.

[0136] For the description of the features in the embodiment corresponding to the data deduplication device, reference can be made to the relevant description of the embodiment corresponding to the data deduplication method, which will not be repeated here.

[0137] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data deduplication method embodiments.

[0138] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned data deduplication method embodiments when running.

[0139] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0140] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned data deduplication method embodiments are implemented.

[0141] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned data deduplication method embodiments.

[0142] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0143] The above is a detailed introduction to the data deduplication method, device, equipment, and medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core ideas of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A data deduplication method, characterized in that: Applied to a storage system; the method comprises: When receiving an input / output request from the host side, monitoring hardware resource utilization according to a first preset period; wherein the hardware resource utilization includes at least central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization and backend disk utilization; Obtaining a hardware resource utilization threshold corresponding to each of the hardware resource utilizations; A deduplication strategy is determined according to each of the hardware resource utilizations and the corresponding hardware resource utilization thresholds, and data deduplication is performed according to the deduplication strategy and the input / output request.

2. The data deduplication method according to claim 1, wherein: Before monitoring the hardware resource utilization in the system according to the first preset period, the method further includes: Configuring initial hardware resource utilization thresholds corresponding to the hardware resource utilizations; wherein the initial hardware resource utilization thresholds include at least an initial CPU utilization threshold, an initial memory utilization threshold, an initial inter-node interconnection bandwidth utilization threshold, and an initial backend disk utilization threshold; Determine the overall status and hardware type of the system; wherein the overall status includes at least a single-controller cluster, a multi-controller cluster, a common cluster, and an active-active cluster; the hardware type includes at least the CPU model and the backend disk type; The initial hardware resource utilization thresholds are adjusted according to the overall state and the hardware type to generate the hardware resource utilization thresholds.

3. The data deduplication method according to claim 2, wherein: Adjusting each of the initial hardware resource utilization thresholds according to the overall state and the hardware type includes: When the overall state is the single-control cluster, increasing the initial node-to-node interconnection bandwidth utilization threshold; When the overall state is the active-active cluster, reducing the initial CPU utilization threshold; When the CPU model is the target model, increasing the initial CPU utilization threshold; When the backend disk type is a mechanical hard disk, the initial backend disk utilization threshold is reduced.

4. The data deduplication method according to claim 1, wherein: Determining a deduplication strategy according to each of the hardware resource utilizations and the corresponding hardware resource utilization thresholds includes: Determine the deduplication timing, deduplication range, block strategy, hash algorithm, and hot and cold data indexes according to each of the hardware resource utilization rates and the corresponding hardware resource utilization thresholds; Among them, the deduplication timing includes online deduplication and background deduplication; the deduplication range includes global deduplication and local deduplication; the blocking strategy includes fixed blocking and variable blocking; the hash algorithm includes a first hash algorithm and a second hash algorithm, and the collision rate and calculation speed of the first hash algorithm are respectively lower than the collision rate and calculation speed of the second hash algorithm.

5. The data deduplication method according to claim 4, wherein: Determining the deduplication timing, deduplication range, block strategy, hash algorithm, and hot and cold data indexes based on the hardware resource utilization rates and the corresponding hardware resource utilization thresholds, including: When the CPU utilization is less than a first CPU utilization threshold, and the memory utilization is less than a first memory utilization threshold, determining that the deduplication opportunity is the online deduplication; When the CPU utilization is not less than a first CPU utilization threshold, and / or the memory utilization is not less than a first memory utilization threshold, determining that the deduplication opportunity is the background deduplication; When the CPU utilization is less than a third CPU utilization threshold, the memory utilization is less than a first memory utilization threshold, the inter-node interconnection bandwidth utilization is less than an inter-node interconnection bandwidth utilization threshold, and the backend disk utilization is less than a backend disk utilization threshold, determining that the deduplication range is the global deduplication; When the CPU utilization is not less than a third CPU utilization threshold, and / or the memory utilization is not less than a first memory utilization threshold, and / or the inter-node interconnection bandwidth utilization is not less than an inter-node interconnection bandwidth utilization threshold, and / or the backend disk utilization is not less than a backend disk utilization threshold, determining that the deduplication range is the local deduplication; When the CPU utilization is less than a second CPU utilization threshold, and the memory utilization is less than a second memory utilization threshold, determining that the block strategy is the variable block strategy; When the CPU utilization is not less than a second CPU utilization threshold, and / or the memory utilization is not less than a second memory utilization threshold, determining that the block strategy is the fixed block strategy; When the CPU utilization is less than a second CPU utilization threshold, the memory utilization is less than a third memory utilization threshold, and the backend disk utilization is less than a backend disk utilization threshold, determining that the hash algorithm is the first hash algorithm; When the CPU utilization is not less than a second CPU utilization threshold, and / or the memory utilization is not less than a third memory utilization threshold, and / or the backend disk utilization is not less than a backend disk utilization threshold, determining that the hash algorithm is the second hash algorithm; When the memory utilization is less than a first memory utilization threshold, adding a hot data index; When the memory utilization is not less than a third memory utilization threshold, reducing the hot data index; Among them, the first CPU utilization threshold is less than the second CPU utilization threshold, and the second CPU utilization threshold is less than the third CPU utilization threshold; the first memory utilization threshold is less than the second memory utilization threshold, and the second memory utilization threshold is less than the third memory utilization threshold.

6. The data deduplication method according to any one of claims 1 to 5, characterized in that: After performing data deduplication according to the deduplication policy and the input / output request, the method further includes: Monitoring the hardware resource utilization according to a second preset period, and returning to the step of obtaining the hardware resource utilization threshold corresponding to each hardware resource utilization; The first preset period is smaller than the second preset period.

7. The data deduplication method according to claim 6, wherein: After performing data deduplication according to the deduplication policy and the input / output request, the method further includes: Determine the deduplication rate and system performance change information for this data deduplication; Determining whether the current data deduplication meets a preset condition based on the deduplication rate and the system performance change information; If not, reconfigure the hardware resource utilization threshold corresponding to each of the hardware resource utilizations, and return to the step of monitoring the hardware resource utilization according to the first preset period; If so, end.

8. A data deduplication device, characterized in that: Applicable to a storage system; the device comprises: A monitoring module, configured to monitor hardware resource utilization according to a first preset period upon receiving an input / output request from the host side; wherein the hardware resource utilization includes at least central processing unit utilization, memory utilization, inter-node interconnection bandwidth utilization, and backend disk utilization; An acquisition module, configured to acquire a hardware resource utilization threshold corresponding to each of the hardware resource utilizations; The deduplication module is configured to determine a deduplication strategy according to each of the hardware resource utilizations and the corresponding hardware resource utilization thresholds, and perform data deduplication according to the deduplication strategy and the input / output request.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data deduplication method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data deduplication method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data deduplication method, device and equipment and readable medium

    CN114138198A

  • Data processing method and device

    CN115599591A

  • Redistribution of processing groups between server nodes based on hardware resource utilization

    US20220222113A1

  • Data reduction method and related system

    WO2025055390A1