Cloud hard disk backup and recovery system based on storage snapshot
By introducing reference strength analysis module and data block replication and drainage module in the cloud hard disk backup and recovery system, the reference strength of data blocks is monitored and regulated in real time, and the problem of performance degradation in the backup process in the prior art is solved, and the use of CPU, IO and network bandwidth is reduced while ensuring backup efficiency.
Patent Information
- Application Number
- CN202510655279.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing cloud hard disk backup solution needs to occupy a large amount of CPU, IO and network bandwidth during the backup process, resulting in a degradation of performance.
By introducing reference strength analysis module and data block replication and drainage module in the cloud hard disk backup and recovery system, the reference strength of data blocks can be monitored and regulated in real time, reducing the consumption of CPU, IO and network bandwidth.
By quantifying the reference intensity of data blocks, we target the data blocks of high frequency and high complexity reference, reducing the multi-level backtracking pressure of the original data block, reducing CPU metadata analysis and IO operation consumption, and ensuring stable cloud host performance during backup.
Smart Images

Figure CN120179467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and particularly to a cloud hard disk backup and recovery system based on storage snapshots. Background Art
[0002] Existing cloud hard disk backup solutions usually rely on the storage system to generate complete cloud hard disk snapshots, and the backup software reads the snapshots in full on the cloud host, calculates the hash value, and calculates the differential data by comparing with the previous records, and then transfers the differential data to the backup storage, resulting in a large amount of CPU, IO, and network bandwidth being occupied during the backup process, thereby causing a decline in performance during backup. Summary of the Invention
[0003] Technical Problems to be Solved In view of the deficiencies of the prior art, the present invention provides a cloud hard disk backup and recovery system based on storage snapshots, which solves the problem of performance decline during the cloud hard disk backup process.
[0004] Technical Solutions To achieve the above object, the present invention is realized through the following technical solutions: A cloud hard disk backup and recovery system based on storage snapshots, including the following specific modules: Cloud platform: used to formulate rules for snapshot generation, data backup, and data recovery; Main storage module: stores data to obtain data blocks, generates snapshots through the cloud platform, and backs up the data blocks before being tampered with to the secondary storage module, and the snapshots are associated with the data blocks; Reference intensity analysis module: obtains the reference count and reference depth of the snapshots in real time, calculates the reference frequency according to the time series for the reference count, and performs normalization processing and comprehensive calculation based on the number of data blocks, reference count, reference depth, and reference frequency to obtain the reference intensity; Data block replication drainage module: used to monitor whether the reference intensity causes performance decline of CPU, IO, and network bandwidth and take measures; Secondary storage module: under the condition that the CPU, IO, and network bandwidth are normal, used to receive the backed-up data blocks or restore the backed-up data blocks to the main storage module through the cloud platform.
[0005] Further, the specific method for obtaining the reference intensity is as follows: ; where represents the reference intensity, represents the number of data blocks, represents the reference breadth, represents the th reference frequency of the th data block for the th object,
[0006] Further, the specific steps for monitoring whether the reference intensity causes a decline in CPU, IO, and network bandwidth performance and taking measures are as follows: Set a reference threshold, and compare the reference intensity with the reference threshold in real time. If the reference intensity is less than the reference threshold, continue the monitoring. If the reference intensity is greater than or equal to the reference threshold, perform regulation through the data drainage algorithm until it is monitored that the CPU, IO, and network bandwidth return to normal.
[0007] Further, the specific steps for performing regulation through the data drainage algorithm are as follows: Assign weights to all referenced data blocks according to the reference intensity to obtain the weight of each referenced data block, denoted as the first reference weight. Sort the first reference weights in ascending order to obtain the largest first reference weight. For the data block corresponding to the largest first reference weight, denoted as A, and make a copy to obtain a copied data block, denoted as B. Back up B to the secondary storage module. Then, assign weights to the objects that reference A according to the reference frequency and reference depth to obtain the second reference weight. Sort the second reference weights in ascending order, and associate the metadata of the object corresponding to the largest second reference weight with B. For other objects corresponding to the second reference weight, and so on. Continue to compare the reference intensity with the reference threshold in real time. If the existence of A and B cannot make the reference intensity less than the reference threshold, continue to copy B, and so on.
[0008] Further, the specific method for obtaining the second reference weight is as follows: ; where represents the second reference weight, represents the reference frequency, represents the reference depth.
[0009] Further, the specific method for backing up the data block before being tampered with to the secondary storage module is as follows: Store this data block in the main storage module at a new physical address through the COW technology, and mark and associate the new logical address with this data block, denoted as the new data block. And back up the data before being tampered with to the secondary storage module, and associate it with a snapshot.
[0010] Further, the specific method for setting the reference threshold is as follows: Obtain the historical normal data of the CPU, IO, and network bandwidth, and perform comprehensive calculation on the historical normal data of the CPU, IO, and network bandwidth to obtain the reference threshold.
[0011] Further, the specific steps for performing comprehensive calculation on the historical normal data of the CPU, IO, and network bandwidth are as follows: Perform average calculation on the historical normal data of the CPU, IO, and network bandwidth to obtain normal equilibrium data, and then perform variance method calculation according to the historical normal data and the normal equilibrium data of the CPU, IO, and network bandwidth.
[0012] Further, the specific method for waiting until the CPU, IO, and network bandwidth return to normal is as follows: If the reference intensity is less than the reference threshold, the cloud platform sequentially deletes according to the number of replicated data blocks in the secondary storage module until the CPU, IO, and network bandwidth return to normal.
[0013] Further, the specific steps for the cloud platform to sequentially delete according to the number of replicated data blocks in the secondary storage module are as follows: Each time, one replicated data block is deleted, and then the reference intensity is continuously compared with the reference threshold in real time. If the reference intensity is greater than or equal to the reference threshold after deleting this replicated data block, this replicated data block is restored until the reference intensity is less than the reference threshold. If the reference intensity is still less than the reference threshold after deleting this replicated data block, then another replicated data block is deleted, and so on, until all the replicated data blocks in the secondary storage module are deleted.
[0014] Beneficial Effects Compared with the prior art, the embodiments of the present invention at least have the following advantages or beneficial effects: 1. By quantifying the reference intensity of data blocks, specifically replicating data blocks with high-frequency and high-complexity references, diverting the reference paths of high-frequency accessed objects to the replicated data blocks, reducing the multi-level backtracking pressure of the original data blocks, thereby reducing the CPU metadata parsing and IO operation consumption, and ensuring the stable performance of the cloud host during backup.
[0015] 2. Dynamically setting the reference threshold based on historical resource data, diverting the pressure of replicated data blocks during overload, gradually deleting the replicated data blocks after the load decreases and monitoring in real time, while improving the system response ability, avoiding redundant storage, and balancing the resource utilization rate and storage cost.
[0016] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. Brief Description of the Drawings
[0017] Figure 1 This is for the present invention: A structural diagram of a cloud hard disk backup and recovery system based on storage snapshots.
[0018] Figure 2 This is for the present invention: A line graph showing the influence of reference intensity on CPU resource utilization rate. Detailed Embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0021] As Figure 1 shown, the embodiments of the present invention provide a cloud disk backup and recovery system based on storage snapshots, including the following specific modules: Cloud platform: used to formulate rules for snapshot generation, data backup and data recovery, automatically generate snapshots according to the set time. Snapshots are used to record the state of data blocks at a certain moment, and mark and associate the location and state of data blocks through metadata. Metadata includes the logical address and physical address of data blocks, the attributes of data blocks, the association relationship of data blocks, the status information of data blocks, etc.
[0022] Main storage module: The cloud platform stores data through the main storage module to obtain data blocks. The data blocks are stored in physical addresses and associated through logical addresses. The data blocks are shared by the main storage module and snapshots. The reference count of the COW technology is used to count the number of shared objects of this data block. If the data block is tampered with, the main storage module stores this data block in a new physical address and updates the metadata to mark and associate the new logical address with this data block, denoted as a new data block. The data block before being tampered with is no longer shared by the main storage module, but only shared by the associated snapshot, that is, the snapshot associates the metadata of this data block before it is tampered with, denoted as the original data block. And the cloud platform backs up the original data block to the secondary storage module. When the subsequent new data block is tampered with, and so on, to avoid the backup software from fully reading the snapshot on the cloud platform, calculating the hash value, and calculating the differential data by comparing with the original data block, and then transmitting the differential data to the backup storage, preventing a large amount of CPU, IO, and network bandwidth from being occupied during the backup process, resulting in a decline in backup performance.
[0023] Reference intensity analysis module: directly obtain the reference count of the snapshot through the storage monitoring API or metadata interface of the cloud platform to get the reference count of the data block, which is recorded as the reference breadth. Automatically construct a hierarchical tree by parsing the dependency relationship through the snapshot management API of the cloud platform to obtain the reference data block hierarchy, that is, the object of the reference data block can reference the data block only by traversing multiple levels, which is recorded as the reference depth. The cloud platform tracks the objects that reference this data block through metadata and calculates the quotient based on the reference count of this object during the system running time to obtain the reference frequency of this object referencing this data block, which is recorded as the reference frequency. Since there are n data blocks in the main storage module and all these data blocks can be referenced, standardize the reference breadth, reference depth, reference frequency, and the number of data blocks to eliminate the dimension, and perform comprehensive calculation to obtain the reference intensity, which reflects the intensity of the snapshot associated with the original data block or new data block being referenced. If the intensity is greater, more CPU, IO, and network bandwidth will be consumed.
[0024] The specific method for obtaining the reference intensity is as follows: ; Among them, represents the reference intensity, and it reflects whether a large amount of CPU or IO and network bandwidth are consumed through the reference intensity. represents the number of data blocks. represents the reference breadth. represents the th reference frequency of the th object of the th data block.
[0025] Table 1 Influence table of a reference intensity on CPU resource utilization
[0026] As shown in Table 1, in Group 1, when the reference intensity is 93, the CPU resource utilization rate is 54%. In Group 2, when the reference intensity is 126, the CPU resource utilization rate is 61%. In Group 3, when the reference intensity is 322, the CPU resource utilization rate is 74%. As Figure 2 shown, this indicates that when the complexity of the data block reference relationship processed by the system increases, the CPU resource utilization rate shows a non-linear growth characteristic with the increase of the reference intensity.
[0027] Data block replication and drainage module: It is used to monitor whether the reference intensity causes a decrease in CPU, IO, and network bandwidth performance and take measures, set a reference threshold, and compare the reference intensity with the reference threshold in real time. If the reference intensity is less than the reference threshold, continue to monitor. If the reference intensity is greater than or equal to the reference threshold, it means that more CPU, IO, and network bandwidth are occupied. Therefore, it is regulated through the data drainage algorithm until it is monitored that the CPU, IO, and network bandwidth return to normal. Only when the CPU, IO, and network bandwidth return to normal can the data block be normally backed up or restored.
[0028] The specific steps for regulation through the data drainage algorithm are as follows: Assign weights to all referenced data blocks according to the reference intensity to obtain the weight of each referenced data block, denoted as the first reference weight; Sort the first reference weights in ascending order through bubble sort to obtain the largest first reference weight; For the data block corresponding to the largest first reference weight, denoted as A, make a copy to obtain a copied data block, denoted as B, and back it up and store it in the secondary storage module to avoid occupying the storage space of the main storage module; Then assign weights to the objects that reference A according to the reference frequency and reference depth to obtain the second reference weight; Sort the second reference weights in ascending order through quick sort, and associate the metadata of the object corresponding to the largest second reference weight with B. Thus, the association with A is reduced and the accuracy of the referenced data block remains unchanged. Since the second reference weight of this object when originally referencing A was the largest, after switching to referencing B, the computational complexity of the reference frequency and reference depth when this object references A is reduced. The same applies to other objects corresponding to the second reference weight. The existence of B also reduces the reference breadth and reference frequency of A, further reducing the reference breadth and reference frequency when the object references A; Continue to compare the reference intensity with the reference threshold in real time. If the existence of A and B cannot make the reference intensity less than the reference threshold, continue to copy B, and so on; The same applies to the data blocks corresponding to other first reference weights except A.
[0029] The specific method for obtaining the second reference weight is as follows: ; Among them, represents the second reference weight, represents the reference frequency, represents the reference depth.
[0030] The specific method for obtaining the set reference threshold is as follows: Obtain the historical normal data of the CPU, IO, and network bandwidth, and perform comprehensive calculations on the historical normal data of the CPU, IO, and network bandwidth to obtain the reference threshold.
[0031] The specific steps for calculating the historical normal data of the CPU, IO, and network bandwidth by the variance method are as follows: Perform an average calculation on the historical normal data of the CPU, IO, and network bandwidth to obtain the normal equilibrium data, which is used as the standard for measuring the historical normal data of the CPU, IO, and network bandwidth. Perform variance method calculations based on the historical normal data of the CPU, IO, and network bandwidth and the normal equilibrium data, and reflect the range of the reference threshold through the volatility of the historical normal data of the CPU, IO, and network bandwidth on the normal equilibrium data.
[0032] The specific steps until it is detected that the CPU, IO, and network bandwidth return to normal are as follows: If the reference intensity is less than the reference threshold, the cloud platform deletes the replicated data blocks in sequence according to the number of replicated data blocks in the secondary storage module, deleting one replicated data block each time, and then continues to perform real-time comparison between the reference intensity and the reference threshold. If the reference intensity is greater than or equal to the reference threshold after deleting this replicated data block, restore this replicated data block until the reference intensity is less than the reference threshold. If the reference intensity is still less than the reference threshold after deleting this replicated data block, continue to delete one replicated data block, and so on, until all the replicated data blocks are deleted in the secondary storage module to save storage space. When the reference intensity is still less than the reference threshold, it indicates that the CPU, IO, and network bandwidth have returned to normal.
[0033] Secondary storage module: Under the condition that the CPU, IO, and network bandwidth are normal, it is used to receive the original data blocks backed up by the cloud platform, and the cloud platform also references the original data blocks through the secondary storage module and restores them to the main storage module, because only when the CPU, IO, and network bandwidth are normal can the efficiency and integrity of data block backup or restoration be guaranteed.
[0034] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principle and practical application of the present invention, so that those skilled in the relevant technical field can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A cloud hard disk backup and recovery system based on storage snapshots, characterized in that: Includes the following specific modules: Cloud platform: used to formulate rules for snapshot generation, data backup, and data recovery; Primary storage module: stores data to obtain data blocks, generates snapshots through the cloud platform and backs up the data blocks before tampering to the secondary storage module, and the snapshots are associated with the data blocks; Citation intensity analysis module: obtains the number of citations and citation depth of snapshots in real time, calculates the number of citations according to the time series, obtains the citation frequency, and performs standardization and comprehensive calculation based on the number of data blocks, number of citations, citation depth and citation frequency to obtain the citation intensity; Data block replication and drainage module: used to monitor whether the reference intensity causes CPU, IO and network bandwidth performance to decline and take measures; Secondary storage module: When the CPU, IO and network bandwidth are normal, it is used to receive the backed-up data blocks or restore the backed-up data blocks to the primary storage module through the cloud platform.
2. The cloud hard disk backup and recovery system based on storage snapshot according to claim 1, characterized in that: The specific method for obtaining the citation strength is as follows: ; in, Indicates the strength of the citation. Indicates the number of data blocks, Indicates the breadth of citations, Indicates The first data block The frequency of references to an object, Indicates The first data block The reference depth of an object.
3. The cloud hard disk backup and recovery system based on storage snapshot according to claim 1, characterized in that: The specific steps for monitoring whether the reference intensity causes CPU, IO and network bandwidth performance degradation and taking measures are as follows: Set a reference threshold and compare the reference strength with the reference threshold in real time. If the reference strength is less than the reference threshold, continue monitoring. If the reference strength is greater than or equal to the reference threshold, adjust and control it through the data diversion algorithm until the CPU, IO, and network bandwidth are monitored to return to normal.
4. The cloud hard disk backup and recovery system based on storage snapshot according to claim 3 is characterized in that: The specific steps of regulating through the data diversion algorithm are as follows: All referenced data blocks are weighted according to the reference strength to obtain the weight of each referenced data block, which is recorded as the first reference weight. The first reference weights are sorted in ascending order to obtain the largest first reference weight. The data block corresponding to the largest first reference weight is recorded as A and copied to obtain a copied data block, which is recorded as B. The backup of B is stored in the secondary storage module, and then the objects referenced by A are weighted according to the reference frequency and reference depth to obtain the second reference weight. The second reference weights are sorted in ascending order, and the metadata of the object corresponding to the largest second reference weight is associated with B. The same is true for other objects corresponding to the second reference weight. The reference strength and the reference threshold are continued to be compared in real time. If the existence of A and B cannot make the reference strength less than the reference threshold, B continues to be copied, and so on.
5. The cloud hard disk backup and recovery system based on storage snapshot according to claim 4 is characterized in that: The specific method for obtaining the second reference weight is as follows: ; in, represents the second citation weight, Indicates the frequency of citations, Indicates the reference depth.
6. The cloud hard disk backup and recovery system based on storage snapshot according to claim 1, characterized in that: The specific method of backing up the data blocks before tampering to the secondary storage module is as follows: The data block in the primary storage module is stored at a new physical address through the COW technology, and the new logical address is marked and associated with the data block, recorded as a new data block, and the data before being tampered is backed up to the secondary storage module and associated with it through a snapshot.
7. The cloud hard disk backup and recovery system based on storage snapshot according to claim 3 is characterized in that: The specific method of setting the reference threshold is as follows: The historical normal data of CPU, IO and network bandwidth are obtained, and the historical normal data of CPU, IO and network bandwidth are comprehensively calculated to obtain the reference threshold.
8. The cloud hard disk backup and recovery system based on storage snapshot according to claim 7, characterized in that: The specific steps of comprehensively calculating the historical normal data of CPU, IO and network bandwidth are as follows: The historical normal data of CPU, IO and network bandwidth are averaged to obtain normal balanced data, and then the variance method is used to calculate the historical normal data of CPU, IO and network bandwidth and the normal balanced data.
9. The cloud hard disk backup and recovery system based on storage snapshot according to claim 3, characterized in that: The specific method until the CPU, IO and network bandwidth are monitored to return to normal is as follows: If the reference strength is less than the reference threshold, the cloud platform deletes the duplicate data blocks in sequence according to the number of duplicate data blocks in the secondary storage module until the CPU, IO, and network bandwidth return to normal.
10. The cloud hard disk backup and recovery system based on storage snapshot according to claim 9, characterized in that: The specific steps of deleting the duplicate data blocks in sequence according to the number of duplicate data blocks in the secondary storage module are as follows: Each time a replicated data block is deleted, the reference strength is compared with the reference threshold in real time. If the reference strength is greater than or equal to the reference threshold after deleting the replicated data block, the replicated data block is restored until the reference strength is less than the reference threshold. If the reference strength is still less than the reference threshold after deleting the replicated data block, a replicated data block is deleted, and so on, until all replicated data blocks are deleted in the secondary storage module.
Citation Information
Patent Citations
Cache data processing method and system and readable storage medium
CN110147331A
Space management method, device and system of cloud storage space, electronic equipment and computer readable storage medium
CN112764663A
Data backup method and device, equipment, storage medium and program product
CN119645728A
Range-based deletion of snapshots archived in cloud / object storage
US10534750B1
Batch-based deletion of snapshots archived in cloud / object storage
US20200019620A1