Information resource multi-level data recovery method and system and medium

By adopting a multi-level data recovery method in multi-snapshot, multi-copy and distributed storage environments, using Merkel tree and blockchain hash links, the problem of difficulty in achieving fast and efficient recovery in traditional technologies is solved, and efficient and reliable data recovery and integrity verification are achieved.

CN120216260APending Publication Date: 2025-06-27AVIC STAR BEIDOU CHONGQING TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510303310.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional snapshot technology is difficult to achieve fast and holistic verification in multi-snapshot, multi-copy and distributed storage environments, and lacks an efficient collaborative recovery process, making it difficult to safely compare and pull data between damaged nodes and healthy nodes.

Method used

Using a multi-level data recovery method, by constructing snapshot metadata block index, Merkel tree data structure and blockchain hash link, trusted chains and data integrity verification between snapshots, and voting at the block level to obtain the correct data.

Benefits of technology

It significantly improves the data recovery effect in multi-snapshot and multi-copy environments, reduces the number of invalid reads and repeated verifications, can record data evolution throughout the process, timely detect and locate tampering or abnormalities, and meets the needs of auditing and evidence collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216260A_ABST
    Figure CN120216260A_ABST
Patent Text Reader

Abstract

The invention provides an information resource multi-level data recovery method and system and a medium, and the method comprises the steps: generating a snapshot header for each snapshot, which comprises a snapshot ID, a previous snapshot hash value, a current snapshot hash value and an optional digital signature; meanwhile, hash calculation is carried out on all the blocks in the snapshot, and a Merkel tree is constructed to obtain a global hash value; when a new or updated block is detected, carrying out Hash calculation on the new or updated block, adding the new or updated block into a Merkel tree, generating a new snapshot Hash value, and writing the new snapshot Hash value into a snapshot head; and taking the hash value of the previous snapshot as a predecessor reference of the new snapshot to form chained association. In a multi-copy environment, when data recovery needs to be carried out, firstly, related snapshots are screened and determined through a snapshot index and a filter; then verifying the snapshot chain to ensure that the front and back references of each snapshot are correctly matched with the Merkel tree hash value; for block-level data, if copies are inconsistent, a correct version is discriminated through weighted consensus; and finally reconstructing the damaged block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data backup and recovery, and particularly to a multi-level data recovery method for information resources. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data, and the Internet of Things, distributed storage systems are increasingly undertaking the tasks of storing and managing massive amounts of data. To meet the requirements of business continuity and data security, it is usually necessary to perform periodic snapshots and long-term archiving of the data in the system. Although traditional snapshot technologies can quickly save the data state at a certain moment, they still have obvious deficiencies in the following aspects:

[0003] Most traditional snapshot schemes only rely on single-point or centralized verification means, and cannot perform fast and holistic verification on historical versions. If data corruption caused by malicious tampering or hardware failures occurs, it is difficult to detect or trace back in a timely manner.

[0004] Conventional data backups and snapshots mainly target individual storage nodes or clusters. When the system expands to a multi-copy deployment across clouds or regions, there is a lack of an efficient collaborative recovery process, and it is difficult to safely compare and pull data between damaged nodes and healthy nodes.

[0005] In a large-scale distributed environment, in the face of frequent write operations and version iterations, how to quickly verify the overall integrity of any historical version and ensure the traceability of data evolution across multiple snapshots is a difficult problem that the industry urgently needs to solve. Summary of the Invention

[0006] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a multi-level data recovery method for information resources, including:

[0007] S1, Data Structure and Hash Link: Construct a snapshot metadata block index / list, a Merkle tree data structure, and form a trusted chain by concatenating snapshots;

[0008] S2, Snapshot Creation and Linking Process: Through data flushing, block hashing, constructing a Merkle tree, writing a snapshot header, and storing snapshot information in a blockchain or distributed ledger;

[0009] S3, Data Recovery and Verification: Identify the required data and verify the integrity of the snapshot chain, vote at the block level to obtain the correct data, and finally perform end-to-end integrity checking and complete the recovery.

[0010] In a preferred embodiment of the present invention, the data structure and hash link include:

[0011] In the snapshot metadata, each snapshot S i, including:

[0012] Snapshot header;

[0013] snapshotID: used to uniquely identify the snapshot, using a globally incrementing serial number;

[0014] prevRootHash: stored in snapshot S i+1 and pointing to the hash value thisRootHash of the previous snapshot S i+1 of S i to establish a snapshot chain;

[0015] thisRootHash: the hash value of the current snapshot, calculated based on the hashes of all data blocks;

[0016] Each snapshot maintains a table or database record corresponding to the set of data blocks under this snapshot:

[0017] blockID: uniquely identifies the location of a data block within a logical volume or file system;

[0018] blockHash: stores the encrypted hash value of the block, used for integrity verification and cross-version comparison;

[0019] size: block size information;

[0020] replicaLocations: indicates how many replicas or storage nodes this block exists on, facilitating the location of available copies in case of failures;

[0021] Build a Merkle tree using all blockHashes in the snapshot, and its hash value is thisRootHash;

[0022] Given a data block B i , its leaf node hash can be expressed as:

[0023] H(B i );

[0024] where H(B i ) is the corresponding hash value of the i-th data block;

[0025] H(*) is a hash function used to ensure that if the data block B i is tampered with, H i will also change accordingly, thus detecting data corruption or tampering at the leaf level;

[0026] The link logic between snapshots is:

[0027] S i+1 →S i , that is, each new snapshot Si+1 Record its previous snapshot, i.e., S i The hash value of → represents the link relationship between the previous and subsequent snapshots;

[0028] It is reflected in the metadata as:

[0029] prevRootHash i+1 = thisRootHash i ;

[0030] Among them, thisRootHash i is the hash value calculated based on all the blocks within snapshot S i ;

[0031] prevRootHash i+1 is the hash value of the previous snapshot S i+1 stored in snapshot S i , and a continuous hash chain is established in this way to ensure the relevance and traceability between snapshots; i When verifying snapshot S

[0032] i , first verify its previous snapshot S i-1 ;

[0033] Confirm that prevRootHash i in snapshot Smatches thisRootHash i in S i-1 ; i-1 Among them, thisRootHash

[0034] i-1 is the hash value calculated based on all the blocks within snapshot S i-1 ;

[0035] prevRootHash i is the hash value of the previous snapshot S i stored in snapshot S i-1 ; i

[0036] If they do not match, it means that the information of the previous snapshot S i-1 referred to by snapshot Sis inconsistent with the actual situation, and snapshot S i is directly determined to be invalid;

[0037] If they match, proceed to the next verification;

[0038] Recalculate the Merkle tree of snapshot S i and confirm that the calculated hash value is the same as the previously recorded prevRootHash​i+1 Consistent;

[0039] If the data is tampered with, the recalculated hash value will not match the originally recorded value;

[0040] If all checks pass, the snapshot S i is considered not to have been tampered with, and its own data is complete and valid.

[0041] In a preferred embodiment of the present invention, the snapshot creation and linking process includes:

[0042] When a snapshot needs to be created, the system determines the blocks that have changed since the last snapshot and calculates the hashes of the newly added or modified blocks;

[0043] Determine the previous snapshot S i of snapshot S i-1 and obtain the set of blocks that have changed relative to S i-1 , denoted as ΔB, and perform local updates by determining which blocks have changed relative to the previous version;

[0044] For each updated block, calculate the hash value blockHash k =H(blockData k );

[0045] ΔB is the set of block data for this round of demand processing;

[0046] blockData k is the content of the kth updated block;

[0047] blockHash k is the hash value of the content of the kth block;

[0048] Collect the hash values of all blocks in the updated snapshot S i for constructing a Merkle tree;

[0049] The Merkle tree adopts a bottom-up hash calculation method. First, calculate the hash value of each data block as a leaf node, and then merge the hashes of the left and right child nodes layer by layer:

[0050] h parent =H(h left_child ||h right_child );

[0051] where h parent is the hash value of the parent node;

[0052] h left_child and h right_childrespectively represent the hash values of the left and right child nodes of the parent node, and || represents the concatenation operation, which is used to maintain the order to ensure that the left and right child nodes will not be tampered with and cause hash errors;

[0053] H(*) is a hash function, which is used for node hash combination calculation, connecting data blocks to form a complete Merkle tree structure;

[0054] When all internal nodes are calculated, the hash value of the highest level is the hash value of the root node, which is used as the integrity summary of the data contained in the entire snapshot;

[0055] Sign with the private key and write the snapshot header containing the signature into persistent storage;

[0056] Mark the snapshot S i as completed;

[0057] Future snapshots will use the thisRootHash of this snapshot as their prevRootHash.

[0058] In a preferred embodiment of the present invention, the data recovery and verification include:

[0059] Master metadata node: Process requests for new snapshots, sharding, or time interval allocation, and write the updated content into the distributed consensus log;

[0060] Slave metadata node: Receive the updates from the master metadata node and synchronize them, which is used for read operation load balancing or quick switching in case of failure;

[0061] Maintain a multi-layer index structure, where each layer corresponds to a time range or data partition, so that specific data segments or relevant snapshots within a time range can be found more quickly:

[0062] RelevantSnapshots(t min ,t max ) = {S i |snapshotTime(S i ) ∈ [t min ,t max};

[0063] Among them, t min and t max represent the start and end times of the queried data or time range;

[0064] RelevantSnapshots(t min ,t max ) represents the set of all relevant snapshots returned within the time interval [t min ,t max ;

[0065] S i is the i-th snapshot;

[0066] snapshotTime(S i ) is the timestamp of snapshot S i ;

[0067] Store a compact filter for each snapshot to indicate which blocks or data IDs are included. During the recovery process, these filters can be queried to quickly skip snapshots that do not contain the required data:

[0068]

[0069] where filter(S i ) is the filter associated with snapshot S i to determine whether block x exists in snapshot S i ;

[0070] query(x) is the query algorithm, which means querying the filter filter(S i );

[0071] x is the index of the block to be queried;

[0072] The return value is a boolean: true means exists, and false means "does not exist";

[0073] Block-level voting and selection:

[0074] If a conflict occurs, that is, the weighted scores of multiple block versions are the same:

[0075] Use the database log and file system log to check which version of the block is consistent with the higher-level data:

[0076] If a certain version can match the transaction record in the log, it is preferred;

[0077] If a certain version can be verified on the hash chain of the front and back data blocks, select that version;

[0078] For blocks that cannot be determined, use weighted consensus and assign weights according to replication reliability and trust scores:

[0079]

[0080] argmax is the optimal block selection based on weighted voting, which is used to select the version with the highest matching degree from all possible block versions;

[0081] Valid(B x ) is the data of the selected correct block;

[0082] {versions of B x} represents the set of multiple versions that the same block may have in different replicas;

[0083] replicas is the set of replicas participating in the voting {r1, r2,....., r n}, where n is the number of replicas currently storing and participating in the voting;

[0084] ω r is the weight of replica r, dynamically adjusted based on historical data;

[0085] H(b) is the hash value of data b, used for integrity verification, snapshot hash matching, and consensus voting to ensure data consistency;

[0086] I(*) is the indicator function, taking the value 1 if the hash values match and 0 if they don't;

[0087] blockHash x,r is the hash of block B recorded by replica r x ;

[0088] The weight ω r combined with the bit error rate ε r , maps reliability to a weighting factor:

[0089] ω r = e -βεr ;

[0090] where β is the sensitivity parameter, used to adjust the influence of the bit error rate on the weight;

[0091] When the bit error rate ε r approaches 0, the weight ω r approaches 1. When the bit error rate ε r is large, the weight ω r rapidly decreases, reducing the influence of this replica;

[0092] Request block data from multiple replicas simultaneously. Once the block returned by a certain replica passes the verification, the process of obtaining this block stops. This can reduce the recovery time in a high-latency or partially faulty network. Send the retrieval request to the set of replicas {r1, r2,....., r n};

[0093] Once a replica returns a block B whose hash matches blockHash x,r x , accept this block;

[0094] ​If the same block or data segment has different timestamps in multiple snapshots, the system can re-verify whether the block has indeed not changed during these time intervals; this can capture replay attacks or mislabeled data.

[0095] The present invention also discloses a computer system, including:

[0096] A memory for storing processor-executable instructions;

[0097] Wherein, the processor is configured to implement the disclosed multi-level data recovery method for information resources when executing the executable instructions.

[0098] The present invention also discloses a computer-readable storage medium, including:

[0099] A memory having a computer program stored thereon;

[0100] A processor for executing the program in the memory to implement the disclosed multi-level data recovery method for information resources.

[0101] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0102] By utilizing the redundant information between multiple copies and multiple time-point snapshots, it is possible to find intact data blocks even when some snapshots or copies are partially damaged, thereby significantly improving the recovery effect.

[0103] Through deduplication and differential scanning, the number of invalid reads and repeated verifications is significantly reduced; this is particularly obvious in large-scale storage scenarios.

[0104] The blockchain-style hash link and security log mechanism can record the data evolution throughout the process, promptly detect and locate tampering or anomalies, and meet the requirements of auditing and evidence collection.

[0105] It can be applied to various environments such as cloud storage, hybrid cloud, enterprise-level distributed storage, database clusters, etc. It is suitable for both conventional data deletion or disk failures and can also achieve unified snapshot management on containerized or virtualized platforms.

[0106] Through the above technical solutions, the present invention realizes efficient information integration and interaction between the physical layer, logical layer, and application layer, breaks the dependence of existing data recovery technologies on a single snapshot or a single file system, and provides a feasible and innovative path for multi-level and multi-version data protection and recovery.

[0107] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. Description of the Drawings

[0108] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0109] Figure 1 is the system architecture diagram of the present invention. Detailed Embodiments

[0110] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as a limitation of the present invention.

[0111] As Figure 1 shown, the present invention aims to address the many deficiencies faced by existing data recovery solutions in multi-snapshot, multi-copy, and distributed storage environments, and proposes a "multi-level data recovery method". By introducing a snapshot management strategy based on blockchain-style hash chaining, combined with cross-timepoint duplicate data detection and fine-grained data consistency voting mechanisms, comprehensive data recovery and verification for multiple versions and multiple copies are achieved.

[0112] 1. Overall Concept

[0113] 1.1 Snapshot Hash-Chaining

[0114] Each snapshot is associated with a hash value (e.g., Merkle Root) to represent all blocks (or files) in that snapshot.

[0115] Subsequent snapshots reference the previous snapshot through its hash value, forming a hash chain.

[0116] 1.2 Tamper-Evident Verification

[0117] If any block in an earlier snapshot is modified, its hash value will not match and propagate backward, causing the entire chain to break.

[0118] By storing the hash values of these snapshots in a secure log or distributed ledger, it can be verified that no tampering has occurred over time.

[0119] 1.3 Recovery Use-Case

[0120] When performing data recovery, the system cross-references the block hashes of multiple snapshots.

[0121] If a block in a certain snapshot is found to be damaged or inconsistent, the system can obtain the correct block from other snapshots or replicas.

[0122] Using the "hash chain" can quickly confirm the integrity of the block without having to perform a full scan from start to end.

[0123] 2. Overview of System Architecture

[0124] Figure 1 Shows a simplified hierarchical architecture integrating blockchain-style snapshot links:

[0125] Recovery Coordination / Snapshot Comparison Layer: The comparison logic between snapshots, fine-grained block "voting", and determining which block version is correct are implemented here.

[0126] Blockchain-style Hash Link Layer: The core of the present invention, responsible for organizing the hash structure of snapshots, managing the link references between snapshots, and providing a verification interface to the upper layer.

[0127] 3. Data Structure and Hash Link

[0128] 3.1 Snapshot Metadata

[0129] For each snapshot S i , it contains the following:

[0130] Snapshot Header

[0131] snapshotID: Used to uniquely identify the snapshot, using a globally increasing serial number.

[0132] prevRootHash: Stored in snapshot S i+1 , pointing to the hash value thisRootHash of the previous snapshot S i+1 of S i , used to establish the snapshot chain.

[0133] thisRootHash: The hash value of the current snapshot, calculated based on the hashes of all data blocks (or higher-level indexes).

[0134] 3.2 Snapshot Block Index

[0135] Each snapshot maintains a table or database record corresponding to the set of data blocks under this snapshot:

[0136] blockID: Uniquely identifies the location of a certain data block within the logical volume or file system (such as offset, block number, or object ID).

[0137] blockHash: Stores the cryptographic hash value of the block, used for integrity verification and cross-version comparison.

[0138] size: Information about the block size (such as 4KB, 64KB).

[0139] replicaLocations: Indicates how many replicas (Replica Node) or storage nodes this block exists on, facilitating the location of available copies in case of failure.

[0140] 3.3 Merkle Tree Structure

[0141] Build a Merkle tree using all blockHashes in the snapshot, and its hash value is thisRootHash.

[0142] 3.3.1 Overview of Merkle Tree:

[0143] A Merkle tree (sometimes called a hash tree) is a tree-shaped data structure where leaf nodes store data blocks or block-level hashes, and intermediate nodes combine the hash values of two (or more) child nodes and then hash them again, thus calculating layer by layer upwards until a hash value (Merkle Root) is finally obtained.

[0144] 3.3.2 Hash Calculation of Leaf Nodes and Intermediate Nodes

[0145] Leaf Node:

[0146] Given data block B i , its leaf node hash can be expressed as:

[0147] H(B i );

[0148] where H(B i ) is the corresponding hash value of the i-th data block;

[0149] H(*) is a hash function used to ensure that if the data block B i is tampered with, H i will also change accordingly, thereby detecting data corruption or tampering at the leaf level;

[0150] Internal Node:

[0151] For the hashes H left and H right of two child nodes, calculate the hash of their parent node:

[0152] H parent = Hash(H left ||H right );

[0153] Among them, || represents the concatenation operation of strings. For a k-ary tree, it is extended to:

[0154] H parent = Hash(H1||H2||…||H k );

[0155] Hash value

[0156] If the entire tree is calculated from bottom to top until only one top-level node remains, its hash is the hash value RootHash. In the current snapshot, this hash value is thisRootHash.

[0157] 3.4 Construction process

[0158] Data block sorting: Collect all data blocks to be included in the verification in this snapshot and generate leaf hashes;

[0159] Hierarchical hashing: Group the leaf hashes in pairs or k at a time in order and calculate their upper-level hashes until the final hash value is obtained;

[0160] Record the hash value: Write the hash value into the thisRootHash field of the snapshot header;

[0161] Reference the previous snapshot: Write the hash value of the previous snapshot into prevRootHash (or the set in the multi-parent scenario) to complete the chain link.

[0162] 3.5 Chaining Logic

[0163] The chaining logic between snapshots is as follows:

[0164] S i+1 → S i , that is, each new snapshot S i+1 records its previous snapshot, that is, S i 's hash value, and → represents the link relationship between the front and back snapshots;

[0165] It is reflected in the metadata as:

[0166] prevRootHash i+1 = thisRootHash i ;

[0167] Among them, thisRootHash i is the hash value calculated based on all blocks within snapshot S i ;

[0168] prevRootHash i+1 is stored in snapshot S i+1The previous snapshot S in i of the hash value, snapshot S i To establish a continuous hash chain to ensure the relevance and traceability between snapshots;

[0169] When verifying snapshot S i first verify its previous snapshot S i-1 ;

[0170] Confirm the prevRootHash in snapshot S i matches the thisRootHash in S i and S i-1 ; i-1 match;

[0171] Among them, thisRootHash i-1 is the hash value calculated based on all blocks within snapshot S i-1 ;

[0172] prevRootHash i is the hash value of the previous snapshot S stored in snapshot S i ; i-1 of the hash value;

[0173] If they do not match, it means that the information of the previous snapshot S i referred to by snapshot S i-1 does not match the actual situation, and directly determine that snapshot S i is invalid;

[0174] If they match, proceed to the next verification;

[0175] Recalculate the Merkle tree of snapshot S i to confirm that the calculated hash value is consistent with the previously recorded prevRootHash i+1 ;

[0176] If the data is tampered with, the recalculated hash value will not match the previously recorded value;

[0177] If all checks pass, it can be determined that snapshot S i has not been tampered with, and its own data is complete and valid.

[0178] 4. Snapshot creation and linking process

[0179] 4.1 Data refresh and block hash calculation

[0180] When a snapshot needs to be created, the system determines the blocks that have changed since the last snapshot and calculates (or reuses the cached) hashes of the new / modified blocks.

[0181] Determine snapshot Si The previous snapshot S i-1 of its hash value, and obtain the set of blocks that have changed relative to S i-1 is denoted as ΔB. By determining which blocks have changed relative to the previous version, local updates are performed;

[0182] For each updated block, calculate the hash value blockHash k = H(blockData k );

[0183] ΔB is the set of block data processed for this round of requirements;

[0184] blockData k is the content of the k-th updated block;

[0185] blockHash k is the hash value of the content of the k-th block;

[0186] 4.2 Merkle Tree Construction

[0187] Collect the hash values of all blocks in the updated snapshot S i for constructing a Merkle tree (including newly added and unchanged blocks that still logically belong to this snapshot).

[0188] The Merkle tree adopts a bottom-up hash calculation method. First, calculate the hash value of each data block as a leaf node, and then merge the hash values of the left and right child nodes layer by layer::

[0189] h parent = H(h left_child ||h right_child );

[0190] where h parent is the hash value of the parent node;

[0191] h left_child and h right_child represent the hash values of the left and right child nodes of the parent node respectively. || represents the concatenation operation, which is used to maintain the order to ensure that the left and right child nodes will not be tampered with and cause hash errors;

[0192] H(*) is a hash function, which is used for node hash combination calculation, connecting data blocks to form a complete Merkle tree structure;

[0193] When all internal nodes are calculated, the hash value of the highest level is the hash value of the root node, which is used as the integrity summary of the data contained in the entire snapshot;

[0194] Sign with the private key and write the snapshot header containing the signature to persistent storage;

[0195] 4.3 Release Snapshot

[0196] Mark snapshot S i as completed.

[0197] Future snapshots will use the thisRootHash of this snapshot as their prevRootHash.

[0198] 5. Data Recovery and Verification

[0199] 5.1 Hierarchical Snapshot Index

[0200] Primary Metadata Node: Processes requests for new snapshots, shard or time interval allocation, and writes updates to the distributed consensus log;

[0201] Replica Metadata Node: Receives updates from the primary metadata node and synchronizes them for read operation load balancing or quick failover in case of failures;

[0202] Maintains a multi - level index structure where each level corresponds to a time range or data partition, enabling faster finding of relevant snapshots for a specific data segment or time range:

[0203] RelevantSnapshots(t min ,t max ) = {S i | snapshotTime(S i ) ∈ [t min ,t max};

[0204] Where, t min and t max represent the start and end times of the data or time range being queried;

[0205] RelevantSnapshots(t min ,t max ) represents the set of all relevant snapshots returned within the time interval [t min ,t max ;

[0206] S i is the i - th snapshot;

[0207] snapshotTime(S i ) is the timestamp of snapshot S i ;

[0208] Stores a compact filter for each snapshot to indicate which blocks or data IDs it contains. During the recovery process, these filters can be queried to quickly skip snapshots that do not contain the required data:

[0209]

[0210] Among them, filter(S i ) is a filter associated with snapshot S i for determining whether block x exists in snapshot S i ;

[0211] query(x) is a query algorithm, which means querying the filter filter(S i );

[0212] x is the block index of the area to be queried;

[0213] The return value is a boolean value: true represents "exists", and false represents "does not exist".

[0214] 5.2 Obtain the snapshot header and verify the chain

[0215] Each snapshot can store a minimum Merkle proof instead of transmitting the entire snapshot data, enabling verification with a small amount of hash data without downloading the entire block during data recovery; the snapshot references the root of the previous snapshot, reducing bandwidth and accelerating chain verification.

[0216] 5.3 Block-level voting and selection

[0217] Block-level voting and selection:

[0218] If a conflict occurs, that is, the weighted scores of multiple block versions are the same:

[0219] Use the database log and file system log to check which version of the block is consistent with the higher-level data:

[0220] If a certain version can match the transaction record in the log, it is preferred;

[0221] If a certain version can be verified on the hash chain of the front and back data blocks, select that version;

[0222] For blocks that cannot be determined, weighted consensus is adopted, and weights are assigned according to replication reliability and trust scores:

[0223]

[0224] argmax is the optimal block selection based on weighted voting, which is used to select the version with the highest matching degree from all possible block versions;

[0225] Valid(B x ) is the data of the selected correct block;

[0226] {versions of B x} represents the set of multiple versions that the same block may have in different replicas;

[0227] replicas is the set of replicas participating in the vote {r1, r2,....., r n}, where n is the number of replicas currently storing and participating in the vote;

[0228] ω r is the weight of replica r, dynamically adjusted based on historical data;

[0229] H(b) is the hash value of data b, used for integrity verification, snapshot hash matching, and consensus voting to ensure data consistency;

[0230] I(*) is the indicator function, which takes the value 1 if the hash values match and 0 if they do not match;

[0231] blockHash x,r is the hash of block B recorded by replica r x ;

[0232] The weight ω r combines with the bit error rate ε r , mapping reliability into the weighting factor:

[0233]

[0234] where β is the sensitivity parameter, used to adjust the influence of the bit error rate on the weight;

[0235] When the bit error rate ε r approaches 0, the weight ω r approaches 1. When the bit error rate ε r is large, the weight ω r rapidly decreases, reducing the influence of this replica.

[0236] 5.4 Reconstructing Missing or Damaged Blocks

[0237] Request block data from multiple replicas simultaneously. Once the block returned by a certain replica passes the verification, the process of obtaining this block stops. This can reduce the recovery time in a high-latency or partially faulty network. Send the retrieval request to the replica set {r1, r2,....., r n};

[0238] Once a replica returns a block B whose hash matches blockHash x,r , accept this block. x

[0239] 5.5 End-to-End Integrity Check​

[0240] If the same block or data segment has different timestamps in multiple snapshots, the system can re-verify whether the block has indeed not changed during these time intervals. This can catch replay attacks or mislabeled data.

[0241] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for recovering multi-level data of information resources, characterized in that: include: S1, data structure and hash link: Hash links are used to form data structures, so that snapshots are not only sorted by time, but also form a verifiable hash chain; S2, snapshot creation and linking process: by flushing data, block hashing, building a Merkle tree, writing snapshot headers, and storing snapshot information in blockchain or distributed ledgers, combined with block hashing and deduplication storage, redundant data usage is reduced; S3, data recovery and verification: Identify the required data and verify the integrity of the snapshot chain, use the weighted consensus voting mechanism to perform multi-level data recovery, vote at the block level to obtain the correct data, and finally perform end-to-end integrity check and complete recovery.

2. A method for recovering multi-level data of information resources according to claim 1, characterized in that: The data structure and hash link include: In snapshot metadata, each snapshot S i ,Include: Snapshot Header; snapshotID: used to uniquely identify the snapshot, using a globally increasing sequence number; prevRootHash: stored in snapshot S i+1 , point to snapshot S i+1 The previous snapshot S i The hash value thisRootHash is used to establish the snapshot chain; thisRootHash: The hash value of the current snapshot, calculated based on the hash of all data blocks; Each snapshot maintains a table or database record, corresponding to the set of data blocks under this snapshot: blockID: uniquely identifies the location of a data block in a logical volume or file system; blockHash: stores the encrypted hash value of the block, which is used for integrity verification and cross-version comparison; size: block size information; replicaLocations: Indicates how many replicas or storage nodes this block exists on, to facilitate locating available copies in case of failure; Given a data block B i , its leaf node hash value can be expressed as: H(B i ); Among them, H(B i ) is the corresponding hash value of the i-th data block; H(*) is a hash function used to ensure that if data block B i When tampered, H(B i ) will also change accordingly, thereby detecting data corruption or tampering at the leaf level; The link logic between snapshots is: S i+1 →S i , that is, each new snapshot S i+1 Record the previous snapshot S i The hash value of → indicates the link relationship between the previous and next snapshots; This is reflected in the metadata as follows: prevRootHash i+1 =thisRootHash i ; Among them, thisRootHash i is based on snapshot S i The hash value calculated by all blocks in the prevRootHash i+1 is stored in snapshot S i+1 The previous snapshot S in i The hash value of snapshot S i This creates a continuous hash chain to ensure the association and traceability between snapshots; When verifying the snapshot S i When , first verify its previous snapshot S i-1 , confirm that the connection relationship between the before and after snapshots is correct; Confirm Snapshot S i prevRootHash in i With S i-1 thisRootHash i-1 Match; Among them, thisRootHash i-1 is based on snapshot S i-1 The hash value calculated by all blocks in the prevRootHash i is stored in snapshot S i The previous snapshot S in i-1 The hash value of If they do not match, then snapshot S i The previous snapshot S referenced i-1 The information does not match the actual situation, and the snapshot S is directly judged i invalid; If they match, the next step of verification is performed; Recalculate snapshot S i Merkle tree to confirm that the calculated hash value is consistent with the previously recorded prevRootHash i+1 Consistency; If the data is tampered with, the recalculated hash value will not match the originally recorded value; If all checks pass, snapshot S can be identified. i It has not been tampered with, and its own data is complete and valid.

3. The method for recovering multi-level data of information resources according to claim 1, characterized in that: The snapshot creation and linking process includes: When a snapshot needs to be created, the system determines the blocks that have changed since the last snapshot and calculates the hash of the new or modified blocks; Determine the snapshot S i The previous snapshot S i-1 The hash value of S i-1 The set of blocks that have changed is denoted as ΔB. We perform local updates by determining which blocks have changed relative to the previous version. For each updated block, calculate the hash value blockHash k =H(blockData k ); ΔB is the block data set that needs to be processed in this round; blockData k is the content of the kth update block; blockHash k is the hash value of the content of the kth block; Collect the updated snapshot S i The hash values ​​of all blocks in the network are used to build the Merkle tree; The Merkle tree uses a bottom-up hash calculation method. First, the hash value of each data block is calculated as a leaf node, and then the hashes of the left and right child nodes are merged layer by layer: h parent =H(h left_child ||h right_child ); Among them, h parent is the hash value of the parent node; h left_child and h right_child They represent the hash values ​​of the left and right child nodes of the parent node respectively. || represents the concatenation operation, which is used to maintain the order and ensure that the left and right child nodes will not be tampered with, resulting in hash errors. H(*) is a hash function; When all internal nodes are calculated, the highest-level hash is the hash value of the root node, which serves as the integrity summary of the data contained in the entire snapshot; Sign with the private key and write the snapshot header containing the signature to persistent storage; Snapshot S i Mark as done; Future snapshots will use this snapshot's thisRootHash as their prevRootHash.

4. The method for recovering multi-level data of information resources according to claim 1, characterized in that: The data recovery and verification includes: Master metadata node: handles requests for new snapshots, shards, or time interval allocations, and writes updates to the distributed consensus log; Slave metadata node: receives updates from the master metadata node and synchronizes them for load balancing of read operations or fast switching in case of failure. Maintain a multi-layer index structure, where each layer corresponds to a time range or data partition, so that relevant snapshots within a specific data segment or time range can be found more quickly: RelevantSnapshots(t min ,t max )={S i |snapshotTime(S i )∈[t min ,t max ]}; Among them, t min and t max Indicates the start and end time of the queried data or time range; RelevantSnapshots(t min ,t max ) means returning the min ,t max ] a collection of all relevant snapshots; S i is the i-th snapshot; snapshotTime(S i ) is the snapshot S i timestamp; A compact filter is stored for each snapshot, indicating which blocks or data IDs are contained in it. During recovery, these filters can be queried to quickly skip snapshots that do not contain the required data: Among them, filter(S i ) is the same as snapshot S i The associated filter is used to determine whether block x exists in snapshot S i middle; query(x) is a query algorithm, which indicates the filter(S i ) to make inquiries; x is the block index to be queried; The return value is a Boolean value: true means "existence", false means "non-existence"; Block-level voting and selection: In case of a conflict, i.e. multiple block versions have the same weighted score: Use database logs and file system logs to check which version of the block is consistent with higher-level data: If a version can match the transaction record in the log, it will be preferred; If a version can be verified on the hash chain of the previous and next data blocks, then this version is selected; For blocks that cannot be determined, a weighted consensus is used to assign weights based on replication reliability and trust score: Argmax is the optimal block selection based on weighted voting, which is used to select the version with the highest matching degree from all possible block versions; Valid(B x ) is the data of the selected correct block; {versions of B x } represents a set of multiple versions of the same block that may exist in different replicas; replicas is the set of replicas participating in the vote {r1,r2,.....,r n }, n is the number of replicas currently stored and participating in voting; ω r is the weight of replica r, which is dynamically adjusted based on historical data; H(b) is the hash value of data b, which is used for integrity verification, snapshot hash matching, and consensus voting to ensure data consistency; I(*) is an indicator function, which takes the value 1 if the hash values ​​match and 0 if they do not match; blockHash x,r It is block B recorded by replica r x Hash of Weight ω r Combined bit error rate ε r , mapping the reliability into weighting factors: Among them, β is the sensitivity parameter, which is used to adjust the influence of bit error rate on weight; When the bit error rate ε r When it approaches 0, the weight ω r Approaching 1, when the bit error rate ε r When the weight ω is large, r Rapidly reduce the impact of this copy; Request block data from multiple replicas at the same time. Once a block returned by a replica passes verification, the acquisition process of the block stops. This can reduce recovery time in a high-latency or partially faulty network. The retrieval request is sent to the replica set {r1,r2,.....,r n }; Once a copy is returned, a hash is returned with blockHash x,r Matching block B x , then accept the block; If the same block or piece of data has different timestamps in multiple snapshots, the system can revalidate that the block has indeed not changed between those time intervals; this can catch replay attacks or mislabeled data.

5. A computer system, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement a multi-level data recovery method for information resources as described in any one of claims 1 to 4 when executing the executable instructions.

6. A computer-readable storage medium, characterized in that: include: a memory having a computer program stored thereon; A processor is used to execute the program in the memory to implement a multi-level data recovery method for information resources as described in any one of claims 1 to 4.

Citation Information

Cited By

  • Data storage method for artificial intelligence learning mode

    CN120723945A

  • File digitalization full life cycle encryption integrity verification method and system

    CN121118091A

  • Block chain storage optimization method, system and equipment based on hierarchical compression and dynamic fragmentation and medium

    CN121350146A

  • Metadata processing method based on data snapshot and electronic equipment

    CN121542090A