A method for recovering electronic information data from archives

By constructing a structured metadata-driven version sequence and hierarchical repair strategy, the problems of version management, heterogeneous data repair, and security auditing in archival data recovery technology have been solved, achieving efficient data recovery and secure retrieval, and meeting the stringent requirements of the judicial and government sectors.

CN120631871BActive Publication Date: 2025-10-28GUIZHOU BLUESKY INNOVATIVE SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511151136.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-10-28
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing archival data recovery technologies have shortcomings in data integrity and version management. They cannot effectively manage historical versions, their heterogeneous data repair strategies are crude, image restoration is prone to pixel distortion, access control and security auditing are insufficient, file integration and catalog reconstruction are inefficient, and they cannot generate dynamic hyperlinks, resulting in long retrieval times.

Method used

By constructing a version sequence driven by structured metadata, structured metadata such as version number, timestamp, and file association is generated. A linear timeline and tree-like branch structure are constructed to record the editing history of multiple users. A hierarchical repair strategy is adopted, combined with text semantic repair and image interpolation technology. Blockchain notarization and dual verification are integrated to achieve access control and security auditing.

Benefits of technology

It achieves systematic management of data object versions, meets version traceability requirements, solves problems such as formatting errors, pixel distortion, and broken metadata associations, ensures that audit records are tamper-proof, improves retrieval efficiency and security, and meets network security level protection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631871B_ABST
    Figure CN120631871B_ABST
Patent Text Reader

Abstract

This invention discloses a method for recovering electronic archival information, relating to the field of archival management and data recovery technology. The method includes the following steps: acquiring original archival data; extracting first data and generating second data according to format characteristics; the second data including hierarchical identifiers and structured metadata; generating a version sequence of data objects based on the structured metadata; matching third data according to a preset strategy; decrypting unextracted encrypted data blocks in the first data as fourth data; generating fifth data according to format characteristics; if there are unextracted unencrypted damaged data blocks in the first data, generating supplementary repair data; integrating the first, third, and fifth data and the repair data according to the hierarchical identifiers; reconstructing the association between the catalog and metadata; generating a complete file; constructing a structured metadata version sequence; hierarchically repairing heterogeneous data; and integrating blockchain dual verification to form a three-in-one system suitable for archival recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of archives management and data recovery technology, and in particular to a method for recovering electronic information data from archives. Background Technology

[0002] In recent years, with the deepening of digital transformation, electronic archives, as the core carrier of information records, have been widely used in scenarios such as judicial case file management, government document archiving, and enterprise contract storage. The core value of electronic archives lies in their integrity, usability, and security. However, factors such as aging storage media, virus attacks, and human error have led to frequent problems such as loss of text data byte streams, pixel damage in image scans, and broken metadata associations. Traditional recovery technologies rely on single-version storage and extensive repair strategies, which are difficult to effectively manage multiple version histories in complex branch editing scenarios and have poor repair effects on data with high damage rates. They can no longer meet the stringent requirements of "original document uniqueness" for judicial evidence collection, "long-term readability" for government archives, and "audit traceability" for enterprise compliance. Against this background, building an efficient data recovery system that takes into account version management, intelligent repair, and security auditing has become a key technical path to solve the pain points of electronic archive management.

[0003] Existing archival data recovery technologies face numerous unresolved issues. Regarding data integrity and version management, traditional methods rely on single-version storage, lacking systematic management of historical versions. For example, early version control systems like SVN only supported linear backtracking, failing to record branch modifications and merge conflicts during concurrent editing by multiple users, resulting in low version matching accuracy in complex branch scenarios. Heterogeneous data repair strategies are crude; text repair relies solely on CRC checksum redundancy, lacking semantic association processing, leading to increased formatting errors. Image repair, under high damage rates, is prone to pixel distortion and lacks OCR text layer restoration. Original metadata often suffers from inconsistent formats, missing fields, and a lack of standardized processing, resulting in broken relationships after repair. Access control and security auditing have significant deficiencies. Key management relies on single-factor authentication with static passwords, with insufficient integration of biometrics and dynamic tokens. Decryption logs stored in centralized databases are easily tampered with. File consolidation and catalog reconstruction are inefficient, with catalog reconstruction relying solely on "Year-Institution-..." The static hierarchical mapping of "security level" does not combine version information to quickly locate the latest valid version. The metadata association is limited to physical path mapping and cannot generate dynamic hyperlinks for file association relationships, resulting in a long average retrieval time for volumes. Summary of the Invention

[0004] The technical problem addressed by this invention is that existing archival data recovery technologies suffer from several unresolved issues. Regarding data integrity and version management, traditional methods rely on single-version storage, lacking systematic management of historical versions. For example, early version control systems like SVN only supported linear backtracking, failing to record branch modifications and merge conflicts during concurrent editing by multiple users, resulting in low version matching accuracy in complex branch scenarios. Heterogeneous data repair strategies are crude; text repair relies solely on CRC checksum redundancy, lacking semantic association processing, leading to increased formatting errors. Image repair, under high damage rates, is prone to pixel distortion and lacks OCR text layer repair. Original metadata often suffers from inconsistent formats, missing fields, and a lack of standardized processing, resulting in broken relationships after repair. Access control and security auditing have significant deficiencies; key management relies on single-factor authentication with static passwords, with insufficient integration of biometrics and dynamic tokens. Decryption logs stored in centralized databases are easily tampered with. File integration and catalog reconstruction are inefficient, with catalog reconstruction relying solely on "Year-Institution-..." The static hierarchical mapping of "security level" does not combine version information to quickly locate the latest valid version. The metadata association is limited to physical path mapping and cannot generate dynamic hyperlinks for file association relationships, resulting in a long average retrieval time for volumes.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A method for recovering electronic information data of archives includes the following steps:

[0006] Step S1: Obtain the original archive data, extract the first data, and generate the second data according to the format characteristics of the file extension and byte encoding rules;

[0007] The second data includes hierarchy identifiers and structured metadata;

[0008] Step S2: Generate a version sequence of data objects based on the version number and timestamp of the structured data, and match the third data according to a preset strategy;

[0009] Step S3: The encrypted data block that was not extracted from the first data is the fourth data, which is then decrypted, and the fifth data is generated according to the format characteristics of the file extension and byte encoding rules.

[0010] If there are unextracted, corrupted, unencrypted data blocks in the first data, then supplementary repair data will be generated;

[0011] Step S4: Integrate the first data, third data, fifth data and repair data according to the hierarchical identifier of the second data, and rebuild the association between the directory and metadata to generate a complete dossier.

[0012] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, the original archival data consists of two or more data objects;

[0013] The first data includes text file data, image scan data, and original file metadata;

[0014] The structured metadata is obtained by standardizing the format and filling in missing fields of the original archive metadata;

[0015] The structured metadata includes version number, modification timestamp, and file association relationships;

[0016] The version sequence is generated based on the version number and the modification timestamp;

[0017] The third data is the version of a valid data object that has been verified.

[0018] The third data includes the complete content of the data object and the corresponding structured metadata snapshot;

[0019] The complete content of the data object includes the binary data of the text file and the binary data of the image file;

[0020] The structured metadata snapshot corresponding to the complete content of the data object includes a version number, timestamp, associated file ID, and content hash value.

[0021] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, the step of generating a version sequence of data objects based on structured metadata and matching third data according to a preset strategy specifically includes:

[0022] Based on the version number and modification timestamp in the structured metadata, a linear version timeline and a tree-like version branch structure are constructed.

[0023] The linear version timeline is arranged in ascending order of modification timestamps to form a complete modification history axis of the data object.

[0024] The tree-like version branch structure takes the initial version as the root node and forms the main branches. When a data object undergoes branch modification, a version node is generated according to the version number to record the branch source and merging relationship.

[0025] The branch modification includes simultaneous editing by different users;

[0026] Each version node includes a structured metadata snapshot and a content hash value, used to quickly verify the complete content consistency of the data object;

[0027] The verification specifically includes:

[0028] When generating a version node, a hash value is calculated for the complete content of the data object of the current version corresponding to the node, and stored in the structured metadata snapshot;

[0029] During verification, the hash value of the complete content of the actual data object stored in the current system is recalculated and compared with the hash value of the content recorded in the version node;

[0030] If the hash values ​​match, it is determined that the complete content of the data object has not been tampered with or damaged.

[0031] If the hash values ​​are inconsistent, it is determined that the complete content of the data object has been altered or corrupted, and a repair process is triggered.

[0032] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, the application scenarios of the verification specifically include:

[0033] After repairing encrypted or unencrypted corrupted data blocks, the correctness of the repair result is verified by comparing the hash values.

[0034] When matching third data according to a preset strategy, the undamaged version is selected by filtering out the hash value;

[0035] When merging parallel branches in the tree-like version branch structure, content conflicts are identified by comparing hash values.

[0036] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, the preset strategy for matching third data specifically includes:

[0037] A version-based priority matching strategy includes selecting the version of the data object that has passed the content hash value verification and has the latest modification timestamp.

[0038] Matching strategies based on version association include filtering data object versions based on file associations and version number continuity in structured metadata;

[0039] Conflict resolution strategies based on branch structure include prioritizing the main branch and merging strategies for multi-user edits;

[0040] Special matching strategies based on data types, including text format feature verification and image metadata verification;

[0041] Abnormal situation handling strategies, including version missing rollback and hash value conflict verification;

[0042] The multi-user editing and merging strategy includes:

[0043] Sort by user permission level, and prioritize the modified versions modified by users with higher permissions;

[0044] If permissions are the same, merge them according to the modification timestamp, retaining the latest modification content;

[0045] The abnormal situation handling strategy includes:

[0046] When all candidate versions are corrupted, automatically roll back to the last available version of the data object;

[0047] When there are cases where the hash values ​​are the same but the contents are different, a second verification is performed;

[0048] The secondary verification verifies the data object by comparing the file size and modification timestamp.

[0049] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, natural language processing technology is used to identify the encoding format, paragraph structure and font features of the text archival data;

[0050] Extract resolution, color mode, file format, and OCR-recognized text layer data from scanned image data;

[0051] The format features include the file extension of the original archive data, byte encoding rules, and the definition specifications of the original archive metadata fields.

[0052] As a preferred embodiment of the method for recovering electronic archival information data according to the present invention, the method includes: verifying user permissions through a key management module, obtaining the corresponding decryption key according to the hierarchical identifier, and decrypting the fourth data, specifically including:

[0053] The key management module adopts a hierarchical key system, hierarchical identifiers, and decryption keys, which are associated through access control lists.

[0054] Each of the aforementioned hierarchical identifiers corresponds to key permissions;

[0055] The key permissions include browsing permissions, editing permissions, and management permissions;

[0056] The user permission verification adopts a dual verification mechanism, and the key acquisition process generates an operation log and stores it in encrypted form on the blockchain node;

[0057] The dual verification mechanism includes biometric authentication and dynamic token verification;

[0058] The biometric authentication verifies the consistency between the user's identity and the biometric information recorded in the permission database through a biometric collection device.

[0059] The biometrics include fingerprint recognition, facial recognition, or iris scanning;

[0060] The dynamic token verification involves the user inputting a dynamic token based on time synchronization or event triggering, and the system verifying whether the token matches the real-time token generated by the server.

[0061] If the biometric authentication and dynamic token verification pass simultaneously, the decryption key acquisition process for the corresponding level identifier is triggered.

[0062] An operation log is generated during the key acquisition process and stored in encrypted form via a blockchain node.

[0063] As a preferred embodiment of the method for recovering electronic information data of archives according to the present invention, wherein: if there are unextracted unencrypted damaged data blocks in the first data, supplementary repair data is generated;

[0064] The repaired data is repaired using a repair algorithm, specifically including:

[0065] A byte stream completion algorithm is used to repair text archive data;

[0066] An interpolation repair algorithm is used to repair the image scan data.

[0067] The repair strategy for unencrypted corrupted data blocks is tiered according to the degree of corruption, specifically including:

[0068] Repair directly when the damage rate is less than the first threshold;

[0069] When the corruption rate reaches the second to third threshold, historical version difference data is used for completion.

[0070] When the corruption rate exceeds the fourth threshold, version sequence matching of the data object to which the currently corrupted data block belongs is triggered to obtain an alternative version.

[0071] As a preferred embodiment of the method for recovering electronic information data of archives according to the present invention, in the integration process of step S4, the hierarchical identifier corresponds to the tree structure of the file directory, and the directory index is reconstructed by the depth-first traversal algorithm.

[0072] The hierarchical identifier corresponds to the tree structure of the dossier directory, with the root node being the dossier number and the version node being the year label, organization label, and security classification label.

[0073] The structured metadata association mechanism includes converting file associations in structured metadata into directory hyperlinks, and generating a unique storage path for data objects using version numbers and modification timestamps.

[0074] In a preferred embodiment of the method for recovering electronic archival information data according to the present invention, after generating a complete file, an integrity verification process is automatically executed:

[0075] The content hash value of all data objects in the dossier is recalculated and compared with the version node record value. A second repair is triggered when the verification pass rate is lower than the fifth threshold.

[0076] Simultaneously, a digital fingerprint of the dossier is generated and matched with the creation fingerprint in the original archive metadata for verification. Once the verification is successful, a timestamp certificate is added and the dossier is archived to the read-only storage area.

[0077] The beneficial effects of this invention are as follows: By constructing a version sequence driven by structured metadata, the original archive metadata is extracted and standardized to generate structured metadata containing version numbers, timestamps, and file associations. A linear timeline and tree-like branch structure are constructed to record the editing history of multiple users, achieving systematic management of data object versions and meeting the version traceability requirements of the State Archives Administration. For heterogeneous data repair, a hierarchical strategy is designed, employing direct repair, version difference completion, and alternative version matching methods according to the degree of damage. Combined with text semantic repair, image interpolation, and OCR text layer association technology, the problems of format disorder, pixel distortion, and broken metadata associations are solved. The security system integrates blockchain evidence storage and dual verification, using biometrics and dynamic tokens for dual authentication, binding key permissions and case file levels, and encrypting and storing decryption logs on the blockchain in real time to ensure that audit records are tamper-proof and meet the requirements of network security level protection. The above technologies are integrated to construct a three-in-one system of "version management - intelligent repair - security audit," breaking through the bottlenecks of traditional solutions in terms of integrity, efficiency, and security, and providing an innovative path for the recovery of electronic archives in the judicial and government fields, from the underlying algorithm to the upper-level application. Attached Figure Description

[0078] Figure 1 This is a schematic diagram of the basic process of a method for recovering electronic information data of archives provided in one embodiment of the present invention. Detailed Implementation

[0079] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0080] Example 1, referring to Figure 1 As an embodiment of the present invention, a method for recovering electronic information data of archives is provided, comprising the following steps:

[0081] Step S1: Obtain the original archive data, extract the first data, and generate the second data according to the format characteristics of the file extension and byte encoding rules;

[0082] The second data includes hierarchical identifiers and structured metadata;

[0083] Step S2: Generate a version sequence of data objects based on the version number and timestamp of the structured data, and match the third data according to a preset strategy;

[0084] Step S3: The encrypted data block that was not extracted from the first data is the fourth data, which is then decrypted, and the fifth data is generated according to the format characteristics of the file extension and byte encoding rules.

[0085] If there are unextracted, corrupted, unencrypted data blocks in the first data, then supplementary repair data will be generated;

[0086] Step S4: Integrate the first data, third data, fifth data and repair data according to the hierarchical identifier of the second data, and rebuild the association between the directory and metadata to generate a complete dossier.

[0087] In one embodiment, the original archive data is first acquired through a data acquisition interface. Text archive data, scanned image data, and original archive metadata are extracted as first data. Simultaneously, the original archive metadata is formatted and missing fields are filled in according to file extensions and byte encoding rules, generating second data including hierarchical identifiers (such as case file number and year tag) and structured metadata (version number, timestamp, and file association). Then, based on the version number and timestamp of the structured metadata, a linear version timeline and a tree-like version branch structure are constructed to form a complete version sequence of the data object. Finally, using a preset strategy combined with content hash value verification, the version sequence is matched... The valid version that has passed verification is used as the third data; then, the encrypted data blocks that have not been extracted from the first data are defined as the fourth data. After verifying the user's permissions through the key management module, the decryption key is obtained for decryption. After decryption, the fifth data is generated according to the format characteristics. If there are any unextracted, non-encrypted, corrupted data blocks, supplementary repair data is generated according to the degree of corruption. Finally, based on the hierarchical identifier of the second data, the dossier directory index is reconstructed using a depth-first traversal algorithm. Various types of data are integrated according to the hierarchy, and the file associations in the structured metadata are converted into directory hyperlinks. The unique storage path of the data object is generated through the version number and timestamp, completing the reconstruction of the directory and metadata association and generating a complete dossier.

[0088] The original archive data consists of two or more data objects;

[0089] The first data includes text archive data, image scan data, and original archive metadata;

[0090] Structured metadata is obtained by standardizing the format of the original archive metadata and filling in missing fields;

[0091] Structured metadata includes version number, modification timestamp, and file associations;

[0092] The version sequence is generated based on the version number and modification timestamp;

[0093] The third data is the version of the valid data object verified;

[0094] The third data includes the complete content of the data object and a snapshot of the corresponding structured metadata;

[0095] The complete content of the data object includes the binary data of the text file and the binary data of the image file;

[0096] A structured metadata snapshot corresponding to the complete content of the data object. The structured metadata snapshot includes the version number, timestamp, associated file ID, and content hash value.

[0097] In one embodiment, during implementation, the original archive data includes two or more data objects. Text archive data, image scan data, and original archive metadata are extracted as the first data. For the original archive metadata, a structured metadata is generated through standardized format processing (such as standardized field naming and data type conversion) and missing field completion (such as filling in default values ​​based on business rules or retrieving missing information by association). This structured metadata includes a version number (identifying the version iteration of the data object), a modification timestamp (recording the version update time), and file association relationships (marking the master-slave or dependency relationships between data objects). Based on the version number and modification timestamp in the structured metadata, a version sequence of data objects is generated in chronological order, visually presenting the complete modification of the data objects. Historically, through preset verification rules, such as content hash value comparison and format integrity checks, each version in the version sequence is verified, and the valid data object version that passes the verification is selected as the third data. The third data consists of two parts: first, the complete content of the data object, that is, the binary data of the text file, which retains the original encoding format and content, and the binary data of the image file, including pixel information and metadata; second, the corresponding structured metadata snapshot, which records the version number, timestamp (update time accurate to the second), associated file ID, unique code and content hash value of the associated data object (a unique fingerprint generated by calculating the complete content of the data object using the SHA-256 algorithm), used to quickly verify data integrity and version association.

[0098] Version sequences of data objects are generated based on structured metadata, and third-party data is matched according to a preset strategy, specifically including:

[0099] Based on the version number and modification timestamp in the structured metadata, a linear version timeline and a tree-like version branch structure are constructed.

[0100] A linear version timeline, arranged in ascending order of modification timestamps, forms a complete modification history axis for data objects;

[0101] The tree-like version branch structure takes the initial version as the root node and forms the main branches. When a data object is modified by a branch, a version node is generated according to the version number to record the branch source and merging relationship.

[0102] Branch modifications include simultaneous editing by different users;

[0103] Each version node includes a structured metadata snapshot and a content hash value, used to quickly verify the complete content consistency of data objects;

[0104] The verification specifically includes:

[0105] When generating a version node, a hash value is calculated for the complete content of the data object of the current version corresponding to the node, and stored in the structured metadata snapshot;

[0106] During verification, the hash value of the complete content of the actual data object stored in the current system is recalculated and compared with the hash value of the content recorded in the version node;

[0107] If the hash values ​​match, it is determined that the complete content of the data object has not been tampered with or damaged;

[0108] If the hash values ​​are inconsistent, it is determined that the complete content of the data object has been altered or corrupted, and a repair process is triggered.

[0109] In one embodiment, when generating version sequences and matching them with third-party data, a linear version timeline and a tree-like version branch structure are constructed based on the version number and modification timestamp in the structured metadata: the linear version timeline arranges all versions in ascending order of modification timestamp, forming a complete modification history axis of the data object; the tree-like version branch structure forms a trunk branch with the initial version (version number 1.0) as the root node. When multiple users edit simultaneously, i.e., when a branch modification occurs, version nodes are generated according to version number rules (such as main version and sub-version) and the branch source and merging relationship are recorded. Each version node includes a snapshot of structured metadata and a content hash value. When generating nodes, the SHA-256 algorithm is used to calculate the hash value of the complete content of the data object and store it in the snapshot. During verification, the hash value of the actual stored content is recalculated and compared with the recorded value. If they match, it is determined that the content has not been tampered with or damaged. If they do not match, it is determined that the content has been changed or damaged and a repair process is triggered. This mechanism realizes systematic version management through timelines and branch structures, and ensures data integrity by combining hash value verification, meeting the version traceability and conflict resolution needs in multi-user concurrent editing scenarios.

[0110] The application scenarios for verification include:

[0111] After repairing encrypted or unencrypted corrupted data blocks, the correctness of the repair result is verified by comparing the hash values.

[0112] When matching third-party data according to a preset strategy, the undamaged version is filtered out by hash value;

[0113] When merging parallel branches in a tree-like branching structure, content conflicts are identified by comparing hash values.

[0114] In one embodiment, the application scenarios of the verification mechanism include: after decrypting encrypted data blocks or repairing unencrypted damaged data blocks, the content hash value of the repaired data object is recalculated using the SHA-256 hash algorithm and compared with the original hash value recorded in the structured metadata snapshot. If they match, the repair is deemed correct; otherwise, a secondary repair process is triggered. When matching third data (valid version) according to a preset strategy, the version nodes in the version sequence are traversed to filter out versions whose content hash value verification passes (i.e., the hash value matches the recorded value), ensuring that the complete content of the selected version's data object has not been tampered with or damaged. Parallel branch merging is performed in the tree-like version branch structure. For example, when merging versions edited by multiple users, the content hash values ​​of the version nodes of the branches to be merged are compared. If the hash values ​​are different, a content conflict is determined, requiring manual intervention or automatic merging rules, such as retaining the latest modifications, resolving the conflict, and then performing the merging operation. This verification mechanism achieves efficient verification of data integrity and version accuracy in repair verification, version filtering, and branch merging scenarios through the consistency judgment of hash value comparison.

[0115] The preset strategy matches third-party data, specifically including:

[0116] A version-based priority matching strategy includes selecting the version of the data object that has passed the content hash value verification and has the latest modification timestamp.

[0117] Matching strategies based on version association include filtering data object versions based on file associations and version number continuity in structured metadata;

[0118] Conflict resolution strategies based on branch structure include prioritizing the main branch and merging strategies for multi-user edits;

[0119] Special matching strategies based on data types, including text format feature verification and image metadata verification;

[0120] Abnormal situation handling strategies, including version missing rollback and hash value conflict verification;

[0121] Multi-user editing and merging strategies include:

[0122] Sort by user permission level, and prioritize the modified versions modified by users with higher permissions;

[0123] If permissions are the same, merge them according to the modification timestamp, retaining the latest modification content;

[0124] Abnormal situation handling strategies include:

[0125] When all candidate versions are corrupted, automatically roll back to the last available version of the data object;

[0126] When there are cases where the hash values ​​are the same but the contents are different, a second verification is performed;

[0127] Secondary verification verifies data objects by comparing file size and modification timestamps.

[0128] In one embodiment, when implementing a preset strategy to match third data, the following core strategies are included: Version integrity priority strategy: Prioritize the version with the content hash value verification passed (i.e., it has not been tampered with or damaged) and the latest modification timestamp to ensure that the latest valid version of the data object is obtained;

[0129] Version association matching strategy: Based on the file association relationship (such as primary and secondary file association) and version number continuity (such as the incrementing logic of 1.0→1.1→1.2) in the structured metadata, filter the versions that are closely associated with the current data object and have a consistent version evolution.

[0130] Branch structure conflict resolution strategy: The main branch (the main branch derived from the initial version) is the priority to be merged. When multiple users are editing, they are sorted by permission level (e.g., administrators are higher than ordinary users). The modified version of the user with higher permission is selected first. If the permissions are the same, the latest content is retained according to the modification timestamp.

[0131] Special matching strategies for data types: Validate format features (such as encoding format and paragraph structure integrity) for text files, and validate metadata (such as resolution and color mode consistency) for image files.

[0132] Anomaly handling strategy: When all candidate versions are corrupted, automatically roll back to the previous usable version that passed the hash check; if a rare conflict occurs with the same hash value but different content, perform a secondary check by comparing file size and modification timestamp to ensure the accuracy of version matching.

[0133] By combining the above strategies, we can achieve intelligent filtering and conflict resolution of effective versions in complex scenarios, ensuring the integrity and accuracy of data object version matching.

[0134] Natural language processing technology is used to identify encoding formats, paragraph structures, and font features in text archive data.

[0135] Extract resolution, color mode, file format, and OCR-recognized text layer data from scanned image data;

[0136] The format features include the file extension of the original archive data, byte encoding rules, and the definition specifications of the original archive metadata fields.

[0137] In one embodiment, for text archive data, basic natural language processing techniques are used for format parsing: character frequency statistics and encoding feature matching, such as detecting the BOM header of UTF-8 and GBK. The system analyzes the double-byte character distribution to identify file encoding formats. It parses paragraph structure and hierarchical relationships based on text delimiter rules (such as newlines and consecutive blank lines) and hierarchical markers (such as the title prefix "Chapter 1" and bold font styles). It directly extracts built-in font metadata (such as the font type field in .doc files) or identifies font types (such as SimSun and HeiTi) and font sizes through image pixel distribution patterns (such as font stroke width and serif features). For scanned image data, it extracts underlying metadata based on resolution (such as 300 DPI), color mode (such as RGB and CMYK), and file format (such as JPEG and PNG) using image processing algorithms. It then uses mature OCR tools (such as Tesseract and Baidu OCR Engine) to generate text layer data and sets an OCR recognition confidence check (threshold set to 90%, below which a manual review process is triggered). Format features specifically cover the original archive data's file extensions (such as .doc and .jpg), byte encoding rules (such as UTF-8 byte order markers), and original archive metadata field definition specifications, such as timestamps using ISO. The 8601 format ensures format compatibility and feature integrity for different types of data through a standardized parsing process.

[0138] The key management module verifies user permissions, retrieves the corresponding decryption key based on the hierarchy identifier, and decrypts the fourth data, specifically including:

[0139] The key management module adopts a hierarchical key system, hierarchical identifiers, and decryption keys, which are associated through access control lists;

[0140] Each level identifier corresponds to a key permission;

[0141] Key permissions include browsing permissions, editing permissions, and management permissions;

[0142] User permission verification employs a dual verification mechanism, and the key acquisition process generates an operation log which is then encrypted and stored on the blockchain node.

[0143] The dual verification mechanism includes biometric authentication and dynamic token verification;

[0144] Biometric authentication verifies the consistency between a user's identity and the biometric information recorded in the permissions database through biometric collection devices.

[0145] Biometrics include fingerprint recognition, facial recognition, or iris scanning;

[0146] Dynamic token verification: The user inputs a dynamic token based on time synchronization or event triggering, and the system verifies whether the token matches the real-time token generated by the server.

[0147] If both biometric authentication and dynamic token verification pass simultaneously, the decryption key acquisition process for the corresponding level identifier is triggered.

[0148] An operation log is generated during the key acquisition process and stored in encrypted form via blockchain nodes.

[0149] In one embodiment, when implementing key management and permission verification, a hierarchical key system is constructed through the key management module: hierarchical identifiers (such as case file classification, departmental permissions) are associated with decryption keys through access control lists (ACLs). Each level corresponds to three types of key permissions: browsing, editing, and management. User permission verification adopts a dual verification mechanism: first, the consistency between the user's biometric features (fingerprint, face, or iris information) and the permission database records is verified through biometric collection devices (such as fingerprint sensors, facial recognition cameras). At the same time, the user inputs a dynamic token based on time synchronization (such as the TOTP algorithm) or event triggering. The system matches the user token with a token generated by the server in real time. Only when biometric authentication and dynamic token verification pass simultaneously is the decryption key acquisition process for the corresponding hierarchical identifier triggered. The operation log generated during the key acquisition process (including user ID, operation time, and key level) is encrypted and stored through blockchain nodes. The immutability of the blockchain ensures the security and traceability of the operation records, realizing hierarchical permission control and security auditing for the decryption of encrypted data blocks.

[0150] If there are unextracted, corrupted, unencrypted data blocks in the first data, then supplementary repair data will be generated;

[0151] Data repair is performed using repair algorithms, specifically including:

[0152] A byte stream completion algorithm is used to repair text archive data;

[0153] An interpolation repair algorithm is used to repair the image scan data.

[0154] The repair strategy for unencrypted corrupted data blocks is tiered according to the degree of corruption, specifically including:

[0155] Repair directly when the damage rate is less than the first threshold;

[0156] When the corruption rate reaches the second to third threshold, historical version difference data is used for completion.

[0157] When the corruption rate exceeds the fourth threshold, version sequence matching of the data object to which the currently corrupted data block belongs is triggered to obtain an alternative version.

[0158] In one embodiment, when repairing unencrypted corrupted data blocks, if unextracted corrupted data blocks are detected in the first data, the data block corruption rate (i.e., the proportion of corrupted bytes to the total number of bytes in the data block) is first calculated, and then processed according to the degree of corruption: when the corruption rate is less than 10% (the first threshold, based on the conventional minor corruption judgment standard in the field of data repair), a byte stream completion algorithm is used for text archive data, such as directly repairing by filling missing bytes through pattern matching of adjacent byte sequences; for image scan data, bilinear interpolation or neighbor pixel copying algorithms are used to repair local damage; when the corruption rate is between 10% and 30% (the second to third thresholds, verified to be within this range through version difference), ... When performing heterogeneous completion (balancing efficiency and accuracy), the difference data between adjacent versions (i.e., the content difference between the current damaged version and the previous valid version) is extracted from the version sequence of the data object, and the damaged part is completed by merging the difference data; when the damage rate is greater than 30% (the fourth threshold, defined as severe damage, with a direct repair success rate of less than 50%), the version sequence matching mechanism is triggered. Based on the version number and timestamp in the structured metadata, the most recent valid replacement version of the same type of data object is selected to cover the current damaged data block. Through the above hierarchical repair strategy, combined with text byte stream completion and image interpolation repair technology, efficient repair of data with different degrees of damage is achieved, ensuring the integrity and usability of electronic archive data.

[0159] During the integration process in step S4, the hierarchical identifier corresponds to the tree structure of the dossier directory, and the directory index is reconstructed using a depth-first traversal algorithm.

[0160] The hierarchical identifier corresponds to the tree structure of the dossier directory, with the root node being the dossier number and the version node being the year label, organization label, and security classification label.

[0161] The structured metadata association mechanism includes converting file associations in structured metadata into directory hyperlinks, and generating a unique storage path for data objects using version numbers and modification timestamps.

[0162] In one embodiment, when implementing data integration and directory reconstruction, a tree structure of the dossier directory is constructed based on hierarchical identifiers: the dossier number is used as the root node, and the dossier is expanded level by level according to the year label (such as "2023"), the organization label (such as "Finance Department"), and the security level label (such as "Confidential") to form a multi-level directory system. A depth-first traversal algorithm is used to recursively generate a directory index file, ensuring the integrity and traceability of the dossier structure. A structured metadata association mechanism is executed simultaneously: file relationships recorded in the structured metadata, such as the hierarchical relationship between the main file and attachments, are converted into directory hyperlinks, allowing users to jump to the associated files by clicking on directory nodes. A unique storage path for data objects is generated by combining version numbers (e.g., "1.0") with modification timestamps (time series accurate to the second) (format: "Dossier Number / Year / Institution / Security Class / Version Number-Timestamp.Extension"), ensuring that each version of the data object has a unique identifier in the storage system. The first, third, and fifth data, as well as the repaired data, are integrated hierarchically to complete the reconstruction of the dossier directory and metadata association, ultimately generating a logically clear and quickly searchable electronic archive dossier.

[0163] After the complete dossier is generated, the integrity verification process will be executed automatically:

[0164] The content hash value of all data objects in the dossier is recalculated and compared with the version node record value. A second repair is triggered when the verification pass rate is lower than the fifth threshold.

[0165] Simultaneously, a digital fingerprint of the dossier is generated and matched with the creation fingerprint in the original archive metadata for verification. Once the verification is successful, a timestamp certificate is added and the dossier is archived to the read-only storage area.

[0166] In one embodiment, after a complete dossier is generated, the system automatically triggers an integrity verification process: First, the SHA-256 hash value of the complete content of all data objects in the dossier (including text binary data and image binary data) is recalculated and compared one by one with the hash value recorded in the corresponding version node's structured metadata snapshot. If the verification pass rate is less than 95% (a preset fifth threshold, based on the high requirements of the archives management industry for data integrity), it is determined that there is a risk of batch data corruption, and a secondary repair process is automatically triggered (calling historical version difference data or alternative versions to complete the data). At the same time, a unique digital fingerprint is generated for the overall data of the dossier using the SHA-256 algorithm and matched with the creation fingerprint recorded in the original archive metadata. After the verification is successful, the timestamp server adds a timestamp certificate conforming to the RFC 3161 standard to the dossier and migrates the dossier to a read-only storage area (such as a CD library or cloud storage read-only bucket) for archiving, ensuring that the dossier data is tamper-proof and traceable.

[0167] This invention constructs a version sequence driven by structured metadata, extracts and standardizes original archival metadata, generates structured metadata containing version numbers, timestamps, and file associations, builds a linear timeline and tree-like branch structure, records the editing history of multiple users, and achieves systematic management of data object versions, meeting the version traceability requirements of the State Archives Administration. For heterogeneous data repair, a hierarchical strategy is designed, employing direct repair, version difference completion, and alternative version matching methods according to the degree of damage. Combined with text semantic repair, image interpolation, and OCR text layer association technologies, it solves problems such as formatting errors, pixel distortion, and broken metadata associations. The security system integrates blockchain evidence storage and dual verification, using biometrics and dynamic tokens for dual authentication, binding key permissions and case file levels, and encrypting and storing decryption logs on the blockchain in real time to ensure that audit records are tamper-proof and meet network security level protection requirements. The above technologies are integrated to construct a three-in-one system of "version management - intelligent repair - security audit," breaking through the bottlenecks of traditional solutions in terms of integrity, efficiency, and security, and providing an innovative end-to-end path for the recovery of electronic archives in the judicial and government fields, from underlying algorithms to upper-level applications.

[0168] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for recovering electronic information data from archives, characterized in that, Includes the following steps: Step S1: Obtain the original archive data, extract the first data, and generate the second data according to the format characteristics of the file extension and byte encoding rules; The second data includes hierarchy identifiers and structured metadata; Step S2 involves generating a version sequence of data objects based on the version number and timestamp of the structured data, and matching it with third-party data according to a preset strategy. Specifically, this includes: Based on the version number and modification timestamp in the structured metadata, a linear version timeline and a tree-like version branch structure are constructed. The linear version timeline is arranged in ascending order of modification timestamps to form a complete modification history axis of the data object. The tree-like version branch structure takes the initial version as the root node and forms the main branches. When a data object undergoes branch modification, a version node is generated according to the version number to record the branch source and merging relationship. The branch modification includes simultaneous editing by different users; Each version node includes a structured metadata snapshot and a content hash value, used to quickly verify the complete content consistency of the data object; Step S3: The encrypted data block that was not extracted from the first data is the fourth data, which is then decrypted, and the fifth data is generated according to the format characteristics of the file extension and byte encoding rules. If there are unextracted, corrupted, unencrypted data blocks in the first data, then supplementary repair data will be generated; Step S4: Integrate the first data, third data, fifth data and repair data according to the hierarchical identifier of the second data, and rebuild the association between the directory and metadata to generate a complete volume. The original archive data consists of two or more data objects; The first data includes text file data, image scan data, and original file metadata; The structured metadata is obtained by standardizing the format and filling in missing fields of the original archive metadata; The structured metadata includes version number, modification timestamp, and file association relationships; The version sequence is generated based on the version number and the modification timestamp; The third data is the version of a valid data object that has been verified. The third data includes the complete content of the data object and the corresponding structured metadata snapshot; The complete content of the data object includes the binary data of the text file and the binary data of the image file; The structured metadata snapshot corresponding to the complete content of the data object includes a version number, timestamp, associated file ID, and content hash value.

2. The method for recovering electronic information data of archives as described in claim 1, characterized in that: The verification specifically includes: When generating a version node, a hash value is calculated for the complete content of the data object of the current version corresponding to the node, and stored in the structured metadata snapshot; During verification, the hash value of the complete content of the actual data object stored in the current system is recalculated and compared with the hash value of the content recorded in the version node; If the hash values ​​match, it is determined that the complete content of the data object has not been tampered with or damaged. If the hash values ​​are inconsistent, it is determined that the complete content of the data object has been altered or corrupted, and a repair process is triggered.

3. The method for recovering electronic information data of archives as described in claim 2, characterized in that: The application scenarios for the verification specifically include: After repairing encrypted or unencrypted corrupted data blocks, the correctness of the repair result is verified by comparing the hash values. When matching third data according to a preset strategy, the undamaged version is selected by filtering out the hash value; When merging parallel branches in the tree-like version branch structure, content conflicts are identified by comparing hash values.

4. The method for recovering electronic information data of archives as described in claim 3, characterized in that: The preset strategy for matching third data specifically includes: A version-based priority matching strategy includes selecting the version of the data object that has passed the content hash value verification and has the latest modification timestamp. Matching strategies based on version association include filtering data object versions based on file associations and version number continuity in structured metadata; Conflict resolution strategies based on branch structure include prioritizing the main branch and merging strategies for multi-user edits; Special matching strategies based on data types, including text format feature verification and image metadata verification; Abnormal situation handling strategies, including version missing rollback and hash value conflict verification; The multi-user editing and merging strategy includes: Sort by user permission level, and prioritize the modified versions modified by users with higher permissions; If permissions are the same, merge them according to the modification timestamp, retaining the latest modification content; The abnormal situation handling strategy includes: When all candidate versions are corrupted, automatically roll back to the last available version of the data object; When there are cases where the hash values ​​are the same but the contents are different, a second verification is performed; The secondary verification verifies the data object by comparing the file size and modification timestamp.

5. The method for recovering electronic information data of archives as described in claim 4, characterized in that: Natural language processing techniques are used to identify the encoding format, paragraph structure, and font features of the text archive data. Extract resolution, color mode, file format, and OCR-recognized text layer data from scanned image data; The format features include the file extension of the original archive data, byte encoding rules, and the definition specifications of the original archive metadata fields.

6. The method for recovering electronic information data of archives as described in claim 5, characterized in that: The key management module verifies user permissions, obtains the corresponding decryption key based on the hierarchy identifier, and decrypts the fourth data, specifically including: The key management module adopts a hierarchical key system, hierarchical identifiers, and decryption keys, which are associated through access control lists. Each of the aforementioned hierarchical identifiers corresponds to key permissions; The key permissions include browsing permissions, editing permissions, and management permissions; The user permission verification adopts a dual verification mechanism, and the key acquisition process generates an operation log and stores it in encrypted form on the blockchain node; The dual verification mechanism includes biometric authentication and dynamic token verification; The biometric authentication verifies the consistency between the user's identity and the biometric information recorded in the permission database through a biometric collection device. The biometrics include fingerprint recognition, facial recognition, or iris scanning; The dynamic token verification involves the user inputting a dynamic token based on time synchronization or event triggering, and the system verifying whether the token matches the real-time token generated by the server. If the biometric authentication and dynamic token verification pass simultaneously, the decryption key acquisition process for the corresponding level identifier is triggered. An operation log is generated during the key acquisition process and stored in encrypted form via a blockchain node.

7. The method for recovering electronic information data of archives as described in claim 6, characterized in that: If there are unextracted, unencrypted, corrupted data blocks in the first data, supplementary repair data is generated. The repaired data is repaired using a repair algorithm, specifically including: A byte stream completion algorithm is used to repair text archive data; An interpolation repair algorithm is used to repair the image scan data. The repair strategy for unencrypted corrupted data blocks is tiered according to the degree of corruption, specifically including: Repair directly when the damage rate is less than the first threshold; When the corruption rate reaches the second to third threshold, historical version difference data is used for completion. When the corruption rate exceeds the fourth threshold, version sequence matching of the data object to which the currently corrupted data block belongs is triggered to obtain an alternative version.

8. The method for recovering electronic information data of archives as described in claim 7, characterized in that: During the integration process in step S4, the hierarchical identifier corresponds to the tree structure of the dossier directory, and the directory index is reconstructed using a depth-first traversal algorithm. The hierarchical identifier corresponds to the tree structure of the dossier directory, with the root node being the dossier number and the version node being the year label, organization label, and security classification label. The structured metadata association mechanism includes converting file associations in structured metadata into directory hyperlinks, and generating a unique storage path for data objects using version numbers and modification timestamps.

9. The method for recovering electronic information data of archives as described in claim 8, characterized in that: After the complete dossier is generated, the integrity verification process is automatically executed: The content hash value of all data objects in the dossier is recalculated and compared with the version node record value. A second repair is triggered when the verification pass rate is lower than the fifth threshold. Simultaneously, a digital fingerprint of the dossier is generated and matched with the creation fingerprint in the original archive metadata for verification. Once the verification is successful, a timestamp certificate is added and the dossier is archived to the read-only storage area.

Citation Information

Patent Citations

  • Block chain data management method and system

    CN119397578A

  • Archive management system based on artificial intelligence

    CN120216747A