Electronic archive data management control method and system and electronic equipment

By combining distributed storage nodes and blockchain technology with the hash value and access strategy of archival metadata, the problems of immutability and traceability in electronic record management systems are solved, realizing secure management and reliable traceability of electronic record data, and meeting the needs of complex permissions and long-term preservation.

CN121880307APending Publication Date: 2026-04-17BEIJING HESI HUIZHI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HESI HUIZHI INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing electronic records management systems suffer from several problems, including a lack of authenticity and immutability, insufficient content addressing and integrity verification, weak operation recording and traceability capabilities, high risks associated with long-term preservation and migration, lack of cross-departmental collaboration and process tracking, and a single access control model. These issues result in insufficient reliability and traceability of electronic records management.

Method used

By employing distributed storage nodes combined with blockchain technology, and using the hash value of archival metadata and access policy parameters, the electronic archival data can be made tamper-proof and efficiently traceable. By utilizing content tag values ​​and operation record policies, combined with lifecycle policies and dynamic control policies, the secure management and traceability of archival data throughout its entire lifecycle can be ensured.

Benefits of technology

It achieves the immutability and efficient traceability of electronic archive data, supports cross-departmental collaboration and complex access control, improves the reliability and auditability of archive management, and meets the stability requirements for long-term preservation and migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880307A_ABST
    Figure CN121880307A_ABST
Patent Text Reader

Abstract

The invention provides an electronic archive data management control method and system and electronic equipment, and relates to the field of archive data management control. According to the method, electronic archive data corresponding to an original archive document are stored in distributed storage nodes through archive metadata corresponding to the electronic archive data; a credible closed loop of formation-circulation-evidence storage-backup-calling-filing-destruction / long-term storage is realized; by combining the non-tampering and traceability of the block chain and the decentralization, high redundancy and tampering prevention capabilities of the content addressing storage, the management control of the whole life cycle of the electronic archive data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of archival data management and control, and in particular to an electronic archival data management and control method, system, and electronic device. Background Technology

[0002] With the accelerated digital transformation of government and enterprises, electronic records have gradually replaced paper records as the main form of documentation, and are widely used in key scenarios such as internal auditing, compliance supervision, business traceability, and judicial litigation. The current mainstream model for electronic record management is a database file system + access control platform. The core process involves uploading files and storing them on a file server, storing metadata in the database, and writing operation logs to the system log. Basic management is mainly achieved through a centralized database architecture and RBAC access control mechanism.

[0003] However, the following problems still exist in the management of electronic record data in the existing technology: Lack of authenticity and immutability: Administrators can directly modify files or clear and tamper with database logs. There is no mechanism of "recording what happens and making the records immutable". Files cannot be detected after being replaced, nor can the actual existence of files at a specific moment be verified.

[0004] Insufficient content addressing and integrity verification: Traditional path storage is used, there is no content fingerprinting mechanism, file content changes are difficult to detect, integrity cannot be verified, and there are problems of duplicate storage and chaotic multi-version management.

[0005] Weak operation record and traceability capabilities: System logs can be deleted or modified by the super administrator, making it impossible to trace internal violations, lacking non-repudiation capabilities, and making it difficult to support the chain of evidence.

[0006] High risks in long-term preservation and migration: Affected by factors such as bit corrosion, disk lifespan, hardware retirement, and file format aging, the long-term preservation stability of electronic archives is poor, and existing manual migration or vendor proprietary format solutions are unsustainable.

[0007] Lack of cross-departmental collaboration and process tracking: There are no records of the process of document approval, transfer, etc., making it impossible to identify the responsible party, and making it difficult to trace the behavior of document possession, download, and external distribution, and making it difficult to hold people accountable for unauthorized access.

[0008] The traditional RBAC mechanism cannot meet the dynamic authorization requirements under complex spatiotemporal conditions such as device, time period, region, and approval status, and cannot achieve zero-trust access control.

[0009] Insufficient support for certification and regulatory audits: The system is unable to provide functional departments with key evidence such as that the documents have not been tampered with, their existence time, transmission links, and version evolution, resulting in insufficient evidentiary value; the lack of a traceable behavioral chain makes it difficult for regulatory audits to determine situations such as unauthorized access, transmission, and destruction of archives. Summary of the Invention

[0010] In view of this, the purpose of the present invention is to provide an electronic archive data management and control method, system and electronic device. The method utilizes the electronic archive data corresponding to the original archive document and saves it to a distributed storage node through its corresponding archive metadata. It can be combined with the dynamic control strategy corresponding to blockchain to realize the management and control of the entire life cycle of electronic archive data, fundamentally realizing the immutability and efficient traceability of electronic archive data, thereby solving the above-mentioned problems mentioned in the background art.

[0011] In a first aspect, embodiments of the present invention provide an electronic archive data management and control method, the method comprising: Obtain the electronic archival data corresponding to the original archival documents, and use the electronic archival data to determine the archival metadata corresponding to the original archival documents; The content tag value corresponding to the electronic archive data is determined by using the hash value of the archive metadata, and the archive metadata is saved in the preset distributed storage node by the content tag value; Based on the distributed storage nodes, the access policy parameters corresponding to the archive metadata are determined, and the operation record policy corresponding to the distributed storage nodes is determined using the archive metadata, hash value, content tag value and access policy parameters. The stage type of the archive metadata is determined based on the rule parameters corresponding to the archive metadata under the distributed storage node. The lifecycle strategy corresponding to the archive metadata under the stage type is determined by the archive type data, confidentiality level data and archive requirement data contained in the archive metadata. The access control parameters and encryption control parameters corresponding to the distributed storage nodes are determined by using the user identity data, department and position data, business scenario data and access data corresponding to the archive metadata. The dynamic control strategy corresponding to the archive metadata is then determined by using the access control parameters and encryption control parameters. Based on operation log strategy, lifecycle strategy and dynamic control strategy, the behavior control strategy corresponding to the archive metadata is determined, and the archive metadata is controlled to flow in the distributed storage nodes through the behavior control strategy.

[0012] Optionally, the steps of obtaining electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document include: Original archive documents are obtained based on preset business system interfaces, scanner interfaces, mobile data interfaces, and department push interfaces. The parsing strategy for the original archive document is determined based on the format type parameter of the original archive document. The text extraction results and layout structure results corresponding to the original archival documents are obtained by using the parsing strategy. The electronic archival data corresponding to the original archival documents are determined by the text extraction results and layout structure results. Obtain the field data corresponding to the file fields in the original file document by using the file fields corresponding to the electronic file data; The fingerprint data corresponding to the original archival document is determined by the hash calculation result corresponding to the electronic archival data. The original archive document's metadata is determined based on field data and fingerprint data.

[0013] Optionally, the step of determining the content tag value corresponding to the electronic archive data using the hash value of the archive metadata, and saving the archive metadata in a preset distributed storage node using the content tag value, includes: Obtain the timestamp data and structure data corresponding to the archive metadata, and use the timestamp data and structure data to calculate the hash value corresponding to the archive metadata; The content identifier corresponding to the hash value is determined based on the content addressing strategy, and the content tag value corresponding to the electronic archive data is determined through the content identifier. The data writing strategy corresponding to the archive metadata is determined by the slicing strategy, erasure coding strategy, node backup strategy and encrypted storage strategy corresponding to the distributed storage nodes; The archive metadata is written and saved to the preset distributed storage nodes using content tag values ​​and in accordance with the data writing strategy.

[0014] Optionally, the step of determining the access policy parameters corresponding to the file metadata based on the distributed storage nodes, and determining the operation record policy corresponding to the distributed storage nodes using the file metadata, hash value, content tag value, and access policy parameters, includes: Obtain the blockchain corresponding to the distributed storage node, and determine the access policy parameters corresponding to the file metadata based on the operation parameters corresponding to the blockchain. The trusted fingerprint corresponding to the blockchain is determined by the archive metadata, hash value, content tag value, and access policy parameters; The time parameters corresponding to the archive metadata are determined based on the first timestamp corresponding to the blockchain and the second timestamp corresponding to the electronic archive data. The node record parameters corresponding to the archive metadata are determined by utilizing the operation parameters corresponding to the creation, modification, access, authorization, transfer, and archiving of archive metadata in the blockchain; The operation record strategy for distributed storage nodes is determined based on trusted fingerprints, time parameters, and node record parameters.

[0015] Optionally, the steps of determining the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and determining the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data, and archive requirement data contained in the archive metadata, include: Based on the formation rules, countersigning rules, circulation rules, archiving rules, retrieval rules, borrowing rules, external distribution rules, transfer rules, and destruction rules corresponding to the archive metadata under the distributed storage nodes, determine the rule parameters corresponding to the archive metadata; The target stage corresponding to the electronic archive data is determined by rule parameters, and the stage type corresponding to the target stage is obtained; among them, the target stage includes at least: formation stage, circulation and approval stage, on-chain evidence storage stage, encrypted storage stage, backup stage, periodic verification stage, archiving stage, borrowing and external distribution stage, migration and format reconstruction node, retention period expiration stage, destruction stage and preservation stage. Obtain the archive type data, security level data, and archive requirement data contained in the archive metadata, and use the stage type, archive type data, security level data, and archive requirement data to determine the target strategy corresponding to the archive metadata; among which, the target strategy includes: retention period, migration strategy, periodic verification strategy, re-backup mechanism, archiving conditions, destruction conditions, access rules, and one or more of the above strategy types. The lifecycle strategy corresponding to the archive metadata is determined based on the target strategy.

[0016] Optionally, the steps of determining the access control parameters and encryption control parameters corresponding to the distributed storage nodes based on the user identity data, department and position data, business scenario data, and access data corresponding to the file metadata, and determining the dynamic control policy corresponding to the file metadata based on the access control parameters and encryption control parameters, include: Obtain user identity data, department / position data, business scenario data, and access data corresponding to the archive metadata; use the user identity data, department / position data, business scenario data, and access data to determine the user attributes, archive attributes, environment attributes, and behavioral attributes corresponding to the distributed storage nodes. The on-chain permission policy, verification policy, encrypted access policy, and key management policy corresponding to the distributed storage node are determined based on user attributes, file attributes, environment attributes, and behavioral attributes. Access control parameters for distributed storage nodes are determined by on-chain permission and verification policies, and encryption control parameters for distributed storage nodes are determined by encryption access and key management policies. The dynamic control strategy corresponding to the archive metadata is determined by using access control parameters and encryption control parameters.

[0017] Optionally, a behavior control strategy corresponding to the archive metadata is determined based on the operation log strategy, lifecycle strategy, and dynamic control strategy. The behavior control strategy controls the steps of archive metadata flowing through distributed storage nodes, including: Use lifecycle strategies to determine the action events corresponding to archive metadata; The flow relationship diagram corresponding to action events is determined by the operation recording strategy; Based on the dynamic control strategy, obtain the flow behavior corresponding to the flow relationship diagram, and obtain the behavior control strategy corresponding to the flow behavior; Based on the flow relationship diagram, behavior control strategies are used to control the flow of archive metadata among distributed storage nodes.

[0018] Optionally, after the steps of obtaining the electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document, the method further includes: Based on the data migration parameters corresponding to the archive metadata, determine the classification and preservation strategy, format migration strategy, consistency verification strategy, and data reconstruction strategy corresponding to the archive metadata; The data update strategy corresponding to the archive metadata is determined by using classification and preservation strategies, format migration strategies, consistency verification strategies, and data reconstruction strategies; Update the archive metadata using a data update strategy.

[0019] Secondly, the present invention provides an electronic archive data management and control system, the system comprising: The archival data acquisition unit is used to acquire electronic archival data corresponding to the original archival documents and to use the electronic archival data to determine the archival metadata corresponding to the original archival documents. The storage control unit is used to determine the content tag value corresponding to the electronic archive data using the hash value of the archive metadata, and to save the archive metadata in a preset distributed storage node using the content tag value; The operation record policy determination unit is used to determine the access policy parameters corresponding to the file metadata based on the distributed storage node, and to determine the operation record policy corresponding to the distributed storage node using the file metadata, hash value, content tag value and access policy parameters. The lifecycle strategy determination unit is used to determine the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and to determine the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data and archive requirement data contained in the archive metadata. The dynamic control strategy determination unit is used to determine the access control parameters and encryption control parameters corresponding to the distributed storage nodes through the user identity data, department and position data, business scenario data and access data corresponding to the archive metadata, and to determine the dynamic control strategy corresponding to the archive metadata through the access control parameters and encryption control parameters. The archive data flow control unit is used to determine the behavior control strategy corresponding to the archive metadata based on the operation record strategy, life cycle strategy and dynamic control strategy, and to control the flow of archive metadata in the distributed storage nodes through the behavior control strategy.

[0020] Thirdly, embodiments of the present invention also provide an electronic device, which includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, and the processor executing the computer-executable instructions to implement the steps of the electronic record data management and control method provided in the first aspect.

[0021] Fourthly, embodiments of the present invention also provide a storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the steps of the electronic record data management and control method provided in the first aspect.

[0022] This invention provides an electronic archive data management and control method, system, and electronic device. In the process of managing and controlling archive data, the method first acquires the electronic archive data corresponding to the original archive document and uses this data to determine the archive metadata corresponding to the original archive document. Then, it uses the hash value of the archive metadata to determine the content tag value corresponding to the electronic archive data and saves the archive metadata in a preset distributed storage node using the content tag value. Next, it determines the access policy parameters corresponding to the archive metadata based on the distributed storage node, and uses the archive metadata, hash value, content tag value, and access policy parameters to determine the operation record policy corresponding to the distributed storage node. Then, it determines the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and determines the lifecycle policy corresponding to the archive metadata under the stage type using the archive type data, confidentiality level data, and archive requirement data contained in the archive metadata. Then, it determines the access control parameters and encryption control parameters corresponding to the distributed storage node using the user identity data, department / position data, business scenario data, and access data corresponding to the archive metadata, and determines the dynamic control policy corresponding to the archive metadata using the access control parameters and encryption control parameters. Finally, it determines the behavior control policy corresponding to the archive metadata based on the operation record policy, lifecycle policy, and dynamic control policy, and controls the flow of the archive metadata within the distributed storage node using the behavior control policy. This method utilizes the electronic archive data corresponding to the original archive documents and saves it to distributed storage nodes through its corresponding archive metadata. It can be combined with the dynamic control strategy corresponding to blockchain to realize the management and control of the entire life cycle of electronic archive data, fundamentally realizing the immutability and efficient traceability of electronic archive data.

[0023] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0026] Figure 1A flowchart of an electronic archive data management and control method provided in an embodiment of the present invention; Figure 2 A flowchart of step S101 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 3 A flowchart of step S102 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 4 A flowchart of step S103 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 5 A flowchart of step S104 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 6 A flowchart of step S105 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 7 A flowchart of step S106 in an electronic archive data management and control method provided in an embodiment of the present invention; Figure 8 A flowchart following step S101 of an electronic archive data management and control method provided in an embodiment of the present invention; Figure 9 A flowchart of another electronic archive data management and control method provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of an electronic archive data management and control system provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0027] icon: 1010 - Archive data acquisition unit; 1020 - Storage control unit; 1030 - Operation record strategy determination unit; 1040 - Lifecycle strategy determination unit; 1050 - Dynamic control strategy determination unit; 1060 - Archive data flow control unit; 101 - Processor; 102 - Memory; 103 - Bus; 104 - Communication interface. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] To facilitate understanding of this embodiment, an electronic archive data management and control method disclosed in this embodiment of the invention will first be introduced, such as... Figure 1 As shown, the method includes: Step S101: Obtain the electronic archive data corresponding to the original archive document, and use the electronic archive data to determine the archive metadata corresponding to the original archive document.

[0030] We acquire original archival documents from multiple channels, including business systems (OA, ERP, etc.), manual uploads, scanner scans, IoT device logs, and external organization pushes, and extract the corresponding electronic archival data. We perform a standardized preprocessing workflow on the electronic archival data: first, we identify the file format (Office, PDF, image, video, etc.) and perform preprocessing operations such as noise reduction, skew correction, and format parsing; then, we extract text content using multi-model combined OCR technology, and combine this with layout structure analysis to identify paragraphs, tables, seals, and other areas to generate a structured document tree; subsequently, based on the unified field standards of the master data model, we extract complete archive metadata containing basic information (title, generation time, etc.), ownership information (responsible person, signatory, etc.), security attributes (security level, anonymization level, etc.), technical attributes (hash digest, file size, etc.), and lifecycle attributes (retention period, whether to retain for a long time, etc.); finally, we detect sensitive information using an entity recognition model and perform anonymization or encryption processing, while simultaneously calculating the file's SHA-256 / Blake3 hash digest to provide a basis for subsequent verification.

[0031] Step S102: Use the hash value of the archive metadata to determine the content tag value corresponding to the electronic archive data, and save the archive metadata in the preset distributed storage node through the content tag value.

[0032] Based on the hash digest of the archive metadata, a unique Content Identifier (CID) is generated by combining the file content, metadata structure, and timestamp. This identifier serves as the unique "content fingerprint" of the electronic archive. The archive package, containing the original text, multiple format derivatives (such as PDF / A), structured metadata, and hash digest, is written to a pre-defined distributed storage node (such as IPFS or Ceph cluster) using Content Addressed Storage (CAS). During storage, file fragmentation and Reed-Solomon erasure coding are automatically performed. After generating checksum fragments, redundant backups are implemented across physical data centers, multiple regions, or multiple cloud vendors. All original archive text is encrypted using AES-256-GCM symmetric encryption to ensure storage security. Simultaneously, the system periodically performs CID recalculation and comparison, fragment health scans, and other operations to ensure data integrity.

[0033] Step S103: Determine the access policy parameters corresponding to the file metadata based on the distributed storage nodes, and determine the operation record policy corresponding to the distributed storage nodes using the file metadata, hash value, content tag value and access policy parameters.

[0034] Based on the deployment architecture of distributed storage nodes, the confidentiality level of archives, and the requirements of business scenarios, the access policy parameters, such as the access subject, permission scope, and operation restrictions, corresponding to the archive metadata are clearly defined. Archive metadata, hash digests, content identifiers (CIDs), and access policy parameters are synchronously written into the consortium blockchain to construct an operation record policy: for the entire process of archive creation, modification, access, authorization, transfer, and archiving, an immutable record is generated containing the operation subject (user ID, department, device fingerprint), operation time, and operation type; a dual timestamp mechanism of blockchain system timestamps and external trusted time sources (TSPs) is introduced to generate time credentials; and a version chain is established through the previous version ID to achieve full-link traceability of the archive evolution process.

[0035] Step S104: Determine the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and determine the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data and archive requirement data contained in the archive metadata.

[0036] Based on the flow status of archival metadata, the system determines the stage type of the archive, including its formation, flow approval, blockchain-based evidence storage, encrypted storage, archiving, and destruction. Using a rules engine, and integrating archival type, security level, and business requirements data from the archival metadata, a targeted lifecycle strategy is generated: specifying the archival retention period (e.g., 30 years for financial archives, 10 years for contract archives), migration strategies (e.g., migrating Office documents to PDF / A-3 format), periodic verification frequency (e.g., hash verification of classified documents every 3 months), backup mechanisms, automatic archiving conditions, and destruction approval processes. Simultaneously, it integrates with a task scheduling framework to achieve automated execution and dynamic adjustment of the strategy.

[0037] Step S105: Determine the access control parameters and encryption control parameters corresponding to the distributed storage nodes through the user identity data, department and position data, business scenario data and access data corresponding to the archive metadata, and determine the dynamic control strategy corresponding to the archive metadata through the access control parameters and encryption control parameters.

[0038] Extract user identity data, department / position data, business scenario data, and access data such as access time, location, and device from the archive metadata to construct a multi-dimensional attribute system. Based on this system, determine access control parameters and encryption control parameters: adopt the ABAC (Attribute-Based Access Control) model, integrating four dimensions of attributes—user, archive, environment, and behavior—to achieve dynamic authorization; embed a zero-trust architecture, implementing a "default denial, continuous verification" mechanism, requiring identity re-verification and behavioral anomaly analysis for each access; encryption control uses the AES-256-GCM algorithm, implementing key sharing and short-term key issuance (e.g., 10-minute validity) through blockchain smart contracts, ensuring that keys are never directly accessed by administrators. The final result is a dynamic control strategy encompassing dynamic authorization, encrypted access, collaborative authorization (sensitive archives require approval from at least two people), and on-chain permission storage.

[0039] Step S106: Determine the behavior control strategy corresponding to the archive metadata based on the operation record strategy, lifecycle strategy and dynamic control strategy, and control the flow of archive metadata in the distributed storage nodes through the behavior control strategy.

[0040] This strategy integrates operation log strategies, lifecycle strategies, and dynamic control strategies to construct a behavior control strategy covering the entire process of document circulation. It uses a graph database to build a circulation graph, with document version nodes, department nodes, and personnel nodes as the core, and operations such as "submit to" and "approved" as related edges, enabling visual tracking of document approval, transfer, countersigning, and external distribution. Specifically, it embeds an AI intelligent audit module and uses models such as LSTM and graph neural networks to identify abnormal behaviors such as unauthorized access and concentrated external distribution late at night, triggering automatic alarms. All circulation behaviors are synchronously written to the blockchain to form an immutable audit log, supporting access control and compliance checks. It also outputs risk scores and compliance reports, ensuring that the entire circulation of documents in distributed storage nodes is controllable, searchable, and auditable.

[0041] Optionally, step S101 involves obtaining the electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document, such as... Figure 2 As shown, it includes: Step S201: Obtain the original archive documents according to the preset business system interface, scanner interface, mobile terminal data interface and department push interface.

[0042] The system acquires original archival documents through a pre-defined standardized interface system, covering data sources across all scenarios: it connects with business system interfaces such as OA, ERP, HR, and production systems to automatically synchronize approval documents, contract PDFs, resumes, and other documents generated by these systems; it accesses scanned copies of paper archives through scanner interfaces to achieve automatic import of physical archives into digital form; it supports users to manually upload Word, PDF, and image files via web and mobile devices through mobile data interfaces; and it receives archival data pushed by internal departments, external partners, and relevant regulatory departments through departmental push interfaces, while also being compatible with log-type machine data generated by production equipment and IoT devices, achieving one-stop collection of archives from multiple channels.

[0043] Step S202: Determine the parsing strategy corresponding to the original archive document based on the format type parameter of the original archive document.

[0044] First, identify the format type parameters of the original archive documents to determine the file category (e.g., Office documents, PDFs, JPG / PNG images, multi-page TIFFs, videos, audios, CAD files, etc.). Based on the format characteristics, develop targeted parsing strategies: For document files (Office, PDF), use the OpenXML parsing engine and PDFBox tool for text extraction and layer parsing, respectively; for image files, use the OpenCV algorithm for noise reduction, skew correction, and grayscale preprocessing; for video archives, configure keyframe extraction rules for subsequent retrieval; for special formats such as multi-page TIFFs, develop parsing logic for pagination and reordering to ensure that different file formats can be adapted to the subsequent processing flow.

[0045] Step S203: Use the parsing strategy to obtain the text extraction results and layout structure results corresponding to the original archive document, and determine the electronic archive data corresponding to the original archive document through the text extraction results and layout structure results.

[0046] This system employs multi-model OCR technology to extract text from scanned documents and image-based archives, encompassing dedicated OCR models for Chinese printed text, handwritten text, and tables to accurately extract text content. A layout structure analysis module identifies paragraphs, tables, seals, and signature areas within the document, constructing a structured document tree (DOM). For archives containing handwritten annotations, a natural language processing model identifies annotation images and converts them into structured information. For non-text archives (such as videos and audio), key feature information (such as video keyframes and audio duration) is extracted according to preset rules. Finally, the text extraction results, layout structure information, original document content, and feature data are integrated to form complete electronic archive data.

[0047] Step S204: Obtain the field data corresponding to the file fields in the original file document through the file fields corresponding to the electronic file data.

[0048] To address the scattered archival fields in electronic archival data, an Archival Master Data Management (ADM) model is adopted to unify and standardize these fields, resolving inconsistencies in field naming across different business systems (e.g., unifying "Document Title / Manuscript Title / Document Name" into "Archival Title"). Various fields are precisely extracted from the electronic archival data, including basic information fields (title, generation time, business type, department), ownership information fields (responsible person, issuer, countersigners), security attribute fields (security classification, anonymization level), and lifecycle-related fields (retention period requirements, whether long-term storage is required), ensuring the standardization and completeness of the field data.

[0049] Step S205: Determine the fingerprint data corresponding to the original document through the hash calculation result corresponding to the electronic document data.

[0050] The original content of the original archival document undergoes cryptographic hash calculation using secure hash algorithms such as SHA-256 or Blake3 to generate a unique and irreversible hash digest, which serves as the fingerprint data of the original archival document. This fingerprint data will act as a unique identifier for the file content, not only for subsequent detection of whether the file has been tampered with, but also to support functions such as multi-version conflict detection and duplicate file identification, laying the foundation for trusted management of electronic archives.

[0051] Step S206: Determine the archive metadata corresponding to the original archive document based on field data and fingerprint data.

[0052] The standardized field data extracted in step S204 is integrated with the fingerprint data (hash digest) generated in step S205, and key information such as technical attribute fields (version number, file size, format identifier, CID placeholder) and sensitive information processing identifiers (e.g., whether it is de-identified, encryption type) are added to form complete structured archive metadata. The archive metadata covers five core categories: basic information, ownership information, security attributes, technical attributes, and lifecycle attributes. It serves as the core data support for subsequent blockchain notarization, distributed storage, and access control, ensuring that the metadata can comprehensively reflect the core characteristics and management needs of the archive.

[0053] Optionally, the content tag value corresponding to the electronic archive data is determined using the hash value of the archive metadata, and the archive metadata is saved in a preset distributed storage node using the content tag value, as in step S102. Figure 3 As shown, it includes: Step S301: Obtain the timestamp data and structure data corresponding to the archive metadata, and use the timestamp data and structure data to calculate the hash value corresponding to the archive metadata.

[0054] Two core auxiliary data types are obtained corresponding to the archival metadata: first, timestamp data, including the system creation time of the archive and the authoritative timestamp obtained from an external trusted time source (TSP); second, metadata structure data, namely the field organization, field types, and relationships of the archival metadata (such as the hierarchical structure of fields like basic information, ownership information, and security attributes). Combining the original content of the archival metadata, secure hash algorithms such as SHA-256 or Blake3 are used to perform encrypted hash calculations on the combined data of "file content + metadata structure + timestamp," generating a unique and irreversible hash value. This hash value serves as the core verification basis for the integrity of the archival metadata and also provides basic data support for the subsequent generation of content identifiers.

[0055] Step S302: Determine the content identifier corresponding to the hash value based on the content addressing strategy, and determine the content tag value corresponding to the electronic archive data through the content identifier.

[0056] Based on the Content Addressed Storage (CAS) strategy, a unique Content Identifier (CID) is generated for the electronic archive data, using the hash value generated in step S301 as the core basis. This CID serves as the content tag value for the electronic archive. Unlike traditional path-based storage, this content tag value is directly associated with the actual content of the archive and possesses three core characteristics: the same content always corresponds to the same CID, enabling rapid deduplication of duplicate archives and reducing storage redundancy; if any tampering occurs with the archive content, the hash value will change synchronously, resulting in a change in the CID, allowing for immediate detection of tampering; and when sharing across systems and departments, there is no need to rely on file paths, as the archive can be located solely through the CID, adapting to multi-party collaboration scenarios. Simultaneously, the CID will serve as the link between on-chain evidence storage and off-chain storage, providing identification support for subsequent archive location and verification.

[0057] Step S303: Determine the data writing strategy corresponding to the file metadata through the slicing strategy, erasure coding strategy, node backup strategy and encrypted storage strategy corresponding to the distributed storage nodes.

[0058] By integrating the underlying architectural features of distributed storage nodes with the requirements for secure file storage, a data writing strategy comprising four core tactics is constructed: Slicing strategy: Use a slicing algorithm with fixed or adaptive size to slice the archive metadata and corresponding original text into several independent data slices; Erasure coding strategy: Based on the Reed-Solomon(k,m) scheme, the fragmented data is encoded to generate m parity fragments, ensuring that the original text can be completely recovered by retaining any k fragments from the k+m data fragments, achieving storage reliability of >99.999%; Node backup strategy: Define the distributed storage rules for sharded data and parity shards, requiring data shards to be deployed across multiple physical data centers, different regions, or multiple cloud vendors (such as local IDC + private cloud + government cloud) to form a redundant backup architecture with multiple regions and multiple nodes, and configure a node health check mechanism to monitor node availability in real time. Encryption storage strategy: The AES-256-GCM symmetric encryption algorithm is used to encrypt the archive metadata and original text. Key management adopts a combination of KMS (Key Management System) and blockchain smart contract to ensure that the key is secure and can only be obtained through the authorized process.

[0059] Step S304: Use the content tag value and follow the data writing strategy to write and save the file metadata to the preset distributed storage node.

[0060] Using the Content Identifier (CID) as the unique index identifier for the archive in the distributed storage system, and following the data writing strategy determined in step S303, the encrypted archive metadata, corresponding shard data, and checksum shards are written to preset distributed storage nodes (such as IPFS, Ceph clusters, or enterprise internal distributed clusters). During the writing process, the storage node location of each shard, the mapping relationship between the CID and the node, and the data are automatically recorded and synchronized to the blockchain for evidence storage. After storage is completed, a periodic maintenance mechanism is initiated: data integrity is verified through operations such as CID recalculation and comparison, shard health scanning, and erasure coding verification; if node anomalies or shard corruption are detected, the shard reconstruction and migration process is automatically triggered to migrate the data to healthy nodes, ensuring the long-term stable preservation of archive metadata in the distributed storage environment.

[0061] Optionally, step S103, which determines the access policy parameters corresponding to the file metadata based on the distributed storage nodes and uses the file metadata, hash value, content tag value, and access policy parameters to determine the operation record policy corresponding to the distributed storage nodes, is as follows: Figure 4 As shown, it includes: Step S401: Obtain the blockchain corresponding to the distributed storage node, and determine the access policy parameters corresponding to the file metadata based on the operation parameters corresponding to the blockchain.

[0062] Acquire a consortium blockchain (jointly maintained by multiple departments / institutions, including nodes from archives, business departments, audit departments, and notary offices) deployed in collaboration with distributed storage nodes. Based on the core operational parameters of the blockchain, derive access policy parameters for the archive metadata. These blockchain operational parameters include node type division of labor (e.g., archive nodes responsible for supervision, audit nodes for read-only monitoring), high-throughput consensus algorithms such as PBFT / PoA, and cross-node data synchronization rules. Access policy parameters must be clearly defined in conjunction with the archive's security classification and business needs: permission levels (read / write / external distribution / download), usage period limits, usage scenario constraints (e.g., internal collaboration only / external distribution to partners), collaborative authorization requirements (e.g., classified archives require dual approval), and access behavior restrictions (e.g., prohibiting printing / allowing only online browsing). All parameters must be registered and stored on the blockchain to ensure that subsequent modifications require on-chain approval, guaranteeing audit transparency.

[0063] Step S402: Determine the trusted fingerprint corresponding to the blockchain using the archive metadata, hash value, content tag value, and access policy parameters.

[0064] By integrating archival metadata, hash values, content identifiers (CIDs), and access policy parameters, a "trusted fingerprint" that can be stored on the blockchain is generated. Specifically, this includes: the MetadataHash corresponding to the archival metadata (ensuring fields have not been tampered with), the hash value (SHA-256 / Blake3) of the electronic archival content, the content identifier (CID, used to locate distributed storage files off-chain), and the hash digest of the access policy parameters. It also includes supplementary key information such as the operation type (creation / modification / access, etc.) and the operation subject (user ID, department, device fingerprint). The trusted fingerprint is organized using a MerkleTree data structure, supporting both rapid verification of single archives and efficient processing for batch certification scenarios, avoiding the storage of sensitive original text on the blockchain, thus balancing security and verification efficiency.

[0065] Step S403: Determine the time parameters corresponding to the archive metadata based on the first timestamp corresponding to the blockchain and the second timestamp corresponding to the electronic archive data.

[0066] The first timestamp corresponding to the blockchain and the second timestamp corresponding to the electronic archival data are linked and integrated to form a time parameter with legal validity. The first timestamp is the system time when the blockchain transaction is written into the block header, recording the on-chain sequence of the operation. The second timestamp is an authoritative timestamp obtained from an external Trusted Time Source (TSP). By sending a document digest to the TSP service and obtaining signature evidence, the objectivity and unforgeability of the time are ensured. This dual timestamp mechanism mutually corroborates each other, clarifying the temporal relationship of archival operations and providing legal proof that "the archive truly existed at a certain moment," meeting the requirements of judicial proceedings and regulatory audits for time validity.

[0067] Step S404: Determine the node record parameters corresponding to the archive metadata using the operation parameters corresponding to the creation, modification, access, authorization, transfer, and archiving of archive metadata in the blockchain.

[0068] Based on a blockchain multi-node consensus mechanism, node record parameters are extracted for key operations (creation, modification, access, authorization, transfer, and archiving) throughout the entire lifecycle of archive metadata. Specifically, these include: the blockchain transaction ID (TxID) for each operation, the identity identifier of the operating entity (user ID, department code, device fingerprint), the approval process record triggered by the operation (approver, approval opinion, approval time), the version association information corresponding to the operation (previous version TxID), and the operation result status (success / failure / abnormal). Node record parameters must be synchronously verified through multiple nodes on the consortium blockchain to ensure that no single node can tamper with them, achieving "operation traceability, accountability, and non-repudiation," providing core evidence for internal violation tracing and judicial evidence collection.

[0069] Step S405: Determine the operation record strategy corresponding to the distributed storage node based on the trusted fingerprint, time parameter, and node record parameter.

[0070] By integrating the trusted fingerprint from step S402, the dual timestamp time parameters from step S403, and the node record parameters from step S404, an operation record strategy corresponding to the distributed storage nodes is constructed. This strategy comprises three core modules: first, an "immutable operation timeline," which strings together all operation records chronologically to form a trusted trajectory throughout the entire lifecycle of the archive; second, a "version chain evolution mechanism," which associates each operation node with the previous version's TxID, supporting historical version backtracking and modification difference comparison; and third, "multi-node consensus verification rules," which rely on the consortium blockchain's PBFT / PoA consensus algorithm to ensure that operation records require confirmation from a majority of nodes to take effect, preventing individual departments or nodes from tampering with them. Simultaneously, the strategy incorporates privacy protection mechanisms, employing hash substitution for the original text and encrypted metadata to prevent the leakage of sensitive content while ensuring operational traceability, thus balancing auditability and data security.

[0071] Optionally, step S104 involves determining the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and determining the lifecycle strategy corresponding to the archive metadata under the stage type using the archive type data, confidentiality level data, and archive requirement data contained in the archive metadata. Figure 5 As shown, it includes: Step S501: Determine the rule parameters corresponding to the archive metadata based on the formation rules, countersigning rules, circulation rules, archiving rules, retrieval rules, borrowing rules, external distribution rules, transfer rules, and destruction rules corresponding to the archive metadata under the distributed storage node.

[0072] By integrating the control requirements of all key stages in the entire lifecycle of electronic records, and extracting the rule parameters corresponding to the record metadata, these rule parameters need to be deeply adapted to industry regulations and internal management systems, specifically covering nine core rule categories: Formation rules: standard format for file creation (e.g., PDF / A-3), required metadata fields (e.g., responsible person, security classification), and requirements for de-identifying sensitive information; Joint signing rules: Participating entities, approval order, signature confirmation process, and archiving requirements for joint signing results among multiple departments; Flow rules: Requirements for cross-departmental transfer of permissions, flow time limits, and node handover records; Archiving rules: archiving trigger conditions (such as approval completion, business closure), archiving storage location (such as distributed storage cold nodes), and archiving metadata integrity verification standards; Invocation rules: Invocation permission verification dimensions, post-invocation operation restrictions (such as prohibiting secondary forwarding), and invocation record retention requirements; Borrowing rules: borrowing duration, renewal process, overdue reminder mechanism, and document protection requirements during the borrowing period; External release rules: approval levels for external release (e.g., sensitive files require dual authorization from department head and file administrator), anonymization standards for external release content, and qualification verification requirements for external release recipients; Transfer rules: Data integrity verification when transferring data to external organizations or archiving departments, transfer link records, and rules for adjusting permissions after transfer; Destruction rules: Destruction trigger conditions (such as the expiration of the storage period), destruction approval process (requires double review), data backup requirements before destruction, and rules for uploading destruction records to the blockchain.

[0073] By standardizing and organizing the above rules, structured archival metadata rule parameters are formed, providing a basis for determining the stage type.

[0074] Step S502: Determine the target stage corresponding to the electronic archive data through rule parameters, and obtain the stage type corresponding to the target stage; wherein, the target stage includes at least: formation stage, circulation and approval stage, on-chain evidence storage stage, encrypted storage stage, backup stage, periodic verification stage, archiving stage, borrowing and external distribution stage, migration and format reconstruction node, retention period expiration stage, destruction stage and preservation stage.

[0075] Based on the rule parameters constructed in step S501, and combined with the current circulation status of the archive metadata and the business scenario, the system automatically matches the target stage corresponding to the electronic archive and clarifies its stage type. The target stages cover key aspects of the entire archive lifecycle, and the types and core characteristics of each stage are as follows: Formation stage: The archive creation and structured processing stage, the core of which is to complete the collection of original documents, format parsing, metadata extraction and sensitive information processing; The circulation and approval stage: The process of submitting, countersigning and approving files among multiple departments needs to be recorded, including the processing behavior and results at each stage. On-chain evidence storage stage: Key information such as file metadata, hash digest, and CID are written into the consortium blockchain to generate a trusted timestamp and version chain; Encryption storage stage: The original file is encrypted using the AES-256-GCM algorithm, stored in a distributed storage node, and redundant backup is performed on multiple nodes; Backup phase: Regular backups are performed based on erasure coding technology to ensure rapid recovery in case of data corruption; Periodic verification phase: Perform hash consistency verification, CID verification, and storage node health checks according to a preset cycle; Archiving phase: After the archiving conditions are met, the files are migrated to long-term storage nodes and write protection settings are executed; Borrowing and distribution phase: Provide borrowing services to internal users or distribute to external institutions in compliance with regulations, and record the entire operation process; Migration and format reconstruction phase: For long-term archives, complete format upgrades (such as Office to PDF / A-3, CAD to DWF) and cross-storage media migration; Retention period expires: When the archives reach the preset retention period, the destruction assessment or long-term preservation determination process is triggered. Destruction phase: After compliance approval, the archive data is completely deleted, and the destruction record is permanently stored on the blockchain as evidence; Long-term preservation stage: For legally mandated long-term or permanent archives, implement multi-dimensional protection (anti-corrosion, encryption algorithm upgrades) and regular maintenance.

[0076] Step S503: Obtain the archive type data, security level data, and archive requirement data contained in the archive metadata, and determine the target strategy corresponding to the archive metadata using the stage type, archive type data, security level data, and archive requirement data; wherein, the target strategy includes: retention period, migration strategy, periodic verification strategy, re-backup mechanism, archiving conditions, destruction conditions, access rules, or one or more of the above strategy types.

[0077] First, three types of core decision-making data are extracted from the archive metadata: File type data: Clearly define the category to which the file belongs (e.g., contracts, financial documents, classified documents, personnel files, etc.); Security classification data: Defines the security level of files (e.g., public, internal sensitive, secret, confidential, top secret). Data on archival requirements includes business value, frequency of use (cold / hot archives), retention period requirements, and cross-departmental collaboration needs.

[0078] Based on the stage type determined in step S502, the corresponding target strategy is automatically matched using Drools or a self-developed rule engine. The target strategy must align with the file characteristics and stage requirements, specifically including: Retention period: 10 years for contract files, 30 years for financial documents, and classified files shall be handled in accordance with relevant regulations. There is no time limit for permanently retained files (such as major project data). Migration strategy: Regular files undergo format migration periodically, while long-term stored files undergo additional cross-storage media migration; Regular verification strategy: Public files are verified once a year, sensitive files every 6 months, and classified files every 3 months for hash consistency testing; Backup mechanism: Hot files are backed up once a month, cold files are backed up once a quarter, and backup data is stored across regions; Archiving criteria: closed-loop business process, all approvals passed, complete metadata with no risk of sensitive information leakage; Destruction conditions: The retention period has expired, the case has been approved and confirmed by two people, there are no outstanding related business transactions, and a backup has been completed before destruction; Access rules: Set access scope (e.g., financial files are only accessible to the Finance Department and Audit Department), operation permissions (read / write / download / print), and environmental restrictions (office equipment, designated regions, working hours) according to the confidentiality level and stage type.

[0079] Depending on the actual situation of the archives, one or more of the above-mentioned strategy types can be flexibly combined to form a targeted strategy.

[0080] Step S504: Determine the lifecycle strategy corresponding to the archive metadata based on the target strategy.

[0081] Based on the target strategy determined in step S503, and combining a task scheduling framework (such as Quartz or Celery) with a strategy versioning management mechanism, a complete archive metadata lifecycle strategy is constructed. This strategy has three core characteristics: Automated execution: The actions such as periodic verification, format migration, backup, and expiration reminder in the target strategy are transformed into automated tasks, which are executed automatically according to preset cycles or trigger conditions. All operation records are written to the blockchain in real time to ensure auditability. Dynamic adaptation: Supports versioned policy management. When industry regulations or internal systems change (such as adjustments to retention periods), newly generated files will be executed according to the new policy, while archived files can retain the old policy or be upgraded as needed. Policy change records are stored on the chain. Collaborative Integration: Deeply integrated with access control strategies. For example, when a file is in the approval stage, only the approver has access rights; once it enters the archiving stage, write protection is automatically enabled, prohibiting modification operations; when the retention period is about to expire, a reminder is automatically pushed to the administrator, triggering the decision process for destruction or long-term preservation.

[0082] The resulting lifecycle strategy enables standardized, automated, and traceable management of the entire process from the creation of archives to their destruction / long-term preservation, avoiding inconsistencies and compliance risks associated with manual management.

[0083] Optionally, step S105 involves determining the access control parameters and encryption control parameters corresponding to the distributed storage nodes based on the user identity data, department / position data, business scenario data, and access data corresponding to the archive metadata, and determining the dynamic control strategy corresponding to the archive metadata based on the access control parameters and encryption control parameters. Figure 6 As shown, it includes: Step S601: Obtain user identity data, department / position data, business scenario data, and access data corresponding to the file metadata, and use the user identity data, department / position data, business scenario data, and access data to determine the user attributes, file attributes, environment attributes, and behavioral attributes corresponding to the distributed storage nodes.

[0084] The system acquires four core data categories associated with the archive metadata and extracts four core attributes of the distributed storage nodes based on this data, forming the basis for access control and encryption decisions. User identity data includes user ID, name, job level, organization, and identity authentication status (such as whether real-name authentication is required), which are used to construct corresponding user attributes (identity, position, job level, department). Departmental and job data: This includes the responsibilities of the user's department, the scope of job authority, cross-departmental collaboration permissions, etc., supplementing and improving the job-related dimensions of the user's attributes, and also linking the ownership information of the file's department; Business scenario data: This includes the business type corresponding to the file (such as contract approval, financial reimbursement, project archiving), business process stage (such as approval in progress, closed loop, archiving in progress), and purpose of use (such as internal collaboration, external audit, judicial evidence), and corresponding file attributes (security level, file type, version, business scenario). Access data includes access time (e.g., office hours / office hours), access location (e.g., office address / remote location), access device (e.g., corporate office computer / personal device, device fingerprint), and access action (e.g., reading, downloading, printing, sending, modifying), which are respectively constructed with environmental attributes (access time, access location, device type) and behavioral attributes (access action type, operation frequency, data interaction scope).

[0085] By accurately mapping data to attributes, a comprehensive attribute system covering "people, files, environment, and behavior" is formed, providing underlying support for subsequent strategy formulation.

[0086] Step S602: Determine the on-chain permission policy, verification policy, encrypted access policy, and key management policy corresponding to the distributed storage node based on user attributes, file attributes, environment attributes, and behavioral attributes.

[0087] Based on the four-dimensional core attributes constructed in step S601, the four core strategies corresponding to the distributed storage nodes are determined: On-chain permission policy: Match permission levels (read / write / external / download / print) based on user attributes and file attributes, clearly define the authorizer, authorized person, validity period of permission, and usage scenario restrictions (such as only available in internal collaboration scenarios), support cross-departmental collaborative authorization rules (such as classified files requiring dual authorization from the file administrator and department leader), and all permission rules must be registered and stored on the consortium blockchain; Verification strategy: Based on environmental and behavioral attributes, a zero-trust architecture is implemented, and "default denial and continuous verification" rules are set. Each access requires re-verification of identity (such as password + secondary verification). User and entity behavior analysis (UEBA) is used to determine whether the access behavior is abnormal (such as accessing high-security files from a different location outside of office hours or downloading files in batches). Abnormal behavior triggers additional verification or direct denial. Encryption access strategy: Determine the encryption strength based on file attributes (such as security level), use the AES-256-GCM symmetric encryption algorithm to encrypt and store the original file and metadata, specify the encryption scope (such as encrypting sensitive fields separately or encrypting the whole text), and set security requirements for the encryption link (such as SSL encryption during transmission). Key management strategy: Key allocation rules are designed based on user attributes and permission levels. A key management model of "KMS (Key Management System) + Blockchain Smart Contract" is adopted to clarify the whole process of key generation, distribution, use and invalidation, limit the validity period of keys (such as one-time keys, 10-minute short keys), and prohibit administrators from directly accessing keys.

[0088] Step S603: Determine the access control parameters corresponding to the distributed storage node through the on-chain permission policy and verification policy, and determine the encryption control parameters corresponding to the distributed storage node through the encryption access policy and key management policy.

[0089] The four core strategies determined in step S602 are further broken down, and the corresponding access control parameters and encryption control parameters are extracted respectively: Based on on-chain permission and verification policies, access control parameters are extracted, including dynamic authorization matching rules (such as "high-secret files are only allowed to be accessed by users of the same or higher job level and corresponding business department"), zero-trust verification trigger conditions (such as access from different locations or access during non-office hours), abnormal behavior judgment thresholds (such as triggering an alarm by downloading more than 50 files in a single day), collaborative authorization approval nodes (such as the specific role configuration for dual authorization), access log recording dimensions (user ID, device fingerprint, operation time, CID association), and permission change approval process parameters. Based on the encrypted access policy and key management policy, the following encrypted control parameters are extracted: including encryption algorithm type (AES-256-GCM), encryption granularity (full-text encryption / field-level encryption), key generation algorithm, key validity period, key distribution trigger conditions (such as automatic distribution after permission verification), key eviction rules (such as immediate invalidation after access ends, automatic eviction upon permission expiration), and encrypted metadata storage specifications (only the encryption policy hash is stored on the chain, and sensitive key information is not stored).

[0090] Both types of parameters need to be linked with the blockchain evidence storage mechanism, and parameter configuration and change records should be uploaded to the blockchain in real time to ensure auditability and traceability.

[0091] Step S604: Determine the dynamic control strategy corresponding to the file metadata using access control parameters and encryption control parameters.

[0092] By integrating access control parameters and encryption control parameters, a dynamic control strategy for file metadata covering the entire process of "authorization-access-encryption-auditing" is constructed, which includes five core policy categories: Dynamic authorization strategy: Based on the ABAC model and access control parameters, it realizes the automated process of "attribute matching-permission determination-authorization" and supports dynamic adjustment of permissions according to business scenarios and environmental conditions (such as only the approver can access during the approval stage, and only the authorized queryer can read after archiving). Encrypted access strategy: Based on encryption control parameters, the original file is stored in encryption. When a user accesses the file, they need to pass the authorization verification. The smart contract will then issue a short-term key for decryption. The decryption behavior is recorded on the blockchain in real time to prevent the risk of key leakage. Zero-trust authentication strategy: Authentication rules are embedded in access control parameters. Identity verification, environment verification, and behavior verification are performed on every access. Even on the corporate intranet, it is not trusted by default, and it continuously prevents unauthorized access. On-chain audit strategy: All permission changes, access operations, and key usage behaviors are synchronously written to the blockchain to form an immutable audit log according to the access control parameters, supporting permission retrospection and compliance checks; Collaborative authorization strategy: For highly sensitive files, a multi-person approval process is triggered according to the collaborative authorization rules in the access control parameters. Access permissions and decryption keys are only issued after the approval record is uploaded to the blockchain to ensure access compliance.

[0093] This dynamic control strategy achieves deep integration of permissions and encryption, which not only meets the fine-grained management needs of complex government and enterprise scenarios, but also ensures the credibility and traceability of strategy execution through blockchain evidence storage.

[0094] Optionally, step S106 involves determining the behavior control strategy corresponding to the archive metadata based on the operation log strategy, lifecycle strategy, and dynamic control strategy, and controlling the flow of archive metadata in the distributed storage nodes through the behavior control strategy. Figure 7 As shown, it includes: Step S701: Use the lifecycle strategy to determine the action events corresponding to the file metadata.

[0095] Based on the lifecycle strategy corresponding to the archival metadata, and combined with its current stage (such as creation, circulation approval, archiving, long-term preservation, etc.), standardized action events covering the entire archival lifecycle are generated. These action events must strictly adhere to the automated scheduling rules and compliance requirements outlined in the strategy, specifically including: Formation stage: Format recognition, OCR text extraction, metadata standardization, sensitive information de-identification, hash digest generation, etc. The circulation and approval stage includes actions such as submitting for approval, co-signing and collaboration, approval / rejection, and departmental handover. Storage and backup phase: file fragmentation, erasure coding, multi-node redundancy backup, AES-256-GCM encrypted storage, etc. Long-term storage phase: Regular hash consistency checks, format migration (such as Office to PDF / A-3), cross-media migration, anti-corrosion repair, etc. Calling and certification stage: actions such as permission verification, short-term key acquisition, file decryption, and evidence chain report generation; Archiving and destruction phase: actions such as write protection settings, archive migration, destruction approval, complete data deletion, and uploading destruction records to the blockchain.

[0096] All action events are associated with corresponding execution conditions (such as "the archiving action is triggered after approval") and responsible entities, providing a foundation for the construction of the workflow diagram.

[0097] Step S702: Determine the flow relationship diagram corresponding to the action event through the operation record strategy.

[0098] Based on the immutable on-chain records (including operation subject, timestamp, version TxID, etc.) in the operation record strategy, and with action events as the core, a graph of archive flow relationships is constructed using a graph database (such as Neo4j). The core components of this graph are as follows: Node design includes: document version node (associated with CID, hash digest, and version number), department node (recording the business departments, audit departments, etc. involved in the process), personnel node (associated with user ID, position, and device fingerprint), and action event node (clearly defining the action type and execution result). Edge relationship definition: Edges are formed by flow operations such as "submit to", "approved", "transmitted to", and "authorized access", which connect each node and mark key information such as operation timestamp and on-chain transaction ID; Version chain integration: By connecting file nodes of different versions through "previous version TxID", a complete version evolution chain is formed, clearly showing the flow trajectory of file modification and iteration.

[0099] The flow relationship graph needs to be synchronized with the on-chain operation records in real time to ensure that the graph is consistent with the actual flow behavior and to provide support for visual tracking.

[0100] Step S703: Obtain the flow behavior corresponding to the flow relationship diagram according to the dynamic control strategy, and obtain the behavior control strategy corresponding to the flow behavior.

[0101] Combining ABAC dynamic authorization, zero-trust authentication, encrypted access, and intelligent auditing rules in the dynamic control strategy, this approach identifies various flow behaviors (such as cross-departmental transfer, external distribution, download, and batch access) from the flow relationship diagram and extracts targeted behavior control strategies. The core strategies include four main approaches: Permission verification strategy: For each type of transfer behavior, dynamic permission matching is performed based on user attributes, file attributes, and environment attributes (e.g., the external transfer of high-security files requires dual authorization from "file administrator + department leader"), and transfer behaviors that do not meet the permission requirements are rejected; Anomaly detection strategy: Embed an AI anomaly detection model based on LSTM and graph neural network (GNN) to identify unauthorized circulation behavior (such as batch downloading from different locations outside of office hours, low-level personnel accessing high-level files, and abnormal circulation paths), and trigger automatic alarms; Encryption control strategy: For transactions involving external transmission and cross-departmental transfer, sensitive information is de-identified, files are encrypted during transmission, and short-term keys are issued through blockchain smart contracts to ensure data security during the transfer process; Auditing and Traceability Strategy: Clearly define the on-chain recording requirements for all flow activities, including the operating entity, action details, timestamps, device information, etc., to generate tamper-proof audit logs and support access control and compliance checks.

[0102] Step S704: Based on the flow relationship diagram, control the flow of archive metadata in the distributed storage nodes through behavior control strategies.

[0103] Using the constructed flow relationship diagram as a visual support, behavior control strategies are employed to implement end-to-end management of the flow of archive metadata across distributed storage nodes: Real-time permission verification: When a workflow is triggered, the permission rules in the behavior control policy are automatically matched, and the verification is completed through the ABAC model and zero-trust verification mechanism. Only after the verification is passed can the next workflow be executed. Abnormal Behavior Interception: The AI ​​anomaly detection module monitors the flow trajectory in real time. If it identifies any violations (such as unauthorized transmission or abnormal download), it immediately interrupts the flow process and pushes an alarm message to the administrator. On-chain synchronous recording: All circulation activities (including operation results, responsible parties, and timestamps) are written to the consortium blockchain in real time, updating the circulation relationship graph and forming a closed loop of "circulation activities - on-chain recording - graph update" to ensure full traceability; Compliance visualization and supervision: Through a full-process visualization console, information such as the document flow path, processing time of each department, bottleneck nodes, and list of responsible persons is displayed. At the same time, risk scores and compliance reports are automatically generated to help managers keep track of the flow status in real time and ensure that the entire document flow is controllable, compliant, and auditable.

[0104] Optionally, after step S101, which involves obtaining the electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document, the following steps are performed: Figure 8 As shown, the method also includes: Step S801: Determine the classification and preservation strategy, format migration strategy, consistency verification strategy, and data reconstruction strategy corresponding to the archive metadata based on the data migration parameters corresponding to the archive metadata.

[0105] Based on data migration parameters associated with archival metadata (including retention period, security classification, current storage media status, format type, business value, risk level, etc.), four core strategies are formulated: Categorized Preservation Strategy: Archives are classified into regular archives (3-10 years), legally long-term preservation archives (≥30 years), and permanent preservation archives (such as major project data) based on their retention period, and each is matched with a differentiated storage solution; regular archives are stored on distributed hot nodes for easy retrieval, while long-term / permanent archives adopt multi-media redundant storage of "local IDC + government cloud + optical storage" and are configured with a higher priority health check mechanism; classified storage is implemented according to the security level, and classified archives must be encrypted and the access range of storage nodes is restricted.

[0106] Format migration strategy: Standardized migration rules are formulated for different original formats to ensure long-term readability; Office documents are migrated to PDF / A-3 format, images are converted to TIFFG4 or PNG lossless format, CAD files are migrated to PDF / A-3 or DWF format, and media files are converted to H.265MP4 format; old and unpopular formats (such as old PPT and Flash files) are given priority for migration, and the entire migration process is recorded on the chain to ensure auditability.

[0107] Consistency verification strategy: Set periodic verification rules. Regular files are recalculated and CID verified once a year. Long-term / permanent files are verified once every 3 to 6 months. Classified files are additionally checked for fragment health and erasure coding every 3 months. The verification content includes file hash consistency, metadata field integrity, and storage node redundancy compliance. The verification results are written to the blockchain for evidence storage in real time.

[0108] Data reconstruction strategy: Based on the Reed-Solomon erasure coding (k,m) scheme, when storage node damage, data fragment loss, or bit erosion (BitRot) is detected, any k valid data fragments are automatically called to reconstruct the complete file; for files with inconsistent hash verification, historical trusted data is traced back through the blockchain version chain, and data repair is completed by combining multi-node backup; the reconstruction process does not require manual intervention, and the reconstruction result must pass the consistency verification again before it can take effect.

[0109] Step S802: Determine the data update strategy corresponding to the archive metadata using the classification and preservation strategy, format migration strategy, consistency verification strategy, and data reconstruction strategy.

[0110] This strategy integrates the core requirements of classification and storage strategies, format migration strategies, consistency verification strategies, and data reconstruction strategies. It combines task scheduling frameworks (such as Quartz and Celery) with strategy versioning management mechanisms to construct a full-process data update strategy, which includes: Update trigger conditions are divided into two categories: periodic triggers (based on the preset cycle of the consistency verification strategy) and event triggers (storage media retirement, format facing obsolescence risk, data corruption detected, and policy changes leading to adjustments in the storage strategy), ensuring that update actions accurately respond to needs.

[0111] Update execution logic: Execute in the order of "format migration - category storage adjustment - consistency verification - data reconstruction (in case of anomalies)". After format migration, the format identifier and CID in the archive metadata are updated synchronously; category storage adjustment requires synchronous update of the topology mapping relationship of distributed storage nodes; if the consistency verification fails, the data reconstruction process is started immediately.

[0112] Version and permission coordination: After data is updated, a new version number is generated, and historical versions are associated with the "previous version TxID" to build a complete version chain; the updated archive metadata, new CID, and migration / reconstruction records are synchronously written to the consortium blockchain to ensure that the update process is tamper-proof; if the update involves a change in security level, the permission policy is synchronously updated and access control rules are reconfigured.

[0113] Anomaly handling mechanism: When format migration fails, it automatically rolls back to the original version and triggers an alarm; when data reconstruction fails to meet consistency requirements, it initiates multi-node data comparison and prioritizes the use of trusted data stored on the chain for secondary reconstruction; all anomalies and processing results are recorded on the chain for easy traceability and investigation.

[0114] Step S803: Update the archive metadata using the data update strategy.

[0115] According to the established data update strategy, automated update operations are performed on the archive metadata and associated original data: Trigger detection and preparation: Real-time monitoring of update trigger conditions. When the periodic or event trigger rules are met, the target file's metadata, current storage status, and on-chain evidence records are automatically extracted to form an update task package.

[0116] Format migration and storage adjustment: Complete file format conversion according to the format migration strategy and generate multi-format derived versions; according to the classification and preservation strategy, migrate the updated archive data to the target storage node, complete the cross-media or cross-node storage adjustment, and synchronously update the storage node information digest to the blockchain.

[0117] Consistency verification and data repair: After migration and storage adjustments, execute the consistency verification strategy, recalculate the file SHA-256 / Blake3 hash and CID, and verify whether the metadata fields are complete; if the verification passes, update the hash digest, CID, format identifier, version number and other fields in the metadata; if the verification fails, immediately start the data reconstruction strategy, complete the data repair and then execute the verification again until the standard is met.

[0118] On-chain synchronization and archiving: After the update is completed, the new archive metadata, update operation records (operation subject, timestamp, update type, comparison of before and after versions), and verification results are written to the consortium blockchain to extend the version chain; an update report is generated and synchronized to the archive management console for administrators to view; write protection is applied to the updated archive to prevent unauthorized modification and ensure the credibility and integrity of the updated data.

[0119] like Figure 9 The flowchart of another electronic record data management and control method shown mainly includes the following steps: Step S901: File collection.

[0120] As the starting point of the process, it enables unified access to original archival documents from multiple channels: it synchronizes approval documents and contracts generated by the system through business system interfaces (OA, ERP, etc.); it receives scanned copies of paper archives through scanner interfaces; it supports users to manually upload Word, PDF and other files through mobile / web interfaces; and it receives archival data from internal and external institutions through department push interfaces, covering the entire scenario of system generation + physical digitization + manual uploading, providing original materials for subsequent processing.

[0121] Step S902: Format recognition / OCR / metadata extraction.

[0122] Standardized preprocessing was performed on the collected original archives: First, identify the file format (Office, PDF, image, video, etc.) and match the corresponding parsing strategy; It uses multi-model OCR technology (dedicated models for printed text / handwritten text / tables) to extract text content, while also analyzing the page layout (paragraphs, tables, and stamp areas). Based on the master data model of archives, structured archive metadata containing basic information (title, generation time), ownership information (responsible person, signatory), security attributes (classification), and technical attributes (hash digest) is extracted, and a hash fingerprint (such as SHA-256) of the file is generated to provide a basis for subsequent integrity verification.

[0123] Step S903: Content-addressable storage.

[0124] Abandoning the traditional path-based storage model, a content-addressable storage (CAS) mechanism is adopted: using the hash value of the file content as the core, a unique content identifier (CID) is generated, binding the file to the content rather than the path; the same content corresponds to the same CID, which not only solves the problem of not being able to detect file content tampering but file name changes, but also quickly identifies duplicate files and reduces storage redundancy.

[0125] Step S904: CID generation / fragmentation / erasure coding / redundancy.

[0126] After generating a unique CID based on the archive's metadata hash value, timestamp, and structure data, distributed storage preprocessing is performed on the archive: The archive is divided into data fragments according to adaptive rules; The Reed-Solomon erasure coding algorithm is used to generate check fragments to ensure that the original text can still be completely recovered even if some fragments are lost. Redundant backups of shards and verification shards are performed across distributed storage nodes in multiple physical data centers and cloud vendors. At the same time, AES-256-GCM encryption is applied to the original file to ensure data security and reliability at the storage layer.

[0127] Step S905: Blockchain-based evidence storage.

[0128] Write the core information of the archive (metadata, CID, hash digest, trusted timestamp) into the consortium blockchain: Synchronously access an external trusted time source (TSP) to generate authoritative timestamps, ensuring that the archive truly exists at a certain moment; Record the creation, modification and other operation traces of the file, and build an immutable version chain through the previous version ID; By relying on the multi-node consensus mechanism of the consortium blockchain, the non-repudiation of operation records is achieved, providing a credible foundation for subsequent traceability and certification.

[0129] Step S906: Lifecycle strategy generation.

[0130] Based on the type of archives, their security classification, and business needs, generate personalized full lifecycle management strategies: Based on rules and parameters such as formation, circulation, and archiving, archives are classified into stages (formation, approval, archiving, etc.). Clearly define the control requirements for each stage: such as retaining contract files for 10 years, performing hash verification on classified files every 3 months, and migrating Office documents to PDF / A format, etc. This strategy supports dynamic adaptation (such as adjustments when policies change) and serves as the basis for control of subsequent processes (the upward arrows in the process correspond to the iterative updates of the strategy).

[0131] Step S907: Access Control / Encrypted Access.

[0132] Based on four dimensions of attributes—user, profile, environment, and behavior—fine-grained dynamic control is achieved. The ABAC (Attribute-Based Access Control) model is adopted, combined with a zero-trust mechanism, to verify identity, environment (office hours / device), and behavior (operation type) for each access. When accessing files, short-term keys (e.g., valid for 10 minutes) are issued via blockchain smart contracts, and AES-256-GCM encryption is used for transmission and storage, which satisfies complex permission scenarios while ensuring data security.

[0133] Step S908: File transfer / approval.

[0134] Supporting collaborative processing of archives across multiple departments: Files are submitted, countersigned, and approved between departments according to business needs, and permission policies are automatically matched in the process (e.g., classified files can only be accessed by the approver). Record the person in charge, time, and opinions at each stage to solve the problem of unclear responsibility in traditional cross-departmental collaboration.

[0135] Step S909: Flow Map / Intelligent Audit.

[0136] Achieving visualization and compliance supervision of the circulation process: A flow graph is constructed using a graph database (nodes represent archives, departments, and personnel, and edges represent operations such as submission and approval) to visually present the archive flow trajectory. An AI-powered intelligent audit module is embedded to identify abnormal behaviors such as unauthorized access and bulk downloads late at night, triggering real-time alerts; all audit logs are synchronously written to the blockchain, providing traceable evidence for compliance checks.

[0137] Step S910: Long-term preservation and format migration.

[0138] Addressing the risks of long-term storage of traditional electronic records: Perform format migrations (such as Office to PDF / A-3) and cross-storage media migrations periodically according to lifecycle strategies; Periodically perform hash consistency checks and storage node health checks. If bit erosion or fragment damage is detected, automatically reconstruct the data using erasure coding algorithms to ensure that the archives are readable and intact for a long time.

[0139] Step S911: Retrieve / Issue Certificate / Authenticity Verification.

[0140] To meet the certificate issuance needs in judicial, notary, and other scenarios: When it is necessary to retrieve files or provide evidence, the CID and hash digest of the blockchain are used to verify that the files have not been tampered with. Based on on-chain timestamps and operation records, an evidence chain report is generated, which includes the existence time of the file, the circulation path, and the version evolution. This directly supports judicial litigation and notarized evidence collection, and solves the pain point of weak traditional evidence issuance capabilities.

[0141] Step S912: Destroy or store for a long period of time.

[0142] As a final step in the process, the archives are either legally terminated or continuously maintained. If the retention period for the archives expires, the destruction operation will be carried out after approval by two people, and the destruction record will be permanently written into the blockchain. For permanently stored archives, a long-term preservation strategy (regular format migration and consistency verification) is continuously implemented to ensure the long-term reliable storage of the archives.

[0143] The results of destruction or long-term storage ultimately flow to the lifecycle policy S906, reflecting the dynamic adaptability of the policy; feedback from long-term storage, auditing and other links will drive the iterative optimization of the lifecycle policy, realizing closed-loop management of the entire process.

[0144] As can be seen from the above-described electronic record data management and control method, this method has the following technical effects: Guarantee of document authenticity: Blockchain immutable records + content hash verification, achieving credibility from technology to system. Any tampering of documents can be detected, and cannot be hidden by internal or external personnel.

[0145] Transparent and traceable throughout the entire lifecycle: All events are recorded on the blockchain, providing a complete archive timeline that can be used for auditing, supervision, and investigation.

[0146] Enhanced long-term archival preservation capabilities: Content-addressable storage + multi-node redundancy + format migration + one-to-one consistency verification shifts long-term archival preservation from relying on experience to relying on technology, with preservation periods reaching 30–50 years or more.

[0147] Reduce the risk of data leakage: ABAC dynamic access control, short-lived keys, and zero-trust architecture significantly reduce unauthorized access and internal leaks.

[0148] Improved efficiency of cross-departmental collaboration: The document flow map enables users to find responsible persons, check status, and visualize processes, significantly reducing communication costs.

[0149] Strong judicial evidence-issuing capabilities: It can output on-chain evidence chain reports to provide legally valid proof of authenticity and completeness, thus giving electronic records legal effect.

[0150] Improved management automation: Automatic lifecycle scheduling, automatic verification, automatic migration, and automatic alarms reduce the workload of record management personnel by 50%–70%.

[0151] By utilizing the electronic archival data corresponding to the original archival documents and saving it to distributed storage nodes through its corresponding archival metadata, a trusted closed loop of "formation-circulation-evidence-backup-retrieval-archiving-destruction / long-term preservation" can be achieved. Combined with the immutability and traceability of blockchain and the decentralized, highly redundant, and tamper-proof capabilities of content-addressed storage, the management and control of electronic archival data throughout its entire lifecycle can be realized.

[0152] Corresponding to the above embodiments of the electronic archive data management and control method, this invention also provides an electronic archive data management and control system, such as... Figure 10 As shown, the system includes: The archive data acquisition unit 1010 is used to acquire electronic archive data corresponding to the original archive document and use the electronic archive data to determine the archive metadata corresponding to the original archive document. The storage control unit 1020 is used to determine the content tag value corresponding to the electronic archive data using the hash value of the archive metadata, and to save the archive metadata in a preset distributed storage node using the content tag value; The operation record policy determination unit 1030 is used to determine the access policy parameters corresponding to the file metadata based on the distributed storage node, and to determine the operation record policy corresponding to the distributed storage node using the file metadata, hash value, content tag value and access policy parameters. The lifecycle strategy determination unit 1040 is used to determine the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and to determine the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data and archive requirement data contained in the archive metadata. The dynamic control strategy determination unit 1050 is used to determine the access control parameters and encryption control parameters corresponding to the distributed storage node through the user identity data, department and position data, business scenario data and access data corresponding to the archive metadata, and to determine the dynamic control strategy corresponding to the archive metadata through the access control parameters and encryption control parameters. The archive data flow control unit 1060 is used to determine the behavior control strategy corresponding to the archive metadata based on the operation record strategy, life cycle strategy and dynamic control strategy, and control the flow of archive metadata in the distributed storage nodes through the behavior control strategy.

[0153] As can be seen from the above-mentioned electronic archive data management and control system, the system uses the electronic archive data corresponding to the original archive documents and saves it to the distributed storage nodes through its corresponding archive metadata, realizing a trusted closed loop of "formation-circulation-evidence-backup-retrieval-archiving-destruction / long-term preservation"; it can combine the immutability and traceability of blockchain with the decentralized, highly redundant and tamper-proof capabilities of content-addressed storage to realize the management and control of the entire life cycle of electronic archive data.

[0154] The electronic archive data management and control system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned electronic archive data management and control method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned electronic archive data management and control method embodiment.

[0155] This embodiment also provides an electronic device, the structural schematic diagram of which is shown below. Figure 11 As shown, the device includes a processor 101 and a memory 102; wherein, the memory 102 is used to store one or more computer instructions, which are executed by the processor to implement the steps of the above-described electronic record data management and control method.

[0156] Figure 11 The electronic device shown also includes a bus 103 and a communication interface 104, with the processor 101, communication interface 104 and memory 102 connected via the bus 103.

[0157] The memory 102 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. The bus 103 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0158] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and to send encapsulated IPv4 packets or IPv4 packets to the user terminal through the network interface.

[0159] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 102, and processor 101 reads the information in memory 102 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0160] This invention also provides a storage medium storing a computer program, which, when executed by a processor, performs the steps of the electronic archive data management and control method described in the foregoing embodiments.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, devices, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0164] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0165] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An electronic file data management control method characterized by, The method includes: Obtain electronic archive data corresponding to the original archive document, and use the electronic archive data to determine the archive metadata corresponding to the original archive document; The content tag value corresponding to the electronic archive data is determined by using the hash value of the archive metadata, and the archive metadata is saved in a preset distributed storage node using the content tag value; Based on the distributed storage node, the access policy parameters corresponding to the file metadata are determined, and the operation record policy corresponding to the distributed storage node is determined using the file metadata, the hash value, the content tag value, and the access policy parameters; The stage type of the archive metadata is determined based on the rule parameters corresponding to the archive metadata under the distributed storage node, and the lifecycle strategy corresponding to the archive metadata under the stage type is determined through the archive type data, confidentiality level data and archive requirement data contained in the archive metadata. The access control parameters and encryption control parameters corresponding to the distributed storage node are determined by the user identity data, department and position data, business scenario data and access data corresponding to the archive metadata, and the dynamic control strategy corresponding to the archive metadata is determined by the access control parameters and the encryption control parameters. Based on the operation record strategy, the lifecycle strategy, and the dynamic control strategy, a behavior control strategy corresponding to the archive metadata is determined, and the archive metadata is controlled to flow in the distributed storage nodes through the behavior control strategy.

2. The electronic archival data management control method of claim 1, wherein, The steps of obtaining electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document include: Original archive documents are obtained based on preset business system interfaces, scanner interfaces, mobile data interfaces, and department push interfaces. The parsing strategy corresponding to the original archive document is determined based on the format type parameter of the original archive document; The parsing strategy is used to obtain the text extraction results and layout structure results corresponding to the original archive document, and the electronic archive data corresponding to the original archive document is determined by the text extraction results and the layout structure results. Obtain the field data corresponding to the archive field in the original archive document by using the archive field corresponding to the electronic archive data; The fingerprint data corresponding to the original document is determined by the hash calculation result corresponding to the electronic document data. Based on the field data and the fingerprint data, the archive metadata corresponding to the original archive document is determined.

3. The electronic archival data management control method of claim 1, wherein, The steps of determining the content tag value corresponding to the electronic archive data using the hash value of the archive metadata, and saving the archive metadata in a preset distributed storage node using the content tag value, include: Obtain the timestamp data and structure data corresponding to the archive metadata, and use the timestamp data and structure data to calculate the hash value corresponding to the archive metadata; The content identifier corresponding to the hash value is determined based on the content addressing strategy, and the content tag value corresponding to the electronic archive data is determined through the content identifier. The data writing strategy corresponding to the archive metadata is determined by the slicing strategy, erasure coding strategy, node backup strategy, and encrypted storage strategy corresponding to the distributed storage nodes; The file metadata is written and saved to a preset distributed storage node using the content tag value and in accordance with the data writing strategy.

4. The electronic archival data management control method of claim 1, wherein, The steps of determining the access policy parameters corresponding to the file metadata based on the distributed storage node, and determining the operation record policy corresponding to the distributed storage node using the file metadata, the hash value, the content tag value, and the access policy parameters, include: Obtain the blockchain corresponding to the distributed storage node, and determine the access policy parameters corresponding to the file metadata based on the operation parameters corresponding to the blockchain. The trusted fingerprint corresponding to the blockchain is determined by the file metadata, the hash value, the content tag value, and the access policy parameters; The time parameter corresponding to the archive metadata is determined based on the first timestamp corresponding to the blockchain and the second timestamp corresponding to the electronic archive data. The node record parameters corresponding to the archive metadata are determined by using the operation parameters corresponding to the creation, modification, access, authorization, transfer, and archiving of the archive metadata in the blockchain; The operation record strategy corresponding to the distributed storage node is determined based on the trusted fingerprint, the time parameter, and the node record parameter.

5. The electronic archival data management control method of claim 1, wherein, The steps of determining the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and determining the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data, and archive requirement data contained in the archive metadata, include: Based on the formation rules, countersigning rules, circulation rules, archiving rules, retrieval rules, borrowing rules, external distribution rules, transfer rules, and destruction rules corresponding to the archive metadata under the distributed storage node, determine the rule parameters corresponding to the archive metadata; The target stage corresponding to the electronic archive data is determined by the rule parameters, and the stage type corresponding to the target stage is obtained; wherein, the target stage includes at least: formation stage, circulation and approval stage, on-chain evidence storage stage, encrypted storage stage, backup stage, periodic verification stage, archiving stage, borrowing and external distribution stage, migration and format reconstruction node, retention period expiration stage, destruction stage and preservation stage. The archive type data, security level data, and archive requirement data contained in the archive metadata are obtained. The target strategy corresponding to the archive metadata is determined using the stage type, the archive type data, the security level data, and the archive requirement data. The target strategy includes one or more of the following strategy types: retention period, migration strategy, periodic verification strategy, re-backup mechanism, archiving conditions, destruction conditions, and access rules. The lifecycle strategy corresponding to the archive metadata is determined based on the target strategy.

6. The electronic archival data management control method of claim 1, wherein, The steps of determining the access control parameters and encryption control parameters corresponding to the distributed storage node based on the user identity data, department / position data, business scenario data, and access data corresponding to the archive metadata, and determining the dynamic control strategy corresponding to the archive metadata based on the access control parameters and the encryption control parameters, include: Obtain user identity data, department / position data, business scenario data, and access data corresponding to the file metadata; use the user identity data, department / position data, business scenario data, and access data to determine the user attributes, file attributes, environmental attributes, and behavioral attributes corresponding to the distributed storage node. Based on the user attributes, the file attributes, the environment attributes, and the behavior attributes, determine the on-chain permission policy, verification policy, encrypted access policy, and key management policy corresponding to the distributed storage node; The access control parameters corresponding to the distributed storage node are determined by the on-chain permission policy and the verification policy, and the encryption control parameters corresponding to the distributed storage node are determined by the encryption access policy and the key management policy. The dynamic control strategy corresponding to the file metadata is determined using the access control parameters and the encryption control parameters.

7. The electronic archival data management control method of claim 1, wherein, The steps of determining the behavior control policy corresponding to the archive metadata based on the operation record policy, the lifecycle policy, and the dynamic control policy, and controlling the flow of the archive metadata in the distributed storage nodes through the behavior control policy, include: The lifecycle strategy is used to determine the action events corresponding to the archive metadata; The operation recording strategy is used to determine the flow relationship diagram corresponding to the action event; The flow behavior corresponding to the flow relationship diagram is obtained according to the dynamic control strategy, and the behavior control strategy corresponding to the flow behavior is obtained. Based on the aforementioned flow relationship diagram, the behavior control strategy controls the flow of the archive metadata among the distributed storage nodes.

8. The electronic archive data management and control method according to claim 1, characterized in that, After the steps of obtaining the electronic archival data corresponding to the original archival document and using the electronic archival data to determine the archival metadata corresponding to the original archival document, the method further includes: Based on the data migration parameters corresponding to the archive metadata, determine the classification and preservation strategy, format migration strategy, consistency verification strategy, and data reconstruction strategy corresponding to the archive metadata; The data update strategy corresponding to the archive metadata is determined using the classification and preservation strategy, the format migration strategy, the consistency verification strategy, and the data reconstruction strategy. The archive metadata is updated using the data update strategy.

9. An electronic archive data management and control system, characterized in that, The system includes: An archive data acquisition unit is used to acquire electronic archive data corresponding to the original archive document and use the electronic archive data to determine the archive metadata corresponding to the original archive document. The storage control unit is used to determine the content tag value corresponding to the electronic archive data using the hash value of the archive metadata, and to save the archive metadata in a preset distributed storage node using the content tag value; An operation record policy determination unit is used to determine the access policy parameters corresponding to the file metadata based on the distributed storage node, and to determine the operation record policy corresponding to the distributed storage node using the file metadata, the hash value, the content tag value, and the access policy parameters; The lifecycle strategy determination unit is used to determine the stage type of the archive metadata based on the rule parameters corresponding to the archive metadata under the distributed storage node, and to determine the lifecycle strategy corresponding to the archive metadata under the stage type through the archive type data, confidentiality level data and archive requirement data contained in the archive metadata; The dynamic control strategy determination unit is used to determine the access control parameters and encryption control parameters corresponding to the distributed storage node through the user identity data, department and position data, business scenario data and access data corresponding to the file metadata, and to determine the dynamic control strategy corresponding to the file metadata through the access control parameters and encryption control parameters; The archive data flow control unit is used to determine the behavior control strategy corresponding to the archive metadata based on the operation record strategy, the life cycle strategy and the dynamic control strategy, and control the flow of the archive metadata in the distributed storage node through the behavior control strategy.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the steps of the electronic record data management and control method according to any one of claims 1 to 8.