A blockchain-based digital archive full-process credible management platform
By building a digital archive management platform based on the improved Merkle mountain algorithm, the problems of broken archive version chains and unverifiable status in cross-organizational collaboration were solved. It achieved a unified and reliable expression and full-process verification of archive content, metadata, version and events, thereby improving the credibility of archive management and system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI AIJI TECH CO LTD
- Filing Date
- 2025-11-25
- Publication Date
- 2026-04-28
AI Technical Summary
Existing digital record management systems cannot provide a unified and reliable technical foundation in cross-organizational collaboration scenarios, resulting in broken record version chains, unverifiable status, scattered event records, and inconsistent results across multiple organizations. Traditional Merkle trees have weak dynamic data processing capabilities, high computational costs, and difficulty in achieving unified confirmation and consistency verification of cross-organizational record results.
A trusted management platform for the entire digital archive process based on the improved Merkle mountain algorithm is constructed. Through state information serialization processing, improved Merkle mountain structure and time anchor nodes, a sustainable and expandable multi-peak structure is generated, cross-organizational conflict fields and organizational identifiers are recorded to form a traceable cross-organizational structure, and a cumulative proof chain is generated and written to the blockchain through a delay function.
It enables unified and reliable representation and full-process verification of digital archives in multi-organizational, multi-stage, and high-frequency update environments, ensuring the authenticity, integrity, and traceability of archives, and improving the credibility and system performance of archive management.
Smart Images

Figure CN121561982B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of blockchain data structure and digital archive management technology, and in particular to a blockchain-based trusted management platform for the entire process of digital archives. Background Technology
[0002] With the increasing digitalization of government services and the online availability of archival resources, the use of digital archives in cross-departmental, cross-regional, and cross-organizational collaborative processing has grown significantly. The requirements for authenticity, integrity, and traceability are becoming increasingly stringent throughout the entire process of archive generation, transfer, archiving, utilization, and multi-entity processing. However, existing digital archive management systems generally rely on traditional database structures or centralized log recording methods to maintain archive content, metadata, and processing records, failing to provide a unified and reliable technical foundation for complex business chains involving multiple organizations. In scenarios where multiple entities independently process the same archive, differences in processing time sequence, processing granularity, and recording standards can easily lead to gaps, loss, or conflicts in the archive version chain, making it difficult to fully reconstruct the archive history and verify the consistency of archives across organizations.
[0003] Existing archival evidence preservation methods based on hash chains or ordinary Merkle tree structures are typically only suitable for single-organization scenarios. When synchronizing, comparing, or merging archival data across organizations, they cannot effectively handle differences in multiple root nodes arising from different organizations, nor can they record the source of these differences at the structural level. This makes it difficult to form a unified and reliable archival chain structure when multiple organizations jointly process or transfer archival content. Furthermore, traditional Merkle trees have weak dynamic data processing capabilities. When archives frequently undergo version updates or processing events, the tree structure needs to be frequently rebuilt, resulting in high computational costs and impacting system performance. Simultaneously, traditional evidence preservation methods lack a unified structured representation of multi-dimensional information such as archival content, metadata, event identifiers, and operator identities, making cross-system data integration difficult and hindering the construction of an archival state representation model covering the entire lifecycle.
[0004] In cross-organizational collaboration scenarios, due to differences in archival processing rules, system environments, and operational sequences among different organizations, the trusted structures generated for the same digital archive may be inconsistent. Traditional technologies lack a structured mechanism that can explicitly record the source of discrepancies, the discrepancy data, and the rules for resolving discrepancies, making it difficult to achieve unified confirmation of cross-organizational archival results. Furthermore, existing solutions lack methods for structurally representing cross-organizational discrepancies and writing them into a unified trusted structure, failing to ensure consistent discrepancy tracing capabilities in subsequent state verification. Additionally, traditional archival time recording methods mostly rely on database timestamps, lacking strong immutability and failing to support the construction of a chronological chain for each event in archival processing. They also cannot structurally describe the relationship between time anchoring, processing nodes, and archival status, affecting the reliable recovery capability of the archival timeline.
[0005] Therefore, how to provide a blockchain-based digital archive end-to-end trusted management platform is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a blockchain-based end-to-end trusted management platform for digital archives. This invention constructs a trusted archive structure based on an improved Merkle mountain algorithm, unifying the serialization of archive content, metadata, versions, events, and operator information to generate a sustainably expandable multi-peak structure. In cross-organizational scenarios, it introduces cross-organizational conflict fields to record differing root nodes and organizational identifiers, forming a traceable cross-organizational structure. By constructing an event sequence chain through time-anchored nodes and generating a cumulative proof chain using a delay function, it writes this chain to the blockchain, thereby achieving end-to-end verification of the authenticity, integrity, and cross-organizational consistency of digital archives.
[0007] According to an embodiment of the present invention, a blockchain-based end-to-end trusted management platform for digital archives includes:
[0008] The status information generation module is used to parse the file body, file header information, additional description information and processing event records of digital archives, and generate a status information sequence in a fixed order;
[0009] The leaf node construction module is used to segment, complete, and encode the state information sequence, generate state-embedded leaf nodes, and add them to the underlying node set of the improved Merkle mountain algorithm.
[0010] The node tree generation module is used to perform node pairing, splicing, encoding alignment and node construction on the underlying node set according to the improved Merkle mountain algorithm, and generate an improved Merkle mountain structure containing multiple mountain root nodes;
[0011] The version reference node module is used to obtain the root node identifier of the previous version and the node sequence identifier of the updated leaf node when the digital archive version is updated, perform field concatenation and encoding processing, generate version reference nodes and form an archive version chain structure.
[0012] The cross-organizational node processing module is used to collect the root nodes of the mountain peaks of different organizations, form a cross-organizational root node set, and perform node construction based on the cross-organizational root node set to generate a second-layer improved Merkle mountain structure.
[0013] The time anchor node module is used to obtain the event time information and the node sequence number of the corresponding leaf node when an event occurs, generate time anchor nodes, and form a time sequence chain in chronological order.
[0014] The cumulative proof generation module is used to obtain the root node field of the improved Merkle mountain structure, perform delayed function operations and construct cumulative proofs, write them into the blockchain and form a cumulative proof chain in the order of generation.
[0015] Optionally, modules can be integrated using the following methods:
[0016] S1. Perform data parsing on digital archives to generate archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier, and combine them in a preset order to form a status information sequence;
[0017] S2. Construct state-embedded leaf nodes based on the state information sequence, and write the state-embedded leaf nodes into the bottom node set of the improved Merkle mountain algorithm;
[0018] S3. Generate a node tree according to the mountain structure of the improved Merkle mountain algorithm, and perform iterative combination operations on the bottom node set to form an improved Merkle mountain structure containing multiple mountain root nodes.
[0019] S4. When the digital archive version is updated, a version reference node is generated and written into the improved Merkle mountain structure. The archive version chain structure is determined based on the version reference node.
[0020] S5. When processing archives across organizations, collect the root nodes of the improved Merkle mountain structure of multiple organizations, generate a cross-organization root node set, and construct a second-layer improved Merkle mountain structure based on the cross-organization root node set.
[0021] S6. When the processing event of the digital archive is recorded, a time anchor node is generated, the time anchor node is written into the corresponding improved Merkle mountain structure, and a time sequence chain is constructed based on the time anchor node.
[0022] S7. Combine the root node of the improved Merkle mountain structure with the output of the delay function to form a cumulative proof, and write the cumulative proof into the blockchain to generate a cumulative proof chain.
[0023] Optionally, S1 specifically includes:
[0024] S11. Perform data parsing on the digital archive's file body, extract text content, structured field content, and embedded object content data, and generate an archive content summary according to preset encoding rules;
[0025] S12. Parse the header information, additional description information and archive directory information of the digital archive, and generate an archive metadata digest according to the parsed information and preset encoding rules.
[0026] S13. Obtain the file number information and version sequence information of the digital file, and combine the file number information and version sequence information according to the preset field order to generate a file version identifier;
[0027] S14. Extract operation type information, operation time information and operation initiating terminal identification information from the recorded file processing operations, and combine them according to the preset field order to generate file processing event identifiers.
[0028] S15. Obtain the operator account identifier and organization identifier of the operator performing the processing operation, and generate the file operator identity identifier by combining them according to the preset field order;
[0029] S16. The archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier are concatenated in a fixed order and serialization processing is performed according to a preset encoding format to generate a status information sequence.
[0030] Optionally, S2 specifically includes:
[0031] S21. The state information sequence is segmented according to a preset length, and the last segment that is less than the preset length is padded according to the padding rule to obtain the data block to be processed.
[0032] S22. Connect the data block to be processed with the node type identifier and the node sequence number identifier in a fixed order to form leaf node input data; the node type identifier indicates that the leaf node input data corresponds to the bottom node of the improved Merkle mountain structure; the node sequence number identifier indicates the position of the leaf node input data in the bottom node set;
[0033] S23. Perform encoding processing on the leaf node input data. The encoding processing includes character set conversion processing and byte length alignment processing to generate leaf node encoded data.
[0034] S24. Generate state-embedded leaf nodes based on the leaf node encoding data, add the state-embedded leaf nodes to the bottom node set of the improved Merkle mountain algorithm, and determine the storage order of the state-embedded leaf nodes in the bottom node set according to the node sequence number identifier.
[0035] Optionally, S3 specifically includes:
[0036] S31. Obtain the state-embedded leaf nodes arranged by node number in the bottom node set, and use the state-embedded leaf nodes as the initial node set.
[0037] S32. The initial node set is combined in an adjacent pairing manner, and every two adjacent leaf nodes are spliced together in a fixed connection order to generate intermediate node input data.
[0038] S33. Perform encoding alignment processing on the intermediate node input data, and use the processed intermediate node input data as the basis for constructing intermediate nodes to generate corresponding intermediate nodes.
[0039] S34. Arrange the intermediate nodes in ascending order according to the node sequence number to form an intermediate node set, and determine whether to form the root node of the current level based on the number of nodes in the intermediate node set.
[0040] S35. If no mountain root node is formed in the current level, the set of intermediate nodes is taken as a new set of nodes to be processed, and the adjacent pairing, splicing and encoding processes are repeatedly performed on the new set of nodes to be processed to generate the set of intermediate nodes of the next level.
[0041] S36. If a mountain root node is formed at the current level, add the mountain root node to the list of mountain root nodes of the improved Merkle mountain structure, and continue to build a new mountain structure starting from the next unprocessed leaf node.
[0042] S37. After all the underlying node sets have been processed, an improved Merkle mountain structure containing multiple mountain root nodes is generated.
[0043] Optionally, S4 specifically includes:
[0044] S41. When a digital archive is updated, obtain the root node identifier of the previous version of the digital archive before the update, and use the root node identifier of the previous version as version association information.
[0045] S42. Obtain the node sequence number identifier of the state embedded leaf node corresponding to the state information sequence generated by the updated digital archive, and use the node sequence number identifier as the update version identifier.
[0046] S43. Connect the version association information with the updated version identifier according to the preset field arrangement order to form version reference input data;
[0047] S44. Perform encoding processing on the version reference input data. The encoding processing includes character encoding uniform processing of the field content and alignment processing according to a preset byte length to generate version reference encoded data.
[0048] S45. Construct version reference nodes based on the version reference encoding data, and add the version reference nodes to the bottom node set of the improved Merkle mountain structure;
[0049] S46. Arrange the version reference nodes in order according to the correspondence between version association information and update version identifier to form an archive version chain structure composed of multiple version reference nodes.
[0050] Optionally, S5 specifically includes:
[0051] S51. When a digital archive is processed by multiple organizations, obtain the mountain root nodes of the improved Merkle mountain structure constructed by different organizations for the same digital archive, and record the mountain root nodes as cross-organization root nodes.
[0052] S52. Obtain the conflict resolution results generated by different organizations in digital archive processing, and record the conflict resolution results as a cross-organizational conflict field.
[0053] S53. The cross-organization root node and the cross-organization conflict field are arranged into a cross-organization root node set according to a preset field order, wherein the preset order is arranged in the byte order of the organization identifier;
[0054] S54. The cross-organization root node set is grouped according to the adjacent pairing method, and every two adjacent cross-organization root nodes are connected according to a fixed field order to form cross-organization intermediate node input data.
[0055] S55. Perform encoding alignment processing on the cross-organization intermediate node input data. The encoding alignment processing includes uniform encoding processing of character content and length alignment using preset byte boundaries to generate cross-organization intermediate node encoded data.
[0056] S56. Construct cross-organization intermediate nodes based on the cross-organization intermediate node encoded data, and arrange them in the order of the cross-organization root node set to form a cross-organization intermediate node set;
[0057] S57. Repeatedly perform connection, encoding alignment and node construction processing on the cross-organization intermediate node set according to the adjacent pairing method until a single cross-organization root node is obtained;
[0058] S58. The single cross-organizational root node is used as the root node of the second-level improved Merkle mountain structure to form a cross-organizational level improved Merkle mountain structure.
[0059] Optionally, S6 specifically includes:
[0060] S61. When a processing event occurs in a digital archive, obtain the event time information corresponding to the processing event, and record the event time information as a time identifier field.
[0061] S62. Obtain the node sequence number identifier corresponding to the processing event from the state-embedded leaf node of the digital archive, and record the node sequence number identifier as an event node field;
[0062] S63. Connect the time identifier field and the event node field according to the preset field order to form time-anchored input data;
[0063] S64. Perform encoding processing on the time-anchored input data. The encoding processing includes character encoding unification processing and alignment processing according to preset byte alignment rules to generate time-anchored encoded data.
[0064] S65. Construct time anchor nodes based on the time anchor coding data, and add the time anchor nodes to the bottom node set of the improved Merkle mountain structure corresponding to the processing event;
[0065] S66. Arrange multiple time anchor nodes according to the order of their time identifier fields to form a time sequence chain.
[0066] Optionally, S7 specifically includes:
[0067] S71. When generating a root node in the improved Merkle mountain structure, obtain the node identifier corresponding to the root node and record the node identifier as a root node field.
[0068] S72. Obtain the input parameters of the delay function. The input parameters include the occurrence time field of the processing event, the node number field, and the byte sequence corresponding to the root node field. Connect the input parameters in a preset order to form the input data of the delay function. Perform multiple rounds of byte perturbation processing and cyclic shift processing on the input data of the delay function according to the calculation rules of the delay function to generate the output data of the delay function. Record the output data of the delay function as the delay field.
[0069] S73. Connect the root node field and the delay field according to the preset field order to form cumulative proof input data;
[0070] S74. Perform encoding processing on the cumulative proof input data. The encoding processing includes uniform encoding processing of character content and length alignment processing according to a preset byte alignment rule to generate cumulative proof encoded data.
[0071] S75. Write the cumulative proof code data into the blockchain, and record the block identifier after writing into the blockchain as the cumulative proof identifier;
[0072] S76. Arrange multiple cumulative proof identifiers in sequence according to the generation order of the cumulative proof identifiers to form a cumulative proof chain.
[0073] The beneficial effects of this invention are:
[0074] This invention constructs a trusted structure system for the entire digital archive process based on an improved Merkle mountain algorithm. Addressing existing issues in cross-organizational digital archive processing, such as broken version chains, unverifiable status, scattered processing event records, and inconsistent results across multiple organizations, this invention employs a unified rule for generating state information sequences. It structurally expresses archive content summaries, metadata summaries, version identifiers, processing event identifiers, and operator identifiers, and serializes them in a fixed order to form a consistent underlying archive state input across different systems. In the structure construction phase, the improved Merkle mountain algorithm is introduced. By performing encoding alignment, node pairing, and multi-level node tree construction on the segmented state information, a multi-peak archive structure capable of dynamic updates and continuous expansion is generated, enabling a chain-like expression of archive version, event, and state evolution. In cross-organizational collaboration scenarios, by collecting and structurally combining the root nodes of different organizations, and introducing cross-organizational conflict fields that record the root nodes of differences, organizational identifiers, and difference rule identifiers, this information is written into a second-layer improved Merkle mountain structure. This enables the processing differences formed by different organizations to have a traceable and verifiable structured expression. In the time series construction stage, by encoding and combining event time information with state node identifiers to generate time anchor nodes, a time sequence chain for the entire archive processing process is further constructed, achieving an immutable record of every state change of the archive. In the cumulative proof generation stage, by combining the root node field with the delay field obtained by multi-round byte perturbation and cyclic shift based on a delay function and encoding it into the blockchain, a proof chain with cumulative characteristics is formed, significantly increasing the difficulty of tampering and enhancing the credible verification capability throughout the entire lifecycle of the archive. Ultimately, this achieves a unified and credible expression and full-process verifiable management of content, metadata, versions, events, and cross-organizational differences in digital archives under a multi-organizational, multi-stage, and high-frequency update business environment, providing structural support for the authenticity, integrity, and traceability of archives. Attached Figure Description
[0075] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0076] Figure 1 This is a schematic diagram of the structure of a blockchain-based end-to-end trusted management platform for digital archives proposed in this invention.
[0077] Figure 2 This is a flowchart illustrating the workflow of a blockchain-based end-to-end trusted management platform for digital archives proposed in this invention.
[0078] Figure 3 This is a flowchart of the cross-organization node processing module in this invention. Detailed Implementation
[0079] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0080] refer to Figure 1-3 A blockchain-based end-to-end trusted management platform for digital archives includes:
[0081] The status information generation module is used to parse the file body, file header information, additional description information and processing event records of digital archives, and generate a status information sequence in a fixed order;
[0082] The leaf node construction module is used to segment, complete, and encode the state information sequence, generate state-embedded leaf nodes, and add them to the underlying node set of the improved Merkle mountain algorithm.
[0083] The node tree generation module is used to perform node pairing, splicing, encoding alignment and node construction on the underlying node set according to the improved Merkle mountain algorithm, and generate an improved Merkle mountain structure containing multiple mountain root nodes;
[0084] The version reference node module is used to obtain the root node identifier of the previous version and the node sequence identifier of the updated leaf node when the digital archive version is updated, perform field concatenation and encoding processing, generate version reference nodes and form an archive version chain structure.
[0085] The cross-organizational node processing module is used to collect the root nodes of the mountain peaks of different organizations, form a cross-organizational root node set, and perform node construction based on the cross-organizational root node set to generate a second-layer improved Merkle mountain structure.
[0086] The time anchor node module is used to obtain the event time information and the node sequence number of the corresponding leaf node when an event occurs, generate time anchor nodes, and form a time sequence chain in chronological order.
[0087] The cumulative proof generation module is used to obtain the root node field of the improved Merkle mountain structure, perform delayed function operations and construct cumulative proofs, write them into the blockchain and form a cumulative proof chain in the order of generation.
[0088] In this embodiment, the modules are interconnected using the following method:
[0089] S1. Perform data parsing on digital archives to generate archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier, and combine them in a preset order to form a status information sequence;
[0090] S2. Construct state-embedded leaf nodes based on the state information sequence, and write the state-embedded leaf nodes into the bottom node set of the improved Merkle mountain algorithm;
[0091] S3. Generate a node tree according to the mountain structure of the improved Merkle mountain algorithm, and perform iterative combination operations on the bottom node set to form an improved Merkle mountain structure containing multiple mountain root nodes.
[0092] S4. When the digital archive version is updated, a version reference node is generated and written into the improved Merkle mountain structure. The archive version chain structure is determined based on the version reference node.
[0093] S5. When processing archives across organizations, collect the root nodes of the improved Merkle mountain structure of multiple organizations, generate a cross-organization root node set, and construct a second-layer improved Merkle mountain structure based on the cross-organization root node set.
[0094] S6. When the processing event of the digital archive is recorded, a time anchor node is generated, the time anchor node is written into the corresponding improved Merkle mountain structure, and a time sequence chain is constructed based on the time anchor node.
[0095] S7. Combine the root node of the improved Merkle mountain structure with the output of the delay function to form a cumulative proof, and write the cumulative proof into the blockchain to generate a cumulative proof chain.
[0096] In this embodiment, S1 specifically includes:
[0097] S11. Perform data parsing on the digital archive's file body, extract text content, structured field content, and embedded object content data, and generate an archive content summary according to preset encoding rules;
[0098] S12. Parse the header information, additional description information and archive directory information of the digital archive, and generate an archive metadata digest according to the parsed information and preset encoding rules.
[0099] S13. Obtain the file number information and version sequence information of the digital file, and combine the file number information and version sequence information according to the preset field order to generate a file version identifier;
[0100] S14. Extract operation type information, operation time information and operation initiating terminal identification information from the recorded file processing operations, and combine them according to the preset field order to generate file processing event identifiers.
[0101] S15. Obtain the operator account identifier and organization identifier of the operator performing the processing operation, and generate the file operator identity identifier by combining them according to the preset field order;
[0102] S16. The archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier are concatenated in a fixed order and serialization processing is performed according to a preset encoding format to generate a status information sequence.
[0103] In this embodiment, S2 specifically includes:
[0104] S21. The state information sequence is segmented according to a preset length, and the last segment that is less than the preset length is padded according to the padding rule to obtain the data block to be processed.
[0105] S22. Connect the data block to be processed with the node type identifier and the node sequence number identifier in a fixed order to form leaf node input data; the node type identifier indicates that the leaf node input data corresponds to the bottom node of the improved Merkle mountain structure; the node sequence number identifier indicates the position of the leaf node input data in the bottom node set;
[0106] S23. Perform encoding processing on the leaf node input data. The encoding processing includes character set conversion processing and byte length alignment processing to generate leaf node encoded data.
[0107] S24. Generate state-embedded leaf nodes based on the leaf node encoding data, add the state-embedded leaf nodes to the bottom node set of the improved Merkle mountain algorithm, and determine the storage order of the state-embedded leaf nodes in the bottom node set according to the node sequence number identifier.
[0108] In this invention, the state information sequence is a continuous byte sequence that has undergone serialization before segmentation. To ensure the comparability of sequences generated by different files or different processing events in terms of length, the sequence is segmented according to a preset length, and padding values are added according to a fixed padding rule when the last segment is less than the preset length, so that all segments are structurally consistent. When constructing the leaf node input data, the node type identifier is represented by a pre-configured byte value indicating the underlying node attribute, and the node sequence number identifier is derived from the generation order of the state information sequence, forming a complete leaf node input structure after being connected to the data block to be processed. During the encoding process, a unified character encoding scheme is used for character set conversion to avoid byte differences generated by different environments, and byte alignment processing is performed according to fixed byte boundaries, so that the leaf node encoded data has a consistent alignment length. The finally generated state-embedded leaf nodes are positioned in the underlying node set according to the node sequence number identifier, providing a well-defined source of basic data for the subsequent generation of the improved Merkle mountain structure.
[0109] In this embodiment, S3 specifically includes:
[0110] S31. Obtain the state-embedded leaf nodes arranged by node number in the bottom node set, and use the state-embedded leaf nodes as the initial node set.
[0111] S32. The initial node set is combined in an adjacent pairing manner, and every two adjacent leaf nodes are spliced together in a fixed connection order to generate intermediate node input data.
[0112] S33. Perform encoding alignment processing on the intermediate node input data, and use the processed intermediate node input data as the basis for constructing intermediate nodes to generate corresponding intermediate nodes.
[0113] S34. Arrange the intermediate nodes in ascending order according to the node sequence number to form an intermediate node set, and determine whether to form the root node of the current level based on the number of nodes in the intermediate node set.
[0114] S35. If no mountain root node is formed in the current level, the set of intermediate nodes is taken as a new set of nodes to be processed, and the adjacent pairing, splicing and encoding processes are repeatedly performed on the new set of nodes to be processed to generate the set of intermediate nodes of the next level.
[0115] S36. If a mountain root node is formed at the current level, add the mountain root node to the list of mountain root nodes of the improved Merkle mountain structure, and continue to build a new mountain structure starting from the next unprocessed leaf node.
[0116] S37. After all the underlying node sets have been processed, an improved Merkle mountain structure containing multiple mountain root nodes is generated.
[0117] In this invention, the node type identifier consists of a fixed-length byte sequence, used to indicate that the current node belongs to the underlying node set. The node sequence number is generated incrementally based on the order of leaf nodes in the underlying node set, starting from zero and arranged sequentially according to the order of addition, thus ensuring that the position of each node in the set is uniquely identifiable. The adjacent pairing method is defined as grouping two nodes with adjacent sequence numbers together; when the number of nodes in the current level is odd, the node with the highest sequence number is directly retained as a candidate peak root node. The fixed connection order is to concatenate nodes according to the order of their sequence numbers, ensuring the traceability of the concatenated data structure. Encoding alignment processing uses a preset byte alignment rule, aligning to eight-byte boundaries, and extending the length of the concatenated data to meet the alignment conditions. The formation condition of the peak root node is defined as follows: when only one node remains in a certain level, that node becomes the peak root node of that level and is added to the peak root node list. The next unprocessed leaf node is determined according to the order of the underlying node set, that is, the next leaf node with a higher sequence number than the already processed nodes that did not participate in the current peak construction.
[0118] In this embodiment, S4 specifically includes:
[0119] S41. When a digital archive is updated, obtain the root node identifier of the previous version of the digital archive before the update, and use the root node identifier of the previous version as version association information.
[0120] S42. Obtain the node sequence number identifier of the state embedded leaf node corresponding to the state information sequence generated by the updated digital archive, and use the node sequence number identifier as the update version identifier.
[0121] S43. Connect the version association information with the updated version identifier according to the preset field arrangement order to form version reference input data;
[0122] S44. Perform encoding processing on the version reference input data. The encoding processing includes character encoding uniform processing of the field content and alignment processing according to a preset byte length to generate version reference encoded data.
[0123] S45. Construct version reference nodes based on the version reference encoding data, and add the version reference nodes to the bottom node set of the improved Merkle mountain structure;
[0124] S46. Arrange the version reference nodes in order according to the correspondence between version association information and update version identifier to form an archive version chain structure composed of multiple version reference nodes.
[0125] In this invention, the root node identifier of the previous version is obtained by selecting the root node with the highest sequence number from the list of mountain peak root nodes corresponding to the improved Merkle mountain structure of the digital archive before the update, to ensure that the version association information is consistent with the structural correspondence of the previous version. The node sequence number identifier corresponding to the updated version is generated incrementally according to the addition order when the state-embedded leaf node is added to the underlying node set, so that each version update can be located in a unique leaf node position in the set. After generating the version reference node, a new node sequence number identifier is assigned to the node, and it is inserted into the set in ascending order of the existing number of nodes in the underlying node set, so that the position of the version reference node in the structure is traceable. When multiple version reference nodes form a version chain structure, they are sorted according to the direct connection relationship between the version association information and the updated version identifier, so that the version reference nodes form a linear arrangement structure from the earlier version to the updated version, which is used to represent the continuous version order of the digital archive.
[0126] In this embodiment, S5 specifically includes:
[0127] S51. When a digital archive is processed by multiple organizations, obtain the mountain root nodes of the improved Merkle mountain structure constructed by different organizations for the same digital archive, and record the mountain root nodes as cross-organization root nodes.
[0128] S52. Obtain the conflict resolution results generated by different organizations in digital archive processing, and record the conflict resolution results as a cross-organizational conflict field.
[0129] S53. The cross-organization root node and the cross-organization conflict field are arranged into a cross-organization root node set according to a preset field order, wherein the preset order is arranged in the byte order of the organization identifier;
[0130] S54. The cross-organization root node set is grouped according to the adjacent pairing method, and every two adjacent cross-organization root nodes are connected according to a fixed field order to form cross-organization intermediate node input data.
[0131] S55. Perform encoding alignment processing on the cross-organization intermediate node input data. The encoding alignment processing includes uniform encoding processing of character content and length alignment using preset byte boundaries to generate cross-organization intermediate node encoded data.
[0132] S56. Construct cross-organization intermediate nodes based on the cross-organization intermediate node encoded data, and arrange them in the order of the cross-organization root node set to form a cross-organization intermediate node set;
[0133] S57. Repeatedly perform connection, encoding alignment and node construction processing on the cross-organization intermediate node set according to the adjacent pairing method until a single cross-organization root node is obtained;
[0134] S58. The single cross-organizational root node is used as the root node of the second-level improved Merkle mountain structure to form a cross-organizational level improved Merkle mountain structure.
[0135] In this invention, different organizations generate their own improved Merkle mountain structure when processing the same digital archive. Each organization can form distinct root nodes based on its own business processes, data sources, or processing order. When aggregating data generated by multiple organizations, to maintain the consistency of the cross-organizational structure, the conflict resolution results generated by each organization during archive processing are used as cross-organizational conflict fields in subsequent node construction, enabling the cross-organizational structure to fully record the processing differences between different organizations. The cross-organizational root nodes and cross-organizational conflict fields can be combined into a unified input sequence according to the byte order of the organization identifier to maintain consistency in the input order during subsequent node construction. For node construction at the cross-organizational level, the same pairing, connection, encoding alignment, and node construction methods as for a single-organizational structure can be adopted to ensure structural consistency between the cross-organizational level and the internal organizational level. The final generated single cross-organizational root node can serve as the top-level node of the second-layer improved Merkle mountain structure, providing a unified root node source for the trusted cross-organizational archive structure.
[0136] In this embodiment, S6 specifically includes:
[0137] S61. When a processing event occurs in a digital archive, obtain the event time information corresponding to the processing event, and record the event time information as a time identifier field.
[0138] S62. Obtain the node sequence number identifier corresponding to the processing event from the state-embedded leaf node of the digital archive, and record the node sequence number identifier as an event node field;
[0139] S63. Connect the time identifier field and the event node field according to the preset field order to form time-anchored input data;
[0140] S64. Perform encoding processing on the time-anchored input data. The encoding processing includes character encoding unification processing and alignment processing according to preset byte alignment rules to generate time-anchored encoded data.
[0141] S65. Construct time anchor nodes based on the time anchor coding data, and add the time anchor nodes to the bottom node set of the improved Merkle mountain structure corresponding to the processing event;
[0142] S66. Arrange multiple time anchor nodes according to the order of their time identifier fields to form a time sequence chain.
[0143] In this invention, when a digital archive processing event is recorded within the system, a data item containing the event occurrence time is simultaneously generated. This data item directly serves as the object for obtaining event time information and is used to form a time identifier field. The node sequence number identifier corresponding to the event is located from the bottom-level node set of the improved Merkle mountain structure. The location method is based on the order in which the state-embedded leaf nodes corresponding to the processing event are added. After concatenating the time identifier field with the event node field, to ensure that the time anchoring input data meets the unified format requirements of the improved Merkle mountain structure for bottom-level nodes, character encoding and byte alignment processing are performed on it to maintain consistency in the encoding structure of different time anchoring nodes. The encoded data serves as the basis for node construction to generate time anchoring nodes, and is added to the corresponding bottom-level node set of the improved Merkle mountain structure according to the structural position of the processing event. For time anchoring nodes generated by different processing events, they are arranged sequentially according to the time order of the time identifier field to form a time sequence chain arranged in chronological order, thereby giving the arrangement relationship of each processing event in the chain structure a clear chronological order.
[0144] In this embodiment, S7 specifically includes:
[0145] S71. When generating a root node in the improved Merkle mountain structure, obtain the node identifier corresponding to the root node and record the node identifier as a root node field.
[0146] S72. Obtain the input parameters of the delay function. The input parameters include the occurrence time field of the processing event, the node number field, and the byte sequence corresponding to the root node field. Connect the input parameters in a preset order to form the input data of the delay function. Perform multiple rounds of byte perturbation processing and cyclic shift processing on the input data of the delay function according to the calculation rules of the delay function to generate the output data of the delay function. Record the output data of the delay function as the delay field.
[0147] S73. Connect the root node field and the delay field according to the preset field order to form cumulative proof input data;
[0148] S74. Perform encoding processing on the cumulative proof input data. The encoding processing includes uniform encoding processing of character content and length alignment processing according to a preset byte alignment rule to generate cumulative proof encoded data.
[0149] S75. Write the cumulative proof code data into the blockchain, and record the block identifier after writing into the blockchain as the cumulative proof identifier;
[0150] S76. Arrange multiple cumulative proof identifiers in sequence according to the generation order of the cumulative proof identifiers to form a cumulative proof chain.
[0151] In this invention, when performing the delay function operation, the byte sequences corresponding to the occurrence time field, node sequence number field, and root node field of the processed event are concatenated in a fixed order to form the input data of the delay function. This input data is then subjected to multiple rounds of byte perturbation and cyclic shift processing according to the predetermined rules of the delay function. Byte perturbation processing includes XOR operations based on round constants and reverse arrangement processing based on fixed segments. Cyclic shift processing includes cyclic left shift or cyclic right shift based on a preset bit width, ensuring that the output data of the delay function maintains a fixed structural difference. After the generated delay function output data forms the delay field, it can be concatenated with the root node field in a fixed field order to generate the cumulative proof input data. This input data undergoes character encoding unification processing and byte alignment processing to ensure that the cumulative proof encoded data meets the structural format requirements for writing to the blockchain. After the cumulative proof encoded data is written to the blockchain, the block identifier formed after writing can be obtained from the blockchain record, and this block identifier is used as the cumulative proof identifier. Multiple cumulative proof identifiers can be arranged in the order of their generation to form a cumulative proof chain with an order relationship, so that the cumulative proof identifiers generated by subsequent processing events present a continuous sequence relationship in the chain structure.
[0152] Example 1:
[0153] To verify the feasibility of this invention in practice, it was applied to a cross-departmental collaborative workflow scenario for digital archives in a government service center. This center involves multiple departments, including civil affairs, education, human resources and social security, medical insurance, housing provident fund, and justice, with over 30,000 digital archives transferred daily across departments. During business processing, digital archives undergo multiple modifications, reviews, corrections, archiving, and transfers, forming a complex multi-source, multi-organizational processing chain. Under traditional technology, each department uses independent archive systems, resulting in different archive record formats, scattered version information, and a lack of verifiable chain structures. Inconsistencies frequently arise in the root node of the version generated by different departments for the same archive, making it difficult to restore the true state of the archives during subsequent business reviews and accountability traceability—a persistent and prominent problem.
[0154] In this experiment, after deploying the digital archive end-to-end trusted management platform of this invention in the government affairs center, a unified status information sequence is first generated for each digital archive. The content summary, metadata summary, version identifier, processing event identifier, and operator identity identifier are serialized in a fixed order, solving the cross-system incompatibility problem caused by inconsistent fields and inconsistent time records in the past. When a citizen applies for household registration transfer, the archive needs to be reviewed by four departments: civil affairs, public security, social security, and medical insurance. Each department will generate new processing events and version information during the review process. This invention automatically constructs a multi-peak structure through an improved Merkle mountain algorithm, so that the archive can still maintain structural continuity even with frequent updates. Unlike traditional Merkle trees, it does not require rebuilding the entire tree structure, greatly reducing the amount of computation.
[0155] When archives enter the cross-departmental circulation stage, the root nodes of the mountain peaks generated by different departments are collected and combined into a cross-organizational root node set. If there are differences in the root nodes generated by different departments, this invention will automatically generate a cross-organizational conflict field to record the differing root nodes, the corresponding organizational identifier, and the difference resolution rule identifier, and write it into the second-layer improved Merkle mountain structure to realize the structured recording of cross-organizational differences, enabling the archives to achieve true traceability in a multi-entity processing chain.
[0156] After the household registration transfer is completed, this invention generates a time-anchored node for each review operation, concatenating and encoding the effective time of the processed event, the corresponding leaf node sequence number, and the conflict field to provide a complete and tamper-proof time sequence chain for subsequent auditing departments. Finally, a cumulative proof is generated through a delay function, and the root node field and the delay field are written into the blockchain to ensure that no department can bypass the evidence storage structure to modify the file history.
[0157] To verify the effectiveness of the method of this invention, three consecutive months of business data from the government service center were selected for testing. The performance of the traditional system and the platform of this invention was compared in four key indicators: file consistency, conflict identification, evidence storage cost, and verification efficiency. The comparative experiment is shown in Table 1.
[0158] Table 1. Comparison of the credibility of cross-organizational government archive chains between the method of this invention and traditional methods.
[0159] Comparison indicators Method of the present invention Traditional methods File root node consistency rate (%) 99.87 95.27 Average number of cross-organizational conflicts per month (times) 3,322 40,812 Cumulative average time to generate proofs (ms) 38 11 Time sequence chain event alignment accuracy (%) 99.92 93.54 Archive version chain breakage rate (%) 0.02 2.31 End-to-end audit verification time (s) 0.7 2.4
[0160] As can be seen from the comparison results in Table 1 above, this invention has significant advantages in solving problems such as cross-organizational data discrepancies, broken archival version chains, unverifiable event records, and the vulnerability of evidence storage structures to parallel attacks. It maintains stable structure generation capabilities even with millions of archives, comprehensively improving the credibility of government digital archives in terms of authenticity, integrity, and traceability, providing strong and reliable technical support for the cross-organizational application of large-scale digital archives. This invention collected 2,561,923 cross-organizational archives within three months. In the traditional system, the inconsistency rate of archive root nodes between different departments was as high as 4.73%, with an average of over 40,000 conflicts per month. After introducing this invention, the consistency rate of archive root nodes was increased to 99.87% by recording the source of differences through cross-organizational conflict fields. The average time consumed by the delay function in generating the cumulative proof chain is 38ms, which is 12 times more effective against parallel attacks compared to the traditional hash direct connection method. Meanwhile, in a random sample of 10,000 files, the time sequence chain constructed by this invention achieved an accuracy of 99.92% in comparing the consistency of processed events, which is higher than the 93.54% of the traditional system, and the integrity of the verification chain was improved by about 6.38%.
[0161] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A blockchain-based end-to-end trusted management platform for digital archives, characterized in that: include: The status information generation module is used to parse the file body, file header information, additional description information and processing event records of digital archives, and generate a status information sequence in a fixed order; The leaf node construction module is used to segment, complete, and encode the state information sequence, generate state-embedded leaf nodes, and add them to the underlying node set of the improved Merkle mountain algorithm. The node tree generation module is used to perform node pairing, splicing, encoding alignment and node construction on the underlying node set according to the improved Merkle mountain algorithm, and generate an improved Merkle mountain structure containing multiple mountain root nodes; The version reference node module is used to obtain the root node identifier of the previous version and the node sequence identifier of the updated leaf node when the digital archive version is updated, perform field concatenation and encoding processing, generate version reference nodes and form an archive version chain structure. The cross-organizational node processing module is used to collect the root nodes of the mountain peaks of different organizations, form a cross-organizational root node set, and perform node construction based on the cross-organizational root node set to generate a second-layer improved Merkle mountain structure. The time anchor node module is used to obtain the event time information and the node sequence number of the corresponding leaf node when an event occurs, generate time anchor nodes, and form a time sequence chain in chronological order. The cumulative proof generation module is used to obtain the root node field of the improved Merkle mountain structure, perform delayed function operations, construct cumulative proofs, write them to the blockchain, and form a cumulative proof chain in the order of generation. Specifically, it includes: When generating the root node in the improved Merkle mountain structure, the node identifier corresponding to the root node is obtained and recorded as the root node field; Obtain the input parameters of the delay function, which include the occurrence time field of the processed event, the node number field, and the byte sequence corresponding to the root node field. Concatenate the input parameters in a preset order to form the input data of the delay function. Perform multiple rounds of byte perturbation processing and cyclic shift processing on the input data of the delay function according to the calculation rules of the delay function to generate the output data of the delay function. Record the output data of the delay function as the delay field. The root node field and the delay field are connected in a preset field order to form cumulative proof input data; The cumulative proof input data is subjected to encoding processing, which includes uniform encoding of character content and length alignment processing according to a preset byte alignment rule to generate cumulative proof encoded data. Write the cumulative proof encoded data into the blockchain, and record the block identifier after writing into the blockchain as the cumulative proof identifier; Multiple cumulative proof identifiers are arranged in sequence according to the generation order of the cumulative proof identifiers to form a cumulative proof chain.
2. The blockchain-based digital archive end-to-end trusted management platform according to claim 1, characterized in that, The modules are connected in the following way: S1. Perform data parsing on digital archives to generate archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier, and combine them in a preset order to form a status information sequence; S2. Construct state-embedded leaf nodes based on the state information sequence, and write the state-embedded leaf nodes into the bottom node set of the improved Merkle mountain algorithm; S3. Generate a node tree according to the mountain structure of the improved Merkle mountain algorithm, and perform iterative combination operations on the bottom node set to form an improved Merkle mountain structure containing multiple mountain root nodes. S4. When the digital archive version is updated, a version reference node is generated and written into the improved Merkle mountain structure. The archive version chain structure is determined based on the version reference node. S5. When processing archives across organizations, collect the root nodes of the improved Merkle mountain structure of multiple organizations, generate a cross-organization root node set, and construct a second-layer improved Merkle mountain structure based on the cross-organization root node set. S6. When the processing event of the digital archive is recorded, a time anchor node is generated, the time anchor node is written into the corresponding improved Merkle mountain structure, and a time sequence chain is constructed based on the time anchor node. S7. Combine the root node of the improved Merkle mountain structure with the output of the delay function to form a cumulative proof, and write the cumulative proof into the blockchain to generate a cumulative proof chain.
3. The blockchain-based digital archive end-to-end trusted management platform according to claim 2, characterized in that, S1 specifically includes: S11. Perform data parsing on the digital archive's file body, extract text content, structured field content, and embedded object content data, and generate an archive content summary according to preset encoding rules; S12. Parse the header information, additional description information and archive directory information of the digital archive, and generate an archive metadata digest according to the parsed information and preset encoding rules. S13. Obtain the file number information and version sequence information of the digital file, and combine the file number information and version sequence information according to the preset field order to generate a file version identifier; S14. Extract operation type information, operation time information and operation initiating terminal identification information from the recorded file processing operations, and combine them according to the preset field order to generate file processing event identifiers. S15. Obtain the operator account identifier and organization identifier of the operator performing the processing operation, and generate the file operator identity identifier by combining them according to the preset field order; S16. The archive content summary, archive metadata summary, archive version identifier, archive processing event identifier, and archive operator identity identifier are concatenated in a fixed order and serialization processing is performed according to a preset encoding format to generate a status information sequence.
4. The blockchain-based digital archive end-to-end trusted management platform according to claim 2, characterized in that, S2 specifically includes: S21. The state information sequence is segmented according to a preset length, and the last segment that is less than the preset length is padded according to the padding rule to obtain the data block to be processed. S22. Connect the data block to be processed with the node type identifier and the node sequence number identifier in a fixed order to form leaf node input data; the node type identifier indicates that the leaf node input data corresponds to the bottom node of the improved Merkle mountain structure; the node sequence number identifier indicates the position of the leaf node input data in the bottom node set; S23. Perform encoding processing on the leaf node input data. The encoding processing includes character set conversion processing and byte length alignment processing to generate leaf node encoded data. S24. Generate state-embedded leaf nodes based on the leaf node encoding data, add the state-embedded leaf nodes to the bottom node set of the improved Merkle mountain algorithm, and determine the storage order of the state-embedded leaf nodes in the bottom node set according to the node sequence number identifier.
5. A blockchain-based end-to-end trusted management platform for digital archives according to claim 2, characterized in that, S3 specifically includes: S31. Obtain the state-embedded leaf nodes arranged by node number in the bottom node set, and use the state-embedded leaf nodes as the initial node set. S32. The initial node set is combined in an adjacent pairing manner, and every two adjacent leaf nodes are spliced together in a fixed connection order to generate intermediate node input data. S33. Perform encoding alignment processing on the intermediate node input data, and use the processed intermediate node input data as the basis for constructing intermediate nodes to generate corresponding intermediate nodes. S34. Arrange the intermediate nodes in ascending order according to the node sequence number to form an intermediate node set, and determine whether to form the root node of the current level based on the number of nodes in the intermediate node set. S35. If no mountain root node is formed in the current level, the set of intermediate nodes is taken as a new set of nodes to be processed, and the adjacent pairing, splicing and encoding processes are repeatedly performed on the new set of nodes to be processed to generate the set of intermediate nodes of the next level. S36. If a mountain root node is formed at the current level, add the mountain root node to the list of mountain root nodes of the improved Merkle mountain structure, and continue to build a new mountain structure starting from the next unprocessed leaf node. S37. After all the underlying node sets have been processed, an improved Merkle mountain structure containing multiple mountain root nodes is generated.
6. A blockchain-based end-to-end trusted management platform for digital archives according to claim 2, characterized in that, S4 specifically includes: S41. When a digital archive is updated, obtain the root node identifier of the previous version of the digital archive before the update, and use the root node identifier of the previous version as version association information. S42. Obtain the node sequence number identifier of the state embedded leaf node corresponding to the state information sequence generated by the updated digital archive, and use the node sequence number identifier as the update version identifier. S43. Connect the version association information with the updated version identifier according to the preset field arrangement order to form version reference input data; S44. Perform encoding processing on the version reference input data. The encoding processing includes character encoding uniform processing of the field content and alignment processing according to a preset byte length to generate version reference encoded data. S45. Construct version reference nodes based on the version reference encoding data, and add the version reference nodes to the bottom node set of the improved Merkle mountain structure; S46. Arrange the version reference nodes in order according to the correspondence between version association information and update version identifier to form an archive version chain structure composed of multiple version reference nodes.
7. A blockchain-based digital archive end-to-end trusted management platform according to claim 2, characterized in that, S5 specifically includes: S51. When a digital archive is processed by multiple organizations, obtain the mountain root nodes of the improved Merkle mountain structure constructed by different organizations for the same digital archive, and record the mountain root nodes as cross-organization root nodes. S52. Obtain the conflict resolution results generated by different organizations in digital archive processing, and record the conflict resolution results as a cross-organizational conflict field. S53. The cross-organization root node and the cross-organization conflict field are arranged into a cross-organization root node set according to a preset field order, wherein the preset order is arranged in the byte order of the organization identifier; S54. The cross-organization root node set is grouped according to the adjacent pairing method, and every two adjacent cross-organization root nodes are connected according to a fixed field order to form cross-organization intermediate node input data. S55. Perform encoding alignment processing on the cross-organization intermediate node input data. The encoding alignment processing includes uniform encoding processing of character content and length alignment using preset byte boundaries to generate cross-organization intermediate node encoded data. S56. Construct cross-organization intermediate nodes based on the cross-organization intermediate node encoded data, and arrange them in the order of the cross-organization root node set to form a cross-organization intermediate node set; S57. Repeatedly perform connection, encoding alignment and node construction processing on the cross-organization intermediate node set according to the adjacent pairing method until a single cross-organization root node is obtained; S58. The single cross-organizational root node is used as the root node of the second-level improved Merkle mountain structure to form a cross-organizational level improved Merkle mountain structure.
8. A blockchain-based digital archive end-to-end trusted management platform according to claim 2, characterized in that, S6 specifically includes: S61. When a processing event occurs in a digital archive, obtain the event time information corresponding to the processing event, and record the event time information as a time identifier field. S62. Obtain the node sequence number identifier corresponding to the processing event from the state-embedded leaf node of the digital archive, and record the node sequence number identifier as an event node field; S63. Connect the time identifier field and the event node field according to the preset field order to form time-anchored input data; S64. Perform encoding processing on the time-anchored input data. The encoding processing includes character encoding unification processing and alignment processing according to preset byte alignment rules to generate time-anchored encoded data. S65. Construct time anchor nodes based on the time anchor coding data, and add the time anchor nodes to the bottom node set of the improved Merkle mountain structure corresponding to the processing event; S66. Arrange multiple time anchor nodes according to the order of their time identifier fields to form a time sequence chain.
Citation Information
Patent Citations
Digital twinning body data management method for full life cycle of product
CN110503290A
Digital archive security management method and system
CN118013493A