Archive storage and management system based on block chain technology
Through the archive proof storage and management system based on blockchain technology, the problems of data tampering and loss in centralized storage are solved, and the secure and reliable storage and management of archives are realized, the value of data assets is enhanced, and digital transformation and intelligent applications are supported.
Patent Information
- Application Number
- CN202510494753.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing archive management system relies on centralized storage, and there are problems such as data tampering, information loss and imperfect permission management. Especially after the popularization of electronic archives, long-term preservation and authenticity verification are difficult to guarantee.
The archive proof storage and management system based on blockchain technology is adopted to remove redundant information through digital processing and cleaning rules engines, and hash sharding and parallel processing are used to generate unique and tamper-proof digital fingerprints. Combined with smart contract management and encryption strategies, the secure storage and management of archive data is achieved.
It improves the security and reliability of archive storage, prevents tampering and loss, provides decentralized, automated, and highly trusted management solutions, enhances the value of data assets, and provides a high-quality data foundation for digital transformation and intelligent applications.
Smart Images

Figure CN120371806A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and in particular to an archival deposit and management system based on blockchain technology. Background Art
[0002] At present, the file management system mainly relies on the centralized storage mode, and there are problems such as data tampering, information loss, and imperfect permission management. In addition, the long-term preservation and authenticity verification of files face challenges. Especially in the context of the increasing popularity of electronic files, how to ensure their immutability and long-term availability has become the focus of the industry's attention. Summary of the Invention
[0003] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose an archival deposit and management system based on blockchain technology.
[0004] To achieve the above object, the present invention adopts the following technical solutions: An archival deposit and management system based on blockchain technology, including the following steps: S1: The operator completes the digitization and blockchain preparation of the files; Including the following sub-steps: S11: The operator digitizes the files; Use a professional scanning device with a resolution of more than 600 dpi to perform high-precision digital output on the files, and the output format uses the PDF / A-2u long-term preservation format. At the same time, use professional digital devices to read special carriers, implement optical character recognition and retain the original image layer, and establish a digital quality control system S12: The file system establishes a data cleaning rule engine to remove redundant information; Including the following steps: S121: The cleaning rule engine transfers the digitized file data to the preprocessing module; The preprocessing module performs hash sharding and parallel processing on the file data; ensure uniform data distribution S122: The cleaning rule engine performs duplicate data detection on the preprocessed file data; First, unify the format of the preprocessed file data. The file data includes text data, image data, and numerical data; for the text data: uniformly encode it as UTF-8 and process special characters; for the image data: uniformly convert it to the RGB mode and standardize the resolution; for the numerical data: unify the units; Then perform structured processing: use pandas to process heterogeneous data; Then the cleaning rule engine extracts features, including text features, image features, and spatio-temporal features; the text feature extraction includes BERT word vector extraction and statistical feature extraction; The image feature extraction includes image preprocessing, model loading, and feature extraction; the spatio-temporal feature extraction undergoes geohash encoding and time window bucketing; Next, the cleaning rule engine performs LSH bucketing operations; the cleaning rule engine obtains candidate pair generation through cross-bucket retrieval, similarity screening, and distributed computing; Finally, the cleaning rule engine performs precise similarity calculation and clears duplicate data through repeated determination; S123: The cleaning rule engine normalizes the metadata of the archival data after detecting duplicate data; S124: The cleaning rule engine converts the format of the archival data with normalized metadata; S125: The cleaning rule engine processes and stores sensitive information of the archival data with normalized metadata; S126: The cleaning rule engine sends the cleaned archival data to the blockchain platform; S2: Blockchain platform construction; The blockchain platform selects the consortium chain solution, uses the Orderer node architecture, uses TLS1.3 for encrypted communication, deploys a blockchain browser in the network configuration, and implements a zero-trust security architecture; S3: Upload the archival data to the blockchain platform for on-chain and notarization operations; The specific steps of step S3 include the following sub-steps: S31: Perform hash processing on the archival data to generate a unique and tamper-proof digital fingerprint; S311: Archival data preprocessing specifications; The preprocessing specifications include standardized encoding, normalization processing, and large file processing; The standardized encoding: The archival data is forced to use UTF-8 encoding, the line break character is unified as LF, and the BOM header is removed; The so-called normalization processing: The archival data performs unicode normalization, removes invisible characters, and unifies space processing; Large file processing: Split the archival data into N data blocks, calculate the hashes H1, H2, H3, H4... H N ; Then merge the hashes of two data blocks pairwise: Hash(H1 + H2), Hash(H3 + H4),..., recursively merge until the root hash is generated; S312: Generate a unique and tamper-proof digital fingerprint in the archival data; Adopt the dual hash construction of content features and structural features: The content hash module uses the SHA3-256 algorithm to generate a data fingerprint, with the complete digital signature of the input byte data, collision probability: approximately 1.0 × 10^-77, and output: a 256-bit binary digest value; The structural feature hash module includes data volume features, information entropy features, and file header features; The data volume feature: precisely records the length value L of the input byte stream, verification range: 0 ≤ L ≤ 2^64 - 1, precision: measured in byte-level units; The calculation algorithm of the information entropy feature: H(X) = -Σ(P(x_i) * log2(P(x_i))), where x_i represents each byte value (0 - 255), and P(x_i) is the byte occurrence frequency; The file header feature: extraction mechanism: captures the first 4-byte magic number, encoding method: represented as a hexadecimal string; Feature serialization processing: 1) Construct a dictionary structure: {'length': L, 'entropy': E, 'header_magic': H}; 2) JSON serialization: ensure ASCII encoding and sorting consistency; 3) Secondary hashing (SHA3-256): generate a structured feature digest; Perform composite hashing on the content hash module and the structural feature hash module: 1) The first layer of hashing operation: perform SHA3-256 hashing calculation on the original data data, and the output result is a 256-bit binary hash value, usually represented as a 64-bit hexadecimal string; 2) The second layer of hashing operation, convert the feature data features into a standardized JSON string, and perform SHA3-256 hashing calculation on the JSON string; 3) Binary concatenation operation, concatenate the two hash values in binary format, with a total length of 64 bytes; 4) The third layer of hashing operation, perform the final SHA3-256 hashing on the concatenated binary data.
[0005] S32: Upload the hashed archive data to the blockchain and record the relevant metadata at the same time; The upload rules include: intelligent contract template design to control the upload granularity Intelligent contract template design; The intelligent contract template includes an access control module, data life cycle management, and audit log rules; The access control module includes a role definition layer and a permission inheritance mechanism; The role definition layer includes a system administrator, an archive administrator, an auditor, and a general user; The permission inheritance mechanism adopts a tree-like role inheritance structure, where upper-level roles automatically inherit lower-level permissions, and dynamic permission adjustment: automatic permission recovery is achieved through time locks.
[0006] The design of the data life cycle management state machine is as follows: Initial state: Waiting to be uploaded to the chain Effective state: Verifiable Archived state: Read-only mode Destroyed state: Trigger zero-knowledge proof verification Version control: Adopt a chained version structure. Each new version contains the hash of the previous version. Version rollback requires two-thirds node consensus.
[0007] The log storage strategy of the audit log rules: Core operation logs are stored on the chain, detailed operation logs are stored using IPFS distributed storage, and log signatures use the SM2 algorithm to ensure integrity.
[0008] Controlling the on-chain granularity: Control the on-chain granularity through file-level processing, directory-level processing, and fonds-level architecture; The file-level processing uses the SM3 algorithm to calculate the file content hash, and attaches a timestamp and the digital signature of the creator.
[0009] Adopt an encryption policy; The encryption policy includes an encryption architecture, a key management scheme, and a key rotation mechanism; The encryption architecture includes a content encryption layer and a metadata encryption layer; The content encryption layer uses the SM2 algorithm to encrypt the file content and realizes fine-grained access based on attribute-based encryption; The metadata encryption layer uses SM3 to generate a hash fingerprint, and sensitive metadata fields are encrypted using SM4; The key management scheme adopts a hierarchical key system: Master key: Protected by the HSM hardware module; Data encryption key: Generated based on the key derivation function; Session key: Generated temporarily and destroyed immediately after use The key rotation mechanism: The content encryption key is automatically rotated monthly, and the master key is rotated annually, adopting a threshold signature scheme.
[0010] S4: The blockchain platform manages and utilizes archival data; The archival system constructs a system architecture that meets the requirements of Equal Protection 2.0, realizes controllable data sovereignty, establishes an on-chain data audit and tracking mechanism, adopts cross-chain technology to achieve multi-chain interoperability, develops standardized data interfaces, and supports the W3C verifiable credential standard.
[0011] The data management mainly includes permission management, and the data utilization is cross-chain query and trusted sharing.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: By using the cleaning rule engine to remove duplicate data in the archive data, not only the current data redundancy problem is solved, but also the data asset value of the organization is enhanced at the strategic level, providing a high-quality data foundation for digital transformation and intelligent applications (such as AI model training).
[0013] The method proposed by the present invention; by using smart contracts to manage and generate unique and immutable data fingerprints, the security of archive storage is improved, preventing tampering and loss, and providing a decentralized, automated, and highly trusted technical solution for archive management. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a step flowchart of an archive deposit and management system based on blockchain technology of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0015] To further understand the purpose, structure, features, and functions of the present invention, the following is a detailed description in conjunction with the embodiments.
[0016] As Figure 1 shown, an archive deposit and management system based on blockchain technology includes the following steps: S1: The operator completes the digitization and blockchain preparation of the archive; It includes the following sub-steps: S11: The operator digitizes the archive; Use a professional scanning device with a resolution of more than 600 dpi to perform high-precision digital output on the archive, and use the PDF / A-2u long-term preservation format for the output format. At the same time, use professional digital devices to read special carriers, perform optical character recognition and retain the original image layer, and establish a digital quality control system S12: The archive system establishes a data cleaning rule engine to remove redundant information; It includes the following steps: S121: The cleaning rule engine transfers the digitized archive data to the preprocessing module; The preprocessing module performs hash sharding and parallel processing on the archive data; ensuring uniform data distribution S122: The cleaning rule engine performs duplicate data detection on the preprocessed archive data; First, standardize the format of the preprocessed archive data. The archive data includes text data, image data, and numerical data; for the text data: uniformly encode it as UTF-8 and process special characters; for the image data: unify it to the RGB mode and standardize the resolution; for the numerical data: unify the units; Then perform structured processing: use pandas to process heterogeneous data; Then the cleaning rule engine extracts features, including text features, image features, and spatio-temporal features; the text feature extraction includes BERT word vector extraction and statistical feature extraction; The image feature extraction includes image preprocessing, model loading, and feature extraction; the spatio-temporal feature extraction goes through geohash encoding and time window bucketing; Next, the cleaning rule engine performs LSH bucketing operations; the cleaning rule engine obtains candidate pair generation through cross-bucket retrieval, similarity screening, and distributed computing; Finally, the cleaning rule engine performs exact similarity calculation and clears duplicate data through repeated determination; S123: The cleaning rule engine standardizes the metadata of the archival data after detecting duplicate data; Implement a unified data dictionary, automated data type inference and conversion, and use JSON Schema to standardize the metadata.
[0017] S124: The cleaning rule engine converts the format of the archival data with standardized metadata; Build-in common data format converters to handle character encoding conversion, automatic detection and automatic repair of abnormal formats.
[0018] S125: The cleaning rule engine processes and stores sensitive information of the archival data with standardized metadata; Sensitive data recognition based on regular expressions and NLP, supports multiple desensitization strategies, and uses columnar storage and integrates archival data encryption.
[0019] S126: The cleaning rule engine sends the cleaned archival data to the blockchain platform; S2: Blockchain platform construction; The blockchain platform selects the consortium chain solution, selects the Orderer node architecture, uses TLS1.3 for encrypted communication, deploys a blockchain browser in the network configuration, and implements a zero-trust security architecture; S3: Upload the archival data to the blockchain platform for on-chain and notarization operations; The step S3 specifically includes the following sub-steps: S31: Perform hash processing on the archival data to generate a unique and tamper-proof digital fingerprint; S311: Archival data preprocessing specifications; The preprocessing specifications include standardized encoding, normalization processing, and large file processing; The standardized encoding: The archival data is forced to use UTF-8 encoding, the line feed character is unified as LF, and the BOM header is deleted; The above normalization process: Normalize the archive data to unicode, remove invisible characters, and unify whitespace handling; Large file processing: Split the archive data into N data blocks, calculate the hashes H1, H2, H3, H4... H of each data block N ; Then merge the hashes of two data blocks pairwise: Hash(H1 + H2), Hash(H3 + H4),..., recursively merge until the root hash is generated; S312: Generate a unique and tamper-proof digital fingerprint in the archive data; Adopt a dual hash construction of content features and structural features: The content hash module uses the SHA3-256 algorithm to generate a data fingerprint, where the complete digital signature of the input byte data, collision probability: approximately 1.0 × 10^-77, output: 256-bit binary digest value; The structural feature hash module includes data volume features, information entropy features, and file header features; The data volume feature: Precisely record the length value L of the input byte stream, verification range: 0 ≤ L ≤ 2^64 - 1, precision: measured in byte-level units; The calculation algorithm of the information entropy feature: H(X) = -Σ(P(x_i) * log2(P(x_i))), where x_i represents each byte value (0 - 255), and P(x_i) is the byte occurrence frequency; The file header feature: Extraction mechanism: Capture the first 4-byte magic number, encoding method: represented as a hexadecimal string; Feature serialization processing: 1) Construct a dictionary structure: {'length': L, 'entropy': E, 'header_magic': H}; 2) JSON serialization: Ensure ASCII encoding and sorting consistency; 3) Secondary hash (SHA3-256): Generate a structured feature digest; Perform composite hashing on the content hash module and the structural feature hash module: 1) The first layer of hashing operation: Perform SHA3-256 hashing calculation on the original data data, and the output result is a 256-bit binary hash value, usually represented as a 64-bit hexadecimal string; 2) The second layer of hashing operation, convert the feature data features into a standardized JSON string, and perform SHA3-256 hashing calculation on the JSON string; 3) Binary concatenation operation, concatenate the two hash values in binary format, with a total length of 64 bytes; 4) The third layer of hashing operation, perform the final SHA3-256 hashing on the concatenated binary data.
[0020] S32: Upload the hashed archive data to the blockchain while recording the relevant metadata; The upload rules include: intelligent contract template design to control the upload granularity Intelligent contract template design; The intelligent contract template includes an access control module, data life cycle management, and audit log rules; The access control module includes a role definition layer and a permission inheritance mechanism; The role definition layer includes a system administrator, an archive administrator, an auditor, and a general user; The system administrator: has the permission to configure contract parameters and assign roles; The archive administrator: is responsible for fonds structure management and metadata maintenance; The auditor: has the permission to query and verify logs; The general user: fine-grained access control based on attributes; The permission inheritance mechanism adopts a tree-like role inheritance structure, where the superior role automatically inherits the permissions of the subordinate roles, and dynamic permission adjustment: the permissions are automatically recovered through a time lock.
[0021] The design of the data life cycle management state machine: Initial state: To be uploaded Effective state: Verifiable Archived state: Read-only mode Destroyed state: Trigger zero-knowledge proof verification Version control: Adopt a chained version structure, each new version contains the hash of the previous version, and version rollback requires two-thirds node consensus.
[0022] The log storage policy of the audit log rules: The core operation logs are stored on the blockchain, the detailed operation logs are stored using IPFS distributed storage, and the log signatures use the SM2 algorithm to ensure integrity.
[0023] Control the upload granularity: Control the upload granularity through file-level processing, directory-level processing, and fonds-level architecture; The file-level processing uses the SM3 algorithm to calculate the file content hash, and appends a timestamp and the digital signature of the creator.
[0024] The directory-level processing: Adopt SM3 as the hash function, and its verification mechanism provides a verification interface that includes path proof and supports batch verification.
[0025] The fonds-level architecture mainly includes blockchain structure design: main chain: storing metadata hashes and directory structures; side chain: storing detailed file data. Cross-chain verification is carried out through SPV verification by light nodes. Among them, version management: uses the blockchain forking mechanism to achieve version iteration, and the retention period of historical versions is the legal retention period.
[0026] Adopt encryption strategies; The encryption strategies include an encryption architecture, a key management scheme, and a key rotation mechanism; The encryption architecture includes a content encryption layer and a metadata encryption layer; the content encryption layer uses the SM2 algorithm to encrypt file content and realizes fine-grained access based on attribute-based encryption; the metadata encryption layer uses SM3 to generate hash fingerprints, and sensitive metadata fields are encrypted using SM4; The key management scheme adopts a hierarchical key system: master key: protected by the HSM hardware module; data encryption key: generated based on a key derivation function; session key: generated temporarily and destroyed immediately after use.
[0027] The key rotation mechanism: the content encryption key is automatically rotated monthly, and the master key is rotated annually, adopting a threshold signature scheme.
[0028] S4: The blockchain platform manages and utilizes archival data; The archival system constructs a system architecture that meets the requirements of Equal Protection 2.0, realizes controllable data sovereignty, establishes an on-chain data audit and tracking mechanism, uses cross-chain technology to achieve multi-chain interoperability, develops standardized data interfaces, and supports the W3C verifiable credential standard.
[0029] The data management mainly includes permission management, and the data utilization is cross-chain query and trusted sharing.
[0030] The present invention has been described by the above related embodiments. However, the above embodiments are only examples for implementing the present invention. It must be pointed out that the disclosed embodiments do not limit the scope of the present invention. On the contrary, modifications and refinements made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention.
Claims
1. An archival deposit and management system based on blockchain technology, characterized in that: It includes the following steps: S1: The operator completes the digitization and blockchain preparation of the files; It includes the following sub-steps: S11: The operator digitizes the files; S12: The file system establishes a data cleaning rule engine to remove redundant information; It includes the following steps: S121: The cleaning rule engine transfers the digitized file data to the preprocessing module; The preprocessing module performs hash sharding and parallel processing on the file data to ensure uniform data distribution S122: The cleaning rule engine performs duplicate data detection on the preprocessed file data; First, unify the format of the preprocessed file data. The file data includes text data, image data, and numerical data. For the text data: uniformly encode it as UTF-8 and process special characters; for the image data: unify it to the RGB mode and standardize the resolution; for the numerical data: unify the units; Then perform structured processing: use pandas to process heterogeneous data; Then the cleaning rule engine extracts features, including text features, image features, and spatio-temporal features. The text feature extraction includes BERT word vector extraction and statistical feature extraction; The image feature extraction includes image preprocessing, model loading, and feature extraction; the spatio-temporal feature extraction undergoes geohash encoding and time window bucketing; Then the cleaning rule engine performs LSH bucketing operations; the cleaning rule engine generates candidate pairs through cross-bucket retrieval, similarity screening, and distributed computing; Finally, the cleaning rule engine performs precise similarity calculation and clears duplicate data through repeated determination; S123: The cleaning rule engine standardizes the metadata of the file data after duplicate data detection; S124: The cleaning rule engine converts the format of the file data with standardized metadata; S125: The cleaning rule engine processes and stores sensitive information in the file data with standardized metadata; S126: The cleaning rule engine sends the cleaned file data to the blockchain platform; S2: Construction of the blockchain platform; The blockchain platform selects the consortium chain solution, chooses the Orderer node architecture, uses TLS1.3 for encrypted communication, deploys a blockchain browser in the network configuration, and implements a zero-trust security architecture; S3: The file data is uploaded to the blockchain platform for blockchain and evidence storage operations; S4: The blockchain platform manages and utilizes the file data; The data management mainly includes permission management, and the data utilization is cross-chain query and trusted sharing.
2. The archival deposit and management system based on blockchain technology according to claim 1, characterized in that: The specific content of step S11 is: Use a professional scanning device with a resolution of over 600 dpi to perform high-precision digital output on the files. The output format uses the PDF / A-2u long-term preservation format. At the same time, use professional digital devices to read special carriers, perform optical character recognition, and retain the original image layer, and establish a digital quality control system.
3. The blockchain technology-based file deposit and management system according to claim 1, characterized in that: The specific content of the blockchain rule in step S13 is: S131: Design of the smart contract template; The smart contract template includes an access control module, data life cycle management, and audit log rules; The access control module includes a role definition layer and a permission inheritance mechanism, S132: Control the blockchain granularity; Control the blockchain granularity through file-level processing, directory-level processing, and fonds-level architecture control; The file-level processing S133: Adopt an encryption policy; The encryption policy includes an encryption architecture, a key management scheme, and a key rotation mechanism; The encryption architecture includes a content encryption layer and a metadata encryption layer; the content encryption layer uses the SM2 algorithm to encrypt the file content and achieves fine-grained access based on attribute-based encryption; the metadata encryption layer uses SM3 to generate a hash fingerprint, and sensitive metadata fields are encrypted using SM4.
4. The blockchain technology-based file deposit and management system according to claim 1, characterized in that: The specific steps of step S3 include the following sub-steps: S31: Perform a hash process on the archival data to generate a unique and tamper-proof digital fingerprint; S32: Upload the archival data after the hash process and record the relevant metadata at the same time.
5. The archival deposit and management system based on blockchain technology according to claim 3, characterized in that: The step S31 includes the following steps: S311: Archival data preprocessing specifications; The preprocessing specifications include standard encoding, normalization processing, and large file processing; The standard encoding: The archival data is forced to use UTF-8 encoding, the line break is unified as LF, and the BOM header is deleted; The so-called normalization processing: The archival data normalizes unicode, removes invisible characters, and unifies space processing; Large file processing: Split the archive data into N data blocks, and calculate the hashes H1, H2, H3, H4... H of each data block. N ; Then merge the hashes of every two data blocks: Hash(H1 + H2), Hash(H3 + H4),..., recursively merge until the root hash is generated; S312: Generate a unique and tamper-proof digital fingerprint in the archival data; Adopt a dual hash construction of content features and structural features; The content hash module uses the SHA3-256 algorithm to generate a data fingerprint, where the complete digital signature of the input byte data, the collision probability: approximately 1.0 × 10^-77, and the output: a 256-bit binary digest value; The structural feature hash module includes data volume features, information entropy features, and file header features; The data volume feature: Accurately record the length value L of the input byte stream, the verification range: 0 ≤ L ≤ 2^64-1, and the accuracy: measured in byte-level units; The calculation algorithm of the information entropy feature: H(X) = -Σ(P(x_i) * log2(P(x_i))), where x_i represents each byte value (0-255), and P(x_i) is the byte occurrence frequency; The file header feature: Extraction mechanism: Capture the first 4-byte magic number, encoding method: represented by a hexadecimal string; Feature serialization processing: 1) Construct a dictionary structure: {'length': L, 'entropy': E, 'header_magic':H}; 2) JSON serialization: Ensure ASCII encoding and sorting consistency; 3) Secondary hash SHA3-256: Generate a structured feature digest; Perform composite hashing on the content hashing module and the structural feature hashing module: 1) The first-layer hashing operation: perform SHA3-256 hashing calculation on the original data data, and the output result is a 256-bit binary hash value, usually represented as a 64-bit hexadecimal string; 2) The second-layer hashing operation, convert the feature data features into a standardized JSON string, and perform SHA3-256 hashing calculation on the JSON string; 3) Binary splicing operation, splice the two hash values in binary format, and the total length is 64 bytes; 4) The third-layer hashing operation, perform the final SHA3-256 hashing on the spliced binary data.
6. The archival deposit and management system based on blockchain technology according to claim 1, wherein: The specific content of step S4 includes the following: The file system constructs a system architecture that meets the requirements of Equal Protection 2.0, realizes controllable data sovereignty, establishes an on-chain data audit and tracking mechanism, uses cross-chain technology to achieve multi-chain interoperability, develops standardized data interfaces, and supports the W3C verifiable credential standard.