Electronic archive intelligent retrieval method and system based on block chain
Through the combination of blockchain and attribute encryption, a secure and reliable electronic archive management system has been built, which solves the problems of high security risks, lagging authority control and low retrieval efficiency in existing technologies, and realizes instant authority updates and efficient, privacy-protected retrieval.
Patent Information
- Application Number
- CN202510808808.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing electronic archive management system has problems such as high security risks, lagging authority control, low retrieval efficiency and insufficient privacy protection. In particular, it is difficult to achieve dynamic updates and cross-domain collaboration in a distributed storage environment.
It adopts a method that combines blockchain technology and attribute-based encryption (ABE) to achieve dynamic permission control and efficient retrieval through sharded encrypted storage, Merkle tree index and smart contracts.
It achieves data immutability, instant permission updates and efficient retrieval in a distributed environment, reduces retrieval latency, protects user privacy, and resolves the technical contradictions between security, efficiency and privacy protection.
Smart Images

Figure CN120653788A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of archive retrieval technology, and in particular to a blockchain-based intelligent retrieval method and system for electronic archives. Background Art
[0002] Current electronic archive management systems generally rely on centralized storage architectures and static access control mechanisms, which present systemic flaws. This centralized storage model centralizes archive texts and search indexes on a single server cluster, leading to a high concentration of security risks. Once attackers breach the system's defenses, they can access vast amounts of sensitive data, posing a particular threat to classified government or financial archives. Traditional permission control utilizes predefined role-based policies. Dynamic changes to user attributes require administrators to manually update policies, and response delays often lead to unauthorized access incidents. Furthermore, data credibility mechanisms are weak, with existing solutions relying on third-party auditing organizations to verify archive integrity. This makes it difficult to track data tampering in real time in a distributed storage environment.
[0003] Existing intelligent retrieval technology for electronic archives faces bottlenecks in dynamic update efficiency and cross-domain collaboration. Although the semantic retrieval system has introduced keyword expansion technology, it lacks in-depth analysis of contextual semantics, resulting in a high rate of missed detection of high-value archives. For example, the semantic association between "stroke" and "infarction" in medical archives is lost. The index update mechanism is rigid. When the access policy is adjusted or the metadata is updated, the inverted index needs to be fully rebuilt, causing minute-level interruptions to the retrieval service of the PB-level archive. Symmetric key pre-distribution schemes are still widely used in the field of encrypted retrieval, which cannot support the immediate effectiveness of complex attribute strategies. In addition, in multi-institutional collaboration scenarios, there is a lack of trusted audit links between independent storage systems. Existing cross-domain authentication requires multiple public key signature operations, which significantly increases the retrieval delay to hundreds of milliseconds, making it difficult to meet real-time access needs. Summary of the Invention
[0004] In order to solve the above technical problems, a method and system for intelligent retrieval of electronic archives based on blockchain is provided. This technical solution solves at least one of the technical problems mentioned in the above background technology.
[0005] In order to achieve the above objects, the technical solution adopted by the present invention is: A blockchain-based intelligent retrieval method for electronic archives, comprising: Extract text features of electronic archives and generate metadata including subject keywords, semantic tags and confidentiality level identification; The original electronic file is divided and encrypted and stored in a distributed storage system to generate a content hash value; Write metadata, content hash values, and ABE access policies into blockchain smart contracts; Build an inverted index based on metadata semantic tags, and associate index items with blockchain storage addresses; Use Merkle tree structure to organize indexes, calculate root hash and anchor it to the blockchain; Set an index update trigger to automatically rebuild the index when the smart contract captures permission changes; Parsing keywords and digital identity credentials in user search requests; Call the smart contract to verify whether the user attributes meet the ABE access policy of the target profile; After verification, the encrypted hash list of matching files is retrieved from the distributed index; The user terminal obtains the encrypted file fragments from the distributed storage system according to the hash list; Compare the file shard hash with the blockchain storage value to verify data integrity; The shards are combined and decrypted using the attribute private key to generate the final retrieval result.
[0006] Preferably, the ABE access strategy is specifically: Define archive access rules as Boolean logic expressions and bind them to user attribute certificates; Compile access policies into on-chain executable code through smart contracts; When user attributes change, policy updates and index rebuilds are automatically triggered.
[0007] Preferably, the inverted index is constructed as follows: Use NLP models to analyze archive content, expand synonym databases, and standardize subject headings; Construct a hierarchical Merkle Patricia Trie index, with leaf nodes storing archive hash pointers; Periodically write the index root hash in batches to the blockchain to generate a timestamp proof.
[0008] Preferably, the verification of whether the user attributes conform to the ABE access policy of the target profile specifically includes: The user submits a zero-knowledge proof commitment to prove the property; Smart contracts verify the validity of zero-knowledge proofs and output permission determination results without exposing attribute details; Generate a dynamic access token, the validity period of which is controlled by a smart contract countdown.
[0009] Preferably, the step of retrieving the encrypted hash list of matching archives in the distributed index specifically includes: Based on the keywords in the user's search request, a shard search mechanism is used to divide the user's search request into several search request shards, each of which contains at least one keyword; Each retrieval request shard is stored in a different blockchain shard network; Each retrieval request shard is routed to the target shard network for parallel execution, and the results are aggregated and returned.
[0010] Preferably, routing each retrieval request shard to the target shard network for parallel execution specifically includes: Use the BERT model to parse contextual semantics based on the keywords corresponding to each retrieval segment; Associate the semantic tag library in the metadata to generate an expanded query vector; Sort the search results by semantic similarity and output the top K related files as the search results for the keyword.
[0011] Preferably, comparing the archive shard hash with the blockchain storage value to verify data integrity specifically includes: The terminal calculates and obtains the SHA-256 hash value of the file fragment; Call the smart contract to compare the original hash stored on the chain; If there are any differences, the audit contract is triggered to verify the consistency of the distributed storage nodes.
[0012] Preferably, the retrieval method further includes an edge cache acceleration mechanism, and the edge cache acceleration mechanism is specifically: Monitor frequently accessed files and cache their encrypted copies in edge CDN nodes; Verify the latest status of the blockchain before caching to ensure data validity; When users search, they give priority to accessing edge nodes.
[0013] Furthermore, a blockchain-based electronic archive intelligent retrieval system is proposed, which is used to implement the above-mentioned blockchain-based electronic archive intelligent retrieval method, including: A metadata generation module configured to extract text features of electronic archives and generate metadata including subject keywords, semantic tags, and confidentiality level identifiers; A distributed storage module is configured to divide and encrypt the original electronic file and store it in a distributed storage system, and generate a content hash value; A blockchain contract module configured to write the metadata, content hash value, and ABE access policy into a blockchain smart contract; The index construction module is configured to build an inverted index based on metadata semantic tags, associate index items with blockchain storage addresses, organize the index using a Merkle tree structure, and calculate the root hash anchored to the blockchain; The trigger module is configured to respond to permission change events captured by the smart contract and automatically trigger index reconstruction; The authentication module is configured to parse the keywords and digital identity credentials in the user's search request and call the smart contract to verify whether the user's attributes meet the ABE access policy of the target archive; a retrieval execution module configured to retrieve a cryptographic hash list of matching archives in the distributed index; a data acquisition and verification module configured to obtain encrypted archive shards from the distributed storage system based on the hash list and compare the shard hash values with the blockchain stored values to verify data integrity; The result generation module is configured to combine the shards and decrypt using the attribute private key to generate the final retrieval result.
[0014] Optionally, the system further includes an edge cache acceleration module for implementing the retrieval method according to claim 8, wherein the edge cache acceleration module is configured to: monitor frequently accessed archives and cache their encrypted copies to edge CDN nodes; verify the latest status of the blockchain before activating the cache to confirm the validity of the data; and preferentially route user retrieval requests to edge nodes for retrieval.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs an electronic archive management system that takes into account both security and intelligent retrieval through the deep coupling of blockchain and attribute encryption. At the security and trust level, a dual-track mechanism of archive sharding encryption storage and blockchain hash anchoring is adopted to ensure the immutability of data in a distributed environment. Any abnormal tampering behavior can be exposed in real time through the on-chain self-verification mechanism, significantly improving the archive credibility level. At the level of dynamic authority control, through the on-chain compilation and automatic triggering mechanism of the ABE access policy, the policy takes effect immediately when the user attribute changes and the index is updated synchronously, fundamentally solving the hidden danger of unauthorized access caused by delayed authority updates. At the retrieval performance level, the inverted index based on semantic extension is combined with edge cache acceleration to significantly reduce the retrieval delay of high-frequency access archives. At the same time, the parallel processing of the sharded network effectively copes with the concurrent pressure of cross-domain access of massive data. At the privacy protection level, the application of zero-knowledge proof technology eliminates the need to disclose user sensitive information during the attribute verification process, providing access control without privacy exposure risks for high-confidentiality scenarios. This system has made a breakthrough in solving the long-standing technical contradiction in the field of electronic archive management, which is difficult to coordinate and optimize security control, retrieval efficiency and privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of the blockchain-based electronic archive intelligent retrieval method proposed in Example 1; Figure 2 A flowchart of the ABE access policy method proposed in Example 1; Figure 3 This is the process of constructing an inverted index proposed in Example 1; Figure 4 This is the method flow for verifying whether the user attributes conform to the ABE access policy of the target file proposed in Example 1; Figure 5 The method flow for retrieving an encrypted hash list of matching files in a distributed index proposed in Example 1; Figure 6 This is the method flow proposed in Example 1 for routing each retrieval request shard to the target shard network for parallel execution; Figure 7 This is the method flow for verifying data integrity proposed in Example 1; Figure 8 This is the method flow of the edge cache acceleration mechanism proposed in Example 2. DETAILED DESCRIPTION
[0017] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0018] Example 1 Reference Figure 1 As shown, a blockchain-based electronic archive intelligent retrieval method includes: Extract text features of electronic archives and generate metadata including subject keywords, semantic tags and confidentiality level identification; By automatically extracting multi-dimensional metadata, subject keywords, semantic tags, and confidentiality level identifiers, we build a machine-understandable archival semantic graph, breaking through the limitations of traditional file name-based retrieval and significantly improving the accuracy of capturing high-value information, especially for cross-disciplinary professional archives. The original electronic file is divided and encrypted and stored in a distributed storage system to generate a content hash value; Sharded encryption technology is used to break the single point of failure risk of centralized storage. Combined with cryptographic hashing, it generates a unique data fingerprint, laying the foundation for the subsequent establishment of distributed trust verification, eliminating the risk of entire file leaks from the root and ensuring data traceability. Write metadata, content hash values, and ABE access policies into blockchain smart contracts; Leveraging the blockchain's immutable nature to solidify access control logic, this technology enables trusted on-chain storage of ABE policies and metadata, resolving vulnerabilities in traditional systems where policy files are susceptible to malicious tampering. This also provides a decentralized trust anchor for cross-institutional collaboration. Build an inverted index based on metadata semantic tags, and associate index items with blockchain storage addresses; Converting unstructured semantic tags into indexable machine codes can improve the recall rate of fuzzy semantic retrieval, such as the association of "cardiovascular disease" with "myocardial infarction", by several times. At the same time, the binding of index items with blockchain addresses can ensure the verifiability of retrieval results. Use Merkle tree structure to organize indexes, calculate root hash and anchor it to the blockchain; By establishing a self-verifying index system through the Merkle tree structure, any tampering with the index node will cause an abnormal change in the root hash value, which improves the verification efficiency of index consistency in a distributed environment by two orders of magnitude compared to traditional replica comparison. Set an index update trigger to automatically rebuild the index when the smart contract captures permission changes; Solve the problem of tightly coupled permissions updates and index maintenance, achieve instant index synchronization when permissions policies are dynamically adjusted, and completely eliminate the service interruption window caused by manual index rebuilding; Parsing keywords and digital identity credentials in user search requests; Separating the authentication and retrieval process prevents overall service stagnation that can occur in traditional systems when a single module lags, providing flexible processing capabilities for high-concurrency scenarios. Call the smart contract to verify whether the user attributes meet the ABE access policy of the target profile; Based on the attribute matching mechanism of verifiable calculation on the chain, it realizes dynamic judgment of fine-grained permissions and solves complex attribute combination scenarios that traditional RBAC models cannot handle; After verification, the encrypted hash list of matching files is retrieved from the distributed index; By distributing the search load through a distributed architecture, query latency for archives with hundreds of millions of files no longer increases linearly with data volume, solving the performance bottleneck problem of centralized index servers during large-scale searches. The user terminal obtains the encrypted file fragments from the distributed storage system according to the hash list; A peer-to-peer data acquisition mechanism based on hash lists prevents archive content from flowing through third-party servers, cutting off the path of man-in-the-middle attacks in the transmission process; Compare the file shard hash with the blockchain storage value to verify data integrity; Establish an end-to-end data integrity closed-loop verification chain, and any tampering at the shard level will be immediately identified, meeting the rigid requirement of the judicial evidence chain that electronic files are "tamper-free"; The shards are combined and decrypted using the attribute private key to generate the final retrieval result.
[0019] The dynamic decryption mechanism using attribute private keys ensures that sensitive information is only delivered to authorized terminals, eliminating the risk of secondary leakage caused by centralized decryption of the entire text in traditional systems.
[0020] Reference Figure 2 As shown, the ABE access policy in this embodiment is specifically as follows: Define archive access rules as Boolean logic expressions and bind them to user attribute certificates; Compile access policies into on-chain executable code through smart contracts; When user attributes change, policy updates and index rebuilds are automatically triggered.
[0021] By compiling access policies into on-chain executable smart contracts, a dynamic permission control paradigm that unifies policy flexibility and execution rigidity is constructed. Access rules are defined using Boolean logic expressions, breaking through the static limitations of the traditional role-binding model, supporting fine-grained control of any attribute combination, and achieving precise authorization of highly sensitive archives. Policy codes solidified based on the blockchain's tamper-proof nature completely eliminate the risk of malicious tampering with access control lists in centralized storage, establishing a mandatory trusted mechanism for policy execution. When user attributes change, automatically triggered index reconstruction and policy updates form a responsive closed loop, enabling instant synchronization of permission revocation and termination of data access capabilities, eliminating the inherent time window period of traditional manual synchronization mechanisms and the risk of over-privilege leakage caused by excessive permissions. This design, for the first time, achieves a deep unification of flexible policy definition and real-time permission management in highly dynamic scenarios such as medical data sharing and cross-institutional collaboration.
[0022] Reference Figure 3 As shown, in this embodiment, the inverted index is constructed as follows: Use NLP models to analyze archive content, expand synonym databases, and standardize subject headings; Construct a hierarchical Merkle Patricia Trie index, with leaf nodes storing archive hash pointers; Periodically write the index root hash in batches to the blockchain to generate a timestamp proof.
[0023] A semantically enhanced hierarchical indexing architecture provides a dual guarantee for retrieval accuracy and data trustworthiness. By employing an NLP model to deeply analyze the semantics of archival content and expand the synonym network, the coverage of the subject vocabulary transcends the limitations of traditional keyword matching. For example, "myocardial infarction" is automatically linked to "heart attack" and "ST-segment elevation myocardial infarction," completely resolving the problem of missed detections caused by terminology variants. The hierarchical Merkle Patricia Trie index structure forms a tamper-proof data link by binding leaf nodes to archival hash pointers. Any local index tampering is captured in real time due to hierarchical hash verification failures, improving the sensitivity of data anomaly detection compared to traditional B+ tree indexes. The index root hash, periodically anchored to the blockchain, provides verifiable timestamp proof while establishing a lightweight evidence storage mechanism for the full index. This improves the efficiency of index integrity verification by several orders of magnitude in distributed environments, providing technical support for judicial traceability of large-scale archival access. This design fundamentally unifies the dual core capabilities of high-precision semantic retrieval and verifiable data trustworthiness.
[0024] Reference Figure 4As shown, in this embodiment, the ABE access policy for verifying whether the user attributes conform to the target profile specifically includes: The user submits a zero-knowledge proof commitment to prove the property; Smart contracts verify the validity of zero-knowledge proofs and output permission determination results without exposing attribute details; Generate a dynamic access token, the validity period of which is controlled by a smart contract countdown.
[0025] Through the collaborative mechanism of zero-knowledge proof and on-chain dynamic tokens, an access verification paradigm that prioritizes both privacy and security, as well as fine-grained control, has been established. The user's zero-knowledge commitment to submitting proof of attributes enables smart contracts to verify the validity of attributes without obtaining specific certificate numbers or institution names, completely eliminating the risk of sensitive identity information leakage in scenarios such as cross-hospital access to medical data. The dynamic access token generated based on on-chain verification uses a smart contract countdown control to precisely limit the validity of permissions, breaking through the security flaws of traditional static tokens that are long-lasting, and achieving time-boxed security isolation for highly sensitive operations such as emergency file transfers in the operating room. This design uses cryptographic technology to reconstruct the attribute verification logic, eliminating privacy exposure risks while establishing a dynamic security boundary with precise control over the time dimension for electronic file access.
[0026] Reference Figure 5 As shown, in this embodiment, searching for an encrypted hash list of matching files in a distributed index specifically includes: Based on the keywords in the user's search request, a shard search mechanism is used to divide the user's search request into several search request shards, each of which contains at least one keyword; Each retrieval request shard is stored in a different blockchain shard network; Each retrieval request shard is routed to the target shard network for parallel execution, and the results are aggregated and returned.
[0027] Through the parallel search architecture of the blockchain sharding network, a search paradigm has been established that simultaneously increases throughput and response speed in massive data environments. The sharding mechanism decouples user search requests into independent keyword subsets. For example, in medical archive retrieval, "electrocardiogram" and "ventricular premature beats" are separated into different shards. Each search shard is then executed in parallel within a dedicated blockchain sharding network, completely eliminating the single-point performance bottleneck of traditional centralized indexing. Based on the precise routing strategy of the target sharding network, multi-chain concurrent computing resources are fully utilized, such as the hospital's internal chain processing professional terminology shards and the public chain processing general word shards, significantly reducing cross-domain search latency for petabyte-level archives. The verifiable merging mechanism of search results at the aggregation layer not only ensures the integrity of distributed execution, but also avoids the risk of network congestion caused by the transmission of full data sets in traditional solutions. This design, for the first time, realizes the high-concurrency, low-latency, and verifiable paradigm-level reconstruction of distributed index retrieval in scenarios such as medical big data platforms and cross-border financial archives.
[0028] Reference Figure 6 As shown, in this embodiment, routing each retrieval request shard to the target shard network for parallel execution specifically includes: Use the BERT model to parse contextual semantics based on the keywords corresponding to each retrieval segment; Associate the semantic tag library in the metadata to generate an expanded query vector; The search results are sorted by semantic similarity, and the first K related files are output as the search results of the keyword. K is the set search output value, which is used to control the number of files output by the search. In some embodiments, the value of K is set to 5.
[0029] Through BERT-driven contextual semantic parsing and multimodal association mechanisms, a high-precision intelligent retrieval paradigm has been established for complex professional scenarios. The BERT model deeply deconstructs the contextual semantics of keywords (for example, in medical scenarios, "conduction block" automatically associates with clinical terms such as "atrioventricular conduction delay" and "bundle branch block"), addressing the drawback of traditional keyword matching, which homogenizes term variants and synonyms. By dynamically linking with metadata semantic tag libraries to generate extended query vectors, the search scope extends beyond the literal limitations of the user's original query, enabling the structured capture of implicit knowledge in professional fields. A mechanism for intelligently sorting search results by semantic similarity and outputting highly relevant archives improves the first-screen hit rate of target data to a practical level in scenarios such as judicial archive investigations and interdisciplinary scientific research retrieval. This design fundamentally reconstructs the depth of semantic understanding in electronic archive retrieval, providing expert-like intelligent interpretation capabilities for professional-level knowledge discovery.
[0030] Reference Figure 7 As shown, in this embodiment, comparing the file shard hash with the blockchain storage value to verify data integrity specifically includes: The terminal calculates and obtains the SHA-256 hash value of the file fragment; Call the smart contract to compare the original hash stored on the chain; If there are any differences, the audit contract is triggered to verify the consistency of the distributed storage nodes.
[0031] Through a closed-loop verification mechanism that collaborates on-chain and with terminals, a real-time, self-feedback data integrity assurance paradigm has been established. The terminal calculates the hash value of the archive shard and compares it in real time with the original value solidified on the blockchain, allowing single-shard-level tampering to be accurately located within milliseconds, improving efficiency by two orders of magnitude compared to traditional full-volume archive verification. When a hash anomaly is detected, a smart audit contract is automatically triggered, rapidly executing multi-copy consistency verification via distributed nodes, eliminating manual intervention while providing automated tracking capabilities for the source of tampering. This design, for the first time, achieves end-to-end, real-time self-verification of data integrity in zero-fault-tolerance scenarios such as real-time surgical file transfers and judicial evidence chain verification, upgrading passive auditing to a proactive defense trust infrastructure.
[0032] Example 2 Reference Figure 8 As shown, based on the first embodiment, this embodiment proposes an edge cache acceleration mechanism to further improve the retrieval efficiency. The edge cache acceleration mechanism is specifically as follows: Monitor frequently accessed files and cache their encrypted copies in edge CDN nodes; Verify the latest status of the blockchain before caching to ensure data validity; When users search, they give priority to accessing edge nodes.
[0033] Frequently accessed archives are determined by using the access data of archives in the previous access cycle, combined with the time decay mechanism. By calculating the proportion of the number of times each archive is retrieved and accessed in each access cycle, the access coefficient of the archive in the access cycle is determined; The specific calculation formula for the access coefficient is: ; in, is the access coefficient of the i-th file in the access cycle, is the number of times the i-th file is accessed during the access cycle, U is the set of files accessed during the access cycle, is the number of visits to the jth element in U; The access frequency coefficient of the access file is calculated by combining the access coefficients of all past access cycles. The calculation formula of the access frequency coefficient is: ; in, is the access frequency coefficient of the i-th access file in the t-th period, is the access frequency coefficient of the i-th access file in the t-1th cycle, α and β are time decay coefficients, and α + β = 1. If the archive system has high timeliness, increase the value of α to increase the weight of the access data in the most recent cycle; If the access frequency coefficient of the file in the most recent access cycle is higher than the threshold, the file is classified as a high-frequency access file.
[0034] Through the edge cache architecture driven by blockchain verification, an efficient retrieval paradigm that dynamically unifies acceleration performance and data trust has been established. It intelligently monitors frequently accessed archives and caches encrypted copies to edge CDN nodes, providing near-local response capabilities for cross-regional retrieval. For example, the response speed for emergency medical record retrieval is improved, completely eliminating the cross-network delay bottleneck of centralized storage. Before activating the cache, the latest blockchain status is forcibly verified to ensure that edge data and on-chain records remain synchronized, maintaining strict on-chain and off-chain data consistency while utilizing cache acceleration. The intelligent routing mechanism during user retrieval prioritizes access to edge nodes, shortening the transmission path of core medical data to the optimal topological level in network congestion scenarios. This design has achieved a breakthrough in resolving the historical contradiction between cache acceleration and data trust, building a core technical foundation for "acceleration without reducing trust" for time-sensitive scenarios such as remote consultations and emergency rescue.
[0035] Example 3 This embodiment proposes a blockchain-based electronic archive intelligent retrieval system, which is used to implement the blockchain-based electronic archive intelligent retrieval method proposed in Example 1, including: A metadata generation module configured to extract text features of electronic archives and generate metadata including subject keywords, semantic tags, and confidentiality level identifiers; A distributed storage module is configured to divide and encrypt the original electronic file and store it in a distributed storage system, and generate a content hash value; The blockchain contract module is configured to write metadata, content hash values, and ABE access policies into the blockchain smart contract; The index construction module is configured to build an inverted index based on metadata semantic tags, associate index items with blockchain storage addresses, organize the index using a Merkle tree structure, and calculate the root hash anchored to the blockchain; The trigger module is configured to respond to permission change events captured by the smart contract and automatically trigger index reconstruction; The authentication module is configured to parse the keywords and digital identity credentials in the user's search request and call the smart contract to verify whether the user's attributes meet the ABE access policy of the target archive; a retrieval execution module configured to retrieve a cryptographic hash list of matching archives in the distributed index; a data acquisition and verification module configured to obtain encrypted archive shards from the distributed storage system based on the hash list and compare the shard hash values with the blockchain stored values to verify data integrity; The result generation module is configured to combine the shards and decrypt using the attribute private key to generate the final retrieval result.
[0036] Example 4 Furthermore, based on the third embodiment, in order to further implement the blockchain-based electronic archive intelligent retrieval method proposed in the second embodiment, the system proposed in this embodiment further includes: The edge cache acceleration module is configured to: monitor frequently accessed files and cache their encrypted copies to edge CDN nodes; verify the latest status of the blockchain to confirm data validity before activating the cache; and prioritize routing user search requests to edge nodes for search execution.
[0037] In summary, the advantages of the present invention are: through the deep coupling of blockchain and attribute encryption, an electronic archive management system that takes into account both security and intelligent retrieval is constructed. At the security and trust level, a dual-track mechanism of archive sharding encryption storage and blockchain hash anchoring is adopted to ensure the immutability of data in a distributed environment. Any abnormal tampering behavior can be exposed in real time through the on-chain self-verification mechanism, significantly improving the archive credibility level; at the level of dynamic authority control, through the on-chain compilation and automatic triggering mechanism of the ABE access policy, the policy is immediately effective and the index is updated synchronously when the user attribute is changed, fundamentally solving the hidden danger of unauthorized access caused by delayed authority update; at the level of retrieval performance, the inverted index based on semantic extension is combined with edge cache acceleration to greatly reduce the retrieval delay of high-frequency access archives, and at the same time, the parallel processing of the sharded network effectively copes with the concurrent pressure of cross-domain access of massive data; at the level of privacy protection, the application of zero-knowledge proof technology makes it unnecessary to disclose user sensitive information during the attribute verification process, providing access control without privacy exposure risk for high-confidentiality scenarios. This system has made a breakthrough in solving the long-standing technical contradiction in the field of electronic archive management, which is difficult to coordinate and optimize security management, retrieval efficiency and privacy protection.
[0038] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A blockchain-based electronic archive intelligent retrieval method, characterized in that: include: Extract text features of electronic archives and generate metadata including subject keywords, semantic tags and confidentiality level identification; The original electronic file is divided and encrypted and stored in a distributed storage system to generate a content hash value; Write metadata, content hash values, and ABE access policies into blockchain smart contracts; Build an inverted index based on metadata semantic tags, and associate index items with blockchain storage addresses; Use Merkle tree structure to organize indexes, calculate root hash and anchor it to the blockchain; Set an index update trigger to automatically rebuild the index when the smart contract captures permission changes; Parsing keywords and digital identity credentials in user search requests; Call the smart contract to verify whether the user attributes meet the ABE access policy of the target profile; After verification, the encrypted hash list of matching files is retrieved from the distributed index; The user terminal obtains the encrypted file fragments from the distributed storage system according to the hash list; Compare the file shard hash with the blockchain storage value to verify data integrity; The shards are combined and decrypted using the attribute private key to generate the final retrieval result.
2. The blockchain-based electronic archive intelligent retrieval method according to claim 1 is characterized in that: The ABE access policy is specifically as follows: Define archive access rules as Boolean logic expressions and bind them to user attribute certificates; Compile access policies into on-chain executable code through smart contracts; When user attributes change, policy updates and index rebuilds are automatically triggered.
3. The blockchain-based electronic archive intelligent retrieval method according to claim 1 is characterized in that: The specific steps of constructing the inverted index are as follows: Use NLP models to analyze archive content, expand synonym databases, and standardize subject headings; Construct a hierarchical Merkle Patricia Trie index, with leaf nodes storing archive hash pointers; Periodically write the index root hash in batches to the blockchain to generate a timestamp proof.
4. The blockchain-based electronic archive intelligent retrieval method according to claim 1 is characterized in that: The verification of whether the user attributes conform to the ABE access policy of the target profile specifically includes: The user submits a zero-knowledge proof commitment to prove the property; Smart contracts verify the validity of zero-knowledge proofs and output permission determination results without exposing attribute details; Generate a dynamic access token, the validity period of which is controlled by a smart contract countdown.
5. The blockchain-based electronic archive intelligent retrieval method according to claim 1 is characterized in that: The method of retrieving the encrypted hash list of matching files in the distributed index specifically includes: Based on the keywords in the user's search request, a shard search mechanism is used to divide the user's search request into several search request shards, each of which contains at least one keyword; Each retrieval request shard is stored in a different blockchain shard network; Each retrieval request shard is routed to the target shard network for parallel execution, and the results are aggregated and returned.
6. The blockchain-based electronic archive intelligent retrieval method according to claim 5 is characterized in that: The routing of each retrieval request shard to the target shard network for parallel execution specifically includes: Use the BERT model to parse contextual semantics based on the keywords corresponding to each retrieval segment; Associate the semantic tag library in the metadata to generate an expanded query vector; Sort the search results by semantic similarity and output the top K related files as the search results for the keyword.
7. The blockchain-based electronic archive intelligent retrieval method according to claim 1 is characterized in that: Comparing the archive shard hash with the blockchain storage value to verify data integrity specifically includes: The terminal calculates and obtains the SHA-256 hash value of the file fragment; Call the smart contract to compare the original hash stored on the chain; If there are any differences, the audit contract is triggered to verify the consistency of the distributed storage nodes.
8. The blockchain-based electronic archive intelligent retrieval method according to any one of claims 1 to 7, characterized in that: The retrieval method further includes an edge cache acceleration mechanism, which is specifically: Monitor frequently accessed files and cache their encrypted copies in edge CDN nodes; Verify the latest status of the blockchain before caching to ensure data validity; When users search, they give priority to accessing edge nodes.
9. A blockchain-based electronic archive intelligent retrieval system, characterized by: A method for implementing the blockchain-based electronic archive intelligent retrieval method according to any one of claims 1 to 8, comprising: A metadata generation module configured to extract text features of electronic archives and generate metadata including subject keywords, semantic tags, and confidentiality level identifiers; A distributed storage module is configured to divide and encrypt the original electronic file and store it in a distributed storage system, and generate a content hash value; A blockchain contract module configured to write the metadata, content hash value, and ABE access policy into a blockchain smart contract; The index construction module is configured to build an inverted index based on metadata semantic tags, associate index items with blockchain storage addresses, organize the index using a Merkle tree structure, and calculate the root hash anchored to the blockchain; The trigger module is configured to respond to permission change events captured by the smart contract and automatically trigger index reconstruction; The authentication module is configured to parse the keywords and digital identity credentials in the user's search request and call the smart contract to verify whether the user's attributes meet the ABE access policy of the target archive; a retrieval execution module configured to retrieve a cryptographic hash list of matching archives in the distributed index; a data acquisition and verification module configured to obtain encrypted archive shards from the distributed storage system based on the hash list and compare the shard hash values with the blockchain stored values to verify data integrity; The result generation module is configured to combine the shards and decrypt using the attribute private key to generate the final retrieval result.
10. The blockchain-based electronic archive intelligent retrieval system according to claim 9 is characterized in that: It also includes an edge cache acceleration module for implementing the retrieval method according to claim 8, wherein the edge cache acceleration module is configured to: monitor frequently accessed archives and cache their encrypted copies to edge CDN nodes; verify the latest status of the blockchain to confirm the validity of the data before activating the cache; and preferentially route user retrieval requests to edge nodes for retrieval.
Citation Information
Patent Citations
Block chain electronic accounting archive construction method
CN115510474A
Electronic archive management method and system based on block chain
CN116304265A
Digital archive security management method and system
CN118013493A
Archive data protection method based on block chain
CN118228312A
Archive management method based on AI and encrypted storage
CN119961216A
Cited By
AI application center application rapid delivery and full life cycle management method and system
CN120832890A
AI Application Center: Methods and Systems for Rapid Application Delivery and Full Lifecycle Management
CN120832890B
File storage system and method based on digital object architecture
CN121051805A
Data processing method and system based on artificial intelligence, storage device and storage medium
CN121502188A
Accounting file management method and system based on block chain technology and large language model
CN121542223A