Archive management method based on AI and encrypted storage

By introducing AI technology to archive management for data preprocessing and multi-level classification, and combining encryption and blockchain technology to ensure data security and integrity, it solves the problem that traditional archive management technology is difficult to achieve efficient, secure and intelligent archive management, and realizes efficient data retrieval and cross-institutional sharing.

CN119961216AInactive Publication Date: 2025-05-09INST OF GEOLOGY CHINESE ACAD OF GEOLOGICAL SCI

Patent Information

Application Number
CN202510435041.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional archive management technology is difficult to achieve efficient, secure and intelligent archive management, especially in terms of data preprocessing, storage security, retrieval efficiency and cross-organization sharing.

Method used

The natural language processing algorithm based on AI is used to pre-process archive information, and combined with dynamic classification tree model and hybrid encryption technology, the archival data is structured, multi-level classification and encryption process. Use distributed databases and blockchain technology for storage, use multimodal retrieval models and dynamic decryption strategies to improve retrieval efficiency and security, and realize cross-institutional archive sharing through alliance chains and smart contracts.

Benefits of technology

It significantly improves the quality and processing efficiency of archival data, ensures the security and integrity of data, improves retrieval efficiency and accuracy, and solves the trust and permission management problems shared by cross-institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961216A_ABST
    Figure CN119961216A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of archive management, and discloses an AI and encrypted storage-based archive management method, which comprises the following steps of: preprocessing original archive information by using a natural language processing algorithm to generate structured archive data, and performing multi-level classification through a dynamic classification tree model. Data are encrypted by adopting a hybrid encryption technology, and encrypted archive data blocks are constructed and stored in a distributed database. And performing parallel retrieval by using a multi-modal retrieval model, performing decryption according to user permission and scenes based on a dynamic decryption strategy, and finally transmitting data through a secure channel. In addition, the system also has the functions of file data updating and version control, log access and anomaly detection, cross-mechanism file sharing and the like. The efficiency, safety and intelligent level of archive management are improved, and the problems existing in traditional archive management are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archive management, and in particular to an archive management method based on AI and encrypted storage. Background Art

[0002] In the digital age, archive management faces unprecedented challenges and opportunities. With the rapid development of information technology, various types of archive data have exploded, and traditional archive management methods can no longer meet the needs of modern society for efficient, secure and intelligent archive management. From the perspective of data processing, traditional archive management lacks intelligent data preprocessing capabilities. Original archive information is often in various formats and complex content, containing a large amount of redundant information and noise data. For example, in the process of digitizing some historical archives, the scanned text has problems such as format disorder and character recognition errors. Manual processing is not only inefficient, but also prone to omissions. Moreover, in the face of massive archival texts, traditional methods are difficult to accurately extract key information and cannot quickly establish effective data indexes, making subsequent retrieval and utilization extremely inconvenient.

[0003] In terms of storage security, the importance of archival data is self-evident. Once the sensitive information contained in it is leaked, it may cause serious losses to individuals, enterprises or institutions. However, traditional storage methods mostly rely on centralized databases, which have the risk of single point failure, and the data is vulnerable to threats such as hacker attacks, malicious tampering and natural disasters. For example, if a centralized database is hacked, a large amount of archival data may be stolen or destroyed. In addition, in the application of encryption technology, a single encryption algorithm was mostly used in the past, which made it difficult to balance encryption efficiency and security, and could not meet the differentiated encryption requirements of different types of archival data.

[0004] Retrieval efficiency is another pain point of traditional archive management. When users need to find specific archives, the search method based on simple keyword matching often leads to inaccurate and incomplete search results because of the inability to understand semantic associations. For example, when searching for archives related to "Application of Artificial Intelligence in the Medical Field", only entering keywords may miss some important materials with similar expressions but different wording. Moreover, as the number of archives increases, the search speed will become slower and slower, seriously affecting work efficiency.

[0005] There are many obstacles to cross-institutional file sharing in the traditional model. There is a lack of unified sharing standards and security mechanisms between different institutions, and data sharing is often restricted by problems such as chaotic authority management and lack of trust. For example, in the medical industry, it is difficult to share patient files between different hospitals, which hinders the coordinated development of medical services; in the government sector, the poor flow of file information between departments affects the efficiency of public affairs. Summary of the invention

[0006] The purpose of the present invention is to provide an archive management method based on AI and encrypted storage to solve the problems raised in the above background technology.

[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solution: an archive management method based on AI and encrypted storage, the method comprising: Preprocessing the original archival information through a natural language processing algorithm, the preprocessing includes text segmentation, semantic entity recognition and redundant information elimination, to generate structured archival data; multi-level classification of the structured archival data is performed based on a dynamic classification tree model, the dynamic classification tree model constructs classification nodes through a hierarchical clustering algorithm, and dynamically adjusts the classification level based on the similarity of archival topics; The classified structured archive data is encrypted using a hybrid encryption technology, wherein the hybrid encryption technology includes a combination of a symmetric encryption algorithm and an asymmetric encryption algorithm, wherein the archive metadata is encrypted using an AES algorithm, and the archive access permission key is encrypted using an RSA algorithm; an encrypted archive data block is constructed, each data block including the encrypted archive content, the encrypted metadata signature, and a hash chain verification code; The encrypted archive data blocks are stored in a distributed database, which adopts a sharding storage mechanism, allocates data blocks in shards based on archive classification labels and access frequencies, and introduces a metadata index table based on blockchain to achieve location mapping and integrity verification of data blocks through smart contracts; According to the keyword information in the user access request, the distributed database is searched in parallel using a multimodal search model, the multimodal search model combines text keywords, semantic vectors and classification labels to build a joint index, and matches the location of the target archive data block through an approximate nearest neighbor search algorithm; The retrieved encrypted archive data block is decrypted based on a dynamic decryption strategy, wherein the dynamic decryption strategy adaptively selects a decryption process according to the user authority level and the access scenario, including single-layer decryption, multi-layer decryption and time-limited decryption; the decrypted archive data is generated and transmitted to the user terminal through a secure channel.

[0008] Preferably, preprocessing the original archival information by a natural language processing algorithm includes: Construct a text cleaning model, which uses a combination of regular expression matching and deep learning to identify and remove irrelevant characters, repeated paragraphs, and format noise in archive texts; The attention-enhanced bidirectional long short-term memory network model is used for semantic entity recognition to extract key entities, timestamps, and topic tags from archival texts and generate entity-relationship graphs. Based on the graph convolutional network, redundant nodes are pruned on the entity-relationship graph, and low-importance nodes and edges are removed by calculating the semantic similarity and connection weights between nodes to generate streamlined semantic structure data; The semantic structure data is aligned and mapped with the original archive text to generate structured archive data containing metadata description.

[0009] Preferably, multi-level classification of structured archival data based on a dynamic classification tree model includes: Initialize the root node of the classification tree and use TF-IDF weighted word vectors and Doc2Vec document vectors to construct archive feature representation; The initial classification level is generated by the hierarchical agglomerative clustering algorithm, the number of clusters is optimized based on the silhouette coefficient, and the topic similarity between adjacent levels is calculated; Introducing a dynamic split and merge mechanism. When the newly added archive data causes the level similarity to be lower than the preset threshold, the node split operation is triggered; when the similarity between levels is higher than the preset threshold, the node merge operation is triggered. A unique classification code is generated for each classification node, wherein the classification code includes a hierarchical path, a subject identifier, and a version number, and the classification code is embedded into the metadata of the structured archive data.

[0010] Preferably, the use of hybrid encryption technology to encrypt the classified structured archive data includes: The archive content is processed in blocks, and encryption priorities are divided according to the size and sensitivity of the data blocks. High-priority data blocks are encrypted using AES-256, and low-priority data blocks are encrypted using AES-128. Generate a random initialization vector and key derivation parameters for each data block, and generate a data block-specific encryption key based on the PBKDF2 algorithm and the user master key; The archive metadata is encrypted using the RSA-OAEP algorithm, an encrypted metadata signature is generated, and the public key hash value is embedded in the data block header; Construct a hash chain verification code, iteratively calculate the data block content and the hash value of the adjacent data block through the SHA-3 algorithm, and generate a chain integrity check code.

[0011] Preferably, storing the encrypted archive data blocks in a distributed database comprises: Build a data sharding strategy based on classification labels and access frequency, allocate high-frequency access data blocks to low-latency storage nodes, and low-frequency access data blocks to high-capacity storage nodes; Deploy a metadata index smart contract in the blockchain network, which records the data block hash value, storage location, and access permission policy, and implements privacy-preserving location query through zero-knowledge proof; Erasure coding technology is used to redundantly encode data blocks, and the encoded data blocks are stored in multiple geographically distributed storage nodes. Data availability is maintained through regular heartbeat detection.

[0012] Preferably, performing parallel retrieval using a multimodal retrieval model includes: Construct a joint index structure, wherein the joint index includes an inverted index, a vector index, and a graph index, corresponding to text keywords, semantic vectors, and classification labels, respectively; Perform semantic expansion on user query keywords and generate synonym sets and related concept sets based on the Word2Vec model and knowledge graph; The quantized approximate nearest neighbor algorithm is used to quickly search the vector index, and a multi-level search tree is constructed through product quantization and hierarchical clustering; Integrate the exact matching results of the inverted index, the similarity ranking results of the vector index, and the association path results of the graph index to generate a comprehensive retrieval score and filter the top N target data block identifiers.

[0013] Preferably, decrypting the encrypted archive data block based on the dynamic decryption strategy includes: Decryption permissions are divided according to user permission levels. Users with high permissions trigger a single-layer decryption process and directly use the master key to decrypt data blocks. Users with intermediate permissions trigger a multi-layer decryption process and need to verify the temporary token and dynamically generated session key in turn. In the time-limited decryption scenario, the decryption key is bound to a timeliness parameter and a time-lock-based encryption scheme is adopted. The key automatically becomes invalid after the preset time. Distributed decryption is achieved through a secure multi-party computing protocol, which divides the decryption task into multiple participating nodes. Each node collaborates to complete the decryption operation based on a secret sharing protocol.

[0014] Preferably, the method further comprises the steps of updating and versioning archive data: When the archive data is modified, an incremental update package is generated based on the differential encoding algorithm, and the updated part is re-encrypted and the hash chain is rebuilt; The version snapshot mechanism is used to record historical versions. Each snapshot contains a timestamp, a modifier's signature, and a summary of version differences. The Merkle tree structure is used to achieve fast verification between versions. Deploy a version synchronization protocol in a distributed database and coordinate data consistency among multiple nodes based on a distributed consistency algorithm.

[0015] Preferably, the method further comprises the steps of accessing logs and detecting anomalies: Record detailed logs of user access operations, including access time, request parameters, decryption operations, and data transmission paths, and encrypt and store the logs in an independent audit database; Build an anomaly detection model based on the combination of isolation forest and long short-term memory network to analyze the access frequency, permission abuse mode and decryption failure events in log data in real time and generate anomaly scores; When the anomaly score exceeds the preset threshold, an adaptive response mechanism is triggered, including temporarily freezing the account, increasing the key rotation frequency, and initiating a data self-destruction protocol.

[0016] Preferably, the method further comprises the step of cross-institutional archive sharing: Construct a consortium chain network, where each participating institution joins the network as a node and defines data sharing rules and permission strategies through smart contracts; Attribute-based encryption technology is used to achieve fine-grained access control, matching user attributes with archive access policies and dynamically generating decryption credentials; The introduction of homomorphic encryption technology in the sharing process allows third parties to perform statistical analysis on encrypted archive data without decryption, generate aggregated results and return them to the requester.

[0017] Compared with the prior art, the present invention has the following beneficial effects: In data processing, the original archival information is preprocessed with the help of natural language processing algorithms, which greatly improves the data quality. The text cleaning model can accurately remove irrelevant characters, repeated paragraphs and format noise. For example, regular expressions combined with deep learning can effectively handle garbled characters and typesetting errors in scanned archives. The bidirectional long short-term memory network model enhanced by the attention mechanism is used for semantic entity recognition, and an entity-relationship graph is constructed to accurately extract key information, providing a solid foundation for subsequent classification and retrieval. At the same time, the redundant node pruning operation based on the graph convolutional network further optimizes the data structure and improves data processing efficiency.

[0018] In terms of storage security, the application of hybrid encryption technology provides double protection. Symmetric encryption algorithms (such as AES) encrypt archive content and metadata to ensure data confidentiality; asymmetric encryption algorithms (such as RSA) are used to encrypt archive access rights keys, enhancing the security of key management. In addition, encryption priorities are divided according to the sensitivity and size of data blocks, and higher-level encryption algorithms (such as AES-256) are used for highly sensitive data, balancing encryption efficiency and security. The constructed hash chain verification code and encrypted metadata signature effectively prevent data tampering and ensure data integrity. The distributed database combines blockchain technology with a shard storage mechanism and metadata index table, which not only improves storage reliability, but also realizes location mapping and integrity verification of data blocks through smart contracts, reducing the risk of data loss and attacks.

[0019] Retrieval efficiency has been greatly improved. The multimodal retrieval model combines text keywords, semantic vectors and classification labels to build a joint index, and uses the approximate nearest neighbor search algorithm to quickly and accurately match the location of the target archive data block. Through semantic expansion technology, synonym sets and associated concept sets are generated based on the Word2Vec model and knowledge graph, which effectively expands the scope of retrieval and improves the accuracy and comprehensiveness of retrieval. For example, when searching for complex subject archives, relevant materials can be located more accurately, saving a lot of search time.

[0020] The dynamic decryption strategy adaptively selects the decryption process according to the user's permission level and access scenario, enhancing the flexibility and security of access control. The single-layer decryption process for high-level users improves work efficiency, while the multi-layer decryption process for intermediate users meets the needs of users at different levels while ensuring security. The time-limited decryption mechanism effectively prevents the risks brought by long-term exposure of keys, and the distributed decryption implemented by the secure multi-party computing protocol further ensures the security of the decryption process.

[0021] In terms of archive data update and version control, the difference encoding algorithm generates incremental update packages, which reduces the amount of data transmission and storage burden and improves update efficiency. The version snapshot mechanism combined with the Merkle tree structure facilitates the rapid verification of version differences and realizes the effective management and tracing of archive historical versions. The version synchronization protocol in the distributed database ensures the consistency of data between multiple nodes based on the distributed consistency algorithm.

[0022] The access log and anomaly detection functions provide security monitoring and early warning capabilities for the archive management system. The user access operation log is recorded in detail and encrypted for easy post-audit and tracking. The anomaly detection model based on the combination of the isolation forest and the long short-term memory network can analyze access behavior in real time, detect anomalies in a timely manner, such as abuse of authority and abnormal access frequency, and trigger an adaptive response mechanism, including temporary freezing of accounts, increasing the frequency of key rotation, and starting the data self-destruction protocol, effectively protecting the security of archive data.

[0023] Cross-institutional archive sharing clarifies data sharing rules and permission strategies by building alliance chain networks and smart contracts, solving trust and permission management issues. Attribute-based encryption technology implements fine-grained access control, and homomorphic encryption technology allows third parties to perform statistical analysis without decryption, which not only protects data privacy, but also promotes the rational use of data and promotes cross-institutional collaboration, such as achieving more efficient information sharing and collaborative work in the medical and government fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a working principle diagram of the archive management method of the present invention; Figure 2Workflow diagram for preprocessing of archival information; Figure 3 It is the workflow diagram of multi-level classification of dynamic classification tree model; Figure 4 Workflow diagram for encrypted archive data block storage. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] See also Figure 1-4 The present invention provides an archive management method based on AI and encrypted storage, and its overall implementation scheme is as follows: The original archival information is processed using natural language processing algorithms, including text segmentation, semantic entity recognition, and redundant information elimination, to generate structured archival data.

[0027] With the help of dynamic classification tree model, structured archival data is classified into multiple levels. The model constructs classification nodes through hierarchical clustering algorithm and dynamically adjusts the classification level according to the similarity of archival themes.

[0028] Hybrid encryption technology, that is, a combination of symmetric encryption algorithm and asymmetric encryption algorithm, is used to encrypt the classified structured archive data. Among them, the archive metadata is encrypted using the AES algorithm, and the archive access permission key is encrypted using the RSA algorithm. At the same time, an encrypted archive data block is constructed, each of which contains the encrypted archive content, encrypted metadata signature, and hash chain verification code.

[0029] The encrypted archive data blocks are stored in a distributed database. The database adopts a sharding storage mechanism, allocates data blocks in shards according to archive classification labels and access frequencies, introduces a metadata index table based on blockchain, and uses smart contracts to achieve location mapping and integrity verification of data blocks.

[0030] Based on the keyword information in the user's access request, a multimodal retrieval model is used to perform parallel retrieval on the distributed database. The model combines text keywords, semantic vectors, and classification labels to build a joint index, and matches the location of the target archive data block through an approximate nearest neighbor search algorithm. Afterwards, the retrieved encrypted archive data block is decrypted based on a dynamic decryption strategy. The decryption strategy adaptively selects the decryption process according to the user's permission level and access scenario, including single-layer decryption, multi-layer decryption, and time-limited decryption. Finally, the decrypted archive data is transmitted to the user terminal through a secure channel.

[0031] The implementation of the present invention is further described below in conjunction with Examples 1 to 6. Example

[0032] This embodiment describes in detail the specific process of archival information preprocessing, which aims to improve the accuracy and efficiency of preprocessing, so that the generated structured archival data can more accurately reflect the archival content and provide a high-quality data foundation for subsequent classification, encryption, retrieval and other operations.

[0033] When building a text cleaning model, the combination of regular expression matching and deep learning plays a key role. Regular expressions can quickly identify and remove irrelevant characters in archival texts. For example, when processing electronic documents containing a large number of special formatting marks, such as HTML tags, LaTeX commands, etc., by writing targeted regular expressions, these formatting characters that are irrelevant to the core content of the archives can be quickly removed. Deep learning models are used to handle more complex situations, such as identifying repeated paragraphs and formatting noise. Taking the convolutional neural network (CNN) as an example, a large number of archival texts containing different formatting noises are trained. CNN can learn the characteristic patterns of noise, thereby accurately detecting and removing formatting noises such as extra blank lines, inconsistent indentation, and typesetting errors, to ensure the purity of the archival text.

[0034] The use of the attention mechanism-enhanced bidirectional long short-term memory network model (BiLSTM-Attention) for semantic entity recognition can effectively extract key information from archival texts. When processing historical archives, there are many complex information such as people, events, and time. The BiLSTM-Attention model can focus on key parts through the attention mechanism and accurately extract key entities such as person names, event names, and timestamps. At the same time, it can also generate topic tags. For example, in an archive about ancient wars, a topic tag such as "history-war-ancient battle" is generated, and an entity-relationship map is constructed to clearly show the relationship between people and events, and events and time, such as whether a general participated in a battle and the specific time when the battle occurred, making the structure of the archival information clearer.

[0035] The redundant nodes of the entity-relationship graph are pruned based on the graph convolutional network (GCN). When calculating the semantic similarity between nodes, the cosine similarity algorithm is used, and the formula is: ,in and Represents the feature vectors of two nodes respectively. After calculating the similarity between nodes through this formula, combined with the connection weight, remove low-importance nodes and edges. For example, in an entity-relationship graph of an enterprise archive, some nodes with weak descriptiveness and little relationship with the core business process, after calculating their similarity and connection weight with other nodes, if it is found that they contribute little to the information expression of the overall graph, these nodes and their connecting edges will be deleted to generate streamlined semantic structure data, improving the clarity and processing efficiency of the graph.

[0036] Align and map the semantic structure data with the original archival text to generate structured archival data containing metadata descriptions. In this process, add detailed metadata descriptions for each entity in the semantic structure data, such as the entity's type, source, and other information. Taking a medical record as an example, for the entity "patient name", add the metadata description "patient basic information-identity identification", and for the "disease diagnosis time", add "medical record-diagnosis timestamp", etc., so that the structured archival data not only contains key information, but also has clear metadata annotations, which is convenient for subsequent management and utilization. Example

[0037] This embodiment revolves around multi-level classification based on a dynamic classification tree model, and its function is to achieve intelligent and dynamic classification of structured archival data, improve the accuracy and adaptability of archival classification, and facilitate users to quickly find and manage archives.

[0038] When initializing the root node of the classification tree, the TF-IDF weighted word vector and the Doc2Vec document vector are used to construct the archive feature representation. For a large amount of news archive data, the TF-IDF algorithm can calculate the importance of each word in the document, highlighting those words that appear frequently in a specific document but are uncommon in the entire document collection. For example, in a news article about scientific and technological achievements, the TF-IDF values ​​of words such as "quantum computing" and "artificial intelligence chips" are relatively high, indicating that they are of great significance to the topic expression of the document. The Doc2Vec model maps the entire document into a vector of fixed dimension, retaining the overall semantic information of the document. Combining these two vectors can more comprehensively describe the characteristics of the archive and provide richer data support for subsequent classification.

[0039] The initial classification level is generated by the hierarchical agglomerative clustering algorithm, and the number of clusters is optimized based on the silhouette coefficient. The calculation formula of the silhouette coefficient is: ,in It is a sample The average distance to other samples in the same cluster, It is a sample The minimum average distance to samples in other clusters. When clustering a batch of educational archives, try different numbers of clusters and calculate the silhouette coefficient for each number of clusters. The closer the silhouette coefficient is to 1, the better the clustering effect is. At this time, the corresponding number of clusters is the optimal number of clusters. At the same time, calculate the topic similarity between adjacent levels. KL divergence and other methods can be used to measure the degree of difference between topics at different levels to ensure the rationality of classification.

[0040] A dynamic splitting and merging mechanism is introduced. When new archive data is added, its similarity with the existing hierarchy is calculated. If the hierarchy similarity is lower than the preset threshold (such as set to 0.5), the node splitting operation is triggered. For example, in an e-commerce product archive, when a new type of product archive is added, such as "virtual reality equipment", there is no suitable category in the original classification hierarchy to accommodate it. At this time, through the node splitting operation, a new sub-node "virtual reality equipment" is created under the relevant electronic product classification to reasonably classify the archive. Conversely, when the similarity between levels is higher than the preset threshold, the node merging operation is triggered to optimize the classification structure, reduce redundancy, and improve the compactness and logic of the classification.

[0041] Generate a unique classification code for each classification node, including the hierarchical path, subject identifier, and version number. For example, the hierarchical path of a classification node is "1-2-3", which means that it is located at the third grandchild node under the second child node of the first layer under the root node; the subject identifier is "electronic products-mobile phones-smart phones", which clarifies the subject of the node; the version number is "V1.1", which is used to record the changes of the classification node. This classification code is embedded in the metadata of the structured archival data, which facilitates the rapid and accurate positioning and identification of the category to which the archive belongs in the subsequent storage, retrieval and management process. Example

[0042] After the file content is divided into blocks, the encryption priority is divided according to the size and sensitivity of the data block. For example, when processing the customer files of financial institutions, the data blocks containing sensitive information such as customer ID numbers and bank card passwords are classified as high priority and use the AES-256 encryption algorithm; while some descriptive information, such as customer occupations, hobbies, etc., are classified as low priority and use the AES-128 encryption algorithm. This can not only ensure the high security of sensitive data, but also improve encryption efficiency to a certain extent.

[0043] Generate a random initialization vector and key derivation parameters for each data block, and generate a data block-specific encryption key based on the PBKDF2 algorithm combined with the user master key. The PBKDF2 algorithm increases the security of the key through multiple iterative calculations. In actual applications, the user master key is processed by the PBKDF2 algorithm, combined with a randomly generated salt value and the number of iterations, to generate a high-strength, hard-to-crack data block-specific encryption key, providing independent encryption protection for each data block.

[0044] The RSA-OAEP algorithm is used to encrypt the archive metadata, generate an encrypted metadata signature, and embed the public key hash value into the data block header. During the encryption process, the RSA-OAEP algorithm enhances the security of encryption through a padding mechanism. For example, when encrypting metadata such as the creator and creation time of an archive, the RSA-OAEP algorithm is used to encrypt and generate an encrypted metadata signature to ensure the integrity and immutability of the metadata. The public key hash value is embedded in the data block header to facilitate quick verification of the legitimacy of the public key during decryption.

[0045] Construct a hash chain verification code, and iteratively calculate the data block content and the hash value of the adjacent data block through the SHA-3 algorithm to generate a chain integrity check code. Assume that there is a data block , , , first calculate ,Then , , This is the final hash chain verification code. In this way, any change in the content of any data block will cause the subsequent hash chain verification code to change, thereby effectively detecting whether the data has been tampered with during storage or transmission. Example

[0046] The data sharding strategy is built based on classification tags and access frequency. High-frequency access data blocks are allocated to low-latency storage nodes, and low-frequency access data blocks are allocated to high-capacity storage nodes. For example, in an enterprise's sales archive management system, recent sales order data has a high access frequency, and these data blocks are allocated to low-latency storage nodes equipped with high-performance solid-state drives (SSDs) to ensure that users can quickly obtain the latest sales information; while historical sales data from many years ago has a low access frequency, so it is allocated to large-capacity mechanical hard disk storage nodes to fully utilize storage resources and reduce storage costs.

[0047] A metadata index smart contract is deployed in the blockchain network. The smart contract records the hash value, storage location, and access permission policy of the data block, and implements privacy-protected location query through zero-knowledge proof. Taking the medical record sharing scenario as an example, different medical institutions serve as blockchain nodes, and the hash value, storage location, and access permission policy of each medical institution for these data of the patient's medical record data block are recorded in the smart contract. When a medical institution needs to query the storage location of a specific patient's file, it can verify whether it has the query permission and obtain the storage location through zero-knowledge proof technology without leaking any other information, thus protecting the patient's privacy and data security.

[0048] Erasure coding technology is used to redundantly encode data blocks, and the encoded data blocks are dispersed and stored in multiple geographically distributed storage nodes, and data availability is maintained through regular heartbeat detection. Assume that a data block is encoded into multiple redundant blocks and stored on storage nodes in different regions. When a node fails, other nodes can restore the lost data through the erasure coding algorithm. At the same time, the regular heartbeat detection mechanism can detect node failures in time, notify the system to repair data and replace nodes, ensure that data is always available, and improve the fault tolerance of the system. Example

[0049] This embodiment describes in detail the specific process of parallel retrieval using a multimodal retrieval model, which is used to improve the accuracy and speed of retrieving target archive data blocks in a distributed database and meet the needs of users to quickly obtain required archive information.

[0050] Construct a joint index structure, including inverted index, vector index and graph index, corresponding to text keywords, semantic vectors and classification labels respectively. In a system containing a large number of academic document archives, the inverted index can quickly locate documents containing specific keywords. For example, if you enter "artificial intelligence algorithm", the inverted index can quickly find all document data blocks containing these keywords; the vector index is based on semantic vectors, and by calculating semantic similarity, it can find documents with similar semantics to the query. For some synonymous or semantically related queries, it can provide more comprehensive search results; the graph index uses classification labels to build association paths. For example, in a document classification system, the graph index can find related documents in different subcategories under the same topic, broaden the search scope, and improve the comprehensiveness of the search.

[0051] The user's query keywords are semantically expanded, and a synonym set and a related concept set are generated based on the Word2Vec model and the knowledge graph. When a user enters "electric vehicle" for search, the Word2Vec model can find similar words, such as "new energy vehicle" and "pure electric vehicle", and the knowledge graph further provides related concepts, such as "battery technology" and "charging pile". Adding these expanded words and concepts to the search conditions can more comprehensively search for relevant archive data blocks and avoid missing important information.

[0052] The quantized approximate nearest neighbor algorithm is used to quickly search the vector index, and a multi-level search tree is constructed through product quantization and hierarchical clustering. Product quantization decomposes high-dimensional vectors into multiple low-dimensional vectors for quantization, reducing storage space and computation. For example, a 100-dimensional vector is decomposed into 10 10-dimensional vectors for quantization. Hierarchical clustering constructs a multi-level search tree. When searching, you can start from the root node and search downward step by step, quickly narrowing the search range and improving search efficiency. In this way, when processing large-scale vector indexes, you can quickly find the vector most similar to the query vector, thereby locating the target archive data block.

[0053] Integrate the exact match results of the inverted index, the similarity ranking results of the vector index, and the association path results of the graph index to generate a comprehensive search score and filter the top N target data block identifiers. For example, when searching for scientific and technological archives, the inverted index finds archives containing keywords, the vector index sorts these archives by semantic similarity, and the graph index finds archives under related topics. Set weights according to the importance of different index results, calculate the comprehensive search score, select the top N archive data block identifiers with the highest scores, and present the archives that best meet user needs to users first. Example

[0054] Decrypt the encrypted archive data blocks based on the dynamic decryption strategy, and divide the decryption permissions according to the user's permission level. In the internal archive management system of the enterprise, high-level users such as corporate executives trigger a single-layer decryption process and directly use the master key to decrypt the data block, which facilitates them to quickly obtain important information; intermediate-level users such as department managers trigger a multi-layer decryption process and need to verify the temporary token and the dynamically generated session key in turn to increase the security of decryption. In the time-limited decryption scenario, the decryption key is bound to the timeliness parameter, and an encryption scheme based on time lock is adopted. For example, the decryption key of a certain archive data block is set to be valid within 30 minutes. After the preset time, the key will automatically expire to prevent data security risks caused by key leakage. Distributed decryption is achieved through a secure multi-party computing protocol, and the decryption task is divided into multiple participating nodes. Each node collaborates to complete the decryption operation based on the secret sharing protocol, which improves the decryption efficiency while ensuring data security.

[0055] In terms of archive data update and version control, when archive data is modified, an incremental update package is generated based on the differential coding algorithm, and the updated part is re-encrypted and the hash chain is rebuilt. For example, in the document archive management of a software development project, after each code update, only the changed part of the code is recorded through the differential coding algorithm, and an incremental update package is generated to reduce data transmission and storage. At the same time, the updated part is re-encrypted to ensure the security of the data, and the hash chain verification code is rebuilt to ensure the integrity of the data. The version snapshot mechanism is used to record historical versions. Each snapshot contains a timestamp, a modifier signature, and a summary of version differences, and a Merkle tree structure is used to achieve rapid verification between versions. In terms of version synchronization protocols, a version synchronization protocol based on a distributed consistency algorithm (such as the Paxos algorithm) is deployed in a distributed database to coordinate data consistency between multiple nodes and ensure that the archive data versions on different nodes are consistent.

[0056] In terms of access logs and anomaly detection, detailed logs of user access operations are recorded, including access time, request parameters, decryption operations, and data transmission paths, and the logs are encrypted and stored in an independent audit database. In a government archive management system, every user access operation is recorded in detail and encrypted to prevent log tampering. An anomaly detection model based on the combination of isolation forest and long short-term memory network is constructed to analyze the access frequency, permission abuse mode, and decryption failure events in the log data in real time to generate anomaly scores. When the anomaly score exceeds the preset threshold, an adaptive response mechanism is triggered, such as temporarily freezing the account to prevent illegal users from continuing to access; increasing the key rotation frequency to improve system security; and starting the data self-destruction protocol to protect important data from being leaked in extreme cases.

[0057] In terms of cross-institutional archive sharing, a consortium chain network is built, and each participating institution joins the network as a node, and defines data sharing rules and permission strategies through smart contracts. In the cross-hospital archive sharing scenario in the medical industry, different hospitals act as consortium chain nodes, and use smart contracts to specify which archives can be shared and the access rights of different hospitals to different types of archives. Attribute-based encryption technology is used to achieve fine-grained access control, match user attributes with archive access policies, and dynamically generate decryption credentials. For example, for a patient's medical record, the corresponding decryption credentials are generated based on the attributes of the doctor's hospital, department, title, and the confidentiality level of the archive. Only qualified doctors can access specific medical records. Introducing homomorphic encryption technology in the sharing process allows third parties to perform statistical analysis on encrypted archive data without decryption. For example, when conducting medical big data research, third-party research institutions can perform disease incidence statistics and other analyses on encrypted medical record data without obtaining patient privacy information, and generate aggregated results and return them to the requesting party, which not only protects patient privacy but also realizes data value mining.

[0058] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0059] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. The archive management method based on AI and encrypted storage is characterized by: include: Preprocessing the original archival information through a natural language processing algorithm, including text segmentation, semantic entity recognition, and redundant information elimination, to generate structured archival data; Performing multi-level classification on the structured archive data based on a dynamic classification tree model, wherein the dynamic classification tree model constructs classification nodes through a hierarchical clustering algorithm and dynamically adjusts the classification level based on archive subject similarity; The classified structured archive data is encrypted using a hybrid encryption technology, wherein the hybrid encryption technology includes a combination of a symmetric encryption algorithm and an asymmetric encryption algorithm, wherein the archive metadata is encrypted using an AES algorithm, and the archive access permission key is encrypted using an RSA algorithm; an encrypted archive data block is constructed, each data block including the encrypted archive content, the encrypted metadata signature, and a hash chain verification code; The encrypted archive data blocks are stored in a distributed database, which adopts a sharding storage mechanism, allocates data blocks in shards based on archive classification labels and access frequencies, and introduces a metadata index table based on blockchain to achieve location mapping and integrity verification of data blocks through smart contracts; According to the keyword information in the user access request, the distributed database is searched in parallel using a multimodal search model, the multimodal search model combines text keywords, semantic vectors and classification labels to build a joint index, and matches the location of the target archive data block through an approximate nearest neighbor search algorithm; Decrypting the retrieved encrypted archive data block based on a dynamic decryption strategy, wherein the dynamic decryption strategy adaptively selects a decryption process according to the user authority level and the access scenario, including single-layer decryption, multi-layer decryption and time-limited decryption; The decrypted archive data is generated and transmitted to the user terminal through a secure channel.

2. The method according to claim 1, characterized in that Preprocessing of original archival information through natural language processing algorithms includes: Construct a text cleaning model, which uses a combination of regular expression matching and deep learning to identify and remove irrelevant characters, repeated paragraphs, and format noise in archive texts; The attention-enhanced bidirectional long short-term memory network model is used for semantic entity recognition to extract key entities, timestamps, and topic tags from archival texts and generate entity-relationship graphs. Based on the graph convolutional network, redundant nodes are pruned on the entity-relationship graph, and low-importance nodes and edges are removed by calculating the semantic similarity and connection weights between nodes to generate streamlined semantic structure data; The semantic structure data is aligned and mapped with the original archive text to generate structured archive data containing metadata description.

3. The method according to claim 1, characterized in that Multi-level classification of structured archival data based on dynamic classification tree model includes: Initialize the root node of the classification tree and use TF-IDF weighted word vectors and Doc2Vec document vectors to construct archive feature representation; The initial classification level is generated by the hierarchical agglomerative clustering algorithm, the number of clusters is optimized based on the silhouette coefficient, and the topic similarity between adjacent levels is calculated; Introducing a dynamic split and merge mechanism. When the newly added archive data causes the level similarity to be lower than the preset threshold, the node split operation is triggered; when the similarity between levels is higher than the preset threshold, the node merge operation is triggered. A unique classification code is generated for each classification node, wherein the classification code includes a hierarchical path, a subject identifier, and a version number, and the classification code is embedded into the metadata of the structured archive data.

4. The method according to claim 1, characterized in that: The hybrid encryption technology is used to encrypt the classified structured archive data, including: The archive content is processed in blocks, and encryption priorities are divided according to the size and sensitivity of the data blocks. High-priority data blocks are encrypted using AES-256, and low-priority data blocks are encrypted using AES-128. Generate a random initialization vector and key derivation parameters for each data block, and generate a data block-specific encryption key based on the PBKDF2 algorithm and the user master key; The archive metadata is encrypted using the RSA-OAEP algorithm, an encrypted metadata signature is generated, and the public key hash value is embedded into the data block header; Construct a hash chain verification code, and iteratively calculate the data block content and the hash value of the adjacent data block through the SHA-3 algorithm to generate a chain integrity check code.

5. The method according to claim 1, characterized in that Storing encrypted archive data blocks in a distributed database includes: Build a data sharding strategy based on classification labels and access frequency, allocate high-frequency access data blocks to low-latency storage nodes, and low-frequency access data blocks to high-capacity storage nodes; Deploy a metadata index smart contract in the blockchain network, which records the data block hash value, storage location, and access permission policy, and implements privacy-preserving location query through zero-knowledge proof; Erasure coding technology is used to redundantly encode data blocks, and the encoded data blocks are stored in multiple geographically distributed storage nodes. Data availability is maintained through regular heartbeat detection.

6. The method according to claim 1, characterized in that Parallel retrieval using multimodal retrieval models includes: Constructing a joint index structure, wherein the joint index includes an inverted index, a vector index, and a graph index, corresponding to text keywords, semantic vectors, and classification labels, respectively; Perform semantic expansion on user query keywords and generate synonym sets and related concept sets based on the Word2Vec model and knowledge graph; The quantized approximate nearest neighbor algorithm is used to quickly search the vector index, and a multi-level search tree is constructed through product quantization and hierarchical clustering; Integrate the exact matching results of the inverted index, the similarity ranking results of the vector index, and the association path results of the graph index to generate a comprehensive retrieval score and filter the top N target data block identifiers.

7. The method according to claim 1, characterized in that Decrypting encrypted archive data blocks based on dynamic decryption strategies includes: Decryption permissions are divided according to user permission levels. Users with high permissions trigger a single-layer decryption process and directly use the master key to decrypt data blocks. Users with intermediate permissions trigger a multi-layer decryption process and need to verify the temporary token and dynamically generated session key in turn. In the time-limited decryption scenario, the decryption key is bound to a timeliness parameter and a time-lock-based encryption scheme is adopted. The key automatically becomes invalid after the preset time. Distributed decryption is achieved through a secure multi-party computing protocol, which divides the decryption task into multiple participating nodes. Each node collaborates to complete the decryption operation based on a secret sharing protocol.

8. The method according to claim 1, characterized in that The method also includes the steps of updating and versioning archive data: When the archive data is modified, an incremental update package is generated based on the differential encoding algorithm, and the updated part is re-encrypted and the hash chain is rebuilt; The version snapshot mechanism is used to record historical versions. Each snapshot contains a timestamp, a modifier's signature, and a summary of version differences. The Merkle tree structure is used to achieve fast verification between versions. Deploy a version synchronization protocol in a distributed database and coordinate data consistency among multiple nodes based on a distributed consistency algorithm.

9. The method according to claim 1, characterized in that: The method also includes the steps of accessing logs and detecting anomalies: Record detailed logs of user access operations, including access time, request parameters, decryption operations, and data transmission paths, and encrypt and store the logs in an independent audit database; Build an anomaly detection model based on the combination of isolation forest and long short-term memory network to analyze the access frequency, permission abuse mode and decryption failure events in log data in real time and generate anomaly scores; When the anomaly score exceeds the preset threshold, an adaptive response mechanism is triggered, including temporarily freezing the account, increasing the key rotation frequency, and initiating a data self-destruction protocol.

10. The method according to claim 1, characterized in that The method also includes the step of cross-institutional archive sharing: Construct a consortium chain network, where each participating institution joins the network as a node and defines data sharing rules and permission strategies through smart contracts; Attribute-based encryption technology is used to achieve fine-grained access control, matching user attributes with archive access policies and dynamically generating decryption credentials; The introduction of homomorphic encryption technology in the sharing process allows third parties to perform statistical analysis on encrypted archive data without decryption, generate aggregated results and return them to the requester.

Citation Information

Patent Citations

  • Method and system for realizing safe sharing of user patient data based on skin database

    CN117786756A

  • File processing method and system based on digital information security

    CN118114301A

  • File management system and working method thereof

    CN118427157A

  • Systems for mandatory access control of secured hierarchical documents and related methods

    EP4361872A1

Cited By

  • Archive classification and identification method and system based on artificial intelligence

    CN120371789A

  • Energy storage cabin terminal remote upgrading method and system based on communication protocol optimization

    CN120474910A

  • Energy storage cabin terminal remote upgrading method and system based on communication protocol optimization

    CN120474910B

  • File encryption and decryption system and method based on message queue

    CN120512303A

  • Safety management system and method for electronic archive data

    CN120579197A