File Management System, Method and Equipment for Intelligent Miniature Archives
By introducing OCR identification, multi-dimensional coding, collaborative filtering algorithms and multi-level encryption technologies into the archive management system, the existing system's problems in intelligence, space utilization, ease of use and security are solved, and efficient and secure archive management is achieved.
Patent Information
- Application Number
- CN202411135330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-08-19
AI Technical Summary
The existing archive management system has shortcomings in terms of intelligence, space utilization, ease of use and security, which cannot meet the needs of small organizations or space-constrained scenarios, and poses potential security risks when facing complex cyber threats.
An archive management system for intelligent micro archive rooms is designed to determine the storage location of archives through OCR identification and clustering technology, and a multi-dimensional coding and improved collaborative filtering algorithm are used to search and recommend archives, and data security is ensured through multi-level encryption and access control mechanisms.
It realizes automatic classification, marking and intelligent retrieval of archives, improves the degree of automation and efficiency of the system, ensures efficient storage and fast access of archives, and provides a high level of data security.
Smart Images

Figure CN119226226B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of file management, and in particular, to a file management system, method, and device for an intelligent micro archive room. Background Art
[0002] File management systems have become an essential tool in modern organizations, playing a key role in information management, historical record preservation, and decision-making support. The main functions of a file management system are to systematically collect, organize, store, and retrieve various types of documents, records, and data to ensure the integrity, security, and accessibility of information. With the rapid development of information technology, traditional paper-based file management has gradually shifted towards digital and intelligent management, which not only improves work efficiency but also greatly enhances the information sharing and utilization capabilities. The importance of file management systems is reflected in multiple aspects: it can effectively protect important historical documents and data, provide a reliable basis for organizational decision-making, and is also an important tool to ensure organizational compliance and transparency. Existing file management system technologies mainly include digital storage, metadata management, full-text retrieval, access control, etc. The working principle of these technologies is to digitize files, add descriptive information (metadata), establish an indexing system, and set access permissions to achieve efficient management and rapid retrieval of files. However, existing technologies still face some problems, such as low efficiency in processing massive data, insufficient intelligence, difficulty in system integration, and inadequate security and privacy protection.
[0003] Currently, to solve the above problems, the industry has developed some improvement methods. For example, distributed storage and cloud computing technologies are adopted to improve data processing capabilities, artificial intelligence algorithms are introduced to enhance the intelligence of retrieval and classification, and blockchain technology is used to strengthen data security and traceability. However, these methods still have some obvious defects. First, most existing solutions still require a large amount of manual intervention and cannot achieve true intelligent and automated management. Second, these systems are often large-scale, requiring complex hardware facilities and professional technical support, and are not suitable for small organizations or scenarios with limited space. Third, existing systems still have deficiencies in data integration and cross-platform compatibility and are difficult to meet the growing needs of information sharing and collaboration. Finally, although some systems have introduced advanced security technologies, there are still potential security risks when facing increasingly complex network threats.
[0004] Therefore, there is an urgent need for a technical solution that can effectively solve the problems of existing systems in terms of intelligence, space utilization, usability, and security. Summary of the Invention
[0005] To address the deficiencies of the prior art, embodiments of the present application provide an archive management system, method, and device for an intelligent micro archive room. The present application solves the technical problems in the aspects of intelligence, space utilization, ease of use, and security that cannot be effectively solved by the prior art.
[0006] Embodiments of the present application provide an archive management system for an intelligent micro archive room, including: a storage location unit, an archive coding unit, an archive retrieval unit, an archive recommendation unit, and an encryption unit; wherein, the storage location unit is used to identify the content of the archive to be stored through OCR, and calculate the similarity between the archive to be stored and multiple target clustering centers to determine the storage location of the archive; the archive coding unit is used to construct a multi-dimensional code of the target archive according to the archive attribute information including archive type, storage location, security level, and date number; the archive retrieval unit is used to calculate the content similarity between the retrieval keyword and all archives, and output the information of one or more candidate archives based on user permissions and the multi-dimensional code of the archive; the archive recommendation unit is used to recommend associated archives of the selected archives by using an improved collaborative filtering algorithm; the encryption unit is used to encrypt the data including the archive content identified by OCR and the multi-dimensional code of the archive for storage in an offline database.
[0007] In a possible implementation, identifying the content of the archive to be stored through OCR and calculating the similarity between the archive to be stored and multiple target clustering centers to determine the storage location of the archive includes: preprocessing the image containing the content of the archive to be stored, and separately detecting the text area and the image area to obtain the structured text content of the archive; performing a primary clustering on the vectors of the structured text content of all archives based on an adaptive threshold and a dynamic merging strategy to determine the optimal number of clusters; constructing an archive relationship graph based on the primary clustering, and using a graph neural network to optimize the nodes; performing a secondary clustering on the optimized node representations through a density-based initial center selection strategy and an adaptive clustering number adjustment mechanism to obtain multiple target clustering centers; including:
[0008]
[0009] Wherein, C * represents the final optimal clustering center set, C represents the current clustering center set, z i represents the optimized node representation, c j represents the j-th clustering center, A j represents the set of nodes belonging to the j-th cluster, k represents the number of clusters, ρ(z i ) represents the density function of the node z i , σ represents the scale parameter of the density function, C initDenote the initial set of cluster centers; calculate the similarity between the file to be stored and multiple target cluster centers to select the cluster with the highest similarity to determine the storage location.
[0010] In an implementable way, calculate the content similarity between the retrieval keyword and all files, and output the information of one or more candidate files based on user permissions and the multi-dimensional encoding of the files, including: using a natural language processing model to perform intent recognition and entity extraction on the retrieval keyword; based on the recognized intent and entity, expand the retrieval query through a knowledge graph and word embedding; construct a multi-level index structure including an inverted index, a B+ tree index, and / or a locality-sensitive hashing index; calculate the comprehensive similarity with all files based on text content similarity, attribute similarity, and time correlation; define an access control policy based on user permissions and the security level of the files, and filter the retrieval results to output the information of one or more candidate files.
[0011] In an implementable way, recommend associated files of the selected files by using an improved collaborative filtering algorithm, including: constructing an interaction matrix and performing matrix factorization to introduce multiple target parameters to optimize the interaction matrix; generating a personalized file recommendation list based on the optimized interaction matrix.
[0012] In an implementable way, encrypt the data including the OCR-recognized file content and the multi-dimensional encoding of the file to store it in an offline database, including: dividing the file content into multiple blocks of a fixed size to independently encrypt each block using an improved AES-256 algorithm, and inserting padding data of a random length between each block; generating a key using a dynamic key generation mechanism and distributing the key through an identity-based encryption scheme; performing access control encryption and homomorphic encryption on data reading based on a preset scheme.
[0013] In an implementable way, calculate the comprehensive similarity with all files based on text content similarity, attribute similarity, and time correlation, including:
[0014] Among them, Similarity(d,q) represents the comprehensive similarity between document d and query q, and α, β, γ represent adjustable weight parameters used to balance the importance of the three components, IDF(q i ) represents the inverse document frequency of query term q i , tf(q i ,d) represents query term q iThe word frequency in document d, where k and b represent the parameters of the BM25+ algorithm, |d| represents the length of document d, avgdl represents the average document length, δ represents the additional parameter in the BM25+ algorithm, HammingDistance() represents calculating the Hamming distance between two encodings, Encoding() represents converting a document or query into a multi-dimensional encoding, MaxDistance represents the maximum possible distance between two encodings, λ represents the time decay factor, current_time represents the current time, and document_time represents the creation or last modification time of the document.
[0015] In a possible implementation, an interaction matrix is constructed and matrix factorization is performed to introduce multiple target parameters to optimize the interaction matrix, including: Among them, R ij represents the predicted interaction matrix strength of user i for profile j, μ represents the global average interaction strength, b i represents the bias term of user i, b j represents the bias term of profile j, p i represents the latent vector of user i, q j represents the latent vector of profile j, T ij represents the time decay factor, γ represents the weight coefficient of content similarity, S ij represents the content similarity between profile i and profile j, δ represents the weight coefficient of context features, c represents the context feature vector, w j represents the context weight vector of profile j.
[0016] In a possible implementation, it further includes a remote supervision unit; where the remote supervision unit is used to provide the superior archives room with information on the real-time operation status, environmental parameters, and / or personnel operations of the subordinate archives room, and remotely approve the corresponding operations of the subordinate archives room.
[0017] The embodiments of the present application also provide a file management method for an intelligent micro archives room, including: identifying the content of the file to be stored through OCR, and calculating the similarity between the file to be stored and multiple target clustering centers to determine the storage location of the file; constructing a multi-dimensional encoding of the target file according to the file attribute information including file type, storage location, security level, and date number; calculating the content similarity between the retrieval keyword and all files, and outputting the information of one or more candidate files based on user permissions and the multi-dimensional encoding of the file; recommending associated files of the selected file using an improved collaborative filtering algorithm; encrypting the data including the file content identified by OCR and the multi-dimensional encoding of the file for storage in an offline database.
[0018] The embodiment of the present application further provides a file management device for an intelligent micro archive room, including: a processor, a memory, and a system bus; wherein, the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the method described in the above embodiment.
[0019] In the file management system, method and device for an intelligent micro archive room provided above, the embodiment of the present application can realize automatic classification, marking and intelligent retrieval of files by deeply integrating artificial intelligence and machine learning technologies, greatly improving the automation degree and efficiency of the system; by introducing advanced data compression and storage technologies, efficient storage and fast access of massive data can be realized within a limited physical space; by adopting a multi-level encryption and access control mechanism, the absolute security of files can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic block diagram of a file management system for an intelligent micro archive room provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic flowchart of a file retrieval method provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Now, various exemplary embodiments of the present application will be described in detail with reference to the drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0024] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present application are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them. It should also be understood that in the embodiments of the present application, "a plurality" may refer to two or more, and "at least one" may refer to one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, without clear limitation or contrary indication in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after. It should also be understood that the description of each embodiment of the present application emphasizes the differences between the embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be described one by one.
[0025] Meanwhile, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and in no way constitutes a limitation to the present application and its application or use. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification. It should be noted that like reference numerals and letters in the following drawings denote like items, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0026] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0027] Figure 1A schematic block diagram of an archive management system for an intelligent micro archive provided by an embodiment of the present application. It should be understood that the system shown in the figure is exemplary rather than restrictive. This means that the involved system architecture is not limited to a specific form or design, but is presented as an example. In other words, the architecture shown in the figure can be regarded as a way of expression to clearly describe relevant concepts and relationships, and does not exclude other forms of architecture. Therefore, when interpreting the architecture in the said picture, it should be understood that the model has flexibility and diversity, and its purpose is to provide an exemplary description rather than a restrictive regulation of a specific form. An embodiment of the present application includes a storage location unit 101, an archive encoding unit 102, an archive retrieval unit 103, an archive recommendation unit 104, and an encryption unit 105.
[0028] The storage location unit 101 is used to identify the content of the archive to be stored through OCR, and calculate the similarity between the archive to be stored and multiple target clustering centers to determine the storage location of the archive.
[0029] First, preprocess the image containing the content of the archive to be stored, and detect the text area and the image area respectively to obtain the structured text content of the archive. It should be noted that the multi-modal fusion OCR recognition algorithm can be adopted in the present application. By combining three modal information of image, text, and voice, and further realizing end-to-end recognition through a deep learning model. Preprocess the input archive image, including denoising, correction, and enhancement. Use an adaptive Gaussian filter to remove noise, and then use a tilt correction algorithm based on the Hough transform to correct the image. Then, apply the contrast-limited adaptive histogram equalization (CLAHE) method to enhance the image quality.
[0030] The preprocessed image is input into an improved YOLOv5 model for text area detection. The backbone network of the YOLOv5 model adopts CSPDarknet53, and a spatial pyramid pooling (SPP) module is introduced to enhance the feature extraction ability. The detected text area is then cropped and input into the recognition model. The recognition model can adopt a bidirectional LSTM-Transformer architecture. First, use CNN to extract image features, and then input them into a bidirectional LSTM network to capture context information. The output of the LSTM passes through the Transformer encoder of the multi-head attention mechanism, and finally the recognition result is obtained through CTC decoding. At the same time, the system can also integrate a voice recognition module to recognize the voice annotations in the archive. By weighted fusion of the recognition results of the three modalities, the final OCR recognition content is obtained. The fusion weights are dynamically adjusted by a reinforcement learning algorithm to adapt to different types of archives.
[0031] Next, based on the adaptive threshold and dynamic merging strategy, perform a clustering on the vectors of the structural text content of all archives to determine the optimal number of clusters. Preprocess the archive content obtained by OCR recognition, including word segmentation, stop word removal, and stemming. Then use the Word2Vec model to convert the text into vector representations. Adopt an improved hierarchical clustering algorithm to perform a preliminary clustering on the archives. In one embodiment, it may include:
[0032]
[0033] Among them, C * represents the result of the first clustering, k represents the number of clusters, C i represents the i-th cluster, x represents a data point in the cluster, μ i represents the center point of the i-th cluster, var(C i ) represents the variance of the i-th cluster, β represents the parameter controlling the change rate of the adaptive weight, θ represents the variance threshold, Euclidean(x, μ i ) represents the Euclidean distance, DTW(x, μ i ) represents the dynamic time warping distance, n represents the total number of data points, γ represents the regularization strength parameter, k * represents the optimal number of clusters determined by the elbow method, Merge(C i , C j ) represents the operation of merging clusters C i and C j , diam(C i ) represents the diameter of cluster C i , and δ represents the merging threshold. This algorithm introduces an adaptive threshold and a dynamic merging strategy, and can automatically determine the optimal number of clusters.
[0034] Construct an archive relationship graph based on the first clustering, and use a graph neural network to optimize the nodes. On the basis of the preliminary clustering, construct an archive relationship graph, with each archive as a node in the graph, and the edge weights between the nodes are determined by the archive content similarity. Then apply a graph neural network (GNN) for further feature learning and clustering optimization. The GNN model adopts the GraphSAGE architecture, which includes an attention mechanism and skip connections.
[0035] Secondly, based on the node representations learned by the GNN, use an improved K-Means++ algorithm to obtain the final clustering result. The improvement points include performing a secondary clustering on the optimized node representations through a density-based initial center selection strategy and an adaptive clustering number adjustment mechanism to obtain multiple target cluster centers; among which include:
[0036]
[0037] Among them, C * represents the set of final optimal clustering centers, C represents the set of current clustering centers, and z i represents the optimized node representation, c j represents the j-th clustering center, A j represents the set of nodes belonging to the j-th cluster, k represents the number of clusters, ρ(z i ) represents the density function of node z i , σ represents the scale parameter of the density function, and C init represents the set of initial clustering centers;
[0038] After obtaining the clustering results, calculate the similarity between the file to be stored and multiple target clustering centers to select the cluster with the highest similarity to determine the storage location. The similarity calculation can adopt a weighted combination of cosine similarity and edit distance: Similarity(doc,cluster) = α·cos(doc,cluster_center)+(1-α)·(1-normalized_edit_distance(doc,cluster_center)), where α is an adjustable weight parameter.
[0039] The file encoding unit 102 is used to construct a multi-dimensional encoding of the target file according to the file attribute information including file type, storage location, security level, and date number. Specifically, use improved Huffman coding to encode the file type to obtain a type encoding; perform position stratification based on the clustering results obtained by the storage positioning unit to obtain a position encoding; calculate the information entropy of the file content and map it to the corresponding security level to obtain a security level encoding; use improved Gray code to encode the date to obtain a date encoding; combine the type encoding, position encoding, security level encoding, and date encoding, and compress the combined encoding based on the context to obtain a multi-dimensional encoding.
[0040] In an implementation scenario, based on the storage location information obtained by the storage positioning unit, combined with other attributes of the file, an innovative multi-dimensional encoding scheme is constructed. This scheme adopts a hierarchical encoding structure and includes the following levels:
[0041] Type encoding: Use improved Huffman coding to encode the file type. First, count the occurrence frequencies of each type, and then construct a Huffman tree. Different from traditional Huffman coding, this scheme introduces a dynamic adjustment mechanism that can automatically update the encoding according to newly added file types.
[0042] Storage location encoding: Based on the clustering results obtained from the storage location units, a hierarchical location encoding scheme is designed. First, the clusters are numbered, and then further subdivision is carried out within each cluster. The encoding structure is: [cluster number][sub-region number][specific location].
[0043] Security level encoding: An encoding scheme based on information entropy is adopted. First, the information entropy of the file content is calculated, and then the entropy value is mapped to different security levels. The encoding length is proportional to the security level. Files with a high security level have longer encodings to increase security.
[0044] Date number encoding: An improved Gray code is used to encode the date. Compared with traditional binary encoding, the Gray code has only one bit difference between adjacent dates, which is convenient for date range queries.
[0045] Finally, the above layers of encoding are combined into the final multi-dimensional encoding. To improve the compression ratio of the encoding, a context-based compression algorithm can be adopted. This algorithm utilizes the correlation between file attributes and reduces redundant information through a prediction model:
[0046] Final_encoding = Compress(Concatenate(E type ,E location ,E security ,E date )); Context-based compression algorithm: Compress(x) = x - Predict(x|context).
[0047] The file retrieval unit 103 is used to calculate the content similarity between the retrieval keywords and all files, and output the information of one or more candidate files based on the user permissions and the multi-dimensional encoding of the files. It includes: using a natural language processing model to perform intent recognition and entity extraction on the retrieval keywords; based on the recognized intent and entities, expanding the retrieval query through a knowledge graph and word embedding; constructing a multi-level index structure including an inverted index, a B+ tree index, and / or a locality-sensitive hashing index; calculating the comprehensive similarity with all files based on text content similarity, attribute similarity, and time correlation; defining an access control policy based on user permissions and the security level of the files, and screening the retrieval results to output the information of one or more candidate files (as Figure 2 shown, the specific method will be elaborated Figure 2 here).
[0048] The file recommendation unit 104 is used to recommend associated files of the selected files by adopting an improved collaborative filtering algorithm. In one embodiment, it includes: constructing an interaction matrix and performing matrix decomposition to introduce multiple objective parameters to optimize the interaction matrix; including: Among them, R ij represents the predicted interaction matrix strength of user i for profile j, μ represents the global average interaction strength, and b i represents the bias term of user i, and b j represents the bias term of profile j, p i represents the latent vector of user i, and q j represents the latent vector of profile j, T ij represents the time decay factor, γ represents the weight coefficient of content similarity, and S ij represents the content similarity between profile i and profile j, δ represents the weight coefficient of context features, c represents the context feature vector, and w j represents the context weight vector of profile j. Then, a personalized profile recommendation list is generated based on the optimized interaction matrix.
[0049] Specifically, this unit proposes a multi-factor dynamic collaborative filtering algorithm, which combines content analysis, user behavior, and time factors to achieve personalized profile recommendations. First, a user-profile interaction matrix R is constructed, where R ij represents the interaction strength of user i for profile j. The interaction strength considers multiple factors: R ij = w1·View_count + w2·Download_count + w3·Edit_time + w4·Cite_count, where w1, w2, w3, w4 are weight parameters that are dynamically adjusted through machine learning algorithms. An improved matrix factorization algorithm is adopted, and the bias terms of users and profiles are introduced: where μ is the global average, and b i and b j are the bias terms of users and profiles respectively, and p i and q j are the latent vectors of users and profiles. The objective function can be set as: The above objective function is optimized using Stochastic Gradient Descent (SGD).
[0050] Next, a time decay factor is introduced to make recent interaction behaviors have higher weights: T ij = exp(-β·(current_time - interaction_time)), and the time factor is integrated into the matrix factorization model: Based on the content features of profiles, the similarity matrix S between profiles is calculated. A hybrid representation of TF-IDF and Word2Vec is adopted:
[0051] Doc_vector = α·TF-IDF(doc) + (1 - α)·Word2Vec(doc);
[0052] S ij = cosine_similarity(Doc_vector i , Doc_vector j ),
[0053] Integrate the content similarity into the recommendation model:
[0054] Secondly, perform context awareness, consider the user's current context, such as work tasks, time, etc. Construct the context feature vector c and expand the model: where w j is the context weight vector of file j.
[0055] During the training process, multi-task learning can be used to optimize multiple objectives simultaneously, including rating prediction, file classification, and label prediction. Construct the multi-task loss function: L = λ1·L rating + λ2·L classification + λ3·L tagging , where L rating is the mean squared error of rating prediction, L classification is the cross-entropy loss of file classification, and L tagging is the F1 score of label prediction.
[0056] The incremental learning algorithm can also be used to update the model parameters in real time. When there is new interaction data, the parameters are updated using the stochastic gradient descent method: where η is the learning rate, which is adaptively adjusted by the Adam optimizer.
[0057] Based on the above model, a personalized file recommendation list is generated for each user. Adopt a strategy that balances diversity and novelty: score(i, j) = α·R ij + β·diversity(j) + γ·novelty(j), where diversity(j) measures the difference between file j and the recommended files, and novelty(j) measures the novelty of file j.
[0058] In addition, recommendation explanations can also be generated to help users understand the reasons for the recommendations. In one embodiment, an explanation generation model based on the attention mechanism is adopted:
[0059] explanation = Attention(user_embedding, item_embedding, context).
[0060] The encryption unit 105 is used to encrypt the data including the OCR-recognized file content and the multi-dimensional encoding of the file, so as to store it in the offline database.
[0061] First, the file content is divided into multiple blocks of a fixed size, and each block is independently encrypted using an improved AES-256 algorithm, and padding data of a random length is inserted between each block. A dynamic key generation mechanism is used to generate keys, and key distribution is performed through an identity-based encryption scheme. The key is generated by the following formula: K = H(user_id ∥ timestamp ∥ random_seed), where H is a secure hash function, and || represents string concatenation. Padding data of a random length is inserted between each block to increase the randomness of the ciphertext. The padding length L can be determined by the following formula: L = (block_index · 17 + timestamp) mod max_padding_length.
[0062] Second, access control encryption and homomorphic encryption are performed on data reading. For access control encryption, an attribute-based encryption (ABE) scheme can be used to embed the access policy into the ciphertext. Define a set of attribute sets A, and specify the access structure T during encryption. Only users who meet T can decrypt:
[0063] CT = ABE.Encrypt(PK, M, T);
[0064]
[0065] For homomorphic encryption, to support data analysis in the encrypted state, a partially homomorphic encryption scheme can be introduced. An improved Paillier encryption algorithm is used to support additive homomorphism and finite-degree multiplicative homomorphism:
[0066] E(m1 + m2) = E(m1) · E(m2);
[0067] E(k · m) = E(m) k 。
[0068] Through the above multi-level and multi-dimensional encryption scheme, the system can achieve all-round protection of file data, ensuring the confidentiality, integrity and availability of the data.
[0069] In addition, the embodiment of the present application may further include configuring a remote supervision unit 106; wherein the remote supervision unit 106 is used to provide the superior archives with information on the real-time operation status, environmental parameters and / or personnel operations of the subordinate archives, and remotely approve the corresponding operations of the subordinate archives. This unit can be configured according to relevant levels.
[0070] Specifically, the remote supervision function aims to achieve effective supervision and management of subordinate units (such as township archives) by superior units (such as district archives bureaus). Through network connection, this function enables the superior unit to grasp the operation status and file management situation of the subordinate archives in real time and conduct necessary intervention and guidance.
[0071] For real-time monitoring of the operation status, environmental parameters, and personnel operations of subordinate archives. Real-time video monitoring can be carried out through cameras installed in the archives, environmental sensors can be deployed to monitor parameters such as temperature, humidity, and light, and the operation logs of the file management system can be recorded. A time series analysis model can also be used to detect abnormal situations. For example, the autoregressive moving average model (ARMA) can be used to predict the normal values of relevant parameters. If the difference between the actual observed value and the predicted value exceeds the preset threshold, an alarm is triggered. In one embodiment, the district archives bureau discovers through the monitoring system that the temperature in the archives of Township A has risen abnormally. The system automatically triggers an alarm, and the district archives bureau immediately notifies Township A to conduct an inspection, promptly discovers and repairs the air conditioner failure, and avoids possible damage to the files.
[0072] Furthermore, remote approval is carried out for important operations of subordinate archives. The types of operations that require remote approval can be set, such as access to confidential files and destruction of a large number of files. The superior administrator receives the approval request through the system and conducts remote review, and the approval result is feedback to the subordinate archives in real time.
[0073] The approval process model can include, but is not limited to, using a finite state machine (FSM) to control the approval process:
[0074] The state set S = {to be submitted, to be approved, in preliminary review, in review, in final review, approved, rejected, withdrawn, expired};
[0075] The input set I = {submit application, withdraw application, preliminary review passed, preliminary review rejected, review passed, review rejected, final review passed, final review rejected, timeout};
[0076] The transition function δ(s, i) can be defined as follows:
[0077] δ(to be submitted, submit application) = to be approved;
[0078] δ(to be approved, withdraw application) = withdrawn;
[0079] δ(to be approved, preliminary review passed) = in review;
[0080] δ(to be approved, preliminary review rejected) = rejected;
[0081] δ(in review, review passed) = in final review;
[0082] δ(In review, review rejected) = Rejected;
[0083] δ(In final review, final review passed) = Passed;
[0084] δ(In final review, final review rejected) = Rejected;
[0085] δ(Any status, timeout) = Expired.
[0086] Combined with the above model, the approval process can be configured to allow multiple approvers to review simultaneously, and only when all approvals are passed can it enter the next stage. According to the application content, different approval paths are selected. Different roles (such as applicant, approver, administrator) and their operation permissions are defined. Text comments and attachments can be added to each approval step. Automatic reminders and follow-ups are set, and reminders are automatically sent for applications approaching timeout. A delegation approval mechanism is introduced, and the approver can temporarily delegate the permission to others. Approval history records are added to record information such as the operator, time, and comments of each step. The approval rollback function can also be set to allow the application to be rolled back to a previous step for re-approval. At the same time, approval templates are introduced, and common approval process templates are preset to facilitate quick creation. For example, Township C applies to destroy a batch of expired archives. The District Archives Bureau can review the destruction list through remote approval and supervise the entire destruction process through video monitoring to ensure compliance operations.
[0087] Through the remote supervision function, the supervision efficiency can be improved, 24 / 7 all-weather supervision is achieved, greatly improving the timeliness and coverage of supervision. Further, the supervision cost can be reduced, the frequency of on-site inspections is reduced, and manpower and travel costs are saved. Further, standardized management can be promoted. Through a unified remote supervision platform, the standardization and regularization of file management work are promoted. The emergency response ability can also be strengthened, rapid response and collaborative disposal are achieved, and the ability to respond to emergencies is improved. Optimize resource allocation. Based on data analysis, more accurate resource allocation and management decisions are achieved. Improve the quality of personnel. Through remote guidance, the professional ability of grass-roots file management personnel is continuously improved.
[0088] Figure 2 It is a flow schematic diagram of a file retrieval method provided by an embodiment of the present application. This unit adopts a multi-stage retrieval algorithm, combines semantic understanding, similarity calculation, and permission control to achieve efficient and accurate file retrieval. As Figure 2 shown, at step S201, a natural language processing model is used to perform intent recognition and entity extraction on the retrieval keywords.
[0089] First, semantic understanding is performed on the retrieval keywords input by the user. An improved BERT model is used for intent recognition and entity extraction. The improvement points include but are not limited to introducing domain-specific pre-training tasks and multi-task learning frameworks.
[0090] At step S202, based on the recognized intent and entity, the retrieved query is expanded through a knowledge graph and word embeddings. Specifically, a hybrid method of the knowledge graph and word embedding is used for the expansion of synonyms and hypernyms / hyponyms. The knowledge graph is constructed using Neo4j and contains professional knowledge in the archival field. Word embedding uses the FastText model, which can handle out-of-vocabulary words. The mathematical expression for query expansion is as follows:
[0091] Q expanded = Q original ∪ KG expand (Q original ) ∪ WE expand (Q original ),
[0092] where Q expanded represents the expanded query set, which contains the original query terms and the related terms obtained through expansion. Q original represents the set of original query terms input by the user. KG expand (Q original ) represents the query expansion based on the knowledge graph. The knowledge graph of the archival field constructed using Neo4j is used to find related terms such as synonyms, hypernyms, and hyponyms through the graph relationships. WE expand (Q original ) represents the query expansion based on word embedding. The FastText model is used for word embedding to find the words similar to the original query terms in the vector space. ∪ represents the set union operation, which combines the original query terms with the terms obtained by the two expansion methods.
[0093] The above formula describes the process of query expansion, that is, combining the original query with the related terms expanded by the two methods of the knowledge graph and word embedding to form the final expanded query set.
[0094] At step S203, a multi-level index structure including an inverted index, a B+ tree index, and / or a locality-sensitive hashing (LSH) index is constructed. In one embodiment:
[0095] Inverted index: Used to quickly locate the archives containing the query terms.
[0096] B+ tree index: Constructed based on the multi-dimensional encoding of the archives and supports range queries.
[0097] LSH index: Used for approximate nearest neighbor search to accelerate similarity calculation.
[0098] The retrieval process adopts a strategy of filtering layer by layer. First, the inverted index is used to screen out the candidate set, then the B+ tree index is used for encoding range filtering, and finally the LSH index is used to calculate the similarity.
[0099] At step S204, the comprehensive similarity with all files is calculated based on the text content similarity, attribute similarity, and time correlation. In one embodiment, it includes:
[0100]
[0101] Among them, Similarity(d,q) represents the comprehensive similarity between document d and query q, and α, β, γ represent adjustable weight parameters used to balance the importance of the three components. IDF(q i ) represents the inverse document frequency of query term q i , tf(q i ,d) represents the term frequency of query term q i in document d, k, b represent the parameters of the BM25+ algorithm, |d| represents the length of document d, avgdl represents the average document length, δ represents the additional parameter in the BM25+ algorithm, HammingDistance() represents calculating the Hamming distance between two encodings, Encoding() represents converting a document or query into a multi-dimensional encoding, MaxDistance represents the maximum possible distance between two encodings, λ represents the time decay factor, current_time represents the current time, and document_time represents the creation or last modification time of the document.
[0102] Specifically, the similarity calculation method that fuses multiple features in the embodiments of the present application considers the text content similarity, attribute similarity, and time correlation.
[0103] The text similarity adopts an improved BM25 algorithm:
[0104]
[0105] The attribute similarity is calculated based on the multi-dimensional encoding of the file:
[0106]
[0107] The time correlation considers the timeliness of the file:
[0108] Time_relevance(d) = exp(-λ·(current_time - document_time)),
[0109] The final similarity is the weighted sum of the three:
[0110] Similarity(d,q) =
[0111] α·BM25 + (d,q) + β·Sim attr (d,q) + γ·Time_relevance(d),
[0112] where α, β, and γ are adjustable weight parameters.
[0113] At step S205, an access control policy is defined based on user permissions and the security level of the file, and the retrieval results are filtered to output the information of one or more candidate files. Specifically, the Attribute-Based Access Control (ABAC) model is used to define the access control policy by combining the advantages of Role-Based Access Control (RBAC):
[0114] Policy(user,document) =
[0115] f(user_attributes,document_attributes,environment_attributes), where f is a decision function that comprehensively considers user attributes, document attributes, and environmental attributes. Then, the Learning to Rank method is used to rank the retrieval results. The LambdaMART algorithm is used to train the ranking model, and the features include similarity scores, document attributes, user historical behaviors, etc. Finally, according to the ranking results and user permissions, the final retrieval results are filtered and displayed.
[0116] The file management system of an intelligent micro archive room disclosed in the above content of this application is expected to effectively solve the problems of the existing system in terms of intelligence, space utilization, usability, and security. It will provide a highly intelligent, secure, and easy-to-use file management solution for small organizations and scenarios with limited space, bringing significant progress to the field of file management. Such a system can not only meet the current needs but also lay a solid foundation for future information management and knowledge sharing, promoting the entire industry to develop in a more intelligent and efficient direction.
[0117] Furthermore, the embodiment of this application also provides a file management device for an intelligent micro archive room, including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any of the above methods.
[0118] Furthermore, an embodiment of the present application also provides a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute any one of the above methods.
[0119] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0120] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0121] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0122] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent micro-archive room file management system, characterized in that: include: storage positioning unit, file encoding unit, file retrieval unit, file recommendation unit and encryption unit; wherein, The storage positioning unit is used to identify the content of the archive to be stored through OCR, and calculate the similarity between the archive to be stored and multiple target cluster centers to determine the storage location of the archive, including: Preprocessing the image containing the archive content to be stored, and detecting the text area and the image area respectively, so as to obtain the structural text content of the archive; Cluster the vectors of the structured text content of all archives based on adaptive thresholds and dynamic merging strategies to determine the optimal number of clusters; Construct an archive relationship graph based on primary clustering and use graph neural network to optimize nodes; The optimized node representation is clustered again through the density-based initial center selection strategy and the adaptive cluster quantity adjustment mechanism to obtain multiple target cluster centers; including: Among them, C * represents the final optimal cluster center set, C represents the current cluster center set, z i represents the optimized node representation, c j represents the jth cluster center, A j represents the set of nodes belonging to the jth cluster, k represents the number of clusters, ρ(z i ) represents node z i The density function, σ represents the scale parameter of the density function, C init represents the initial cluster center set; Calculate the similarity between the archive to be stored and multiple target cluster centers, and select the cluster with the highest similarity to determine the storage location; The archive encoding unit is used to construct a multi-dimensional code of the target archive according to the archive attribute information including archive type, storage location, security level and date number; The archive retrieval unit is used to calculate the similarity between the retrieval keyword and the content of all archives, and output the information of one or more archives to be selected based on the user authority and the multi-dimensional code of the archive; The file recommendation unit is used to recommend related files of the selected file by using an improved collaborative filtering algorithm; The encryption unit is used to encrypt data including the archive content recognized by OCR and the archive multi-dimensional encoding, so as to store them in an offline database.
2. The file management system according to claim 1, characterized in that: in, Calculate the similarity between the search keyword and the content of all archives, and output the information of one or more archives to be selected based on user permissions and the multi-dimensional coding of the archives, including: Use natural language processing models to identify the intent and extract entities of search keywords; Based on the identified intents and entities, the retrieved queries are expanded through knowledge graphs and word embeddings; Construct a multi-level index structure including inverted index, B+ tree index and / or locality sensitive hash index; Calculate the comprehensive similarity with all archives based on text content similarity, attribute similarity and time correlation; The access control policy is defined based on the user's authority and the security level of the archive, and the search results are filtered to output the information of one or more archives to be selected.
3. The file management system according to claim 1, characterized in that: in, An improved collaborative filtering algorithm is used to recommend related files of the selected files, including: constructing an interaction matrix and performing matrix decomposition to introduce multiple objective parameters to optimize the interaction matrix; Generate a personalized profile recommendation list based on the optimized interaction matrix.
4. The file management system according to claim 1, characterized in that: in, Encrypt the data including the OCR-recognized archive content and the archive multi-dimensional encoding to store in the offline database, including: Divide the archive content into multiple fixed-size blocks, encrypt each block independently using the modified AES-256 algorithm, and insert padding data of random length between each block; A dynamic key generation mechanism is used to generate keys, and key distribution is performed through an identity-based encryption scheme; Access control encryption and homomorphic encryption are performed on data reading based on preset schemes.
5. The file management system according to claim 2, characterized in that: in, The comprehensive similarity with all archives is calculated based on text content similarity, attribute similarity and time correlation, including: Among them, Similarity(d,q) represents the comprehensive similarity between document d and query q, α, β, and γ represent adjustable weight parameters used to balance the importance of the three components, and IDF(q i ) represents the query word q i The inverse document frequency, tf(q i ,d) represents the query word q i The term frequency in document d, k, b represent the parameters of the BM25+ algorithm, |d| represents the length of document d, avgdl represents the average document length, δ represents the additional parameters in the BM25+ algorithm, HammingDistance() represents the calculation of the Hamming distance between two encodings, Encoding() represents the conversion of a document or query into a multidimensional encoding, MaxDistance represents the maximum possible distance between two encodings, λ represents the time decay factor, current\_time represents the current time, and document\_time represents the creation or last modification time of the document.
6. The file management system according to claim 3, characterized in that: in, An interaction matrix is constructed and matrix decomposition is performed to introduce multiple objective parameters to optimize the interaction matrix, including: Among them, R ij represents the predicted interaction matrix strength of user i to profile j, μ represents the global average interaction strength, and b i represents the bias term of user i, b j represents the bias term of file j, p i represents the latent vector of user i, q j represents the latent vector of file j, T ij represents the time decay factor, γ represents the weight coefficient of content similarity, S ij represents the content similarity between archive i and archive j, δ represents the weight coefficient of context feature, c represents the context feature vector, w j represents the context weight vector of profile j.
7. The file management system according to claim 1, characterized in that: It also includes a remote monitoring unit; wherein the remote monitoring unit is used to provide the superior archive room with the real-time operating status, environmental parameters and / or personnel operation information of the subordinate archive room, and remotely approve the corresponding operations of the subordinate archive room.
8. A method for managing archives in an intelligent micro-archive room, characterized in that: include: Identify the content of the archive to be stored through OCR, and calculate the similarity between the archive to be stored and multiple target cluster centers to determine the storage location of the archive, including: Preprocessing the image containing the archive content to be stored, and detecting the text area and the image area respectively, so as to obtain the structural text content of the archive; Cluster the vectors of the structured text content of all archives based on adaptive thresholds and dynamic merging strategies to determine the optimal number of clusters; Construct an archive relationship graph based on primary clustering and use graph neural network to optimize nodes; The optimized node representation is clustered again through the density-based initial center selection strategy and the adaptive cluster quantity adjustment mechanism to obtain multiple target cluster centers; including: Among them, C * represents the final optimal cluster center set, C represents the current cluster center set, z i represents the optimized node representation, c j represents the jth cluster center, A j represents the set of nodes belonging to the jth cluster, k represents the number of clusters, ρ(z i ) represents node z i The density function, σ represents the scale parameter of the density function, C init represents the initial cluster center set; Calculate the similarity between the archive to be stored and multiple target cluster centers, and select the cluster with the highest similarity to determine the storage location; Constructing a multi-dimensional code of a target archive based on archive attribute information including archive type, storage location, security level and date number; Calculate the similarity between the search keyword and the content of all archives, and output the information of one or more archives to be selected based on the user's authority and the multi-dimensional coding of the archives; An improved collaborative filtering algorithm is used to recommend related files of the selected files; The data including the archive content recognized by OCR and the archive multi-dimensional coding are encrypted and stored in the offline database.
9. An intelligent micro-archive room file management device, characterized in that: include: A processor, a memory, and a system bus; wherein the processor and the memory are connected via the system bus; The memory is used to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to perform the method of claim 8.
Citation Information
Patent Citations
Archive data storage system based on big data
CN117725283A
Intelligent archive indexing and retrieval system
CN117909440A