File management system and method based on vertical large model

Through the archive management system based on vertical large models, multimodal data analysis and semantic analysis technology, the problems of low security and unsatisfactory query efficiency of the existing electronic archive management system are solved, and efficient and secure archive management and retrieval are achieved.

CN120371779AInactive Publication Date: 2025-07-25GUANGZHOU LONGJIANDA ELECTRONICS CO LTD

Patent Information

Application Number
CN202510864274.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing electronic file management system has low security, unsatisfactory query management efficiency, and is prone to errors in file information correction.

Method used

The archive management system based on vertical large models is adopted, including automatic sorting module, cleaning and archiving module, knowledge graph module, intelligent editing module, semantic search module, face recognition module and prediction management module. Through multimodal data analysis and semantic analysis technology, the efficiency and security of archive management are improved.

Benefits of technology

It improves the efficiency of archive sorting and management, reduces manual intervention, provides diversified data processing and analysis functions, enhances the accuracy and speed of archive retrieval, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371779A_ABST
    Figure CN120371779A_ABST
Patent Text Reader

Abstract

The invention relates to the field of archive management, in particular to an archive management system and method based on a vertical large model. The system comprises an automatic arrangement module, a cleaning and archiving module, a knowledge graph module, a knowledge base module, an intelligent editing and research module, a semantic retrieval module, a face recognition module, an association module and a prediction management module. Through multiple functional modules, the efficiency of archive arrangement and management can be effectively improved, and manual intervention is reduced. And diversified data processing and analysis functions are provided, and deep research and data mining are supported. And the accuracy and rapidity of file retrieval are enhanced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of file management, and more specifically, to a file management system and method based on a vertical large model. Background Art

[0002] Electronic files refer to a collection of general electronic image files that are stored through devices such as computer disks, corresponding to paper files and interconnected, usually in units of volumes. With the rapid development of information technology, electronic files have become an indispensable information carrier for modern people, enterprises, and institutions. In order to manage electronic files more conveniently, currently, electronic files are managed through an electronic file management system, including management operations such as precise search, correction, and deletion of electronic files. However, the current electronic file management system has low security. After simple account password authentication by the management personnel, they can enter the electronic file management system to perform management operations on electronic files. When the account password is known to others, others can smoothly enter the electronic file management system for operations. Moreover, the current electronic file management system only allows management personnel with relevant permissions to perform modification and update operations on electronic files. The management personnel modify and update the electronic files according to the collected file information. When the file information is large, it is easy to make correction errors.

[0003] The prior art discloses a security management system and method for electronic files, including: an information data collection unit, an information data storage unit, a random code generation unit, a random code sending unit, an identity authentication unit, an electronic file operation unit, a duration automatic assignment unit, an operation duration real-time monitoring unit, and a system exit unit; the information data collection unit: used to collect information data. However, the query management efficiency of this method for files is not very ideal. Summary of the Invention

[0004] The purpose of the present invention is to disclose a file management system and method based on a vertical large model with higher query management efficiency.

[0005] To achieve the above purpose, the present invention provides a file management system based on a vertical large model, including: An automatic sorting module: automatically sort files using a vertical large model; A cleaning and archiving module: clean and archive the data in the automatic sorting module; A knowledge graph module: construct a file knowledge graph for the data in the cleaning and archiving module and store it in the knowledge base module; A knowledge base module: store files; An intelligent compilation and research module: intelligently extract relevant materials of files in the knowledge base module, perform multimodal data analysis, and automatically generate corresponding compilation and research theme materials; Semantic Retrieval Module: Utilize natural language understanding and semantic parsing technologies to convert user queries into machine - understandable representations, and sort and recommend retrieval results by calculating the semantic similarity or relevance between the queries and the content of the files in the knowledge base module; Face Recognition Module: Extract image objects from the files in the knowledge base module and perform face recognition. Manage the archival content containing face images by identifying and saving the face information in the images, as well as associating the coordinate positions and page numbers in the files; Association Module: Read and store the valid information of the files into the knowledge base, and perform matching and association to ensure the consistency and integrity of the knowledge base and the file content; Prediction and Management Module: Predict and manage the confidentiality period of the files in the knowledge base module.

[0006] Furthermore, the vertical large - model is the LLM large - language model: OCR Recognition: Convert the text in the image into machine - encoded text through optical character recognition; File Number Extraction: Extract the file number from the text recognized by optical character recognition; the file number is used to identify and track files or documents; Title Extraction: Automatically extract the title of the document; the title is important metadata of the document, which helps to understand and search for the document; Issuing Time: Identify and record the release or creation date of the document; Automatic Annotation: Based on the previous steps, automatically mark or classify the document; Metadata Storage: The file number, title, issuing time, and the content of automatic annotation will be saved as metadata.

[0007] Furthermore, the cleaning and archiving module includes: Data Collection: Data collection is the first step in the data analysis process, which involves the process of obtaining the required information; Noise Removal: The collected data often contains useless or irrelevant information, called "noise". Noise removal means deleting these unnecessary parts from the data to improve the data quality and accuracy; Text Normalization: Text normalization is the process of converting text into a unified format; it includes converting text to lowercase, deleting punctuation marks, and replacing abbreviations; Document Classification: Group the documents according to specific criteria or features; including classifying the documents according to the theme, author, date, etc.; Tag Extraction and Generation: Tags are keywords or phrases used to describe data attributes; automatically extract or generate tags from the documents to better understand and organize the data; Data Duplication Removal: If there are duplicate entries in the dataset, data duplication removal is required to ensure the accuracy of the analysis results; Data verification and evaluation: Verify and evaluate the processed data to ensure its quality and integrity; including checking data consistency, integrity, and accuracy; The text summarization algorithm generates a summary by weighted extraction of important sentences; Importance score calculation Given a document D consisting of n sentences S1, S2,..., Sn, the weight W(Si) of a sentence is calculated based on sentence length, keyword density, and similarity to the title: W(Si)=α·Length(Si)+β·KeywordDensity(Si)+γ·Sim(Si,T) Where α, β, γ are adjustment coefficients, Length(Si) is the sentence length, KeywordDensity(Si) is the keyword density, and Sim(Si,T) is the similarity between the sentence and the title T; Summary generation αβγκ The summary Ssummary consists of the top k sentences with the highest weights: Ssummary = Top-k{Si|W(Si)}.

[0008] Furthermore, the knowledge graph module includes: Entity recognition: The process of identifying specific types of words or phrases in the text, including: personal names, locations, or organization names; Relationship extraction: Extracting relationships between entities from the text; Attribute filling: Adding detailed information to entities to make them more rich and detailed; Triple generation: A triple is a structure consisting of three elements used to represent facts in the knowledge graph, including: subject, predicate, object; generating triples based on the previously extracted entities and their relationships; Graph storage: Storing the generated triples to form a knowledge graph for convenient subsequent querying and analysis; Data import: Importing other relevant data sources for combination with the current knowledge graph; Multi-source data fusion: Combining multiple data sources to obtain a more comprehensive view of information; Graph visualization: Displaying the knowledge graph in a graphical way to make it easier for users to understand the relationships between data; Graph evaluation: Evaluating the quality and effectiveness of the knowledge graph to see if it meets the requirements and whether improvement is needed.

[0009] Furthermore, the semantic retrieval module includes: Polymorphic LLM: A polymorphic large model that receives various types of data inputs, such as documents, PPTs, DOCs, images, etc.; Prompt Engineering: Prompt engineering is a method of guiding the model to perform specific tasks, which can direct the model to focus on specific information or questions; Multimodal Analysis: Multimodal analysis is to combine information from different modalities (such as text, images, audio, etc.) for analysis; Generate Necessary Information: In this step, the system generates key information about the input data; Storage: The generated information will be stored for subsequent use; Vector Database: The stored information will be put into a vector database, which is a space for storing high-dimensional data; Question Answering: Users can access the stored information by asking questions; Similarity Search: Users can also find the answers that best match their questions through similarity search.

[0010] Furthermore, the face recognition module includes: Extract image features using a Convolutional Neural Network (CNN) and perform matching through vectorized representation; Feature Extraction: Extract the feature vector f from the input image I: f = CNN(I) Similarity Matching: Given the existing face feature vectors f1, f2, …, fn in the database, calculate the similarity between the input image feature f and each face feature in the database, and return the result with the highest similarity; In the semantic retrieval module, rank the retrieval results by calculating the semantic similarity between the query Q and the profile content D; Vector Representation First, represent the query and the document as vectors q and d: q = Embedding(Q) d = Embedding(D) Similarity Calculation Calculate the semantic similarity between the query and the document using cosine similarity:

[0011] where q·d represents the dot product of the two vectors, and |q| and |d| are the norms of the vectors.

[0012] Furthermore, the prediction management module includes: Profile Classification: Classify the files; Sensitivity Analysis: Analyze by evaluating the sensitivity of the model to parameter changes;​ Model Selection and Training: Select a suitable model based on the sensitivity analysis results and perform training; Secrecy Period Attribution Analysis: Conduct attribution analysis to obtain the factors affecting the results and their weighting techniques; Secrecy Period Prediction: Predict the secrecy period of the file; LLM: The large model is responsible for processing file classification and prediction; Output Influence Factors and Weights: Output the factors affecting the secrecy period and their importance levels; Output Secrecy Period Prediction: Output the secrecy period of the file.

[0013] Furthermore, the semantic retrieval module further includes: Metadata Extraction: Based on template matching, extract metadata information from the file: Design a set of metadata templates; Parse the file using regular expressions, rule engines, or pre-trained text classification models; The extracted metadata will be associated with the source file content and stored as structured data; Metadata Vectorization: Convert the metadata into vector representations and store the associated information with the source file; Use a pre-trained embedding model to vectorize the metadata; The vectorization process will maintain the semantic integrity of the metadata while introducing the associated information of the source file as context; The vector data is stored in an efficient vector database; Dynamic Retrieval Weight Assignment: Dynamically assign retrieval weights according to the question type or intent; Retrieval Methods and Default Weights: Metadata Vector Recall: The default weight is set to 0.2; Source File Vector Recall: The default weight is set to 0.3; Metadata Inverted Index Recall: The default weight is set to 0.5; Analyze the characteristics of the input question through the question intent recognition model; Dynamically adjust the weights of the above three retrieval methods to match the retrieval target; Hybrid Retrieval: Simultaneously adopt three methods, namely metadata vector, source file vector, and metadata inverted index, for retrieval and perform result fusion based on weights; The three recall methods run independently and each returns a result set; Use a weighted fusion algorithm to combine the three result sets, considering the recall scores and weights; Adopt a re-ranking model to further optimize the combined results to ensure that highly relevant results are ranked at the top; Result Return: Return the final result set based on hybrid retrieval; Provide the top N sorted results along with their relevance scores; Mark the relevance of metadata, content, and inverted index in the returned results; Combine weighted fusion and re - ranking to ensure that the returned result set is both comprehensive and accurate.

[0014] Furthermore, the semantic retrieval module further includes: Automatically select the recall method through question intent recognition: Question metadata recognition Question metadata extraction: Extract the metadata of the user's question; Metadata structuring: Structurally process the metadata; Perform intent recognition based on metadata Intent recognition: Identify the intent of the question through a machine learning model; if it is a factual query question, use RAG vector retrieval; if it involves multi - entity relationships or complex reasoning questions, use graph RAG; for complex questions, use the recall results of both traditional RAG and graph RAG, and then comprehensively generate the final answer through a generative model.

[0015] In addition, the present invention also provides an archive management method based on a vertical large model, which is characterized by including: Automatically organize the archives using a vertical large model through an automatic organization module; Clean and archive the data from the automatic organization module through a cleaning and archiving module; Construct an archive knowledge graph for the data from the cleaning and archiving module through a knowledge graph module and store it in a knowledge base module; Store files through a knowledge base module; Intelligently extract relevant materials of the files in the knowledge base module through an intelligent compilation and research module, perform multi - modal data analysis, and automatically generate corresponding compilation and research theme materials; Through a semantic retrieval module, use natural language understanding and semantic parsing technologies to convert the user's query into a machine - understandable representation form, and sort and recommend the retrieval results by calculating the semantic similarity or relevance between the query and the file content in the knowledge base module; Extract image objects from the files in the knowledge base module through a face recognition module and perform face recognition. Manage the archive content containing face images by identifying and saving the face information in the image, as well as associating the coordinate positions and page numbers in the file; Read and store the valid information of the files into the knowledge base through an association module, and perform matching and association to ensure the consistency and integrity of the knowledge base and the file content; Predict and manage the confidentiality period of the files in the knowledge base module through a prediction management module.

[0016] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows: Through multiple functional modules, the present invention can effectively improve the efficiency of file sorting and management and reduce manual intervention. And it provides diversified data processing and analysis functions to support in-depth research and data mining. In addition, it enhances the accuracy and speed of file retrieval and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The file management system based on a vertical large model described in Embodiment 1; Figure 2 The flowchart of the file management method based on a vertical large model described in Embodiment 3; DETAILED DESCRIPTION OF THE EMBODIMENTS The drawings are only for illustrative purposes and should not be construed as a limitation of this patent; The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0018] Embodiment 1: This embodiment provides a file management system based on a vertical large model as shown in Figure 1 and includes: An automatic sorting module: automatically sorting files by using a vertical large model; A cleaning and archiving module: cleaning and archiving the data from the automatic sorting module; A knowledge graph module: constructing a file knowledge graph for the data from the cleaning and archiving module and storing it in the knowledge base module; A knowledge base module: storing files; An intelligent compilation and research module: intelligently extracting relevant materials of files from the knowledge base module, performing multi-modal data analysis, and automatically generating corresponding compilation and research theme materials; A semantic retrieval module: using natural language understanding and semantic parsing technologies to convert user queries into machine-understandable representations, and sorting and recommending retrieval results by calculating the semantic similarity or relevance between the queries and the file contents in the knowledge base module; A face recognition module: extracting image objects from files in the knowledge base module and performing face recognition, managing file contents containing face images by identifying and saving face information in the images, as well as associating coordinate positions and page numbers in the files; An association module: reading and storing the valid information of files into the knowledge base, and performing matching and association to ensure the consistency and integrity of the knowledge base and file contents; A prediction management module: predicting and managing the confidentiality period of files in the knowledge base module.

[0019] This embodiment can effectively improve the efficiency of file sorting and management through multiple functional modules, reducing manual intervention. And it provides diversified data processing and analysis functions, supporting in-depth research and data mining. In addition, it enhances the accuracy and speed of file retrieval, improving the user experience.

[0020] Embodiment Two: This embodiment further discloses on the basis of Embodiment One: Furthermore, the vertical large model is the LLM large language model: OCR Recognition: Convert the text in the image into machine-coded text through optical character recognition; File Number Extraction: Extract the file number from the text recognized by optical character recognition; the file number is used to identify and track files or documents; Title Extraction: Automatically extract the title of the document; the title is important metadata of the document, which helps to understand and search for the document; Issuing Time: Identify and record the release or creation date of the document; Automatic Annotation: Based on the previous steps, automatically mark or classify the document; Metadata Storage: The file number, title, issuing time, and the content of automatic annotation will all be saved as metadata.

[0021] Furthermore, the cleaning and archiving module includes: Data Collection: Data collection is the first step in the data analysis process, which involves the process of obtaining the required information; Noise Removal: The collected data often contains useless or irrelevant information, which is called "noise". Removing noise means deleting these unnecessary parts from the data to improve the data quality and accuracy; Text Normalization: Text normalization is the process of converting text into a unified format; it includes converting text to lowercase, deleting punctuation marks, and replacing abbreviations; Document Classification: Group documents according to specific criteria or features; including: classifying documents according to the theme, author, date, etc.; Tag Extraction and Generation: Tags are keywords or phrases used to describe data attributes; automatically extract or generate tags from the document to better understand and organize the data; Data Duplication Removal: If there are duplicate entries in the dataset, data duplication removal is required to ensure the accuracy of the analysis results; Data Verification and Evaluation: Verify and evaluate the processed data to ensure its quality and integrity; including checking the consistency, integrity, and accuracy of the data; The text summarization generation algorithm generates a summary by weighted extraction of important sentences; Importance Score Calculation Given a document D consisting of n sentences S1, S2, ..., Sn, the weight W(Si) of a sentence is calculated based on the sentence length, keyword density, and similarity to the title: W(Si) = α·Length(Si) + β·KeywordDensity(Si) + γ·Sim(Si, T) where α, β, and γ are adjustment coefficients, Length(Si) is the sentence length, KeywordDensity(Si) is the keyword density, and Sim(Si, T) is the similarity between the sentence and the title T; Abstract generation αβγκ The abstract Ssummary consists of the top k sentences with the highest weights: Ssummary = Top-k{Si|W(Si)}.

[0022] Furthermore, the knowledge graph module includes: Entity recognition: The process of identifying specific types of words or phrases in the text, including: personal names, locations, or organization names; Relationship extraction: Extracting the relationships between entities from the text; Attribute filling: Adding detailed information to entities to make them more rich and detailed; Triple generation: A triple is a structure consisting of three elements used to represent facts in the knowledge graph, including: subject, predicate, and object; generating triples based on the previously extracted entities and their relationships; Graph storage: Storing the generated triples to form a knowledge graph for convenient subsequent querying and analysis; Data import: Importing other relevant data sources to integrate with the current knowledge graph; Multi-source data fusion: Merging multiple data sources together to obtain a more comprehensive view of the information; Graph visualization: Displaying the knowledge graph in a graphical way to make it easier for users to understand the associations between data; Graph evaluation: Evaluating the quality and effectiveness of the knowledge graph to see if it meets the requirements and whether improvement is needed.

[0023] Furthermore, the semantic retrieval module includes: Polymorphic LLM: A polymorphic large model that receives various types of data inputs, such as documents, PPTs, DOCs, pictures, etc.; Prompt engineering: Prompt engineering is a method of guiding the model to perform specific tasks, which can direct the model to focus on specific information or questions; Multimodal Analysis: Multimodal analysis combines information from different modalities (such as text, images, audio, etc.) for analysis; Generate Necessary Information: In this step, the system generates key information about the input data; Storage: The generated information is stored for subsequent use; Vector Database: The stored information is placed in a vector database, which is a space for storing high-dimensional data; Question Answering: Users can access the stored information by asking questions; Similarity Search: Users can also find the answers that best match their questions through similarity search.

[0024] Furthermore, the face recognition module includes: Extract image features using a Convolutional Neural Network (CNN) and perform matching through vectorized representation; Feature Extraction: Extract the feature vector f from the input image I: f = CNN(I) Similarity Matching: Given the existing face feature vectors f1, f2, …, fn in the database, calculate the similarity between the input image feature f and each face feature in the database and return the result with the highest similarity; In the semantic retrieval module, the retrieval results are sorted by calculating the semantic similarity between the query Q and the archive content D; Vector Representation First, represent the query and the document as vectors q and d: q = Embedding(Q) d = Embedding(D) Similarity Calculation Use cosine similarity to calculate the semantic similarity between the query and the document:

[0025] where q·d represents the dot product of the two vectors, and |q| and |d| are the norms of the vectors.

[0026] Furthermore, the prediction management module includes: Archive Classification: Classify the files; Sensitivity Analysis: Analyze by evaluating the sensitivity of the model to parameter changes; Model Selection and Training: Select a suitable model based on the sensitivity analysis results and perform training; Secrecy Period Attribution Analysis: Conduct attribution analysis to obtain the techniques of the factors affecting the results and their weights; Secrecy Period Prediction: Predict the secrecy period of the files; LLM: The large language model is responsible for processing file classification and prediction; Output impact factors and weights: Output the factors affecting the confidentiality period and their importance levels; Output confidentiality period prediction: Output the confidentiality period of the file.

[0027] Furthermore, the semantic retrieval module also includes: Metadata extraction: Based on template matching, extract metadata information from files: Design a set of metadata templates; Parse files using regular expressions, rule engines, or pre-trained text classification models; The extracted metadata will be associated with the source file content and stored as structured data; Metadata vectorization: Convert metadata into vector representations and store the associated information with the source file; Use a pre-trained embedding model to vectorize the metadata; The vectorization process will maintain the semantic integrity of the metadata while introducing the associated information of the source file as context; The vector data is stored in an efficient vector database; Dynamic retrieval weight assignment: Dynamically assign retrieval weights according to the question type or intent; Retrieval methods and default weights: Metadata vector recall: The weight is default set to 0.2; Source file vector recall: The weight is default set to 0.3; Metadata inverted index recall: The weight is default set to 0.5; Analyze the characteristics of the input question through the question intent recognition model; Dynamically adjust the weights of the above three retrieval methods to match the retrieval target; Hybrid retrieval: Simultaneously adopt three methods of metadata vector, source file vector, and metadata inverted index for retrieval, and fuse the results based on weights; The three recall methods run independently and each return a result set; Use a weighted fusion algorithm to combine the three result sets, considering the recall scores and weights comprehensively; Adopt a re-ranking model to further optimize the combined results to ensure that highly relevant results are ranked at the top; Result return: Return the final result set based on hybrid retrieval; Provide the top N sorted results with their relevance scores attached; Mark the metadata relevance, content relevance, and inverted index relevance in the returned results; Combining weighted fusion and re-ranking ensures that the returned result set is both comprehensive and accurate.

[0028] Furthermore, the semantic retrieval module further includes: Automatically select the recall method through question intention recognition: Question metadata recognition Question metadata extraction: Extract the metadata of the user's question; Metadata structuring: Perform structuring processing on the metadata; Perform intention recognition based on the metadata Intention recognition: Identify the intention of the question through a machine learning model; if it is a fact query question, use RAG vector retrieval; if it involves questions of multi-entity relationships and complex reasoning, use graph RAG; for complex questions, use the recall results of both traditional RAG and graph RAG, and then comprehensively generate the final answer through a generative model.

[0029] In this embodiment, multiple functional modules can effectively improve the efficiency of file sorting and management, reduce manual intervention. And provide diverse data processing and analysis functions to support in-depth research and data mining. As well as enhance the accuracy and speed of file retrieval and improve the user experience.

[0030] Embodiment 3: This embodiment provides a Figure 2 file management method based on a vertical large model as shown in Use the vertical large model to automatically sort the files through the automatic sorting module; Clean and file the data from the automatic sorting module through the cleaning and filing module; Construct a file knowledge graph for the data from the cleaning and filing module through the knowledge graph module and store it in the knowledge base module; Store files through the knowledge base module; Intelligently extract relevant materials of the files from the knowledge base module through the intelligent compilation and research module, perform multi-modal data analysis, and automatically generate corresponding compilation and research theme materials; Use natural language understanding and semantic parsing technology through the semantic retrieval module to convert the user's query into a machine-understandable representation form, and sort and recommend the retrieval results by calculating the semantic similarity or relevance between the query and the file content in the knowledge base module; Extract image objects from the files in the knowledge base module through the face recognition module and perform face recognition. Manage the file content containing face images by identifying and saving the face information in the image, as well as associating the coordinate position and page number information in the file; Read and store the valid information of the files into the knowledge base through the association module, and perform matching and association to ensure the consistency and integrity of the knowledge base and the file content; Predict and manage the confidentiality period of files in the knowledge base module through the prediction management module.

[0031] In this embodiment, multiple functional modules can effectively improve the efficiency of file sorting and management, reduce manual intervention, provide diversified data processing and analysis functions, support in-depth research and data mining, and enhance the accuracy and speed of file retrieval, thus improving the user experience.

[0032] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. An archive management system based on a vertical large model, characterized in that Including: Automatic sorting module: Automatically sort files using a vertical large model; Cleaning and archiving module: Clean and archive the data from the automatic sorting module; Knowledge graph module: Construct an archive knowledge graph for the data from the cleaning and archiving module and store it in the knowledge base module; Knowledge base module: Store files; Intelligent compilation and research module: Intelligently extract relevant materials of the files in the knowledge base module, conduct multimodal data analysis, and automatically generate corresponding compilation and research theme materials; Semantic retrieval module: Utilize natural language understanding and semantic parsing technologies to convert user queries into machine-understandable representations, and rank and recommend retrieval results by calculating the semantic similarity or relevance between the queries and the file contents in the knowledge base module; Facial recognition module: Extract image objects from the files in the knowledge base module and conduct facial recognition. By identifying and saving the facial information in the images, as well as associating the coordinate positions and page numbers information in the files, manage the archive contents containing facial images; Association module: Read and store the valid information of the files into the knowledge base, and conduct matching and association to ensure the consistency and integrity of the knowledge base and the file contents; Prediction and management module: Predict and manage the confidentiality periods of the files in the knowledge base module.

2. The archival management system based on a vertical large model according to claim 1, characterized in that, The vertical large model is the LLM large language model: OCR recognition: Convert the text in the image into machine-coded text through optical character recognition; File number extraction: Extract the file number from the text recognized by optical character recognition; The file number is used to identify and track files or documents; Title extraction: Automatically extract the title of the document; The title is important metadata of the document, which helps to understand and search for the document; Issuing time: Identify and record the release or creation date of the document; Automatic annotation: Based on the previous steps, automatically mark or classify the document; Metadata storage: The file number, title, issuing time, and the content of automatic annotation will all be saved as metadata.

3. The archival management system based on a vertical large model according to claim 2, wherein The cleaning and archiving module includes: Data collection: Data collection is the first step in the data analysis process, which involves the process of obtaining the required information; Noise removal: The collected data often contains useless or irrelevant information, which is called noise. Noise removal refers to deleting the useless or irrelevant information from the data to improve the data quality and accuracy; Text normalization: Text normalization is the process of converting text into a unified format; It includes converting text to lowercase, deleting punctuation marks, and replacing abbreviations; Document classification: Group documents according to specific criteria or features; It includes: classifying documents according to the theme, author, and date; Tag extraction and generation: Tags are keywords or phrases used to describe the attributes of data; Automatically extract or generate tags from documents to better understand and organize the data; Data deduplication: If there are duplicate entries in the dataset, then conduct data deduplication to ensure the accuracy of the analysis results; Data verification and evaluation: Verify and evaluate the processed data to ensure its quality and integrity; It includes checking the consistency, integrity, and accuracy of the data; The text summarization generation algorithm generates a summary by weighted extraction of important sentences; Importance score calculation Given a document D consisting of n sentences S1, S2, ..., Sn, the weight W(Si) of a sentence is calculated based on the sentence length, keyword density, and similarity to the title: W(Si) = α·Length(Si) + β·KeywordDensity(Si) + γ·Sim(Si, T) where α, β, and γ are adjustment coefficients, Length(Si) is the sentence length, KeywordDensity(Si) is the keyword density, and Sim(Si, T) is the similarity between the sentence and the title T; Abstract generation for αβγκ The abstract Ssummary consists of the top k sentences with the highest weights: Ssummary = Top-k{Si|W(Si)}.

4. The archival management system based on a vertical large model according to claim 1, wherein The knowledge graph module includes: Entity recognition: The process of identifying specific types of words or phrases in the text, including: personal names, locations, or organization names; Relationship extraction: Extracting the relationships between entities from the text; Attribute filling: Adding detailed information to entities to make them more rich and detailed; Triple generation: A triple is a structure consisting of three elements used to represent facts in the knowledge graph, including: subject, predicate, and object; Generate triples based on the previously extracted entities and their relationships; Graph storage: Store the generated triples to form a knowledge graph for convenient subsequent querying and analysis; Data import: Import relevant data sources to combine with the current knowledge graph; Multi-source data fusion: Merge multiple data sources together to obtain a more comprehensive view of information; Graph visualization: Display the knowledge graph in a graphical way to make it easier for users to understand the associations between data; Graph evaluation: Evaluate the quality and effectiveness of the knowledge graph to see if it meets the requirements and if improvements are needed.

5. The archival management system based on a vertical large model according to claim 1, wherein The semantic retrieval module includes: Polymorphic LLM: A polymorphic large model that receives various types of data inputs, including: documents, PPTs, DOCs, pictures; Prompt engineering: Prompt engineering is a method of guiding the model to perform specific tasks, guiding the model to focus on specific information or questions; Multimodal analysis: Multimodal analysis is to combine information from different modalities, including: text, image, audio for analysis; Generate necessary information: Generate key information about the input data; Storage: The generated information is stored for subsequent use; Vector database: The stored information is put into a vector database, which is a space for storing high-dimensional data; Question answering: Users access the stored information by asking questions; Similarity search: Users find the answer that best matches the question through a similar search.

6. The archival management system based on a vertical large model according to claim 1, wherein The face recognition module includes: Extract image features using a convolutional neural network and perform matching through vectorized representation; Feature extraction: Extract the feature vector f from the input image I: f = CNN(I) Similarity matching: Given the existing face feature vectors f1, f2, …, fn in the database, calculate the similarity between the input image feature f and each face feature in the database and return the result with the highest similarity; In the semantic retrieval module, the retrieval results are sorted by calculating the semantic similarity between the query Q and the archive content D; Vector Representation First, represent the query and the document as vectors q and d: q = Embedding(Q) d = Embedding(D) Similarity Calculation Use cosine similarity to calculate the semantic similarity between the query and the document: where q·d represents the dot product of the two vectors, and |q| and |d| are the norms of the vectors.

7. The archival management system based on a vertical large model according to claim 1, wherein The prediction management module includes: Archive Classification: Classify the files; Sensitivity Analysis: Analyze by evaluating the sensitivity of the model to parameter changes; Model Selection and Training: Select a suitable model according to the sensitivity analysis results and train it; Confidentiality Period Attribution Analysis: Conduct attribution analysis to obtain the factors affecting the results and their weights; Confidentiality Period Prediction: Predict the confidentiality period of the file; LLM: The large model is responsible for processing archive classification and prediction; Output Influence Factors and Weights: Output the factors affecting the confidentiality period and their importance; Output Confidentiality Period Prediction: Output the confidentiality period of the file.

8. The archival management system based on a vertical large model according to claim 1, wherein, The semantic retrieval module also includes: Metadata Extraction: Based on template matching, extract metadata information from the file: Design a set of metadata templates; Parse the file using regular expressions, rule engines, or pre-trained text classification models; The extracted metadata will be associated with the source file content and stored as structured data; Metadata Vectorization: Convert the metadata into vector representation and store its association information with the source file; Use a pre-trained embedding model to vectorize the metadata; The vectorization process will maintain the semantic integrity of the metadata and introduce the association information of the source file as context; The vector data is stored in an efficient vector database; Dynamic Retrieval Weight Assignment: Dynamically assign retrieval weights according to the question type or intention; Retrieval Method and Default Weight: Metadata Vector Recall: The weight is default set to 0.2; Source File Vector Recall: The weight is default set to 0.3; Metadata Inverted Index Recall: The weight is default set to 0.5; Analyze the characteristics of the input question through the question intention recognition model; Dynamically adjust the weights of the above three retrieval methods to match the retrieval target; Hybrid Retrieval: Simultaneously use three methods, namely metadata vector, source file vector, and metadata inverted index, for retrieval and fuse the results based on weights; The three recall methods run independently and each returns a result set; Use a weighted fusion algorithm to merge the three result sets, considering the recall score and weight; Use a re-ranking model to further optimize the merged results to ensure that highly relevant results are ranked at the top; Result Return: Return the final result set based on hybrid retrieval; Provide the top N sorted results with their relevance scores; Mark the metadata relevance, content relevance, and inverted index relevance in the returned results; Combine weighted fusion and re-ranking.

9. The archival management system based on a vertical large model according to claim 1, wherein The semantic retrieval module also includes: Automatically select the recall method through question intention recognition: Question Metadata Recognition Question Metadata Extraction: Extract the metadata of the user question; Metadata Structuring: Structurally process the metadata; Intention Recognition Based on Metadata Intent recognition: Identify the intent of the question through a machine learning model; use RAG vector retrieval for factual query questions; use graph RAG for questions involving multi-entity relationships and complex reasoning; for complex questions, use the recall results of both traditional RAG and graph RAG, and then synthesize the final answer through a generative model.

10. An archive management method based on a vertical large model, applied to the archive management system based on the vertical large model described in claim 1, characterized in that, It includes: Automatically organize the archives using a vertical large model through an automatic organization module; Clean and archive the data from the automatic organization module through a cleaning and archiving module; Construct an archive knowledge graph for the data from the cleaning and archiving module through a knowledge graph module and store it in the knowledge base module; Store files through the knowledge base module; Intelligently extract relevant materials from the files in the knowledge base module through an intelligent compilation and research module, perform multi-modal data analysis, and automatically generate corresponding compilation and research theme materials; Through a semantic retrieval module, use natural language understanding and semantic parsing technologies to convert user queries into machine-understandable representations, and sort and recommend retrieval results by calculating the semantic similarity or relevance between the query and the file content in the knowledge base module; Extract image objects from the files in the knowledge base module through a face recognition module and perform face recognition. Manage the archive content containing face images by identifying and saving the face information in the image, as well as associating the coordinate positions and page numbers in the file; Read and store the valid information of the file into the knowledge base through an association module, and perform matching and association to ensure the consistency and integrity of the knowledge base and the file content; Predict and manage the confidentiality period of the files in the knowledge base module through a prediction management module.

Citation Information

Patent Citations

  • Archive intelligent auxiliary editing and research method and system and related equipment

    CN115730119A

  • File digitization management method and device

    CN116343210A

  • Archive knowledge management method and system based on large language model

    CN119180330A

  • Knowledge management system and method for constructing large voice model

    CN119476458A

Cited By

  • File digital governance method and system based on large model

    CN121092759A