Large-model private knowledge base oriented to congruent content optimization and retrieval processing method and system
Through deeply optimized topic model and markdown format conversion, combined with semantic matching algorithm and topic association retrieval, the problem of low document management and knowledge retrieval in the private knowledge base of big models is solved, and efficient and accurate knowledge management and summary content answers are achieved.
Patent Information
- Application Number
- CN202510414679.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
When the existing technology is applied to a private knowledge base, document processing lacks intelligent topic extraction and format conversion, resulting in insufficient storage and management, knowledge retrieval depends on simple matching strategies, and the accuracy of summary content answers is poor, which affects user experience.
The deep-optimized theme model is used to extract the core theme, convert it into markdown format, and optimized reverse order index is constructed, combining semantic matching algorithms and topic association search to generate efficient and accurate knowledge retrieval results.
It improves the efficiency and accuracy of document management and knowledge retrieval, improves the quality of answers to summarized content, and enhances user experience and satisfaction.
Smart Images

Figure CN120336476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and data processing technology, and in particular to a large-model private knowledge base and a retrieval processing method and system for summarizing content optimization. Background Art
[0002] With the rapid development of artificial intelligence technology, large models have been widely used in knowledge processing and question-answering systems. However, there are currently some significant problems when large models are applied to private knowledge base scenarios. For example, the processing method for user-uploaded documents is relatively simple, lacking intelligent subject extraction and format conversion mechanisms, resulting in inefficient storage and management of documents in the knowledge base. In terms of knowledge retrieval, existing technologies often rely only on simple keyword matching or a single retrieval strategy, which is unable to accurately locate and extract the knowledge required by users, especially for answers to summary content. The accuracy is poor, which seriously affects the user experience and the effect of knowledge application.
[0003] Therefore, how to improve the efficiency and accuracy of knowledge management and retrieval, and thus improve the quality of answers to summary content, is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The technical task of the present invention is to provide a large-model private knowledge base and retrieval processing method and system optimized for summary content, so as to solve the problem of how to improve the efficiency and accuracy of knowledge management and retrieval, and thus improve the quality of answers to summary content.
[0005] The technical task of the present invention is achieved in the following way: a large model private knowledge base and retrieval processing method for summarizing content optimization, the method is as follows:
[0006] Document upload processing: After the user uploads the document, the built-in large model is called to extract the core theme using the deeply optimized theme model. At the same time, the document format is converted to markdown format with the help of the format conversion tool, and the document directory structure is extracted. The theme, directory, file name and markdown document content are structured and stored in the database and vector library;
[0007] Knowledge retrieval: Build an optimized reverse index, filter documents by subject, directory and name matching, and then use the semantic matching algorithm trained with parameter optimization (such as cosine similarity, BM25 algorithm) to match document fragments: If the match is not good, filter out topics with high matching degrees through an intelligent judgment mechanism, match related documents again, and if it is still not ideal, perform a subject-related search; finally, submit the search results and user questions to the big model to generate answers.
[0008] As a preferred embodiment, the document upload process is as follows:
[0009] Topic extraction: Use topic models and deep learning technology to conduct a comprehensive semantic analysis of the document content uploaded by users and accurately extract the core topics. And use the extracted core topics to comprehensively and accurately summarize the document content to form a clear document theme, laying the foundation for the subsequent accurate grasp and efficient use of document information, making the subsequent document processing process more targeted.
[0010] Format conversion: Call the document format conversion tool pandoc, and with the help of the Apache Tika open source framework, convert word documents (carrying rich typesetting styles and editing information), pdf documents (page presentation is stable and typesetting is complex) and other niche but commonly used document formats into markdown format; the markdown format is concise and easy to read, with simple syntax and easy editing. After conversion, the convenience of subsequent storage, editing and display of documents is significantly improved; when storing, the markdown format takes up little space, has a simple structure, and is easy to manage; when editing, the concise syntax allows users to focus on content creation without having to pay attention to complex typesetting; when displaying, the clear structure allows readers to quickly understand the content. This format conversion improves the versatility of document processing and provides great convenience for subsequent operations;
[0011] Document directory structure extraction: After the document is successfully converted to markdown format, the directory structure extraction process is started according to the markdown format structure characteristics, and regular expression matching and syntax analysis technology based on natural language processing are used to identify the key information of the titles and chapters in the document, and a complete and clear document directory structure is sorted out; the document directory structure intuitively reflects the overall framework of the document, providing an important basis for subsequent document retrieval and content positioning, and facilitating users to quickly browse and locate specific content;
[0012] Synchronous information storage: Uploaded documents are stored in the database in a structured storage method, and the document name, subject, document directory, and markdown document content are saved in the vector library. The Faiss (Facebook AI SimilaritySearch) efficient vector search engine is used to vectorize the document information through vector representation, providing strong support for efficient retrieval based on vector similarity.
[0013] Preferably, the database should be a relational database, such as MySQL, PostgreSQL, etc., which uses a rigorous and standardized table structure to accurately store information and ensure data consistency and integrity; or a document database, such as MongoDB, which uses a flexible document storage mode to adapt to different document information storage requirements, achieve orderly information storage, and lay the foundation for fast retrieval and efficient call;
[0014] Create a document information table in the database. Among them, the fields of the document information table include document ID, subject field, directory field, and file name. Each field cooperates with each other to ensure the integrity and availability of information from multiple dimensions. Among them, the document ID is used as the unique identifier to accurately locate each document and ensure the accuracy and uniqueness of the data. The subject field records the key themes extracted, intuitively reflecting the core content of the document. The directory field presents the structure context of the document in detail, facilitating users to understand the overall framework. The file name intuitively shows the original document name, enabling users to clearly identify the document source.
[0015] Preferably, the knowledge retrieval is as follows:
[0016] Subject, directory, and name matching: After the user submits a knowledge retrieval request, perform precise comparison on the extracted subject, document directory, and document name. Use the inverted index technology (like a book index) to establish a mapping relationship between the document keywords and the corresponding documents. At the same time, use the distributed cache technology (such as Redis) to cache the frequently accessed index data to speed up the query speed.
[0017] Document fragment matching: Perform refined segmentation on the selected documents.
[0018] Subject-related retrieval: If no suitable fragment is found or the matching degree is lower than the preset threshold (such as 0.5) during the document fragment matching process, automatically start the subject-related retrieval mechanism. Among them, the core of the subject-related retrieval mechanism is to deeply retrieve in the documents corresponding to the subject.
[0019] Answer generation: Submit the retrieval results and the user's question to the large model. As the core of the knowledge retrieval and answer generation system, the large model, relying on its powerful language understanding and generation capabilities, comprehensively integrates and deeply analyzes the retrieval results.
[0020] Preferably, the document fragment matching is as follows:
[0021] According to the logical relationship of the document, divide the document into logically coherent paragraphs or relatively independent chapters with clearly defined fragments.
[0022] Process the segmented document fragments using semantic matching algorithms; among them, the semantic matching algorithms include the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering factors such as document length and term frequency; or, introduce a pre-trained language model based on the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), to perform more accurate semantic encoding and matching on the document fragments and the retrieval request;
[0023] Match the user's retrieval request with the document fragments one by one using semantic matching algorithms or pre-trained language models based on the Transformer architecture, perform precise calculation and comparison, find the document fragment that best fits the user's needs, and provide more targeted and accurate retrieval results to meet the user's need to obtain detailed information.
[0024] Preferably, the large model deeply understands the internal relationships among the retrieval results, sorts out and summarizes the key information from different document fragments and the content obtained through topic-related retrieval, extracts and summarizes the key information of multiple relevant document fragments, and organizes them into accurate and comprehensive answers in a logically clear and concise language; when generating answers, fully consider the user's question background and intention to ensure that the answers fit the user's needs and have high readability and practicality;
[0025] The large model timely feeds back the generated answers to the user, helps the user quickly acquire knowledge, solve practical problems, and provide an efficient and convenient knowledge service experience.
[0026] A large model private knowledge base and retrieval processing system optimized for summary content, the system includes:
[0027] A document upload and processing module, which, when a user uploads a document, calls the built-in large model, uses the topic model algorithm to perform semantic analysis on the document content, and extracts the core topic; at the same time, calls a professional format conversion tool (such as pandoc) to convert the document into markdown format, and extracts the document directory structure according to the characteristics of the markdown format; finally, saves the extracted topic, document directory, file name, and the converted markdown document content in a structured manner to a database (either a relational database or a document database), and synchronously saves it to the vector library;
[0028] A knowledge retrieval module, which is used to build an inverted index when a user submits a retrieval request, perform matching and screening based on the extracted topics, document directories, and document names; and segment the screened documents, and use a semantic matching algorithm to match the user's retrieval request with the segmented document fragments: if the matching degree is lower than the preset threshold, then screen out the topics with higher matching degrees, and perform re-matching based on the documents corresponding to the screened topics; if a satisfactory result is still not achieved, then perform topic association retrieval in the documents corresponding to the topic and their related documents; finally, submit the retrieval result and the user's question to the large model, and the large model generates an answer and feedbacks it to the user.
[0029] Preferably, a relational database or a document database is selected as the database;
[0030] A document information table is created in the database. Among them, the fields of the document information table include document ID, topic field, directory field, and file name. Each field cooperates with each other to ensure the integrity and availability of information from multiple dimensions; among them, the document ID is used as the unique identifier to accurately locate each document and ensure the accuracy and uniqueness of the data; the topic field records the key topics refined, intuitively reflecting the core content of the document; the directory field presents the document structure context in detail, facilitating users to understand the overall framework; the file name intuitively shows the original document name, allowing users to clearly identify the document source.
[0031] Preferably, the process of using a semantic matching algorithm (such as the cosine similarity algorithm, BM25 algorithm) to match the user's retrieval request with the segmented document fragments is as follows:
[0032] According to the logical relationship of the document, the document is segmented into logically coherent paragraphs or relatively independent chapters with clearly defined meanings;
[0033] Use a semantic matching algorithm to process the segmented document fragments; among them, the semantic matching algorithm includes the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures the text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering factors such as document length and word frequency; or, introduce a pre-trained language model based on the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), to perform more accurate semantic encoding and matching on the document fragments and the retrieval request;
[0034] The user's search request and the document fragment are matched one by one using a semantic matching algorithm or a pre-trained language model based on the Transformer architecture, and the comparison is accurately calculated to find the document fragment that best meets the user's needs, providing more targeted and accurate search results to meet the user's needs for detailed information.
[0035] Better yet, the big model deeply understands the internal connections between the search results, sorts out and summarizes the key information from different document fragments and the content obtained based on thematic association retrieval, extracts and summarizes the key information of multiple related document fragments, and organizes them into accurate and comprehensive answers in clear logic and concise language; when generating answers, it fully considers the user's question background and intention to ensure that the answers meet user needs and are highly readable and practical;
[0036] The big model will promptly feed back the generated answers to users, helping them to quickly acquire knowledge, solve practical problems, and provide an efficient and convenient knowledge service experience.
[0037] The large-model private knowledge base and retrieval processing method and system for summarizing content optimization of the present invention have the following advantages:
[0038] (1) The present invention improves the efficiency of document management: through automatic subject extraction and format conversion, the storage of documents in the knowledge base is more standardized and orderly, which is convenient for management and maintenance; at the same time, the synchronously saved subject, directory and file name information provides a multi-dimensional index for subsequent retrieval, which greatly improves the efficiency and accuracy of knowledge management and retrieval, especially improves the quality of answers to summary content;
[0039] (ii) The present invention improves the accuracy of knowledge retrieval: the innovative three-step retrieval strategy, from topic, document fragment to topic-related retrieval, gradually narrows the retrieval scope, accurately locates the knowledge required by users, greatly improves the efficiency and accuracy of document management and knowledge retrieval, and effectively solves the problem of poor answers to summary content;
[0040] (III) The present invention enhances user experience: it provides users with more accurate and comprehensive knowledge answers, meets users' knowledge needs in different scenarios, and improves users' satisfaction and usage frequency of the large model private knowledge base system;
[0041] (iii) In the document upload processing module of the present invention, the topic model algorithm used when extracting topics from documents is a deeply optimized algorithm to improve the accuracy and efficiency of topic extraction;
[0042] (iv) After converting the document into the markdown format, the document upload processing module of the present invention uses specific algorithms and rules to extract the directory structure of the markdown format document to accurately identify information such as titles at all levels and chapter identifiers, so as to form a complete and clear document directory structure;
[0043] (5) In the knowledge retrieval module of the present invention, when performing theme, directory, and name matching, the constructed inverted index is optimized to improve the retrieval efficiency and quickly screen out the document scope related to the retrieval request;
[0044] (6) When the knowledge retrieval module of the present invention performs document fragment matching, the semantic matching algorithm adopted is optimized in parameters and trained to improve the matching accuracy between the document fragment and the retrieval request;
[0045] (7) When the knowledge retrieval module of the present invention performs theme-related retrieval, it has an intelligent judgment mechanism that can accurately determine the corresponding theme document scope that needs to be deeply retrieved based on the retrieval request and the existing retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be further described below with reference to the accompanying drawings.
[0047] Attached Figure 1 is a flow block diagram of knowledge retrieval. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The large model private knowledge base and retrieval processing method and system optimized for summary content of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0049] Embodiment 1:
[0050] This embodiment provides a large model private knowledge base and retrieval processing method optimized for summary content. The method is as follows:
[0051] S1. Document upload processing: After the user uploads a document, the built-in large model is called to extract the core theme using a deeply optimized theme model. At the same time, the document format is converted to the markdown format with the help of a format conversion tool, and the document directory structure is extracted. The theme, directory, file name, and markdown document content are stored in a structured manner in the database and vector library;
[0052] S2. Knowledge retrieval: An optimized inverted index is constructed, and documents are screened according to theme, directory, and name matching. Then, a semantic matching algorithm (such as cosine similarity, BM25 algorithm) optimized in parameters is used to perform document fragment matching: If the matching degree is not good, the themes with high matching degrees are screened through the intelligent judgment mechanism, and the relevant documents are matched again. If it is still not ideal, theme-related retrieval is performed; finally, the retrieval result and the user's question are submitted to the large model to generate an answer.
[0053] The document upload processing in step S1 of this embodiment is specifically as follows:
[0054] S101. Topic Extraction: Use topic models and deep learning techniques to conduct a comprehensive semantic analysis of the content of the documents uploaded by users, and accurately extract the core topics. For example, for a document on "Enterprise Financial Management Strategies", the model can keenly capture key information and identify topics such as "Cost Control", "Cash Flow Management", and "Budget Planning"; and comprehensively and accurately summarize the content of the document through the extracted core topics to form a clear document topic, laying a foundation for accurately grasping and efficiently utilizing the document information subsequently, making the subsequent document processing process more targeted; Topic models include classic Latent Dirichlet Allocation (LDA) models, as well as emerging Neural Topic Models (NTMs); Deep learning techniques include Recurrent Neural Networks (RNNs) and their variants Long Short-Term Memory Networks (LSTMs) and Gated Recurrent Units (GRUs);
[0055] S102. Format Conversion: Call the document format conversion tool pandoc and, with the help of the Apache Tika open-source framework, uniformly convert word documents (carrying rich typesetting styles and editing information), pdf documents (with stable page presentation and complex typesetting), and other less common but frequently used document formats into markdown format; The markdown format is concise and easy to read, with simple syntax and convenient for editing, significantly improving the convenience of subsequent storage, editing, and display of the document; When storing, the markdown format occupies less space and has a simple structure, facilitating management; When editing, the simple syntax allows users to focus on content creation without having to worry about complex typesetting; When displaying, the clear structure enables readers to quickly understand the content. This format conversion improves the generality of document processing and provides great convenience for subsequent operations;
[0056] S103. Document Table of Contents Structure Extraction: After the document is successfully converted to markdown format, start the table of contents structure extraction process based on the characteristics of the markdown format structure, and use regular expression matching and syntactic analysis techniques based on natural language processing to identify the key information of the headings and chapter identifiers at all levels in the document, and sort out a complete and clear document table of contents structure; Among them, the document table of contents structure intuitively reflects the overall framework of the document, providing an important basis for subsequent document retrieval and content positioning, and facilitating users to quickly browse and locate specific content;
[0057] S104. Information Synchronous Saving: Store the uploaded document in the database using a structured storage method, and save the document name, topic, document table of contents, and markdown document content to a vector library. Use the Faiss (Facebook AI Similarity Search) efficient vector search engine to vectorize the document information through vector representation, providing strong support for efficient retrieval based on vector similarity.
[0058] The database in this embodiment uses a relational database, such as MySQL, PostgreSQL, etc., which uses a rigorous and standardized table structure to accurately store information to ensure data consistency and integrity; or a document database, such as MongoDB, which uses a flexible document storage mode to adapt to different document information storage requirements, realize orderly storage of information, and lay the foundation for fast retrieval and efficient call;
[0059] A document information table is created in the database. The fields of the document information table include document ID, subject field, directory field and file name. The fields work together to ensure that the information is complete and available from multiple dimensions. The document ID is used as a unique identifier to accurately locate each document and ensure that the data is accurate and unique. The subject field records the refined key topics and intuitively reflects the core content of the document. The directory field presents the document structure in detail to facilitate users to understand the overall framework. The file name intuitively displays the original document name, allowing users to clearly identify the source of the document.
[0060] As attached Figure 1 As shown, the knowledge retrieval in step S2 of this embodiment is specifically as follows:
[0061] S201, subject, directory and name matching: After the user submits a knowledge search request, the extracted subject, document directory and document name are accurately compared, and the inverted index technology (similar to the book index) is used to establish a mapping relationship between the document keywords and the corresponding documents. At the same time, the distributed cache technology (such as Redis) is used to cache the frequently accessed index data to speed up the query speed; for example, when the user searches for "cost control related documents", the system can quickly filter out documents whose themes include the keyword "cost control" based on the inverted index, greatly narrowing the subsequent search scope, avoiding blind search, and significantly improving the search efficiency, allowing users to obtain a list of highly relevant documents in a short time;
[0062] S202, document fragment matching: finely segment the screened documents;
[0063] S203, subject-related retrieval: If no suitable fragment is found during the document fragment matching process or the matching degree is lower than a preset threshold (such as 0.5), the subject-related retrieval mechanism is automatically started; the core of the subject-related retrieval mechanism is to conduct in-depth retrieval in the documents of the corresponding subject; for example, when a user searches for "cost analysis of new marketing channels", the existing document fragments are not accurately matched, but with the help of subject-related search, the subject document of "new marketing channels" is found, and then the document and its related documents are further searched, and the search is conducted from a broader perspective and more detailed level, and the upstream and downstream industrial chain analysis, market trend research and other information related to "new marketing channels" in the document are comprehensively considered to explore the potential "cost analysis" related content; through this mechanism, the system effectively responds to the complex and diverse retrieval needs of users and provides comprehensive and accurate knowledge retrieval services as much as possible;
[0064] S204, Answer Generation: The retrieval results and the user's question are submitted to the large model together. As the core of the knowledge retrieval and answer generation system, the large model comprehensively integrates and deeply analyzes the retrieval results with its powerful language understanding and generation capabilities.
[0065] The document fragment matching in step S202 of this embodiment is specifically as follows:
[0066] S20201. According to the logical relationship of the document, the document is segmented into logically coherent paragraphs or relatively independent chapters with clearly defined meaningful fragments;
[0067] S20202. Use semantic matching algorithms to process the segmented document fragments; among them, the semantic matching algorithms include the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures the text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering factors such as document length and term frequency; or, introduce a pre-trained language model based on the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), to perform more accurate semantic encoding and matching on the document fragments and the retrieval request;
[0068] S20203. Match the user's retrieval request with the document fragments one by one using semantic matching algorithms or pre-trained language models based on the Transformer architecture, perform accurate calculation and comparison, find the document fragment that best meets the user's needs, provide more targeted and accurate retrieval results, and meet the user's need to obtain detailed information.
[0069] The large model in this embodiment deeply understands the internal relationships among the retrieval results, sorts out and summarizes the key information from different document fragments and the content obtained through topic-related retrieval, extracts and summarizes the key information of multiple relevant document fragments, and organizes them into accurate and comprehensive answers in a logically clear and concise language; when generating answers, fully consider the background and intention of the user's question to ensure that the answers meet the user's needs and have high readability and practicality;
[0070] The large model timely feeds back the generated answers to the user, helps the user quickly acquire knowledge, solve practical problems, and provide an efficient and convenient knowledge service experience.
[0071] Embodiment 2:
[0072] This embodiment provides a large model private knowledge base and retrieval processing system for optimizing summary content. The system includes:
[0073] A document upload processing module, which is used to call a built-in large model when a user uploads a document, use a topic model algorithm to perform semantic analysis on the document content, and extract the core topic; at the same time, call a professional format conversion tool (such as pandoc) to convert the document into markdown format, and extract the document directory structure according to the characteristics of the markdown format; finally, save the extracted topic, document directory, file name, and the converted markdown document content in a structured manner to a database (either a relational database or a document database), and synchronously save it to a vector library.
[0074] A knowledge retrieval module, which is used to build an inverted index when a user submits a retrieval request, and perform matching and screening based on the extracted topics, document directories, and document names; and segment the screened documents, and use a semantic matching algorithm to match the user's retrieval request with the segmented document fragments: if the matching degree is lower than the preset threshold, then screen out the topics with higher matching degrees, and perform re-matching based on the documents corresponding to the screened topics; if the satisfactory result is still not achieved, then perform topic association retrieval in the documents corresponding to the topic and their related documents; finally, submit the retrieval result and the user's question to the large model, and the large model generates an answer and feedbacks it to the user.
[0075] In this embodiment, the database selects a relational database or a document database.
[0076] A document information table is created in the database. Among them, the fields of the document information table include document ID, topic field, directory field, and file name. Each field cooperates with each other to ensure the integrity and availability of information from multiple dimensions; among them, the document ID is used as the unique identifier to accurately locate each document and ensure the accuracy and uniqueness of the data; the topic field records the key topics refined, intuitively reflecting the core content of the document; the directory field presents the document structure context in detail, facilitating users to understand the overall framework; the file name intuitively shows the original document name, allowing users to clearly identify the document source.
[0077] In this embodiment, the specific process of using a semantic matching algorithm (such as cosine similarity algorithm, BM25 algorithm) to match the user's retrieval request with the segmented document fragments is as follows:
[0078] According to the logical relationship of the document, the document is segmented into logically coherent paragraphs or relatively independent chapters with clearly defined fragments.
[0079] Process the segmented document fragments using semantic matching algorithms; among them, the semantic matching algorithms include the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering factors such as document length and word frequency; or, introduce a pre-trained language model based on the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), to perform more accurate semantic encoding and matching on the document fragments and the retrieval request;
[0080] Match the user's retrieval request with the document fragments one by one using semantic matching algorithms or pre-trained language models based on the Transformer architecture, perform accurate calculation and comparison, find the document fragment that best meets the user's needs, and provide more targeted and accurate retrieval results to meet the user's need to obtain detailed information.
[0081] The large model in this embodiment deeply understands the internal relationships among the retrieval results, sorts out and summarizes the key information from different document fragments and the content obtained through topic-related retrieval, extracts and summarizes the key information of multiple relevant document fragments, and organizes them into accurate and comprehensive answers in a logically clear and concise language; when generating answers, fully consider the user's question background and intention to ensure that the answers meet the user's needs and have high readability and practicality;
[0082] The large model timely feeds back the generated answers to the user, helps the user quickly acquire knowledge, solve practical problems, and provides an efficient and convenient knowledge service experience.
[0083] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A large model private knowledge base and retrieval processing method optimized for summary content, characterized in that The method is as follows: Document upload processing: After the user uploads a document, the built-in large model is called to extract the core theme using a deeply optimized topic model. At the same time, a format conversion tool is used to convert the document format to the markdown format, and the document table of contents structure is extracted. The theme, table of contents, file name, and markdown document content are stored in a structured manner in the database and vector library; Knowledge retrieval: An optimized inverted index is constructed to match and filter documents according to the theme, table of contents, and name. Then, a semantic matching algorithm optimized by parameters is used for document fragment matching: If the matching degree is poor, a smart judgment mechanism is used to screen for themes with a high matching degree and re-match the relevant documents. If it is still not ideal, a theme-related retrieval is performed; Finally, the retrieval results and the user's question are submitted to the large model to generate an answer.
2. The large model private knowledge base and retrieval processing method optimized for summary content according to claim 1, characterized in that, The document upload processing is as follows: Theme extraction: The content of the document uploaded by the user is comprehensively semantically analyzed using a topic model and deep learning technology to accurately extract the core theme; and the core theme is used to comprehensively and accurately summarize the document content to form a clear document theme; Format conversion: Call the document format conversion tool pandoc and use the Apache Tika open-source framework to uniformly convert word documents, pdf documents, and other document formats to the markdown format; Document table of contents structure extraction: After the document is successfully converted to the markdown format, the table of contents structure extraction process is started according to the characteristics of the markdown format structure, and regular expression matching and syntactic analysis technology based on natural language processing are used to identify the key information of each level of title and chapter identifier in the document, and a complete and clear document table of contents structure is sorted out; Information synchronization and storage: The uploaded document is stored in the database using a structured storage method, and the document name, theme, document table of contents, and markdown document content are stored in the vector library. The Faiss efficient vector search engine is used to vectorize the document information through vector representation, providing strong support for efficient retrieval based on vector similarity.
3. The large model private knowledge base and retrieval processing method optimized for summary content according to claim 2, characterized in that, The database selects a relational database or a document database A document information table is created in the database. Among them, the fields of the document information table include document ID, theme field, table of contents field, and file name. Each field cooperates with each other to ensure the integrity and availability of information from multiple dimensions; among them, the document ID is used as a unique identifier to accurately locate each document and ensure the accuracy and uniqueness of the data; the theme field records the key themes extracted, intuitively reflecting the core content of the document; the table of contents field presents the document structure context in detail, facilitating users to understand the overall framework; the file name intuitively shows the original document name, allowing users to clearly identify the document source.
4. The large model private knowledge base and retrieval processing method optimized for summary content according to claim 1, characterized in that, The knowledge retrieval is as follows: Theme, table of contents, and name matching: After the user submits a knowledge retrieval request, precise comparison is made for the extracted theme, document table of contents, and document name. The inverted index technology is used to establish a mapping relationship between the document keywords and the corresponding documents. At the same time, the distributed cache technology is used to cache the frequently accessed index data; Document fragment matching: The selected documents are finely segmented; Subject-related retrieval: If no suitable fragment is found during the document fragment matching process or the matching degree is lower than the preset threshold, the subject-related retrieval mechanism is automatically activated; among them, the core of the subject-related retrieval mechanism is to deeply retrieve in the documents corresponding to the subject; Answer generation: The retrieval results and the user's question are submitted to the large model together. As the core of the knowledge retrieval and answer generation system, the large model comprehensively integrates and deeply analyzes the retrieval results with its powerful language understanding and generation capabilities.
5. The large model private knowledge base and retrieval processing method optimized for summary content according to claim 4, characterized in that The specific process of document fragment matching is as follows: According to the logical relationship of the document, the document is segmented into logically coherent paragraphs or relatively independent chapters with clearly defined fragments; Use semantic matching algorithms to process the segmented document fragments; among them, semantic matching algorithms include the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures the text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering factors such as document length and word frequency; alternatively, a pre-trained language model based on the Transformer architecture is introduced to perform more accurate semantic encoding and matching on the document fragments and the retrieval request; The user's retrieval request and the document fragments are matched one by one using semantic matching algorithms or a pre-trained language model based on the Transformer architecture, and the comparison is accurately calculated to find the document fragment that best meets the user's needs, providing more targeted and accurate retrieval results to meet the user's need to obtain detailed information.
6. The large model private knowledge base and retrieval processing method optimized for summary content according to claim 4, characterized in that, The large model deeply understands the internal relationships between the retrieval results, sorts out and summarizes the key information from different document fragments and the content obtained through subject-related retrieval, extracts and summarizes the key information of multiple relevant document fragments, and organizes it into an accurate and comprehensive answer in a clear and concise language; The large model timely feedbacks the generated answer to the user, helps the user quickly acquire knowledge, solve practical problems, and provides an efficient and convenient knowledge service experience.
7. A large model private knowledge base and retrieval processing system optimized for summary content, characterized in that, The system includes: A document upload and processing module, which is used to, when a user uploads a document, call the built-in large model, use the topic model algorithm to perform semantic analysis on the document content, and extract the core topic; at the same time, call a professional format conversion tool to convert the document into markdown format, and extract the document directory structure according to the characteristics of the markdown format; finally, save the extracted topic, document directory, file name, and the converted markdown document content in a structured manner to the database and synchronously save it to the vector library; A knowledge retrieval module, which is used to build an inverted index when a user submits a retrieval request, perform matching and screening based on the extracted topics, document directories, and document names; and segment the screened documents, and use a semantic matching algorithm to match the user's retrieval request with the segmented document fragments: if the matching degree is lower than the preset threshold, then screen out the topics with higher matching degrees, and perform secondary matching based on the documents corresponding to the screened topics; if a satisfactory result is still not achieved, then perform topic-related retrieval in the documents corresponding to the topic and their related documents; finally, submit the retrieval results and the user's question to a large model, and the large model generates an answer and feedbacks it to the user.
8. The large model private knowledge base and retrieval processing system optimized for summary content according to claim 7, wherein, The database selects a relational database or a document database; A document information table is created in the database. Among them, the fields of the document information table include document ID, topic field, directory field, and file name. Each field cooperates with each other to ensure the integrity and availability of information from multiple dimensions; among them, the document ID is used as the unique identifier to accurately locate each document and ensure the accuracy and uniqueness of the data; the topic field records the extracted key topics, intuitively reflecting the core content of the document; the directory field presents the document structure context in detail, facilitating users to understand the overall framework; the file name intuitively displays the original document name, enabling users to clearly identify the document source.
9. The large model private knowledge base and retrieval processing system optimized for summary content according to claim 7, characterized in that, The specific process of using the semantic matching algorithm to match the user's retrieval request with the segmented document fragments is as follows: According to the logical relationship of the document, the document is segmented into logically coherent paragraphs or relatively independent chapters with clearly defined meanings; Use the semantic matching algorithm to process the segmented document fragments; among them, the semantic matching algorithm includes the cosine similarity algorithm and the BM25 algorithm; the cosine similarity algorithm measures the text similarity by calculating the cosine value of the angle between two text vectors in the vector space, and the closer the value is to 1, the more similar the semantics; the BM25 algorithm calculates the relevance between the document and the retrieval request by comprehensively considering the document length and word frequency factors; or, introduce a pre-trained language model based on the Transformer architecture to perform more accurate semantic encoding and matching on the document fragments and the retrieval request. Match the user's retrieval request with the document fragments one by one using the semantic matching algorithm or the pre-trained language model based on the Transformer architecture, perform accurate calculation and comparison, find the document fragments that best meet the user's needs, and provide more targeted and accurate retrieval results to meet the user's need to obtain detailed information.
10. The large model private knowledge base and retrieval processing system optimized for summary content according to any one of claims 7-9, characterized in that, The large model deeply understands the internal relationships among the retrieval results, sorts out and summarizes the key information from different document fragments and the content obtained through topic-related retrieval, extracts and summarizes the key information of multiple relevant document fragments, and organizes them into accurate and comprehensive answers in a clear and concise language. The large model timely feedbacks the generated answer to the user, helps the user quickly acquire knowledge, solve practical problems, and provides an efficient and convenient knowledge service experience.