Document storage method, intelligent question and answer method, system, device and medium
By performing document type identification, semantic similarity calculation, and chapter division, the problem of structured storage of multi-source heterogeneous documents is solved, enabling accurate document content extraction and intelligent question-answering services, and improving the adaptability and accuracy of the knowledge base management system.
Patent Information
- Application Number
- CN202511059738.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to achieve accurate paragraph segmentation and semantic recognition when processing multi-source heterogeneous documents, resulting in ineffective structured extraction and impacting the accuracy of knowledge aggregation and intelligent question-answering services.
By identifying document types, calculating semantic similarity, dividing chapters and performing semantic recognition, the document content is segmented into multiple semantic fragments and stored in a structured manner based on hierarchical relationships. This is then combined with intelligent question-answering methods for querying and responding.
It enables unified processing and precise structured extraction of multi-source heterogeneous documents, improves the adaptability of the knowledge base management system and the accuracy of intelligent question answering, and supports multimodal information needs and permission management.
Smart Images

Figure CN120950458A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of text information extraction technology, and in particular to a document storage method, intelligent question answering method, system, device and medium. Background Technology
[0002] With the continuous improvement of enterprise informatization, the number of documents such as business system documents, process specifications, and product information accumulated in daily work is growing rapidly. These documents are not only diverse in format (such as PDF, Word, images, scanned copies, etc.), but also scattered across heterogeneous systems such as system platforms, departmental cloud drives, and personal terminals. Due to the different sources of these documents, their content structures are complex, their quality varies, and they commonly contain nested tables and multi-level headings. Currently, the content extraction of these documents usually relies on primitive data cleaners, which are based on template matching or simple text matching rules, lacking a deep understanding of the semantic structure of the documents. When processing documents with strong semantic hierarchical relationships, it is difficult to accurately segment and semantically recognize the documents, thus failing to accurately extract the structured content of the documents, which in turn affects subsequent knowledge collection and intelligent question answering services. Therefore, there is a need to provide a document storage method, an intelligent question answering method, a system, equipment, and media. Summary of the Invention
[0003] This invention provides a document storage method, an intelligent question-answering method, a system, an apparatus, and a medium to solve the technical problem that existing data cleaners, which rely solely on raw data cleaners, cannot accurately extract structured document content due to the heterogeneity of document formats.
[0004] This invention provides a document storage method, comprising: acquiring a document to be analyzed; performing type identification on the document, and extracting document content from the document based on the identified document type; calculating the semantic similarity between the document content and multiple pre-stored standard classification contents, and determining the content category corresponding to the document based on the semantic similarity; wherein each standard classification content corresponds to one content category; dividing the document content into chapters and determining a chapter summary for each chapter; wherein each chapter corresponds to a chapter summary; performing semantic identification on each chapter, and segmenting each chapter into multiple semantic segments; and storing the document in a structured manner according to the hierarchical relationship between the content category, chapter summary, and semantic segments.
[0005] In one embodiment of the present invention, document type identification is performed, and document content is extracted from the document based on the identified document type, including: performing document type identification to determine the document type: if the document is in editable text format, a text parsing engine is invoked to extract document content from the document; if the document is in image format, an optical character recognition engine or a visual analysis model is invoked to extract document content from the document.
[0006] In one embodiment of the present invention, calculating the semantic similarity between document content and multiple pre-stored standard classification contents, and determining the content category corresponding to the document based on the semantic similarity results, includes: parsing the document content to determine whether there is a text region that satisfies preset summary features; if so, using the text content of the text region as the document summary of the document content; otherwise, inputting the document content into a summary extraction model to generate a document summary of the document content; wherein, the summary extraction model is a generative model; calculating the semantic similarity between the document summary and each standard classification content, and selecting the content category with a semantic similarity greater than a preset similarity threshold as the content category corresponding to the document.
[0007] In one embodiment of the present invention, the document content is divided into chapters, and a chapter summary for each chapter is determined, including: determining whether the document content has a chapter structure based on a preset chapter division rule; if so, dividing the document content into multiple chapters based on the hierarchical relationship of the chapters; otherwise, inputting the document content into a text segmentation model, performing semantic analysis on the document content, and dividing the document content into multiple chapters based on the analysis results; wherein, the text segmentation model is a temporal model; and inputting each chapter into a summary generation model to generate a corresponding chapter summary; wherein, the summary generation model is a generative model or an extractive model.
[0008] This invention also provides an intelligent question-answering method, the method comprising: acquiring and saving a query request; calling a first language model to perform intent recognition on the query request, and semantically completing the query request based on the recognition result to generate a standard query request; calculating the similarity between the standard query and each semantic fragment pre-stored in the knowledge base, and selecting the semantic fragment with the highest similarity as the target semantic fragment; wherein, each semantic fragment in the knowledge base is content obtained according to the document processing method of any of the above items; inputting the query request and the target semantic fragment into a second language model, and generating and saving the response result for the query request based on the retrieval enhancement generation mechanism.
[0009] In one embodiment of the present invention, a first language model is invoked to perform intent recognition on the query request, and semantic completion is performed on the query request based on the recognition result to generate a standard query request. This includes: determining whether the query request is the first query request; if so, the query request is input into the first language model to perform intent recognition on the query request, and semantic completion is performed according to preset business rules based on the recognized intent to generate a standard query request; otherwise, the following process is executed: retrieving historical dialogue information associated with the query request; wherein, the historical dialogue information includes at least one historical query request and its corresponding historical response result; inputting the query request and historical dialogue information together into the first language model to identify the content association between the query request and the historical information, and performing referential resolution and omission completion on the query request based on the identified content association to generate a standard query request.
[0010] In one embodiment of the present invention, after generating a response result for a query request, the method further includes: extracting and displaying the document and / or the link to the document and the corresponding chapter from the knowledge base.
[0011] This invention also provides a document storage system, comprising: a document acquisition module for acquiring documents to be analyzed; a content extraction module for identifying document types and extracting document content based on the identified document types; a category determination module for calculating the semantic similarity between the document content and multiple pre-stored standard classification contents, and determining the content category corresponding to the document based on the semantic similarity; wherein each standard classification content corresponds to one content category; a chapter summary extraction module for dividing the document content into chapters and determining the chapter summary of each chapter; wherein each chapter corresponds to one chapter summary; a semantic recognition module for performing semantic recognition on each chapter and dividing each chapter into multiple semantic segments; and a storage module for structurally storing the document according to the hierarchical relationship between content categories, chapter summaries, and semantic segments.
[0012] The present invention also provides an electronic device, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device enables the document storage method or the intelligent question-answering method described above.
[0013] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer processor, causes the computer to execute any of the document storage methods or intelligent question-answering methods described above.
[0014] The beneficial effects of this invention are as follows: This invention proposes a document storage method, intelligent question-answering method, system, device, and medium. By identifying the type of the document to be analyzed and extracting its content in a corresponding manner, it can adapt to various document formats such as PDF, Word, images, and scanned documents, achieving unified processing of multi-source heterogeneous documents. By calculating the semantic similarity between the document content and multiple standard classification contents, the document can be categorized holistically. Furthermore, by dividing the document content into chapters and determining corresponding chapter summaries, combined with semantic recognition of each chapter, each chapter is split into multiple semantically coherent semantic fragments. Based on the hierarchical structure between content categories, chapter summaries, and semantic fragments, the entire document is stored in a multi-layered structured manner. This approach not only improves the adaptability to multi-source heterogeneous documents but also enables more accurate and fine-grained structured information extraction from complex document content. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0016] In the attached diagram:
[0017] Figure 1 A flowchart illustrating a document storage method according to an embodiment of the present invention;
[0018] Figure 2 This is a flowchart illustrating an intelligent question-answering method provided in an embodiment of the present invention;
[0019] Figure 3 This is a structural block diagram of a document storage system provided in one embodiment of the present invention;
[0020] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0021] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0022] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0023] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0024] The inventors discovered that current mainstream knowledge base management systems (such as customer service Q&A systems) typically include three core modules: indexing, retrieval, and generation. The indexing module is responsible for converting raw document data into vectors for storage in the knowledge base. The retrieval module vectorizes user queries and retrieves relevant document content from the knowledge base based on a similarity matching mechanism. The generation module utilizes a large language model to combine user queries with retrieval results to generate the final answer.
[0025] However, the aforementioned knowledge base management systems still have many shortcomings in practical applications: For example, in the data preprocessing stage, relying solely on raw data cleaners is insufficient to handle complex document structures (such as nested tables, multi-level headings, or semantically hierarchical paragraphs), easily leading to inaccurate data parsing and unreasonable content segmentation. In the retrieval stage, irrelevant documents may interfere with result generation, thus reducing the accuracy of the generated results. In the result generation stage, the lack of a source tracing mechanism after answer generation results in low credibility and is prone to illusion problems. Furthermore, current knowledge base management systems generally lack paragraph-level source tracing capabilities, only providing general answers and failing to directly link to specific clauses in the original text. Regarding multimodal processing, traditional Retrieval-augmented Generation (RAG) mechanisms generally only support text input and text output, making it difficult to support the multimodal information needs of intelligent question-answering scenarios. On the other hand, in dialogue scenarios involving continuous follow-up questions, traditional RAG methods struggle to fully utilize contextual information and accurately understand the user's true intent. This results in insufficient question mining and inaccurate answer generation, significantly reducing the user's information acquisition efficiency and overall user experience. Furthermore, in terms of access control, most current knowledge base management systems fail to implement dynamic access control based on user identity or organizational structure, making it difficult to meet the hierarchical access needs of different business lines within an enterprise, such as Party and Mass Affairs or Business.
[0026] To address the aforementioned problems, this invention provides a document storage method. By identifying the type of the document to be analyzed and extracting its content according to the corresponding method, it can adapt to various document formats such as PDF, Word, images, and scanned documents, achieving unified processing of multi-source heterogeneous documents. Semantic similarity is calculated between the document content and multiple standard classification content to categorize the document holistically. Furthermore, the document content is divided into chapters and corresponding chapter summaries are determined. Combined with semantic recognition of each chapter, each chapter is split into multiple semantically coherent segments. Based on the hierarchical structure between content categories, chapter summaries, and semantic segments, the document content is mapped into a set of structured knowledge units, which are then written into a knowledge base hierarchically. This approach not only improves the adaptability to multi-source heterogeneous documents but also enables more accurate and fine-grained structured extraction of complex document content.
[0027] like Figure 1 As shown, the document storage method proposed in this invention, applied to an enterprise's internal knowledge base question-and-answer platform, includes the following steps:
[0028] S11. Obtain the document to be analyzed.
[0029] This invention's knowledge base question-and-answer platform connects to various data sources through a unified document input interface, supporting multiple document access methods. It can automatically acquire documents from internal enterprise policy publishing platforms, departmental cloud drives, shared directories, and personal local files, and also allows users to upload relevant documents themselves; the specific method is not limited. This knowledge base question-and-answer platform supports diverse methods such as online real-time acquisition and offline upload to ensure effective management of internal enterprise knowledge. During document acquisition, metadata related to the document can be simultaneously obtained, including but not limited to the document's department, creation time, and version number, facilitating subsequent document classification and management.
[0030] S12. Perform document type identification and extract document content from the document based on the identified document type.
[0031] Since this invention supports various document formats such as PDF, Word, Excel, TXT, scanned documents, and images, it is necessary to perform type identification on the acquired documents to determine their type in order to effectively process these documents. It is understood that type identification can be based on file extensions. Based on the identified document type, the document content is extracted using the corresponding content extraction strategy.
[0032] In an optional embodiment of the present invention, step S12 includes the following data processing procedure: performing type identification on the document to determine the document type: if the document is in editable text format, then calling the text parsing engine to extract the document content from the document; if the document is in image format, then calling the optical character recognition engine or visual analysis model to extract the document content from the document.
[0033] Specifically, the acquired documents undergo type identification. After determining the document type, if the document is in editable text format, the corresponding text parsing engine is invoked based on the specific document type to extract the structured content. Editable text formats include, but are not limited to, Word, TXT, Excel, or editable PDF. Specifically, for Word or TXT documents, paragraph content and their hierarchical structure can be extracted by invoking a text parsing engine; for Excel documents, the table's row and column structure, cell ranges, and merging relationships can be identified, and a text extraction strategy can be applied to each cell to extract the structured row and column information; for editable PDFs, the main text content can be extracted using open-source language model processing tools such as MinerU to achieve document structure parsing. To enhance subsequent processing efficiency, the extracted document content can be uniformly saved in Markdown format to preserve the document's semantic hierarchy. Furthermore, considering the potential for line breaks, header and footer interference, etc., in the original document, for editable text formats, a text parsing engine is invoked to extract document content. This process includes: text cleaning of the document; and invoking the text parsing engine to extract document content from the cleaned document. Text cleaning methods include, but are not limited to, removing redundant spaces, line breaks, headers, and footers. The resulting document is then processed using a text parsing engine to extract structured information, ensuring the accuracy of the parsing results.
[0034] For image-based formats, text information can be extracted by either using an optical character recognition (OCR) engine to identify characters within the image content or by analyzing the document's layout using a visual analysis model to identify text regions, thus making the extracted content more accurate. The visual analysis model can be any model with image-text layout understanding capabilities, including but not limited to DBNet, PGNet, or TextSpotter models. Furthermore, to facilitate subsequent original image display and source tracing, the original image path will be saved.
[0035] Furthermore, considering that PDF files include editable PDF documents and image-based PDF documents, if a document is identified as PDF format, the document type will be determined by parsing the character stream or analyzing whether the information in the document is vector text. For image-based PDFs, the strategy for image formats will be applied; for editable PDFs, the strategy for editable text formats will be applied. This dual-channel processing strategy ensures good compatibility when dealing with different PDF document formats.
[0036] Furthermore, in another optional embodiment of the present invention, considering that a single document may contain multiple content formats such as text, images, and tables, a layout structure analysis is performed on the document after document type identification to identify and extract different content areas within the document. Specifically, layout analysis algorithms identify content blocks such as text areas, image areas, and table areas in the document. The aforementioned editable text format strategy is applied to text and table areas, while the aforementioned image format strategy is applied to image areas. Through multi-channel parallel processing, different content types are uniformly structured and output, forming a complete and semantically clear expression of document content, thereby providing a more accurate data foundation for subsequent chapter summary generation and semantic segmentation. It should be noted that the aforementioned layout analysis algorithm refers to a method for identifying different content areas within a page of a document, including but not limited to connected component analysis, projection methods, or deep learning-based layout understanding models (such as the LayoutLM model, the Donut model, etc.).
[0037] S13. Calculate the semantic similarity between the document content and multiple pre-stored standard classification contents, and determine the content category corresponding to the document based on the semantic similarity results; wherein, the standard classification contents correspond one-to-one with the content category.
[0038] To automate document categorization, document content can first be encoded to generate semantic vectors. The knowledge base question-answering platform pre-stores a set of standard classification content, with each standard classification corresponding to a content category. The content category represents the semantic features of that type of document. The standard classification content is pre-constructed by the knowledge base question-answering platform through automatic collection or manual input. Encoding the standard classification content using the same encoding method as the document content yields its corresponding standard semantic vector. By calculating the semantic similarity between the semantic vector and the standard semantic vector, the content category matching the document to be analyzed is determined, thus achieving automated semantic categorization of the document. The calculation methods for semantic similarity include, but are not limited to, cosine similarity, Euclidean distance, or dot product similarity. It should be noted that determining the document's content category based on semantic similarity results can be achieved by selecting content categories with semantic similarity greater than or equal to a preset similarity threshold, or by selecting the content category with the highest semantic similarity. Those skilled in the art can adapt the specific method to the actual scenario, and no limitations are imposed here. Therefore, a document to be analyzed can have one content category or multiple content categories.
[0039] Furthermore, the knowledge base question-answering platform can jointly model the metadata related to documents (such as the document's department, creation time, and version number) acquired during document collection with the corresponding document content. This metadata can be used as additional features and co-encoded with the document content to enhance the accuracy of document content category judgment. For manually uploaded documents to be analyzed, after extracting their semantic vectors, semantic similarity calculations can be performed between the semantic vectors and various standard semantic vectors to obtain preliminary classification suggestions. The results are then manually confirmed and labeled, and the confirmed document content and its classification tags are written into the knowledge base, thereby improving the accuracy of structured document storage.
[0040] In an optional embodiment of the present invention, step S13 includes the following data processing steps: parsing the document content and determining whether there is a text region that satisfies the preset summary features; if there is, then the text content of the text region is used as the document summary of the document content; otherwise, the document content is input into the summary extraction model to generate the document summary of the document content; wherein, the summary extraction model is a generative model; calculating the semantic similarity between the document summary and each standard classification content, and selecting the content category with a semantic similarity greater than a preset similarity threshold as the content category corresponding to the document.
[0041] Encoding the entire document is time-consuming and inaccurate. Calculating semantic similarity between the encoded result and standard classification content leads to poor accuracy and speed in semantic similarity calculations. To improve processing speed and accuracy, considering that some documents have explicitly marked summary information during typesetting, a rule engine can be used to analyze the layout structure of the document content on a preset target page. The document content of the target page is divided into multiple text regions, where the target page represents the page in the document where summary information is frequently displayed. For each text region, it is judged according to preset summary features. These features include, but are not limited to, positional features (such as the top of the page or above the main text), font style features (such as bold or centered), and keyword features (such as summary, overview, etc.). If a text region matching the summary features is detected, the text content of that region is taken as the document summary. Conversely, if no text region satisfying the summary features is detected, it indicates that the current document does not explicitly provide summary information. In this case, the entire document content can be input into a pre-trained generative summary extraction model to extract semantic features from the document content and generate a document summary representing the core content of the current document based on the extracted semantic features. After obtaining the document summary, it is encoded to generate a semantic vector. The semantic similarity of the semantic vector is calculated with multiple pre-stored standard semantic vectors, and the standard classification content with a semantic similarity greater than a preset similarity threshold is selected as the content category of the current document. Preferably, considering that different types of documents have different page layouts, the target page is different for different types of documents.
[0042] S14. Divide the document content into chapters and determine the chapter summary for each chapter; each chapter corresponds to one chapter summary.
[0043] In an optional embodiment of the present invention, step S14 includes the following process: based on a preset chapter division rule, determine whether the document content has a chapter structure; if so, divide the document content into multiple chapters based on the hierarchical relationship of the chapters; otherwise, input the document content into a text segmentation model, perform semantic analysis on the document content, and divide the document content into multiple chapters based on the analysis results; wherein, the text segmentation model is a temporal model; input each chapter into a summary generation model to generate a corresponding chapter summary; wherein, the summary generation model is a generative model or an extractive model.
[0044] The document content is structured and parsed according to preset chapter division rules to determine if a chapter structure exists. If it does, the document content is segmented according to the identified hierarchical relationship between chapters, dividing a complete document into multiple hierarchical chapters. For example, for a Markdown document containing chapter identifiers such as "Chapter 1" and "Chapter 2," chapter boundaries can be identified based on rules such as heading numbering, font style, and indentation structure, and the complete document content can be segmented into multiple hierarchical chapters based on the heading hierarchy. Conversely, if the document does not have a chapter structure, the entire document content is input into a pre-trained text segmentation model. Based on each text statement and its context, semantic breakpoints are identified, and the document content is divided into multiple semantically coherent chapters. For each of the aforementioned chapters, the following processing is performed: the chapter is input into a summary generation model to generate a chapter summary. When the summary generation model is a generative model, it encodes the input chapters into text and generates chapter summaries based on the encoding results. When the summary generation model is an extractive model, it scores each text statement in the chapter and selects a preset number of text statements with the highest scores to arrange them in order to obtain chapter summaries.
[0045] S15. Perform semantic recognition on each chapter and divide each chapter into multiple semantic segments.
[0046] To further enhance the semantic understanding of the document, after completing the chapter division, each chapter needs to be further subdivided at the semantic level. Specifically, each chapter is processed as follows: the content of the chapter is segmented into sentences to obtain multiple text statements, and each text statement is encoded. A corresponding text statement vector is obtained by combining the context information of the text statement. After obtaining the text statement vectors of all text statements in the chapter in the above manner, the semantic similarity between adjacent text statement vectors is calculated to detect the semantic continuity between adjacent text statements. When the semantic similarity is less than a preset semantic similarity threshold, the adjacent text statements are determined to be semantic boundaries. Based on these semantic boundaries, the entire chapter is divided into multiple semantically coherent and appropriately sized semantic segments. For example, for the i-th text statement and the (i+1)-th text statement in a chapter, if their semantic similarity is less than the semantic similarity threshold, the i-th text statement is determined to be the ending sentence of the current semantic segment, and the (i+1)-th text statement is determined to be the starting sentence of a new semantic segment. Among them, semantic fragments consist of several semantically related text statements, which can express a complete information point or business element relatively independently, and are the most basic semantic units for building a knowledge base.
[0047] For example, for a certain financial document, its content category is financial system, and the corresponding three chapter summaries are, in order, detailed description of expense reimbursement process, explanation of approval authority, and requirements for filing financial vouchers. For the first chapter summary (i.e., detailed description of expense reimbursement process), its three semantic fragments are as follows: Semantic fragment 1: When employees submit expense reimbursements for business trips, they need to provide invoices, expense lists, and business trip approval forms, and provide these to the finance department; Semantic fragment 2: When the finance department reviews invoices, they need the signature of the superior manager before they can be entered into the accounts; Semantic fragment 3: All original vouchers should be filed to the financial archives within 5 working days after the accounting processing is completed, ensuring that the filing is complete and without omissions.
[0048] S16. Store the document in a structured manner according to the hierarchical relationship between content categories, chapter summaries, and semantic fragments.
[0049] To achieve structured management of document content, this invention stores document information hierarchically based on the aforementioned three-layer semantic relationship of content category, chapter summary, and semantic fragment. Specifically, the content category serves as the first-level index node in the knowledge base to achieve overall document classification and archiving. Under the first-level index node of the knowledge base, a set of chapter summaries corresponding to the content category is established, and each chapter summary serves as a second-level index node to represent the semantic segmentation result of the corresponding document. Under each second-level index node, each semantic fragment corresponding to the chapter summary serves as a third-level index node to represent a finer-grained semantic unit in the document. Furthermore, to support subsequent semantic retrieval and intelligent question-answering services, the third-level index node stores the original text content of the semantic fragment, the corresponding semantic vector, and the location information of the document and chapter, etc., to establish the semantic association between the semantic fragment and its context. Thus, each document corresponds to at least one content category and at least one chapter summary, and each chapter summary corresponds to at least one semantic fragment. Through the above multi-level association, each semantic fragment is implicitly associated with its respective chapter summary, content category, and document. Therefore, through this invention's three-layer semantic structure of content category-chapter summary-semantic fragment, a mapping relationship from fine-grained structured semantic units to the document level is achieved. This constructs a semantic indexing system in the knowledge base, comprising macro-level content classification, meso-level chapter classification, and micro-level semantic fragment classification. This ensures that each semantic fragment can be traced back to its respective chapter, content category, and document, thereby improving the accuracy of subsequent precise semantic retrieval.
[0050] like Figure 2 As shown, the present invention also provides an intelligent question-answering method, comprising the following process:
[0051] S21. Obtain and save the query request.
[0052] After receiving a user's query request in natural language, the knowledge base question-answering platform needs to encapsulate the request for easier retrieval later. The query request contains company-related information the user wants to obtain, such as "What are the company's travel expense reimbursement standards?". Specifically, upon receiving the user's query request, a unique query identifier is assigned, and the user's identity identifier is obtained. This query identifier, identity identifier, query content, and the time the query was retrieved are then encapsulated and saved.
[0053] S22. Call the first major language model to perform intent recognition on the query request, and perform semantic completion on the query request based on the recognition results to generate a standard query request.
[0054] The query request is input into the pre-trained first language model, which identifies the query intent expressed in the query request. Based on the identification results and preset business rules, the model completes any omitted or ambiguous information in the query request, thereby generating a semantically clear standard query request.
[0055] In an optional embodiment of the present invention, step S22 includes the following process: determining whether the query request is the first query request; if so, inputting the query request into the first language model, performing intent recognition on the query request, and performing semantic completion based on the recognized intent according to preset business rules to generate a standard query request; otherwise, performing the following process: retrieving historical dialogue information associated with the query request; wherein, the historical dialogue information includes at least one historical query request and its corresponding historical response result; inputting the query request and historical dialogue information together into the first language model, identifying the content association between the query request and the historical information, and performing referential resolution and omission completion on the query request based on the identified content association to generate a standard query request.
[0056] Specifically, considering the potential for multiple rounds of continuous dialogue, for each user's query request, it is determined whether the query request is the first query request in the current dialogue. If it is the first query request, it can be directly input into the first language model for intent recognition. Based on the recognition results, and according to preset business rules, semantic completion is performed on missing or ambiguous parts of the query request to generate a semantically clear standard query request. Conversely, if the query request is not the first query request in the current dialogue, historical dialogue information associated with it can be retrieved. Historical dialogue information refers to query requests and corresponding responses from a preset number of rounds (e.g., 5 rounds) prior to the current query request. The current query request and historical dialogue information are input into the first language model together to analyze the semantic relationship between them. This identifies any pronouns, omissions, or incomplete statements in the current query request. Based on the historical dialogue information, the current non-standard query request is semantically expanded and completed. By combining the context vector matching results and fusing semantic cues from multiple rounds of dialogue, a standard query request that fully and accurately expresses the user's true intent is obtained.
[0057] For example, in the first round, the user's query is "What are the characteristics of a sale-leaseback product?", and in the second round, the follow-up query is "What are its applicable scenarios?". Based on historical dialogue information, it is identified that "it" refers to "sale-leaseback," and the second-round query is semantically completed, resulting in the standard query "What are the applicable scenarios for sale-leaseback?". Continuing the dialogue, in the third round, the user's query is "The kind that requires prepayment." This query is semantically incomplete. Combining this with the previous round's "applicable scenarios for sale-leaseback," the semantically completed standard query is "What are the requirements for the leased asset in a sale-leaseback?". In the fourth round, the user's query is "How is the payment schedule designed for the model we discussed last time?". Based on historical dialogue information, it is identified that "the model we discussed last time" refers to "sale-leaseback," therefore the standard query is "How is the payment schedule designed for a sale-leaseback?".
[0058] S23. Calculate the similarity between the standard query and each semantic fragment pre-stored in the knowledge base, and select the semantic fragment with the highest similarity as the target semantic fragment; wherein, each semantic fragment in the knowledge base is the content obtained according to the document processing method of any of the above items.
[0059] An embedding model can be used to encode a standard query, generating a standard query encoding vector. This vector is then compared with the semantic fragment vectors corresponding to various semantic fragments in the knowledge base to calculate their similarity, determining the matching degree. The semantic fragment with the highest similarity is selected as the target semantic fragment for subsequent question-answering generation. The semantic fragments are structured content obtained through document content classification, chapter summary extraction, and chapter semantic segmentation, as described in this invention. Embedding models include, but are not limited to, BERT, ELMo, and Word2Vec; any model capable of encoding text statements into vectors is acceptable and is not limited here.
[0060] In an optional embodiment of the present invention, after generating the response result for the query request, the method further includes: extracting and displaying the document and / or the document and corresponding chapter link corresponding to the response result from the knowledge base. Specifically, based on the aforementioned matched target semantic fragment, its structured storage path in the knowledge base is traced back to determine the original document, chapter, and location corresponding to the semantic fragment. Document identification information associated with the semantic fragment is extracted from the knowledge base, and corresponding document links, or document links and chapter links, are generated according to the document storage path. The response result, the source of the corresponding answer, and the traced document or document and corresponding chapter link are then displayed to the user. Through this original text tracing mechanism, users can view the original materials on which the response is based when they receive the response result. This reduces the model illusion problem and facilitates user verification of information, thereby improving the interpretability and credibility of the question and answer.
[0061] S24. Input the query request and target semantic fragment into the second language model, and generate and save the response results for the query request based on the retrieval enhancement generation mechanism.
[0062] The query request and the target semantic fragment obtained from the previous step are input into the second large language model. Using a retrieval enhancement generation mechanism, the query request is subjected to content understanding, and deep semantic fusion is performed by combining the contextual information carried by the target semantic fragment to generate a response result that semantically matches the query request. The response result is then bound to and saved with the corresponding query request for subsequent multi-turn dialogue queries. It should be noted that the first and second large language models of this invention can be the same large language model, or they can be two large language models with different parameters or network structures; no specific limitation is made. The first and second large language models of this invention include, but are not limited to, Deepseek, GPT series models, Baichuan series models, etc., and no specific limitation is made.
[0063] Furthermore, in the stage where the second large language model generates the response result, after obtaining the user's query request, semantic fragments corresponding to the query request can be retrieved and recalled from the vector database of the knowledge base Q&A platform through semantic vector matching as the semantic support for the response. The query request and the corresponding semantic fragments are jointly input into the second large language model, combined with the pre-configured prompt optimization strategy to guide the response result generated by the second large language model to be more standardized. Specifically, the prompt optimization strategy is a combination of multiple methods such as role binding, thought chain guidance, few-shot examples, and knowledge base enhanced prompts. Furthermore, personalized prompts can also be generated in combination with specific enterprise scenarios. Exemplarily, business term standardization binding can be performed, that is, the key terms used in the generation of the second large language model are unified, such as standardizing "customer" to "lessee", "our company" to "lessor", and "equipment" to "leased item" to ensure that the term expressions conform to the company's specifications. Numerical expression standardization can also be performed, that is, it is required that the second large language model strictly follows the company's numerical expression specifications, such as the amount unit is unified as "ten thousand yuan" (for example, 15.6 million yuan), the proportion is expressed as "30%" instead of "thirty percent", and the date format is unified as "YYYY-MM-DD" to ensure the consistency of the response result format. Through the above prompt optimization strategy, combined with the original text tracing mechanism, the response result generated by the second large language model can be more accurate and professional.
[0064] The present invention takes into account the requirements of enterprise permission control, so it combines the preset access permissions of the second large language model and the permission mechanism preset in the knowledge base content to implement the access strategy according to the user's identity identifier, so as to ensure that different user groups can only access the content within their permission scope, thereby ensuring the security and compliance of data. Furthermore, the present invention also sends a feedback report to the user after generating the response result for collecting the evaluation information of the user on the response result. After receiving the feedback report filled in and submitted by the user, by analyzing the feedback content, feedback information such as the satisfaction score and error correction suggestions are extracted, and according to these feedback information, the knowledge base content is updated regularly or irregularly, or the generation strategies of the first large language model and / or the second large language model are optimized to improve the accuracy of subsequent Q&A services, thereby greatly improving user satisfaction.
[0065] Such as Figure 3As shown, the document storage system 300 includes: a document acquisition module 310, a content extraction module 320, a category determination module 330, a chapter summary extraction module 340, a semantic recognition module 350, and a storage module 360. The document acquisition module 310 acquires the document to be analyzed. The content extraction module 320 performs type identification on the document and extracts document content based on the identified document type. The category determination module 330 calculates the semantic similarity between the document content and multiple pre-stored standard classification contents, and determines the content category corresponding to the document based on the semantic similarity; each standard classification content corresponds to one content category. The chapter summary extraction module 340 divides the document content into chapters and determines the chapter summary for each chapter; each chapter corresponds to one chapter summary. The semantic recognition module 350 performs semantic recognition on each chapter, dividing each chapter into multiple semantic segments. The storage module 360 performs structured storage of the document according to the hierarchical relationship between content categories, chapter summaries, and semantic segments.
[0066] Specific limitations regarding the document storage system can be found in the limitations on document storage methods described above, and will not be repeated here. Each module in the aforementioned document storage system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware format, or stored in the memory of a computer device in software format, so that the processor can invoke the corresponding operations of each module.
[0067] It should be noted that, in order to highlight the innovative aspects of this invention, this embodiment does not include modules that are not closely related to solving the technical problems proposed by this invention, but this does not mean that there are no other modules in this embodiment.
[0068] like Figure 4 As shown, electronic device 4 may include memory 41, processor 42 and bus, and may also include computer programs stored in memory 41 and executable on processor 42, such as document storage programs or intelligent question-and-answer programs.
[0069] The memory 41 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the electronic device 4, such as a portable hard drive. In other embodiments, the memory 41 can be an external storage device of the electronic device 4, such as a plug-in portable hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 4. Furthermore, the memory 41 can include both internal and external storage units of the electronic device 4. The memory 41 can be used not only to store application software and various types of data installed on the electronic device 4, such as document storage code or intelligent question-and-answer code, but also to temporarily store data that has been output or will be output.
[0070] In some embodiments, the processor 42 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 42 is the control unit of the electronic device 4, connecting various components of the electronic device 4 via various interfaces and lines. It executes programs or modules stored in the memory 41 (such as document storage programs or intelligent question-and-answer programs) and calls data stored in the memory 41 to perform various functions and process data in the electronic device 4.
[0071] Processor 42 executes the operating system of electronic device 4 and various installed applications. Processor 42 executes applications to implement the steps in the document storage method or intelligent question-and-answer method described above.
[0072] For example, a computer program can be divided into one or more modules, one or more of which are stored in memory 41 and executed by processor 42 to complete the present invention. One or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in electronic device 4. For example, the computer program can be divided into a document acquisition module 310, a content extraction module 320, a category determination module 330, a chapter summary extraction module 340, a semantic recognition module 350, and a storage module 360.
[0073] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The software functional module stored in the storage medium includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute part of the functions of the document storage method or intelligent question-answering method of the various embodiments of the present invention.
[0074] In summary, this invention, by displaying relevant original texts and their links when generating response results, generates results based on facts, significantly reducing the "illusion" problem of the second major language model. Compared to traditional generative models, this invention's original text tracing mechanism improves the accuracy of answer generation by more than 30%. Furthermore, this method can dynamically update the knowledge base by acquiring documents to be analyzed in real time, enabling rapid response to business changes such as internal policy adjustments and product iterations, enhancing the real-time nature and adaptability of question-and-answer results. Further, in handling complex query tasks, this invention's hierarchical document storage method supports multi-document linked retrieval and semantic fusion generation, intelligently integrating cross-departmental and multi-dimensional information queries, thereby improving the ability to respond to comprehensive questions. In addition, this automated data processing flow reduces manual intervention, significantly lowering the maintenance cost of the knowledge base and improving knowledge update efficiency by more than 50%. For the user end, it supports multimodal output formats presenting text answers and images, combined with natural language interaction and answer source tracing mechanisms, improving information credibility and user convenience. Furthermore, to ensure data security, a dynamic access control mechanism based on departments and roles has been introduced to achieve precise isolation of information access and prevent unauthorized behavior.
[0075] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for storing documents, characterized in that, The method includes: Obtain the document to be analyzed; The document is type-identified, and the document content is extracted from the document based on the identified document type; Calculate the semantic similarity between the document content and multiple pre-stored standard category contents, and determine the content category corresponding to the document based on the semantic similarity; wherein, each standard category content corresponds to one content category; The document content is divided into chapters, and a chapter summary is determined for each chapter; wherein, each chapter corresponds to one chapter summary; Semantic recognition is performed on each chapter, and each chapter is divided into multiple semantic segments; The document is stored in a structured manner based on the hierarchical relationship between content categories, chapter summaries, and semantic fragments.
2. The document storage method according to claim 1, characterized in that, The document is type-identified, and based on the identified document type, document content is extracted from the document, including: Perform type identification on the document to determine its document type: If the document is in an editable text format, then the text parsing engine is invoked to extract the document content from the document; If the document is in image format, then an optical character recognition engine or visual analysis model is invoked to extract the document content from the document.
3. The document storage method according to claim 1, characterized in that, Calculate the semantic similarity between the document content and multiple pre-stored standard category contents, and determine the content category corresponding to the document based on the semantic similarity results, including: The document content is parsed to determine whether there are text regions that meet preset summary characteristics: If it exists, the text content of the text region will be used as the document summary of the document content; Otherwise, the document content is input into the summary extraction model to generate a document summary of the document content; wherein, the summary extraction model is a generative model; Calculate the semantic similarity between the document summary and the content of each standard classification, and select the content category with a semantic similarity greater than a preset similarity threshold as the content category corresponding to the document.
4. The document storage method according to claim 1, characterized in that, The document content is divided into chapters, and a chapter summary for each chapter is determined, including: Based on preset chapter division rules, determine whether the document content has a chapter structure: If so, the document content will be divided into multiple chapters based on the hierarchical relationship between chapters; Otherwise, the document content is input into a text segmentation model to perform semantic analysis on the document content, and the document content is divided into multiple chapters based on the analysis results; wherein, the text segmentation model is a time-series model; Each chapter is input into the summary generation model to generate a corresponding chapter summary; wherein, the summary generation model is a generative model or an extractive model.
5. An intelligent question-answering method, characterized in that, The method includes: Retrieve and save the query request; The first major language model is invoked to perform intent recognition on the query request, and semantic completion is performed on the query request based on the recognition results to generate a standard query request; Calculate the similarity between the standard query and each semantic fragment pre-stored in the knowledge base, and select the semantic fragment with the highest similarity as the target semantic fragment; wherein, each semantic fragment in the knowledge base is the content obtained by the document processing method according to any one of claims 1-4; The query request and the target semantic fragment are input into the second language model, and a response result for the query request is generated and saved based on the retrieval enhancement generation mechanism.
6. The intelligent question-answering method according to claim 5, characterized in that, The process involves invoking the first major language model to perform intent recognition on the query request, and then semantically completing the query request based on the recognition results to generate a standard query request, including: Determine whether the query request is the first query request: If so, the query request is input into the first large language model, the intent of the query request is identified, and semantic completion is performed according to the preset business rules based on the identified intent to generate a standard query request; Otherwise, proceed as follows: Retrieve historical dialogue information associated with the query request; wherein, the historical dialogue information includes at least one historical query request and its corresponding historical response results; The query request and the historical dialogue information are input into the first large language model to identify the content association between the query request and the historical information. Based on the identified content association, the query request is resolved by resolving the reference and completing the omitted statements to generate a standard query request.
7. The intelligent question-answering method according to claim 5, characterized in that, After generating the response result for the query request, the method further includes: extracting and displaying the document and / or the link to the document and corresponding chapter from the knowledge base.
8. A document storage system, characterized in that, The system includes: The document acquisition module is used to acquire the documents to be analyzed. The content extraction module is used to identify the document type and extract the document content from the document based on the identified document type. The category determination module is used to calculate the semantic similarity between the document content and multiple pre-stored standard category contents, and determine the content category corresponding to the document based on the semantic similarity; wherein, each standard category content corresponds to one content category; The chapter summary extraction module is used to divide the document content into chapters and determine the chapter summary of each chapter; wherein, each chapter corresponds to one chapter summary; The semantic recognition module is used to perform semantic recognition on each chapter and divide each chapter into multiple semantic segments. The storage module is used to perform structured storage of the document according to the hierarchical relationship between content categories, chapter summaries, and semantic fragments.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the document storage method of any one of claims 1 to 4 or the intelligent question-answering method of any one of claims 5 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the document storage method of any one of claims 1 to 4 or the intelligent question-answering method of any one of claims 5 to 7.
Citation Information
Patent Citations
Knowledge filing method, device and equipment, storage medium and computer program product
CN119293109A
Multi-modal intelligent question-answering system based on large model and construction method and device
CN119783819A
Question and answer method and system for structured long document
CN119848223A
Financial intelligent question and answer method and system based on mixed retrieval and dynamic query
CN120256574A
Cited By
Heterogeneous information-based large language model extension method and system, and storage medium
CN121683963A