Knowledge question and answer method, system and device, electronic equipment and computer program product

By constructing a knowledge question-answering system to clean and graph documents in a specified field, and combining it with a knowledge retrieval enhancement module, the problem of insufficient training corpus for large language models in the telecommunications industry is solved, and efficient knowledge query and processing of non-standard natural language data are achieved.

CN120929559APending Publication Date: 2025-11-11CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410578998.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Large-scale language models suffer from insufficient corpus in specific vertical fields such as the telecommunications industry, resulting in inaccurate answers and low information extraction efficiency when faced with highly specialized questions and insufficient training samples.

Method used

A pre-built knowledge question-answering system is used to perform data cleaning, text slicing, and tabular knowledge graph construction on documents to be processed in a specified domain. Text vector representations are generated by combining text slices, and question retrieval is performed through a knowledge retrieval enhancement module. Table data is stored and queried using a knowledge base to generate recommended answers.

Benefits of technology

It enables effective knowledge retrieval in a specified domain even with insufficient training corpus, enhances the model's ability to process non-standard natural language data, and improves the accuracy and efficiency of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929559A_ABST
    Figure CN120929559A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge question answering method, system and device, electronic equipment and a computer program product, relates to the technical field of natural language processing, and is applied to scenes for providing matched answers for user questions in different fields. The method comprises the steps that a pre-constructed knowledge question-answering system is adopted to conduct knowledge retrieval processing on a target question, a target answer corresponding to the target question is obtained, and the knowledge question-answering system is obtained by training recommended question-answering pairs of a specified field, the recommended question and answer pair in the specified field is a feedback result of a to-be-processed document in the specified field after knowledge retrieval enhancement processing, and the to-be-processed document in the specified field comprises table data. Through the knowledge retrieval enhancement scheme, when the problem of insufficient training corpora is faced, knowledge query in a specified field can still be effectively solved; and knowledge graph storage is carried out on table data, so that the model has the capability of processing non-standard natural language data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of natural language processing technology, and more specifically, to a knowledge question answering method, a knowledge question answering system, a knowledge question answering device, an electronic device, and a computer program product. Background Technology

[0002] In today's era of rapid information technology development, electronic communication, as the core of information communication, has accumulated a wealth of industry data and professional knowledge. This knowledge, especially documents such as operation manuals and case studies in tabular form, is of immense value for business decision-making, service improvement, and problem-solving. However, the efficient extraction and application of this knowledge, particularly in non-standard natural language contexts, poses a significant challenge to traditional knowledge retrieval technologies. For example, large language models perform remarkably well with standard language input, but in specific vertical fields (such as the telecommunications industry), they are often limited by their highly specialized nature and insufficient training samples, resulting in insufficient accuracy and low information extraction efficiency in practical applications.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this disclosure is to provide a knowledge question answering method, knowledge question answering system, knowledge question answering device, electronic device, and computer program product, thereby overcoming, to at least a certain extent, the problem that the model cannot learn relevant knowledge due to insufficient corpus during model training, and the lack of a knowledge retrieval enhancement scheme for tabular data.

[0005] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to a first aspect of this disclosure, a knowledge question answering method is provided, comprising: using a pre-constructed knowledge question answering system to perform knowledge retrieval processing on a target question to obtain a target answer corresponding to the target question, wherein the knowledge question answering system is obtained by training on recommended question-answer pairs in a specified domain, wherein the recommended question-answer pairs in the specified domain are feedback results of documents to be processed in the specified domain after knowledge retrieval enhancement processing, and the documents to be processed in the specified domain include tabular data.

[0007] In one exemplary embodiment of this disclosure, the method further includes: performing data cleaning and text slicing processing on the acquired original documents in a specified domain to obtain a document to be processed, the document to be processed including text slices; constructing a tabular knowledge graph based on the document to be processed; performing graph textization processing on the tabular knowledge graph; and generating a text vector representation corresponding to the document to be processed by combining the text slices.

[0008] In one exemplary embodiment of this disclosure, the step of constructing a table knowledge graph based on the document to be processed includes: performing table extraction processing on the document to be processed to obtain table data, the table data including table context information, the table context information including table content and content hierarchy relationship; performing normalization processing on the table data to obtain normalized table data; and performing entity-relationship mapping processing on the normalized table data to obtain the table knowledge graph.

[0009] In one exemplary embodiment of this disclosure, the step of performing graph textification processing on the tabular knowledge graph and generating a text vector representation corresponding to the document to be processed in combination with the text slice includes: retrieving entity and relation placeholders from the tabular knowledge graph according to a predefined text template; filling the entity and relation placeholders based on the retrieved entity types and context information to obtain natural language table data; and performing text vectorization processing on the natural language table data and the text slice to obtain the text vector representation.

[0010] In one exemplary embodiment of this disclosure, the method further includes: extracting and vectorizing the acquired user questions to obtain question keywords and question vector representations; performing retrieval processing based on the question keywords and the question vector representations respectively to obtain question retrieval results, wherein the question retrieval results are used as training data for a large language model.

[0011] In one exemplary embodiment of this disclosure, the step of performing retrieval processing based on the question keywords and the question vector representation to obtain question retrieval results includes: performing graph retrieval processing based on the question keywords to obtain graph retrieval results; performing sparse retrieval processing based on the question keywords to obtain keyword text slices; performing vector retrieval processing based on the question vector representation to obtain relevant text vector representations; and performing deduplication and reordering processing on the graph retrieval results, the keyword text slices, and the relevant text vector representations to obtain the question retrieval results.

[0012] In one exemplary embodiment of this disclosure, the method further includes: querying a memory database based on a received user question to obtain a question query result corresponding to the user question, wherein the memory database is used to store recommended question-answer pairs; when the question query result indicates that a matching answer has been found, returning the matching answer; when the question query result indicates that no matching answer has been found, performing knowledge retrieval enhancement processing based on the user question to obtain a recommended answer corresponding to the user question, wherein the recommended answer is used to generate a new recommended question-answer pair with the user question after annotation processing.

[0013] In one exemplary embodiment of this disclosure, the step of performing knowledge retrieval enhancement processing based on the user question to obtain a recommended answer corresponding to the user question includes: sending the user question to the knowledge retrieval enhancement module by a large language model; and the knowledge retrieval enhancement module performing knowledge retrieval processing based on the user question to obtain the recommended answer.

[0014] In one exemplary embodiment of this disclosure, the step of obtaining the recommended answer by the knowledge retrieval enhancement module based on the user question includes: performing a retrieval operation on a knowledge base based on the user question to obtain multiple initial recommended answers, wherein the knowledge base is used to store one or more of text slices, vector representations, knowledge graphs, and recommended question-answer pairs, and the knowledge base includes a memory database; and performing summary processing and reordering processing on the multiple initial recommended answers to obtain the recommended answer.

[0015] In one exemplary embodiment of this disclosure, the method further includes: returning the recommended answer to the user interface layer, receiving user annotation operations through the user interface layer; generating the new recommended question-answer pair based on the user question and the recommended answer after the user annotation operation, and storing the new recommended question-answer pair in the memory database.

[0016] According to a second aspect of this disclosure, a knowledge question answering system is provided, comprising: a large language model module for determining the target answer to a target question, wherein the large language model module is obtained by adjusting the parameters of a pre-trained language model using recommended question-answer pairs in a specified domain; a knowledge retrieval enhancement module for performing knowledge retrieval processing on knowledge retrieval data according to the target question to obtain a question retrieval result, and returning the question retrieval result to the large language model module, wherein the knowledge retrieval data is generated based on a document to be processed in a specified domain, and the knowledge retrieval data includes one or more of text slices, text vector representations, and tabular knowledge graphs; and a knowledge base module for storing document data in a specified domain, wherein the document data includes one or more of the document to be processed, text slices, text slice vectors, recommended question-answer pairs, question-answer pair vectors, and tabular knowledge graphs.

[0017] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module includes a text slicing unit, a text vectorization unit, and a graph generation unit; the text slicing unit is used to perform text segmentation processing on the document to be processed to obtain the text slices; the text vectorization unit is used to perform vectorization processing on the text slices to obtain the text slice vectors; the graph generation unit is used to create a knowledge graph based on the document to be processed, the knowledge graph including a tabular knowledge graph.

[0018] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module includes a knowledge retrieval unit and an answer determination unit; the knowledge retrieval unit is used to perform knowledge retrieval enhancement processing based on the target question to obtain the question retrieval result, the knowledge retrieval enhancement processing including one or more of graph retrieval, vector retrieval, and sparse retrieval; the answer determination unit is used to perform deduplication and reordering processing on the question retrieval result to obtain the recommended question-answer pair.

[0019] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module further includes a metadata management unit, a question understanding unit, and a retrieval evaluation unit; the metadata management unit is used to manage the metadata of the document to be processed; the question understanding unit is used to determine the question intent of the target question and provide a matching target answer, wherein the question intent is used to determine the target answer matching the target question; and the retrieval evaluation unit is used to evaluate the recommendation effect of the recommended question-answer pair.

[0020] In one exemplary embodiment of this disclosure, the knowledge base module includes a file storage unit, an index storage unit, and a graph storage unit; the file storage unit is used to store the document to be processed; the index storage unit is used to store one or more of the text slices, the text slice vectors, the recommended question-answer pairs, and the question-answer pair vectors, wherein the text slices are obtained by performing text segmentation processing on the document to be processed; and the graph storage unit is used to store the tabular knowledge graph.

[0021] In one exemplary embodiment of this disclosure, the knowledge base module includes a data cleaning unit; the data cleaning unit is used to collect original documents and recommended question-and-answer pairs in a specified field, and to perform page number and table of contents deletion and parsing and recognition processing on the original documents after deduplication to obtain table recognition results and paragraph recognition results.

[0022] According to a third aspect of this disclosure, a knowledge question answering device is provided, comprising: a knowledge question answering module, configured to perform knowledge retrieval processing on a target question using a pre-built knowledge question answering system to obtain a target answer corresponding to the target question, wherein the knowledge question answering system is obtained by training on recommended question-answer pairs in a specified domain, wherein the recommended question-answer pairs in the specified domain are feedback results of documents to be processed in the specified domain after knowledge retrieval enhancement processing, and the documents to be processed in the specified domain include tabular data.

[0023] In one exemplary embodiment of this disclosure, the knowledge question answering device includes a text processing module, which is used to perform data cleaning and text slicing processing on the acquired original documents in a specified domain to obtain a document to be processed, the document to be processed including text slices; construct a tabular knowledge graph based on the document to be processed, perform graph textification processing on the tabular knowledge graph, and generate a text vector representation corresponding to the document to be processed by combining the text slices.

[0024] In one exemplary embodiment of this disclosure, the text processing module includes a table graph construction unit, configured to: extract table data from the document to be processed, the table data including table context information, the table context information including table content and content hierarchy relationship; normalize the table data to obtain normalized table data; and perform entity-relationship mapping processing on the normalized table data to obtain the table knowledge graph.

[0025] In one exemplary embodiment of this disclosure, the text processing module includes a table vectorization processing unit, configured to: retrieve entity and relation placeholders from the table knowledge graph according to a predefined text template; fill the entity and relation placeholders based on the retrieved entity types and context information to obtain natural language table data; and perform text vectorization processing on the natural language table data and the text slices to obtain the text vector representation.

[0026] In one exemplary embodiment of this disclosure, the knowledge question answering device further includes a question retrieval module, configured to: extract keywords and vectorize the acquired user questions to obtain question keywords and question vector representations; perform retrieval processing based on the question keywords and the question vector representations respectively to obtain question retrieval results, wherein the question retrieval results are used as training data for a large language model.

[0027] In one exemplary embodiment of this disclosure, the question retrieval module includes a question retrieval unit, configured to: perform graph retrieval processing based on the question keywords to obtain graph retrieval results; perform sparse retrieval processing based on the question keywords to obtain keyword text slices; perform vector retrieval processing based on the question vector representation to obtain relevant text vector representations; and perform deduplication and reordering processing on the graph retrieval results, the keyword text slices, and the relevant text vector representations to obtain the question retrieval results.

[0028] In one exemplary embodiment of this disclosure, the knowledge question-answering device further includes a recommended answer determination module, configured to: perform a query operation on a memory database based on a received user question to obtain a question query result corresponding to the user question, wherein the memory database is used to store recommended question-answer pairs; when the question query result indicates that a matching answer has been found, return the matching answer; when the question query result indicates that no matching answer has been found, perform knowledge retrieval enhancement processing based on the user question to obtain a recommended answer corresponding to the user question, wherein the recommended answer is used to generate a new recommended question-answer pair with the user question after annotation processing.

[0029] In one exemplary embodiment of this disclosure, the recommended answer determination module includes a recommended answer determination unit, configured to: send the user question to the knowledge retrieval enhancement module by the large language model; and have the knowledge retrieval enhancement module perform knowledge retrieval processing based on the user question to obtain the recommended answer.

[0030] In one exemplary embodiment of this disclosure, the recommended answer determination unit includes a recommended answer determination subunit, configured to: perform a retrieval operation on the knowledge base based on the user question to obtain multiple initial recommended answers, wherein the knowledge base is used to store one or more of text slices, vector representations, knowledge graphs, and recommended question-answer pairs, and the knowledge base includes a memory database; and perform summary processing and reordering processing on the multiple initial recommended answers to obtain the recommended answer.

[0031] In one exemplary embodiment of this disclosure, the knowledge question-answering device further includes a question-answer pair generation module, configured to: return the recommended answer to the user interface layer, receive user annotation operations through the user interface layer; generate the new recommended question-answer pair based on the user question and the recommended answer after the user annotation operation, and store the new recommended question-answer pair in the memory database.

[0032] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions that, when executed by the processor, implement the knowledge question-answering method according to any one of the preceding claims.

[0033] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the knowledge question-answering method described in any of the preceding embodiments.

[0034] The technical solution provided in this disclosure may include the following beneficial effects:

[0035] The knowledge question answering method in the exemplary embodiments of this disclosure, on the one hand, implements the construction of knowledge retrieval enhancement techniques for a specified domain, and can still effectively achieve knowledge querying for a specified domain even when faced with insufficient training corpus. On the other hand, the knowledge question answering system enables knowledge querying of tabular data, enhancing the model's ability to process non-standard natural language data.

[0036] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0038] Figure 1 A flowchart illustrating a knowledge question-answering method according to an exemplary embodiment of the present disclosure is shown schematically;

[0039] Figure 2 A flowchart illustrating the knowledge retrieval enhancement process of a large language model according to an exemplary embodiment of the present disclosure is shown.

[0040] Figure 3 The flowchart illustrating the process of determining recommended question-answer pairs based on a user feedback mechanism according to an exemplary embodiment of the present disclosure is shown in the illustration.

[0041] Figure 4 The diagram schematically illustrates the overall architecture of a knowledge question-answering system according to an exemplary embodiment of the present disclosure;

[0042] Figure 5 A block diagram of a knowledge question-answering device according to an exemplary embodiment of the present disclosure is shown schematically;

[0043] Figure 6 The illustration schematically shows a computer-readable storage medium according to an exemplary embodiment of the present disclosure;

[0044] Figure 7 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown schematically. Detailed Implementation

[0045] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0046] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details described, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known structures, methods, apparatuses, implementations, materials, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0047] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, or in one or more software-hardened modules, or in different network and / or processor devices and / or microcontroller devices.

[0048] The mainstream approach to building domain-specific large language models is to combine corpus pre-training with fine-tuning based on specific question-answering instructions. During pre-training, the model learns from a large amount of text data, mastering the basic structure and general knowledge of the language. Subsequently, in the fine-tuning phase, the model is optimized through specific question-answering system (QA) tasks, enabling it to more accurately understand and respond to domain-related questions, aiming to achieve optimal performance within that specific domain.

[0049] Furthermore, to further enhance the model's domain adaptability and knowledge depth, the industry has also explored strategies using knowledge retrieval-augmented generation (RAG) as an external knowledge base. However, the above approaches have the following drawbacks: both corpus pre-training and QA instruction fine-tuning require sufficiently rich and high-quality text or labeled question-answer pairs. If the corpus is not rich enough or not updated in a timely manner, the model may not be able to understand the latest knowledge.

[0050] The telecommunications industry has a large number of operation manuals and case studies in tabular form. These tables are rich in professional knowledge, but there is currently a lack of an effective knowledge retrieval and recall solution for this type of tabular data.

[0051] High-quality question-and-answer pairs tagged by users are of great value in alleviating the long-tail problem of large models and improving the overall performance of the models. At present, there is a lack of a complete mechanism and solution to systematically collect and optimize the response quality of the models based on high-quality question-and-answer feedback from users.

[0052] Based on this, this disclosure proposes a knowledge question answering method, a knowledge question answering system, a knowledge question answering device, an electronic device, and a computer program product.

[0053] In this article, it's important to understand that the terms used, such as Domain-Specific Large Language Models (DSLMs), refer to large-scale natural language processing models tailored to specific industries or knowledge domains. They are pre-trained within a specific domain (such as healthcare) to deeply understand and generate specialized text and information within that domain. Compared to general-purpose language models, DSLMs can handle specialized terminology and complex concepts more accurately, providing high-quality prediction, analysis, and text generation services, significantly improving performance and practicality in specific application scenarios.

[0054] A knowledge graph is a structured semantic knowledge base used to describe concepts and their relationships in the physical world in symbolic form. Its basic building blocks are the "entity-relationship-entity" triple, as well as entities and their associated attribute-value pairs. Entities are interconnected through relations, forming a network-like knowledge structure.

[0055] Knowledge retrieval-augmented generation significantly improves the output quality of language models by combining information retrieval with text generation. This method first retrieves information relevant to the input query from a large dataset or knowledge base, and then incorporates this information as additional context into the text generation process.

[0056] In this example embodiment, a knowledge question answering method is first provided. The knowledge question answering method of this disclosure can be implemented using a server or using a terminal device. The terminal described in this disclosure may include mobile terminals such as mobile phones, tablets, laptops, handheld computers, and personal digital assistants (PDAs), as well as fixed terminals such as desktop computers. Figure 1The illustration shows a schematic diagram of a knowledge question-answering method flow according to some embodiments of the present disclosure. Reference Figure 1 This knowledge-based question-and-answer method may include the following steps:

[0057] Step S110: Use a pre-built knowledge question answering system to perform knowledge retrieval processing on the target question to obtain the target answer corresponding to the target question. The knowledge question answering system is obtained by training on recommended question-answer pairs in a specified domain. The recommended question-answer pairs in the specified domain are the feedback results of the documents to be processed in the specified domain after knowledge retrieval enhancement processing. The documents to be processed in the specified domain include tabular data.

[0058] According to some exemplary embodiments of this disclosure, a knowledge question answering system can be a system that performs knowledge retrieval processing on current input data to complete processing tasks such as prediction, classification, and question answering of the input data. The target question can be a currently acquired user question. Knowledge retrieval processing can be a retrieval operation performed in a knowledge base based on different data sources during the knowledge reasoning stage. The target answer can be the question answer obtained after knowledge retrieval processing based on the target question. Recommended question-answer pairs for a specified domain can be question-answer pairs obtained after knowledge retrieval enhancement processing of the document to be processed in a specified domain. Feedback results can be the retrieval results returned by the knowledge retrieval enhancement module in the knowledge question answering system after performing knowledge retrieval enhancement processing on the document to be processed in a specified domain. Tabular data can be document content in non-standard natural language stored in tabular form.

[0059] In a knowledge-based question-answering scenario, the system retrieves the target question that needs to be answered. This target question can be a user question entered through a graphical user interface, or it can be a question-answering task received by the system, which may include one or more target questions. After retrieving the target question, it is input into the pre-built knowledge-based question-answering system. The system then performs knowledge retrieval processing based on the target question and outputs the target answer that matches it.

[0060] The knowledge question answering system in this embodiment includes a Large Language Model (LLM), a Knowledge Retrieval Enhancement (RAG) component, and a knowledge base. The system uses documents from a specific domain as training data. Since these documents can take various forms, they may include natural language text as well as non-standard natural language data such as tabular data. The documents are stored in the knowledge base using a corresponding storage format; for example, tabular data can be stored in the form of a knowledge graph.

[0061] During the training process of the knowledge question-answering system, the knowledge retrieval enhancement part, through the reasoning part, performs knowledge retrieval operations based on user questions contained in the documents to be processed within a specified domain, obtaining feedback results (i.e., recommended answers) corresponding to the user questions. Then, based on the user questions and recommended answers, it generates recommended question-answer pairs for the specified domain. The documents to be processed within the specified domain can be various documents obtained from that domain; for example, they may include user manuals, technical manuals, operation cases, and other types of documents. The recommended question-answer pairs within the specified domain are used as training data for a large language model, allowing for parameter tuning to improve the model's performance in understanding questions and generating correct answers.

[0062] Because the knowledge retrieval enhancement section stores tabular data in the form of a knowledge graph, relevant information can be quickly retrieved during the reasoning stage. The reasoning section can retrieve knowledge from the knowledge base in real time when a user asks a question, thus providing the large language model with the latest knowledge support to generate more accurate answers. Furthermore, the model has the ability to process non-standard natural language data such as tabular data.

[0063] According to the knowledge question answering method in this example embodiment, on the one hand, the knowledge question answering system implements the construction of knowledge retrieval enhancement technology for a specified domain, and can still effectively achieve knowledge querying for a specified domain even when faced with insufficient training corpus. On the other hand, the knowledge question answering system enables knowledge querying of tabular data, enhancing the model's ability to process non-standard natural language data.

[0064] The knowledge question-and-answer method in this example embodiment will be further explained below.

[0065] In one exemplary embodiment of this disclosure, the original documents obtained in a specified domain are subjected to data cleaning and text slicing to obtain documents to be processed, which include text slices; a tabular knowledge graph is constructed based on the documents to be processed, the tabular knowledge graph is subjected to graph textification, and text vector representations corresponding to the documents to be processed are generated by combining the text slices.

[0066] The original documents in the specified domain can be various documents directly obtained from that domain. Data cleaning is the process of removing noise and irrelevant information from the original documents in the specified domain, ensuring that the processed data is accurate and usable. Text slicing is the process of segmenting text using natural language processing, identifying and splitting complex Chinese sentence structures, ensuring that each paragraph, sentence, and phrase is accurately identified. A text slice is a text fragment obtained after the document to be processed has undergone text slicing, such as a paragraph or phrase. A tabular knowledge graph is a knowledge graph constructed by extracting content from tabular data. Graph textification is the process of filling in entity and relation placeholders in the tabular knowledge graph according to the queried entity types and contextual content to generate natural language that can be understood by large language models. Text vector representation is the process of vectorizing text data to obtain a vector representation.

[0067] All industries have a large amount of production data. The data formats of production data vary in different technical fields. For example, in the telecommunications industry, various operation cases, operation manuals and other documents include not only text data stored in standard formats, but also non-standard language data stored in tabular form. This tabular data is rich in knowledge value and stores a lot of useful information.

[0068] To address the issue that the presentation of table content, often not in standard natural language, negatively impacts the comprehension capabilities of large language models, this implementation proposes a method using knowledge graph technology to effectively transform and store table data from documents within a specified domain. Furthermore, during knowledge retrieval, the data in the knowledge graph can be queried, and the query results are ultimately fed back to the large language model. For details, please refer to [link to relevant documentation]. Figure 2 As shown, Figure 2 A flowchart illustrating the knowledge retrieval enhancement process of a large language model according to an exemplary embodiment of the present disclosure is shown.

[0069] from Figure 2 As can be seen, the knowledge retrieval enhancement module mainly consists of an offline component and a reasoning component. The offline component is primarily responsible for preprocessing, slicing, vectorizing, generating a knowledge graph, and storing the original documents so that relevant information can be quickly retrieved in the reasoning component. The reasoning component is mainly responsible for retrieving knowledge from the knowledge base in real time when a user asks a question, thereby providing the latest knowledge support for the large model to generate more accurate answers. Specifically, the offline processing flow is as follows:

[0070] Step 1: Perform data cleaning on the original documents. Clean various types of documents, such as operation cases and manuals, within a specified field (e.g., telecommunications). For example, use the data cleaning module in a knowledge base to remove noise and irrelevant information from the original documents in the specified field, ensuring that the remaining data after cleaning is accurate and usable, resulting in the document to be processed.

[0071] Step Two: File Storage. The cleaned documents in various formats are ultimately saved as text (txt) files, and then stored on a File Transfer Protocol (FTP) server for later use.

[0072] Step 3: Text Slicing. Read the stored text from the FTP server and use a word segmenter (such as the ChineseRecursiveTextSplitter algorithm) to slice the cleaned document, obtaining the text slices corresponding to the document to be processed. The ChineseRecursiveTextSplitter is a segmentation technology specifically designed for Chinese text, capable of effectively identifying and splitting complex Chinese sentence structures. This algorithm can recursively segment text, ensuring that each paragraph, sentence, and phrase is accurately identified.

[0073] Step 4: Construct a tabular knowledge graph. Telecommunications documents contain a large amount of tabular data, which aggregates key information and metrics, such as network configuration parameters, service quality data, and user statistics. Extracting the tabular content from these documents and constructing a knowledge graph yields the tabular knowledge graph.

[0074] After constructing the tabular knowledge graph, it undergoes graph textification. Since the documents to be processed may also contain standard natural language data stored in text format, text slices can be combined to generate text vector representations corresponding to the documents. By converting tabular data into knowledge graph format, the ability of large models to process non-standard natural language data can be enhanced.

[0075] In one exemplary embodiment of this disclosure, constructing a table knowledge graph based on a document to be processed includes: performing table extraction processing on the document to be processed to obtain table data, the table data including table context information, the table context information including table content and content hierarchy relationship; performing normalization processing on the table data to obtain normalized table data; and performing entity-relationship mapping processing on the normalized table data to obtain a table knowledge graph.

[0076] The table context information can be relevant information about the table data in the document to be processed, including the document title, multi-level directory titles, and table titles. The table content can be the text content contained within the table data. The content hierarchy can be the structural relationship between the text content in the table data. The normalized table data can be the table data after normalization processing. The entity-relationship mapping process can be the process of mapping the normalized table data to a knowledge graph through entity recognition and relationship extraction schemes.

[0077] Extracting table content from documents and constructing a knowledge graph involves a multi-stage processing flow. Specifically: (1) Table recognition and extraction. During the document cleaning stage, advanced text parsing tools are used to accurately identify and extract tables from the document. At the same time, title recognition technology is used to capture the contextual information of the tables, including the document title, multi-level directory titles, and table titles. This information is stored together in JavaScript ObjectNotation (JSON) format to ensure that the correlation between the table content and its hierarchical relationship is preserved.

[0078] (2) Data Normalization. The extracted tabular data needs to be normalized. The specific process of normalization is to convert data of different formats into a unified format, such as unifying different units of measurement. At this stage, the tabular data can be cleaned, ambiguities and errors removed, ensuring high data quality and consistency.

[0079] (3) Entity-Relationship Mapping. Normalized data is mapped to a knowledge graph using entity recognition and relationship extraction techniques. This step involves constructing an ontology structure using table titles and identifying the main columns of each table as central entity types. Then, based on the table titles and content, relationships between the central entities and other fields are established, forming the table graph structure. After obtaining the table knowledge graph, step five, knowledge graph storage, is executed. The constructed table knowledge graph is stored in a non-relational database (Not Only SQL NoSQL) graph database (Neo4j). Through knowledge graph technology, table data can be effectively transformed and stored in the knowledge graph, enabling subsequent knowledge retrieval processes to use the knowledge graph data as a data source.

[0080] In one exemplary embodiment of this disclosure, a table knowledge graph is subjected to graph textification processing, and text vector representations of the document to be processed are generated by combining text slices. The process includes: retrieving entity and relation placeholders from the table knowledge graph according to a predefined text template; filling the entity and relation placeholders based on the retrieved entity types and context information to obtain natural language table data; and performing text vectorization processing on the natural language table data and text slices to obtain text vector representations.

[0081] The text template can be a pre-defined content template based on the text content. Entity and relation placeholders can be placeholders used to determine the specific location of entities and relations in the text. Entity types can be the specific types of entities contained in the knowledge graph. Contextual information can be the contextual content related to the entities and relations. Natural language tabular data can be tabular data obtained by populating the entities and relations in the knowledge graph.

[0082] After constructing the tabular knowledge graph, step six, graph textification, is performed. Using predefined text templates, placeholders for entities and relationships retrieved from the tabular knowledge graph stored in Neo4j are filled in according to the entity types and context, generating natural language that can be understood by large models, thus obtaining natural language tabular data.

[0083] Steps seven and eight involve text vectorization. Vector transformation algorithms (such as Dmeta-embedding) are used to convert text data into numerical vectors. This includes vectorization of slices constructed from documents such as case manuals, vectorization of question-answer pairs, and vectorization of natural language generated from knowledge graphs. These vectors are used to calculate the similarity between texts during subsequent RAG retrieval, retrieving relevant information from massive amounts of text.

[0084] Step nine, vectorization and storage. The vectorized slice vector representations are stored in a distributed search (Elasticsearch) database for retrieval during subsequent reasoning. Vectorizing the data in the tabular knowledge graph facilitates content retrieval during the reasoning process.

[0085] In one exemplary embodiment of this disclosure, the obtained user questions are subjected to keyword extraction and vectorization to obtain question keywords and question vector representations; retrieval processing is performed based on the question keywords and question vector representations respectively to obtain question retrieval results, which are used as training data for a large language model.

[0086] Keyword extraction can be the process of extracting keywords or phrases from user questions. Vectorization can be the process of vectorizing the user question into text. Question keywords can be used in subsequent retrieval processes. Question vector representation can be the vector representation obtained after vectorizing the user question. Question retrieval results can be the feedback results obtained after knowledge retrieval enhancement processing based on the user question.

[0087] Continue to refer to Figure 2 The reasoning process includes the following steps: Step 1: Keyword extraction. Based on the user's question, a minimal keyword extraction algorithm (KeyBert algorithm) based on Bidirectional Encoder Representations from Transformers (BERT) is used to extract keywords or phrases from the user's question. These keywords are used to guide the subsequent retrieval process.

[0088] The KeyBERT algorithm works by finding the word in a document that is most similar to the document itself. The specific steps include: first, extracting document embeddings using BERT to obtain a document-level representation; then, extracting word embeddings from the Chinese language model (N-gram); and finally, using cosine similarity to find the word / phrase most similar to the document. Through this process, the most similar word can be identified as the word that best describes the entire document.

[0089] Step Two: Query Vectorization. Using the same Dmeta-embedding algorithm as in the offline section, user questions are vectorized. After obtaining the question keywords and question vector representations, the knowledge base can be retrieved based on these keywords and vector representations to obtain the retrieval results. These results include recommended answers matching the user question. The recommended question-answer pairs generated based on the user question and recommended answers can be used as training data for the large language model. Retrieving based on user questions allows the feedback results to be used as training data for the large language model. This provides an effective way for the knowledge retrieval enhancement part to bridge the information between the internal knowledge of the large language model and the external world. This not only reduces the burden on the large model from trying to remember all possible information but also significantly improves the performance when processing specific, complex, or up-to-date queries.

[0090] In one exemplary embodiment of this disclosure, retrieval processing is performed based on question keywords and question vector representations to obtain question retrieval results, including: performing graph retrieval processing based on question keywords to obtain graph retrieval results; performing sparse retrieval processing based on question keywords to obtain keyword text slices; performing vector retrieval processing based on question vector representations to obtain relevant text vector representations; and performing deduplication and reordering processing on the graph retrieval results, keyword text slices, and relevant text vector representations to obtain question retrieval results.

[0091] Specifically, the knowledge graph retrieval results can be obtained by searching the knowledge graph based on entities identified in the tabular data. Keyword text slicing can be obtained by sparse retrieval based on question keywords. Relevant text vector representations can be obtained by vector retrieval based on question vector representations.

[0092] After obtaining the question keywords and question vector representations, perform the following retrieval operations. Step 3: Knowledge Graph Retrieval. Based on the extracted keywords, perform entity recognition and entity linking, that is, match the identified entities with the corresponding entities in the knowledge graph, and then retrieve related subgraphs based on the queried entities. Based on the subgraphs, generate answers in natural language that the large model can understand, as the knowledge graph retrieval results.

[0093] Step 4: Sparse Search. Based on the extracted keywords, search the Elasticsearch database for document slices containing these keywords, and use them as keyword text slices.

[0094] Step 5: Vector Retrieval. Based on the query's vector representation, the similarity between the query vector representation and the text slice vector representation is calculated using the cosine similarity algorithm in the Elasticsearch database. The most relevant text fragments or documents are then retrieved as the relevant text vector representations.

[0095] Step Six: Slicing and Deduplication. The search results from graph retrieval, sparse retrieval, and vector retrieval are summarized and deduplicated. Results from different retrieval methods are merged together, and duplicates are removed to ensure the uniqueness of the final result.

[0096] Step 7: Reordering. Based on the reordering algorithm (bge-reranker-large algorithm), the merged results are reordered to determine which results are most relevant to the user's question.

[0097] Step 8: Select the top n text fragments from the re-ranked results and provide them as input to the large language model. The large language model further generates or refines the precise answer to the query. Through the above knowledge retrieval enhancement process, the obtained question retrieval results are used as training data for the large language model, increasing the model's adaptability and flexibility, enabling it to better adapt to changing information needs and constantly evolving knowledge domains.

[0098] Those skilled in the art will readily understand that the three retrieval methods—graph retrieval, sparse retrieval, and vector retrieval—can be executed in parallel, and this embodiment does not impose any special limitation on the specific execution order of the three retrieval methods. During the knowledge retrieval process, any one retrieval method or multiple parallel retrieval methods can be used for knowledge retrieval enhancement processing.

[0099] In one exemplary embodiment of this disclosure, a query operation is performed on the memory database based on the received user question to obtain the question query result corresponding to the user question. The memory database is used to store recommended question-answer pairs. When the question query result shows that a matching answer has been found, the matching answer is returned. When the question query result shows that no matching answer has been found, knowledge retrieval enhancement processing is performed based on the user question to obtain the recommended answer corresponding to the user question. The recommended answer is used to generate a new recommended question-answer pair with the user question after annotation processing.

[0100] The memory database can be a database used to store recommended question-answer pairs. Question query results can be retrieved from the memory database based on the user's question. A recommended question-answer pair can be a pair consisting of a user's question and its corresponding matching answer. The matching answer can be the question-answer retrieved from the memory database that corresponds to the user's question. The recommended answer can be the question-answer obtained after knowledge retrieval enhancement processing that corresponds to the user's question. Adding a new recommended question-answer pair can be done by generating a question-answer pair based on the user's question and the recommended answer.

[0101] To address the long-tail and illusion problems that may exist in large language models, this disclosure proposes a user feedback mechanism. The overall process of the user feedback mechanism is as follows: Figure 3 As shown, Figure 3 A flowchart illustrating a user feedback-based process for determining recommended question-answer pairs according to an exemplary embodiment of this disclosure is shown. Figure 3 As can be seen, the user feedback mechanism mainly includes five modules: user interface layer, task distribution layer, knowledge base, RAG, and large language model. Details are as follows:

[0102] The user interface layer is the interface through which users interact with the system. Users submit questions and provide feedback on the quality of the results generated by the large model. The task distribution layer receives query commands from the large model, retrieves information from the knowledge base, invokes the large model's generation capabilities, and stores high-quality question-and-answer pairs provided by the user. The knowledge base stores all original document texts, document slices, vector representations, and knowledge graphs, and records user feedback and user-inputted recommended question-and-answer pairs in the memory database. RAG (Reference Aggregate Graph) is used for knowledge retrieval-enhanced generation. When the large language model is invoked at the task distribution layer, it uses the RAG function to retrieve relevant knowledge from the knowledge base to generate answers to user questions. The large language model, when invoked at the task distribution layer, generates relevant answers for the user based on the user's question and the similar document slice results returned by the RAG.

[0103] The user feedback mechanism's processing steps include: Step 1: The question-answer pairs stored in the memory database mainly come from two sources: one is user feedback generated by the large language model during use; the other is high-quality questions and answers recorded by the user themselves. If there are initial high-quality question-answer pairs, they are entered into the memory database through this step. Step 2: Users submit questions to the system through the user interface. Step 3: After receiving a user's question, the task distribution layer first retrieves similar questions and answers that have been verified through user feedback in the memory database.

[0104] Step four: The knowledge base returns the search results. If a similar question and answer are found in the memory database, the results are returned; otherwise, nothing is returned. Step five: The task distribution layer receives the results returned by the knowledge base. If a matching question-answer pair exists, the corresponding answer is directly returned to the user. Step six: If no matching item is found in the knowledge base, the task distribution layer activates the large language model result generation module and sends the query to the large language model. This triggers the LLM+RAG module to interact. Simultaneously, high-quality question-answer pairs, along with the knowledge graph and text slices, are used as search objects for RAG, retrieved, reordered, and the results are fed back to the large model. This method not only improves the performance and efficiency of the large model's answer generation but also addresses the long-tail problem and the illusion problem to some extent.

[0105] In one exemplary embodiment of this disclosure, knowledge retrieval enhancement processing is performed based on the user question to obtain a recommended answer corresponding to the user question, including: the user question is sent to the knowledge retrieval enhancement module by the large language model; and the knowledge retrieval enhancement module performs knowledge retrieval processing based on the user question to obtain a recommended answer.

[0106] Among them, the knowledge retrieval enhancement module is a module structure that can be used to perform knowledge retrieval enhancement processing.

[0107] The specific steps for the interaction between the large language model and the RAG module to determine recommended question-answer pairs include: Step 7, referring to the existing knowledge stored in the knowledge base, the large language model calls the RAG interface and passes the query to RAG. Step 8, RAG uses various retrieval algorithms to search the knowledge base for document slices, vector representations, knowledge graphs, and question-answer pairs, obtaining multiple retrieval results. The most matching question answer is determined from the retrieval results as the recommended answer, and the result is fed back to the large language model, which can improve the model's accuracy and efficiency in information retrieval.

[0108] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module performs knowledge retrieval processing based on the user's question to obtain recommended answers, including: performing a retrieval operation on the knowledge base based on the user's question to obtain multiple initial recommended answers, wherein the knowledge base is used to store one or more of text slices, vector representations, knowledge graphs, and recommended question-answer pairs, and the knowledge base includes a memory database; and performing summary processing and reordering processing on the multiple initial recommended answers to obtain recommended answers.

[0109] The initial recommended answer can be a recommended answer corresponding to the user's question, determined through various retrieval methods. The re-ranking process can be a procedure that sorts multiple initial recommended answers based on their relevance to the user's question.

[0110] The specific process of knowledge retrieval processing by the knowledge retrieval enhancement module includes: Step nine, RAG receives the multi-path retrieval results returned by the knowledge base, i.e., multiple initial recommended answers. The knowledge base can include data from various data sources such as text slices, vector representations, knowledge graphs, and recommended question-and-answer pairs. The multiple initial recommended answers returned after the retrieval operation are summarized and re-ranked to obtain the final document fragment ranking.

[0111] Step 10: RAG returns the top N most similar document fragments to the large language model. Step 11: The large language model constructs answer prompts based on the top N most similar document fragments returned by RAG and the user query, generates answers, and returns them to the task distribution layer. The number of generated recommended answers can be one or more. Combining the large language model with the knowledge retrieval enhancement module to feed the retrieval results back to the large language model can improve its answer generation performance and efficiency.

[0112] In one exemplary embodiment of this disclosure, the recommended answer is returned to the user interface layer, and the user annotation operation is received through the user interface layer; based on the user question and the recommended answer after the user annotation operation, a new recommended question-answer pair is generated, and the new recommended question-answer pair is stored in the memory database.

[0113] The user interface layer can be used to receive user input questions and return corresponding answers to the user. User annotation operations can be processing operations that annotate recommended answers.

[0114] After obtaining recommended answers through knowledge retrieval enhancement, in step 12, the task distribution layer feeds back the final answer to the user interface layer for presentation. In step 13, the user annotates the final result generated by the large language model through the interface to determine whether a recommended answer is the ideal answer, and this is fed back to the task distribution layer. In step 14, if the user determines that the current answer is the ideal answer, the recommended answer that has been determined to be the ideal answer by the user is obtained. The user's question and the recommended answer after user annotation are used together to generate a new recommended question-answer pair, and the new recommended question-answer pair is stored in the memory database. Introducing a user feedback mechanism in the knowledge retrieval enhancement process, and using the user-annotated question-answer pairs as training data for subsequent large language models, can, to some extent, solve the long-tail problem and illusion problem of the model.

[0115] It will be readily understood by those skilled in the art that the designated field in this disclosure can be any technical field that includes tabular production data, and this disclosure does not impose any special limitation on the specific field type of the designated field.

[0116] In summary, the knowledge-based question-answering method disclosed herein employs a pre-constructed knowledge-based question-answering system to perform knowledge retrieval processing on the target question, obtaining the target answer corresponding to the target question. The knowledge-based question-answering system is trained on recommended question-answer pairs within a specified domain. These recommended question-answer pairs are feedback results from documents in the specified domain that have undergone knowledge retrieval enhancement processing, including tabular data. On one hand, the knowledge-based question-answering system implements the construction of knowledge retrieval enhancement techniques for a specified domain, effectively enabling knowledge queries within that domain even when faced with insufficient training corpus. On the other hand, the system enables knowledge queries on tabular data, enhancing the model's ability to process non-standard natural language data. Furthermore, by introducing a user feedback mechanism, recommended question-answer pairs that have been user-annotated and approved are stored in a memory database. These recommended question-answer pairs can be used to train large language models, thus addressing the long-tail problem and illusion problem of the model to some extent.

[0117] It should be noted that although the steps of the method in this invention are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0118] Furthermore, a knowledge-based question-answering system is also provided in this example embodiment. (See reference) Figure 4 The knowledge question answering system 400 may include: a large language model module 410, a knowledge retrieval enhancement module 420, and a knowledge base module 430.

[0119] Specifically, the large language model module 410 is used to determine the target answer to the target question. The large language model module uses recommended question-answer pairs from a specified domain to adjust the parameters of a pre-trained language model. The knowledge retrieval enhancement module 420 is used to process the knowledge retrieval data according to the target question, obtain the question retrieval result, and return the question retrieval result to the large language model module. The knowledge retrieval data is generated based on the documents to be processed in a specified domain. The knowledge retrieval data includes one or more of the following: text slices, text vector representations, and tabular knowledge graphs. The knowledge base module 430 is used to store document data in a specified domain. The document data includes one or more of the following: documents to be processed, text slices, text slice vectors, recommended question-answer pairs, question-answer pair vectors, and tabular knowledge graphs.

[0120] The large language model module 410 can include three task processing stages: corpus pre-training, question-answering inference fine-tuning, and large model inference. Specifically, corpus pre-training involves training the language model using large-scale text data so that the model can understand the grammatical structure and semantic information of natural language. QA inference fine-tuning involves specifically adjusting the model parameters for the question-answering system to improve the model's performance in understanding questions and generating correct answers. Large model inference involves using the trained model to process new input data, performing tasks such as prediction, classification, and question answering.

[0121] This disclosed knowledge question answering system implements a large model scheme based on knowledge graph enhancement, which can enhance the model's ability to understand non-standard natural language data and improve the accuracy and efficiency of information retrieval.

[0122] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module 420 includes a text slicing unit, a text vectorization unit, and a graph generation unit; the text slicing unit is used to perform text segmentation processing on the document to be processed to obtain text slices; the text vectorization unit is used to perform vectorization processing on the text slices to obtain text slice vectors; the graph generation unit is used to create a knowledge graph based on the document to be processed, the knowledge graph including a tabular knowledge graph.

[0123] The text slicing unit in the Knowledge Retrieval Enhancement Module (RAG) is primarily used for text slicing, dividing the original document into smaller, more manageable text fragments. The text vectorization unit converts these text slices into high-dimensional vectors representing the meaning of the text using vector transformation algorithms such as Dmeta-embedding. This facilitates text comparison and ranking during subsequent retrieval.

[0124] The knowledge graph generation unit creates a knowledge graph based on the document to be processed. Since the document can include both structured and unstructured data sources, the knowledge graph generation unit creates a knowledge graph based on these data sources, converting the data into a collection of interconnected entities and relationships. In this embodiment, it primarily transforms table content in the text that is difficult to parse using large models into a knowledge graph. The knowledge retrieval enhancement module provides various data generation units to store data from different data sources for subsequent knowledge retrieval.

[0125] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module 420 includes a knowledge retrieval unit and an answer determination unit; the knowledge retrieval unit is used to perform knowledge retrieval enhancement processing based on the target question to obtain the question retrieval result, and the knowledge retrieval enhancement processing includes one or more of graph retrieval, vector retrieval, and sparse retrieval; the answer determination unit is used to perform deduplication and reordering processing on the question retrieval result to obtain recommended question-answer pairs.

[0126] The knowledge retrieval unit is primarily used for knowledge retrieval enhancement based on the target question, mainly including three parallel retrieval methods: graph retrieval, vector retrieval, and sparse retrieval. During the inference phase, these three retrieval methods are invoked based on different data sources. Graph retrieval retrieves the most relevant paths or subgraphs in the knowledge graph based on the query information, thereby generating context. Vector retrieval uses vectors generated in the previous text vectorization steps to retrieve information, searching the vector library for the document slice vector that is closest to the user's question vector. Compared to vector retrieval, sparse retrieval relies more heavily on traditional keyword matching techniques, such as Boolean search or the bag-of-words model.

[0127] Before performing knowledge retrieval based on user questions, keyword extraction can be performed on the user questions. Keyword extraction refers to extracting the most important words from the user query and then performing sparse retrieval and graph retrieval based on the extracted keywords.

[0128] After obtaining the search results through the above retrieval operations, a segmentation and deduplication process is performed. Three retrieval methods are used to obtain more relevant segment content. To avoid providing users with duplicate content during the retrieval process, identical or similar text fragments are identified, merged, or removed. Following the segmentation and deduplication process, the search results are reordered. After segmentation and deduplication, the currently obtained text segments are reordered based on their relevance, accuracy, or other metrics to improve the final output quality. The knowledge retrieval enhancement module performs inference operations, increasing the model's adaptability and flexibility, enabling it to better suit changing information needs and evolving knowledge domains.

[0129] In one exemplary embodiment of this disclosure, the knowledge retrieval enhancement module 420 further includes a metadata management unit, a question understanding unit, and a retrieval evaluation unit; the metadata management unit is used to manage the metadata of the document to be processed; the question understanding unit is used to determine the question intent of the target question and provide a matching target answer, wherein the question intent is used to determine the target answer that matches the target question; and the retrieval evaluation unit is used to evaluate the recommendation effect of the recommended question-answer pair.

[0130] The knowledge retrieval enhancement module also provides metadata management functionality. The metadata management unit is responsible for managing metadata related to various data and text fragments, such as source, creation date, and author information, which is crucial for maintaining the system's transparency and traceability. The question understanding unit, upon receiving a user's question, accurately understands the user's intent and provides relevant answers or results, including question generalization, sub-question generation, multi-step question generation, and answer merging. The retrieval evaluation unit is responsible for evaluating the effectiveness of the retrieval enhancement, including answer relevance, contextual relevance, and answer fidelity. By managing metadata and evaluating answer retrieval, the accuracy of the model can be effectively improved.

[0131] In one exemplary embodiment of this disclosure, the knowledge base module 430 includes a file storage unit, an index storage unit, and a graph storage unit; the file storage unit is used to store the document to be processed; the index storage unit is used to store one or more of text slices, text slice vectors, recommended question-answer pairs, and question-answer pair vectors, wherein the text slices are obtained by performing text segmentation processing on the document to be processed; and the graph storage unit is used to store a tabular knowledge graph.

[0132] The knowledge base module provides three data storage methods, implemented through file storage, index storage, and graph storage units: File storage for storing raw document data, such as text files, PDFs, and Word documents; Index storage (Elasticsearch storage, ES storage) for storing, retrieving, and analyzing text data using Elasticsearch as a search engine, including fine-grained text segmentation and its vectorized representation for large-scale text searches; and Graph storage for storing structured information about entities and relationships, using the Neo4j database for complex queries and data associations. By providing multiple storage methods, the knowledge base module can store various types of data.

[0133] In one exemplary embodiment of this disclosure, the knowledge base module 430 includes a data cleaning unit; the data cleaning unit is used to collect original documents and recommended question-and-answer pairs in a specified field, perform page number and directory deletion processing and parsing and recognition processing on the original documents after deduplication, and obtain table recognition results and paragraph recognition results.

[0134] The knowledge base module extracts and identifies content from original documents in a specified domain through a data cleaning unit. Specifically, this includes: document acquisition and editing, which collects and edits original documents to ensure the accuracy and usability of information; and question-answer pair acquisition and editing, which collects and edits question-answer pair data, including targeted acquisition and individual correction of single question-answer instances, as well as batch acquisition and collective editing of large numbers of question-answer pairs. This process aims to provide accurate question-answer pair data for large models.

[0135] Document deduplication is used to detect and remove duplicate documents, maintaining the uniqueness of the knowledge base content. Knowledge auditing is used by professionals to review the accuracy and applicability of information, ensuring the high quality of the knowledge base content. Sensitive word desensitization is used to identify and process sensitive information from documents, ensuring that inappropriate content is not leaked. Layout analysis is used to analyze the document's layout, thereby understanding the position and relationships of various text areas, images, tables, headings, paragraphs, and other layout elements.

[0136] Table recognition identifies and structures the table content in the document for later input into the knowledge graph for retrieval. Paragraph recognition identifies paragraphs at various levels in the document for further content understanding and information extraction. Page number and table of contents removal removes page numbers and tables of contents from the document to avoid interference during data analysis. Through these data cleaning and recognition operations, the content knowledge contained in the document can be extracted, facilitating subsequent knowledge retrieval operations.

[0137] Furthermore, in this example embodiment, a knowledge question-answering device is also provided. (See reference...) Figure 5 The knowledge question answering device 500 may include: a knowledge question answering module 510.

[0138] Specifically, the knowledge question answering module 510 is used to perform knowledge retrieval processing on the target question using a pre-built knowledge question answering system to obtain the target answer corresponding to the target question. The knowledge question answering system is obtained by training on recommended question-answer pairs in a specified domain. The recommended question-answer pairs in the specified domain are the feedback results of the documents to be processed in the specified domain after knowledge retrieval enhancement processing. The documents to be processed in the specified domain include tabular data.

[0139] In one exemplary embodiment of this disclosure, the knowledge question answering device 500 includes a text processing module, which is used to perform data cleaning and text slicing processing on the acquired original documents in a specified domain to obtain a document to be processed, the document to be processed including text slices; construct a tabular knowledge graph based on the document to be processed, perform graph textification processing on the tabular knowledge graph, and combine the text slices to generate a text vector representation corresponding to the document to be processed.

[0140] In one exemplary embodiment of this disclosure, the text processing module includes a table graph construction unit, configured to: extract table data from the document to be processed, the table data including table context information, the table context information including table content and content hierarchy relationship; normalize the table data to obtain normalized table data; and perform entity-relationship mapping processing on the normalized table data to obtain a table knowledge graph.

[0141] In one exemplary embodiment of this disclosure, the text processing module includes a table vectorization processing unit, configured to: retrieve entity and relation placeholders from a table knowledge graph based on a predefined text template; fill in the entity and relation placeholders based on the retrieved entity types and context information to obtain natural language table data; and perform text vectorization processing on the natural language table data and text slices to obtain text vector representations.

[0142] In one exemplary embodiment of this disclosure, the knowledge question answering device 500 further includes a question retrieval module, which is used to: extract keywords and vectorize the acquired user questions to obtain question keywords and question vector representations; perform retrieval processing based on the question keywords and question vector representations respectively to obtain question retrieval results, and use the question retrieval results as training data for a large language model.

[0143] In one exemplary embodiment of this disclosure, the question retrieval module includes a question retrieval unit, configured to: perform graph retrieval processing based on question keywords to obtain graph retrieval results; perform sparse retrieval processing based on question keywords to obtain keyword text slices; perform vector retrieval processing based on question vector representations to obtain relevant text vector representations; and perform deduplication and reordering processing on the graph retrieval results, keyword text slices, and relevant text vector representations to obtain question retrieval results.

[0144] In one exemplary embodiment of this disclosure, the knowledge question-answering device 500 further includes a recommended answer determination module, configured to: perform a query operation on a memory database based on a received user question to obtain a question query result corresponding to the user question, wherein the memory database is used to store recommended question-answer pairs; when the question query result finds a matching answer, return the matching answer; when the question query result does not find a matching answer, perform knowledge retrieval enhancement processing based on the user question to obtain a recommended answer corresponding to the user question, wherein the recommended answer is used to generate a new recommended question-answer pair with the user question after annotation processing.

[0145] In one exemplary embodiment of this disclosure, the recommended answer determination module includes a recommended answer determination unit, configured to: send the user question to the knowledge retrieval enhancement module by the large language model; and have the knowledge retrieval enhancement module perform knowledge retrieval processing based on the user question to obtain a recommended answer.

[0146] In one exemplary embodiment of this disclosure, the recommended answer determination unit includes a recommended answer determination subunit, which is used to: perform a retrieval operation on the knowledge base based on the user's question to obtain multiple initial recommended answers, wherein the knowledge base is used to store one or more of text slices, vector representations, knowledge graphs, and recommended question-answer pairs, and the knowledge base includes a memory database; and perform summary processing and reordering processing on the multiple initial recommended answers to obtain recommended answers.

[0147] In one exemplary embodiment of this disclosure, the knowledge question answering device 500 further includes a question-answer pair generation module, configured to: return recommended answers to the user interface layer, receive user annotation operations through the user interface layer; generate new recommended question-answer pairs based on user questions and recommended answers after user annotation operations, and store the new recommended question-answer pairs in a memory database.

[0148] The specific details of the virtual modules of each knowledge question-answering device mentioned above have been described in detail in the corresponding knowledge question-answering methods. For any undisclosed details, please refer to the implementation methods in the method section, and therefore will not be repeated here.

[0149] It should be noted that although several modules or units of the knowledge question-answering device have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0150] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0151] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be embodied in the following forms: a completely hardware embodiment, a completely software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0152] Exemplary embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the aforementioned knowledge-based question-and-answer method.

[0153] In one implementation, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. (See reference...) Figure 6 , Figure 6 The illustration schematically depicts a computer-readable storage medium according to an exemplary embodiment of the present disclosure. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, a computer program product can be implemented as a non-volatile storage medium storing a computer program, such as read-only memory, NAND flash memory, etc.

[0154] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.

[0155] Computer program code can be written in one or more programming languages. Examples of programming languages ​​include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).

[0156] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic fields, and infrared radiation. Electronic devices can convert the signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, to execute by the processor of the electronic device) the method steps of various exemplary embodiments of this disclosure, such as the knowledge-answering method described above.

[0157] Exemplary embodiments of this disclosure also provide an electronic device, which may include a processor and a memory. The memory stores executable instructions of the processor, such as a computer program. The processor executes the executable instructions to perform the method steps of various exemplary embodiments of this disclosure. Furthermore, the electronic device may also include a display for displaying a graphical user interface.

[0158] The following is for reference. Figure 7 The electronic device is illustrated by way of a general-purpose computing device. It should be understood that... Figure 7 The electronic device 700 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0159] like Figure 7 As shown, the electronic device 700 may include: a processor 710, a memory 720, a bus 730, an I / O (input / output) interface 740, a network adapter 750, and a display 760.

[0160] The memory 720 may include volatile memory, such as RAM 721 and cache unit 722, and may also include non-volatile memory, such as ROM 723. The memory 720 may also include one or more program modules 724, including but not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. For example, program module 724 may include the modules described above.

[0161] The processor 710 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit).

[0162] The processor 710 can be used to execute executable instructions stored in the memory 720, such as the aforementioned knowledge question and answer method.

[0163] Bus 730 is used to connect different components of electronic device 700 and may include a data bus, an address bus and a control bus.

[0164] Electronic device 700 can communicate with one or more external devices 800 (such as keyboard, mouse, external controller, etc.) through I / O interface 740.

[0165] Electronic device 700 can communicate with one or more networks via network adapter 750. For example, network adapter 750 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication. Network adapter 750 can communicate with other modules of electronic device 700 via bus 730.

[0166] Electronic device 700 can display a graphical user interface, such as a user question and answer interface, through display 760.

[0167] although Figure 7 As not shown in the diagram, other hardware and / or software modules may also be configured in the electronic device 700, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0168] As can be seen from the above, the technical solutions disclosed herein can be implemented as methods, apparatus, systems, computer program products, storage media, electronic devices, etc. Those skilled in the art will understand that various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be referred to as "circuit," "module," or "system," respectively.

[0169] It should be understood that this disclosure is not limited to the specific methods, steps, or structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. Those skilled in the art will readily conceive of other embodiments based on the specific implementations provided in this disclosure. Therefore, the specific implementations provided in this disclosure are merely exemplary, and the scope and spirit of this disclosure are indicated by the claims, and should cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary technical means in the art not disclosed in this disclosure.

Claims

1. A knowledge-based question-and-answer method, characterized in that, include: A pre-built knowledge question-answering system is used to perform knowledge retrieval processing on the target question to obtain the target answer corresponding to the target question. The knowledge question-answering system is obtained by training on recommended question-answer pairs in a specified domain. The recommended question-answer pairs in the specified domain are the feedback results of the documents to be processed in the specified domain after knowledge retrieval enhancement processing. The documents to be processed in the specified domain include tabular data.

2. The method according to claim 1, characterized in that, The method further includes: The original documents in the specified domain are subjected to data cleaning and text slicing to obtain the document to be processed, which includes text slices; A tabular knowledge graph is constructed based on the document to be processed. The tabular knowledge graph is then processed into a graph text, and the text slices are combined to generate a text vector representation of the document to be processed.

3. The method according to claim 2, characterized in that, The construction of a tabular knowledge graph based on the document to be processed includes: The document to be processed is subjected to table extraction processing to obtain the table data, the table data including table context information, the table context information including table content and content hierarchy relationship; The table data is normalized to obtain normalized table data; The normalized tabular data is subjected to entity-relation mapping to obtain the tabular knowledge graph.

4. The method according to claim 2, characterized in that, The step of performing graph-to-text processing on the table knowledge graph and generating a text vector representation of the document to be processed by combining the text slices includes: Based on a predefined text template, retrieve entity and relation placeholders from the table knowledge graph; Based on the retrieved entity types and context information, the entity and relation placeholders are filled to obtain natural language table data; The natural language table data and the text slices are processed into text vectors to obtain the text vector representation.

5. The method according to claim 1, characterized in that, The method further includes: The obtained user questions are processed by keyword extraction and vectorization to obtain question keywords and question vector representations; The retrieval process is performed based on the question keywords and the question vector representation respectively to obtain the question retrieval results, which are used as training data for the large language model.

6. The method according to claim 5, characterized in that, The process of performing retrieval based on the question keywords and the question vector representation to obtain question retrieval results includes: Based on the aforementioned problem keywords, a graph retrieval process is performed to obtain the graph retrieval results; Sparse retrieval processing is performed based on the keywords of the problem to obtain keyword text slices; Vector retrieval processing is performed based on the aforementioned question vector representation to obtain relevant text vector representations; The graph retrieval results, the keyword text slices, and the related text vector representations are deduplicated and reordered to obtain the question retrieval results.

7. The method according to claim 1, characterized in that, The method further includes: Based on the received user question, a query operation is performed on the memory database to obtain the question query result corresponding to the user question. The memory database is used to store recommended question-answer pairs. When the query result for the question shows a matching answer, the matching answer is returned. When the query result for the question is no matching answer found, knowledge retrieval enhancement processing is performed based on the user question to obtain a recommended answer corresponding to the user question. The recommended answer is used to generate a new recommended question-answer pair with the user question after annotation processing.

8. The method according to claim 7, characterized in that, The step of performing knowledge retrieval enhancement processing based on the user question to obtain a recommended answer corresponding to the user question includes: The large language model sends the user's question to the knowledge retrieval enhancement module; The knowledge retrieval enhancement module performs knowledge retrieval processing based on the user's question to obtain the recommended answer.

9. The method according to claim 8, characterized in that, The step of obtaining the recommended answer by the knowledge retrieval enhancement module based on the user's question includes: Based on the user's question, a retrieval operation is performed on the knowledge base to obtain multiple initial recommended answers. The knowledge base is used to store one or more of the following: text slices, vector representations, knowledge graphs, and recommended question-answer pairs. The knowledge base includes a memory database. The multiple initial recommended answers are aggregated and reordered to obtain the recommended answer.

10. The method according to claim 7, characterized in that, The method further includes: The recommended answer is returned to the user interface layer, through which user annotation operations are received; Based on the user's question and the recommended answer after the user's annotation operation, the newly added recommended question-answer pair is generated and stored in the memory database.

11. A knowledge-based question-and-answer system, characterized in that, include: The large language model module is used to determine the target answer to the target question. The large language model module is obtained by adjusting the parameters of the pre-trained language model using recommended question-answer pairs in a specified domain. The knowledge retrieval enhancement module is used to perform knowledge retrieval processing on the knowledge retrieval data according to the target question, obtain the question retrieval result, and return the question retrieval result to the large language model module. The knowledge retrieval data is generated based on the document to be processed in a specified domain. The knowledge retrieval data includes one or more of text slices, text vector representations, and tabular knowledge graphs. The knowledge base module is used to store document data in a specified domain. The document data includes one or more of the following: the document to be processed, text slices, text slice vectors, recommended question-answer pairs, question-answer pair vectors, and tabular knowledge graphs.

12. The system according to claim 11, characterized in that, The knowledge retrieval enhancement module includes a text slicing unit, a text vectorization unit, and a graph generation unit; The text slicing unit is used to perform text segmentation processing on the document to be processed to obtain the text slices; The text vectorization unit is used to perform vectorization processing on the text slice to obtain the text slice vector; The graph generation unit is used to create a knowledge graph based on the document to be processed, and the knowledge graph includes a tabular knowledge graph.

13. The system according to claim 11 or 12, characterized in that, The knowledge retrieval enhancement module includes a knowledge retrieval unit and an answer determination unit; The knowledge retrieval unit is used to perform knowledge retrieval enhancement processing based on the target question to obtain the question retrieval result. The knowledge retrieval enhancement processing includes one or more of graph retrieval, vector retrieval, and sparse retrieval. The answer determination unit is used to perform deduplication and reordering on the question retrieval results to obtain the recommended question-answer pair.

14. The system according to claim 11, characterized in that, The knowledge retrieval enhancement module also includes a metadata management unit, a problem understanding unit, and a retrieval evaluation unit; The metadata management unit is used to manage the metadata of the document to be processed; The question understanding unit is used to determine the question intent of the target question and provide a matching target answer, wherein the question intent is used to determine the target answer that matches the target question; The retrieval evaluation unit is used to evaluate the recommendation effect of the recommended question-answer pair.

15. The system according to claim 11, characterized in that, The knowledge base module includes a file storage unit, an index storage unit, and a graph storage unit; The file storage unit is used to store the document to be processed; The index storage unit is used to store one or more of the text slices, the text slice vectors, the recommended question-answer pairs, and the question-answer pair vectors. The text slices are obtained by performing text segmentation on the document to be processed. The graph storage unit is used to store the table knowledge graph.

16. The system according to claim 11 or 14, characterized in that, The knowledge base module includes a data cleaning unit; The data cleaning unit is used to collect original documents and recommended question-and-answer pairs in a specified field, and to perform page number and table of contents deletion and parsing and recognition processing on the original documents after deduplication to obtain table recognition results and paragraph recognition results.

17. A knowledge-based question-and-answer device, characterized in that, include: The knowledge question answering module is used to perform knowledge retrieval processing on a target question using a pre-built knowledge question answering system to obtain the target answer corresponding to the target question. The knowledge question answering system is obtained by training on recommended question-answer pairs in a specified domain. The recommended question-answer pairs in the specified domain are the feedback results of the documents to be processed in the specified domain after knowledge retrieval enhancement processing. The documents to be processed in the specified domain include tabular data.

18. An electronic device, characterized in that, include: processor; as well as A memory storing computer-readable instructions that, when executed by the processor, implement the knowledge question-answering method according to any one of claims 1 to 10.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the knowledge question-answering method according to any one of claims 1 to 10.

Citation Information

Cited By

  • A method, apparatus, equipment, and medium for corpus mining of a domain-wide large model

    CN122412676A

  • A corpus mining method, device and equipment of a domain large model and a medium

    CN122412676B