Construction method of letter intelligent question-answering system based on adaptive mapping knowledge domain enhanced LLM
By using an adaptive knowledge graph to enhance LLM, we have solved the illusion problem of large language models in complex text processing and the problem of inaccurate answers in professional domains, and achieved a highly accurate intelligent question answering system.
Patent Information
- Application Number
- CN202511100068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-07
AI Technical Summary
Large language models are prone to illusions when processing complex text, require extensive training, and are inaccurate in answering questions in specialized domains. Traditional question answering systems struggle to effectively handle questions involving rare or long-tailed entities.
By acquiring and preprocessing text data from the target region, using a fine-tuned sentence converter to convert it into vector format and constructing a vector database, and combining it with a pre-built knowledge graph for joint reasoning verification, the final answer is generated.
It improves the accuracy and reliability of the question-answering system, enabling it to adaptively handle complex questions and generate logically sound answers.
Smart Images

Figure CN120910216A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent question answering, and in particular to a construction method of a letter intelligent question answering system based on adaptive knowledge graph enhanced LLM. BACKGROUND
[0002] Large language models are an important achievement of deep learning technology in the field of natural language processing. They have a large number of parameters and complex neural network structures, and can learn human language by analyzing large amounts of text data. The main advantage of large language models is their strong language understanding and generation capabilities, which are due to pre-training on large-scale datasets, allowing large language models to master the structure, grammar, and semantics of language, thereby performing well in various natural language processing tasks, including text generation, sentiment analysis, and machine translation.
[0003] A knowledge graph (KG) is a knowledge base that represents and stores entities and their relationships through a graph structure. It organizes information into nodes (representing entities) and edges (representing relationships between entities), making the connections and attributes between different entities more clear through the form of a graph.
[0004] In the field of artificial intelligence technology, intelligent question answering systems have become an important tool for enterprises and organizations to efficiently utilize data resources. Traditional question answering systems mainly rely on rule-based methods or simple retrieval techniques. However, as data scales expand and user demands become more complex, large language models are prone to hallucinations when dealing with rare or long-tail entities. When large language models lack sufficient knowledge, they still attempt to provide answers, even if they are inaccurate or completely fabricated.
[0005] There are now methods that combine knowledge graphs with large language models for question answering to overcome the problem of long-tail entities and reduce the risk of hallucinations by large language models. However, they often require a large amount of training or inefficiently handle complex graph structures. SUMMARY
[0006] The present application aims to provide a construction method of a letter intelligent question answering system based on adaptive knowledge graph enhanced LLM, which solves the problem of hallucinations of large language models when dealing with complex text, the need for large amounts of training of large language models, and inaccurate answers in specialized fields.
[0007] To achieve the above-mentioned purpose, the following technical solutions are adopted in the present application: A construction method of a letter intelligent question answering system based on adaptive knowledge graph enhanced LLM, comprising the following steps: Obtain text data of a target area, and after preprocessing the text data, convert it to vector format through a fine-tuned sentence converter and construct a vector database; Vectorize the user question text and identity information, calculate the similarity with the vector database to generate context information, input the context information and user question into the generative large language model, and output multiple candidate answers; Based on the pre-constructed target domain knowledge graph, the multiple candidate answers are jointly inferred and verified, and the candidate answers are sorted and the final answer is selected through a scoring mechanism. In the above scheme, the text data of the target area is obtained by grabbing text data related to the target domain from different data sources, and the preprocessing includes data cleaning and chunking preprocessing.
[0008] In the above scheme, the text data related to the target domain is grabbed from different data sources, specifically: grabbing text data of the corresponding field from multiple data sources, identifying key information of the field from file text data, and outputting the key information in json format, with the field as the key and the corresponding key information as the value.
[0009] In the above scheme, the sentence converter is converted into a vector format and a vector database is constructed through fine-tuning, specifically: The preprocessed target domain text data and file text data are converted into a vector form and a vector database is constructed using the fine-tuned All-Mpnet-Base-V2 sentence converter model. In the above scheme, the short question and answer pairs required for training the sentence converter are generated using GenQ technology, and the All-Mpnet-Base-V2 sentence converter model is fine-tuned.
[0010] In the above scheme, the similarity is cosine similarity, and the calculation formula is: , Where query is the vector form of the user question, and chunk is the text block vector in the knowledge base.
[0011] In the above scheme, based on the pre-constructed target domain knowledge graph, the multiple candidate answers are jointly inferred and verified, and the candidate answers are sorted and the final answer is selected through a scoring mechanism, specifically: Extract the relevant subgraph from the knowledge graph according to the user question, which contains entity nodes and relationship edges directly connected to the question entity. Collect the label text of all nodes in the subgraph and concatenate to generate subgraph text in topological order in the graph. Calculate the ROUGE-L F1 score between multiple potential answers and the subgraph text, and select the highest score answer as the final output In the above scheme, the relevant subgraph is extracted from the knowledge graph according to the user question, specifically including: (1) Extract key entities from user questions, locate them through named entity recognition, link them to knowledge graph nodes, and when there are multiple matching nodes, sort and filter them based on semantic similarity scores; (2) Analyze the complexity of the question using a language model, generate a SPARQL query framework, and predict the maximum exploration jump number K required; (3) Starting from the linked entity nodes, iterate within K jumps: The ith jump obtains all the connection nodes and relationships of the current node; score the semantic relevance of the new nodes and relationships; keep the nodes with scores higher than the threshold τ as the starting point for the next jump; (4) When K jumps or no qualified nodes are reached, output the connected subgraph containing the answer entity.
[0012] As can be seen from the above scheme, the construction method of the letter intelligent question answering system based on adaptive knowledge graph enhanced LLM can realize the answering of complex questions, improve the accuracy and reliability of the question and answer; without being based on a preset answer template, the answer can be accurately generated through reasoning, ensuring the accuracy and logic of the answer. BRIEF DESCRIPTION OF DRAWINGS Fig. 1 is a method flowchart of the present application; Fig. 2 is a knowledge graph construction flowchart of the present application; Fig. 3 is a flowchart of the converter execution. DETAILED DESCRIPTION
[0013] The present application will be further described below in conjunction with the drawings: As Figs. 1-3 shown, the present application provides a construction method of a letter intelligent question answering system based on adaptive knowledge graph enhanced LLM, the specific steps are as follows: Step 1: Aggregate corpus from multiple heterogeneous data sources, dynamically crawl text data related to the target domain, preprocess the file and target domain text data, and use a fine-tuned sentence converter to convert the preprocessed text data into item vector format and store it as a database.
[0014] Specifically, first, determine the corresponding field of the constructed knowledge base according to the file, crawl the text data of the corresponding field from multiple data sources, and pre-use a large language model trained using a data set containing text samples of multiple fields and labeled with corresponding field tags to extract key information of the recognized field from the file text data. Obtain the corresponding field text data from different data sources. Use the prompt word engineering to generate the prompt word of the large language model, and output the key information in json format, with key as the field and value as the corresponding key information.
[0015] The text data in a specific field is preprocessed, the text information is cleaned, irrelevant special characters are removed, and the text is divided into blocks to obtain the preprocessed text.
[0016] All-Mpnet-Base-V2 is a Transformer model based on the MPNet architecture that encodes sentences, paragraphs, and documents into fixed-length vectors. By fine-tuning on sentence-pair tasks, it can improve its ability to represent text in specific domains.
[0017] The preprocessed target domain text data and file text data are converted into vector form using the fine-tuned All-Mpnet-Base-V2 sentence converter model and stored in the knowledge base, so that other semantically similar text documents can be found in the database using cosine similarity and other metrics.
[0018] Fine-tuning the sentence converter model on a specific domain can improve its ability to represent text in that domain. The All-Mpnet-Base-V2 sentence converter model is fine-tuned by generating short question-answer pairs required for training using GenQ technology.
[0019] To implement GenQ, a large Google T5 LLM is used, which has been trained to generate question-answer pairs from a given text body to fine-tune our sentence converter model. Preprocessed text data is used for fine-tuning. First, multiple paragraphs are generated from all preprocessed text, and these paragraphs are passed to the model BeIR / query-gen-msmarco-t5-large-v17. The T5 model then generates question-answer pairs, which are used to fine-tune the All-Mpnet-Base-V2 model.
[0020] Step 2: Vectorize the user question text and identity information, calculate the similarity with the vectors in the knowledge base, generate context information, and pass it to a large generative LLM along with the user question to get 5 candidate answers generated by the LLM.
[0021] Convert the user question into vector format by fine-tuning the sentence converter model, then search and extract the 5 most similar answer contexts from the knowledge base using cosine similarity, and pass the question and the extracted 5 contexts to a large generative LLM to generate answers.
[0022] The formula for calculating the cosine similarity is: (1) Where query is the vector form of the user question, and chunk is the vector of the text block in the knowledge base.
[0023] Step 3: Based on the knowledge graph, use the method of joint reasoning to determine the final answer instead of the preset answer template, apply the scoring mechanism, and sort the 5 candidate answers to select a final answer to answer the user's question.
[0024] First, the relevant subgraph is obtained from the overall knowledge graph according to the question raised by the user. Then, the labels of all nodes in the subgraph are collected to form a string, called subgraph text. Finally, the ROUGE-L score between each potential answer and the subgraph text is calculated. The final answer selects the answer with the highest ROUGE-L F1 score and displays it to the user.
[0025] For the acquisition process of the relevant subgraph, an adaptive strategy is used to explore the knowledge graph to obtain a more refined subgraph, the steps of which are as follows: First, the key entities are extracted from the user's input question, which is completed through the named entity recognition (NER) based on the Transformer SpaCy model.
[0026] Then, by matching with the entities in the knowledge graph, the extracted entities are linked to the possible entity nodes in the knowledge graph. If an entity has multiple matching items in the knowledge graph, the matching items are sorted by using semantic similarity scoring. Semantic similarity includes two parts: 1. Question-Entity Similarity: By calculating the similarity between the input question and the description of the entity in the knowledge graph, it is obtained which entities are most relevant to the query.
[0027] 2. Question-Relationship Similarity: By calculating the similarity between the question and the relationship description in the knowledge graph, it is obtained which relationship between entities is most relevant to answering the question.
[0028] After identifying the relevant entities, the number of hops required to answer the question is predicted by the language model according to the complexity of the question. First, a prompt is constructed, requiring the language model to create a SPARQL query for the given question and entities. The generated SPARQL query is then analyzed to calculate the required number of hops, which serves as an initial estimate of the depth of exploration required.
[0029] In the knowledge graph exploration process, for each hop, first extract all directly connected nodes of the current entity set. These connected nodes and their relationships are sorted according to their semantic similarity to the original question. From the remaining candidates, we select the top N entities, where N is a user-defined parameter. This process is repeated for each hop until the maximum depth predicted.
[0030] By utilizing the semantic similarity scores, the method can adaptively focus on the most relevant paths in the knowledge graph. The exploration process can be fine-tuned by several key parameters. This strategy filters out entities that are not sufficiently relevant to the question, even if they are connected through relevant relations. These thresholds work in concert to refine the exploration process, ensuring that both the relations and entities in the subgraph remain highly relevant to the original question.
[0031] After obtaining the relevant subgraph based on the above method, collect the labels of all nodes in the subgraph to form a string, i.e. subgraph text, and calculate the ROUGE-L score between each potential answer and the subgraph text. The final answer selects the answer with the highest ROUGE-L F1 score. The calculation method of ROUGE-L is as follows: ROUGE-L is an evaluation metric used for automatically evaluating the quality of text generation tasks in natural language processing. It focuses specifically on the longest common subsequence (LCS) between the generated text and the reference text.
[0032] (2) where Recall represents the proportion of the LCS in the generated text that matches the LCS in the reference text.
[0033] (3) where Precision represents the proportion of the LCS in the generated text that matches the LCS in the reference text.
[0034] (4) where, ROUGE-L F1 score is the harmonic mean of precision and recall, which comprehensively measures the similarity between the generated text and the reference text.
[0035] The above-described embodiments are merely descriptions of the preferred embodiments of the present application and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application given by those skilled in the art should fall within the scope of protection determined by the claims of the present application.
Claims
1. A method for constructing a letter intelligent question and answer system based on adaptive knowledge graph enhanced LLM, characterized in that, The method comprises the following steps: Obtaining text data of a target region, preprocessing the text data, converting the text data into a vector format through a fine-tuned sentence converter, and constructing a vector database; Vectorizing user question text and identity information, calculating the similarity between the vectorized user question text and the vector database to generate context information, inputting the context information and the user question into a generative large language model, and outputting multiple candidate answers; Based on the pre-constructed target domain knowledge graph, the multiple candidate answers are jointly reasoned and verified, and the candidate answers are sorted and the final answer is selected through a scoring mechanism.
2. The method according to claim 1, wherein the method is characterized in that: The text data of the target region is obtained by grabbing text data related to the target domain from different data sources, and the preprocessing includes data cleaning and chunking preprocessing.
3. The method according to claim 2, wherein the method is characterized in that: The text data related to the target domain is grabbed from different data sources, specifically as follows: Text data of the corresponding field is grabbed from multiple data sources, and key information of the field is identified from the text data of the file, wherein the key information is output in json format, the key is the field, and the value is the corresponding key information.
4. The method according to claim 3, wherein, The fine-tuned sentence converter is used to convert the preprocessed target domain text data and file text data into a vector format and construct a vector database, specifically as follows: The fine-tuned All-Mpnet-Base-V2 sentence converter model is used to convert the preprocessed target domain text data and file text data into a vector format and construct a vector database.
5. The method according to claim 1, wherein the method is characterized by, The fine-tuned sentence converter is used to convert the preprocessed target domain text data and file text data into a vector format and construct a vector database, specifically as follows:
6. The method according to claim 1, wherein the method is characterized by, The similarity is a cosine similarity, and the calculation formula is as follows: , Wherein, query is the vector form of the user question, and chunk is the text block vector in the knowledge base.
7. The method according to claim 1, wherein the method is characterized by, Based on the pre-constructed target domain knowledge graph, the multiple candidate answers are jointly reasoned and verified, and the candidate answers are sorted and the final answer is selected through a scoring mechanism, specifically as follows: According to the user question, a related subgraph is extracted from the knowledge graph, and the subgraph includes entity nodes and relationship edges directly connected to the problem entity; Collecting the label text of all nodes in the subgraph, and concatenating to generate subgraph text in topological order in the graph; Calculate the ROUGE-L F1 score between the multiple potential answers and the subgraph text, and select the highest score answer as the final output.
8. The method according to claim 4, wherein the method is characterized by, According to the user question, a related subgraph is extracted from the knowledge graph, specifically as follows: (1) Extract key entities from the user question, locate them through named entity recognition, link them to knowledge graph nodes, and when there are multiple matching nodes, sort and filter based on semantic similarity score; (2) Analyze the complexity of the question using a language model, generate a SPARQL query framework, and predict the maximum exploration jump number K; (3) Starting from the linked entity node, iterate within K jumps: Get all connected nodes and relationships of the current node at the ith jump; score the semantic relevance of the new nodes and relationships; keep the nodes with a score higher than the threshold τ as the starting point for the next jump; (4) When K-hop or no qualified nodes are reached, output the connected subgraph containing the answer entity.
Citation Information
Cited By
Large model routing data synthesis method with diversified expression styles
CN121092680A
Intelligent conversion method and system for data units
CN121412399A
Intelligent conversion method and system of data units
CN121412399B
Dizziness inquiry information processing method and device based on knowledge graph constraint
CN121528579A