Multi-round knowledge-guided question and answer method and system fusing large language model and knowledge graph

Through a multi-round knowledge-guided question-and-answer method that integrates large language models and knowledge graphs, the shortcomings of existing question-and-answer systems in semantic understanding and natural language generation are solved, and user questions are answered efficiently and accurately in complex fields, improving the accuracy of user needs analysis and comprehensiveness of answers.

CN120069068AActive Publication Date: 2025-05-30GUANGZHOU SEMANTIC TECH CO LTD

Patent Information

Application Number
CN202510126681.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

The existing Q&A system has shortcomings in semantic understanding and natural language generation, especially in multiple rounds of Q&A scenarios, and it is difficult to accurately extract the user's real needs, resulting in poor accuracy and reliability of answers, especially in complex areas such as social security policies and regulations.

Method used

Through the multi-round knowledge-guided question-and-answer method that integrates large language models and knowledge graphs, the path guidance ability of the knowledge graph and the semantic ability of the pre-trained large language model are deeply integrated to achieve multiple iterative extraction of user needs, approaching the user's final intention layer by layer, and generating accurate answers.

Benefits of technology

It realizes efficient and accurate answers to user questions in complex fields, improves the accuracy of user needs analysis and comprehensiveness of answers, and is efficient and scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069068A_ABST
    Figure CN120069068A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-round knowledge-guided question-answering method and system fusing a large language model and a knowledge graph. The method comprises the following steps: S1, constructing the knowledge graph; s2, performing retrieval and clarification; s3, extracting labels and classifying questions; s4, performing scene guidance and dynamic retrieval; and S5, answer generation and formatting output. The method aims at solving the fuzzy problem in complex policy and regulation questions and answers, and the accuracy of user demand analysis and the comprehensiveness of answers are improved. According to the method, a knowledge graph is taken as a core, and multi-path semantic analysis and problem guidance capabilities are provided by constructing a hierarchical structure covering first-level items, second-level items, scene categories and keyword nodes. In combination with strong semantic comprehension and generation capabilities of the pre-trained large language model, the system can dynamically track context information in user input, iteratively access related nodes along different paths of the knowledge graph, and gradually clarify user intentions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a multi-round knowledge-guided question-answering method and system that integrates large language models and knowledge graphs. Background Art

[0002] With the rapid development of information technology, people's demand for information acquisition is increasing day by day, and question-answering systems, as a convenient information interaction tool, have received extensive attention. Traditional question-answering systems often rely on a single knowledge source or technical means and have many limitations. For example, a question-answering system based solely on a knowledge graph can accurately use structured knowledge for answering, but is relatively weak in semantic understanding and natural language generation; while a question-answering system that relies solely on a large language model, although excellent in language understanding and generation, may not be able to extract the true needs of users due to untimely knowledge updates or lack of structured knowledge support, resulting in poor accuracy and reliability of the answers and being unable to be applied to question-answering in complex fields. In multi-round question-answering scenarios, these problems are even more prominent, and the gradual refinement of user needs and the requirement of context coherence pose higher challenges to question-answering systems. Summary of the Invention

[0003] Aiming at the problems of the prior art, the purpose of the present invention is to provide a multi-round knowledge-guided question-answering method and system that integrates large language models and knowledge graphs. Through the deep integration of the path guidance ability of the knowledge graph and the semantic ability of the pre-trained large language model, it can iteratively extract the true needs of users in multiple rounds, gradually approach the user's ultimate intention, achieve accurate answers to complex questions, and have high efficiency and scalability, and is applicable to intelligent question-answering in complex fields such as social security policies and regulations.

[0004] To achieve the above objectives, the present invention provides the following technical solutions:

[0005] The present invention provides a multi-round knowledge-guided question-answering method that integrates large language models and knowledge graphs, including:

[0006] S1. Construct a knowledge graph: Perform semantic parsing and hierarchical processing on the input document content, extract and label the first-level matters, second-level matters, scenario categories, relevant keywords, and questions and answers of the corresponding knowledge in the document; through the entity alignment and relationship extraction module, construct the extracted content into a knowledge graph with a hierarchical structure;

[0007] S2. Retrieve and clarify: Based on the user input, perform semantic parsing through the pre-trained large language model, and perform hybrid retrieval in combination with the knowledge graph and knowledge base constructed in step S1. If the retrieval result is not clear, trigger the clarification process;

[0008] S3. Tag Extraction and Question Classification: Extract the first-level matters, second-level matters, scenario categories, and keywords from the user input in step S2 through the tag extraction module, and classify and process the questions according to the extraction results;

[0009] S4. Scenario Guidance and Dynamic Retrieval: Based on the knowledge graph constructed in step S1 and the tag information extracted in step S3, determine the details of the scenario category, dynamically generate retrieval conditions, and perform further accurate queries on the knowledge base;

[0010] S5. Answer Generation and Formatted Output: Combine the context information obtained in steps S1 to S4, and through the Retrieval-Augmented Generation (RAG) module, use the pre-trained large language model to generate answers, and optimize the answer output through the formatting module.

[0011] In step S1, the hierarchical knowledge graph includes knowledge label matter paths, knowledge questions, knowledge question vectors, knowledge answers, and knowledge answer vectors.

[0012] In a possible implementation manner, in step S1, the specific process of entity alignment and relationship extraction is as follows:

[0013] S11. Perform text cleaning and context semantic enhancement on the input document content D = {d 1 , d 2 ,..., d n}, and convert it into a normalized text representation; based on the sentence embedding model E s , generate the vectorized representation v i = E s (d i ) for subsequent entity and relationship analysis; for the document question and answer parts, use the extraction model K(z) to extract the sets K Q and K A to ensure the extraction accuracy of Q&A knowledge;

[0014] S12. Adopt a relationship extraction method based on a pre-trained large language model to construct entity and relationship Prompt templates; analyze the document context, and adopt a combination of semantic parsing and nested hierarchy extraction to construct a hierarchical knowledge graph, which includes:

[0015] S121. First-level matter nodes T 1 : Extract the main matters or topics involved in the document, represented as a node set These nodes usually represent high-level policy categories or domains;

[0016] S122. Second-level matter nodes T 2: Further refine and extract secondary matters from the context of primary matters to form a set of child nodes associated with the primary matters to more precisely represent the refined content;

[0017] S123. Scene category node S: According to the description of the document and the application background, extract the set of nodes related to the scene S = {s 1 , s 2 ,..., s k}, which is used to represent the usage scenario of knowledge;

[0018] S124. Keyword node K: Extract the core keywords K = {k 1 , k 2 ,..., k p} in the document, and associate them with the relevant primary matters, secondary matters, and scene category nodes.

[0019] In a possible implementation manner, in step S1, the specific process of constructing the knowledge graph is as follows:

[0020] 1. Use the pre-trained large language model to perform in-depth parsing on the semantic meaning at the sentence level of the document and generate the context vector representation Combine with the syntactic analysis model to extract the hierarchical entity nodes and their relationships;

[0021] 2. Design the recursive relationship extraction function RelExtract(·) to associate the primary node T 1 with the keywords, secondary nodes in its context:

[0022]

[0023] Among them, Rih represents the relationship type between the primary matter and the secondary matter;

[0024] 3. The extraction of the scene category node S is through context nested analysis, combined with the association probability P(s|t i , k j ), to generate the scene branch structure of the hierarchical graph;

[0025] Finally, all nodes and relationships will be organized into a hierarchical knowledge graph in the following form:

[0026] G = (N, E), N = T 1 ∪T 2 ∪S∪K, E = {(n i , n j , R ij )|n i , n j ∈N}, (2)

[0027] Among them, N is a set of nodes, E is a set of relationships between nodes, and R ij represents the relationship between any two nodes;

[0028] The construction process can dynamically generate hierarchical nodes and relationships according to the document content, ensuring the accuracy and scalability of the knowledge graph, and providing structured semantic support for subsequent retrieval and clarification (step S2).

[0029] In step S2, the clarification process includes clarification of first-level label matters, clarification of second-level label matters, scenario-guided clarification, and rejection process.

[0030] In a possible implementation, in step S2, the specific process of semantic parsing is as follows:

[0031] S211. Vectorization of user input:

[0032] For the question Q u input by the user, use the pre-trained large language model to generate an embedding vector generated by the following formula:

[0033]

[0034] Among them, Embedding(·, θ) represents the semantic embedding generation function, and θ is the model parameter;

[0035] If the user input involves multi-round context, combine the current input with the historical conversation content, and generate a context semantic vector by splicing

[0036]

[0037] Among them, H t-1 represents the previous round of conversation history;

[0038] S212. Using the semantic structure information of the knowledge graph G = (N, E) constructed in step S12, compare the user input with the node embedding vectors in the constructed knowledge graph G for preliminary matching, and calculate the similarity of each node:

[0039]

[0040] According to the similarity threshold τ 1 , extract the set of highly relevant nodes including first-level matters T 1 , second-level matters T 2 , scenario categories C s and keywords K;

[0041] Preliminarily determine the knowledge nodes related to the user input and their association relationships, providing a basis for the subsequent retrieval module.

[0042] In a possible implementation manner, in step S2, the specific process of retrieval is as follows:

[0043] S221. High-confidence quick matching:

[0044] If among the node set N matched by the knowledge graph, nodes that match similar vectors are retrieved through k-NN retrieval of Elasticsearch, and the similarity Sim of the key question in a certain node is ≥ 0.98 or the re-ranking correlation Rerank of the answer is ≥ 0.95, directly call the relevant document as the preliminary answer and enter the formatting module, skipping the clarification process;

[0045] S222. Vector retrieval expansion:

[0046] If the high-confidence nodes are not hit, use the knowledge graph in step S1 and the embedding vector of the user question to construct a retrieval condition:

[0047] Q r = ConstructQuery(T 1 ,T 2 ,C s ,K), (6)

[0048] where Q r is a retrieval condition dynamically generated in combination with the knowledge graph nodes, used to call external vector retrieval (such as k-NN retrieval of Elasticsearch);

[0049] Retrieve the candidate document set D, and score the relevance of the candidate documents through the re-ranking module, filtering out low-relevance documents:

[0050] D filtered = {d|RelScore(d)≥τ 2 ,d∈D}, (7)

[0051] where D filtered is the retrieved knowledge document; d is a single knowledge document, which is an element of the original knowledge document set D filtered ; RelScore is the document relevance score, indicating the matching degree between d and the user query, usually calculated by a certain vector similarity, such as cosine similarity or dot product; τ 2 is the set similarity or relevance threshold;

[0052] S223. Prompt construction:

[0053] Use the matched node information in the knowledge graph (such as first-level matters, scenario categories, etc.) and the content of D filtered to construct a Prompt and call a pre-trained large language model to generate a preliminary answer A raw .

[0054] In a possible implementation manner, in step S2, the specific process of triggering the clarification process is as follows:

[0055] S231. Triggering condition:

[0056] If D filtered is empty, or the uncertainty of the preliminary answer A raw (according to the confidence of the generation model) is higher than the threshold α, trigger the clarification process;

[0057] S232. Clarification question generation:

[0058] Based on the knowledge graph nodes and relationships in step S1, generate a clarification question:

[0059] 1. If the first-level matter T 1 is not clear, first retrieve all first-level label matters and the currently obtained knowledge graph according to facts, and generate a clarification question for T 1 through a pre-trained large language model;

[0060] 2. If the second-level matter T 2 is not clear, first retrieve all second-level label matters and the currently obtained knowledge graph according to facts, and generate a clarification question for T 2 through a pre-trained large language model;

[0061] 3. If the keyword set K is incomplete, generate a clarification question for the keyword;

[0062] S233. User interaction and update:

[0063] Send the clarification question to the user, receive the user's feedback, and update the current session context Context and the knowledge graph node set;

[0064] S234. Re-retrieval after clarification:

[0065] Re-execute the retrieval process according to the user's feedback until all key nodes are clear or the rejection process is triggered.

[0066] In step S3, classifying and processing the problem according to the extraction result means judging which type of elements are missing according to the slot filling extracted by the user, and performing a follow-up question process according to the specific missing elements.

[0067] In a possible implementation manner, in step S3, the specific process of label extraction is as follows:

[0068] S311. Input data:

[0069] Input the question Q entered by the user u , the retrieved knowledge document D filtered , the clarified context information C u and the knowledge graph G;

[0070] S312. Tag extraction process:

[0071] Use the large model-based tag extraction method to generate the tag extraction Prompt:

[0072] P extract = ConstructPrompt(Q u , C u , D filtered , G), (8)

[0073] where P extract is the final extracted Prompt, representing the specific task description generated by combining the user query and context information, and is used to guide the pre-trained large model to generate answers;

[0074] The Prompt contains: user questions, context, retrieval results, and knowledge graph node information;

[0075] Extract the following information from the Prompt through the pre-trained large language model:

[0076] The first-level matter T 1 : The set of first-level node tags most relevant to the user question;

[0077] The second-level matter T 2 : The set of specific sub-node tags further restricted based on T 1 ;

[0078] The scenario category C s : The classification tag describing the semantic background of the question;

[0079] The keyword K: The set of words highly relevant to the user question;

[0080] S313. Tag identification generation:

[0081] Package the extraction result into the tag structure Tags:

[0082] Tags = {T 1 , T 2 , C s , K}, (9)

[0083] S314. Multiple authentication:

[0084] Verify the extracted tags using the knowledge graph G:

[0085] If T 1 or T 2 does not exist in the knowledge graph G, trigger the tag completion process (return to the clarification module in step S2);

[0086] If the extracted tag is ambiguous: confidence Conf < τ 3 , generate a clarification prompt to obtain more explicit information.

[0087] In a possible implementation manner, in step S3, the specific process of classifying and processing the problem is as follows:

[0088] S321. Classification rules:

[0089] According to the extracted tag information Tags, divide the user's question into the following categories:

[0090] 1. Direct answer type: The tag extraction is complete and clear, and the answer in the knowledge graph or retrieved document can be directly matched;

[0091] 2. Clarification requirement type: The tag extraction is incomplete (such as lacking primary or secondary matters), and further clarification is required;

[0092] 3. Dynamic retrieval type: The tag is clear, but refined retrieval needs to be combined with the scenario category and keywords;

[0093] S322. Classification discrimination formula:

[0094] Use a classification model or rule to determine:

[0095]

[0096] Among them, Complete(Tags) indicates whether the tag is complete; is when the scenario category C s or the keyword K is an empty set, that is, no tags are extracted;

[0097] S323. Classification output:

[0098] According to the classification result, adjust the subsequent process:

[0099] 1. For the direct answer type, enter the Retrieval-Augmented Generation (RAG) module to generate an answer;

[0100] 2. For the clarification requirement type, return to step S2 and enter the clarification Q&A process to supplement the missing information;

[0101] 3. For the dynamic retrieval type, the Tags structure (including primary matters, secondary matters, scenario categories, and keywords) is used as the input for scenario guidance in step S4, and scenario guidance and dynamic retrieval are performed.

[0102] In step S4, the retrieval conditions are dynamically generated based on the lack of scenario guidance type. Specifically, corresponding reference clarification statements are retrieved dynamically according to the scenario type for counter-questioning to accurately extract the specific scenario guidance details of the user.

[0103] In a possible implementation manner, in step S4, the specific process is as follows:

[0104] S41. Scenario condition initialization and dynamic expansion

[0105] S411. Input initialization:

[0106] Obtain the tag T extracted in step S3 1 , T 2 , C s , K and the preliminary retrieval result D filtered ;

[0107] Load the nodes and context information related to the current session from the knowledge graph G;

[0108] S412. Condition expansion:

[0109] If some tags are missing, such as T 2 or C s , then dynamically call the knowledge graph for condition expansion:

[0110] Find the set of possible tags associated with T 1 or K:

[0111]

[0112] Include in the scenario derivation process;

[0113] S413. Generate dynamic retrieval conditions:

[0114] Initial conditions:

[0115]

[0116] Condition optimization:

[0117] Use the retrieval history and feedback to adjust the conditions, such as increasing the priority of specific tags:

[0118] Q optimized = Q init + {"boost": TagsPriority(T 1 , T2 )}, (13)

[0119] S42. Dynamic Retrieval Execution and Feedback Iteration

[0120] S421. Initial Retrieval:

[0121] Based on Q optimized and the Elasticsearch configuration, perform a retrieval:

[0122] D scene = Elasticsearch(Q optimized ), (14)

[0123] Sort the returned results by relevance score and filter D filtered :

[0124] D filtered = TopN(D scene , N = 10, score > 0.8), (15)

[0125] S422. Feedback Iteration:

[0126] Generate a rhetorical Prompt for the pending tags:

[0127]

[0128] S423. Dynamically Supplement Pending Tags:

[0129] For the undefined C s , search for relevant expressions in the knowledge graph G;

[0130]

[0131] Use the expressions to clarify the scenario category until C s is defined or the user aborts the session;

[0132] S43. Scenario Refinement and Tag Confirmation

[0133] S431. Tag Confirmation Logic:

[0134] If T 1 , T 2 and K are defined, but C s is pending, use the following logic to refine;

[0135] 1. If the number of candidates for C s is 1, directly confirm it as the final tag:

[0136]

[0137] 2. If there are more than one candidate, dynamically generate a clarification option Prompt:

[0138]

[0139] Prompt the user to make a selection;

[0140] S432. Label priority sorting:

[0141] Sort the labels by priority based on the context information and user feedback to guide the subsequent process; S44. Confirm the output and perform multi-round optimization

[0142] S441. Result output:

[0143] Output the final determined label set T 1 , T 2 , K and the exact retrieval result D final ;

[0144] Ensure that D final is filtered and the content is highly relevant to the current context;

[0145] S442. Interface with the subsequent module:

[0146] Provide for the answer generation in step S5:

[0147] 1. The complete scenario category

[0148] 2. The high-quality knowledge document set D final ;

[0149] 3. The undetermined labels are used as feedback and returned to step S3 to form a closed loop.

[0150] In a possible implementation manner, in step S5, the specific process is as follows:

[0151] S51. Answer generation

[0152] S511. Context integration;

[0153] Construct the context information C based on the final scenario category, primary matter, secondary matter, and keyword information obtained in step S4, combined with the user's historical input and the current question user ;

[0154] S512. Knowledge retrieval integration:

[0155] Use the knowledge document K retrieved in step S4 retrieved as external knowledge support and input it together with the context information C user into the retrieval-enhanced generation module to construct the generative Prompt PRAG :

[0156] P RAG = f(C user , K retrieved ), (20)

[0157] Generate the preliminary answer A through a pre-trained large language model raw ;

[0158] The pre-trained large language model uses Qwen2.5 with 72B parameters;

[0159] S52. Answer optimization

[0160] S521. Consistency check:

[0161] Compare the generated preliminary answer A raw with the paths and nodes in the knowledge graph to ensure that the answer logic is consistent with the knowledge base; for parts that deviate from the knowledge graph, trigger supplementary retrieval and iteratively optimize to generate the answer;

[0162] S522. Formatting:

[0163] Input the preliminary answer A raw into the formatting module to optimize its language fluency, structural clarity, and expression of professional terms, and output the final answer A final ;

[0164] S53. Answer output

[0165] The system returns the final answer A final to the user, along with interactive prompt information, allowing the user to initiate further clarification or modification requests for the generated answer, thus realizing cyclic knowledge Q&A.

[0166] The present invention also provides a multi-round knowledge-guided Q&A system integrating a large language model and a knowledge graph, including an entity alignment and relation extraction module, a semantic parsing module, a retrieval module, a clarification Q&A module, a label extraction module, a question classification module, a scenario guidance and dynamic retrieval module, and an answer generation and formatted output module.

[0167] The beneficial technical effects of the present invention are:

[0168] (1) The present invention provides a multi-round knowledge-guided question-answering method that integrates large language models and knowledge graphs, aiming to solve the ambiguity problems in complex policy and regulation question-answering and improve the accuracy of user demand parsing and the comprehensiveness of answers. This method takes the knowledge graph as the core and provides multi-path semantic parsing and question-guiding capabilities by constructing a hierarchical structure covering first-level matters, second-level matters, scenario categories, and keyword nodes. Combining the powerful semantic understanding and generation capabilities of pre-trained large language models, the system can dynamically track the context information in the user input, iteratively access relevant nodes along different paths of the knowledge graph, and gradually clarify the user's intention.

[0169] (2) The present invention provides a multi-round knowledge-guided question-answering system that integrates large language models and knowledge graphs. The user input is matched with the knowledge graph through a deep semantic parsing module, triggering a dynamic retrieval or intention clarification process; the label extraction module extracts key semantic information to construct multi-dimensional question labels; the clarification module combines the knowledge graph path and the pre-trained large language model to generate intelligent follow-up questions, gradually guiding the user to provide more explicit input; in multi-round conversations, it dynamically combines the context with relevant knowledge nodes to generate accurate and semantically logical answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0170] Figure 1 is the system flow chart;

[0171] Figure 2 is the knowledge graph storage structure;

[0172] Figure 3 is the instance diagram of the accurate recognition effect of popularized expression (Application Example 1);

[0173] Figure 4 is the instance diagram of the guiding follow-up question effect for ambiguous intentions (Application Example 2);

[0174] Figure 5 is the instance diagram of the actual application effect in the production environment (Application Example 3);

[0175] Figure 6 is the instance diagram of the actual application effect in the production environment (Application Example 4);

[0176] Figure 7 is the instance diagram of the actual application effect in the production environment (Application Example 5). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0177] The technical solution of the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0178] Embodiment

[0179] Referring to the attached Figures 1 - 2 , a multi-round knowledge-guided question-answering method that integrates large language models and knowledge graphs includes the following steps:

[0180] Step 1, construct a knowledge graph: perform semantic parsing and hierarchical processing on the input document content, extract and label the first-level matters, second-level matters, scenario categories, relevant keywords, and questions and answers of the corresponding knowledge; through the entity alignment and relationship extraction module, construct the extracted content into a knowledge graph with a hierarchical structure; the knowledge graph with a hierarchical structure includes the knowledge label matter path, knowledge question, knowledge question vector, knowledge answer, and knowledge answer vector;

[0181] The specific process of entity alignment and relationship extraction is as follows:

[0182] S11. Perform text cleaning and context semantic enhancement on the input document content D = {d 1 , d 2 ,..., d n}, and convert it into a normalized text representation; based on the sentence embedding model E s , generate the vectorized representation v i = E s (d i ) for subsequent entity and relationship analysis; for the document question and answer part, use the extraction model K(z) to extract the sets K Q and K A to ensure the extraction accuracy of Q&A knowledge;

[0183] S12. Adopt a relationship extraction method based on a pre-trained large language model to construct entity and relationship Prompt templates; analyze the document context, and adopt a combination of semantic parsing and nested hierarchy extraction to construct a hierarchical knowledge graph, including:

[0184] S121. First-level matter node T 1 : Extract the main matters or topics involved in the document, represented as a node set These nodes usually represent high-level policy categories or domains;

[0185] S122. Second-level matter node T 2 : Further refine from the context of the first-level matter, extract the second-level matter, and form a set of sub-nodes associated with the first-level matter to more accurately represent the refined content;

[0186] S123. Scenario category node S: According to the description and application background of the document, extract the node set S = {s 1 , s 2 ,..., s k} related to the scenario to represent the usage scenario of the knowledge;

[0187] S124. Keyword Node K: Extract the core keywords K = {k 1 , k 2 ,..., k p} in the document, and associate them with relevant first-level matters, second-level matters, and scenario category nodes;

[0188] The specific process of constructing the knowledge graph is as follows:

[0189] 1. Use a pre-trained large language model to deeply analyze the semantic meaning at the sentence level of the document and generate a context vector representation Combined with a syntactic analysis model extract hierarchical entity nodes and their relationships;

[0190] 2. Design a recursive relationship extraction function RelExtract(·) to associate the first-level node T 1 with the keywords and second-level nodes in its context:

[0191]

[0192] where Rij represents the relationship type between the first-level matter and the second-level matter;

[0193] 3. The extraction of the scenario category node S is through context nesting analysis, combined with the association probability P(s|t i , k j ), to generate the scenario branch structure of the hierarchical graph;

[0194] Finally, all nodes and relationships will be organized into a hierarchical knowledge graph in the following form:

[0195] G = (N, E), N = T 1 ∪T 2 ∪S∪K, E = {(n i , n j , R ij )|n i , n j ∈N}, (2)

[0196] where N is the set of nodes, E is the set of relationships between nodes, and R ij represents the relationship between any two nodes;

[0197] The construction process can dynamically generate hierarchical nodes and relationships according to the document content, ensuring the accuracy and scalability of the knowledge graph, and providing structured semantic support for subsequent retrieval and clarification (step 2);

[0198] Step 2, Retrieval and Clarification: Based on the user input, perform semantic parsing through a pre-trained large language model, and execute hybrid retrieval in combination with the knowledge graph and knowledge base constructed in Step 1. If the retrieval result is unclear, trigger the clarification process; the clarification process includes first-level label item clarification, second-level label item clarification, scenario-guided clarification, and rejection process;

[0199] Among them, the specific process of semantic parsing is as follows:

[0200] S211, Vectorization of User Input:

[0201] For the question Q input by the user u , generate an embedding vector using a pre-trained large language model Generated through the following formula:

[0202]

[0203] Among them, Embedding(·, θ) represents the semantic embedding generation function, and θ is the model parameter;

[0204] If the user input involves multi-round context, combine the current input with the historical conversation content and generate a context semantic vector through splicing

[0205]

[0206] Among them, H t-1 represents the previous round of conversation history;

[0207] S212, Using the semantic structure information of the knowledge graph G=(N, E) constructed in Step 12, compare the user input with the node embedding vectors in the constructed knowledge graph G for preliminary matching, and calculate the similarity of each node:

[0208]

[0209] According to the similarity threshold τ 1 , extract the set of highly relevant nodes including first-level matters T 1 , second-level matters T 2 , scenario categories C s and keywords K;

[0210] Preliminarily determine the knowledge nodes related to the user input and their association relationships, providing a basis for the subsequent retrieval module;

[0211] Among them, the specific process of retrieval is as follows:

[0212] S221, High-confidence Quick Matching:

[0213] If among the set of nodes N matched by the knowledge graph, nodes with similar vectors are matched through k-NN retrieval of Elasticsearch, and the similarity Sim of the key question in a certain node is ≥ 0.98 or the re-rank correlation Rerank of the answer is ≥ 0.95, directly call the relevant document as the preliminary answer and enter the formatting module, skipping the clarification process;

[0214] S222. Vector Retrieval Extension:

[0215] If high-confidence nodes are not hit, use the knowledge graph and the embedding vector of the user question in step 1 to construct a retrieval condition:

[0216] Q r = ConstructQuery(T 1 , T 2 , C s , K), (6)

[0217] where Q r is a retrieval condition dynamically generated in combination with knowledge graph nodes, used to call external vector retrieval (such as k-NN retrieval of Elasticsearch);

[0218] Retrieve the candidate document set D, and score the relevance of the candidate documents through the re-rank module, filtering out low-relevance documents:

[0219] D filtered = {d|RelScore(d) ≥ τ 2 , d ∈ D}, (7)

[0220] where D filtered is the retrieved knowledge document; d is a single knowledge document, which is an element of the original knowledge document set D filtered ; RelScore is the document relevance score, indicating the matching degree between d and the user query, usually calculated by a certain vector similarity, such as cosine similarity or dot product; τ 2 is the set similarity or relevance threshold;

[0221] S223. Prompt Construction:

[0222] Use the information of the nodes already matched in the knowledge graph (such as the first-level matters, scenario categories, etc.) and the content of D filtered to construct a Prompt and call the pre-trained large language model to generate a preliminary answer A raw ;

[0223] Among them, the specific process of triggering the clarification process is:

[0224] S231. Trigger condition:

[0225] If D filtered is empty, or the uncertainty of the preliminary answer A raw (according to the confidence of the generation model) is higher than the threshold α, trigger the clarification process;

[0226] S232. Clarification question generation:

[0227] Generate clarification questions based on the knowledge graph nodes and relationships in step 1:

[0228] 1. If the primary matter T 1 is not clear, first retrieve all primary label matters and the current knowledge graph obtained according to facts, and generate clarification questions for T 1 through the pre-trained large language model;

[0229] 2. If the secondary matter T 2 is not clear, first retrieve all secondary label matters and the current knowledge graph obtained according to facts, and generate clarification questions for T 2 through the pre-trained large language model;

[0230] 3. If the keyword set K is incomplete, generate clarification questions for the keywords;

[0231] S233. User interaction and update:

[0232] Send the clarification questions to the user, receive the user's feedback, and update the current session context Context and the knowledge graph node set;

[0233] S234. Re-retrieval after clarification:

[0234] Re-execute the retrieval process according to the user's feedback until all key nodes are clear or the rejection process is triggered;

[0235] Step 3. Label extraction and question classification: Extract the primary matter, secondary matter, scenario category, and keywords from the user input in step 2 through the label extraction module, and classify and process the questions according to the extraction results; Classifying and processing the questions according to the extraction results means judging which type of elements are missing according to the slot filling extracted by the user, and performing a follow-up question process according to the specific missing elements;

[0236] Among them, the specific process of label extraction is:

[0237] S311. Input data:

[0238] Input the question Q u input by the user, the retrieved knowledge document D filtered , the clarified context information C u and the knowledge graph G;

[0239] S312. Label extraction process:

[0240] Use the label extraction method based on the large model to generate the label extraction Prompt:

[0241] P extract = ConstructPrompt(Q u , C u , D filtered , G), (8)

[0242] where P extract is the final extracted Prompt, which represents the specific task description generated by combining the user query and context information and is used to guide the pre-trained large model to generate answers;

[0243] The Prompt contains: user questions, context, retrieval results, and knowledge graph node information;

[0244] Extract the following information from the Prompt through the pre-trained large language model:

[0245] Primary matter T 1 : The set of primary node labels most relevant to the user question;

[0246] Secondary matter T 2 : The set of specific sub-node labels further restricted based on T 1 ;

[0247] Scenario category C s : The classification label describing the semantic background of the question;

[0248] Keywords K: The set of words highly relevant to the user question;

[0249] S313. Label identification generation:

[0250] Package the extraction result into the label structure Tags:

[0251] Tags = {T 1 , T 2 , C s , K}, (9)

[0252] S314. Multi-factor authentication:

[0253] Use the knowledge graph G to verify the extracted labels:

[0254] If T 1 or T 2 does not exist in the knowledge graph G, trigger the label completion process (return to the clarification module in step 2);

[0255] If the extracted tags are ambiguous: confidence Conf < τ 3 , generate a clarification prompt to obtain more explicit information;

[0256] Among them, the specific process of classifying and processing problems is as follows:

[0257] S321. Classification rules:

[0258] According to the extracted tag information Tags, divide the user's question into the following categories:

[0259] 1. Direct answer type: The tags are completely and clearly extracted, and the answer in the knowledge graph or retrieved document can be directly matched;

[0260] 2. Clarification requirement type: The tag extraction is incomplete (such as missing the first-level or second-level matters), and further clarification is required;

[0261] 3. Dynamic retrieval type: The tags are clear, but refined retrieval needs to be combined with the scenario category and keywords;

[0262] S322. Classification discrimination formula:

[0263] Use the classification model or rules to judge:

[0264]

[0265] Among them, Complete(Tags) indicates whether the tags are complete; It is when the scenario category C s or the keyword K is an empty set, that is, when no tags are extracted;

[0266] S323. Classification output:

[0267] According to the classification result, adjust the subsequent process:

[0268] 1. For the direct answer type, enter the Retrieval-Augmented Generation (RAG) module to generate an answer;

[0269] 2. For the clarification requirement type, return to step 2 and enter the clarification Q&A process to supplement the missing information;

[0270] 3. For the dynamic retrieval type, the Tags structure (including the first-level matter, second-level matter, scenario category, and keyword) is used as the input for scenario guidance in step 4, and perform scenario guidance and dynamic retrieval;

[0271] Step 4, Scenario Guidance and Dynamic Retrieval: Based on the knowledge graph constructed in Step 1 and the tag information extracted in Step 3, determine the details of the scenario category, dynamically generate retrieval conditions, and perform further precise queries on the knowledge base; Dynamically generating retrieval conditions is to dynamically retrieve the corresponding reference clarification words according to the missing scenario guidance type and ask a rhetorical question to precisely extract the specific scenario guidance details of the user.

[0272] S41. Initialization and Dynamic Expansion of Scenario Conditions

[0273] S411. Input Initialization:

[0274] Obtain the tag T extracted in Step 3 1 , T 2 , C s , K and the preliminary retrieval result D filtered ;

[0275] Load the nodes and context information related to the current session from the knowledge graph G;

[0276] S412. Condition Expansion:

[0277] If some tags are missing, such as T 2 or C s , then dynamically call the knowledge graph for condition expansion:

[0278] Find the possible tag sets associated with T 1 or K:

[0279]

[0280] Include in the scenario derivation process;

[0281] S413. Generate Dynamic Retrieval Conditions:

[0282] Initial Conditions:

[0283]

[0284] Condition Optimization:

[0285] Use the retrieval history and feedback to adjust the conditions, such as increasing the priority of specific tags:

[0286] Q optimized = Q init + {"boost": TagsPriority(T 1 , T 2 )}, (13)

[0287] S42. Dynamic Retrieval Execution and Feedback Iteration

[0288] S421. Preliminary Retrieval:

[0289] Execute a retrieval based on Q optimized and the Elasticsearch configuration:

[0290] D scene = Elasticsearch(Q optimized ), (14)

[0291] Sort the returned results by the relevance score and filter D filtered :

[0292] D filtered = TopN(D scene , N = 10, score > 0.8), (15)

[0293] S422. Feedback Iteration:

[0294] Generate a rhetorical Prompt for the pending tags:

[0295]

[0296] S423. Dynamically Supplement Pending Tags:

[0297] For the undefined C s , search for relevant expressions in the knowledge graph G;

[0298]

[0299] Use the expressions to clarify the scenario category until C s is clarified or the user aborts the session;

[0300] S43. Scenario Refinement and Tag Confirmation

[0301] S431. Tag Confirmation Logic:

[0302] If T 1 , T 2 and K are clear, but C s is pending, use the following logic to refine;

[0303] 1. If the number of candidates for C s is 1, directly confirm it as the final tag:

[0304]

[0305] 2. If the number of candidates is greater than 1, dynamically generate a clarification option Prompt:

[0306]

[0307] Prompt the user to select;

[0308] S432. Tag priority sorting:

[0309] Sort the tags by priority based on context information and user feedback to guide the subsequent process;

[0310] S44. Confirm the output and multi-round optimization

[0311] S441. Result output:

[0312] Output the final determined tag set T 1 ,T 2 , K and the exact retrieval result D final ;

[0313] Ensure that D final is filtered and the content is highly relevant to the current context;

[0314] S442. Interface with subsequent modules:

[0315] Provide for the answer generation in Step 5:

[0316] 1. Complete scenario categories

[0317] 2. High-quality knowledge document set D final ;

[0318] 3. Pending tags as feedback, return to Step 3 to form a closed loop;

[0319] Step 5. Answer generation and formatted output: Combine the context information obtained in Steps 1-4, and through the Retrieval-Augmented Generation (RAG) module, use the pre-trained large language model to generate answers and optimize the answer output through the formatting module;

[0320] S51. Answer generation

[0321] S511. Context integration;

[0322] According to the final scenario categories, primary matters, secondary matters, and keyword information obtained in Step 4, combine the user's historical input and the current question to construct the context information C user ;

[0323] S512. Knowledge retrieval integration:

[0324] Use the knowledge document K retrieved in Step 4 retrieved as external knowledge support, and together with the context information C user input into the retrieval augmented generation module to construct the generative Prompt P RAG :

[0325] P RAG = f(C user , K retrieved ), (20)

[0326] Generate the preliminary answer A through a pre-trained large language model raw ;

[0327] The pre-trained large language model uses Qwen 2.5 with 72B parameters;

[0328] S52. Answer optimization

[0329] S521. Consistency check:

[0330] Compare the generated preliminary answer A raw with the paths and nodes in the knowledge graph to ensure that the answer logic is consistent with the knowledge base; for parts that deviate from the knowledge graph, trigger supplementary retrieval and iteratively optimize the generated answer;

[0331] S522. Formatting:

[0332] Input the preliminary answer A raw into the formatting module to optimize its language fluency, structural clarity, and expression of professional terms, and output the final answer A final ;

[0333] S53. Answer output

[0334] The system returns the final answer A final to the user, along with interactive prompt information, allowing the user to initiate further clarification or modification requests for the generated answer, thus realizing a cyclic knowledge Q&A.

[0335] Application Example 1

[0336] In complex policy document queries, the user expresses their needs in popular language (such as the unemployment benefit issue in the social security policy). Traditional systems are difficult to directly understand such natural language, while the present invention combines a knowledge graph and a large language model to accurately identify the user's intention.

[0337] Figure 3 It is an example diagram of the accurate identification effect of popular expression, showing that after the user inputs a popular question, the system can obtain the user's intention from the user's question without a strict format, and how the system uses the knowledge graph and semantic parsing module to map the user's natural language to specific first-level matters, second-level matters, and keyword nodes. Screen documents through formula (4) (semantic similarity calculation), optimize the retrieval results in combination with formula (7), and finally generate an accurate answer.

[0338] Actual effect:

[0339] When the user asks the question "I came here from another province to work. I was recently fired by the company. Can I get any unemployment benefits?", the system successfully maps the question in plain language to the node of "Unemployment Insurance - Application for Unemployment Insurance Benefits - Non-local Household Registration - Unemployed Persons". When the user's plain-language question information is relatively comprehensive, the first-level matter label and the second-level matter label can already be extracted, but there is still one condition lacking in the scenario guidance. At this time, the system will ask a clarifying question: "May I ask whether you participated in unemployment insurance as an enterprise employee or as a flexible employee before?" At this time, the system determines that the necessary nodes for scenario guidance have been completed, and a clear policy interpretation is generated.

[0340] Application Example 2

[0341] When the user's question intention is not clear (such as "pension insurance problem"), the system needs to guide the user to clarify specific requirements (such as business type, handling conditions). Traditional retrieval methods are prone to failure in this scenario, while the present invention accurately clarifies the user's intention through a dynamically generated rhetorical question mechanism.

[0342] Figure 4 The figure shows an example of the rhetorical question effect for guiding fuzzy intentions. It shows that after the user inputs a fuzzy question, the system triggers the clarification of the first-level matter (such as "To better assist you, may I ask whether you want to know about the handling process of pension insurance, unemployment insurance, or work-related injury insurance?") based on the knowledge graph and the clarification module. By constructing dynamic retrieval conditions through formulas (5) and (8), the user's requirements are gradually guided to be clear.

[0343] Actual Effect:

[0344] The user initially inputs "I want to handle social insurance." In the first round of conversation, the user's question information is relatively fuzzy. The first-level label extracted is social insurance, but the second-level label and the scenario guidance are both unknown. In the second round of conversation, the user clarifies the first-level matter label again according to the rhetorical question. After two rounds of rhetorical questions from the system, it is finally clarified as "pension insurance". In the third round of conversation, the user continues to clarify the second-level matter label according to the rhetorical question. At this time, the system can map the question to the node of "Pension Insurance - Relationship Transfer", but there is still a lack of conditions for scenario guidance, so the system continues to ask the user for details about the scenario guidance. Finally, after the user clarifies the scenario type, the system generates an accurate answer in the retrieval enhancement generation module.

[0345] Application Example 3

[0346] Figure 5This is an example diagram of the actual application effect in the production environment. It shows a real system dialogue scenario in which users use colloquial language to express their needs (such as unemployment benefits in social security policies) in complex policy document queries. In the figure, each round of dialogue can explain the workflow process of the system logic, and finally generates clarification questions according to step (S232) for counter-asking.

[0347] Actual effect:

[0348] After the user inputs a popular question, the system can obtain the user's intention through the user's question without strict format. The system visualizes each step of the decision-making and judgment process, and generates a clarification question based on the missing conditions, "Did you participate in unemployment insurance as an enterprise employee or as a flexible employment person before?", which has the effect of asking the user a question back.

[0349] Application Example 4

[0350] Figure 6 This is an example diagram of the actual application effect in the production environment. It shows that in a complex policy document query, the user uses popular language to express his needs (such as the unemployment benefit issue in the social security policy). After accurately clarifying the user's intention, the answer is generated through formula (20) (retrieval enhancement generation module), and the result of accurately answering the user's intention is obtained through step (S522) (answer formatting).

[0351] Actual effect:

[0352] After the user inputs a popular question, the system can obtain the user's intention through the user's unformatted questions. After multiple rounds of clarification and questioning, the knowledge graph of the user's intention is fully constructed, and the answer is generated through the retrieval enhancement generation module. The answer is formatted and summarized according to the point-by-point logic to answer the user's intention in an organized and standardized manner.

[0353] Application Example 5

[0354] When the user's question intention has been clarified and the answer has been accurately replied, if the user needs to modify a certain scenario guidance condition or a certain knowledge link, the system can dynamically modify the link based on the historical messages of the context, allowing the user to perceive the result of the ongoing conversation instead of repeating the clarification. Traditional conversation retrieval methods are prone to fail in this scenario, while the present invention can dynamically adjust the graph nodes through dynamically generated knowledge graph maintenance and revise the answers to subsequent questions.

[0355] Figure 7This is an example diagram of the actual application effect in the production environment, showing the scenario where the user continues to ask questions after the system has accurately answered the question after multiple rounds of clarification. When the user asks about other nodes in the knowledge graph again, the system will dynamically adjust the secondary item tags and reply with answers for the neighbor knowledge links. When the user continues to ask in the second round and changes or adds the conditions guided by the scenario, the system can also automatically switch answers according to the condition change and in combination with the context.

[0356] Actual effect:

[0357] Based on the system's final intention answer to the previous round ".. procedures for applying for unemployment insurance benefits...", the user inputs again "I want to receive it all at once. What materials do I need to apply?" In the first round of conversation, the secondary label extracted from the citizen's question information is modified to the application for one-time unemployment insurance benefits. The system continues to judge that the knowledge graph construction is completed and conducts a precise search and reply: "If you are an enterprise employee with a non-local household registration and wish to receive unemployment insurance benefits all at once, you need to prepare the following materials:......".

[0358] In the second round of conversation, the citizen adds and adjusts the scenario guidance node of the knowledge graph again: "Online application". Finally, after the system judges that the citizen has clearly defined the scenario type, the system generates an accurate answer in the retrieval enhancement generation module according to the context information: "Of course. You can apply for one-time receipt of unemployment insurance benefits online through your mobile phone. The following are the specific online application channels:......".

[0359] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A multi-round knowledge-guided question-answering method that integrates a large language model and a knowledge graph, characterized in that include: S1. Constructing a knowledge graph: semantically parse and hierarchically process the input document content, extract and annotate the primary items, secondary items, scenario categories, relevant keywords, and corresponding knowledge questions and answers in the document; construct the extracted content into a knowledge graph with a hierarchical structure through the entity alignment and relationship extraction modules; S2, retrieval and clarification: Based on user input, semantic analysis is performed through a pre-trained large language model, and hybrid retrieval is performed in combination with the knowledge graph and knowledge base constructed in step S1. If the retrieval result is unclear, the clarification process is triggered; S3, label extraction and question classification: extracting primary matters, secondary matters, scenario categories and keywords from the user input in step S2 through the label extraction module, and classifying and processing questions according to the extraction results; S4, scenario guidance and dynamic retrieval: Based on the knowledge graph constructed in step S1 and the tag information extracted in step S3, the details of the scenario category are determined, the retrieval conditions are dynamically generated, and the knowledge base is further accurately queried; S5, answer generation and formatted output: Combined with the context information obtained in steps S1 to S4, the answer is generated using the pre-trained large language model through the retrieval enhancement generation module, and the answer output is optimized through the formatting module.

2. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S1, the specific process of entity alignment and relationship extraction is as follows: S11, input document content D = {d1, d2, ..., d n }Perform text cleaning and contextual semantic enhancement to convert it into a standardized text representation; Based on the sentence embedding model E s , generate a vectorized representation v for each sentence in the document i =E s (d i ), which is used for subsequent entity and relationship analysis; for the document question and answer part, the extraction model K(z) is used to extract the set K Q and K A , ensure the accuracy of question and answer knowledge extraction; S12, using the relation extraction method based on the pre-trained large language model to build entity and relation prompt templates; Analyze the document context and build a hierarchical knowledge graph by combining semantic parsing with nested hierarchical extraction, which includes: S121, first-level item node T1: extract the main items or topics involved in the document and represent them as a node set These nodes typically represent high-level policy categories or areas; S122, Secondary matter node T2: further refine the context of the primary matter, extract the secondary matter, and form a set of child nodes associated with the primary matter To express detailed content more accurately; S123, scene category node S: According to the description of the document and the application background, extract the scene-related node set S = {s1, s2, ..., s k }, used to represent the usage scenarios of knowledge; S124, keyword node K: extract the core keywords K in the document = {k1, k2, ..., k p } and associate it with the relevant first-level matters, second-level matters and scenario category nodes.

3. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S1, the hierarchical knowledge graph includes knowledge label item paths, knowledge questions, knowledge question vectors, knowledge answers and knowledge answer vectors.

4. A multi-round knowledge guided question answering method integrating a large language model and a knowledge graph according to claim 1 or 3, characterized in that: The specific process of building a knowledge graph is as follows: 1) Using pre-trained large language models Deeply analyze the semantics of the document sentence level and generate context vector representation Combined with syntactic analysis model Extract hierarchical entity nodes and their relationships; 2) Design a recursive relation extraction function RelExtract(·) to associate the first-level node T1 with the keywords and second-level nodes in its context: Among them, Rij represents the relationship type between the first-level matters and the second-level matters; 3) The scene category node S is extracted through context nesting analysis and combined with the association probability P(s|t i ,k j ), generate the scene branch structure of the hierarchical graph; Ultimately, all nodes and relationships will be organized into a hierarchical knowledge graph in the following form: G=(N,E),N=T1∪T2∪S∪K,E={(n i ,n j ,R ij )∣n i ,n j ∈N}, (2) Among them, N is the node set, E is the relationship set between nodes, R ij Represents the relationship between any two nodes; The construction process can dynamically generate hierarchical nodes and relationships based on the document content, ensuring the accuracy and extensibility of the knowledge graph and providing structured semantic support for subsequent retrieval and clarification.

5. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S2, the specific process of semantic analysis is as follows: S211, User input vectorization: Questions Q u , using a pre-trained large language model to generate embedding vectors Generated by the following formula: Where, Embedding(·,θ) represents the semantic embedding generation function, and θ is the model parameter; If the user input involves multiple rounds of context, the current input is combined with the historical conversation content to generate a context semantic vector by splicing Among them, H t-1 Represents the previous round of dialogue history; S212, using the semantic structure information of the knowledge graph G = (N, E) constructed in step S12, the user input And the node embedding vector in the constructed knowledge graph G Perform preliminary matching and calculate the similarity of each node: According to the similarity threshold τ1, extract the set of highly correlated nodes Including first-level matters T1, second-level matters T2, and scenario category C s and keyword k; preliminarily determine the knowledge nodes and their associations related to the user input, providing a basis for subsequent retrieval modules; In step S2, the specific process of retrieval is: S221, high confidence fast matching: If, in the node set N matched by the knowledge graph, nodes matching similar vectors are retrieved through Elasticsearch's k-NN, and the similarity Sim of the key question in a certain node is ≥ 0.98 or the reranking relevance Perank of the answer is ≥ 0.95, the relevant document is directly called as the preliminary answer and enters the formatting module, skipping the clarification process; S222, vector search extension: If the high confidence node is not hit, use the knowledge graph in step S1 and the embedding vector of the user question Construct search conditions: Q r =ConstructQuery(T1,T2,C s ,K), (6) Among them, Q r It is a search condition dynamically generated by combining knowledge graph nodes and is used to call external vector search (such as Elasticsearch's k-NN search); Retrieve the candidate document set D, and score the relevance of the candidate documents through the re-ranking module to filter out low-relevance documents: D filtered ={d|RelScore(d)≥τ2,d∈D}, (7) Among them, D filtered is the retrieved knowledge document; d is a single knowledge document, belonging to the original knowledge document set D filtered An element in ; RelScore is the document relevance score, which indicates the degree of match between d and the user query, usually calculated by some vector similarity, such as cosine similarity or dot product; τ2 is the threshold for setting similarity or relevance; S223, Prompt build: Use the matched node information in the knowledge graph (such as first-level matters, scene categories, etc.) and D filtered , construct Prompt, and call the pre-trained large language model to generate a preliminary answer A raw ; In step S2, the specific process of triggering the clarification process is: S231, trigger conditions: If D filtered Empty, or the initial answer is A raw The uncertainty (according to the confidence of the generative model) is higher than the threshold α, triggering the clarification process; S232, clarification question generation: Generate clarification questions based on the knowledge graph nodes and relationships in step S1: 1) If the first-level item T1 is not clear, first retrieve all the first-level label items and the knowledge graph currently obtained based on the facts, and generate clarification questions for T1 through the pre-trained large language model; 2) If the secondary item T2 is not clear, first retrieve all secondary label items and the knowledge graph currently obtained based on the facts, and generate clarification questions for T2 through the pre-trained large language model; 3) If the keyword set K is incomplete, generate clarification questions for the keywords; S233, User Interaction and Update: Send clarifying questions to the user, receive user feedback, and update the current session context and knowledge graph node set; S234, Re-search after clarification: Re-execute the search process based on user feedback until all key nodes are clarified or the rejection process is triggered.

6. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S3, the specific process of label extraction is: S311. Input data: Enter the question Q entered by the user u , retrieved knowledge documents D filtered , clarified context information C u and knowledge graph G; S312, label extraction process: Use the large model-based tag extraction method to generate a tag extraction prompt: P extract =ConstructPrompt(Q u ,C u ,D filtered ,G), (8) Among them, P extract is the final Prompt extracted, which is represented by a specific task description generated by combining user query and context information, and is used to guide the pre-trained large model to generate answers; Prompt contains: user question, context, search results and knowledge graph node information; The following information is extracted from Prompt using a pre-trained large language model: First-level items T1: the first-level node label set most relevant to the user's question; Secondary matter T2: a specific sub-node label set further defined on the basis of T1; Scenario Category C s : Classification label describing the semantic context of the question; Keyword K: A set of words that are highly relevant to the user's question; S313, label identification generation: Encapsulate the extraction results into a tag structure Tags: Tags={T1,T2,C s ,K}, (9) S314, Multiple Authentication: Use the knowledge graph G to verify the extracted tags: If T1 or T2 does not exist in the knowledge graph g, the tag completion process is triggered (return to the clarification module in step S2); If the extracted label is ambiguous: confidence Conf < τ3, a clarification hint is generated to obtain more explicit information.

7. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S3, the specific process of classifying and processing the problem is as follows: S321. Classification rules: Based on the extracted tag information Tags, user questions are divided into the following categories: 1) Direct answer type: The label extraction is complete and clear, and the answer in the knowledge graph or retrieval document can be directly matched; 2) Clarification Requirement: Label extraction is incomplete (e.g., missing primary or secondary items), and further clarification is required; 3) Dynamic retrieval type: The labels are clear, but the search needs to be refined by combining scene categories and keywords; S322, classification and discrimination formula: Use classification models or rules to determine: Among them, Complete(Tags) indicates whether the tags are complete; When the scene category C s Or whether the keyword K is an empty set, that is, no label is extracted; S323, classification output: Adjust the subsequent process according to the classification results: 1) For direct answer type, enter the Retrieval Augmentation Generation (RAG) module to generate answers; 2) For clarification needs, return to step S2 and enter the clarification question and answer process to supplement the missing information; 3) For dynamic retrieval type, the Tags structure (including primary items, secondary items, scene categories and keywords) is used as the input of the scene guidance in step S4 to perform scene guidance and dynamic retrieval. In step S4, the search condition is dynamically generated based on the lack of scene guidance type, and the corresponding reference clarification words are dynamically retrieved based on the scene type to ask questions in order to accurately extract the user's specific scene guidance details.

8. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S4, the specific process is: S41. Scene condition initialization and dynamic expansion S411, input initialization: Get the tags T1, T2, C extracted in step S3 s , K and preliminary search results D filtered ; Load nodes and context information related to the current session from the knowledge graph G; S412, Conditional Extension: If some tags are missing, such as T2 or C s , then dynamically call the knowledge graph for conditional expansion: Find the possible set of labels associated with T1 or K: Will Incorporate into the scenario derivation process; S413, generating dynamic search conditions: Initial conditions: Condition optimization: Use search history and feedback to refine your criteria, such as increasing the priority of specific tags: Q optimized =Q init +{"boost":TagsPriority(T1,T2)}, (13)S42, dynamic retrieval execution and feedback iteration S421, Preliminary Search: Based on Q optimized And the Elasticsearch configuration to perform the search: D scene =Elasticsearch(Q optimized ), (14) The returned results are sorted by relevance score, and D filtered : D filtered =TopN(D scene ,N=10,score>0.8), (15) S422, Feedback Iteration: Generate a question prompt for pending tags: S423, Dynamically supplement pending tags: For unspecified C s , search for relevant words in the knowledge graph G; Use words to clarify the scene category until C s Explicit or user termination of the session; S43. Scene refinement and label confirmation S431, tag confirmation logic: If T1, T2 and K are clear, but C s Undecided, use the following logic to refine; 1) If C s The number of candidate items is 1, which is directly confirmed as the final label: 2) If the number of candidate items is greater than 1, dynamically generate a clarification option prompt: Prompt the user to make a choice; S432, tag priority sorting: Based on contextual information and user feedback, prioritize tags to guide subsequent processes; S44. Confirm output and multiple rounds of optimization S441, result output: Output the finalized label set T1, T2, K and accurate retrieval results D final ; Ensure D final After filtering, the content is highly relevant to the current context; S442, connecting with subsequent modules: The answer generation for step S5 is provided as follows: 1) Complete scene categories 2) High-quality knowledge document collection D final ; 3) The pending tags are used as feedback and return to step S3 to form a closed loop.

9. According to claim 1, a multi-round knowledge-guided question-answering method integrating a large language model and a knowledge graph is characterized in that: In step S5, the specific process is: S51. Answer generation S511, context integration; According to the final scene category, primary items, secondary items and keyword information obtained in step S4, combined with the user's historical input and current question, the context information C is constructed. user ; S512, Knowledge Retrieval Integration: The knowledge document K retrieved from step S4 retrieved As external knowledge support, and context information C user Input them into the retrieval enhancement generation module to construct the generative formula Prompt P RAG : P RAG =f(C user ,K retrieved ), (20) Generate a preliminary answer A through a pre-trained large language model raw ; The pre-trained large language model uses Qwen2.5, with 72B parameters; S52. Answer optimization S521, consistency check: Compare the generated preliminary answer A raw Ensure that the answer logic is consistent with the knowledge base and the paths and nodes in the knowledge graph; trigger supplementary retrieval and iteratively optimize the generated answers for the parts that deviate from the knowledge graph; S522, Formatting: The initial answer A raw Input the formatting module to optimize its language fluency, structural clarity and professional terminology expression, and output the final answer A final ; S53, answer output The system will give the final answer A final The answer is returned to the user with interactive prompts, allowing the user to initiate further clarification or modification requests for the generated answer, thus realizing a circular knowledge question and answer.

10. A multi-round knowledge-guided question-answering system integrating a large language model and a knowledge graph, characterized in that: It includes entity alignment and relationship extraction module, semantic parsing module, retrieval module, clarification question and answer module, label extraction module, question classification module, scenario guidance and dynamic retrieval module, answer generation and formatted output module.

Citation Information

Patent Citations

  • Multi-terminal task collaboration method and device for robot exhibition hall scene based on cloud side-end architecture

    CN116627637A

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Medical auxiliary question and answer method and system based on knowledge calibration and retrieval enhancement

    CN117573843A

  • Personalized question and answer method based on general large language model and knowledge graph

    CN119149719A

  • Reasonable language model learning for text generation from a knowledge graph

    US20230134798A1

Cited By

  • Data retrieval method, device and system

    CN120316119A

  • Government affair hotline service knowledge graph construction method and system based on large language model

    CN120354923A

  • Report interpretation method and related device

    CN120407815A

  • A report interpretation method and related apparatus

    CN120407815B

  • Modularized knowledge graph and retrieval enhanced large model fusion interaction method and system oriented to financial branch mechanism

    CN120448510A