Multi-modal man-machine question and answer guiding system based on large language model
Through a multimodal human-computer question-and-answer guidance system, user intentions are dynamically analyzed and structured follow-up questions are conducted. Combined with authoritative knowledge traceability and compliance verification, this solves the parsing difficulties faced by existing systems when faced with non-standard language, achieves high-precision and efficient question-and-answer interaction, and improves user experience and the credibility of answers.
Patent Information
- Application Number
- CN202510713601.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-09
AI Technical Summary
Existing question-answering systems driven by large language models have difficulty parsing user needs when faced with non-standard, vague, incomplete or ambiguous language. They lack the ability to actively correct and optimize, resulting in low interaction efficiency and an inability to meet the compliance and security requirements of high-precision demand scenarios.
A multimodal human-machine question-answering guidance system is adopted to achieve accurate understanding of user intentions and compliant answers through dynamic intent analysis, structured question generation, authoritative knowledge tracing and compliance verification, combined with reinforcement learning to optimize interaction.
It improves the accuracy and explainability of the question-answering system, reduces the number of interaction rounds, and enhances the user experience and the authority and security of the answers.
Smart Images

Figure CN120611025A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence natural language processing, and specifically relates to a dynamic guidance system based on a large language model (LLM). Through a five-layer architecture of multi-dimensional intent analysis - structured questioning - knowledge tracing - compliance verification - interaction optimization, it solves core problems such as irregular user questions, system comprehension deviations, unreliable answers, and low interaction efficiency. It is suitable for high-precision demand scenarios such as intelligent customer service, medical diagnosis, legal consultation, and educational guidance, and improves the accuracy, explainability, and user experience of the question-and-answer system. Background Art
[0002] Existing large language model-driven question-answering systems (such as ChatGPT, DeepSeek, and Wenxinyiyan) have the following technical flaws:
[0003] User questions are not standardized: colloquial, vague, incomplete or ambiguous language, and unstructured expressions (such as "What should I do if that thing breaks?") make it difficult for the system to parse requirements and lack the ability to proactively correct and optimize.
[0004] Static guidance mechanism: Existing systems rely on preset templates or keyword matching and cannot adapt to the dynamic questioning needs of complex scenarios (such as intent drift in multi-round conversations).
[0005] Inefficient guidance: Users need to revise their questions multiple times to get accurate answers, and the system lacks the ability to proactively correct errors and trace knowledge.
[0006] Lack of knowledge traceability: The answers lack authoritative knowledge sources and compliance verification, making it difficult to meet the rigorous requirements of medical, legal and other fields (for example, the source of the citation is not marked in the answer).
[0007] Security risks: It is easy to generate false information or sensitive content, and there is a lack of controllable content filtering mechanism. Summary of the Invention
[0008] In view of the deficiencies of the existing technology, the present invention provides a multimodal human-computer question-answering guidance system and a multi-dimensional intention optimization method to make up for the shortcomings of the existing technology.
[0009] The core innovation of this invention is:
[0010] Dynamic intent parsing and entity filling: Combines explicit intent (such as query, confirmation), implicit intent (such as emotional tendency) and entity slot (such as time, place, name of person or product) analysis to achieve accurate understanding of complex questions.
[0011] Structured follow-up question generation: Generates clarification questions (such as "Which product do you mean?") and normative questions (such as "Do you need to add a contract number?") based on the ambiguity of intent and the entity missing rate, guiding users to supplement structured information.
[0012] Authoritative knowledge tracing and compliance verification: Real-time access to authoritative databases (such as medical guidelines, laws and regulations), matching answer basis through vector retrieval, and compliance verification (such as sensitive word filtering and logical consistency checking) based on rule engines and model fine-tuning.
[0013] Multimodal interaction optimization: Through reinforcement learning (PPO algorithm) and user feedback mechanism, the follow-up strategy is dynamically adjusted to reduce the number of interaction rounds.
[0014] The core architecture of the present invention includes: [user multimodal input layer] → [dynamic intent analysis engine] → [structured question generation module] → [authoritative knowledge traceability and compliance verification layer] → [answer generation layer] → [multimodal interaction optimization engine].
[0015] User multimodal input layer: supports text, voice, and structured form input, and achieves semantic alignment and feature fusion through cross-modal encoders (such as ViT+T5).
[0016] Dynamic intent parsing engine: includes explicit intent and entity extraction, implicit intent mining, fuzziness and missing rate calculation.
[0017] Among them, explicit intent and entity extraction: identify intent type (such as query, confirmation) and entity slot (such as time, place) based on the BERT-CRF model.
[0018] Among them, implicit intent mining: analyzing context dependencies through graph neural networks (GNN) to mine emotional tendencies (such as anxiety, satisfaction) and potential needs (such as price comparison).
[0019] Among them, ambiguity and missing rate calculation: the clarity of the question is evaluated by entropy calculation (such as the intention entropy value > 0.8 is judged as ambiguity), and the missing rate is calculated by entity coverage (such as missing rate > 30% triggers follow-up questions).
[0020] Structured follow-up question generation module: includes generation of clarifying questions, recommendation of normative questions, and adjustment of follow-up question strategies.
[0021] Among them, clarification question generation: generating questions that guide users to provide additional information based on the ambiguity of intent (such as "Which time period do you need to query data for?").
[0022] Among them, normative questions are recommended: combining the domain knowledge graph to recommend structured questions that need to be supplemented (such as "Is it necessary to supplement the contract number to accurately match the terms?").
[0023] Among them, the follow-up strategy adjustment: optimize the generation order and content of follow-up questions through reinforcement learning (such as giving priority to asking questions about entities with high missing rates).
[0024] Authoritative knowledge traceability and compliance verification layer: including authoritative database access, vector retrieval and matching, and compliance verification.
[0025] Among them, authoritative database access: real-time connection to authoritative data sources such as medical guidelines, laws and regulations, and support dynamic updates.
[0026] Among them, vector retrieval and matching: match the answer basis through dense vector retrieval (such as FAISS) to generate answer content containing reference links (such as Article X of the "XX Law").
[0027] Among them, compliance verification: verify the compliance of the answer content through the rule engine (such as regular expression filtering sensitive words) and model fine-tuning (such as RoBERTa-based logical verification).
[0028] Multimodal interaction optimization engine: includes reinforcement learning sub-module, user feedback sub-module, and model optimization sub-module.
[0029] Among them, the reinforcement learning sub-module: dynamically adjusts the questioning strategy through the PPO algorithm to maximize long-term rewards (such as the reduction in interaction rounds and the increase in user satisfaction).
[0030] Among them, the user feedback submodule collects users' satisfaction scores on answers (1-5 points) and correction behavior data (such as the number of times entities are supplemented).
[0031] Among them, the model optimization sub-module updates the parameters of the intent recognition and question generation models through contrastive learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION
[0033] The principles and features of the present application are further described in detail below with reference to the accompanying drawings. It should be understood that the specific methods and embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0034] Example 1: Medical consultation scenario
[0035] User input: Voice input "I have a headache, what should I do?"
[0036] System processing:
[0037] Speech recognition and cross-modal encoding: Convert speech to text and generate semantic vectors through the ViT+T5 model.
[0038] Dynamic intent resolution:
[0039] Explicit intent: identified as "query";
[0040] Implicit intent: GNN analyzes the context and finds the sentiment tendency as “anxiety” (confidence level 0.75);
[0041] Entity slots: Missing key information such as "location" (such as "forehead") and "duration".
[0042] Structured question generation:
[0043] Generate clarifying questions: "Is your headache in your forehead, temple, or somewhere else?";
[0044] Combined with the medical knowledge graph, the normative question is recommended: "Is it necessary to supplement the allergy history to accurately recommend drugs?"
[0045] User supplement: Enter "forehead, no history of allergies".
[0046] Authoritative knowledge source:
[0047] Match relevant entries of "forehead headache" in the "Guidelines for the Diagnosis and Treatment of Headache" through vector search;
[0048] Generate an answer: "According to the Guidelines for the Diagnosis and Treatment of Headaches, forehead headaches may be related to migraines. It is recommended to take non-steroidal anti-inflammatory drugs (such as ibuprofen) and pay attention to rest. (Quote: Article 3.2 of the Guidelines for the Diagnosis and Treatment of Headaches)".
[0049] Compliance verification:
[0050] Sensitive word filtering: No sensitive words were detected;
[0051] Logical consistency check: Verify that the drug recommendation is consistent with the knowledge graph (ibuprofen → migraine).
[0052] Interaction optimization:
[0053] User satisfaction rating: 5 points;
[0054] Reinforcement learning adjustment strategy: Increase the priority of questions related to "body part" and reduce the frequency of questions about "allergy history" (because the user has clearly stated that he has no history of allergies).
[0055] Example 2: Legal Consultation Scenario
[0056] User input: Text input "If my company fires me, how much compensation will I get?"
[0057] System processing:
[0058] Dynamic intent resolution:
[0059] Explicit Intent: Identified as “compensation calculation”;
[0060] Physical slot: Key information such as "years of work experience" and "monthly salary" is missing.
[0061] Structured question generation:
[0062] Generate clarifying questions: "How long have you worked? What is your monthly salary?";
[0063] Combined with the legal knowledge map, the normative question is recommended: "Is an open-ended employment contract signed?"
[0064] User supplement: Enter "Working for 5 years, monthly salary of 20,000, signed an open-ended contract".
[0065] Authoritative knowledge source:
[0066] Vector search matches Article 47 of the Labor Contract Law: "Economic compensation shall be paid to the employee based on the number of years the employee has worked in the unit, at the rate of one month's salary for each year of service..."
[0067] Generated answer: "According to Article 47 of the Labor Contract Law, you can receive economic compensation of 100,000 yuan (5 years × 20,000 yuan / month × 1 times). (Quote: Article 47 of the Labor Contract Law)".
[0068] Compliance verification:
[0069] Sensitive word filtering: No sensitive words were detected;
[0070] Logical consistency check: Verify that the calculation method is consistent with the legal provisions (5 years × 20,000 × 1 times = 100,000).
[0071] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
[0072] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A multimodal human-machine question-answering guidance system based on a large language model, characterized in that: include: The multimodal input module receives and processes text, voice, and structured form input, generating a unified semantic representation through a cross-modal encoder. Dynamic intent parsing engine, used to extract explicit intent, implicit intent, and entity slots of user questions, and calculate intent ambiguity and entity missing rate; A structured question generation module is used to generate clarifying and normative questions based on the ambiguity of intent and the missing entity rate; Authoritative knowledge traceability and compliance verification layer, used to access authoritative databases in real time and generate answers containing reference links and compliance verification results; A multimodal interaction optimization engine is used to dynamically adjust follow-up strategies based on reinforcement learning and user feedback.
2. The system according to claim 1, wherein: The dynamic intent parsing engine includes: The explicit intent and entity extraction submodule is used to identify intent types (such as query and confirmation) and entity slots (such as time and place) through the BERT-CRF model; The implicit intent mining submodule is used to analyze contextual dependencies through graph neural networks to mine emotional tendencies and potential needs; The fuzziness and missing rate calculation submodule is used to evaluate the clarity and entity completeness of the question through entropy calculation.
3. The system according to claim 1, wherein: The structured question generation module includes: The clarification question generation submodule is used to generate questions that guide users to provide additional information based on the ambiguity of the intent; The normative question recommendation submodule is used to recommend structured questions that need to be supplemented based on the domain knowledge graph; The follow-up question strategy adjustment submodule is used to optimize the generation order and content of follow-up questions through reinforcement learning.
4. The system according to claim 1, wherein: The authoritative knowledge traceability and compliance verification layer includes: The authoritative database access submodule is used to connect to authoritative data sources such as medical guidelines, laws and regulations in real time; Vector retrieval and matching submodule, used to retrieve and match answer evidence through dense vectors; The compliance verification submodule is used to verify the compliance of the answer content (such as sensitive word filtering and logical consistency checking) through the rule engine and model fine-tuning.
5. The system according to claim 1, wherein: The multimodal interaction optimization engine includes: Reinforcement learning submodule, used to dynamically adjust the questioning strategy through the PPO algorithm to reduce the number of interaction rounds; User feedback submodule is used to collect user satisfaction evaluation of answers and correction behavior data; The model optimization submodule is used to update the parameters of the intent recognition and question generation models through contrastive learning.
6. A method for guiding standard questioning based on the system of claim 1, characterized in that: The following steps are involved: Receive multimodal data input by users and perform semantic analysis; Extract the explicit intent, implicit intent, and entity slots of the user's question, and calculate the intent ambiguity and entity missing rate; Generate clarification questions and normative questions based on intent ambiguity and entity missing rate; Real-time access to authoritative databases to generate structured answers and return them to users; Optimize guidance strategies based on user feedback and reinforcement learning.
7. The method according to claim 6, characterized in that The step of "extracting the explicit intent, implicit intent, and entity slots of the user's question" includes: The BERT-CRF model identifies intent types (e.g., query, confirmation, clarification) and entity slots (e.g., time, location, object). Entity slots include explicit entities (e.g., "January 1, 2025") and implicit entities (e.g., "recently" is parsed as "past 30 days" using the temporal reasoning engine). Graph neural networks (GNNs) are used to analyze contextual dependencies and mine emotional tendencies (such as anxiety and satisfaction) and potential needs (such as price comparison and after-sales service). Emotional tendencies are quantified using pre-trained emotional dictionaries (such as the NRC Emotion Lexicon), and potential needs are recommended using demand association graphs (such as "headache" → "recommended department").
8. The method according to claim 6, characterized in that The step of "generating clarification questions and normative questions based on intent ambiguity and entity missing rate" includes: If the intent ambiguity (entropy value) is higher than a threshold (such as 0.8) or the entity missing rate is higher than a threshold (such as 30%), a clarification question (such as "Which time period do you need to query data for?") is generated. The question types include: Entity supplement type: such as "Please add contract number"; Intent clarification: such as "Do you need to compare prices of different models?"; Combined with the domain knowledge graph to recommend normative questions (such as "Is it necessary to supplement the allergy history to accurately recommend drugs?" in the medical scenario), the knowledge graph is stored through triples (entity-relationship-entity), such as ("headache"-"related diseases"-"migraine").
9. The method according to claim 6, characterized in that The step of "real-time access to an authoritative database to generate a structured answer" includes: Use vector search (such as FAISS) to match relevant entries in authoritative databases. The matching strategies include: Semantic matching: Convert user questions and database entries into dense vectors and calculate cosine similarity; Keyword matching: Extract keywords from questions (such as "legal terms") and quickly locate them through inverted indexing; Generates a response containing a reference link (e.g., Article X of the XX Law) and compliance verification results. Compliance verification includes: Sensitive word filtering: Double verification through regular expressions and pre-trained models (such as BERT-based sensitive word detection); Logical consistency check: Verify the consistency of the answer with the knowledge graph through comparative learning (e.g. "Drug A and Drug B have no interaction" needs to be compared with the drug interaction database).
10. The method according to claim 6, characterized in that The step of "optimizing guidance strategy based on user feedback and reinforcement learning" includes: Collect user satisfaction ratings (1-5 points) and correction behavior data (such as the number of additional entities and question corrections). User feedback is obtained through active inquiries (such as "Did this answer solve your question?") and passive monitoring (such as whether the user continues to ask questions); The PPO algorithm is used to adjust the generation strategy of follow-up questions. The strategy adjustment goals include: Maximizing long-term rewards: For example, the weighted sum of the reduction in interaction rounds (weight 0.6) and the increase in user satisfaction (weight 0.4); Minimize system load: such as reducing redundant queries (e.g., avoiding repeated queries for already supplemented entities); Update the parameters of the intent recognition and question generation models through contrastive learning. The contrastive learning strategy includes: Positive sample pairs: the matching pair between the user's revised question and the system's understood intent; Negative sample pairs: random combinations of ambiguous questions and incorrect intents before user correction.
Citation Information
Cited By
Interaction method and system combining domain knowledge graph and dynamic intention clarification
CN120832369A
Intelligent deep questioning method and system
CN121542399A
An intelligent deep interrogation method and system
CN121542399B