Adaptive learning question-answering system and method based on multi-modal interaction

By constructing a modal memory map and user portrait across rounds, and combining an adaptive learning question-and-answer system with multimodal interaction, the problem of context information loss in multiple rounds of multimodal interaction is solved, high-quality and personalized question-and-answer generation is achieved, and user experience and system adaptability is improved.

CN120256575APending Publication Date: 2025-07-04LIAOCHENG UNIV
View PDF 0 Cites 15 Cited by

Patent Information

Application Number
CN202510364855.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing Q&A systems are prone to losing context information in multiple rounds of multi-modal interactions, resulting in incorrect answers or logical confusion, affecting user experience.

Method used

Build an adaptive learning question-and-answer system for multimodal interaction, and establish a modal memory map across rounds through modal recognition, chronological coding and graph modeling, and dynamically update the question-and-answer strategy to generate personalized, context-related answers through modal recognition, chronological coding and graph modeling.

Benefits of technology

It significantly improves the semantic understanding and continuous dialogue capabilities of the Q&A system in complex interactive scenarios, improves the interactive experience and user satisfaction, and has the ability to quickly adapt to changes in user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256575A_ABST
    Figure CN120256575A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive learning question answering system and method based on multi-modal interaction, and particularly relates to the technical field of self-adaptive learning question answering. By constructing a modal recognition and preprocessing module, standardization and structuralization of multi-modal input of texts, voices, images and the like are realized; establishing a cross-round modal memory map through time sequence coding and map modeling; dynamically updating a user portrait in combination with user interaction history and current modal characteristics; in the question and answer generation process, current input, a user portrait and a historical graph are fused, intermediate semantic representation is generated through a context enhancement module, and the intermediate semantic representation is combined with a knowledge base to generate answers; meanwhile, a modal confidence degree dynamic evaluation mechanism is introduced, and weights are distributed according to the input quality, the user adaptation degree and the context correlation; and finally, incremental optimization is carried out on the graph structure, the user portrait and the question and answer strategy through a user feedback driving system, and multi-round, multi-mode and self-adaptive intelligent question and answer interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of adaptive learning question answering, and particularly relates to an adaptive learning question answering system and method based on multimodal interaction. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the breakthroughs in the fields of natural language processing, computer vision, and speech recognition, question answering systems have gradually shifted from simple rule-based queries to interactive question answering systems that integrate multimodal information and intelligently understand user intentions. Especially in scenarios such as educational tutoring, medical assistance, and enterprise knowledge services, the ways users pose questions to the system show a diverse trend, no longer limited to traditional text input, but gradually developing into a combination of various forms such as text, speech, images, and even videos, namely the so-called multimodal interaction. However, in multi-round conversations, if the information input by the user in different rounds involves multiple modalities, the system may lose context information, resulting in incorrect answers or logical confusion. This seriously affects the user experience and makes it impossible for the system to conduct natural, multimodal continuous conversations. Summary of the Invention

[0003] The purpose of the present invention is to provide an adaptive learning question answering system and method based on multimodal interaction to solve the deficiencies in the background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions: An adaptive learning question answering method based on multimodal interaction, including: Receiving various modal input information from the user, and respectively performing format standardization and preprocessing on each modal data through a modal recognition module to obtain a structured modal input sequence; Encoding the time sequence of each round of modal input generated during the multi-round conversation process, and constructing a cross-round modal memory graph; Updating the user profile based on the user's interaction history behavior and the current modal input characteristics; Based on the current round of modal input, user profile, and historical modal memory graph, generating a semantically consistent intermediate semantic representation through a context enhancement module; and then generating an answer content by a question answering generation module in combination with the semantic representation and knowledge base information; During the question answering generation process, dynamically allocating modal confidence weights according to the quality of each modal input, user profile, and context relevance; Receiving user feedback information, and incrementally optimizing the modal memory graph, user profile, and question answering strategy according to the feedback information to achieve multi-round adaptive update.

[0005] Preferably, constructing the cross-round modal memory graph includes attaching a timestamp and a round mark to each round of modal input, and establishing a graph edge structure based on the reference relationship, supplementary relationship or semantic reference relationship existing between modalities. The edge structure supports dynamic update and semantic confidence weighting.

[0006] Preferably, updating the user profile includes incrementally updating the user profile using a sliding window mechanism and a time decay function, so that new input behaviors have a higher weight in adjusting the profile.

[0007] Preferably, the context enhancement module uses a graph neural network or a multi-head attention mechanism to model the semantic path between the current modal input and the historical nodes in the modal memory graph, and generates a context-enhanced semantic representation vector.

[0008] Preferably, for each modality m, it is first necessary to calculate the modal input quality, including: calculating the signal-to-noise ratio, content length and modal recognition error rate of the modal signal respectively; after normalizing the obtained signal-to-noise ratio, content length and modal recognition error rate of the modal signal, perform weighted average summation calculation to obtain the modal input quality.

[0009] Preferably, judge whether the current modality m belongs to the set of frequently used modalities Mpref in the user profile, and define a boolean function: ; is the set of modal preferences in the user profile, is the historical usage frequency of modality m; is the preference frequency threshold; is the boolean value indicating whether the current modality is a preferred modality, and the comprehension ability scoring function , score the user's comprehension ability in the current modality, and use an S-shaped mapping function: ; In the formula, is the mean value of the historical answer satisfaction of the user in modality m, a is the function steepness coefficient, controlling the slope of the S-shaped curve, and μ is the empirically set median value of the comprehension ability; is the mapped comprehension ability score; use a non-linear combination strategy to calculate the image fitness .

[0010] Preferably, it is determined whether the current modality is explicitly referred to by the user through natural language processing. If there is an explicit reference, the marked value Refm = 1 is assigned; the semantic features of the current modality and the historical modality nodes are compared to calculate the similarity score Simm; the shortest path length L is calculated based on the modality memory graph, and the path matching score Pathm is assigned accordingly. According to the logical rules, when there is an explicit reference and the path is close, Simm is directly used as the context relevance; when there is no reference but there is an indirect path, the minimum value of Simm and Pathm is taken; if there is no effective semantic path, the relevance is set to 0.

[0011] Preferably, the modality input quality, image adaptation degree, and context relevance are converted into a comprehensive feature vector, and the comprehensive feature vector is used as the input of the machine learning model. The machine learning model takes predicting the comprehensive confidence score label of each modality for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the comprehensive confidence score labels of all each modality as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence and then the model training is stopped. The comprehensive confidence score of each modality is determined according to the model output result, where the machine learning model is a polynomial regression model.

[0012] Preferably, the comprehensive confidence score of each obtained modality is compared with a predetermined threshold. If the comprehensive confidence score of each modality is less than the predetermined threshold, the modality is down-weighted or ignored during the generation process; if the comprehensive confidence score of each modality is greater than or equal to the predetermined threshold, the modality remains unchanged.

[0013] The present invention also provides an adaptive learning Q&A system based on multimodal interaction, including a multimodal input processing module, a modality memory graph construction and management module, a user profile modeling and updating module, a Q&A generation module, and a modality confidence evaluation module; Multimodal input processing module: Receives various modality input information from the user, and respectively performs format standardization and preprocessing on each modality data through the modality recognition module to obtain a structured modality input sequence; Modality memory graph construction and management module: Encodes the modality inputs generated in each round of the multi-round dialogue process in chronological order, and constructs a cross-round modality memory graph; User profile modeling and updating module: Updates the user profile based on the user's interaction history behavior and the current modality input features; Q&A generation module: Based on the modality input of the current round, the user profile, and the historical modality memory graph, generates a semantically consistent intermediate semantic representation through the context enhancement module; then the Q&A generation module combines the semantic representation with the knowledge base information to generate the answer content; Modal confidence evaluation module: During the question and answer generation process, dynamically allocate modal confidence weights according to the input quality of each modality, user profile, and context relevance. Adaptive optimization module: Receive user feedback information, and incrementally optimize the modal memory graph, user profile, and question and answer strategy according to the feedback information to achieve multi-round adaptive updates.

[0014] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. By introducing a multi-modal recognition and structured processing mechanism, the present invention realizes the unified parsing of various modal inputs such as text, speech, images, and videos, constructs a modal memory graph encoded in chronological order, and effectively retains the context information in multi-round multi-modal interactions through semantic relationship modeling, thereby significantly improving the semantic understanding ability and continuous dialogue ability of the question and answer system in complex interaction scenarios. Combined with a dynamically updated user profile model, the system can generate personalized and context-related high-quality answers according to the user's knowledge level, modal preferences, and behavior habits, effectively improving the interaction experience and user satisfaction.

[0015] 2. The present invention proposes a comprehensive confidence calculation mechanism based on modal input quality, user profile adaptability, and context relevance, and dynamically allocates modal weights in combination with a machine learning model, realizing refined control of the question and answer strategy and high-robustness output. By introducing a user feedback-driven incremental optimization mechanism, the system can continuously learn and self-adjust, while ensuring stable performance, it has the ability to quickly adapt to changes in user needs, and is widely applicable to various high-demand scenarios such as education tutoring, intelligent customer service, and medical consultation, with strong practical value and promotion prospects. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is the method flow chart of the present invention.

[0018] Figure 2 It is the system module diagram of the present invention. Detailed Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Example 1. Refer to Figure 1 As shown, the adaptive learning Q&A method based on multi-modal interaction in this embodiment includes: Receiving various modal input information from the user, and respectively performing format standardization and preprocessing on the data of each modality through a modality recognition module to obtain a structured modal input sequence; Encoding the time sequence of each round of modal input generated during the multi-round dialogue process, and constructing a cross-round modal memory graph; Updating the user profile based on the user's interaction history behavior and current modal input features; Based on the current round of modal input, user profile, and historical modal memory graph, generating a semantically consistent intermediate semantic representation through a context enhancement module; and then generating an answer content by a Q&A generation module in combination with the semantic representation and knowledge base information; During the Q&A generation process, dynamically allocating modal confidence weights according to the quality of each modal input, user profile, and context relevance; Receiving user feedback information, and incrementally optimizing the modal memory graph, user profile, and Q&A strategy according to the feedback information to achieve multi-round adaptive update.

[0021] When the user interacts with the Q&A system, problem information can be input in any one or more modal ways such as text, voice, image, and video. To ensure that the system has unified processing capabilities in the face of input scenarios with different modal combinations, after receiving the user input, the present invention first automatically discriminates the type of input data by a modality recognition module, and then performs targeted preprocessing and format standardization operations respectively. Finally, the processing result is converted into a unified structured modal input sequence, providing a data basis for subsequent multi-modal semantic modeling and context reasoning.

[0022] Specifically, when the user input is in text form, the system first performs text cleaning, including removing invalid characters, punctuation standardization, case unification, etc.; then uses a word segmentation model (such as a word segmentation algorithm based on a dictionary or sub-word units) to segment the text, and then extracts semantic features such as keywords, entity names, and problem types through a syntax analysis tool, and finally forms a structured text feature vector representation.

[0023] When the user input is in the form of speech, the system first uses a speech recognition engine to transcribe the speech signal into text information, while retaining optional metadata such as the speaker's timbre characteristics, intonation changes, and pause information. The transcribed text information will enter the same preprocessing process as text input to further extract semantic features. In addition, the system can also evaluate parameters such as the background noise level and signal-to-noise ratio of the speech input, which serves as the basis for subsequent modality confidence adjustment.

[0024] When the user input is in the form of an image, the system uses an image recognition module to perform format unification and size standardization processing on the uploaded image, such as scaling the image to a specified resolution, unifying color channels, removing redundant borders, etc. Subsequently, a pre-trained image feature extraction network is used to extract key region features, and optionally, techniques such as object detection and image segmentation are combined to annotate entity, region, or text information in the image, generating a feature tensor and semantic annotation corresponding to the image.

[0025] When the user input is in the form of a video, the system first performs frame processing on the video content and selects key frames based on the frame change rate; for the selected key frames, an image processing process is executed, and at the same time, the audio track of the video is extracted for speech recognition and sentiment analysis. The video input is thus disassembled into a joint input of the image modality and the speech modality, and then the corresponding preprocessing processes are executed respectively.

[0026] After all modalities have completed their respective data cleaning and feature extraction, the system maps them into a unified modality input sequence through timestamp and dialogue turn marking, ensuring that the modality information can accurately correspond to the specific user input context in the subsequent processing stage, and avoiding modality misalignment, semantic repetition, or loss.

[0027] To solve the problems of easy loss of user context information and lack of semantic continuity between modalities in the multi-round multi-modal question-and-answer process, after completing the standardization and structuring processing of each round of modality input, the system further establishes a cross-round, multi-modal memory expression structure, namely the "modality memory graph", through a chronological encoding mechanism and a modality graph construction strategy.

[0028] Specifically, the system first attaches timestamp information and dialogue turn marking to the structured modality data of each round of user input. The timestamp information is recorded based on the time when the system receives the input, accurate to the millisecond level, and is used to characterize the relative or absolute order of the modality data in the entire conversation process; while the turn marking is used to label which round of interaction in the entire conversation this modality input belongs to, such as "the first round of image input" "the third round of speech input", etc., thus forming a modality input stream sorted in time series.

[0029] Next, the system summarizes and organizes all-round modal input information through the modal memory graph construction module to construct a cross-round semantic structure graph. The modal memory graph consists of nodes and edges: Nodes represent each modal input unit, which can include text fragments, speech transcriptions, image objects, video frames, etc. Each node contains attributes such as feature representations of the modality, context semantic summaries, round information, confidence scores, etc.; Edges represent the association relationships between different modalities or different rounds. The types of edges include "semantic reference relationship", "context reference relationship", "modal complement relationship", etc.

[0030] For example, when the user inputs an image in the first round and asks "Is there any problem here?" in the speech of the third round, the system, through the natural language reference resolution and modal alignment mechanism, recognizes that "here" has a reference relationship with a certain key area in the first-round image, and then establishes a reference edge from the "image node A" to the "speech node C" in the modal memory graph.

[0031] After the graph is constructed, the system can, in each subsequent round of question and answer, combine the current modal input with the historical modal nodes in the graph for dynamic semantic fusion and context reasoning. Through graph traversal, subgraph extraction, and edge weight analysis, the system can find a set of modal nodes semantically related to the current user question and generate more accurate and contextually coherent answer content accordingly.

[0032] In addition, to prevent redundant growth or semantic drift of the memory graph, the present invention also has a graph compression mechanism and an expired node elimination strategy. The system periodically evaluates the activity of nodes and the usage frequency of edges, and merges or removes modal nodes that have not been referenced for a long time or have repeated semantics, ensuring that the graph remains efficient and semantically clear during continuous multi-round interactions.

[0033] To improve the performance of the question and answer system in personalized response and continuous learning, the system introduces a dynamic user profile update mechanism during multi-round conversations. This mechanism models and continuously updates the user's knowledge state, preference tendency, interaction habits, etc. by comprehensively analyzing the user's interaction history behavior data and the multi-modal features of the current round of input, so as to provide personalized strategy support for subsequent question and answer generation.

[0034] Specifically, the user profile consists of multiple attribute dimensions, mainly including but not limited to the following types: Knowledge level attribute: Represents the user's knowledge mastery level in a specific field, which can be evaluated by the complexity of the user's questions, the professionalism of the terms used, the satisfaction of answer feedback, etc.; Content preference attributes: Record which types of modalities the user prefers in questions (such as preferring a combination of text and images or voice interaction), content styles (such as concise or explanatory), and areas of concern (such as education, healthcare, technology, etc.); Interaction behavior attributes: Include the frequency of user questions, the tendency for multi-round follow-up questions, whether there are error correction behaviors (such as repeatedly modifying questions), etc.; Modality usage habit attributes: Record which modality the user more often uses to ask questions and the historical evaluation of the input quality under that modality (such as speech recognition accuracy, image clarity, etc.).

[0035] After each round of question-and-answer interaction is completed, the system will fuse the semantic features extracted from the current modality input with the user's historical portrait. The fusion strategies include: Feature extraction: Extract key features such as semantic keywords, sentiment tendencies, and visual themes from the current input modality data. For example, if the text contains technical terms or the image contains technical drawings, the system can judge the user's professional level based on this; Behavior recognition: Analyze the user's behavior, such as immediately adding an explanation after the previous question, choosing to rephrase the question, actively uploading an auxiliary modality (such as an attached drawing), etc., to judge their initiative in obtaining accurate answers or their trust in the system; Portrait update mechanism: Adopt a sliding window mechanism or a weighted average strategy to incrementally update the existing user portrait, ensuring that new behavior information reasonably affects the user modeling result without completely covering the old data. At the same time, the system adjusts parameters such as the "satisfaction response threshold" or "preference for understanding depth" in the portrait according to user feedback (such as the score for the answer, whether to continue asking questions).

[0036] For example, if a user continuously uses voice input for three rounds, repeatedly mentions basic medical common sense questions in the voice, and is accompanied by the upload of an image (such as a skin photo), the system can mark the user portrait with features such as "non-professional medical field user", "prefers voice interaction", and "has a certain willingness for self-health management". In subsequent answers, the system will automatically use a more popular and explanatory language style, combined with graphic and text-based auxiliary information for answering, to improve user satisfaction.

[0037] To ensure the long-term effectiveness of the user portrait and the system's adaptability, the present invention also has a time decay mechanism, that is, the historical interaction behaviors are weighted according to time to ensure that the most recent behaviors have a greater influence on the user portrait, thereby reflecting changes in the user's knowledge state or interests.

[0038] After the system receives the user's modal input for the current round and completes the update of the user profile, to generate Q&A content with accurate semantics and coherent context, it further introduces a context enhancement module. This module comprehensively utilizes the current input modal information, the historical modal memory graph, and the user profile data to generate a highly consistent and targeted intermediate semantic representation, which serves as the core basis for Q&A generation.

[0039] The system first calls the context enhancement module to process the modal information in the current input round and link it with the relevant nodes stored in the modal memory graph in the historical rounds. This module uses a multi-head attention mechanism or a graph neural network (GNN) to model the semantic relationship between the current node and the historical semantic nodes.

[0040] For example, if the current user uploads an experimental video and asks a voice question: "What's the difference between this part and the previous experimental results?" The system will automatically recognize that "this part" is related to the current video modality, while "the previous experimental results" need to retrieve the previous round's image or text content from the graph. Through context graph traversal, the system can construct a semantic path, extract the embedding representation of the historical nodes, and fuse it with the current modal semantics to form an enhanced context expression.

[0041] Meanwhile, the system aligns the fused semantic vector with the preference features extracted from the user profile, such as identifying whether the user prefers a concise description or a detailed explanation, so as to adjust the style or depth of the semantic representation to better match the user's cognitive level.

[0042] Finally, the context enhancement module outputs a set of semantically consistent intermediate semantic representation vectors, which have encoded the current modal information, historical context references, and user preference tendencies.

[0043] After obtaining the intermediate semantic representation, the system enters the Q&A generation module, which can select retrieval-based, generation-based, or hybrid Q&A methods according to the application scenario.

[0044] For a retrieval-based Q&A system, the semantic representation vector is used as the query input, and an embedded vector retrieval engine is called to match semantically similar knowledge fragments from a structured knowledge base or a domain document set. The matching results are extracted by the summary module to obtain the key content and fused into the answer content.

[0045] For a generation-based Q&A system, a neural generation model (such as a text generation network based on the Transformer architecture) is called based on the current semantic vector to generate natural language answers. The generation process can be controlled by the user profile, such as introducing control tags to control the output length, tone, language complexity, etc.

[0046] For the hybrid Q&A mechanism, the system can first retrieve relevant content as a generation prompt, and then execute answer generation in combination with semantic representation to improve accuracy and controllability.

[0047] For example, if a high school student uploads an image of a chemical experiment device and asks, "What's wrong with this device?" The system will identify the structural features in the image (such as incorrect connection order of the conduits), and combine labels such as "high school student" and "preference for non-technical terms" in the user profile to generate the answer: "The connection order of the conduits and the condenser in this device is incorrect, which may cause the reaction gas to not condense smoothly, affecting the experimental results." To achieve adaptive weight allocation of multi-modal information in the Q&A generation process, the system calculates the modal input quality, user profile adaptability, and context relevance respectively based on the multi-modal input content received in each round of conversation, in combination with the degree of association in the user profile and the context graph; the system normalizes them and constructs a comprehensive confidence function to dynamically adjust the weight of each modality in Q&A generation, so as to enhance the decision-making ability of high-quality modalities and reduce the impact of noisy or irrelevant modalities.

[0048] For each modality m (such as text, speech, image, video), the system first needs to calculate the modal input quality, including: calculating the signal-to-noise ratio, content length, and modal recognition error rate of the modal signal respectively; the signal-to-noise ratio of the modal signal is applicable to the speech / image / video modality; the content length reflects the amount of information in the text modality or the number of objects detected in the image; the modal recognition error rate is the ASR recognition error rate; after normalizing the obtained signal-to-noise ratio, content length, and modal recognition error rate of the modal signal, a weighted average summation calculation is performed to obtain the modal input quality.

[0049] Judge whether the current modality m belongs to the set of commonly used modalities Mpref in the user profile, and define a boolean function: ; is the set of modal preferences in the user profile (such as {text, image}), is the historical usage frequency of modality m; is the preference frequency threshold (set by the system, for example, more than 3 times is considered a preference); is the boolean value (1 or 0) indicating whether the current modality is a preferred modality, and the comprehension ability scoring function , scores the user's comprehension ability in the current modality, using an S-shaped mapping function: ; In the formula, is the mean value of the historical answer satisfaction of the user in modality m (between 0 and 1), a is the function steepness coefficient, controlling the slope of the S-shaped curve, and μ is the empirically set median value of the comprehension ability (such as 0.5); is the comprehension ability score after mapping, ranging from 0 to 1. This function ensures that the score decays rapidly under low comprehension ability and converges quickly to 1 under high comprehension ability.

[0050] To determine whether the interaction style of the current modality matches the user's preferred style, a discrete scoring function is used: ={1.0, complete match (e.g. preference for "illustration and text" and the current one is image + text); 0.7, partial match (e.g. preference for concise style but the current one is image); 0.3, obvious mismatch (e.g. preference for voice and the current one is long text); 0.0, unknown or conflicting modal style; is the style matching factor, with a value range of [0, 1]. Judgment basis: joint inference based on the expression preference dimension recorded in the user portrait + modal content style label.

[0051] The system uses a more logical nonlinear combination strategy to calculate image adaptability : ; If the modal is not in the preference list , even if the comprehension ability and style match well, it will not be given a high degree of adaptation; if the mode is a preferred mode The system takes the minimum value of comprehension ability and style matching as the final adaptability of the modality, ensuring that a high score can be obtained only when both dimensions are met; avoiding excessively high scores in one dimension to cover up the defects of another dimension, ensuring the rationality and robustness of the score.

[0052] First, the system determines whether there is an explicit reference relationship in the current modal input, that is, whether the user explicitly refers to a previous modal content through language expression. For example, if the user uses expressions such as "this picture" or "the video just now", the system confirms whether the modality is explicitly mentioned by the user through natural language processing and matching the modal graph timestamp. If a reference relationship is identified, the system sets the reference mark value Refm of the modality to 1, otherwise it is 0.

[0053] Next, the system calculates the semantic similarity between the current modality and the most relevant modality node in the historical context. The system first embeds the features of the current input modality and compares the semantics with the nodes in the historical modality memory graph. The similarity score Simm is calculated through cosine similarity or vector distance. The value range is 0 to 1. The closer to 1, the closer the semantics.

[0054] Subsequently, the system analyzes whether there is an effective semantic path connection between the current modality and the relevant nodes in the graph, which is used to measure the contextual structure relationship between modalities. The system calculates the shortest path length L in the graph and sets the corresponding path matching score Pathm according to the path length: if the path length does not exceed 2, it is considered that the modalities are strongly associated and assigned a value of 1.0; when the path is between 3 and 4, it is considered an indirect association and assigned a value of 0.5; if the path exceeds 4 or no connection can be established, a value of 0 is assigned, indicating no effective contextual connection. Finally, the system combines the above three factors and calculates the context relevance parameter through the following logical rules: When the modality is explicitly referenced (i.e., Refm = 1) and there is a direct path connection with the historical modality (Pathm = 1.0), the system directly uses the semantic similarity score Simm as the context relevance; When the modality is not explicitly referenced (Refm = 0), but there is an indirect path connection (Pathm > 0), the system takes the smaller value of the semantic similarity and the path matching degree as a more conservative relevance score; if the current modality cannot establish an effective semantic path connection with any historical node, the relevance score is directly set to 0, indicating no contextual semantic connection.

[0055] Convert the modality input quality, image adaptation degree, and context relevance into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes predicting the comprehensive confidence score label of each modality as the prediction target, and takes minimizing the sum of the prediction errors of the comprehensive confidence score labels of all modalities as the training target. Train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training. Determine the comprehensive confidence score of each modality according to the model output result, where the machine learning model is a polynomial regression model.

[0056] Compare the obtained comprehensive confidence score of each modality with a predetermined threshold. If the comprehensive confidence score of each modality is less than the predetermined threshold, the modality is down-weighted or ignored during the generation process to avoid incorrect judgments caused by low-quality inputs; if the comprehensive confidence score of each modality is greater than or equal to the predetermined threshold, the modality remains unchanged.

[0057] To achieve the continuous learning and dynamic adaptation of the multi-modal interactive Q&A system, after each round of Q&A interaction, the system will actively receive and analyze user feedback information, and based on this, incrementally update the key components in the system, including the modality memory graph, user profile, and Q&A strategy model, so as to achieve real-time response to changes in user behavior, knowledge state, and preferences, and improve the overall Q&A accuracy and interaction intelligence of the system.

[0058] The sources of feedback information can include but are not limited to the following forms: Explicit evaluation, such as the user's satisfaction score for the answer result, or selecting options like "Is it helpful?" Implicit behaviors, such as whether the user asks follow-up questions, repeats the question, modifies the question, or changes the modality input; System-side behaviors, such as whether the answer is interrupted or the response times out.

[0059] After the user completes a round of Q&A interaction, the system will dynamically annotate and adjust the nodes and edges in the knowledge graph based on the feedback information. For example, if the user is not satisfied with a certain answer and clearly corrects the misunderstanding of the previous modality in the follow-up question, the system will automatically reduce the semantic confidence of the corresponding node, or set the weight of the incorrect semantic edge to zero to avoid its incorrect reference in subsequent conversations.

[0060] In addition, the system will also mark "the modality area actively referred to by the user" or "the image segment corrected by the user" as highly sensitive areas, and establish a dedicated historical correction path in the knowledge graph to improve the accuracy of answering such questions in the future.

[0061] The feedback information is also used to update the user profile in real time, especially in the following dimensions: If the user frequently modifies the question or repeats the question, the system will infer that the user has a relatively low level of knowledge in this field, appropriately lower the "knowledge level" score, and adjust the subsequent answers to a more basic style; If the user has multiple high-satisfaction interactions in the image or voice modality, the system will increase its "modality preference weight" and predict that this modality will be preferentially used for recommendations in the future; If the user has a repeated preference for the answer style (such as preferring concise answers), the system will record its expression style preference and make matching adjustments in subsequent generations.

[0062] The update method adopts an incremental modeling approach, that is, using the current feedback as the basis for fine-tuning, without clearing the original profile information, but dynamically iterating the profile parameters through a sliding window or a weighted time decay mechanism.

[0063] According to the feedback results, the system will make local adjustments to the inference path, modality combination strategy, language style template, etc. of the Q&A generation module. For example: If the system finds that the current modality combination continuously leads to user dissatisfaction in a certain type of question, it will reduce the triggering probability of this modality combination in similar questions; If the user continuously supplements detailed information in multiple rounds of follow-up questions, the system will actively switch the Q&A strategy from a simple answer style to a guiding Q&A style to encourage the user to provide more context; The system can also mark certain specific question forms as "vague questions" or "context-dependent questions", and preferentially use a multi-round clarification strategy when similar expressions appear later.

[0064] The adjustment process of the question-and-answer strategy is not based on overall model retraining, but adopts a module-level parameter fine-tuning mechanism to achieve low-cost and fast-response local optimization, improving system stability and interaction efficiency.

[0065] Example 2, please refer to Figure 2 As shown, the adaptive learning question-and-answer system based on multimodal interaction in this embodiment includes a multimodal input processing module, a modal memory graph construction and management module, a user portrait modeling and updating module, a question-and-answer generation module, and a modal confidence evaluation module; Multimodal input processing module: Receives various modal input information from the user, and through the modal recognition module, standardizes and preprocesses the format of each modal data respectively to obtain a structured modal input sequence; Modal memory graph construction and management module: Encodes the modal inputs generated in multiple rounds of conversations in chronological order and constructs a cross-round modal memory graph; User portrait modeling and updating module: Updates the user portrait based on the user's interaction history behavior and current modal input features; Question-and-answer generation module: Based on the modal input of the current round, the user portrait, and the historical modal memory graph, generates a semantically consistent intermediate semantic representation through the context enhancement module; then the question-and-answer generation module combines the semantic representation with the knowledge base information to generate the answer content; Modal confidence evaluation module: During the question-and-answer generation process, dynamically assigns modal confidence weights according to the quality of each modal input, the user portrait, and the context relevance; Adaptive optimization module: Receives user feedback information and incrementally optimizes the modal memory graph, the user portrait, and the question-and-answer strategy according to the feedback information to achieve multi-round adaptive updates.

[0066] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0067] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0068] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0069] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.

Claims

1. An adaptive learning Q&A method based on multimodal interaction, characterized in that: It includes: Receiving various modal input information from the user, and respectively performing format standardization and preprocessing on each modal data through a modal recognition module to obtain a structured modal input sequence; Encoding the modal inputs in each round generated during the multi-round conversation process in chronological order, and constructing a cross-round modal memory graph; Updating the user profile based on the user's interaction history behavior and the current modal input features; Based on the current round of modal input, user profile, and historical modal memory graph, generating a semantically consistent intermediate semantic representation through a context enhancement module; and then generating an answer content by a question and answer generation module in combination with the semantic representation and knowledge base information; During the question and answer generation process, dynamically allocate modal confidence weights according to the quality of each modal input, user profile, and context relevance; Receiving user feedback information, and incrementally optimizing the modal memory graph, user profile, and question and answer strategy according to the feedback information to achieve multi-round adaptive updates.

2. The adaptive learning Q&A method based on multimodal interaction according to claim 1, wherein: The construction of the cross-round modal memory graph includes attaching a timestamp and a round mark to each round of modal input, and establishing a graph edge structure based on the existing reference relationship, complementary relationship, or semantic reference relationship between modalities, and the edge structure supports dynamic update and semantic confidence weighting.

3. The adaptive learning Q&A method based on multimodal interaction according to claim 1, wherein: The update of the user profile includes using a sliding window mechanism and a time decay function to incrementally update the user profile, so that new input behaviors have a higher weight for profile adjustment.

4. The adaptive learning Q&A method based on multimodal interaction according to claim 1, characterized in that: The context enhancement module adopts a graph neural network or a multi-head attention mechanism to model the semantic path between the current modal input and historical nodes in the modal memory graph, and generates a context-enhanced semantic representation vector.

5. The adaptive learning Q&A method based on multimodal interaction according to claim 4, characterized in that: For each modality m, it is first necessary to calculate the modal input quality, including: calculating the signal-to-noise ratio, content length, and modal recognition error rate of the modal signal respectively; after normalizing the obtained signal-to-noise ratio, content length, and modal recognition error rate of the modal signal, performing weighted average summation calculation to obtain the modal input quality.

6. The adaptive learning Q&A method based on multimodal interaction according to claim 5, characterized in that: Determine whether the current modality m belongs to the set of frequently used modalities Mpref in the user profile, and define a Boolean function: ; is the set of modality preferences in the user profile, is the historical usage frequency of modality m; is the preference frequency threshold; is the Boolean value indicating whether the current modality is a preferred modality, and the comprehension ability scoring function , scores the user's comprehension ability in the current modality, using an S-shaped mapping function: ; In the formula, is the mean historical answer satisfaction of the user in modality m, a is the function steepness coefficient, controlling the slope of the S-shaped curve, and μ is the empirically set median of the comprehension ability; is the mapped comprehension ability score; Calculate the image fitness using a non-linear combination strategy .

7. The adaptive learning Q&A method based on multimodal interaction according to claim 6, characterized in that: Identifying whether the current modality is explicitly referred to by the user through natural language processing. If there is an explicit reference, assign a marker value Refm = 1; compare the semantic features of the current modality with historical modal nodes, and calculate the similarity score Simm; calculate the shortest path length L based on the modal memory graph, and accordingly assign a path matching score Pathm. According to the logical rule, when there is an explicit reference and the path is close, directly use Simm as the context relevance; when there is no reference but there is an indirect path, take the minimum value of Simm and Pathm; if there is no effective semantic path, the relevance is set to 0.

8. The adaptive learning Q&A method based on multimodal interaction according to claim 7, wherein: Convert the modal input quality, image fitness, and context relevance into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes predicting the comprehensive confidence score label of each modality for each group of comprehensive feature vectors as the prediction target, and minimizing the sum of the prediction errors of the comprehensive confidence score labels of all modalities as the training target. Train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training. Determine the comprehensive confidence score of each modality according to the model output result, where the machine learning model is a polynomial regression model.

9. The adaptive learning Q&A method based on multimodal interaction according to claim 8, characterized in that: Compare the obtained comprehensive confidence score of each modality with a predetermined threshold. If the comprehensive confidence score of each modality is less than the predetermined threshold, the modality is down-weighted or ignored during generation; If the comprehensive confidence score of each modality is greater than or equal to the predetermined threshold, the modality remains unchanged.

10. An adaptive learning Q&A system based on multimodal interaction, which is used to implement the multimodal interaction-based adaptive learning Q&A method according to any one of claims 1-9, and is characterized in that: It includes a multi-modal input processing module, a modal memory graph construction and management module, a user profile modeling and updating module, a question and answer generation module, and a modal confidence evaluation module; Multi-modal input processing module: Receive various modal input information from the user, and respectively perform format standardization and preprocessing on the data of each modality through the modal recognition module to obtain a structured modal input sequence; Modal memory graph construction and management module: Encode the modal inputs of each round generated during the multi-round conversation process in chronological order, and construct a cross-round modal memory graph; User profile modeling and updating module: Update the user profile based on the user's interaction history behavior and current modal input features; Question and answer generation module: Based on the current round of modal input, user profile, and historical modal memory graph, generate a semantically consistent intermediate semantic representation through the context enhancement module; then the question and answer generation module combines the semantic representation with the knowledge base information to generate the answer content; Modal confidence evaluation module: During the question and answer generation process, dynamically allocate modal confidence weights according to the input quality of each modality, user profile, and context relevance; Adaptive optimization module: Receive user feedback information, and incrementally optimize the modal memory graph, user profile, and question and answer strategy according to the feedback information to achieve multi-round adaptive update.

Citation Information

Cited By

  • Intelligent question and answer recommendation method and system based on metro survey data large model

    CN120429413A

  • Man-machine interaction system and method based on AI large model

    CN120560519A

  • Method and system for shooting and translating by applying smart watch, and smart watch

    CN120745659A

  • A method and system for taking a picture by using a smart watch and a smart watch

    CN120745659B

  • AI model decision-making method and device based on multi-modal data and electronic equipment

    CN120781082A