Dynamic confidence coefficient weight distribution method and system for multi-modal emotion interaction

By acquiring user profiles and calculating multimodal information biases in real time, generating effective cognitive gaps and adjusting weights, the problem of decision-making misjudgment in intelligent navigation systems under multimodal signal conflict is solved, enabling adaptive personalized content output and improving the system's intelligence and user experience.

CN121934720APending Publication Date: 2026-04-28SHENZHEN WEILIAN ELEPHANT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN WEILIAN ELEPHANT TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing intelligent tour guide systems struggle to accurately assess a user's cognitive state when faced with conflicting multimodal signals, leading to incorrect decisions. Furthermore, they lack self-optimization capabilities and are unable to adapt to individual differences.

Method used

By acquiring user profile data, establishing initial confidence configuration parameters, calculating the deviation of multimodal information in real time, generating effective cognitive gaps, adjusting weights through remedial game tasks, and dynamically reconstructing confidence based on user operation accuracy, a closed-loop self-optimization mechanism is established.

Benefits of technology

It improves the system's accuracy in understanding user intent, enhances user engagement and the system's intelligence level, enables adaptive personalized content output, and improves the robustness and intelligence of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934720A_ABST
    Figure CN121934720A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer man-machine interaction, and particularly discloses a dynamic confidence coefficient weight distribution method and system for multi-modal emotional interaction, and the method comprises the steps: capturing and positioning possible cognitive conflict points of a user through monitoring the confidence coefficient deviation of multi-modal information, such as expressions and voice, in real time; the user intention understanding accuracy of the system is improved; when a cognitive gap is detected, a remedial game task is generated, and the operation accuracy of the user is used as a more objective behavior data source to calibrate and reconstruct the confidence coefficient weight in real time, so that knowledge blind spots of the user are effectively compensated; a closed-loop self-optimization mechanism is established, the effectiveness of the previous intervention measure can be evaluated by analyzing the secondary interaction behavior of the user on the self-adaptive output content, and the core judgment parameters are reversely corrected accordingly, so that the system can gradually adapt to the unique emotion expression and cognitive habits of each user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer human-computer interaction technology, and relates to a dynamic confidence weight allocation method and system for multimodal emotional interaction. Background Technology

[0002] In the field of human-computer interaction technology, particularly in scenarios such as smart tour guides and intelligent education, multimodal emotional interaction has become a core technology for enhancing user experience. This technology aims to understand a user's real-time emotional and cognitive state by comprehensively analyzing various physiological and behavioral signals, such as facial expressions, voice tone, and body posture, thereby providing more humanized and intelligent feedback and services. Immersive tour guide systems in smart museums are a typical application of this technology, striving to go beyond traditional voice explanations and create an interactive experience that fosters deep emotional and cognitive resonance with visitors.

[0003] In existing technologies, some intelligent tour guide systems have begun to attempt to integrate multimodal information to determine users' interests and comprehension levels. These systems typically incorporate separate facial expression recognition and voice analysis modules. For example, they use cameras to capture users' facial expressions to determine whether they are confused or interested, while simultaneously analyzing users' voice commands or responses through microphones to assess their knowledge acquisition. Some systems use preset static weights to weight and fuse the analysis results from different modalities to arrive at a comprehensive user status score, adjusting the level of detail in the explanations accordingly. Furthermore, some systems incorporate gamification elements, such as question-and-answer challenges, to enhance the fun of the tour.

[0004] However, existing technical solutions still have significant shortcomings in practical applications. First, when user multimodal signals conflict—for example, a user's facial expression shows confusion but their verbal response indicates understanding—mechanisms based on static weight fusion struggle to accurately determine the user's true cognitive state, easily leading to incorrect system decisions and the delivery of inappropriate navigation content. Second, gamification elements in existing technologies are often disconnected from the core cognitive state assessment system, existing primarily as independent entertainment or testing modules. They fail to integrate organically with multimodal information conflict resolution mechanisms, and cannot dynamically intervene and provide effective calibration the moment a user experiences cognitive impairment. Finally, the interaction logic of existing systems is typically fixed, lacking the ability to self-optimize based on interaction history and failing to adapt to individual differences among users. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above objectives, the present invention proposes the following technical solution: a dynamic confidence weight allocation method for multimodal emotional interaction, comprising: A1. acquiring user profile data, and establishing a confusion confidence threshold for the facial expression channel and a mastery confidence threshold for the voice channel based on the user profile data, and generating initial confidence configuration parameters.

[0006] A2. During the interaction, acquire the user's facial expression data and voice data in real time, calculate the deviation between facial expression confidence and voice confidence based on the initial confidence configuration parameters, and compare the deviation with the preset collaborative tolerance interval. When the deviation exceeds the collaborative tolerance interval, generate an effective cognitive gap.

[0007] A3. In response to the effective cognitive gap, generate remedial game tasks and obtain the user's operation accuracy for the remedial game tasks. Use the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction.

[0008] A4. Generate a knowledge confidence upgrade factor based on the comparison results between the operation accuracy rate and the preset achievement threshold.

[0009] A5. Integrate knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages.

[0010] A6. Obtain the user's secondary interaction behavior with the adaptive output content, verify the effectiveness of the remedy based on the secondary interaction behavior, and use the verification results to reverse the collaborative tolerance range.

[0011] The second aspect of the present invention provides a dynamic confidence weight allocation system for multimodal emotional interaction, comprising: a data thresholding module, which acquires user profile data, establishes a confusion confidence threshold for the facial expression channel and a mastery confidence threshold for the voice channel based on the user profile data, and generates initial confidence configuration parameters.

[0012] The gap identification module acquires the user's facial expression data and voice data in real time during the interaction. Based on the initial confidence configuration parameters, it calculates the deviation between the facial expression confidence score and the voice confidence score, and compares the deviation with the preset collaborative tolerance interval. When the deviation exceeds the collaborative tolerance interval, an effective cognitive gap is generated.

[0013] The weight reconstruction module, in response to the effective cognitive gap, generates remedial game tasks and obtains the user's operation accuracy for the remedial game tasks. It uses the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction.

[0014] The factor generation module generates knowledge confidence upgrade factors based on the comparison results between the operation accuracy rate and the preset achievement threshold.

[0015] The content output module integrates knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages.

[0016] The verification and correction module obtains the user's secondary interaction behavior with the adaptive output content, verifies the effectiveness of the remedy based on the secondary interaction behavior, and uses the verification results to reverse the collaborative tolerance range.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) By monitoring the confidence deviation of multimodal information such as facial expressions and voice in real time, the present invention can effectively capture and locate the cognitive conflict points that users may have, solve the problem of misjudgment caused by the reliance on a single information channel or static weight fusion in traditional methods, thereby significantly improving the accuracy of the system's understanding of user intentions and providing a reliable basis for subsequent personalized content output.

[0018] (2) This invention forms a collaborative interaction mode by deeply integrating gamification intervention mechanism with confidence weight adjustment. When a cognitive gap is detected, the system does not simply provide supplementary information, but generates remedial game tasks. The user's operation accuracy is used as a more objective behavioral data source to calibrate and reconstruct confidence weight in real time. This method not only effectively makes up for the user's knowledge blind spots, but also transforms the interaction process into a dynamic diagnosis and correction closed loop, which greatly enhances user participation and system interaction effect.

[0019] (3) The present invention establishes a closed-loop self-optimization mechanism, which enables the system to have the ability to continuously learn and adapt itself. By analyzing the user's secondary interaction behavior with the adaptive output content, the effectiveness of the previous intervention measures can be evaluated, and the core judgment parameters can be corrected accordingly. This feedback loop enables the system to gradually adapt to each user's unique emotional expression and cognitive habits, realizing the leap from one-time interaction to long-term personalized adaptation, thereby fundamentally improving the intelligence level and robustness of the human-computer interaction system. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention.

[0022] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1

[0025] Please see Figure 1 As shown, the dynamic confidence weight allocation method for multimodal emotional interaction proposed in this invention includes: A1. Obtaining user profile data, and establishing a confusion confidence threshold for the facial expression channel and a mastery confidence threshold for the voice channel based on the user profile data, and generating initial confidence configuration parameters.

[0026] In a preferred embodiment, generating initial confidence configuration parameters includes: parsing user profile data to determine the knowledge base weight level.

[0027] Based on the knowledge base weight hierarchy, confidence thresholds for confusion and mastery are configured respectively.

[0028] The initial confidence configuration parameters are generated by combining the knowledge base weight level, the confusion confidence threshold, and the mastery confidence threshold.

[0029] Specifically, the system obtains pre-defined user profile data through a system interface. This user profile data is a structured collection of information describing user characteristics, including at least the user's age and knowledge level. After obtaining the user profile data, the system uses this data as input to query a pre-defined weight mapping table. This weight mapping table stores the correspondence between different user profile data and knowledge base weight levels. For example, user profile data with the age group "child" and knowledge level "beginner" is mapped to the "basic" knowledge base weight level, and user profile data with the knowledge level "expert" is mapped to the "advanced" knowledge base weight level, thereby determining an initial knowledge base weight level for the current user.

[0030] Next, based on the established knowledge base weight hierarchy, the system differentiates the confusion confidence threshold for the facial expression channel and the mastery confidence threshold for the voice channel. The confusion confidence threshold is a numerical benchmark for determining whether a user is confused. The mastery confidence threshold is a numerical benchmark for determining whether a user has mastered the knowledge.

[0031] The relationship can be determined by the following function: and In the formula, Represents the weight hierarchy of the knowledge base; User's age; Based on the user's knowledge level; , and These are preset weighting coefficients used to balance the influence of age and knowledge level on the confidence threshold for confusion or the confidence threshold for mastery. , and Preset sub-functions related to user age or user knowledge level, for example, can be piecewise functions that set different base thresholds for different age groups or different knowledge levels; , These are constant terms used to fine-tune the body perplexity confidence threshold or the body mastery confidence threshold, set based on experience; Indicates the confidence threshold for confusion; Indicates the confidence threshold of the degree of mastery; The hierarchical classification is obtained by querying the weight mapping table. and This serves as the numerical benchmark used by the system to subsequently determine emotional states. For example, for the "basic" knowledge base weight level, the system sets a lower confidence threshold for confusion and a higher confidence threshold for mastery, making the system more sensitive to the confused expressions of beginner users, while requiring users to provide more explicit verbal affirmation in order to determine that they have mastered the knowledge, thereby avoiding misjudgment.

[0032] Finally, the system combines the determined knowledge base weight levels, confusion confidence threshold, and mastery confidence threshold into a structured data package to generate initial confidence configuration parameters. These parameters serve as the basic input for subsequent real-time confidence conflict detection steps, providing a personalized starting point for the entire dynamic weight adjustment process. By combining these parameters, the system can make personalized judgments based on the user's prior knowledge level and cognitive habits at the beginning of the interaction, rather than using uniform, undifferentiated initial parameters.

[0033] This invention achieves personalized configuration during the navigation system's startup phase by executing the aforementioned multimodal sentiment confidence initialization steps, rather than using uniform, undifferentiated initial parameters. This method dynamically generates unique initial confidence configuration parameters based on each user's specific profile data, enabling the system to make judgments based on the user's prior knowledge and cognitive habits from the very beginning of the interaction. This targeted initialization significantly improves the accuracy and efficiency of subsequent multimodal sentiment conflict detection, laying a solid foundation for subsequent dynamic weight adjustments and adaptive content output, thereby enhancing the overall intelligence level of the navigation system and the user experience's fit.

[0034] A2. During the interaction, acquire the user's facial expression data and voice data in real time, calculate the deviation between facial expression confidence and voice confidence based on the initial confidence configuration parameters, and compare the deviation with the preset collaborative tolerance interval. When the deviation exceeds the collaborative tolerance interval, generate an effective cognitive gap.

[0035] In a preferred embodiment, generating an effective cognitive gap includes: extracting facial features from facial expression data to calculate facial expression confidence.

[0036] Speech features are extracted from speech data to calculate speech confidence.

[0037] The confidence scores for facial expressions and speech are standardized using the initial confidence configuration parameters to generate standardized confidence scores.

[0038] Calculate the deviation between standardized confidence levels, and generate an effective cognitive gap when the deviation exceeds the collaborative tolerance range.

[0039] Specifically, the system collects users' facial expression data in real time through built-in or external image acquisition devices, and calls the preset facial expression recognition algorithm to extract facial key points, muscle movement units and other facial expression features. These facial expression features are then input into the trained model to calculate the facial expression confidence score, which represents the user's level of confusion.

[0040] Meanwhile, the system collects users' voice data in real time through sound acquisition devices such as microphones. This voice data includes at least the user's voice responses to a set of questions and answers explaining basic knowledge. The system analyzes and extracts voice features such as pitch, speech rate, and keywords from the collected voice data. These features can reflect the user's voice state and emotional tendency. For example, pitch features reflect the user's vocal range and are related to the user's emotional state; speech rate features reflect the speed at which the user answers, suggesting the depth of their thinking or familiarity with the content; and volume features are related to the user's level of confidence or emotional intensity.

[0041] A pre-defined algorithm or model (such as a rule-based system or machine learning model) is used to calculate a speech confidence score that represents the user's level of knowledge mastery. This score reflects the user's understanding of the content being explained. For example, if the user's speaking speed is within a certain range and the keyword matching is high, their level of knowledge mastery is considered high.

[0042] After obtaining the original facial expression confidence and speech confidence, the system calls the initial confidence configuration parameters and uses the confusion confidence threshold and mastery confidence threshold to standardize the facial expression confidence and speech confidence respectively, mapping them to a unified and comparable metric space.

[0043] Subsequently, the system calculates the deviation between the standardized facial expression confidence score and the speech confidence score; this process can be represented as follows: Where D represents the deviation value, This represents the standardized confidence level of facial expressions. This represents the standardized confidence level of the speech.

[0044] Finally, the system compares the calculated deviation value D with a preset co-tolerance interval, which defines the acceptable degree of inconsistency between the two modalities. Once the deviation value D exceeds the upper limit of this interval, indicating a significant conflict between the cognitive states conveyed by the user's facial expressions and speech, the system generates a valid cognitive gap marker. This marker serves as a trigger signal to initiate the subsequent dynamic weight adjustment mechanism.

[0045] This invention, by executing the aforementioned real-time confidence conflict detection steps, can accurately and dynamically capture cognitive contradictions in users during the learning process. Compared to traditional methods that rely solely on single-modal information or simple weighted fusion, this solution achieves a deep understanding of the user's true cognitive state by calculating the deviation of multimodal confidence in real time and comparing it with the collaborative tolerance interval. This effectively identifies complex situations such as "saying one thing but meaning another" or "seemingly understanding but not truly understanding." This mechanism transforms the system from a passive information provider into an intelligent diagnostic tool that proactively discovers and locates user comprehension obstacles, providing a reliable basis for subsequent precise and effective personalized interventions, and greatly enhancing the intelligence and humanization of the entire navigation system's interaction.

[0046] In a further preferred embodiment, the step of extracting facial features from the facial expression data to calculate the facial expression confidence score includes: capturing a sequence of facial images from the facial expression data.

[0047] Key point features and expression classification features are extracted simultaneously from facial image sequences to form a combined feature vector.

[0048] The combined feature vectors are input into the confidence model to output the expression confidence.

[0049] Specifically, dynamic image frames containing the user's face are continuously captured through a camera on the user's device or in the environment, forming a real-time updated sequence of facial images.

[0050] Next, the system processes each frame in the facial image sequence, identifying and extracting two complementary feature types: keypoint features and expression classification features. Keypoint features are the two-dimensional or three-dimensional coordinates of facial key points located using a facial landmark detection algorithm, such as the eyebrows, corners of the eyes, and corners of the mouth. These features accurately describe subtle geometric changes in facial muscles. Expression classification features, on the other hand, are semantic labels related to cognitive states, such as "confused," "neutral," and "focused," output after analyzing the entire facial region or region of interest using a pre-trained expression recognition network. These labels are typically presented as probability vectors.

[0051] Finally, the system combines the extracted keypoint features and facial expression classification features into a single feature vector, which is then input into a pre-defined confidence model. This confidence model, such as a regression network or support vector machine, is specifically trained to learn the complex nonlinear mapping relationship between these two features and the true state of confusion. Its final output is a continuous numerical value, namely the facial expression confidence value.

[0052] The confidence value of facial expressions can be obtained by a function. It means that, among them, This is a pre-defined confidence level model; This is a keypoint feature vector containing the coordinates of all keypoints. These correspond to the coordinates of key points such as eyebrows and corners of the mouth, obtained through facial landmark detection algorithms. This is a feature vector for facial expression classification that contains the probabilities of all facial expressions. These correspond to the probability distributions of cognitive states such as confusion and neutrality output by the pre-trained facial expression recognition network. The final expression confidence value includes relative position normalization of key point feature vectors and confusion weight enhancement of expression classification feature vectors (i.e., multiplying the probability distribution value of the cognitive state being in a confused state to reduce the influence of the primary emotional factor that causes learning interruption based on the confused state in the tour guide scenario). After fusing the two types of features, a nonlinear mapping is performed to map the original feature space data to high-dimensional feature space data, and then the Sigmoid function is used to normalize its value range to a continuous scale of [0,1].

[0053] This invention employs a method combining keypoint features and facial expression classification features to calculate facial expression confidence, achieving a more accurate and robust perception of the user's cognitive state. Compared to techniques using only a single type of feature, this approach exhibits a synergistic enhancement effect. Keypoint features provide fine-grained, objective geometric motion information, capable of capturing the user's subconscious micro-expressions; while facial expression classification features provide macroscopic, semantic-level judgments. The combination of these two features effectively distinguishes between similar but distinctly different expressions, such as frowns caused by focus and frowns caused by confusion, significantly reducing the misjudgment rate. This high-fidelity facial expression confidence output provides a highly reliable data foundation for subsequent real-time confidence conflict detection steps, ensuring that the system only triggers intervention mechanisms when truly necessary, thereby improving the accuracy and efficiency of the entire dynamic weight allocation method.

[0054] A3. In response to the effective cognitive gap, generate remedial game tasks and obtain the user's operation accuracy for the remedial game tasks. Use the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction.

[0055] In a preferred embodiment, the execution weight reconstruction includes: identifying exhibit explanations associated with effective cognitive gaps.

[0056] Increase the confidence weight of voice interaction to the dominant level and generate weight adjustment parameters.

[0057] Remedial game tasks are generated based on the content of the exhibit explanations.

[0058] Monitor user interactions with remedial game tasks to calculate operational accuracy in real time.

[0059] By using operational accuracy as real-time input to replace facial expression confidence, weight reconstruction is completed.

[0060] Specifically, the system obtains the timestamps of valid cognitive gap markers, traces back and identifies the exhibit explanations that occurred simultaneously with these markers, and precisely locates the specific knowledge points that cause cognitive conflict for users, such as "methods for dating dinosaur fossils" or "the physical principles of quantum entanglement." Then, it retrieves relevant text descriptions, images, videos, or 3D models from the exhibit database to form structured explanation content segments. Example: If a user appears confused and gives an incorrect verbal response when hearing "ecological characteristics of Triassic dinosaurs," the system will identify this segment and extract fossil images, ecological reconstruction diagrams, and explanations of key terms as related content.

[0061] After recognizing a segment of the narration, the system immediately adjusts the weights, raising the confidence weight of voice interaction from the default value to the dominant level, while lowering the confidence weight of facial expressions. This means that in subsequent comprehensive judgments, information from the voice channel will be given greater decisiveness, and the adjusted weight parameters are stored in the user's session state to provide a basis for subsequent calculations.

[0062] Next, based on the identified exhibit explanations, the system uses a content analysis engine to extract core concepts and selects or dynamically generates a remedial game task from a pre-set gamification template library. This task can take the form of an interactive Q&A or a hands-on challenge targeting that knowledge point, aiming to verify and reinforce the user's understanding through active practice. Examples: A sorting task where users drag and drop dinosaur fossils in chronological order; an interactive Q&A where users can choose the correct food type from a "dinosaur diet" game using voice or touch controls.

[0063] After the user begins performing the remedial game task, the system monitors each of the user's actions in real time and compares them with the preset correct path or answer, thereby continuously calculating the accuracy rate. This accuracy rate can be calculated using the formula... Definition, where For operational accuracy; The number of correct operation steps completed by the user, such as dragging and dropping to the correct category area; This represents the total number of operation steps performed by the user.

[0064] Finally, the system uses this dynamically changing operational accuracy as a new and more objective source of confidence, replacing the unstable facial expression confidence output by the facial expression channel and the one that conflicts with the voice channel in real time. This completes the weight reconstruction and forms a new confidence evaluation system composed of dominant voice confidence and objective game behavior data.

[0065] This invention addresses the decision-making dilemma in multimodal information conflict by executing the aforementioned dynamic weight gamification reorganization steps. It goes beyond simply choosing to trust a particular channel; instead, it introduces gamified tasks as a third-party, behavioral evaluation method. This approach transforms a passive, ambiguous emotional signal (facial expression) problem into an active, quantifiable cognitive behavior (game operation) problem. The game task plays a dual role here: it serves as an educational tool to help users fill knowledge gaps and a measurement tool to provide reliable confidence input to the system. This closed-loop reorganization mechanism of detection, intervention, and measurement enables the system to self-calibrate when encountering uncertainty, greatly enhancing its robustness and adaptability. It achieves a deep, dynamic, and accurate grasp of the user's true cognitive state, with an overall effect far exceeding the level achievable by simply combining gamification and multimodal analysis as independent functions.

[0066] In a further preferred embodiment, generating remedial game tasks based on exhibit explanations includes: parsing exhibit explanations to identify knowledge blind spots that lead to effective cognitive gaps.

[0067] Game task content is designed based on knowledge gaps, and the difficulty level of game tasks is set according to user profile data.

[0068] Integrate 3D visualization elements into game mission content to enhance interactivity and generate remedial game missions.

[0069] Specifically, after receiving a valid cognitive gap marker, the system will design content based on the knowledge blind spot corresponding to the marker.

[0070] Specifically, the system performs semantic analysis on the exhibit explanations that cause conflict, extracts core concepts, logical relationships, or key facts, and based on this, selects the most matching interaction paradigm from a preset game template library to design game task content.

[0071] Once the game content and logic are determined, the system will further utilize existing user profile data to personalize the difficulty level of the game tasks. This step dynamically adjusts task parameters based on the user's age and knowledge level. For example, for beginners, the number of distracting options may be reduced, reaction time extended, or progressive hints provided. For advanced users, the complexity or depth of the tasks may be increased to ensure a balance between challenge and feasibility.

[0072] Subsequently, to enhance user engagement and comprehension, the system integrates 3D visualization elements. It retrieves high-precision 3D models related to the current exhibit from the digital asset library and uses them as the core interactive object of the game. This allows users to intuitively explore and learn in 3D space through operations such as rotation, scaling, and clicking, greatly enhancing the game's interactivity. Finally, the system packages all elements after content design, difficulty adaptation, and 3D integration to generate a standalone executable game task instance, which is then pushed to the user's terminal device interface via a communication interface for direct user operation.

[0073] This invention, by executing the steps described above to generate remedial game tasks, constructs not merely a simple entertainment element, but a highly customized, contextualized, and visualized miniature learning tool. First, the task content directly addresses knowledge gaps, ensuring targeted and efficient intervention. Second, the dynamic matching of difficulty levels ensures a smooth learning experience, avoiding user frustration or boredom caused by difficulty mismatches. Finally, the integration of 3D visualization elements transforms abstract knowledge into a tangible and interactive experience, significantly reducing cognitive load and enhancing learning interest and memory retention. The synergistic effect of these three aspects allows this remedial game task to bridge effective cognitive gaps in the most efficient and user-friendly way, with an overall effect far exceeding traditional text-based question-and-answer games or general games unrelated to the topic.

[0074] A4. Generate a knowledge confidence upgrade factor based on the comparison results between the operation accuracy rate and the preset achievement threshold.

[0075] In a preferred embodiment, generating a knowledge confidence upgrade factor based on the comparison result between the operation accuracy rate and the preset achievement threshold includes: obtaining the difficulty of exhibit knowledge points marked as effective cognitive gaps and user profile data, setting an achievement threshold, and then comparing the operation accuracy rate with the preset achievement threshold to generate a comparison result.

[0076] When the comparison results indicate that the accuracy of the operation exceeds the preset achievement threshold, a positive knowledge confidence upgrade factor is generated.

[0077] When the comparison results indicate that the accuracy of the operation has failed to meet the standard for several consecutive times, the three-dimensional visualization compensation mechanism is activated, and a compensation-type knowledge confidence upgrade factor is generated.

[0078] Update the user's knowledge base weight level by using positive knowledge confidence upgrade factors or compensatory knowledge confidence upgrade factors.

[0079] Specifically, by identifying the knowledge difficulty of exhibits to ensure user experience adaptability, tiered achievement thresholds are set according to the user's knowledge level. For example, for beginner users, a basic threshold of 70% indicates completion of a basic task, and an advanced threshold of 85% indicates completion of a challenge task; for advanced users, the basic threshold is 80%, and the advanced threshold is 95%.

[0080] Subsequently, the system compares the user's operational accuracy after completing the remedial game task with a preset achievement threshold for the knowledge point of the exhibit, generating a clear comparison result.

[0081] When the comparison results show that the user's operation accuracy exceeds the preset achievement threshold, the system determines that the user has successfully overcome the previous cognitive gap and generates a positive knowledge confidence upgrade factor. This factor acts as a gain signal, representing that the user's cognitive level on the knowledge point of the exhibit has been effectively improved.

[0082] Conversely, if the comparison results indicate that the user's operational accuracy consistently falls short of the target during the interaction, the system will activate the 3D visualization compensation mechanism and generate a compensatory knowledge confidence upgrade factor. This compensation mechanism aids understanding by providing a more intuitive and simplified 3D model display, and the generated compensatory knowledge confidence upgrade factor serves as an adjustment signal, indicating that the user currently needs basic knowledge compensation rather than in-depth knowledge expansion.

[0083] Finally, based on the generated positive or compensatory knowledge confidence upgrade factors, the system dynamically updates the knowledge base weight levels in the user profile. This update process can be represented as follows: ,in, The updated knowledge base weight hierarchy; The preset achievement threshold; For operational accuracy; This is the preset achievement gain coefficient, which is usually greater than 1, such as the default value of 1.2; This is the preset compensation strength coefficient, which is usually less than 1, such as the default value of 0.6; The knowledge base weight hierarchy before the update; The generated knowledge confidence upgrade factor determines which formula to use to update the knowledge base weight level; A positive knowledge confidence upgrade factor; This is a compensatory knowledge confidence upgrade factor.

[0084] This invention establishes a closed-loop feedback mechanism from behavioral performance to cognitive model updates by executing the aforementioned game achievement confidence transformation steps. This method not only evaluates a user's individual performance in a remedial task but also transforms that performance into a knowledge confidence upgrade factor that can influence the system's decisions over the long term. This transformation mechanism enables the system to distinguish between a user's momentary mastery and persistent confusion, and accordingly dynamically and precisely adjust the user's knowledge base weight level. It achieves self-learning and evolution of the navigation system, continuously optimizing its internal model based on the user's actual learning curve, ensuring that subsequent content is always highly matched to the user's true cognitive level. This achieves a leap from single-point problem-solving to long-term personalized adaptation, significantly enhancing the educational effect and intelligent depth of the navigation.

[0085] A5. Integrate knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages.

[0086] In a preferred embodiment, generating and outputting adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages includes: parsing knowledge confidence upgrade factors to determine the weight upgrade state or confidence compensation state.

[0087] Under the weight upgrade state, deep academic derivative modules are retrieved from the deep academic database and generated.

[0088] Under the confidence compensation state, a gamified simplified knowledge package is retrieved from the gamified knowledge base and generated.

[0089] Deep academic derivative modules or gamified concise knowledge packages are delivered to the user interface as adaptive output content.

[0090] Specifically, the received knowledge confidence upgrade factors are analyzed to determine whether they are positive or compensatory knowledge confidence upgrade factors: positive knowledge confidence upgrade factors correspond to weight upgrade states, while compensatory knowledge confidence upgrade factors correspond to confidence compensation states.

[0091] When the system determines that it has entered the weight upgrade state, it will use the core knowledge points of the current exhibit as an index to send a request to the in-depth academic database in the background. The in-depth academic database is a knowledge base that stores in-depth information such as professional papers, archaeological reports, and academic monograph abstracts related to the exhibit. Based on the returned data, the system automatically compiles and generates a rich in-depth academic derivative module, which is designed to meet the user's needs for in-depth exploration.

[0092] Conversely, when the system determines that it has entered a confidence compensation state, it will invoke the gamified knowledge base. The gamified knowledge base stores simplified knowledge points and media materials characterized by fun and interactivity. Based on the previously identified effective cognitive gaps, the system extracts and combines these into an easy-to-understand gamified concise knowledge package, such as an animated explanation or a simple interactive diagram.

[0093] Finally, whether it's the generated in-depth academic derivative modules or the gamified concise knowledge packages, the system encapsulates them into standard adaptive output content data packages and passes them to the user interface for display through the data interface, thus completing a complete adaptive content push cycle.

[0094] This invention constructs an intelligent and highly differentiated content response mechanism by executing the aforementioned weighted adaptive output steps. The core effect of this method is that it enables the system to provide two distinct subsequent content paths based on the user's verified actual cognitive level. For users demonstrating mastery, the system effectively expands their knowledge boundaries and maintains their interest in exploration by outputting in-depth academic derivative modules, avoiding boredom caused by overly simplistic content. For users still experiencing comprehension difficulties, the system reduces cognitive load and provides a more user-friendly learning path by generating gamified, concise knowledge packages, avoiding frustration caused by overly complex content. This precise content distribution strategy based on dynamic evaluation results transforms each interaction into an effective personalized teaching intervention, thereby elevating the guided tour experience from a "one-to-many" broadcast mode to a "one-to-one" intelligent tutoring mode, significantly improving the efficiency of knowledge transfer and overall user satisfaction.

[0095] A6. Obtain the user's secondary interaction behavior with the adaptive output content, verify the effectiveness of the remedy based on the secondary interaction behavior, and use the verification results to reverse the collaborative tolerance range.

[0096] In a preferred embodiment, the reverse correction cooperative tolerance interval includes: monitoring the user's secondary interaction behavior with the adaptive output content to collect feedback data.

[0097] Analyze feedback data to assess the effectiveness of the remedy and generate an effectiveness score.

[0098] Based on the effectiveness score, the range parameter of the collaborative tolerance interval is adjusted to generate an updated collaborative tolerance interval.

[0099] Replace the collaborative tolerance interval in the initial confidence configuration parameters with the updated collaborative tolerance interval.

[0100] Specifically, the system continuously monitors users' secondary interaction behaviors after receiving in-depth academic derivative modules or gamified concise knowledge packages, and quantifies them as feedback data. These secondary interaction behaviors include, but are not limited to, the duration of user dwell on the content, the frequency of operation on interactive elements, subsequent voice questions, and facial expression changes accompanying the interaction. The system collects these behaviors through a log module and a multimodal perception module to form structured feedback data.

[0101] Next, the system comprehensively analyzes the collected feedback data to evaluate the actual effectiveness of the remedial measures taken in the previous step and generates a quantitative effectiveness score. For example, the system may use a weighted evaluation model to assign different weights to indicators such as dwell time, interaction completion rate, and positive emotional feedback, and finally calculate a comprehensive effectiveness score by summing them up.

[0102] After generating an effectiveness score, the system dynamically adjusts the range parameter of the cooperative tolerance interval used for conflict detection based on the score. The adjustment logic is as follows: if the effectiveness score is high, it indicates that the previous conflict detection and intervention were successful, and the system can appropriately tighten the cooperative tolerance interval to improve the sensitivity of future detections; if the effectiveness score is low, it may indicate that the previous detection was too sensitive or the intervention was ineffective, and the system will moderately widen the cooperative tolerance interval to reduce false positives.

[0103] The adjustment process can be performed by a function. Definition, where and These are the parameters representing the cooperative tolerance interval range before and after adjustment, respectively. For effectiveness scoring, It is a preset adjustment function based on the scoring results, such as a piecewise function that compares the validity score with a preset scoring threshold. When the score is higher than the threshold, the adjustment magnitude is increased positively; when the score is lower than or equal to the threshold, the adjustment magnitude is decreased negatively. Finally, the system updates the adjusted collaborative tolerance interval into the user's initial confidence configuration parameters, ensuring that in subsequent guided interactions, the system will use this optimized and more personalized parameter for real-time confidence conflict detection.

[0104] This invention, by executing the aforementioned closed-loop confidence optimization steps, endows the entire system with the ability to self-evolve and adapt to individual preferences. This allows the system to move beyond fixed, universal rules and gradually learn and adapt to each user's unique cognitive and emotional expression habits based on continuous interaction. As the number of interactions increases, the collaborative tolerance range becomes increasingly aligned with the characteristics of individual users, thereby significantly improving the accuracy of identifying future effective cognitive gaps and reducing the false positive rate. This continuous, automated reverse correction process greatly enhances the system's robustness and intelligence, enabling truly dynamic personalization and continuous optimization of the navigation experience.

[0105] For example, to verify the effectiveness of the present invention, the museum recruited 100 visitors for a one-month experience test and recorded data on the system under different interactive scenarios. This embodiment will explain in detail how the system initializes its configuration for different user profiles, how it detects cognitive conflicts in real time, and how it ultimately achieves personalized content output through gamification intervention and closed-loop optimization.

[0106] In this embodiment, the system first performs a multimodal emotion confidence initialization step. Before the tour begins, the system obtains a user profile through visitor registration information. For user Xiaoming, the system obtains his age as "child" and his knowledge level as "beginner," and queries a preset weight mapping table to determine his knowledge base weight level as "beginner." Based on this level, the system differentially sets a lower confusion confidence threshold (…). ) and a high confidence threshold of mastery ( This makes the system more sensitive to Xiaoming's confused expressions. For Professor Wang, whose profile is "adult" and "expert," the system determines his knowledge base weight level to be "advanced" and sets a high confidence threshold for confusion. ) and a lower confidence threshold for mastery ( These parameters are combined into initial confidence configuration parameters, providing a personalized starting point for subsequent interactions.

[0107] During the guided tour, the system enters a real-time confidence conflict detection step. When the guided tour system explains the "casting method of the Simuwu Ding," the camera captures Xiaoming's facial image in real time. By combining key point features (such as a furrowed brow) with facial expression classification features (a high probability of the "confused" label), the facial expression confidence score is calculated to be 0.82. Simultaneously, the microphone captures Xiaoming saying "I understand" to show politeness, and the speech analysis yields a speech confidence score of 0.91. After standardization using the initial confidence configuration parameters, the system calculates the deviation value D between the two to be 0.35. This deviation value exceeds the upper limit of the pre-set collaborative tolerance interval [0.1, 0.3] for novice users, therefore the system generates a valid cognitive gap marker.

[0108] Immediately, the system initiated a dynamic weighted gamification restructuring process. The system identified "mold casting method" as the content associated with the cognitive gap and immediately elevated the confidence weight of the voice interaction to the dominant level. Next, the system invoked the gamification template library and, based on the knowledge blind spot of "mold casting method" and Xiaoming's "beginner" profile, generated a moderately challenging 3D visual remedial game task: "Bronze Casting Workshop." This task required Xiaoming to drag and combine virtual molds and cores in the correct order. The system monitored his actions in real time and calculated the operation accuracy A. During this process, this dynamically changing operation accuracy A replaced the previously conflicting facial expression confidence level, completing the weight restructuring.

[0109] After Xiaoming finished the game, the system performed the game achievement confidence conversion step. Xiaoming's final operation accuracy rate was 85%, which exceeded the preset achievement threshold of 80% for this task. Based on this, the system determined that Xiaoming had overcome cognitive obstacles and generated a positive knowledge confidence upgrade factor. This factor was used to update Xiaoming's user profile, and his knowledge base weight level received a slight increase.

[0110] Next, the system enters the weighted adaptive output step. Since it receives a positive knowledge confidence upgrade factor, the system determines it has entered the weight upgrade state. Using "mold casting method" as an index, the system sends a request to the deep academic database and automatically generates a deep academic derivative module with the content "Explore other ancient artifacts that used similar casting techniques," which is then pushed to Xiaoming in the form of a short animation to satisfy his interest in further exploration.

[0111] Finally, the system performs a closed-loop confidence optimization step. After Xiaoming watched the short animation, the system detected that he spent a relatively long time on the content and actively clicked on an interactive point in the animation. The system quantifies and analyzes these secondary interactive behaviors, resulting in a high effectiveness score. Based on this high score, the system reverse-corrects Xiaoming's cooperative tolerance interval, fine-tuning its range from [0.1, 0.3] to [0.1, 0.28], making the system more sensitive and accurate in detecting Xiaoming's cognitive conflicts in the future.

[0112] Through the application of this invention, the museum's guided tour system has achieved a transformation from passive explanation to proactive diagnosis and intelligent intervention.

[0113] Example 2

[0114] Please see Figure 2 As shown, based on Embodiment 1, the second aspect of the present invention provides a dynamic confidence weight allocation system for multimodal emotional interaction, comprising: a data thresholding module, a gap identification module, a weight reconstruction module, a factor generation module, a content output module, and a verification and correction module.

[0115] The data thresholding module, gap identification module, weight reconstruction module, factor generation module, content output module, and verification and correction module are connected in sequence.

[0116] The data thresholding module acquires user profile data and establishes a confusion confidence threshold for the facial expression channel and a mastery confidence threshold for the voice channel based on the user profile data, generating initial confidence configuration parameters.

[0117] The gap identification module acquires the user's facial expression data and voice data in real time during the interaction, calculates the deviation between the facial expression confidence level and the voice confidence level based on the initial confidence configuration parameters, and compares the deviation with the preset collaborative tolerance range. When the deviation exceeds the collaborative tolerance range, an effective cognitive gap is generated.

[0118] The weight reconstruction module, in response to the effective cognitive gap, generates a remedial game task and obtains the user's operation accuracy for the remedial game task. It then uses the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction.

[0119] The factor generation module generates knowledge confidence upgrade factors based on the comparison results between the operation accuracy rate and the preset achievement threshold.

[0120] The content output module integrates knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified simplified knowledge packages.

[0121] The verification and correction module acquires the user's secondary interaction behavior with the adaptive output content, verifies the effectiveness of the remedy based on the secondary interaction behavior, and uses the verification results to reverse the collaborative tolerance range.

[0122] It should be noted that the formulas described above, through the principle of dimensional consistency and mathematical standardization methods (such as normalization, dimensionless parameter conversion, or unit system unification), can translate physical quantities with different properties into unitless standard values ​​or superimposed parameters of the same dimension. This eliminates the interference of different dimensions on the computational logic, allowing the formulas to retain the original data distribution characteristics while possessing mathematical rationality and adaptability to objective laws. The descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the invention.

Claims

1. A dynamic confidence weight allocation method for multimodal emotional interaction, characterized in that, include: A1. Obtain user profile data, and based on the user profile data, establish the confusion confidence threshold for the facial expression channel and the mastery confidence threshold for the voice channel, and generate initial confidence configuration parameters. A2. During the interaction, acquire the user's facial expression data and voice data in real time, calculate the deviation between facial expression confidence and voice confidence based on the initial confidence configuration parameters, compare the deviation with the preset collaborative tolerance interval, and generate an effective cognitive gap when the deviation exceeds the collaborative tolerance interval. A3. In response to the effective cognitive gap, generate remedial game tasks and obtain the user's operation accuracy of the remedial game tasks. Use the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction. A4. Generate a knowledge confidence upgrade factor based on the comparison results between the operation accuracy rate and the preset achievement threshold; A5. Integrate knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages; A6. Obtain the user's secondary interaction behavior with the adaptive output content, verify the effectiveness of the remedy based on the secondary interaction behavior, and use the verification results to reverse the collaborative tolerance range.

2. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The generation of initial confidence configuration parameters includes: Analyze user profile data to determine the knowledge base weight hierarchy; Based on the knowledge base weight hierarchy, confidence thresholds for confusion and mastery are configured separately; The initial confidence configuration parameters are generated by combining the knowledge base weight level, the confusion confidence threshold, and the mastery confidence threshold.

3. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The generation of effective cognitive gaps includes: Extract facial expression features from facial expression data to calculate facial expression confidence. Speech features are extracted from speech data to calculate speech confidence. The confidence scores for facial expressions and speech are standardized using the initial confidence configuration parameters to generate standardized confidence scores. Calculate the deviation between standardized confidence levels, and generate an effective cognitive gap when the deviation exceeds the collaborative tolerance range.

4. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 3, characterized in that, The step of extracting facial expression features from facial expression data to calculate facial expression confidence includes: Capture facial image sequences from facial expression data; Key point features and expression classification features are extracted simultaneously from facial image sequences to form a combined feature vector; The combined feature vectors are input into the confidence model to output the expression confidence.

5. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The execution weight reconstruction includes: Identify exhibit explanations that are associated with effective cognitive gaps; Increase the confidence weight of voice interaction to the dominant level and generate weight adjustment parameters; Generate remedial game tasks based on the content of the exhibit explanations; Monitor user interactions with remedial game tasks to calculate operational accuracy in real time; By using operational accuracy as real-time input to replace facial expression confidence, weight reconstruction is completed.

6. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 5, characterized in that, The method of generating remedial game tasks based on the exhibit explanation content includes: Analyze the content of the exhibits to identify knowledge blind spots that lead to effective cognitive gaps; Game task content is designed based on knowledge gaps, and the difficulty level of game tasks is set according to user profile data; Integrate 3D visualization elements into game mission content to enhance interactivity and generate remedial game missions.

7. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The step of generating a knowledge confidence upgrade factor based on the comparison between the operation accuracy rate and a preset achievement threshold includes: Acquire the difficulty of exhibit knowledge points marked as effective cognitive gaps and user profile data, preset achievement thresholds, and then compare the operation accuracy with the preset achievement thresholds to generate comparison results; When the comparison results indicate that the operation accuracy exceeds the preset achievement threshold, a positive knowledge confidence upgrade factor is generated. When the comparison results indicate that the operation accuracy has failed to meet the standard for several consecutive times, the three-dimensional visualization compensation mechanism is activated, and a compensation-type knowledge confidence upgrade factor is generated. Update the user's knowledge base weight level by using positive knowledge confidence upgrade factors or compensatory knowledge confidence upgrade factors.

8. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The adaptive output content generated and output in the form of in-depth academic derivative modules or gamified concise knowledge packages includes: Analyze the knowledge confidence upgrade factor to determine the weight upgrade state or confidence compensation state; Under the weight upgrade state, retrieve and generate deep academic derivative modules from the deep academic database; Under the confidence compensation state, retrieve and generate a gamified simplified knowledge package from the gamified knowledge base; Deep academic derivative modules or gamified concise knowledge packages are delivered to the user interface as adaptive output content.

9. The dynamic confidence weight allocation method for multimodal emotional interaction according to claim 1, characterized in that, The reverse correction cooperative tolerance interval includes: Monitor users' secondary interactions with the adaptive output content to collect feedback data; Analyze feedback data to assess the effectiveness of the remedy and generate an effectiveness score; Based on the effectiveness score, the range parameter of the collaborative tolerance interval is adjusted to generate an updated collaborative tolerance interval; Replace the collaborative tolerance interval in the initial confidence configuration parameters with the updated collaborative tolerance interval.

10. A dynamic confidence weight allocation system for multimodal emotional interaction, characterized in that, include: The data threshold building module acquires user profile data and establishes the confusion confidence threshold for the facial expression channel and the mastery confidence threshold for the voice channel based on the user profile data, generating initial confidence configuration parameters. The gap identification module acquires the user's facial expression data and voice data in real time during the interaction. Based on the initial confidence configuration parameters, it calculates the deviation between the facial expression confidence and the voice confidence, and compares the deviation with the preset collaborative tolerance interval. When the deviation exceeds the collaborative tolerance interval, an effective cognitive gap is generated. The weight reconstruction module, in response to the effective cognitive gap, generates remedial game tasks and obtains the user's operation accuracy for the remedial game tasks. It uses the operation accuracy to replace the facial expression confidence input in real time to perform weight reconstruction. The factor generation module generates knowledge confidence upgrade factors based on the comparison results between the operation accuracy rate and the preset achievement threshold. The content output module integrates knowledge confidence upgrade factors to generate and output adaptive output content in the form of in-depth academic derivative modules or gamified concise knowledge packages. The verification and correction module obtains the user's secondary interaction behavior with the adaptive output content, verifies the effectiveness of the remedy based on the secondary interaction behavior, and uses the verification results to reverse the collaborative tolerance range.