Dynamic explanation accompaniment method and system for learning emotion recognition

By collecting learners' multimodal data and combining it with a dual-core decision-making mechanism, the problem of inaccurate emotion recognition in existing learning support tools has been solved, enabling precise intervention in personalized teaching and prediction of learning setbacks, thus improving the effectiveness of learning assistance.

CN121502701BActive Publication Date: 2026-05-01JIANGSU HAOHAN INFORMATION TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU HAOHAN INFORMATION TECH
Filing Date
2026-01-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing learning support tools are one-sided in their ability to capture emotions, lack personalized intervention, and cannot predict learning setbacks, resulting in inaccurate emotion recognition results and failing to meet the needs of personalized teaching.

Method used

By collecting learners' multimodal data in real time, including facial videos, audio recordings, and interactive text, and combining it with a dual-core decision-making mechanism, teaching intervention strategies are generated. The dual-core decision-making mechanism includes white-box decision-making and black-box decision-making. White-box decision-making is based on preset rules, while black-box decision-making is based on a time-series prediction model.

Benefits of technology

It achieves precise matching of personalized emotional states, avoids learning setbacks in advance, and improves the accuracy and effectiveness of learning assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502701B_ABST
    Figure CN121502701B_ABST
Patent Text Reader

Abstract

The application discloses a dynamic explanation accompanying training method and system for learning emotion recognition, relates to the technical field of human-computer interaction education, and comprises the following steps: collecting multi-modal data of learners in a learning process in real time, including facial video data, voice audio data and interactive text data; performing emotion state recognition based on the multi-modal data, and generating an emotion state signal carrying a confidence mark; generating a teaching intervention strategy through a dual-core decision mechanism based on the emotion state signal and the confidence, in combination with a learning task context of the learner; and the dual-core decision mechanism comprises a white-box decision based on a preset rule and a black-box decision based on a time sequence prediction model. The application solves the technical problems of one-sided emotion capture, lack of individualized intervention and inability to predict learning setbacks of existing learning accompanying training tools, and achieves the technical effects of making individualized learning accompanying training more suitable for the needs of learners and improving the accuracy and effectiveness of learning assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction education application technology, and in particular to a dynamic explanation and practice method and system for learning emotion recognition. Background Technology

[0002] In learning accompaniment, accurate capture and adaptive intervention of learners' emotional states are crucial for improving learning efficiency, and the matching degree between emotion recognition results and teaching strategies directly affects learning outcomes. Existing technologies mostly use single-dimensional data for emotion judgment, relying on fixed intervention logic. They lack the ability to fuse multi-source information in data processing and fail to achieve personalized adaptation in human-computer interaction. These methods can play a certain role in simple learning scenarios, but with the increasing demand for personalized learning, their shortcomings become apparent when applied to dynamic learning scenarios. Traditional detection methods lack precise data processing mechanisms and adapted human-computer interaction modes, failing to comprehensively capture multimodal emotion data such as facial expressions, voice, and text. This results in inaccurate emotion recognition results, rigid intervention strategies, and data that is insufficient to support personalized teaching, failing to meet the needs of accurate assessment and efficient management in learning accompaniment. Summary of the Invention

[0003] This application provides a dynamic explanation and training method and system for learning emotion recognition, which solves the technical problems of existing learning training tools, such as one-sided emotion capture, lack of personalized intervention, and inability to predict learning setbacks.

[0004] The first aspect of this application provides a dynamic explanation and practice method for learning emotion recognition. The method includes: real-time acquisition of multimodal data of learners during the learning process, including facial video data, voice audio data, and interactive text data; emotion state recognition based on the multimodal data to generate an emotion state signal carrying a credibility marker; and generating a teaching intervention strategy based on the emotion state signal and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism, wherein the dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.

[0005] A second aspect of this application provides a dynamic explanation and practice system for learning emotion recognition. The system includes: a multimodal data acquisition module for real-time acquisition of multimodal data of learners during the learning process, including facial video data, voice audio data, and interactive text data; an emotion state signal acquisition module for emotion state recognition based on the multimodal data, generating emotion state signals carrying credibility markers; and a teaching intervention strategy acquisition module for generating teaching intervention strategies based on the emotion state signals and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application acquires multimodal data, including facial videos, audio recordings, and interactive text, during the learner's learning process in real time. Through multimodal feature extraction, fusion model computation, credibility scoring allocation, and personalized emotional baseline comparison, it obtains emotional state signals, calculates emotional state deviation, inflection point risk probability, and uncertainty estimates. Combining a dual-core decision-making mechanism (white-box and black-box decision-making), it dynamically adjusts teaching intervention strategies to accurately adapt to learners' different emotional states, such as anxiety and immersion, and proactively avoids learning setbacks. This makes the personalized adaptation and precise intervention of learning accompaniment more reliable and efficient, achieving the technical effect of making personalized learning accompaniment more aligned with learners' needs and improving the accuracy and effectiveness of learning assistance. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating the dynamic explanation and practice method for learning emotion recognition provided in the embodiments of this application.

[0010] Figure 2 This is a schematic diagram of the structure of the dynamic explanation and tutoring system for learning emotion recognition provided in the embodiments of this application.

[0011] Figure labeling: Multimodal data acquisition module 1, emotional state signal acquisition module 2, teaching intervention strategy acquisition module 3. Detailed Implementation

[0012] This application provides a dynamic explanation and training method and system for learning emotion recognition, which solves the technical problems of existing learning training tools, such as one-sided emotion capture, lack of personalized intervention, and inability to predict learning setbacks.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] Example 1, as Figure 1 As shown, a dynamic explanation and practice method for learning emotion recognition is described, wherein the method includes:

[0016] It collects multimodal data of learners in real time during the learning process, including facial video data, voice and audio data, and interactive text data.

[0017] Specifically, the first step is to build a multimodal data acquisition environment adapted to the learning scenario, and deploy hardware devices with real-time capture capabilities, such as mobile phones, smart tablets, educational smart terminals or computer accessories, including high-definition cameras for capturing facial images, high-fidelity microphones for recording voice, and terminal devices for learners to interact with. At the same time, the acquisition devices and data processing systems are adapted and connected to ensure smooth data transmission channels.

[0018] For facial video data, a high-definition camera is activated to capture the dynamic facial images of learners during the learning process in real time. The frame rate is kept stable during the acquisition process to fully record changes in facial expressions. By processing the acquired continuous video frames, the geometric features of key facial points are extracted frame by frame to obtain basic data that can reflect the learner's facial emotional state.

[0019] For voice audio data, a high-fidelity microphone is turned on to continuously record the learner's voice information during the learning process, covering various voice content such as answering questions and expressing doubts. The real-time acquired audio signals are pre-processed to reduce noise and remove environmental noise interference. Then, the acoustic features in the voice are extracted in real time to form the basic voice data that can be used for subsequent emotion recognition.

[0020] For interactive text data, learners' text input content, including answer texts and question texts, is recorded in real time through the learning interaction terminal. At the same time, the time node and content type of the text input are recorded synchronously. The collected interactive text is parsed to identify the preset learning behavior events and complete the collection and preliminary processing of interactive text data.

[0021] The collection of the above three types of data is carried out simultaneously. During the collection process, all types of data are transmitted to the backend storage module in real time to form a complete learner multimodal raw dataset, providing comprehensive and continuous data support for subsequent emotion state recognition.

[0022] Based on the multimodal data, emotion state recognition is performed to generate emotion state signals carrying credibility markers.

[0023] In this embodiment, the emotional state signal is a concrete assessment result obtained by analyzing the learner's multi-dimensional performance, including facial expressions, tone of voice, and interactive text, and includes a credibility score, reflecting the deviation of the learner's current emotion from their typical emotion.

[0024] Optionally, firstly, geometric features and acoustic features of facial key points are extracted from three types of multimodal data: facial video, audio, and interactive text, along with pre-defined behavioral events, to construct a multimodal feature recognition result. Next, this result, along with the context information of the current learning task, is input into a multimodal fusion model to obtain an initial emotional state probability distribution. Subsequently, signal consistency recognition is performed on the multimodal data to assign a credibility score to the initial emotional state probability distribution. Finally, the initial emotional state probability distribution is compared with the learner's personalized emotional baseline, and the emotional state deviation bound to the credibility score is output as the emotional state signal.

[0025] Based on the emotional state signals and their credibility, and combined with the learner's learning task context, a teaching intervention strategy is generated through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.

[0026] In this embodiment, white-box decision-making involves predefined teaching intervention rules, including rules for anxiety states and rules for immersive fluency states. It uses emotional state deviation with credibility scores and learning task context parameters as inputs to logical judgment units. When all judgment units are satisfied, corresponding analysis is triggered, and a first teaching intervention strategy is generated. Black-box decision-making involves inputting an emotional state deviation sequence with credibility scores and a learning task context sequence into a time-series prediction model trained on an LSTM with Dropout applied to all layers. After multiple forward propagations, the inflection point risk probability and uncertainty estimate are obtained. These are then combined to adjust the knowledge density and generate a second teaching intervention strategy.

[0027] In one embodiment of this application, the white-box decision-making method first predefines a set of teaching intervention rules. The triggering condition of each rule depends on the joint output of multiple logical judgment units. The inputs of these logical judgment units include the deviation of the emotional state with a credibility score and learning task context parameters. Only when all logical judgment units in the triggering condition of a teaching intervention rule are satisfied will the intervention strategy analysis of the corresponding rule be triggered, thereby generating the corresponding first teaching intervention strategy.

[0028] Next, a black-box decision-making method collects the emotional state deviation sequence with credibility scores and the learning task context sequence, which are then input into a time-series prediction model trained on LSTM and using Dropout before all layers. After the model outputs T prediction results through T forward propagations, the average of the results is used to obtain the inflection point risk probability, and the standard deviation of the results is calculated to obtain the uncertainty estimate. Finally, these two values ​​are combined to adjust the knowledge density, thereby generating the second teaching intervention strategy.

[0029] Furthermore, the method provided in this application embodiment includes:

[0030] Geometric features of facial key points are extracted from the facial video data, acoustic features are extracted in real time from the speech audio data, and preset behavioral events are identified from the interactive text data to establish a multimodal feature recognition result. The multimodal feature recognition result and the current learning task context information are input into a multimodal fusion model to obtain an initial emotional state probability distribution. Signal consistency recognition is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution. The initial emotional state probability distribution is compared with a personalized emotional baseline established for the learner, and an emotional state deviation with a credibility score is output as the emotional state signal.

[0031] Specifically, feature extraction operations are first performed on the three types of multimodal data to form multimodal feature recognition results. For facial video data, continuous video frames are processed frame by frame to accurately extract geometric features of key facial points such as coordinates, spacing, and angles of key areas like the eyes, eyebrows, and mouth. For speech audio data, the original audio signal is first preprocessed with noise reduction, and then acoustic features such as fundamental frequency, speech rate, volume, and spectrum are extracted in real time. For interactive text data, the learner's input, including answers and inquiries, is processed to identify pre-defined behavioral events such as frequent questioning, prolonged periods without text input, and incorrect answers. The extracted features from the three types are then integrated to establish a complete multimodal feature recognition result.

[0032] Subsequently, a multimodal fusion model was constructed and trained as a neural network model to output the initial emotional state probability distribution. During model construction, a neural network architecture was built, comprising an input layer, a feature fusion layer, a hidden layer, and an output layer. The input layer was configured with independent input channels corresponding to the three types of features mentioned above: the geometric feature input channel for facial key points was set to 128 dimensions, the acoustic feature input channel for speech was set to 64 dimensions, and the interactive text behavior event feature input channel was set to 32 dimensions. The feature fusion layer used a 4-head multi-head attention mechanism to weightedly fuse the three types of features, with each attention head uniformly mapped to a 64-dimensional feature dimension. The hidden layer consisted of two fully connected network layers: the first layer had 256 neurons, and the second layer had 128 neurons. Both layers incorporated the ReLU activation function to improve the model's fitting ability. A Dropout layer was added between the two fully connected network layers, with a dropout probability of 0.3 to prevent overfitting. The output layer contained nodes corresponding to various emotional types such as anxiety, immersion, calmness, irritability, and confusion, and used the Softmax activation function to output the probability values ​​of each emotion.

[0033] During model training, a dataset containing 100,000 valid samples was selected as the training samples. This dataset covers the facial key point geometric features, speech acoustic features, and interactive text behavior event features of learners from different educational levels and subjects in various learning scenarios. Each sample is labeled with a corresponding real emotion tag. The training set, validation set, and test set are divided in a 7:1:2 ratio. The stochastic gradient descent algorithm is used to iteratively adjust the network weights and bias parameters. The initial learning rate is set to 0.001, and the learning rate decays to 0.5 every 10 iterations. The batch size is set to 64. The cross-entropy loss function is used as the model optimization objective. Iterative training continues until the loss function value on the validation set no longer decreases after 5 consecutive iterations, and the model's emotion recognition accuracy on the test set reaches a preset standard of 90%. The model training is then complete. The multimodal feature recognition results established in the above steps, along with contextual information such as the difficulty of the current learning task, subject type, and learning stage, are input into the trained multimodal fusion model. After model calculation, the initial emotion state probability distribution containing the proportion of different emotion types is output.

[0034] Next, the geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into the pre-trained corresponding unimodal emotion recognition models to obtain three types of unimodal emotion probability distributions. The divergence between these three distributions and the initial emotion state probability distribution is calculated to generate a divergence matrix. Then, the divergence matrix is ​​weighted and the result is mapped to a credibility score. The larger the divergence value, the lower the corresponding credibility score. This step will be explained in detail later.

[0035] Finally, when generating the emotional state signal, a personalized emotional baseline containing the probability distribution of typical emotional states under different learning task contexts is first obtained. Then, the statistical distance between the initial emotional state probability distribution and the probability distribution of typical emotional states corresponding to the current learning task context is calculated to obtain the emotional state deviation value. Finally, this value is bound to the credibility score to generate the emotional state deviation signal. This step will be explained in detail in the following content.

[0036] Furthermore, the method provided in this application embodiment includes:

[0037] The geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into pre-trained facial monomodal emotion recognition models, speech monomodal emotion recognition models, and interactive text monomodal emotion recognition models, respectively, to obtain facial emotion probability distributions, speech emotion probability distributions, and interactive text emotion probability distributions. The divergence between each pair of the facial emotion probability distribution, the speech emotion probability distribution, the interactive text emotion probability distribution, and the initial emotion state probability distribution is calculated to generate a divergence matrix. The divergence matrix is ​​weighted and the calculation result is mapped to the credibility score, where the larger the divergence value, the lower the credibility score.

[0038] Optionally, a facial monomodal emotion recognition model is first constructed and trained. This model adopts a CNN architecture with an input layer dimension of 128, adapted to the dimensionality of facial key point geometric features. A feature extraction network with two convolutional layers is constructed. The first convolutional layer has a kernel size of 3×3 and a number of 64, while the second convolutional layer has a kernel size of 3×3 and a number of 128. Each convolutional layer is followed by a max pooling layer with a pooling window of 2×2. The hidden layer is set to a fully connected network with 256 neurons, introducing the ReLU activation function and adding a dropout layer with a dropout probability of 0.25 to prevent overfitting. The number of nodes in the output layer corresponds to the number of emotions such as anxiety, immersion, calmness, irritability, and confusion, and the Softmax activation function is used. The model training used tens of thousands of facial feature samples, each labeled with a corresponding emotion tag. The training, validation, and test sets were divided in a 7:1:2 ratio. The Adam optimizer was used with an initial learning rate of 0.001 and a batch size of 32. Training was iterated for 50 epochs until the model achieved a 90% accuracy rate in emotion recognition on the test set. The model's input is the geometric features of facial key points, and its output is the probability distribution of facial emotions.

[0039] Next, a monomodal speech emotion recognition model was constructed and trained, employing a combined CNN and RNN architecture. The input layer dimension was set to 64 dimensions to match the dimensions of the speech acoustic features. The feature extraction part consisted of two convolutional layers, each with a 3×3 kernel size and 32 and 64 kernels respectively, followed by a bidirectional RNN layer with 128 hidden units to capture the temporal features of the speech signal. The hidden layer was a fully connected network with 128 neurons, incorporating the ReLU activation function. The output layer was consistent with the monomodal facial emotion recognition model, setting multiple emotion nodes and using the Softmax activation function. During training, tens of thousands of speech acoustic feature samples were selected, covering speech data with different speaking speeds and volumes. The ratio of training, validation, and test sets was 7:1:2. The Adam optimizer was used with an initial learning rate of 0.001, a batch size of 32, and 50 iterations. Training stopped when the validation set loss no longer decreased for three consecutive iterations and the test set accuracy reached 90%. The model's input is the speech acoustic features, and its output is the speech emotion probability distribution.

[0040] Then, an interactive text unimodal emotion recognition model was constructed and trained using the TextCNN architecture. The input layer transforms the interactive text behavioral event features into 64-dimensional vectors. The feature extraction layer uses three different sizes of convolutional kernels: 2×64, 3×64, and 4×64, with 32 kernels of each size, to extract semantic features of different lengths. After convolution, the features are processed by a max-pooling layer. The hidden layer is a single fully connected network with 128 neurons, incorporating the ReLU activation function and adding a dropout layer with a dropout probability of 0.3. The output layer follows a 5-node design and uses the Softmax activation function. Tens of thousands of interactive text feature samples were selected for training. The training, validation, and test sets were divided in a 7:1:2 ratio. The Adam optimizer was used with an initial learning rate of 0.001, a batch size of 32, and 50 iterations of training. Training was considered complete when the test set accuracy reached 90%. The model's input is the interactive text behavioral event recognition result, and its output is the interactive text emotion probability distribution.

[0041] The preprocessed facial key point geometric features, acoustic features, and interactive behavior event recognition results are then input into the corresponding trained single-modal emotion recognition models. After the models perform their respective calculations, the probability distributions of facial emotions, speech emotions, and interactive text emotions, which include the proportions of the five emotion types, are obtained.

[0042] Next, the KL divergence calculation method is used to perform pairwise divergence calculations on the four types of probability distributions. The KL divergence calculation method is adapted to the scenario of probability distribution difference quantification. During the calculation, two types of probability distributions are used as inputs, and the divergence value reflecting the distribution difference is obtained through the formula. Specifically, six sets of divergences are calculated: the first divergence between the facial emotion probability distribution and the voice emotion probability distribution; the second divergence between the facial emotion probability distribution and the interactive text emotion probability distribution; the third divergence between the facial emotion probability distribution and the initial emotion state probability distribution; the fourth divergence between the voice emotion probability distribution and the interactive text emotion probability distribution; the fifth divergence between the voice emotion probability distribution and the initial emotion state probability distribution; and the sixth divergence between the interactive text emotion probability distribution and the initial emotion state probability distribution. A 6×6 divergence matrix is ​​constructed based on these six sets of divergence values.

[0043] Finally, the divergence matrix is ​​weighted and calculated. Considering the importance of the three types of unimodal data in emotion recognition, weights are assigned to the six groups of divergences. The weighted total score is obtained by multiplying each divergence value by its corresponding weight and summing the results. A linear mapping method is used to map the weighted total score to a confidence score of 0-100. The maximum weighted total score corresponds to 0 points, the minimum to 100 points, and intermediate scores are proportionally adjusted to achieve a mapping relationship where a larger divergence value corresponds to a lower confidence score.

[0044] By constructing a single-modal model with specific parameters, using KL divergence to calculate pairwise distribution differences, and weighting and mapping scores, a series of steps were taken to achieve accurate quantification of the reliability of the probability distribution of the initial emotional state, providing reliable data support for the generation of subsequent teaching intervention strategies.

[0045] Furthermore, the method provided in this application embodiment includes:

[0046] A personalized emotional baseline is established for the learner, which includes the probability distribution of the learner's typical emotional states in different learning task contexts; the statistical distance between the initial emotional state probability distribution and the probability distribution of the typical emotional states corresponding to the current learning task context is calculated to obtain an emotional state deviation value; the emotional state deviation value is bound to the credibility score to generate the emotional state deviation signal.

[0047] Specifically, the probability distribution of typical emotional states is structured data that quantifies learners’ common emotional expressions in a specific learning task context. It includes the probability values ​​of various emotional types such as anxiety, immersion, calmness, irritability, and confusion. The sum of the probability percentages of all emotional types is 1, which can accurately reflect the learner’s inherent emotional tendencies in the task scenario.

[0048] Next, a personalized emotional baseline is obtained. This baseline is a core reference continuously built and dynamically updated based on the learner's past learning data. It includes the probability distribution of typical emotional states of the learner in different learning task contexts (covering subject type, task difficulty, learning stage, etc.). The backend database uses a hierarchical storage structure to manage the baseline data. The learner's unique identifier is used as the top-level classification index. Under the index, there are first-level subdirectories divided by subject type, such as mathematics, Chinese, and English. Each first-level subdirectory is further divided into second-level subdirectories according to the task difficulty (basic, intermediate, advanced) and the learning stage (preview, in-class exercises, after-class review, stage tests, etc.). Each second-level subdirectory stores the probability distribution data file of typical emotional states in the corresponding scenario.

[0049] Subsequently, by associating the learner's currently logged-in account with their unique identifier, the complete personalized emotion baseline dataset corresponding to that identifier is retrieved. Then, specific contextual parameters for the current learning task are extracted, including the explicit subject type, task difficulty level, and learning stage information. These parameters are precisely matched against the directory hierarchy within the dataset, filtering out typical emotion state probability distribution files that perfectly fit the current task scenario. If multiple historical emotion data records exist for the same task scenario, the average of the corresponding emotion type probability values ​​from each record is taken to form a comprehensive typical emotion state probability distribution, which serves as the benchmark data for subsequent comparisons with the initial emotion state probability distribution.

[0050] Next, the Euclidean distance is calculated to obtain the emotional state deviation value. First, the dimensions of the initial emotional state probability distribution and the selected typical emotional state probability distribution are clarified to ensure that the number and order of emotional types in the two distributions are completely consistent. The probability values ​​of each emotional type in each distribution constitute the corresponding dimension vector. By performing operations on the probability values ​​at the same position in the two vectors one by one, the difference between corresponding elements is calculated first, then all differences are squared and summed, and finally the square root of the sum is taken. The obtained value is the emotional state deviation value between the initial emotional state and the typical emotional state. The larger the value, the more significant the difference between the learner's current emotion and their typical emotion.

[0051] Finally, the emotional state deviation value and the credibility score are linked. A fixed information combination format is first determined, with the emotional state deviation value as the core data and the credibility score as auxiliary verification information, and the two are integrated in a character-separated format. After integration, a structured emotional state deviation signal is generated. This signal contains both the core value reflecting the learner's current emotional abnormality and a credibility score measuring the reliability of that value, providing complete and accurate data support for the subsequent dual-core decision-making mechanism to generate teaching intervention strategies.

[0052] Furthermore, the method provided in this application embodiment includes:

[0053] The personalized emotion baseline is dynamically updated by means of the following mechanism: collecting the probability distribution of the learner's emotional state in different learning task contexts when the confidence score is greater than a preset confidence threshold, and using the cluster center of the emotional state probability distribution to update the corresponding typical emotional state probability distribution in different learning task contexts.

[0054] Specifically, a preset credibility threshold is first set. Combining the actual needs of emotion recognition scenarios with industry standards, this threshold is set to 80 points. This threshold serves as the core benchmark for filtering valid emotion data, eliminating probability distributions of emotion states with low credibility and little reference value. Next, the data collection process is initiated. Through the backend data acquisition module, the probability distributions of learners' emotion states in various learning tasks are captured in real time, while simultaneously associating the corresponding credibility scores with learning task context information. The learning task context includes key parameters such as subject type, task difficulty, and learning stage. Emotion state probability distribution data with credibility scores greater than 80 points are extracted from the stored data and categorized according to the learning task context, forming multiple sets of valid emotion datasets adapted to different scenarios.

[0055] Next, the effective emotion dataset after classification was clustered using the K-means clustering algorithm. Before clustering, the number of clusters corresponding to each learning task context was determined to be 1, ensuring that only one core cluster was formed in each scenario, focusing on reflecting the typical emotional characteristics of the learner in that scenario. When performing the clustering operation, firstly, one emotion state probability distribution in the data set was randomly selected as the initial cluster center. Then, the Euclidean distance between each other emotion state probability distribution in the group and the initial cluster center was calculated, and all data were grouped into one cluster. Then, the average value of the probability values ​​of each emotion type in all emotion state probability distributions in the cluster was calculated iteratively, and the cluster center was updated. This iterative process was repeated until the change in the cluster center was less than the preset convergence threshold of 0.001, at which point the iteration was stopped and the final cluster center was determined.

[0056] Finally, the final cluster centers obtained through clustering under each learning task context are directly used as the updated typical emotion state probability distribution for that scenario. The old data for the corresponding learning task context stored in the backend database is replaced with the new typical emotion state probability distribution, and the update time and data change logs before and after the update are recorded for subsequent traceability and verification. For newly added learning task contexts, if their effective emotion dataset meets the clustering conditions, the corresponding typical emotion state probability distribution is generated through the same process described above and added to the personalized emotion baseline to improve the baseline's scenario coverage.

[0057] By setting a reliable threshold to filter valid data, using the K-means clustering algorithm to determine cluster centers and update baseline data, the personalized emotion baseline is dynamically optimized, ensuring that the baseline continuously matches the learner's emotional expression characteristics and improving the accuracy of subsequent emotion state recognition and deviation calculation.

[0058] Furthermore, the method provided in this application embodiment includes:

[0059] A set of predefined teaching intervention rules is provided, wherein the triggering condition of each rule depends on the joint output of multiple logical judgment units. The input of the logical judgment units includes the deviation of the emotional state with a credibility score and the learning task context parameters. The intervention strategy analysis of the corresponding rule is triggered and the corresponding first teaching intervention strategy is generated only when all logical judgment units in the triggering condition of any one of the teaching intervention rules are satisfied.

[0060] In this embodiment, the logic judgment unit is a digital circuit or computer system component constructed from basic logic gates. It can perform logical operations on input data and output true or false Boolean results, providing a basis for judgment in system decision-making or control processes.

[0061] Specifically, a set of teaching intervention rules is predefined, including rules for anxiety states and rules for immersive fluency states. For the anxiety state rule, its triggering condition depends on the joint output of multiple logical judgment units. The inputs to the logical judgment units are the emotional state deviation with a credibility score and learning task context parameters. Specific conditions can be found in the example of the first rule, i.e., the emotional deviation is greater than a threshold. Credibility greater than the threshold 1. The number of consecutive errors is ≥2 and the current thinking / pause time is greater than the threshold. When all these conditions are met, the cognitive scaffolding mode is activated to break down the problem-solving path and provide growth mindset feedback.

[0062] For the immersive / fluid state rule, the triggering condition depends on the joint output of multiple logical judgment units. The input is the emotional state deviation with a credibility score and the learning task context parameters. Referring to the example of the second rule, the emotional state (immersive / fluid) deviation is greater than the threshold. When the current task completion accuracy rate is greater than 95% and the task completion speed is greater than the average speed, interdisciplinary challenge questions are dynamically generated to stimulate exploration potential.

[0063] Subsequently, each logical judgment unit compares the deviation of the emotional state with the credibility score with the preset threshold, and matches the learning task context parameters (such as task type, difficulty, etc.) with the rule conditions, and outputs a boolean value indicating whether the condition is met or not.

[0064] When all logical judgment units in a certain teaching intervention rule are satisfied, the corresponding rule's intervention strategy analysis is triggered. This step will be explained in detail later. For example, when the anxiety state rule is satisfied, the system decomposes the problem-solving path according to a predefined process and generates growth mindset feedback; when the immersive fluency state rule is satisfied, the system dynamically constructs interdisciplinary challenge tasks based on subject attributes and knowledge nodes.

[0065] Finally, corresponding primary teaching intervention strategies are generated, such as pushing out broken-down problem-solving steps, displaying growth mindset prompts, or presenting interdisciplinary challenge questions, to intervene in the learner's emotional state and help them adjust their learning state.

[0066] Furthermore, the method provided in this application embodiment includes:

[0067] The teaching intervention rules include at least the anxiety state rule and the immersion and fluency state rule; when the anxiety state rule is triggered, the cognitive scaffolding mode is activated to decompose the problem-solving path and implant growth mindset feedback; when the immersion and fluency state rule is triggered, interdisciplinary challenge questions are dynamically generated.

[0068] In one embodiment, the core components of teaching intervention rules are first clarified, and a structured teaching intervention rule base is constructed, incorporating anxiety state rules and immersion fluency state rules as core entries. Each rule in the rule base is assigned an independent identifier code, associated with corresponding triggering condition parameters and execution instructions. This clarifies that the two rules respectively cater to the learners' different needs for negative emotion regulation and positive learning state expansion, providing stable data support for subsequent precise triggering of intervention strategies.

[0069] When the anxiety state rule is triggered, the cognitive scaffolding mode is activated to break down the problem-solving path. First, subject knowledge point extraction technology is used to locate the core difficulties and related basic concepts of the current problem-solving task. Using these core difficulties as anchors, the complete problem-solving path is broken down into 3-5 consecutive, actionable steps according to the cognitive hierarchy of basic cognition – deep understanding – practical application. Each step clearly defines the corresponding knowledge points, problem-solving methods, and key operational points. Simultaneously, a visual problem-solving path diagram is generated, clearly presenting the progressive relationship between steps and reducing the learner's cognitive load.

[0070] Simultaneously, growth-oriented feedback is integrated using a growth-oriented feedback template matching approach. A pre-built library of growth-oriented feedback templates covering different problem-solving scenarios is constructed, containing standardized feedback texts for dimensions such as effort process, strategy adjustment, and overcoming difficulties. The system extracts key behavioral data from the learner's problem-solving process, such as the number of attempts, the problem-solving methods used, and the types of errors. Based on these behavioral data characteristics, corresponding feedback content from the template library is matched to generate personalized thinking feedback that focuses on the process, guiding learners to pay attention to their efforts and methods in the problem-solving process.

[0071] When the immersive, fluid state rule is triggered, interdisciplinary challenge questions are dynamically generated using a cross-disciplinary knowledge graph retrieval method. Based on a pre-constructed multidisciplinary knowledge graph, core knowledge points from the current learning task are extracted as retrieval anchors. A graph neural network is then used to retrieve related concepts and application scenarios of these knowledge points in other disciplines, establishing interdisciplinary knowledge connections. Next, according to Bloom's Taxonomy of Educational Objectives, the difficulty level of the questions is set slightly above the learner's current level, ensuring that the questions are within their zone of proximal development. The question scenario is adjusted in conjunction with the context of the current learning task to generate complete challenge questions that integrate multidisciplinary knowledge, both aligning with existing knowledge reserves and stimulating exploration potential.

[0072] Furthermore, the method provided in this application embodiment includes:

[0073] A sequence of emotional state deviation with credibility scores and a sequence of learning task context are collected and input into the temporal prediction model, which is trained based on LSTM, with Dropout used before all layers of the LSTM. The temporal prediction model performs T forward propagation and outputs T prediction results. The average of the T prediction results is taken to obtain the inflection point risk probability. The standard deviation of the T prediction results is calculated to obtain the uncertainty estimate. The knowledge density is adjusted by combining the inflection point risk probability and the uncertainty estimate to generate a second teaching intervention strategy.

[0074] Optionally, a time-series prediction model is first constructed and trained. Specifically, an LSTM architecture is used, with two LSTM layers and 128 hidden units per layer. A Dropout layer is added before each LSTM layer with a Dropout ratio of 0.2. By randomly shutting down some neurons during training, overfitting is reduced and generalization ability is improved. The output layer is a fully connected layer with an output dimension of 1. During training, learners' historical emotional state deviation sequences with credibility scores and learning task context sequences are collected. The dataset is divided into time windows, with each window containing the sequence data of the past 5 minutes as input and the corresponding subsequent emotional changes as labels. The Adam optimizer is used with a learning rate of 0.001 and a batch size of 32. Iterative training is performed for 200 epochs until the model's loss on the validation set converges. The model's input consists of a historical and real-time emotional state deviation sequence with credibility scores and a learning task context sequence, specifically covering data from the current moment to the past 5 minutes. The output is the predicted inflection point risk probability value for a single forward propagation.

[0075] Next, the emotional state deviation sequence with credibility score and the learning task context sequence from the current time to the past 5 minutes are collected. They are preprocessed according to the input format during model training to ensure that the sequence length and feature dimension are consistent with the training data before being input into the trained time series prediction model.

[0076] Then, T forward propagations are performed, for example, T=100. Since Dropout is active during prediction in the model, different combinations of neurons are randomly turned off during each forward propagation. Therefore, each propagation will output a different prediction result, resulting in a total of T prediction results.

[0077] Then, the average of the T prediction results is taken, and all the results are added together and divided by T to obtain the inflection point risk probability. This probability reflects the average risk level of learners experiencing an emotional inflection point in the future.

[0078] Then, the standard deviation of the T prediction results is calculated. The dispersion of the results is quantified by the standard deviation calculation formula to obtain an uncertainty estimate. This value reflects the fluctuation of the model prediction results and reflects the uncertainty of the prediction.

[0079] Finally, the knowledge density is adjusted by combining the probability of inflection point risk and the estimate of uncertainty, and a second teaching intervention strategy is generated. This step will be explained in detail in the following content.

[0080] By constructing an LSTM time series prediction model with Dropout, and through multiple forward propagation steps and calculation of average risk and uncertainty, the model achieves accurate prediction of learners' emotional inflection points and intelligent generation of intervention strategies, thereby improving the foresight and accuracy of teaching interventions.

[0081] Furthermore, the method provided in this application embodiment includes:

[0082] If the inflection point risk probability is greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, a teaching intervention strategy to reduce knowledge density is generated, including breaking down problem-solving steps and embedding review materials; if the inflection point risk probability is less than the preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, the current knowledge density is maintained, and a help prompt message to dynamically increase knowledge density is generated, awaiting user response; if the uncertainty estimate is greater than the preset uncertainty threshold, the current knowledge density is maintained, and an optional help prompt message is generated, awaiting user response.

[0083] In one embodiment, two core preset thresholds are first set using the percentile method: a preset risk probability threshold and a preset uncertainty threshold. Historical data on inflection point risk probability and uncertainty estimates of learners over the past three months are collected and sorted in ascending order. The 70th percentile of the historical inflection point risk probability data is used as the preset risk probability threshold, which covers most high-risk scenarios. The 20th percentile of the historical uncertainty estimates data is used as the preset uncertainty threshold, ensuring that only low-volatility data is considered reliable, providing a clear quantitative standard for subsequent strategy judgments.

[0084] Next, a dual assessment is performed on the inflection point risk probability and uncertainty estimate. When the inflection point risk probability is detected to be greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, the process of generating a teaching intervention strategy to reduce knowledge density is triggered. First, the problem-solving steps are broken down using a hierarchical knowledge point decomposition method. Using the core knowledge points of the current learning task as anchors, and following the cognitive pattern of basic cognition – deep understanding – practical application, the complete problem-solving path is divided into 3-5 consecutive executable small steps. Each step is clearly labeled with the corresponding knowledge points and operational key points, reducing the difficulty of understanding. Then, a related review material push method is used to embed review materials. Based on the core knowledge points in the current problem-solving steps, related past basic knowledge point review content is retrieved from the backend knowledge base, and exercises and concept explanations suitable for the current task are selected and simultaneously pushed to the learner.

[0085] When the probability of an inflection point risk is less than a preset risk probability threshold, and the estimated uncertainty value is less than or equal to a preset uncertainty threshold, the current knowledge presentation density is maintained, without changing the depth of explanation of knowledge points, the difficulty of questions, or the pace of content progression. Simultaneously, a prompt message template generation method is used to create help prompt messages that dynamically increase knowledge density. A preset prompt message template library is invoked, containing standardized prompt text adapted to different subjects and knowledge points. Based on the subject type and knowledge point attributes of the current learning task, the corresponding template is matched, and content such as "Would you like to try more difficult extension questions?" is filled in to generate personalized help prompt messages. These messages are then pushed to the learner's interface and await user response.

[0086] For the third scenario, when the detected uncertainty estimate exceeds a preset uncertainty threshold, the current knowledge density is maintained to avoid intervention errors due to ambiguous state judgment. A two-way option prompt generation method is used to create optional help prompts. Two standardized option texts are constructed: one to increase knowledge density, with content expanding on advanced topics related to the current knowledge point; and the other to decrease knowledge density, focusing on basic consolidation of the current knowledge point. These two options are integrated into a unified prompt, which is then pushed to learners for self-selection, fully adapting to scenarios with uncertain states.

[0087] Through the above steps, the teaching intervention strategy was accurately matched with the learner's real-time emotional state and data reliability, thereby improving the rationality and practicality of the second teaching intervention strategy.

[0088] In summary, the dynamic explanation and practice method for learning emotion recognition provided in this application has the following technical effects:

[0089] This application collects multimodal data of learners' facial videos, audio, and interactive text in real time. Through feature extraction, unimodal model operation, and divergence weighting, it obtains the deviation of emotional state with credibility scores. Combined with dynamic updates of personalized emotional baselines, it generates intervention strategies through a dual-core decision-making mechanism of white-box and black-box, accurately identifies learning emotions, and intervenes in a timely manner. This achieves the technical effect of making personalized learning tutoring more in line with learners' needs and improving the accuracy and effectiveness of learning assistance.

[0090] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a dynamic explanation and practice system for learning emotion recognition, the system comprising:

[0091] Multimodal data acquisition module 1 is used to collect learners' multimodal data in real time during the learning process, including facial video data, voice and audio data, and interactive text data.

[0092] The emotional state signal acquisition module 2 performs emotional state recognition based on the multimodal data and generates an emotional state signal carrying a credibility marker.

[0093] The teaching intervention strategy acquisition module 3 generates teaching intervention strategies based on the emotional state signals and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.

[0094] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps:

[0095] Geometric features of facial key points are extracted from the facial video data, acoustic features are extracted in real time from the speech audio data, and preset behavioral events are identified from the interactive text data to establish a multimodal feature recognition result. The multimodal feature recognition result and the current learning task context information are input into a multimodal fusion model to obtain an initial emotional state probability distribution. Signal consistency recognition is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution. The initial emotional state probability distribution is compared with a personalized emotional baseline established for the learner, and an emotional state deviation with a credibility score is output as the emotional state signal.

[0096] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps:

[0097] The geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into pre-trained facial monomodal emotion recognition models, speech monomodal emotion recognition models, and interactive text monomodal emotion recognition models, respectively, to obtain facial emotion probability distributions, speech emotion probability distributions, and interactive text emotion probability distributions. The divergence between each pair of the facial emotion probability distribution, the speech emotion probability distribution, the interactive text emotion probability distribution, and the initial emotion state probability distribution is calculated to generate a divergence matrix. The divergence matrix is ​​weighted and the calculation result is mapped to the credibility score, where the larger the divergence value, the lower the credibility score.

[0098] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps:

[0099] A personalized emotional baseline is established for the learner, which includes the probability distribution of the learner's typical emotional states in different learning task contexts; the statistical distance between the initial emotional state probability distribution and the probability distribution of the typical emotional states corresponding to the current learning task context is calculated to obtain an emotional state deviation value; the emotional state deviation value is bound to the credibility score to generate the emotional state deviation signal.

[0100] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps:

[0101] The personalized emotion baseline is dynamically updated by means of the following mechanism: collecting the probability distribution of the learner's emotional state in different learning task contexts when the confidence score is greater than a preset confidence threshold, and using the cluster center of the emotional state probability distribution to update the corresponding typical emotional state probability distribution in different learning task contexts.

[0102] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps:

[0103] A set of predefined teaching intervention rules is provided, wherein the triggering condition of each rule depends on the joint output of multiple logical judgment units. The input of the logical judgment units includes the deviation of the emotional state with a credibility score and the learning task context parameters. The intervention strategy analysis of the corresponding rule is triggered and the corresponding first teaching intervention strategy is generated only when all logical judgment units in the triggering condition of any one of the teaching intervention rules are satisfied.

[0104] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps:

[0105] The teaching intervention rules include at least the anxiety state rule and the immersion and fluency state rule; when the anxiety state rule is triggered, the cognitive scaffolding mode is activated to decompose the problem-solving path and implant growth mindset feedback; when the immersion and fluency state rule is triggered, interdisciplinary challenge questions are dynamically generated.

[0106] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps:

[0107] A sequence of emotional state deviation with credibility scores and a sequence of learning task context are collected and input into the temporal prediction model, which is trained based on LSTM, with Dropout used before all layers of the LSTM. The temporal prediction model performs T forward propagation and outputs T prediction results. The average of the T prediction results is taken to obtain the inflection point risk probability. The standard deviation of the T prediction results is calculated to obtain the uncertainty estimate. The knowledge density is adjusted by combining the inflection point risk probability and the uncertainty estimate to generate a second teaching intervention strategy.

[0108] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps:

[0109] If the inflection point risk probability is greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, a teaching intervention strategy to reduce knowledge density is generated, including breaking down problem-solving steps and embedding review materials; if the inflection point risk probability is less than the preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, the current knowledge density is maintained, and a help prompt message to dynamically increase knowledge density is generated, awaiting user response; if the uncertainty estimate is greater than the preset uncertainty threshold, the current knowledge density is maintained, and an optional help prompt message is generated, awaiting user response.

[0110] The dynamic explanation and practice system for learning emotion recognition provided in this embodiment of the invention can execute the dynamic explanation and practice method for learning emotion recognition provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0111] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0112] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A dynamic explanation and practice method for learning emotion recognition, characterized by: include: Real-time collection of learners' multimodal data during the learning process, including facial video data, voice and audio data, and interactive text data; Based on the multimodal data, emotion state recognition is performed to generate an emotion state signal carrying a credibility marker; Based on the emotional state signals and their credibility, and combined with the learner's learning task context, a teaching intervention strategy is generated through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model. Among them, white-box decision-making based on preset rules includes: A predefined set of teaching intervention rules is provided, wherein the triggering condition of each rule depends on the joint output of multiple logical judgment units, and the input of the logical judgment units includes the deviation of the emotional state with a credibility score and the learning task context parameters. The intervention strategy analysis of the corresponding rule is triggered and the corresponding first teaching intervention strategy is generated only when all logical judgment units in the triggering condition of any one of the teaching intervention rules are satisfied. The analysis of intervention strategies that trigger the corresponding rules includes: The teaching intervention rules include at least the anxiety state rule and the immersion and fluency state rule; When the aforementioned anxiety state rule is triggered, the cognitive scaffolding mode is activated to decompose the problem-solving path and implant growth mindset feedback. When the immersive and fluid state rule is triggered, interdisciplinary challenge questions are dynamically generated; Among them, black-box decision-making based on time-series prediction models includes: The emotional state deviation sequence with credibility score and the learning task context sequence are collected and input into the temporal prediction model, wherein the temporal prediction model is trained based on LSTM and Dropout is used before all layers of LSTM. The time-series prediction model performs T forward propagations and outputs T prediction results. The inflection point risk probability is obtained by averaging the T prediction results. The standard deviation of the T prediction results is calculated to obtain the uncertainty estimate; By combining the inflection point risk probability and the uncertainty estimate, the knowledge density is adjusted to generate a second teaching intervention strategy.

2. The dynamic explanation and practice method for learning emotion recognition as described in claim 1, characterized in that, Based on the multimodal data, emotion state recognition is performed to generate an emotion state signal carrying a confidence level label, including: The facial key point geometric features are extracted from the facial video data, acoustic features are extracted in real time from the voice audio data, and preset behavioral events are identified from the interactive text data to establish multimodal feature recognition results. The multimodal feature recognition results and the current learning task context information are input into the multimodal fusion model to obtain the initial emotional state probability distribution; Signal consistency identification is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution; The initial emotional state probability distribution is compared with the personalized emotional baseline established for the learner, and the emotional state deviation with a credibility score is output as the emotional state signal.

3. The dynamic explanation and practice method for learning emotion recognition as described in claim 2, characterized in that, Signal consistency identification is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution, including: The geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into the pre-trained facial monomodal emotion recognition model, speech monomodal emotion recognition model, and interactive text monomodal emotion recognition model, respectively, to obtain the facial emotion probability distribution, speech emotion probability distribution, and interactive text emotion probability distribution. Calculate the divergence between each pair of the facial emotion probability distribution, the voice emotion probability distribution, the interactive text emotion probability distribution, and the initial emotion state probability distribution to generate a divergence matrix; The divergence matrix is ​​weighted and the result is mapped to the credibility score, where the larger the divergence value, the lower the credibility score.

4. The dynamic explanation and practice method for learning emotion recognition as described in claim 2, characterized in that, The initial emotional state probability distribution is compared with a personalized emotional baseline established for the learner, and the emotional state deviation with a credibility score is output as the emotional state signal, including: Obtain a personalized emotional baseline established for the learner, the personalized emotional baseline including the probability distribution of the learner's typical emotional states in different learning task contexts; Calculate the statistical distance between the initial emotional state probability distribution and the typical emotional state probability distribution corresponding to the current learning task context to obtain the emotional state deviation value; The emotional state deviation value is bound to the credibility score to generate the emotional state deviation signal.

5. The dynamic explanation and practice method for learning emotion recognition as described in claim 4, characterized in that, The personalized emotion baseline is dynamically updated, and the update mechanism is as follows: Collect the probability distribution of the learner's emotional state in different learning task contexts when the credibility score is greater than a preset credibility threshold, and use the cluster center of the emotional state probability distribution to update the corresponding typical emotional state probability distribution in different learning task contexts.

6. The dynamic explanation and practice method for learning emotion recognition as described in claim 1, characterized in that, By combining the inflection point risk probability and the uncertainty estimate, knowledge density is adjusted to generate a second teaching intervention strategy, including: If the inflection point risk probability is greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, a teaching intervention strategy to reduce knowledge density is generated, including decomposing problem-solving steps and embedding review materials. If the inflection point risk probability is less than a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, then the current knowledge density is maintained, and a help prompt message to dynamically increase the knowledge density is generated, waiting for the user's response. If the estimated uncertainty value is greater than the preset uncertainty threshold, the current knowledge density is maintained, and an optional help prompt is generated, waiting for the user's response.

7. A dynamic explanation and practice system for learning emotion recognition, characterized in that: The system is used to implement the dynamic explanation and practice method for learning emotion recognition according to any one of claims 1-6, the system comprising: The multimodal data acquisition module is used to collect learners' multimodal data in real time during the learning process, including facial video data, voice and audio data, and interactive text data; The emotional state signal acquisition module identifies emotional states based on the multimodal data and generates emotional state signals carrying confidence markers. The teaching intervention strategy acquisition module generates teaching intervention strategies based on the emotional state signals and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.

Citation Information

Patent Citations

  • Psychological accompanying method based on multi-modal emotion recognition

    CN120690390A

  • Multi-scene self-adaptive man-machine interaction system and method based on emotion recognition

    CN121255014A