Dynamic explanation partner training method and system for learning emotion recognition
By collecting learners' multimodal data and combining it with a dual-core decision-making mechanism, the problems of inaccurate emotion recognition and rigid intervention strategies in existing learning support tools have been solved. This has enabled precise adaptation to learners' emotional states and personalized intervention, thereby improving the effectiveness of learning assistance.
Patent Information
- Application Number
- CN202610039170.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2046-01-13
AI Technical Summary
Existing learning support tools are one-sided in their ability to capture emotions, lack personalized intervention, and cannot predict learning setbacks, resulting in inaccurate emotion recognition results and rigid intervention strategies that fail to meet the needs of personalized teaching.
By collecting learners' multimodal data in real time, including facial videos, audio recordings, and interactive text, and combining this with a dual-core decision-making mechanism to generate teaching intervention strategies, the system achieves accurate identification of emotional state signals and personalized intervention. The dual-core decision-making mechanism includes white-box decision-making and black-box decision-making. White-box decision-making is based on preset rules, while black-box decision-making is based on a time-series prediction model.
It achieves precise matching of learners' emotional states, avoids learning setbacks in advance, improves the personalization and accuracy of learning accompaniment, and enhances the effectiveness of learning assistance.
Smart Images

Figure CN121502701A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction education, in particular to a dynamic explanation and practice method and system for learning emotion recognition. BACKGROUND
[0002] In learning practice, accurate capture and adaptive intervention of the emotional state of learners are crucial to improving learning efficiency, and the matching degree of emotion recognition results and teaching strategies directly affects the learning effect. The existing technology mainly uses single-dimensional data for emotion judgment, relies on fixed intervention logic, lacks multi-source information fusion capability in data processing, and also fails to achieve personalized adaptation in human-computer interaction. These methods can play a certain role in simple learning scenarios, but as the demand for personalized learning increases, they expose obvious shortcomings when applied to dynamic learning scenarios. Traditional detection methods lack accurate data processing mechanisms and adaptive human-computer interaction modes, and cannot comprehensively capture multi-modal emotional data such as faces, voices, and texts, resulting in inaccurate emotion recognition results, rigid intervention strategies, and difficulty in obtaining data to support personalized teaching, which cannot meet the needs of accurate evaluation and efficient control of learning practice. SUMMARY
[0003] The present application provides a dynamic explanation and practice method and system for learning emotion recognition, which solves the technical problems of one-sided emotion capture, lack of personalized intervention, and inability to predict learning setbacks in existing learning practice tools.
[0004] In a first aspect of the present application, a dynamic explanation and practice method for learning emotion recognition is provided, which includes: real-time collection of multi-modal data of learners during the learning process, including facial video data, voice audio data, and interactive text data; emotion state recognition based on the multi-modal data to generate an emotion state signal carrying a confidence mark; generation of a teaching intervention strategy through a dual-core decision mechanism based on the emotion state signal and confidence, combined with the learning task context of the learner, the dual-core decision mechanism including white-box decision based on preset rules and black-box decision based on a time series prediction model.
[0005] In a second aspect of the present application, a dynamic explanation and practice system for learning emotion recognition is provided, which includes: a multi-modal data acquisition module for real-time collection of multi-modal data of learners during the learning process, including facial video data, voice audio data, and interactive text data; an emotion state signal acquisition module for emotion state recognition based on the multi-modal data to generate an emotion state signal carrying a confidence mark; a teaching intervention strategy acquisition module for generating a teaching intervention strategy through a dual-core decision mechanism based on the emotion state signal and confidence, combined with the learning task context of the learner, the dual-core decision mechanism including white-box decision based on preset rules and black-box decision based on a time series prediction model.
[0006] The one or more technical solutions provided in the present application have at least the following technical effects or advantages: The present application acquires an emotional state signal by collecting facial video, speech audio, interactive text and other multi-modal data in the learning process of learners in real time, through multi-modal feature extraction, fusion model operation, credibility score allocation, personalized emotional baseline comparison and other processing, calculates emotional state deviation, inflection point risk probability and uncertainty estimate value, dynamically adjusts teaching intervention strategies by combining white box decision and black box decision of the double-core decision mechanism, thereby accurately adapting to different emotional states of learners such as anxiety and immersion, avoiding learning setbacks in advance, making the personalized adaptation and precise intervention effect of learning accompaniment more reliable and efficient, and achieving the technical effect of making personalized learning accompaniment more suitable for the needs of learners and improving learning assistance accuracy and effectiveness. BRIEF DESCRIPTION OF DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0008] Figure 1 is a flowchart of a dynamic explanation and accompaniment method for learning emotion recognition provided by the embodiments of the present application.
[0009] Figure 2 is a structural schematic diagram of a dynamic explanation and accompaniment system for learning emotion recognition provided by the embodiments of the present application.
[0010] Explanation of reference signs: multi-modal data acquisition module 1, emotional state signal acquisition module 2, teaching intervention strategy acquisition module 3. DETAILED DESCRIPTION
[0011] The present application provides a dynamic explanation and accompaniment method and system for learning emotion recognition, which solves the technical problems of one-sided emotion capture, lack of personalized intervention and inability to predict learning setbacks of existing learning accompaniment tools.
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the present application.
[0013] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.
[0014] Embodiment one, as shown, a dynamic explanation and practice method for learning emotion recognition, wherein the method comprises: Figure 1 Real-time collection of multi-modal data of learners during the learning process, including facial video data, speech audio data and interactive text data.
[0015] Specifically, first, a multi-modal data collection environment suitable for the learning scene is built, and hardware devices with real-time capture capability are deployed, such as mobile phones, smart tablets, education-specific smart terminals or computer accessory device combinations, including high-definition cameras for collecting facial pictures, high-fidelity microphones for recording speech, and terminal devices for carrying out interactive operations of learners, while completing the adaptive connection of the collection device and the data processing system to ensure smooth data transmission channels.
[0016] For facial video data, start the high-definition camera to capture the facial dynamic picture of the learner during the learning process in real time, maintain the picture frame rate stable during the collection process to record the facial expression changes completely, process the collected continuous video frames, extract the geometric features of the facial key points frame by frame, and obtain the basic data that can reflect the facial emotional state of the learner.
[0017] For speech audio data, turn on the high-fidelity microphone to continuously record the speech information of the learner during the learning process, covering various speech contents such as answer speech and question expression, perform noise reduction preprocessing on the real-time collected audio signal to remove environmental noise interference, and then extract the acoustic features in the speech in real time to form the speech basic data that can be used for subsequent emotion recognition.
[0018] For interactive text data, the text input content of the learner is recorded in real time through the learning interactive terminal, including question answering text, question consulting text, etc., and the time node, content type, etc. of the text input are recorded synchronously, the collected interactive text is analyzed, the preset learning behavior events are identified, and the collection and preliminary arrangement of the interactive text data are completed.
[0019] The three types of data are collected synchronously, and during the collection process, the various types of data are transmitted to a backend storage module in real time to form a complete learner multi-modal raw data set, providing comprehensive and continuous data support for subsequent emotional state recognition.
[0020] An emotional state signal carrying a confidence score is generated based on the multi-modal data.
[0021] In the embodiments of the present application, the emotional state signal is an objectified evaluation result obtained by combining the multi-dimensional performance of the learner's facial expression, speech tone, interactive text, etc., and reflecting the deviation of the current emotion from the typical emotional bias of the learner, with a confidence score.
[0022] Optionally, first, facial key point geometric features, acoustic features, and preset behavior events are extracted from the three types of multi-modal data, i.e., facial video, speech audio, and interactive text, to construct multi-modal feature recognition results. Then, the results and the current learning task context information are input into a multi-modal fusion model to obtain an initial emotional state probability distribution. Subsequently, signal consistency recognition is performed on the multi-modal data to assign a confidence score to the initial emotional state probability distribution. Finally, the initial emotional state probability distribution is compared with the learner's personalized emotional baseline, and an emotional state deviation degree with a confidence score is output as the emotional state signal.
[0023] Based on the emotional state signal and the confidence score, a teaching intervention strategy is generated by a dual-core decision mechanism in combination with the learning task context of the learner, the dual-core decision mechanism including a white-box decision based on preset rules and a black-box decision based on a time series prediction model.
[0024] In the embodiments of the present application, the white-box decision is a pre-defined teaching intervention rule including an anxiety state rule, an immersive and fluent state rule, etc., and the emotional state deviation degree with a confidence score and the learning task context parameters are input as logical judgment units. When all the judgment units are satisfied, the corresponding analysis is triggered and a first teaching intervention strategy is generated. The black-box decision is a time series prediction model based on LSTM training and all layers are pre-processed with Dropout. The emotional state deviation degree sequence with a confidence score and the learning task context sequence are input into the model, and after multiple forward propagations, a turning point risk probability and an uncertainty estimate value are obtained. The two are combined to adjust the knowledge density to generate a second teaching intervention strategy.
[0025] In an embodiment of the present application, the white-box decision method first defines a set of teaching intervention rules, and the triggering condition of each rule depends on the joint output of multiple logical judgment units, and the inputs of these logical judgment units include the emotion state deviation degree with credibility score and learning task context parameters. Only when all the logical judgment units in the triggering condition of a certain teaching intervention rule are met, the intervention strategy analysis of the corresponding rule is triggered, and then the corresponding first teaching intervention strategy is generated.
[0026] Next, the black-box decision method collects the emotion state deviation degree sequence with credibility score and the learning task context sequence, and inputs them into the time series prediction model trained based on LSTM and using Dropout before all layers. After the model outputs T times of prediction results through T times of forward propagation, the inflection point risk probability is obtained by averaging the results, and the uncertainty estimate value is obtained by calculating the standard deviation of the results, and finally the knowledge density is adjusted combined with the two values, and then the second teaching intervention strategy is generated.
[0027] Further, the method provided in the embodiments of the present application comprises: extracting facial key point geometric features from the facial video data, extracting acoustic features from the speech audio data in real time, identifying preset behavior events from the interactive text data, and establishing a multi-modal feature recognition result; inputting the multi-modal feature recognition result and the current learning task context information into a multi-modal fusion model to obtain an initial emotion state probability distribution; performing signal consistency recognition on the multi-modal data to assign a credibility score to the initial emotion state probability distribution; comparing the initial emotion state probability distribution with the personalized emotion baseline established for the learner to output an emotion state deviation degree with a credibility score as the emotion state signal.
[0028] Specifically, first, feature extraction operations are carried out on the three types of multi-modal data respectively to form a multi-modal feature recognition result. For facial video data, continuous video frames are processed frame by frame to accurately extract facial key point geometric features such as coordinates, distances, and angles of key regions such as eyes, eyebrows, and mouth; for speech audio data, the original audio signal is first preprocessed by noise reduction, and then acoustic features such as fundamental frequency, speech rate, volume, and spectrum are extracted in real time; for interactive text data, the answer content and consultation statements input by the learner are processed to identify preset behavior events such as frequent questioning, long time without text input, and incorrect answers, and the three types of features extracted are integrated to establish a complete multi-modal feature recognition result.
[0029] Subsequently, a multi-modal fusion model is constructed and trained as a neural network model for outputting an initial emotional state probability distribution. In model construction, a neural network architecture including an input layer, a feature fusion layer, a hidden layer, and an output layer is built, and the input layer is set to have independent input channels corresponding to the three types of features described above, wherein the face key point geometric feature input channel dimension is set to 128, the speech acoustic feature input channel dimension is set to 64, and the interactive text behavior event feature input channel dimension is set to 32; the feature fusion layer adopts a 4-head multi-head attention mechanism to weight and fuse the three types of features, and the feature dimension of each head of attention is uniformly mapped to 64; the hidden layer is set to have two fully connected networks, the first layer has 256 neurons, and the second layer has 128 neurons, both layers have ReLU activation functions to improve the model fitting capability, and a Dropout layer is added between the two fully connected networks with a dropout probability of 0.3 to prevent overfitting; the output layer is set to have nodes corresponding to various emotional types such as anxiety, immersion, calmness, irritability, and confusion, and a Softmax activation function is used to output probability values of each type of emotion.
[0030] In model training, a dataset containing 100,000 valid samples is selected as training samples, which cover face key point geometric features, speech acoustic features, and interactive text behavior event features of learners in different learning stages and subjects in various learning scenarios, and each sample is labeled with a corresponding true emotional label. The training set, validation set, and test set are divided in a ratio of 7:1:2; the random gradient descent algorithm is used to iteratively adjust the network weight and bias parameters, the initial learning rate is set to 0.001, the learning rate is decayed to 0.5 of the original value every 10 iterations, and the batch size is set to 64; the cross-entropy loss function is selected as the model optimization objective, and the model is continuously iteratively trained until the validation set loss function value no longer decreases for 5 consecutive times, and the emotional recognition accuracy of the model on the test set reaches the preset standard of 90%, and the model training is completed. The multi-modal feature recognition results established in the above steps are input into the trained multi-modal fusion model together with the context information such as the difficulty of the current learning task, the subject type, and the learning stage, and the model is operated to output an initial emotional state probability distribution including the proportion of different emotional types.
[0031] Next, the face key point geometric feature, acoustic feature, and interactive behavior event recognition results are respectively input into the corresponding pre-trained single-modal emotion recognition model to obtain three types of single-modal emotion probability distributions, and the divergence between each two of the three types of distributions and the initial emotional state probability distribution is calculated to generate a divergence matrix. Then, the divergence matrix is weighted and calculated, and the result is mapped to a credibility score. The greater the divergence value, the lower the corresponding credibility score. This step is described in detail in the subsequent content.
[0032] Finally, in generating the emotional state signal, a personalized emotional baseline containing the typical emotional state probability distribution of the learner in different learning task contexts is obtained first, then the statistical distance between the initial emotional state probability distribution and the typical emotional state probability distribution corresponding to the current learning task context is calculated to obtain the emotional state deviation value, and finally the value is bound with the credibility score to generate the emotional state deviation signal. This step is described in detail in the subsequent content.
[0033] Further, the method provided by the embodiment of the present application comprises: The facial key point geometric features, acoustic features, and interactive behavior event recognition results are input into the pre-trained facial single-modal emotion recognition model, speech single-modal emotion recognition model, and interactive text single-modal emotion recognition model respectively to obtain facial emotion probability distribution, speech emotion probability distribution, and interactive text emotion probability distribution; the divergences between the facial emotion probability distribution, speech emotion probability distribution, interactive text emotion probability distribution, and two of the initial emotional state probability distributions are calculated to generate a divergence matrix; the divergence matrix is subjected to weighted calculation, and the calculation result is mapped to the credibility score, wherein the greater the divergence value, the lower the credibility score.
[0034] Optionally, first, the facial single-modal emotion recognition model is constructed and trained. The model adopts a CNN architecture, the dimension of the input layer is set to 128, and the dimension specification of the facial key point geometric features is adapted; a feature extraction network containing 2 convolution layers is constructed, the first layer has a convolution kernel size of 3x3 and a number of 64, and the second layer has a convolution kernel size of 3x3 and a number of 128, a maximum pooling layer is connected after each convolution, and the pooling window is 2x2; a hidden layer is set to a fully connected network with a neuron number of 256, a ReLU activation function is introduced, and a dropout layer is added to prevent overfitting, with a dropout probability of 0.25; the number of nodes of the output layer is set to correspond to the number of emotions such as anxiety, immersion, calmness, restlessness, and confusion, and a Softmax activation function is adopted. Model training selects tens of thousands of facial feature samples, each sample is labeled with a corresponding emotion label, the training set, validation set, and test set are divided according to 7:1:2, an Adam optimizer is adopted, the initial learning rate is 0.001, the batch size is set to 32, and the model is iteratively trained for 50 rounds until the emotional recognition accuracy of the model on the test set reaches 90%, and the training is completed. The model input is the facial key point geometric features, and the output is the facial emotion probability distribution.
[0035] Next, a monomodal speech emotion recognition model was constructed and trained, employing a combined CNN and RNN architecture. The input layer dimension was set to 64 dimensions to match the dimensions of the speech acoustic features. The feature extraction part consisted of two convolutional layers, each with a 3×3 kernel size and 32 and 64 kernels respectively, followed by a bidirectional RNN layer with 128 hidden units to capture the temporal features of the speech signal. The hidden layer was a fully connected network with 128 neurons, incorporating the ReLU activation function. The output layer was consistent with the monomodal facial emotion recognition model, setting multiple emotion nodes and using the Softmax activation function. During training, tens of thousands of speech acoustic feature samples were selected, covering speech data with different speaking speeds and volumes. The ratio of training, validation, and test sets was 7:1:2. The Adam optimizer was used with an initial learning rate of 0.001, a batch size of 32, and 50 iterations. Training stopped when the validation set loss no longer decreased for three consecutive iterations and the test set accuracy reached 90%. The model's input is the speech acoustic features, and its output is the speech emotion probability distribution.
[0036] Then, an interactive text unimodal emotion recognition model was constructed and trained using the TextCNN architecture. The input layer transforms the interactive text behavioral event features into 64-dimensional vectors. The feature extraction layer uses three different sizes of convolutional kernels: 2×64, 3×64, and 4×64, with 32 kernels of each size, to extract semantic features of different lengths. After convolution, the features are processed by a max-pooling layer. The hidden layer is a single fully connected network with 128 neurons, incorporating the ReLU activation function and adding a dropout layer with a dropout probability of 0.3. The output layer follows a 5-node design and uses the Softmax activation function. Tens of thousands of interactive text feature samples were selected for training. The training, validation, and test sets were divided in a 7:1:2 ratio. The Adam optimizer was used with an initial learning rate of 0.001, a batch size of 32, and 50 iterations of training. Training was considered complete when the test set accuracy reached 90%. The model's input is the interactive text behavioral event recognition result, and its output is the interactive text emotion probability distribution.
[0037] The preprocessed facial key point geometric features, acoustic features, and interactive behavior event recognition results are then input into the corresponding trained single-modal emotion recognition models. After the models perform their respective calculations, the probability distributions of facial emotions, speech emotions, and interactive text emotions, which include the proportions of the five emotion types, are obtained.
[0038] Next, the KL divergence calculation method is used to calculate the divergence between the four types of probability distributions. The KL divergence calculation method is suitable for probability distribution difference quantization scenarios. When calculating, two types of probability distributions are input, and the divergence value reflecting the distribution difference is calculated through the formula. Six groups of divergence are calculated, including the first divergence between the facial emotion probability distribution and the speech emotion probability distribution, the second divergence between the facial emotion probability distribution and the interactive text emotion probability distribution, the third divergence between the facial emotion probability distribution and the initial emotion state probability distribution, the fourth divergence between the speech emotion probability distribution and the interactive text emotion probability distribution, the fifth divergence between the speech emotion probability distribution and the initial emotion state probability distribution, and the sixth divergence between the interactive text emotion probability distribution and the initial emotion state probability distribution. A 6x6 divergence matrix is constructed based on the six groups of divergence values.
[0039] Finally, the divergence matrix is calculated by weighting. According to the importance of the three types of single-modal data in emotion recognition, weights are assigned to the six groups of divergence. The weighted total score is obtained by multiplying each group of divergence value by its corresponding weight and summing. The linear mapping method is used to map the weighted total score to a reliability score of 0-100. The maximum value of the weighted total score corresponds to 0, the minimum value corresponds to 100, and the intermediate values are converted in a linear proportion. The mapping relationship between the larger divergence value and the lower reliability score is realized.
[0040] Through the series of steps of constructing a single-modal model with specific parameters, calculating the divergence between two distributions by KL divergence, and weighting and mapping the score, the reliability of the initial emotion state probability distribution is accurately quantified, providing reliable data support for the generation of subsequent teaching intervention strategies.
[0041] Further, the method provided in the embodiments of the present application comprises: obtaining a personalized emotion baseline established for the learner, the personalized emotion baseline comprising a typical emotion state probability distribution of the learner in different learning task contexts; calculating a statistical distance between the initial emotion state probability distribution and a typical emotion state probability distribution corresponding to a current learning task context, to obtain an emotion state deviation value; and binding the emotion state deviation value with the reliability score to generate an emotion state deviation signal.
[0042] Specifically, first, the typical emotion state probability distribution is defined for a specific learning task context, which is structured data quantifying the regular emotional performance of the learner, containing probability values of various emotion types such as anxiety, immersion, calmness, restlessness, and confusion, and the sum of the probability proportions of all emotion types is 1, which can accurately reflect the inherent emotional tendency of the learner in the task scenario.
[0043] Then the personalized emotion baseline is obtained, which is a core reference basis continuously constructed and dynamically updated based on the learning data of the learner in the past, and contains the typical emotion state probability distribution of the learner under different learning task contexts (covering subject types, task difficulties, learning stages, etc.). The backend database uses a hierarchical storage structure to manage baseline data, with the unique identifier of the learner as the top-level classification index, and the index is divided into a first-level subdirectory by subject type, such as mathematics, Chinese, and English. Each first-level subdirectory is further divided into a second-level subdirectory by task difficulty (basic, medium, advanced) and learning stage (preparation, classroom practice, after-school review, stage test). Each second-level subdirectory stores the typical emotion state probability distribution data file under the corresponding scene.
[0044] Then the unique identifier of the learner is associated with the account currently logged in by the learner, the complete personalized emotion baseline data set corresponding to the identifier is called, and then the specific context parameters of the current learning task are extracted, including the specific subject type, task difficulty level, and learning stage information. These parameters are precisely matched with the directory levels in the data set one by one, and the typical emotion state probability distribution file that completely fits the current task scene is filtered out. If there are multiple historical emotion data records under the same task scene, the average value of the probability values of the corresponding emotion types in each record is taken to form a comprehensive typical emotion state probability distribution, which is used as the reference data for subsequent comparison with the initial emotion state probability distribution.
[0045] Then the Euclidean distance is calculated to obtain the emotion state deviation value. First, the dimensions of the initial emotion state probability distribution and the filtered typical emotion state probability distribution are determined to ensure that the number and arrangement of emotion types in the two distributions are completely consistent, and the probability values of each emotion type in each distribution form a corresponding dimension vector. The probability values at the same position in the two vectors are calculated one by one, the difference between the corresponding elements is calculated first, then all the differences are squared and summed, and finally the square root of the sum is taken to obtain the emotion state deviation value between the initial emotion state and the typical emotion state. The larger the value, the more significant the difference between the current emotion of the learner and the typical emotion of the learner.
[0046] Finally, the emotion state deviation value and the credibility score are bound. First, a fixed information combination format is determined, with the emotion state deviation value as the core data and the credibility score as the auxiliary verification information, and the two are integrated in the form of character separation. After integration, a structured emotion state deviation signal is generated, which contains not only the core value reflecting the degree of learner's current emotional abnormality, but also the credibility score measuring the reliability of the value, providing complete and accurate data support for the subsequent dual-core decision mechanism to generate teaching intervention strategies.
[0047] Further, the method provided by the embodiment of the present application comprises: The personalized emotion baseline is dynamically updated, and the updating mechanism is: collecting the emotion state probability distribution of the learner in different learning task contexts under the condition that the credibility score is greater than the preset credibility threshold, and using the cluster center of the emotion state probability distribution to update the corresponding typical emotion state probability distribution in different learning task contexts.
[0048] Specifically, first, a preset credibility threshold is set, which is set to 80 points in combination with the actual needs of the emotion recognition scene and the industry conventional standard. The threshold is used as the core reference for screening effective emotion data, for eliminating emotion state probability distributions with low credibility and small reference value. Subsequently, the data collection process is started, and the emotion state probability distribution of the learner in various different learning tasks is real-time captured through the backend data acquisition module, while the corresponding credibility score and learning task context information are associated, wherein the learning task context covers key parameters such as subject type, task difficulty, and learning stage. The emotion state probability distribution data with a credibility score greater than 80 points is extracted from the stored data, and is classified and arranged according to the learning task context, forming multiple groups of effective emotion data sets adapted to different scenes.
[0049] Then, the classified effective emotion data sets are respectively subjected to clustering processing using the K-means clustering algorithm. Before clustering, the number of clusters corresponding to each learning task context is determined to be 1, so as to ensure that only one core cluster is formed in each scene, focusing on reflecting the typical emotion characteristics of the learner in the scene. When performing the clustering operation, first, an emotion state probability distribution in the group of data is randomly selected as the initial cluster center, then the Euclidean distance between each emotion state probability distribution in the group and the initial cluster center is calculated, all data are classified into a cluster, then the average value of the probability values of each emotion type in all emotion state probability distributions in the cluster is calculated through iteration, the cluster center is updated, and the iteration process is repeated until the change amplitude of the cluster center is less than the preset convergence threshold 0.001, the iteration is stopped, and the final cluster center is determined.
[0050] Finally, the final cluster center obtained through clustering in each learning task context is directly used as the updated typical emotion state probability distribution in the scene. The new typical emotion state probability distribution is used to replace the old data stored in the backend database under the corresponding learning task context, and the update time and the data change log before and after the update are recorded, so as to facilitate subsequent tracing and verification. For the newly added learning task context, if the effective emotion data set thereof meets the clustering condition, the corresponding typical emotion state probability distribution is generated through the same process as above, and is supplemented to the personalized emotion baseline, so as to perfect the scene coverage range of the baseline.
[0051] By setting a credible threshold to screen effective data, using a K-means clustering algorithm to determine the clustering center and updating the baseline data, the dynamic optimization of the personalized emotion baseline is realized, ensuring that the baseline can continuously fit the emotional performance characteristics of the learners, and improving the accuracy of subsequent emotion state recognition and deviation calculation.
[0052] Further, the method provided by the embodiments of the present application comprises: A set of teaching intervention rules are predefined, wherein the trigger condition of each rule depends on the joint output of multiple logical judgment units, and the input of the logical judgment units includes the emotion state deviation with a credibility score and learning task context parameters; when and only when all the logical judgment units in the trigger condition of any one of the teaching intervention rules are satisfied, the intervention strategy analysis of the corresponding rule is triggered to generate the corresponding first teaching intervention strategy.
[0053] In the embodiments of the present application, the logical judgment unit is a digital circuit or a computer system component constructed by a basic logic gate, which can perform logical operations on input data and output Boolean results of true or false, providing a basis for system decision or control flow.
[0054] Specifically, a set of teaching intervention rules are first predefined, including rules for anxiety state and rules for immersion and fluency state. For the anxiety state rule, the trigger condition depends on the joint output of multiple logical judgment units, and the input of the logical judgment units includes the emotion state deviation with a credibility score and learning task context parameters. The specific conditions can refer to the first rule example, i.e., the emotion deviation is greater than a threshold , the credibility is greater than a threshold , the number of consecutive errors is greater than or equal to 2 and the current thinking / residence time is greater than a threshold When these conditions are all met, the cognitive scaffolding mode is started, the problem solving path is decomposed, and growth type thinking feedback is provided.
[0055] For the immersion and fluency state rule, the trigger condition depends on the joint output of multiple logical judgment units, and the input is the emotion state deviation with a credibility score and learning task context parameters. Referring to the second rule example, the emotion state (immersion / fluency) deviation is greater than a threshold , the current task completion accuracy is greater than 95% and the task completion speed is greater than the average speed, when these conditions are all met, a cross-disciplinary challenge question is dynamically generated to stimulate exploration potential.
[0056] Subsequently, each logical judgment unit compares the emotion state deviation with a credibility score with a preset threshold, and matches the learning task context parameters (such as task type, difficulty, etc.) with the rule conditions, and outputs a Boolean value of whether it is satisfied or not.
[0057] When all the logical judgment units in a certain teaching intervention rule output satisfaction, the intervention strategy analysis of the corresponding rule is triggered, which is described in detail in the subsequent content. For example, when the anxiety state rule is satisfied, the system decomposes the problem solving path according to the pre-defined process and generates growth-type thinking feedback; when the immersion flow state rule is satisfied, the system dynamically constructs a cross-disciplinary challenge task in combination with the subject attribute and knowledge node.
[0058] Finally, the corresponding first teaching intervention strategy is generated, such as pushing the decomposed problem solving steps, displaying the growth-type thinking prompt, or presenting the cross-disciplinary challenge question, to realize the intervention on the learner's emotional state and assist them in adjusting their learning state.
[0059] Further, the method provided by the embodiments of the present application comprises: The teaching intervention rules at least include an anxiety state rule and an immersion flow state rule; when the anxiety state rule is triggered, a cognitive scaffolding mode is started, the problem solving path is decomposed, and growth-type thinking feedback is implanted; when the immersion flow state rule is triggered, a cross-disciplinary challenge question is dynamically generated.
[0060] In one embodiment, the core components of the teaching intervention rules are first defined, a structured teaching intervention rule library is constructed, and the anxiety state rule and the immersion flow state rule are included as core entries in the library. An independent identification code is set for each rule in the rule library, the corresponding trigger condition parameters and execution instructions are associated, and the different needs of the two rules for adjusting the negative emotions of learners and expanding the positive learning state are clarified, thereby providing stable data support for the subsequent precise triggering of intervention strategies.
[0061] When the anxiety state rule is triggered, the cognitive scaffolding mode is started to disassemble the problem solving path. The core difficulty of the current problem solving task and the associated basic concepts are located through the subject knowledge point extraction technology, the complete problem solving path is disassembled into 3-5 executable small steps according to the cognitive levels of basic cognition-depth understanding-practice application with the core difficulty as the anchor point, each step is clear about the corresponding knowledge point, problem solving method and operation key points, and a visual problem solving path diagram is generated to clearly present the progressive relationship between the steps and reduce the cognitive load of the learners.
[0062] Synchronously, growth-type thinking feedback is implanted by matching the growth-type feedback template. A growth-type thinking feedback template library covering different problem solving scenarios is constructed in advance, the template library includes standardized feedback texts for effort process, strategy adjustment, difficulty breakthrough and other dimensions. The system extracts key behavior data in the problem solving process of the learner, such as the number of attempts, the adopted problem solving method, the error type, etc., matches the corresponding feedback content in the template library according to the behavior data characteristics, generates personalized thinking feedback focusing on the process, and guides the learner to focus on the effort and method in the problem solving process.
[0063] When the immersive fluency state rule is triggered, a cross-disciplinary challenge question is dynamically generated by adopting cross-disciplinary knowledge graph retrieval. Relying on the pre-constructed multi-disciplinary knowledge graph, the core knowledge points in the current learning task are first extracted as the retrieval anchor point, the associated concepts and application scenarios of the knowledge point in other disciplines are retrieved through the graph neural network, and the cross-disciplinary knowledge connection is established. Then according to the Bloom's Taxonomy of Educational Objectives, the difficulty of the question is set to be slightly higher than the current level of the learner, so as to ensure that the question is in the nearest development zone of the learner. Combined with the context of the current learning task, a complete challenge question integrating multi-disciplinary knowledge is generated, which not only fits the existing knowledge reserve but also stimulates the exploration potential.
[0064] Further, the method provided by the embodiment of the application comprises: The emotion state deviation sequence with the credibility score and the learning task context sequence are collected and input into the time series prediction model, wherein all layers of the LSTM are trained based on the LSTM, and a Dropout is used before each layer of the LSTM; the time series prediction model is forward propagated T times, and T prediction results are output; the T prediction results are averaged to obtain a turning point risk probability; the standard deviation of the T prediction results is calculated to obtain an uncertainty estimation value; the turning point risk probability and the uncertainty estimation value are combined for knowledge density adjustment to generate a second teaching intervention strategy.
[0065] Optionally, a time series prediction model is first constructed and trained. Specifically, an LSTM architecture is adopted, 2 layers of LSTM are set, and the number of hidden units in each layer is 128. A Dropout layer is added before each LSTM, and the Dropout ratio is set to 0.2. Through training, part of the neurons is randomly turned off to reduce model overfitting and improve generalization ability. The output layer is a fully connected layer, and the output dimension is 1. During training, the emotion state deviation sequence with the credibility score and the learning task context sequence of the learner in the history are collected, the data set is divided according to the time window, each window contains the sequence data in the past 5 minutes as input, and the subsequent emotion change is taken as the label. The Adam optimizer is used, the learning rate is set to 0.001, the batch size is 32, and the model is iteratively trained for 200 rounds until the loss on the validation set converges. The model input is the emotion state deviation sequence with the credibility score and the learning task context sequence composed of the history and the real time, specifically including the data from the current time to the past 5 minutes, and the output is the turning point risk probability prediction value of a single forward propagation.
[0066] Then, the emotion state deviation sequence with the credibility score and the learning task context sequence from the current time to the past 5 minutes are collected, preprocessed according to the input format during model training, and input into the trained time series prediction model after ensuring that the sequence length and feature dimension are consistent with the training data.
[0067] Then T times of forward propagation are performed, for example, T = 100. Since Dropout in the model is activated during prediction, different combinations of neurons are randomly turned off each time, so each propagation outputs a prediction result that differs from the previous one, and a total of T prediction results are obtained.
[0068] Then the T prediction results are averaged, and the inflection point risk probability is obtained by dividing the sum of all results by T. This probability reflects the average risk level of the learner's emotional inflection point in the future period of time.
[0069] Subsequently, the standard deviation of the T prediction results is calculated, and the uncertainty estimate is obtained by quantifying the dispersion of the results through the standard deviation calculation formula. This value reflects the volatility of the model's prediction results and reflects the uncertainty of the prediction.
[0070] Finally, the inflection point risk probability and the uncertainty estimate are combined to adjust the knowledge density, and the second teaching intervention strategy is generated. This step is described in detail in the subsequent content.
[0071] By constructing an LSTM time series prediction model with Dropout, and through multiple forward propagations and the calculation of average risk and uncertainty, the precise prediction of the learner's emotional inflection point and the intelligent generation of intervention strategies are achieved, improving the forward-looking and accuracy of teaching intervention.
[0072] Further, the method provided by the embodiments of the present application comprises: If the inflection point risk probability is greater than or equal to the preset risk probability threshold, and the uncertainty estimate is less than or equal to the preset uncertainty threshold, a teaching intervention strategy that reduces the knowledge density is generated, including decomposing the problem-solving steps and implanting review materials; if the inflection point risk probability is less than the preset risk probability threshold, and the uncertainty estimate is less than or equal to the preset uncertainty threshold, the current knowledge density is maintained, and a help prompt information that dynamically increases the knowledge density is generated, waiting for user response; if the uncertainty estimate is greater than the preset uncertainty threshold, the current knowledge density is maintained, and a selectable help prompt information is generated, waiting for user response.
[0073] In one embodiment, the percentile method is first used to set two core preset thresholds to determine the preset risk probability threshold and the preset uncertainty threshold. The historical data of the learner's inflection point risk probability and the historical data of the uncertainty estimate in the past three months are collected, and the two sets of data are sorted in ascending order. The value corresponding to the 70th percentile of the inflection point risk probability historical data is taken as the preset risk probability threshold, which can cover most high-risk scenarios; the value corresponding to the 20th percentile of the uncertainty estimate historical data is taken as the preset uncertainty threshold, which ensures that only low-volatility data is determined to be reliable, providing clear quantitative standards for subsequent strategy judgment.
[0074] Then, the inflection point risk probability and the uncertainty estimation value are double judged, when it is detected that the inflection point risk probability is greater than or equal to a preset risk probability threshold value, and the uncertainty estimation value is less than or equal to a preset uncertainty threshold value, a teaching intervention strategy generation process of reducing knowledge density is triggered. First, the problem solving steps are decomposed by a knowledge point level decomposition method, taking the core knowledge point of the current learning task as an anchor point, following the cognitive law of basic cognition-depth understanding-practice application, the complete problem solving path is divided into 3-5 continuous executable small steps, each step is marked with the corresponding knowledge point and operation key points to reduce the understanding difficulty. Then, the associated review materials are implanted by using the associated review material pushing method, based on the core knowledge point in the current problem solving step, the associated past basic knowledge point review content in the backend knowledge base is searched, the exercises and concept analysis adapted to the current task are selected and pushed to the learner.
[0075] When the inflection point risk probability is less than the preset risk probability threshold value, and the uncertainty estimation value is less than or equal to the preset uncertainty threshold value, the current knowledge presentation density is maintained, and the explanation depth of the knowledge point, the question difficulty and the content advancing speed are not changed. At the same time, the help prompt information with dynamically increased knowledge density is made by using the prompt information template generation method, the preset prompt information template library is called, the library contains standardized prompt texts adapted to different subjects and knowledge points, according to the subject type and knowledge point attribute of the current learning task, the corresponding template is matched, the content such as “whether to try higher difficulty extension questions” is filled, the personalized help prompt information is generated, and the information is pushed to the learner interface and waits for the user response.
[0076] For the third scenario, when it is detected that the uncertainty estimation value is greater than the preset uncertainty threshold value, the current knowledge density is maintained to avoid intervention errors caused by ambiguous state judgment. The selectable help prompt information is made by using the bidirectional option prompt generation method, two standardized option texts are constructed, one is an option for increasing the knowledge density, and the content is extended around the advanced extension of the current knowledge point; the other is an option for reducing the knowledge density, and the content focuses on the basic consolidation of the current knowledge point. The two options are integrated into unified prompt content, which is pushed and waits for the learner to choose, fully adapting to the scenario of uncertain state.
[0077] Through the above steps, the real-time emotional state of the learner and the data reliability are accurately adapted to the teaching intervention strategy, and the rationality and practicality of the second teaching intervention strategy are improved.
[0078] In summary, the learning emotion recognition dynamic explanation training method provided by the embodiment of the application has the following technical effects: This application collects multimodal data of learners' facial videos, audio, and interactive text in real time. Through feature extraction, unimodal model operation, and divergence weighting, it obtains the deviation of emotional state with credibility scores. Combined with dynamic updates of personalized emotional baselines, it generates intervention strategies through a dual-core decision-making mechanism of white-box and black-box, accurately identifies learning emotions, and intervenes in a timely manner. This achieves the technical effect of making personalized learning tutoring more in line with learners' needs and improving the accuracy and effectiveness of learning assistance.
[0079] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a dynamic explanation and practice system for learning emotion recognition, the system comprising: Multimodal data acquisition module 1 is used to collect learners' multimodal data in real time during the learning process, including facial video data, voice and audio data, and interactive text data.
[0080] The emotional state signal acquisition module 2 performs emotional state recognition based on the multimodal data and generates an emotional state signal carrying a credibility marker.
[0081] The teaching intervention strategy acquisition module 3 generates teaching intervention strategies based on the emotional state signals and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.
[0082] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps: Geometric features of facial key points are extracted from the facial video data, acoustic features are extracted in real time from the speech audio data, and preset behavioral events are identified from the interactive text data to establish a multimodal feature recognition result. The multimodal feature recognition result and the current learning task context information are input into a multimodal fusion model to obtain an initial emotional state probability distribution. Signal consistency recognition is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution. The initial emotional state probability distribution is compared with a personalized emotional baseline established for the learner, and an emotional state deviation with a credibility score is output as the emotional state signal.
[0083] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps: The geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into pre-trained facial monomodal emotion recognition models, speech monomodal emotion recognition models, and interactive text monomodal emotion recognition models, respectively, to obtain facial emotion probability distributions, speech emotion probability distributions, and interactive text emotion probability distributions. The divergence between each pair of the facial emotion probability distribution, the speech emotion probability distribution, the interactive text emotion probability distribution, and the initial emotion state probability distribution is calculated to generate a divergence matrix. The divergence matrix is weighted and the calculation result is mapped to the credibility score, where the larger the divergence value, the lower the credibility score.
[0084] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps: A personalized emotional baseline is established for the learner, which includes the probability distribution of the learner's typical emotional states in different learning task contexts; the statistical distance between the initial emotional state probability distribution and the probability distribution of the typical emotional states corresponding to the current learning task context is calculated to obtain an emotional state deviation value; the emotional state deviation value is bound to the credibility score to generate the emotional state deviation signal.
[0085] Furthermore, the emotional state signal acquisition module 2 is used to perform the following steps: The personalized emotion baseline is dynamically updated by means of the following mechanism: collecting the probability distribution of the learner's emotional state in different learning task contexts when the confidence score is greater than a preset confidence threshold, and using the cluster center of the emotional state probability distribution to update the corresponding typical emotional state probability distribution in different learning task contexts.
[0086] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps: A set of predefined teaching intervention rules is provided, wherein the triggering condition of each rule depends on the joint output of multiple logical judgment units. The input of the logical judgment units includes the deviation of the emotional state with a credibility score and the learning task context parameters. The intervention strategy analysis of the corresponding rule is triggered and the corresponding first teaching intervention strategy is generated only when all logical judgment units in the triggering condition of any one of the teaching intervention rules are satisfied.
[0087] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps: The teaching intervention rules include at least the anxiety state rule and the immersion and fluency state rule; when the anxiety state rule is triggered, the cognitive scaffolding mode is activated to decompose the problem-solving path and implant growth mindset feedback; when the immersion and fluency state rule is triggered, interdisciplinary challenge questions are dynamically generated.
[0088] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps: A sequence of emotional state deviation with credibility scores and a sequence of learning task context are collected and input into the temporal prediction model, which is trained based on LSTM, with Dropout used before all layers of the LSTM. The temporal prediction model performs T forward propagation and outputs T prediction results. The average of the T prediction results is taken to obtain the inflection point risk probability. The standard deviation of the T prediction results is calculated to obtain the uncertainty estimate. The knowledge density is adjusted by combining the inflection point risk probability and the uncertainty estimate to generate a second teaching intervention strategy.
[0089] Furthermore, the teaching intervention strategy acquisition module 3 is used to perform the following steps: If the inflection point risk probability is greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, a teaching intervention strategy to reduce knowledge density is generated, including breaking down problem-solving steps and embedding review materials; if the inflection point risk probability is less than the preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, the current knowledge density is maintained, and a help prompt message to dynamically increase knowledge density is generated, waiting for user response; if the uncertainty estimate is greater than the preset uncertainty threshold, the current knowledge density is maintained, and an optional help prompt message is generated, waiting for user response.
[0090] The dynamic explanation and practice system for learning emotion recognition provided in this embodiment of the invention can execute the dynamic explanation and practice method for learning emotion recognition provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0091] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A dynamic explanation and practice method for learning emotion recognition, characterized by: include: Real-time collection of learners' multimodal data during the learning process, including facial video data, voice and audio data, and interactive text data; Based on the multimodal data, emotion state recognition is performed to generate an emotion state signal carrying a credibility marker; Based on the emotional state signals and their credibility, and combined with the learner's learning task context, a teaching intervention strategy is generated through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.
2. The dynamic explanation and practice method for learning emotion recognition as described in claim 1, characterized in that, Based on the multimodal data, emotion state recognition is performed to generate an emotion state signal carrying a confidence level label, including: The facial key point geometric features are extracted from the facial video data, acoustic features are extracted in real time from the voice audio data, and preset behavioral events are identified from the interactive text data to establish multimodal feature recognition results. The multimodal feature recognition results and the current learning task context information are input into the multimodal fusion model to obtain the initial emotional state probability distribution; Signal consistency identification is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution; The initial emotional state probability distribution is compared with the personalized emotional baseline established for the learner, and the emotional state deviation with a credibility score is output as the emotional state signal.
3. The dynamic explanation and practice method for learning emotion recognition as described in claim 2, characterized in that, Signal consistency identification is performed on the multimodal data, and a credibility score is assigned to the initial emotional state probability distribution, including: The geometric features of facial key points, acoustic features, and the results of interactive behavior event recognition are input into the pre-trained facial monomodal emotion recognition model, speech monomodal emotion recognition model, and interactive text monomodal emotion recognition model, respectively, to obtain the facial emotion probability distribution, speech emotion probability distribution, and interactive text emotion probability distribution. Calculate the divergence between each pair of the facial emotion probability distribution, the voice emotion probability distribution, the interactive text emotion probability distribution, and the initial emotion state probability distribution to generate a divergence matrix; The divergence matrix is weighted and the result is mapped to the credibility score, where the larger the divergence value, the lower the credibility score.
4. The dynamic explanation and practice method for learning emotion recognition as described in claim 2, characterized in that, The initial emotional state probability distribution is compared with a personalized emotional baseline established for the learner, and the emotional state deviation with a credibility score is output as the emotional state signal, including: Obtain a personalized emotional baseline established for the learner, the personalized emotional baseline including the probability distribution of the learner's typical emotional states in different learning task contexts; Calculate the statistical distance between the initial emotional state probability distribution and the typical emotional state probability distribution corresponding to the current learning task context to obtain the emotional state deviation value; The emotional state deviation value is bound to the credibility score to generate the emotional state deviation signal.
5. The dynamic explanation and practice method for learning emotion recognition as described in claim 4, characterized in that, The personalized emotion baseline is dynamically updated, and the update mechanism is as follows: Collect the probability distribution of the learner's emotional state in different learning task contexts when the credibility score is greater than a preset credibility threshold, and use the cluster center of the emotional state probability distribution to update the corresponding typical emotional state probability distribution in different learning task contexts.
6. The dynamic explanation and practice method for learning emotion recognition as described in claim 1, characterized in that, White-box decision-making based on pre-defined rules includes: A predefined set of teaching intervention rules is provided, wherein the triggering condition of each rule depends on the joint output of multiple logical judgment units, and the input of the logical judgment units includes the deviation of the emotional state with a credibility score and the learning task context parameters. The intervention strategy analysis of the corresponding rule is triggered and the corresponding first teaching intervention strategy is generated only when all logical judgment units in the triggering condition of any one of the teaching intervention rules are satisfied.
7. The dynamic explanation and practice method for learning emotion recognition as described in claim 6, characterized in that, Analysis of intervention strategies that trigger the corresponding rules, including: The teaching intervention rules include at least the anxiety state rule and the immersion and fluency state rule; When the aforementioned anxiety state rule is triggered, the cognitive scaffolding mode is activated to decompose the problem-solving path and implant growth mindset feedback. When the immersive and fluid state rule is triggered, interdisciplinary challenge questions are dynamically generated.
8. The dynamic explanation and practice method for learning emotion recognition as described in claim 1, characterized in that, Black-box decision-making based on time-series forecasting models includes: The emotional state deviation sequence with credibility score and the learning task context sequence are collected and input into the temporal prediction model, wherein the temporal prediction model is trained based on LSTM and Dropout is used before all layers of LSTM. The time-series prediction model performs T forward propagations and outputs T prediction results. The inflection point risk probability is obtained by averaging the T prediction results. The standard deviation of the T prediction results is calculated to obtain the uncertainty estimate; By combining the inflection point risk probability and the uncertainty estimate, the knowledge density is adjusted to generate a second teaching intervention strategy.
9. The dynamic explanation and practice method for learning emotion recognition as described in claim 8, characterized in that, By combining the inflection point risk probability and the uncertainty estimate, knowledge density is adjusted to generate a second teaching intervention strategy, including: If the inflection point risk probability is greater than or equal to a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, a teaching intervention strategy to reduce knowledge density is generated, including decomposing problem-solving steps and embedding review materials. If the inflection point risk probability is less than a preset risk probability threshold, and the uncertainty estimate is less than or equal to a preset uncertainty threshold, then the current knowledge density is maintained, and a help prompt message to dynamically increase the knowledge density is generated, waiting for the user's response. If the estimated uncertainty value is greater than the preset uncertainty threshold, the current knowledge density is maintained, and an optional help prompt is generated, waiting for the user's response.
10. A dynamic explanation and practice system for learning emotion recognition, characterized in that, The system is used for implementing the dynamic explanation and practice method for learning emotion recognition according to any one of claims 1-9, the system comprising: The multimodal data acquisition module is used to collect learners' multimodal data in real time during the learning process, including facial video data, voice and audio data, and interactive text data; The emotional state signal acquisition module identifies emotional states based on the multimodal data and generates emotional state signals carrying confidence markers. The teaching intervention strategy acquisition module generates teaching intervention strategies based on the emotional state signals and credibility, combined with the learner's learning task context, through a dual-core decision-making mechanism. The dual-core decision-making mechanism includes white-box decision-making based on preset rules and black-box decision-making based on a time-series prediction model.
Citation Information
Patent Citations
Emotion recognition method and system for teaching assistant robot
CN119128651A
Online teaching optimization method and system based on emotion recognition
CN120355539A
Psychological accompanying method based on multi-modal emotion recognition
CN120690390A
Multi-scene self-adaptive man-machine interaction system and method based on emotion recognition
CN121255014A