A method for identifying students' emotional state in digital teacher teaching situations
By collecting and integrating students' facial, interactive and voice information in the teaching situation of digital teachers, and combining historical emotion analysis to generate comprehensive emotional representations, the shortcomings of existing emotion recognition methods in information fusion and time dimensions are solved, and accurate identification and personalized adjustment of students' emotional state are achieved.
Patent Information
- Application Number
- CN202510065606.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing emotion recognition methods have obvious shortcomings in the consideration of information fusion and time dimensions, which affect the accuracy and comprehensiveness of emotion recognition.
A method for identifying students' emotional states in the teaching situation of digital teachers is proposed. By inputting students' facial information, platform interaction information and voice information into the trained emotion recognition model, collecting and preprocessing three-dimensional data information in real time, inputting data into the analysis module at preset times, generating comprehensive emotional representations through cross-attention fusion, and optimizing emotional result predictions in combination with historical emotion representations.
It realizes accurate identification of students' emotional state in the digital teacher teaching scenario, not only considers the students' instantaneous emotions, but also integrates historical emotion analysis, making emotional recognition more accurate, highlighting individual differences, and helping the digital teacher platform make targeted adjustments to students' emotions.
Smart Images

Figure CN119516594B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for identifying a student's emotional state in a digital teacher teaching situation. Background Art
[0002] Emotion recognition refers to the process of identifying an individual's inner emotional state by analyzing their physiological, behavioral and language characteristics. With the development of artificial intelligence and big data technology, emotion recognition has been widely used in education, psychology, medicine and other fields. In the context of digital teacher teaching, students' emotional state directly affects their learning effect and participation. Therefore, accurately identifying students' emotional state not only helps teachers adjust teaching strategies in a timely manner, but also provides students with personalized learning support and improves the overall learning experience. The goal of emotion recognition is to accurately capture and understand students' emotional changes in the learning process through comprehensive analysis of multiple information sources, thereby providing a scientific basis for educational decision-making.
[0003] However, existing emotion recognition methods have many defects, which limit their application effect in actual teaching. On the one hand, many existing methods rely too much on a single information source, such as facial expressions or voice features, and fail to effectively integrate information from multiple dimensions. For example, facial information can reflect students' immediate emotional reactions, while interactive information can reveal students' interactions with teachers and peers during the learning process, and voice information provides important clues such as emotional intonation and speech speed. The lack of fusion of multi-dimensional information affects the accuracy and comprehensiveness of emotion recognition. On the other hand, existing methods often ignore the relationship between instantaneous emotions and historical emotions. Students' emotional states are dynamically changing, and current emotions may be affected by historical emotions, while the accumulation of historical emotions will affect students' understanding and response to current learning content. Therefore, the time series characteristics of emotional states are not considered, which makes the results of emotion recognition lack depth and coherence. In summary, existing emotion recognition methods have obvious deficiencies in information fusion and consideration of the time dimension. A new method is urgently needed to improve the accuracy and practicality of emotion recognition in order to better serve student emotion management in digital teaching environments. Summary of the invention
[0004] Based on the technical problems existing in the background technology, the present invention proposes a method for identifying the emotional state of students in a digital teacher teaching situation, which can accurately identify the emotional state of students in a digital teacher teaching scenario.
[0005] The present invention proposes a method for identifying the emotional state of students in a digital teacher teaching situation, which inputs the student's facial information, platform interaction information and voice information into a trained emotion recognition model to obtain the emotion recognition result of the student;
[0006] Collect and pre-process three-dimensional data information in real time, including face, platform interaction and voice;
[0007] At every preset time, the pre-processed data information within the past set time interval is input into the analysis module to obtain The comprehensive parameter vector group at the moment is obtained by cross-attention fusion of the comprehensive parameter vector group Comprehensive emotional representation of the moment; after optimizing the comprehensive emotional representation, output the final emotional result prediction;
[0008] Construct a loss function to train the emotion recognition model.
[0009] Furthermore, real-time collection and preprocessing of three-dimensional data information specifically include:
[0010] A high-resolution camera is used to monitor students’ faces in real time, and the video is stored in mp4 format;
[0011] Record the platform interaction information of students when using the platform. The platform interaction information includes the number of abnormal keyboard and mouse interactions, question answering status, learning progress and classroom performance, which is stored in the form of key-value pairs;
[0012] The microphone is used to collect students' voice information in real time, and the audio is stored in wav format;
[0013] Use the target detection algorithm provided by OpenCV to detect the face position of the face information, extract the face, and obtain a new video to construct pre-processed data information, thereby avoiding interference information;
[0014] Use the functions provided by the audio processing tool librosa to remove noise from the audio and enhance the clarity of the human voice, and save it as a new audio to construct preprocessed data information.
[0015] Furthermore, at every preset time, the pre-processed data information within the past set time interval is input into the analysis module, specifically: starting from the 5th second of the class, every 1 second, the pre-processed data information within the past 5 seconds is intercepted, and the information obtained includes facial information , platform interaction information and voice messages .
[0016] Furthermore, the comprehensive parameter vector group includes a facial feature vector , interaction feature vector and speech feature vector , the generation process of the comprehensive parameter vector group is as follows:
[0017] Use OpenCV's VideoCapture interface to read new videos frame by frame, use the DeepFace model to process each frame, extract facial emotion features, calculate the average emotion features in the past 5 seconds, and save the processed results as facial feature vectors ;
[0018] Organize the platform interaction information, and send the number of abnormal keyboard and mouse interactions, question answering, learning progress and classroom performance in the past 5 seconds into the large language model in the form of tokens to form a series of vectors. Send the series of vectors into a trainable transformer layer and feedforward neural network to convert them into a high-dimensional interaction feature vector. ;
[0019] Use the audio processing library to analyze the new audio frame by frame and output it in the form of waveform images. Use the pre-trained VGG16 model to extract features from each waveform image and average the feature vectors of all frames to obtain the speech feature vector. .
[0020] Furthermore, in step 2, the comprehensive sentiment representation The generation process is as follows:
[0021] are facial feature vectors , interaction feature vector , speech feature vector Set up a cross-modal query, key, and value calculation;
[0022] Calculate the cross-modal attention weight of each feature vector relative to the other two feature vectors, and sum the cross-modal attention weights with the corresponding values to obtain the three modal information;
[0023] Calculate the features of the three modal information interactions and sum them up to obtain a comprehensive emotional representation .
[0024] Furthermore, the cross-modal attention weight of each feature vector relative to the other two feature vectors is calculated as follows:
[0025] Calculate facial feature vectors Attention weights for the other two feature vectors: facial feature vector As a query, the interaction feature vector As the key, get the facial feature vector The first modality attention weight ; With facial feature vector As a query, take the speech feature vector As the key, get the facial feature vector The second modality attention weight ;
[0026] Compute interaction eigenvectors Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the interaction feature vector The first modality attention weight ; With interactive feature vector As a query, take the speech feature vector As the key, get the interaction feature vector The second modality attention weight ;
[0027] Calculate speech feature vector Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the speech feature vector The first modality attention weight ; Based on speech feature vector As a query, the interaction feature vector As the key, get the speech feature vector The second modality attention weight .
[0028] Furthermore, comprehensive emotional representation The calculation formula is as follows:
[0029] ;
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] in, , , are all weights, i.e. trainable parameters, is the facial feature vector The cross-modal attention weights of is the interaction feature vector The cross-modal attention weights of is the speech feature vector The cross-modal attention weights of and They are trainable feedforward neural networks and normalization layers, , and They are the indexes of facial information, platform interaction information, and voice information. is the facial feature vector The cross-modal value of is the interaction feature vector The cross-modal value of is the speech feature vector The cross-modal value of is the facial modality information, is the interactive modal information, is the voice modality information, is the modal information of relative facial interaction, is the modal information of face relative to speech, is the modal information of the interactive relative face, is the modal information of the interaction relative to the speech, is the modal information of speech relative to face, It is the modal information of speech relative interaction.
[0037] Furthermore, the final sentiment result prediction generation process is as follows:
[0038] Use the attention mechanism to dynamically assign weights to historical sentiment representations so that the current comprehensive sentiment representation Refer to the historical sentiment representation to obtain the optimized sentiment representation ;
[0039] The optimized sentiment representation Send it to the trainable fully connected layer and normalization layer to get the final emotional result prediction .
[0040] Furthermore, emotion representation The specific formula is as follows:
[0041] ;
[0042] ;
[0043] in, is the attention weight of historical sentiment representation, for The comprehensive emotional representation of the moment, for A comprehensive emotional representation of the moment.
[0044] Furthermore, based on the sentiment prediction With actual emotions The loss function is constructed by the cross entropy between .
[0045] The advantages of the method for identifying the emotional state of students in a digital teacher teaching situation provided by the present invention are: the multi-dimensional classroom performance of students is effectively applied to the emotion recognition of students, and the behavioral information of students is standardized and integrated from multiple angles; not only the instantaneous emotions of students are considered, but also the historical emotions of students in this class are incorporated into the analysis, so that the emotions of students can be more accurately and effectively identified, highlighting the differences between individuals. In this way, not only can the emotional state of students be accurately identified in the digital teacher teaching scenario, but it also helps the digital teacher platform to make targeted adjustments to the different emotions of students, providing favorable support for AI education and digital teacher classroom effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0047] Below, the technical solution of the present invention is described in detail through specific embodiments. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific implementation disclosed below.
[0048] like Figure 1 As shown, the present invention proposes a method for identifying the emotional state of students in a digital teacher teaching situation, which inputs the student's facial information, platform interaction information and voice information into a trained emotion recognition model to obtain the emotion recognition result of the student;
[0049] The training process of the emotion recognition model is as follows:
[0050] Step 1: Collect and pre-process three-dimensional data information in real time, including face, platform interaction and voice;
[0051] The specific preprocessing process of data information is as follows: S11 to S16:
[0052] S11, use a high-resolution camera to monitor students’ faces in real time, and the video is stored in mp4 format;
[0053] S12. Record the platform interaction information of students when using the platform. The platform interaction information includes four aspects, namely, the number of abnormal keyboard and mouse interactions (for example, clicking the wrong button or area, quickly repeating keys or pressing invalid keys, etc.), question answering status (such as multiple-choice answer results, time spent on answering, etc.), learning progress (such as question completion rate, page browsing time, etc.) and classroom performance (class ranking, homework submission status, group discussion status, number of questions asked to the teacher, etc.), which are stored in the form of key-value pairs;
[0054] S13, using a microphone to collect students' voice information in real time, and storing the audio in wav format;
[0055] S14. Use the target detection algorithm (Haar Cascades model) provided by OpenCV to detect the face position of the face information, extract the face, and obtain a new video to construct pre-processed data information, thereby avoiding interference information;
[0056] S15. Use the functions provided by the audio processing tool librosa to remove noise from the audio and enhance the clarity of the human voice, and save it as a new audio to construct pre-processed data information;
[0057] S16, starting from the 5th second of the class, every 1 second, intercept the pre-processed data information in the past 5 seconds, and obtain information including facial information , platform interaction information and voice messages .
[0058] Through steps S11 to S16, the collected data information is standardized and sorted, dirty data is avoided, and valid data is sorted. Every 1 second, the valid information in the past 5 seconds is sent to the analysis module of step 2.
[0059] Step 2: At every preset time, the pre-processed data information within the past set time interval is input into the analysis module to obtain The comprehensive parameter vector group at the moment, the comprehensive parameter vector group includes a facial feature vector , interaction feature vector and speech feature vector , for facial feature vector , interaction feature vector , speech feature vector Based on cross attention fusion Comprehensive emotional representation of the moment ;
[0060] It should be noted that Facial feature vector at the moment , interaction feature vector , speech feature vector , which are the emotional representation results of students in three dimensions in the past 5 seconds.
[0061] Step 2 specifically includes S21 to S26:
[0062] S21. Use OpenCV's VideoCapture interface to read new videos frame by frame, use the DeepFace model to process each frame, extract facial emotion features, calculate the average emotion features in the past 5 seconds, and save the processed results as facial feature vectors. , which is the representation of the student's facial emotion information in the past 5 seconds at time t;
[0063] S22. Organize the platform interaction information, and send the number of abnormal keyboard and mouse interactions, question answering, learning progress and classroom performance in the past 5 seconds into the large language model in the form of tokens to form a series of vectors. Send the series of vectors into a trainable transformer layer and feedforward neural network to convert them into a high-dimensional interaction feature vector. , which is the representation of the platform interaction information in the past 5 seconds at time t;
[0064] S23. Use the audio processing library to analyze the new audio frame by frame, output it in the form of waveform images, use the pre-trained VGG16 model to extract features from each waveform image, average the feature vectors of all frames, and obtain the speech feature vector at time t. , which is the representation of the student's speech information in the past 5 seconds;
[0065] S24, respectively, are facial feature vectors , the interaction feature vector , speech feature vector Set up a cross-modal query, key, and value calculation;
[0066] ;
[0067] ;
[0068] ;
[0069] in, , , is a learnable weight matrix, are the cross-modal query, key, and value of the facial feature vector, respectively. are the cross-modal query, key, and value of the interaction feature vector, respectively. are the cross-modal query, key, and value of the speech feature vector, respectively.
[0070] S25. Calculate the cross-modal attention weight of each feature vector relative to the other two feature vectors, and weighted sum the cross-modal attention weights with the corresponding values to obtain three modal information;
[0071] Calculate facial feature vectors Attention weights for the other two feature vectors: facial feature vector As a query, the interaction feature vector As the key, get the facial feature vector The first modality attention weight ; Taking facial feature vector As a query, take the speech feature vector As the key, get the facial feature vector The second modality attention weight ;
[0072] , .
[0073] Compute interaction eigenvectors Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the interaction feature vector The first modality attention weight ; With interactive feature vector As a query, take the speech feature vector As the key, get the interaction feature vector The second modality attention weight ;
[0074] , .
[0075] Calculate speech feature vector Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the speech feature vector The first modality attention weight ; Based on speech feature vector As a query, the interaction feature vector As the key, get the speech feature vector The second modality attention weight ;
[0076] , .
[0077] Finally, the cross-modal attention weights are used to weight the corresponding values:
[0078] ;
[0079] ;
[0080] ;
[0081] S26. Calculate the features after the interaction of the three modal information and sum them up to obtain a comprehensive emotional representation , which means the comprehensive emotional representation of students in the past 5 seconds;
[0082] The features after the interaction of the three modal information are calculated:
[0083] ;
[0084] ;
[0085] ;
[0086] The normalized features are weighted and summed to obtain a comprehensive sentiment representation , which is the moment The comprehensive emotional representation of students:
[0087] ;
[0088] The parameters in the above formula are: , , are all weights, i.e. trainable parameters, is the facial feature vector The cross-modal attention weights of is the interaction feature vector The cross-modal attention weights of is the speech feature vector The cross-modal attention weights of and They are trainable feedforward neural networks and normalization layers, , and They are the indexes of facial information, platform interaction information, and voice information. is the facial feature vector The cross-modal value of is the interaction feature vector The cross-modal value of is the speech feature vector The cross-modal value of is the facial modality information, is the interactive modal information, is the voice modality information, is the modal information of relative facial interaction, is the modal information of face relative to speech, is the modal information of the interactive relative face, is the modal information of the interaction relative to the speech, is the modal information of speech relative to face, It is the modal information of speech relative interaction.
[0089] Step 3: Represent the comprehensive emotion Combined with historical sentiment representation to obtain optimized sentiment representation , after the fully connected layer and the activation layer, the final emotional result prediction is obtained , that is, the probability distribution of 6 states: peaceful, anxious, distracted, angry, disappointed, and proud;
[0090] Step 3 specifically includes S31 to S32:
[0091] S31. Use the attention mechanism to dynamically assign weights to historical sentiment representations so that the current comprehensive sentiment representation Able to flexibly refer to historical sentiment representations to obtain optimized sentiment representations :
[0092] ;
[0093] ;
[0094] in, is the attention weight of historical sentiment representation, for The comprehensive emotional representation of the moment, for A comprehensive emotional representation of the moment.
[0095] S32. Represent the optimized emotion Send it to the trainable fully connected layer and normalization layer to get the final emotional result prediction , It is the probability distribution representation of multi-classification:
[0096] ;
[0097] in, are the trainable weight matrices for the normalization layer and the fully connected layer, is the fully connected layer, is the normalization layer, which means the activation function.
[0098] Step 4: Predict the emotional results With actual emotions The cross entropy between is used as the loss function to train the emotion recognition model until the emotion recognition model reaches a convergence state.
[0099] Step 4 includes S41 to S42:
[0100] S41, sort out parameters and set hyperparameters;
[0101] In this step, you need to organize all the parameters required for the emotion recognition model, including the weights and biases of the emotion recognition model. At the same time, you also need to set some hyperparameters, such as learning rate, batch size, number of training rounds, etc.
[0102] S42. Predicting emotional results With actual emotions Do cross entropy to train the entire model parameters, specific loss function for:
[0103] ;
[0104] in, It is a trainable parameter, and 6 represents 6 states: peaceful, anxious, distracted, angry, disappointed, and proud. is one of the 6 states. For the The actual emotion of a state, Represents the predicted emotional outcome output by the emotion recognition model; is a hyperparameter that controls the regularization strength.
[0105] Update the model parameters of the emotion recognition model to minimize this loss function ,This process will be repeated until the performance of the emotion recognition model reaches a satisfactory level or the preset number of training rounds is reached.
[0106] Through steps one to four, this embodiment records students' multi-dimensional information at all times in the digital teacher classroom, sends the information to respective analysis modules to obtain three-dimensional information representation, and then sends it to the interactive attention fusion module to obtain a comprehensive emotional representation result, and combines it with the historical emotional representation to finally calculate the predicted result of the emotional state.
[0107] This embodiment effectively applies students' multi-dimensional classroom performance to students' emotion recognition, standardizes and integrates students' behavioral information from multiple angles; it not only considers students' instantaneous emotions, but also incorporates students' historical emotions in this class into the analysis, so that students' emotions can be more accurately and effectively identified, highlighting the differences between individuals. In this way, not only can the emotional state of students be accurately identified in the digital teacher teaching scenario, but it also helps the digital teacher platform to make targeted adjustments to students' different emotions, providing favorable support for AI education and digital teacher classroom effects.
[0108] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for identifying the emotional state of students in a digital teacher teaching situation, characterized in that: Input the student's facial information, platform interaction information, and voice information into the trained emotion recognition model to obtain the emotion recognition result of the student; The training process of the emotion recognition model is as follows: Collect and pre-process three-dimensional data information in real time, including face, platform interaction and voice; At every preset time, the pre-processed data information within the past set time interval is input into the analysis module to obtain The comprehensive parameter vector group at the moment is obtained by cross-attention fusion of the comprehensive parameter vector group The comprehensive emotional representation of the moment; after optimizing the comprehensive emotional representation, the final emotional result prediction is output, and the comprehensive parameter vector group includes the facial feature vector , interaction feature vector and speech feature vector ; Construct loss function to train the emotion recognition model; The comprehensive emotional representation The generation process is as follows: are facial feature vectors , interaction feature vector , speech feature vector Set up a cross-modal query, key, and value calculation; Calculate the cross-modal attention weight of each feature vector relative to the other two feature vectors, and sum the cross-modal attention weights with the corresponding values to obtain the three modal information; Calculate the features of the three modal information interactions and sum them up to obtain a comprehensive emotional representation ; Loss Function for: ; in, is a trainable parameter, is one of the 6 states. For the The actual emotion of a state, Represents the predicted emotional outcome output by the emotion recognition model; is a hyperparameter.
2. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 1 is characterized in that: Real-time collection and preprocessing of three-dimensional data information, including: A high-resolution camera is used to monitor students’ faces in real time, and the video is stored in mp4 format; Record the platform interaction information of students when using the platform. The platform interaction information includes the number of abnormal keyboard and mouse interactions, question answering status, learning progress and classroom performance, which is stored in the form of key-value pairs; The microphone is used to collect students' voice information in real time, and the audio is stored in wav format; Use the target detection algorithm provided by OpenCV to detect the face position of the face information, extract the face, and obtain a new video to construct pre-processed data information, thereby avoiding interference information; Use the functions provided by the audio processing tool to remove noise from the audio and enhance the clarity of the human voice, and save it as a new audio to construct pre-processed data information.
3. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 2 is characterized in that: At every preset time, the pre-processed data information within the past set time interval is input into the analysis module, specifically: starting from the 5th second of the class, every 1 second, the pre-processed data information within the past 5 seconds is intercepted, and the information obtained includes facial information , platform interaction information and voice messages .
4. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 1, characterized in that: The generation process of the comprehensive parameter vector group is as follows: Use OpenCV's interface to read new videos frame by frame, use the DeepFace model to process each frame, extract facial emotion features, calculate the average emotion features in the past 5 seconds, and save the processed results as facial feature vectors ; Organize the platform interaction information, and send the number of abnormal keyboard and mouse interactions, question answering, learning progress and classroom performance in the past 5 seconds into the large language model in the form of tokens to form a series of vectors. Send the series of vectors into a trainable transformer layer and feedforward neural network to convert them into a high-dimensional interaction feature vector. ; Use the audio processing library to analyze the new audio frame by frame and output it in the form of waveform images. Use the pre-trained VGG16 model to extract features from each waveform image and average the feature vectors of all frames to obtain the speech feature vector. .
5. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 1, characterized in that: Calculate the cross-modal attention weight of each feature vector relative to the other two feature vectors, specifically: Calculate facial feature vectors Attention weights for the other two feature vectors: facial feature vector As a query, the interaction feature vector As the key, get the facial feature vector The first modality attention weight ; Taking facial feature vector As a query, take the speech feature vector As the key, get the facial feature vector The second modality attention weight ; Compute interaction eigenvectors Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the interaction feature vector The first modality attention weight ; With interactive feature vector As a query, take the speech feature vector As the key, get the interaction feature vector The second modality attention weight ; Calculate speech feature vector Attention weights for the other two feature vectors: As a query, take the facial feature vector As the key, get the speech feature vector The first modality attention weight ; Based on speech feature vector As a query, the interaction feature vector As the key, get the speech feature vector The second modality attention weight .
6. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 5, characterized in that: Comprehensive emotion representation The calculation formula is as follows: ; ; ; ; ; ; ; in, , , are all weights, i.e. trainable parameters, is the facial feature vector The cross-modal attention weights of is the interaction feature vector The cross-modal attention weights of is the speech feature vector The cross-modal attention weights of and They are trainable feedforward neural networks and normalization layers, , and They are the indexes of facial information, platform interaction information, and voice information. is the facial feature vector The cross-modal value of is the interaction feature vector The cross-modal value of is the speech feature vector The cross-modal value of is the facial modality information, is the interactive modal information, is the voice modality information, is the modal information of relative facial interaction, is the modal information of face relative to speech, is the modal information of the interactive relative face, is the modal information of the interaction relative to the speech, is the modal information of speech relative to face, It is the modal information of speech relative interaction.
7. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 1, characterized in that: The final emotional result prediction generation process is as follows: Use the attention mechanism to dynamically assign weights to historical sentiment representations so that the current comprehensive sentiment representation Refer to the historical sentiment representation to obtain the optimized sentiment representation ; The optimized sentiment representation Send it to the trainable fully connected layer and normalization layer to get the final emotional result prediction .
8. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 7, characterized in that: Emotional Representation The specific formula is as follows: ; ; in, is the attention weight of historical sentiment representation, for The comprehensive emotional representation of the moment, for A comprehensive emotional representation of the moment.
9. The method for identifying the emotional state of students in a digital teacher teaching situation according to claim 1, characterized in that: Prediction based on sentiment results With actual emotions The loss function is constructed by the cross entropy between .
Citation Information
Patent Citations
Detainee emotion recognition method for multi-modal feature fusion based on Transformer, equipment, and medium
CN113822192A
Cross-modal video description model based on dynamic memory network
CN117496388A
Personalized educational experience AI enabled student card system
CN118691432A