Anger emotion recognition method based on video stream

The method improves anger emotion recognition by integrating facial, skeletal, and audio features through advanced neural networks and attention mechanisms, addressing the limitations of traditional methods in capturing multi-modal temporal and spatial correlations.

CN120318891AActive Publication Date: 2025-07-15SICHUAN CANCER HOSPITAL

Patent Information

Application Number
CN202510769064.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-15
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Traditional anger emotion recognition methods fail to fully explore the space-time correlation of facial, limb, and voice modalities, resulting in poor accuracy of anger emotion recognition and inability to cope with emotional recognition in complex scenarios.

Method used

By preprocessing the video stream, standardized facial image sequences, 3D skeleton sequences and voice fragments with space-time alignment are extracted, facial features are extracted using 3D-ResNet34 network and micro-expression optical flow enhancement algorithm, and limb motion features are extracted, and feature fusion is combined with multi-head space-time cross attention mechanism and expansion timing convolution algorithm to classify the weights.

Benefits of technology

Cross-modal time and space-time cross-fusion is achieved, improving the accuracy and robustness of recognition of anger emotions, enhancing the ability to recognize instantaneous micro-expressions, and being able to perform high-precision anger emotions recognition in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318891A_ABST
    Figure CN120318891A_ABST
Patent Text Reader

Abstract

The invention discloses an angry emotion recognition method based on a video stream, and belongs to the technical field of emotion recognition. The method comprises the following steps: preprocessing a video stream of a visitor to obtain a time-space aligned standardized facial image sequence, a 3D skeleton sequence and a voice segment; performing feature extraction on the standardized facial image sequence by using a 3D-ResNet34 network and a micro-expression optical flow enhancement algorithm to obtain optical flow facial features; performing feature extraction on the 3D skeleton sequence by using a space-time diagram convolutional network to obtain limb action features; performing feature extraction on the voice segments to obtain voice features; and performing fusion processing on the facial features, the body movement features and the voice features by using a multi-head space-time cross attention mechanism and an expansion time sequence convolution algorithm, and outputting an angry emotion recognition result through dynamic weight distribution classification. According to the method, the angry emotion recognition accuracy of the system is greatly improved on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and particularly to a method for recognizing angry emotions based on a video stream. Background Art

[0002] In the technical field of emotion recognition, traditional recognition methods mostly adopt feature splicing or simple weighted fusion, resulting in insufficient exploration of the spatio-temporal correlation of facial, limb, and speech modalities. For example, angry expressions are often accompanied by arm waving and sudden changes in speech fundamental frequency. Traditional recognition methods are difficult to capture such cross-modal spatio-temporal synchronization features, resulting in poor recognition accuracy of angry emotions and inability to handle emotion recognition in complex scenarios. Summary of the Invention

[0003] The main purpose of the present invention is to provide a method for recognizing angry emotions based on a video stream, aiming to solve the technical problem of poor recognition accuracy of angry emotions in related technologies.

[0004] To achieve the above object, the present invention provides a method for recognizing angry emotions based on a video stream, which includes the following steps:

[0005] S1, preprocess the video stream of the visiting person to obtain a spatio-temporally aligned standardized facial image sequence, a 3D skeleton sequence, and a speech segment;

[0006] S2, use a 3D-ResNet34 network and a micro-expression optical flow enhancement algorithm to extract features from the standardized facial image sequence to obtain optical flow facial features;

[0007] S3, use a spatio-temporal graph convolutional network to extract features from the 3D skeleton sequence to obtain limb motion features;

[0008] S4, extract features from the speech segment to obtain speech features;

[0009] S5, use a multi-head spatio-temporal cross-attention mechanism and a dilated temporal convolutional algorithm to fuse the facial features, limb motion features, and speech features, and then output the angry emotion recognition result through dynamic weight allocation classification.

[0010] The present invention realizes cross-modal spatio-temporal cross-fusion through fine-grained spatio-temporal correlation modeling of facial, limb, and speech modalities, so as to be able to fully capture multi-modal collaborative features. At the same time, combined with dynamic weighting of the optical flow field, the recognition ability of the system for instantaneous micro-expressions can be enhanced. On this basis, combined with the stability evaluation of temporal features and the gated update mechanism, a dynamic classification weight is established to realize continuous quantitative analysis of emotions. Therefore, the present invention greatly improves the accuracy of the system for recognizing angry emotions as a whole. Description of the Drawings

[0011] Figure 1 This is a schematic flowchart of an embodiment of the method for recognizing angry emotions based on video streams according to the present invention;

[0012] Figure 2 This is a detailed flowchart of an embodiment of the method for recognizing angry emotions based on video streams according to the present invention.

[0013] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0014] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0015] The inventive concept of the present application will be further elaborated below in conjunction with some specific embodiments and specific implementation manners.

[0016] An embodiment of the present invention provides a method for recognizing angry emotions based on video streams. Referring to Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of a method for recognizing angry emotions based on video streams according to the present invention.

[0017] In this embodiment, the method for recognizing angry emotions based on video streams includes:

[0018] Step S1: Preprocess the video stream of the visiting person to obtain a spatio-temporally aligned standardized facial image sequence, a 3D skeleton sequence, and a voice segment.

[0019] The specific steps of step S1 are as follows:

[0020] Step S11: Segment the video stream into video segments of a fixed length. Each video segment contains 16 frames of original images, and the overlapping window for each segmentation is 8 frames.

[0021] Anger can be a short-lived outburst or persistent. Therefore, maintaining overlap during segmentation can capture more coherent changes and avoid missing key frames.

[0022] Step S12: Extract the facial ROI for the original images of each video segment to obtain a standardized facial image sequence .

[0023] Specifically, use the RetinaFace detector to locate the facial region in the original image, align the key points through affine transformation, and obtain a standardized facial image:

[0024] ;

[0025] Among them, denote the original image of the t-th frame in the video clip; denote the normalized facial image corresponding to the original image of the t-th frame; denote the alignment operation; denote the affine transformation parameters; denote the translation parameter in the x-axis direction; denote the translation parameter in the y-axis direction; denote the rotation angle; denote the scaling ratio; denote the two-dimensional Euclidean group, that is, the transformation set including translation, rotation, and scaling.

[0026] After processing all the original images, a sequence of normalized facial images of the video clip is obtained .

[0027] In step S12, the normalized facial image can ensure the consistency of facial positions and sizes in different videos and reduce noise. RetinaFace detects key points, and the aligned facial features will be clearer (such as the changes in eyebrows and mouth), which helps to capture features such as widened eyes and downturned corners of the mouth when angry.

[0028] Step S13: For the original images of each video clip, use AlphaPose to detect 2D skeleton key points and generate a 3D skeleton sequence through inverse kinematics optimization .

[0029] Assume the number of joint points J = 17. Through the joint points, limb movements such as arm waving and body leaning forward when angry can be captured.

[0030] Step S14: Extract the audio stream from the video stream, and successively perform pre-emphasis filtering, frame segmentation (such as a 25ms window, 10ms overlap), and silent segment filtering, and retain 16 speech segments synchronized with the video clip.

[0031] By pre-emphasis filtering and filtering silent segments, high-energy speech frames matching the speech burst characteristics of anger can be retained.

[0032] In the whole step S1, the video stream is preprocessed through spatio-temporal segmentation, facial ROI extraction, limb skeleton extraction, and speech processing to achieve multi-modal decoupling, so as to systematically extract and standardize facial, limb, and speech features, and combined with the temporal alignment mechanism, significantly improve the robustness and accuracy of anger emotion recognition.

[0033] Step S2: Use the 3D-ResNet34 network and the micro-expression optical flow enhancement algorithm to extract features from the sequence of normalized facial images to obtain optical flow facial features .

[0034] The specific steps of step S2 are as follows:

[0035] Step S21: Input the standardized facial image sequence into the 3D-ResNet34 network, and use 3D convolutional kernels to extract the spatio-temporal feature sequence :

[0036] ;

[0037] Among them, the spatio-temporal feature sequence has a dimension of 16×512; represents the 3D convolution operation, including 4 residual blocks connected in sequence, and each residual block is composed of a 3D convolutional layer, a batch normalization layer, and a ReLU activation layer connected in sequence.

[0038] Step S22: Perform micro-expression optical flow enhancement processing on the standardized facial image sequence to obtain the optical flow feature sequence.

[0039] The specific steps of step S22 are as follows:

[0040] Step S22-1: Calculate the dense optical flow fields of adjacent frames in the standardized facial image sequence in sequence :

[0041] ;

[0042] Among them, represents the FlowNet model; represents the t-th frame of the standardized facial image; represents the (t + 1)-th frame of the standardized facial image, and the dense optical flow field has a dimension of 112×112×2.

[0043] Step S22-2: Detect micro-expression instants based on the dense optical flow field to obtain the micro-expression weight coefficient :

[0044] ;

[0045] Among them, represents the optical flow amplitude; represents the average pooling operation; represents a learnable convolutional kernel with a dimension of 1×1×1, which is used to map the global average pooling result of the optical flow amplitude to a scalar weight; represents the Sigmoid function.

[0046] Step S22-3: Use a 1D convolutional kernel for the dense optical flow field Perform dimensionality reduction processing to obtain optical flow features; return to execute step S22-1 until all normalized facial images are processed to obtain an optical flow feature sequence :

[0047] ;

[0048] Among them, represents the optical flow feature corresponding to the t-th frame of the normalized facial image; represents the convolution operation.

[0049] Step S23: After concatenating the optical flow feature sequence and the spatio-temporal feature sequence frame by frame, use the micro-expression weight coefficient to perform weighting to obtain the optical flow facial feature :

[0050] ;

[0051] Among them, represents the optical flow facial feature corresponding to the t-th frame of the normalized facial image, with a dimension of 567; represents the spatio-temporal feature corresponding to the t-th frame of the normalized facial image; represents the concatenation operation.

[0052] Finally, concatenate the 16-frame optical flow facial features to obtain the optical flow facial feature .

[0053] In the entire step S2, by combining the 3D-ResNet34 network and the optical flow enhancement technology, the efficient extraction of facial dynamic features is achieved. Thus, not only can the subtle changes that are difficult to detect in static images be captured, but also the sensitivity to micro-expressions is enhanced through the optical flow method, thereby improving the accuracy and robustness of micro-expression recognition.

[0054] Step S3: Use the spatio-temporal graph convolutional network to extract features from the 3D skeleton sequence to obtain limb motion features .

[0055] The specific steps of step S3 are as follows:

[0056] Step S31: Based on the 3D skeleton sequence, construct a spatio-temporal graph through spatial edges and temporal edges.

[0057] Specifically, the spatial edges connect adjacent joint points, and the temporal edges connect the same joint points in adjacent frames, thereby forming a complete spatio-temporal graph.

[0058] Step S32: Perform spatial graph convolution on the skeleton features of each time frame in the spatio-temporal graph to extract local spatial features.

[0059] Step S33: Perform one-dimensional convolution on the local spatial features along the time dimension to capture the temporal dynamic changes of the action.

[0060] Step S34: Based on the joint attention mechanism, calculate the dot product of the query, key, and value for the convolution result and apply the softmax function to output the limb action features. 。

[0061] In the entire Step S3, by simultaneously modeling the spatial pose and temporal dynamics through the spatio-temporal graph, the action patterns unique to anger can be captured. On this basis, combining the attention mechanism can highlight the key joints, reduce noise interference, distinguish subtle action differences, and improve the accuracy of anger action recognition.

[0062] Step S4: Extract features from the speech segment to obtain speech features. 。

[0063] The specific steps of Step S4 include the following steps:

[0064] Step S41: Calculate and extract 28-dimensional static MFCC parameters based on the speech segment, and then obtain 28-dimensional dynamic MFCC features through first-order difference processing. Concatenate the static MFCC parameters and the dynamic MFCC features to obtain a 56-dimensional MFCC feature vector.

[0065] Step S42: Use the SWT-HS algorithm to detect the fundamental frequency of the speech segment, and then calculate the jitter coefficient of the speech segment based on the fundamental frequency:

[0066] ;

[0067] where, represents the jitter coefficient; represents the fundamental frequency value of the speech data in the t-th frame; represents the fundamental frequency value of the speech data in the (t + 1)-th frame.

[0068] Step S43: Calculate the global average spectral centroid based on the speech segment:

[0069] ;

[0070] ;

[0071] where, represents the spectral centroid of the speech data in the t-th frame; represents the discrete Fourier transform coefficient of the k-th frequency component of the speech data in the t-th frame; represents the frequency value of the k-th frequency component; represents the global average spectral centroid.

[0072] Step S44: Based on the MFCC feature vectors, jitter coefficients, and global average spectral centroid, splice them to obtain speech feature vectors, and then align them to 16 frames through dynamic time warping to obtain speech features .

[0073] In the entire Step S4, by combining multiple acoustic features, multiple-dimensional manifestations of anger emotion in speech can be captured, including spectrum, fundamental frequency perturbation, energy, and time-domain characteristics, thereby improving the accuracy of anger emotion recognition. At the same time, the time alignment process can enhance the system's adaptability to different speaking patterns, making the feature expressions more consistent and contributing to learning the key patterns of anger.

[0074] Step S5: Use the multi-head spatio-temporal cross-attention mechanism and dilated temporal convolutional algorithm to fuse the facial features, limb movement features, and speech features, and then output the anger emotion recognition result through dynamic weight allocation classification.

[0075] As Figure 2 shown, Step S5 specifically includes the following steps:

[0076] Step S51: Map the facial features , limb movement features , and speech features to a unified space to obtain the facial feature vector , limb movement feature vector , and speech feature vector :

[0077] ;

[0078] where represents the input feature m, which is the facial features, limb movement features, and speech features; represents the feature vector of feature m; represents the weight matrix corresponding to feature m; represents the bias term corresponding to feature m; represents the dimension of feature m; represents the normalization operation.

[0079] Based on the above feature mapping method, finally obtain the facial feature vector , limb movement feature vector , and speech feature vector .

[0080] Step S52: Based on the facial feature vector , the limb movement feature vector and the speech feature vector perform cross-attention calculation to obtain the attention of facial features to limb movement features, the attention of facial features to speech features, the attention of limb movement features to facial features, the attention of limb movement features to speech features, the attention of speech features to facial features, and the attention of speech features to limb movement features:

[0081] ;

[0082] ;

[0083] Among them, represents the attention of facial features to limb movement features; represents the query vector; represents the key vector; represents the value vector; , and both represent linear transformation matrices with dimensions of 512×64; represents the key vector dimension; represents the key vector transpose.

[0084] Similarly, the attention of facial features to speech features can be obtained , the attention of limb movement features to facial features , the attention of limb movement features to speech features , the attention of speech features to facial features and the attention of speech features to limb movement features .

[0085] Step S53: Use 4-layer dilated causal convolution to fuse the facial feature vector , the limb movement feature , the speech feature vector and the cross-attention calculation result to obtain the facial fusion feature, the limb movement fusion feature, and the speech fusion feature; the dilation rates of the 4-layer dilated causal convolution are 1, 2, 4, and 8 respectively:

[0086] ;

[0087] Among them, represents the facial fusion feature; represents the concatenation operation; represents the dilated causal convolution operation.

[0088] Similarly, the limb movement fusion feature And voice fusion features 。

[0089] Step S54: Calculate the temporal variances of facial features, limb movement features, and voice features:

[0090] ;

[0091] Among them, represents the variance of feature m; represents the feature of the t-th frame of feature m; represents the feature mean of feature m.

[0092] Step S55: Establish dynamic classification weights based on the temporal variances to obtain the classification weights at the current moment:

[0093] The specific steps of step S55 include the following steps:

[0094] Step S55-1: Set the initial classification weights 。

[0095] Step S55-2: Dynamically update the classification weights using a gated recurrent unit (GRU) to obtain the classification weights at the current moment:

[0096] First, calculate the forget gate and the update gate :

[0097] ;

[0098] Among them, represents the weight matrix of the forget gate, with a dimension of 2×64; represents the weight matrix of the update gate, with a dimension of 2×64; represents the classification loss at the previous moment.

[0099] Then, calculate the candidate state :

[0100] ;

[0101] Among them, represents the hyperbolic tangent function, which is used to introduce non-linearity; represents the mapping matrix, with a dimension of 2×64, which is used to map the input signal to the hidden state space; represents the Hadamard product (element-wise multiplication).

[0102] Finally, calculate the classification weights at the current moment :

[0103] ;

[0104] Among them, represents the classification weight at the previous moment.

[0105] Step S56: Based on the classification weight at the current moment, weighted aggregate the facial fusion feature, limb movement fusion feature, and speech fusion feature, output the classification probability, and obtain the angry emotion recognition result:

[0106] ;

[0107] Among them, represents the classification probability, including two emotion categories (angry / non-angry); represents the fusion feature of feature m.

[0108] Finally, take the emotion category corresponding to the maximum probability value as the angry emotion recognition result.

[0109] In the whole step S5, through unified feature space projection and multi-modal cross-attention mechanism, dynamically integrate the spatio-temporal features of face, limb and speech, effectively capture the discriminative information of multi-modal collaborative evolution in angry emotion (such as the synchronization of facial distortion and limb movement). On this basis, use dilated convolution to model long-term temporal dependencies to locate the emotion outbreak point, and combine GRU dynamic weight allocation to adaptively balance the modal contributions under noise interference, so as to achieve highly robust angry recognition in complex scenarios (occlusion, noise), and significantly improve the detection accuracy of micro-expression capture and rapid emotion change.

[0110] In this embodiment, through fine-grained spatio-temporal correlation modeling of face, limb and speech modalities, cross-modal spatio-temporal cross-fusion is realized, so as to be able to fully capture multi-modal collaborative features. At the same time, combined with the dynamic weighting of the optical flow field, the recognition ability of the system for instantaneous micro-expressions can be enhanced. On this basis, combined with the temporal feature stability evaluation and gating update mechanism, a dynamic classification weight is established to realize the continuous quantitative analysis of emotions. Therefore, the present invention greatly improves the accuracy of the system for angry emotion recognition as a whole.

[0111] The above serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0112] The above is only the preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for recognizing angry emotions based on video streams, characterized in that, The method includes the following steps: S1. Preprocess the video stream of the visitor to obtain a spatio-temporally aligned standardized facial image sequence, a 3D skeleton sequence, and a speech segment; S2. Use the 3D-ResNet34 network and the micro-expression optical flow enhancement algorithm to extract features from the standardized facial image sequence to obtain optical flow facial features; S3. Use the spatio-temporal graph convolutional network to extract features from the 3D skeleton sequence to obtain limb motion features; S4. Extract features from the speech segment to obtain speech features; S5. Use the multi-head spatio-temporal cross-attention mechanism and the dilated temporal convolutional algorithm to fuse the facial features, limb motion features, and speech features, and then output the anger emotion recognition result through dynamic weight assignment classification.

2. The method for recognizing angry emotions based on video streams according to claim 1, wherein The specific content of S1 includes: S11. Split the video stream into video segments of a fixed length. Each video segment contains 16 frames of original images, and the overlapping window for each split is 8 frames; S12. For the original images of each video segment, perform facial ROI extraction to obtain the standardized facial image sequence of this video segment; S13. For the original images of each video segment, use AlphaPose to detect 2D skeleton joints, and generate a 3D skeleton sequence through inverse kinematics optimization; S14. Extract the audio stream from the video stream, and perform pre-emphasis filtering, framing, and silent segment filtering in sequence, and retain 16 frames of speech segments synchronized with the video segment.

3. The method for recognizing angry emotions based on video streams according to claim 1, wherein, The specific content of S2 includes: S21, input the standardized facial image sequence into the 3D-ResNet34 network, and use 3D convolutional kernels to extract the spatio-temporal feature sequence ; S22. Perform micro-expression optical flow enhancement processing on the standardized facial image sequence to obtain an optical flow feature sequence ; S23. Concatenate the optical flow feature sequence and the spatio-temporal feature sequence frame by frame, and then use the micro-expression weight coefficient to perform weighting to obtain the optical flow facial feature .

4. The method for recognizing angry emotions based on video streams according to claim 3, wherein The specific content of S22 includes: S22-1, successively calculate the dense optical flow fields of adjacent frames in the standardized facial image sequence : ; Among them, represents the FlowNet model; represents the normalized facial image of the t-th frame; represents the normalized facial image of the (t + 1)-th frame, and the dense optical flow field has a dimension of 112×112×2; S22-2, based on the dense optical flow field Detect micro-expression moments and obtain micro-expression weight coefficients : ; Among them, represents the optical flow amplitude; represents the average pooling operation; represents a learnable convolutional kernel with a dimension of 1×1×1, which is used to map the global average pooling result of the optical flow amplitude to a scalar weight; represents the Sigmoid function; S22-3, use a 1D convolutional kernel to process the dense optical flow field for dimensionality reduction to obtain optical flow features; return to execute step S22-1 until all standardized facial image processing is completed to obtain an optical flow feature sequence : ; Among them, represents the optical flow feature corresponding to the normalized facial image of the t-th frame; represents a convolution operation.

5. The method for recognizing angry emotions based on a video stream according to claim 1, wherein The specific content of S3 includes: S31. Based on the 3D skeleton sequence, construct a spatio-temporal graph through spatial edges and temporal edges; S32. Perform spatial graph convolution on the skeleton features of each time frame in the spatio-temporal graph to extract local spatial features; S33. Perform one-dimensional convolution on the local spatial features along the time dimension to capture the temporal dynamic changes of the action; S34, based on the joint attention mechanism, calculate the dot product of the query, key, and value for the convolution result and apply the softmax function to output the limb action features .

6. The method for recognizing angry emotions based on video stream according to claim 1, characterized in that, The specific content of S4 includes: S41. Calculate and extract 28-dimensional static MFCC parameters based on the speech segment, and then obtain 28-dimensional dynamic MFCC features through first-order difference processing. Concatenate the static MFCC parameters and the dynamic MFCC features to obtain a 56-dimensional MFCC feature vector; S42. Use the SWT-HS algorithm to detect the fundamental frequency of the speech segment, and then calculate the jitter coefficient of the speech segment according to the fundamental frequency; ; Among them, represents the jitter coefficient; represents the fundamental frequency value of the t-th frame of speech data; represents the fundamental frequency value of the (t + 1)-th frame of speech data; S43. Based on the speech segment, calculate the global average spectral centroid; ; ; Among them, represents the spectral centroid of the t-th frame of speech data; represents the discrete Fourier transform coefficient of the k-th frequency component of the t-th frame of speech data; represents the frequency value of the k-th frequency component; represents the global average spectral centroid; S44. Based on the MFCC feature vectors, jitter coefficients, and global average spectral centroids, the speech feature vectors are spliced, and then aligned to 16 frames through dynamic time warping to obtain the speech features .

7. The method for recognizing angry emotions based on video stream according to claim 1, wherein, The specific content of S5 includes: S51, map facial feature , limb movement feature and voice feature to a unified space to obtain a facial feature vector , a limb movement feature vector and a voice feature vector ; S52, based on the facial feature vector , the limb movement feature vector and the voice feature vector perform cross-attention calculations to obtain the attention of facial features to limb movement features, the attention of facial features to voice features, the attention of limb movement features to facial features, the attention of limb movement features to voice features, the attention of voice features to facial features, and the attention of voice features to limb movement features; S53, using 4-layer dilated causal convolution, for facial feature vectors , limb movement features , speech feature vectors and the cross-attention calculation results are fused to obtain facial fusion features, limb movement fusion features, and speech fusion features; the dilation rates of the 4-layer dilated causal convolution are 1, 2, 4, and 8 respectively; S54. Calculate the temporal variances of the facial features, limb motion features, and speech features; S55. Establish dynamic classification weights based on the temporal variances to obtain the classification weights at the current moment; S56. Based on the classification weights at the current moment, weighted aggregate the facial fusion features, limb motion fusion features, and speech fusion features, output the classification probability, and obtain the anger emotion recognition result.

8. The method for recognizing angry emotions based on video streams according to claim 7, wherein, The specific content of S55 includes: S55-1. Set the initial classification weights; S55-2. Use the gated recurrent unit to dynamically update the classification weights to obtain the classification weights at the current moment.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on micro-expressions, body movements and voices

    CN113469153A

  • Video face emotion recognition method based on frame attention mechanism

    CN115393933A

  • Micro-expression recognition method based on multi-mode double-branch space-time motion feature fusion

    CN118262397A

  • Multi-granularity layer spectrum fusion occasion cognition method

    CN119226921A

  • Dynamic micro-expression recognition method based on frame weight and related device

    CN119580326A

Cited By

  • Video emotion recognition method based on multiple modes

    CN122157382A

  • Emotion recognition method and system, computer equipment and storage medium

    CN122266035A

  • High-precision micro-expression recognition method and device, electronic equipment and storage medium

    CN122266038A