A method and system for recognizing dynamic facial expressions
Through the method of combining ResNet18 and self-attention ST-Former model, combined with conditional random field and cross-scale fusion module, the problem of difficult to capture timing dependence and local detail changes in dynamic face expression recognition is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510446458.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The prior art is difficult to effectively capture the timing dependence between video frames and changes in facial local details in dynamic facial expression recognition, resulting in poor recognition effect.
The ResNet18 network model is used for convolution and frame differential processing, combined with the self-attention ST-Former model for spatial and temporal attention weighting processing, and feature communication and fusion are enhanced through the conditional random field fusion module and the cross-scale cross-fusion module, and finally the expression recognition results are output through the full connection layer.
It significantly improves the performance and accuracy of dynamic facial expression recognition, can more accurately capture the changes and timing dependence of facial expressions, and improves the robustness and adaptability of the system.
Smart Images

Figure CN119964224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a method and system for recognizing dynamic facial expressions. Background Art
[0002] With the rapid development of deep learning, convolutional neural networks have achieved remarkable results in static expression recognition and have gradually been applied to dynamic expression recognition. However, related technologies still face many challenges in information fusion and edge processing between video frames. First, dynamic expression recognition involves the continuity and change of temporal information between video frames, and when related technologies process the temporal dependence between video frames, it is difficult to effectively capture the subtle expression changes over a long time span. Second, in cases where the expression changes greatly or the expression details are more complex, related technologies often ignore the local facial details and the subtle changes in the edge regions, resulting in poor expression recognition effects. Summary of the Invention
[0003] The main objective of the present invention is to provide a method and system for recognizing dynamic facial expressions, aiming to solve the technical problem of poor recognition effect of related technologies on dynamic facial expressions.
[0004] To achieve the above objective, the present invention provides a method for recognizing dynamic facial expressions, and the method includes the following steps:
[0005] S1, after performing convolution processing on multiple consecutive color images through a ResNet18 network model, perform frame difference processing to obtain three first feature map sets of different scales; wherein, each first feature map set includes a first dynamic feature map and a first static feature map corresponding to each frame of color image;
[0006] S2, respectively perform spatial attention weighting processing on each feature map in the three first feature map sets through a self-attention ST-Former model, and then perform temporal attention weighting processing to obtain three second feature map sets; wherein, each second feature map set includes a second dynamic feature map and a second static feature map corresponding to each frame of color image;
[0007] S3, respectively perform feature communication enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of color image in the three second feature map sets through a conditional random field fusion module to obtain three third feature map sets;
[0008] S4, fuse the three third feature map sets in order from largest to smallest scale through a cross-scale cross-fusion module to obtain a fifth feature map set;
[0009] S5, process each fifth feature map in the fifth feature map set through a fully connected layer and output an expression recognition result.
[0010] In addition, to achieve the above object, the present invention also provides a dynamic facial expression recognition system, which includes:
[0011] The ResNet18 network model is used to perform convolutional processing on multiple consecutive color images and then perform frame difference processing to obtain three first feature map sets of different scales;
[0012] The self-attention ST-Former model is used to perform spatial attention weighting processing on each feature map in the three first feature map sets respectively, and then perform temporal attention weighting processing to obtain three second feature map sets;
[0013] The conditional random field fusion module is used to perform feature interaction enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of color image in the three second feature map sets respectively to obtain three third feature map sets;
[0014] The cross-scale cross-fusion module is used to fuse the three third feature map sets in order from largest to smallest scale to obtain the fifth feature map set;
[0015] The fully connected layer is used to process the fifth feature map set and output the expression recognition result.
[0016] Through multi-level convolutional processing and difference operations, the present invention can refine the extraction of dynamic and static features of facial expressions, providing rich input information for subsequent processing. Secondly, by combining the spatial attention mechanism and the temporal attention mechanism, it is possible to weight features simultaneously in the spatial and temporal dimensions, thereby effectively capturing the changes and temporal dependencies of facial expressions. In addition, through conditional random fields and cross-scale fusion, the dependence between features can be enhanced, enabling the system to more accurately capture the subtle changes of dynamic expressions. Furthermore, by combining various feature enhancement and fusion means, the robustness and accuracy of system recognition are improved, and the system can maintain high recognition performance under different expressions, illuminations, and angles. Thus, it can be seen that the present invention can significantly improve the performance and accuracy of dynamic facial expression recognition as a whole, ensuring strong adaptability and accuracy in complex real-world environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic structural diagram of a dynamic facial expression recognition system of the present invention;
[0018] Figure 2 It is a schematic flowchart of a dynamic facial expression recognition method of the present invention;
[0019] Figure 3 It is a schematic processing flowchart of the self-attention ST-Former model;
[0020] Figure 4 It is a schematic diagram of the processing flow of the conditional random field fusion module;
[0021] Figure 5 It is a schematic diagram of the processing flow of the cross-scale cross-fusion module.
[0022] The realization, functional features and advantages of the purpose of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0023] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The inventive concept of the present application will be further elaborated below in conjunction with some specific embodiments and specific implementation manners.
[0024] Referring to Figure 1 , an embodiment of the present invention provides a dynamic facial expression recognition system, and the dynamic facial expression recognition system may include:
[0025] The ResNet18 network model is used to perform convolutional processing on multiple consecutive color images, and then perform frame difference processing to obtain three first feature map sets of different scales;
[0026] The self-attention ST-Former model is used to perform spatial attention weighting processing on each feature map in each of the three first feature map sets first, and then perform temporal attention weighting processing to obtain three second feature map sets;
[0027] The conditional random field fusion module is used to perform feature communication enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of color image in each of the three second feature map sets to obtain three third feature map sets;
[0028] The cross-scale cross-fusion module is used to fuse the three third feature map sets in order from largest to smallest scale to obtain a fifth feature map set;
[0029] The fully connected layer is used to perform dimensionality reduction conversion processing on the fifth feature map set and output the expression recognition result.
[0030] Furthermore, the self-attention ST-Former model may specifically include:
[0031] The spatial attention unit is used to obtain the spatial attention weight of each feature map in each of the three first feature map sets through the spatial attention mechanism, and use the spatial attention weight to perform spatial attention weighting processing on the feature map to obtain three spatially weighted feature map sets;
[0032] A temporal attention unit is used to obtain the temporal attention matrix of each feature map in each of the three spatially weighted feature map sets through a temporal attention mechanism, and use the temporal attention matrix to perform temporal attention weighting on the feature maps in each spatially weighted feature map set to obtain three second feature map sets.
[0033] Referring Figure 2 , an embodiment of the present invention provides a dynamic facial expression recognition method, which is applied to the above-mentioned dynamic facial expression recognition system, and specifically includes the following steps:
[0034] Step S1: After performing convolutional processing on multiple consecutive color images through a ResNet18 network model, perform frame difference processing to obtain three first feature map sets of different scales.
[0035] Specifically, first continuously extract color images with a size of 224*224 and having RGB channels from a face video stream, such as 16 frames. There are three convolutional kernels of different scales in the ResNet18 network model.
[0036] The specific steps of step S1 include the following steps:
[0037] Step S1-1: Use the convolutional kernel of the first scale in the ResNet18 network model to perform image feature extraction on 16 color images respectively to obtain the first feature map set R1; then perform frame difference processing on the first feature map set R1 to obtain the first dynamic feature map and the first static feature map corresponding to each frame of color image therein;
[0038] Step S1-2: Use the convolutional kernel of the second scale in the ResNet18 network model to perform image feature extraction on 16 color images in the first feature map set R1 to obtain the first feature map set R2; then perform frame difference processing on the first feature map set R2 to obtain the first dynamic feature map and the first static feature map corresponding to each frame of color image therein;
[0039] Step S1-3: Use the convolutional kernel of the third scale in the ResNet18 network model to perform image feature extraction on 16 color images in the first feature map set R2 to obtain the first feature map set R3; then perform frame difference processing on the first feature map set R3 to obtain the first dynamic feature map and the first static feature map corresponding to each frame of color image therein.
[0040] In step S1, static feature information and dynamic feature information of different scales in the color image can be captured, which helps the system to capture local information, global information, and subtle change information of facial expressions simultaneously.
[0041] Step S2: For each feature map in each of the three first feature map sets, perform spatial attention weighting processing through the self-attention ST-Former model first, and then perform temporal attention weighting processing to obtain three second feature map sets.
[0042] For ease of understanding the solution, here each feature map represents any dynamic feature map or static feature map in the three first feature map sets in Step S1.
[0043] Further, as Figure 3 shown, Step S2 specifically includes:
[0044] Step S21: Through the spatial attention mechanism, obtain the spatial attention weights of each feature map in each of the three first feature map sets, and use the spatial attention weights to perform spatial attention weighting processing on the feature maps to obtain three spatially weighted feature map sets.
[0045] Step S21 specifically includes:
[0046] Step S21-1: Perform spatial embedding processing on each feature map in each of the three first feature map sets to obtain the spatial feature vector corresponding to each feature map; among them, the spatial embedding processing includes segmentation processing, flattening processing, and convolutional projection processing.
[0047] Step S21-2: Use the multi-head attention mechanism to process the spatial feature vectors, and after performing normalization processing using the Softmax function, obtain the spatial attention weights corresponding to each feature map.
[0048] Step S21-3: After using the spatial attention weights to perform weighting processing on the corresponding feature maps, then pass each feature map through layer normalization 1, a feed-forward network, layer normalization 2, and average pooling processing in sequence, so as to obtain three spatially weighted feature map sets.
[0049] Step S22: Through the temporal attention mechanism, obtain the temporal attention matrix of each feature map in each of the three spatially weighted feature map sets, and use the temporal attention matrix to perform temporal attention weighting processing on the feature maps in each of the three spatially weighted feature map sets to obtain three second feature map sets, namely the second feature map set R21, the second feature map set R22, and the second feature map set R23.
[0050] Step S22 specifically includes:
[0051] Step S22-1: Perform temporal embedding processing on each feature map in each of the three spatially weighted feature map sets to obtain the temporal feature vector corresponding to each feature map; among them, the temporal embedding processing includes flattening processing and projection encoding processing.
[0052] Step S22-2: Process the time feature vector using the multi-head attention mechanism, and after normalization using the Softmax function, obtain the time attention matrix corresponding to each feature map.
[0053] Step S22-3: After weighting the corresponding feature maps using the time attention matrix, pass each feature map through layer normalization 3, feed-forward network 2, and layer normalization 4 in sequence, thereby obtaining three second feature map sets.
[0054] Among them, each second feature map set contains a second dynamic feature map and a second static feature map corresponding to each frame of the color image, which are simply referred to as "feature maps" in the description of step S2, that is, each second feature map set contains 32 feature maps.
[0055] In step 2, by combining spatial attention weighting processing and time attention weighting processing, joint modeling of spatio-temporal features is performed. While effectively fusing local and global information of facial expressions, it can also effectively capture dependencies within a long time span, avoiding the problem of gradient disappearance encountered by traditional RNN models in long sequence modeling, being able to handle long sequence facial expression changes, and recognizing complex expression evolutions.
[0056] Step S3: Through the conditional random field fusion module, perform feature communication enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets, obtaining three third feature map sets.
[0057] Further, as Figure 4 shown, step S3 specifically includes:
[0058] Step S31: For the second dynamic feature map and the second static feature map corresponding to each frame of the color image, construct a unary potential energy.
[0059] Step S32: Through the binary potential energy function, perform feature communication processing on the unary potential energy to obtain the dynamic feature map D1 and the static feature map D2 respectively.
[0060] Further, step S32 specifically includes:
[0061] Step S32-1: Construct a binary potential energy function, which includes convolution processing, Prelu activation function, feature connection processing, and Relu activation function.
[0062] Step S32-2: Use the binary potential energy function to perform feature communication processing on the second dynamic feature map and the second static feature map in the unary potential energy respectively, obtaining the dynamic feature map D1 and the static feature map D2.
[0063] Further, the feature communication process in step S32-2 specifically includes:
[0064] Step S32-2-1: After performing convolution processing on the second dynamic feature map and the second static feature map respectively, and then using the PReLU activation function for processing, the third dynamic feature map and the third static feature map are obtained.
[0065] Step S32-2-2: Perform feature connection processing on the second dynamic feature map and the third dynamic feature map, and obtain the dynamic feature map D1 after passing through the Relu activation function; at the same time, perform feature connection processing on the second static feature map and the third static feature map, and obtain the static feature map D2 after passing through the Relu activation function.
[0066] Step S33: Based on the mean fusion strategy, fuse the dynamic feature map D1 and the static feature map D2 to obtain the third feature map corresponding to this frame of color image.
[0067] Step S34: Until all frames of color images have completed steps S31 to S33, three third feature map sets are obtained, namely the third feature map set R31, the third feature map set R32, and the third feature map set R33.
[0068] In step S3, the conditional random field (CRF) models local dependencies through the binary potential function and the mean fusion strategy, and can effectively capture the fine-grained information interaction between the dynamic and static feature maps. Specifically, facial expressions are not only affected by the current frame, but also have temporal dependencies with the previous and subsequent frames. The conditional random field can efficiently perform the communication of dynamic and static features, thus helping to identify subtle expression changes.
[0069] Step S4: The three third feature map sets are successively fused from large to small in scale through the cross-scale cross-fusion module to obtain the fifth feature map set.
[0070] The three third feature map sets obtained in step S3 are, from large to small in scale, the third feature map set R31, the third feature map set R32, and the third feature map set R33. The specific steps of step S4 include:
[0071] Step S41: Extract the third feature maps corresponding to the same frame of color image from the third feature map set R31 and the third feature map set R32, and input them into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed, thereby obtaining the fourth feature map set R4.
[0072] Such as Figure 5As shown in the figure, the cross-scale cross-fusion module includes a first multi-head attention unit and a second multi-head attention unit with the same structure. Among them, the first multi-head attention unit includes a multi-head attention mechanism layer, a splicing layer, a feature connection layer, a linear layer, and a non-linear transformation layer. The multi-head attention mechanism layer has three attention heads. The second multi-head attention unit has the same structure as the first multi-head attention unit, so it will not be described in detail.
[0073] Step S41 specifically includes:
[0074] Step S41-1: Perform layer normalization on the two extracted third feature maps respectively. After obtaining the query matrix, key matrix, and value matrix corresponding to each third feature map, perform cross-attention processing using six attention heads.
[0075] Specifically, perform layer normalization on the third feature map corresponding to any frame of color image in the third feature map set R31 to obtain the query matrix Q1, key matrix K1, and value matrix V1 corresponding to this third feature map. Perform layer normalization on the third feature map corresponding to the same frame of color image in the third feature map set R32 to obtain the query matrix Q2, key matrix K2, and value matrix V2 corresponding to this third feature map. Then, jointly input the query matrix Q1, key matrix K2, and value matrix V2 into the three attention heads in the first multi-head attention unit for multi-head attention processing. That is to say, jointly input the query matrix Q1, key matrix K2, and value matrix V2 into the first attention head in the first multi-head attention unit, also jointly input them into the second attention head in the first multi-head attention unit, and also jointly input them into the third attention head in the first multi-head attention unit. At the same time, jointly input the query matrix Q2, key matrix K1, and value matrix V1 into the three attention heads in the second multi-head attention unit for multi-head attention processing.
[0076] Step S41-2: The feature maps output by the three attention heads in the first multi-head attention unit are spliced to obtain a feature map T1; the feature maps output by the three attention heads in the second multi-head attention unit are spliced to obtain a feature map T2.
[0077] Step S41-3: Feature-connect the third feature map in the third feature map set R31 with the feature map T1 to obtain a feature map Z1; feature-connect the third feature map in the third feature map set R32 with the feature map T2 to obtain a feature map Z2.
[0078] Step S41-4: Input the feature map Z1 into the linear layer and non-linear transformation layer of the first multi-head attention unit in sequence to obtain the feature map Y1; input the feature map Z2 into the linear layer and non-linear transformation layer of the second multi-head attention unit in sequence to obtain the feature map Y2; after performing feature connection on the feature maps Y1 and Y2, perform mean fusion to obtain the fourth feature map.
[0079] Step S41-5: Return to Step S41-1 until all frames are processed to obtain the fourth feature map set R4. The fourth feature map set R4 contains the fourth feature maps corresponding to 16 frames of color images.
[0080] Step S42: Input the third feature maps corresponding to each frame of color image in the third feature map set R33 and the fourth feature maps corresponding to the same frame of color image in the fourth feature map set R4 into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed to obtain the fifth feature map set R5.
[0081] The cross-scale cross-fusion module used in Step S42 is the same as that in Step S41, and the cross-fusion processing process is also the same as that in Step S41, so it will not be elaborated here. This step finally obtains the fifth feature map set R5, which contains the fifth feature maps corresponding to 16 frames of color images.
[0082] In Step S4, through the multi-head cross-attention mechanism, it is possible to focus on different types of feature relationships in feature maps of different scales, which can promote the effective information transmission and interaction between features. Thus, in the process of dynamic expression recognition, it can effectively improve the recognition ability of the model, enhance the accuracy and robustness of feature information transmission, and better capture complex expression changes and time dependencies.
[0083] Step S5: Process each fifth feature map in the fifth feature map set through a fully connected layer to output the expression recognition result.
[0084] Further, Step S5 specifically includes:
[0085] Step S51: Normalize each fifth feature map in the fifth feature map set to obtain 16 corresponding groups of feature vectors.
[0086] Step S52: Use the weight matrix obtained by deep learning to perform mapping transformation processing on each group of feature vectors to obtain the scores of multiple expression categories corresponding to each fifth feature map.
[0087] Specifically, if there are 7 expression categories corresponding to the final output result, then in the processing result of Step S52, there are 7 expression scores corresponding to each feature map, totaling 16×7 scores.
[0088] Step S53: Convert the scores into predicted probabilities corresponding to each expression category through the Softmax activation function.
[0089] Step S54: Determine the expression category with the highest predicted probability as the expression recognition result.
[0090] In this embodiment, through multi-level convolutional processing and differential operations, the dynamic and static features of facial expressions can be refined to provide rich input information for subsequent processing. Secondly, by combining the spatial attention mechanism and the temporal attention mechanism, features can be weighted simultaneously in both the spatial and temporal dimensions, thus effectively capturing the changes and temporal dependencies of facial expressions. In addition, through conditional random fields and cross-scale fusion, the dependencies between features can be enhanced, enabling the system to more accurately capture the subtle changes in dynamic expressions. Furthermore, by combining various feature enhancement and fusion means, the robustness and accuracy of system recognition are improved, and high recognition performance can be maintained under different expressions, illuminations, and angles. It can be seen that the present invention can significantly improve the performance and accuracy of dynamic human face expression recognition as a whole, ensuring strong adaptability and accuracy in complex real-world environments.
[0091] It should be noted that the functions that can be realized by each module in the dynamic human face expression recognition system provided in this embodiment and the corresponding technical effects achieved can refer to the descriptions of the specific implementation manners in the various embodiments of the dynamic human face expression recognition method of the present invention. For the sake of brevity of the specification, they will not be elaborated here.
[0092] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0093] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for recognizing dynamic facial expressions, characterized in that: The method comprises: S1, after convolution processing is performed on multiple frames of continuous color images through a ResNet18 network model, frame difference processing is performed to obtain three first feature atlas sets of different scales; wherein each of the first feature atlas sets includes a first dynamic feature map and a first static feature map corresponding to each frame of the color image; S2, performing spatial attention weighted processing on each feature map in the three first feature map sets through the self-attention ST-Former model, and then performing temporal attention weighted processing to obtain three second feature map sets; wherein each of the second feature map sets includes a second dynamic feature map and a second static feature map corresponding to each frame of the color image; S3, performing feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets through a conditional random field fusion module to obtain three third feature map sets; S4, fusing the three third feature atlases in order from large to small scales through a cross-scale cross fusion module to obtain a fifth feature atlas; S5, processing each fifth feature map in the fifth feature map set through a fully connected layer, and outputting an expression recognition result.
2. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The S2 specifically includes: S21, respectively obtaining the spatial attention weight of each feature map in each of the first feature map sets through the spatial attention mechanism, and performing spatial attention weighted processing on the feature maps using the spatial attention weights to obtain three spatial weighted feature map sets; S22, obtaining the temporal attention matrix of each feature map in each of the spatially weighted feature maps through the temporal attention mechanism, and performing temporal attention weighted processing on the feature maps in each of the spatially weighted feature maps using the temporal attention matrix to obtain three second feature maps, namely, second feature map set R21, second feature map set R22 and second feature map set R23.
3. The method for recognizing dynamic facial expressions as claimed in claim 2, characterized in that: The S21 specifically includes: S21-1, performing spatial embedding processing on each feature map in each of the first feature map sets to obtain a spatial feature vector corresponding to each feature map; wherein the spatial embedding processing includes segmentation processing, flattening processing, and convolution projection processing; S21-2, using a multi-head attention mechanism to process the spatial feature vector, and after normalizing it using a Softmax function, obtain a spatial attention weight corresponding to each feature map; S21-3, after using the spatial attention weight to perform weighted processing on the corresponding feature map, each feature map is processed in turn through layer normalization 1, feedforward network, layer normalization 2 and average pooling, so as to obtain the three spatial weighted feature map sets.
4. The method for recognizing dynamic facial expressions as claimed in claim 2, characterized in that: The S22 specifically includes: S22-1, performing a time embedding process on each feature map in each of the spatial weighted feature map sets to obtain a time feature vector corresponding to each feature map; wherein the time embedding process includes a flattening process and a projection coding process; S22-2, using a multi-head attention mechanism to process the temporal feature vector, and after normalizing it using a Softmax function, obtain a temporal attention matrix corresponding to each feature map; S22-3, after using the temporal attention matrix to perform weighted processing on the corresponding feature map, each feature map is processed in turn through layer normalization 3, feedforward network 2 and layer normalization 4 to obtain three sets of the second feature maps.
5. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The S3 specifically includes: S31, constructing a unary potential energy for the second dynamic feature map and the second static feature map corresponding to each frame of the color image; S32, performing characteristic exchange processing on the one-dimensional potential energy through a binary potential energy function to obtain a dynamic characteristic graph D1 and a static characteristic graph D2 respectively; S33, based on the mean fusion strategy, the dynamic feature map D1 and the static feature map D2 are fused to obtain a third feature map corresponding to the frame color image; S34, until all frames of color images have completed steps S31 to S33, thereby obtaining the three third feature atlases, namely, the third feature atlas R31, the third feature atlas R32, and the third feature atlas R33.
6. The method for recognizing dynamic facial expressions as claimed in claim 5, characterized in that: The S32 specifically includes: S32-1, constructing a binary potential energy function, wherein the binary potential energy function includes convolution processing, Prelu activation function, feature connection processing, and Relu activation function; S32-2, using a binary potential energy function to perform feature exchange processing on the second dynamic feature graph and the second static feature graph in the univariate potential energy respectively, to obtain the dynamic feature graph D1 and the static feature graph D2.
7. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The three third feature atlases are, in descending order of scale, a third feature atlas R31, a third feature atlas R32, and a third feature atlas R33, and S4 specifically includes: S41, extracting the third feature map corresponding to the same frame color image from the third feature map set R31 and the third feature map set R32, and inputting the third feature map into the cross-scale cross fusion module for cross fusion processing until all frames are processed, thereby obtaining the fourth feature map set R4; S42, input the third feature map corresponding to each frame color image in the third feature atlas R33 and the fourth feature map corresponding to the same frame color image in the fourth feature atlas R4 into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed, thereby obtaining the fifth feature atlas.
8. The method for recognizing dynamic facial expressions as claimed in claim 7, characterized in that: The cross-scale cross-fusion module includes a first multi-head attention unit and a second multi-head attention unit with the same structure; the first multi-head attention unit includes a multi-head attention mechanism layer, a splicing layer, a feature connection layer, a linear layer and a nonlinear transformation layer; the multi-head attention mechanism layer has three attention heads, and S41 specifically includes: S41-1, performing layer normalization processing on the two extracted third feature maps respectively, so as to obtain the query matrix, key matrix and value matrix corresponding to each third feature map, and then performing cross attention processing using six attention heads; S41-2, the feature maps output by the three attention heads in the first multi-head attention unit are concatenated to obtain a feature map T1; the feature maps output by the three attention heads in the second multi-head attention unit are concatenated to obtain a feature map T2; S41-3, performing multi-head self-attention processing on three first query matrices, three second key matrices, and three second value matrices through the first multi-head attention unit to obtain a feature map Z1; S41-4, performing multi-head self-attention processing on the three second query matrices, the three first key matrices, and the three first value matrices through the second multi-head attention unit to obtain a feature map Z2; S41-5, after performing feature connection processing on the feature map Z1 and the feature map Z2, perform mean fusion processing to obtain a feature map Z3.
9. A dynamic facial expression recognition system, characterized in that: The system comprises: The ResNet18 network model is used to perform convolution processing on multiple frames of continuous color images and then perform frame difference processing to obtain three first feature atlases of different scales; A self-attention ST-Former model is used to perform spatial attention weighted processing on each feature map in the three first feature map sets, and then perform temporal attention weighted processing to obtain three second feature map sets; A conditional random field fusion module is used to perform feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets, so as to obtain three third feature map sets; A cross-scale cross-fusion module, used to fuse the three third feature atlases in order from large to small scales to obtain a fifth feature atlas; The fully connected layer is used to process the fifth feature atlas and output the expression recognition result.
10. The dynamic facial expression recognition system as claimed in claim 9, characterized in that: The self-attention ST-Former model includes: A spatial attention unit, used for obtaining a spatial attention weight of each feature map in each of the first feature map sets through a spatial attention mechanism, and performing spatial attention weighted processing on the feature maps using the spatial attention weights to obtain three spatial weighted feature map sets; A temporal attention unit is used to obtain the temporal attention matrix of each feature map in the three spatially weighted feature map sets through the temporal attention mechanism, and use the temporal attention matrix to perform temporal attention weighted processing on the feature maps in each of the spatially weighted feature map sets to obtain the three second feature map sets.
Citation Information
Patent Citations
CNN (Convolutional Neural Network) and Transform improved lightweight monocular depth estimation method
CN117710429A
Dynamic micro-expression recognition method based on frame weight and related device
CN119580326A