Dynamic facial expression recognition method and system

Through the combination of ResNet18 and self-attention ST-Former model, the problem of difficult to capture timing dependence and local detail changes in dynamic face expression recognition is solved, and higher recognition accuracy and robustness are achieved.

CN119964224AActive Publication Date: 2025-05-09XIHUA UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510446458.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the timing dependence between video frames and changes in facial local details in dynamic facial expression recognition, resulting in poor recognition effect.

Method used

The ResNet18 network model is used for convolution and frame differential processing, combined with the self-attention ST-Former model for spatial and temporal attention weighting processing, and feature communication and fusion are enhanced through the conditional random field fusion module and the cross-scale cross-fusion module, and finally the expression recognition results are output through the full connection layer.

Benefits of technology

It significantly improves the performance and accuracy of dynamic facial expression recognition, can more accurately capture subtle changes and timing dependence of dynamic expressions, and improves the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964224A_ABST
    Figure CN119964224A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic facial expression recognition method and system, and belongs to the technical field of image recognition. The method comprises the following steps: performing convolution and frame difference processing on multiple frames of continuous color images through a ResNet18 network model to obtain three first feature image sets with different scales; performing space attention weighting and time attention weighting on each feature map in the three first feature map sets through a self-attention ST-Former model to obtain three second feature map sets; performing feature exchange enhancement on a second dynamic feature map and a second static feature map corresponding to each frame of color image in the three second feature map sets through a conditional random field fusion module to obtain three third feature map sets; fusing the three third feature image sets in sequence from large to small according to scales through a cross-scale cross fusion module to obtain a fifth feature image set; and processing the fifth feature image set through a full connection layer, and outputting an expression recognition result. According to the invention, the recognition capability of the system on dynamic facial expressions can be effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method and system for recognizing dynamic facial expressions. Background Art

[0002] With the rapid development of deep learning, convolutional neural networks have achieved remarkable results in static expression recognition and have gradually been applied to dynamic expression recognition. However, related technologies still face many challenges in information fusion and edge processing between video frames. First, dynamic expression recognition involves the continuity and changes of temporal information between video frames, and related technologies have difficulty in effectively capturing subtle changes in expression over a long time span when processing the temporal dependencies between video frames. Second, in cases where the expression changes are large or the expression details are complex, related technologies often ignore local facial details and subtle changes in edge areas, resulting in poor expression recognition results. Summary of the invention

[0003] The main purpose of the present invention is to provide a method and system for recognizing dynamic facial expressions, aiming to solve the technical problem that the related technology has poor recognition effect on dynamic facial expressions.

[0004] To achieve the above object, the present invention provides a method for recognizing dynamic facial expressions, the method comprising the following steps:

[0005] S1, after convolution processing is performed on multiple frames of continuous color images through the ResNet18 network model, frame difference processing is performed to obtain three first feature atlas sets of different scales; wherein each first feature atlas set includes a first dynamic feature map and a first static feature map corresponding to each frame of color image;

[0006] S2, using the self-attention ST-Former model to perform spatial attention weighted processing on each feature map in the three first feature map sets, and then perform temporal attention weighted processing to obtain three second feature map sets; wherein each second feature map set includes a second dynamic feature map and a second static feature map corresponding to each frame of the color image;

[0007] S3, performing feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets through a conditional random field fusion module to obtain three third feature map sets;

[0008] S4, fusing the three third feature atlases in order from large to small scales through a cross-scale cross fusion module to obtain a fifth feature atlas;

[0009] S5, processing each fifth feature map in the fifth feature map set through a fully connected layer, and outputting the expression recognition result.

[0010] In addition, to achieve the above object, the present invention also provides a dynamic facial expression recognition system, the system comprising:

[0011] The ResNet18 network model is used to perform convolution processing on multiple frames of continuous color images and then perform frame difference processing to obtain three first feature atlases of different scales;

[0012] The self-attention ST-Former model is used to perform spatial attention weighted processing on each feature map in the three first feature map sets, and then perform temporal attention weighted processing to obtain three second feature map sets;

[0013] A conditional random field fusion module is used to perform feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets, so as to obtain three third feature map sets;

[0014] A cross-scale cross-fusion module is used to fuse the three third feature atlases in order from large to small scales to obtain a fifth feature atlas;

[0015] The fully connected layer is used to process the fifth feature atlas and output the expression recognition result.

[0016] The present invention can extract the dynamic and static features of facial expressions in a refined manner through multi-level convolution processing and differential operations, providing rich input information for subsequent processing. Secondly, by combining the spatial attention mechanism and the temporal attention mechanism, it is possible to weight features in both spatial and temporal dimensions, thereby effectively capturing the changes and temporal dependencies of facial expressions. In addition, through conditional random fields and cross-scale fusion, the dependencies between features can be enhanced, so that the system can more accurately capture subtle changes in dynamic expressions. Furthermore, by combining a variety of feature enhancement and fusion methods, the robustness and accuracy of system recognition are improved, and high recognition performance can be maintained under different expressions, lighting, and angles. It can be seen that the present invention can significantly improve the performance and accuracy of dynamic facial expression recognition as a whole, ensuring strong adaptability and accuracy in complex real-world environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A structural schematic diagram of a dynamic facial expression recognition system of the present invention;

[0018] Figure 2 A schematic diagram of a flow chart of a method for recognizing dynamic facial expressions of the present invention;

[0019] Figure 3 Schematic diagram of the processing flow of the self-attention ST-Former model;

[0020] Figure 4 This is a schematic diagram of the processing flow of the conditional random field fusion module;

[0021] Figure 5 Schematic diagram of the cross-scale cross-fusion module processing flow.

[0022] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0023] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The inventive concept of the present application is further described below in conjunction with some specific embodiments and specific implementation methods.

[0024] Reference Figure 1 The embodiment of the present invention provides a dynamic facial expression recognition system, and the dynamic facial expression recognition system may include:

[0025] The ResNet18 network model is used to perform convolution processing on multiple frames of continuous color images and then perform frame difference processing to obtain three first feature atlases of different scales;

[0026] The self-attention ST-Former model is used to perform spatial attention weighted processing on each feature map in the three first feature map sets, and then perform temporal attention weighted processing to obtain three second feature map sets;

[0027] A conditional random field fusion module is used to perform feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets, so as to obtain three third feature map sets;

[0028] A cross-scale cross-fusion module is used to fuse the three third feature atlases in order from large to small scales to obtain a fifth feature atlas;

[0029] The fully connected layer is used to perform dimensionality reduction transformation on the fifth feature atlas and output the expression recognition result.

[0030] Furthermore, the self-attention ST-Former model can specifically include:

[0031] A spatial attention unit, used for obtaining a spatial attention weight of each feature map in each first feature map set through a spatial attention mechanism, and performing spatial attention weighted processing on the feature map using the spatial attention weight to obtain three spatial weighted feature map sets;

[0032] The temporal attention unit is used to obtain the temporal attention matrix of each feature map in the three spatial weighted feature map sets through the temporal attention mechanism, and use the temporal attention matrix to perform temporal attention weighted processing on the feature maps in each spatial weighted feature map set to obtain three second feature map sets.

[0033] Reference Figure 2 The embodiment of the present invention provides a dynamic facial expression recognition method, which is applied to the above-mentioned dynamic facial expression recognition system, and specifically includes the following steps:

[0034] Step S1: After performing convolution processing on multiple frames of continuous color images through the ResNet18 network model, frame difference processing is performed to obtain three first feature atlases of different scales.

[0035] Specifically, first, color images with a size of 224*224 and RGB channels are continuously extracted from the face video stream, for example, the number of frames is 16. There are three convolution kernels of different scales in the ResNet18 network model.

[0036] The step S1 specifically includes the following steps:

[0037] Step S1-1: Use the convolution kernel of the first scale in the ResNet18 network model to extract image features from 16 frames of color images respectively to obtain a first feature atlas R1; then perform frame difference processing on the first feature atlas R1 to obtain a first dynamic feature map and a first static feature map corresponding to each frame of the color image;

[0038] Step S1-2: Use the convolution kernel of the second scale in the ResNet18 network model to extract image features from the 16 frames of color images in the first feature atlas R1 to obtain the first feature atlas R2; then perform frame difference processing on the first feature atlas R2 to obtain the first dynamic feature map and the first static feature map corresponding to each frame of the color image;

[0039] Step S1-3: Use the third scale convolution kernel in the ResNet18 network model to extract image features of the 16 frames of color images in the first feature atlas R2 to obtain the first feature atlas R3; then perform frame difference processing on the first feature atlas R3 to obtain the first dynamic feature map and the first static feature map corresponding to each frame of the color image.

[0040] In step S1, static feature information and dynamic feature information of different scales in the color image can be captured, which helps the system to simultaneously capture local information, global information, and subtle change information of facial expressions.

[0041] Step S2: Through the self-attention ST-Former model, each feature map in the three first feature map sets is first subjected to spatial attention weighting processing, and then subjected to temporal attention weighting processing to obtain three second feature map sets.

[0042] To facilitate understanding of the solution, each feature map here represents any dynamic feature map or static feature map in the three first feature map sets in step S1.

[0043] Further, such as Figure 3 As shown, step S2 specifically includes:

[0044] Step S21: Obtain the spatial attention weight of each feature map in each first feature map set through the spatial attention mechanism, perform spatial attention weighted processing on the feature map using the spatial attention weight, and obtain three spatial weighted feature map sets.

[0045] Step S21 specifically includes:

[0046] Step S21-1: perform spatial embedding processing on each feature map in each first feature map set to obtain a spatial feature vector corresponding to each feature map; wherein the spatial embedding processing includes segmentation processing, flattening processing and convolution projection processing.

[0047] Step S21-2: Use the multi-head attention mechanism to process the spatial feature vector, and use the Softmax function to normalize it to obtain the spatial attention weight corresponding to each feature map.

[0048] Step S21-3: After weighting the corresponding feature maps using the spatial attention weights, each feature map is processed in turn through layer normalization 1, feedforward network, layer normalization 2, and average pooling to obtain three spatial weighted feature map sets.

[0049] Step S22: Obtain the temporal attention matrix of each feature map in each spatial weighted feature map set through the temporal attention mechanism, and use the temporal attention matrix to perform temporal attention weighted processing on the feature maps in each spatial weighted feature map set to obtain three second feature map sets, namely, the second feature map set R21, the second feature map set R22 and the second feature map set R23.

[0050] Step S22 specifically includes:

[0051] Step S22-1: Perform time embedding processing on each feature map in each spatial weighted feature map set to obtain a time feature vector corresponding to each feature map; wherein the time embedding processing includes flattening processing and projection coding processing.

[0052] Step S22-2: Use the multi-head attention mechanism to process the temporal feature vector, and use the Softmax function to normalize it to obtain the temporal attention matrix corresponding to each feature map.

[0053] Step S22-3: After weighting the corresponding feature maps using the temporal attention matrix, each feature map is processed in turn through layer normalization 3, feedforward network 2, and layer normalization 4 to obtain three second feature map sets.

[0054] Among them, each second feature map set contains a second dynamic feature map and a second static feature map corresponding to each frame of color image, which are referred to as "feature maps" in the description of step S2, that is, each second feature map set has 32 feature maps.

[0055] In step 2, by combining spatial attention weighted processing and temporal attention weighted processing, joint modeling of spatiotemporal features is performed. While effectively integrating local information and global information of facial expressions, it can also effectively capture dependencies over long time spans, avoiding the gradient vanishing problem encountered by traditional RNN models in long sequence modeling. It can handle long sequences of facial expression changes and recognize complex expression evolutions.

[0056] Step S3: Perform feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets through the conditional random field fusion module to obtain three third feature map sets.

[0057] Further, such as Figure 4 As shown, step S3 specifically includes:

[0058] Step S31: constructing a unary potential energy for the second dynamic feature map and the second static feature map corresponding to each frame of color image.

[0059] Step S32: Perform characteristic exchange processing on the one-dimensional potential energy through the binary potential energy function to obtain a dynamic characteristic graph D1 and a static characteristic graph D2 respectively.

[0060] Further, step S32 specifically includes:

[0061] Step S32-1: construct a binary potential energy function, which includes convolution processing, Prelu activation function, feature connection processing, and Relu activation function.

[0062] Step S32-2: Use the binary potential energy function to perform feature exchange processing on the second dynamic feature graph and the second static feature graph in the univariate potential energy respectively, to obtain the dynamic feature graph D1 and the static feature graph D2.

[0063] Furthermore, the feature communication processing process in step S32-2 specifically includes:

[0064] Step S32-2-1: After performing convolution processing on the second dynamic feature map and the second static feature map respectively, they are processed using the Prelu activation function to obtain a third dynamic feature map and a third static feature map.

[0065] Step S32-2-2: perform feature connection processing on the second dynamic feature map and the third dynamic feature map, and obtain the dynamic feature map D1 after passing the Relu activation function; at the same time, perform feature connection processing on the second static feature map and the third static feature map, and obtain the static feature map D2 after passing the Relu activation function.

[0066] Step S33: Based on the mean fusion strategy, the dynamic feature map D1 and the static feature map D2 are fused to obtain a third feature map corresponding to the frame color image.

[0067] Step S34: until all frames of color images have completed steps S31 to S33, three third feature atlases are obtained, namely, the third feature atlas R31, the third feature atlas R32, and the third feature atlas R33.

[0068] In step S3, the conditional random field (CRF) models local dependencies through a binary potential function and a mean fusion strategy, which can effectively capture the fine-grained information interaction between dynamic and static feature maps. Specifically, facial expressions are not only affected by the current frame, but also have temporal dependencies with the previous and next frames. The conditional random field can efficiently communicate dynamic and static features, thereby helping to identify subtle changes in expression.

[0069] Step S4: The three third feature atlases are fused in sequence from large to small scales through a cross-scale cross fusion module to obtain a fifth feature atlas.

[0070] The three third feature atlases obtained in step S3 are, in descending order of scale, the third feature atlas R31, the third feature atlas R32, and the third feature atlas R33. The step S4 specifically includes:

[0071] Step S41: extract the third feature map corresponding to the same frame color image from the third feature atlas R31 and the third feature atlas R32, and input it into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed, thereby obtaining the fourth feature atlas R4.

[0072] like Figure 5As shown, the cross-scale cross-fusion module includes a first multi-head attention unit and a second multi-head attention unit with the same structure; wherein the first multi-head attention unit includes a multi-head attention mechanism layer, a splicing layer, a feature connection layer, a linear layer, and a nonlinear transformation layer; the multi-head attention mechanism layer has three attention heads. The structure of the second multi-head attention unit is the same as that of the first multi-head attention unit, so it is not repeated here.

[0073] The step S41 specifically includes:

[0074] Step S41-1: perform layer normalization processing on the two extracted third feature maps respectively, so as to obtain the query matrix, key matrix and value matrix corresponding to each third feature map, and then use six attention heads to perform cross attention processing.

[0075] Specifically, the third feature map corresponding to any frame of color image in the third feature map set R31 is subjected to layer normalization processing, thereby obtaining the query matrix Q1, key matrix K1 and value matrix V1 corresponding to the third feature map; the third feature map corresponding to the same frame of color image in the third feature map set R32 is subjected to layer normalization processing, thereby obtaining the query matrix Q2, key matrix K2 and value matrix V2 corresponding to the third feature map. Then, the query matrix Q1, key matrix K2 and value matrix V2 are input into the three attention heads in the first multi-head attention unit for multi-head attention processing, that is, the query matrix Q1, key matrix K2 and value matrix V2 are input into the first attention head in the first multi-head attention unit, and are also input into the second attention head in the first multi-head attention unit, and are also input into the third attention head in the first multi-head attention unit. At the same time, the query matrix Q2, key matrix K1 and value matrix V1 are input into the three attention heads in the second multi-head attention unit for multi-head attention processing.

[0076] Step S41-2: The feature maps output by the three attention heads in the first multi-head attention unit are concatenated to obtain the feature map T1; the feature maps output by the three attention heads in the second multi-head attention unit are concatenated to obtain the feature map T2.

[0077] Step S41 - 3 : Feature-connect the third feature map in the third feature map set R31 with the feature map T1 to obtain the feature map Z1 ; feature-connect the third feature map in the third feature map set R32 with the feature map T2 to obtain the feature map Z2 .

[0078] Step S41-4: Input the feature map Z1 into the linear layer and the nonlinear transformation layer of the first multi-attention unit in sequence to obtain the feature map Y1; input the feature map Z2 into the linear layer and the nonlinear transformation layer of the second multi-attention unit in sequence to obtain the feature map Y2; perform feature connection on the feature map Y1 and the feature map Y2, and then perform mean fusion to obtain the fourth feature map.

[0079] Step S41 - 5 : Return to step S41 - 1 until all frames are processed, thereby obtaining a fourth feature atlas set R4 . The fourth feature atlas set R4 contains fourth feature maps corresponding to 16 frames of color images.

[0080] Step S42: The third feature map corresponding to each frame color image in the third feature atlas R33 and the fourth feature map corresponding to the same frame color image in the fourth feature atlas R4 are input into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed, thereby obtaining the fifth feature atlas R5.

[0081] The cross-scale cross-fusion module used in step S42 is the same as that in step S41, and the cross-fusion processing performed is also the same as that in step S41, so it will not be repeated here. This step finally obtains the fifth feature atlas R5, which has the fifth feature maps corresponding to 16 frames of color images.

[0082] In step S4, the multi-head cross attention mechanism can focus on different types of feature relationships in feature maps of different scales, which can promote effective information transfer and interaction between features. Therefore, in the process of dynamic expression recognition, the recognition ability of the model can be effectively improved, the accuracy and robustness of feature information transfer can be enhanced, and complex expression changes and time dependencies can be better captured.

[0083] Step S5, processing each fifth feature map in the fifth feature map set through a fully connected layer, and outputting the expression recognition result.

[0084] Furthermore, step S5 specifically includes:

[0085] Step S51: normalize each fifth feature map in the fifth feature map set to obtain corresponding 16 groups of feature vectors.

[0086] Step S52: Use the weight matrix obtained by deep learning to perform mapping conversion processing on each group of feature vectors respectively, and obtain the scores of each fifth feature map corresponding to multiple expression categories.

[0087] Specifically, if there are 7 expression categories corresponding to the final output result, then in the processing result of step S52, there are 7 expression scores corresponding to each feature map, including a total of 16×7 scores.

[0088] Step S53: converting the score into a prediction probability corresponding to each expression category through a Softmax activation function.

[0089] Step S54: Determine the expression category with the highest prediction probability as the expression recognition result.

[0090] In this embodiment, through multi-level convolution processing and differential operations, the dynamic and static features of facial expressions can be refined and extracted, providing rich input information for subsequent processing. Secondly, by combining the spatial attention mechanism and the temporal attention mechanism, the features can be weighted in both the spatial and temporal dimensions, thereby effectively capturing the changes and temporal dependencies of facial expressions. In addition, through conditional random fields and cross-scale fusion, the dependencies between features can be enhanced, so that the system can more accurately capture the subtle changes in dynamic expressions. Furthermore, by combining a variety of feature enhancement and fusion methods, the robustness and accuracy of system recognition are improved, and high recognition performance can be maintained under different expressions, lighting, angles, etc. It can be seen that the present invention can significantly improve the performance and accuracy of dynamic facial expression recognition as a whole, ensuring strong adaptability and accuracy in complex real-world environments.

[0091] It should be noted that the functions that can be realized by each module in the dynamic facial expression recognition system provided in this embodiment and the corresponding technical effects achieved can refer to the description of the specific implementation methods in each embodiment of the dynamic facial expression recognition method of the present invention. For the sake of brevity of the specification, they will not be repeated here.

[0092] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0093] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for recognizing dynamic facial expressions, characterized in that: The method comprises: S1, after convolution processing is performed on multiple frames of continuous color images through a ResNet18 network model, frame difference processing is performed to obtain three first feature atlas sets of different scales; wherein each of the first feature atlas sets includes a first dynamic feature map and a first static feature map corresponding to each frame of the color image; S2, performing spatial attention weighted processing on each feature map in the three first feature map sets through the self-attention ST-Former model, and then performing temporal attention weighted processing to obtain three second feature map sets; wherein each of the second feature map sets includes a second dynamic feature map and a second static feature map corresponding to each frame of the color image; S3, performing feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets through a conditional random field fusion module to obtain three third feature map sets; S4, fusing the three third feature atlases in order from large to small scales through a cross-scale cross fusion module to obtain a fifth feature atlas; S5, processing each fifth feature map in the fifth feature map set through a fully connected layer, and outputting an expression recognition result.

2. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The S2 specifically includes: S21, respectively obtaining the spatial attention weight of each feature map in each of the first feature map sets through the spatial attention mechanism, and performing spatial attention weighted processing on the feature maps using the spatial attention weights to obtain three spatial weighted feature map sets; S22, obtaining the temporal attention matrix of each feature map in each of the spatially weighted feature maps through the temporal attention mechanism, and performing temporal attention weighted processing on the feature maps in each of the spatially weighted feature maps using the temporal attention matrix to obtain three second feature maps, namely, second feature map set R21, second feature map set R22 and second feature map set R23.

3. The method for recognizing dynamic facial expressions as claimed in claim 2, characterized in that: The S21 specifically includes: S21-1, performing spatial embedding processing on each feature map in each of the first feature map sets to obtain a feature vector corresponding to each feature map; wherein the spatial embedding processing includes segmentation processing, flattening processing, and convolution projection processing; S21-2, using a multi-head attention mechanism to process the spatial feature vector, and after normalizing it using a Softmax function, obtain a spatial attention matrix corresponding to each feature map; S21-3, after using the spatial attention weight to perform weighted processing on the corresponding feature map, each feature map is processed in turn through layer normalization 1, feedforward network, layer normalization 2 and average pooling, so as to obtain the three spatial weighted feature map sets.

4. The method for recognizing dynamic facial expressions as claimed in claim 2, characterized in that: The S22 specifically includes: S22-1, performing a time embedding process on each feature map in each of the spatial weighted feature map sets to obtain a time feature vector corresponding to each feature map; wherein the time embedding process includes a flattening process and a projection coding process; S22-2, using a multi-head attention mechanism to process the temporal feature vector, and after normalizing it using a Softmax function, obtain a temporal attention matrix corresponding to each feature map; S22-3, after using the temporal attention matrix to perform weighted processing on the corresponding feature map, each feature map is processed in turn through layer normalization 3, feedforward network 2 and layer normalization 4 to obtain three sets of the second feature maps.

5. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The S3 specifically includes: S31, constructing a unary potential energy for the second dynamic feature map and the second static feature map corresponding to each frame of the color image; S32, performing characteristic exchange processing on the one-dimensional potential energy through a binary potential energy function to obtain a dynamic characteristic graph D1 and a static characteristic graph D2 respectively; S33, based on the mean fusion strategy, the dynamic feature map D1 and the static feature map D2 are fused to obtain a third feature map corresponding to the frame color image; S34, until all frames of color images have completed steps S31 to S33, thereby obtaining the three third feature atlases, namely, the third feature atlas R31, the third feature atlas R32, and the third feature atlas R33.

6. The method for recognizing dynamic facial expressions as claimed in claim 5, characterized in that: The S32 specifically includes: S32-1, constructing a binary potential energy function, wherein the binary potential energy function includes convolution processing, Prelu activation function, feature connection processing, and Relu activation function; S32-2, using a binary potential energy function to perform feature exchange processing on the second dynamic feature graph and the second static feature graph in the univariate potential energy respectively, to obtain the dynamic feature graph D1 and the static feature graph D2.

7. The method for recognizing dynamic facial expressions as claimed in claim 1, characterized in that: The three third feature atlases are, in descending order of scale, a third feature atlas R31, a third feature atlas R32, and a third feature atlas R33, and S4 specifically includes: S41, extracting the third feature map corresponding to the same frame color image from the third feature map set R31 and the third feature map set R32, and inputting the third feature map into the cross-scale cross fusion module for cross fusion processing until all frames are processed, thereby obtaining the fourth feature map set R4; S42, input the third feature map corresponding to each frame color image in the third feature atlas R33 and the fourth feature map corresponding to the same frame color image in the fourth feature atlas R4 into the cross-scale cross-fusion module for cross-fusion processing until all frames are processed, thereby obtaining the fifth feature atlas.

8. The method for recognizing dynamic facial expressions as claimed in claim 7, characterized in that: The cross-scale cross-fusion module includes a first multi-head attention unit and a second multi-head attention unit with the same structure; the first multi-head attention unit includes a multi-head attention mechanism layer, a splicing layer, a feature connection layer, a linear layer and a nonlinear transformation layer; The multi-head attention mechanism layer has three attention heads, and S41 specifically includes: S41-1, performing layer normalization processing on the two extracted third feature maps respectively, so as to obtain the query matrix, key matrix and value matrix corresponding to each third feature map, and then performing cross attention processing using six attention heads; S41-2, the feature maps output by the three attention heads in the first multi-head attention unit are concatenated to obtain a feature map T1; the feature maps output by the three attention heads in the second multi-head attention unit are concatenated to obtain a feature map T2; S41-3, performing multi-head self-attention processing on three first query matrices, three second key matrices, and three second value matrices through the first multi-head attention unit to obtain a feature map Z1; S41-4, performing multi-head self-attention processing on the three second query matrices, the three first key matrices, and the three first value matrices through the second multi-head attention unit to obtain a feature map Z2; S41-5, after performing feature connection processing on the feature map Z1 and the feature map Z2, perform mean fusion processing to obtain a feature map Z3.

9. A dynamic facial expression recognition system, characterized in that: The system comprises: The ResNet18 network model is used to perform convolution processing on multiple frames of continuous color images and then perform frame difference processing to obtain three first feature atlases of different scales; A self-attention ST-Former model is used to perform spatial attention weighted processing on each feature map in the three first feature map sets, and then perform temporal attention weighted processing to obtain three second feature map sets; A conditional random field fusion module is used to perform feature exchange enhancement processing on the second dynamic feature map and the second static feature map corresponding to each frame of the color image in the three second feature map sets, so as to obtain three third feature map sets; A cross-scale cross-fusion module, used to fuse the three third feature atlases in order from large to small scales to obtain a fifth feature atlas; The fully connected layer is used to process the fifth feature atlas and output the expression recognition result.

10. The dynamic facial expression recognition system as claimed in claim 9, characterized in that: The self-attention ST-Former model includes: A spatial attention unit, used for obtaining a spatial attention weight of each feature map in each of the first feature map sets through a spatial attention mechanism, and performing spatial attention weighted processing on the feature maps using the spatial attention weights to obtain three spatial weighted feature map sets; A temporal attention unit is used to obtain the temporal attention matrix of each feature map in the three spatially weighted feature map sets through the temporal attention mechanism, and use the temporal attention matrix to perform temporal attention weighted processing on the feature maps in each of the spatially weighted feature map sets to obtain the three second feature map sets.

Citation Information

Patent Citations

  • CNN (Convolutional Neural Network) and Transform improved lightweight monocular depth estimation method

    CN117710429A

  • Expression recognition method and system based on multi-scale features and space attention

    CN118298491A

  • Dynamic micro-expression recognition method based on frame weight and related device

    CN119580326A

  • Facial expression recognition method and system combined with attention mechanism

    US20230298382A1

  • Emotional evolution method and terminal for virtual avatar in educational metaverse

    US20250014470A1