Facial action unit recognition method based on image-text pre-training model
By combining image and text pre-training models with joint training of visual and text features, the problems of insufficient semantic understanding and neglect of spatiotemporal information in facial action unit recognition are solved, thereby improving recognition accuracy.
Patent Information
- Application Number
- CN202511184866.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies for facial action unit recognition suffer from problems such as insufficient semantic understanding, large transfer gaps, and neglect of spatiotemporal information, resulting in low recognition accuracy.
A facial action unit recognition method based on an image-text pre-trained model is adopted. By combining visual and text features through a multimodal encoder, a temporal encoder and a multimodal AU recognition module, image-text alignment loss, video-text alignment loss and multi-label classification loss are jointly trained to capture semantic and spatiotemporal information.
It enhances semantic understanding and feature alignment, captures spatiotemporal dynamic information, and improves the accuracy of facial action unit recognition.
Smart Images

Figure CN121095993A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of deep learning and visual emotion recognition, and in particular to a facial action unit recognition method based on a graph-text pre-training model. BACKGROUND
[0002] In recent years, artificial intelligence research has also paid more and more attention to the concept of "people-oriented, serving people". Facial expression, as the most direct manifestation of human emotion, has great research value. However, in daily life, people often convey emotional states through subtle local changes rather than large facial movements. Among them, facial action units (Action Units, AU) are basic facial movements in local facial regions defined by the Facial Action Coding System (FACS), which describe the fine-grained changes of human facial expressions. Each AU involves one or more local muscles and represents a local facial movement. Each AU has six intensity levels, namely 0, 1, 2, 3, 4 and 5, where 0 means not appearing and 5 means the strongest intensity. It is because facial action units can objectively and quantitatively describe the subtle expressions of the face that the application of facial action units in recent years has important value and broad prospects in the fields of healthcare, depression screening, games and animated movies, etc.
[0003] The action unit recognition task usually includes two sub-tasks of AU recognition and AU intensity estimation. AU recognition, as a recognition of micro-expression actions in local facial regions, mainly faces three challenges: label scarcity due to labeling difficulty, label imbalance and difficulty in feature extraction. In the past few years, facial action unit recognition research has used auxiliary information (such as facial landmarks, the relationship between AU and emotion, the spatial relationship of AU, etc.) to improve the performance of AU recognition. There are also transfer learning methods that use AU dataset fine-tuning on pre-trained models on other datasets. The document "Libreface: An open-source toolkit for deep facial expression analysis, CoRR, 2023" discloses a facial action unit recognition method based on pre-training and fine-tuning. This method combines large-scale pre-trained networks, feature knowledge distillation and task-specific fine-tuning, etc., effectively and accurately analyzes facial expressions using visual information, thereby facilitating the realization of real-time interactive applications.
[0004] There is a significant difference between the pre-training data features used by the widely used large-scale pre-training model and the facial action unit data features, which leads to a large migration gap between the two. Secondly, this method only uses visual information, however, facial action units are not just visual representations, they also contain rich semantic information when expressing human emotions. Some researchers have also paid attention to the semantic information of AU, using semantic embedding as attention weight to fuse visual features, but there is no method that directly uses AU semantics as supervision information. In addition, AU datasets are usually presented in video form, and the method of using static models to predict each frame ignores the rich temporal dynamic information contained in the video, which may limit the prediction ability of the model. SUMMARY
[0005] The purpose of the present application is to provide a facial action unit recognition method based on a graph-text pre-training model, which solves the problems of the prior art in semantic understanding, capturing spatio-temporal information, recognition accuracy, etc.
[0006] In order to achieve the above task, the present application adopts the following technical solutions: The facial action unit recognition method based on a graph-text pre-training model comprises: A video segment to be recognized is obtained and sampled to obtain an image sequence, which is input into a trained facial action unit recognition model to obtain a facial action unit recognition result; wherein the facial action unit recognition model comprises a multi-modal encoder, a temporal encoder and a multi-modal AU recognition module; The multi-modal encoder comprises a visual encoder and a text encoder; the visual encoder is used for image feature extraction on the image sequence obtained by sampling the video segment, to obtain an image feature sequence; the text encoder is used to generate a global text feature based on the global dynamic semantic cues of the video segment for training to calculate a global contrast loss, and is also used to generate an initial text feature sequence based on the frame-by-frame semantic cues of the video segment; The temporal encoder comprises a temporal image encoder and a temporal text encoder; the temporal image encoder is used for position encoding and modeling on the image feature sequence to obtain a final visual feature sequence; the temporal text encoder is used for temporal modeling on the initial text feature sequence to obtain a final text feature sequence; The multi-modal AU recognition module is used for nonlinear transformation and fusion of the final visual feature sequence and the final text feature sequence, and then the fused feature vector is classified by a classifier to obtain the recognition result of each facial action unit class in each frame of the image sequence.
[0007] Further, the temporal image encoder receives the image feature sequence and adds position encoding After that, the time interaction of adjacent frames is modeled through a self-attention mechanism, and a final visual feature sequence that integrates temporal context information is output :
[0008] wherein, is the final visual feature sequence, is the image feature sequence, is the additive position encoding, denotes a real matrix with dimensions , is the sequence length, is the feature dimension, is the real space, denotes the processing procedure of the temporal image encoder.
[0009] Further, the temporal text encoder performs temporal modeling on the initial text feature sequence to obtain the final text feature sequence :
[0010] wherein, is the final text feature sequence, is the initial text feature sequence, is the additive position encoding, denotes the processing procedure of the temporal text encoder.
[0011] Further, the procedure of performing nonlinear transformation on the final visual feature sequence and the final text feature sequence is as follows: An intermediate feature transformation layer is introduced to perform nonlinear transformation, and is denoted as follows:
[0012] wherein, is the input final visual feature sequence or final text feature sequence, is the transformed feature vector; and are the learnable weight matrices of two linear layers, and ReLU is the rectified linear unit activation function; is the combination of double-layer linear transformation and activation function, is a hyperparameter that controls the strength of residual connection.
[0013] Further, in the training of the facial action unit recognition model, the video clips in the training set include facial expression videos and corresponding AU labels. For each frame of an image in a video clip, a descriptive text is constructed based on its corresponding AU tag, denoted as a frame-by-frame semantic prompt; For the changes in AU over time in the entire video clip, a global text describing this dynamic change is constructed, denoted as the global dynamic semantic cue.
[0014] Furthermore, the facial action unit recognition model employs a joint loss function during training, which consists of a weighted average of three parts: image-text alignment loss, video-text alignment loss, and multi-label classification loss; where: The image-text alignment loss is calculated by taking the cosine similarity between each final visual feature in the final visual feature sequence and each final text feature in the final text feature sequence; then, the difference between the predicted matching probability distribution and the true matching probability distribution is minimized by using KL divergence. The video-text alignment loss is achieved by performing average pooling along the time dimension on the final visual feature sequence to obtain a global video representation; then, the contrast loss between the global video representation and the global text features is calculated. The multi-label classification loss is calculated by using the binary cross-entropy loss function to calculate the multi-label classification loss during the classification process of the multi-label classifier.
[0015] Furthermore, when the facial action unit recognition model is applied: A video segment to be identified is acquired, and the video segment is sampled to obtain a multi-frame image sequence. The image sequence is input into a trained facial action unit recognition model, which outputs the predicted probability vector of each AU category in each frame. The predicted probability of each AU category is compared with a pre-set threshold to make a final decision and obtain the AU recognition result.
[0016] A terminal device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, it implements the facial action unit recognition method based on a pre-trained image and text model.
[0017] A computer-readable storage medium storing a computer program; when executed by a processor, the computer program implements the facial action unit recognition method based on a pre-trained image and text model.
[0018] Compared with the prior art, the present invention has the following technical features: 1. Enhanced Semantic Understanding and Feature Alignment: Rich semantic information is introduced into model training through carefully designed frame-by-frame AU semantic cues and global dynamic change descriptions. By utilizing joint contrastive learning of image-text and video-text, the alignment of visual features and text features in the embedding space is effectively guided, strengthening the semantic consistency of feature representations and reducing the transfer gap of large-scale pre-trained models on AU recognition tasks.
[0019] 2. Capturing Spatiotemporal Dynamic Information: The model constructed in this scheme not only utilizes the text-image pre-trained model (CLIP) to extract powerful static frame features, but also captures the spatiotemporal evolution patterns of AUs in the video through a temporal encoder. This comprehensive utilization of dynamic information enables the model to more accurately identify subtle facial expressions and movements that depend on time changes.
[0020] 3. Improved recognition accuracy: In the multimodal classification module, this solution not only considers image information, but also further integrates text features for joint classification, which enhances the feature representation of AU and thus improves the accuracy of recognition. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the network structure of the facial action unit recognition model constructed in this invention; Figure 2 Generate example graphs for AU semantic prompts; where blue represents facial regions / locations, red represents actions, and green represents the intensity of actions. Detailed Implementation
[0022] This invention provides a facial action unit (AMU) recognition method based on an image-text pre-trained model. Based on this model, the invention designs frame-by-frame AMU semantic description and global AMU dynamic change description, introduces a Transformer temporal encoder, and proposes a facial action unit recognition model that jointly integrates image text alignment and video text alignment. The method includes: The video segment to be identified is acquired and sampled to obtain an image sequence, which is then input into a trained facial action unit recognition model to obtain the facial action unit recognition result; wherein, the facial action unit recognition model includes a multimodal encoder, a temporal encoder and a multimodal AU recognition module.
[0023] 1. Multimodal encoder.
[0024] The multimodal encoder includes a visual encoder and a text encoder. The visual encoder is used to extract image features from the image sequence after sampling and processing video clips to obtain an image feature sequence. The text encoder generates global text features based on global dynamic semantic cues of video clips for calculating global contrast loss during training, and generates an initial text feature sequence based on frame-by-frame semantic cues of video clips.
[0025] (1) Visual encoder .
[0026] Visual encoder The image encoder in the pre-trained CLIP (Contrastive Language-Image Pre-Training) model is used; after sampling, video segments are processed to obtain an image sequence, and each frame in the image sequence is processed. As a visual encoder The input is used to extract its high-dimensional visual features; therefore, for a given input containing... Image sequence of frames Its output is the initial frame-level image feature sequence. ;in ,in This represents the processing procedure of the visual encoder.
[0027] (2) Text Encoder .
[0028] Text Encoder A text encoder from a pre-trained CLIP model is used; this encoder performs two tasks: first, receiving global dynamic semantic cues from video segments. Output global text features Secondly, receiving Frame-by-frame semantic hints for each frame in a frame image sequence Output the initial text feature sequence at the frame level. ;in This describes the processing procedure of the text encoder.
[0029] 2. Timing encoder.
[0030] Temporal encoders include temporal image encoders and temporal text encoders; temporal image encoders perform positional encoding and modeling on image feature sequences to obtain the final visual feature sequence; temporal text encoders are used to perform temporal modeling on initial text feature sequences to obtain the final text feature sequence.
[0031] In one embodiment of the present invention, the TransformerEncoder structure is used as a temporal image encoder and a temporal text encoder.
[0032] (1) Timing image encoder .
[0033] Timing Image Encoder Received image feature sequence And add position encoding Subsequently, a self-attention mechanism is used to model the temporal interaction between adjacent frames, outputting a final visual feature sequence that incorporates temporal context information. The calculation process can be expressed as follows:
[0034] in, For the final visual feature sequence, For image feature sequences, For additive positional encoding, The dimension is A real matrix, For sequence length (number of frames), For feature dimension, For the real number space, This represents the processing procedure of a time-series image encoder.
[0035] (2) Temporal text encoder .
[0036] Timing text encoder For the initial text feature sequence Temporal modeling is performed to obtain the final text feature sequence. The calculation process is as follows:
[0037] in, This is the final text feature sequence. This is the initial text feature sequence. This is an additive positional encoding.
[0038] 3. Multimodal AU (Facial Action Unit) recognition module.
[0039] The multimodal AU recognition module is used to perform nonlinear transformation and fusion on the final visual feature sequence and the final text feature sequence, and then use a classifier to classify the fused feature vector to obtain the recognition results of each facial action unit category in each frame of the image sequence.
[0040] This module aims to fuse multimodal features and make a final classification decision, specifically including: (1) Intermediate feature transformation layer.
[0041] To better adapt to the AU recognition task, this invention introduces an intermediate feature transformation layer to perform a nonlinear transformation on the input final visual feature sequence and final text feature sequence; the calculation formula is as follows:
[0042] in, It is the final visual feature sequence or the final text feature sequence input. It is the transformed feature vector; and It is a learnable weight matrix of two linear layers, and ReLU is the modified linear unit activation function; It is a combination of two-level linear transformation and activation function. It is a hyperparameter that controls the strength of residual connections.
[0043] (2) Multi-label classifier.
[0044] The transformed feature vectors corresponding to the final visual feature sequence and the final text feature sequence are fused (e.g., by concatenation or addition) and then input into a multi-label classifier (e.g., a fully connected layer followed by a Sigmoid activation function). The output is the predicted probability of each AU category in each frame of the image sequence, and the recognition result is obtained after making a decision.
[0045] 4. Dataset construction.
[0046] (1) Data acquisition and sample definition.
[0047] Obtain a public dataset containing videos of facial expressions and corresponding AU annotations, such as DISFA or BP4D. Each sample in the dataset is a video clip, which is accompanied by annotation information on the AU category and its intensity appearing in each video frame.
[0048] (2) Semantic prompt generation.
[0049] To incorporate rich semantic information to guide model learning, this invention generates two types of text prompts for each video segment of a sample.
[0050] Frame-by-frame semantic cues: For each frame of an image in a video clip, a descriptive text is constructed based on its corresponding AU tag, denoted as a frame-by-frame semantic cues. The sentence template used in this embodiment is "Aphotoofaface,[AUdescription]", where the [AUdescription] part is a natural language description concatenated based on all AUs appearing in the frame and their intensity information; for example... Figure 2As shown, if a frame contains AU1 (inner brow raised, intensity is very high), AU6 (cheek raised, intensity is significant), and AU12 (lip corner stretched, intensity is slight), the spliced description is "The innerbrow was raised extremely. The cheek was raised markedly. The lip corner was stretched slightly."
[0051] Global Dynamic Semantic Cue: For the changes in AU (Active Aspect) over time throughout the entire video clip, a global text describing these dynamic changes is constructed, denoted as the Global Dynamic Semantic Cue. For example, if the AU1 label changes from 0 to 1 in a video clip, representing the occurrence of the action of "inner brow raiser", then the following dynamic description text can be generated: "There was a merit of activity in AU1, indicative of inner brow raiser."
[0052] (3) Sampling processing.
[0053] For each video segment, a sparse sampling rule is applied to obtain samples. Each video frame constitutes an image sequence. and the corresponding frame-by-frame semantic cue sequence .
[0054] 5. Training of the facial action unit recognition model.
[0055] During training, a joint loss function is designed to perform end-to-end optimization training on the entire facial action unit recognition model; the joint loss function consists of three weighted parts: image-text alignment loss, video-text alignment loss, and multi-label classification loss.
[0056] (1) Image-text alignment loss (local contrast loss) .
[0057] This loss is used to align the final visual features of each frame in the image sequence after the video clip has been sampled with its corresponding final text features.
[0058] First, calculate the final visual feature sequence. The first in One final visual feature With the final text feature sequence The first in One final text feature Cosine similarity between Then, the difference between the predicted matching probability distribution and the true matching probability distribution (1 for positive sample pairs and 0 for negative sample pairs) is minimized using the KL divergence (Kullback-Leibler divergence); its calculation formula is as follows:
[0059] in, This indicates the calculation of KL divergence. and Based on cosine similarity The predicted matching probability distributions from image to final text features and from final text features to image are calculated using the softmax function. and It is the true matching probability distribution.
[0060] (2) Video-text alignment loss (global contrast loss) .
[0061] This loss is used to align the features of the global visual representation of the entire video segment with its global dynamic semantic description.
[0062] First, the final visual feature sequence Mean pooling is performed along the time dimension to obtain the global video representation. Then, the global video representation is calculated using InfoNCE loss. With global text features Comparative loss between:
[0063] in, It is a natural exponential function. It is the global video representation of the current sample. It is a global text feature; the summation term in the denominator iterates through all the data in a batch. One sample, It is the first in the batch Global video representation of each sample, It is the first in the batch Global text features of each sample It is a temperature hyperparameter.
[0064] (3) Multi-label classification loss .
[0065] During the classification process of the multi-label classifier, the binary cross-entropy loss function with Logits (BCEWithLogitsLoss) is used to calculate the multi-label classification loss of the classification task.
[0066] (4) Joint loss function and optimization.
[0067] The final joint loss function is the weighted sum of the three: ,in These are the weighting coefficients for each type of loss.
[0068] Using optimization algorithms such as Adam, based on The training is completed by updating all trainable parameters in the network model through backpropagation until the model converges.
[0069] 6. Application reasoning of facial action unit recognition model.
[0070] Once the facial action unit recognition model is trained, its network structure and parameters are fixed, and it can be used to perform real-time facial action unit recognition on new, unlabeled video clips.
[0071] (1) Obtain a video segment to be identified, and perform the same sampling process on the video segment as in the training phase to obtain... Frame image sequence.
[0072] (2) Input the image sequence into the trained facial action unit recognition model. The model then passes the image sequence through a visual encoder. Timing image encoder The final visual feature sequence is obtained. Then, it is processed by a multimodal AU recognition module to output the predicted probability vector of each AU category in each frame of the image. ,in This represents the total number of categories in AU.
[0073] (3) Predicted probability for each AU category , with a pre-set threshold The comparisons are made to reach a final decision and obtain the AU identification result; the decision rules are as follows:
[0074] in, Indicates the judgment of the first An AU exists. This indicates that it does not exist. Finally, the AU recognition result corresponding to each frame of the video is output.
[0075] In summary, this invention designs semantic cues for 15 common Active Anomalies (AUs) based on the definitions and descriptions of AUs in the FACS manual. These cues include information such as facial location, facial movement type, and movement intensity, providing effective semantic information support for the model. The facial image and the text information of the AU semantic cues are input into the image encoder and text encoder of the CLIP model for feature extraction. Then, the CLIP language-image contrastive learning method is used to further enhance the alignment and similarity of features from different modalities in the embedding space.
[0076] To address changes in Active Characters (AUs) within the video, text cues describing AU dynamics were defined, and global text encoding features were obtained using CLIP's text encoder. Simultaneously, using the pre-trained CLIP model's image and text encoder, features were extracted from a set of consecutive image frames and their corresponding AU semantic text sequences. Then, a Transformer temporal encoder was used to obtain the image and text encoding features for each frame.
[0077] A complete multimodal facial action unit (AU) recognition model is proposed. During the training phase, three loss functions are considered: 1) local contrast loss for aligning the image and semantic text of each frame of the video; 2) global contrast loss for aligning the video representation obtained by averaging the image features of each frame in the video sequence with the dynamic semantic features of AU; and 3) classification loss for AU recognition by using a multi-label classifier to fuse the image and text encoding features of each frame.
[0078] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A facial action unit recognition method based on an image-text pre-trained model, characterized in that, include: The video segment to be identified is acquired and sampled to obtain an image sequence, which is then input into a trained facial action unit recognition model to obtain the facial action unit recognition result; wherein, the facial action unit recognition model includes a multimodal encoder, a temporal encoder, and a multimodal AU recognition module; The multimodal encoder includes a visual encoder and a text encoder. The visual encoder is used to extract image features from the image sequence after sampling and processing video clips to obtain an image feature sequence. The text encoder generates global text features based on global dynamic semantic cues of video clips for calculating global contrast loss during training, and generates an initial text feature sequence based on frame-by-frame semantic cues of video clips. Temporal encoders include temporal image encoders and temporal text encoders; temporal image encoders perform positional encoding and modeling on image feature sequences to obtain the final visual feature sequence; temporal text encoders are used to perform temporal modeling on initial text feature sequences to obtain the final text feature sequence. The multimodal AU recognition module is used to perform nonlinear transformation and fusion on the final visual feature sequence and the final text feature sequence, and then use a classifier to classify the fused feature vector to obtain the recognition results of each facial action unit category in each frame of the image sequence.
2. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, Timing Image Encoder Received image feature sequence And add position encoding Subsequently, a self-attention mechanism is used to model the temporal interaction between adjacent frames, outputting a final visual feature sequence that incorporates temporal context information. : in, For the final visual feature sequence, For image feature sequences, For additive positional encoding, The dimension is A real matrix, For sequence length, For feature dimension, For the real number space, This represents the processing procedure of a time-series image encoder.
3. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, Timing text encoder For the initial text feature sequence Temporal modeling is performed to obtain the final text feature sequence. : in, This is the final text feature sequence. This is the initial text feature sequence. For additive positional encoding, This describes the processing procedure of a time-series text encoder.
4. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, The process of performing nonlinear transformation on the final visual feature sequence and the final text feature sequence is as follows: An intermediate feature transformation layer is introduced for nonlinear transformation, as shown below: in, It is the final visual feature sequence or the final text feature sequence input. It is the transformed feature vector; and It is a learnable weight matrix of two linear layers, and ReLU is the modified linear unit activation function; It is a combination of two-level linear transformation and activation function. It is a hyperparameter that controls the strength of residual connections.
5. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, During training, the video clips in the training set of the facial action unit recognition model include facial expression videos and corresponding AU annotations; For each frame of an image in a video clip, a descriptive text is constructed based on its corresponding AU tag, denoted as a frame-by-frame semantic prompt; For the changes in AU over time in the entire video clip, a global text describing this dynamic change is constructed, denoted as the global dynamic semantic cue.
6. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, The facial action unit recognition model employs a joint loss function during training, which consists of a weighted average of three parts: image-text alignment loss, video-text alignment loss, and multi-label classification loss; where: The image-text alignment loss is calculated by taking the cosine similarity between each final visual feature in the final visual feature sequence and each final text feature in the final text feature sequence; then, the difference between the predicted matching probability distribution and the true matching probability distribution is minimized by using KL divergence. The video-text alignment loss is achieved by performing average pooling along the time dimension on the final visual feature sequence to obtain a global video representation; then, the contrast loss between the global video representation and the global text features is calculated. The multi-label classification loss is calculated by using the binary cross-entropy loss function to calculate the multi-label classification loss during the classification process of the multi-label classifier.
7. The facial action unit recognition method based on an image-text pre-trained model according to claim 1, characterized in that, When the facial action unit recognition model is applied: A video segment to be identified is acquired, and the video segment is sampled to obtain a multi-frame image sequence. The image sequence is input into a trained facial action unit recognition model, which outputs the predicted probability vector of each AU category in each frame. The predicted probability of each AU category is compared with a pre-set threshold to make a final decision and obtain the AU recognition result.
8. A terminal device, comprising a processor, a memory, and a computer program stored in the memory; characterized in that, When the processor executes the computer program, it implements the facial action unit recognition method based on the image and text pre-trained model according to any one of claims 1-7.
9. A computer-readable storage medium storing a computer program; characterized in that, When the computer program is executed by the processor, it implements the facial action unit recognition method based on the image-text pre-trained model according to any one of claims 1-7.