CLIP-based multi-modal dynamic facial expression recognition method

Through the CLIP model combining multimodal data mining and adaptive fusion strategies, the data acquisition and model generalization problems in dynamic facial emotion recognition are solved, high-precision emotion recognition is achieved, and application scenarios are expanded.

CN120356253AInactive Publication Date: 2025-07-22HUNAN UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510814911.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as difficulty in dynamic facial emotions recognition, insufficient model generalization ability and limited feature extraction ability in dynamic facial emotions recognition, resulting in a decrease in recognition accuracy.

Method used

A multimodal dynamic facial expression recognition method based on CLIP is adopted to generate positive-negative text supervision by constructing a tag enhancement module, combining the multimodal data mining module to extract facial expressions, audio and fine-grained text features from the video, and feature fusion is used to perform feature fusion, and finally emotional classification is performed through cosine similarity calculation.

Benefits of technology

It improves the recognition accuracy of dynamic facial emotion recognition and the adaptability of models, and can identify new categories without the need for complex classifiers, expanding application scenarios such as human-computer interaction and medical and healthcare.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356253A_ABST
    Figure CN120356253A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal dynamic facial expression recognition method based on CLIP, and the method comprises the following steps: constructing a label enhancement module, generating positive-negative text supervision, and obtaining positive text features and negative text features; constructing a multi-modal data mining module, and mining different levels of feature information from the video; fusing the facial expression features, the audio features and the fine-grained text description features by using an adaptive fusion strategy to obtain fused feature representation; and performing cosine similarity calculation on the fused feature representation, the positive text features and the negative text features to obtain final emotion classification. According to the method, class label enhancement is introduced, class labels are converted into positive-negative text supervision, and label enhancement is performed through P-N descriptors, so that fuzzy categories which are originally difficult to distinguish can be distinguished; and the similarity between correct image-text pairs is maximized by utilizing a CLIP contrast learning mechanism, so that the classification and retrieval precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and particularly relates to a multi-modal dynamic facial expression recognition method based on CLIP. Background Art

[0002] As a core part of affective computing, emotion recognition plays an important role in the field of human-computer interaction. In recent years, video-based dynamic facial emotion recognition (DFER) has become an important research direction in the field of emotion recognition, attracting more and more scholars and research institutions. With the rapid development of deep learning, computer vision, and natural language processing technologies, researchers have gradually realized the limitations of single-modal (such as audio or vision) in dynamic facial emotion recognition. Therefore, multi-modal DFER has gradually become the mainstream research method, which combines multiple disciplines such as artificial intelligence, cognitive neuroscience, psychology, biology, and computer science, and the research content is rich and complex.

[0003] In the application of emotion recognition, a computer analyzes information of multiple modalities such as speech, facial expressions, body postures, electrocardiograms, electroencephalograms, electromyograms, etc. to judge a person's emotional state. The signal sources that can reflect human emotion changes can be divided into two categories: physical signals and physiological signals. Physiological signals can directly reflect individual emotion changes, such as changes in brain waves or heart rates, etc. These changes are usually not affected by an individual's subjective consciousness, so they can objectively reflect the emotional state and have the characteristic of being difficult to hide, being very objective and real. However, because currently collecting physiological signals usually requires specific instruments and equipment for measurement, these devices are often not easily available and are relatively complex to use. In addition, the popularity of wearable devices in the market is still limited, resulting in restricted data collection. Relatively speaking, the acquisition of physical signal sources is much simpler and can be easily obtained through common devices (such as mobile phone cameras or cameras), thereby reducing the research threshold and cost. This not only improves the flexibility of data collection but also makes emotion recognition technology easier to be widely applied and promoted. Therefore, using physical signals for emotion recognition has a very broad prospect.

[0004] Although some progress has been made in video-based dynamic facial emotion recognition tasks, there are still several challenges to be addressed. In previous mainstream research, some methods directly use 3D CNN to extract joint spatio-temporal features from raw videos. 3D CNN usually requires a large amount of labeled data for training to effectively capture spatio-temporal features, and obtaining high-quality video data and labeling it is a time-consuming and costly process. In the case of insufficient data, the performance of the model may be affected and it is difficult to generalize to new scenarios or populations. There are also some methods that combine 2D CNN with RNN for feature extraction and sequence modeling. The expression of facial emotions is very complex, and different emotions may show subtle differences in facial features. 2D CNN usually extracts features through convolutional operations, but its feature extraction ability is limited by the designed convolutional kernels and number of layers, and may not be able to fully capture all key emotional features. For example, certain micro-expressions or subtle facial muscle changes may not be effectively recognized and extracted, resulting in a decrease in the accuracy of the model for emotion classification. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a CLIP-based multi-modal dynamic facial expression recognition method with simple algorithm and high recognition accuracy.

[0006] The technical solution of the present invention to solve the above technical problems is: a CLIP-based multi-modal dynamic facial expression recognition method, comprising the following steps:

[0007] S1: Construct a label enhancement module, perform label enhancement through a PN descriptor to generate positive-negative text supervision, so as to obtain positive text features and negative text features;

[0008] S2: Construct a multi-modal data mining module to mine different levels of feature information from the video. The different levels of feature information include facial expression features, audio features extracted from the video, and fine-grained text description features obtained through a multi-modal large language model M-LLM;

[0009] S3: Use an adaptive fusion strategy to complete the fusion of facial expression features, audio features, and fine-grained text description features to obtain a fused feature representation;

[0010] S4: Calculate the cosine similarity between the fused feature representation and the positive text features and negative text features obtained in step S1 to obtain the final emotion classification.

[0011] In the above-mentioned CLIP-based multi-modal dynamic facial expression recognition method, in step S1, the original class labels are extended from positive and negative perspectives, so that the original class labels are converted into text supervision C, and then the text supervision C is divided into two different sets: CP and CN. CP represents the positive text supervision set, and CN represents the negative text supervision set. CP and CN are tokenized and projected into the word embedding layer to obtain the word embedding representations of CP and the word embedding representations of CN , 、 ∈ , where represents the real number field, represents the text length, represents the dimension of the text embedding; the word embedding representations are further constructed as:

[0012]

[0013]

[0014] where represent the position encodings of positive and negative tokenizations respectively, 、 represent the initial positive and negative text inputs of the text encoder, that is, the initial word embedding representations of positive and negative texts;

[0015] To further encode and , the text encoder of the pre-trained vision-language model VLM is used as the first text encoder. VLM is a model with pre-trained Transformer layers. The first text encoder is represented by { represents retaining the original weights of the trained Transformer layers in the first text encoder, represents the th frozen Transformer layer in the first text encoder, is the total number of frozen Transformer layers in the first text encoder; after , a trainable lightweight adapter is introduced. The th positive text adapter and the th negative text adapter are represented as and respectively; then, the encoded and are obtained through the following method:

[0016]

[0017]

[0018] Among them, represents the positive text features output after being processed by the th positive text adapter, represents the negative text features output after being processed by the th negative text adapter.

[0019] For the above-mentioned CLIP-based multi-modal dynamic facial expression recognition method, in step S1, all adapters adopt the basic adapter structure Adapter. Adapter performs fine-tuning of tasks in the architecture. Adapter performs dimension elevation, non-linear transformation, and dimension reduction on the input features through three components: the ascending linear transformation layer FC UP, the Gaussian activation function GELU, and the descending linear transformation layer FC DOWN. Finally, through the residual connection, the result of dimension elevation, non-linear transformation, and dimension reduction of the input features is added to the input features to obtain an enhanced feature representation; in step S1, the positive text features and negative text features are obtained through the following method:

[0020]

[0021]

[0022] Among them, is the last Token representation of , which is used to represent the semantics of the entire positive text supervision sentence, is the last Token representation of , which is used to represent the semantics of the entire negative text supervision sentence, is the projection layer.

[0023] For the above-mentioned CLIP-based multi-modal dynamic facial expression recognition method, in step S2, the process of mining facial expression features is as follows:

[0024] Perform video frame segmentation on the input video. Use OpenCV to detect whether there is a face in the video frame and extract the face from the video frame, so that each extracted video frame contains face information. Perform facial alignment and cropping on the video frames containing face information, and perform data augmentation; select 16 key frames from the video frames containing face information, that is, extract facial feature points from the processed images containing face information, arrange the image sequence in descending order of the number of feature points, and select the first 16 images as key frames. The facial feature points provide the accurate positions indicating different facial parts.

[0025] In the above CLIP-based multimodal dynamic facial expression recognition method, in step S2, the specific process of obtaining the facial expression features is as follows:

[0026] Given a video clip , , where represents the real number field, , are the height and width of the space respectively, is the time length, and the video frame sequence is regarded as an image sequence in the time dimension;

[0027] There are a total of 16 key frames, and the resolution of each frame is , so the input is represented as 16 images of . If the th video frame is represented as , then ;

[0028] For the th video frame, that is, the th image, the image is spatially segmented into fixed-size and non-overlapping image patches , represents the th image patch of the th image, ∈ , , where is the total number of image patches in the th image, , represents the size of each image patch; each is mapped to a high-dimensional space through a linear projection to obtain the corresponding patch embedding , , where is the patch embedding dimension, that is, the feature dimension after mapping for each image patch; the final picture sequence patch embedding is represented as:

[0029]

[0030] For the image sequence in the time dimension, the time information refers to the dynamic change features between frames in the video, which is a kind of dynamic dependence relationship across frames; in order to fully capture the time information, a time adapter module is introduced to model the time context in the frame sequence; the time adapter also performs dimension elevation, non-linear transformation, and dimension reduction on the input features through three components: the ascending linear transformation layer FC UP, the Gaussian activation function GELU, and the descending linear transformation layer FC DOWN, and obtains the enhanced feature representation through the residual connection;

[0031] The image encoder of the pre-trained vision-language model VLM is trained. The image encoder is represented by { Preserve the original weights of the trained Transformer layers in the image encoder, Denote the th frozen Transformer layer in the image encoder, is the total number of frozen Transformer layers in the image encoder; After , a trainable lightweight spatio-temporal adapter is introduced. Denote the th spatio-temporal adapter as ; After undergoing processing, a time feature representation containing time information is obtained; The spatial feature representation after undergoing the spatio-temporal adapter processing is derived through the following process :

[0032]

[0033]

[0034] Among them, Denote the output image features of the th frozen Transformer layer in the image encoder;

[0035] Input the result after layer normalization of into the MLP layer and the th spatio-temporal adapter operating in parallel with the MLP layer at the same time, and then add it to the spatial feature representation to obtain the feature representation output by the th frozen Transformer layer in the image encoder as follows:

[0036]

[0037] Among them, MLP is a feed-forward neural network for non-linear feature mapping, which includes a fully connected layer and an activation function; LN is a processing operation for normalizing each dimension of the input features; is the spatio-temporal scaling factor;

[0038] Finally, the facial expression feature is derived as , is the projection layer, is the The feature representation output by a frozen Transformer layer, that is, the feature representation finally output by the image encoder.

[0039] In the above CLIP-based multi-modal dynamic facial expression recognition method, in step S2, when extracting audio features from a video, first, the input video is used to perform an audio extraction operation through Ffmpeg to obtain the original audio signal; then, the original audio signal is pre-emphasized; then, the audio signal is framed. Framing is to divide the continuous audio signal into frames of a fixed length to facilitate the extraction of short-time features. If the corresponding video frame rate is FPS, when framing, the audio frame shift is set to 1 / FPS, and the window size, that is, the frame length, is set to 2 / FPS, that is, to be synchronized with the video, and then the Hamming window function is applied to reduce spectral leakage; finally, the number of audio frames is adjusted, and zero-padding or truncation operations are performed to make the number of audio frames consistent with the number of video frames.

[0040] In the above CLIP-based multi-modal dynamic facial expression recognition method, in step S2, the specific process of obtaining audio features is as follows:

[0041] Let the audio segment extracted from the video be , where is the number of audio channels, is the number of audio frames;

[0042] Determine the key frames in the audio according to the key frame picture sequence extracted from the video frames, then it can be known that the number of audio key frames corresponds to 16 frames. If the th audio key frame is represented as , then ;

[0043] Regard the audio as a continuous waveform signal in the time dimension. For , convert it to the corresponding Mel spectrogram :

[0044]

[0045] where is the number of frequency bands of the Mel spectrogram, is the time step, and the Mel spectrogram is divided into non-overlapping time-frequency blocks along the time dimension , represents the Mel spectrogram 's th time-frequency block, , = 1, 2… where represents The total number of time-frequency blocks divided = / , denotes the time step size of the time-frequency block; each time-frequency block is mapped to a high-dimensional space through a linear projection to obtain the corresponding block embedding , , where is the feature dimension after mapping for each time-frequency block, then the audio sequence block embedding is expressed as:

[0046]

[0047] Audio sequence block embedding uses the Vit-B encoder in the audio pre-training model AudioMAE as the audio encoder for encoding. The audio encoder consists of { denotes, and retains the original weights of the trained Transformer layers in the audio encoder, denotes the th frozen Transformer layer in the audio encoder, is the total number of frozen Transformer layers in the audio encoder; after , a trainable lightweight time-frequency adapter is introduced. The th time-frequency adapter is denoted as ; After going through the time-frequency adapter, a time feature representation containing time information is obtained ; The frequency feature representation after going through the time-frequency adapter is derived through the following process :

[0048]

[0049]

[0050] where denotes the output audio feature of the th frozen Transformer layer in the audio encoder;

[0051] The result after layer normalization is simultaneously input into the MLP layer and which operates in parallel with the MLP layer, and then added to to obtain the feature representation output by the th frozen Transformer layer in the audio encoder as follows:

[0052]

[0053] Among them, is the time-frequency scaling factor;

[0054] Finally, the audio feature is exported as , is the projection layer, is the feature representation output by the th frozen Transformer layer in the audio encoder, that is, the feature representation finally output by the audio encoder.

[0055] In the above CLIP-based multi-modal dynamic facial expression recognition method, in step S2, for each video segment , use Video-LLaMA2 to generate a detailed text description, and guide the model to provide details of facial dynamic changes through prompts, and then refine it. The refinement process is as follows:

[0056] First, adopt a rule-based method to remove redundant or irrelevant text information through a preset filter;

[0057] Then, use the NLTK tool to process the text to eliminate noise; all data is manually reviewed to screen out abnormal descriptions;

[0058] Next, select Long-CLIP as the fine-grained text encoder. Let the refined text description be , further tokenize and project it onto the word embedding layer to obtain the word embedding representation , and then further construct the input ;

[0059]

[0060] Among them is the positional encoding of the tokens of the refined text description; is the initial text input of the fine-grained text encoder, that is, the initial word embedding representation of the refined text description;

[0061] Use the pre-trained Long-CLIP as the second text encoder part. The second text encoder is represented by { represents, and retain the original weights of the trained Transformer layers in the second text encoder, represents the th frozen Transformer layer in the second text encoder, is the total number of frozen Transformer layers in the second text encoder; After introduce a trainable lightweight text adapter, and denote the th text adapter as ; Then, obtain :

[0062]

[0063] where, represents the refined text feature representation output after being processed by the th text adapter;

[0064] Finally, obtain the fine-grained text description feature :

[0065]

[0066] where, is the last Token representation of, which is used to represent the semantics of the entire fine-grained text description sentence; is the projection layer.

[0067] For the above CLIP-based multi-modal dynamic facial expression recognition method, the specific process of step S3 is as follows:

[0068] Given a video clip , assuming any feature representation is m, m ∈ { , , }, the similarity between m and , and the similarity between m and are defined by calculating the cosine similarity:

[0069]

[0070]

[0071] where, represents the cosine similarity between the positive text feature of the th class of text supervision and m, represents the cosine similarity between the negative text feature of the th class of text supervision and m; represents the vector norm; represents the text supervision category, the value of corresponds to the emotion category, Indicates the total number of emotion categories;

[0072] Further obtain:

[0073]

[0074] Represents the similarity between the text supervision of the

[0075] category and m; ; Indicates the maximum similarity value between the text supervision of a certain category and m; since m ∈ { , , }, so , , ; , and respectively represent the maximum similarity values between the text supervision of a certain category and , and ; Normalize , and to obtain the weight corresponding to m as:

[0076]

[0077] For different features , , , calculate the corresponding feature weights; , and respectively represent , and weights; Then, perform weighted summation on different features to obtain the overall multi-modal fusion feature of the multi-modal encoder as follows:

[0078]

[0079] The above-mentioned CLIP-based multi-modal dynamic facial expression recognition method, the specific process of step S4 is:

[0080] If a batch of video segments are given, take the th video segment as the th sample, , and the The multi-modal fusion features of one sample are , then , , denotes the cosine similarity between ; denotes the cosine similarity between ; , denotes the similarity score between and the text supervision of the th class; thus, denotes the maximum similarity score between and the text supervision of the

[0081] th class; The probability that the th sample belongs to the th class of emotion labels is:

[0082]

[0083] denotes temperature, which is used to control the smoothness of the softmax function;

[0084] When , denotes the predicted probability that the th sample belongs to the true class; where denotes the true class label of the th sample;

[0085] The total loss function is expressed as:

[0086]

[0087] where denotes the cross-entropy loss function, which measures the difference between and the predicted probability p; , , respectively denote the predicted th sample's facial expression feature , the th sample's audio feature , and the th sample's fine-grained text description feature as the input, predicting the samples belong to the true emotion category probability.

[0088] The beneficial effects of the present invention are:

[0089] 1. The present invention proposes a multimodal dynamic facial expression recognition method based on CLIP. First, a new method of expanding class labels into positive-negative text descriptions is introduced to make the originally vague and difficult to distinguish categories more discriminative; then a multimodal large language model is used to generate a detailed description of cross-frame facial changes through prompt words, thereby enhancing the expressiveness of features; then an adaptive weighted fusion method is used to perform feature fusion on facial expression features, audio features, and fine-grained text description features; finally, the CLIP model is used to perform emotion recognition by calculating the similarity between image and text feature vectors; the temporal and spatial information in the video source data can be fully mined, and the face and audio features in the video data can be effectively captured in combination with a multiple attention mechanism.

[0090] 2. The present invention introduces a class label enhancement method, which mainly converts class labels into positive-negative text supervision and enhances labels through P (Positive)-N (Negative) descriptors, so that fuzzy categories that were originally difficult to distinguish can be distinguished; in the process of combining with the CLIP model, this class label enhancement method can utilize the contrastive learning mechanism of CLIP to maximize the similarity between correct image-text pairs, thereby improving the accuracy of classification and retrieval.

[0091] 3. The present invention can identify new categories without special training, which greatly enhances the adaptability and generalization ability of the model. It only needs to calculate the similarity between features without the need for complex classifiers. It is more efficient than traditional classification methods, making dynamic facial emotion recognition technology more applicable in scenarios such as human-computer interaction, medical health, and intelligent driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 It is the overall flow chart of the present invention. DETAILED DESCRIPTION

[0093] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0094] like Figure 1 As shown, a multimodal dynamic facial expression recognition method based on CLIP includes the following steps:

[0095] S1: Construct a label enhancement module, perform label enhancement through PN descriptors, generate positive-negative text supervision, and thus obtain positive text features and negative text features.

[0096] In the step S1, the original class label is extended from positive and negative perspectives, so that the original class label is converted into text supervision C, and then the text supervision C is divided into two different sets: CP and CN. CP represents the positive text supervision set, and CN represents the negative text supervision set. CP and CN are tokenized and projected into the word embedding layer to obtain the word embedding representations of CP and the word embedding representation of CN , 、 ∈ , where represents the real number field, represents the text length, represents the dimension of the text embedding; the word embedding representation is further constructed as:

[0097]

[0098]

[0099] where, represent the positional encodings of positive and negative tokenizations respectively, 、 represent the initial positive and negative text inputs of the text encoder, that is, the initial word embedding representations of positive and negative texts;

[0100] To further encode and , the text encoder of the pre-trained vision-language model VLM is used as the first text encoder. VLM is a model with pre-trained Transformer layers. The first text encoder is represented by { represents retaining the original weights of the trained Transformer layers in the first text encoder, represents the th frozen Transformer layer in the first text encoder, is the total number of frozen Transformer layers in the first text encoder; after , a trainable lightweight adapter is introduced. The th positive text adapter and the th negative text adapter are represented as and respectively; then, the encoded and are obtained through the following method:

[0101]

[0102]

[0103] Among them, represents the positive text features output after being processed by the th positive text adapter, represents the negative text features output after being processed by the th negative text adapter.

[0104] For all adapters, the basic adapter structure Adapter is adopted. Adapter performs fine-tuning of tasks in the architecture. Adapter performs dimensionality increase, non-linear transformation, and dimensionality reduction on the input features through three components: the ascending linear transformation layer FC UP, the Gaussian activation function GELU, and the descending linear transformation layer FC DOWN. Finally, through the residual connection, the result of dimensionality increase, non-linear transformation, and dimensionality reduction of the input features is added to the input features to obtain an enhanced feature representation; in the step S1, the positive text features and negative text features are obtained through the following method:

[0105]

[0106]

[0107] Among them, is the last Token representation of , used to represent the semantics of the entire positive text supervision sentence, is the last Token representation of , used to represent the semantics of the entire negative text supervision sentence, is the projection layer.

[0108] S2: Construct a multi-modal data mining module to mine different levels of feature information from the video. The different levels of feature information include facial expression features, audio features extracted from the video, and fine-grained text description features obtained through the multi-modal large language model M-LLM.

[0109] The mining process of facial expression features is as follows:

[0110] Since the main task of the framework is to identify face information, the input video is selected for video frame segmentation. OpenCV is used to detect whether there is a face in the video frame and extract the face from the video frame, so that each extracted video frame contains face information. The video frames containing face information are subjected to facial alignment and cropping, and data augmentation (rotation, flipping, brightness adjustment) is performed to eliminate noise and background interference in the image; in order to reduce and unify the subsequent face feature dimensions, 16 key frames are selected from the video frames containing face information, that is, facial feature points are extracted from the processed images containing face information, and the image sequence is arranged in descending order of the number of feature points, and the first 16 images are selected as key frames. The facial feature points provide the accurate positions indicating different facial parts (for example, eyes, nose, etc.).

[0111] The specific process of obtaining the facial expression features is as follows:

[0112] In the step S2, the specific process of obtaining the facial expression features is as follows:

[0113] Given a video clip , where represents the real number field, , are the height and width of the space respectively, is the time length, and the video frame sequence is regarded as an image sequence in the time dimension;

[0114] There are a total of 16 key frames, and the resolution of each frame is , then the input is represented as 16 images of . If the th video frame is represented as , then ;

[0115] For the th video frame, that is, the th image, the image is spatially segmented into fixed-size and non-overlapping image patches , represents the th image patch of the th image, ∈ , where is the total number of image patches in the th image, , represents the size of each image patch; each is mapped to a high-dimensional space through a linear projection to obtain the corresponding patch embedding ,​ , where is the block embedding dimension, that is, the feature dimension after mapping each image block; the final picture sequence block embedding is expressed as:

[0116]

[0117] For the image sequence in the time dimension, the time information refers to the dynamic change characteristics between frames in the video, which is a kind of dynamic dependence relationship across frames; in order to fully capture the time information, a time adapter module is introduced to model the time context in the frame sequence; the time adapter also performs dimension elevation, non-linear transformation, and dimension reduction on the input features through three components: the rising linear transformation layer FC UP, the Gaussian activation function GELU, and the falling linear transformation layer FC DOWN, and obtains the enhanced feature representation through the residual connection;

[0118] The image encoder of the pre-trained vision-language model VLM is trained. The image encoder consists of { is represented, and the original weights of the trained Transformer layers in the image encoder are retained. represents the th frozen Transformer layer in the image encoder. is the total number of frozen Transformer layers in the image encoder; after , a trainable lightweight spatio-temporal adapter is introduced, and the th spatio-temporal adapter is represented as ; After experiencing , the time feature representation containing time information is obtained ; the spatial feature representation after being processed by the spatio-temporal adapter is derived through the following process :

[0119]

[0120]

[0121] Among them, represents the output image feature of the th frozen Transformer layer in the image encoder;

[0122] The after layer normalization processing is input into the MLP layer and the th spatio-temporal adapter that operates in parallel with the MLP layer at the same time, and then added to the spatial feature representation to obtain the The feature representation output by a frozen Transformer layer is as follows:

[0123]

[0124] Among them, MLP is a feed-forward neural network for non-linear mapping of features, including a fully connected layer and an activation function; LN is a processing operation for normalizing each dimension of the input features, which helps to improve training stability and expressive ability; is the spatio-temporal scaling factor;

[0125] Finally, the facial expression feature is exported as , is the projection layer, is the feature representation output by the th frozen Transformer layer in the image encoder, that is, the feature representation finally output by the image encoder.

[0126] When extracting audio features from a video, first perform an audio extraction operation on the input video through Ffmpeg to obtain the original audio signal; then perform pre-emphasis processing on the original audio signal; pre-emphasis processing is actually passing the speech signal through a high-pass filter: enhancing high-frequency signals and compensating for the energy loss of the high-frequency part in the speech signal. The pre-emphasis calculation formula is as follows, where the pre-emphasis coefficient is usually taken as 0.97:

[0127]

[0128] Among them, represents the time point, is the original audio value of is the audio value at the previous time point ( );

[0129] The purpose of pre-emphasis is to enhance the high-frequency part, make the spectrum of the signal flat, and ensure that the spectrum can be obtained with the same signal-to-noise ratio throughout the frequency band from low frequency to high frequency. At the same time, it is also to eliminate the effects of the vocal cords and lips during the generation process, compensate for the high-frequency part of the speech signal suppressed by the pronunciation system, and highlight the high-frequency formants. Then, the audio signal is framed. Framing is to divide the continuous audio signal into frames of a fixed length to facilitate the extraction of short-term features. The framing parameters will affect the temporal alignment between the audio features and the video frames. Therefore, the audio needs to be framed according to the frame rate of the corresponding video. If the frame rate of the corresponding video is FPS, when framing, the audio frame shift is set to 1 / FPS, and the window size, i.e., the frame length, is set to 2 / FPS, so as to be synchronized with the video. Then, the Hamming window function is applied to reduce spectral leakage. Finally, the number of audio frames is adjusted by zero-padding or truncation operations to make the number of audio frames consistent with the number of video frames. After the audio data preprocessing operation, the audio frames can be temporally synchronized with the video frames, which is conducive to extracting more effective audio features subsequently and can effectively improve the performance of model training;

[0130] The specific process of obtaining audio features is as follows:

[0131] Let the audio segment extracted from the video be , where is the number of audio channels, is the number of audio frames;

[0132] Determine the key frames in the audio according to the key frame picture sequence extracted from the video frames. It can be known that the number of audio key frames corresponds to 16 frames. If the th audio key frame is denoted as , then ;

[0133] Regard the audio as a continuous waveform signal in the time dimension. For , convert it to the corresponding Mel spectrogram :

[0134]

[0135] where is the number of frequency bands of the Mel spectrogram, is the time step. The Mel spectrogram is divided into non-overlapping time-frequency blocks along the time dimension , represents the th time-frequency block of the Mel spectrogram, , = 1, 2… where denote the total number of time-frequency blocks divided, = / , denote the time step size of the time-frequency block; each time-frequency block is mapped to a high-dimensional space through a linear projection to obtain the corresponding block embedding , , where is the feature dimension after mapping each time-frequency block, then the audio sequence block embedding is denoted as:

[0136]

[0137] Audio sequence block embedding uses the Vit-B encoder in the audio pre-training model AudioMAE as the audio encoder for encoding. The audio encoder consists of { denote, and retain the original weights of the trained Transformer layers in the audio encoder, denote the th frozen Transformer layer in the audio encoder, is the total number of frozen Transformer layers in the audio encoder; after , introduce a trainable lightweight time-frequency adapter, and denote the th time-frequency adapter as ; After going through the time-frequency adapter, obtain the time feature representation containing time information ; Derive the frequency feature representation after going through the time-frequency adapter through the following process :

[0138]

[0139]

[0140] where, denote the output audio features of the th frozen Transformer layer in the audio encoder;

[0141] Input the result after layer normalization of into the MLP layer and operating in parallel with the MLP layer at the same time, and then add it to to obtain the feature representation output by the th frozen Transformer layer in the audio encoder As follows:

[0142]

[0143] Among them, is the time-frequency scaling factor;

[0144] Finally, the audio feature is exported as , is the projection layer, is the feature representation output by the th frozen Transformer layer in the audio encoder, that is, the feature representation finally output by the audio encoder.

[0145] For each video clip , use Video-LLaMA2 to generate a detailed text description, and guide the model to provide details of facial dynamic changes through prompts. To ensure that the description is more accurate and detailed, the prompts set granularity requirements and clearly specify specific actions involving different facial regions. Specifically, the following requirements are made for the design of the prompts:

[0146] 1. In terms of format: detailed feature and behavior analysis; 2. Focus on the detailed changes and implicit emotional actions of multiple facial parts; 3. Analyze subtle facial and action clues; 4. Use manually annotated videos; 5. Do not contain emotion words; 6. At most 100 words;

[0147] However, the generated text may still contain emotion words or redundant information related to class labels. Such text is called low-quality text. Specifically, there are two types of low-quality text: 1) Directly expressing emotions. For example, the statement "The person in the video has a happy expression..." may lead to data leakage during the training process. 2) Indirectly implying emotions. For example, the statement "The person in the video has the corners of their mouth slightly downturned and their eyes squinted, suggesting a feeling of sadness or depression." Although it does not explicitly contain label information, these descriptions still pose a potential risk of data leakage. Therefore, refinement is required to obtain a concise and high-quality overview. The refinement process is as follows:

[0148] First, adopt a rule-based method to remove redundant or irrelevant text information through preset filters;

[0149] Then, use the NLTK tool to process the text to eliminate noise; all data is manually reviewed to screen out abnormal descriptions;

[0150] Next, select Long-CLIP as the fine-grained text encoder, and set the refined text description as , and further perform Perform word segmentation and project it onto the word embedding layer to obtain the word embedding representation , and then further construct the input ;

[0151]

[0152] where is the positional encoding for the refined text description word segmentation; is the initial text input of the fine-grained text encoder, i.e., the initial word embedding representation of the refined text description;

[0153] Use the pre-trained Long-CLIP as the second text encoder part. The second text encoder is represented by { indicating to retain the original weights of the trained Transformer layers in the second text encoder, indicating the th frozen Transformer layer in the second text encoder, being the total number of frozen Transformer layers in the second text encoder; After , introduce a trainable lightweight text adapter, and represent the th text adapter as ; Then, obtain through the following:

[0154]

[0155] where, represents the refined text feature representation output after being processed by the th text adapter;

[0156] Finally, obtain the fine-grained text description feature through the following:

[0157]

[0158] where, is 's last Token representation, used to represent the semantics of the entire fine-grained text description sentence; is the projection layer.

[0159] S3: Use an adaptive fusion strategy to complete the fusion of face expression features, audio features, and fine-grained text description features to obtain the fused feature representation.

[0160] The specific process of the step S3 is as follows:

[0161] Given the video clip , assuming any feature representation is m, m ∈ { , , }, m and The similarity between, and the similarity between m and are defined by calculating the cosine similarity:

[0162]

[0163]

[0164] where, represents the cosine similarity between the positive text features of the -th class of text supervision and m, and represents the cosine similarity between the negative text features of the -th class of text supervision and m; represents the vector norm; represents the text supervision category, The value of corresponds to the emotion category, ; represents the total number of emotion categories;

[0165] Furthermore, we obtain:

[0166]

[0167] represents the similarity between the -th class of text supervision and m;

[0168] Then, by finding the category with the maximum similarity to m among all categories, we obtain ; represents the value of the maximum similarity between a certain category of text supervision and m; since m ∈ { , , }, thus we get , , ; , and respectively represent the values of the maximum similarity between a certain category of text supervision and , and ; Normalize , and to obtain the weight corresponding to m as:

[0169]

[0170] For different features 、 、 , calculate the corresponding feature weights; 、 and respectively represent 、 and 's weights; then, perform weighted summation on different features to obtain the overall multimodal fusion feature of the multimodal encoder as follows:

[0171]

[0172] S4: Calculate the cosine similarity between the fused feature representation and the positive text feature and negative text feature obtained in step S1 to obtain the final sentiment classification.

[0173] The specific process of the said step S4 is:

[0174] If a batch of video segments are given, take the th video segment as the th sample, , and the multimodal fusion feature of the th sample is , then , , represents and 's cosine similarity, represents and 's cosine similarity; , represents and the th class text supervision's similarity score; thus obtaining , represents and the th class text supervision's maximum similarity score;

[0175] Then the probability that the th sample belongs to the th class of emotion labels is:

[0176]

[0177] Represents the temperature, which is used to control the smoothness of the softmax function;

[0178] When is the case, represents the predicted probability that the model assigns the -th sample to the true class; among them, represents the true class label of the -th sample;

[0179] The total loss function is expressed as:

[0180]

[0181] where represents the cross-entropy loss function, which measures the difference between and the predicted probability p; , , respectively represent the probability of predicting that the -th sample belongs to the true emotion class when using only the facial expression features of the -th sample , the audio features of the -th sample , and the fine-grained text description features of the -th sample as inputs, and the calculation process is the same as that of calculating . Similarly.

Claims

1. A multi-modal dynamic facial expression recognition method based on CLIP, characterized in that, It includes the following steps: S1: Construct a label enhancement module to perform label enhancement through PN descriptors, generate positive-negative text supervision, and thus obtain positive text features and negative text features; S2: Construct a multi-modal data mining module to mine different levels of feature information from the video. The different levels of feature information include facial expression features, audio features extracted from the video, and fine-grained text description features obtained through the multi-modal large language model M-LLM; S3: Use an adaptive fusion strategy to complete the fusion of facial expression features, audio features, and fine-grained text description features to obtain a fused feature representation; S4: Calculate the cosine similarity between the fused feature representation and the positive text features and negative text features obtained in step S1 to obtain the final emotion classification.

2. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 1, wherein In the step S1, the original class label is extended from positive and negative perspectives, so that the original class label is converted into text supervision C, and then the text supervision C is divided into two different sets: CP and CN. CP represents the positive text supervision set, and CN represents the negative text supervision set. CP and CN are tokenized and projected into the word embedding layer to obtain the word embedding representation of CP and the word embedding representation of CN , 、 ∈ , where represents the real number field, represents the text length, represents the dimension of the text embedding; the word embedding representation is further constructed as: ; ; Among them, represent the position encodings of positive and negative participles respectively, , represent the initial positive and negative text inputs of the text encoder, that is, the initial word embedding representations of positive and negative texts; For further encoding and , the text encoder of a pre-trained vision-language model VLM is used as the first text encoder, where the VLM is a model with pre-trained Transformer layers. The first text encoder is represented by { The original weights of the trained Transformer layers in the first text encoder are retained, denotes the th frozen Transformer layer in the first text encoder, is the total number of frozen Transformer layers in the first text encoder; After , a trainable lightweight adapter is introduced. The th positive text adapter and the th negative text adapter are denoted as and respectively; Then, the encoded and are obtained by the following: ; ; Among them, represents the positive text features output after being processed by the th positive text adapter, represents the negative text features output after being processed by the th negative text adapter.

3. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 2, characterized in that In the step S1, a basic adapter structure Adapter is adopted for all adapters. The Adapter performs fine-tuning of tasks in the architecture. The Adapter performs dimensionality increase, non-linear transformation, and dimensionality reduction on the input features through three components: an ascending linear transformation layer FC UP, a Gaussian activation function GELU, and a descending linear transformation layer FC DOWN. Finally, through a residual connection, the result of dimensionality increase, non-linear transformation, and dimensionality reduction of the input features is added to the input features to obtain an enhanced feature representation; in the step S1, positive text features and negative text features are obtained in the following manner: ; ; Among them, is the last Token representation, which is used to represent the semantics of the entire positive text supervision sentence, is the last Token representation, which is used to represent the semantics of the entire negative text supervision sentence, is the projection layer.

4. The method for multi-modal dynamic facial expression recognition based on CLIP according to claim 3, wherein In step S2, the process of mining facial expression features is as follows: Perform video frame splitting on the input video, use OpenCV to detect whether there is a face in the video frame and extract the face from the video frame, so that each extracted video frame contains face information. Perform facial alignment and cropping on the video frame containing face information and perform data augmentation; Select 16 key frames from the video frames containing face information, that is, extract facial feature points from the processed images containing face information, arrange the image sequence in descending order of the number of feature points, and select the first 16 images as key frames. The facial feature points provide the accurate positions indicating different facial parts.

5. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 4, wherein In step S2, the specific process of obtaining facial expression features is as follows: Given video clip , , where represents the real number field, , are the height and width of the space respectively, is the time length, and the video frame sequence is regarded as an image sequence in the time dimension; There are 16 key frames, and the resolution of each frame is , so the input is represented as 16 images. If the th video frame is represented as , then ; For the th video frame, i.e., the th image, the image is spatially segmented into fixed-size and non-overlapping image patches , denotes the th image patch of the th image, ∈ , , where is the total number of image patches in the th image, , denotes the size of each image patch; each is mapped to a high-dimensional space through a linear projection to obtain the corresponding patch embedding , , where is the patch embedding dimension, i.e., the feature dimension after mapping for each image patch; the final image sequence patch embedding is expressed as: ; For the image sequence in the time dimension, the time information refers to the dynamic change features between frames in the video, which is a kind of dynamic dependence relationship across frames; To fully capture the time information, a time adapter module is introduced to model the time context in the frame sequence; The time adapter also performs dimensionality increase, non-linear transformation, and dimensionality reduction on the input features through three components: an ascending linear transformation layer FC UP, a Gaussian activation function GELU, and a descending linear transformation layer FC DOWN, and obtains an enhanced feature representation through a residual connection; The image encoder of the pre-trained vision-language model VLM is trained. The image encoder is represented by { It is stated that the original weights of the trained Transformer layers in the image encoder are retained. It represents the th frozen Transformer layer in the image encoder. is the total number of frozen Transformer layers in the image encoder; After , a trainable lightweight spatio-temporal adapter is introduced. The th spatio-temporal adapter is represented as ; After undergoing , a time feature representation containing time information is obtained; Derive the spatial feature representation after being processed by the spatio-temporal adapter through the following process : ; ; Among them, represents the output image features of the th frozen Transformer layer in the image encoder; The result after layer normalization is simultaneously input into the MLP layer and the th spatio-temporal adapter that operates in parallel with the MLP layer, and then added to the spatial feature representation to obtain the feature representation output by the th frozen Transformer layer in the image encoder, as follows: the th frozen Transformer layer in the image encoder as follows: ; Among them, the MLP is a feedforward neural network for non-linear feature mapping, including a fully connected layer and an activation function; LN is a processing operation for normalizing each dimension of the input features; is the spatio-temporal scaling factor; Finally, the facial expression features are exported as , which is the projection layer, and is the feature representation output by the th frozen Transformer layer in the image encoder, that is, the feature representation finally output by the image encoder.

6. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 5, characterized in that, In step S2, when extracting audio features from the video, first perform an audio extraction operation on the input video through Ffmpeg to obtain the original audio signal; Then perform pre-emphasis processing on the original audio signal; Next, perform frame splitting on the audio signal. Frame splitting is to divide the continuous audio signal into frames of a fixed length to facilitate the extraction of short-time features. If the corresponding video frame rate is FPS, when performing frame splitting, the audio frame shift is set to 1 / FPS, and the window size, that is, the frame length, is set to 2 / FPS, that is, to be synchronized with the video, and then apply the Hamming window function to reduce spectral leakage; Finally, adjust the number of audio frames and perform zero-padding or truncation operations to make the number of audio frames consistent with the number of video frames.

7. The method for multi-modal dynamic facial expression recognition based on CLIP according to claim 6, wherein In step S2, the specific process of obtaining audio features is as follows: Let the audio segment extracted from the video be , , where is the number of audio channels, is the number of audio frames; Determine the key frames in the audio based on the sequence of key frame pictures extracted from the video frames. It can be known that the number of audio key frames corresponds to 16 frames. If the th audio key frame is represented as , then ; Regarding the audio as a continuous waveform signal in the time dimension, for , convert it into the corresponding Mel spectrogram : ; Among them is the number of frequency bands of the mel spectrogram, is the time step, which divides the mel spectrogram along the time dimension into non-overlapping time-frequency blocks , represents the mel spectrogram the th time-frequency block, , = 1, 2… , where represents the total number of time-frequency blocks divided, = / , represents the time step size of the time-frequency block; each time-frequency block is mapped to a high-dimensional space through a linear projection to obtain the corresponding block embedding , , where is the feature dimension after mapping of each time-frequency block, then the audio sequence block embedding is expressed as: ; Audio sequence block embedding Use the Vit-B encoder in the audio pre-trained model AudioMAE as the audio encoder for encoding. The audio encoder consists of { denoted as, and retain the original weights of the trained Transformer layers in the audio encoder. denotes the th frozen Transformer layer in the audio encoder. is the total number of frozen Transformer layers in the audio encoder; After introduce a trainable lightweight time-frequency adapter, and denote the th time-frequency adapter as ; After passing through the time-frequency adapter, obtain the time feature representation containing time information ; The frequency feature representation after passing through the time-frequency adapter processing is derived through the following process : ; ; Among them, represents the output audio features of the th frozen Transformer layer in the audio encoder; The result after layer normalization is simultaneously input into the MLP layer and into the which operates in parallel with the MLP layer, and then added to to obtain the feature representation output by the th frozen Transformer layer in the audio encoder as follows: ; wherein, is the time-frequency scaling factor; Finally, the audio features are exported as , which is the projection layer and is the feature representation output by the th frozen Transformer layer in the audio encoder, that is, the feature representation finally output by the audio encoder.

8. The method for multi-modal dynamic facial expression recognition based on CLIP according to claim 7, wherein, In the step S2, for each video clip , use Video-LLaMA2 to generate a detailed text description, and guide the model to provide details of facial dynamic changes through prompts, and then refine it. The refining process is as follows: First, adopt a rule-based method to remove redundant or irrelevant text information through a preset filter; Then, use the NLTK tool to process the text to eliminate noise; All data has been manually reviewed to screen out abnormal descriptions; Next, Long-CLIP is selected as the fine-grained text encoder, and the refined text description is set as , and further is tokenized and projected onto the word embedding layer to obtain the word embedding representation , and then the input is further constructed as follows: ; Among them is the position encoding for tokenizing the refined text description; is the initial text input of the fine-grained text encoder, that is, the initial word embedding representation of the refined text description; Use the pre-trained Long-CLIP as the second text encoder part, and the second text encoder is represented by { indicating that the original weights of the trained Transformer layers in the second text encoder are retained, indicating the th frozen Transformer layer in the second text encoder, being the total number of frozen Transformer layers in the second text encoder; after introduce a trainable lightweight text adapter, and represent the th text adapter as ; then, obtain through the following: ; Among them, represents the refined text feature representation output after being processed by the th text adapter; Finally, fine-grained text description features are obtained in the following manner : ; Among them, is the last Token representation, used to represent the semantics of the entire fine-grained text description sentence; is the projection layer.

9. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 8, characterized in that, The specific process of step S3 is as follows: Given video clip , assuming any feature is represented as m, m ∈ { , , }, the similarity between m and , and the similarity between m and are defined by calculating the cosine similarity: ; ; Among them, represents the cosine similarity between the positive text features of the th type of text supervision and m, represents the cosine similarity between the negative text features of the th type of text supervision and m; represents the vector norm; represents the text supervision category, whose value corresponds to the emotion category, ; represents the total number of emotion categories; Further obtain: ; Represent the similarity between class text supervision and m; Then, by finding the category with the maximum similarity to m among all categories, ; represents the maximum similarity value between a certain category text supervision and m; since m ∈ { , , }, thus obtaining , , ; , and respectively represent the maximum similarity values between a certain category text supervision and , and ; normalizing , and gives the weight corresponding to m as: ; For different features , , , the corresponding feature weights are calculated; , and respectively represent , and weights; then, the different features are weighted and summed to obtain the overall multimodal fusion feature of the multimodal encoder as follows: 。 10. The multi-modal dynamic facial expression recognition method based on CLIP according to claim 9, wherein, The specific process of step S4 is as follows: If a given batch video clips, take the th video clip as the th sample, , and the multi-modal fusion feature of the th sample is , then , , denotes and the cosine similarity between them, denotes and the cosine similarity between them; , denotes and the th class text supervision similarity score; thus obtaining , denotes and the th class text supervision maximum similarity score; Then the th sample belongs to the th emotion label with a probability of being: ; Represents the temperature, which is used to control the smoothness of the softmax function; When it is the predicted probability that the model assigns the -th sample to the true class; where represents the true class label of the -th sample. Total loss function is expressed as: ; Among them represents the cross-entropy loss function, measuring the difference from the predicted probability p; , , respectively represent the probabilities of predicting that the th sample belongs to the true emotion category when using only the facial expression features of the th sample, the audio features of the th sample, and the fine-grained text description features of the th sample as inputs. ​​​

Citation Information

Cited By

  • Electroencephalogram emotion estimation method and device based on time-frequency domain characteristics and language prompts

    CN120983051A

  • A brainwave emotion estimation device based on time-frequency domain features and language cues

    CN120983051B

  • Image tag identification method based on multi-modal feature fusion

    CN121458992A