A Cross-Domain Facial Expression Action Unit Detection Method Based on Text Bridging
The text-bridged cross-domain AU detection method aligns visual and textual features using cross-modal attention and negative sample contrastive learning, addressing data scarcity and performance degradation issues to enhance AU detection accuracy in educational settings.
Patent Information
- Application Number
- CN202510634229.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When the existing technology faces cross-domain application of AU detection models in primary and secondary education scenarios, there are problems such as scarcity of data, high labeling costs and insufficient generalization capabilities of the model, resulting in a decrease in detection accuracy.
Using a cross-domain expression motion unit detection method based on text bridging, a multimodal AU data set is constructed, combined with a visual encoder and a text encoder, a cross-modal attention and graph neural network are used for feature interaction and fusion, and a comparative learning of negative sample AU description is introduced to narrow the difference in the feature distribution of the source domain and the target domain.
It significantly improves the generalization performance and detection accuracy of the AU detection model in the target domain, and improves the robustness and accuracy of cross-domain applications.
Smart Images

Figure CN120145166B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of facial expression action unit detection, and mainly relates to a cross-domain facial expression action unit detection method based on text bridging. Background Art
[0002] Facial expressions are the intuitive external expressions of emotions and can directly convey a person's emotional state. Early facial expression analysis methods only classified expressions into basic emotions such as happiness, sadness, and surprise, and could not recognize complex and subtle expression changes. To comprehensively and objectively analyze facial expressions, psychologist Ekman proposed the Facial Action Coding System (FACS) from the perspective of human face anatomy. The FACS system defines 44 facial expression action units (ActionUnit, AU), and each AU corresponds to a movement of a local muscle of the human face. For example, AU1 (Inner Brow Raiser) represents the movement of raising the inner part of the eyebrows. Facial expressions can be represented as a combination of a series of AUs. For example, a happy expression is a combination of AU6 (Cheek Raiser) and AU12 (Lip Corner Puller). The goal of AU detection is to build a detection system that can automatically recognize facial AUs in images or videos. With the rapid development of deep learning technology, AU detection technology has been further developed and shown great application potential in fields such as human-computer interaction, medical health, and auxiliary education.
[0003] In the context of primary and secondary school education, the adoption of AU detection technology helps teachers promptly understand the learning status and emotional changes of students, enabling them to more precisely adjust teaching plans and improve the overall teaching quality. However, two core challenges limit the application of AU detection technology in educational scenarios:
[0004] First, the labeled AU dataset for primary and secondary school students is extremely scarce, and AU annotation is time-consuming and costly.
[0005] Second, when current AU detection models are applied to cross-domain scenarios such as smart education, the data in these scenarios is significantly different from the training data used by the original AU models, and the AU detection models show a significant decline in performance.
[0006] To address these challenges, it is urgent to study and implement a robust and efficient cross-domain AU detection method that uses the labeled AU data in the source domain (public dataset) and the unlabeled data in the target domain (educational scenario) to greatly reduce the difference in feature distributions between the source domain and the target domain, so as to improve the detection accuracy of the AU model in the smart education scenario.
[0007] The publication number is CN117765596A, and the name is an invention patent application for a method for establishing a facial action unit detection model based on multi-task learning. Its main technical means is: combining facial AU detection, key point detection, and emotion recognition for multi-task joint learning, and using a graph attention network to model the dependencies between AUs to improve the accuracy of AU detection. The publication number is CN117576765A, and the name is an invention patent application for a method for constructing a facial action unit detection model based on hierarchical feature alignment. Its main technical means is: performing consistency alignment within AU categories, between AU categories, and at the sample level to improve the accuracy of AU detection. Although these methods have achieved performance improvements in the AU detection task, the generalization ability of the model is insufficient, making it difficult to effectively address the challenges of cross-domain AU detection. Summary of the Invention
[0008] The present invention provides a cross-domain expression motion unit detection method based on text bridging to effectively improve the generalization ability and robustness of the model and address the challenges of cross-domain AU detection.
[0009] To achieve the above objective, the solution of the present invention is as follows:
[0010] A cross-domain expression motion unit detection method based on text bridging, comprising:
[0011] Step 1, obtaining a multi-modal expression motion unit AU dataset including an image dataset and a text dataset;
[0012] Step 2, constructing a cross-domain AU detection network;
[0013] Step 3, dividing the image dataset into a source domain dataset and a target domain dataset, and training and optimizing the cross-domain AU detection network in combination with the text dataset to obtain a cross-domain AU detection model;
[0014] Step 4, using the cross-domain AU detection model to implement cross-domain AU detection.
[0015] Preferably, the obtaining step of the multi-modal AU dataset includes:
[0016] Step 1.1, obtaining an image dataset D including face images and corresponding labels V ;
[0017] Step 1.2, based on the Facial Action Coding System (FACS) manual, collecting text descriptions of AUs to obtain a text dataset D composed of the text descriptions of AUs L ;
[0018] Step 1.3, integrating the image dataset D V and the text dataset D L , an AU dataset D including an image dataset and a text dataset is formed.
[0019] Preferably, the steps for constructing the cross-domain AU detection network include:
[0020] Construct a visual encoder to extract features from image data. Among them, the visual encoder is divided into multiple stages, each stage corresponding to a different receptive field size, obtaining AU visual features of different scales output by each stage, and constructing an independent adapter for the AU visual features output by each stage to obtain AU visual features with the same spatial dimension;
[0021] Construct a tokenizer, take text data as input to obtain corresponding text vectors; concatenate a set of learnable prompt vectors with the text vectors to obtain text input;
[0022] Construct a text encoder to extract features from the text input to obtain corresponding AU text features;
[0023] Use the cross-modal attention mechanism to calculate interaction features with the AU visual features as queries, and the AU text features as keys and values;
[0024] Take the interaction features of each AU as the nodes of the graph neural network, and define the similarity between any two nodes as the edges of the graph neural network to construct the graph neural network; adopt a multi-round message passing mechanism to update the node features, and aggregate the neighbor node information through graph convolution to obtain the final feature representation of each node; concatenate the final feature representations of each node to obtain the fusion feature;
[0025] Construct an AU classifier, take the fusion feature as input to obtain the corresponding prediction probability.
[0026] Preferably, there are two directed edges between any two nodes , and the directed edge represents pointing from node to node , ; the directed edge represents pointing from node to node , .
[0027] Preferably, based on the source domain dataset, a multi-label binary cross-entropy loss function is used to supervise the training of the cross-domain AU detection network:
[0028]
[0029] Among them, represents the number of image data in the source domain dataset, Indicates the number of AUs, and respectively represent the label and predicted value of the -th AU in the -th image data of the source domain dataset, represents that the -th AU is activated in the -th image data of the source domain dataset, represents that the -th AU is not activated in the -th image data of the source domain dataset.
[0030] Preferably, contrast learning between AU visual features and AU text features is introduced to construct a contrast loss function , and based on the source domain dataset and the text dataset, the multi-label binary cross-entropy loss function is used to supervise the training of the cross-domain AU detection network:
[0031]
[0032]
[0033] Among them, is the overall loss function of the cross-domain AU detection network, is the weight, , , is the temperature hyperparameter, is the concatenated fusion result of the AU visual features at each stage corresponding to the -th image data of the source domain dataset, is the -th AU text feature.
[0034] Preferably, contrast learning of negative sample AU descriptions is introduced to construct a contrast learning loss function , and based on the image dataset and the text dataset, the multi-label binary cross-entropy loss function is used to supervise the optimization of the cross-domain AU detection network:
[0035]
[0036]
[0037]
[0038] Among them, is the overall loss function of the cross-domain AU detection network, 、 are respectively , The weight of represents the number of image data in the target domain dataset, represents the number of AUs, and respectively represent the pseudo-label and predicted value of the th AU in the th image data of the target domain dataset. If then represents that the th AU is activated in the th image data of the target domain dataset. If then represents that the th AU is not activated in the th image data of the target domain dataset. , represents the confidence threshold, , is the temperature hyperparameter, is the fusion feature corresponding to the th image data of the target domain dataset obtained through the graph neural network, is the AU text feature corresponding to the th AU, is the text feature of the th negative AU text description corresponding to the th AU extracted by the text encoder, is the number of negative AU text descriptions.
[0039] Preferably, training and optimizing the cross-domain AU detection network includes two stages:
[0040] In the first stage, the cross-domain AU detection network is trained using the source domain dataset and the text dataset;
[0041] In the second stage, the cross-domain AU detection network is optimized using the image dataset and the text dataset.
[0042] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0043] The present invention also provides an electronic device, including:
[0044] A memory for storing a computer program;
[0045] A processor for implementing the steps of the above method when executing the computer program.
[0046] Compared with the prior art, the significant advantages of the present invention are as follows: The present invention uses a visual encoder to extract AU visual features at different scales, introduces a set of learnable prompt vectors, combines the AU text description and inputs it into the text encoder to obtain AU text features. At the same time, contrastive learning is used to align the source domain features and text features to enhance the discriminability of the source domain AU visual features; a cross-modal attention mechanism is adopted to achieve the initial interaction between vision and text to obtain an AU interaction feature representation; through a graph neural network, the deep interaction and fusion of text and visual information are further promoted to obtain a fused AU feature; further, a contrastive learning method based on negative sample AU descriptions is adopted to align the target domain data and text features, indirectly narrowing the feature distribution difference between the source domain and the target domain. Finally, the method uses the domain invariance of the text to unify the visual features of the source domain and the target domain and align them with the text features, thereby narrowing the feature distribution difference between the source domain and the target domain, and significantly improving the generalization performance of the AU detection model and the accuracy on the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of a cross-domain facial expression action unit detection method based on text bridging.
[0048] Figure 2 is a schematic diagram of the training of the first stage of the cross-domain AU detection network structure.
[0049] Figure 3 is a schematic diagram of the training of the second stage of the cross-domain AU detection network structure.
[0050] Figure 4 is a schematic diagram of the test stage of the cross-domain AU detection network structure. DETAILED DESCRIPTION OF THE INVENTION
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] As Figure 1 shown, the specific process of the cross-domain facial expression action unit detection method based on text bridging of the present invention is as follows:
[0053] 1. Data preparation stage
[0054] 1.1 Collect a dataset of facial expression action units (AU)
[0055] Collect video clips containing diverse facial expressions and extract image frames from the videos. Then, use a face detection model to detect and crop the extracted image frames to obtain face images with a size of 256x256 pixels. According to the dataset label file, add the corresponding AU annotation information to the face images to construct an AU image dataset 。
[0056] 1.2 Collect AU Descriptions
[0057] Sort out the AU text descriptions from the Facial Action Coding System manual. These descriptions contain rich semantic information, covering the facial muscles, face regions, intensities, etc. involved in the AUs. By integrating the text data composed of the original AU descriptions and the AU image dataset composed of image data obtain a multimodal AU dataset 。In the cross-domain AU detection task, the AU image dataset is further divided into source domain data and target domain data 。
[0058] 2. In the model design stage, the specific design of the two-stage model is as follows:
[0059] 2.1 Denote the overall model as , which specifically includes a visual encoder , a text encoder , an adapter , cross-modal attention , a graph neural network and an AU classifier . The model input is the multimodal AU dataset , including source domain data , target domain data and AU descriptions .
[0060] 2.2 Extract multi-scale AU visual features. Different AUs cover different ranges of face regions, resulting in scale differences among different AUs. For example, AU6 (Cheek Raiser) usually covers a larger facial area, while AU12 (Lip Corner Puller) is limited to a smaller lip corner area. To comprehensively capture the visual cues of different scale AUs, use a frozen pre-trained CLIP visual encoder to extract AU visual features. The visual encoder is organized into four main parts in the network structure, and these parts are called stages. Each stage It is usually composed of multiple network modules (e.g., residual blocks) corresponding to receptive fields of different sizes. Let the input in the initial stage be the source domain data . Then, the AU visual features output in four stages of the visual encoder are extracted from shallow to deep, which can be expressed as:
[0061]
[0062] Among them, , represent the AU visual features extracted in the , th stage. To facilitate subsequent cross-modal interaction and feature fusion, an independent adapter is constructed for the AU visual features of each scale. This adapter consists of a one-dimensional convolutional layer and a pooling layer. The AU visual features are input into the corresponding adapter to obtain AU visual features with the same spatial dimension: .
[0063] 2.3 Extract AU text features. To improve the generalization ability of the AU detection model, the AU text is used to bridge the source domain and the target domain, and its domain invariance is utilized to reduce the feature distribution difference between the source domain and the target domain. First, to learn high-quality text feature representations, a set of learnable prompt vectors are introduced to learn task-specific guiding information, enabling the CLIP text encoder to generate more discriminative AU text features. The original text description of the th AU is encoded into the corresponding text vector using a pre-trained CLIP tokenizer. Then, the learnable prompt vectors are combined with this sequence to construct a text input in the following form , and it is input into the text encoder to obtain the text features corresponding to the original description of the th original AU:
[0064] .
[0065] To further improve the discriminability of the visual features, contrastive learning between vision and text is first introduced in the source domain. For the th input image in the source domain, the AU visual features of different scales are concatenated and fused through a fully connected layer to obtain the AU visual feature representation . Each of the th input images in the source domain has a set of corresponding AU labels, where the value of each label is 1 or 0, indicating whether the kth AU is "activated" in the image respectively. Let Denote the AU indices marked as active for the th input image in the source domain. The positive sample pairs are composed of visual features and text features corresponding to all active AUs, while the negative sample pairs are composed of the same visual features and text features corresponding to those non-active AUs. Define , for the th input image in the source domain, the contrastive learning loss can be expressed as:
[0066]
[0067] where is the cosine similarity, and is the temperature hyperparameter. Through contrastive learning, the model is prompted to better distinguish between active and non-active AUs in the feature space, align visual and text features, and thus improve the discriminative power of visual and text features.
[0068] 2.4 Constructing a Graph Neural Network to achieve sufficient interaction and fusion of visual and text information. First, cross-modal attention is adopted, using the AU visual feature as the query, and the AU text feature as the key and value to achieve the initial interaction between AU visual and text features, and calculate the AU interaction feature :
[0069]
[0070] where , , , , , is a learnable parameter matrix, and is the scaling factor ( usually takes the dimension of the key ). This cross-modal attention precisely semantically guides the vision through the text, thus effectively fusing semantic information, establishing a semantic association between vision and text, and obtaining more discriminative AU interaction features.
[0071] On this basis, to further promote the deep fusion of visual and text modalities, a graph neural network is constructed as . Among them, the node set of the graph consists of four AU interaction features ; the edge set is composed of the similarities between all nodes. For example, for any two different nodes , there are two directed edges reflects the correlation between node features. The directed edge represents from node pointing to , and its calculation formula is . Similarly, the directed edge can be calculated by the formula . In the graph neural network, a multi-round message passing mechanism is adopted to update node features to achieve deep fusion of visual and text information. In the t-th iteration, the message passed from node to node can be expressed as: . Among them, the function is used to perform row-wise normalization on , and and respectively represent the features of node and directed edge at the -th iteration. The aggregated message received by node from the neighbor node set can be expressed as: . Among them, the is used to adjust the shape of the features. Finally, the features of node are updated through a convolutional gated recurrent unit (ConvGRU):
[0072] .
[0073] After a total of T rounds of message passing, the final feature representation of the -th node is . Concatenating the node features along the channel dimension, the fused feature for AU detection is obtained. Using the multi-label binary cross-entropy loss to supervise the model training, it can be expressed as:
[0074]
[0075] Among them, represents the number of image data in the source domain dataset, represents the number of AUs, and respectively represent the label and predicted value of the -th AU in the -th image data of the source domain dataset, represents that the -th AU is activated in the -th image data of the source domain dataset, Indicates the th AU is not activated in the th image data of the source domain dataset.
[0076] 2.5 Contrastive learning based on negative sample AU descriptions. To reduce the difference in feature distributions between the source domain and the target domain, the present invention proposes a contrastive learning method based on negative sample AU descriptions, which indirectly reduces the difference in the feature distributions of the source domain and the target domain by aligning the visual features and text features of the target domain. First, relying on the powerful text understanding and generation capabilities of a large language model (such as DeepSeek), for each original AU description, the keywords (such as verbs, adjectives) that describe facial actions, states, or degrees of change are replaced with words having opposite meanings to generate negative AU descriptions with opposite meanings. Each original AU description can generate corresponding negative AU descriptions. For example, for the original text description of AU1: "The inner corners of the eyebrows are lifted slightly, the skin of the glabella and forehead above it is lifted slightly and wrinkles deepen slightly and a trace of new ones form in the center of the forehead.", the keywords such as "lifted", "deepen", and "slightly" can be replaced with "lowered", "smoothed", and corresponding antonyms respectively to generate a negative description with the opposite semantics. Then, the negative AU descriptions corresponding to the th AU are respectively input into the text encoder to obtain a set of negative AU text features . Secondly, considering that the target domain data lacks AU annotations, the model trained in the first stage is used to infer the th image of the target domain, and the graph neural network is used to extract the fused AU features , and the prediction probabilities of each AU are obtained through a classifier . Set a confidence threshold . For the th image of the target domain and the th AU, if , then assign a pseudo-label , indicating that this AU is in an activated state; if , then assign a pseudo-label , indicating that the AU is not activated. Finally, on the target domain, combining pseudo-labels, visual features of images, and all AU text descriptions, contrastive learning based on negative samples is constructed. For the th image on the target domain, its positive sample pair consists of visual features and AU text features corresponding to all activated AUs; the negative samples include two parts: one is visual features and AU text features corresponding to unactivated AUs, and the other is all negative AU text features corresponding to activated AUs. On this basis, for the th image on the target domain, its contrastive learning loss can be expressed as:
[0077]
[0078] where represents the number of AUs, represents the number of negative AU descriptions corresponding to each AU, represents the set of AU indices predicted to be activated in the th image on the target domain. The similarity function is defined as where represents the cosine similarity, and is the temperature hyperparameter. Through the above contrastive learning mechanism, the visual features of the target domain are effectively aligned to the corresponding text semantic space, thereby reducing the feature distribution difference between the source domain and the target domain to improve the generalization ability of the AU model.
[0079] 2.6 Model The training is completed in two stages. In the first stage, the source domain data and text data are used to train the model, and the overall loss function is used for model training, where represents the AU classification loss on the source domain, and represents the contrastive learning loss on the source domain. In this stage, the model is jointly optimized, including the visual encoder , the text encoder , the adapter , the cross-modal attention , the graph neural network and the AU classifier . On this basis, in the second stage, the source domain data, target domain data, and text data are used to train the model, and the overall loss function is used to jointly optimize the model , where represents the AU classification loss on the target domain, and represents the contrastive learning loss based on negative samples. The model Update is performed by the gradient descent method. Steps 2.1, 2.2, 2.3, and 2.4 are executed in sequence to complete the training of the first stage. Subsequently, step 2.5 is executed to complete the training of the second stage until the model converges. The parameter The update follows the following strategy:
[0080]
[0081] where represents the learning rate.
[0082] 2.7 The above steps are unified into a two-stage deep neural network framework to achieve optimized training of the model.
[0083] 3. Model Training Stage
[0084] 3.1 Input the source domain dataset in step 1 , the target domain data , and the text data into the network model designed in step 2 . Use the batch stochastic gradient descent method for model training. There are 4 loss functions, namely the source domain AU classification loss , the target domain AU classification loss , the source domain contrastive learning loss , and the contrastive learning loss based on negative samples . The target domain dataset is used during the training stage to verify the model training effect, that is, when the model obtains good cross-domain AU detection results on the target domain dataset , and the performance no longer improves in subsequent training iterations, then the model training converges and the training process stops.
[0085] 3.2 The final trained model is obtained.
[0086] 4. Model Testing Stage
[0087] 4.1 The input data is the target domain dataset and the text data . The model is used during the testing stage, including the visual encoder , the text encoder , the cross-modal attention , the graph neural network , and the AU classifier .
[0088] 4.2 Input the target domain data into the model obtained in step 3.2 The facial action unit detection results of the target domain can be obtained therefrom.
[0089] Table 1 shows the performance comparison of the method proposed in the present invention with the Source and Target methods in four different cross-domain facial action unit detection tasks, involving the BP4D, BP4D+ and GFT datasets. The values in the table are F1 scores, and the arrows indicate the direction of data migration. For example, B→G means using the BP4D dataset as the source domain and the GFT dataset as the target domain for cross-domain AU detection. Among them, the "Source" method means directly applying the model trained only on the source domain data to the target domain data for evaluation, serving as a baseline without cross-domain adaptation; while the "Target" method means performing in-domain training and evaluation on the target domain data, serving as the theoretical upper limit of cross-domain adaptation. The experimental results show that the method proposed in the present invention has achieved significant performance improvement in all four cross-domain scenarios, demonstrating the effectiveness of the proposed method in cross-domain AU detection.
[0090] Table 1 Cross-Domain Facial Action Unit Detection Results
[0091] Method B→G G→B B+→G G→B+ Source 38.1 46.3 36.5 41.3 The proposed method 42.1 52.6 39.7 46.1 Target 52.0 61.0 52.0 56.2
[0092] To solve the problem of insufficient robustness in the prior art due to over-reliance on limited labeled datasets, a cross-domain facial action unit detection method based on text bridging proposed in the present invention has the following prominent key points:
[0093] 1) Multi-scale AU visual feature extraction: Different AUs cover different ranges of the human face area, that is, there are scale differences between different facial AUs. For example, AU6 (Cheek Raiser) covers a large facial area, while AU12 (LipCorner Puller) is limited to a smaller corner-of-mouth area. To effectively represent the visual features of AUs at different scales, a frozen pre-trained CLIP visual encoder is used for multi-scale AU feature extraction. This encoder consists of four stages, corresponding to receptive fields of different sizes, and can extract visual features of AUs at different scales from shallow to deep, thus comprehensively capturing AU visual cues. To promote subsequent cross-modal interaction and fusion, an adapter composed of one-dimensional convolutional and pooling layers is constructed for the visual features of AUs at each scale, unifying the multi-scale AU visual features extracted at different stages to the same spatial dimension.
[0094] 2) AU Text Feature Extraction: To utilize the rich semantic information with domain-invariant properties contained in AU text descriptions, they are encoded into high-quality AU text features. First, AU text descriptions are sorted out from the Facial Action Coding System manual. These text descriptions not only contain rich emotional semantic information but also have domain invariance, and their semantics are not affected by specific visual interference factors such as age, lighting, and pose. Next, to learn high-quality text feature representations, a set of learnable prompt vectors are introduced. The prompt vectors are combined with AU descriptions to form a text input sequence. Then, this sequence is input into a frozen pre-trained CLIP text encoder to obtain AU text features. Finally, contrastive learning between vision and text is conducted on the source domain to align the source domain visual features to the text features, enhancing the AU discriminability of visual features. At the same time, it lays a foundation for using domain-invariant text as a bridge in the future and achieving cross-domain adaptation of visual features.
[0095] 3) Constructing a Graph Neural Network to Facilitate the Interaction and Fusion of Visual and Text Modalities: First, cross-modal attention is adopted, using AU visual features as the query (Query), and AU text features as the key (Key) and value (Value). This cross-modal attention uses text to provide precise semantic guidance for vision, effectively fusing semantic information, establishing semantic associations between vision and text, and obtaining more discriminative AU interaction features. Second, to further promote the deep fusion of visual and text modalities, a graph neural network with AU interaction features as nodes is constructed. This network realizes the dynamic update of nodes through a message passing mechanism. Each node continuously aggregates information from other nodes in multiple rounds of iteration, further deepening the deep fusion of multi-scale AU visual features and text information, and thus obtaining enhanced AU features with richer semantics. Finally, all enhanced AU feature representations are concatenated and fused using a fully connected layer to obtain fused AU features for AU detection.
[0096] 4) Contrastive learning based on negative sample AU descriptions: To reduce the feature distribution difference between source-domain images and target-domain images, contrastive learning based on negative sample AU descriptions is adopted to align the visual features and text features of the target domain, indirectly reducing the feature distribution and semantic differences between the source domain and the target domain. First, the large language model (such as Chatgpt, DeepSeek) is used to replace the key verbs or adjectives in the original AU description with words having opposite meanings, thus generating several negative AU descriptions with opposite meanings for each AU. Then, with the help of the AU detection model trained in the first stage, the target-domain images are inferred to obtain the activation probabilities of each AU. By setting a confidence threshold, AU "pseudo-labels" are generated for the target-domain data to indicate the activation status of the corresponding AU in the target-domain images. For the target-domain images, contrastive learning is implemented in combination with the pseudo-labels, AU visual features, the original AU description, and the generated negative sample AU descriptions. This contrastive learning mechanism can promote the model to deeply understand the AU activation status described by the text, and align the visual features of the target domain with the AU text features in terms of features and semantics, establishing a more accurate correspondence between visual AUs and text descriptions. Since the source-domain features and text features have been aligned in the first stage, and the target-domain features are also aligned with the text features in this stage, the feature distribution difference between the source domain and the target domain is indirectly reduced, thus improving the cross-domain generalization ability of the AU model.
[0097] 5) Cross-domain expression action unit detection method based on text bridging: As Figures 2 to 4 shown, the input data are AU descriptions, source-domain and target-domain data, and a two-stage training model is used. For the first stage: First, a frozen pre-trained CLIP visual encoder is used for multi-scale AU feature extraction. Then, different-scale AU visual features are mapped to a unified spatial dimension through an adapter. Next, a set of learnable prompt vectors are introduced and combined with the AU text description and input into the frozen pre-trained CLIP text encoder to obtain AU text features. At the same time, contrastive learning is implemented between vision and text in the source domain to align the visual features and text features of the source domain. Further, sufficient interaction and fusion, and alignment of the visual and text modalities are carried out. First, cross-modal attention is used for the initial interaction between vision and text, and then the deep fusion of visual and text information is realized through a graph neural network. Finally, the enhanced AU feature representations corresponding to different nodes are subjected to feature concatenation and fusion to obtain the fused AU features for expression action unit detection. The source-domain AU classification loss and the source-domain contrastive learning loss Perform model training. For the second stage: First, use a large language model (such as DeepSeek) to rewrite the original AU description to obtain negative sample AU descriptions. Then, input the target domain samples into the model of the first stage to obtain pseudo-labels for the target domain samples. Next, through contrast learning based on negative samples on the target domain, align the visual features and text features of the target domain, thereby indirectly reducing the difference in feature distributions between the source domain and the target domain and improving the AU detection performance of the model on the target domain. The second stage uses the source domain AU classification loss , the target domain AU classification loss , and the contrast learning loss based on negative samples to jointly optimize the model.
[0098] In summary, the present invention designs a cross-domain facial expression action unit detection method network based on text bridging. This network uses a frozen pre-trained CLIP visual encoder to extract AU visual features at different scales, and introduces a set of learnable prompt vectors. The AU text description is combined and input into the frozen CLIP text encoder to obtain AU text features. At the same time, contrast learning is used to align the source domain features and text features to improve the discriminability of visual features and text features; cross-modal attention is used to achieve the initial interaction between vision and text to obtain the AU interaction feature representation; the graph neural network is further used to promote the deep interaction and fusion of text and visual information to obtain the fused AU features; a contrast learning method based on negative samples is used to align the target domain data and text features, indirectly reducing the difference in feature distributions between the source domain and the target domain. Finally, this method uses the domain invariance of the text to align both the visual features of the source domain and the target domain with the text features, thereby reducing the difference in feature distributions between the source domain and the target domain and significantly improving the generalization performance of the AU detection model and the accuracy on the target domain.
[0099] Based on the same technical solution, the present invention also proposes an electronic device, including:
[0100] A memory for storing a computer program;
[0101] A processor for implementing the steps of the above-mentioned cross-domain facial expression action unit detection method based on text bridging when executing the computer program.
[0102] Based on the same technical solution, the present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned cross-domain expression motion unit detection method based on text bridging are implemented. The computer-readable storage medium may include various media capable of storing program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memories (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memories. Volatile memories can include random access memories (RAM) or external cache memories. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
Claims
1. A cross-domain facial expression action unit detection method based on text bridging, characterized in that Including: Step 1: Obtain a multimodal expression Action Unit (AU) dataset including an image dataset and a text dataset; Step 2: Construct a cross-domain AU detection network; The construction steps of the cross-domain AU detection network include: Construct a visual encoder to extract features from image data. The visual encoder is divided into multiple stages, each stage corresponding to a different receptive field size, obtaining AU visual features of different scales output by each stage, and constructing an independent adapter for the AU visual features output by each stage to obtain AU visual features with the same spatial dimension; Construct a tokenizer, take text data as input, and obtain corresponding text vectors; concatenate a set of learnable prompt vectors with the text vectors to obtain a text input; Construct a text encoder to extract features from the text input to obtain corresponding AU text features; Use the AU visual features as queries, the AU text features as keys and values, and calculate interaction features using a cross-modal attention mechanism; Use the interaction features of each AU as nodes of a graph neural network, define the similarity between any two nodes as the edges of the graph neural network, and construct a graph neural network; adopt a multi-round message passing mechanism to update the node features, and aggregate neighbor node information through graph convolution to obtain the final feature representation of each node; concatenate the final feature representations of each node to obtain a fused feature; Construct an AU classifier, take the fused feature as input, and obtain corresponding prediction probabilities; Step 3: Divide the image dataset into a source domain dataset and a target domain dataset, and combine the text dataset to train and optimize the cross-domain AU detection network to obtain a cross-domain AU detection model; Based on the source domain dataset, the multi-label binary cross-entropy loss function is adopted. Training of the supervised cross-domain AU detection network: , Among them, represents the number of image data in the source domain dataset, represents the number of AUs, and respectively represent the label and predicted value of the th AU in the th image data of the source domain dataset, represents that the th AU is activated in the th image data of the source domain dataset, represents that the th AU is not activated in the th image data of the source domain dataset; Introduce contrastive learning with negative sample AU descriptions and construct a contrastive learning loss function , based on the image dataset and the text dataset, jointly use the multi-label binary cross-entropy loss function , supervise the optimization of the cross-domain AU detection network: , , , Among them, is the overall loss function of the cross-domain AU detection network, and are respectively and the weights of, represents the number of image data in the target domain dataset, represents the number of AUs, and respectively represent the pseudo-label and predicted value of the th AU in the th image data of the target domain dataset. If then represents that the th AU is activated in the th image data of the target domain dataset. If then represents that the th AU is not activated in the th image data of the target domain dataset, , represents the confidence threshold, , is the temperature hyperparameter, is the fusion feature corresponding to the th image data of the target domain dataset obtained via the graph neural network, is the AU text feature corresponding to the th AU, is the text feature of the th negative AU text description corresponding to the th AU extracted via the text encoder, is the number of negative AU text descriptions; Step 4: Use the cross-domain AU detection model to achieve cross-domain AU detection.
2. The method according to claim 1, wherein The steps for obtaining the multimodal expression Action Unit (AU) dataset include: Step 1.1, obtain an image dataset D including face images and corresponding labels V ; Step 1.2, based on the Facial Action Coding System (FACS) manual, collect the text descriptions of AUs to obtain a text dataset D composed of the text descriptions of AUs L ; Step 1.3, integrate the image dataset D V and the text dataset D L , to form the AU dataset D that includes the image dataset and the text dataset.
3. The method according to claim 1, wherein Any two nodes There are two directed edges between them , and the directed edge indicates from node pointing to node , ; the directed edge indicates from node pointing to node , .
4. The method according to claim 1, characterized in that Contrastive learning between AU visual features and AU text features to construct a contrastive loss function , based on the source domain dataset and the text dataset, jointly use the multi-label binary cross-entropy loss function Supervise the training of the cross-domain AU detection network: , , Among them, is the overall loss function of the cross-domain AU detection network, is 's weight, , , is the temperature hyperparameter, is the concatenated fusion result of the AU visual features at each stage corresponding to the th image data of the source domain dataset, is the th AU text feature.
5. The method according to claim 1, wherein Training and optimizing the cross-domain AU detection network includes two stages: The first stage: Train the cross-domain AU detection network using the source domain dataset and the text dataset; The second stage: Optimize the cross-domain AU detection network using the image dataset and the text dataset.
6. A computer-readable storage medium, characterized in that, A computer program is stored on the computable readable storage medium, and when the computer program is executed by a processor, the steps of the method as described in any one of claims 1 to 5 are implemented.
7. An electronic device, including: A memory for storing a computer program; A processor for implementing the steps of the method as described in any one of claims 1 to 5 when executing the computer program.
Citation Information
Patent Citations
Facial action unit detection model construction method based on hierarchical feature alignment
CN117576765A
Facial action unit detection model establishment method based on multi-task learning
CN117765596A
Speaking face generation method and device, electronic equipment and storage medium
CN116844215A
Audio driven facial animation using machine learning
CN119494894A