Cross-domain expression motion unit detection method based on text bridging

Through the cross-domain AU detection method based on text bridging, using multimodal data sets and cross-domain AU detection network, AU features are extracted and fused, and the problems of scarcity of labeled data and differences in cross-domain feature distribution when AU detection technology is applied in educational scenarios, achieving higher detection accuracy and generalization capabilities.

CN120145166AActive Publication Date: 2025-06-13JIANGSU SECOND NORMAL UNIVERSITY +2

Patent Information

Application Number
CN202510634229.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

When used in primary and secondary education scenarios, existing AU detection technology faces the problems of scarcity of labeled data and differences in cross-domain data feature distribution, resulting in a decrease in detection accuracy.

Method used

A cross-domain AU detection method based on text bridging is adopted to acquire multimodomain AU data sets, and a cross-domain AU detection network is built, and AU features are extracted using visual encoder and text encoder, and feature fusion and comparison learning are performed through cross-modal attention and graph neural networks to narrow the difference in feature distribution between the source domain and the target domain.

Benefits of technology

It significantly improves the generalization performance and accuracy on the target domain of the AU detection model, and enhances the robustness and cross-domain adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145166A_ABST
    Figure CN120145166A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain expression motion unit detection method based on text bridging, and the method comprises the steps: extracting AU visual features of different scales through employing a visual encoder, introducing a group of learnable prompt vectors, and combining with AU text description to obtain AU text features; the source domain features and the text features are aligned through comparative learning, and the discrimination of the source domain AU visual features is improved; a cross-modal attention mechanism is adopted to realize preliminary interaction of vision and texts, and AU interaction feature representation is obtained; further promoting deep interaction and fusion of the text and the visual information through a graph neural network to obtain a fused AU feature; and on the basis of a negative sample AU description-based comparative learning method, target domain features are aligned with text features, and feature distribution differences between a source domain and a target domain are reduced. According to the method, the field invariance of the text is utilized, the visual features of the source domain and the target domain are unified, the text features are aligned, and the generalization performance of an AU detection model and the accuracy of the AU detection model on the target domain are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of facial expression action unit detection, and mainly relates to a cross-domain facial expression action unit detection method based on text bridging. Background Art

[0002] Facial expressions are the intuitive external expressions of emotions and can directly convey a person's emotional state. Early facial expression analysis methods only classified expressions into basic emotions such as happiness, sadness, and surprise, and could not recognize complex and subtle expression changes. To comprehensively and objectively analyze facial expressions, psychologist Ekman proposed the Facial Action Coding System (FACS) from the perspective of human face anatomy. The FACS system defines 44 facial expression action units (ActionUnit, AU), and each AU corresponds to a movement of a local muscle of the human face. For example, AU1 (Inner Brow Raiser) represents the movement of raising the inner part of the eyebrows. Facial expressions can be represented as a combination of a series of AUs. For example, a happy expression is a combination of AU6 (Cheek Raiser) and AU12 (Lip Corner Puller). The goal of AU detection is to build a detection system that can automatically recognize facial AUs in images or videos. With the rapid development of deep learning technology, AU detection technology has been further developed and shows great application potential in fields such as human-computer interaction, medical health, and auxiliary education.

[0003] In the context of primary and secondary school education, the adoption of AU detection technology helps teachers promptly understand the learning status and emotional changes of students, enabling them to more precisely adjust teaching plans and improve the overall teaching quality. However, two core challenges limit the application of AU detection technology in educational scenarios:

[0004] First, the labeled AU dataset for primary and secondary school students is extremely scarce, and AU annotation is time-consuming and costly.

[0005] Second, when the current AU detection model is applied to cross-domain scenarios such as intelligent education, the data in these scenarios is significantly different from the training data used by the original AU model, and the AU detection model shows a significant decline in performance.

[0006] To address these challenges, it is urgent to research and implement a robust and efficient cross-domain AU detection method that uses the labeled AU data in the source domain (public dataset) and the unlabeled data in the target domain (educational scenario) to greatly reduce the difference in feature distributions between the source domain and the target domain, so as to improve the detection accuracy of the AU model in the intelligent education scenario.

[0007] The publication number is CN117765596A, and the name is a patent application for an invention method for establishing a facial action unit detection model based on multi-task learning. Its main technical means are: combining facial AU detection, key point detection, and emotion recognition for multi-task joint learning, and using a graph attention network to model the dependencies between AUs to improve the accuracy of AU detection. The publication number is CN117576765A, and the name is a patent application for an invention method for constructing a facial action unit detection model based on hierarchical feature alignment. Its main technical means are: performing consistency alignment within AU categories, between AU categories, and at the sample level to improve the accuracy of AU detection. Although these methods have achieved performance improvements in the AU detection task, the generalization ability of the model is insufficient, and it is difficult to effectively cope with the challenges of cross-domain AU detection. Summary of the Invention

[0008] The present invention provides a cross-domain expression motion unit detection method based on text bridging to effectively improve the generalization ability and robustness of the model and cope with the challenges of cross-domain AU detection.

[0009] To achieve the above object, the solution of the present invention is:

[0010] A cross-domain expression motion unit detection method based on text bridging, comprising:

[0011] Step 1, obtaining a multi-modal expression motion unit AU dataset including an image dataset and a text dataset;

[0012] Step 2, constructing a cross-domain AU detection network;

[0013] Step 3, dividing the image dataset into a source domain dataset and a target domain dataset, and training and optimizing the cross-domain AU detection network in combination with the text dataset to obtain a cross-domain AU detection model;

[0014] Step 4, using the cross-domain AU detection model to achieve cross-domain AU detection.

[0015] Preferably, the obtaining step of the multi-modal AU dataset includes:

[0016] Step 1.1, obtaining an image dataset D including face images and corresponding labels V ;

[0017] Step 1.2, based on the Facial Action Coding System (FACS) manual, collecting text descriptions of AUs to obtain a text dataset D composed of text descriptions of AUs L ;

[0018] Step 1.3, integrating the image dataset D V and the text dataset D L ​​, form the AU dataset D including the image dataset and the text dataset.

[0019] Preferably, the steps for constructing the cross-domain AU detection network include:

[0020] Construct a visual encoder to extract features from the image data. Among them, the visual encoder is divided into multiple stages, each stage corresponding to a different receptive field size, obtain the AU visual features of different scales output by each stage, and construct an independent adapter for the AU visual features output by each stage to obtain AU visual features with the same spatial dimension;

[0021] Construct a tokenizer, take the text data as input to obtain the corresponding text vector; concatenate a set of learnable prompt vectors with the text vector to obtain the text input;

[0022] Construct a text encoder to extract features from the text input to obtain the corresponding AU text features;

[0023] Use the cross-modal attention mechanism to calculate the interaction features with the AU visual features as the query and the AU text features as the key and value;

[0024] Use the interaction features of each AU as the nodes of the graph neural network, and define the similarity between any two nodes as the edges of the graph neural network to construct the graph neural network; adopt a multi-round message passing mechanism to update the node features, and aggregate the neighbor node information through graph convolution to obtain the final feature representation of each node; concatenate the final feature representations of each node to obtain the fusion feature;

[0025] Construct an AU classifier, take the fusion feature as input to obtain the corresponding prediction probability.

[0026] Preferably, there are two directed edges between any two nodes The directed edge means from node pointing to node pointing to node , ; the directed edge means from node pointing to node pointing to node .

[0027] Preferably, based on the source domain dataset, use the multi-label binary cross-entropy loss function to supervise the training of the cross-domain AU detection network:

[0028]

[0029] Among them, represents the number of image data in the source domain dataset, Indicates the number of AUs, and respectively represent the label and predicted value of the th AU in the th image data of the source domain dataset, represents that the th AU is activated in the th image data of the source domain dataset, represents that the th AU is not activated in the th image data of the source domain dataset.

[0030] Preferably, contrast learning between AU visual features and AU text features is introduced to construct a contrast loss function , and based on the source domain dataset and the text dataset, a multi-label binary cross-entropy loss function is used to supervise the training of the cross-domain AU detection network:

[0031]

[0032]

[0033] Among them, is the overall loss function of the cross-domain AU detection network, is 's weight, , , is the temperature hyperparameter, is the concatenated fusion result of the AU visual features at each stage corresponding to the th image data of the source domain dataset, is the th AU text feature.

[0034] Preferably, contrast learning of negative sample AU descriptions is introduced to construct a contrast learning loss function , and based on the image dataset and the text dataset, a multi-label binary cross-entropy loss function is used to supervise the optimization of the cross-domain AU detection network:

[0035]

[0036]

[0037]

[0038] Among them, is the overall loss function of the cross-domain AU detection network, , are respectively , The weight of represents the number of image data in the target domain dataset, represents the number of AUs, and respectively represent the pseudo-label and predicted value of the th AU in the th image data of the target domain dataset. If then represents that the th AU is activated in the th image data of the target domain dataset. If then represents that the th AU is not activated in the th image data of the target domain dataset. , represents the confidence threshold, , is the temperature hyperparameter, is the fused feature corresponding to the th image data of the target domain dataset obtained via the graph neural network, is the AU text feature corresponding to the th AU, is the text feature of the th negative AU text description corresponding to the th AU extracted via the text encoder, is the number of negative AU text descriptions.

[0039] Preferably, training and optimizing the cross-domain AU detection network includes two stages:

[0040] The first stage is to train the cross-domain AU detection network using the source domain dataset and the text dataset;

[0041] The second stage is to optimize the cross-domain AU detection network using the image dataset and the text dataset.

[0042] The present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0043] The present invention also proposes an electronic device, including:

[0044] A memory for storing a computer program;

[0045] A processor for implementing the steps of the above method when executing the computer program.

[0046] Compared with the prior art, the remarkable advantages of the present invention are as follows: The present invention uses a visual encoder to extract AU visual features at different scales, introduces a set of learnable prompt vectors, combines the AU text description and inputs it into the text encoder to obtain AU text features. At the same time, contrastive learning is used to align the source domain features and text features to enhance the discriminability of the source domain AU visual features; a cross-modal attention mechanism is adopted to achieve the initial interaction between vision and text to obtain an AU interaction feature representation; through a graph neural network, the in-depth interaction and fusion of text and visual information are further promoted to obtain a fused AU feature; further, a contrastive learning method based on negative sample AU descriptions is adopted to align the target domain data and text features, indirectly reducing the difference in feature distributions between the source domain and the target domain. Finally, the method utilizes the domain invariance of the text to unify the visual features of the source domain and the target domain and align them with the text features, thereby reducing the difference in feature distributions between the source domain and the target domain, and significantly improving the generalization performance of the AU detection model and the accuracy on the target domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flowchart of a cross-domain facial expression action unit detection method based on text bridging.

[0048] Figure 2 is a schematic diagram of the training of the first stage of the cross-domain AU detection network structure.

[0049] Figure 3 is a schematic diagram of the training of the second stage of the cross-domain AU detection network structure.

[0050] Figure 4 is a schematic diagram of the test stage of the cross-domain AU detection network structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] As Figure 1 shown, the specific process of the cross-domain facial expression action unit detection method based on text bridging of the present invention is as follows:

[0053] 1. Data preparation stage

[0054] 1.1 Collect a facial expression action unit (AU) dataset

[0055] Collect video clips containing diverse facial expressions and extract image frames from the videos. Then, use a face detection model to perform face detection and cropping on the extracted image frames to obtain face images with a size of 256x256 pixels. According to the dataset label file, add corresponding AU annotation information to the face images to construct an AU image dataset 。

[0056] 1.2 Collect AU Descriptions

[0057] Sort out the AU text descriptions from the Facial Action Coding System manual. These descriptions contain rich semantic information, covering the facial muscles, face regions, intensities, etc. involved in the AUs. By integrating the text data composed of the original AU descriptions and the AU image dataset composed of image data ,a multi-modal AU dataset is obtained 。In the cross-domain AU detection task, the AU image dataset is further divided into source domain data and target domain data 。

[0058] 2. In the model design stage, the specific design of the two-stage model is as follows:

[0059] 2.1 The overall model is denoted as and specifically includes a visual encoder , a text encoder , an adapter , cross-modal attention , a graph neural network and an AU classifier 。The model input is the multi-modal AU dataset , including source domain data , target domain data and AU descriptions 。

[0060] 2.2 Extract multi-scale AU visual features. Different AUs cover different ranges of face regions, resulting in scale differences among different AUs. For example, AU6 (Cheek Raiser) usually covers a larger facial area, while AU12 (Lip Corner Puller) is limited to a smaller lip corner area. To comprehensively capture the visual cues of different scale AUs, a frozen pre-trained CLIP visual encoder is used to extract AU visual features. The visual encoder is organized into four main parts in the network structure, and these parts are called stages. Each stage It is usually composed of multiple network modules (e.g., residual blocks) corresponding to receptive fields of different sizes. Let the input in the initial stage be the source domain data . Then, the AU visual features output at four stages of the visual encoder are extracted from shallow to deep, which can be expressed as:

[0061]

[0062] Among them, and represent the AU visual features extracted at the -th and -th stages. To facilitate subsequent cross-modal interaction and feature fusion, an independent adapter is constructed for the AU visual features of each scale. This adapter consists of a one-dimensional convolutional layer and a pooling layer. The AU visual features are input into the corresponding adapter to obtain AU visual features with the same spatial dimension: .

[0063] 2.3 Extract AU text features. To improve the generalization ability of the AU detection model, the AU text is used to bridge the source domain and the target domain, and its domain invariance is utilized to reduce the difference in feature distributions between the source domain and the target domain. First, to learn high-quality text feature representations, a set of learnable prompt vectors are introduced to learn task-specific guiding information, enabling the CLIP text encoder to generate more discriminative AU text features. The original text description of the -th AU is encoded into the corresponding text vector using a pre-trained CLIP tokenizer. Then, the learnable prompt vectors are combined with this sequence to construct a text input in the following form and input into the text encoder to obtain the text features corresponding to the original description of the -th AU:

[0064] .

[0065] To further improve the discriminability of visual features, contrastive learning between vision and text is first introduced in the source domain. For the -th input image in the source domain, the AU visual features at different scales are concatenated and fused through a fully connected layer to obtain the AU visual feature representation . Each input image in the source domain has a set of corresponding AU labels, where the value of each label is 1 or 0, indicating whether the k-th AU is "activated" in the image. Let ​​Denote the AU indices marked as activated for the th input image in the source domain. The positive sample pairs are composed of visual features and text features corresponding to all activated AUs, while the negative sample pairs are composed of the same visual features and text features corresponding to those non-activated AUs. Define , for the th input image in the source domain, the contrastive learning loss can be expressed as:

[0066]

[0067] where is the cosine similarity, and is the temperature hyperparameter. Through contrastive learning, the model is encouraged to better distinguish activated and non-activated AUs in the feature space, align visual and text features, and thus improve the discriminative power of visual and text features.

[0068] 2.4 Constructing the Graph Neural Network to achieve sufficient interaction and fusion of visual and text information. First, cross-modal attention is adopted, using the AU visual feature as the query, and the AU text feature as the key and value to achieve the initial interaction of AU visual and text features, and calculate the AU interaction feature :

[0069]

[0070] where , , , , , is the learnable parameter matrix, and is the scaling factor ( usually takes the dimension of the key ). This cross-modal attention precisely semantically guides the vision through the text, thus effectively fusing semantic information, establishing semantic associations between vision and text, and obtaining more discriminative AU interaction features.

[0071] On this basis, to further promote the deep fusion of visual and text modalities, construct the graph neural network as . Among them, the node set of the graph consists of four AU interaction features ; the edge set consists of the similarities between all nodes. For example, for any two different nodes , there are two directed edges , reflecting the correlation between node features. The directed edge indicates starting from node pointing to , and its calculation formula is . Similarly, the directed edge can be calculated by the formula . In the graph neural network, a multi-round message passing mechanism is adopted to update node features to achieve deep fusion of visual and text information. In the t-th iteration, the message passed from node to node can be expressed as: . Among them, the function is used to perform row-wise normalization on , and and respectively represent the features of node and directed edge at the -th iteration. The aggregated message received by node from the neighbor node set can be expressed as: . Among them, is used to adjust the shape of the feature. Finally, the feature of node is updated through a convolutional gated recurrent unit (ConvGRU):

[0072] .

[0073] After a total of T rounds of message passing, the final feature representation of the -th node is . Concatenating the node features along the channel dimension, the fused feature for AU detection is obtained. Using the multi-label binary cross-entropy loss to supervise model training can be expressed as:

[0074]

[0075] Among them, represents the number of image data in the source domain dataset, represents the number of AUs, and respectively represent the label and predicted value of the -th AU in the -th image data of the source domain dataset, represents that the -th AU is activated in the -th image data of the source domain dataset, Indicates the th AU is not activated in the th image data of the source domain dataset.

[0076] 2.5 Contrastive learning based on negative sample AU descriptions. To reduce the difference in feature distributions between the source domain and the target domain, the present invention proposes a contrastive learning method based on negative sample AU descriptions, which indirectly reduces the difference in the feature distributions of the source domain and the target domain by aligning the visual features and text features of the target domain. First, leveraging the powerful text understanding and generation capabilities of a large language model (such as DeepSeek), for each original AU description, the keywords (such as verbs, adjectives) that describe facial actions, states, or degrees of change are replaced with words having opposite meanings to generate negative AU descriptions with opposite meanings. Each original AU description can generate corresponding negative AU descriptions. For example, for the original text description of AU1: "The inner corners of the eyebrows are lifted slightly, the skin of the glabella and forehead above it is lifted slightly and wrinkles deepen slightly and a trace of new ones form in the center of the forehead.", the keywords such as "lifted", "deepen", and "slightly" can be replaced with "lowered", "smoothed", and corresponding antonyms respectively to generate a negative description with the opposite semantics. Then, the negative AU descriptions corresponding to the th AU are respectively input into the text encoder to obtain the negative AU text feature set . Secondly, considering that the target domain data lacks AU annotations, the model trained in the first stage is used to infer the th image of the target domain, and the graph neural network is used to extract the fused AU features , and the predicted probabilities of each AU are obtained through the classifier . A confidence threshold is set. For the th image of the target domain and the th AU, if , then assign the pseudo-label , indicating that this AU is in the activated state; if , then assign the pseudo-label , indicating that the AU is not activated. Finally, on the target domain, combining pseudo-labels, visual features of images, and all AU text descriptions, contrastive learning based on negative samples is constructed. For the th image on the target domain, its positive sample pair consists of visual features and AU text features corresponding to all activated AUs; the negative samples include two parts: one is visual features and AU text features corresponding to unactivated AUs, and the other is all negative AU text features corresponding to activated AUs. On this basis, for the th image on the target domain, its contrastive learning loss can be expressed as:

[0077]

[0078] where represents the number of AUs, represents the number of negative AU descriptions corresponding to each AU, represents the set of AU indices predicted to be activated in the th image on the target domain. The similarity function is defined as where represents the cosine similarity, and is the temperature hyperparameter. Through the above contrastive learning mechanism, the visual features of the target domain are effectively aligned to the corresponding text semantic space, thereby reducing the feature distribution difference between the source domain and the target domain to improve the generalization ability of the AU model.

[0079] 2.6 Model The training is completed in two stages. In the first stage, the model is trained using source domain data and text data, and the overall loss function is used for model training, where represents the AU classification loss on the source domain, and represents the contrastive learning loss on the source domain. In this stage, the model , including the visual encoder , the text encoder , the adapter , the cross-modal attention , the graph neural network , and the AU classifier , is jointly optimized. On this basis, in the second stage, the model is trained using source domain data, target domain data, and text data, and the overall loss function is used to jointly optimize the model , where represents the AU classification loss on the target domain, and represents the contrastive learning loss based on negative samples. The model Update is performed by the gradient descent method. Steps 2.1, 2.2, 2.3, and 2.4 are executed in sequence to complete the training in the first stage. Subsequently, step 2.5 is executed to complete the training in the second stage until the model converges. The parameters The update follows the following strategy:

[0080]

[0081] where represents the learning rate.

[0082] 2.7 The above steps are unified into a two-stage deep neural network framework to achieve optimized training of the model.

[0083] 3. Model Training Stage

[0084] 3.1 Input the source domain dataset in step 1 , the target domain data and the text data into the network model designed in step 2 , and use the batch stochastic gradient descent method for model training. There are 4 loss functions, namely the source domain AU classification loss , the target domain AU classification loss , the source domain contrastive learning loss , and the contrastive learning loss based on negative samples . The target domain dataset is used during the training stage to verify the model training effect, that is, when the model obtains good cross-domain AU detection results on the target domain dataset , and the performance no longer improves during subsequent training iterations, then the model training converges and the training process stops.

[0085] 3.2 The final trained model is obtained.

[0086] 4. Model Testing Stage

[0087] 4.1 The input data is the target domain dataset and the text data . The model is used during the testing stage, including the visual encoder , the text encoder , the cross-modal attention , the graph neural network and the AU classifier .

[0088] 4.2 Input the target domain data into the model obtained in step 3.2 The facial action unit detection results of the target domain can be obtained therefrom.

[0089] Table 1 shows the performance comparison of the method proposed in the present invention with the Source and Target methods in four different cross-domain facial action unit detection tasks, involving the BP4D, BP4D+ and GFT datasets. The values in the table are F1 scores, and the arrows indicate the direction of data migration. For example, B→G means using the BP4D dataset as the source domain and the GFT dataset as the target domain for cross-domain AU detection. Among them, the "Source" method means directly applying the model trained only on the source domain data to the target domain data for evaluation, serving as a baseline without cross-domain adaptation; while the "Target" method means performing in-domain training and evaluation on the target domain data, serving as the theoretical upper limit of cross-domain adaptation. The experimental results show that the method proposed in the present invention has achieved significant performance improvement in all four cross-domain scenarios, demonstrating the effectiveness of the proposed method in cross-domain AU detection.

[0090] Table 1 Cross-domain facial action unit detection results Method B→G G→B B+→G G→B+ Source 38.1 46.3 36.5 41.3 The proposed method 42.1 52.6 39.7 46.1 Target 52.0 61.0 52.0 56.2

[0091] To solve the problem of insufficient robustness caused by the excessive dependence on limited labeled datasets in the prior art, a cross-domain facial action unit detection method based on text bridging proposed in the present invention has the following prominent key points:

[0092] 1) Multi-scale AU visual feature extraction: Different AUs cover different ranges of the human face area, that is, there are scale differences between different facial AUs. For example, AU6 (Cheek Raiser) covers a large facial area, while AU12 (LipCorner Puller) is limited to a smaller corner of the mouth area. To effectively represent the visual features of AUs at different scales, a frozen pre-trained CLIP visual encoder is used for multi-scale AU feature extraction. This encoder consists of four stages, corresponding to different receptive fields, and can extract visual features of AUs at different scales from shallow to deep, so as to comprehensively capture AU visual cues. To promote subsequent cross-modal interaction and fusion, an adapter composed of one-dimensional convolution and pooling layers is constructed for the visual features of AUs at each scale, unifying the multi-scale AU visual features extracted at different stages to the same spatial dimension.

[0093] 2) AU Text Feature Extraction: To utilize the rich semantic information with domain-invariant properties contained in the AU text description, it is encoded into high-quality AU text features. First, the AU text descriptions are sorted out from the Facial Action Coding System manual. These text descriptions not only contain rich emotional semantic information but also have domain invariance, and their semantics are not affected by specific visual interference factors such as age, lighting, and pose. Next, to learn high-quality text feature representations, a set of learnable prompt vectors are introduced. The prompt vectors are combined with the AU descriptions to form a text input sequence. Then, this sequence is input into the frozen pre-trained CLIP text encoder to obtain AU text features. Finally, contrastive learning between vision and text is performed on the source domain to align the source domain visual features to the text features, enhancing the AU discriminability of the visual features. At the same time, it lays a foundation for subsequent use of domain-invariant text as a bridge to achieve cross-domain adaptation of visual features.

[0094] 3) Construct a Graph Neural Network to Promote the Interaction and Fusion of Visual and Text Modalities: First, cross-modal attention is adopted, using the AU visual features as the query (Query), and the AU text features as the key (Key) and value (Value). This cross-modal attention uses the text to provide precise semantic guidance for the vision, effectively fusing semantic information, establishing semantic associations between vision and text, and obtaining more discriminative AU interaction features. Second, to further promote the deep fusion of visual and text modalities, a graph neural network with AU interaction features as nodes is constructed. This network realizes the dynamic update of nodes through a message passing mechanism. Each node continuously aggregates information from other nodes in multiple rounds of iteration, further deepening the deep fusion of multi-scale AU visual features and text information, thereby obtaining enhanced AU features with richer semantics. Finally, all enhanced AU feature representations are concatenated and fused using a fully connected layer to obtain the fused AU features for AU detection.

[0095] 4) Contrastive learning based on negative sample AU descriptions: To reduce the feature distribution differences between source domain images and target domain images, contrastive learning based on negative sample AU descriptions is adopted to align the visual features and text features of the target domain, indirectly reducing the feature distribution and semantic differences between the source domain and the target domain. First, the key verbs or adjectives in the original AU descriptions are replaced with words having opposite meanings using large language models (such as Chatgpt, DeepSeek), thereby generating several negative AU descriptions with opposite meanings for each AU. Then, with the help of the AU detection model trained in the first stage, the target domain images are inferred to obtain the activation probabilities of each AU. By setting a confidence threshold, AU "pseudo-labels" are generated for the target domain data to indicate the activation status of the corresponding AUs in the target domain images. For the target domain images, contrastive learning is implemented in combination with the pseudo-labels, AU visual features, the original AU descriptions, and the generated negative sample AU descriptions. This contrastive learning mechanism can promote the model to deeply understand the AU activation status described by the text, and align the visual features of the target domain with the AU text features in terms of features and semantics, establishing a more accurate correspondence between visual AUs and text descriptions. Since the source domain features and text features have been aligned in the first stage, and this stage also aligns the target domain features with the text features, the feature distribution differences between the source domain and the target domain are indirectly reduced, thereby improving the cross-domain generalization ability of the AU model.

[0096] 5) Cross-domain expression action unit detection method based on text bridging: As Figures 2 to 4 shown, the input data are AU descriptions, source domain and target domain data, through a two-stage training model. For the first stage: First, a frozen pre-trained CLIP visual encoder is used for multi-scale AU feature extraction. Then, the AU visual features of different scales are mapped to a unified spatial dimension through an adapter. Next, a set of learnable prompt vectors are introduced and combined with the AU text descriptions and input into the frozen pre-trained CLIP text encoder to obtain AU text features. At the same time, contrastive learning is implemented between vision and text in the source domain, thereby aligning the visual features and text features of the source domain. Further, full interaction and fusion, alignment between the visual and text modalities are implemented. First, cross-modal attention is used for the initial interaction between vision and text, and then the deep fusion of visual and text information is achieved through a graph neural network. Finally, the enhanced AU feature representations corresponding to different nodes are subjected to feature concatenation and fusion to obtain the fused AU features for expression action unit detection. The source domain AU classification loss and the source domain contrastive learning loss Perform model training. For the second stage: First, use a large language model (such as DeepSeek) to rewrite the original AU description to obtain negative sample AU descriptions. Then, input the target domain samples into the model of the first stage to obtain the pseudo-labels of the target domain samples. Next, through contrast learning based on negative samples on the target domain, align the visual features and text features of the target domain, thereby indirectly reducing the feature distribution difference between the source domain and the target domain and improving the AU detection performance of the model on the target domain. The second stage uses the source domain AU classification loss , the target domain AU classification loss , and the contrast learning loss based on negative samples to jointly optimize the model.

[0097] In summary, the present invention designs a cross-domain expression action unit detection method network based on text bridging. This network uses a frozen pre-trained CLIP visual encoder to extract AU visual features at different scales, and introduces a set of learnable prompt vectors. Combined with the AU text description, it is input into the frozen CLIP text encoder to obtain AU text features. At the same time, contrast learning is used to align the source domain features and text features to improve the discriminability of visual features and text features; cross-modal attention is used to achieve the initial interaction between vision and text to obtain the AU interaction feature representation; through the graph neural network, the deep interaction and fusion of text and visual information are further promoted to obtain the fused AU features; a contrast learning method based on negative samples is used to align the target domain data and text features, indirectly reducing the feature distribution difference between the source domain and the target domain. Finally, this method uses the domain invariance of text to align both the visual features of the source domain and the target domain with the text features, thereby reducing the feature distribution difference between the source domain and the target domain and significantly improving the generalization performance of the AU detection model and the accuracy on the target domain.

[0098] Based on the same technical solution, the present invention also proposes an electronic device, including:

[0099] A memory for storing a computer program;

[0100] A processor for implementing the steps of the above-mentioned cross-domain expression action unit detection method based on text bridging when executing the computer program.

[0101] Based on the same technical solution, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned cross-domain expression motion unit detection method based on text bridging are implemented. The computer-readable storage medium may include various media capable of storing program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0102] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memories (ROMs), programmable ROMs (PROMs), electrically programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), or flash memories. Volatile memories can include random access memories (RAMs) or external cache memories. By way of illustration and not limitation, RAMs are available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

Claims

1. A cross-domain expression motion unit detection method based on text bridging, characterized in that: include: Step 1, obtaining a multimodal expression motion unit AU dataset including an image dataset and a text dataset; Step 2: Build a cross-domain AU detection network; Step 3: divide the image dataset into a source domain dataset and a target domain dataset, and train and optimize the cross-domain AU detection network in combination with the text dataset to obtain a cross-domain AU detection model; Step 4: Use the cross-domain AU detection model to implement cross-domain AU detection.

2. The method according to claim 1, characterized in that The step of acquiring the multimodal expression motion unit AU data set comprises: Step 1.1: Obtain an image dataset D including face images and corresponding labels V ; Step 1.2: Based on the Facial Action Coding System (FACS) manual, collect text descriptions of AUs to obtain a text dataset D consisting of text descriptions of AUs. L ; Step 1.3, integrate image dataset D V And the text dataset D L , forming an AU dataset D including an image dataset and a text dataset.

3. The method according to claim 1, characterized in that The steps of constructing the cross-domain AU detection network include: Construct a visual encoder to extract features from image data, wherein the visual encoder is divided into multiple stages, each stage corresponds to a receptive field of different sizes, and obtains AU visual features of different scales output by each stage, and constructs an independent adapter for the AU visual features output by each stage to obtain AU visual features with the same spatial dimension; Build a word segmenter, take text data as input, and get the corresponding text vector; concatenate a set of learnable prompt vectors with the text vector to get the text input; Construct a text encoder to extract features from text input and obtain corresponding AU text features; Using AU visual features as queries and AU text features as keys and values, the cross-modal attention mechanism is used to calculate the interactive features; The interaction features of each AU are used as nodes of the graph neural network, and the similarity between any two nodes is defined as the edge of the graph neural network to construct the graph neural network; a multi-round message passing mechanism is used to update node features, and neighbor node information is aggregated through graph convolution to obtain the final feature representation of each node; the final feature representations of each node are spliced ​​to obtain the fused features; Construct an AU classifier, take the fused features as input, and get the corresponding prediction probability.

4. The method according to claim 3, characterized in that Any two nodes There are two directed edges between , directed edge Represents a slave node Point to Node , ; directed edge Represents a slave node Point to Node , .

5. The method according to claim 3, characterized in that: Based on the source domain dataset, a multi-label binary cross entropy loss function is used Supervised training of cross-domain AU detection network: , in, Represents the number of image data in the source domain dataset, represents the number of AUs, and They represent the source domain datasets. The image data The labels and predicted values ​​of AUs, Indicates AU in the source domain dataset The image data is activated. Indicates AU in the source domain dataset The image data is not activated.

6. The method according to claim 5, characterized in that Comparative learning between AU visual features and AU text features, constructing a contrast loss function , based on the source domain dataset and the text dataset, the joint multi-label binary cross entropy loss function Supervised training of cross-domain AU detection network: , , in, is the overall loss function of the cross-domain AU detection network, yes The weight of , , is the temperature hyperparameter, is the source domain dataset The splicing and fusion results of the AU visual features at each stage corresponding to the image data, It is AU text features.

7. The method according to claim 3, characterized in that Introduce contrastive learning of negative sample AU description and construct contrastive learning loss function , based on image datasets and text datasets, joint multi-label binary cross entropy loss function , supervised optimization of cross-domain AU detection network: , , , in, is the overall loss function of the cross-domain AU detection network, , They are , The weight of Represents the number of image data in the target domain dataset, represents the number of AUs, and They represent the target domain dataset. The image data The pseudo label and predicted value of AU, if but Indicates AU in the target domain dataset image data is activated if but Indicates AU in the target domain dataset The image data is not activated. , represents the confidence threshold, , is the temperature hyperparameter, is the target domain dataset obtained by graph neural network. The fusion features corresponding to the image data are It is AU text features corresponding to AUs, is extracted by the text encoder AU corresponds to The text features of negative AU text descriptions, is the number of negative AU text descriptions.

8. The method according to claim 1, characterized in that Training and optimizing the cross-domain AU detection network consists of two stages: In the first stage, the cross-domain AU detection network is trained using the source domain dataset and the text dataset; In the second stage, the cross-domain AU detection network is optimized using image datasets and text datasets.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. An electronic device comprising: Memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Facial action unit detection model construction method based on hierarchical feature alignment

    CN117576765A

  • Facial action unit detection model establishment method based on multi-task learning

    CN117765596A

  • Speaking face generation method and device, electronic equipment and storage medium

    CN116844215A

  • Skeleton action recognition method based on prompt type contrast learning and storage medium

    CN118865490A

  • Audio driven facial animation using machine learning

    CN119494894A

Cited By

  • Expression motion unit detection method based on double cross-modal attention

    CN120147358A

  • Expression Movement Unit Detection Method Based on Dual Cross-Modal Attention

    CN120147358B

  • Few-sample leather anomaly detection method based on domain confrontation and multi-scale fusion

    CN120375109A

  • Few-shot leather anomaly detection method based on domain adversarial and multi-scale fusion

    CN120375109B