Expression motion unit detection method based on double cross-modal attention

By adopting a dual cross-modal attention method in the detection of expression motion unit, combining multi-grained visual features and multi-level text features, the problem of insufficient robustness caused by relying on limited annotation data sets in the prior art is solved, and more efficient expression motion unit detection performance is achieved.

CN120147358AActive Publication Date: 2025-06-13JIANGSU SECOND NORMAL UNIVERSITY

Patent Information

Application Number
CN202510634232.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing expression motion unit detection methods rely too much on limited labeled data sets, resulting in insufficient robustness and difficulty in effectively dealing with individual differences and environmental interference.

Method used

The expression motion unit detection method based on dual cross-modal attention is adopted, and multi-grained visual features are extracted through visual encoder, local perceptual attention matrix and graph convolution network, and combined with multi-level text feature extraction and dual cross-modal attention mechanism, deep interaction between AU visual mode and text mode is achieved.

Benefits of technology

It significantly enhances the semantic fidelity and robustness of AU feature representation, improves the performance of expression motion unit detection, and can more effectively understand the complex semantic associations between visual and text modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147358A_ABST
    Figure CN120147358A_ABST
Patent Text Reader

Abstract

The invention discloses an expression motion unit detection method based on double cross-modal attention. The method comprises the following steps: obtaining refined multi-granularity visual feature representation through a visual encoder, a local perception attention matrix and a graph convolutional network; the method comprises the following steps: firstly, modeling semantic association among words in AU descriptions by utilizing a multi-level coding process, and then modeling sentence-level semantic association among different AU descriptions, so that rich semantic information in the AU descriptions is effectively mined, and the expression ability of AU text features is remarkably enhanced; a global and local collaborative dual cross-modal attention strategy is designed to realize deep interaction of visual and text modals, help the model to more comprehensively understand complex semantic association between the visual and text modals, and enhance AU feature representation. Finally, a powerful deep learning framework is constructed by combining the multi-granularity visual features and the multi-level text features with the synergistic effect of double cross-modal attention, and the expression motion unit detection performance is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of facial expression action unit detection, and mainly relates to a facial expression action unit detection method based on dual cross-modal attention. Background Technique

[0002] Facial expressions are an important way for humans to convey emotions. To comprehensively and objectively describe human facial expressions, psychologists such as Ekman et al. constructed the Facial Action Coding System (FACS). Based on anatomical principles, this system decomposes complex and diverse facial expressions into quantifiable local facial muscle movements, namely facial expression action units (Action Unit, AU). Through different combination methods, these action units can accurately depict various facial expressions, thereby revealing the emotional state of an individual. Therefore, facial expression action unit detection is a research focus in the fields of computer vision and affective computing. Its goal is to automatically identify the activation state of facial expression action units from images, and then analyze the emotional information conveyed by facial expressions. This automated facial expression detection technology shows broad application prospects in fields such as human-computer interaction and mental health assessment.

[0003] The facial expression action unit detection task faces challenges such as insufficient labeled data and significant individual differences. Specifically: First, the data annotation of facial expression action units requires professional knowledge and is time-consuming, resulting in a limited number of labeled samples available for model training, severely restricting the model performance; Second, there are differences in facial features among different individuals, and they are easily affected by factors such as lighting, pose, and occlusion, further increasing the difficulty of facial expression action unit detection.

[0004] Existing facial expression action unit detection methods can be mainly divided into two categories: "correlation learning-based" and "region learning-based", as follows:

[0005] 1) Correlation learning-based facial expression action unit detection methods:

[0006] Patent application for invention with publication numbers CN114842542A, CN116416667A, and CN117765596A, named "Facial Action Unit Recognition Method and Device Based on Adaptive Attention and Spatiotemporal Correlation", "Facial Action Unit Detection Method Based on Dynamic Correlation Information Embedding", and "Method for Establishing a Facial Action Unit Detection Model Based on Multi-Task Learning", respectively. Their main technical means include: constructing a graph attention network, a dynamically updated AU correlation map, and an adaptive spatiotemporal graph convolutional neural network to capture the dependence relationship between AUs. These methods use graph neural networks to model the correlation between AUs, and rely heavily on the label distribution in the dataset, with low model generalization ability.

[0007] 2) Facial Expression Action Unit Detection Method Based on Region Learning:

[0008] Patent application for invention with publication numbers CN117576765A, CN115862120A, and CN111626113A, titled "Method for Constructing a Facial Action Unit Detection Model Based on Hierarchical Feature Alignment", "Facial Action Unit Recognition Method and Device Decoupled by a Separable Variational Autoencoder", and "Facial Expression Recognition Method and Device Based on Facial Action Units". Their main technical means include enhancing the model's perception ability of key local information through technologies such as convolutional neural networks and attention mechanisms. Due to significant differences in facial features among different individuals, and the rapid and subtle movement changes in local facial regions, while these methods improve the ability to capture key local information, they are prone to introducing noise, resulting in a decline in detection accuracy. Summary of the Invention

[0009] To solve the problem of insufficient robustness in the prior art due to over-reliance on limited labeled datasets, the present invention provides a facial expression action unit detection method based on dual cross-modal attention.

[0010] To achieve the above objective, the solution of the present invention is:

[0011] A facial expression action unit detection method based on dual cross-modal attention, comprising:

[0012] Step 1, obtain a multi-modal facial expression action unit AU dataset D including image data and text data;

[0013] Step 2, construct a multi-modal AU detection network;

[0014] Step 3, divide the AU dataset D into a training set and a validation set, train and test the multi-modal AU detection network to obtain a multi-modal AU detection model;

[0015] Step 4, use the multi-modal AU detection model to achieve AU detection.

[0016] Preferably, the steps for obtaining the multi-modal facial expression action unit AU dataset include:

[0017] Step 1.1, obtain an image dataset V including facial images and corresponding labels;

[0018] Step 1.2, based on the Facial Action Coding System (FACS) manual, collect text descriptions of AUs to obtain a text dataset T composed of text descriptions of AUs;

[0019] Step 1.3, integrate the image dataset V and the text dataset T to form an AU dataset D including images and text data.

[0020] Preferably, the steps for constructing the multi-modal AU detection network include:

[0021] Construct a visual encoder to extract features from image data and obtain global visual features;

[0022] Construct N locally perceptive attention matrices with independent parameters, using the global visual features as input to obtain the corresponding N local visual features;

[0023] Construct an AU internal encoder, using text data as input to obtain word-level feature representations; perform a pooling operation on the word-level feature representations to obtain sentence-level feature representations; input the sentence-level feature representations into the AU interactive semantic encoder to obtain text feature representations;

[0024] Using the global visual features as queries, the text feature representations as keys and values, and calculate the global interaction features using the cross-modal attention mechanism;

[0025] Using the local visual features as queries, the text feature representations as keys and values, and calculate the local interaction features using the cross-modal attention mechanism;

[0026] Fuse the global interaction features, local interaction features, and local visual features to obtain fused features;

[0027] Construct an AU classifier, using the fused features as input to obtain the corresponding prediction probabilities.

[0028] Preferably, calculate the loss using a cosine similarity difference loss function:

[0029]

[0030] where I represents the identity matrix, represents the number of words in the word-level feature representation, represents the word-level feature representation.

[0031] Preferably, the cross-modal attention mechanism is defined as:

[0032] ,

[0033] where, is the modal feature as the query, is the modal feature as the key and value, , , is a learnable parameter matrix, is the scaling factor.

[0034] Preferably, the steps for constructing the multi-modal AU detection network further include:

[0035] Perform global average pooling operation on N local visual features to obtain a feature set;

[0036] Construct a graph neural network, use the local visual features of each AU as the nodes of the graph neural network, and define the cosine similarity between any two nodes as the edge of the graph neural network;

[0037] For each node of the graph neural network, select the K neighbor nodes with the largest similarity, and aggregate the neighbor node information through graph convolution to obtain a refined feature representation for each node;

[0038] Construct an AU classifier, use the refined feature representation of the nodes as the input, and obtain the corresponding prediction probability.

[0039] Preferably, use a weighted asymmetric loss function Calculate the loss:

[0040]

[0041] where, 、 and respectively represent the prediction probability, true value and loss weight of the th AU.

[0042] Preferably, The calculation formula of is: , where represents the occurrence frequency of the th AU in the training set, represents the occurrence frequency of the th AU in the training set.

[0043] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned expression action unit detection method based on dual cross-modal attention are realized.

[0044] The present invention also provides an electronic device, including:

[0045] A memory for storing a computer program;

[0046] A processor for realizing the steps of the above-mentioned expression action unit detection method based on dual cross-modal attention when executing the computer program.

[0047] Compared with the prior art, the significant advantages of the present invention are as follows: The present invention realizes the deep interaction between the AU visual modality and the text modality. In particular, refined multi-granularity visual feature representations are obtained through a visual encoder, a local perception attention matrix, and a graph convolutional network; by using a multi-level encoding process, the semantic associations between words in the AU description are first modeled, and then the sentence-level semantic associations between different AU descriptions are modeled, effectively mining the rich semantic information in the AU description and significantly enhancing the expressive power of the AU text features; a dual cross-modal attention strategy that coordinates global and local is designed to achieve deep interaction between the visual and text modalities, helping the model to more comprehensively understand the complex semantic associations between the visual and text modalities and enhancing the AU feature representation. Finally, by combining multi-granularity visual features, multi-level text features, and the synergistic effect of dual cross-modal attention, a powerful deep learning framework is constructed, effectively improving the performance of expression action unit detection. Brief Description of the Drawings

[0048] Figure 1 It is a flowchart of the expression action unit detection method based on dual cross-modal attention.

[0049] Figure 2 It is a schematic diagram of the multi-modal AU detection network structure, where (A) is the training stage and (B) is the testing stage. Detailed Embodiments

[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] As Figure 1 shown, the specific process of the expression action unit detection method based on dual cross-modal attention of the present invention is as follows:

[0052] 1. Data preparation stage:

[0053] 1.1 Collect the expression action unit dataset:

[0054] To construct an expression action unit (AU) dataset, first perform face detection and alignment processing on the video segments containing human faces in the dataset. Specifically, for each frame of the image, a face detection and alignment model is used for processing to ensure that the facial region in the image is accurately recognized and aligned. Subsequently, the processed facial images are cropped to a standard size of 256×256 pixels to unify the input size. According to the label file in the dataset, each image is annotated, thereby forming an AU dataset composed of image data V, providing training and validation data for subsequent AU detection tasks.

[0055] 1.2 Collection and Expansion of Facial Action Unit Descriptions:

[0056] Based on the Facial Action Coding System (FACS) manual, text descriptions of facial action units were collected and organized. These descriptions cover rich semantic information such as the facial regions, intensities, categories, and interdependencies related to AUs. However, the number of AU descriptions is scarce, with only one original AU description for each AU. To further expand and diversify the AU descriptions, combining the original AU descriptions with carefully designed prompts, large language models (e.g., DeepSeek, GPT-4) were used to generate 50 or more large model augmented AU descriptions for each AU, and the quality of the generated text was ensured through manual screening. Finally, the text data and the AU dataset were integrated to form a multi-modal AU dataset including image and text data .

[0057] 2. In the model design stage, the specific design of the two-stage model is as follows:

[0058] 2.1 The overall model is denoted as , including a visual encoder , an AU internal encoder , an AU interaction encoder , a graph convolutional network , a dual cross-modal attention module , and an AU classifier C. The model input is the multi-modal AU dataset , including image data and text data .

[0059] 2.2 Extract multi-granularity visual features. First, the visual encoder is used to extract features from the input image data , and the global visual feature is output. To accurately capture the subtle expressions in the local regions of the human face, parameter-independent local perception attention matrices are constructed, with each AU corresponding to one local perception attention matrix. The global visual feature is input into these local perception attention matrices to extract the local visual features corresponding to the respective AUs. This process can be formally expressed as:

[0060] .

[0061] Local visual features It is used to perceive the local area of the human face, laying a foundation for subsequent AU relationship modeling and cross-modal interaction.

[0062] 2.3 Construct an AU relationship modeling module based on the graph convolutional network . This module constructs a graph neural network to model the dependence relationship between different AU visual features, thereby enhancing the semantic expression ability of local visual features. Specifically, first, perform global average pooling operation on local visual features to obtain the feature set , where represents the local visual feature of the th AU. Secondly, take the local visual feature of each AU as the node of the graph neural network, and define the cosine similarity and between any two nodes as the edge of the graph neural network, thereby constructing the graph neural network. Then, for each node , calculate and the similarity of all other nodes , find the first nodes with the largest similarity, and form the neighbor node set of node , which can be expressed as . When node belongs to , then the connection relationship and between nodes is set to 1; otherwise, when node does not belong to , then is set to 0. Using the connection relationship between nodes, aggregate neighbor node information through graph convolution to obtain the refined feature representation of the th node. This process can be expressed as:

[0063]

[0064] where, is a non-linear activation function, and represent differentiable functions of the graph convolutional network layer, represents the connection relationship between nodes and , 1 indicates the existence of a connection, and 0 indicates no connection.

[0065] Input the node features into the AU classifier C to obtain the predicted probability of each AU . Specifically, an AU classifier is constructed for each AU, and the classifier consists of a learnable vector with the same dimension . Then the predicted probability of the -th AU can be expressed as:

[0066]

[0067] where ReLU is a non-linear activation function. Then, a weighted asymmetric loss function is used to calculate the loss:

[0068]

[0069] where , and represent the predicted probability, true value, and loss weight of the -th AU respectively. is used to alleviate the label imbalance problem in the dataset, and its calculation formula is: , where represents the occurrence frequency of the -th AU in the training set.

[0070] 2.4 Construct the AU internal encoder , the AU interaction encoder . The present invention first constructs an AU internal encoder for modeling the semantic relationship between words in the AU description. Input the AU description into the AU internal encoder to extract the word-level feature representation ={ }, where represents the number of words in the AU description, and each represents the feature vector of the word. To further enhance the model's ability to identify different AU text features, a cosine similarity difference loss function is introduced:

[0071]

[0072] where I represents the identity matrix. This loss function effectively enhances the specificity between AU text features by forcing the similarity matrix calculated from the text features to exhibit a diagonally dominant structure, that is, the values on the non-diagonal of the matrix approach 0, and prompts the model to learn AU description features with significant distinguishability. To obtain the sentence-level feature representation, a pooling operation is performed on the word-level features to obtain the sentence-level feature representation of each AU description , where represents the feature of the -th word. On this basis, to further integrate the information between different AU descriptions, the sentence-level feature is input into the AU interactive semantic encoder to fuse the semantic information from other AU descriptions and further enhance the text features of each AU, finally obtaining a robust text feature representation: ( ).

[0073] 2.5 Constructing a dual cross-modal attention module to achieve precise cross-modal interaction. This module consists of global and local cross-modal attention working together to model multi-granularity cross-modal interaction information and enhance the AU feature representation. Given the modal feature as the query, and the modal features as the key and value, the cross-modal attention mechanism can be defined as:

[0074]

[0075] where , , is a learnable parameter matrix, and is the scaling factor ( usually takes the dimension of the key-value ).

[0076] Global cross-modal attention is used to model the semantic dependence between global visual features and AU description features. Specifically, using the global visual feature as the query, and the text feature of the -th AU as the key and value, the global interaction feature is calculated using the cross-modal attention mechanism: . This design can eliminate the interference of other AU description information and focus on establishing the exclusive semantic association between global visual features and the current AU text features. While local cross-modal attention focuses on the interaction between the visual and text features of a single AU. Specifically, for the -th AU, using the local visual feature as the query, and the AU text feature as the key and value, the local interaction feature is calculated using the cross-modal attention mechanism: . The local cross-modal attention module is guided by local visual features to fuse the semantic information of the AU description text that is uniquely corresponding to the visual features of this AU, thereby achieving fine-grained cross-modal feature interaction and information query. Through the collaborative work of global and local cross-modal attention, multi-granularity cross-modal information interaction and fusion are achieved: the global module captures the semantic association between each AU description and the overall face, and this process can couple and utilize the mutual exclusion and co-occurrence relationships between AUs to achieve a preliminary query of the AU activation state; the local module focuses on the fine interaction between the specific visual features of the AU and the text features of the AU description. Through strict one-to-one input constraints, this process ensures that there is no mixing of other AU features during the cross-modal information fusion process, which helps to improve the semantic fidelity of the single AU feature representation and achieve an accurate query of the AU activation state.

[0077] 2.6 Fuse the global interaction features , local interaction features and local visual features to obtain the fused features for multi-modal expression action unit detection . Subsequently, use the AU classifier C described in Section 2.3 to calculate the predicted probability of the fused features of the -th AU, and use the weighted asymmetric loss function to calculate the detection task loss.

[0078] 2.7 Model training is divided into two stages: in the first stage, use the asymmetric loss function to specifically train the visual encoder , the graph neural network G and the AU classifier C to establish the preliminary visual feature extraction ability; on the basis of the first stage training, in the second stage, further introduce three modules: the AU internal encoder , the AU interaction encoder , and the dual cross-modal attention module . In this stage, use the overall loss function to train the model M, where represents the newly added differential loss function in this stage. In the second stage, jointly optimize the newly introduced three modules, the visual encoder and the AU classifier C. Update the model parameters by the gradient descent method. First, execute steps 2.1, 2.2, and 2.3 to complete the training of the first stage. Then execute steps 2.4, 2.5, and 2.6 to complete the training of the second stage until convergence. The parameter update strategy is as follows:

[0079]

[0080] Among them represents the learning rate.

[0081] The above steps 2.8 are unified into a two-stage deep neural network framework to achieve optimized training of the model.

[0082] 3. Model training stage:

[0083] 3.1 Divide the multi-modal sentiment recognition dataset obtained in step 1.1 into a training set and a validation set ;

[0084] 3.2 Input the training set into the network model designed in step 2 and use the batch random gradient descent method to train the model. There are two loss functions, namely the weighted asymmetric loss , and the dissimilarity loss . The validation set is used simultaneously during the training stage to verify the training effect of the model, that is, when the model obtains good AU detection results on the validation set and the performance no longer improves in subsequent training iterations, the model training converges and the training process stops;

[0085] 3.3 Finally, the model is obtained after training is completed.

[0086] 4. Model testing stage:

[0087] 4.1 The input data is a multi-modal sentiment recognition dataset D = {V, T} similar to that processed in step 1.1. A test set is divided from it , and the large model augmented AU descriptions in the text data are removed. The model used in the testing stage is the visual encoder in the model , the AU internal encoder , the AU interaction encoder , the dual cross-modal attention module , and the AU classifier .

[0088] 4.2 Input the test set into the model obtained in step 3.3 to obtain the test set The detection results of facial action units. Among them, the F1 scores of the facial action unit detection results of the BP4D and DISFA datasets are shown in Table 1 and Table 2 respectively. The detection results show that the multi-modal facial action unit detection method based on visual and text information proposed by the present invention is effective and has achieved good detection performance on both the BP4D and DISFA datasets.

[0089] Table 1 Detection Results of Facial Action Units in BP4D Dataset AU1 AU2 AU4 AU6 AU7 AU10 AU12 AU14 AU15 AU17 AU23 AU24 Avg 54.1 49.7 63.3 79.3 79.8 84.5 88.8 68.5 57.0 62.6 53.1 56.8 66.5 Table 2 Detection Results of Facial Action Units in DISFA Dataset AU1 AU2 AU4 AU6 AU9 AU12 AU25 AU26 Avg 55.7 58.4 75.4 51.0 56.5 74.8 93.9 63.8 66.2

[0090] To solve the problem of insufficient robustness caused by the over-reliance of existing technologies on limited labeled datasets, the multi-modal facial action unit detection method based on visual and text information proposed by the present invention has the following prominent key points:

[0091] 1) Multi-granularity visual feature extraction: First, a visual encoder is used to extract features from the input image to obtain global visual features containing overall facial information. Then, to accurately capture the subtle expressions in the local areas of the human face, a set of parameter-independent local perception attention matrices are introduced. Each attention matrix is designed for a specific AU and focuses on the local area related to the corresponding AU. By inputting the global visual features into these matrices, the local visual features corresponding to the AUs can be independently extracted. To further enhance the expressive power of these local visual features, a graph convolutional network is constructed to model the dependencies between AUs. In the graph neural network, the local visual features of each AU are regarded as nodes, and the similarity between the local visual features is regarded as edges. For each node, the top K neighbor nodes with the highest similarity to it are selected, and the information of these neighbor nodes is aggregated through graph convolution, so as to fully explore the mutual dependencies between AUs and finally obtain a refined representation of each node feature. In the above multi-granularity visual feature extraction process, the global visual features provide macroscopic information, while the local visual features focus on the microscopic details of the local areas of the face. The two complement each other and jointly describe the multi-granularity information of the face.

[0092] 2) Multi-level text feature extraction: To effectively learn robust AU text features, a multi-level text feature extraction strategy is adopted, gradually modeling from the word level to the sentence level to deeply mine the rich semantic information in AU descriptions. Since each AU has only one original description, direct learning may lead to insufficient expression of text features. Therefore, large language models (such as DeepSeek, GPT-4) are introduced to augment the data of the original AU descriptions, generate more descriptions for each AU, and ensure the text quality through manual screening, thereby enhancing the diversity of AU descriptions. Specifically, first, word-level features are extracted through the AU internal encoder, and then a pooling operation is performed to obtain sentence-level features. To enhance the specificity of text features, a differential loss function based on cosine similarity is introduced. This loss function ensures that the differences between AU text features are significant, and finally, a highly recognizable AU-specific semantic representation is constructed. Then, through the AU interaction encoder, text information from other AUs is further fused to strengthen the sentence-level features, improve the expression ability of text features, and effectively alleviate the problem of data sparsity in AU descriptions.

[0093] 3) Dual cross-modal attention to promote cross-modal interaction: A dual cross-modal attention strategy that combines global and local cooperation is adopted to deeply mine the semantic information in text descriptions, promote precise interaction between the visual modality and the text modality, and thus obtain a more detailed and robust feature representation. In global cross-modal attention, the global visual feature serves as the query (Query), and the text features of each facial action unit (AU) serve as the key (Key) and value (Value). By associating the visual information of the entire facial region with the semantic information of the corresponding AU description, the overall emotional expression of the face image is captured, and the global interaction feature is output. Local cross-modal attention uses the local visual feature of a specific AU as the query, and the text feature corresponding to the AU as the key and value, focusing on the text information related to the specific AU, precisely matching the local visual feature with the fine-grained semantic information to capture the details of the facial expression and query the activation state of the current AU, and finally output the local interaction feature. Through this collaborative and complementary dual cross-modal attention mechanism, the model can comprehensively understand the complex semantic associations between the visual and text modalities, effectively match the AU text definition with the global and local visual features to accurately identify the activation state of the AU.

[0094] 4) Facial action unit detection method based on dual cross-modal attention: As Figure 2As shown in (A) and (B) therein, a robust facial expression action unit (AU) detection is performed by a two-stage training model. The first stage mainly focuses on the extraction of multi-granularity visual features: First, the global visual features are extracted by a visual encoder; Then, a set of local perception attention matrices are used to extract the local visual features of each AU, and the dependencies between AUs are modeled by a graph convolutional network to enhance the expressive power of the local features. In this stage, a weighted asymmetric loss function \(L_{wa}\) is used to optimize the local visual features of AUs to obtain a refined feature representation. The second stage includes multi-level text feature extraction and cross-modal interaction with dual cross-modal attention: First, word-level text features are extracted by an AU internal encoder and sentence-level features are obtained by average pooling; Then, a difference loss function \(L_{dis}\) is introduced to enhance the specificity of the text features; Next, the semantic information of different AUs is fused by an AU interaction encoder to further strengthen the sentence-level features. Under the dual cross-modal attention strategy, the visual features are used as queries, and the AU text features are used as keys and values. The visual and text information is deeply mined in a global and local collaborative manner to achieve precise cross-modal interaction. Finally, the refined local visual features extracted in the first stage are fused with the local and global interaction features obtained in the second stage to generate the final multi-modal features, which are input into an AU classifier for the prediction and classification of AUs. The classification loss is calculated using the asymmetric loss \(L_{wa}\) in the first stage, and the model ensures the effective fusion of multi-modal features through multiple jointly acting loss functions, thereby improving the performance of facial expression action unit detection.

[0095] In summary, the multi-modal facial expression action unit detection model proposed by the present invention realizes the deep interaction between the AU visual modality and the text modality. In particular, refined multi-granularity visual feature representations are obtained through a visual encoder, local perception attention matrices, and a graph convolutional network; By using a multi-level encoding process, the semantic associations between words in the AU description are modeled first, and then the sentence-level semantic associations between different AU descriptions are modeled, effectively mining the rich semantic information in the AU description and significantly enhancing the expressive power of the AU text features; A dual cross-modal attention strategy of global and local collaboration is designed to achieve deep cross-modal interaction between vision and text, helping the model to more comprehensively understand the complex semantic associations between the visual and text modalities and enhancing the AU feature representation. Finally, by combining multi-granularity visual features, multi-level text features, and the collaborative effect of dual cross-modal attention, a powerful deep learning framework is constructed to effectively improve the performance of facial expression action unit detection.

[0096] Based on the same technical solution, the present invention also proposes an electronic device, including:

[0097] A memory for storing a computer program;

[0098] A processor, configured to implement the steps of the above-described facial action unit detection method based on dual cross-modal attention when executing the computer program.

[0099] Based on the same technical solution, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-described facial action unit detection method based on dual cross-modal attention. The computer-readable storage medium may include: various media capable of storing program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0100] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. The non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

Claims

1. The facial expression motion unit detection method based on dual cross-modal attention is characterized by: include: Step 1, obtaining a multimodal expression motion unit AU data set D including image data and text data; Step 2: construct a multimodal AU detection network; Step 3, divide the AU dataset D into a training set and a validation set, train and test the multimodal AU detection network, and obtain a multimodal AU detection model; Step 4: Use the multimodal AU detection model to implement AU detection.

2. The method according to claim 1, characterized in that The step of acquiring the multimodal expression motion unit AU data set comprises: Step 1.1, obtaining an image dataset V including facial images and corresponding labels; Step 1.2, based on the Facial Action Coding System FACS manual, collect text descriptions of AUs to obtain a text dataset T consisting of text descriptions of AUs; Step 1.3, integrate the image dataset V and the text dataset T to form an AU dataset D including image and text data.

3. The method according to claim 1, characterized in that The steps of constructing the multimodal AU detection network include: Construct a visual encoder to extract features from image data and obtain global visual features; Construct N local perception attention matrices with independent parameters, take global visual features as input, and obtain corresponding N local visual features; Construct an AU internal encoder, take text data as input, and obtain word-level feature representation; perform pooling operation on the word-level feature representation to obtain sentence-level feature representation; input the sentence-level feature representation into the AU interactive semantic encoder to obtain text feature representation; Using global visual features as queries and text feature representations as keys and values, the cross-modal attention mechanism is used to calculate global interaction features. Using local visual features as queries and text feature representations as keys and values, the cross-modal attention mechanism is used to calculate local interaction features. The global interaction features, the local interaction features and the local visual features are fused to obtain fused features; Construct an AU classifier, take the fused features as input, and get the corresponding prediction probability.

4. The method according to claim 3, characterized in that The cross-modal attention mechanism is defined as: , in, is the modal feature used as the query, are modal features as keys and values, , , is the learnable parameter matrix, is the scaling factor.

5. The method according to claim 3, characterized in that: The step of constructing the multimodal AU detection network also includes: Perform a global average pooling operation on N local visual features to obtain a feature set; Construct a graph neural network, use the local visual features of each AU as the nodes of the graph neural network, and define the cosine similarity between any two nodes as the edge of the graph neural network; For each node in the graph neural network, select the K neighbor nodes with the greatest similarity, aggregate the neighbor node information through graph convolution, and obtain the refined feature representation of each node; Construct an AU classifier, take the refined feature representation of the node as input, and obtain the corresponding prediction probability.

6. The method according to claim 5, characterized in that For each node in the graph neural network, select the K neighbor nodes with the greatest similarity, and aggregate the neighbor node information through graph convolution to obtain a refined feature representation of each node, including: For each node of the graph neural network, select the K neighbor nodes with the greatest similarity to form the neighbor node set of each node; By using the connection relationship between nodes, the neighbor node information is aggregated through graph convolution to obtain a refined feature representation of each node; the connection relationship between nodes is expressed as: , In the formula, For Node and The connection relationship between For Node The set of neighbor nodes.

7. The method according to claim 3, characterized in that Using a weighted asymmetric loss function Calculate the loss of the multimodal AU detection network: , in, , and Respectively represent The predicted probability, true value and loss weight of each AU.

8. The method according to claim 7, characterized in that The calculation formula is: ,in Indicates the first The frequency of occurrence of AUs, Indicates the first The occurrence frequency of AUs.

9. The method according to claim 5, characterized in that Using a weighted asymmetric loss function Calculate the loss of visual encoder, graph neural network and AU classifier; introduce the difference loss function based on cosine similarity , joint weighted asymmetric loss function , calculate the loss of the multimodal AU detection network : , , , in, , and Respectively represent The predicted probability, true value and loss weight of each AU, I represents the unit matrix, represents the number of words in the word-level feature representation, Represents word-level feature representation.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Facial expression recognition method and device based on facial action unit

    CN111626113A

  • Facial action unit recognition method and device based on adaptive attention and space-time correlation

    CN114842542A

  • Facial action unit recognition method and equipment for separable variational auto-encoder decoupling

    CN115862120A

  • Facial action unit detection method based on dynamic associated information embedding

    CN116416667A

  • Facial action unit detection model construction method based on hierarchical feature alignment

    CN117576765A

Cited By

  • Decision-making method and device based on multi-modal collaborative optimization, equipment and medium

    CN120822176A

  • Decision-making method and device based on multi-modal collaborative optimization, equipment and medium

    CN120822176B

  • Textile dyeing defect intelligent detection method and system based on AI vision

    CN121504871A

  • Cross-identity expression motion unit detection method based on probability prototype double calibration

    CN121884421A