A facial expression recognition method based on multi-stream attention interaction

By introducing a multi-flow attention interaction mechanism into facial expression recognition technology, combining global feature interaction and channel interaction strategies, the noise labels are dynamically updated, and the recognition accuracy problem under the influence of noise labels in the existing technology is solved, achieving higher accuracy and robustness.

CN119229510BActive Publication Date: 2025-05-16SHANDONG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411773421.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-16
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

When existing facial expression recognition technology processes data containing noisy tags, it is difficult to capture nuances between expressions, resulting in a decrease in model overfitting and generalization capabilities, affecting recognition accuracy.

Method used

A facial expression recognition method based on multi-stream attention interaction is proposed. Through global feature interaction and channel interaction strategy, effective interaction between facial image branches and facial key point branches are ensured, and dynamic update of noise tags is guided through dynamic tag adjustment module.

Benefits of technology

It improves the accuracy and robustness of facial expression recognition, effectively suppresses the influence of noise labels, reduces the risk of model overfitting, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229510B_ABST
    Figure CN119229510B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of facial expression recognition technology, and specifically discloses a facial expression recognition method based on multi-stream attention interaction. The present invention builds a facial expression recognition model based on multi-stream attention interaction, which includes modules such as a global feature interaction module, a channel interaction module, and a dynamic label adjustment module. Through the provided global feature interaction module and channel interaction module, the features between the facial image branch and the facial key point branch can be fully integrated, thereby achieving efficient information interaction and improving recognition accuracy. The dynamic label adjustment module, with the guidance of weight distribution, realizes the effective iterative update of noise labels, ensuring the full use of all sample data. The present invention can effectively suppress the influence of noise labels and improve the accuracy and robustness of facial expression recognition through the designed global feature interaction module, channel interaction module and dynamic label adjustment module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of facial expression recognition and relates to a facial expression recognition method based on multi-stream attention interaction. Background Art

[0002] Facial Expression Recognition (FER) usually uses an end-to-end supervised learning method. The effectiveness of this method depends on large-scale and accurately annotated datasets. However, the ambiguity of facial expressions, low image quality, and subjectivity of annotators make it difficult to achieve high-precision annotation of datasets obtained from the Internet. Therefore, most existing facial expression recognition methods refer to inaccurately annotated labels as noisy labels, and improve data quality through label correction so that more accurate dataset labels can be used in subsequent training processes. In addition, the similarity between various facial expressions is very high. For example, for a frowning facial expression, it may be appropriate to interpret it as sad or angry.

[0003] Existing noisy label processing methods have difficulty capturing these subtle differences. Therefore, developing more robust facial expression recognition technology is of great significance both in theory and in practical applications. With the help of large-scale datasets for end-to-end supervised learning, facial expression recognition technology has made remarkable achievements. However, the existence of noisy labels has become a major challenge, which may cause model overfitting and weaken its generalization ability, thereby affecting the model recognition accuracy.

[0004] For end-to-end supervised learning methods with noisy labels, current research mainly adopts two major strategies: one is to estimate the potential true labels and readjust the weights of training samples; the other is to implement sample selection, that is, to select correctly labeled samples from noisy datasets to enhance the robustness of the model. However, this sample selection method may mistakenly discard some clean samples when discarding noisy samples, thereby losing useful information of the data.

[0005] Subtle changes in facial expressions are enough to classify them into different expression categories. For example, while the overall appearance remains the same, just a slight upward or downward movement of the mouth is enough to define two completely different emotional expressions.

[0006] However, current noisy label facial expression recognition techniques cannot cope well with fine-grained differences between expressions, which not only hinders the model training process but also reduces the accuracy of facial expression recognition. In addition, existing sample selection methods have the risk of mistakenly deleting clean samples when dealing with noisy samples, thus losing useful information of the data. Summary of the invention

[0007] The purpose of the present invention is to propose a facial expression recognition method based on multi-stream attention interaction, in which a global feature interaction and channel interaction strategy is proposed to ensure that effective interaction can occur between the facial image branch and the facial key point branch, and through a clever weight distribution mechanism, the dynamic update process of the noise label is guided, thereby further improving the accuracy and robustness of model recognition, and thus helping to improve the accuracy and robustness of facial expression recognition.

[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A multi-stream attention interaction facial expression recognition method comprises the following steps: Step 1. Acquire facial image data and adjust it to a preset size.

[0009] Step 2. Build a facial expression recognition model based on multi-stream attention interaction, which includes a multi-branch neural network module, a global feature interaction module, a channel interaction module, and a dynamic label adjustment module.

[0010] The multi-branch neural network module is used to extract features from the acquired facial image data, including extracting facial features using a backbone network and extracting facial key point features using a face key point detector.

[0011] The global feature interaction module is used to perform global information fusion on the extracted features; wherein the extracted facial features and facial key point features are respectively linearly transformed and mapped into query vectors, key vectors and value vectors.

[0012] The query vector is exchanged between the extracted facial features and facial key point features to obtain a dual-branch feature.

[0013] The channel interaction module is used to perform channel-level information fusion on the dual-branch features to obtain the final features.

[0014] Specifically, the dual-branch features are first cascaded to obtain the fused features.

[0015] At the same time, a multi-layer perceptron is introduced and combined with the softmax function to jointly learn the weight vector, which is used to re-weight the facial features and facial key point features on the channel, and the two channel adaptive weights are used to fuse the features to obtain the final features.

[0016] The dynamic label adjustment module is used to reweight the final features and update the noisy labels.

[0017] By performing full connection and softmax normalization operations on the final feature map, the expression category probability is obtained and output.

[0018] Step 3. Use the facial image data obtained in step 1 to train the facial expression recognition model based on multi-stream attention interaction built in step 2, and use the trained facial expression recognition model to perform facial expression recognition.

[0019] The present invention has the following advantages: As mentioned above, the present invention relates to a multi-stream attention interaction facial expression recognition method for processing facial expression recognition tasks under noise labels. Specifically, the present invention builds a facial expression recognition model based on multi-stream attention interaction, which includes modules such as a global feature interaction module, a channel interaction module, and a dynamic label adjustment module. The global feature interaction module and the channel interaction module provided by the present invention can fully integrate the features between the facial image branch and the facial key point branch, realize efficient information interaction, effectively reduce the similarity between categories, and further improve the recognition accuracy. The dynamic label adjustment module provided by the present invention is used to supervise the network to learn meaningful weights. The contribution of each sample to the training is evaluated based on the reliability, importance and noise level of the sample, and then each sample is weighted, which provides a powerful reference for subsequent label updates. The dynamic label adjustment module uses the guidance of weight allocation to achieve effective iterative updates of noise labels and ensure full utilization of all sample data. The present invention can suppress the influence of noise labels and improve the accuracy and robustness of facial expression recognition through the designed global feature interaction module, channel interaction module and dynamic label adjustment module. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Flow chart of the facial expression recognition method with multi-stream attention interaction in an embodiment of the present invention.

[0021] Figure 2 This is a diagram of the overall network structure of facial expression recognition based on multi-stream attention interaction in an embodiment of the present invention.

[0022] Figure 3 4 is a network structure diagram of the global feature interaction module in an embodiment of the present invention.

[0023] Figure 4 4 is a network structure diagram of a channel interaction module in an embodiment of the present invention.

[0024] Figure 5 FIG. 4 is a network structure diagram of a dynamic label adjustment module in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0026] Example 1

[0027] This embodiment 1 describes a facial expression recognition method based on multi-stream attention interaction, which aims to deeply mine fine-grained facial features while effectively reducing the interference of noise labels on the recognition process. Specifically, a facial expression recognition model is built in the present invention. The model architecture is based on multiple branches, specifically covering two major components: facial image branch and facial key point branch. The facial image branch is responsible for extracting global features, and the facial key point branch is responsible for accurately locating facial salient areas. This information serves as a guide for fine-grained feature attention, which helps to reduce similarity confusion between different categories. Furthermore, by designing three core modules, namely the global feature interaction module, the channel interaction module, and the dynamic label adjustment module, it is ensured that the facial image branch and the facial key point branch can interact effectively, and through a clever weight distribution mechanism, the dynamic update process of the noise label is guided, thereby further improving the accuracy and robustness of recognition.

[0028] like Figure 1 As shown, the facial expression recognition method with multi-stream attention interaction in this embodiment includes the following steps: Step 1. Acquire facial image data and adjust it to a preset size.

[0029] The RAF-DB dataset was selected to perform detailed category classification on the collected facial images, perform face alignment, face normalization and data enhancement preprocessing on the images, and divide them into training and test sets.

[0030] Face alignment refers to adjusting the detected face to a standard posture, including face detection and feature point positioning operations. Face normalization refers to standardizing the pixel values ​​of face images to eliminate the effects of lighting, posture changes, and occlusion, reduce image differences under different environmental conditions, and make it easier for the model to capture important features of the face. Data augmentation operations include image rotation, scaling, cropping, etc., to increase data diversity and improve the model's adaptability to new samples.

[0031] Step 2. Build a facial expression recognition model based on multi-stream attention interaction, such as Figure 2 As shown, the model includes a multi-branch neural network module, a global feature interaction module, a channel interaction module, and a dynamic label adjustment module.

[0032] The multi-branch neural network module is used to extract features from the acquired facial image data, including extracting facial features using a backbone network and extracting facial key point features using a face key point detector.

[0033] The global feature interaction module is used to perform global information fusion on the extracted features, wherein the extracted facial features and facial key point features are linearly transformed and mapped into query vectors, key vectors, and value vectors.

[0034] The query vector is exchanged between the extracted facial features and facial key point features to obtain a dual-branch feature.

[0035] The channel interaction module is used to perform channel-level information fusion on the dual-branch features to obtain the final features.

[0036] Specifically, the dual-branch features are first cascaded to obtain the fused features.

[0037] At the same time, a multi-layer perceptron is introduced and combined with the softmax function to jointly learn the weight vector, which is used to re-weight the facial features and facial key point features on the channel, and the two channel adaptive weights are used to fuse the features to obtain the final features.

[0038] The dynamic label adjustment module is used to reweight the final features and update the noisy labels.

[0039] By performing full connection and softmax normalization operations on the final feature map, the node values ​​are converted into probability values, the expression category probability is obtained and output, and the prediction result of facial expression recognition is obtained.

[0040] In this embodiment, the backbone network adopts a residual network (Residual Network, referred to as ResNet), and the face key point detector adopts a pre-trained face key point detector (Mobile Face Network, referred to as MobileFaceNet).

[0041] The overall processing flow of the facial expression recognition model based on multi-stream attention interaction is as follows: First, for a given batch of facial images with noisy labels, ResNet is used to extract the facial features of the images, and MobileFaceNet is used to accurately obtain the facial key point features to further enrich the representation of facial information.

[0042] Next, the global feature interaction module and the channel interaction module are introduced. These two modules can comprehensively and efficiently fuse the feature information from the two independent branches of the backbone network and the face key point detector.

[0043] In addition, in order to deal with the impact of noisy labels on recognition performance, the present invention designs a dynamic label adjustment module, which captures the importance of samples by reallocating weights to samples.

[0044] Then, these weights are arranged in descending order, and the samples are divided into two subsets according to a certain ratio: subset I contains high-weight samples, i.e., clean samples; subset II contains low-weight samples, i.e., noise samples.

[0045] In order to regularize the two subsets, a preset threshold for dividing the subsets is imposed between the average weights of the two subsets, and the regional control loss (RC-Loss) is used to implement this regularization process.

[0046] This regularization step aims to further emphasize clean and reliable samples and suppress those uncertain ones.

[0047] Next, the labels of low-weight samples are updated. The specific updating process is as follows: by comparing the difference between the maximum predicted probability and the probability of the given label, the samples in subset II are relabeled.

[0048] Specifically, if the maximum predicted probability of a sample is higher than the probability of its given label by more than a preset threshold, then the sample will be assigned a new pseudo-label, which is the category corresponding to the maximum predicted probability.

[0049] like Figure 3 The network structure diagram of the global feature interaction module is shown. The core purpose of the global feature interaction module proposed in the present invention is to effectively capture the global feature correlation between multi-branch data, and it is significantly different from the traditional non-local block in design. The global feature interaction module effectively models the global feature correlation between facial images and facial key points by introducing an innovative attention mechanism. The global feature interaction module comprehensively captures the global correlation between any positions in the input image by constructing multiple non-local attention layers operating in parallel. Specifically, the input feature map X with a dimension of C×H×W is processed, where C represents the channel dimension, and H and W represent the height and width of the image, respectively. First, the input feature map X is mapped to a query vector Q, a key vector K, and a value vector V through three precisely designed linear transformations. Subsequently, the relationship between any two positions in the feature map is encoded using the Q, K, and V triples, so as to deeply mine and capture global context information. The output value of this process is accurately calculated by the weighted sum of the values, where the weight of each value is dynamically assigned according to the similarity between the query vector Q and the corresponding key vector K.

[0050] The input features of the global feature interaction module in this embodiment are not from a single source, but are integrated from multiple different branches. In the processing flow of this module, the Key vector and the Value vector are carefully designed to be extracted from the features of the same branch, while the corresponding Query vector comes from another different branch with unique information.

[0051] Specifically, the processing process of the global feature interaction module is as follows: first, the facial features extracted by the backbone network and the facial key point features extracted using the face key point detector are input into the global feature interaction module, and are defined as input feature one and input feature two respectively.

[0052] Input feature 1 and input feature 2 are linearly transformed and mapped into query vector, key vector and value vector respectively.

[0053] Define the input features: the query vector, key vector and value vector mapped by linear transformation are , and , the query vector, key vector and value vector of input feature 2 after linear transformation are , and .

[0054] Then, the query vector, key vector, and value vector triplet are used to encode the relationship between any two positions in the feature map of each processing branch, so as to deeply mine and capture the global context information.

[0055] For the processing branch where the input feature is located, the key vector and value vector are respectively and , and the query vector is the query vector obtained by linear transformation from the processing branch where the input feature 2 is located .

[0056] For the processing branch where the input feature 2 is located, the key vector and value vector are respectively and , and the query vector is obtained by linear transformation from the processing branch where the input feature 1 is located .

[0057] For the processing branches where input feature 1 and input feature 2 are located, the processing results are feature maps , .

[0058] .

[0059] .

[0060] Among them, a represents facial features, b represents facial key point features, is the normalized scale factor, Represents the embedding dimension, which is composed of feature maps , From the calculation formula, we can see that and Swap between facial features and keypoint features.

[0061] Finally, the feature maps are respectively transformed through a 1×1 convolutional layer , The channel dimensions of are restored to the initial channel dimension C.

[0062] Through the above processing of the global feature interaction module, the facial features are endowed with the salient areas provided by the key point features, and on the other hand, the key point features are provided with global information from the facial features.

[0063] The global feature interaction module integrates and calibrates the information from these two branches to achieve effective fusion and alignment of information.

[0064] In addition, in view of the limitations of non-local blocks widely used in various visual tasks, that is, only focusing on the spatial correlation of features and ignoring the feature recalibration problem at the channel level, this paper proposes a novel channel interaction module, such as Figure 4 As shown in the figure, it aims to efficiently fuse multi-branch features and accurately capture the interactions and influences between channels.

[0065] like Figure 4 As shown, the processing process of the channel interaction module in this embodiment is as follows: first, the dual-branch features output by the global feature interaction module, namely, facial features and facial key point features, are spliced ​​to achieve preliminary fusion of features and generate spliced ​​features. .

[0066] Next, a multi-layer perceptron (Multi-Layer Perceptron, referred to as ), Multi-layer Perceptron Combined with the softmax function, it constitutes a learning mechanism for learning and generating weight vectors .

[0067] .

[0068] in, and is the channel adaptive weight learned by the fusion mechanism, is the channel adaptive weight of facial features, is the channel adaptive weight of facial key point features.

[0069] In order to maximize the utilization efficiency of aggregated information and effectively reduce the noise and redundant components in the features, this paper adopts an innovative feature fusion strategy, using the designed two-channel adaptive weights to fuse the corresponding features. The output result of this fusion process is the final feature , the specific calculation steps and formula are as follows:

[0070] .

[0071] in, and They represent feature one and feature two of the input global feature interaction module, namely the facial features extracted by the backbone network and the facial key point features extracted using the face key point detector.

[0072] The channel interaction module improves the model's sensitivity to information features on important channels and enhances the final feature map.

[0073] In addition, the present invention also designs a dynamic label adjustment module, such as Figure 5 As shown in the figure, the overall concept of the dynamic label adjustment module is to first re-weight the samples and accurately divide the samples into two subsets, namely the clean sample subset and the noise sample subset, based on the evaluation results. Furthermore, the regional control loss is introduced to achieve effective regularization of the grouping. Then, by comparing the differences between different prediction probabilities, the noise labels can be accurately identified and updated.

[0074] Specifically, the process of the dynamic label adjustment module to reweight the final features and update the noisy labels is as follows: Step I. Use a linear fully connected layer and a sigmoid activation function to obtain the importance weight of each sample, and divide them into two subsets, namely, a clean sample subset and a noisy sample subset, and use the regional control loss to regularize the importance weight.

[0075] The sample in step I refers to the final feature output by the channel interaction module The samples in .

[0076] First, all samples are weighted and evaluated, and the contribution of each sample to the training process can be accurately measured through the evaluation. Specifically, the present invention is expected to identify samples with high accuracy and assign them higher weights so that their reliable information can be fully utilized during the training process. On the contrary, for samples with higher uncertainty, their weights will be automatically reduced, thereby reducing their potential interference with the training process.

[0077] Let the final feature output of the channel interaction module be , contains the facial features of N images.

[0078] The reweighting part includes a linear fully connected layer and a sigmoid activation function to obtain the final features after fusion. As input, and output a corresponding importance weight for each feature vector, expressed as:

[0079] .

[0080] in, is the importance weight of the i-th sample, is the parameter of the fully connected layer for importance attention, represents the sigmoid activation function, Represents the final feature The i-th sample in .

[0081] In this way, the model can effectively utilize high-quality samples while reducing the interference caused by noisy labels.

[0082] In order to further strengthen the constraints on the importance of uncertain samples, the present invention introduces a regional control loss, and the subset division strategy ensures that the average weight of the high-weight sample subset can exceed the average weight of the low-weight sample subset by a preset threshold.

[0083] Among them, the high-weight sample subset is the clean sample subset, and the low-weight sample subset is the noise sample subset.

[0084] To achieve the above objectives, the present invention defines the regional control loss (RC-Loss), and its calculation formula is as follows:

[0085] .

[0086] .

[0087] .

[0088] in, represents the regional control loss, is the threshold for dividing subsets, and They are A high-weighted subset of samples and The average value of a low-weighted subset of samples.

[0089] According to the pre-set division ratio , the samples are divided into two subsets: a clean sample subset and a noise sample subset.

[0090] Step II. Based on the softmax function prediction probability, for each sample in the low-weight sample subset, compare the difference between its maximum prediction probability and the prediction probability corresponding to the original given label to achieve accurate identification and update of the noise label.

[0091] If the maximum predicted probability of a sample is higher than the maximum predicted probability of a given label by a preset threshold, a new pseudo-label will be assigned to the sample, which is the category corresponding to the maximum predicted probability.

[0092] Specifically, this step updates the noise labels of the low-weight sample subset; this step focuses on the samples in the low-weight sample subset and performs label update operations based on their softmax prediction probabilities. The process is as follows: For each sample, the dynamic label adjustment module compares its maximum prediction probability with the prediction probability corresponding to the original given label; if the maximum prediction probability of a sample exceeds a preset threshold of the maximum prediction probability of the given label, a new pseudo-label will be assigned to the sample, which is the category corresponding to the maximum prediction probability.

[0093] A label update is defined as:

[0094] .

[0095] in, is the new pseudo label, is the threshold for label update, is the maximum predicted probability, is the predicted probability of a given label, and They are the original given label and the label corresponding to the maximum predicted probability.

[0096] The present invention utilizes a multi-branch structure to comprehensively fuse feature information from two independent branches, effectively suppresses and adjusts noise labels, reduces the impact of noise labels on network training, and further effectively learns robust facial expression features with noise labels.

[0097] Step 3. Use the facial image data obtained in step 1 to train the facial expression recognition model based on multi-stream attention interaction built in step 2, and use the trained facial expression recognition model to perform facial expression recognition.

[0098] During model training, weighted cross entropy loss is used to handle noisy labels in multi-classification problems.

[0099] The weighted cross entropy loss is defined using the following formula: :

[0100] .

[0101] in, is the true category corresponding to the i-th sample The weight vector of is the weight vector of the jth category, represents the total number of categories, Represents the total number of samples.

[0102] In model training, the total loss function for: ,in For the trade-off ratio.

[0103] By minimizing the total loss function of the model, the facial expression recognition model is continuously trained until the model reaches the set number of training times or the model accuracy meets the accuracy requirements on the test set, that is, the training is completed.

[0104] After the model is trained, the trained facial expression recognition model is used for facial expression recognition. Specifically, the facial image to be recognized is processed according to step 1, input into the trained model, and the facial expression recognition result is output.

[0105] The method of the present invention introduces a global feature interaction module, a channel interaction module and a dynamic label adjustment module into the constructed model. Through their cooperation, the negative impact of noise labels on the network training process and generalization ability is effectively reduced, thereby improving the accuracy of facial expression recognition. The method is suitable for scenarios such as human-computer interaction, security monitoring, medical treatment, and education.

[0106] In addition, in order to verify the effectiveness of the method proposed in the present invention, experiments were carried out on the representative public dataset RAF-DB, and the method of the present invention was compared with the two current mainstream methods RAN and DMUE.

[0107] Table 1 Comparison between the method of the present invention and the mainstream method

[0108]

[0109] The experimental results in Table 1 show that, compared with RAN and DMUE, the method of the present invention exhibits excellent performance in dealing with synthetic noise and noise label problems in the real world. Whether in terms of recognition accuracy, robustness or generalization ability, the method of the present invention is significantly better than the comparison method, and is particularly suitable for scenarios such as human-computer interaction, security monitoring, and medical education.

[0110] The method of the present invention solves the problem of noisy labels in facial expression recognition well. By introducing a global feature interaction module, a channel interaction module and a dynamic label adjustment module into the model, the negative impact of noisy labels on the network training process and generalization ability is significantly reduced. The problem that the current noisy label facial expression recognition technology cannot well cope with the fine-grained differences between expressions and the shortcomings of the existing sample selection method is successfully solved, and the accuracy of facial expression recognition is further improved.

[0111] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.

Claims

1. A multi-stream attention interaction facial expression recognition method, characterized in that: The steps include: Step 1. Obtain facial image data and adjust it to a preset size; Step 2. Build a facial expression recognition model based on multi-stream attention interaction, which includes a multi-branch neural network module, a global feature interaction module, a channel interaction module, and a dynamic label adjustment module; The multi-branch neural network module is used to extract features from the acquired facial image data, including extracting facial features using a backbone network and extracting facial key point features using a facial key point detector; The global feature interaction module is used to perform global information fusion on the extracted features; wherein the extracted facial features and facial key point features are respectively linearly transformed and mapped into query vectors, key vectors and value vectors; The query vector is exchanged between the extracted facial features and the facial key point features to obtain a dual-branch feature; The channel interaction module is used to perform channel-level information fusion on the dual-branch features to obtain the final features; First, the dual-branch features output by the global feature interaction module, namely the facial features and the facial key point features, are spliced ​​to achieve the initial fusion of features and generate the spliced ​​feature F c ; Next, a multi-layer perceptron layer MLP is introduced. The multi-layer perceptron layer MLP is combined with the softmax function to form a learning mechanism for learning and generating the weight vector w a ,w b :[w a ,w b ]=softmax(MLP(F c )); Among them, w a and w b is the channel adaptive weight learned by the fusion mechanism, w a is the channel adaptive weight of facial features, w b is the channel adaptive weight of facial key point features; The corresponding features are fused using the designed two-channel adaptive weights. The output of this fusion process is the final feature F fuse , the calculation formula is as follows: F fuse =w a *f1+w b *f2; Among them, f1 and f2 represent the feature 1 and feature 2 of the input global feature interaction module, namely the facial features extracted by the backbone network and the facial key point features extracted using the face key point detector; The dynamic label adjustment module is used to re-weight the final features and update the noisy labels; By performing full connection and softmax normalization operations on the final feature map, the expression category probability is obtained and output; Step 3. Use the facial image data obtained in step 1 to train the facial expression recognition model based on multi-stream attention interaction built in step 2, and use the trained facial expression recognition model to perform facial expression recognition.

2. The multi-stream attention interactive facial expression recognition method according to claim 1, characterized in that: The backbone network adopts a ResNet network, and the face key point detector adopts a pre-trained face key point detector.

3. The multi-stream attention interactive facial expression recognition method according to claim 1, characterized in that: The processing process of the global feature interaction module is as follows: Firstly, the facial features extracted by the backbone network and the facial key point features extracted by the face key point detector are input into the global feature interaction module, and they are defined as input feature 1 and input feature 2 respectively; Input feature 1 and input feature 2 are linearly transformed and mapped into query vector, key vector and value vector respectively; Define the input features: the query vector, key vector and value vector after linear transformation mapping are Q a , K a and V a The query vector, key vector, and value vector of input feature 2 after linear transformation are Q b , K b and V b ; Then, the query vector, key vector, and value vector triples are used to encode the relationship between any two positions in the feature map of each processing branch, so as to deeply mine and capture global context information; For the processing branch where the input feature 1 is located, the key vector and the value vector are K a and V a , and the query vector is the query vector Q obtained by linear transformation from the processing branch where the input feature 2 is located b ; For the processing branch where the input feature 2 is located, the key vector and value vector are K b and V b , and the query vector is the query vector Q obtained by linear transformation from the processing branch where the input feature 1 is located a ; For the processing branches where input feature 1 and input feature 2 are located, the processing results are feature maps Z a , Z b ; Where a is the facial feature, b is the facial key point feature, is the normalized scaling factor, and d represents the embedding dimension; Finally, the feature map Z is transformed into a , Z b The channel dimensions are restored to the initial channel dimensions.

4. The multi-stream attention interactive facial expression recognition method according to claim 1, characterized in that: In the step 1, facial image data is obtained, face alignment, face normalization and data enhancement preprocessing operations are performed on the facial image, and the preprocessed image is uniformly adjusted to 224×224 pixels.

5. The multi-stream attention interactive facial expression recognition method according to claim 1, characterized in that: The process of the dynamic label adjustment module re-weighting the final features and updating the noise labels is as follows: Step I. Use a linear fully connected layer and a sigmoid activation function to obtain the importance weight of each sample, and divide them into two subsets, namely, a clean sample subset and a noise sample subset, and use the regional control loss to regularize the importance weight; The sample in step I refers to the final feature F output by the channel interaction module. fuse Samples in Step II. Based on the softmax function prediction probability, for each sample in the low-weight sample subset, compare the difference between its maximum prediction probability and the prediction probability corresponding to the original given label to achieve accurate identification and update of the noise label; Among them, the low-weight sample subset is the noise sample subset obtained in step I; The specific judgment process is: if the maximum predicted probability of a sample is higher than the maximum predicted probability of a given label by a preset threshold, a new pseudo-label will be assigned to the sample, and this pseudo-label is the category corresponding to the maximum predicted probability.

6. The multi-stream attention interactive facial expression recognition method according to claim 5, characterized in that: The step I is specifically as follows: Suppose the final feature F output by the channel interaction module fuse =[x1,x2,...,x N ], containing facial features of N images; The reweighting part includes a linear fully connected layer and a sigmoid activation function to obtain the final feature F after fusion. fuse As input, and output a corresponding importance weight for each feature vector, expressed as: Among them, α i is the importance weight of the i-th sample, W a is the parameter of the fully connected layer for importance attention, σ represents the sigmoid activation function, x i Represents the final feature F fuse The i-th sample in , i = 1, 2, ..., N; According to the preset division ratio β, the samples are divided into two subsets: a clean sample subset and a noise sample subset.

7. The multi-stream attention interactive facial expression recognition method according to claim 6, characterized in that: In step I, the regional control loss L is defined RC , the calculation formula is as follows: L RC =max{0,δ1-(α H -a L )}; Among them, δ1 is the threshold for dividing subsets, α H and α L are the average values ​​of the high-weight sample subset of M samples and the low-weight sample subset of NM samples, respectively. The high-weight sample subset is the clean sample subset obtained in step I.

8. The multi-stream attention interactive facial expression recognition method according to claim 7, characterized in that: In step II, a label update operation is performed for the softmax predicted probability of samples in the low-weight sample subset; For each sample in the low-weight sample subset, compare its maximum predicted probability with the predicted probability corresponding to the original given label; if the maximum predicted probability of a sample exceeds a preset threshold of the maximum predicted probability of the given label, a new pseudo-label will be assigned to the sample, which is the category corresponding to the maximum predicted probability; A label update is defined as: Among them, y′ is the new pseudo label, δ2 is the threshold for label update, and P max is the maximum predicted probability, P gt is the predicted probability of a given label, l org and l max They are the original given label and the label corresponding to the maximum predicted probability.

9. The multi-stream attention interactive facial expression recognition method according to claim 8, characterized in that: In step 3, during the model training process, weighted cross entropy loss is used to process the noise labels in the multi-classification problem; The weighted cross entropy loss L is defined using the following formula: WCE : Among them, w yi is the true category y corresponding to the i-th sample i The weight vector, w j is the weight vector of the jth category, C represents the total number of categories, and N represents the total number of samples; During training, the total loss function L total For: L total =γL RC +(1-γ)L WCE , where γ is the trade-off ratio.

Citation Information

Patent Citations

  • Facial expression prediction method based on dynamic distribution fusion

    CN116363733A

  • Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism

    CN117370828A

  • Micro-expression recognition method based on noise label self-repairing

    CN117831092A