A multi-scale fusion expression recognition method and system based on visual transformer
Through the multi-scale fusion of expression recognition method and multi-task classifier, the problem of insufficient long-distance dependence and emotional dimension capture in expression recognition in visual Transformer is solved, and the accuracy and generalization ability of expression recognition are improved.
Patent Information
- Application Number
- CN202510186766.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing visual Transformers are difficult to effectively capture long-distance dependencies and introduce continuous emotional dimensions, resulting in insufficient expression recognition accuracy and generalization capabilities.
A multi-scale fusion expression recognition method is adopted to fusion short-distance and long-distance through a multi-head attention mechanism, combined with a visual transformer for expression recognition, and a multi-task classifier is used to predict expression recognition results.
The accuracy and generalization ability of expression recognition have been improved, so that the expression recognition model has more delicate description ability and stronger feature extraction ability for expressions.
Smart Images

Figure CN119672787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to facial expression recognition technology in the field of image processing, and in particular to a multi-scale fusion expression recognition method and system based on a visual transformer. Background Art
[0002] The goal of facial expression recognition is to judge the emotional state of an individual by analyzing various features in facial images (such as the shape and position changes of eyes, mouth, and eyebrows), and to maintain a high recognition accuracy under different lighting conditions, expression changes, occlusion, angle changes, etc. Facial expression recognition tasks face three main challenges: inter-class similarity, intra-class differences, and scale sensitivity. Convolutional neural networks (CNNs) are good at modeling fine-grained local features, while Transformers are better at modeling global contextual information. Obviously, CNNs and Transformers have complementary characteristics, among which CNNs have a stronger ability to capture spatial structural information within the local receptive field. Compared with convolutional neural networks, visual transformers establish long-distance dependencies through self-attention mechanisms, thereby extracting global information more effectively. The computational complexity of traditional transformers grows quadratically (that is, it increases with the square of the sequence length), making it difficult for existing visual transformers to effectively extract image features of different scales and ranges. This is because the input tokens are all generated by fixed-size image blocks. In order to reduce computational complexity, the typical Swin-Transformer abandons the modeling of long-distance dependencies and only extracts attention features in a local range. Therefore, how to capture long-distance dependencies and introduce continuous emotion dimensions to more accurately capture subtle changes in emotions and improve the accuracy and generalization ability of expression recognition has become a key technical problem that needs to be solved urgently. Summary of the invention
[0003] Technical problem to be solved by the present invention: In view of the above-mentioned problems of the prior art, a multi-scale fusion expression recognition method and system are provided. The present invention aims to capture long-distance dependencies and introduce a continuous emotion dimension to more accurately capture subtle changes in emotions, improve the accuracy and generalization ability of expression recognition, and enable the expression recognition model to have a more delicate description ability and stronger feature extraction ability for expressions.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] A multi-scale fusion expression recognition method based on a visual transformer comprises the following steps: extracting multi-scale key point features and facial image features from an input facial expression image; sequentially performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at each scale to obtain fusion features; fusing the fusion features at each scale to obtain multi-scale fusion features; performing expression recognition on the multi-scale fusion features using a visual transformer to obtain an expression recognition result, and sequentially performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at a certain scale, comprising: respectively dividing the key point features and facial image features at the scale into local windows after being mapped by a linear layer, using each attention head of a multi-head attention mechanism to calculate the local windows of the key point features and facial image features, and using the multi-head attention mechanism to extract features as short-distance fusion features; dividing the short-distance fusion features into local windows, extracting feature blocks from the fixed grid positions specified in the local windows of the key point features and the short-distance fusion features, sampling the extracted feature blocks with a step length of r across windows, and then using the multi-head attention mechanism to extract features as the fusion features finally obtained at the scale.
[0006] Optionally, the function expression of the local window in which each attention head of the multi-head attention mechanism calculates key point features and facial image features and extracts features as short-distance fusion features using the multi-head attention mechanism is:
[0007] ,
[0008] in, is the short-distance fusion feature. For splicing operation, ~ They are No. 1~ The output features of the attention head are is the output projection matrix of short-distance fusion, and any The output features of the attention head The calculation function expression is:
[0009] ,
[0010] in, is the softmax activation function, Key point features No. A local window, is the facial image feature No. Local window, key point features Query Q as attention head, facial image features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, The dimension of the key.
[0011] Optionally, the function expression of the fusion feature at this scale obtained by sampling the extracted feature blocks across windows with a step size of r and then extracting features using a multi-head attention mechanism is:
[0012] ,
[0013] in, To fusion features, For splicing operation, ~ They are No. 1~ The output features of the attention head are is the output projection matrix of long-distance fusion, and any The output features of the attention head The calculation function expression is:
[0014] ,
[0015] in, is the softmax activation function, Key point features No. A local window, For short-distance fusion features No. local window, key point features Query Q as attention head, short-distance fusion features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, Express The step length is Cross-window feature sampling, The dimension of the key.
[0016] Optionally, when the key point features and facial image features at the scale are divided into local windows through linear layer mapping, the local window sizes at different scales are different, and the function expression for dividing into local windows is:
[0017] ,
[0018] In the above formula, Key point features No. A local window, is the facial image feature No. A local window, For partition operation, The window size for extracting short-distance fusion features. ,in is the number of local windows divided when extracting short-distance fusion features; when the short-distance fusion features are divided into local windows, the local window sizes at different scales are different, and the function expression for dividing into local windows is:
[0019] ,
[0020] In the above formula, Key point features No. A local window, For short-distance fusion features No. A local window, is the window size used to extract fusion features. ,in The number of local windows divided when extracting fusion features.
[0021] Optionally, when extracting multi-scale key point features and facial image features from the input facial expression image respectively, extracting multi-scale key point features from the input facial expression image respectively includes: extracting three scales of key point features from the input facial expression image respectively using the MobileFaceNet model; when extracting multi-scale key point features and facial image features from the input facial expression image respectively, extracting multi-scale facial image features from the input facial expression image respectively includes: extracting three scales of facial image features from the input facial expression image respectively using the IR50 model; fusing the fused features at each scale to obtain a multi-scale fused feature refers to splicing the fused features at each scale to obtain a fused multi-scale fused feature; before extracting multi-scale key point features and facial image features from the input facial expression image respectively, it includes using pre-trained model parameters for the MobileFaceNet model and freezing the model parameters of the MobileFaceNet model, and optimizing the model parameters of the IR50 model based on sample training of facial expression images.
[0022] In addition, the present invention also provides a multi-scale fusion expression recognition method, comprising the following steps: extracting multi-scale key point features and facial image features from an input facial expression image; performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at each scale in turn to obtain fusion features; fusing the fusion features at each scale to obtain multi-scale fusion features; performing expression recognition on the multi-scale fusion features using a multi-task classifier to obtain an expression recognition result, and performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at a certain scale in turn, including: dividing the key point features and facial image features at the scale into local windows after being mapped by a linear layer, using each attention head of a multi-head attention mechanism to calculate the local windows of the key point features and facial image features, and using the multi-head attention mechanism to extract features as short-distance fusion features; using the short-distance fusion features to obtain the expression recognition result; and performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at the scale in turn. The short-distance fusion feature is divided into local windows, feature blocks are extracted from the fixed grid positions specified in the local windows of the key point features and the short-distance fusion features, and the extracted feature blocks are sampled with a cross-window feature with a step size of r, and then a multi-head attention mechanism is used to extract features as the final fusion features at this scale; the multi-task classifier includes a benefit value estimation module, a connection module, an awakening value estimation module, a connection module and an expression classification module connected in sequence, the benefit value estimation module is used to perform benefit value estimation according to the input fusion feature to obtain a benefit value, the benefit value is input to the first connection module and then connected with the input fusion feature as the input of the awakening value estimation module, so as to obtain the awakening value through the awakening value estimation module, the awakening value is input to the second connection module and then connected with the output result of the first connection module as the input of the expression classification module, so as to obtain the final expression recognition result through the expression classification module.
[0023] Optionally, the potency value estimation module is a single-layer fully connected network layer, and the function expression of the potency value obtained by performing potency value estimation is:
[0024] ,
[0025] in, For the effective value, is the transpose of the learnable weight matrix, is a multi-scale fusion feature. is the bias; the wake-up value estimation module is a two-layer neural network layer, and the function expression for the wake-up value estimation is:
[0026] ,
[0027] ,
[0028] in, is the wake-up value, and are the transpose of the weight matrices of the two neural network layers, and are the biases of the two neural network layers, is the output of the first neural network layer; the expression classification module performs expression classification as follows:
[0029] ,
[0030] in, is the probability of belonging to a certain category of expression, is the transpose of the weight matrix of the expression classification module, is the bias of the expression classification module, for and The splicing result is for and The splicing result.
[0031] In addition, the present invention also provides a multi-scale fusion expression recognition system, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer or the multi-scale fusion expression recognition method.
[0032] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer or the multi-scale fusion expression recognition method through a processor.
[0033] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer or the multi-scale fusion expression recognition method through a processor.
[0034] Compared with the prior art, the present invention mainly has the following advantages: the present invention comprises the following steps: the present invention sequentially performs short-distance fusion and long-distance fusion based on the attention mechanism on the key point features and facial image features at each scale to obtain fusion features, the key point features and facial image features at the scale are divided into local windows after being mapped by a linear layer, each attention head of the multi-head attention mechanism is used to calculate the local windows of the key point features and facial image features, and the multi-head attention mechanism is used to extract features as short-distance fusion features; the short-distance fusion features are divided into local windows, feature blocks are extracted from the fixed grid positions specified in the local windows of the key point features and the short-distance fusion features, the extracted feature blocks are sampled across windows with a step size of r, and then the multi-head attention mechanism is used to extract features as the final fusion features at the scale, which can capture long-distance dependencies and introduce continuous emotion dimensions to more accurately capture subtle changes in emotions, improve the accuracy and generalization ability of expression recognition, and enable the expression recognition model to have a more delicate description ability and a stronger feature extraction ability for expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of the overall network structure of the first embodiment of the present invention.
[0036] Figure 2 The figure is a schematic diagram of the network structure of short-distance fusion and long-distance fusion in the first embodiment of the present invention.
[0037] Figure 3 Schematic diagrams of the principles of short-distance fusion and long-distance fusion in Embodiment 1 of the present invention, wherein (a) is a schematic diagram of the principles of short-distance fusion, and (b) is a schematic diagram of the principles of long-distance fusion.
[0038] Figure 4 This is a schematic diagram of the overall network structure of the second embodiment of the present invention.
[0039] Figure 5 Schematic diagram of the network structure of the multi-task classifier in the second embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] Embodiment 1:
[0042] like Figure 1As shown, this embodiment provides a multi-scale fusion expression recognition method based on a visual transformer, comprising the following steps: extracting multi-scale key point features and facial image features from an input facial expression image; performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at each scale to obtain fusion features; fusing the fusion features at each scale to obtain multi-scale fusion features; and performing expression recognition on the multi-scale fusion features using a visual transformer to obtain an expression recognition result.
[0043] See also Figure 2 In this embodiment, the key point features and facial image features at a certain scale are sequentially subjected to short-distance fusion and long-distance fusion based on the attention mechanism, including: the key point features and facial image features at the scale are divided into local windows after being mapped by a linear layer, and each attention head of the multi-head attention mechanism is used to calculate the local windows of the key point features and facial image features, and the multi-head attention mechanism is used to extract features as short-distance fusion features; the short-distance fusion features are divided into local windows, and feature blocks are extracted from the fixed grid positions specified in the local windows of the key point features and the short-distance fusion features, and the extracted feature blocks are subjected to cross-window feature sampling with a step size of r, and then the multi-head attention mechanism is used to extract features as the final fusion features at the scale.
[0044] See also Figure 2 In this embodiment, each attention head of the multi-head attention mechanism calculates the key point features and the local window of the facial image features, and the multi-head attention mechanism is used to extract features as the function expression of the short-distance fusion feature:
[0045] ,
[0046] in, is the short-distance fusion feature. For splicing operation, ~ They are No. 1~ The output features of the attention heads (where The value can be taken as required, for example, in this embodiment ), is the output projection matrix of short-distance fusion, and any The output features of the attention head The calculation function expression is:
[0047] ,
[0048] in, is the softmax activation function, Key point features No. A local window, is the facial image feature No. Local window, key point features Query Q as attention head, facial image features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, is the dimension of the key, is the scaling factor, Scaling prevents the gradient from being too small. The output features of the attention head The computational function expression of adopts the cross-attention mechanism. Through the cross-scale cross-attention mechanism, the facial key point features and facial image features are integrated locally and globally, so that the model can preferentially learn the detail association pattern in the local area, significantly improving the ability to capture the subtle deformation of facial features. Finally, the multi-head attention mechanism is used to combine the 1st to the 2nd The output features of the attention heads are combined to obtain short-range fusion features.
[0049] See also Figure 2 In this embodiment, the extracted feature blocks are sampled across windows with a step size of r, and then the multi-head attention mechanism is used to extract features as the function expression of the fusion features at this scale finally obtained:
[0050] ,
[0051] in, To fusion features, For splicing operation, ~ They are No. 1~ The output features of the attention heads (where The value can be taken as required, for example, in this embodiment ), is the output projection matrix of long-distance fusion, and any The output features of the attention head The calculation function expression is:
[0052] ,
[0053] in, is the softmax activation function, Key point features No. A local window, For short-distance fusion features No. local window, key point features Query Q as attention head, short-distance fusion features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, Express The step length is Cross-window feature sampling, is the dimension of the key, is the scaling factor. In the long-distance fusion stage, under the premise of keeping the feature map window division unchanged, a cross-window hole sampling strategy is creatively designed. Feature blocks are extracted from fixed grid positions (such as the center point, etc.) of each window, and these cross-window co-location features are reconstructed. The long-range dependencies between multiple windows are captured simultaneously through a single matrix operation. Its computational complexity grows linearly with the local window stage, while the computational complexity of the traditional global attention mechanism increases quadratically with the size of the feature map.
[0054] In this embodiment, when the key point features and facial image features at this scale are divided into local windows through linear layer mapping (multiplied by the weight matrix), the local windows at different scales have different sizes and do not overlap, and the function expression for dividing into local windows is:
[0055] ,
[0056] In the above formula, Key point features No. A local window, is the facial image feature No. A local window, For partition operation, The window size for extracting short-distance fusion features. ,in is the number of local windows divided when extracting short-distance fusion features; when the short-distance fusion features are divided into local windows, the local window sizes at different scales are different, and the function expression for dividing into local windows is:
[0057] ,
[0058] In the above formula, Key point features No. A local window, For short-distance fusion features No. A local window, is the window size used to extract fusion features. ,in The number of local windows divided when extracting fusion features. Figure 3 The schematic diagrams of the principles of short-distance fusion and long-distance fusion in this embodiment, where (a) is a schematic diagram of the principles of short-distance fusion, and (b) is a schematic diagram of the principles of long-distance fusion. It should be noted that the window size and They can be the same or different. For example, as an optional implementation, in this embodiment, when short-distance fusion and long-distance fusion are performed, the window sizes divided at the three scales are and are the same and are set to 28, 14, and 7 respectively.
[0059] See also Figure 1 As an optional implementation, in this embodiment, when extracting multi-scale key point features and facial image features from the input facial expression image, respectively, extracting multi-scale key point features from the input facial expression image includes: extracting three scales of key point features from the input facial expression image using the MobileFaceNet model (existing known model); when extracting multi-scale key point features and facial image features from the input facial expression image, respectively, extracting multi-scale facial image features from the input facial expression image includes: extracting three scales of facial image features from the input facial expression image using the IR50 model (existing known model); the fusion of fusion features at each scale to obtain multi-scale fusion features refers to splicing the fusion features at each scale to obtain the fused multi-scale fusion features. Undoubtedly, other existing feature extraction models can also be used to extract key point features and facial image features as needed. Figure 1 In the figure, the key point features and facial image features at three scales are sequentially fused in short distance and long distance based on the attention mechanism to obtain the fused features, which are represented as stage 1, stage 2 and stage 3 respectively.
[0060] In this embodiment, before extracting multi-scale key point features and facial image features from the input facial expression image, the model parameters of the MobileFaceNet model are pre-trained and the model parameters of the MobileFaceNet model are frozen, and the model parameters of the IR50 model are optimized based on the sample training of the facial expression image. Considering that the key point positioning needs to maintain spatial consistency, the MobileFaceNet model is a pre-trained model with frozen network parameters. The reason is to freeze all network parameters of the MobileFaceNet model to prevent gradient updates from destroying its geometric feature extraction capabilities obtained through large-scale data set training, so as to ensure the stability of the key point features output by the MobileFaceNet model. The IR50 model is initialized with pre-trained network parameters, and the training process also includes continuing to optimize the network parameters of the IR50 model. At the same time, the input image is sent to the IR50 model in parallel for multi-scale feature extraction. This branch adopts a full parameter unfreezing strategy, and globally updates all network layer parameters through an adaptive momentum optimizer (AdamW), sets the basic learning rate according to different data sets, and applies a weight decay coefficient of 1e-4, so as to enhance the network's adaptability to expression detail features, while avoiding the risk of overfitting caused by excessive parameter adjustment of deep networks.
[0061] In order to verify the multi-scale fusion expression recognition method based on the visual transformer in this embodiment, the well-known AffectNet dataset (including AffectNet-7 and AffectNet-8 datasets) and RAF-DB dataset are used for experiments in this embodiment. Two NVIDIA GeForce RTX 2080 Ti graphics cards are used in the experiment, and the model training is implemented through PyTorch. The Adam optimizer is used. On the AffectNet dataset, the initial learning rate is set to 1e-6, the batch size is set to 144, the weight decay is 0.05, and 200 epochs are trained. All experiments use the same data enhancement strategy to avoid overfitting during training. Specifically, during the training process, the original image is randomly cropped to 224×224, and randomly horizontally flipped and erased before input. As a comparison of the method in this embodiment, this embodiment specifically uses advanced methods in recent years (including POSTER++, POSTER, FG-AGR, Multi-task EfficientNet-B2, MFER, MFEFER and FST-MWOS) for comparison, and the final results are shown in Table 1.
[0062] Table 1 Recognition results of the method in this embodiment and other advanced methods in recent years on the AffectNet-8 dataset
[0063]
[0064] As shown in Table 1, the recognition accuracy of the method in this embodiment for each category in the AffectNet-8 dataset is good, especially in the sad category, where the recognition accuracy reaches 66.8%, the highest value. In the happy category, the difference with the highest accuracy is 0.2%, ranking second among all methods; in the surprise category and the fear category, the recognition accuracy is second only to the optimal method. In the overall average recognition accuracy, the proposed method has the highest recognition result, which is 0.585% better than the second place. This fully proves that the method in this embodiment has strong recognition ability and generalization performance in facial expression recognition tasks.
[0065] In addition, in order to verify the effectiveness of short-distance fusion and long-distance fusion in the method of this embodiment, an ablation experiment was designed using the AffectNet-8 dataset and the RAF-DB dataset in this embodiment, and the results obtained are shown in Table 2.
[0066] Table 2 Ablation experiment design and results of the method in this embodiment
[0067]
[0068] As shown in Table 2, on the RAF-DB and AffectNet-7 datasets, when short-distance fusion or long-distance fusion is used alone, the performance is lower than the combined use of the two (increased by 0.3% and 0.12% respectively). This shows that short-distance fusion and long-distance fusion are functionally complementary. Short-distance fusion contributes more in the RAF-DB dataset (expression-dominated) (91.36% vs. 91.30%), while long-distance fusion is more advantageous in the AffectNet-7 dataset (complex scenes) (67.37% vs. 67.20%), verifying the adaptability of joint modeling of short-distance fusion and long-distance fusion.
[0069] In addition, in order to verify the effectiveness of layer-by-layer splicing of multi-scale fusion features in the method of this embodiment, a comparative experiment of direct splicing is also designed in this embodiment, that is, directly splicing the multi-scale key point features and facial image features and then using the visual transformer to perform expression recognition on the multi-scale fusion features to obtain the expression recognition results. The results are shown in Table 3.
[0070] Table 3 Comparative analysis results of multi-task learning architecture performance
[0071]
[0072] As shown in Table 3, the layer-by-layer concatenation architecture (this solution) has an accuracy improvement of 0.22% (64.39% vs 64.17%) compared to the direct concatenation (control group), which verifies the superiority of progressive feature fusion. The baseline model (63.73%) without multi-task learning has the lowest performance, indicating that continuous dimension prediction (valence / arousal) has a significant auxiliary effect on the expression classification task. Layer-by-layer concatenation achieves progressive fusion from low-level visual features to high-level semantic features by retaining intermediate features (such as the 512-dimensional hidden layer features of the arousal module). Compared with direct concatenation, the model is more robust.
[0073] In summary, the multi-scale fusion expression recognition method based on the visual transformer in this embodiment extracts multi-scale key point features and facial image features from the input facial expression image; performs short-distance fusion and long-distance fusion based on the attention mechanism on the key point features and facial image features at each scale to obtain fusion features; fuses the fusion features at each scale to obtain multi-scale fusion features; and uses the visual transformer to perform expression recognition on the multi-scale fusion features to obtain expression recognition results. This embodiment aims at the similarity between facial expression categories and the difference within the category. By introducing facial key point information, the image features are guided to focus on the expression salient area, thereby reducing the influence of skin color, age and illumination changes on the recognition performance. This embodiment adopts short-distance fusion and long-distance fusion to realize the cross-attention fusion of local features and global features, so as to effectively capture local features and global features, and can effectively learn long-distance dependencies, make up for the limitations of traditional convolutional neural networks, and maintain better performance while reducing computational complexity.
[0074] In addition, this embodiment also provides a multi-scale fusion expression recognition system, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the multi-scale fusion expression recognition method based on the visual transformer. This embodiment also provides a computer-readable storage medium, wherein a computer program or instruction is stored in the computer-readable storage medium, wherein the computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method based on the visual transformer through a processor. This embodiment also provides a computer program product, including a computer program or instruction, wherein the computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method based on the visual transformer through a processor.
[0075] Embodiment 2:
[0076] This embodiment is basically the same as the first embodiment, and the main differences are: Figure 4 As shown, in this embodiment, the visual transformer in the first embodiment is replaced by a multi-task classifier. After obtaining the multi-scale fusion features, the multi-task classifier is used to perform expression recognition on the multi-scale fusion features to obtain the expression recognition result.
[0077] Facial expression recognition is a typical problem in image recognition, where the input facial image X should belong to one of the expression categories, such as anger, surprise, etc. Expressions are usually represented in many ways, such as valence (V), arousal (A), and action unit (AU) detection. ,in Valence and arousal, respectively. More advanced representations include a set of discrete action units (AUs) from the Ekman model of the Facial Action Coding System (FACS), and psychologist James Russell's continuous encoding of emotions in the two-dimensional space of arousal-valence. The former reflects the degree of passivity or activity of the emotional state, while the latter characterizes the positive and negative polarity of the emotion. Although individual emotions can be identified through a variety of signals (such as speech, pronunciation, body language, etc.), facial analysis usually provides more accurate results. Multi-task classifiers can effectively explore the relationship between different tasks, improve the discriminative ability of shared backbones, and simultaneously identify facial expressions in static images and predict valence and arousal. Multi-task learning has been widely used in computer vision tasks. The dependence between tasks is mainly reflected in two aspects: first, some tasks share low-level feature representations, which can be used to improve the performance of each task; second, the high-level features of one task may become important inputs for other tasks. For example, since the definition of expression depends to a certain extent on the action units (AUs) of the face, the high-level features in the action unit AU detection task can be used for expression estimation. In order to model the correlation between different tasks (i.e., estimating valence value, estimating arousal value, and expression classification) and realize multi-task sentiment analysis, this embodiment aims to model the correlation through a multi-task classifier, and expands the valence-arousal prediction task of the continuous sentiment model on the basis of the original discrete sentiment model, so that the model has a more delicate description ability and stronger feature extraction ability for expressions. Figure 5 As shown, the multi-task classifier in this embodiment includes a benefit value estimation module, a connection module, an awakening value estimation module, a connection module and an expression classification module which are connected in sequence. The benefit value estimation module is used to perform benefit value estimation according to the input fusion feature to obtain the benefit value. After the benefit value is input into the first connection module, it is connected with the input fusion feature and then used as the input of the awakening value estimation module, so as to obtain the awakening value by performing the awakening value estimation through the awakening value estimation module. After the awakening value is input into the second connection module, it is connected with the output result of the first connection module and then used as the input of the expression classification module, so as to obtain the final expression recognition result by performing expression classification through the expression classification module.
[0078] In this embodiment, the efficacy value estimation module is a single-layer fully connected network layer, and the function expression of the efficacy value obtained by performing efficacy value estimation is:
[0079] ,
[0080] in, For the effective value, is the transpose of the learnable weight matrix, is a multi-scale fusion feature. is the bias; the wake-up value estimation module is a two-layer neural network layer, and the function expression for the wake-up value estimation is:
[0081] ,
[0082] ,
[0083] in, is the wake-up value, and are the transpose of the weight matrices of the two neural network layers, and are the biases of the two neural network layers, is the output of the first neural network layer; the expression classification module performs expression classification as follows:
[0084] ,
[0085] in, is the probability of belonging to a certain category of expression, is the transpose of the weight matrix of the expression classification module, is the bias of the expression classification module, for and The splicing result is for and The splicing result. In the training process of the multi-task classifier in this embodiment, for the regression tasks of the effectiveness value and the arousal value, the consistency correlation coefficient (Concordance Correlation Coefficient, CCC) is used as the loss (consistency correlation loss) in this embodiment. For the classification task, the cross entropy loss (Cross Entropy, CE loss) is used, so the loss function adopted by the multi-task classifier is for:
[0086] ,
[0087] in, and are the consistency-related losses of valence value and arousal value, respectively. is the weight coefficient, is the expression classification loss, where the calculation function expression of the consistency-related loss is:
[0088] ,
[0089] In the above formula, is the consistency-related loss, is the Pearson correlation coefficient between the predicted value and the true value, , is the statistic of the predicted value, and is the statistic of the true value. For the dimension of the efficacy value, substitute , , which can calculate the consistency-related loss of the effectiveness value For the dimension of the awakening value, substitute , , which can calculate the consistency-related loss of the wake-up value .in, and are the predicted value and true value of the efficacy value, and are the predicted value and the true value of the arousal value respectively. The calculation function expression of the expression classification loss is:
[0090] ,
[0091] In the above formula, is the sample size, is the total number of expression categories, and Respectively The samples belong to The true probability and predicted probability of each expression category. The distribution consistency of the continuous dimension prediction is constrained by the consistency correlation loss, and the discriminability of discrete classification is guaranteed by the cross entropy loss. Finally, multi-objective joint optimization can be achieved through weighted summation. Although the visual transformer in Example 1 is replaced by a multi-task classifier in this embodiment, since the multi-scale fusion feature extraction method is the same, the technical effect of the multi-scale fusion feature extraction method in Example 1 can also be achieved.
[0092] In addition, this embodiment also provides a multi-scale fusion expression recognition system, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the multi-scale fusion expression recognition method. This embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method through a processor. This embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method through a processor.
[0093] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention may be in the form of methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that instructions executed by the processor of a computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0094] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A multi-scale fusion expression recognition method based on visual transformer, characterized in that: The method comprises the following steps: extracting multi-scale key point features and facial image features from an input facial expression image; performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at each scale to obtain fusion features; fusing the fusion features at each scale to obtain multi-scale fusion features; performing expression recognition on the multi-scale fusion features using a visual transformer to obtain expression recognition results, and performing short-distance fusion and long-distance fusion based on an attention mechanism on the key point features and facial image features at a certain scale to obtain local windows respectively through linear layer mapping, using a multi-scale fusion transformer to obtain expression recognition results; and performing short-distance fusion and long-distance fusion based on an attention mechanism to obtain expression recognition results; and performing short-distance fusion and long-distance fusion on the key point features and facial image features at a certain scale respectively through a linear layer mapping, and using a multi-scale fusion transformer to obtain expression recognition results; and Each attention head of the head attention mechanism calculates the local window of the key point feature and the facial image feature, and uses the multi-head attention mechanism to extract the feature as the short-distance fusion feature; the short-distance fusion feature is divided into local windows, and the feature blocks are extracted from the fixed grid positions specified in the local windows of the key point feature and the short-distance fusion feature. The extracted feature blocks are sampled across the windows with a step size of r, and then the multi-head attention mechanism is used to extract the feature as the final fusion feature at this scale. The function expression of the local window of the key point feature and the facial image feature calculated by each attention head of the multi-head attention mechanism and the multi-head attention mechanism to extract the feature as the short-distance fusion feature is: , in, is the short-distance fusion feature. For splicing operation, ~ They are No. 1~ The output features of the attention head are is the output projection matrix of short-distance fusion, and any The output features of the attention head The calculation function expression is: , in, is the softmax activation function, Key point features No. A local window, is the facial image feature No. Local window, key point features Query Q as attention head, facial image features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, The dimension of the key.
2. The multi-scale fusion expression recognition method based on visual transformer according to claim 1 is characterized in that: The function expression of the fusion feature at this scale obtained by sampling the extracted feature blocks across windows with a step length of r and then extracting features using a multi-head attention mechanism is: , in, To fusion features, For splicing operation, ~ They are No. 1~ The output features of the attention head are is the output projection matrix of long-distance fusion, and any The output features of the attention head The calculation function expression is: , in, is the softmax activation function, Key point features No. A local window, For short-distance fusion features No. Local window, key point features Query Q as attention head, short-distance fusion features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, Express The step length is Cross-window feature sampling, The dimension of the key.
3. The multi-scale fusion expression recognition method based on visual transformer according to claim 2 is characterized in that: When the key point features and facial image features at the scale are divided into local windows through linear layer mapping, the local window sizes at different scales are different, and the function expression for dividing into local windows is: , In the above formula, Key point features No. A local window, is the facial image feature No. A local window, For partition operation, The window size for extracting short-distance fusion features. ,in is the number of local windows divided when extracting short-distance fusion features; when the short-distance fusion features are divided into local windows, the local window sizes at different scales are different, and the function expression for dividing into local windows is: , In the above formula, Key point features No. A local window, For short-distance fusion features No. A local window, is the window size used to extract fusion features. ,in The number of local windows divided when extracting fusion features.
4. The multi-scale fusion expression recognition method based on visual transformer according to claim 1 is characterized in that: When extracting multi-scale key point features and facial image features from the input facial expression image, respectively, extracting multi-scale key point features from the input facial expression image includes: extracting three scales of key point features from the input facial expression image using the MobileFaceNet model; when extracting multi-scale key point features and facial image features from the input facial expression image, respectively, extracting multi-scale facial image features from the input facial expression image includes: extracting three scales of facial image features from the input facial expression image using the IR50 model; fusing fusion features at each scale to obtain multi-scale fusion features refers to splicing fusion features at each scale to obtain fused multi-scale fusion features; before extracting multi-scale key point features and facial image features from the input facial expression image, respectively, it includes using pre-trained model parameters for the MobileFaceNet model and freezing the model parameters of the MobileFaceNet model, and optimizing the model parameters of the IR50 model based on sample training of facial expression images.
5. A multi-scale fusion expression recognition method, characterized in that: The method comprises the following steps: extracting multi-scale key point features and facial image features from an input facial expression image; performing short-distance fusion and long-distance fusion based on the attention mechanism on the key point features and facial image features at each scale in turn to obtain fusion features; fusing the fusion features at each scale to obtain multi-scale fusion features; performing expression recognition on the multi-scale fusion features using a multi-task classifier to obtain expression recognition results, and performing short-distance fusion and long-distance fusion based on the attention mechanism on the key point features and facial image features at a certain scale in turn, including: dividing the key point features and facial image features at the scale into local windows after being mapped by a linear layer, using each attention head of the multi-head attention mechanism to calculate the local windows of the key point features and facial image features; using the multi-head attention mechanism to extract features as short-distance fusion features; dividing the short-distance fusion features into local windows, extracting features from the fixed grid positions specified in the local windows of the key point features and the short-distance fusion features. The feature block is extracted by adopting cross-window feature sampling with a step length of r, and then adopting a multi-head attention mechanism to extract features as the final fusion features under this scale; the multi-task classifier includes a benefit value estimation module, a connection module, an awakening value estimation module, a connection module and an expression classification module connected in sequence, the benefit value estimation module is used to estimate the benefit value according to the input fusion feature to obtain the benefit value, the benefit value is input into the first connection module and connected with the input fusion feature as the input of the awakening value estimation module, so as to obtain the awakening value through the awakening value estimation module, the awakening value is input into the second connection module and connected with the output result of the first connection module as the input of the expression classification module, so as to obtain the final expression recognition result through the expression classification module, and each attention head of the multi-head attention mechanism calculates the key point features and the local window of the facial image features, and the multi-head attention mechanism is used to extract features as the function expression of the short-distance fusion feature: , in, is the short-distance fusion feature. For splicing operation, ~ They are No. 1~ The output features of the attention head are is the output projection matrix of short-distance fusion, and any The output features of the attention head The calculation function expression is: , in, is the softmax activation function, Key point features No. A local window, is the facial image feature No. Local window, key point features Query Q as attention head, facial image features As the key K and value V of the attention head, , and Respectively The weight matrix of query Q, key K and value V corresponding to each attention head, The dimension of the key.
6. The multi-scale fusion expression recognition method according to claim 5, characterized in that: The efficacy value estimation module is a single-layer fully connected network layer, and the function expression of the efficacy value obtained by performing efficacy value estimation is: , in, For the effective value, is the transpose of the learnable weight matrix, is a multi-scale fusion feature. is the bias; the wake-up value estimation module is a two-layer neural network layer, and the function expression for the wake-up value estimation is: , , in, is the wake-up value, and are the transpose of the weight matrices of the two neural network layers, and are the biases of the two neural network layers, is the output of the first neural network layer; the expression classification module performs expression classification as follows: , in, is the probability of belonging to a certain category of expression, is the transpose of the weight matrix of the expression classification module, is the bias of the expression classification module, for and The splicing result is for and The splicing result.
7. A multi-scale fusion expression recognition system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer as described in any one of claims 1 to 4 or the multi-scale fusion expression recognition method as described in claim 5 or 6.
8. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer described in any one of claims 1 to 4 or the multi-scale fusion expression recognition method described in claim 5 or 6 through a processor.
9. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the multi-scale fusion expression recognition method based on visual transformer described in any one of claims 1 to 4 or the multi-scale fusion expression recognition method described in claim 5 or 6 through a processor.
Citation Information
Patent Citations
Multi-task facial expression recognition method guided by emotional prior topological graph
CN116721457A
Infrared image dynamic range adaptive enhancement method and system based on signal-to-noise ratio perception
CN118396914A
Emotion recognition method based on facial feature and key point fusion network
CN119323815A