A facial expression recognition method based on regional self-attention convolutional neural network
By adopting regional self-attention convolutional neural network in facial expression recognition technology, combining deep global features and regional local multi-valued mode, the problem of insufficient robustness when applied in real environments is solved, and more efficient facial expression recognition performance is achieved.
Patent Information
- Application Number
- CN202210492125.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-07
AI Technical Summary
When existing facial expression recognition technology is applied in real environments, due to the uncontrollability of the demand environment and the complexity of the model algorithm, it is impossible to maintain efficient application efficiency, especially when the texture changes in the significant area of the expression is not robust enough.
The facial expression recognition method based on regional self-attention convolution neural network is adopted, and the depth global features are extracted through VGG network, the regional local multi-valued mode is designed and the improved K-means algorithm is used to extract multi-mode information and grayscale difference information, and the contribution of different regions to expression recognition is quantified through regional self-attention mechanism and rank regularization loss, enhancing the robustness of the significant regional features of expressions.
It improves the performance of facial expression recognition, enhances the robustness and representation ability of features, improves the ability to identify prominent areas of expression, and improves the application efficiency in real environments.
Smart Images

Figure CN114842534B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of pattern recognition and computer vision, and in particular belongs to a method for recognizing facial expressions. Background Art
[0002] In recent years, artificial intelligence (AI) has been booming, and the application of artificial intelligence in the grassroots of society is giving rise to a profound social and technological change. The government and artificial intelligence scholars have recognized the key points of artificial intelligence, which has become the core driving force of a new round of industrial transformation and industrial revolution, and is crucial to the transformation and upgrading of the world's economic structure. With the continuous in-depth exploration of artificial intelligence technology, it has gradually become known to the public. In particular, major technology companies have commercialized related technologies and applied them to various industries such as transportation, medical care, finance, industry, and education. Computer vision (CV), as an important branch of artificial intelligence, is also widely used in various industries. Computer vision, as the name suggests, is a discipline that allows computers to have human visual functions. To be more precise, it enables computers to have the human function of identifying and tracking targets. Among them, facial expression recognition, as a key technology for perceiving and identifying facial biometric information, has been maturely applied in many fields. However, the current facial expression recognition technology is subject to factors such as the uncontrollable demand environment and the high complexity of the model algorithm, which cannot enable the technology to maintain high application efficiency in real environments.
[0003] Early research on facial expression recognition was conducted on static frontal facial images, but with the development of technology and the promotion of applications, there is an increasing demand for the recognition of facial expression images that change under uncontrolled conditions. When the camera's perspective changes relative to the captured facial expression image, the projection deformation will cause different degrees of pulling and deformation of the facial expression, and it is also possible to capture the occlusion of the facial image by different occluders, resulting in low quality of the collected facial expression images. This poses a great challenge to the application of facial expression recognition in real scenes. With the emergence of deep learning, scholars began to train sample images through deep networks to learn facial expression features under big data, thereby increasing the robustness and generalization ability of the model, and further increasing the application of facial expression recognition technology.
[0004] In recent years, facial expression recognition methods are mainly divided into feature design-based methods and feature learning-based methods. Among them, the facial expression recognition method based on feature design mainly designs efficient feature extraction operators to perform robust feature extraction, and then inputs the extracted features into the classifier to classify facial expressions. Well-known feature design operators, such as LBP, Gabor, SIFT, and HOG, are widely used in facial expression recognition. However, the facial expression recognition method based on feature design is not robust enough when texture turbulence changes in the expression-significant area, and has weak generalization and is insensitive to gradient changes of the face. The recognition performance of the method based on feature design depends largely on the quality of the designed feature extraction operator. The limitations of the feature extraction operator will lead to poor performance in practical applications.
[0005] 2012 is known as the "first year of deep learning", which brought a wave of deep learning craze. Due to the significant improvement of chip processing power, deep networks have become a research hotspot due to their powerful learning ability. Well-designed network frameworks have achieved the most advanced recognition performance in facial expression recognition tasks, and have greatly exceeded the recognition performance of feature learning methods. A series of classic networks have been proposed by scholars, such as CNN, DBN, RNN and GAN, which have achieved good performance. However, deep networks usually treat the input image as a whole, and cannot distinguish the significant areas of facial expressions, and cannot tilt resources to the significant areas of expression.
[0006] In order to alleviate the above shortcomings, the attention mechanism has been widely used in the field of computer vision by imitating the human visual mechanism and focusing more resources on some information, which helps to assist judgment. Secondly, many methods combine feature design methods with feature learning methods to enhance the robustness of features while enhancing the representation ability of features. Inspired by the above methods, this method proposes a facial expression recognition method based on regional self-attention convolutional neural network to solve the above problems.
[0007] After searching, the application publication number CN114170666A, a method for facial expression recognition based on regional self-attention convolutional neural network, relates to the field of computer vision, and the technical scheme of the present invention includes 1) extraction of deep global features, extracting deep global features of facial expression images through VGG network; 2) designing a regional texture enhancement module, extracting multi-modal information and grayscale difference information of facial expression images respectively through the designed regional local multi-valued mode and improved K-means algorithm, and enhancing regional texture features; 3) designing a regional self-attention module, quantifying the contribution of different regions to expression recognition through regional self-attention mechanism and rank regularization loss, and constraining the weights of different regions, so that the weight values of different regions are more distinguishable. 4) Using a large-scale facial expression data set to train an attention convolutional neural network, a facial expression recognition model is obtained. The local features and deep global features of the expression-significant region can be extracted at the same time, and the influence weight of the expression-significant region on facial expression recognition is increased, and the performance of facial expression recognition is improved by enhancing the robustness of texture features under expression turbulence transformation. Summary of the invention
[0008] The present invention aims to solve the above problems of the prior art. A method for facial expression recognition based on regional self-attention convolutional neural network is proposed. The technical solution of the present invention is as follows:
[0009] A method for facial expression recognition based on regional self-attention convolutional neural network, comprising the following steps:
[0010] Step 1: Input the original expression image into the feature extraction network based on VGG16 to extract the deep global features of the input expression image;
[0011] Step 2: Design a regional local multi-value pattern, input the original expression image into the regional local multi-value pattern to enhance the regional texture; wherein, the regional local multi-value pattern uses an improved K-means algorithm to dynamically cluster pixels. In the improved K-means algorithm, the distance from each data point to the origin is first calculated. Then, the original data points are sorted according to the sorted distances, and the sorted data points are divided into k equal sets, with the middle point of each group as the initial centroid. These initial centroids can obtain better unique clustering results. The improved K-means algorithm can ensure the robustness of the expression change regional features, expand the binary pattern to k patterns, integrate the grayscale difference information between pixels in the region, and enhance the regional texture features;
[0012] Step 3: Input the enhanced regional texture features into the regional self-attention module, which mainly includes the regional self-attention mechanism and rank regularization loss. The regional self-attention mechanism enhances the weights of the features of the regions with significant expression changes, quantifies the contribution of different regions to expression recognition, and obtains the enhanced regional texture attention features. The rank regularization loss is used to constrain the weights of different regions to make the weight values of different regions more distinguishable.
[0013] Step 4: Fusion the regional weighted features extracted in step 3 with the deep global features extracted by the VGG network.
[0014] Furthermore, the step 1 inputs the original expression image into a feature extraction network based on VGG16 to extract the deep global features of the input expression image, specifically including:
[0015] A1: The facial expression image is detected by the face detection and alignment network MTCNN to obtain facial key points, and the face image is aligned and cropped into an input image I of size 224×224;
[0016] A2: Input image I into the VGG16 network to extract features, and use F g If F g It can be defined as:
[0017] F g =γ(I;θ) (11)
[0018] Where γ(;) is the backbone network, θ is the parameter in the backbone network, F g It is the deep global feature extracted by the backbone network.
[0019] Furthermore, the step 2: designing a regional local multi-valued model, inputting the original expression image into the regional local multi-valued model to enhance the regional texture, specifically comprises the following steps:
[0020] B1: For the input facial expression image, it is evenly cropped into a 3×3 facial expression image area;
[0021] B2: For each region, define the difference m between its gray value and the mean value of the local neighborhood pixels i , and then use the difference as the new pixel map M enhance , defined as follows:
[0022]
[0023]
[0024] Where P c Represents the central pixel value of the pixel map, P i Indicates that Pc Adjacent pixel values; represents the local neighborhood pixel mean, P represents the set of surrounding sampled pixels, and i represents the index of the surrounding sampled pixel set.
[0025] B3: The enhanced feature map M enhance The enhanced pixels are stored in array a and divided into k equal parts to obtain a 1 ,a 2 ,…,a k , define the center value of each class as the calibration point, calculate the distance from each pixel to the calibration point; group the closest pixels into a class, calculate the mean of the pixels in the class, and use the mean as the new calibration point, and finally get the last k calibration points through iteration;
[0026] B4: Binarize the pixel values of each layer to obtain k patterns, and cascade these patterns to obtain a robust feature representation F for each region r .
[0027] Furthermore, the step 3: inputs the enhanced regional texture features into the regional self-attention module, which mainly includes the regional self-attention mechanism and the rank regularization loss. The regional self-attention mechanism enhances the weights of the features of the regions with significant expression changes, quantifies the contributions of different regions to expression recognition, and obtains the enhanced regional texture attention features. The rank regularization loss is used to constrain the weights of different regions, making the weight values of different regions more distinguishable. Specifically, the following steps are included:
[0028] C1: The robust feature representation F obtained in step B4 r Input to the dimensionality reduction convolutional neural network to obtain the depth feature map of each region, and define the input regional texture image as I 1 ,I 2 ,…,I 9 , the definition of the dimensionality reduction convolutional neural network is as follows:
[0029] X=[F 1 ,F 2 ,…,F 9 ]=[V(I 1 ;θ),V(I 2 ;θ),…,V(I 9 ;θ)] (14)
[0030] Where V(·; θ) is a dimensionality reduction convolutional neural network, θ is a parameter in the dimensionality reduction convolutional neural network, and X is a set of regional features extracted by the dimensionality reduction convolutional neural network;
[0031] C2: In order to obtain the contribution of each region in the facial expression recognition task, the self-attention mechanism is used to obtain the weight of each region, and the rough weight of the feature is calculated by FC and Sigmoid function, which is defined as follows:
[0032] W=[a 1 ,a 2 ,…,a k ]=[f(F 1 T q),f(F 2 T q),…,f(F 9 T q)] (15)
[0033] where a i represents the weight of the i-th region, f represents the Sigmoid function, q represents the parameter of the fully connected layer, and all local features with attention weights are summarized as an overall representation F m , which is defined as follows:
[0034]
[0035] Here, s represents the cascade operation between feature blocks;
[0036] C3: Use rank regularization loss RRLoss to constrain the weights of different regions; first, sort the weights of different regions, and then divide them into high-weight groups and low-weight groups according to a certain ratio; secondly, calculate the average weights of the high-weight and low-weight groups, respectively, using a high and a low express.
[0037] A difference value M is added to RRLoss to limit the average weight of these groups, which is defined as follows:
[0038]
[0039]
[0040] L RR =max{0,M-(a high -a low )} (19)
[0041] G high and G low represents the weight mean of the high-weight group and the low-weight group, λ represents the proportion of the high-weight group, and N represents the number of regions. M is a difference value, which can be a fixed learnable parameter or a hyperparameter, L RR Used to enhance the weights of regional attention, encouraging the network to prioritize areas of expression changes during training.
[0042] Furthermore, the step 4: fusing the regional weighted features extracted in step 3 with the deep global features extracted by the VGG network. Specifically, it includes the following steps:
[0043] D1: The enhanced regional texture attention feature F extracted from step C2 m With the deep global feature F g To carry out effective integration;
[0044] D2: The number of channels is added through the Concat operation, and the features used to describe the image itself are increased. Its definition is as follows:
[0045] F = concat(F m ,F g ) (20)
[0046] The fused feature F is then sent to the classifier for expression recognition.
[0047] The advantages and beneficial effects of the present invention are as follows:
[0048] The method of the present invention first proposes a new facial expression recognition method, called a regional self-attention convolutional neural network. The network consists of three parts: a feature extraction module, a regional texture enhancement module and a regional self-attention module. The VGG16 network is used to extract the deep global features of the original image. Then, the regional texture enhancement module is used to enhance the texture features of the expression-significant area, wherein the improved K-means algorithm is used to dynamically cluster the regional pixel difference information, and the multi-scale texture information is supplemented by the designed regional local multi-value pattern. In addition, the regional self-attention mechanism is used to calculate the weight contributed by each region in expression recognition, and finally the rank regularization loss is used to constrain the weight. The main advantages and beneficial effects are as follows:
[0049] 1. A facial expression recognition method based on a new regional self-attention convolutional neural network is designed. Different from previous methods, this method fully utilizes the regional texture information and deep global information of the face through the effective fusion of deep global features and regional weighted features, enhances the robustness of the features, enhances the expressiveness of the features of the expression-significant regions, and improves the performance of facial expression recognition.
[0050] 2. The regional texture enhancement module designed by the present invention makes full use of the grayscale difference information between pixels by combining the regional local multi-valued pattern and the improved K-means algorithm, thereby enhancing the representation ability of regional texture. Secondly, the multi-modal information of the region is taken into consideration, and the improved K-means algorithm is cleverly combined with the regional local multi-valued pattern to dynamically cluster the pixels in the region, expand the local binary pattern into multiple patterns, and improve the robustness of the feature to the interference of expression changes.
[0051] 3. The present invention designs a regional self-attention module to adaptively quantify the contribution of different regions to the facial expression recognition task, and introduces a rank regularization loss to constrain the weight of each region, thereby enhancing the influence weight of the expression-significant region on facial expression recognition, and improving the performance of facial expression recognition by enhancing the robustness of texture features under expression turbulence changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is an overall framework diagram of the method of the preferred embodiment provided by the present invention. DETAILED DESCRIPTION
[0053] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.
[0054] The technical solution of the present invention to solve the above technical problems is:
[0055] As attached Figure 1 As shown, a facial expression recognition method based on regional self-attention convolutional neural network includes the following steps:
[0056] 1. As attached Figure 1 As shown, the original expression image is input into the feature extraction network based on VGG16 to extract the deep global features of the input expression image, including:
[0057] A1: The facial expression image is detected by the face detection and alignment network MTCNN to obtain facial key points, and the face image is aligned and cropped into an input image I of size 224×224;
[0058] A2: Input image I into the VGG16 network to extract features, and use F g If F g It can be defined as:
[0059] F g =γ(I;θ) (21)
[0060] Where γ(;) is the backbone network, θ is the parameter in the backbone network, F g It is the deep global feature extracted by the backbone network.
[0061] 2. As attached Figure 1 As shown, a regional local multi-value pattern is designed, and the original expression image is input into the regional local multi-value pattern to enhance the regional texture, which specifically includes the following steps:
[0062] B1: For the input facial expression image, it is evenly cropped into a 3×3 facial expression image area;
[0063] B2: For each region, define the difference m between its gray value and the mean value of the local neighborhood pixels i , and then use the difference as the new pixel map M enhance , defined as follows:
[0064]
[0065]
[0066] Where P c Represents the central pixel value of the pixel map, P i Indicates that P c Adjacent pixel values; represents the local neighborhood pixel mean, P represents the set of surrounding sampled pixels, and i represents the index of the surrounding sampled pixel set.
[0067] B3: The enhanced feature map M enhance The enhanced pixels are stored in array a and divided into k equal parts to obtain a 1 ,a 2 ,…,a k , define the center value of each class as the calibration point, calculate the distance from each pixel to the calibration point; group the closest pixels into a class, calculate the mean of the pixels in the class, and use the mean as the new calibration point, and finally get the last k calibration points through iteration;
[0068] B4: Binarize the pixel values of each layer to obtain k patterns, and cascade these patterns to obtain a robust feature representation F for each region r .
[0069] 3. As attached Figure 1 As shown in the figure, the enhanced regional texture features are input into the regional self-attention module. The regional self-attention module mainly includes the regional self-attention mechanism and the rank regularization loss. The regional self-attention mechanism enhances the weights of the features of the regions with significant expression changes, quantifies the contributions of different regions to expression recognition, and obtains the enhanced regional texture attention features. The rank regularization loss is used to constrain the weights of different regions, making the weight values of different regions more distinguishable. Specifically, it includes the following steps:
[0070] C1: The robust feature representation F obtained in step B4 r Input to the dimensionality reduction convolutional neural network to obtain the depth feature map of each region, and define the input regional texture image as I 1 ,I 2 ,…,I 9 , the definition of the dimensionality reduction convolutional neural network is as follows:
[0071] X=[F1 ,F 2 ,…,F 9 ]=[V(I 1 ;θ),V(I 2 ;θ),…,V(I 9 ;θ)] (24)
[0072] Where V(·; θ) is a dimensionality reduction convolutional neural network, θ is a parameter in the dimensionality reduction convolutional neural network, and X is a set of regional features extracted by the dimensionality reduction convolutional neural network;
[0073] C2: In order to obtain the contribution of each region in the facial expression recognition task, the self-attention mechanism is used to obtain the weight of each region, and the rough weight of the feature is calculated by FC and Sigmoid function, which is defined as follows:
[0074] W=[a 1 ,a 2 ,…,a k ]=[f(F 1 T q),f(F 2 T q),…,f(F 9 T q)] (25)
[0075] where a i represents the weight of the i-th region, f represents the Sigmoid function, q represents the parameter of the fully connected layer, and all local features with attention weights are summarized as an overall representation F m , which is defined as follows:
[0076]
[0077] Here, s represents the cascade operation between feature blocks;
[0078] C3: Use rank regularization loss RRLoss to constrain the weights of different regions; first, sort the weights of different regions, and then divide them into high-weight groups and low-weight groups according to a certain ratio; secondly, calculate the average weights of the high-weight and low-weight groups, respectively, using a high and a low express.
[0079] A difference value M is added to RRLoss to limit the average weight of these groups, which is defined as follows:
[0080]
[0081]
[0082] L RR=max{0,M-(a high -a low )} (29)
[0083] G high and G low represents the weight mean of the high-weight group and the low-weight group, λ represents the proportion of the high-weight group, and N represents the number of regions. M is a difference value, which can be a fixed learnable parameter or a hyperparameter, L RR Used to enhance the weights of regional attention, encouraging the network to prioritize areas of expression changes during training.
[0084] 4. As attached Figure 1 As shown in Figure 1, the regional weighted features extracted in step 3 are fused with the deep global features extracted by the VGG network. The specific steps are:
[0085] D1: The enhanced regional texture attention feature F extracted from step C2 m With the deep global feature F g To integrate effectively.
[0086] D2: The number of channels is added through the Concat operation, and the features used to describe the image itself are increased. Its definition is as follows:
[0087] F = concat(F m ,F g ) (30)
[0088] The fused feature F is then sent to the classifier for expression recognition.
[0089] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0090] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A facial expression recognition method based on regional self-attention convolutional neural network, characterized in that: The following steps are involved: Step 1: Input the original expression image into the feature extraction network based on VGG16 to extract the deep global features of the input expression image; Step 2: Design a regional local multi-value pattern, input the original expression image into the regional local multi-value pattern to enhance the regional texture; wherein, the regional local multi-value pattern uses an improved K-means algorithm to dynamically cluster pixels; in the improved K-means algorithm, firstly calculate the distance from each data point to the origin; then, sort the original data points according to the sorted distance, divide the sorted data points into k equal sets, and use the middle point in each group as the initial centroid; these initial centroids obtain better unique clustering results; the improved K-means algorithm can ensure the robustness of the expression change regional features, expand the binary pattern to k patterns, integrate the grayscale difference information between pixels in the region, and enhance the regional texture features; Step 3: Input the enhanced regional texture features into the regional self-attention module, which includes the regional self-attention mechanism and rank regularization loss. The regional self-attention mechanism enhances the weights of the features of the regions with significant expression changes, quantifies the contributions of different regions to expression recognition, and obtains the enhanced regional texture attention features; while the rank regularization loss is used to constrain the weights of different regions to make the weight values of different regions more distinguishable; Step 4: Fusion the regional weighted features extracted in step 3 with the deep global features extracted by the VGG network; The step 3 inputs the enhanced regional texture features into the regional self-attention module, which includes a regional self-attention mechanism and a rank regularization loss. The regional self-attention mechanism enhances the weights of the features of the regions with significant expression changes, quantifies the contributions of different regions to expression recognition, and obtains enhanced regional texture attention features; and the rank regularization loss is used to constrain the weights of different regions to make the weight values of different regions more distinguishable; specifically, the following steps are included: C1: The robust feature representation F obtained in step B4 r Input to the dimensionality reduction convolutional neural network to obtain the deep feature map of each region. The input regional texture image is defined as I1, I2, …, I9. The definition of the dimensionality reduction convolutional neural network is as follows: X=[F1,F2,…,F9]=[V(I1;θ),V(I2;θ),…,V(I9;θ)] (4) Where V(·; θ) is a dimensionality reduction convolutional neural network, θ is a parameter in the dimensionality reduction convolutional neural network, and X is a set of regional features extracted by the dimensionality reduction convolutional neural network; C2: In order to obtain the contribution of each region in the facial expression recognition task, the self-attention mechanism is used to obtain the weight of each region, and the rough weight of the feature is calculated by FC and Sigmoid function, which is defined as follows: W=[a1,a2,…,a k ]=[f(F1 T q),f(F2 T q),…,f(F9 T q)] (5) where a i represents the weight of the i-th region, f represents the Sigmoid function, q represents the parameter of the fully connected layer, and all local features with attention weights are summarized as an overall representation F m , which is defined as follows: Here, s represents the cascade operation between feature blocks; C3: Use rank regularization loss RRLoss to constrain the weights of different regions; first, sort the weights of different regions, and then divide them into high-weight groups and low-weight groups according to a certain ratio; secondly, calculate the average weights of the high-weight and low-weight groups, respectively, using a high and a low express; A difference value M is added to RRLoss to limit the average weight of these groups, which is defined as follows: L RR =max{0,M-(a high -a low )} (9) G high and G low They represent the weight means of the high-weight group and the low-weight group, λ represents the proportion of the high-weight group, N represents the number of regions; M is a difference value, which is a fixed learnable parameter or hyperparameter, L RR Used to enhance the weights of regional attention, encouraging the network to prioritize areas of expression changes during training.
2. A facial expression recognition method based on regional self-attention convolutional neural network according to claim 1, characterized in that: The step 1 inputs the original expression image into a feature extraction network based on VGG16 to extract the deep global features of the input expression image, specifically including: A1: The facial expression image is detected by the face detection and alignment network MTCNN to obtain facial key points, and the face image is aligned and cropped into an input image I of size 224×224; A2: Input image I into the VGG16 network to extract features, and use F g If F g Defined as: F g =γ(I;θ) (1) Where γ(;) is the backbone network, θ is the parameter in the backbone network, F g It is the deep global feature extracted by the backbone network.
3. A facial expression recognition method based on regional self-attention convolutional neural network according to claim 2, characterized in that: The step 2, designing a regional local multi-valued model, inputting the original expression image into the regional local multi-valued model to enhance the regional texture, specifically comprises the following steps: B1: For the input facial expression image, it is evenly cropped into a 3×3 facial expression image area; B2: For each region, define the difference m between its gray value and the mean value of the local neighborhood pixels i , and then use the difference as the new pixel map M enhance , defined as follows: Where P c Represents the central pixel value of the pixel map, P i Indicates that P c Adjacent pixel values; represents the local neighborhood pixel mean, P represents the set of surrounding sampled pixels, and i represents the index of the sum of the surrounding sampled pixel sets; B3: The enhanced feature map M enhance The enhanced pixels are stored in array a and divided into k equal parts to obtain a1, a2, …, a k , define the center value of each class as the calibration point, calculate the distance from each pixel to the calibration point; group the closest pixels into a class, calculate the mean of the pixels in the class, and use the mean as the new calibration point, and finally get the last k calibration points through iteration; B4: Binarize the pixel values of each layer to obtain k patterns, and cascade these patterns to obtain a robust feature representation F for each region r .
4. A method for facial expression recognition based on regional self-attention convolutional neural network according to claim 3, characterized in that: The step 4, fusing the regional weighted features extracted in step 3 with the deep global features extracted by the VGG network, specifically comprises the following steps: D1: The enhanced regional texture attention feature F extracted from step C2 m With the deep global feature F g To carry out effective integration; D2: The number of channels is added through the Concat operation, and the features used to describe the image itself are increased. Its definition is as follows: F=concat(F m ,F g ) (10) The fused feature F is then sent to the classifier for expression recognition.
Citation Information
Patent Citations
Multi-branch feature fusion remote sensing scene image classification method based on attention mechanism
CN112861978A
Facial expression recognition method based on multi-region convolutional neural network
CN114170666A