Expression recognition method and device
Through adaptive area processing and self-attention coding technology, local and global features in the expression recognition model are integrated, and the problem of inaccurate expression recognition in complex scenarios is solved, improving the accuracy and robustness of expression recognition.
Patent Information
- Application Number
- CN202311837227.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-22
AI Technical Summary
Existing expression recognition technology fails to effectively consider the relationship between different local areas in complex scenarios, resulting in inaccurate recognition results.
Through feature extraction, adaptive area processing, self-attention coding and multiplication processing, different scale features and global expression features in the image are captured, important area information is integrated, and the accuracy of expression recognition is improved.
It enhances the accuracy and robustness of expression recognition, and can more accurately recognize facial expressions in complex scenarios.
Smart Images

Figure CN120356247A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method and device for facial expression recognition. Background Art
[0002] With the development of deep learning technologies and the popularization of intelligent devices, image recognition technologies have achieved unprecedented rapid development. As an important part of image recognition, facial expression recognition is an important part for a computer to understand human emotions and an important content of human-computer interaction. Therefore, facial expression recognition tasks have received increasing attention. In recent years, facial expression recognition tasks have received extensive attention and applications in the fields of automation, medicine, communication, and intelligent driving. Facial expression recognition refers to predicting the emotions and psychological changes of the people in static photos or video sequences. Most facial expression recognition tasks classify the facial expression images of people and divide them into different facial expression categories. Existing facial expression recognition algorithms extract global facial expression features through a feature extraction network and classify them based on the global facial expression features. However, the differences between different facial expressions are often reflected in local regions in many cases. The recognition effect of obtaining facial expression category results based on global facial expression features is not good in many complex scenarios. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a method and device for facial expression recognition to solve the problem in the prior art that the relationship between different local regions is not considered in image recognition in complex scenarios, resulting in inaccurate facial expression recognition results.
[0004] In a first aspect of embodiments of the present disclosure, a method for facial expression recognition is provided, including: obtaining an image to be recognized, where the image to be recognized includes a facial image of an object; performing feature extraction on the image to be recognized to obtain a feature map of the image to be recognized; performing adaptive region processing on the feature map of the image to be recognized to obtain a weight feature map of the image to be recognized; determining a first fusion feature map of the image to be recognized according to the feature map of the image to be recognized; performing self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain a global facial expression feature map of the image to be recognized; performing multiplication processing on the weight feature map of the image to be recognized and the global facial expression feature map of the image to be recognized to obtain an enhanced feature map of the image to be recognized; and determining a facial expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
[0005] In a second aspect of the embodiments of the present disclosure, an apparatus for facial expression recognition is provided, including: an acquisition module configured to acquire an image to be recognized, where the image to be recognized includes a facial image of an object; a feature extraction module configured to perform feature extraction on the image to be recognized to obtain a feature map of the image to be recognized; an adaptive processing module configured to perform adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized; a first fusion processing module configured to determine a first fusion feature map of the image to be recognized according to the feature map of the image to be recognized; a self-attention processing module configured to perform self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain a global facial expression feature map of the image to be recognized; a second fusion processing module configured to perform multiplication processing on the weighted feature map of the image to be recognized and the global facial expression feature map of the image to be recognized to obtain an enhanced feature map of the image to be recognized; and a recognition module configured to determine a facial expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
[0006] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.
[0007] In a fourth aspect of the embodiments of the present disclosure, a readable storage medium is provided, where the readable storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0008] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: Feature extraction is performed on the image to be recognized to obtain the feature map of the image to be recognized, and adaptive region processing is performed based on the feature map of the image to be recognized, adaptively learning regional features of different scales. Multiple convolution operations are performed on the feature map of the image to be recognized, so as to obtain regional features at different scales. According to the importance of the above regional features, the weight feature map of the image to be recognized is adaptively learned and obtained, capturing features of different scales in the image to be recognized, and the feature information in the image to be recognized can be represented more comprehensively. According to the feature map of the image to be recognized, the first fusion feature map of the image to be recognized is determined. The first fusion feature map of the image to be recognized contains the semantic features and position features of multiple sub-images corresponding to the image to be recognized. Self-attention encoding processing is performed on the first fusion feature map of the image to be recognized to capture the global features in the image, obtaining the global expression feature map of the image to be recognized, and more accurately representing the content of the image to be recognized and the expression of the object in the image to be recognized. Multiplying the weight feature map of the image to be recognized and the global expression feature map of the image to be recognized can pay more attention to the effective region, fuse the information of the weight feature map of the image to be recognized and the global expression feature map of the image to be recognized, and emphasize the effective region, obtaining the enhanced feature map of the image to be recognized. Based on the enhanced feature map of the image to be recognized, expression recognition is performed to obtain the expression recognition result, improving the accuracy and robustness of expression recognition, and solving the problem that the expression recognition result is not accurate enough in the prior art due to the failure to consider the relationship between different local regions in image recognition in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure;
[0011] Figure 2 is a schematic flowchart of a method for expression recognition provided by the embodiments of the present disclosure;
[0012] Figure 3 is a schematic flowchart of another method for expression recognition provided by the embodiments of the present disclosure;
[0013] Figure 4 is a schematic flowchart of still another method for expression recognition provided by the embodiments of the present disclosure;
[0014] Figure 5It is a schematic structural diagram of a facial expression recognition device provided by an embodiment of the present disclosure;
[0015] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0016] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0017] A method and a device for facial expression recognition according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0018] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the present disclosure. The application scenario may include terminal devices 1, 2, and 3, a server 4, and a network 5.
[0019] The terminal devices 1, 2, and 3 may be hardware or software. When the terminal devices 1, 2, and 3 are hardware, they may be various electronic devices having a display screen and supporting communication with the server 4, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.; when the terminal devices 1, 2, and 3 are software, they may be installed in the above-mentioned electronic devices. The terminal devices 1, 2, and 3 may be implemented as multiple software or software modules, or may also be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications may be installed on the terminal devices 1, 2, and 3, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.
[0020] The server 4 may be a server providing various services. For example, it is a background server that receives requests sent by terminal devices with which it establishes a communication connection. The background server may receive and analyze requests sent by terminal devices and generate processing results. The server 4 may be a single server, or may be a server cluster composed of several servers, or may also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.
[0021] It should be noted that the server 4 can be hardware or software. When the server 4 is hardware, it can be various electronic devices that provide various services for the terminal devices 1, 2, and 3. When the server 4 is software, it can be multiple software or software modules that provide various services for the terminal devices 1, 2, and 3, or it can be a single software or software module that provides various services for the terminal devices 1, 2, and 3. The embodiments of the present disclosure do not limit this.
[0022] The network 5 can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or it can be a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), Infrared, etc. The embodiments of the present disclosure do not limit this.
[0023] The user can establish a communication connection with the server 4 via the network 5 through the terminal devices 1, 2, and 3 to receive or send information, etc. Specifically, the server 4 obtains the image to be recognized from the terminal device 1, 2, or 3, and the image to be recognized includes the facial image of the object; extracts features from the image to be recognized to obtain the feature map of the image to be recognized; performs adaptive region processing on the feature map of the image to be recognized to obtain the weighted feature map of the image to be recognized; determines the first fusion feature map of the image to be recognized according to the feature map of the image to be recognized; performs self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain the global expression feature map of the image to be recognized; multiplies the weighted feature map of the image to be recognized and the global expression feature map of the image to be recognized to obtain the enhanced feature map of the image to be recognized; determines the expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
[0024] It should be noted that the specific types, quantities, and combinations of the terminal devices 1, 2, and 3, the server 4, and the network 5 can be adjusted according to the actual requirements of the application scenario. The embodiments of the present disclosure do not limit this.
[0025] Figure 2 It is a schematic flowchart of a method for expression recognition provided by an embodiment of the present disclosure. Figure 2 The expression recognition method can be executed by Figure 1 the terminal device or the server of Figure 2 As shown in
[0026] Step 201, obtain the image to be recognized, where the image to be recognized includes the facial image of the object.
[0027] In some embodiments, the image to be recognized includes facial images of at least one object. In specific applications, in order to improve the accuracy of expression recognition, the image to be recognized can be selectively preprocessed, including but not limited to cropping and rotating the image to be recognized, so that the preprocessed image to be recognized meets the requirements of the input expression recognition model.
[0028] Step 202: Extract features from the image to be recognized to obtain a feature map of the image to be recognized.
[0029] In some embodiments, the expression recognition model includes a feature extraction network. The image to be recognized is input into the expression recognition model, and the feature extraction network extracts features from the image to be recognized to obtain a feature map of the image to be recognized. The feature extraction network can be ResNet50. ResNet50 is a deep learning model that contains 4 stages, and each stage includes a stack of multiple residual structures. The above-mentioned residual structures can effectively solve the problem of gradient disappearance in deep neural networks. By using ResNet50 to extract features from the image to be recognized, complex image features in the image to be recognized can be captured, and a feature map of the image to be recognized can be obtained, which can be used as the input for subsequent classification and other tasks, thereby realizing the effective recognition and analysis of the image to be recognized and improving the accuracy of the expression recognition task.
[0030] Step 203: Perform adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized.
[0031] In some embodiments, adaptive region processing is a region-based method. By performing convolutions on the feature map of the image to be recognized at different scales multiple times, multiple convolution results are obtained, and the multiple convolution results are added together to fuse region features at different scales, and the fusion result is normalized to obtain a weighted feature map of the image to be recognized. The above-mentioned weighted feature map is an adaptive region attention map, which can reflect the attention degree of the feature map of the image to be recognized at different scales, can further enhance the expression ability of the features and be used in subsequent classification tasks to improve the accuracy of expression recognition.
[0032] Step 204: Determine a first fusion feature map of the image to be recognized according to the feature map of the image to be recognized.
[0033] In some embodiments, for the feature map of the image to be recognized, it is sliced into multiple sub-feature maps. Each sub-feature map represents a part of the feature map of the image to be recognized and can provide some local feature information of the image to be recognized. Then, tensor flattening processing is performed on each sub-feature map to obtain the semantic feature vector of each sub-feature map. Through tensor flattening processing, a feature vector can be obtained, which is beneficial to subsequent self-attention encoding processing. According to the position of each sub-feature map in the corresponding feature map, position encoding is performed on each sub-feature map to obtain the position feature vector of each corresponding sub-feature map. The semantic feature vector of each sub-feature map is fused with the corresponding position feature vector of each sub-feature map to fuse the position information and semantic information of each sub-feature map, obtaining the fused feature vector of each sub-feature map. Then, the fused feature vectors of each sub-feature map are concatenated. The fused feature vectors of the above sub-feature maps are concatenated in the order of the original sub-feature maps to obtain the first fused feature map of the image to be recognized. Determining the first fused feature map of the image to be recognized based on the feature map of the image to be recognized can extract and retain important local and global information in the image to be recognized, providing a richer feature representation for subsequent facial expression recognition tasks. At the same time, by using multiple sub-feature maps and their corresponding position information, the robustness of the facial expression recognition method can be improved.
[0034] Step 105: Perform self-attention encoding processing on the first fused feature map of the image to be recognized to obtain the global facial expression feature map of the image to be recognized.
[0035] In some embodiments, the facial expression recognition model further includes a Transformer encoder. The first fused feature map of the image to be recognized is input into the Transformer encoder of the facial expression recognition model for self-attention encoding processing. The Transformer encoder processes the input first fused feature map of the image to be recognized and generates a set of weight coefficients, and performs weighted summation on the input feature vectors according to the weight coefficients. Through self-attention encoding processing, the model can learn the key features related to facial expressions in the image to be recognized, assign different weights to them, obtain the local facial expression vectors in each image to be recognized, and concatenate the local facial expression vectors in the image to be recognized according to their positions to obtain the global facial expression feature map of the image to be recognized, which helps to more comprehensively understand the content of the image to be recognized and improve the robustness and accuracy of the facial expression recognition model.
[0036] Step 106: Multiply the weight feature map of the image to be recognized by the global facial expression feature map of the image to be recognized to obtain the enhanced feature map of the image to be recognized.
[0037] In some embodiments, the weighted feature map of the image to be recognized and the global expression feature map of the image to be recognized are multiplied, that is, each element of the weighted feature map of the image to be recognized is subjected to a corresponding multiplication operation with each element of the global expression feature map of the image to be recognized, so as to fuse the information of the weighted feature map of the image to be recognized and the global expression feature map of the image to be recognized, enhance the co-occurring features and weaken the unimportant features, and obtain the enhanced feature map of the image to be recognized. The enhanced feature map of the image to be recognized can pay more attention to more effective regions, contain richer and more accurate expression features, contribute to subsequent expression recognition tasks, and improve the accuracy of expression recognition in complex scenarios.
[0038] Step 107, determine the expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
[0039] In some embodiments, a series of corresponding transformation processes are performed on the enhanced feature map of the image to be recognized, and the expression recognition model predicts the expression category of the object in the image to be recognized based on the enhanced feature map of the image to be recognized. The expression recognition result of the image to be recognized can be the probability of the expression category corresponding to the object in the image to be recognized. For example, the expression category of the object in the image 000001 is predicted to be category A, and the calculated prediction probability of category A is 0.8, the prediction probability of the expression category of the object in the image 000001 being predicted as category B is 0.1, and the prediction probability of the expression category of the object in the image 000001 being predicted as category C is 0.1. Specifically, category A can be "happy", category B can be "fear", and category C can be "angry".
[0040] Through the method for facial expression recognition provided by the present disclosure, feature extraction is performed on the image to be recognized to obtain the feature map of the image to be recognized, and adaptive region processing is performed based on the feature map of the image to be recognized to adaptively learn regional features of different scales. Multiple convolutional operations are performed on the feature map of the image to be recognized, so as to obtain regional features at different scales, and according to the importance of the above regional features, the weight feature map of the image to be recognized is adaptively learned and obtained, capturing features of different scales in the image to be recognized, and can more comprehensively represent the feature information in the image to be recognized. According to the feature map of the image to be recognized, the first fusion feature map of the image to be recognized is determined. The first fusion feature map of the image to be recognized includes semantic features and position features of multiple sub-images corresponding to the image to be recognized. And self-attention encoding processing is performed on the first fusion feature map of the image to be recognized to capture the global features in the image, obtaining the global facial expression feature map of the image to be recognized, which more accurately represents the content of the image to be recognized and the facial expressions of the objects in the image to be recognized. Multiplying the weight feature map of the image to be recognized by the global facial expression feature map of the image to be recognized can pay more attention to the effective region, fuse the information of the weight feature map of the image to be recognized and the global facial expression feature map of the image to be recognized together, and emphasize the effective region, obtaining the enhanced feature map of the image to be recognized, and performing facial expression recognition based on the enhanced feature map of the image to be recognized to obtain the facial expression recognition result, improving the accuracy and robustness of facial expression recognition, and solving the problem that in the prior art, the relationship between different local regions is not considered in image recognition in complex scenarios, resulting in inaccurate facial expression recognition results.
[0041] In some embodiments, performing adaptive region processing on the feature map of the image to be recognized to obtain the weight feature map of the image to be recognized includes: performing convolutional processing on the feature map of the image to be recognized through a first convolutional layer to obtain the first convolutional result of the image to be recognized; performing convolutional processing on the feature map of the image to be recognized through a first dilated convolutional layer to obtain the first dilated convolutional result of the image to be recognized; performing convolutional processing on the feature map of the image to be recognized through a second dilated convolutional layer to obtain the second dilated convolutional result of the image to be recognized, where the dilation rates of the first dilated convolutional layer and the second dilated convolutional layer are different; performing fusion processing on the first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized to obtain the weight feature map of the image to be recognized.
[0042] In some embodiments, the facial expression recognition model further includes a first convolutional layer, a first dilated convolutional layer, and a second dilated convolutional layer. The first convolutional layer can be a convolutional layer with various convolutional kernels. The convolutional kernels of the first dilated convolutional layer and the second dilated convolutional layer can be the same. For example, the first convolutional layer can be a 3x3 convolutional layer, the first dilated convolutional layer can be a 3x3 convolutional layer with a dilation rate of 1, and the second dilated convolutional layer can be a 3x3 convolutional layer with a dilation rate of 3. The feature maps of the image to be recognized are respectively input into the first convolutional layer, the first dilated convolutional layer, and the second dilated convolutional layer for convolutional processing to obtain the first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized. While keeping the sizes of the first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized consistent, the receptive field is increased, thereby extracting large features of different scales. In the above convolutional processing process, the first convolutional layer uses a 3x3 convolutional kernel with a relatively small receptive field, the first dilated convolutional layer uses a 3x3 convolutional kernel with a dilation rate of 1 and a slightly larger receptive field, and the second dilated convolutional layer uses a 3x3 convolutional kernel with a dilation rate of 3 and an even larger receptive field. By extracting features of different scales, the details and structural information in the image to be recognized can be better captured, improving the generalization ability and robustness of the facial expression recognition model.
[0043] The first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized are added together to obtain a richer and more comprehensive feature representation, and the result of the fusion processing is normalized to eliminate the scale differences between different features, obtaining the weighted feature map of the image to be recognized. The weighted feature map of the image to be recognized can be used in subsequent classification, recognition, and other tasks to provide a more accurate and robust input.
[0044] In some embodiments, the first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized are subjected to a fusion process to obtain the weighted feature map of the image to be recognized, including: adding the first convolutional result of the image to be recognized, the first dilated convolutional result of the image to be recognized, and the second dilated convolutional result of the image to be recognized to obtain the second fusion feature map of the image to be recognized; performing convolutional processing on the second fusion feature map of the image to be recognized through the second convolutional layer to obtain the second convolutional result of the image to be recognized; normalizing the second convolutional result of the image to be recognized to obtain the weighted feature map of the image to be recognized.
[0045] In some embodiments, the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized are added together to integrate different feature extraction results. The feature maps of the image to be recognized from different convolutional layers are added together to fuse the information between them, obtaining a more comprehensive and rich feature representation (i.e., the second fused feature map of the image to be recognized). The second convolutional layer can be a convolutional layer with various convolutional kernels. For example, the second convolutional layer can be a 1x1 convolutional layer. The second fused feature map of the image to be recognized is convolved through the second convolutional layer to further extract and integrate the information in the second fused feature map of the image to be recognized, obtaining the second convolution result of the image to be recognized. And the second convolution result of the image to be recognized is normalized. The sigmoid function can be used for normalization to eliminate the scale differences between different features, obtaining the weighted feature map of the image to be recognized. The weighted feature map of the image to be recognized can be used in subsequent classification, recognition, and other tasks to provide a more accurate and robust input, improving the accuracy of facial expression recognition.
[0046] Reference Figure 3 , the facial expression recognition model further includes an adaptive region processing module 300. The adaptive region processing module 300 includes a first convolutional layer 301, a first dilated convolutional layer 302, a second dilated convolutional layer 303, an addition processing module 304, a second convolutional layer 305, and a normalization processing module 306. The feature map of the image to be recognized is input into the first convolutional layer 301 for convolution processing to obtain the first convolution result of the image to be recognized. The feature map of the image to be recognized is input into the first dilated convolutional layer 302 for convolution processing to obtain the first dilated convolution result of the image to be recognized. The feature map of the image to be recognized is input into the second dilated convolutional layer 303 for convolution processing to obtain the second dilated convolution result of the image to be recognized. The first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized are input into the addition processing module 304 for addition processing to obtain the second fused feature map of the image to be recognized. The second fused feature map of the image to be recognized is input into the second convolutional layer 305 for convolution processing to obtain the second convolution result of the image to be recognized. And the second convolution result of the image to be recognized is input into the normalization processing module 306 for normalization processing to obtain the weighted feature map of the image to be recognized. It can further enhance the feature representation ability and be used in subsequent classification tasks to improve the accuracy of facial expression recognition.
[0047] In some embodiments, determining a first fusion feature map of an image to be recognized according to the feature map of the image to be recognized includes: performing a splitting process on the feature map of the image to be recognized to obtain a plurality of corresponding sub-feature maps of the image to be recognized; performing a flattening process on each sub-feature map of the image to be recognized to obtain semantic feature vectors of each sub-feature map of the image to be recognized; performing position encoding on each sub-feature map of the image to be recognized to obtain position feature vectors of each sub-feature map of the image to be recognized; performing a fusion process on the semantic feature vectors of each sub-feature map of the image to be recognized and the position feature vectors of each sub-feature map of the image to be recognized to obtain fusion feature vectors of each sub-feature map; and performing splicing on the fusion feature vectors of each sub-feature map to obtain the first fusion feature map of the image to be recognized.
[0048] In some embodiments, a sliding window can be used to move on the feature map of the image to be recognized, and each time it moves a preset distance, the feature map of the image to be recognized is split. Before performing the splitting task, the size of each sub-feature map and the moving step of the sliding window are set, and the feature map of the image to be recognized is split into a plurality of sub-feature maps. The flattening process can be performed on each sub-feature map of the image to be recognized through the method of flatten(·) to obtain semantic feature vectors of each sub-feature map of the image to be recognized. At the same time, according to the position of each sub-feature map in the feature map of the image to be recognized, position encoding is performed on each sub-feature map of the image to be recognized. Through the method of position encoding, the position information of each sub-feature map in the original feature map can be captured to obtain position feature vectors of each sub-feature map of the image to be recognized, which helps to understand the overall structure information of the image to be recognized. The relative position method can be used for position encoding. The semantic feature vectors of each sub-feature map of the image to be recognized and the position feature vectors of each sub-feature map of the image to be recognized are added together to combine the semantic feature information and the position feature information to obtain fusion feature vectors of each sub-feature map. The obtained fusion feature vectors of each sub-feature map contain both position information and semantic information. And according to the position of each sub-feature map in the original feature map, a regular splicing process is performed on the fusion feature vectors of each sub-feature map to obtain a global feature representation, that is, the first fusion feature map of the image to be recognized. By performing splitting, flattening, position encoding, addition, and splicing operations on the feature map of the image to be recognized, a richer and more accurate feature representation of the image to be recognized can be obtained, which can improve the performance and robustness of the model and provide a better input for subsequent expression recognition tasks.
[0049] In some embodiments, performing self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain the global expression feature map of the image to be recognized includes: performing self-attention encoding processing on the first fusion feature map of the image to be recognized, calculating the attention weights, performing weighted summation based on the attention weights to obtain the local expression feature maps of the image to be recognized; splicing the local expression feature maps of the image to be recognized based on the positions of the respective sub-feature maps of the image to be recognized to obtain the global expression feature map of the image to be recognized.
[0050] In some embodiments, the self-attention encoding processing is performed on the first fusion feature map of the input image to be recognized through a transformer encoder, the correlation between different positions in the first fusion feature map of the image to be recognized is calculated to assign the attention weights, and based on the calculated attention weights, the first fusion feature map of the image to be recognized is subjected to weighted summation, and the attention weights are applied to each position of the feature map, so as to obtain a set of local expression feature maps of the image to be recognized, and each local expression feature map represents the expression information in a certain sub-feature map of the image to be recognized. And the local expression feature maps of the image to be recognized are spliced based on the positions of the respective sub-feature maps in the feature map of the image to be recognized to obtain the global expression feature map of the image to be recognized. The global expression feature map of the image to be recognized fuses the local information of each sub-feature map and provides a more comprehensive and accurate feature representation for the subsequent expression recognition task. By performing self-attention encoding processing on the first fusion feature map of the image to be recognized and performing weighted summation and splicing operations based on the attention weights, the global expression feature map of the image to be recognized can be obtained, and the accuracy and robustness of expression recognition can be improved.
[0051] In some embodiments, determining the expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized includes: performing dimensionality reduction processing on the enhanced feature map of the image to be recognized through a global average pooling layer to obtain the global dimensionality reduction processing result of the image to be recognized; performing feature transformation on the global dimensionality reduction processing result of the image to be recognized through a fully connected layer to obtain the fully connected feature vector of the image to be recognized; performing classification processing on the fully connected feature vector of the image to be recognized through a classification layer to obtain the expression recognition result of the image to be recognized.
[0052] In some embodiments, the facial expression recognition module further includes a global average pooling layer, a fully connected layer, and a classification layer. The facial expression recognition result of the image to be recognized can be the probability of the facial expression category corresponding to the object in the image to be recognized. The enhanced feature map of the image to be recognized is input into the global average pooling layer of the facial expression recognition model, and the average value is calculated for each position of the enhanced feature map of the image to be recognized, so as to reduce the size of the feature map, obtain the global dimensionality reduction processing result of the image to be recognized, effectively reduce the dimensionality of the feature map, reduce the number of parameters, and improve the computational efficiency and generalization ability of the model. The global dimensionality reduction processing result of the image to be recognized is input into the fully connected layer of the facial expression recognition model for feature transformation, and linear transformation and weighted summation are performed on the global dimensionality reduction processing of the input image to be recognized, so as to map the input features to a new feature space, obtaining a richer and more effective feature representation, that is, the fully connected feature vector of the image to be recognized, providing a more effective and accurate input for the subsequent facial expression recognition task. The fully connected feature vector of the image to be recognized is input into the classification layer of the facial expression recognition model for classification processing. Classification decisions are made based on the input fully connected feature vector of the image to be recognized. The classification layer will learn the relationship between the category information of the data and the features, so as to be able to predict the facial expression category to which the object in the image to be recognized belongs according to the input fully connected feature vector of the image to be recognized.
[0053] Reference Figure 4 , the specific architecture of the facial expression recognition model is described in detail as follows:
[0054] The facial expression recognition model may specifically include a feature extraction network 401, an adaptive region processing module 300, a fusion processing module 402, a self-attention encoding module 403, a multiplication processing module 404, a global average pooling layer 405, a fully connected layer 406, and a classification layer 407. The image to be recognized is input into the feature extraction network 401 for feature extraction to obtain a feature map of the image to be recognized. The feature map of the image to be recognized is input into the adaptive region processing module 300 for adaptive region processing to obtain a weighted feature map of the image to be recognized. The feature map of the image to be recognized is input into the fusion processing module 402, and the feature map of the image to be recognized is segmented to obtain a plurality of corresponding sub-feature maps of the image to be recognized; each sub-feature map of the image to be recognized is flattened to obtain a semantic feature vector of each sub-feature map of the image to be recognized; each sub-feature map of the image to be recognized is positionally encoded to obtain a position feature vector of each sub-feature map of the image to be recognized; the semantic feature vectors of each sub-feature map of the image to be recognized and the position feature vectors of each sub-feature map of the image to be recognized are fused to obtain a fused feature vector of each sub-feature map; the fused feature vectors of each sub-feature map are concatenated to obtain a first fused feature map of the image to be recognized. The first fused feature map of the image to be recognized is input into the self-attention encoding module 403 for self-attention encoding processing to calculate attention weights, and weighted summation is performed based on the attention weights to obtain local facial expression feature maps of the image to be recognized; the local facial expression feature maps of the image to be recognized are concatenated based on the positions of each sub-feature map of the image to be recognized to obtain a global facial expression feature map of the image to be recognized. The weighted feature map of the image to be recognized and the global facial expression feature map of the image to be recognized are input into the multiplication processing module 404 for multiplication processing to obtain an enhanced feature map of the image to be recognized. The enhanced feature map of the image to be recognized is input into the global average pooling layer 405 for dimensionality reduction processing of the enhanced feature map of the image to be recognized to obtain a global dimensionality reduction processing result of the image to be recognized. The global dimensionality reduction processing result of the image to be recognized is input into the fully connected layer 406 for feature transformation of the global dimensionality reduction processing result of the image to be recognized to obtain a fully connected feature vector of the image to be recognized. The fully connected feature vector of the image to be recognized is input into the classification layer 407 for classification processing of the fully connected feature vector of the image to be recognized to obtain a facial expression recognition result of the image to be recognized, improving the accuracy and robustness of facial expression recognition, and solving the problem in the prior art that the relationship between different local regions is not considered in image recognition in complex scenarios, resulting in inaccurate facial expression recognition results.
[0055] In some embodiments, before performing feature extraction on the image to be recognized, it further includes: obtaining an expression recognition training set, where the expression recognition training set includes multiple training images and labels corresponding to each training image; selecting multiple training images from the expression recognition training set according to a first preset percentage for flipping processing to obtain flipped training images corresponding to each training image, and selecting multiple training images from the expression recognition training set according to a second preset percentage for occlusion processing to obtain occluded training images corresponding to each training image; replacing each flipped training image with the corresponding training image, and replacing each occluded training image with the corresponding training image to obtain an enhanced expression recognition training set; inputting each training image in the enhanced expression recognition training set into an expression recognition model to perform feature extraction on each training image in the enhanced expression recognition training set to obtain feature maps of each training image; performing adaptive region processing on the feature maps of each training image to obtain weighted feature maps of each training image; determining fused feature maps of each training image according to the feature maps of each training image; performing self-attention encoding processing on the fused feature maps of each training image to obtain global expression feature maps of each training image; multiplying the global expression feature maps of each training image and the weighted feature maps of each training image to obtain enhanced feature maps of each training image; determining expression recognition results of each image according to the enhanced feature maps of each training image; determining a loss value according to the expression recognition results of each image and the labels of each training image, and updating the parameters in the expression recognition model based on the loss value.
[0056] In some embodiments, the facial expression recognition training set includes multiple training images and labels corresponding to each training image, and the above-mentioned labels are used to indicate the facial expression categories of the training objects in the training images. For the facial expression recognition training set, there can be multiple categories of facial expressions of the training objects in the training images. Specifically, the facial expression category of the training object in training image 0001 can be category A, the facial expression category of the training object in training image 0002 can be category B, and the facial expression category of the training object in training image 0003 can be category C. For example, category A can be "happy", category B can be "fear", category C can be "angry", and so on. The first preset percentage can be 20%, and the second preset percentage can also be 20%. Randomly flip 20% of the training images in the facial expression recognition training set horizontally, which can be achieved by the method of transforms.RandomHorizontalFlip to obtain the flipped training images corresponding to each training image. The method of RandomErasing can be used to randomly occlude 20% of the training images in the facial expression recognition training set to obtain the occluded training images corresponding to each training image. Data augmentation of the facial expression recognition training set is realized by the methods of random flipping and random occlusion. Replace each flipped training image with the corresponding training image, and replace each occluded training image with the corresponding training image to obtain the enhanced facial expression recognition training set, improving the generalization and robustness of the facial expression recognition model.
[0057] The facial expression recognition model includes a feature extraction network, which can be ResNet50. Input each training image in the enhanced facial expression recognition training set into the feature extraction network of the facial expression recognition model for feature extraction, extract useful features from the training images, and obtain the feature maps of each training image. Perform adaptive region processing on the feature maps of each training image. By performing convolutions of different scales on the feature maps of each training image multiple times, obtain multiple convolution results, add the multiple convolution results, fuse the regional features of different scales, and normalize the fusion result to obtain the weighted feature maps of each training image. For the feature maps of each training image, divide them into multiple sub-feature maps. Each sub-feature map represents a part of the feature map of the training image and can provide some local feature information of the training image. Then perform tensor flattening processing on the sub-feature maps of each training image to obtain the semantic feature vectors of each sub-feature map of each training image. And according to the position of each sub-feature map in the corresponding feature map, perform position encoding on each sub-feature map to obtain the position feature vectors of each sub-feature map of each corresponding training image. Fuse the semantic feature vectors of each sub-feature map of each training image with the corresponding position feature vectors of each sub-feature map of each training image, fuse the position information and semantic information of each sub-feature map of each training image, obtain the fusion feature vectors of each sub-feature map of each training image, and splice the fusion feature vectors of each sub-feature map of each training image. Splice the fusion feature vectors of the sub-feature maps of the above-mentioned each training image in the order of the original sub-feature maps to obtain the fusion feature maps of each training image. Perform encoding processing on the fusion feature maps of each training image. Obtain the local facial expression feature maps of each training image, and splice the local facial expression feature maps of each training image based on the positions of each sub-feature map of each training image to obtain the global facial expression feature maps of each training image. Multiply the global facial expression feature maps of each training image by the weighted feature maps of each training image, fuse the information between the global facial expression feature maps of each training image and the weighted feature maps of each training image, enhance the important features, and obtain the enhanced feature maps of each training image. Perform a series of transformation processes on the enhanced feature maps of each training image, which can include global average pooling, fully connected feature transformation, and classification processing, to obtain the facial expression recognition results of each image. And based on the cross-entropy loss function according to the facial expression recognition results of each image and the labels of each training image, determine the loss value, and update the parameters in the facial expression recognition model through the backpropagation algorithm based on the loss value, and reduce the loss value during the training process. When the loss value is less than or equal to the preset value, obtain the trained facial expression recognition model.
[0058] Any combination of the above all optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated here one by one.
[0059] The following are embodiments of the disclosed apparatus, which can be used to execute the method embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.
[0060] Figure 5 It is a schematic diagram of an apparatus for facial expression recognition provided by an embodiment of the present disclosure. As Figure 5 shown, the facial expression recognition apparatus includes:
[0061] An acquisition module 501, configured to acquire an image to be recognized, where the image to be recognized includes a facial image of an object;
[0062] A feature extraction module 502, configured to perform feature extraction on the image to be recognized to obtain a feature map of the image to be recognized;
[0063] An adaptive processing module 503, configured to perform adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized;
[0064] A first fusion processing module 504, configured to determine a first fusion feature map of the image to be recognized according to the feature map of the image to be recognized;
[0065] A self-attention processing module 505, configured to perform self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain a global facial expression feature map of the image to be recognized;
[0066] A second fusion processing module 506, configured to perform a multiplication process on the weighted feature map of the image to be recognized and the global facial expression feature map of the image to be recognized to obtain an enhanced feature map of the image to be recognized;
[0067] A recognition module 507, configured to determine a facial expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
[0068] According to the technical solution provided by the embodiments of the present disclosure, the feature extraction module 502 extracts features from the image to be recognized to obtain the feature map of the image to be recognized, and the adaptive processing module 503 performs adaptive region processing based on the feature map of the image to be recognized, adaptively learns regional features of different scales, performs multiple convolution operations on the feature map of the image to be recognized, so as to obtain regional features at different scales, and adaptively learns and obtains the weight feature map of the image to be recognized according to the importance of the above regional features, captures features of different scales in the image to be recognized, and can more comprehensively represent the feature information in the image to be recognized. The first fusion processing module 504 determines the first fusion feature map of the image to be recognized according to the feature map of the image to be recognized, and the first fusion feature map of the image to be recognized includes the semantic features and position features of multiple sub-images corresponding to the image to be recognized. The self-attention processing module 505 performs self-attention encoding processing on the first fusion feature map of the image to be recognized, captures the global features in the image, and obtains the global expression feature map of the image to be recognized, which more accurately represents the content of the image to be recognized and the expression of the object in the image to be recognized. The second fusion processing module 506 multiplies the weight feature map of the image to be recognized and the global expression feature map of the image to be recognized, can pay more attention to the effective region, fuses the information of the weight feature map of the image to be recognized and the global expression feature map of the image to be recognized together, emphasizes the effective region, obtains the enhanced feature map of the image to be recognized, and the recognition module 507 performs expression recognition based on the enhanced feature map of the image to be recognized to obtain the expression recognition result, improves the accuracy and robustness of expression recognition, and solves the problem that in the prior art, the relationship between different local regions is not considered in image recognition in complex scenarios, resulting in inaccurate expression recognition results.
[0069] In some embodiments, the adaptive processing module 503 is configured to perform convolution processing on the feature map of the image to be recognized through the first convolutional layer to obtain the first convolution result of the image to be recognized; perform convolution processing on the feature map of the image to be recognized through the first dilated convolutional layer to obtain the first dilated convolution result of the image to be recognized; perform convolution processing on the feature map of the image to be recognized through the second dilated convolutional layer to obtain the second dilated convolution result of the image to be recognized, and the dilation rates of the first dilated convolutional layer and the second dilated convolutional layer are different; perform fusion processing on the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized to obtain the weight feature map of the image to be recognized.
[0070] In some embodiments, the adaptive processing module 503 is configured to add the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized to obtain the second fused feature map of the image to be recognized; perform convolution processing on the second fused feature map of the image to be recognized through the second convolutional layer to obtain the second convolution result of the image to be recognized; perform normalization processing on the second convolution result of the image to be recognized to obtain the weight feature map of the image to be recognized.
[0071] In some embodiments, the first fusion processing module 504 is configured to split the feature map of the image to be recognized to obtain a plurality of corresponding sub-feature maps of the image to be recognized; flatten each sub-feature map of the image to be recognized to obtain the semantic feature vectors of each sub-feature map of the image to be recognized; perform position encoding on each sub-feature map of the image to be recognized to obtain the position feature vectors of each sub-feature map of the image to be recognized; perform fusion processing on the semantic feature vectors of each sub-feature map of the image to be recognized and the position feature vectors of each sub-feature map of the image to be recognized to obtain the fused feature vectors of each sub-feature map; splice the fused feature vectors of each sub-feature map to obtain the first fused feature map of the image to be recognized.
[0072] In some embodiments, the self-attention processing module 505 is configured to perform self-attention encoding processing on the first fused feature map of the image to be recognized, calculate the attention weights, perform weighted summation based on the attention weights to obtain the local expression feature maps of the image to be recognized; splice the local expression feature maps of the image to be recognized based on the positions of the respective sub-feature maps of the image to be recognized to obtain the global expression feature map of the image to be recognized.
[0073] In some embodiments, the recognition module 507 is configured to perform dimensionality reduction processing on the enhanced feature map of the image to be recognized through the global average pooling layer to obtain the global dimensionality reduction processing result of the image to be recognized; perform feature transformation on the global dimensionality reduction processing result of the image to be recognized through the fully connected layer to obtain the fully connected feature vector of the image to be recognized; perform classification processing on the fully connected feature vector of the image to be recognized through the classification layer to obtain the expression recognition result of the image to be recognized.
[0074] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0075] Figure 6 is a schematic diagram of the electronic device 6 provided by the embodiments of the present disclosure. As Figure 6As shown, the electronic device 6 of this embodiment includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, the steps in the above-described method embodiments are implemented. Alternatively, when the processor 601 executes the computer program 603, the functions of the various modules / units in the above-described device embodiments are implemented.
[0076] The electronic device 6 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 6 may include, but is not limited to, the processor 601 and the memory 602. Those skilled in the art can understand that Figure 6 merely examples of the electronic device 6, which do not constitute a limitation on the electronic device 6, may include more or fewer components than shown in the figure, or different components.
[0077] The processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0078] The memory 602 may be an internal storage unit of the electronic device 6, for example, the hard disk or memory of the electronic device 6. The memory 602 may also be an external storage device of the electronic device 6, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 6. The memory 602 may also include both an internal storage unit and an external storage device of the electronic device 6. The memory 602 is used to store computer programs and other programs and data required by the electronic device.
[0079] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0080] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium (such as a computer-readable storage medium). Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0081] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. A method for facial expression recognition, characterized in that, Including: Obtain an image to be recognized, where the image to be recognized includes a facial image of an object; Extract features from the image to be recognized to obtain a feature map of the image to be recognized; Perform adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized; Determine a first fused feature map of the image to be recognized according to the feature map of the image to be recognized; Perform self-attention encoding processing on the first fused feature map of the image to be recognized to obtain a global expression feature map of the image to be recognized; Multiply the weighted feature map of the image to be recognized and the global expression feature map of the image to be recognized to obtain an enhanced feature map of the image to be recognized; Determine an expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
2. The method according to claim 1, wherein The performing adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized includes: Perform convolution processing on the feature map of the image to be recognized through a first convolutional layer to obtain a first convolution result of the image to be recognized; Perform convolution processing on the feature map of the image to be recognized through a first dilated convolutional layer to obtain a first dilated convolution result of the image to be recognized; Perform convolution processing on the feature map of the image to be recognized through a second dilated convolutional layer to obtain a second dilated convolution result of the image to be recognized, where the dilation rates of the first dilated convolutional layer and the second dilated convolutional layer are different; Perform fusion processing on the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized to obtain a weighted feature map of the image to be recognized.
3. The method according to claim 2, characterized in that, The performing fusion processing on the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized to obtain a weighted feature map of the image to be recognized includes: Perform addition processing on the first convolution result of the image to be recognized, the first dilated convolution result of the image to be recognized, and the second dilated convolution result of the image to be recognized to obtain a second fused feature map of the image to be recognized; Perform convolution processing on the second fused feature map of the image to be recognized through a second convolutional layer to obtain a second convolution result of the image to be recognized; Perform normalization processing on the second convolution result of the image to be recognized to obtain a weighted feature map of the image to be recognized.
4. The method according to claim 1, characterized in that, The determining a first fused feature map of the image to be recognized according to the feature map of the image to be recognized includes: Perform splitting processing on the feature map of the image to be recognized to obtain corresponding multiple sub-feature maps of the image to be recognized; Perform flattening processing on each sub-feature map of the image to be recognized to obtain semantic feature vectors of each sub-feature map of the image to be recognized; Perform position encoding on each sub-feature map of the image to be recognized to obtain position feature vectors of each sub-feature map of the image to be recognized; Fuse the semantic feature vectors and the position feature vectors of each sub-feature map of the image to be recognized to obtain the fused feature vector of each sub-feature map; Concatenate the fused feature vectors of each sub-feature map to obtain the first fused feature map of the image to be recognized.
5. The method according to claim 4, wherein The self-attention encoding process for the first fused feature map of the image to be recognized to obtain the global expression feature map of the image to be recognized includes: Perform self-attention encoding on the first fused feature map of the image to be recognized, calculate the attention weights, and perform weighted summation based on the attention weights to obtain the local expression feature maps of the image to be recognized; Concatenate the local expression feature maps of the image to be recognized based on the positions of the sub-feature maps of the image to be recognized to obtain the global expression feature map of the image to be recognized.
6. The method according to claim 1, wherein The determining the expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized includes: Perform dimensionality reduction on the enhanced feature map of the image to be recognized through a global average pooling layer to obtain the global dimensionality reduction result of the image to be recognized; Perform feature transformation on the global dimensionality reduction result of the image to be recognized through a fully connected layer to obtain the fully connected feature vector of the image to be recognized; Perform classification on the fully connected feature vector of the image to be recognized through a classification layer to obtain the expression recognition result of the image to be recognized.
7. The method according to any one of claims 1 to 6, characterized in that Before performing feature extraction on the image to be recognized, it further includes: Obtain an expression recognition training set, which includes multiple training images and labels corresponding to each training image; Select multiple training images from the expression recognition training set according to a first preset percentage for flipping processing to obtain the flipped training images corresponding to each training image, and select multiple training images from the expression recognition training set according to a second preset percentage for occlusion processing to obtain the occluded training images corresponding to each training image; Replace the corresponding training images with the flipped training images and replace the corresponding training images with the occluded training images to obtain an enhanced expression recognition training set; Input each training image in the enhanced expression recognition training set into the expression recognition model to perform feature extraction on each training image in the enhanced expression recognition training set to obtain the feature maps of each training image; Perform adaptive region processing on the feature maps of each training image to obtain the weight feature maps of each training image; Determine the fused feature maps of each training image according to the feature maps of each training image; Perform self-attention encoding on the fused feature maps of each training image to obtain the global expression feature maps of each training image; Multiply the global expression feature maps of each training image by the weight feature maps of each training image to obtain the enhanced feature maps of each training image; Determine the expression recognition results of each image according to the enhanced feature maps of each training image; Determine a loss value according to the facial expression recognition results of each of the images and the labels of each of the training images, and update the parameters in the facial expression recognition model based on the loss value.
8. An expression recognition device, characterized in that, Including: An acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes a facial image of an object; A feature extraction module, configured to perform feature extraction on the image to be recognized to obtain a feature map of the image to be recognized; An adaptive processing module, configured to perform adaptive region processing on the feature map of the image to be recognized to obtain a weighted feature map of the image to be recognized; A first fusion processing module, configured to determine a first fusion feature map of the image to be recognized according to the feature map of the image to be recognized; A self-attention processing module, configured to perform self-attention encoding processing on the first fusion feature map of the image to be recognized to obtain a global facial expression feature map of the image to be recognized; A second fusion processing module, configured to perform a multiplication process on the weighted feature map of the image to be recognized and the global facial expression feature map of the image to be recognized to obtain an enhanced feature map of the image to be recognized; A recognition module, configured to determine a facial expression recognition result of the image to be recognized according to the enhanced feature map of the image to be recognized.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.