Facial expression recognition method and device
The method enhances facial expression recognition accuracy in complex scenarios by combining global and local feature information, addressing issues of occlusion and body posture, and improving user experience.
Patent Information
- Application Number
- JP2024545150
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-17
- Filing Date
- 2022-07-26
- Publication Date
- 2026-01-28
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Existing facial expression recognition technologies struggle in complex scenarios where important face parts are occluded or due to different body postures, leading to inaccurate or inability to recognize facial expressions.
A facial expression recognition method that combines global and local feature information by acquiring an image feature map, determining global and local features, and identifying the expression type based on these features, using neural networks and attention mechanisms to enhance accuracy.
Improves the accuracy of facial expression recognition by effectively combining global and local feature information, restoring more detailed image information and enhancing user experience.
Smart Images

Figure 0007808200000002 
Figure 0007808200000003 
Figure 0007808200000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of computer technology, and more particularly to a method and apparatus for facial expression recognition. [Background technology]
[0002] With the development of image processing technology, image recognition technology is being used in more and more scenarios. Although current image recognition technology has achieved very good results in facial expression recognition, in complex scenarios, for example, important parts of the face are occluded or the complete face cannot be collected due to different body postures (e.g., turning the body sideways), resulting in inability to recognize facial expressions or inaccurate facial expression recognition results. Therefore, a new solution for facial expression recognition is urgently needed. Summary of the Invention
[0003] In view of this, the embodiments of the present disclosure provide a facial expression recognition method, device, computer device and computer-readable storage medium to solve the problem that in complex scenarios in the prior art, the complete face cannot be collected due to, for example, important parts of the face being occluded or due to different body postures (e.g., turning the body to the side), and therefore facial expressions cannot be recognized or the facial expression recognition results are inaccurate.
[0004] In a first aspect of the disclosed embodiment, acquiring a recognition target image; obtaining an image feature map of the recognition target image based on the recognition target image; determining global and local feature information based on the image feature map; and identifying an expression type of the recognition target image based on the global feature information and the local feature information.
[0005] In a second aspect of the disclosed embodiment, an image acquisition module for acquiring a recognition target image; a first feature acquisition module for acquiring an image feature map of the recognition target image based on the recognition target image; a second feature acquisition module for identifying global feature information and local feature information based on the image feature map; and an expression type identification module for identifying the expression type of the recognition target image based on global feature information and local feature information.
[0006] In a third aspect of an embodiment of the present disclosure, there is provided a computing device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, the computing device implementing the steps of the method when the processor executes the computer program.
[0007] In a fourth aspect of an embodiment of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, the computer program implementing the steps of the above method when executed by a processor.
[0008] Beneficial advantages of the disclosed embodiments over the prior art: The disclosed embodiments can first obtain a recognition target image, then obtain an image feature map of the recognition target image based on the recognition target image, then characterize global feature information and local feature information based on the image feature map, and finally identify the facial expression type of the recognition target image based on the global feature information and local feature information. Because the global feature information of the recognition target image reflects the overall facial information and the local feature information of the recognition target image reflects detailed information of each region of the face, an effective combination of the local feature information and the global feature information can restore more detailed image information of facial expressions. Thus, the facial expression type of the recognition target image identified based on the global feature information and the local feature information can be more accurate, i.e., the accuracy of the recognition result of the facial expression type corresponding to the detection target image can be improved, and the user experience can be further improved. [Brief explanation of the drawings]
[0009] In order to more clearly explain the technical solutions in the embodiments of the present disclosure, the following briefly introduces drawings necessary for explaining the embodiments or prior art. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings based on these drawings without the need for creative work. [Figure 1] FIG. 1 is a scenario schematic diagram of an application scenario of an embodiment of the present disclosure. [Figure 2] 1 is a flowchart of a facial expression recognition method provided in an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram of the network architecture of a facial expression recognition model provided in an embodiment of the present disclosure. [Figure 4] 1 is a flowchart of a GoF modulation process provided in an embodiment of the present disclosure. [Figure 5] 1 is a schematic diagram of a network architecture of a Local-Global Attention structure provided in an embodiment of the present disclosure. FIG. [Figure 6]FIG. 1 is a schematic diagram of a network architecture of a global model provided in an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram of the network architecture of a Split-Attention module provided in an embodiment of the present disclosure. [Figure 8] FIG. 1 is a schematic diagram of a weight layer network architecture provided in an embodiment of the present disclosure. [Figure 9] FIG. 1 is a block diagram of a facial expression recognition device provided in an embodiment of the present disclosure. [Figure 10] FIG. 1 is a schematic diagram of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] In the following description, for purposes of explanation, not limitation, specific details, such as particular system structures and techniques, are provided to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary details.
[0011] Hereinafter, a facial expression recognition method and device according to an embodiment of the present disclosure will be described in detail with reference to the drawings.
[0012] In the prior art, in complex scenarios, for example, because important parts of the face are occluded or due to different body postures (e.g., turning the body sideways), the complete face cannot be collected, resulting in inaccurate facial expression recognition or inability to recognize facial expressions. Therefore, a new facial expression recognition solution is urgently needed.
[0013] To solve the above problems, the present invention provides a facial expression recognition method, in which the method provided in this embodiment first obtains a recognition target image, then obtains an image feature map of the recognition target image based on the recognition target image, then obtains global feature information and local feature information based on the image feature map, and finally identifies the facial expression type of the recognition target image based on the global feature information and local feature information. Because the global feature information of the recognition target image reflects the overall facial information, and the local feature information of the recognition target image reflects detailed information of each region of the face, an effective combination of the local feature information and the global feature information can restore more detailed image information of facial expressions, and thus the facial expression type of the recognition target image identified based on the global feature information and the local feature information can be more accurate, that is, the accuracy of the recognition result of the facial expression type corresponding to the detection target image can be improved, and the user experience can be further improved.
[0014] By way of example, an embodiment of the present invention may be applied to the application scenario shown in Figure 1. This scenario may comprise a terminal device 1 and a server 2.
[0015] The terminal device 1 may be hardware or software. If the terminal device 1 is hardware, it may be various electronic devices that have functions of collecting and storing images and support communication with the server 2, including, but not limited to, smartphones, tablets, laptop computers, digital cameras, monitors, video recorders, and desktop computers. If the terminal device 1 is software, it may be installed in, for example, the above-mentioned electronic devices. The terminal device 1 may be realized as multiple software programs or software modules, or as a single software program or software module; the embodiments of the present disclosure are not limited thereto. Furthermore, various apps, such as an image collection app, an image storage app, and a live chat app, may be installed in the terminal device 1.
[0016] The server 2 may be a server that provides various services, such as a back-end server that receives bills sent by terminal devices that establish a communication connection with the server 2, and the back-end server may receive, analyze, and otherwise process the bills sent by the terminal devices to generate processing results. The server 2 may be a single server, a server cluster consisting of several servers, or a cloud computing service center, and the embodiments of the present disclosure are not limited thereto.
[0017] The server 2 may be hardware or software. If the server 2 is hardware, it may be various electronic devices that provide various services to the terminal device 1. If the server 2 is software, it may be multiple software programs or software modules that provide various services to the terminal device 1, or it may be a single software program or software module that provides various services to the terminal device 1, and the embodiments disclosed herein are not limited to these.
[0018] The terminal device 1 and the server 2 may be communicatively connected via a network. The network may be a wired network connected using a coaxial cable, a twisted pair, or an optical fiber, or may be a wireless network that can realize interconnection of various communication devices without the need for wiring, such as Bluetooth, Near Field Communication (NFC), or Infrared, and the embodiments of the present disclosure are not limited thereto.
[0019] Specifically, a user identifies a detection target image using a terminal device 1 and transmits the detection target image to a server 2. After the server 2 receives the detection target image, the server 2 can extract an image feature map of the recognition target image. The server 2 can then identify global feature information and local feature information based on the image feature map. The server 2 can then identify an expression type of the recognition target image based on the global feature information and the local feature information. In this way, the global feature information of the recognition target image reflects overall facial information, and the local feature information of the recognition target image reflects detailed information about each region of the face. Therefore, by effectively combining the local feature information and the global feature information, more detailed image information about facial expressions can be restored. In this way, the expression type of the recognition target image identified based on the global feature information and the local feature information can be more accurate, which improves the accuracy of the expression type recognition result corresponding to the detection target image and further enhances the user experience.
[0020] It should be noted that the specific types, numbers and combinations of the terminal device 1, the server 2 and the network can be adjusted according to the actual needs of the application scenario, and the embodiments disclosed herein are not limited thereto.
[0021] It should be noted that the above application scenarios are merely exemplified for easy understanding of the present disclosure, and the embodiments of the present disclosure are not limited in this respect, but may be used in any application scenario.
[0022] 2 is a flowchart of a facial expression recognition method provided in an embodiment of the present disclosure. The facial expression recognition method of FIG. 2 may be performed by the terminal device and / or the server of FIG. 1. As shown in FIG. 2, the facial expression recognition method includes the following steps:
[0023] S201: Acquire a recognition target image.
[0024] In this embodiment, an image or video frame that requires facial expression recognition may be referred to as a recognition target image, and a face that requires facial expression recognition may be referred to as a target face.
[0025] For example, the terminal device can provide a web page, and the user can upload an image on the web page and click a preset button to trigger facial expression recognition on the image, thereby setting the image as a recognition target image. Of course, the user can also obtain the recognition target image by taking a photo using the terminal device. In addition, the user needs to specify the target face that needs to be recognized, or set the target face that needs to be recognized by the system as a default setting.
[0026] S202: An image feature map of the recognition target image is obtained based on the recognition target image.
[0027] After the target image is obtained, an image feature map of the target image may be extracted for the target image. As can be understood, the image feature map of the target image is a complete feature map extracted from the target image, which can reflect the entire image feature information of the target image.
[0028] S203: Identify global feature information and local feature information based on the image feature map.
[0029] In complex scenarios, for example, due to occlusion of important facial features or differences in body posture (e.g., turning sideways), the complete face cannot be collected, leading to inaccurate or inaccurate recognition of facial expression types. To solve this problem, in this embodiment, after obtaining an image feature map of the target image, global feature information and local feature information of the target face can be obtained from the image feature map based on the image feature map. The global feature information of the target image can reflect the overall information of the face, i.e., the global feature information of the target image can be understood as the global information of the target face, such as skin color, contour, and distribution of facial features. The local feature information of the target image can reflect detailed information of each region of the face, i.e., the local feature information can reflect detailed features of the target face, such as organ features such as scars, moles, dimples, and odd facial features. As can be seen, the global feature information is used for rough matching, and the local feature information is used for detailed matching. Therefore, by effectively combining the local feature information and the global feature information, more detailed image information of facial expressions can be recovered.
[0030] S204: Identify the facial expression type of the recognition target image based on the global feature information and the local feature information.
[0031] In this embodiment, after the global feature information and the local feature information are identified, more detailed image information of facial expressions can be restored using the global feature information and the local feature information, and the detailed image information of the restored facial expressions can then be used to identify the expression type of the target face. In this way, the expression type of the recognition target image identified based on the global feature information and the local feature information can be made more accurate, and the accuracy of the recognition result of the expression type corresponding to the detection target image can be improved.
[0032] Beneficial advantages of the disclosed embodiments over the prior art: The disclosed embodiments can first acquire a target image, then obtain an image feature map of the target image based on the target image, then characterize global feature information and local feature information based on the image feature map, and finally identify the facial expression type of the target image based on the global feature information and local feature information. Because the global feature information of the target image reflects the overall facial information and the local feature information of the target image reflects detailed information about each region of the face, an effective combination of the local feature information and global feature information can restore more detailed image information about facial expressions. Thus, the facial expression type of the target image identified based on the global feature information and local feature information is more accurate, which improves the accuracy of the recognition result of the facial expression type corresponding to the target image and further improves the user experience. In other words, in the facial expression type recognition process, while focusing on local features, context information of global features can also be introduced, significantly improving recognition performance.
[0033] Next, an implementation of S202 of "obtaining an image feature map of the recognition target image based on the recognition target image" will be described. In this embodiment, the corresponding method of FIG. 2 can be applied to a facial expression recognition model (shown in FIG. 3), and the facial expression recognition model may include a first neural network model, and S202 of "obtaining an image feature map of the recognition target image based on the recognition target image" may include the following steps:
[0034] A recognition target image is input to a first neural network model to obtain an image feature map of the recognition target image.
[0035] Here, the first neural network model includes a number of first convolution modules, which are connected in series, and each first convolution module includes a filter, a batch normalization layer, an MP model, and an activation function. For example, as shown in Figure 3, the first neural network model includes three first convolution modules, where each first convolution module includes a filter (e.g., a GC module), a batch normalization layer (i.e., a BN layer), an MP model (i.e., MP, Max Pooling), and an activation function (i.e., ReLU).
[0036] Here, the GC module is a convolution kernel with an added Gabor filter. It should be emphasized that in this embodiment, the GC module can be referred to as a Gabor directional filter (i.e., a Gabor directional filter modified with a normal convolution kernel), and the GoF can enhance the robustness (orientation invariance and scale invariance) of the network against changes in orientation and scale. Note that Gabor filters have excellent properties, not only being able to preserve frequency domain information in an image but also having strong robustness against geometric transformations such as orientation and size, which traditional convolutional neural networks lack. Therefore, Gabor filters are suitable for facial expression recognition in scenarios where orientation and size change frequently. However, facial expression recognition focuses more on local features such as the nose and mouth compared to other facial tasks, and Gabor filters can extract richer local features than traditional convolutional neural networks. As can be seen, the GC module, after the Gabor filter is added, can enhance the robustness of the network against changes in orientation and scale, while reducing the number of network layers and thereby reducing the forward inference time of the model.
[0037] In this embodiment, the modulation process of GoF is as follows:
[0038] Gabor filters have U-direction and V-scale. To integrate the steerable properties into GCN, the direction information is encoded in the training filter, while the scale information is embedded in different convolution layers. In GoF, Gabor filters capture the direction and scale information, so the corresponding convolution functions are enhanced.
[0039] The filter trained by the backpropagation algorithm is called the training filter. Unlike the standard convolutional neural network, the filter trained by the GCN is three-dimensional in order to encode the directional channel. Assuming the size of the training filter is N×W×W, W×W is the size of the filter and N represents the channel. In a traditional CNN, the weights of each layer are C out ×C in When expressed as W × W × W, the GCN is out ×C in × N × W × W, and C out and C in represent the output and input channels of the feature map, respectively. In the forward convolution process, the number of channels of the feature map is matched, and N is selected as U, that is, used to modulate the number of directions of the Gabor filter in this learning filter. On a given scale V, the GoF is obtained through the modulation process. JPEG0007808200000001.jpg62170
[0040] In GoF, the value of ν increases with the number of layers, which means that the ratio of Gabor filters in GoF changes with the number of layers. At each scale, the size of GoF is U×N×W×W. However, for a given Gabor filter, we only need to store N×W×W training filters, which means that we can obtain enhanced functionality without increasing the number of parameters.
[0041] Next, an implementation of S203 of "determining global feature information and local feature information based on the image feature map" will be described. In this embodiment, the corresponding method of Fig. 2 can be applied to a facial expression recognition model, where the facial expression recognition model has a local-global attention structure as shown in Fig. 3, and the local-global attention structure has a global model and a local model as shown in Fig. 5. S203 of "determining global feature information and local feature information based on the image feature map" may include the following steps:
[0042] S203a: Input the image feature map into the global model to obtain global feature information.
[0043] The conv and feature layers in Figure 5 are specifically shown in Figure 6, that is, the global model includes a first convolutional layer, an H-Sigmoid activation function layer (i.e., F in Figure 6), a channel attention module (CAM, i.e., channel attention in Figure 6), a spatial attention module (SAM, i.e., spatial attention in Figure 6), and a second convolutional layer. As can be understood, the channel attention module and the spatial attention module may be understood as a single CBAM attention structure.
[0044] Specifically, an image feature map can be input to a first convolutional layer to obtain a first feature map. The first feature map can then be input to an H-Sigmoid activation function layer to obtain a second feature map. The H-Sigmoid activation function layer can avoid exponential operations and thus improve computation speed. Next, the second feature map can be input to a channel attention module to obtain a channel attention map. The role of the channel attention module is to learn a weight for each channel and then multiply the weight by the corresponding channel, i.e., multiply the features of different channels in the second feature map by their corresponding weights to obtain the channel attention map. Then, the channel attention map can be input to a spatial attention module to obtain a spatial attention map. The spatial attention module uses the channel attention map to perform attention calculation in the feature map space (i.e., the width and height dimensions). That is, each local region in the channel attention map learns a weight, and then multiplies the weight by the channel attention map to obtain the spatial attention map. Finally, the spatial attention map can be input to the second convolutional layer to obtain global feature information.
[0045] S203b: Input the image feature map into the local model to obtain local feature information.
[0046] In this embodiment, as shown in FIG. 5, the local model may include N local feature extraction convolution layers (i.e., a conv layer and a feature layer, where k and N are the same number), and an attention module, where N is a positive integer greater than 1.
[0047] In this embodiment, N local image blocks can be generated based on the image feature map. For example, the image feature map can be randomly cropped to obtain N different local image blocks, each containing local information of a different region. Then, each local image block can be input to each local feature extraction convolution layer to obtain N local feature maps, that is, each local feature extraction convolution layer outputs one local feature map.
[0048] Next, the N local feature maps can be input to an attention module to obtain local feature information. Note that in this embodiment, the attention module may be a Split-Attention module, as shown in FIG. 7, which may include a pooling layer (i.e., Global Pooling in FIG. 7), a second convolution module (i.e., Conv+BN+ReLU in FIG. 7, i.e., the second convolution module includes a convolution layer, a BN layer, and a ReLU function layer), N third convolution layers (i.e., Conv layer in FIG. 7), and a normalization layer (i.e., r-Softmax layer in FIG. 7). Specifically, as shown in FIG. 7, the N local feature maps are fused to obtain a fused feature map. The fused feature map is input to a pooling layer to obtain a pooled feature map. The pooled feature map is input to a second convolution module to obtain a processed local feature map. The processed local feature map is input to each third convolutional layer to obtain N sub-local feature maps, i.e., the processed local feature map is input to each independent third convolutional layer process to extract each sub-local feature map. For each sub-local feature map, the sub-local feature map is input to a normalization layer to obtain a weight value corresponding to the sub-local feature map, and the local feature map corresponding to the sub-local feature map is obtained based on the weight value corresponding to the sub-local feature map and the local feature map. In other words, an r-Softmax operation is performed on the sub-local feature map using the normalization layer to obtain a weight value corresponding to the sub-local feature map, and then the local feature map corresponding to the sub-local feature map is multiplied by the weight value corresponding to the sub-local feature map to obtain a processed local feature map corresponding to the sub-local feature map. The N local feature maps (i.e., the processed local feature maps) are fused to obtain local feature information.
[0049] Next, an implementation of S204 of “inputting global feature information into the first global average pooling layer to obtain global pooling feature information” will be described. In this embodiment, S204 may include the following steps:
[0050] S204a: Input the global feature information into a first global average pooling layer to obtain global pooling feature information.
[0051] In this embodiment, as shown in FIG. 5 , a first global average pooling layer (i.e., the GAP layer in FIG. 5 ) is further connected to the global model. After extracting the global feature information, the first global average pooling layer can perform global average pooling (GAP) processing on the global feature information to obtain the global pooling feature information.
[0052] S204b: Input the local feature information into a second global average pooling layer to obtain local pooling feature information.
[0053] In this embodiment, as shown in FIG. 5 , a second global average pooling layer (i.e., the GAP layer in FIG. 5 ) is further connected to the local model. After extracting the local feature information, the second global average pooling layer performs global average pooling (GAP) processing on the local feature information to obtain local pooled feature information.
[0054] S204c: The global pooling feature information and the local pooling feature information are input to the fully connected layer to obtain the facial expression type of the recognition target image.
[0055] In this embodiment, after obtaining the global pooling feature information and the local pooling feature information, the global pooling feature information and the local pooling feature information are first merged to obtain merged pooling feature information, which is then input to the fully connected layer (i.e., FC in FIG. 3 ) to obtain the facial expression type of the recognition target image output by the fully connected layer.
[0056] In addition, in one implementation of this embodiment, the facial expression recognition model may include a weight layer (i.e., Weight branch in FIG. 3) as shown in FIG. 3, and the weight layer includes two fully connected layers (i.e., FC1 and FC2 in FIG. 8) and one Sigmoid activation layer (i.e., Sigmoid in FIG. 8) as shown in FIG. 8. Specifically, the corresponding method of FIG. 2 further includes the following steps:
[0057] Step 1: Input the global pooling feature information and the local pooling feature information into the weight layer to obtain the true weight value of the facial expression type of the recognition target image.
[0058] Here, the true weight value of the facial expression type of the recognition target image is intended to represent the probability that the facial expression type is the corresponding true facial expression type of the recognition target image, and in one implementation, the true weight value may be a weight value between 0 and 1. As can be seen, the larger the true weight value of the facial expression type of the recognition target image, the greater the probability that the facial expression type is the corresponding true facial expression type of the recognition target image, and conversely, the smaller the true weight value of the facial expression type of the recognition target image, the smaller the probability that the facial expression type is the corresponding true facial expression type of the recognition target image.
[0059] Step 2: Identify the expression type detection result corresponding to the recognition target image based on the truth weight value of the expression type of the recognition target image and the expression type of the recognition target image.
[0060] In this embodiment, if the true weight value of the expression type of the recognition target image satisfies a predetermined condition, for example, if the true weight value of the expression type of the recognition target image is greater than a predetermined threshold, the expression type of the recognition target image can be determined as the corresponding expression type detection result of the recognition target image. The predetermined threshold may be set in advance or, for example, set by a user according to actual needs. Conversely, if the true weight value of the expression type of the recognition target image does not satisfy the predetermined condition, for example, if the true weight value of the expression type of the recognition target image is less than the predetermined threshold, the expression type of the recognition target image does not need to be determined as the corresponding expression type detection result of the recognition target image, and the corresponding expression type detection result of the recognition target image can be determined as no result.
[0061] Furthermore, prior art facial expression recognition datasets suffer from labeling uncertainty. This uncertainty is caused by the subjectivity of the annotators, resulting in differences between the same photo after it is labeled by different annotators. On the other hand, it is also caused by the ambiguity of facial expression photos, i.e., a single facial expression photo may contain a variety of different expressions. Factors such as occlusion, lighting, and image quality issues may also affect the labeling results. Training data with such uncertain labeling is detrimental to learning effective facial expression features and can result in non-convergence in the early stages of network training. Therefore, in this embodiment, a weight layer is trained using an attention-weighted loss loss function. Specifically, after feature extraction, a weight branch (i.e., a weight layer) is added to each sample to learn a weight. This weight represents the degree of certainty of the current sample. When calculating the loss, the loss of the training sample is weighted according to the calculated weight.
[0062] In the model training process, if the calculated weight value of the weight layer is smaller than the predetermined training value, it indicates that the certainty of the current training sample is very low, which is unfavorable to the optimization of the network. Therefore, these samples are filtered by an artificially set threshold, and the calculated weight value of the weight layer is equal to or greater than the predetermined training value. The higher the value, the stricter the requirement for certainty, and some samples between certain and uncertain are removed. The lower the weight value, the more samples will contribute to the loss. Samples with low certainty (i.e., the calculated weight value of the weight layer is smaller than the predetermined training value) will not be used in the loss calculation, which can significantly improve the accuracy of facial expression recognition. After filtering the training samples, the features of all samples and their corresponding weights can be obtained. After calculating the cross-entropy for each sample, a weighted sum is performed using these weights. Alternatively, a different weighting method can be used, i.e., the sample features taken are weighted by the calculated weight sum, and then the loss is calculated for the weighted features.
[0063] In addition, after adding a weight layer to the facial expression recognition model (i.e., adding one weight branch), the branch learns the certainty of each sample, and then uses Attention Weighted Loss to introduce the learned uncertainty into the Attention Weighted Loss, thereby suppressing the impact of uncertainty labeling on the facial expression recognition training process.
[0064] All the above-mentioned optional technical solutions may be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described one by one here.
[0065] The following are apparatus embodiments of the present disclosure for carrying out the method embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.
[0066] 9 is a schematic diagram of a facial expression recognition device provided in an embodiment of the present disclosure. As shown in FIG. 9, the facial expression recognition device includes: an image acquisition module 901 for acquiring a recognition target image; a first feature acquisition module 902 for acquiring an image feature map of the recognition target image based on the recognition target image; a second feature acquisition module 903 for identifying global and local feature information based on the image feature map; and an expression type identification module 904 for identifying the expression type of the recognition target image based on the global feature information and the local feature information.
[0067] In some embodiments, the method is applied to a facial expression recognition model, and the facial expression recognition model includes a first neural network model, and the first feature acquisition module 902 specifically includes: inputting a recognition target image into a first neural network model to obtain an image feature map of the recognition target image; Here, the first neural network model includes a number of first convolution modules, which are connected in series, and each of the first convolution modules includes a filter, a batch normalization layer, an MP model, and an activation function.
[0068] In some embodiments, the method is applied to a facial expression recognition model, and the facial expression recognition model includes a global model and a local model. The second feature acquisition module 903 specifically includes: Input the image feature map into the global model to obtain global feature information, The image feature map is input to the local model to obtain local feature information.
[0069] In some embodiments, the global model includes a first convolution layer, an H-Sigmoid activation function layer, a channel attention module, a spatial attention module, and a second convolution layer, and the second feature acquisition module 903 specifically includes: Input the image feature map into the first convolutional layer to obtain the first feature map, The first feature map is input to the H-Sigmoid activation function layer to obtain the second feature map. Input the second feature map into the channel attention module to obtain a channel attention map; Input the channel attention map into the spatial attention module to obtain the spatial attention map; The spatial attention map is input to the second convolutional layer to obtain global feature information.
[0070] In some embodiments, the local model includes N local feature extraction convolution layers, an attention module, where N is a positive integer greater than 1, and the second feature acquisition module 903 specifically includes: Generate N local image blocks based on the image feature map; Each local image block is input to each local feature extraction convolution layer to obtain N local feature maps. N local feature maps are input to the attention module to obtain local feature information.
[0071] In some embodiments, the attention module includes a pooling layer, a second convolution module, N third convolution layers, and a normalization layer, and the second feature acquisition module 903 specifically includes: Fuse N local feature maps to obtain a fused feature map; The fused feature map is input to the pooling layer to obtain the pooled feature map. The pooled feature map is input to the second convolution module to obtain the processed local feature map; The processed local feature map is input to each third convolutional layer to obtain N sub-local feature maps. For each sub-local feature map, input the sub-local feature map into a normalization layer to obtain a weight value corresponding to the sub-local feature map; and obtain a local feature map corresponding to the sub-local feature map based on the weight value corresponding to the sub-local feature map and the local feature map; It is used to fuse N local feature maps to obtain local feature information.
[0072] In some embodiments, the facial expression type identification module 904 includes: The global feature information is input to the first global average pooling layer to obtain global pooling feature information; The local feature information is input to the second global average pooling layer to obtain local pooling feature information; The global pooling feature information and the local pooling feature information are input to the fully connected layer to obtain the facial expression type of the recognition target image.
[0073] In some embodiments, the facial expression recognition model further comprises a weight layer, the weight layer comprising two fully connected layers and one sigmoid activation layer, and the apparatus further comprises a result identification module; inputting the global pooling feature information and the local pooling feature information into a weight layer to obtain a true weight value of an expression type of a recognition target image, wherein the true weight value of the expression type of the recognition target image represents a probability that the expression type is a corresponding true expression type of the recognition target image; The facial expression type detection result corresponding to the recognition target image is specified based on the true weight value of the facial expression type of the recognition target image and the facial expression type of the recognition target image.
[0074] Beneficial effects of the embodiments of the present disclosure over the prior art: The facial expression recognition device provided in the embodiments of the present disclosure includes an image acquisition module 901 for acquiring a recognition target image, a first feature acquisition module 902 for acquiring an image feature map of the recognition target image based on the recognition target image, a second feature acquisition module 903 for identifying global feature information and local feature information based on the image feature map, and an expression type identification module 904 for identifying an expression type of the recognition target image based on the global feature information and the local feature information. Because the global feature information of the recognition target image reflects the overall facial information and the local feature information of the recognition target image reflects detailed information of each region of the face, an effective combination of the local feature information and the global feature information can restore more image details of facial expressions, and thus the expression type of the recognition target image identified based on the global feature information and the local feature information can be more accurate, i.e., the accuracy of the recognition result of the expression type corresponding to the detection target image can be improved, and the user experience can be further improved.
[0075] It should be understood that the magnitude of the numbers of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and does not arbitrarily limit the implementation process of the embodiments disclosed herein.
[0076] 10 is a schematic diagram of a computer device 10 provided in an embodiment of the present disclosure. As shown in FIG. 10, the computer device 10 of the embodiment includes a processor 1001, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable by the processor 1001. When the processor 1001 executes the computer program 1003, it realizes the steps in each of the above method embodiments. Alternatively, when the processor 1001 executes the computer program 1003, it realizes each module / function of each module in each of the above device embodiments.
[0077] For example, the computer program 1003 may be divided into one or more modules, and one or more modules may be stored in the memory 1002 and executed by the processor 1001 to accomplish the present disclosure. One or more modules may be a series of computer program command sections that can perform specific functions, and the command sections are intended to explain the process of the computer program 1003 being executed on the computer device 10.
[0078] The computer device 10 may be a computing device such as a desktop computer, a laptop computer, a palmtop computer, a cloud server, etc. The computer device 10 may include, but is not limited to, a processor 1001 and a memory 1002. As will be appreciated by those skilled in the art, FIG. 10 is merely an example of the computer device 10 and is not intended to limit the computer device 10, which may include more or fewer components than those shown, or may combine certain components or different components; for example, the computer device may include input / output devices, network access devices, buses, etc.
[0079] The processor 1001 may be a central processing unit (CPU) or other general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. The general-purpose processor may be a microprocessor, any common processor, etc.
[0080] The memory 1002 may be an internal storage module of the computer device 10, such as a hard disk or memory of the computer device 10. The memory 1002 may be an external storage device of the computer device 10, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash card, etc., installed in the computer device 10. Furthermore, the memory 1002 may comprise both an internal storage module of the computer device 10 and an external storage device. The memory 1002 is intended to store computer programs and other programs and data required by the computer device. The memory 1002 may also be used to temporarily store data that has been output or that is to be output.
[0081] Those skilled in the art will clearly understand that, for convenience and brevity, only the above-mentioned functional modules and module divisions have been described as examples. However, in actual applications, the above functions can be assigned to different functional modules as needed, i.e., all or part of the above-described functions can be achieved by dividing the internal structure of the device into different functional modules or modules. The functional modules in the embodiments may be integrated into a single processing module, each module may exist physically alone, or two or more modules may be integrated into a single module. The integrated module may be implemented in the form of hardware or software functional modules. The specific names of the functional modules are provided solely for the purpose of distinguishing them from one another and do not limit the scope of protection of the present disclosure. For the specific operating processes of the modules in the above-mentioned system, reference may be made to the corresponding processes in the above-mentioned method embodiments, and further description is omitted here.
[0082] In the above embodiments, the description of each embodiment has its own emphasis, and for the details or parts not described in an embodiment, reference can be made to the relevant descriptions of other embodiments.
[0083] Those skilled in the art can recognize that the modules and algorithm steps of each example described in the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software is determined by the specific application and design constraints of the technical solution. Those skilled in the art may implement the described functions using different methods for each specific application, but such implementation should not be considered to go beyond the scope of this disclosure.
[0084] It should be understood that the disclosed device / computer apparatus and method in the embodiments provided in this disclosure can be realized in other ways. For example, the device / computer apparatus embodiments described above are merely illustrative, and the division of the modules or modules is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules or components may be combined or integrated into other systems, or some features may be omitted or not implemented. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through several interfaces, devices, or modules, and may be electrical, mechanical, or other types.
[0085] Furthermore, modules described as separate components may or may not be physically separated, and components described as modules may or may not be physical modules, i.e., located in one location or distributed across multiple network modules, some or all of which may be selected according to actual needs to achieve the objectives of the solutions of this embodiment.
[0086] Furthermore, each functional module in each embodiment of the present disclosure may be integrated into one processing module, each module may exist physically independently, or two or more modules may be integrated into one module. The integrated module may be realized in the form of hardware or in the form of a software functional module.
[0087] The integrated module / modules may be realized in the form of a software functional module and stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the present disclosure provides that the realization of all or part of the processes in the above-described method embodiments can be accomplished by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-described method embodiments can be realized. The computer program may include computer program code, which may be in source code format, object code format, an executable file, or some intermediate format. The computer-readable storage medium may include any entity or device capable of carrying computer program code, such as a recording medium, a U-disk, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier wave signal, an electrical communication signal, and a software distribution medium. Furthermore, the content contained on a computer-readable storage medium may be increased or decreased as required by the legislation and patent practice of a jurisdiction. For example, in some jurisdictions, the legislation and patent practice may require that a computer-readable storage medium not include electrical carrier signals and telecommunications signals.
[0088] The above-described examples are merely for the purpose of illustrating the technical solutions of the present disclosure, and are not intended to limit the same. Although the present disclosure has been described in detail with reference to the above-described examples, those skilled in the art may still amend the technical solutions described in the above-described examples or equivalently replace some of the technical features therein, and it should be understood that such amendments or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and all of them should be included in the protection scope of the present disclosure.
Claims
1. A method for facial expression recognition, the method comprising: acquiring a recognition target image; acquiring an image feature map of the recognition target image based on the recognition target image; determining global and local feature information based on the image feature map; and identifying a facial expression type of the recognition target image based on the global feature information and the local feature information, The method is applied to a facial expression recognition model, the facial expression recognition model comprising a global model and a local model; The step of identifying the global feature information and the local feature information based on the image feature map includes: inputting the image feature map into the global model to obtain the global feature information; inputting the image feature map into the local model to obtain the local feature information; The local model includes N local feature extraction convolution layers and an attention module, where N is a positive integer greater than 1; The step of inputting the image feature map into a local model to obtain local feature information includes: generating N local image blocks based on the image feature map; Inputting each of the local image blocks into each of the local feature extraction convolution layers to obtain N local feature maps; inputting the N local feature maps into the attention module to obtain the local feature information; The attention module comprises a pooling layer, a second convolution module, N third convolution layers, and a normalization layer; The step of inputting the N local feature maps to the attention module and obtaining the local feature information includes: fusing the N local feature maps to obtain a fused feature map; inputting the fused feature map into the pooling layer to obtain a pooled feature map; inputting the pooled feature map into the second convolution module to obtain a processed local feature map; inputting the processed local feature map into each of the third convolutional layers to obtain N sub-local feature maps; For each of the sub-local feature maps, input the sub-local feature map to a normalization layer to obtain a weight value corresponding to the sub-local feature map, and obtain the local feature map corresponding to the sub-local feature map based on the weight value corresponding to the sub-local feature map and the local feature map; fusing the N local feature maps to obtain the local feature information. A method characterized by:
2. The step of identifying a facial expression type of the recognition target image based on the global feature information and the local feature information includes: inputting the global feature information into a first global average pooling layer to obtain global pooling feature information; inputting the local feature information into a second global average pooling layer to obtain local pooling feature information; and inputting the global pooling feature information and the local pooling feature information into a fully connected layer to obtain an expression type of the recognition target image.
2. The method of claim 1 .
3. A facial expression recognition method, the method comprising: acquiring a recognition target image; acquiring an image feature map of the recognition target image based on the recognition target image; determining global and local feature information based on the image feature map; and identifying a facial expression type of the recognition target image based on the global feature information and the local feature information, The method is applied to a facial expression recognition model, the facial expression recognition model comprising a global model and a local model; The step of identifying the global feature information and the local feature information based on the image feature map includes: inputting the image feature map into the global model to obtain the global feature information; inputting the image feature map into the local model to obtain the local feature information; The local model includes N local feature extraction convolution layers and an attention module, where N is a positive integer greater than 1; The step of inputting the image feature map into a local model to obtain local feature information includes: generating N local image blocks based on the image feature map; Inputting each of the local image blocks into each of the local feature extraction convolution layers to obtain N local feature maps; inputting the N local feature maps into the attention module to obtain the local feature information; The step of identifying a facial expression type of the recognition target image based on the global feature information and the local feature information includes: inputting the global feature information into a first global average pooling layer to obtain global pooling feature information; inputting the local feature information into a second global average pooling layer to obtain local pooling feature information; inputting the global pooling feature information and the local pooling feature information into a fully connected layer to acquire an expression type of the recognition target image; The facial expression recognition model further comprises a weight layer, the weight layer including two fully connected layers and one sigmoid activation layer; The method further comprises: a step of inputting the global pooling feature information and the local pooling feature information into the weight layer to obtain a true weight value of the expression type of the recognition target image, the true weight value of the expression type of the recognition target image being for representing a probability that the expression type is a corresponding true expression type of the recognition target image; and identifying a facial expression type detection result corresponding to the recognition target image based on the truth weight value of the facial expression type of the recognition target image and the facial expression type of the recognition target image. A method characterized by:
4. 1. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, The processor, when executing the computer program, performs the steps of the method of claim 1.
1. A computer device characterized by:
5. A computer-readable storage medium on which a computer program is stored, The computer program, when executed by a processor, implements the steps of the method of claim 1. A computer-readable storage medium comprising: