Image processing device, image processing method, and device-readable storage medium

The image processing device enhances facial micro-expression recognition by dividing facial images into regions, extracting and classifying action unit features, addressing the challenge of automatic micro-expression detection.

JP7746725B2Active Publication Date: 2025-10-01FUJITSU LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021130023
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-10
Filing Date
2021-08-06
Publication Date
2025-10-01
Estimated Expiration
2041-08-06

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for automatically recognizing facial micro-expressions, which are crucial for understanding psychological and emotional states in various applications.

Method used

An image processing device and method that utilizes an information acquisition unit to divide facial images into regions, extracts action unit features, calculates similarities between these features, and classifies them using global facial features to recognize micro-expressions.

Benefits of technology

The device effectively detects and classifies facial micro-expressions by learning AU relationships, improving detection accuracy through simultaneous and non-simultaneous AU feature recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746725000007
    Figure 0007746725000007
  • Figure 0007746725000008
    Figure 0007746725000008
  • Figure 0007746725000009
    Figure 0007746725000009
Patent Text Reader

Abstract

To provide an image processing apparatus, an image processing method, and a machine readable storage medium.SOLUTION: An image processing apparatus includes: an information acquisition section that divides an input image into a plurality of areas and acquires information on face action units in the plurality of areas; an action unit characteristic extraction section that extracts action unit characteristics from the areas of the face action units based on the information on the face action units; a first calculation section that calculates the similarity between the action unit characteristics of the respective face action units and the action unit characteristics of each face action unit; a second calculation section that re-calculates the action unit characteristics of the respective face action units based on a result of calculation of the similarity; a global face characteristic extraction section that extracts global face characteristics based on the information on the face action units; and a classification section that classifies the face action units based on the re-calculated action unit characteristics and the global face characteristics.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the technical field of image processing, and more particularly to an image processing device, an image processing method, and a machine-readable storage medium for facial micro-expression recognition. [Background technology]

[0002] This section provides background information related to the present disclosure, but is not necessarily prior art.

[0003] Faces always carry a lot of information about people's psychological and emotional states. Therefore, facial micro-expression recognition can quantify mental states. Facial micro-expression recognition has many applications in real life. For example, it can improve work efficiency by paying attention to employees' psychology and improving their motivation, or estimate customer satisfaction to improve purchasing motivation (digital marketing), or estimate a driver's driving behavior.

[0004] Therefore, the development of a device that can effectively recognize facial micro-expressions is of great importance. Summary of the Invention [Problem to be solved by the invention]

[0005] This section provides a general overview of the disclosure and does not fully disclose its entire scope or all of its features.

[0006] The present disclosure aims to provide an image processing device, an image processing method, and a machine-readable storage medium for automatically recognizing facial micro-expressions based on learning of action unit (AU) relationships. [Means for solving the problem]

[0007] In one aspect of the present disclosure, an image processing device is provided that includes an information acquisition unit that divides an input image into multiple regions and acquires information about facial action units in the multiple regions; an action unit feature extraction unit that extracts action unit features from the facial action unit regions based on the acquired information about the facial action units; a first calculation unit that calculates a similarity between the action unit features of each facial action unit among different facial action units and the action unit features of each facial action unit; a second calculation unit that recalculates the action unit features of each facial action unit based on the similarity calculation result; a global facial feature extraction unit that extracts global facial features based on the acquired information about the facial action units; and a classification unit that classifies the facial action units based on the recalculated action unit features of each facial action unit and the global facial features.

[0008] Another aspect of the present disclosure provides an image processing method including the steps of dividing an input image into a plurality of regions and acquiring information about facial action units in the plurality of regions; extracting action unit features from the regions of the facial action units based on the acquired information about the facial action units; calculating a similarity between the action unit features of each facial action unit among different facial action units and the action unit features of each facial action unit; recalculating the action unit features of each facial action unit based on the similarity calculation results; extracting global facial features based on the acquired information about the facial action units; and classifying the facial action units based on the recalculated action unit features of each facial action unit and the global facial features.

[0009] Another aspect of the present disclosure provides a machine-readable storage medium having recorded thereon a program product storing machine-readable instruction code, the program product being capable of causing the computer to execute an image processing method according to the present disclosure when the instruction code is read and executed by the computer.

[0010] By using the image processing device, image processing method, and machine-readable storage medium according to the present disclosure, it is possible to detect the appearance of facial action units corresponding to each local region of the face and recognize micro-expressions.

[0011] The scope of applicability of the present disclosure will become clearer from the description provided herein. The description and specific examples in this section are intended to be illustrative only and are not intended to limit the scope of the present disclosure. [Brief explanation of the drawings]

[0012] The drawings described herein are intended to illustrate preferred embodiments, not all possible embodiments, and are not intended to limit the scope of the present disclosure. [Figure 1] 1 is a block diagram illustrating a configuration of an image processing device according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating a configuration of an information acquisition unit in an image processing device according to an embodiment of the present disclosure. [Figure 3] FIG. 2 is a block diagram illustrating a configuration of an action unit feature extraction unit in an image processing device according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a block diagram illustrating a configuration of a global facial feature extraction unit in the image processing device according to an embodiment of the present disclosure. [Figure 5] 2 is a schematic diagram illustrating exemplary processes performed by a first calculation unit and a second calculation unit in an image processing device according to an embodiment of the present disclosure. FIG. [Figure 6] FIG. 10 is a schematic diagram illustrating another exemplary process performed by a first calculation unit and a second calculation unit in an image processing device according to an embodiment of the present disclosure. [Figure 7]1 is a flowchart illustrating an image processing method according to an embodiment of the present disclosure. [Figure 8] 1 is a block diagram showing an exemplary configuration of a general-purpose personal computer capable of implementing an image processing device and an image processing method according to an embodiment of the present disclosure. While various modifications and alternatives may be made to the present disclosure, specific embodiments thereof will be described in detail with reference to the drawings. The description of the specific embodiments does not limit the present disclosure to the specific aspects of the disclosure, and various modifications, equivalent variations, and alternatives may be made within the spirit and scope of the present disclosure. In the drawings, identical components are designated by the same reference numerals. DETAILED DESCRIPTION OF THE INVENTION

[0013]

[0023] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The following description is merely illustrative and does not limit the present disclosure, applications, and uses.

[0014] The following provides exemplary embodiments to describe the present disclosure in detail and to enable those skilled in the art to fully understand the scope of the present disclosure. In order to provide a detailed understanding of the embodiments of the present disclosure, many specific details, such as examples of specific means, devices, and methods, are described. However, as will be apparent to those skilled in the art, the specific details need not be used, and the exemplary embodiments may be implemented using different methods, and these examples do not limit the scope of the present disclosure. In some exemplary embodiments, well-known processes, well-known configurations, and well-known technologies are not described in detail.

[0015] Hereinafter, a method for recognizing facial micro-expressions using an image processing device according to an embodiment of the present disclosure will be described with reference to FIG.

[0016] FIG. 1 is a block diagram showing a configuration of an image processing device 100 according to an embodiment of the present disclosure. As shown in FIG. 1, the image processing device 100 according to an embodiment of the present disclosure may include an information acquisition unit 110, an action unit feature extraction unit 120, a global facial feature extraction unit 130, a classification unit 140, a first calculation unit 150, and a second calculation unit 160. Assume that there is only one face in the input image. The task of facial micro-expression recognition is to detect the appearance of all facial action units that interest the user. The presence or absence of the appearance of a facial action unit can indicate whether the facial muscles in the facial region corresponding to the facial action unit are moving.

[0017] First, the information acquisition unit 110 may divide the input image into multiple regions and acquire information about the facial action units in the multiple regions. The information acquisition unit 110 may also provide the acquired information about the facial action units in the multiple regions to the action unit feature extraction unit 120 and the global facial feature extraction unit 130.

[0018] Then, the action unit feature extraction unit 120 may extract action unit features from the region of the facial action unit based on the information about the facial action unit provided by the information acquisition unit 110. Furthermore, the action unit feature extraction unit 120 may provide the extracted action unit features to the first calculation unit 150 and the second calculation unit 160.

[0019] Then, the first calculation unit 150 may calculate the similarity between the action unit features of each facial action unit among the different facial action units and the action unit features of each facial action unit. Furthermore, the first calculation unit 150 may provide the calculation result of the similarity to the second calculation unit 160.

[0020] Then, the second calculation unit 160 may recalculate the action unit features of each facial action unit based on the action unit features extracted by the action unit feature extraction unit 120 and the similarity calculation result provided by the first calculation unit 150. Furthermore, the second calculation unit 160 may provide the recalculated action unit features to the classification unit 140.

[0021] Then, the global facial feature extraction unit 130 may extract global facial features based on the information about the facial action units provided by the information acquisition unit 110. Furthermore, the global facial feature extraction unit 130 may provide the extracted global facial features to the classification unit 140.

[0022] Then, the classification unit 140 may classify the facial action units based on the recalculated action unit features of each facial action unit provided by the second calculation unit 160 and the global facial features provided by the global facial feature extraction unit 130.

[0023] As a result, the image processing device 100 according to the embodiment of the present disclosure can provide an effective end-to-end device that can recognize facial micro-expressions by detecting the appearance of facial action units corresponding to each local region of the face. The image processing device 100 can obtain good detection results because it can utilize both action unit features and global facial features.

[0024] In particular, the image processing device 100 automatically recognizes facial micro-expressions by learning the relationships between action units (AUs). Due to physical limitations of facial muscles, when performing a particular facial expression, some AUs may appear simultaneously, while other AUs may not appear simultaneously. Furthermore, since AUs that appear simultaneously have similar AU features and AUs that do not appear simultaneously have different AU features, the similarity between local features (AUs) may be used as a weight for reweighting the local features. The action unit features processed in this way can recognize various facial expressions, thereby achieving better detection results.

[0025] The configurations of the information acquisition unit 110, the action unit feature extraction unit 120, and the global facial feature extraction unit 130 in the image processing device 100 according to an embodiment of the present disclosure will be described below with reference to FIGS.

[0026] 2 is a block diagram showing the configuration of the information acquisition unit 110 in the image processing device 100 according to an embodiment of the present disclosure. In FIG. 2, the information acquisition unit 110 is, for example, an information acquisition unit 200.

[0027] The information acquiring unit 200 has multiple convolution layers for acquiring information about facial action units in multiple regions of the input facial image. Preferably, the information about the facial action units may include information indicating the positions of the facial action units in the image. Preferably, the information acquiring unit 200 may acquire the information about the facial action units by converting the image into a feature map in a low-dimensional space.

[0028] 2, an input image is provided to a two-dimensional convolution (Conv) block 210. The Conv block 210 performs a convolution operation on the input image and may perform representation learning on the input image to abstract it. Furthermore, the image processed by the Conv block 210 is input to a division block 220, which may divide the input image into n regions.

[0029] Each of the n regions is then processed by a set of batch normalization (BatchNorm) blocks 230-1 to 230-n, parametric rectified linear unit (PReLU) blocks 240-1 to 240-n, and Conv blocks 250-1 to 250-n, where i in BatchNorm block 230-i, PReLU block 240-i, and Conv block 250-i in FIG. 2 is an integer greater than 1 and less than n.

[0030] The processing results output by the n sets of BatchNorm, PReLU, and Conv blocks are provided to a splicing block 260, which may splice the processing results of the n regions. For example, the splicing block 260 may splice the processing results of the n regions to the n regions divided by the division block 220. The processing results of the splicing block 260 are then provided to an addition block 270. The addition block 270 may add the image processed by the Conv block 210 and the result processed by the splicing block 260, and provide the added result sequentially to a BatchNorm block 280, a PReLU block 290, and a maxpooling block 295 for subsequent processing.

[0031] Finally, the information acquisition unit 200 provides the feature map as an output to the action unit feature extraction unit 120 and the global facial feature extraction unit 130 shown in FIG.

[0032] For example, the information acquisition unit 200 may be a uniform block learning module. The uniform block learning module may divide the input image into, for example, 8x8 regions, so that n is 64. Each of the 64 regions is processed by a set of BatchNorm, PReLU, and Conv blocks. Then, the output feature maps of each set of BatchNorm, PReLU, and Conv blocks are spliced ​​together and sequentially provided to the BatchNorm block 280, PReLU block 290, and Maxpooling block 295 after performing the addition operation as described above. Here, the final output feature map is also divided into 8x8 units. In this case, the output feature map may be referred to as a uniform feature map.

[0033] 3 is a block diagram showing the configuration of the action unit feature extraction unit 120 in the image processing device 100 according to an embodiment of the present disclosure. In FIG. 3, the action unit feature extraction unit 120 is, for example, an action unit feature extraction unit 300.

[0034] The action unit feature extraction unit 300 may extract action unit features from facial action unit regions based on information about the facial action units provided by the information acquisition unit. The action unit feature extraction unit 300 may be composed of several convolutional layers, and these convolutional layers are only involved in action unit regions. Preferably, the action unit feature extraction unit 300 may extract action unit features from each facial action unit region individually using multiple convolutional layers. The action unit features may be, for example, a final feature map of facial AU regions.

[0035] For example, as shown in FIG. 3 , the action unit feature extraction unit 300 is provided with feature maps of m AU regions of a facial image. More specifically, each of the feature maps of the m AU regions is input to a set of Conv blocks 310-1 to 310-m, BatchNorm blocks 320-1 to 320-m, and PReLU blocks 330-1 to 330-m, respectively. The processing results of the Conv blocks 310-1 to 310-m, BatchNorm blocks 320-1 to 320-m, and PReLU blocks 330-1 to 330-m are input to corresponding Maxpooling blocks 370-1 to 370-m, respectively. After processing by the Maxpooling blocks 370-1 to 370-m, a final feature map of the m AU regions of the face can be obtained. Preferably, m is 12, i.e., the input is an initial feature map of 12 AU regions of the face, and the output is a final feature map of the 12 AU regions of the face.

[0036] Note that for each AU region, multiple sets of Conv blocks, BatchNorm blocks, and PReLU blocks may be used to process the feature map of the input AU region. This allows further optimization of the action unit features, thereby improving the accuracy of the subsequent classification process by the classifier 140. For example, in FIG. 3, a set of Conv blocks 340-1 to 340-m, BatchNorm blocks 350-1 to 350-m, and PReLU blocks 360-1 to 360-m is added.

[0037] 4 is a block diagram showing the configuration of the global facial feature extraction unit 130 in the image processing device 100 according to an embodiment of the present disclosure. In FIG. 4, the global facial feature extraction unit 130 is, for example, a global facial feature extraction unit 400.

[0038] The global facial feature extraction unit 400 in Fig. 4 extracts global facial features based on information about facial action units provided by the information acquisition unit. The learning of global facial features is configured by several convolutional layers across the entire face. Preferably, the global facial feature extraction unit 400 extracts global facial features based on the feature map provided by the information acquisition unit through multiple convolutional layers.

[0039] 4, a feature map of a face image, such as a uniform feature map, is provided to the global facial feature extraction unit 400. Specifically, the feature map is input to a Conv block 410, a BatchNorm block 420, and a PReLU block 430, and the results after processing by the Conv block 410, the BatchNorm block 420, and the PReLU block 430 are input to a Maxpooling block 470. After processing by the Maxpooling block 470, a global feature map of the face can be obtained.

[0040] It should be noted that multiple sets of Conv, BatchNorm, and PReLU blocks may be used to process the input feature map, which can further optimize global facial features and improve the accuracy of the subsequent classifier 140. For example, in Figure 4, a set of Conv, BatchNorm, and PReLU blocks 440, 450, and 460 is added.

[0041] Next, the operation of the classifier 140 in FIG. 1 will be described. The classifier 140 may detect the occurrence of an action unit by combining action unit features and global facial features. For example, the classifier 140 may be a fully connected neural network. Alternatively, the classifier 140 may be configured with a few linear layers with softmax output, and the output may be the occurrence probability of all action units. In other words, the classifier 140 may classify facial action units by analyzing the occurrence probability of all facial action units. Preferably, the classifier 140 may use a linear layer with softmax output to analyze the occurrence probability of all facial action units.

[0042] Note that the use of the image processing device 100 according to the embodiment of the present disclosure has two phases. In the training phase, a network is trained using facial images labeled with the occurrence of action units. In the evaluation phase, the trained network is used to detect actions in test facial images and evaluate performance. For example, in the training phase, facial images, feature information about facial features (e.g., eyebrows, corners of the mouth, etc.), and linkage annotations about facial action units may be input to the image processing device shown in FIG. 1. Through the processing of the image processing device, predicted probabilities of various facial action units can be obtained.

[0043] The AU relationship learning process executed by the first calculation unit 150 and the second calculation unit 160 in the image processing device 100 according to the embodiment of the present disclosure will be described below with reference to FIGS.

[0044] 5 and 6 are schematic diagrams illustrating exemplary processing executed by the first calculation unit 150 and the second calculation unit 160 in the image processing device 100 according to an embodiment of the present disclosure. The first calculation unit 150 and the second calculation unit 160 use the similarity between local AU features as weights for re-weighting the local AU features. Note that FIGS. 5 and 6 only illustrate processing executed for the action unit features of 12 action units AU1 to AU12, and the first calculation unit 150 and the second calculation unit 160 are not limited to processing for 12 AUs, and the number of AUs may be other numbers.

[0045] The dashed frame in Fig. 5 indicates that the action unit feature extraction unit 120 extracts action unit features from the areas AU1 to AU12 of different face action units, and obtains feature maps of AU1 to AU12. (outside 1) TIFF0007746725000001.tif1264 is a process executed by the first calculation unit 150, and is the mark in FIG. (outside 2) TIFF0007746725000002.tif1464 is a process executed by the second calculation unit 160.

[0046] As can be seen from the curved arrow on the right side of the dashed frame, the first calculation unit 150 may calculate the similarity between the action unit features (here, feature maps) of each of the different facial action units AU1 to AU12 and the action unit features (here, feature maps) of each of the facial action units AU1 to AU12. Preferably, the first calculation unit 150 calculates the similarity by calculating the cosine distance between the action unit features. Furthermore, the result of the similarity calculated by the first calculation unit 150 is, for example, a similarity matrix shown in FIG. 5. Since 12 AUs are shown in FIG. 5, the size of the similarity matrix in FIG. 5 is 12×12.

[0047] process (Outside 3) In TIFF0007746725000003.tif1164, for a facial action unit, the second calculation unit 160 may determine the sum of the products of the similarity between the action unit features of the facial action unit and the action unit features of each facial action unit AU1 to AU12 and the action unit features of the corresponding facial action units AU1 to AU12 as the recalculated action unit features of the facial action unit.

[0048] For example, in Fig. 5, the first row of the similarity matrix represents the similarity between the AU1 feature map and each AU feature map. Therefore, as shown by the solid arrow in the lower right part of Fig. 5, for AU1, second calculation unit 160 determines the new feature map of AU1 as the sum of the product of the AU1 feature map and 1.0 (the similarity between the AU1 feature map and itself), the product of the AU2 feature map and 0.1 (the similarity between the AU1 feature map and the AU2 feature map), ..., and the product of the AU12 feature map and 0.7 (the similarity between the AU1 feature map and the AU12 feature map).

[0049] 5, the last row of the similarity matrix represents the similarity between the AU12 feature map and each AU feature map. Therefore, as indicated by the dashed arrow in the lower right part of Fig. 5, for AU12, second calculation unit 160 sets the new feature map of AU12 to the sum of the product of the AU1 feature map and 0.7 (the similarity between the AU12 feature map and the AU1 feature map), the product of the AU2 feature map and 0.4 (the similarity between the AU12 feature map and the AU2 feature map), ..., and the product of the AU12 feature map and 1.0 (the similarity between the AU12 feature map and itself).

[0050] This allows us to obtain a similarity matrix (12x12) by calculating the similarity between local AU features, and then use the similarity values ​​to re-weight the AU feature map to generate a new AU feature map, thereby obtaining new action unit features for each action unit AU1 to AU12. The new action unit features can reflect the features of AUs that appear simultaneously and AUs that do not appear simultaneously, allowing for more effective recognition of facial micro-expressions.

[0051] Fig. 6 is a schematic diagram illustrating another exemplary process performed by the first calculation unit 150 and the second calculation unit 160 in the image processing device 100 according to an embodiment of the present disclosure. The difference between the AU relation learning process performed in Fig. 6 and the AU relation learning process performed in Fig. 5 lies in the higher granularity and the lower granularity. In other words, Fig. 6 illustrates a learning process of high-granularity relations of AU feature maps.

[0052] The dashed lines in Fig. 6 indicate that the action unit feature extraction unit 120 extracts action unit features from the regions AU1 to AU12 of different face action units, thereby obtaining feature maps for AU1 to AU12. (outside 4) TIFF0007746725000004.tif1164 is a process executed by the first calculation unit 150, and is the mark in FIG. (outside 5) TIFF0007746725000005.tif1164 is a process executed by the second calculation unit 160.

[0053] In the process shown in FIG. 6, the action unit features of each facial action unit AU1 to AU12 may be represented by a predetermined number of matrices. The predetermined number may represent the number of channels of an image. For example, as can be seen from the right side of the dashed frame in FIG. 6, the feature map of each of AU1 to AU12 is roughly represented by multiple layers, each layer roughly representing one matrix, and the size of the matrix is, for example, 5×5. Note that although only three layers are shown in FIG. 6, in reality, 160 layers are represented. That is, the feature map of each of AU1 to AU12 in FIG. 6 may be represented by 160 matrices. Furthermore, in other examples, the number of layers (matrices) may be other numbers, and the size of the matrix may be other sizes.

[0054] As can be seen from the curved arrow on the right side of the dashed frame in FIG. 6, the first calculation unit 150 may calculate the similarity between each matrix of the action unit features (here, feature maps) of each facial action unit AU1 to AU12 and each matrix of the action unit features (here, feature maps) of each facial action unit AU1 to AU12. Preferably, the first calculation unit 150 calculates the similarity by calculating the cosine distance between the action unit features. Furthermore, the result of the similarity calculated by the first calculation unit 150 is, for example, the similarity matrix shown in FIG. 6. Since 12 AUs are shown in FIG. 6 and the feature map of each AU is represented by 160 matrices, the size of the similarity matrix in FIG. 6 is (12*160)×(12*160), i.e., 1920×1920.

[0055] process (outside 6) In TIFF0007746725000006.tif1164, for a facial action unit, the second calculation unit 160 may calculate the sum of the products of the similarity between each matrix among a predetermined number of matrices of the action unit features of the facial action unit and the action unit features of each facial action unit AU1 to AU12 and the matrix of the action unit features of the corresponding facial action unit AU1 to AU12 as the action unit feature represented by the recalculated predetermined number of matrices of the facial action unit.

[0056] 6, for example, the first row of the similarity matrix represents the similarity between the first matrix of the AU1 feature map and the matrices of each AU feature map. Therefore, as shown by the solid arrow in the lower right part of Fig. 5, for the first matrix of the AU1 feature map, second calculation unit 160 determines the new first matrix of the AU1 feature map as the sum of the product of the first matrix of the AU1 feature map and 1.0 (the similarity between the first matrix of the AU1 feature map and itself), the product of the second matrix of the AU1 feature map and 0.1 (the similarity between the first matrix of the AU1 feature map and the second matrix of the AU1 feature map), ..., and the product of the 160th matrix of the AU12 feature map and 0.7 (the similarity between the first matrix of the AU1 feature map and the 160th matrix of the AU12 feature map).

[0057] Using a similar method, new matrices 2 to 160 of the AU1 feature map can be calculated, and a new AU1 feature map represented by the recalculated matrices can be obtained.

[0058] As another example, the last row of the similarity matrix represents the similarity between the 160th matrix of the AU12 feature map and the matrices of each AU feature map. Therefore, for the 160th matrix of the AU12 feature map, the second calculation unit 160 sets the new 160th matrix of the AU12 feature map as the sum of the product of the first matrix of the AU1 feature map and 0.7 (the similarity between the 160th matrix of the AU12 feature map and the first matrix of the AU1 feature map), the product of the second matrix of the AU1 feature map and 0.4 (the similarity between the 160th matrix of the AU12 feature map and the second matrix of the AU1 feature map), ..., and the product of the 160th matrix of the AU12 feature map and 1.0 (the similarity between the 160th matrix of the AU12 feature map and itself).

[0059] Therefore, new matrices for the feature maps of AU1 to AU12 can be obtained, and thus feature maps represented by 160 recalculated matrices for each of AU1 to AU12 can be obtained.

[0060] This allows us to obtain a similarity matrix (1920 x 1920) by calculating the similarity between local AU features, and then use the similarity values ​​to re-weight the AU feature map to generate a new AU feature map, thereby obtaining new action unit features for each action unit AU1 to AU12. The new action unit features can reflect the features of AUs that appear simultaneously and AUs that do not appear simultaneously, allowing for more effective recognition of facial micro-expressions.

[0061] The image processing method according to the embodiment of the present disclosure will be described below with reference to FIG.

[0062] 7, the image processing method according to the embodiment of the present disclosure starts from step S110. In step S110, an input image is divided into a plurality of regions, and information about facial action units in the plurality of regions is obtained.

[0063] Then, in step S120, action unit features are extracted from the regions of the facial action units based on the acquired information about the facial action units.

[0064] Then, in step S130, the similarity between the action unit features of each facial action unit and the action unit features of each facial action unit among the different facial action units is calculated.

[0065] Then, in step S140, the action unit features of each facial action unit are recalculated based on the calculation results of the similarity.

[0066] Then, in step S150, global facial features are extracted based on the acquired information about the facial action units.

[0067] Then, in step S160, the facial action units are classified based on the recalculated action unit features of each facial action unit and the global facial features.

[0068] In an embodiment of the present disclosure, the step of recalculating the action unit features of the facial action unit includes a step of calculating the sum of the products of the similarities between the action unit features of the facial action unit and the action unit features of each corresponding facial action unit as the recalculated action unit features of the facial action unit.

[0069] In an embodiment of the present disclosure, the action unit features of each facial action unit are represented by a predetermined number of matrices, and the step of calculating the similarity includes the step of calculating the similarity between each matrix of the action unit features of each facial action unit and each matrix of the action unit features of each facial action unit.

[0070] In an embodiment of the present disclosure, the step of recalculating the action unit features of the facial action unit includes a step of obtaining the action unit features represented by the recalculated predetermined number of matrices of the facial action unit by calculating the sum of the products of the similarity between the matrix of the action unit features of the facial action unit and the matrix of the action unit features of each facial action unit and the matrix of the action unit features of each corresponding facial action unit as the recalculated matrix.

[0071] Therefore, the image processing method according to the embodiment of the present disclosure can recognize facial micro-expressions by detecting the appearance of facial action units corresponding to each local region of the face. In particular, the image processing method according to the embodiment of the present disclosure can automatically recognize facial micro-expressions by learning the relationship between action units (AUs), and can better recognize various facial expressions by using the similarity between local features (AUs).

[0072] In an embodiment of the present disclosure, the information about the facial action unit includes information indicating the position of the facial action unit in the image.

[0073] In an embodiment of the present disclosure, obtaining information about facial action units includes converting the image into a feature map in a low-dimensional space.

[0074] In an embodiment of the present disclosure, extracting action unit features includes extracting action unit features from each facial action unit region individually using multiple convolutional layers.

[0075] In an embodiment of the present disclosure, extracting global facial features includes extracting global facial features based on the feature map using multiple convolutional layers.

[0076] In an embodiment of the present disclosure, the multiple convolution layers include at least a batch normalization unit, a parametric normalized linear unit (PReLU), a two-dimensional convolution unit, and a max pooling unit.

[0077] In an embodiment of the present disclosure, the step of classifying the facial action units includes analyzing the occurrence probability of all facial action units.

[0078] In the embodiment of the present disclosure, a linear layer with softmax output is used to analyze the occurrence probability of all facial action units.

[0079] In an embodiment of the present disclosure, the similarity is calculated by calculating the cosine distance between the action unit features.

[0080] Note that steps S110 to S160 of the image processing method shown in Fig. 7 are not necessarily executed in the order shown in Fig. 7. Multiple steps may be executed in a different order or in parallel. For example, the processing in step S150 may be executed between steps S110 and S120, or may be executed after step S110 in parallel with step S120.

[0081] The specific embodiments of the above steps of the image processing method according to the embodiment of the present disclosure have been described in detail, so the description thereof will be omitted here.

[0082] Each process of the image processing method of the present disclosure may be realized by a computer-executable program stored in a storage medium readable by various devices.

[0083] The object of the present disclosure may also be realized by directly or indirectly providing a storage medium storing the executable program code to a system or device, and having a computer or central processing unit (CPU) in the system or device read and execute the program code. In this case, the system or device may have a function capable of executing a program, and embodiments of the present disclosure are not limited to programs. Furthermore, the program may be in any format, such as an object program, a program executed by an interpreter, or a script program provided to an operating system.

[0084] Storage media that can be read by the above-mentioned devices include, but are not limited to, various types of memories, storage units, semiconductor devices, disks such as optical disks, magnetic disks and magneto-optical disks, and other media that can store information.

[0085] Also, embodiments of the present disclosure can be implemented by connecting a computer to a corresponding website on the Internet and downloading, installing, and executing the computer program code of the present disclosure on the computer.

[0086] FIG. 8 is a block diagram illustrating an exemplary configuration of a general-purpose personal computer capable of implementing an image processing device and an image processing method according to an embodiment of the present disclosure.

[0087] 8, a CPU 1301 executes various processes according to programs stored in a read-only memory (ROM) 1302 or programs loaded from a storage unit 1308 into a random access memory (RAM) 1303. The RAM 1303 stores data necessary for the CPU 1301 to execute various processes as needed. The CPU 1301, ROM 1302, and RAM 1303 are connected to one another via a bus 1304. An input / output interface 1305 is also connected to the bus 1304.

[0088] An input unit 1306 (including a keyboard, a mouse, etc.), an output unit 1307 (including a display, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), and a speaker, etc.), a storage unit 1308 (including, for example, a hard disk, etc.), and a communication unit 1309 (including, for example, a network interface card, such as a LAN card or a modem, etc.) are connected to the input / output interface 1305. The communication unit 1309 performs communication processing via a network, such as the Internet. If necessary, a driver 1310 may be connected to the input / output interface 1305. A removable medium 1311 is, for example, a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, and is set up in the driver 1310 as necessary, and a computer program read from the removable medium 1311 is installed in the storage unit 1308 as necessary.

[0089] When the above processing is performed by software, a program constituting the software is installed via a network, such as the Internet, or a storage medium, such as a removable medium 1311 .

[0090] 8, which stores the program and provides the program to the user separately from the device. Removable medium 1311 includes, for example, a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including an optical disk-read only memory (CD-ROM) and a digital versatile disk (DVD)), a magneto-optical disk (a minidisk (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be ROM 1302, a hard disk included in storage unit 1308, or the like, which stores the program and is provided to the user together with the device containing the program.

[0091] In the system and method of the present disclosure, each unit or step may be decomposed and / or recombined. Such decomposition and / or recombination is considered equivalent to the present disclosure. Furthermore, the method of the present disclosure is not limited to being performed in the chronological order described in the specification, and may be performed sequentially, in parallel, or independently in another chronological order. Therefore, the execution order of the method described in this specification does not limit the technical scope of the present disclosure.

[0092] Although the embodiments of the present disclosure have been described in detail above with reference to the drawings, the above-described embodiments and examples are merely illustrative and do not limit the present disclosure. Those skilled in the art may make various modifications and changes to the present disclosure within the spirit and scope of the claims. These modifications and changes are intended to be included in the scope of protection of the present disclosure.

[0093] Furthermore, the following supplementary notes are disclosed regarding the embodiments including the above-described examples. (Appendix 1) an information acquisition unit that divides an input image into a plurality of regions and acquires information about face action units in the plurality of regions; an action unit feature extraction unit that extracts action unit features from the region of the face action unit based on the acquired information about the face action unit; a first calculation unit that calculates a similarity between an action unit feature of each facial action unit among different facial action units and an action unit feature of each facial action unit; a second calculation unit that recalculates the action unit features of each face action unit based on the calculation result of the similarity; a global facial feature extraction unit that extracts global facial features based on the acquired information about the facial action units; a classification unit that classifies the facial action units based on the recalculated action unit features of each facial action unit and the global facial features. (Appendix 2) The image processing device described in Appendix 1, wherein the second calculation unit calculates the recalculated action unit feature of the facial action unit by the sum of the products of the similarity between the action unit feature of the facial action unit and the action unit feature of each facial action unit and the action unit feature of each corresponding facial action unit. (Appendix 3) The action unit features of each facial action unit are represented by a predetermined number of matrices, The image processing device according to claim 1, wherein the first calculation unit calculates a similarity between each matrix of the action unit features of each facial action unit and each matrix of the action unit features of each facial action unit. (Appendix 4) The image processing device described in Appendix 3, wherein the second calculation unit obtains the action unit features represented by the recalculated predetermined number of matrices of the facial action unit by calculating the sum of the products of the similarity between the matrix of the action unit features of the facial action unit and the matrix of the action unit features of each facial action unit and the matrix of the action unit features of the corresponding facial action unit as a recalculated matrix for each matrix of the predetermined number of matrices of the action unit features of the facial action unit. (Appendix 5) The image processing device according to claim 1, wherein the action unit feature extraction unit individually extracts the action unit features from the region of each facial action unit using multiple convolutional layers. (Appendix 6) 2. The image processing device according to claim 1, wherein the classification unit classifies the facial action units by analyzing the occurrence probability of all facial action units. (Appendix 7) 7. The image processing device of claim 6, wherein the classifier uses a linear layer with a softmax output to analyze the occurrence probability of all the facial action units. (Appendix 8) 2. The image processing device according to claim 1, wherein the first calculation unit calculates the similarity by calculating a cosine distance between action unit features. (Appendix 9) Dividing an input image into a plurality of regions and obtaining information about facial action units in the plurality of regions; extracting action unit features from the regions of the facial action units based on the acquired information about the facial action units; calculating a similarity between the action unit features of each facial action unit and the action unit features of each facial action unit among the different facial action units; recalculating the action unit features of each facial action unit based on the similarity calculation result; extracting global facial features based on the acquired information about the facial action units; and classifying the facial action units based on the recalculated action unit features of each facial action unit and the global facial features. (Appendix 10) An image processing method as described in Appendix 9, wherein the step of recalculating the action unit features of the facial action unit includes a step of calculating the sum of the products of the similarity between the action unit features of the facial action unit and the action unit features of each facial action unit and the action unit features of each corresponding facial action unit as the recalculated action unit features of the facial action unit. (Appendix 11) The action unit features of each facial action unit are represented by a predetermined number of matrices, An image processing method as described in Appendix 9, wherein the step of calculating the similarity includes a step of calculating the similarity between each matrix of the action unit features of each facial action unit and each matrix of the action unit features of each facial action unit. (Appendix 12) The image processing method described in Appendix 11, wherein the step of recalculating the action unit features of the facial action unit includes a step of obtaining the action unit features represented by the recalculated predetermined number of matrices of the facial action unit by calculating the sum of the products of the similarity between the matrix of the action unit feature of the facial action unit and the matrix of the action unit feature of each facial action unit and the matrix of the action unit feature of each corresponding facial action unit as a recalculated matrix. (Appendix 13) 10. The image processing method of claim 9, wherein the step of extracting the action unit features includes a step of extracting the action unit features individually from the region of each facial action unit using multiple convolutional layers. (Appendix 14) 10. The image processing method of claim 9, wherein the step of classifying the facial action units includes a step of analyzing the occurrence probability of all facial action units. (Appendix 15) 15. The image processing method of claim 14, wherein a linear layer with a softmax output is used to analyze the probability of occurrence of all the facial action units. (Appendix 16) 10. The image processing method of claim 9, wherein the similarity is calculated by calculating a cosine distance between action unit features. (Appendix 17) 10. The image processing method of claim 9, wherein the information about the facial action unit includes information indicating the position of the facial action unit in the image. (Appendix 18) 10. The image processing method of claim 9, wherein the step of obtaining information about the facial action units includes the step of converting the image into a feature map in a low-dimensional space. (Appendix 19) 10. The image processing method of claim 9, wherein extracting global facial features comprises extracting the global facial features based on the feature map using multiple convolutional layers. (Appendix 20) A machine-readable storage medium having recorded thereon a program product storing machine-readable instruction codes, the program product being capable of causing the computer to execute an image processing method according to any one of appendices 9 to 19 when the instruction codes are read and executed by the computer.

Claims

1. an information acquisition unit that divides an input image into a plurality of regions and acquires a feature map related to facial action units in the plurality of regions; an action unit feature extraction unit that extracts action unit features from the region of the face action unit based on the acquired feature map related to the face action unit; a first calculation unit that calculates a similarity between an action unit feature of each facial action unit among different facial action units and an action unit feature of each facial action unit; a second calculation unit that recalculates the action unit features of each facial action unit based on the calculation result of the similarity; a global facial feature extraction unit for extracting global facial features based on the acquired feature map relating to the facial action units; a classifier for classifying the facial action units based on the recalculated action unit features of each facial action unit and the global facial features; the action unit feature extraction unit individually extracts the action unit features from the regions of each face action unit using a plurality of convolution layers; the second calculation unit calculates the sum of the products of the similarity between the action unit feature of the facial action unit and the action unit feature of each facial action unit and the action unit feature of each corresponding facial action unit as the recalculated action unit feature of the facial action unit; the global facial feature extraction unit extracts the global facial features based on the feature map using a plurality of convolutional layers; The classification unit classifies the facial action units by combining the action unit features and the global facial features and analyzing the appearance probability of all facial action units.

2. The action unit features of each facial action unit are represented by a predetermined number of matrices, The image processing device according to claim 1 , wherein the first calculation unit calculates a similarity between each matrix of action unit features of each facial action unit and each matrix of action unit features of each facial action unit.

3. 3. The image processing device according to claim 2, wherein the second calculation unit acquires the action unit features represented by the recalculated predetermined number of matrices of the facial action unit by calculating the sum of the products of the similarity between the matrix of the action unit feature of the facial action unit and the matrix of the action unit feature of each facial action unit and the matrix of the action unit feature of the corresponding facial action unit as a recalculated matrix.

4. The image processing device of claim 1 , wherein the classifier uses a linear layer with softmax output to analyze the occurrence probability of all the facial action units.

5. The image processing device according to claim 1 , wherein the first calculation unit calculates the similarity by calculating a cosine distance between action unit features.

6. Dividing an input image into a plurality of regions and obtaining feature maps for facial action units in the plurality of regions; extracting action unit features from the region of the facial action unit based on the obtained feature map for the facial action unit; calculating a similarity between the action unit features of each facial action unit and the action unit features of each facial action unit among the different facial action units; recalculating the action unit features of each facial action unit based on the similarity calculation result; extracting global facial features based on the obtained feature map for the facial action units; and classifying the facial action units based on the recalculated action unit features of each facial action unit and the global facial features; In the step of extracting the action unit features, the action unit features are individually extracted from the regions of each face action unit by a plurality of convolutional layers; In the step of recalculating the action unit feature, the sum of the products of the similarity between the action unit feature of the face action unit and the action unit feature of each face action unit and the action unit feature of each corresponding face action unit is set as the recalculated action unit feature of the face action unit; In the step of extracting global facial features, the global facial features are extracted based on the feature map by a plurality of convolutional layers; In the step of classifying the facial action units, the facial action units are classified by combining the action unit features and the global facial features and analyzing the occurrence probability of all facial action units.

7. A machine-readable storage medium on which a program product storing machine-readable instruction code is recorded, the program product being capable of causing the computer to execute the image processing method described in claim 6 when the instruction code is read and executed by the computer.

Citation Information

Patent Citations

  • Facial expression recognition method and device based on facial action unit

    CN111626113A

  • Expression analysis device and expression analysis program

    JP2015035172A

  • Systems and methods for facial expression recognition and annotation

    JP2019517693A