Image processing apparatus, image processing method and machine-readable storage medium

By using an image processing device based on action unit relationship learning, and leveraging AU feature similarity for automatic micro-facial expression recognition, the problem of recognition in existing technologies is solved, and efficient micro-facial expression detection is achieved.

CN114255488BActive Publication Date: 2025-10-31FUJITSU LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010947694.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-10
Publication Date
2025-10-31
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively recognize micro-facial expressions, and there is a lack of efficient automated recognition devices and methods.

Method used

An image processing device based on action unit (AU) relationship learning is used to perform micro-facial expression recognition by acquiring information, extracting action unit features, extracting global facial features, calculating similarity and recalculating features.

Benefits of technology

It achieves efficient micro-facial expression recognition, improves detection accuracy and automation, and can identify and classify facial action units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255488B_ABST
    Figure CN114255488B_ABST
Patent Text Reader

Abstract

This disclosure relates to an image processing apparatus, an image processing method, and a machine-readable storage medium. The image processing apparatus includes: an information acquisition unit that divides an input image into multiple regions and acquires information about facial action units in the multiple regions; an action unit feature extraction unit that extracts action unit features from regions of action units based on the information about the facial action units; a first calculation unit that calculates the similarity between the action unit features of each action unit and the action unit features of all other action units; a second calculation unit that recalculates the action unit features of each action unit based on the similarity calculation results; a global facial feature extraction unit that extracts global facial features based on the information about the facial action units; and a classification unit that classifies the facial action units based on the recalculated action unit features and the global facial features. The image processing apparatus can perform automatic micro-facial expression recognition based on action unit relationship learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and more particularly to image processing apparatus, image processing methods, and machine-readable storage media for micro-facial expression recognition. Background Technology

[0002] This section provides background information relating to this disclosure, which is not necessarily prior art.

[0003] Our faces carry a wealth of information about our psychological and emotional states at every moment. Therefore, mental state can be quantified based on micro-facial expression recognition. Micro-facial expression recognition has many applications in real life. For example, it can improve labor productivity by caring for employees' mental well-being and enhancing their motivation, or it can be used to estimate customer satisfaction and improve their purchasing motivation (digital marketing), or it can be used to estimate a driver's driving status, etc.

[0004] It is evident that developing an effective device for recognizing micro-facial expressions is of great significance. Summary of the Invention

[0005] This section provides a general overview of this disclosure, rather than a full disclosure of its entire scope or all its features.

[0006] The purpose of this disclosure is to provide an image processing apparatus, image processing method, and machine-readable storage medium for automatic micro-facial expression recognition based on action unit (AU) relationship learning.

[0007] According to one aspect of this disclosure, an image processing apparatus is provided, comprising: an information acquisition unit that divides an input image into multiple regions and acquires information about facial motion units in the multiple regions; a motion unit feature extraction unit that extracts motion unit features from regions of facial motion units based on the acquired information about the facial motion units; a first calculation unit that calculates the similarity between the motion unit features of each facial motion unit and the motion unit features of all facial motion units; a second calculation unit that recalculates the motion unit features of each facial motion unit based on the result of the similarity calculation; a global facial feature extraction unit that extracts global facial features based on the acquired information about the facial motion units; and a classification unit that classifies the facial motion units based on both the recalculated motion unit features of each facial motion unit and the global facial features.

[0008] According to another aspect of this disclosure, an image processing method is provided, comprising: dividing an input image into multiple regions and acquiring information about facial action units in the multiple regions; extracting action unit features of the regions of the facial action units based on the acquired information about the facial action units; calculating the similarity between the action unit features of each facial action unit and the action unit features of each facial action unit; recalculating the action unit features of each facial action unit based on the result of the similarity calculation; extracting global facial features based on the acquired information about the facial action units; and classifying the facial action units based on both the recalculated action unit features of each facial action unit and the global facial features.

[0009] According to another aspect of this disclosure, a machine-readable storage medium is provided that carries a program product including machine-readable instruction code stored thereon, wherein the instruction code, when read and executed by a computer, enables the computer to perform an image processing method according to this disclosure.

[0010] Using the image processing apparatus, image processing method, and machine-readable storage medium according to the present disclosure, micro-expressions can be identified by detecting the occurrence of facial motion units corresponding to each local region of the face.

[0011] Further applicability will become apparent from the description provided herein. The descriptions and specific examples in this summary are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0012] The accompanying drawings described herein are for illustrative purposes only and not for all possible implementations, and are not intended to limit the scope of this disclosure. In the drawings:

[0013] Figure 1 This is a block diagram illustrating the structure of an image processing apparatus according to an embodiment of the present disclosure;

[0014] Figure 2 This is a block diagram illustrating the structure of an information acquisition unit in an image processing apparatus according to an embodiment of the present disclosure;

[0015] Figure 3 This is a block diagram illustrating the structure of the action unit feature extraction unit in an image processing apparatus according to an embodiment of the present disclosure;

[0016] Figure 4 This is a block diagram illustrating the structure of a global facial feature extraction unit in an image processing apparatus according to an embodiment of the present disclosure;

[0017] Figure 5This is a schematic diagram illustrating exemplary processing performed by a first computing unit and a second computing unit in an image processing apparatus according to an embodiment of the present disclosure;

[0018] Figure 6 This is a schematic diagram illustrating another exemplary process performed by a first computing unit and a second computing unit in an image processing apparatus according to an embodiment of the present disclosure;

[0019] Figure 7 A flowchart illustrating an image processing method according to an embodiment of the present disclosure; and

[0020] Figure 8 This is a block diagram of an exemplary structure of a general-purpose personal computer in which image processing apparatus and methods according to embodiments of the present disclosure can be implemented.

[0021] While this disclosure is readily subject to various modifications and substitutions, specific embodiments thereof have been shown by way of example in the accompanying drawings and are described in detail herein. However, it should be understood that the description of specific embodiments herein is not intended to limit this disclosure to the specific forms disclosed, but rather, this disclosure is intended to cover all modifications, equivalents, and substitutions falling within the spirit and scope of this disclosure. It should be noted that throughout the drawings, corresponding reference numerals indicate corresponding parts. Detailed Implementation

[0022] Examples of this disclosure will now be described more fully with reference to the accompanying drawings. The following description is merely exemplary and is not intended to limit the disclosure, its application, or its uses.

[0023] Example embodiments are provided so that this disclosure will become exhaustive and will fully convey its scope to those skilled in the art. Numerous specific details, such as examples of particular components, apparatus, and methods, are set forth to provide a detailed understanding of embodiments of this disclosure. It will be apparent to those skilled in the art that the specific details are not required, and that the example embodiments may be implemented in many different forms, none of which should be construed as limiting the scope of this disclosure. In some example embodiments, well-known processes, well-known structures, and well-known techniques are not described in detail.

[0024] The following is combined Figure 1 This will illustrate how an image processing apparatus according to embodiments of the present disclosure recognizes micro-facial expressions.

[0025] Figure 1 A block diagram illustrating the structure of an image processing apparatus 100 according to an embodiment of the present disclosure is shown. Figure 1As shown, the image processing apparatus 100 according to an embodiment of this disclosure may include an information acquisition unit 110, an action unit feature extraction unit 120, a global facial feature extraction unit 130, a classification unit 140, a first calculation unit 150, and a second calculation unit 160. Here, it is assumed that the input image contains only one face. The micro-facial expression recognition task is to detect the presence of all facial action units of interest to the user. The presence or absence of a facial action unit can indicate whether the facial muscles in the facial region corresponding to that action unit are moving.

[0026] First, the information acquisition unit 110 can divide the input image into multiple regions and acquire information about facial motion units in the multiple regions. Furthermore, the information acquisition unit 110 can provide the acquired information about facial motion units in the multiple regions to the motion unit feature extraction unit 120 and the global facial feature extraction unit 130.

[0027] Furthermore, the motion unit feature extraction unit 120 can extract motion unit features from the regions of the facial motion units based on the information about the facial motion units provided by the information acquisition unit 110. In addition, the motion unit feature extraction unit 120 can provide the extracted motion unit features to the first calculation unit 150 and the second calculation unit 160.

[0028] Furthermore, the first calculation unit 150 can calculate the similarity between the motion unit features of each facial motion unit and the motion unit features of all facial motion units. In addition, the first calculation unit 150 can provide the similarity calculation results to the second calculation unit 160.

[0029] Furthermore, the second calculation unit 160 can recalculate the motion unit features of each facial motion unit based on the motion unit features extracted by the motion unit feature extraction unit 120 and the similarity calculation results provided by the first calculation unit 150. In addition, the second calculation unit 160 can provide the recalculated motion unit features to the classification unit 140.

[0030] Furthermore, the global facial feature extraction unit 130 can extract global facial features based on the information about the facial action unit provided by the information acquisition unit 110. Additionally, the global facial feature extraction unit 130 can provide the extracted global facial features to the classification unit 140.

[0031] Furthermore, the classification unit 140 can classify facial action units based on both the action unit features of each facial action unit recalculated by the second calculation unit 160 and the global facial features provided by the global facial feature extraction unit 130.

[0032] Therefore, the image processing apparatus 100 according to the embodiments of this disclosure can provide an effective end-to-end device that can perform micro-facial expression recognition by detecting the occurrence of facial motion units corresponding to each local region of the face. Since the image processing apparatus 100 can utilize both motion unit features and global facial features, it can obtain good detection results.

[0033] Specifically, the image processing device 100 performs automatic micro-facial expression recognition based on action unit (AU) relationship learning. Due to the physical limitations of facial muscles, some AUs may appear simultaneously while others may not, when making certain expressions. Furthermore, since co-occurring AUs should have similar AU features, and anti-co-occurring AUs should have different AU features, the similarity between local features (AUs) can be used as weights to reweight local features. Action unit features processed in this way are more helpful in recognizing different facial expressions, thereby obtaining better detection results.

[0034] The following is combined Figures 2 to 4 The structure of the information acquisition unit 110, the action unit feature extraction unit 120, and the global facial feature extraction unit 130 in the image processing apparatus 100 according to an embodiment of the present disclosure will be explained.

[0035] Figure 2 A block diagram illustrating the structure of the information acquisition unit 110 in an image processing apparatus 100 according to an embodiment of the present disclosure is shown. Figure 2 In this context, the information acquisition unit 110 is exemplarily an information acquisition unit 200.

[0036] The information acquisition unit 200 has multiple convolutional layers for acquiring information about facial action units in multiple regions of an input face image. Preferably, the information about facial action units may include information indicating the location of the facial action units in the image. Preferably, the information acquisition unit 200 can acquire information about facial action units by converting the image into a feature map in a low-dimensional space.

[0037] like Figure 2 As shown, the input image is fed to a two-dimensional convolution (Conv) block 210. The Conv block 210 performs convolution operations on the input image to learn and abstract its representation. Furthermore, the image processed by the Conv block 210 is input to a partitioning block 220, which divides the input image into n regions.

[0038] Subsequently, each of the n regions is processed through a set of batch normalization blocks 230-1 to 230-n, parametric rectified linear unit (PReLU) blocks 240-1 to 240-n, and Conv blocks 250-1 to 250-n. Figure 2 In the BatchNorm block 230-i, PReLU block 240-i, and Conv block 250-i, i is an integer greater than 1 and less than n.

[0039] The processing results from the n BatchNorm blocks, PReLU blocks, and Conv blocks are provided to the stitching block 260, which stitches together the processing results from these n regions. For example, the stitching block 260 can stitch together the processing results from these n regions corresponding to the division of these n regions by the partitioning block 220. The result processed by the stitching block 260 is then provided to the addition block 270. The addition block 270 adds the image processed by the Conv block 210 to the result processed by the stitching block 260, and sequentially provides the added result to the BatchNorm block 280, the PReLU block 290, and the Maxpooling block 295 for further processing.

[0040] Finally, the information acquisition unit 200 provides the feature map as output. Figure 1 The action unit feature extraction unit 120 and the global facial feature extraction unit 130 are shown.

[0041] For example, the information acquisition unit 200 can be exemplarily a uniform block learning module, which can divide the input image into, for example, 8×8 regions, thus n is 64. Each of the 64 regions is processed by a set of BatchNorm blocks, PReLU blocks, and Conv blocks. Then, the output feature maps of each set of BatchNorm blocks, PReLU blocks, and Conv blocks are concatenated and, after the addition operation described above, are sequentially fed into BatchNorm block 280, PReLU block 290, and Maxpooling block 295. Here, the final output feature map is also divided into 8×8 units. In such a case, the output feature map can also be called a uniform feature map.

[0042] Figure 3 A block diagram illustrating the structure of the motion unit feature extraction unit 120 in an image processing apparatus 100 according to an embodiment of the present disclosure is shown. Figure 3In this context, the action unit feature extraction unit 120 is exemplarily the action unit feature extraction unit 300.

[0043] The action unit feature extraction unit 300 can extract action unit features from regions of facial action units based on information about facial action units provided by the information acquisition unit. The action unit feature extraction unit 300 can consist of several convolutional layers, which only participate in the action unit regions. Preferably, the action unit feature extraction unit 300 can extract action unit features individually for each region of the facial action unit through multiple convolutional layers. The action unit features can be, for example, the final feature map of the AU region of the face.

[0044] For example, such as Figure 3 As shown, the action unit feature extraction unit 300 is provided with feature maps of m AU regions of a face image. More specifically, each of the feature maps of the m AU regions is input to a set of Conv blocks 310-1 to 310-m, BatchNorm blocks 320-1 to 320-m, and PReLU blocks 330-1 to 330-m, respectively. The results of processing by the Conv blocks 310-1 to 310-m, BatchNorm blocks 320-1 to 320-m, and PReLU blocks 330-1 to 330-m are then input to the corresponding Maxpooling blocks 370-1 to 370-m. After processing by the Maxpooling blocks 370-1 to 370-m, the final feature maps of the m AU regions on the face are obtained. Preferably, m is 12, meaning that the input is the initial feature maps of 12 AU regions of the face, and the output is the final feature maps of the 12 AU regions of the face.

[0045] Furthermore, it's important to note that for each AU region, multiple sets of Conv blocks, BatchNorm blocks, and PReLU blocks can be used to process the feature map of the input AU region. This allows for further optimization of the action unit features, thereby improving the accuracy of subsequent classification processing by the classification unit 140. For example, in Figure 3 In the middle, a set of Conv blocks 340-1 to 340-m, BatchNorm blocks 350-1 to 350-m, and PReLU blocks 360-1 to 360-m were added.

[0046] Figure 4 A block diagram illustrating the structure of a global facial feature extraction unit 130 in an image processing apparatus 100 according to an embodiment of the present disclosure is shown. Figure 4 In this context, the global facial feature extraction unit 130 is exemplarily a global facial feature extraction unit 400.

[0047] Figure 4The global facial feature extraction unit 400 extracts global facial features based on information about facial action units provided by the information acquisition unit. Global facial feature learning consists of several convolutional layers that cover the entire face. Preferably, the global facial feature extraction unit 400 extracts global facial features based on feature maps provided by the information acquisition unit through multiple convolutional layers.

[0048] like Figure 4 As shown, the global facial feature extraction unit 400 is provided with feature maps of the face image, such as uniform feature maps. Specifically, the feature maps are input to the Conv block 410, the BatchNorm block 420, and the PReLU block 430. The results after processing by the Conv block 410, the BatchNorm block 420, and the PReLU block 430 are input to the Maxpooling block 470. After processing by the Maxpooling block 470, the global feature map of the face can be obtained.

[0049] Furthermore, it's important to note that multiple sets of Conv blocks, BatchNorm blocks, and PReLU blocks can be used to process the input feature map. This allows for further optimization of global facial features, thereby improving the accuracy of subsequent classification unit 140. For example, in Figure 4 In this process, a set of Conv blocks 440, BatchNorm blocks 450, and PReLU blocks 460 were added.

[0050] Next, we will describe Figure 1 The operation of classification unit 140 is described below. Classification unit 140 can combine action unit features and global facial features to detect the occurrence of action units. For example, classification unit 140 can be a fully connected neural network. Alternatively, classification unit 140 can consist of a few linear layers with softmax outputs, the output of which is the probability of occurrence of all action units. In other words, classification unit 140 can classify facial action units by analyzing the probability of occurrence of all facial action units. Preferably, classification unit 140 can use linear layers with softmax outputs to analyze the probability of occurrence of all facial action units.

[0051] Furthermore, it should be noted that the image processing apparatus 100 according to embodiments of this disclosure has two phases: in the training phase, a network is trained using facial images with action unit appearance labels; in the evaluation phase, the trained network is used to detect actions in test facial images and evaluate performance. For example, in the training phase, facial images, feature information about facial feature regions (e.g., eyebrows, corners of the mouth, etc.), and linkage annotations about facial action units can be input into a system such as... Figure 1The image processing device shown can obtain the predicted probabilities of different facial action units after processing.

[0052] The following is combined Figure 5 and Figure 6 This is a schematic diagram illustrating AU relation learning processing performed by a first computing unit 150 and a second computing unit 160 in an image processing apparatus 100 according to an embodiment of the present disclosure.

[0053] Figure 5 and Figure 6 The illustration shows a schematic diagram of an exemplary process performed by a first computing unit 150 and a second computing unit 160 in an image processing apparatus 100 according to an embodiment of the present disclosure. The first computing unit 150 and the second computing unit 160 are used to reweight local AU features using the similarity between local AU features as weights. It should be noted that... Figure 5 and Figure 6 Only the processing performed on the action unit features of 12 action units AU1-AU12 is shown. It is clear that the first calculation unit 150 and the second calculation unit 160 are not limited to processing 12 AUs, and the number of AUs can be other numbers.

[0054] exist Figure 5 The dashed box in the image shows the action unit feature extraction unit 120 extracting action unit features from regions AU1 to AU12 of different facial action units, thereby obtaining feature maps for AU1 to AU12. Furthermore... Figure 5 The identifier ① in the text indicates the processing performed by the first computing unit 150, while... Figure 5 The identifier ② in the text refers to the processing performed by the second computing unit 160.

[0055] As can be seen from the curved arrow on the right side of the dashed box, the first calculation unit 150 can calculate the similarity between the motion unit features (here, feature maps) of each of the different facial motion units AU1 to AU12 and the motion unit features (here, feature maps) of each of the facial motion units AU1 to AU12. Preferably, the first calculation unit 150 calculates the similarity by calculating the cosine distance between the motion unit features. Furthermore, the similarity results calculated by the first calculation unit 150 are exemplarily shown in... Figure 5 The similarity matrix is ​​shown below. Because... Figure 5 The diagram shows 12 AUs, therefore Figure 5 The size of the similarity matrix is ​​12×12.

[0056] In processing ②, for a certain facial action unit, the second calculation unit 160 can use the sum of the similarity between the action unit features of a certain facial action unit and the action unit features of each facial action unit AU1 to AU12 and the corresponding product of the action unit features of each facial action unit AU1 to AU12 as the recalculated action unit features of a certain facial action unit.

[0057] For example, in Figure 5 In the similarity matrix, the first row represents the similarity between the AU1 feature map and each of the individual AU feature maps. Therefore, as... Figure 5 As shown by the solid arrow in the lower right corner, for AU1, the second calculation unit 160 uses the sum of the product of AU1 feature map and 1.0 (the similarity between AU1 feature map and itself), the product of AU2 feature map and 0.1 (the similarity between AU1 feature map and AU2 feature map), ..., the product of AU12 feature map and 0.7 (the similarity between AU1 feature map and AU12 feature map) as the new feature map of AU1.

[0058] As another example, in Figure 5 In the similarity matrix, the last row represents the similarity between the AU12 feature map and each of the individual AU feature maps. Therefore, as... Figure 5 As shown by the dashed arrow in the lower right corner, for AU12, the second calculation unit 160 takes the product of AU1 feature map and 0.7 (the similarity between AU12 feature map and AU1 feature map), the product of AU2 feature map and 0.4 (the similarity between AU12 feature map and AU2 feature map), ..., the product of AU12 feature map and 1.0 (the similarity between AU12 feature map and itself) as the new feature map of AU12.

[0059] Therefore, a similarity matrix (12×12) is obtained by calculating the similarity between local AU features. Then, the AU feature maps are reweighted using the similarity values ​​to generate new AU feature maps, resulting in new action unit features for each facial action unit AU1 to AU12. The new action unit features can reflect the characteristics of co-occurring AUs and anti-co-occurring AUs, and can more effectively identify micro-facial expressions.

[0060] Figure 6 The illustration shows a schematic diagram of another exemplary process performed by a first computing unit 150 and a second computing unit 160 in an image processing apparatus 100 according to an embodiment of the present disclosure. Figure 6 AU relation learning processing performed in the process Figure 5 The difference in AU relation learning processing performed in the process lies in fine-grained versus coarse-grained approaches. In other words, Figure 6 The illustration shows the fine-grained relation learning process of AU feature maps.

[0061] exist Figure 6 The dashed box in the image shows the action unit feature extraction unit 120 extracting action unit features from regions AU1 to AU12 of different facial action units, thereby obtaining feature maps for AU1 to AU12. Furthermore... Figure 6 The identifier ① in the text indicates the processing performed by the first computing unit 150, while... Figure 6 The identifier ② in the text refers to the processing performed by the second computing unit 160.

[0062] exist Figure 6 In the processing shown, the action unit features of each facial action unit AU1 to AU12 can be represented by a predetermined number of matrices. The predetermined number can represent the number of channels in the image. For example, from... Figure 6 As can be seen to the right of the dashed box, the feature maps of each of AU1 to AU12 are schematically represented by multiple layers, where each layer schematically represents a matrix, for example, 5×5. It is important to note that although... Figure 6 Only 3 layers are shown, but it actually represents 160 layers. That is, Figure 6 The feature maps of each of AU1 to AU12 can be represented by 160 matrices. Furthermore, in other examples, the number of layers (matrices) can be other numbers, and the size of the matrices can also be other sizes.

[0063] from Figure 6 As can be seen from the curved arrow to the right of the dashed box, the first calculation unit 150 calculates the similarity between each matrix of the motion unit features (here, feature maps) of each facial motion unit AU1 to AU12 and each matrix of the motion unit features (here, feature maps) of each facial motion unit AU1 to AU12. Preferably, the first calculation unit 150 calculates the similarity by calculating the cosine distance between the motion unit features. Furthermore, the results of the similarity calculated by the first calculation unit 150 are exemplarily shown in... Figure 6 The similarity matrix is ​​shown below. Because... Figure 6 The diagram shows 12 AUs, and the feature map of each AU is represented by 160 matrices, therefore Figure 6 The size of the similarity matrix is ​​(12*160)×(12*160), which is 1920×1920.

[0064] In processing ②, for a certain facial motion unit, the second calculation unit 160, for each of the predetermined number of matrices of motion unit features of a certain facial motion unit, takes the sum of the similarity between the matrix of motion unit features of a certain facial motion unit and the matrix of motion unit features of each facial motion unit AU1 to AU12 and the product of the corresponding matrix of motion unit features of each facial motion unit AU1 to AU12 as a recalculated matrix, thereby obtaining the motion unit features of a certain facial motion unit represented by the predetermined number of recalculated matrices.

[0065] For example, in Figure 6 In the similarity matrix, the first row represents the similarity between the first matrix of the AU1 feature map and the matrices of each of the individual AU feature maps. Therefore, as... Figure 6 As shown by the solid arrow in the lower right corner, for the first matrix of the AU1 feature map, the second calculation unit 160 takes the product of the first matrix of the AU1 feature map with 1.0 (the similarity between the first matrix of the AU1 feature map and itself), the product of the second matrix of the AU1 feature map with 0.1 (the similarity between the first matrix of the AU1 feature map and the second matrix of the AU1 feature map), ..., the product of the first sixty matrix of the AU12 feature map with 0.7 (the similarity between the first matrix of the AU1 feature map and the first sixty matrix of the AU12 feature map) as the new first matrix of the AU1 feature map.

[0066] Using the same method, new matrices from the second to the 160th of the AU1 feature map can be calculated. Thus, a new AU1 feature map represented by the recalculated matrices can be obtained.

[0067] As another example, the last row of the similarity matrix represents the similarity between the first 160 matrix of the AU12 feature map and the matrices of each AU feature map. Therefore, for the first 160 matrix of the AU12 feature map, the second calculation unit 160 takes the product of the first matrix of the AU1 feature map and 0.7 (the similarity between the first 160 matrix of the AU12 feature map and the first matrix of the AU1 feature map), the product of the second matrix of the AU1 feature map and 0.4 (the similarity between the first 160 matrix of the AU12 feature map and the second matrix of the AU1 feature map), ..., the product of the first 160 matrix of the AU12 feature map and 1.0 (the similarity between the first 160 matrix of the AU12 feature map and itself) as the new first 160 matrix of the AU12 feature map.

[0068] Therefore, each matrix of the new AU1 to AU12 feature maps can be obtained, thus obtaining the feature map of each AU1 to AU12 represented by 160 recalculated matrices.

[0069] Therefore, a similarity matrix (1920×1920) is obtained by calculating the similarity between local AU features. Then, the AU feature maps are reweighted using the similarity values ​​to generate new AU feature maps, resulting in new action unit features for each facial action unit AU1 to AU12. The new action unit features can reflect the characteristics of co-occurring AUs and anti-co-occurring AUs, and can more effectively identify micro-facial expressions.

[0070] The following is combined Figure 7 To describe an image processing method according to embodiments of the present disclosure.

[0071] like Figure 7 As shown, the image processing method according to an embodiment of the present disclosure begins at step S110. In step S110, the input image is divided into multiple regions, and information about facial motion units in the multiple regions is obtained.

[0072] Next, in step S120, action unit features are extracted from the regions of the facial action units based on the acquired information about the facial action units.

[0073] Next, in step S130, the similarity between the motion unit features of each facial motion unit and the motion unit features of each facial motion unit is calculated.

[0074] Next, in step S140, the motion unit features of each facial motion unit are recalculated based on the similarity calculation results.

[0075] Next, in step S150, global facial features are extracted based on the acquired information about the facial action units.

[0076] Next, in step S160, facial motion units are classified based on both the recalculated motion unit features of each facial motion unit and the global facial features. After this, the process ends.

[0077] According to an embodiment of this disclosure, the step of recalculating motion unit features for a certain facial motion unit includes: taking the sum of the similarity between the motion unit features of a certain facial motion unit and the motion unit features of each other facial motion unit and the product of the corresponding motion unit features of each other facial motion unit as the recalculated motion unit features of a certain facial motion unit.

[0078] According to embodiments of this disclosure, the motion unit features of each facial motion unit are represented by a predetermined number of matrices, and the step of calculating similarity includes calculating the similarity between each matrix of the motion unit features of each facial motion unit and each matrix of the motion unit features of each individual facial motion unit.

[0079] According to an embodiment of this disclosure, the step of recalculating motion unit features for a certain facial motion unit includes: for each of a predetermined number of matrices of motion unit features for a certain facial motion unit, taking the sum of the product of the similarity between the matrix of motion unit features of a certain facial motion unit and the matrix of motion unit features of each facial motion unit and the matrix of motion unit features of each corresponding facial motion unit as the recalculated matrix, thereby obtaining the motion unit features of a certain facial motion unit represented by the predetermined number of recalculated matrices.

[0080] Therefore, the image processing method according to embodiments of the present disclosure can perform micro-facial expression recognition by detecting the occurrence of facial action units corresponding to each local region of the face. More advantageously, the image processing method according to embodiments of the present disclosure performs automatic micro-facial expression recognition based on action unit (AU) relationship learning, which can better identify different facial expressions by utilizing the similarity between local features (AU).

[0081] According to embodiments of this disclosure, information about the facial motion unit includes information indicating the position of the facial motion unit in an image.

[0082] According to embodiments of this disclosure, obtaining information about facial motion units includes converting an image into a feature map in a low-dimensional space.

[0083] According to embodiments of this disclosure, the step of extracting action unit features includes: extracting action unit features individually for the region of each facial action unit through multiple convolutional layers.

[0084] According to embodiments of this disclosure, the step of extracting global facial features includes: extracting global facial features based on feature maps through multiple convolutional layers.

[0085] According to embodiments of this disclosure, the multiple convolutional layers include at least batch normalization units, parameterized corrected linear units (PReLU), two-dimensional convolutional units, and max pooling units.

[0086] According to embodiments of this disclosure, the step of classifying facial motion units includes analyzing the probability of occurrence of all facial motion units.

[0087] According to embodiments of this disclosure, a linear layer with softmax output is used to analyze the probability of occurrence of all facial motion units.

[0088] According to embodiments of this disclosure, similarity is calculated by calculating the cosine distance between action unit features.

[0089] It is important to note that Figure 7The individual steps S110 to S160 of the image processing method shown may not necessarily follow the prescribed procedure. Figure 7 The steps are executed in the order shown. Multiple steps may be executed in other orders or in parallel. For example, the processing in step S150 may be executed between steps S110 and S120, or it may be executed after step S110 and in parallel with step S120.

[0090] Various specific embodiments of the above-described steps of the image processing method according to the embodiments of this disclosure have been described in detail above and will not be repeated here.

[0091] Obviously, the various operational processes of the image processing method according to this disclosure can be implemented as a computer executable program stored in various machine-readable storage media.

[0092] Furthermore, the objective of this disclosure can also be achieved by providing a storage medium storing the aforementioned executable program code directly or indirectly to a system or device, and having a computer or central processing unit (CPU) in the system or device read and execute the aforementioned program code. In this case, as long as the system or device has the function of executing a program, the implementation of this disclosure is not limited to a program, and the program can be in any form, such as an object program, a program executed by an interpreter, or a script program provided to an operating system.

[0093] The aforementioned machine-readable storage media include, but are not limited to: various memories and storage units, semiconductor devices, disk units such as optical, magnetic and magneto-optical disks, and other media suitable for storing information.

[0094] Alternatively, the technical solution of this disclosure can also be implemented by connecting to a corresponding website on the Internet, downloading and installing the computer program code according to this disclosure onto the computer, and then executing the program.

[0095] Figure 8 This is a block diagram of an exemplary structure of a general-purpose personal computer in which image processing apparatus and methods according to embodiments of the present disclosure can be implemented.

[0096] like Figure 8 As shown, CPU 1301 executes various processes based on programs stored in read-only memory (ROM) 1302 or programs loaded into random access memory (RAM) 1303 from storage section 1308. RAM 1303 also stores data required as needed when CPU 1301 executes various processes, etc. CPU 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Input / output interface 1305 is also connected to bus 1304.

[0097] The following components are connected to the input / output interface 1305: input section 1306 (including keyboard, mouse, etc.), output section 1307 (including display, such as cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.), storage section 1308 (including hard disk, etc.), and communication section 1309 (including network interface card, such as LAN card, modem, etc.). The communication section 1309 performs communication processing via a network, such as the Internet. Drive 1310 may also be connected to the input / output interface 1305 as needed. Removable media 1311, such as disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1310 as needed, so that computer programs read from them can be installed into storage section 1308 as needed.

[0098] When the above series of processes are implemented by software, the program constituting the software is installed from a network such as the Internet or a storage medium such as removable media 1311.

[0099] Those skilled in the art will understand that such storage media are not limited to Figure 8 The illustrated removable medium 1311 stores a program and is distributed separately from the device to provide the program to the user. Examples of removable media 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including optical disc read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including mini-disk (MD) (registered trademark)), and semiconductor memory. Alternatively, the storage medium may be ROM 1302, a hard disk included in storage section 1308, etc., containing programs and distributed to the user along with the device containing them.

[0100] In the systems and methods of this disclosure, it is apparent that the components or steps can be decomposed and / or recombined. Such decomposition and / or recombination should be considered equivalent to those disclosed. Furthermore, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order. Some steps can be performed in parallel or independently of each other.

[0101] While embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative and do not constitute a limitation thereof. Those skilled in the art can make various modifications and alterations to the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is defined only by the appended claims and their equivalents.

[0102] Regarding the implementation methods including the above embodiments, the following notes are also disclosed:

[0103] Appendix 1. An image processing apparatus, comprising:

[0104] An information acquisition unit divides the input image into multiple regions and acquires information about facial motion units in the multiple regions;

[0105] An action unit feature extraction unit extracts action unit features from regions of the facial action unit based on the acquired information about the facial action unit.

[0106] The first calculation unit calculates the similarity between the motion unit features of each facial motion unit and the motion unit features of each facial motion unit.

[0107] The second calculation unit recalculates the motion unit features of each facial motion unit based on the similarity calculation results.

[0108] A global facial feature extraction unit extracts global facial features based on the acquired information about the facial action unit; and

[0109] The classification unit classifies the facial action units based on both the action unit features of each recalculated facial action unit and the global facial features.

[0110] Appendix 2. The image processing apparatus according to Appendix 1, wherein,

[0111] For a certain facial motion unit, the second calculation unit uses the sum of the similarity between the motion unit feature of the certain facial motion unit and the motion unit features of each other, and the product of the corresponding motion unit features of each other, as the recalculated motion unit feature of the certain facial motion unit.

[0112] Appendix 3. The image processing apparatus according to Appendix 1, wherein,

[0113] Each facial motion unit's motion unit features are represented by a predetermined number of matrices, and the first computing unit calculates the similarity between each matrix of the motion unit features of each facial motion unit and each matrix of the motion unit features of each facial motion unit.

[0114] Appendix 4. The image processing apparatus according to Appendix 3, wherein,

[0115] For a certain facial motion unit, the second calculation unit, for each of the predetermined number of matrices of motion unit features of the certain facial motion unit, takes the sum of the product of the similarity between the matrix of motion unit features of the certain facial motion unit and the matrix of motion unit features of each other and the matrix of motion unit features of each other respectively as a recalculated matrix, thereby obtaining the motion unit features of the certain facial motion unit represented by the recalculated predetermined number of matrices.

[0116] Appendix 5. The image processing apparatus according to Appendix 1, wherein,

[0117] The action unit feature extraction unit extracts the action unit features individually for each facial action unit region through multiple convolutional layers.

[0118] Appendix 6. The image processing apparatus according to Appendix 1, wherein,

[0119] The classification unit classifies the facial motion units by analyzing the probability of occurrence of all facial motion units.

[0120] Note 7. The image processing apparatus according to Note 6, wherein,

[0121] The classification unit uses a linear layer with softmax output to analyze the probability of occurrence of all facial action units.

[0122] Appendix 8. The image processing apparatus according to Appendix 1, wherein,

[0123] The first calculation unit calculates the similarity by calculating the cosine distance between the features of the action units.

[0124] Appendix 9. An image processing method, comprising:

[0125] The input image is divided into multiple regions, and information about facial motion units in the multiple regions is obtained;

[0126] Based on the acquired information about the facial action unit, action unit features are extracted from the region of the facial action unit;

[0127] Calculate the similarity between the motion unit features of each facial motion unit and the motion unit features of all facial motion units;

[0128] The motion unit features of each facial motion unit are recalculated based on the similarity calculation results;

[0129] Global facial features are extracted based on the acquired information about the facial action units; and

[0130] The facial motion units are classified based on both the motion unit features of each recalculated facial motion unit and the global facial features.

[0131] Note 10. According to the method described in Note 9, wherein,

[0132] The step of recalculating the action unit features for a certain facial action unit includes: taking the sum of the similarity between the action unit features of the certain facial action unit and the action unit features of each other facial action unit and the product of the corresponding action unit features of each other facial action unit as the recalculated action unit features of the certain facial action unit.

[0133] Appendix 11. According to the method described in Appendix 9, wherein,

[0134] Each facial motion unit's motion unit features are represented by a predetermined number of matrices, and the step of calculating the similarity includes calculating the similarity between each matrix of the motion unit features of each facial motion unit and each matrix of the motion unit features of each individual facial motion unit.

[0135] Appendix 12. According to the method described in Appendix 11, wherein,

[0136] The step of recalculating the action unit features for a certain facial action unit includes: for each of the predetermined number of matrices of action unit features of the certain facial action unit, taking the sum of the product of the similarity between the matrix of action unit features of the certain facial action unit and the matrix of action unit features of each facial action unit and the matrix of action unit features of each corresponding facial action unit as the recalculated matrix, thereby obtaining the action unit features of the certain facial action unit represented by the recalculated predetermined number of matrices.

[0137] Note 13. According to the method described in Note 9, wherein,

[0138] The step of extracting the action unit features includes: extracting the action unit features individually for the region of each facial action unit through multiple convolutional layers.

[0139] Appendix 14. According to the method described in Appendix 9, wherein,

[0140] The step of classifying the facial motion units includes analyzing the probability of occurrence of all facial motion units.

[0141] Note 15. According to the method described in Note 14, wherein,

[0142] A linear layer with softmax output is used to analyze the probability of occurrence of all facial motion units.

[0143] Note 16. According to the method described in Note 9, wherein,

[0144] The similarity is calculated by calculating the cosine distance between action unit features.

[0145] Note 17. According to the method described in Note 9, wherein,

[0146] Information about the facial motion unit includes information indicating the location of the facial motion unit in the image.

[0147] Note 18. According to the method described in Note 9, wherein,

[0148] Obtaining information about the facial motion unit includes converting the image into a feature map in a low-dimensional space.

[0149] Note 19. According to the method described in Note 9, wherein,

[0150] The step of extracting the global facial features includes: extracting the global facial features based on the feature map through multiple convolutional layers.

[0151] Appendix 20. A machine-readable storage medium carrying a program product including machine-readable instruction code stored thereon, wherein, when read and executed by a computer, the instruction code enables the computer to perform the image processing method according to Appendices 9-19.

Claims

1. An image processing device for micro-facial expression recognition, comprising: An information acquisition unit divides the input image into multiple regions and acquires information about facial motion units in the multiple regions; An action unit feature extraction unit extracts action unit features from regions of the facial action unit based on the acquired information about the facial action unit. The first calculation unit calculates the similarity between the motion unit features of each facial motion unit and the motion unit features of each facial motion unit. The second calculation unit recalculates the motion unit features of each facial motion unit based on the similarity calculation results. A global facial feature extraction unit extracts global facial features based on the information obtained about the facial action unit; as well as The classification unit classifies the facial action units based on both the action unit features of each recalculated facial action unit and the global facial features. Specifically, for a certain facial motion unit, the second calculation unit uses the sum of the similarity between the motion unit features of the certain facial motion unit and the motion unit features of each other, and the product of the corresponding motion unit features of each other, as the recalculated motion unit feature of the certain facial motion unit.

2. The image processing apparatus according to claim 1, wherein, Each facial motion unit's motion unit features are represented by a predetermined number of matrices, and the first computing unit calculates the similarity between each matrix of the motion unit features of each facial motion unit and each matrix of the motion unit features of each facial motion unit.

3. The image processing apparatus according to claim 2, wherein, For a certain facial motion unit, the second calculation unit, for each of the predetermined number of matrices of motion unit features of the certain facial motion unit, takes the sum of the product of the similarity between the matrix of motion unit features of the certain facial motion unit and the matrix of motion unit features of each other and the matrix of motion unit features of each other respectively as a recalculated matrix, thereby obtaining the motion unit features of the certain facial motion unit represented by the recalculated predetermined number of matrices.

4. The image processing apparatus according to claim 1, wherein, The action unit feature extraction unit extracts the action unit features individually for each facial action unit region through multiple convolutional layers.

5. The image processing apparatus according to claim 1, wherein, The classification unit classifies the facial motion units by analyzing the probability of occurrence of all facial motion units.

6. The image processing apparatus according to claim 5, wherein, The classification unit uses a linear layer with softmax output to analyze the probability of occurrence of all facial action units.

7. The image processing apparatus according to claim 1, wherein, The first calculation unit calculates the similarity by calculating the cosine distance between the features of the action units.

8. An image processing method for micro-facial expression recognition, comprising: The input image is divided into multiple regions, and information about facial motion units in the multiple regions is obtained; Based on the acquired information about the facial action unit, action unit features are extracted from the region of the facial action unit; Calculate the similarity between the motion unit features of each facial motion unit and the motion unit features of all facial motion units; The motion unit features of each facial motion unit are recalculated based on the similarity calculation results; Global facial features are extracted based on the information obtained about the facial action unit; as well as The facial motion units are classified based on both the motion unit features of each recalculated facial motion unit and the global facial features. Specifically, for a certain facial motion unit, the sum of the similarity between the motion unit feature of the certain facial motion unit and the motion unit features of each other, and the product of the corresponding motion unit features of each other, is used as the recalculated motion unit feature of the certain facial motion unit.

9. A machine-readable storage medium having thereon carrying a program product including machine-readable instruction code stored thereon, wherein, When the instruction code is read and executed by a computer, it enables the computer to perform the image processing method according to claim 8.

Citation Information

Patent Citations

  • Facial expression recognition method and device based on facial action unit

    CN111626113A

  • Human face identifying method based on structural principal element analysis

    CN1975759A