Learning program, identification program, learning method, and identification method

The learning program improves AU detection accuracy by training a machine learning model to differentiate between images with and without occlusions, addressing the challenge of reduced accuracy due to facial obstructions.

JP7841381B2Active Publication Date: 2026-04-07FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Conventional methods for detecting Action Units (AUs) in facial expressions are hindered by occlusions such as hair or masks, leading to decreased accuracy in identifying the presence or absence of AUs.

Method used

A learning program and method that involves acquiring and classifying facial images with and without occlusions, training a machine learning model to minimize the distance between images with and without occlusions for the same AU, and maximize the distance between images with and without AUs, using a feature calculation and discrimination model to improve AU detection accuracy.

Benefits of technology

Enhances the discrimination accuracy of AUs by mitigating the effects of occlusions, allowing accurate identification of facial expressions even in partially obscured images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841381000003
    Figure 0007841381000003
  • Figure 0007841381000004
    Figure 0007841381000004
  • Figure 0007841381000005
    Figure 0007841381000005
Patent Text Reader

Abstract

To improve AU identification accuracy.SOLUTION: A training program causes a computer to execute the processing of acquisition, classification, calculation, and training. The acquisition processing acquires a plurality of images that includes a face of a person. The classification processing classifies the plurality of images on the basis of a combination of whether or not an AU related to a motion of a specific portion of the face occurs and whether or not an occlusion is included in an image in which the AU occurs. The calculation processing calculates a feature amount of the image by inputting each of the plurality of classified images into a machine learning model. The training processing trains the machine learning model so as to decrease a first distance between feature amounts of an image in which the AU occurs and an image with an occlusion with respect to the image in which the AU occurs and to increase a second distance between feature amounts of the image with the occlusion with respect to the image in which the AU occurs and an image with an occlusion with respect to an image in which the AU does not occur.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention relate to a learning program, an identification program, a learning method, and an identification method. [Background technology]

[0002] Recent advancements in image processing technology have led to the development of systems that can detect subtle changes in a person's psychological state from their facial expressions (surprise, joy, sadness, etc.) and perform processing in accordance with these changes. One of the representative methods for describing changes in facial expressions used in this expression detection is the description of expressions using AUs (Action Units) (an expression includes a combination of multiple AUs).

[0003] AU (Action Unit) is a unit of facial movement that quantifies facial expressions by breaking them down based on facial parts and facial muscles. Dozens of types are defined, corresponding to the movements of facial muscles, such as AU1 (raising the inner part of the eyebrow), AU4 (lowering the eyebrow), and AU12 (lifting both corners of the lips). During facial expression detection, the Occurrence (presence or absence) of these AUs is identified from the target facial image, and subtle changes in facial expression are recognized based on the generated AUs.

[0004] Conventional techniques for identifying the presence or absence of each AU from a facial image include those that use machine learning to recognize facial image data and identify the presence or absence of each AU based on the output obtained. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] JAA-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] However, in the above prior art, when there is occlusion (hereinafter also referred to as "occlusion") due to hair, mask, etc. in a part of the face image, there is a problem that the discrimination accuracy of the presence or absence of each AU deteriorates. For example, in a face image, when a part of the generation site of a certain AU is partially occluded, it becomes difficult to recognize whether that part has moved. As an example, when a part of the space between the eyebrows is hidden by hair, it becomes difficult to recognize the movement of the eyebrows such as AU4 (lowering the eyebrows).

[0007] On one side, an object is to provide a learning program, a discrimination program, a learning method, and a discrimination method that can improve the discrimination accuracy of AUs.

Means for Solving the Problems

[0008] In one proposal, the learning program causes a computer to execute an acquisition process, a classification process, a calculation process, and a learning process. The acquisition process acquires a plurality of images including a person's face. The classification process classifies the plurality of images based on a combination of the presence or absence of the generation of an action unit related to the movement of a specific part of the face and the presence or absence of occlusion for an image with the generation of the action unit. The calculation process inputs each of the classified plurality of images into a machine learning model to calculate the feature amount of the image. The learning process learns the machine learning model so that a first distance between the feature amounts of an image with the generation of the action unit and an image with occlusion for an image with the generation of the action unit becomes small, and a second distance between the feature amounts of an image with occlusion for an image with the generation of the action unit and an image with occlusion for an image without the generation of the action unit becomes large.

Effects of the Invention

[0009] The discrimination accuracy of AUs can be improved.

Brief Description of the Drawings

[0010] [Figure 1]FIG. 1 is an explanatory diagram for explaining an example of a face image. [Figure 2] FIG. 2 is an explanatory diagram for explaining feature amount calculation. [Figure 3] FIG. 3 is an explanatory diagram for explaining learning of feature amount calculation. [Figure 4] FIG. 4 is an explanatory diagram for explaining discrimination learning from feature amounts. [Figure 5] FIG. 5 is a block diagram showing a functional configuration example of an information processing apparatus according to the first embodiment. [Figure 6] FIG. 6 is a flowchart showing an operation example of an information processing apparatus according to the first embodiment. [Figure 7] FIG. 7 is a block diagram showing a functional configuration example of an information processing apparatus according to the second embodiment. [Figure 8] FIG. 8 is a flowchart showing an operation example of an information processing apparatus according to the second embodiment. [Figure 9] FIG. 9 is an explanatory diagram for explaining an example of a computer configuration. [Figure 10] FIG. 10 is a diagram showing an example of an expression recognition rule. [Embodiment for Implementing the Invention]

[0011] Hereinafter, a learning program, a discrimination program, a learning method, and a discrimination method according to the embodiment will be described with reference to the drawings. In the embodiment, components having the same function are denoted by the same reference numerals, and redundant descriptions are omitted. Note that the learning program, the discrimination program, the learning method, and the discrimination method described in the following embodiments are merely examples and do not limit the embodiments. Also, the following embodiments may be appropriately combined within a non - conflicting range.

[0012] [Expression Recognition System] This section describes the overall configuration of the facial expression recognition system used in this implementation. The facial expression recognition system comprises multiple cameras and an information processing device that performs video data analysis. The information processing device recognizes a person's facial expression from facial images captured by the cameras using a facial expression recognition model. The facial expression recognition model is an example of a machine learning model that generates facial expression information, which is an example of a person's features. Specifically, the facial expression recognition model is a machine learning model that estimates Action Units (AUs), which are a method for decomposing and quantifying facial expressions based on facial parts and facial muscles. In response to the input image data, this facial expression recognition model outputs facial expression recognition results such as "AU1:2, AU2:5, AU4:1, ..." which are expressed in terms of the intensity of each AU from AU1 to AU28 set to identify the facial expression (for example, on a 5-point scale).

[0013] Facial expression recognition rules are rules for recognizing facial expressions using the output results of a facial expression recognition model. Figure 10 shows an example of facial expression recognition rules. As shown in Figure 10, facial expression recognition rules store a correspondence between "facial expressions" and "estimated results". "Facial expressions" are the facial expressions to be recognized, and "estimated results" are the intensities of each AU from AU1 to AU28 that correspond to each facial expression. In the example in Figure 10, it is shown that if "AU1 has an intensity of 2, AU2 has an intensity of 5, AU3 has an intensity of 0...", the facial expression is recognized as a "smile". Facial expression recognition rules are data that has been registered in advance by an administrator or other person.

[0014] (Summary of the embodiment) Figure 1 is an explanatory diagram illustrating an example of a face image. As shown in Figure 1, face images 100 and 101 are images that include a person's face 110. When there is no occlusion on the face, as in face image 100's face 110, the presence or absence of frown lines (AU04) can be correctly identified.

[0015] In contrast, as in face image 101, when part of the area between the eyebrows is obscured by the hair 111 of face 110 (occlusion), the occlusion makes it difficult to see wrinkles in the skin between the eyebrows, and the edges of the hair 111 may be mistakenly recognized as wrinkles. Therefore, with conventional recognition models, when there is occlusion in the area between the eyebrows, it becomes difficult to correctly identify whether or not wrinkles are present between the eyebrows (AU04).

[0016] Figure 2 is an explanatory diagram illustrating feature calculation. As shown in Figure 2, in the information processing device according to the embodiment, face images 100a, 100b, and 100c, each classified into several patterns, are input to the feature calculation model M1, and feature quantities (first feature quantity 120a, second feature quantity 120b, and third feature quantity 120c) for each image are calculated. In the following description, unless otherwise distinguished, the feature quantities for each image will be referred to as feature quantity 120.

[0017] Here, the feature extraction model M1 is a machine learning model that calculates and outputs 120 features related to an input image. This feature extraction model M1 can utilize neural networks such as GMCNN (Generative Multi-column Convolutional Neural Networks) or GAN (Generative Adversarial Networks). The input image to this feature extraction model M1 may be a still image or a sequence of images in chronological order. Furthermore, the 120 features calculated by the feature extraction model M1 can be any information that represents the characteristics of the input image, such as vector information indicating the movement of facial muscles in the image, or the intensity of each AU (Auditory Unit).

[0018] Face image 100a is an image in which the action unit (AU) of lifting both corners of the lips (AU15) occurs in face 110 (no occlusion occurs). The first feature 120a is the feature calculated by inputting this face image 100a into the feature calculation model M1. In this embodiment, the presence or absence of lifting both corners of the lips (AU15) is shown as an example, but the AU is not limited to AU15 and can be arbitrary.

[0019] Face image 100b is an image in which occlusion occurs in face 110, where an AU (augmentation) of the corners of the lips is raised (AU15), due to an obstruction 112 around the mouth. The second feature 120b is the feature calculated by inputting this face image 100b into the feature calculation model M1.

[0020] Face image 100c is an image in which occlusion occurs around the mouth due to an obstruction 112, in face 110 where AU (augmentation) of the corners of the lips (AU15) is not occurring. The third feature 120c is the feature calculated by inputting this face image 100c into the feature calculation model M1. In the following explanation, unless otherwise specified, face images 100a, 100b, and 100c will be referred to as face image 100.

[0021] The information processing device according to this embodiment calculates a first distance (d) between a first feature quantity 120a of a face image 100a with AU (no occlusion) and a second feature quantity 120b of a face image 100b with occlusion relative to the face image 100a with AU. o Next, the information processing device according to the embodiment calculates the first distance (d o The feature acquisition model M1 is trained so that ) becomes small.

[0022] Furthermore, the information processing device according to the embodiment measures the second distance (d) between the second feature quantity 120b of the occluded face image 100b relative to the face image 100a with AU occurrence and the third feature quantity 120c of the occluded face image 100c relative to the image without AU occurrence. auObtain (). Next, the information processing apparatus according to the embodiment learns the feature amount calculation model M1 so that the second distance (d au ) increases.

[0023] For example, the information processing apparatus obtains a feature amount from a neural network by inputting a face image 100 into the neural network. Then, the information processing apparatus generates a machine learning model that changes the parameters of the neural network so that the error from the correct data is reduced in the obtained feature amount. The first distance (d o ) decreases, and the feature amount calculation model M1 is learned so that the second distance (d au ) increases.

[0024] FIG. 3 is an explanatory diagram for explaining the learning of feature amount calculation. As shown in FIG. 3, the information processing apparatus according to the embodiment learns the feature amount calculation model M1 so that the first distance (d o ) decreases and the second distance (d au ) increases, thereby reducing the influence of occlusion by the occluder 112 on the feature amount output by the feature amount calculation model M1.

[0025] For example, when face images 100a and 100b with the occurrence of AU are input to the learned feature amount calculation model M1, it becomes difficult for a difference due to the presence or absence of occlusion to occur in the feature amount. Further, when face images 100b and 100c that both have occlusion but differ in the occurrence of AU are input to the learned feature amount calculation model M1, a difference due to the occurrence of AU is likely to occur in the feature amount.

[0026] Note that the information processing apparatus according to the embodiment learns the feature amount calculation model M1 in which the first distance (d o ) decreases and the second distance (d au ) increases based on the following loss function (Loss) of Equation (1). Here, m o , m au are the first distance (d o ), the second distance (dau This is a margin parameter related to the loss function. This margin parameter adjusts the distance margin when calculating the loss function, and can be set to a value arbitrarily determined by the user.

[0027]

number

[0028] In the loss function (Loss) of equation (1), the first distance (d o The loss will be large if the second distance (d) is large and even if AU occurs between the features, there will be a difference between them due to the presence or absence of occlusion. Also, in the loss function (Loss) of equation (1), the second distance (d) au The loss becomes large when the ) is small and there are differences in whether or not AU occurs between them, but occlusion does not create a difference between the features.

[0029] Furthermore, the information processing device according to the embodiment learns a discrimination model that identifies the presence or absence of AUs based on the features obtained by inputting an image to the feature calculation model M1 on which correct information indicating the presence or absence of AUs has been added. This discrimination model may be a machine learning model using a neural network separate from the feature calculation model M1, or it may be a discrimination layer placed after the feature calculation model M1.

[0030] Figure 4 is an explanatory diagram illustrating discriminative learning from features. As shown in Figure 4, the information processing device according to the embodiment inputs a face image 100, to which ground truth information indicating the presence or absence of AUs has been added, into a trained feature calculation model M1 to obtain features 120. Here, the ground truth information is an array (AU1, AU2, ...) indicating the presence or absence of each AU. For example, if the array (1, 0, ...) is added to the face image 100 as ground truth information, then the presence or absence of AU1 is indicated in the face image 100.

[0031] The information processing device according to this embodiment learns the identification model M2 by updating its parameters so that, when the feature quantity 120 is input to the identification model M2, the identification model M2 outputs a value corresponding to the presence or absence of AU as indicated by the correct answer information. The information processing device according to this embodiment can identify the presence or absence of AU in a face from a face image to be identified by using the feature calculation model M1 and the identification model M2 that have been learned in this way.

[0032] For example, the information processing device inputs feature vector 120 into a neural network to obtain a feature vector indicating whether or not an AU (Automatic Output) is generated from the neural network. The information processing device then generates a machine learning model by modifying the parameters of the neural network so that the error with the ground truth data is minimized using the obtained feature vector.

[0033] (First Embodiment) Figure 5 is a block diagram showing an example of the functional configuration of an information processing device according to the first embodiment. As shown in Figure 5, the information processing device 1 includes an image input unit 11, a face region extraction unit 12, a partially obscured image generation unit 13, an AU comparison image generation unit 14, an image database 15, an image set generation unit 16, a feature calculation unit 17, a distance calculation unit 18, a distance learning execution unit 19, an AU recognition learning execution unit 20, and an identification unit 21.

[0034] The image input unit 11 is a processing unit that receives image input from an external source via a communication line or the like. Specifically, when training the feature calculation model M1 and the classification model M2, the image input unit 11 receives input of the source image and correct information indicating whether or not an AU (Article Awareness) occurs. In addition, when performing classification, the image input unit 11 receives input of the image to be classified.

[0035] The face region extraction unit 12 is a processing unit that extracts face regions from an image received by the image input unit 11. The face region extraction unit 12 identifies face regions from the image received by the image input unit 11 using known face recognition processing and sets the identified face regions as face images 100. Then, during the training of the feature calculation model M1 and the discrimination model M2, the face region extraction unit 12 outputs face images 100 to the partially obscured image generation unit 13, the AU comparison image generation unit 14, and the image set generation unit 16. Furthermore, during discrimination, the face region extraction unit 12 outputs face images 100 to the discrimination unit 21.

[0036] The partial occlusion image generation unit 13 is a processing unit that generates images with occlusion (face images 100b, 100c) by partially concealing the face image 100 (without occlusion) output from the face region extraction unit 12 and the AU comparison image generation unit 14. Specifically, the partial occlusion image generation unit 13 generates an image that masks at least a portion of the operational area where AU occurs, as indicated as correct information, for the face image 100 without occlusion. Then, the partial occlusion image generation unit 13 outputs the generated image (image with occlusion) to the image set generation unit 16.

[0037] For example, if the correct information indicates that the corners of the lips should be raised (AU15), the partial censorship image generation unit 13 will generate an image that is masked to hide a part of the mouth area, which corresponds to AU15. The same applies to other action areas corresponding to other AUs. For example, if the correct information indicates that the inner part of the eyebrows should be raised (AU1), the partial censorship image generation unit 13 will generate an image that is masked to hide a part of the eyebrows, which corresponds to AU1.

[0038] Furthermore, masking is not limited to masking only a portion of the operating area; areas other than the operating area may also be masked. For example, the partial masking image generation unit 13 may mask a randomly selected portion of the entire area of ​​the face image 100.

[0039] The AU comparison image generation unit 14 is a processing unit that generates an image opposite to the presence or absence of AU (augmented genocide) indicated by the correct answer information, based on the face image 100 output from the face region extraction unit 12. Specifically, the AU comparison image generation unit 14 refers to an image database 15 that stores multiple face images of a person assigned the presence or absence of AU, and obtains an image opposite to the presence or absence of AU indicated by the correct answer information. The AU comparison image generation unit 14 outputs the obtained image to the partial occlusion image generation unit 13 and the image set generation unit 16.

[0040] Here, the image database 15 is a database that stores multiple face images. Each face image stored in the image database 15 is associated with information indicating the presence or absence of each AU (for example, an array indicating the presence or absence of each AU (AU1, AU2, ...)).

[0041] The AU comparison image generation unit 14 refers to this image database 15 and, for example, if the sequence (1,0,...) indicating the presence of AU1 is correct information, it obtains a face image corresponding to the absence of AU1 (0,*(arbitrary),...). In this way, the AU comparison image generation unit 14 obtains an image in which the presence or absence of AU is reversed compared to the input learning source face image 100.

[0042] In other words, the image input unit 11, face region extraction unit 12, partial occlusion image generation unit 13, and AU comparison image generation unit 14 are examples of acquisition units that acquire multiple images including a person's face.

[0043] The image set generation unit 16 is a processing unit that generates an image set by classifying the face images (face images 100a, 100b, 100c) output from the face region extraction unit 12, the partial occlusion image generation unit 13, and the AU comparison image generation unit 14 into one of two patterns, which is a combination of the presence or absence of AU and the presence or absence of occlusion in the image where AU occurs. In other words, the image set generation unit 16 is an example of a classification unit that classifies each of multiple images.

[0044] Specifically, the image set generation unit 16 calculates the first distance (d o ) and the second distance (dau The images are classified into sets (face images 100a, 100b, 100c) to obtain the desired result.

[0045] As an example, the image set generation unit 16 combines three images: a face image 100a output by the face region extraction unit 12 from an input image to which correct information indicating the presence of AU has been added; a face image 100b output by the partial concealment image generation unit 13 after masking of face image 100a; and a face image 100c generated by the AU comparison image generation unit 14 as an image in which the presence or absence of AU is reversed compared to face image 100a, and output after masking by the partial concealment image generation unit 13.

[0046] The image set generation unit 16 calculates the first distance (d o A set of images (face images 100a, 100b) to obtain the second distance (d au The images may be classified into two categories: a set of images (face images 100b, 100c) for obtaining the desired result.

[0047] The feature calculation unit 17 is a processing unit that calculates feature quantities 120 for each image in the image set generated by the image set generation unit 16. Specifically, the feature calculation unit 17 inputs each image in the image set into the feature calculation model M1 to obtain the output (feature quantities 120) from the feature calculation model M1.

[0048] The distance calculation unit 18 calculates the first distance (d) based on the feature quantities 120 for each image in the image set calculated by the feature quantity calculation unit 17. o ) and the second distance (d au This is a processing unit that calculates the first distance (d) based on the feature quantities obtained from the image set combining the face images 100a and 100b. Specifically, the distance calculation unit 18 calculates the first distance (d) based on the feature quantities obtained from the image set combining the face images 100a and 100b. o Similarly, the distance calculation unit 18 calculates the second distance (d) based on the features obtained from the image set combining the face images 100b and 100c. au Calculate ).

[0049] The distance learning execution unit 19 uses the first distance (d) calculated by the distance calculation unit 18. o) and the second distance (d au Based on ), the first distance (d o As the second distance (d) decreases, au This is a processing unit that trains the feature calculation model M1 so that the value of ) becomes large. Specifically, the distance learning execution unit 19 adjusts the parameters of the feature calculation model M1 using known methods such as backpropagation so that the loss in the loss function of equation (1) described above becomes small.

[0050] The distance learning execution unit 19 stores parameters and other information related to the trained feature calculation model M1 in a storage device (not shown). Therefore, during identification, the identification unit 21 can obtain the trained feature calculation model M1 from the distance learning execution unit 19 by referring to the information stored in the storage device.

[0051] The AU recognition learning execution unit 20 is a processing unit that trains the discrimination model M2 based on the correct information indicating whether or not an AU occurs and the feature quantities 120 calculated by the feature quantity calculation unit 17. Specifically, when the AU recognition learning execution unit 20 inputs the feature quantities 120 to the discrimination model M2, it updates the parameters of the discrimination model M2 so that the discrimination model M2 outputs a value corresponding to whether or not an AU occurs as indicated by the correct information.

[0052] The AU recognition learning execution unit 20 stores parameters and other information related to the trained identification model M2 in a storage device (not shown). Therefore, during identification, the identification unit 21 can obtain the trained identification model M2 from the AU recognition learning execution unit 20 by referring to the information stored in the storage device.

[0053] The identification unit 21 is a processing unit that, during identification, identifies whether or not AU occurs based on the face image 100 extracted by the face region extraction unit 12 from the image to be identified.

[0054] Specifically, the identification unit 21 constructs the feature calculation model M1 and the identification model M2 by referring to information stored in the storage device to obtain parameters related to the feature calculation model M1 and the identification model M2. Next, the identification unit 21 inputs the face image 100 extracted by the face region extraction unit 12 into the feature calculation model M1 to obtain feature quantities 120 related to the face image 100. Next, the identification unit 21 inputs the obtained feature quantities 120 into the identification model M2 to obtain information indicating whether or not an AU has occurred. The identification unit 21 outputs the identification result obtained in this way (presence or absence of AU occurrence) to, for example, a display device.

[0055] Figure 6 is a flowchart showing an example of the operation of the information processing device 1 according to the first embodiment. As shown in Figure 6, when processing starts, the image input unit 11 receives input of an image to be used as the learning source (including correct answer information) (S11).

[0056] Next, the face region extraction unit 12 extracts the region around the face by performing face recognition processing on the input image (S12). Then, the partial occlusion image generation unit 13 superimposes an occlusion mask image onto the image of the region around the face (face image 100) (S13). As a result, the partial occlusion image generation unit 13 generates an occluded image with occlusion for the face image 100 (without occlusion).

[0057] Next, the AU comparison image generation unit 14 selects and acquires an AU comparison image from the image database 15 in which the presence or absence of AU is reversed compared to the face surrounding region image (face image 100). Then, the partial occlusion image generation unit 13 superimposes an occlusion mask image onto the acquired AU comparison image (S14). As a result, the partial occlusion image generation unit 13 generates an image with occlusion relative to the AU comparison image (without occlusion).

[0058] Next, the image set generation unit 16 registers the occluded image, the image before occluding (image of the area around the face (face image 100)), and the AU comparison image (with occlusion) as a pair (S15). Next, the feature calculation unit 17 calculates feature quantities 120 (first feature quantity 120a, second feature quantity 120b, and third feature quantity 120c) from each of the three images in the image pair (S16).

[0059] Next, the distance calculation unit 18 calculates the distance (d) between the feature quantities of the hidden image and the face surrounding region image. o ) and the distance (d) between the features of the occluded image and the AU comparison image (with occlusion). au Calculate (S17).

[0060] Next, the distance learning execution unit 19 calculates the distance (d) obtained by the distance calculation unit 18. o d au ) and the first distance (d o As the second distance (d) decreases, au The feature acquisition model M1 is trained so that ) becomes larger (S18).

[0061] Next, the AU recognition learning execution unit 20 calculates the feature quantities 120 of the hidden image using the feature calculation model M1. Then, the AU recognition learning execution unit 20 performs AU recognition learning (S19) so that when the calculated feature quantities 120 are input to the discrimination model M2, the discrimination model M2 outputs a value corresponding to whether or not an AU occurs as indicated by the correct information, and then terminates the process.

[0062] (Second embodiment) Figure 7 is a block diagram showing an example of the functional configuration of an information processing device according to the second embodiment. As shown in Figure 7, the information processing device 1a according to the second embodiment has a face image input unit 11a that receives input of image data from which face images have been extracted in advance. In other words, the information processing device 1a according to the second embodiment differs from the information processing device 1 according to the first embodiment in that it does not have a face region extraction unit 12.

[0063] Figure 8 is a flowchart showing an example of the operation of the information processing device 1a according to the second embodiment. As shown in Figure 8, in the information processing device 1a, the face image input unit 11a accepts the input of a face image (S11a), so it is not necessary to extract the region around the face (S12).

[0064] (effect) As described above, the information processing devices 1 and 1a acquire multiple images, including human faces. The information processing devices 1 and 1a classify each of the multiple images into one of two patterns, which is a combination of the presence or absence of a specific action unit (AU) related to facial movement and the presence or absence of occlusion in the image where the action unit is present. The information processing devices 1 and 1a input each of the images classified into a pattern into the feature calculation model M1 and calculate the image features. The image input units 11 and 1a train the feature calculation model M1 so that the first distance between the features of the image where the action unit is present and the image where occlusion occurs in the image where the action unit is present becomes smaller, and the second distance between the features of the image where occlusion occurs in the image where the action unit is present and the image where occlusion occurs in the image where the action unit is not present becomes larger.

[0065] Thus, the information processing devices 1 and 1a can train the feature calculation model M1 to mitigate the effects of occlusion and output the magnitude of changes in the face image due to the occurrence of a specific unit of motion (AU) as a feature. Therefore, by inputting the image to be identified into the trained feature calculation model M1 and using the obtained features to identify AUs, it is possible to accurately identify the presence or absence of AUs even when the image to be identified has occlusion.

[0066] Furthermore, the information processing devices 1 and 1a refer to an image database 15 that stores multiple facial images of a person, each assigned whether or not a motion unit is present, based on the input image along with correct information indicating whether or not a motion unit is present. They then obtain an image where the presence or absence of a motion unit is the opposite of the presence or absence of a motion unit in the input image. In this way, the information processing devices 1 and 1a can obtain both images with and without motion units present from the input image.

[0067] Furthermore, the information processing devices 1 and 1a obtain an image with occlusion by obscuring a portion of the image based on the input image and the image database 15. In this way, the information processing devices 1 and 1a can obtain an image with occlusion from the input image, including images with and without the occurrence of motion units.

[0068] Furthermore, when acquiring an image with occlusion, the information processing devices 1 and 1a conceal at least a portion of the operating parts related to the operating unit. As a result, the information processing devices 1 and 1a can obtain an image with occlusion in which at least a portion of the operating parts related to the operating unit are concealed. Therefore, since the information processing devices 1 and 1a can train the feature calculation model M1 using an image with occlusion in which at least a portion of the operating parts related to the operating unit are concealed, they can efficiently learn about cases where the operating parts are concealed.

[0069] Furthermore, the information processing devices 1 and 1a determine the first distance d o , the second distance is d au , the margin parameter for the first distance is m o , the margin parameter for the second distance is m au The feature acquisition model M1 is trained based on the loss function Loss in equation (1). As a result, the information processing devices 1 and 1a can train the feature acquisition model M1 such that the first distance becomes smaller and the second distance becomes larger, according to the loss function Loss.

[0070] Furthermore, information processing devices 1 and 1a learn a discrimination model M2 to output whether or not an action unit has occurred, as indicated by the ground truth information, when an image with ground truth information indicating the presence or absence of an action unit is input to the feature calculation model M1 and the resulting features are input to the information processing device 1 and 1a. In this way, information processing devices 1 and 1a can learn a discrimination model M2 to identify whether or not an action unit has occurred based on the features obtained by inputting to the feature calculation model M1.

[0071] Furthermore, the information processing devices 1 and 1a acquire the trained feature calculation model M1 and input the image to be identified, which includes a person's face, into the acquired feature calculation model M1. Based on the obtained features, they identify whether or not a specific motion unit occurs in the person's face included in the image to be identified. As a result, even if there is occlusion in the image to be identified, the information processing devices 1 and 1a can accurately identify whether or not a specific motion unit occurs based on the features obtained from the feature calculation model M1.

[0072] (others) It should be noted that the components of each illustrated device do not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0073] Furthermore, the various processing functions of the information processing devices 1 and 1a (image input unit 11, face image input unit 11a, face region extraction unit 12, partial occlusion image generation unit 13, AU comparison image generation unit 14, image set generation unit 16, feature calculation unit 17, distance calculation unit 18, distance learning execution unit 19, AU recognition learning execution unit 20, and identification unit 21) may be executed in whole or in any part on a CPU (or a microcomputer such as an MPU or MCU (Micro Controller Unit)). It goes without saying that the various processing functions may also be executed in whole or in any part on a program analyzed and executed by a CPU (or a microcomputer such as an MPU or MCU), or on hardware using wired logic. Additionally, the various processing functions performed by the information processing device 1 may be executed collaboratively by multiple computers using cloud computing.

[0074] By the way, the various processing functions described in the above embodiment can be realized by executing a pre-prepared program on a computer. Therefore, below we will describe an example of a computer configuration (hardware) that executes a program having the same functions as in the above embodiment. Figure 9 is an explanatory diagram illustrating an example of a computer configuration.

[0075] As shown in Figure 9, the computer 200 includes a CPU 201 that performs various calculations, an input device 202 that accepts data input, a monitor 203, and a speaker 204. The computer 200 also includes a media reader 205 that reads programs and the like from a storage medium, an interface device 206 for connecting to various devices, and a communication device 207 for communicating with external devices via wired or wireless connections. The information processing device 1 includes a RAM 208 for temporarily storing various information and a hard disk drive 209. Each part (201 to 209) within the computer 200 is connected to a bus 210.

[0076] The hard disk drive 209 stores a program 211 for executing various processes in the above-mentioned various processing functions (for example, the image input unit 11, face image input unit 11a, face region extraction unit 12, partial occlusion image generation unit 13, AU comparison image generation unit 14, image set generation unit 16, feature quantity calculation unit 17, distance calculation unit 18, distance learning execution unit 19, AU recognition learning execution unit 20, and identification unit 21). The hard disk drive 209 also stores various data 212 that the program 211 refers to. The input device 202 receives operation information from the operator, for example. The monitor 203 displays various screens operated by the operator, for example. The interface device 206 is connected to an interface device, for example, a printer. The communication device 207 is connected to a communication network such as a LAN (Local Area Network) and exchanges various information with external devices via the communication network.

[0077] The CPU 201 reads the program 211 stored in the hard disk drive 209, loads it into the RAM 208, and executes it, thereby performing various processes related to the various processing functions described above. Note that the program 211 does not necessarily have to be stored in the hard disk drive 209. For example, the computer 200 may read and execute the program 211 stored in a storage medium that it can read. Examples of storage media that the computer 200 can read include portable recording media such as CD-ROMs, DVD discs, USB (Universal Serial Bus) memory, semiconductor memory such as flash memory, and hard disk drives. Alternatively, the program 211 may be stored in a device connected to a public network, the internet, a LAN, etc., and the computer 200 may read and execute the program 211 from there.

[0078] The following additional information is disclosed regarding the embodiments described above.

[0079] (Note 1) Obtain multiple images including a person's face, Based on the combination of whether or not an action unit related to the movement of a specific part of the face occurs and whether or not there is occlusion in the image in which the action unit occurs, the plurality of images are classified. Each of the above-mentioned classified images is input into a machine learning model to calculate the feature quantities of the images. The machine learning model is trained such that the first distance between the features of the image with the action unit occurring and the image with occlusion relative to the image with the action unit occurring becomes smaller, and the second distance between the features of the image with occlusion relative to the image with the action unit occurring and the image with occlusion relative to the image without the action unit occurring becomes larger. A learning program that instructs a computer to perform a task.

[0080] (Note 2) The acquisition process involves referring to a storage unit that stores multiple facial images of a person, each assigned whether or not an action unit has occurred, based on an image input along with correct information indicating whether or not an action unit has occurred, and acquiring an image in which the presence or absence of the action unit is the opposite of the presence or absence of the action unit in the input image. The learning program described in Appendix 1, characterized by the features described herein.

[0081] (Note 3) The acquisition process described above involves obtaining an image with occlusion by obscuring a portion of the input image and the acquired image. The learning program described in Appendix 2, characterized by the features described herein.

[0082] (Note 4) The acquisition process described above conceals at least a portion of the operating parts related to the action unit. The learning program described in Appendix 3, characterized by the features described herein.

[0083] (Note 5) The learning process described above uses the first distance d o , the second distance d au, the margin parameter for the first distance is m o , the margin parameter for the second distance is m au The machine learning model is trained based on the loss function Loss in equation (1) when given the following conditions. The learning program described in Appendix 1, characterized by the features described herein.

[0084] (Note 6) When an image with correct information indicating whether or not the action unit has occurred is input to the machine learning model and the resulting feature is input, the computer is further made to perform a process to train the discrimination model so that it outputs whether or not the action unit indicated by the correct information has occurred. The learning program described in Appendix 1, characterized by the features described herein.

[0085] (Note 7) Each of the multiple images classified based on the combination of whether or not an action unit related to the movement of a specific part of a person's face occurs and whether or not there is occlusion in the image in which the action unit occurs is input into a machine learning model to calculate the feature quantities of the images, and the machine learning model is obtained which has been trained so that the distance between the feature quantities of the image in which the action unit occurs and the image in which there is occlusion in the image in which the action unit occurs becomes small, and the distance between the feature quantities of the image in which there is occlusion in the image in which the action unit does not occur becomes large. Based on the features obtained by inputting an image containing a person's face into the acquired machine learning model, the presence or absence of a specific action unit occurring in the person's face included in the image is identified. An identification program that instructs a computer to perform a process.

[0086] (Note 8) Obtain multiple images including a person's face, Based on the combination of whether or not an action unit related to the movement of a specific part of the face occurs and whether or not there is occlusion in the image in which the action unit occurs, the plurality of images are classified. Each of the above-mentioned classified images is input into a machine learning model to calculate the feature quantities of the images. The machine learning model is trained such that the first distance between the features of the image with the action unit occurring and the image with occlusion relative to the image with the action unit occurring becomes smaller, and the second distance between the features of the image with occlusion relative to the image with the action unit occurring and the image with occlusion relative to the image without the action unit occurring becomes larger. A learning method in which a computer performs a process.

[0087] (Note 9) The acquisition process involves referring to a storage unit that stores multiple facial images of a person to which the presence or absence of the action unit has been assigned, based on an image input along with correct information indicating the presence or absence of the action unit, and acquiring an image in which the presence or absence of the action unit is the opposite of the presence or absence of the action unit in the input image. The learning method described in Appendix 8, characterized by the features described above.

[0088] (Note 10) The acquisition process described above involves obtaining an image with occlusion by obscuring a portion of the input image and the acquired image. The learning method described in Appendix 9, characterized by the features described herein.

[0089] (Note 11) The acquisition process described above conceals at least a portion of the operating parts related to the action unit. The learning method described in Appendix 10, characterized by the features described herein.

[0090] (Note 12) The learning process described above uses the first distance d o , the second distance d au , the margin parameter for the first distance is mo , the margin parameter for the second distance is m au The machine learning model is trained based on the loss function Loss in equation (1) when given the following conditions. The learning method described in Appendix 8, characterized by the features described above.

[0091] (Note 13) When an image with correct information indicating whether or not the action unit has occurred is input to the machine learning model and the resulting feature is input, the computer is further made to perform a process to train the discrimination model so that it outputs whether or not the action unit indicated by the correct information has occurred. The learning method described in Appendix 8, characterized by the features described above.

[0092] (Note 14) Each of the multiple images classified based on the combination of whether or not an action unit occurs for movement of a specific part of a person's face and whether or not there is occlusion in the image in which the action unit occurs is input into a machine learning model to calculate the feature quantities of the images, and the machine learning model is obtained which has been trained so that the distance between the feature quantities of the image in which the action unit occurs and the image in which there is occlusion in the image in which the action unit occurs becomes small, and the distance between the feature quantities of the image in which there is occlusion in the image in which the action unit does not occur becomes large. Based on the features obtained by inputting an image containing a person's face into the acquired machine learning model, the presence or absence of a specific action unit occurring in the person's face included in the image is identified. A method of identification by which a computer will perform a process. [Explanation of Symbols]

[0093] 1, 1a... Information Processing Device 11…Image input section 11a... Face image input section 12…Face region extraction section 13...Partially obscured image generation unit 14...AU comparison image generation unit 15…Image Database 16…Image set generation unit 17…Feature calculation unit 18... Distance calculation unit 19… Distance Learning Execution Unit 20...AU Recognition Learning Execution Unit 21...Identification section 100, 100a~100c, 101... Face images 110, 110a... Face 111... Hair 112…shielding object 120…Features 120a...First feature 120b...Second feature 120c...Third feature 200... Computer 201…CPU 202...Input device 203…Monitor 204...Speaker 205... Media reading device 206… Interface device 207...Communication equipment 208...RAM 209... Hard disk drive 210... Bus 211…Program 212... Various data M1...Feature calculation model M2…Discrimination Model

Claims

1. Obtain multiple images, including the faces of people, Based on the combination of whether or not an action unit related to the movement of a specific part of the face occurs and whether or not there is occlusion in the image in which the action unit occurs, the plurality of images are classified. Each of the above-mentioned classified images is input into a machine learning model to calculate the feature quantities of the images. The machine learning model is trained such that the first distance between the feature quantities of the image with the action unit occurring and the image with occlusion relative to the image with the action unit occurring becomes smaller, and the second distance between the feature quantities of the image with occlusion relative to the image with the action unit occurring and the image with occlusion relative to the image without the action unit occurring becomes larger. A learning program that instructs a computer to perform a task.

2. The acquisition process involves referring to a storage unit that stores multiple facial images of a person, each assigned whether or not an action unit has occurred, based on an input image along with correct information indicating whether or not the action unit has occurred, and acquiring an image in which the presence or absence of the action unit is the opposite of the presence or absence of the action unit in the input image. The learning program according to feature 1.

3. The aforementioned acquisition process involves obtaining an image with occlusion by obscuring a portion of the input image and the acquired image. The learning program according to feature 2.

4. The aforementioned acquisition process conceals at least a portion of the operating parts related to the action unit. The learning program according to feature 3.

5. The learning process described above uses the first distance d o , the second distance d au , the margin parameter for the first distance is m o , the margin parameter for the second distance is m au The machine learning model is trained based on the loss function Loss in the following equation (1). [Math 1] The learning program according to feature 1.

6. When an image with correct information indicating whether or not the aforementioned action unit has occurred is input to the machine learning model, and the features obtained from this input are then used to train the discrimination model to output whether or not the action unit indicated by the correct information has occurred. The learning program according to feature 1.

7. Multiple images classified based on the combination of whether or not an action unit related to the movement of a specific part of a person's face occurs, and whether or not there is occlusion in the image with the action unit occurring, are input into a machine learning model to calculate the features of the images, and the machine learning model is obtained which has been trained so that the distance between the features of the image with the action unit occurring and the image with occlusion in the image with the action unit occurring becomes small, and the distance between the features of the image with occlusion in the image with the action unit occurring and the image with occlusion in the image without the action unit occurring becomes large. Based on the features obtained by inputting an image containing a person's face into the acquired machine learning model, the presence or absence of a specific action unit occurring in the person's face included in the image is identified. An identification program that instructs a computer to perform a process.

8. Obtain multiple images, including the faces of people, Based on the combination of whether or not an action unit related to the movement of a specific part of the face occurs and whether or not there is occlusion in the image in which the action unit occurs, the plurality of images are classified. Each of the above-mentioned classified images is input into a machine learning model to calculate the feature quantities of the images. The machine learning model is trained such that the first distance between the feature quantities of the image with the action unit occurring and the image with occlusion relative to the image with the action unit occurring becomes smaller, and the second distance between the feature quantities of the image with occlusion relative to the image with the action unit occurring and the image with occlusion relative to the image without the action unit occurring becomes larger. A learning method in which a computer performs a process.

9. Multiple images classified based on the combination of whether or not an action unit related to the movement of a specific part of a person's face occurs, and whether or not there is occlusion in the image with the action unit occurring, are input into a machine learning model to calculate the features of the images, and the machine learning model is obtained which has been trained so that the distance between the features of the image with the action unit occurring and the image with occlusion in the image with the action unit occurring becomes small, and the distance between the features of the image with occlusion in the image with the action unit occurring and the image with occlusion in the image without the action unit occurring becomes large. Based on the features obtained by inputting an image containing a person's face into the acquired machine learning model, the presence or absence of a specific action unit occurring in the person's face included in the image is identified. A method of identification by which a computer will perform a process.

Citation Information

Patent Citations

  • Image processing method and image processing program and image processing system

    JP2021128476A