Training data generation program, training data generation method, and training data generation device

The training data generation device corrects label distortions in facial expression estimation by using IR cameras for precise marker tracking and image normalization, enhancing the accuracy of AU intensity estimation in machine learning models.

JP7746917B2Active Publication Date: 2025-10-01FUJITSU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022079723
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-10-01
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

The generation of training data for facial expression estimation using machine learning is distorted due to gaps in marker movement between processed face images, leading to inaccurate AU intensity estimation.

Method used

A training data generation device that corrects labels based on marker movement and image processing to ensure accurate correspondence between marker movement and labels, using IR cameras for precise marker tracking and a system to normalize and correct face images.

Benefits of technology

Prevents distortion in training data, ensuring accurate AU intensity estimation by correcting labels according to face size and photographing position, thereby improving the accuracy of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746917000001
    Figure 0007746917000001
  • Figure 0007746917000002
    Figure 0007746917000002
  • Figure 0007746917000003
    Figure 0007746917000003
Patent Text Reader

Abstract

To suppress generation of distorted training data of corresponding relationship between a marker movement on a face image and a label.SOLUTION: A training data generating program according to the present invention makes a computer executing processing of a step of acquiring a captured image including a face of a person added with a marker, a step of changing an image size of a face image of a person extracted from the acquired captured image, a step of specifying a position of a marker included in the acquired, captured image, a step of generating a label indicating generation strength of an action unit formed of units forming an expression of a human face and corresponding to the marker position, a step of correcting the generated label based on a position of capturing a person upon capturing the captured image and a face size of a person on the captured image, and a step of generating a training data for machine learning by giving the corrected label to the training face image with the marker deleted from the face image with the image size changed.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to training data generation techniques. [Background technology]

[0002] Facial expressions play an important role in nonverbal communication. Facial expression estimation technology is important for understanding and sensing people. A method called AU (Action Unit) is known as a tool for facial expression estimation. AU is a method for breaking down and quantifying facial expressions based on facial parts and facial muscles.

[0003] The AU estimation engine is based on machine learning using a large amount of training data, which consists of facial expression image data and the occurrence and intensity of each AU. The occurrence and intensity of the training data are then annotated by experts called coders.

[0004] As described above, if the generation of training data is left to annotation by a coder or the like, it is costly and time-consuming, making it difficult to generate a large amount of training data. In light of this, a generation device that generates training data for AU estimation has been proposed.

[0005] For example, the generating device identifies the position of a marker included in a captured image containing a face and determines the intensity of the AU based on the amount of movement from the marker position in an initial state, e.g., an expressionless state. Meanwhile, the generating device generates a face image by cutting out a face region from the captured image and normalizing the image size. The generating device then generates training data for machine learning by assigning labels, including the intensity of the AU, to the generated face image. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-8949 [Patent Document 2] International Publication No. 2022 / 024272 [Patent Document 3] U.S. Patent Application Publication No. 2021 / 0271862 [Patent Document 4] US Patent Application Publication No. 2019 / 0294868 Summary of the Invention [Problem to be solved by the invention]

[0007] However, in the above-described generation device, when the movement of the same marker is captured, processing such as cropping and normalization of the captured image causes a gap in the movement of the marker between processed face images, while each face image is assigned a label with the same AU intensity. In this way, when training data in which the correspondence between the movement of the marker on the face image and the label is distorted is used for machine learning, the estimated values ​​of AU intensity output by a machine learning model to which captured images capturing similar facial expression changes are input vary, resulting in a decrease in the accuracy of AU estimation.

[0008] In one aspect, the present invention aims to provide a training data generation program, a training data generation method, and a training data generation device that can prevent the generation of training data in which the correspondence between the movement of markers on a face image and the labels is distorted. [Means for solving the problem]

[0009] A training data generation program according to one aspect causes a computer to perform the following processes: acquire a captured image including a person's face, cut out an image of the person's face from the captured image and normalize the image size, identify the position of a marker included in the captured image, generate a label corresponding to the occurrence intensity of the action unit based on the amount of movement of the marker obtained from the reference position of the marker corresponding to the action unit and the identified position of the marker, correct the label based on the photographing position of the person when the captured image was taken or the face size of the person in the captured image, and generate training data for machine learning by assigning the corrected label to a training face image in which the marker has been deleted from the normalized face image. [Effects of the Invention]

[0010] According to one embodiment, it is possible to prevent the generation of training data in which the correspondence between the movements of markers on a face image and labels is distorted. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a schematic diagram showing an example of the operation of the system. [Figure 2] FIG. 2 is a diagram showing an example of camera placement. [Figure 3] FIG. 3 is a schematic diagram showing an example of processing a captured image. [Figure 4] FIG. 4 is a schematic diagram illustrating one aspect of the problem. [Figure 5] FIG. 5 is a block diagram illustrating an example of a functional configuration of a training data generation device. [Figure 6] FIG. 6 is a diagram illustrating an example of movement of a marker. [Figure 7] FIG. 7 is a diagram for explaining a method for determining the occurrence intensity. [Figure 8] FIG. 8 is a diagram illustrating an example of a method for determining the occurrence intensity. [Figure 9] FIG. 9 is a diagram illustrating an example of a method for creating a mask image. [Figure 10]FIG. 10 is a diagram illustrating an example of a method for creating a mask image. [Figure 11] FIG. 11 is a schematic diagram showing an example of imaging of a subject. [Figure 12] FIG. 12 is a schematic diagram showing an example of imaging of a subject. [Figure 13] FIG. 13 is a schematic diagram showing an example of imaging of a subject. [Figure 14] FIG. 14 is a schematic diagram showing an example of imaging of a subject. [Figure 15] FIG. 15 is a flowchart showing the procedure of the overall processing. [Figure 16] FIG. 16 is a flowchart showing the procedure of the determination process. [Figure 17] FIG. 17 is a flowchart showing the procedure of the image processing. [Figure 18] FIG. 18 is a flowchart showing the procedure of the correction process. [Figure 19] FIG. 19 is a schematic diagram showing an example of a camera unit. [Figure 20] FIG. 20 is a diagram showing an example of training data generation. [Figure 21] FIG. 21 is a diagram showing an example of training data generation. [Figure 22] FIG. 22 is a schematic diagram showing an example of imaging of a subject. [Figure 23] FIG. 23 is a diagram showing an example of a face image after correction. [Figure 24] FIG. 24 is a diagram showing an example of a face image after correction. [Figure 25] FIG. 25 is a flowchart showing the procedure of the correction process applied to a camera other than the reference camera. [Figure 26] FIG. 26 illustrates an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of a training data generation program, a training data generation method, and a training data generation device according to the present application will be described with reference to the accompanying drawings. Each embodiment merely illustrates one example or aspect, and does not limit the range of values, functions, or usage scenarios. Furthermore, each embodiment can be appropriately combined within the scope of not causing any contradiction in the processing content. [Example]

[0013] <System configuration> 1 is a schematic diagram illustrating an example of the operation of the system 1. As shown in FIG. 1, the system 1 may include an imaging device 31, a measurement device 32, a training data generation device 10, and a machine learning device 50.

[0014] The imaging device 31 may be realized, for example, by an RGB (Red, Green, Blue) camera or the like. The measuring device 32 may be realized, for example, by an IR (infrared) camera or the like. In this way, the imaging device 31, for example, has spectral sensitivity corresponding to visible light, while also having spectral sensitivity corresponding to infrared light. The imaging device 31 and measuring device 32 may be positioned so as to face the face of a person to whom a marker is attached. Hereinafter, a person to whom a marker is attached is referred to as a subject to be photographed, and such a subject to be photographed may be referred to as a "subject."

[0015] The subject's facial expression changes as the imaging device 31 captures the image and the measurement device 32 measures it. This allows the training data generation device 10 to capture the changes in facial expression over time as captured images 110. The imaging device 31 may also capture a video as the captured images 110. Such a video can also be considered as a plurality of still images arranged in time series. The subject may change their facial expression freely, or may change their facial expression according to a predetermined scenario.

[0016] The markers are realized by, for example, IR reflective (retroreflective) markers. Using the IR reflection from such markers, the measurement device 32 can perform motion capture.

[0017] FIG. 2 is a diagram showing an example of camera arrangement. As shown in FIG. 2, the measurement device 32 is realized by a marker tracking system using multiple IR cameras 32A to 32E. Such a marker tracking system can measure the position of an IR reflective marker by stereo photography. The relative positional relationship between these IR cameras 32A to 32E can be corrected in advance by camera calibration. Note that while FIG. 2 shows an example in which five camera units, IR cameras 32A to 32E, are used in the marker tracking system, any number of IR cameras may be used in the marker tracking system.

[0018] Furthermore, multiple markers are attached to the subject's face so as to cover the target AUs (e.g., AU1 to AU28). The positions of the markers change according to changes in the subject's facial expression. For example, marker 401 is placed near the base of the eyebrows. Furthermore, markers 402 and 403 are placed near the facial line. The markers may be placed on one or more AUs and on skin corresponding to the movement of facial muscles. Furthermore, the markers may be placed to avoid areas of skin where texture changes are significant due to wrinkles, etc. Note that AUs are units that make up a person's facial expression.

[0019] Furthermore, the subject wears an appliance 40 with reference point markers attached. The positions of the reference point markers attached to the appliance 40 are assumed to remain constant even when the subject's facial expression changes. Therefore, the training data generation device 10 can measure changes in the positions of the markers attached to the face based on changes in their relative positions from the reference point markers. By using three or more such reference markers, the training data generation device 10 can identify the positions of the markers in three-dimensional space.

[0020] The device 40 may be, for example, a headband, and may place reference markers outside the contours of the face. Alternatively, the device 40 may be a VR headset, a mask made of a hard material, or the like. In this case, the training data generation device 10 can use the rigid surface of the device 40 as the reference markers.

[0021] The marker position can be determined with high accuracy using the marker tracking system realized using these IR cameras 32A to 32E and the device 40. For example, the position of the marker in three-dimensional space can be measured with an error of 0.1 mm or less.

[0022] Such a measuring device 32 can obtain the position of the markers as well as the position of the subject's head in three-dimensional space as measurement results 120. Hereinafter, the coordinate position in three-dimensional space may be referred to as the "3D position."

[0023] The training data generation device 10 provides a training data generation function that generates training data in which labels including AU occurrence intensities and the like are assigned to training face images 113 generated from captured images 110 of a subject's face. As just one example, the training data generation device 10 acquires captured images 110 captured by an imaging device 31 and measurement results 120 measured by a measuring device 32. Then, the training data generation device 10 determines AU occurrence intensities 121 corresponding to the markers based on the movement amounts of the markers obtained as the measurement results 120.

[0024] The "occurrence intensity" referred to here may be, for example, data in which the intensity at which each AU occurs is expressed using a five-level rating from A to E, and annotated as "AU1:2, AU2:5, AU4:1, ...". Note that the occurrence intensity is not limited to being expressed using a five-level rating, and may be expressed using, for example, a two-level rating (presence or absence of occurrence). In this case, for example, if the rating of the five-level rating is 2 or higher, it may be expressed as "present", while if the rating is less than 2, it may be expressed as "absence".

[0025] In addition to determining the AU generation strength 121, the training data generation device 10 processes the captured image 110 captured by the imaging device 31, such as cutting out a face region, normalizing the image size, and removing markers from the image. In this way, the training data generation device 10 generates training face images 113 from the captured image 110.

[0026] Fig. 3 is a schematic diagram showing an example of processing a captured image. As shown in Fig. 3, face detection is performed on a captured image 110 (S1). As a result, a face area 110A of 726 pixels in height and 726 pixels in width is detected from the captured image 110 of 1920 pixels in height and 1080 pixels in width. A partial image corresponding to the detected face area 110A is cut out from the captured image 110 (S2). As a result, a cut-out face image 111 of 726 pixels in height and 726 pixels in width is obtained.

[0027] Generating the cropped face image 111 in this manner is effective for the following reasons. One aspect is that the markers are used solely to determine the occurrence intensity of AUs, which are labels assigned to training data, and are deleted from the captured image 110 so as not to affect the determination of the occurrence intensity of AUs by the machine learning model m. When deleting a marker, the position of the marker present on the image is searched for. However, narrowing the search area to the face area 110A can reduce the amount of calculations by several to several tens of times compared to when the entire captured image 110 is used as the search area. Another aspect is that when a dataset of training data TR is stored, it is not necessary to store unnecessary areas other than the face area 110A. For example, in the example of the training sample shown in FIG. 3, the image size can be reduced from the captured image 110 of 1920 vertical x 1080 horizontal pixels to the cropped face image 111 of 726 vertical x 726 horizontal pixels.

[0028] Thereafter, the cropped face image 111 is resized to an input size of width and height that is equal to or smaller than the size of the input layer of a machine learning model m, for example, a convolved neural network (CNN). For example, if the input size of the machine learning model m is 512 pixels high by 512 pixels wide, the cropped face image 111, which is 726 pixels high by 726 pixels wide, is normalized to an image size of 512 pixels high by 512 pixels wide (S3). This results in a normalized face image 112, which is 512 pixels high by 512 pixels wide. Furthermore, markers are deleted from the normalized face image 112 (S4). As a result of steps S1 to S4, a training face image 113, which is 512 pixels high by 512 pixels wide, is obtained.

[0029] Then, the training data generation device 10 generates a dataset including training data TR in which the training face images 113 are associated with the occurrence intensities 121 of AUs that are used as correct labels. Then, the training data generation device 10 outputs the dataset of the training data TR to the machine learning device 50.

[0030] The machine learning device 50 provides a machine learning function that executes machine learning using a dataset of training data TR output from the training data generation device 10. For example, the machine learning device 50 sets the training face images 113 as explanatory variables of the machine learning model m, sets the occurrence intensities 121 of AUs that are correct labels as objective variables of the machine learning model m, and trains the machine learning model m according to a machine learning algorithm such as deep learning. This generates a machine learning model M that receives face images obtained from captured images as input and outputs estimated values ​​of the occurrence intensities of AUs.

[0031] <One aspect of the issue> As explained in the background art above, when the captured image is processed, there is an aspect that training data is generated in which the movement of markers on the face image and the correspondence between labels are distorted.

[0032] Examples of such distortion of the correspondence relationship include cases where there are individual differences in the size of the subject's face, cases where the same subject is photographed at different shooting positions, etc. In these cases, even when the same marker movement amount is observed, cut-out face images 111 with different image sizes are cut out from the captured image 110.

[0033] FIG. 4 is a schematic diagram showing one aspect of the problem. In FIG. 4, cut-out images 111a and cut-out face images 111b cut out from two captured images in which the movement amount d of the same marker is captured are shown. It is assumed that the cut-out image 111a and the cut-out face image 111b are captured at the distance between the optical center of the imaging device 31 and the subject's face.

[0034] As shown in FIG. 4, the cut-out image 111a is a partial image obtained by cutting out a face region of 720 pixels in length and 720 pixels in width from the captured image of the subject a with a large face. On the other hand, the cut-out face image 111b is a partial image obtained by cutting out a face region of 360 pixels in length and 360 pixels in width from the captured image of the subject b with a small face.

[0035] These cut-out images 111a and cut-out face images 111b are normalized to an image size of 512 pixels in length and 512 pixels in width, which is the size of the input layer of the machine learning model m. As a result, in the normalized face image 112a, the marker movement amount is reduced from d1 to d11 (<d1). On the other hand, in the normalized face image 112b, the marker movement amount is enlarged from d1 to d12 (>d1). Thus, a gap occurs in the marker movement amount between the normalized face image 112a and the normalized face image 112b.

[0036] On the other hand, in both the subject a and the subject b, the same marker movement amount d1 is obtained as the measurement result 120 by the measuring device 32, so the same AU generation intensity 121 is assigned as a label to the normalized face image 112a and the normalized face image 112b.

[0037] As a result, in the training face image corresponding to the normalized face image 112a, the amount of movement of the marker on the training face image is reduced to d11, which is smaller than the actual measurement value d1 measured by the measurement device 32, while the correct label is assigned the AU generation intensity corresponding to the actual measurement value d1. In addition, in the training face image corresponding to the normalized face image 112b, the amount of movement of the marker on the training face image is enlarged to d12, which is larger than the actual measurement value d1 measured by the measurement device 32, while the correct label is assigned the AU generation intensity corresponding to the actual measurement value d1.

[0038] In this way, training data in which the movement of markers on the face image and the correspondence between labels are distorted can be generated from the normalized face image 112a and the normalized face image 112b. Note that, although an example in which there are individual differences in the size of the subject's face has been given here, a similar problem can arise when the same subject is photographed at shooting positions with different distances from the optical center of the image capture device 31.

[0039] <One aspect of the problem-solving approach> Therefore, the training data generation function of this embodiment corrects the label of the generation intensity of the AU corresponding to the amount of marker movement measured by the measurement device 32 based on the distance between the optical center of the imaging device 31 and the subject's head or the face size on the captured image.

[0040] This allows the labels to be corrected in accordance with the movement of the markers on the face image, which changes due to processing such as cutting out the face region and normalizing the image size.

[0041] Therefore, the training data generation function according to this embodiment can prevent the generation of training data in which the correspondence between the movements of markers on a face image and labels is distorted.

[0042] <Configuration of training data generation device 10> Fig. 5 is a block diagram showing an example of the functional configuration of training data generation device 10. Fig. 5 schematically shows blocks related to the machine learning function of training data generation device 10. As shown in Fig. 5, training data generation device 10 has a communication control unit 11, a storage unit 13, and a control unit 15. Note that Fig. 1 only shows a selection of functional units related to the above-mentioned training data generation function, and training data generation device 10 may be provided with functional units other than those shown.

[0043] The communication control unit 11 is a functional unit that controls communications with other devices, such as the imaging device 31, the measurement device 32, and the machine learning device 50. For example, the communication control unit 11 may be realized by a network interface card such as a LAN (Local Area Network) card. In one aspect, the communication control unit 11 receives captured images 110 captured by the imaging device 31 and measurement results 120 measured by the measurement device 32. In another aspect, the communication control unit 11 outputs to the machine learning device 50 a dataset of training data in which training face images 113 are associated with occurrence intensities 121 of AUs that are to be used as correct labels.

[0044] The storage unit 13 is a functional unit that stores various types of data. As just one example, the storage unit 13 is realized by internal, external, or auxiliary storage of the training data generation device 10. For example, the storage unit 13 can store various types of data such as AU information 13A that indicates the correspondence between markers and AUs. In addition to such AU information 13A, the storage unit 13 can store various types of data such as camera parameters of the image capture device 31 and calibration results.

[0045] The control unit 15 is a processing unit that performs overall control of the training data generation device 10. For example, the control unit 15 is realized by a hardware processor. Alternatively, the control unit 15 may be realized by hardwired logic. As shown in FIG. 5, the control unit 15 includes an identification unit 15A, a determination unit 15B, an image processing unit 15C, a correction coefficient calculation unit 15D, a correction unit 15E, and a generation unit 15F.

[0046] The identification unit 15A is a processing unit that identifies the position of a marker included in a captured image. The identification unit 15A identifies the position of each of multiple markers included in the captured image. Furthermore, when multiple images are acquired in chronological order, the identification unit 15A identifies the position of the marker for each image. In this way, the identification unit 15A can identify the coordinates on a plane or in space, for example, the 3D position, of each marker based on the positional relationship with a reference marker attached to the instrument 40. Note that the identification unit 15A may determine the position of the marker from a reference coordinate system or from a projected position on a reference plane.

[0047] The determination unit 15B is a processing unit that determines whether each of the multiple AUs has occurred based on the AU determination criteria and the positions of the multiple markers. The determination unit 15B determines the occurrence intensity of one or more AUs that have occurred among the multiple AUs. At this time, when the determination unit 15B determines that an AU corresponding to a marker has occurred among the multiple AUs based on the determination criteria and the positions of the marker, it can select the AU corresponding to the marker.

[0048] For example, the determination unit 15B determines the occurrence intensity of the first AU based on the movement amount of the first marker calculated based on the distance between the reference position of the first marker associated with the first AU included in the determination criterion and the position of the first marker identified by the identification unit 15A. Note that the first marker can be one or more markers corresponding to a specific AU.

[0049] The AU determination criteria indicate, for example, one or more markers among the multiple markers used to determine the occurrence intensity of each AU. The AU determination criteria may include reference positions of the multiple markers. The AU determination criteria may also include, for each of the multiple AUs, a relationship (conversion rule) between the movement amount of the marker used to determine the occurrence intensity and the occurrence intensity. Note that the reference positions of the markers may be determined according to the positions of the multiple markers in a captured image in which the subject is expressionless (no AUs are occurring).

[0050] Here, the movement of the marker will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating an example of the movement of the marker. Reference numerals 110-1 to 110-3 in Fig. 6 denote captured images captured by an RGB camera corresponding to an example of the imaging device 31. The captured images are assumed to be captured in the order of reference numerals 110-1, 110-2, and 110-3. For example, the captured image 110-1 is an image of the subject with a neutral expression. The training data generation device 10 can regard the position of the marker in the captured image 110-1 as a reference position where the amount of movement is zero.

[0051] 6, the subject is frowning. At this time, the position of marker 401 moves downward in accordance with the change in facial expression. At this time, the distance between the position of marker 401 and the reference marker attached to device 40 increases.

[0052] The variation values ​​of the distance of marker 401 from the reference marker in the X and Y directions are expressed as shown in Fig. 7. Fig. 7 is a diagram illustrating a method for determining the generation intensity. As shown in Fig. 7, determination unit 15B can convert the variation values ​​into generation intensity. Note that the generation intensity may be quantized into five levels in accordance with FACS (Facial Action Coding System), or may be defined as a continuous quantity based on the amount of variation.

[0053] There are various possible rules for the determination unit 15B to convert the amount of fluctuation into the occurrence intensity. The determination unit 15B may perform the conversion according to one predetermined rule, or may perform the conversion according to multiple rules and use the one with the largest occurrence intensity.

[0054] For example, the determination unit 15B may acquire in advance a maximum variation amount, which is the amount of variation when the subject changes their facial expression to the maximum, and convert the occurrence intensity based on the ratio of the variation amount to the maximum variation amount. Alternatively, the determination unit 15B may determine the maximum variation amount using data tagged by a coder using a conventional method. Alternatively, the determination unit 15B may linearly convert the variation amount to the occurrence intensity. Alternatively, the determination unit 15B may perform the conversion using an approximation formula created from prior measurements of multiple subjects.

[0055] Furthermore, for example, the determination unit 15B can determine the occurrence intensity based on a movement vector of the first marker calculated based on a position previously set as a determination criterion and the position of the first marker identified by the identification unit 15A. In this case, the determination unit 15B determines the occurrence intensity of the first AU based on the degree of match between the movement vector of the first marker and a predetermined vector previously defined for the first AU. The determination unit 15B may also correct the correspondence between the magnitude of the vector and the occurrence intensity using an existing AU estimation engine.

[0056] FIG. 8 is a diagram illustrating an example of a method for determining the generation intensity. For example, assume that the AU4 definition vector corresponding to AU4 is predetermined, such as (-2 mm, -6 mm). In this case, the determination unit 15B calculates the dot product of the movement vector of the marker 401 and the AU4 definition vector, and normalizes it by the magnitude of the AU4 definition vector. If the dot product matches the magnitude of the AU4 definition vector, the determination unit 15B determines the AU4 generation intensity to be 5 out of 5 levels. On the other hand, if the dot product is half the AU4 definition vector, for example, in the case of the linear conversion rule described above, the determination unit 15B determines the AU4 generation intensity to be 3 out of 5 levels.

[0057] 8, for example, it is assumed that the magnitude of the AU11 vector corresponding to AU11 is predetermined to be 3 mm. In this case, if the amount of change in the distance between marker 402 and marker 403 matches the magnitude of the AU11 vector, the determination unit 143 determines the generation intensity of AU11 to be 5 out of 5 levels. On the other hand, if the amount of change in the distance is half that of the AU4 vector, for example, in the case of the linear conversion rule described above, the determination unit 15B determines the generation intensity of AU11 to be 3 out of 5 levels. In this way, the determination unit 15B can determine the generation intensity based on the change in the distance between the position of the first marker and the position of the second marker identified by the identification unit 15A.

[0058] The image processing unit 15C is a processing unit that processes a captured image into a training image. As an example, the image processing unit 15C processes the captured image 110 captured by the imaging device 31, such as cutting out a face region, normalizing the image size, and removing markers from the image.

[0059] As described with reference to FIG. 3, the image processing unit 15C performs face detection on the captured image 110 (S1). As a result, a face region 110A of 726 pixels in height and 726 pixels in width is detected from the captured image 110 of 1920 pixels in height and 1080 pixels in width. The image processing unit 15C then cuts out a partial image from the captured image 110 corresponding to the face region 110A detected by face detection (S2). As a result, a cut-out face image 111 of 726 pixels in height and 726 pixels in width is obtained. The image processing unit 15C then normalizes the cut-out face image 111 of 726 pixels in height and 726 pixels in width to an image size of 512 pixels in height and 512 pixels in width, which corresponds to the input size of the machine learning model m (S3). As a result, a normalized face image 112 of 512 pixels in height and 512 pixels in width is obtained. Furthermore, the image processing unit 15C deletes markers from the normalized face image 112 (S4). As a result of steps S1 to S4, a training face image 113 of 512 pixels in height and 512 pixels in width is obtained from the captured image 110 of 1920 pixels in height and 1080 pixels in width.

[0060] The following provides additional information regarding the removal of such markers. As an example, a mask image can be used to remove the markers. FIG. 9 is a diagram illustrating an example of a method for creating a mask image. Reference numeral 112 in FIG. 9 is an example of a normalized face image. First, the image processing unit 15C extracts the color of a marker that has been intentionally added in advance and defines it as a representative color. Then, as indicated by reference numeral 112d in FIG. 9, the image processing unit 15C generates an area image of a color near the representative color. Furthermore, as indicated by reference numeral 112D in FIG. 9, the image processing unit 15C performs processing such as contraction and expansion on the area of ​​color near the representative color to generate a mask image for removing the markers. Furthermore, the accuracy of extracting the marker color may be improved by setting the marker color to a color that is unlikely to exist as a facial color.

[0061] FIG. 10 is a diagram illustrating an example of a marker deletion method. As shown in FIG. 10, first, the image processing unit 15C applies a mask image to a normalized face image 112 generated from a still image acquired from a video. The image processing unit 15C then inputs the image with the mask image applied to, for example, a neural network to obtain a training face image 113 as a processed image. It is assumed that the neural network has been trained using images of the subject with and without a mask. Acquiring still images from a video has the advantage of being able to obtain intermediate data on facial expression changes and a large amount of data in a short period of time. The image processing unit 15C may also use a generative multi-column convolutional neural network (GMCNN) or a generative adversarial network (GAN) as the neural network.

[0062] The method by which the image processing unit 15C deletes the markers is not limited to the above. For example, the image processing unit 15C may detect the position of the marker based on a predetermined marker shape and generate a mask image. Furthermore, the relative positions of the IR camera 32 and the RGB camera 31 may be calibrated in advance. In this case, the image processing unit 15C can detect the position of the marker from information on marker tracking by the IR camera 32.

[0063] Furthermore, the image processing unit 15C may employ different detection methods depending on the marker. For example, the marker on the nose moves little and its shape is easy to recognize, so the image processing unit 15C may detect its position by shape recognition. On the other hand, the marker next to the mouth moves a lot and its shape is difficult to recognize, so the image processing unit 15C may detect its position by extracting a representative color.

[0064] Returning to the explanation of FIG. 5, the correction coefficient calculation unit 15D is a processing unit that calculates correction coefficients used to correct labels assigned to training face images.

[0065] In one aspect, the correction coefficient calculation unit 15D calculates a "face size correction coefficient" by which the label is multiplied to correct the label according to the face size of the subject. FIGS. 11 and 12 are schematic diagrams showing an example of photographing a subject. In FIGS. 11 and 12, an RGB camera placed in front of the subject's face is shown as a reference camera 31A as an example of the imaging device 31, and both the reference subject e0 and the subject a are shown photographed at a reference position. Note that the "reference position" here refers to a position that is a distance L0 from the optical center of the reference camera 31A.

[0066] As shown in Figure 11, when a reference subject e0, whose actual face size is a reference size S0 in width and height, is photographed by the reference camera 31A, the face size in the captured image is assumed to be width P0 x height P0 pixels. The "face size in the captured image" here corresponds to the size of the face area obtained by performing face detection on the captured image. The face size P0 of the reference subject e0 in the captured image can be obtained as a set value by performing calibration in advance.

[0067] On the other hand, as shown in Fig. 12, when the face size in the captured image of a certain subject a photographed by the reference camera 31A is P1 pixels in width and P1 pixels in height, the ratio of the face size in the captured image of subject a to that of reference subject e0 can be calculated as the face size correction coefficient C1. That is, according to the example shown in Fig. 12, the correction coefficient calculation unit 15D can calculate the face size correction coefficient C1 as "P0 / P1".

[0068] By multiplying the label by this face size correction coefficient C1, the label can be corrected to match the image size to which the captured image of subject a is normalized, even if there is variation in the subject's face size due to individual differences. For example, consider a case in which the movement of the same marker corresponding to a common AU is captured between subject a and reference subject e0. In this case, if subject a's face size is larger than that of reference subject e0, i.e., if "P1 > P0," the movement of the marker on subject a's training face image will be smaller than the movement of the marker on reference subject e0's training face image due in part to the normalization process. Even in such a case, the label assigned to subject a's training face image can be corrected to a smaller size by multiplying the label assigned by the face size correction coefficient C1 = (P0 / P1) < 1.

[0069] As another aspect, the correction coefficient calculation unit 15D calculates a "position correction coefficient" that is multiplied by the label from the aspect of correcting the label according to the head position of the subject. FIG. 13 is a schematic diagram showing an example of photographing a subject. In FIG. 13, as an example of the imaging device 31, an RGB camera arranged in front of the face of the subject a is shown as the reference camera 31A, and the state where the subject a is photographed at different positions including the reference position is shown.

[0070] As shown in FIG. 13, when the subject a is photographed at the photographing position k1, the ratio of the photographing position k1 to the reference position can be calculated as the position correction coefficient C2. For example, since the measuring device 32 can measure not only the position of the marker but also the 3D position of the head of the subject a by motion capture, such a 3D position of the head can be referred to from the measurement result 120. Therefore, the distance L1 between the reference camera 31A and the subject a can be calculated based on the 3D position of the head of the subject a obtained as the measurement result 120. From the distance L1 corresponding to such a photographing position k1 and the distance L0 corresponding to the reference position, the position correction coefficient C2 can be calculated as "L1 / L0".

[0071] By multiplying such a position correction coefficient C2 by the label, even when there is a variation in the photographing position of the subject a, the label can be corrected according to the image size in which the captured image of the subject a is normalized. For example, consider a case where the movement amount of the same marker corresponding to a common AU between the reference position and the photographing position k1 is photographed. At this time, when the distance L1 corresponding to the photographing position k1 is smaller than the distance L0 corresponding to the reference position, that is, when L1 < L0, the movement amount of the marker on the training face image at the photographing position k1 is smaller than the movement amount of the marker on the training face image at the reference position due to the normalization process. Even in such a case, by multiplying the label given to the training face image at the photographing position k1 by the position correction coefficient C2 = (L1 / L0) < 1, the label can be corrected to be smaller.

[0072] <As a further aspect, the correction coefficient calculation unit 15D can also calculate an "integrated correction coefficient C3" that is an integration of the above-mentioned "face size correction coefficient C1" and the above-mentioned "position correction coefficient C2." Fig. 14 is a schematic diagram showing an example of photographing a subject. In Fig. 14, as an example of the imaging device 31, an RGB camera placed in front of the face of subject a is shown as the reference camera 31A, and the subject a is also shown being photographed at different positions including the reference position.

[0073] 14, when subject a is photographed at photographing position k2, the correction coefficient calculation unit 15D can calculate the distance L1 from the optical center of the reference camera 31A based on the 3D position of the head of subject a obtained as measurement result 120. In accordance with such distance L1 from the optical center of the reference camera 31A, the correction coefficient calculation unit 15D can calculate the position correction coefficient C2 as "L1 / L0".

[0074] Furthermore, the correction coefficient calculation unit 15D can obtain the face size P1 of subject a in the captured image obtained as a result of face detection in the captured image of subject a, i.e., width P1 × height P1 pixels. Based on such face size P1 of subject a in the captured image, the correction coefficient calculation unit 15D can calculate an estimated value P1' of subject a's face size at the reference position. For example, from the ratio between the reference position and the shooting position k2, P1' can be calculated as "P1 / (L1 / L0)" according to the derivation of the following equation (1). Furthermore, the correction coefficient calculation unit 15D can calculate the face size correction coefficient C1 as "P0 / P1'" from the ratio of the face sizes at the reference position between subject a and reference subject e0.

[0075] P1′=P1×(L0 / L1) =P1 / (L1 / L0) (1)

[0076] The correction coefficient calculation unit 15D calculates an integrated correction coefficient C3 by integrating the position correction coefficient C2 and the face size correction coefficient C1. That is, the integrated correction coefficient C3 can be calculated as "(P0 / P1) × (L1 / L0)" according to the derivation of the following equation (2).

[0077] C3=P0 / P1′ =P0÷{P1 / (L1 / L0)} =P0×(1 / P1)×(L1 / L0) =(P0 / P1)×(L1 / L0) (2)

[0078] Returning to the explanation of FIG. 5, the correction unit 15E is a processing unit that corrects the label. As merely an example, the correction unit 15E can correct the label by multiplying the occurrence intensity of the AU determined by the determination unit 15B, i.e., the label, by the integrated correction coefficient C3 calculated by the correction coefficient calculation unit 15D, as shown in the following equation (3). Note that although an example of multiplying the label by the integrated correction coefficient C3 has been given here, this is merely an example, and the label may also be multiplied by a face size correction coefficient C1 or a position correction coefficient C2, as shown in equations (4) and (5).

[0079] Example 1: Corrected label = Label x C3 =Label×(P0 / P1)×(L1 / L0)...(3) Example 2: Corrected label = Label x C1 =Label×(P0 / P1) (4) Example 3: Corrected label = Label × C2 =Label×(L1 / L0) (5)

[0080] The generation unit 15F is a processing unit that generates training data. As just one example, the generation unit 15F generates training data for machine learning by assigning labels corrected by the correction unit 15E to training face images generated by the image processing unit 15C. Such training data generation is performed for each captured image captured by the imaging device 31, thereby obtaining a data set of training data.

[0081] For example, when the machine learning device 50 executes the machine learning using a dataset of training data, the training data generated by the training data generation device 10 may be added to existing training data.

[0082] As just one example, the training data can be used for machine learning of an estimation model that uses an image as input and estimates occurring AUs. The estimation model may also be a model specialized for each AU. If the estimation model is specialized for a specific AU, the training data generation device 10 may change the generated training data to training data that uses only information related to the specific AU as a training label. In other words, for an image in which an AU other than the specific AU occurs, the training data generation device 10 can delete information related to the other AU and add information indicating that the specific AU does not occur as a training label.

[0083] According to this embodiment, it is possible to estimate the amount of training data required. Generally, implementing machine learning requires a huge amount of computational cost. The computational cost includes time and the amount of usage of a GPU, etc.

[0084] As the quality and quantity of a dataset improve, the accuracy of a model obtained by machine learning improves. Therefore, if the quality and quantity of a dataset required for a target accuracy can be roughly estimated in advance, the computational cost can be reduced. Here, for example, the quality of a dataset refers to the marker removal rate and removal accuracy. Also, for example, the quantity of a dataset refers to the number of datasets and the number of subjects.

[0085] Among the combinations of AUs, there are some combinations that are highly correlated with each other. Therefore, it is believed that an estimation performed for a certain AU can be applied to other AUs that are highly correlated with that AU. For example, it is known that AU18 and AU22 are highly correlated, and they may share corresponding markers. Therefore, if we can estimate the quality and quantity of a dataset that will achieve the target estimation accuracy for AU18, we can roughly estimate the quality and quantity of a dataset that will achieve the target estimation accuracy for AU22.

[0086] The machine learning model M generated by the machine learning device 50 can be provided to an estimation device (not shown) that estimates the occurrence intensity of AUs. The estimation device actually performs estimation using the machine learning model M generated by the machine learning device 50. The estimation device acquires an image that shows a person's face, in which the occurrence intensity of each AU is unknown, and inputs the acquired image into the machine learning model M, thereby outputting the occurrence intensity of AUs output by the machine learning model M as an AU estimation result to an arbitrary output destination. Such an output destination may be, by way of example only, a device, program, or service that estimates facial expressions using the occurrence intensity of AUs, or calculates a level of understanding or satisfaction.

[0087] <Processing flow> Next, we will explain the processing flow of the training data generation device 10. Here, we will explain (1) overall processing executed by the training data generation device 10, and then explain (2) determination processing, (3) image processing, and (4) correction processing.

[0088] (1) Overall processing Fig. 15 is a flowchart showing the procedure of the overall process. As shown in Fig. 15, an image captured by the imaging device 31 and a measurement result measured by the measuring device 32 are acquired (step S101).

[0089] Next, the specifying unit 15A and the determining unit 15B execute a "determination process" to determine the generation intensity of AUs based on the captured image and measurement results acquired in step S101 (step S102).

[0090] Then, the image processing unit 15C executes an "image processing" to process the captured image acquired in step S101 into a training image (step S103).

[0091] Thereafter, the correction coefficient calculation unit 15D and the correction unit 15E execute a "correction process" to correct the determination strength of the AU determined in step S102, that is, the label (step S104).

[0092] Then, the generation unit 15F generates training data by assigning the labels corrected in step S104 to the training face images generated in step S103 (step S105), and ends the process.

[0093] 15 can be performed at any timing after the extracted face image has been normalized. For example, the timing of step S104 is not limited to after the marker has been deleted, and it may be performed before the marker is deleted.

[0094] (2) Judgment process Fig. 16 is a flowchart showing the procedure of the determination process. As shown in Fig. 16, the identification unit 15A identifies the position of the marker included in the captured image acquired in step S101 based on the measurement result acquired in step S101 (step S301).

[0095] Then, the determination unit 15B determines an occurring AU occurring in the captured image based on the AU determination criteria included in the AU information 13A and the positions of the plurality of markers identified in step S301 (step S302).

[0096] Thereafter, the determination unit 15B executes a loop process 1 in which the processes of steps S304 and S305 are repeated the number of times corresponding to the number M of occurring AUs determined in step S302.

[0097] That is, the determination unit 15B calculates a movement vector of the marker based on the position of the marker assigned to the estimation of the m-th occurring AU among the positions of the markers identified in step S301 and the reference position (step S304).Then, the determination unit 15B determines the occurrence intensity of the m-th occurring AU, i.e., the label, based on the movement vector (step S305).

[0098] The occurrence intensity can be determined for each occurrence AU by repeating this loop process 1. Note that, although an example in which the processes of steps S304 and S305 are executed repeatedly has been given in the flowchart shown in Fig. 16, this is not limiting, and the processes may be executed in parallel for each occurrence AU.

[0099] (3) Image processing 17 is a flowchart showing the procedure of the image processing. As shown in FIG. 17, the image processing unit 15C executes face detection on the captured image acquired in step S101 (step S501). Then, the image processing unit 15C cuts out a partial image corresponding to the face area detected in step S501 from the captured image (step S502).

[0100] Thereafter, the image processing unit 15C normalizes the cut-out face image cut out in step S502 to an image size corresponding to the input size of the machine learning model m (step S503). Then, the image processing unit 15C deletes the markers from the normalized face image normalized in step S503 (step S504), and ends the process.

[0101] As a result of the processing in steps S501 to S504, training face images are obtained from the captured images.

[0102] (4) Correction processing Fig. 18 is a flowchart showing the procedure of the correction process. As shown in Fig. 18, correction coefficient calculation unit 15D calculates distance L1 from reference camera 31A to the head of the subject based on the 3D position of the head of the subject obtained as the measurement result acquired in step S101 (step S701).

[0103] Next, the correction coefficient calculation unit 15D calculates a position correction coefficient according to the distance L1 calculated in step S701 (step S702). Furthermore, the correction coefficient calculation unit 15D calculates an estimated value P1' of the subject's face size at the reference position based on the face size of the subject in the captured image obtained as a result of face detection in the captured image of the subject (step S703).

[0104] Thereafter, correction coefficient calculation unit 15D calculates an integrated correction coefficient from estimated value P1' of the subject's face size at the reference position and the ratio of face sizes at the reference position between the subject and the reference subject (step S704).

[0105] Then, the correction unit 15E corrects the label by multiplying the occurrence intensity of the AU determined in step S304, that is, the label, by the integrated correction coefficient calculated in step S704 (step S705), and ends the process.

[0106] <One aspect of the effect> As described above, the training data generation device 10 according to this embodiment corrects the labels of the generation intensity of AUs corresponding to the amount of marker movement measured by the measurement device 32 based on the distance between the optical center of the imaging device 31 and the subject's head or the face size on the captured image. This makes it possible to correct the labels in accordance with the movement of markers on the face image that varies due to processing such as cutting out the face region and normalizing the image size. Therefore, the training data generation device 10 according to this embodiment can prevent the generation of training data in which the correspondence between the movement of markers on the face image and the labels is distorted. [Example]

[0107] Although the embodiments of the disclosed device have been described above, the present invention may be embodied in various different forms other than the above-described embodiments. Therefore, other embodiments included in the present invention will be described below.

[0108] <Application example of the imaging device 31> In the above-described first embodiment, an RGB camera placed in front of the face of the subject is illustrated as the reference camera 31A as an example of the imaging device 31, but an RGB camera may be placed other than the reference camera 31A. For example, the imaging device 31 may be realized as a camera unit using a plurality of RGB cameras including the reference camera.

[0109] Fig. 19 is a schematic diagram showing an example of a camera unit. As shown in Fig. 19, the imaging device 31 may be realized as a camera unit including three RGB cameras: a base camera 31A, an upper camera 31B, and a lower camera 31C.

[0110] For example, the reference camera 31A is positioned in front of the subject's face, at eye level, at a horizontal camera angle. The upper camera 31B is positioned above the subject's face at a high angle. The lower camera 31C is positioned below the subject's face at a low angle.

[0111] Such a camera unit can capture the changes in facial expressions made by the subject from multiple camera angles, making it possible to generate multiple training face images of the subject with different facial orientations for the same AU.

[0112] 19 is merely an example, and the cameras do not necessarily have to be placed directly in front of the subject's face, but may be placed facing the left front, left side, right front, right side, etc. Also, the number of cameras shown in Fig. 19 is merely an example, and any number of cameras may be placed.

[0113] <One aspect of the issues when applying a camera unit> 20 and 21 are diagrams showing examples of training data generation. Fig. 20 and 21 show examples of training images 113A generated from images captured by the reference camera 31A and training images 113B generated from images captured by the upper camera 31B. It should be noted that the training images 113A and 113B shown in Fig. 20 and 21 are generated from images captured in synchronization with changes in the facial expression of the subject.

[0114] 20, a label A is assigned to a training image 113A, while a label B is assigned to a training image 113B. In this case, different labels are assigned to the same AU captured at different camera angles. As a result, if there is variation in the orientation in which the subject's face is captured, this may be one factor in generating a machine learning model M that outputs different labels even for the same AU.

[0115] 21, label A is assigned to training image 113A, and label A is also assigned to training image 113B. In this case, a single label can be assigned to the same AU captured at different camera angles. As a result, even if there is variation in the orientation in which the subject's face is captured, a machine learning model M that outputs a single label can be generated.

[0116] For this reason, when the same AU is photographed at different camera angles, it is preferable to assign a single label to the training face images generated from each of the captured images taken by the reference camera 31A, the upper camera 31B, and the lower camera 31C.

[0117] In this case, to maintain the correspondence between the movement of markers on the face image and the labels, label value (numerical) conversion is more advantageous than image conversion in terms of the amount of calculation, etc. However, if the labels are corrected for each captured image taken by each of multiple cameras, different labels are assigned to each camera, making it difficult to assign a single label.

[0118] <One aspect of the problem-solving approach> From this perspective, the training data generation device 10 can correct the image sizes of training face images in accordance with the labels instead of correcting the labels. In this case, the image sizes of all normalized face images corresponding to all cameras included in the camera unit can be corrected, or the image sizes of some normalized face images corresponding to some cameras, for example, a group of cameras other than the reference camera, can be corrected.

[0119] A method for calculating such an image size correction coefficient will now be described. As an example, let us generalize the number of cameras included in the camera unit to N, assign the camera number 0 to base camera 31A, assign the camera number 1 to upper camera 31B, and identify the cameras by adding an underscore followed by the camera number.

[0120] Hereinafter, as an example only, a method for calculating a correction coefficient for correcting the image size of a normalized face image corresponding to the upper camera 31B will be described, where the index for identifying the camera number is n=1, but this is not limited to the upper camera 31B. In other words, it goes without saying that the image size correction coefficient can be calculated in the same way when the index n=0 or n is 2 or greater.

[0121] FIG. 22 is a schematic diagram showing an example of photographing a subject. FIG. 22 shows an excerpt of upper camera 31B. As shown in FIG. 22, when subject a is photographed at photographing position k3, correction coefficient calculation unit 15D can calculate distance L1_1 from the optical center of upper camera 31B to the face of subject a based on the 3D position of the head of subject a obtained as measurement result 120. From the ratio of such distance L1_1 to distance L0_1 corresponding to the reference position, correction coefficient calculation unit 15D can calculate the position correction coefficient of the image size as "L1_1 / L0_1."

[0122] Furthermore, the correction coefficient calculation unit 15D can obtain the face size P1_1 of the subject a on the captured image obtained as a result of face detection on the captured image of the subject a, i.e., width P1_1 × height P1_1 pixels. Based on such face size P1 of the subject a on the captured image, the correction coefficient calculation unit 15D can calculate an estimated value P1_1' of the face size of the subject a at the reference position. For example, P1_1' can be calculated as "P1_1 / (L1_1 / L0_1)" from the ratio of the reference position and the shooting position k3.

[0123] Then, the correction coefficient calculation unit 15D calculates the integrated correction coefficient K for the image size as "(P1_1 / P0_1) x (L0_1 / L1_1)" from the estimated value P1_1' of the subject's face size at the reference position and the ratio of the face sizes at the reference position between subject a and reference subject e0.

[0124] Thereafter, correction unit 15E changes the image size of the normalized face image generated from the image captured by upper camera 31B in accordance with the integrated correction coefficient K=(P1_1 / P0_1)×(L0_1 / L1_1) of the image size. For example, the image size of the normalized face image is changed to an image size obtained by multiplying the number of pixels in the width and height of the normalized face image generated from the image captured by upper camera 31B by the integrated correction coefficient K=(P1_1 / P0_1)×(L0_1 / L1_1) of the image size. By changing the image size of the normalized face image in this way, a corrected face image is obtained.

[0125] 23 and 24 are diagrams showing an example of a corrected face image. Each of FIGS. 23 and 24 shows a cropped face image 111B generated from an image captured by upper camera 31B, and a corrected face image 114B obtained by normalizing the cropped face image 111B and changing the image size of the normalized face image based on an integrated correction coefficient K. Furthermore, FIG. 23 shows the corrected face image 114B when the integrated correction coefficient K of the image size is 1 or greater, while FIG. 24 shows the corrected face image 114B when the integrated correction coefficient K of the image size is less than 1. Furthermore, each of FIGS. 23 and 24 shows, with a dashed line, an image size corresponding to 512 pixels vertically by 512 pixels horizontally, which is an example of the input size of machine learning model m.

[0126] As shown in Fig. 23, when the integrated correction coefficient K for image size is 1 or greater, the image size of the corrected face image 114B is larger than the input size of the machine learning model m, which is 512 pixels high by 512 pixels wide. In this case, a training face image 115B is generated by re-cropping an area of ​​512 pixels high by 512 pixels wide, which corresponds to the input size of the machine learning model m, from the corrected face image 114B. Note that, for convenience of explanation, Fig. 23 shows an example in which a face area is detected with a margin included in the face area detected by the face detection engine set to 0%, but by setting the margin to α%, for example, several tens of percent, it is possible to prevent the loss of a face portion from the re-cropped training face image 115B.

[0127] 24, when the integrated correction coefficient K for image size is less than 1, the image size of corrected face image 114B is smaller than the input size of machine learning model m, which is 512 pixels in height and 512 pixels in width. In this case, training face image 115B is generated by adding a margin to corrected face image 114B to make up for the shortage of 512 pixels in height and 512 pixels in width, which corresponds to the input size of machine learning model m.

[0128] Since the correction by image resizing as described above requires more calculations than label correction, it is also possible to perform label correction without performing image correction on normalized images generated from images captured by some cameras, such as reference camera 31A.

[0129] In this case, the correction process shown in Figure 18 is applied to the normalized face image corresponding to the reference camera 31A, while the correction process corresponding to Figure 25 is applied to the normalized face image corresponding to a camera other than the reference camera 31A.

[0130] Fig. 25 is a flowchart showing the procedure of the correction process applied to cameras other than the reference camera 31A. As shown in Fig. 25, the correction coefficient calculation unit 15D executes loop process 1 in which the processes from step S901 to step S907 are repeated a number of times corresponding to the number N-1 of cameras other than the reference camera 31A.

[0131] That is, the correction coefficient calculation unit 15D calculates the distance L1_n from the camera 31n with camera number n to the head of the subject based on the 3D position of the head of the subject obtained as the measurement result acquired in step S101 (step S901).

[0132] Next, the correction coefficient calculation unit 15D calculates a position correction coefficient "L1_n / L0_n" for the image size of the camera number n based on the distance L1_n calculated in step S901 and the distance L0_n corresponding to the reference position (step S902).

[0133] Then, the correction coefficient calculation unit 15D calculates an estimated value of the subject's face size at the reference position, "P1_n'=P1_n / (L1_n / L0_n)", based on the face size of the subject in the captured image obtained as a result of face detection for the captured image of camera number n (step S903).

[0134] Next, the correction coefficient calculation unit 15D calculates the integrated correction coefficient "K = (P1_n / P0_n) x (L0_n / L1_n)" for the image size of camera number n from the estimated value P1_n' of the subject's face size at the reference position and the ratio of the face sizes at the reference position between subject a and reference subject e0 (step S904).

[0135] Then, the correction coefficient calculation unit 15D refers to the integrated correction coefficient of the label of the reference camera 31A, that is, the integrated correction coefficient C3 calculated in step S704 shown in FIG. 18 (step S905).

[0136] Then, the correction unit 15E changes the image size of the normalized face image based on the integrated correction coefficient K of the image size of the camera number n calculated in step S904 and the integrated correction coefficient of the label of the reference camera 31A referenced in step S905 (step S906). For example, the image size of the normalized face image is changed by a factor of (P1_n / P0_n)×(L0_n / L1_n)×(P0_0 / P1_0)×(L1_0 / L0_0). This results in a training face image of the camera number n.

[0137] The training face image of camera number n obtained in step S906 is assigned the following label when the process proceeds to step S105 shown in Fig. 15. That is, the training face image of camera number n is assigned the same label as the corrected label assigned to the training face image (without image resizing) generated from the image captured by the reference camera 31A, i.e., Label × (P0 / P1) × (L1 / L0). This makes it possible to assign a single label to the training face images of all cameras.

[0138] <Application example> In the above-described first embodiment, the training data generation device 10 and the machine learning device 50 are each configured as separate devices. However, the training data generation device 10 may also have the functions of the machine learning device 50.

[0139] In the above embodiment, the determination unit 15B determines the generation intensity of an AU based on the amount of movement of the marker. However, the fact that the marker has not moved can also be a criterion for determining the generation intensity by the determination unit 15B.

[0140] The marker may be surrounded by a color that makes it easier to detect. For example, a round green adhesive sticker with an IR marker in the center may be attached to the subject. In this case, the training data generation device 10 can detect the round green area from the captured image and delete the area along with the IR marker.

[0141] The information, including the processing procedures, control procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. Furthermore, the specific examples, distributions, numerical values, etc. described in the embodiments are merely examples and can be changed as desired.

[0142] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown. In other words, all or part of the devices can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.

[0143] <Hardware> Next, an example of the hardware configuration of the computer described in the first and second embodiments will be described. FIG. 26 is a diagram illustrating an example of the hardware configuration. As shown in FIG. 26, the training data generation device 10 includes a communication device 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components shown in FIG. 26 are connected to each other via a bus or the like.

[0144] The communication device 10a is a network interface card or the like, and communicates with other servers. The HDD 10b stores programs and DBs that operate the functions shown in FIG.

[0145] The processor 10d reads a program that executes the same processing as the processing unit shown in FIG. 5 from the HDD 100b or the like and loads it into the memory 100c, thereby operating a process that executes the functions described in FIG. 5 or the like. For example, this process executes the same functions as the processing units of the training data generation device 10. Specifically, the processor 10d reads a program that has the same functions as the identification unit 15A, the determination unit 15B, the image processing unit 15C, the correction coefficient calculation unit 15D, the correction unit 15E, the generation unit 15F, etc. from the HDD 10b or the like. Then, the processor 10d executes a process that executes the same processing as the identification unit 15A, the determination unit 15B, the image processing unit 15C, the correction coefficient calculation unit 15D, the correction unit 15E, the generation unit 15F, etc.

[0146] In this way, the training data generation device 10 operates as an information processing device that executes a training data generation method by reading and executing a program. The training data generation device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the training data generation device 10. For example, the present invention can also be applied to cases where another computer or server executes the program, or where these execute the program in cooperation with each other.

[0147] The above program can be distributed via a network such as the Internet. The above program can also be recorded on any recording medium and executed by a computer by reading it from the recording medium. For example, the recording medium can be a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), a digital versatile disk (DVD), or the like.

[0148] The following additional notes are provided regarding the embodiments including the above examples.

[0149] (Appendix 1) A captured image including a face of a person with a marker is obtained, changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation program that causes a computer to execute a process.

[0150] (Supplementary Note 2) The correction process includes a process of correcting the label based on a ratio of a photographing position of the person to a reference photographing position, or a ratio of a face size of the person to a reference face size. 2. The training data generation program according to claim 1,

[0151] (Supplementary Note 3) The acquiring process includes acquiring a first captured image and a second captured image in which the face of the person is captured at different camera positions or different camera angles, the correction process includes a process of correcting a label generated from a movement amount of the marker corresponding to the first captured image, and a process of correcting an image size of a face image normalized from an image size of a face image cut out from the second captured image, based on a photographing position of the person at the time of photographing the second captured image or a face size of the person in the second captured image; The process of generating the training data includes a process of generating first training data by assigning the label corrected in the correction process to a first training face image obtained by cutting out a face image of the person from the first captured image, normalizing the image size, and deleting the marker, and a process of generating second training data by assigning the same label as the label assigned to the first training data to a second training face image obtained by deleting the marker from the face image whose image size has been corrected in the correction process. 2. The training data generation program according to claim 1,

[0152] (Appendix 4) The correction process includes a process of, if the image size after correction is larger than the input size of the machine learning model, cutting out an area corresponding to the input size of the machine learning model from the corrected face image, and, if the image size after correction is smaller than the input size of the machine learning model, adding a margin to the corrected face image to make up for the lack of the input size of the machine learning model. 4. The training data generation program according to claim 3,

[0153] (Note 5) The first captured image corresponds to an image captured at an eye level camera position and a horizontal camera angle. The second captured image corresponds to an image captured at a camera position other than eye level or at a camera angle other than horizontal. 4. The training data generation program according to claim 3,

[0154] (Appendix 6) A captured image including the face of a person to which a marker is attached is obtained, changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation method characterized in that processing is executed by a computer.

[0155] (Supplementary Note 7) The correction process includes a process of correcting the label based on a ratio of a photographing position of the person to a reference photographing position, or a ratio of a face size of the person to a reference face size. 7. The training data generation method according to claim 6,

[0156] (Supplementary Note 8) The acquiring process includes acquiring a first captured image and a second captured image in which the face of the person is captured at different camera positions or different camera angles, the correction process includes a process of correcting a label generated from a movement amount of the marker corresponding to the first captured image, and a process of correcting an image size of a face image normalized from an image size of a face image cut out from the second captured image, based on a photographing position of the person at the time of photographing the second captured image or a face size of the person in the second captured image; The process of generating the training data includes a process of generating first training data by assigning the label corrected in the correction process to a first training face image obtained by cutting out a face image of the person from the first captured image, normalizing the image size, and deleting the marker, and a process of generating second training data by assigning the same label as the label assigned to the first training data to a second training face image obtained by deleting the marker from the face image whose image size has been corrected in the correction process. 7. The training data generation method according to claim 6,

[0157] (Appendix 9) The correction process includes a process of, if the image size after correction is larger than the input size of the machine learning model, cutting out an area corresponding to the input size of the machine learning model from the corrected face image, and, if the image size after correction is smaller than the input size of the machine learning model, adding a margin to the corrected face image to make up for the lack of the input size of the machine learning model. 9. The training data generation method according to claim 8,

[0158] (Appendix 10) The first captured image corresponds to an image captured at an eye level camera position and a horizontal camera angle. The second captured image corresponds to an image captured at a camera position other than eye level or at a camera angle other than horizontal. 9. The training data generation method according to claim 8,

[0159] (Appendix 11) A captured image including a face of a person to which a marker is attached is obtained, changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation device including a control unit that executes processing.

[0160] (Supplementary Note 12) The correction process includes a process of correcting the label based on a ratio of a photographing position of the person to a reference photographing position, or a ratio of a face size of the person to a reference face size. 12. The training data generation device according to claim 11,

[0161] (Supplementary Note 13) The acquiring process includes acquiring a first captured image and a second captured image in which the face of the person is captured at different camera positions or different camera angles, the correction process includes a process of correcting a label generated from a movement amount of the marker corresponding to the first captured image, and a process of correcting an image size of a face image normalized from an image size of a face image cut out from the second captured image, based on a photographing position of the person at the time of photographing the second captured image or a face size of the person in the second captured image; The process of generating the training data includes a process of generating first training data by assigning the label corrected in the correction process to a first training face image obtained by cutting out a face image of the person from the first captured image, normalizing the image size, and deleting the marker, and a process of generating second training data by assigning the same label as the label assigned to the first training data to a second training face image obtained by deleting the marker from the face image whose image size has been corrected in the correction process. 12. The training data generation device according to claim 11,

[0162] (Supplementary Note 14) The correction process includes a process of, if the image size after correction is larger than the input size of the machine learning model, cutting out an area corresponding to the input size of the machine learning model from the corrected face image, and, if the image size after correction is smaller than the input size of the machine learning model, adding a margin to the corrected face image to make up for the shortage of the input size of the machine learning model. 14. The training data generation device according to claim 13,

[0163] (Appendix 15) The first captured image corresponds to an image captured at an eye level camera position and a horizontal camera angle. The second captured image corresponds to an image captured at a camera position other than eye level or at a camera angle other than horizontal. 14. The training data generation device according to claim 13, [Explanation of symbols]

[0164] 1 System 10 Training data generator 11 Communication control section 13 Storage section 13A AU information 15 Control Unit 15A Specific part 15B Judgment section 15C Image processing department 15D Correction coefficient calculation section 15E Correction Unit 15F Generation section 31 Imaging device 32 Measuring equipment 50 Machine Learning Device

Claims

1. A captured image including the face of the person to which the marker is attached is acquired; changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation program that causes a computer to execute a process.

2. the correcting process includes a process of correcting the label based on a ratio of a photographing position of the person to a reference photographing position, or a ratio of a face size of the person to a reference face size.

2. The training data generation program according to claim 1, wherein:

3. the acquiring process includes acquiring a first captured image and a second captured image in which the face of the person is captured at different camera positions or different camera angles; the correction process includes a process of correcting a label generated from a movement amount of the marker corresponding to the first captured image, and a process of correcting an image size of a face image normalized from an image size of a face image cut out from the second captured image, based on a photographing position of the person at the time of photographing the second captured image or a face size of the person in the second captured image; The process of generating the training data includes a process of generating first training data by assigning the label corrected in the correction process to a first training face image obtained by cutting out a face image of the person from the first captured image, normalizing the image size, and deleting the marker, and a process of generating second training data by assigning the same label as the label assigned to the first training data to a second training face image obtained by deleting the marker from the face image whose image size has been corrected in the correction process.

2. The training data generation program according to claim 1, wherein:

4. The correction process includes a process of, when the image size after correction is larger than the input size of the machine learning model, cutting out an area corresponding to the input size of the machine learning model from the corrected face image, and, when the image size after correction is smaller than the input size of the machine learning model, adding a margin portion that is insufficient for the input size of the machine learning model to the corrected face image.

4. The training data generation program according to claim 3.

5. The first captured image corresponds to an image captured with a camera positioned at eye level and at a horizontal angle. the second captured image corresponds to an image captured at a camera position other than eye level or at a camera angle other than horizontal; 4. The training data generation program according to claim 3.

6. A captured image including the face of the person to which the marker is attached is acquired; changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation method characterized in that processing is executed by a computer.

7. A captured image including the face of the person to which the marker is attached is acquired; changing the image size of the face image of the person extracted from the acquired captured image; Identifying the position of the marker included in the acquired captured image; generating a label indicating the intensity of occurrence of an action unit that is composed of units constituting the facial expression of the person and corresponds to the position of the marker; correcting the generated label based on a photographing position of the person at the time of photographing the captured image or a face size of the person on the captured image; generating training data for machine learning by assigning the corrected labels to training face images obtained by deleting the markers from the face images whose image sizes have been changed; A training data generation device including a control unit that executes processing.

Citation Information

Patent Citations

  • Face expression amplification device, expression recognition device, face expression amplification method, expression recognition method and program

    JP2012008949A

  • Expression analysis device and expression analysis program

    JP2015035172A

  • Image normalization for facial analysis

    JP2021043960A

  • Learning data generating program and learning data generation method and estimation device

    JP2021111114A

  • System and method for recognition and annotation of facial expressions

    US20190294868A1