Face orientation estimation device and face orientation estimation method
The face direction estimation device addresses erroneous facial orientation issues by detecting facial feature points, estimating skeletal attributes, and adjusting three-dimensional models to match individual facial structures, achieving precise facial direction estimation.
Patent Information
- Application Number
- PCT/JP2024/015795
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional face direction estimation methods using three-dimensional models fail to account for individual skeletal attributes, leading to erroneous estimations of facial orientation, particularly in the pitch direction, due to varying positional relationships of facial features like eyes, nose, and mouth among individuals.
A face direction estimation device that detects facial feature points, estimates skeletal attributes, selects a three-dimensional model matching the individual's skeletal attributes, and adjusts the model's position and orientation to accurately estimate facial direction, using machine learning and image processing techniques.
Prevents erroneous facial direction estimation by accounting for individual skeletal attributes, ensuring highly accurate face direction estimation in three-dimensional space.
Smart Images

Figure JP2024015795_30102025_PF_FP_ABST
Abstract
Description
Facial direction estimation device and facial direction estimation method
[0001] The present disclosure relates to a face direction estimation device and a face direction estimation method.
[0002] A known technique uses a prepared three-dimensional model (hereinafter referred to as a "three-dimensional model") representing a person's face to estimate the facial orientation of a person (hereinafter referred to as an "estimation target") captured in an image in three-dimensional space. In this technique, the facial orientation of the estimation target is estimated by comparing, on the image, a plurality of feature points (hereinafter referred to as "facial feature points") representing facial features of the estimation target detected based on the image with a plurality of feature points (hereinafter referred to as "model facial feature points") in the three-dimensional model corresponding to the plurality of feature points. For example, Patent Literature 1 discloses a facial orientation estimation device that detects a plurality of facial feature points (e.g., the tip of the nose, the chin, the left corner of the left eye, the right corner of the right eye, the left corner of the mouth, and the right corner of the mouth) of a person's face in imaging data captured by an imaging device, detects the facial feature points from three-dimensional data of the face shape stored in a storage device, and performs calculations such as matrix transformation between the facial feature points detected from the imaging data and the facial feature points detected from the three-dimensional data to estimate the angle of the face in the imaging data.
[0003] International Publication No. 2020 / 255645
[0004] In a technology that uses a three-dimensional model to estimate the facial orientation of an object captured in an image in three-dimensional space, the facial feature points generally include feature points corresponding to, for example, the eyes, nose, and mouth, in order to prevent a decrease in the estimation accuracy of the facial orientation of the object. Herein, the positional relationship between the eyes, nose, and mouth of a person facing straight ahead when viewed from the side may differ from person to person. This is because the positional relationship between these facial features depends on each person's skeleton, including the shape of their skull or the way the skull is attached to their neck bones. Hereinafter, attributes such as the shape of their skeleton that define the positional relationship of these facial features are referred to as "skeletal attributes." The above-mentioned conventional technology does not take into account that these skeletal attributes may differ from person to person, resulting in a problem of erroneously estimating the facial orientation of the object, particularly the facial orientation in the pitch direction. The face direction estimation device disclosed in Patent Literature 1 compares facial feature points detected from imaging data and 3D data using a method such as projective transformation to detect changes, and then adjusts the position coordinates of the facial feature points in the 3D data based on the changes, thereby updating the 3D data of an average face shape to a face shape that approximates that of a person captured by the imaging device. However, this adjustment only adjusts the position coordinates of the facial feature points in the 3D data, and does not adjust the positional relationship of the eyes, nose, and mouth when the 3D data is viewed from the side. Therefore, even if the face direction estimation device disclosed in Patent Literature 1 updates the 3D data to a face shape that approximates that of a person captured by the imaging device, the positional relationship of the eyes, nose, and mouth when viewed from the side of the 3D data may not approximate the positional relationship of the eyes, nose, and mouth when viewed from the side of the person captured by the imaging device, which may result in an erroneous estimation of the face direction of the person, and the above problem remains unresolved.
[0005] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a face direction estimation device that, compared to conventional techniques, prevents erroneous estimation of face direction in three-dimensional space of an estimation target captured in a captured image using a three-dimensional model.
[0006] a feature point detection unit that detects a plurality of facial feature points, including the eyes, nose, and mouth, from the facial image acquired by the image acquisition unit; an attribute estimation unit that estimates skeletal attributes of the occupant based on the facial image acquired by the image acquisition unit; a 3D model selection unit that selects an estimation 3D model to be used for estimating the facial direction from among a plurality of 3D models representing human faces corresponding to different skeletal attributes, based on the skeletal attributes of the occupant estimated by the attribute estimation unit; a facial direction estimation unit that estimates the facial direction of the occupant based on the plurality of facial feature points detected by the feature point detection unit and the estimation 3D model selected by the 3D model selection unit; and a 3D model adjustment unit that adjusts the positions of model facial feature points corresponding to the facial feature points in the estimation 3D model, based on the facial direction estimation result regarding the facial direction of the occupant estimated by the facial direction estimation unit and information about the plurality of facial feature points detected by the feature point detection unit.
[0007] According to the present disclosure, in estimating the face direction in three-dimensional space of an estimation target captured in a captured image using a three-dimensional model, erroneous estimation of the face direction can be prevented compared to conventional techniques.
[0008] 1 is a diagram illustrating an example of the configuration of a face direction estimation device according to a first embodiment. 2A is a diagram illustrating an example of a captured image of a person from the side, the person having skeletal attributes, such as a high nose, eyes and mouth positioned approximately the same in the anterior-posterior direction, and a tucked-in mouth and chin; in other words, a mouth and chin that do not protrude; FIG. 2B is a diagram illustrating the skull of the person shown in FIG. 2A as viewed from the front and from the side; FIG. 2C is a diagram illustrating an example of a captured image of a person from the side, the person having skeletal attributes, such as a nose that is not as high as that of the person shown in FIG. 2A but is slightly upward-pointing, eyes and mouth positioned approximately the same in the anterior-posterior direction, a flat forehead, and a tucked-in mouth and chin; FIG. 2D is a diagram illustrating the skull of the person shown in FIG. 2C as viewed from the front and from the side; FIG. 2E is a diagram illustrating an example of a captured image of a person from the side, the person having skeletal attributes, such as a mouth and chin that protrude forward more than the nose; and FIG. 2F is a diagram illustrating the skull of the person shown in FIG. 2E as viewed from the front and from the side. 7A and 7B are diagrams illustrating an example of the result of estimating the skeletal attributes of an occupant by the attribute estimation unit in accordance with the attribute estimation conditions in embodiment 1. FIG. 7B is a diagram for explaining an example of a method for estimating the face direction of an occupant by the face direction estimation device according to embodiment 1. FIG. 7C is a flowchart for explaining the operation of the face direction estimation device according to embodiment 1. FIG. 7D is a flowchart for explaining in detail an example of processing performed by the attribute estimation unit in the skeletal attribute estimation processing in step ST4 of FIG. 5. FIG. 7A and FIG. 7B are diagrams illustrating an example of the hardware configuration of the face direction estimation device according to embodiment 1.
[0009] Embodiments of the present disclosure will be described in detail below with reference to the drawings. Embodiment 1. FIG. 1 is a diagram illustrating an example configuration of a face direction estimation device 1 according to Embodiment 1. The face direction estimation device 1 according to Embodiment 1 estimates the face direction in three-dimensional space of a person captured in an image (hereinafter referred to as the "estimation target"). In the following Embodiment 1, the estimation target is a vehicle occupant. Hereinafter, the vehicle occupant will also be simply referred to as the "occupant." The face direction estimation device 1 estimates the face direction of the occupant using a three-dimensional model representing the person's face, based on an image (hereinafter referred to as the "face image") of the occupant captured by an imaging device 2 mounted on the vehicle, for example. More specifically, the face direction estimation device 1 places the three-dimensional model in a virtual three-dimensional space represented by a camera coordinate system (hereinafter referred to as the "virtual three-dimensional space"). The face direction estimation device 1 then compares, on the face image, a plurality of feature points (hereinafter referred to as "facial feature points") that indicate the facial features of the occupant, detected based on the facial image, with a plurality of feature points (hereinafter referred to as "model facial feature points") in a three-dimensional model that correspond to the plurality of feature points, and estimates the facial direction of the occupant in the actual three-dimensional space by changing the position and orientation of the three-dimensional model in the virtual three-dimensional space. In the following description, "estimating the facial direction of the occupant" means "estimating the facial direction of the occupant in three-dimensional space." From the viewpoint of preventing a decrease in the estimation accuracy of the facial direction of the occupant, the plurality of facial feature points used to estimate the facial direction of the occupant generally include feature points corresponding to the eyes, nose, and mouth. This is because the feature points corresponding to the eyes, nose, and mouth are main features of a human face, are generally easy to recognize, and are estimated to be feature points that can be detected relatively stably. In the first embodiment, the face direction estimation device 1 detects a plurality of facial feature points corresponding to the eyes, nose, and mouth.
[0010] As described above, the relative positions of the eyes, nose, and mouth when a person facing straight ahead is viewed from the side may differ from person to person. This is because the relative positions of these facial features depend on each person's skeleton, including the shape of their skull or the way the skull is attached to the neck bones. Hereinafter, attributes such as the shape of the skeleton that define the relative positions of these facial features will be referred to as "skeletal attributes." Note that in the first embodiment, "directly ahead" is not limited to strictly being directly ahead, but also includes being approximately directly ahead.
[0011] FIG. 2 is a diagram illustrating an example of differences in the positional relationship of the eyes, nose, and mouth when viewed from the side, which differ from person to person. FIG. 2A is a diagram illustrating an example of a captured image of a person captured from the side, with skeletal attributes including a high nose, eyes, and mouth positioned approximately the same in the anterior-posterior direction, and a tucked-in mouth and chin, in other words, a mouth and chin that are not protruding. FIG. 2B is a diagram illustrating the skull of the person shown in FIG. 2A as viewed from the front and side. Skeletal attributes such as those shown in FIGS. 2A and 2B are said to be common among Caucasians, for example. FIG. 2C is a diagram illustrating an example of a captured image of a person captured from the side, with a nose that is not as high as that of the person shown in FIG. 2A but is slightly upward-facing, eyes, and mouth positioned approximately the same in the anterior-posterior direction, a flat forehead, and a tucked-in mouth and chin. FIG. 2D is a diagram illustrating the skull of the person shown in FIG. 2C as viewed from the front and side. Skeletal attributes such as those shown in FIGS. 2C and 2D are said to be common among Mongoloids, for example. Fig. 2E is a diagram showing an example of an image captured from the side of a person with a skeletal attribute in which the mouth and chin protrude further forward than the nose, and Fig. 2F is a diagram showing the skull of the person shown in Fig. 2E as viewed from the front and side. Skeletal attributes such as those shown in Fig. 2E and Fig. 2F are said to be common among Negroid people, for example. The person shown in Fig. 2A, the person shown in Fig. 2C, and the person shown in Fig. 2E are all looking straight ahead.
[0012] The coordinates of multiple facial feature points detected from a facial image cannot determine the positional relationship of multiple faces when viewed from the side. For example, the positional relationship of the feature points representing the eyes, nose, and mouth detected from a facial image cannot be determined when viewed from the side. Therefore, if a three-dimensional model corresponds to a skeletal attribute in which the mouth does not protrude further forward than the nose when viewed from the side, such as the skeletal attributes of a person shown in FIGS. 2A and 2C , in other words, a three-dimensional model having such skeletal attributes, while an occupant has a skeletal attribute in which the mouth protrudes further forward than the nose when viewed from the side, such as the skeletal attributes of a person shown in FIG. 2E , comparing the facial feature points with the model facial feature points on the facial image and changing the position and orientation of the three-dimensional model in a virtual three-dimensional space to estimate the occupant's facial orientation may result in an erroneous estimation of the occupant's facial orientation. More specifically, in the above example, the occupant's facial orientation may be estimated to have a higher pitch angle than the actual one, in other words, to be pointing upward. 2A and 2B, the above-described problem does not occur. The face direction estimation device 1 according to the first embodiment makes it possible to switch between three-dimensional models used for estimating the face direction of the occupant, taking into account skeletal attributes that are attributes of the occupant's skeleton that define the positional relationship between the eyes, nose, and mouth when the occupant's face is viewed from the side, thereby achieving highly accurate face direction estimation that addresses the above-described problem.
[0013] As shown in FIG. 1, the face direction estimation device 1 is connected to an imaging device 2 and includes an image acquisition unit 11, a feature point detection unit 12, an accessory estimation unit 13, an attribute estimation unit 14, a three-dimensional model selection unit 15, a face direction estimation unit 16, a three-dimensional model adjustment unit 17, and a memory unit 18.
[0014] The imaging device 2 is a camera or the like installed for the purpose of monitoring the interior of the vehicle, and is installed so as to be able to capture at least an image of the face of the occupant. The imaging device 2 may be, for example, a device shared with a so-called "driver monitoring system (DMS)." The imaging device 2 is installed, for example, near the center of the dashboard inside the vehicle. Note that in the first embodiment, the center of the dashboard is not limited to the exact center, but also includes an approximate center. The imaging device 2 outputs an image of the occupant's face to the face direction estimation device 1. Note that the imaging device 2 outputs the face image to the face direction estimation device 1 frame by frame.
[0015] The image acquisition unit 11 acquires a facial image from the imaging device 2. The image acquisition unit 11 acquires the facial image frame by frame. The image acquisition unit 11 outputs the acquired facial image to the feature point detection unit 12.
[0016] The feature point detection unit 12 detects facial feature points indicating facial features of the occupant from the facial image acquired by the image acquisition unit 11. In the first embodiment, the feature point detection unit 12 detects facial feature points indicating the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the corners of the mouth. Note that this is merely an example, and the feature point detection unit 12 may further detect facial feature points indicating facial features other than the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the corners of the mouth. For example, the feature point detection unit 12 may detect facial feature points of the nostrils. It is sufficient that the feature point detection unit 12 detects facial feature points indicating at least the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the corners of the mouth from the facial image. The facial feature points indicating the inner corners and outer corners of the eyes are feature points corresponding to the eyes, the facial feature point indicating the tip of the nose is a feature point corresponding to the nose, and the facial feature points indicating the corners of the mouth are feature points corresponding to the mouth. The facial feature points are represented by coordinates on the facial image.
[0017] The feature point detection unit 12 may detect facial feature points using an appropriate method. For example, the feature point detection unit 12 may detect facial feature points using a known feature point detection technique based on machine learning. Alternatively, for example, the feature point detection unit 12 may detect facial feature points using a known image recognition technique such as edge extraction. Alternatively, for example, the feature point detection unit 12 may detect facial feature points by template matching. The feature point detection unit 12 outputs information about the detected facial feature points (hereinafter referred to as "facial feature point information") to the accessory estimation unit 13 and the face direction estimation unit 16. The facial feature point information is, for example, information in which information that can identify facial feature points is associated with the coordinates of the facial feature points in a face image.
[0018] The accessory estimation unit 13 estimates whether the occupant is wearing a mask based on the facial image acquired by the image acquisition unit 11. For example, the accessory estimation unit 13 estimates whether the occupant is wearing a mask using a trained model in machine learning (hereinafter referred to as a "machine learning model") that receives a facial image as input and outputs information indicating whether the person captured in the facial image is wearing a mask. In the first embodiment, this machine learning model is referred to as a "mask determination machine learning model." The mask determination machine learning model is generated in advance by an administrator or the like and stored in an internal buffer or the like of the accessory estimation unit 13. The accessory estimation unit 13 inputs the facial image acquired by the image acquisition unit 11 into the mask determination machine learning model, and if information indicating that a mask is being worn is obtained, estimates that the occupant is wearing a mask. Note that the facial image acquired by the image acquisition unit 11 is included in the facial feature point information. Furthermore, for example, the accessory estimation unit 13 may estimate whether the occupant is wearing a mask using a known image recognition technique or template matching for the facial image acquired by the image acquisition unit 11. The accessory estimation unit 13 outputs information indicating whether the occupant is wearing a mask (hereinafter referred to as "mask wearing information") to the attribute estimation unit 14. At this time, the accessory estimation unit 13 outputs the facial feature point information output from the feature point detection unit 12 to the attribute estimation unit 14 together with the mask wearing information.
[0019] The attribute estimation unit 14 estimates the skeletal attribute of the occupant based on the face image acquired by the image acquisition unit 11. In the first embodiment, the attribute estimation unit 14 estimates whether the skeletal attribute of the occupant is a skeletal attribute in which the mouth does not protrude beyond the nose when the face is viewed from the side, as shown in Figures 2A to 2D above (hereinafter referred to as the "first skeletal attribute"), or a skeletal attribute in which the mouth protrudes beyond the nose when the face is viewed from the side, as shown in Figures 2E and 2F above (hereinafter referred to as the "second skeletal attribute").
[0020] For example, the attribute estimation unit 14 estimates the skeletal attributes of the occupant using a machine learning model. In the first embodiment, the machine learning model used by the attribute estimation unit 14 when estimating the skeletal attributes of the occupant is referred to as an "attribute estimation machine learning model." The attribute estimation machine learning model has been trained in advance to receive a facial image as input and output information indicating the skeletal attributes of the person captured in the facial image. Here, the information indicating the person's skeletal attributes is, for example, information indicating a first skeletal attribute or a second skeletal attribute. For example, an administrator or the like generates the attribute estimation machine learning model in advance and stores the generated attribute estimation machine learning model in a buffer or the like inside the attribute estimation unit 14. The attribute estimation unit 14 estimates the skeletal attributes of the occupant based on the facial image acquired by the image acquisition unit 11 and the attribute estimation machine learning model. More specifically, the attribute estimation unit 14 inputs the facial image acquired by the image acquisition unit 11 into a machine learning model for attribute estimation, and obtains information indicating skeletal attributes, thereby estimating the skeletal attributes of the occupant, i.e., whether the skeletal attributes of the occupant are the first skeletal attributes or the second skeletal attributes.
[0021] Furthermore, for example, the attribute estimation unit 14 may calculate feature amounts (hereinafter referred to as "facial feature amounts") based on a plurality of facial feature points detected by the feature point detection unit 12 based on a facial image acquired by the image acquisition unit 11, and estimate the skeletal attributes of the occupant based on the calculated facial feature amounts in accordance with set conditions (hereinafter referred to as "conditions for attribute estimation"). The facial feature amounts calculated by the attribute estimation unit 14 are also referred to as "feature amounts for attribute estimation" here, since they are used to estimate skeletal attributes. Note that the attribute estimation unit 14 can identify information on the plurality of facial feature points detected by the feature point detection unit 12 from the feature point information.
[0022] The attribute estimation conditions are conditions that define whether the skeletal attribute of an occupant is estimated to belong to the first skeletal attribute or the second skeletal attribute when the attribute estimation feature has a certain value. The attribute estimation conditions are set in advance by an administrator or the like. After setting the attribute estimation conditions, the administrator or the like stores information indicating the attribute estimation conditions (hereinafter referred to as "attribute estimation condition information") in a buffer or the like of the attribute estimation unit 14. The attribute estimation conditions include conditions (hereinafter referred to as "feature amount calculation conditions") that define which facial feature points are used to calculate which attribute estimation feature amounts. In other words, the attribute estimation condition information includes information indicating the feature amount calculation conditions (hereinafter referred to as "feature amount calculation condition information").
[0023] In the first embodiment, as an example, the feature calculation conditions are set as follows: "Two attribute estimation feature quantities are calculated using the area of a triangle on a face image connecting a facial feature point indicating the corners of the mouth and a feature point indicating the tip of the nose as a first feature quantity, and the ratio of the length of a perpendicular line dropped from the feature point indicating the tip of the nose to the line segment connecting the feature points indicating the corners of the mouth (the mouth corner line segment) to the line segment connecting the feature points indicating the corners of the mouth (the mouth corner line segment) as a second feature quantity." Also, as an example, the attribute estimation conditions include the following: "Regarding attribute estimation feature quantities (here, the first feature quantity and the second feature quantity) calculated from a plurality of facial feature points detected based on a set number of frames (hereinafter referred to as the "number of estimation frames") of face images, if the average or median of the first feature quantity is equal to or greater than a predetermined threshold (first threshold) and the average or median of the second feature quantity is equal to or less than a predetermined threshold (second threshold), the skeletal attribute of the occupant is estimated to be the second skeletal attribute." It is assumed that the following condition is set: "For feature quantities for attribute estimation calculated from a plurality of facial feature points detected based on face images for the number of estimation frames, if the average or median value of the first feature quantity is less than a first threshold value, or the average or median value of the second feature quantity is greater than a second threshold value, the skeletal attribute of the occupant is estimated to be the first skeletal attribute."
[0024] The attribute estimation unit 14 refers to attribute estimation condition information stored in a buffer or the like inside the attribute estimation unit 14, calculates feature quantities for attribute estimation in accordance with the attribute estimation conditions, and estimates the skeletal attributes of the occupant based on the calculated feature quantities for attribute estimation. More specifically, the attribute estimation unit 14 first calculates feature quantities for attribute estimation (first feature quantities and second feature quantities) from a plurality of facial feature points detected by the feature point detection unit 12 based on the facial image acquired by the image acquisition unit 11, in accordance with feature quantity calculation conditions included in the attribute estimation conditions. The attribute estimation unit 14 stores the calculated feature quantities for attribute estimation in a buffer or the like inside the attribute estimation unit 14. When the attribute estimation features for the number of estimation frames are stored in an internal buffer or the like, the attribute estimation unit 14 compares the attribute estimation features for the number of estimation frames with a first threshold and a second threshold, and if the average value or median value of the first features is equal to or greater than the first threshold and the average value or median value of the second features is equal to or less than a predetermined second threshold, the attribute estimation unit 14 estimates the skeletal attribute of the occupant to be the second skeletal attribute. If the average value or median value of the first features is less than the first threshold or the average value or median value of the second features is greater than the second threshold, the attribute estimation unit 14 estimates the skeletal attribute of the occupant to be the first skeletal attribute.
[0025] 3 is a diagram illustrating an image of the result of estimating the skeletal attribute of the occupant according to the attribute estimation conditions by the attribute estimation unit 14 in Embodiment 1. For example, if the average value or median value of the attribute estimation feature falls within the shaded range in FIG. 3, the skeletal attribute of the occupant is estimated to belong to the second skeletal attribute.
[0026] The attribute estimation unit 14 outputs information indicating the estimated skeletal attributes of the occupant (hereinafter referred to as “skeletal attribute information”) to the three-dimensional model selection unit 15 .
[0027] In the first embodiment, when the accessory estimation unit 13 estimates that the occupant is wearing a mask, the attribute estimation unit 14 does not need to perform the skeletal attribute estimation process for estimating the skeletal attributes of the occupant as described above. When the accessory estimation unit 13 estimates that the occupant is wearing a mask, in other words, when mask wearing information indicating that the occupant is wearing a mask is output from the accessory estimation unit 13, the attribute estimation unit 14 does not perform the skeletal attribute estimation process and outputs the mask wearing information output from the accessory estimation unit 13 to the three-dimensional model selection unit 15.
[0028] The three-dimensional model selection unit 15 selects a three-dimensional model (hereinafter referred to as an "estimation three-dimensional model") to be used for estimating the facial orientation from among a plurality of three-dimensional models representing human faces, based on the estimation result of whether or not the occupant is wearing a mask by the accessory estimation unit 13 or the skeletal attributes of the occupant estimated by the attribute estimation unit 14. The storage unit 18 pre-stores a plurality of three-dimensional models representing human faces corresponding to different skeletal attributes, and a three-dimensional model representing the face of a person wearing a mask (hereinafter referred to as a "mask three-dimensional model"). In the first embodiment, the storage unit 18 stores at least a three-dimensional model corresponding to a first skeletal attribute (hereinafter referred to as a "first skeletal attribute-corresponding three-dimensional model"), a three-dimensional model corresponding to a second skeletal attribute (hereinafter referred to as a "second skeletal attribute-corresponding three-dimensional model"), and a mask three-dimensional model. The storage unit 18 may be configured to store a three-dimensional model corresponding to a skeletal attribute that can be estimated by the attribute estimation unit 14, and a mask three-dimensional model. More specifically, the first skeletal attribute-corresponding three-dimensional model stored in the storage unit 18 is a three-dimensional model in which the mouth does not protrude beyond the nose when the first skeletal attribute-corresponding three-dimensional model is viewed from the side while facing directly ahead. Also, the second skeletal attribute-corresponding three-dimensional model stored in the storage unit 18 is a three-dimensional model in which the mouth protrudes beyond the nose when the second skeletal attribute-corresponding three-dimensional model is viewed from the side while facing directly ahead. For example, an administrator or the like generates the first skeletal attribute-corresponding three-dimensional model, the second skeletal attribute-corresponding three-dimensional model, and a mask three-dimensional model in advance and stores them in the storage unit 18.
[0029] The three-dimensional model selection unit 15 selects, as an estimation three-dimensional model, a three-dimensional model corresponding to the estimated skeletal attribute of the occupant or a three-dimensional mask model from the storage unit 18, based on the estimation result of whether or not the occupant is wearing a mask by the accessory estimation unit 13 or the skeletal attribute of the occupant estimated by the attribute estimation unit 14. The three-dimensional model selection unit 15 selects the three-dimensional mask model as the estimation three-dimensional model when the accessory estimation unit 13 estimates that the occupant is wearing a mask, in other words, when mask wearing / non-wearing information indicating that the occupant is wearing a mask is output from the accessory estimation unit 13 via the attribute estimation unit 14. The three-dimensional model selection unit 15 selects the three-dimensional mask model as the estimation three-dimensional model when the skeletal attribute of the occupant estimated by the attribute estimation unit 14 is the first skeletal attribute, in other words, when skeletal attribute information indicating that the estimated skeletal attribute of the occupant is the first skeletal attribute is output from the attribute estimation unit 14. Furthermore, when the skeletal attribute of the occupant estimated by the attribute estimation unit 14 is the second skeletal attribute, in other words, when skeletal attribute information indicating that the skeletal attribute of the estimated occupant is the second skeletal attribute is output from the attribute estimation unit 14, the three-dimensional model selection unit 15 selects a three-dimensional model corresponding to the second skeletal attribute as the three-dimensional model for estimation.
[0030] The three-dimensional model selection unit 15 outputs the selected three-dimensional model for estimation to the face direction estimation unit 16 .
[0031] The face direction estimation unit 16 estimates the face direction of the occupant based on the multiple facial feature points detected by the feature point detection unit 12 and the 3D model for estimation selected by the 3D model selection unit 15. More specifically, the face direction estimation unit 16 first detects multiple model face feature points that indicate facial features in the 3D model for estimation. Note that the facial features indicated by the facial feature points and the facial features indicated by the model face feature points are assumed to be the same type of features. On the 3D model, the model face feature points corresponding to the facial feature points can be identified. Here, the facial feature points detected by the feature point detection unit 12 are facial feature points that indicate the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the corners of the mouth, and therefore the face direction estimation unit 16 detects the model face feature points that indicate the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the corners of the mouth. Next, the face direction estimation unit 16 compares the plurality of facial feature points with the plurality of model facial feature points corresponding to each of the plurality of facial feature points on the face image, and changes the position and orientation of the estimation 3D model in the virtual 3D space so that the error between the positions of the plurality of facial feature points on the face image and the positions of the plurality of model facial feature points is within a predetermined threshold (hereinafter referred to as the "error determination threshold"). In Embodiment 1, the error between the positions of the plurality of facial feature points on the face image and the positions of the plurality of model facial feature points being within the error determination threshold means that the errors between the positions of the plurality of facial feature points on the face image and the plurality of model facial feature points are each at their minimum values. Note that in the face image, the positions of the facial feature points and the positions of the model facial feature points are represented by coordinates on the face image.
[0032] Specifically, the face direction estimation unit 16 calculates the coordinates of the model face feature points on the face image using, for example, the DLT (Direct Linear Transform) method. The DLT method is an algorithm that derives a perspective projection matrix that perspectively projects three-dimensional coordinates onto a two-dimensional image when the two-dimensional coordinates of feature points on a two-dimensional image and the three-dimensional coordinates of feature points corresponding to the feature points on the two-dimensional image are known. The face direction estimation unit 16 perspectively projects the coordinates of multiple model face feature points on the three-dimensional model for estimation onto the face image using, for example, the DLT method. Then, the face direction estimation unit 16 rotates or translates the estimation 3D model in the virtual 3D space so that the error between the coordinates of the model facial feature points (hereinafter referred to as "projected model facial feature points") perspectively projected onto the face image and the coordinates of the facial feature points corresponding to the projected model facial feature points on the face image is minimized, that is, so that the error between the positions of the multiple facial feature points on the face image and the positions of the multiple model facial feature points is within an error determination threshold. The face direction estimation unit 16 estimates a rotation matrix or translation matrix that derives the position and orientation of the estimation 3D model in the virtual 3D space that minimizes the error between the coordinates of the projected model facial feature points perspectively projected onto the face image and the coordinates of the facial feature points corresponding to the projected model facial feature points. The face direction estimation unit 16 estimates the rotation matrix or translation matrix, for example, by the least squares method. The face direction estimation unit 16 uses a rotation matrix or a translation matrix to change the position and orientation of the 3D estimation model that is placed at an initial position and in an initial orientation in a virtual 3D space.
[0033] FIG. 4 is a diagram illustrating an example of a method for estimating the face direction of an occupant by the face direction estimation device 1 according to the first embodiment. The concept of estimating the face direction of an occupant by the face direction estimation unit 16 as described above will be described using FIG. 4 . It should be noted that in FIG. 4 , the occupant whose face direction is to be estimated is assumed to be not wearing a mask. In FIG. 4 , I indicates a face image. The diagram on the left side of FIG. 4 illustrates a plurality of facial feature points and model face feature points corresponding to the plurality of facial feature points, with the occupant's face on the face image indicated by Dr1 and the three-dimensional estimation model indicated by E. In the diagram, the facial feature points and the corresponding model face feature points are connected by dotted lines. In FIG. 4 , the facial feature points are the inner corners and outer corners of both eyes, the tip of the nose, and the corners of the mouth. The face direction estimation unit 16 rotates or translates the estimation 3D model in virtual 3D space so that the error between the coordinates of multiple facial feature points on the face image and the coordinates of multiple projection model facial feature points corresponding to the multiple facial feature points is minimized. In Figure 4, the right side of the figure is a diagram illustrating a face image in which the error between the coordinates of multiple facial feature points and the coordinates of multiple projection model facial feature points is minimized. Note that the estimation 3D model is not shown in the right side of Figure 4. The face direction estimation unit 16 estimates the facial direction of the occupant from the posture of the estimation 3D model after rotating or translating it in virtual 3D space.
[0034] As described above, when the accessory estimation unit 13 estimates that the occupant is not wearing a mask, the attribute estimation unit 14 estimates the skeletal attribute of the occupant, and the 3D model selection unit 15 selects a 3D model corresponding to the skeletal attribute of the occupant estimated by the attribute estimation unit 14 as the 3D model for estimation. For example, if the skeletal attribute of the occupant is estimated to be a first skeletal attribute, the 3D model corresponding to the first skeletal attribute is selected as the 3D model for estimation. If the skeletal attribute of the occupant is estimated to be a second skeletal attribute, the 3D model corresponding to the second skeletal attribute is selected as the 3D model for estimation. The face direction estimation unit 16 estimates the face direction of the occupant using an estimation 3D model in which the positional relationship of the eyes, nose, and mouth when viewed from the side matches the positional relationship of the eyes, nose, and mouth when viewed from the side. This allows the face direction estimation device 1 to prevent erroneous estimation of the face direction of the occupant, for example, estimating the pitch angle of the face direction of the occupant to be higher than the actual pitch angle.
[0035] If the 3D model for estimation selected by the 3D model selection unit 15 is a 3D model for a mask, the face direction estimation unit 16 estimates the face direction of the occupant based on the multiple facial feature points detected by the feature point detection unit 12 and the 3D model for a mask selected by the 3D model selection unit 15 as the 3D model for estimation. In this case, the face direction estimation unit 16, for example, does not use facial feature points indicating the corners of the mouth that are estimated to be hidden by the mask in estimating the face direction, and estimates the face direction of the occupant using facial feature points indicating the inner corners of the eyes, facial feature points indicating the outer corners of the eyes, and facial feature points indicating the tip of the nose estimated from the hidden state. For example, if facial feature points indicating the inner corners of the eyebrows or the outer corners of the eyebrows are detected, the face direction estimation unit 16 may use these facial feature points estimated not to be hidden by the mask in estimating the face direction. Alternatively, the mask itself may be recognized and feature points indicating the edge portions of the mask may be used in estimating the face direction.
[0036] The face direction estimation unit 16 outputs information regarding the estimated face direction of the occupant (hereinafter referred to as the "face direction estimation result") to the 3D model adjustment unit 17. The face direction estimation result includes information indicating the estimated face direction of the occupant, information regarding projected model facial feature points, which are points on the face image corresponding to multiple model facial feature points in the 3D model for estimation after the position and posture are changed by the face direction estimation unit 16, and the 3D model for estimation. Note that the information regarding the projected model facial feature points is specifically the coordinates of the projected model facial feature points on the face image. For example, the face direction estimation result may include information (e.g., an ID) that can identify the 3D model for estimation instead of the 3D model for estimation.
[0037] Furthermore, the face direction estimation unit 16 outputs information indicating the estimated face direction of the occupant to an external device (not shown) connected to the face direction estimation device 1. The external device performs various controls based on the information indicating the face direction of the occupant output from the face direction estimation unit 16. For example, the external device is an inattentive driving judgment device that determines whether the occupant is looking aside based on the information indicating the face direction of the occupant output from the face direction estimation unit 16. For example, the inattentive driving judgment device determines that the occupant is looking aside if the face direction of the occupant is equal to or greater than a predetermined threshold (hereinafter referred to as the "inattentive driving judgment threshold"). If the inattentive driving judgment device determines that the occupant is looking aside, it outputs a sound or the like to warn the occupant about inattentive driving from, for example, a speaker (not shown) mounted in the vehicle.
[0038] The three-dimensional model adjustment unit 17 adjusts the positions of model face feature points corresponding to the face feature points in the estimation three-dimensional model, based on the face direction estimation result output from the face direction estimation unit 16 and face feature point information regarding the plurality of face feature points detected by the feature point detection unit 12. The three-dimensional model adjustment unit 17 may acquire the face feature point information from the feature point detection unit 12, for example, via the accessory estimation unit 13, the attribute estimation unit 14, the three-dimensional model selection unit 15, and the face direction estimation unit 16. The three-dimensional model adjustment unit 17 may also acquire the face feature point information directly from the feature point detection unit 12. Note that the arrow from the feature point detection unit 12 to the three-dimensional model adjustment unit 17 is omitted in FIG. 1 .
[0039] More specifically, the three-dimensional model adjustment unit 17 adjusts the positions of the model facial feature points in the estimation three-dimensional model based on the face direction estimation result and the facial feature point information and on errors between the plurality of facial feature points detected by the feature point detection unit 12 and the plurality of projected model facial feature points corresponding to each of the plurality of facial feature points. For example, the three-dimensional model adjustment unit 17 compares the position coordinates of the plurality of facial feature points with the position coordinates of the projected model facial feature points corresponding to the plurality of facial feature points on the face image, and if the difference between the position coordinates of the projected model facial feature points and the position coordinates of the facial feature points is equal to or greater than a predetermined threshold (hereinafter referred to as the "difference determination threshold"), adjusts the position of the model facial feature point corresponding to the projected model facial feature point so that it becomes a position corresponding to the difference.
[0040] After adjusting the estimation 3D model, the 3D model adjustment unit 17 updates the corresponding 3D model stored in the storage unit 18 with the adjusted estimation 3D model. For example, if the estimation 3D model is a first skeletal attribute-corresponding 3D model, the 3D model adjustment unit 17 adjusts the positions of the model facial feature points in the first skeletal attribute-corresponding 3D model and updates the first skeletal attribute-corresponding 3D model stored in the storage unit 18 with the adjusted first skeletal attribute-corresponding 3D model. As a result, the 3D model adjustment unit 17 can create a 3D model stored in the storage unit 18 in which the positions of the model facial feature points in the 3D model are close to the positions of the occupant's facial feature points, in other words, the facial shape of the 3D model is close to the occupant's facial shape. As a result, the face direction estimation device 1 can improve the accuracy of face direction estimation by the face direction estimation unit 16.
[0041] The storage unit 18 stores, for example, a three-dimensional model. Note that, in Fig. 1, the storage unit 18 is shown to be provided in the face direction estimation device 1, but this is merely an example. The storage unit 18 may be provided in a location outside the face direction estimation device 1 that can be referenced by the face direction estimation device 1.
[0042] The operation of face direction estimation device 1 according to embodiment 1 will now be described. Fig. 5 is a flowchart for explaining the operation of face direction estimation device 1 according to embodiment 1. For example, when an occupant gets into a vehicle and the vehicle engine is turned on, face direction estimation device 1 starts the operation shown in the flowchart of Fig. 5 and repeatedly performs the operation shown in the flowchart of Fig. 5 until the vehicle engine is turned off.
[0043] The image acquisition unit 11 acquires a face image from the imaging device 2 (step ST1). The image acquisition unit 11 outputs the acquired face image to the feature point detection unit 12.
[0044] The feature point detection unit 12 detects facial feature points indicating facial features of the occupant from the facial image acquired by the image acquisition unit 11 in step ST1 (step ST2). The feature point detection unit 12 outputs facial feature point information to the accessory estimation unit 13 and the face direction estimation unit 16.
[0045] The accessory estimation unit 13 estimates whether the occupant is wearing a mask based on the facial image acquired by the image acquisition unit 11 in step ST1 (step ST3). The accessory estimation unit 13 outputs mask wearing information to the attribute estimation unit 14. At this time, the accessory estimation unit 13 outputs the facial feature point information output from the feature point detection unit 12 to the attribute estimation unit 14 together with the mask wearing information.
[0046] The attribute estimation unit 14 performs a skeleton attribute estimation process (step ST4) to estimate the skeleton attributes of the occupant based on the face image acquired by the image acquisition unit 11 in step ST1. The attribute estimation unit 14 outputs the skeleton attribute information to the three-dimensional model selection unit 15.
[0047] The three-dimensional model selection unit 15 selects a three-dimensional model for estimation from among the plurality of three-dimensional models stored in the storage unit 18 based on the result of estimation by the accessory estimation unit 13 in step ST3 as to whether the occupant is wearing a mask or not, or the skeletal attributes of the occupant estimated by the attribute estimation unit 14 in step ST4 (step ST5). The three-dimensional model selection unit 15 outputs the selected three-dimensional model for estimation to the face direction estimation unit 16.
[0048] The face direction estimation unit 16 estimates the face direction of the occupant based on the plurality of facial feature points detected by the feature point detection unit 12 in step ST2 and the 3D model for estimation selected by the 3D model selection unit 15 in step ST5 (step ST6). The face direction estimation unit 16 outputs the face direction estimation result to the 3D model adjustment unit 17. The face direction estimation unit 16 also outputs information indicating the estimated face direction of the occupant to an external device connected to the face direction estimation device 1.
[0049] Based on the face direction estimation result output from face direction estimation unit 16 in step ST6 and the face feature point information regarding the plurality of face feature points detected by feature point detection unit 12 in step ST2, 3D model adjustment unit 17 adjusts the positions of model face feature points corresponding to the face feature points in the 3D model for estimation (step ST7). After adjusting the 3D model for estimation, 3D model adjustment unit 17 updates the corresponding 3D model stored in storage unit 18 with the adjusted 3D model for estimation.
[0050] FIG. 6 is a flowchart for explaining in detail an example of the processing performed by the attribute estimation unit 14 in the skeletal attribute estimation processing in step ST4 of FIG. 5 . Note that FIG. 6 is a flowchart illustrating details of the skeletal attribute estimation processing when the attribute estimation unit 14 estimates the skeletal attribute of an occupant based on attribute estimation feature quantities and in accordance with attribute estimation conditions. As an example, the attribute estimation conditions are set to the following condition: "For attribute estimation feature quantities (here, first feature quantities and second feature quantities) calculated from a plurality of facial feature points detected based on the number of frames of face images for estimation, if the average value or median of the first feature quantities is equal to or greater than a first threshold and the average value or median of the second feature quantities is equal to or less than a second threshold, the skeletal attribute of the occupant is estimated to be the second skeletal attribute. For attribute estimation feature quantities calculated from a plurality of facial feature points detected based on the number of frames of face images for estimation, if the average value or median of the first feature quantities is less than the first threshold or the average value or median of the second feature quantities is greater than the second threshold, the skeletal attribute of the occupant is estimated to be the first skeletal attribute."
[0051] 5, the attribute estimation unit 14 determines whether the accessory estimation unit 13 has estimated that the occupant is wearing a mask (step ST41). If it is determined in step ST41 that the accessory estimation unit 13 has estimated that the occupant is wearing a mask ("YES" in step ST41), the attribute estimation unit 14 outputs the mask wearing information output from the accessory estimation unit 13 to the three-dimensional model selection unit 15. Then, the operation of the face direction estimation device 1 proceeds to step ST5 in FIG.
[0052] If it is determined in step ST41 that the accessory estimation unit 13 has estimated that the occupant is not wearing a mask (if "NO" in step ST41), the attribute estimation unit 14 calculates feature quantities for attribute estimation (first feature quantity and second feature quantity) from the plurality of facial feature quantities detected by the feature quantity detection unit 12 in step ST2 of FIG. 5 based on the facial image acquired by the image acquisition unit 11 in step ST1 of FIG. 5 in accordance with the feature quantity calculation conditions included in the attribute estimation conditions (step ST42).
[0053] Then, the attribute estimation unit 14 stores the feature amounts for attribute estimation calculated in step ST42 in an internal buffer or the like of the attribute estimation unit 14 (step ST43).
[0054] The attribute estimation unit 14 determines whether or not the attribute estimation features for the number of estimation frames have been stored in a buffer or the like inside the attribute estimation unit 14, in other words, whether or not the attribute estimation features to be stored have become full (step ST44).
[0055] If it is determined in step ST44 that the attribute estimation features for the number of estimation frames have been stored (if "YES" in step ST44), the attribute estimation unit 14 estimates the skeletal attribute of the occupant (step ST45). Specifically, the attribute estimation unit 14 compares the attribute estimation features for the number of estimation frames with a first threshold and a second threshold. If the average or median of the first features is equal to or greater than the first threshold and the average or median of the second features is equal to or less than a predetermined second threshold, the attribute estimation unit 14 estimates the skeletal attribute of the occupant to be the second skeletal attribute. If the average or median of the first features is less than the first threshold or the average or median of the second features is greater than the second threshold, the attribute estimation unit 14 estimates the skeletal attribute of the occupant to be the first skeletal attribute. The attribute estimation unit 14 outputs the skeletal attribute information to the three-dimensional model selection unit 15.
[0056] If the attribute estimation unit 14 determines in step ST44 that the number of attribute estimation feature amounts stored is not sufficient for the number of estimation frames (NO in step ST44), the operation of the face direction estimation device 1 returns to the processing of step ST1 in Fig. 5. Then, the face direction estimation device 1 acquires the next frame of face image from the imaging device 2, and performs each process shown in the flowchart in Fig. 5 on the acquired next frame of face image.
[0057] Here, the description has been given with reference to an example in which the attribute estimation unit 14 estimates the skeletal attributes of the occupant based on the attribute estimation features and in accordance with the attribute estimation conditions. However, the attribute estimation unit 14 may estimate the skeletal attributes of the occupant using, for example, an attribute estimation machine learning model. In this case, the processes of steps ST42 to ST44 of the operation of the attribute estimation unit 14 shown in the flowchart of FIG. 6 are omitted, and in step ST45, the attribute estimation unit 14 inputs the facial image acquired by the image acquisition unit 11 in step ST1 of FIG. 5 into the attribute estimation machine learning model to obtain information indicating the skeletal attributes, thereby estimating the skeletal attributes of the occupant, i.e., whether the skeletal attribute of the occupant is the first skeletal attribute or the second skeletal attribute. The attribute estimation unit 14 outputs the skeletal attribute information to the three-dimensional model selection unit 15.
[0058] In this way, the face direction estimation device 1 detects multiple facial feature points, including the eyes, nose, and mouth, from the facial image of the occupant. The face direction estimation device 1 also estimates the skeletal attributes of the occupant based on the facial image of the occupant, and selects a 3D model for estimation from multiple 3D models representing human faces corresponding to different skeletal attributes based on the estimated skeletal attributes of the occupant. The face direction estimation device 1 then estimates the facial direction of the occupant based on the multiple facial feature points and the 3D model for estimation. The face direction estimation device 1 also adjusts the positions of model facial feature points corresponding to the facial feature points in the 3D model for estimation based on the facial direction estimation result and information on the multiple facial feature points. Therefore, the face direction estimation device 1 can prevent erroneous estimation of the facial direction in 3D space of an estimation target captured in a captured image using a 3D model, compared to conventional techniques. For example, even if the 3D estimation model is a 3D model corresponding to a skeletal attribute in which the mouth does not protrude further forward than the nose when viewed from the side, while the occupant has a skeletal attribute in which the mouth protrudes further forward than the nose when viewed from the side, the face direction estimation device 1 can prevent erroneous estimation such as estimating the occupant's face direction to be at a pitch angle higher than the actual angle. In this way, the face direction estimation device 1 realizes highly accurate face direction estimation by being able to switch the 3D estimation model used to estimate the occupant's face direction in consideration of the skeletal attribute, which is an attribute of the occupant's skeleton that defines the positional relationship between the eyes, nose, and mouth when the occupant's face is viewed from the side.
[0059] In the first embodiment described above, the face direction estimation device 1 includes the accessory estimation unit 13, but this is merely an example. The face direction estimation device 1 does not necessarily include the accessory estimation unit 13, and can be configured without the accessory estimation unit 13. When the face direction estimation device 1 does not include the accessory estimation unit 13, the 3D model selection unit 15 in the face direction estimation device 1 selects a 3D model for estimation from among multiple 3D models representing human faces corresponding to different skeletal attributes based on the skeletal attributes of the occupant estimated by the attribute estimation unit 14. The storage unit 18 does not need to store a 3D mask model. In this case, with regard to the operation of the face direction estimation device 1 shown in the flowcharts of FIGS. 5 and 6 , the processing of step ST3 in FIG. 5 and the processing of step ST41 in FIG. 6 can be omitted.
[0060] Furthermore, in the above-described first embodiment, face direction estimation device 1 is described as including three-dimensional model adjustment unit 17, but this is merely an example. For example, the function of three-dimensional model adjustment unit 17 may be provided in a device external to face direction estimation device 1. In this case, the processing of step ST7 in FIG. 5 can be omitted from the operation of face direction estimation device 1 shown in the flowcharts of FIGS. 5 and 6 .
[0061] Furthermore, in the above-described first embodiment, examples of skeletal attributes in which the mouth protrudes further forward than the nose and skeletal attributes in which the mouth does not protrude further forward than the nose have been given as examples of erroneous estimation of the facial orientation of an occupant that may occur when the fact that skeletal attributes may differ from person to person is not taken into consideration, but this is merely an example. For example, if a three-dimensional model corresponds to a skeletal attribute in which the mouth and chin do not protrude further forward than the nose when viewed from the side and the positions of the eyes and mouth are approximately the same in the anterior-posterior direction, while an occupant has a skeletal attribute in which the mouth is large and recessed and is positioned further back in the anterior-posterior direction than the eyes when viewed from the side, when the facial feature points and model facial feature points are compared on a face image and the facial orientation of the occupant is estimated by changing the position and posture of the three-dimensional model in a virtual three-dimensional space, the facial orientation of the occupant may be erroneously estimated to be at a lower pitch angle than the actual one, in other words, facing downward than the actual one. In this case, by preparing in advance a three-dimensional model corresponding to a skeletal attribute in which the mouth is deeply recessed and positioned further back in the front-to-back direction than the eyes when viewed from the side, the face direction estimation device 1 can prevent the above-mentioned erroneous estimation. By making it possible to switch between three-dimensional models corresponding to various skeletal attributes, the face direction estimation device 1 can perform highly accurate face direction estimation taking into account the skeletal attributes, which are attributes of the occupant's skeleton that define the positional relationship between the eyes, nose, and mouth when the occupant's face is viewed from the side.
[0062] In the first embodiment described above, the face direction estimation device 1 is an in-vehicle device mounted on a vehicle, and the image acquisition unit 11, the feature point detection unit 12, the accessory estimation unit 13, the attribute estimation unit 14, the 3D model selection unit 15, the face direction estimation unit 16, and the 3D model adjustment unit 17 are provided in the in-vehicle device. However, without being limited to this, for example, some of the image acquisition unit 11, the feature point detection unit 12, the accessory estimation unit 13, the attribute estimation unit 14, the 3D model selection unit 15, the face direction estimation unit 16, and the 3D model adjustment unit 17 may be mounted in the in-vehicle device of the vehicle, and the rest may be provided in a server connected to the in-vehicle device via a network, and a face direction estimation system may be configured by the in-vehicle device and the server. In addition, the image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17 may all be provided on the server.
[0063] 7A and 7B are diagrams illustrating an example of the hardware configuration of face direction estimation device 1 according to embodiment 1. In embodiment 1, the functions of image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17 are realized by processing circuitry 101. That is, face direction estimation device 1 includes processing circuitry 101 for controlling estimation of the face direction of an estimation target in consideration of a skeletal attribute that indicates that the positional relationship between the eyes, nose, and mouth of a person facing directly ahead when viewed from the side may differ from person to person. Processing circuit 101 may be dedicated hardware as shown in FIG. 7A or a processor 104 that executes a program stored in memory as shown in FIG. 7B.
[0064] When processing circuitry 101 is dedicated hardware, processing circuitry 101 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.
[0065] When the processing circuit is a processor 104, the functions of the image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17 are realized by software, firmware, or a combination of software and firmware. The software or firmware is written as a program and stored in memory 105. The processor 104 reads and executes the program stored in memory 105 to execute the functions of the image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17. In other words, the face direction estimation device 1 includes memory 105 for storing a program that, when executed by the processor 104, results in the execution of steps ST1 to ST7 of FIG. 5 described above. In addition, the program stored in memory 105 can also be said to cause a computer to execute the processing procedures or methods of the image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17. Here, the memory 105 may be, for example, a non-volatile or volatile semiconductor memory such as a RAM, a ROM (Read Only Memory), a flash memory, an EPROM (Erasable Programmable Read Only Memory), or an EEPROM (Electrically Erasable Programmable Read-Only Memory), or a magnetic disk, a flexible disk, an optical disk, a compact disk, a mini disk, or a DVD (Digital Versatile Disc).
[0066] Note that the functions of the image acquisition unit 11, feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17 may be partially implemented by dedicated hardware and partially implemented by software or firmware. For example, the image acquisition unit 11 may be implemented by a processing circuit 101 as dedicated hardware, while the functions of the feature point detection unit 12, accessory estimation unit 13, attribute estimation unit 14, 3D model selection unit 15, face direction estimation unit 16, and 3D model adjustment unit 17 may be implemented by a processor 104 reading and executing programs stored in a memory 105. The storage unit 18 may be configured, for example, with the memory 105. The face direction estimation device 1 also includes an input interface device 102 and an output interface device 103 that perform wired or wireless communication with the imaging device 2 or an external device, etc.
[0067] As described above, according to the first embodiment, the face direction estimation device 1 includes an image acquisition unit 11 that acquires a face image of a vehicle occupant, a feature point detection unit 12 that detects a plurality of face feature points including the eyes, nose, and mouth from the face image acquired by the image acquisition unit 11, an attribute estimation unit 14 that estimates a skeletal attribute of the occupant based on the face image acquired by the image acquisition unit 11, and an attribute estimation unit 15 that selects a 3D model for estimation to be used for estimating the face direction from a plurality of 3D models representing human faces that correspond to different skeletal attributes based on the skeletal attribute of the occupant estimated by the attribute estimation unit 14. The face direction estimation device 1 is configured to include a 3D model selection unit 15 that selects a model, a face direction estimation unit 16 that estimates the face direction of the occupant based on the plurality of facial feature points detected by the feature point detection unit 12 and the estimation 3D model selected by the 3D model selection unit 15, and a 3D model adjustment unit 17 that adjusts the positions of model face feature points corresponding to the facial feature points in the estimation 3D model based on the face direction estimation result regarding the face direction of the occupant estimated by the face direction estimation unit 16 and information on the plurality of facial feature points detected by the feature point detection unit 12. Therefore, the face direction estimation device 1 can prevent erroneous estimation of the face direction in 3D space of an estimation target captured in a captured image using a 3D model, compared to conventional techniques. The face direction estimation device 1 can prevent erroneous estimation of the face direction of the occupant due to skeletal attributes of the occupant differing from skeletal attributes of the 3D model.
[0068] Furthermore, according to the first embodiment, in the face direction estimation device 1, the skeletal attributes of the occupant estimated by the attribute estimation unit 14 are attributes of the occupant's skeleton that define the positional relationship of the occupant's eyes, nose, and mouth when the occupant's face is viewed from the side, and the multiple three-dimensional models include two three-dimensional models that have different positional relationships of the eyes, nose, and mouth when viewed from the side. Therefore, the face direction estimation device 1 can prevent erroneous estimation of the face direction in the three-dimensional space of the estimation target captured in the captured image using the three-dimensional models, compared to conventional techniques. The face direction estimation device 1 can prevent erroneous estimation of the face direction of the occupant due to the skeletal attributes of the occupant differing from the skeletal attributes of the three-dimensional model.
[0069] According to the first embodiment, the plurality of three-dimensional models may further include a three-dimensional mask model representing the face of a person wearing a mask, and the face direction estimation device 1 may include an accessory estimation unit 13 that estimates whether the occupant is wearing a mask based on the face image acquired by the image acquisition unit 11, and the three-dimensional model selection unit 15 may be configured to select the three-dimensional mask model as the estimation three-dimensional model when the accessory estimation unit 13 estimates that the occupant is wearing a mask. This allows the face direction estimation device 1 to determine the face direction of the occupant assuming that the occupant is wearing a mask.
[0070] Any of the components of the embodiments may be modified or omitted.
[0071] The face direction estimation device of the present disclosure can prevent erroneous estimation of face direction in a three-dimensional space of an estimation target captured in a captured image using a three-dimensional model, compared to conventional techniques.
[0072] 1 Face direction estimation device, 11 Image acquisition unit, 12 Feature point detection unit, 13 Accessory estimation unit, 14 Attribute estimation unit, 15 Three-dimensional model selection unit, 16 Face direction estimation unit, 17 Three-dimensional model adjustment unit, 18 Storage unit, 2 Imaging device, 101 Processing circuit, 102 Input interface device, 103 Output interface device, 104 Processor, 105 Memory.
Claims
1. A face direction estimation device comprising: an image acquisition unit that acquires a facial image of a vehicle occupant; a feature point detection unit that detects a plurality of facial feature points, including the eyes, nose, and mouth, from the facial image acquired by the image acquisition unit; an attribute estimation unit that estimates skeletal attributes of the occupant based on the facial image acquired by the image acquisition unit; a 3D model selection unit that selects a 3D model for estimation to be used for facial direction estimation from a plurality of 3D models representing human faces corresponding to different skeletal attributes, based on the skeletal attributes of the occupant estimated by the attribute estimation unit; a face direction estimation unit that estimates the facial direction of the occupant based on the plurality of facial feature points detected by the feature point detection unit and the 3D model for estimation selected by the 3D model selection unit; and a 3D model adjustment unit that adjusts the positions of model facial feature points corresponding to the facial feature points in the 3D model for estimation, based on the face direction estimation result regarding the facial direction of the occupant estimated by the face direction estimation unit and information about the plurality of facial feature points detected by the feature point detection unit.
2. The face direction estimation device according to claim 1, characterized in that the skeletal attributes of the occupant estimated by the attribute estimation unit are attributes of the occupant's skeleton that define the relative positions of the eyes, nose, and mouth when the occupant's face is viewed from the side, and the multiple three-dimensional models include two three-dimensional models that have different relative positions of the eyes, nose, and mouth when viewed from the side.
3. The face direction estimation device according to claim 1, characterized in that the plurality of three-dimensional models further include a three-dimensional model for a mask showing the face of a person wearing a mask, and the device is equipped with an accessory estimation unit that estimates whether or not the occupant is wearing the mask based on the face image acquired by the image acquisition unit, and the three-dimensional model selection unit selects the three-dimensional model for a mask as the three-dimensional model for estimation when the accessory estimation unit estimates that the occupant is wearing the mask.
4. The face direction estimation device according to claim 1, characterized in that the attribute estimation unit estimates the skeletal attributes of the occupant based on the facial image acquired by the image acquisition unit and a machine learning model that inputs the facial image and outputs information indicating the skeletal attributes.
5. The face direction estimation device according to claim 1, characterized in that the attribute estimation unit estimates the skeletal attributes of the occupant in accordance with attribute estimation conditions based on facial feature amounts calculated from a plurality of facial feature points detected by the feature point detection unit based on the facial image acquired by the image acquisition unit.
6. The face direction estimation device according to claim 1, characterized in that the face direction estimation unit compares, on the face image, the plurality of face feature points detected by the feature point detection unit with the plurality of model face feature points corresponding to each of the plurality of face feature points in the estimation three-dimensional model, and estimates the face direction of the occupant by changing the position and attitude of the estimation three-dimensional model in virtual three-dimensional space so that the error between the positions of the plurality of face feature points on the face image and the positions of the plurality of model face feature points is within an error judgment threshold.
7. The face direction estimation device of claim 6, characterized in that the face direction estimation result includes information about projected model face feature points, which are points on the face image corresponding to the plurality of model face feature points in the estimation 3D model after the face direction estimation unit has changed the position and posture, and the 3D model adjustment unit adjusts the positions of the model face feature points in the estimation 3D model based on errors between the plurality of face feature points detected by the feature point detection unit and the plurality of projected model face feature points corresponding to each of the plurality of face feature points.
8. An image acquisition unit acquires a facial image of a vehicle occupant; a feature point detection unit detects a plurality of facial feature points including eyes, nose, and mouth from the facial image acquired by the image acquisition unit; an attribute estimation unit estimates skeletal attributes of the occupant based on the facial image acquired by the image acquisition unit; a three-dimensional model selection unit selects a three-dimensional model for estimation to be used for estimating a facial direction from a plurality of three-dimensional models representing human faces corresponding to different skeletal attributes based on the skeletal attributes of the occupant estimated by the attribute estimation unit; a facial direction estimation unit estimates the facial direction of the occupant based on the plurality of facial feature points detected by the feature point detection unit and the three-dimensional model for estimation selected by the three-dimensional model selection unit; a step of adjusting positions of model facial feature points corresponding to the facial feature points in the estimation three-dimensional model, based on a facial direction estimation result regarding the facial direction of the occupant estimated by the facial direction estimation unit and information regarding the plurality of facial feature points detected by the feature point detection unit, by a three-dimensional model adjustment unit.
Citation Information
Patent Citations
Image recognition apparatus, method and program
JP2007004767A
Face recognition device
JP2018163481A
Image processing device, image processing method, and program
WO2015029982A1