Gaze estimation device, gaze estimation method, and gaze estimation program

The gaze estimation device addresses the challenge of estimating the gaze of a moving subject by integrating gazes from multiple cameras and prioritizing frontal and closer-distance views, achieving accurate gaze estimation over wide areas.

WO2025104924A1PCT designated stage expired Publication Date: 2025-05-22MITSUBISHI ELECTRIC CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/041528
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Conventional gaze estimation technologies struggle to accurately estimate the gaze of a person who is moving over a wide area, as they typically assume the subject is stationary and close to the camera.

Method used

A gaze estimation device that uses multiple cameras to capture image data, integrates the estimated gazes based on how the subject appears in the image data, and prioritizes gazes from frontal and closer-distance views to improve accuracy.

Benefits of technology

Enables accurate gaze estimation even when the subject moves over a wide area by integrating gazes from multiple cameras and prioritizing frontal and closer-distance views, thereby enhancing the reliability of gaze detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023041528_22052025_PF_FP_ABST
    Figure JP2023041528_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A gaze estimation unit (23) sets each of a plurality of cameras (31) as a target camera (31). The gaze estimation unit (23) estimates the gaze of a subject from image data (41) imaged by the target camera (31). A gaze synthesis unit (24) defines the gaze of the subject by synthesizing the gazes that were estimated by the gaze estimation unit (23) treating each of the plurality of cameras (31) as the target camera (31), in accordance with how the subject appears in the image data (41) imaged by the target camera (31).
Need to check novelty before this filing date? Find Prior Art

Description

Gaze estimation device, gaze estimation method, and gaze estimation program

[0001] The present disclosure relates to a technology for estimating the gaze of a person captured in image data.

[0002] Human gaze is one of the important pieces of information for understanding a person's behavioral intentions. Therefore, technology for estimating human gaze is used in a variety of situations.

[0003] A conventional gaze estimation technique is a technique for estimating the gaze direction of a stationary person from a color image of the person's face. Patent Document 1 describes a conventional gaze estimation technique.

[0004] In particular, Patent Document 1 describes a technology for solving the problem that the accuracy of gaze estimation may decrease when the face orientation, lighting brightness, etc. are under certain conditions. In Patent Document 1, the gaze is estimated using multiple estimators for a single face image. The gaze is then determined by integrating the results of the multiple estimators using a weighted average or the like based on a comparison of the distance between the camera and the person when capturing the face image, the camera's angle of view, and data collected in advance. This improves the accuracy of gaze estimation under various conditions.

[0005] Japanese Patent Application Laid-Open No. 2018-547082

[0006] Conventional gaze estimation technologies often assume that the subject of gaze estimation is stationary and near a camera, as in the following cases (1) and (2): (1) A case in which the gaze of a computer operator is estimated from image data captured by a camera attached to a display device; and (2) A case in which the gaze of a vehicle driver is estimated from image data captured by an in-vehicle camera. The technology described in Patent Document 1 may also function effectively when the subject of gaze estimation is stationary and near a camera. However, when the subject of gaze estimation moves around a city or the like and over a wide area, it is difficult to appropriately estimate the gaze. The present disclosure aims to enable appropriate gaze estimation even when the subject of gaze estimation moves over a wide area.

[0007] The gaze estimation device according to the present disclosure includes a gaze estimation unit that estimates the gaze of a subject from image data captured by each of a plurality of cameras as a target camera, and a gaze integration unit that integrates the gazes estimated by the gaze estimation unit for each of the plurality of cameras as the target camera according to how the subject appears in the image data captured by the target camera, to identify the gaze of the subject.

[0008] In the present disclosure, the gaze of a subject is identified by integrating the gaze depending on how the subject appears in the image data, thereby enabling appropriate gaze estimation even when the subject of gaze estimation moves over a wide area.

[0009] 1 is a configuration diagram of a gaze estimation device 10 according to embodiment 1. FIG. 2 is a flowchart showing the operation of the gaze estimation device 10 according to embodiment 1. FIG. 3 is a flowchart showing the operation of the gaze estimation device 10 according to embodiment 2. FIG. 4 is an explanatory diagram of gaze correction processing according to embodiment 2. FIG. 5 is an explanatory diagram of gaze correction processing according to modification 1. FIG. 6 is a configuration diagram of a gaze estimation device 10 according to embodiment 3. FIG. 7 is an explanatory diagram of a camera 32 and gaze position estimation processing according to embodiment 3. FIG. 8 is a flowchart showing the operation of the gaze estimation device 10 according to embodiment 3. FIG. 9 is a configuration diagram of a gaze estimation device 10 according to embodiment 4. FIG. 10 is a flowchart showing the operation of the gaze estimation device 10 according to embodiment 4. FIG. 11 is a flowchart showing the generation processing of an obstruction removal model according to embodiment 4. FIG. 12 is a configuration diagram of a gaze estimation device 10 according to modification 3.

[0010] Embodiment 1. ***Description of Configuration*** The configuration of a gaze estimation device 10 according to embodiment 1 will be described with reference to Fig. 1. The gaze estimation device 10 is a computer. The gaze estimation device 10 includes hardware such as a processor 11, a memory 12, a storage 13, and a communication interface 14. The processor 11 is connected to other hardware via signal lines and controls the other hardware.

[0011] The processor 11 is an IC that performs processing. IC stands for Integrated Circuit. Specific examples of the processor 11 include a CPU, a DSP, and a GPU. CPU stands for Central Processing Unit. DSP stands for Digital Signal Processor. GPU stands for Graphics Processing Unit.

[0012] The memory 12 is a storage device that temporarily stores data. Specific examples of the memory 12 include SRAM and DRAM. SRAM stands for Static Random Access Memory. DRAM stands for Dynamic Random Access Memory.

[0013] The storage 13 is a storage device that stores data. A specific example of the storage 13 is an HDD. HDD is an abbreviation for Hard Disk Drive. The storage 13 may also be a portable recording medium such as an SD (registered trademark) memory card, CompactFlash (registered trademark), NAND flash, a flexible disk, an optical disk, a compact disk, a Blu-ray (registered trademark) disk, or a DVD. SD is an abbreviation for Secure Digital. DVD is an abbreviation for Digital Versatile Disk.

[0014] The communication interface 14 is an interface for communicating with external devices. Specific examples of the communication interface 14 include Ethernet (registered trademark), USB, and HDMI (registered trademark) ports. USB stands for Universal Serial Bus. HDMI stands for High-Definition Multimedia Interface.

[0015] The gaze estimation device 10 is connected to multiple cameras 31 via a communication interface 14. In the following description, it is assumed that N cameras 31 are connected. The cameras 31 may be RGB cameras, RGB-D cameras, stereo cameras, etc. RGB stands for Red Green Blue. RGB-D stands for Red Green Blue-Depth.

[0016] The gaze estimation device 10 includes, as functional components, a photographing unit 21, a face detection unit 22, a gaze estimation unit 23, and a gaze integration unit 24. The functions of each functional component of the gaze estimation device 10 are realized by software. A program that realizes the function of each functional component of the gaze estimation device 10 is stored in the storage 13. This program is read into the memory 12 by the processor 11 and executed by the processor 11. In this way, the function of each functional component of the gaze estimation device 10 is realized.

[0017] ***Description of Operation*** The operation of the gaze estimation device 10 according to the first embodiment will be described with reference to Fig. 2. The operation procedure of the gaze estimation device 10 according to the first embodiment corresponds to the gaze estimation method according to the first embodiment. Furthermore, the program that realizes the operation of the gaze estimation device 10 according to the first embodiment corresponds to the gaze estimation program according to the first embodiment.

[0018] The processes from step S11 to step S13 are executed for each of the plurality of cameras 31 as a target camera 31.

[0019] (Step S11: Photographing Process) The photographing unit 21 photographs the photographing area using the target camera 31. In this way, the photographing unit 21 acquires image data 41 photographed by the target camera 31.

[0020] (Step S12: Face detection processing) The face detection unit 22 detects the face of the subject from the image data 41 acquired in step S11. The subject is a person whose gaze is to be estimated. In this way, the face detection unit 22 acquires three-dimensional coordinates of the facial feature points of the subject. Note that if the face of the subject is not detected by the face detection unit 22, processing for the target camera 31 ends. The face detection method can be realized using existing technology. As a specific example, the face detection unit 22 may detect the face of the subject by inputting the image data 41 into a face detection model generated by machine learning such as supervised learning.

[0021] (Step S13: Gaze Estimation Process) The gaze estimation unit 23 estimates the gaze of the subject from the image data 41. The gaze is represented by the position of the subject's eyes and a vector indicating the direction from the position of the subject's eyes to the gaze position. The position of the subject's eyes is the three-dimensional coordinate of the midpoint of the line segment connecting the subject's eyes. A method for estimating the gaze from image data can be realized using existing technology. As a specific example, the gaze estimation unit 23 may estimate the gaze of the subject by inputting the image data 41 into a gaze estimation model generated by machine learning such as supervised learning. Note that when estimating the gaze from image data, the estimation is often performed by focusing on the movement of the pupil within the eyeball. If the camera 31 is an RGB-D camera or a stereo camera, depth information can be obtained. In this case, the gaze estimation unit 23 may use the depth information to estimate the position of the subject's eyes, which is the origin of the gaze. Using the depth information can improve the accuracy of estimating the position of the subject's eyes.

[0022] (Step S14: Gaze Integration Processing) The gaze integration unit 24 integrates the gazes estimated from each of the multiple cameras 31 as the target camera 31 to identify the gaze of the subject. Specifically, the gaze integration unit 24 integrates the gazes according to how the subject appears in the image data 41 captured by the target camera, to identify the gaze of the subject. The gaze integration unit 24 outputs the identified gaze to an external device via the communication interface 14. For example, the gaze integration unit 24 outputs the identified gaze to a display device for display. At this time, the gaze integration unit 24 may also output the three-dimensional coordinates of the facial feature points obtained when the face was detected in step S12.

[0023] When estimating the gaze in step S13, the gaze detection accuracy tends to be higher when the subject's face is photographed from a more frontal position. Furthermore, the gaze detection accuracy tends to be higher when the subject's face is photographed from a closer position. Therefore, the gaze integration unit 24 prioritizes the gaze estimated from image data 41 photographed from the front of the subject during integration. Furthermore, the gaze integration unit 24 prioritizes the gaze estimated from image data 41 photographed from a closer distance during integration. "Prioritizing integration" refers to, for example, integrating by assigning a heavier weight to the gaze estimated from the image data 41 photographed from the subject at a closer distance. Specifically, the gazes are integrated by multiplying the gazes estimated from each of the multiple cameras 31 by their weights and dividing the sum of the weights by the sum of the weights. In other words, the subject's eye position is identified by dividing the sum of the weighted values ​​by which the gaze vector is multiplied by the weights by the sum of the weights. Furthermore, the gaze vector is identified by dividing the sum of the weighted values ​​by which the gaze vector is multiplied by the weights by the sum of the weights.

[0024] More specifically, in the second embodiment, the gaze integrating unit 24 calculates the reliability of the gazes by at least one of the following two methods (1) and (2). Then, the gaze integrating unit 24 integrates the gazes by a weighted average weighted by the calculated reliability. Here, the higher the reliability, the heavier the weight.

[0025] (1) Reliability Based on Iris The gaze integration unit 24 calculates the reliability so that the more clearly the subject's iris is photographed, the higher the reliability. The evaluation indexes for determining whether the iris is clear or not are the shape of the iris and the size of the iris area. The closer the iris shape is to a circle, the clearer the iris is judged to be. The larger the size of the iris area, the clearer the iris is judged to be. The more the subject's face is photographed from the front, the closer the iris shape becomes to a circle. Also, the closer the subject's face is photographed from a position, the larger the size of the iris area becomes. In other words, the clearer the iris is photographed and the higher the reliability, the higher the gaze detection accuracy tends to be.

[0026] (2) Reliability based on distance from camera 31 The gaze integration unit 24 uses the center position C eye is the center position C of the target camera 31. cam For example, the line-of-sight integration unit 24 calculates the reliability such that the closer the line of sight is to the distance |C cam -C eye α / (|C cam -C eye |) is calculated as the reliability. cam -C eye If | is 0, the weight may be set to 1. The closer the subject's face is photographed from the front and from a closer position, the smaller the distance |C cam -C eye | becomes shorter. In other words, the distance |C cam -C eye The shorter | is and the higher the reliability is, the higher the accuracy of gaze detection tends to be.

[0027] The gaze integrating unit 24 may exclude gazes with a reliability (weight) below a reference value from the integration targets. This is because it is thought that integrating gazes with low reliability will tend to result in a large error. Note that if the reliability is low, the weight is small, so even if the gaze is included as an integration target, the impact on the result is small, so it is not necessary to exclude it from the integration targets.

[0028] Here, when integrating the gazes, it is necessary to integrate the gazes estimated from the image data 41 captured by each camera 31 at the same timing. The gaze integration unit 24 may perform synchronization processing using the image capture times to identify the gazes estimated from the image data 41 captured by each camera 31 at the same timing. The synchronization processing may be performed based on the surrounding conditions captured in the image data 41. Alternatively, the synchronization processing may be performed using a synchronization signal.

[0029] In step S14, the line-of-sight integration unit 24 may output the weight used for integration together with the line of sight.

[0030] ***Effects of Embodiment 1*** As described above, gaze estimation device 10 according to Embodiment 1 identifies the gaze of the subject by integrating the gazes depending on how the subject appears in image data 41. This makes it possible to appropriately estimate the gaze even in cases where the subject of gaze estimation moves over a wide area, which is difficult to estimate from image data 41 captured by a single camera 31.

[0031] In particular, the gaze estimation device 10 according to the first embodiment prioritizes and integrates gazes estimated from image data 41 in which the subject is photographed from the front. Furthermore, the gaze estimation device 10 according to the first embodiment prioritizes and integrates gazes estimated from image data 41 in which the subject is photographed from a closer distance. This makes it possible to estimate the gaze with high accuracy when taking into account image data 41 photographed by multiple cameras 31.

[0032] Furthermore, the gaze estimation device 10 according to the first embodiment identifies whether the image data 41 shows a subject photographed from the front or a subject photographed from a close distance, by using at least one of the distance between the iris and the camera 31. This makes it possible to easily and accurately identify whether the image data 41 shows a subject photographed from the front or a subject photographed from a close distance.

[0033] Furthermore, the gaze estimation device 10 according to embodiment 1 identifies the gaze from image data 41 captured by multiple cameras 31, and is therefore capable of identifying the gaze even if some of the cameras 31 fail.

[0034] Embodiment 2. Embodiment 2 differs from embodiment 1 in that the estimated line of sight is corrected. In embodiment 2, this difference will be explained, and explanation of the same points will be omitted.

[0035] ***Description of Operation*** The operation of the gaze estimation device 10 according to embodiment 2 will be described with reference to Fig. 3. The processes from step S21 to step S23 are the same as the processes from step S11 to step S13 in Fig. 2. The process of step S25 is the same as the process of step S14 in Fig. 2.

[0036] (Step S24: Gaze Correction Process) The gaze estimation unit 23 corrects the gaze estimated in step S23. Specifically, the gaze estimation unit 23 calculates a projection transformation matrix for calibration in advance. Then, the gaze estimation unit 23 corrects the estimated gaze by transforming it using the projection transformation matrix.

[0037] This will be described in more detail with reference to FIG. 4 . The gaze is estimated by focusing on the movement of the pupil within the eyeball. The farther the gaze position is from side to side or up and down, the greater the movement of the pupil. The greater the movement of the pupil, the greater the error in the estimated gaze, resulting in distortion. Therefore, the gaze estimation unit 23 estimates the gaze when gazing at a calibration point. The gaze estimation unit 23 identifies the gaze position determined from the estimated gaze as the estimated position. The actual gaze position of the gazed calibration point is called the actual position. The gaze estimation unit 23 calculates a projective transformation matrix such that the estimated position matches the actual position. Specifically, as shown in FIG. 4 , the gaze estimation unit 23 defines a rectangular area on the wall as a gaze area 43, and sets the four corners of the gaze area 43 as calibration points. The gaze estimation unit 23 targets each of the four corner points, estimates the gaze when gazing at the target point, and identifies the estimated position. This identifies an estimated area 44, which is a rectangular area connecting the estimated positions. The correction from the estimated region 44 to the gaze region 43 can be regarded as a projective transformation using the four corner points of the gaze region 43 and the estimated region 44. Therefore, the gaze estimation unit 23 calculates a projective transformation matrix for correcting the estimated region 44 to the gaze region 43, as shown in Equation 1.

[0038] For example, assume that the estimated position identified from the gaze estimated in step S25 is position (x, y) in the estimation area 44 in Fig. 4. Then, the position (x, y) is corrected to position (x', y') in the gaze area 43 by the projective transformation matrix of Equation 1. As a result, the gaze position becomes position (x', y'), and the gaze direction is from the position of the subject's eyes to position (x', y').

[0039] ***Effects of Embodiment 2*** As described above, the gaze estimation device 10 according to Embodiment 2 corrects the estimated gaze using a projective transformation matrix prepared in advance, thereby making it possible to reduce errors contained in the estimated gaze.

[0040] ***Other Configurations*** <Variation 1> The shape of the human eye is asymmetric. Therefore, the magnitude of the error and the way it is distorted vary depending on the direction of gaze. For example, when you move your eyeball to look up, your eyelids rise and your eyes open wide. On the other hand, when you move your eyeball to look down, your eyelids drop and your eyes close slightly.

[0041] Therefore, the gaze estimation unit 23 prepares multiple projective transformation matrices according to the gaze direction. Specifically, the field of view is divided into multiple divided regions 45 based on the position directly in front of the subject, and a projective transformation matrix is ​​prepared for each divided region 45. For example, as shown in FIG. 5 , the field of view is divided into four divided regions 45, (1) upper left, (2) upper right, (3) lower left, and (4) lower right, based on the position directly in front of the subject, and a projective transformation matrix is ​​prepared for each of the four divided regions 45. The gaze estimation unit 23 then corrects the estimated gaze using the projective transformation matrix corresponding to the divided region 45 into which the gaze is estimated to fall. For example, if the gaze is estimated to fall into the (2) upper right divided region 45, the gaze estimation unit 23 corrects the gaze using the projective transformation matrix corresponding to the (2) upper right divided region 45.

[0042] Here, the divided area 45 into which the gaze is estimated to fall means that the estimated gaze falls into the estimated area 44 when the divided area 45 is set as the gaze area 43. In other words, the gaze estimation unit 23 sets each divided area 45 as the gaze area 43. The gaze estimation unit 23 sets each of the four corner points of the gaze area 43 as calibration points. The gaze estimation unit 23 estimates the gaze when the calibration points are gazed upon. This specifies the estimated area 44 corresponding to each divided area 45. Then, the gaze estimation unit 23 specifies the estimated area 44 into which the gaze estimated in step S23 falls, thereby determining which projective transformation matrix to use for correction.

[0043] As described above, by correcting the line of sight using a projective transformation matrix according to the line of sight direction, it is possible to further reduce errors.

[0044] Embodiment 3. Embodiment 3 differs from Embodiments 1 and 2 in that it estimates the gaze position that exists ahead of the line of sight. In Embodiment 3, this difference will be explained, and explanation of the same points will be omitted. In Embodiment 3, a case where a function is added to Embodiment 1 will be explained. However, it is also possible to add a function to Embodiment 2.

[0045] ***Description of Configuration*** The configuration of the gaze estimation device 10 according to embodiment 3 will be described with reference to Fig. 6. The gaze estimation device 10 differs from the gaze estimation device 10 shown in Fig. 1 in that it includes a gaze position estimation unit 25 as a functional component. The function of the gaze position estimation unit 25 is realized by software, like the other functional components.

[0046] Furthermore, the gaze estimation device 10 differs from the gaze estimation device 10 shown in FIG. 1 in that it is connected to one or more cameras 32 for space recognition via a communication interface 14. The cameras 32 are, for example, RGB-D cameras and stereo cameras. Here, the camera 31 is a camera for capturing an image of the subject. In contrast, the camera 32 is a camera for capturing an image of the space in which the subject is present. For example, as shown in FIG. 7 , the camera 31 is installed so as to capture an image of the subject from the front side. On the other hand, the camera 32 is installed so as to capture an image from the rear side of the subject or from a bird's-eye view so as to capture an image of the space in which the subject is present.

[0047] ***Description of Operation*** The operation of the gaze estimation device 10 according to embodiment 3 will be described with reference to Fig. 8. The processes from step S31 to step S34 are the same as the processes from step S11 to step S14 in Fig. 2.

[0048] (Step S35: Gaze Position Estimation Process) The gaze position estimation unit 25 estimates the gaze position located at the line of sight identified in step S34. Specifically, the gaze position estimation unit 25 generates a three-dimensional model of the space where the subject is located from image data captured by the camera 32. The three-dimensional model of the space where the subject is located may also be generated from known data, such as design data of the space where the subject is located. As shown in FIG. 7 , the gaze position estimation unit 25 sets the line of sight identified in step S34 on the three-dimensional model. That is, the gaze position estimation unit 25 sets a vector indicating the direction from the eye position of the subject indicated by the line of sight to the gaze position in the three-dimensional model. The gaze position estimation unit 25 extends the set vector and identifies the intersection with the object in the three-dimensional model as the gaze position. The gaze position estimation unit 25 needs to align the three-dimensional model with the image data 41 captured by each camera 31.

[0049] ***Effects of Embodiment 3*** As described above, the gaze estimation device 10 according to embodiment 3 can estimate the gaze position by setting the gaze in a three-dimensional model. When the subject moves over a wide area, the positional relationship of the subject as seen from the camera 31 changes. This changes the positional relationship between the subject and the object being gazed at. For this reason, it is difficult to identify the gaze position simply by identifying the gaze. However, the gaze estimation device 10 according to embodiment 3 can estimate the gaze position even when the subject moves over a wide area by using a three-dimensional model of the space in which the subject is located.

[0050] Embodiment 4. Embodiment 4 differs from Embodiments 1 to 3 in that the gaze is estimated after removing any obstructions, such as glasses or a mask, worn on the subject's face. In Embodiment 4, this difference will be explained, and explanations of the same points will be omitted. In Embodiment 4, a case where a function is added to Embodiment 1 will be explained. However, it is also possible to add a function to Embodiments 2 and 3.

[0051] ***Description of Configuration*** The configuration of the gaze estimation device 10 according to embodiment 4 will be described with reference to Fig. 9. The gaze estimation device 10 differs from the gaze estimation device 10 shown in Fig. 1 in that it includes an image generation unit 26 as a functional component. The function of the image generation unit 26 is realized by software, like the other functional components.

[0052] ***Description of Operation*** The operation of the gaze estimation device 10 according to the fourth embodiment will be described with reference to Fig. 10. The processes from step S41 to step S42 are the same as the processes from step S11 to step S12 in Fig. 2. Furthermore, the process of step S46 is the same as the process of step S14 in Fig. 2.

[0053] (Step S43: Obstruction Determination Process) The image generation unit 26 determines whether or not an obstruction is attached to the face of the subject detected in step S42. An obstruction is an object that obscures the face of the subject. Specific examples of the obstruction include glasses or a mask. The method for determining whether or not an obstruction is attached can be realized using existing technology. Specific examples include the image generation unit 26 determining whether or not an obstruction is attached to the face of the subject by inputting the image data 41 and three-dimensional coordinates of facial feature points into an obstruction presence / absence determination model generated by machine learning such as supervised learning. If an obstruction is attached, the image generation unit 26 proceeds to step S44. On the other hand, if an obstruction is not attached, the image generation unit 26 skips step S44 and proceeds to step S45.

[0054] Here, if the user is wearing glasses, the accuracy of gaze estimation may be reduced due to reflection or distortion from the lenses, etc. Also, if the user is wearing a mask, the facial features required for gaze estimation may not be detected properly, and gaze estimation may not be performed properly.

[0055] (Step S44: Image Generation Processing) The image generation unit 26 generates an unworn image 42 from the worn image, which is image data 41 determined to contain an occluder in step S43, by removing the occluder. Specifically, the image generation unit 26 uses a generative adversarial network (GAN) to generate the unworn image 42 from the worn image. GAN stands for Generative Adversarial Networks. A generative adversarial network is a type of generative model. A generative adversarial network is a generative model that can generate non-existent data in accordance with the characteristics of input data by learning features from data. Here, the image generation unit 26 uses an occluder removal model, which is a generative adversarial network that generates an unworn image 42 from the face by removing the occluder. That is, the image generation unit 26 provides the worn image as input to the occluder removal model to generate the unworn image 42 from the face by removing the occluder.

[0056] 2, the gaze estimation unit 23 estimates the gaze of the subject. At this time, if it is determined in step S43 that no obstruction is being worn, the gaze estimation unit 23 estimates the gaze from the image data 41. On the other hand, if it is determined in step S43 that a obstruction is being worn, the gaze estimation unit 23 estimates the gaze from the uncovered image 42.

[0057] The process of generating an object-removed model according to the fourth embodiment will be described with reference to Fig. 11. The process of generating an object-removed model is executed before the process of Fig. 10 is executed. Therefore, the object-removed model is generated before the process of Fig. 10 is executed.

[0058] (Step S51: Obstruction object image generation process) The image generation unit 26 rotates a 3D model of an obstruction to generate an obstruction image, which is an image of the obstruction for each facial direction. The image generation unit 26 prepares a 3D model for each type of obstruction, such as glasses and a mask. The image generation unit 26 may also prepare multiple 3D models for each shape of the same type of obstruction. The image generation unit 26 generates an obstruction image for each facial direction for each 3D model.

[0059] (Step S52: Face Direction Estimation Process) The image generation unit 26 sets each of a plurality of face images in which no obstruction is attached as a processing target. The image generation unit 26 estimates the face direction in the face image to be processed. Estimation of the face direction can be realized using existing technology.

[0060] (Step S53: Wearing Image Generation Processing) The image generation unit 26 sets each face image and each 3D model as a processing target. The image generation unit 26 identifies an obstruction image generated from the 3D model of the processing target in step S51, which corresponds to the face orientation estimated in step S52 for the face image of the processing target. The image generation unit 26 generates a wearing training image by superimposing the identified obstruction image on the face image of the processing target. At this time, the image generation unit 26 superimposes the obstruction image at a position according to the type of obstruction. Note that if the obstruction type is glasses, reflection or distortion may occur due to the lenses. Therefore, the image generation unit 26 may reflect possible reflection and distortion in the face image of the processing target. The reflection and distortion are reflected only in portions of the face image that are affected by the lenses. The image generation unit 26 then generates a wearing training image by superimposing the obstruction image on the face image with the reflection and distortion reflected.

[0061] (Step S54: Learning process) The image generation unit 26 sets each face image as a processing target. The image generation unit 26 sets a pair of the non-wearing learning image, which is the face image to be processed, and each wearing learning image generated from the face image to be processed, as learning data. The image generation unit 26 performs learning using the learning data and generates an occlusion removal model.

[0062] ***Effects of Embodiment 4*** As described above, when an obstruction is worn, the gaze estimation device 10 according to Embodiment 4 generates an unworn image from which the obstruction has been removed and estimates the gaze from the unworn image. This makes it possible to appropriately estimate the gaze even when the subject is wearing an obstruction.

[0063] ***Other Configurations*** <Modification 2> In the second to fourth embodiments, similar to the first embodiment, gazes estimated from image data 41 captured by a plurality of cameras 31 are integrated. However, in the second to fourth embodiments, the number of cameras 31 may be one. A configuration may also be adopted in which the gaze is estimated from image data 41 captured by one camera 31. In this case, the function of gaze integration unit 24 is unnecessary.

[0064] <Modification 3> In the first to fourth embodiments, only one processor 11 is illustrated. However, there may be multiple processors 11, and the multiple processors 11 may execute programs that realize each function in cooperation with each other. For example, as shown in FIG. 12 , the gaze estimation device 10 may include a processor 11 for each camera 31, and the gaze estimation may be performed by the processor 11 corresponding to each camera 31. The gaze estimation device 10 may then perform processing common to each camera 31, such as integrating the estimated gazes, using a processor 11 other than the processor 11 provided for each camera 31. That is, among the functional components described in the first to fourth embodiments, the photographing unit 21, face detection unit 22, gaze estimation unit 23, and image generation unit 26 are realized by the processor 11 provided for each camera 31. The remaining gaze integration unit 24 and gaze position estimation unit 25 are realized by a processor 11 other than the processor 11 provided for each camera 31. Note that the memory 12, storage 13, and communication interface 14 are omitted from FIG. 12 .

[0065] <Modification 4> In the first to fourth embodiments, each functional component is realized by software. However, in Modification 4, each functional component may be realized by hardware. The following describes the differences between Modification 4 and the first to fourth embodiments.

[0066] When each functional component is realized by hardware, the gaze estimation device 10 includes an electronic circuit instead of the processor 11, the memory 12, and the storage 13. The electronic circuit is a dedicated circuit that realizes the functions of each functional component, the memory 12, and the storage 13.

[0067] Possible electronic circuits include a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, a logic IC, a GA, an ASIC, and an FPGA. GA stands for Gate Array. ASIC stands for Application Specific Integrated Circuit. FPGA stands for Field-Programmable Gate Array. Each functional component may be realized by a single electronic circuit, or each functional component may be distributed across multiple electronic circuits.

[0068] <Modification 5> As a modification 5, some of the functional components may be realized by hardware, and other functional components may be realized by software.

[0069] The processor 11, the memory 12, the storage 13, and the electronic circuitry are collectively referred to as a processing circuit. In other words, the functions of the respective functional components are realized by the processing circuit.

[0070] Furthermore, the term "unit" in the above description may be read as a "circuit," "step," "procedure," "process," or "processing circuit."

[0071] The embodiments and modifications of the present disclosure have been described above. Some of these embodiments and modifications may be combined and implemented. Furthermore, one or more of them may be implemented partially. Note that the present disclosure is not limited to the above embodiments and modifications, and various modifications are possible as needed.

[0072] 10 Gaze estimation device, 11 Processor, 12 Memory, 13 Storage, 14 Communication interface, 15 Electronic circuit, 21 Photography unit, 22 Face detection unit, 23 Gaze estimation unit, 24 Gaze integration unit, 25 Gaze position estimation unit, 26 Image generation unit, 31 Camera, 32 Camera, 41 Image data, 42 Non-wearing image, 43 Gaze area, 44 Estimation area, 45 Segmented area.

Claims

1. A gaze estimation device comprising: a gaze estimation unit that estimates the gaze of a subject from image data captured by a plurality of cameras, each of which is a target camera; and a gaze integration unit that identifies the gaze of the subject by integrating the gazes estimated by the gaze estimation unit for each of the plurality of cameras, each of which is a target camera, in accordance with how the subject appears in the image data captured by the target camera.

2. The gaze estimation device according to claim 1, wherein the gaze integration unit prioritizes and integrates gazes estimated from image data in which the subject is photographed from the front.

3. The gaze estimation device according to claim 1 or 2, wherein the gaze integration unit prioritizes and integrates gazes estimated from image data captured from a closer distance of the subject.

4. The gaze estimation device according to any one of claims 1 to 3, wherein the gaze integration unit performs integration based on at least one of the size and shape of the subject's iris.

5. The gaze estimation device according to any one of claims 1 to 4, wherein the gaze integration unit performs integration according to the distance between the positions of the subject's eyes and the position of the camera.

6. The gaze estimation device according to any one of claims 1 to 5, wherein the gaze estimation unit corrects the gaze estimated from the image data using a projective transformation matrix that matches the gaze position identified when gazing at a calibration point with the actual gaze position.

7. The gaze estimation device according to claim 6, wherein a plurality of the projective transformation matrices are prepared according to a gaze direction, and the gaze estimation unit corrects the gaze using the projective transformation matrix corresponding to the gaze direction of the subject estimated from the image data.

8. The gaze estimation device described in any one of claims 1 to 7, further comprising a gaze position estimation unit that uses a three-dimensional model of the space in which the subject is located to estimate the gaze position located at the end of the gaze identified by the gaze integration unit.

9. The gaze estimation device according to any one of claims 1 to 8, further comprising an image generation unit that, when the subject captured in the image data captured by the target camera is wearing a mask covering his or her face, generates a non-wearing image from a wearing image, which is image data showing the subject wearing the mask, in which the mask has been removed, and the gaze estimation unit estimates the gaze from the non-wearing image generated by the image generation unit.

10. The gaze estimation device described in claim 9, wherein the image generation unit is an occlusion removal model obtained by training a pair of a non-wearing training image, which is a face image without an occlusion object, and a wearing training image in which the non-wearing training image is worn with an occlusion object attached, as training data, and generates the non-wearing image by inputting the wearing image into the occlusion removal model which takes an input of a face image with an occlusion object attached and outputs image data from which the occlusion has been removed.

11. A gaze estimation method in which a computer estimates the gaze of a subject from image data captured by each of a plurality of cameras as a target camera, and integrates the gazes estimated by the computer for each of the plurality of cameras as the target camera in accordance with how the subject appears in the image data captured by the target camera, thereby identifying the gaze of the subject.

12. A gaze estimation program that causes a computer to function as a gaze estimation device that performs a gaze estimation process that estimates the gaze of a subject from image data captured by each of a plurality of cameras as a target camera, and a gaze integration process that integrates the gazes estimated by the gaze estimation process, each of the plurality of cameras as the target camera, in accordance with how the subject appears in the image data captured by the target camera, to identify the gaze of the subject.

Citation Information

Patent Citations

  • Eye movement tracking method, device and equipment and storage medium

    CN111966219A

  • Device for analyzing human action and recording medium for recording human action analytic program

    JP2001008197A

  • Eye gaze tracking using binocular fixation constraints

    US20150277553A1