Line-of-sight estimation device, gaze training system, and input interface device
The gaze estimation device simplifies the gaze estimation process by using a general-purpose camera and subject-specific training, improving accuracy and reducing costs while maintaining user comfort.
Patent Information
- Application Number
- PCT/JP2025/006376
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2025-02-25
- Publication Date
- 2025-09-04
AI Technical Summary
Existing gaze estimation technologies are complex and require additional mechanisms for estimating the positional relationship between a monitor and the face, making the process cumbersome and costly.
A gaze estimation device that uses a general-purpose camera to estimate the gaze point on a display surface based on captured images, employing a learning model trained with additional data from the subject's face, simplifying the estimation process and reducing costs by eliminating the need for expensive tracking devices.
The device simplifies gaze estimation by accurately determining gaze points without complex tracking systems, using a learning model trained with subject-specific data to enhance estimation accuracy and reduce user discomfort.
Smart Images

Figure JP2025006376_04092025_PF_FP_ABST
Abstract
Description
Gaze estimation device, gaze training system, and input interface device
[0001] The present disclosure relates to a gaze estimation device, a gaze training system, and an input interface device.
[0002] For example, Patent Document 1 discloses a gaze estimation system that includes a display control means for displaying the gaze point at which the subject is gazing so that it moves in a predetermined movement pattern, a detection means for detecting the movement of the subject's eyes from an image of the subject, and a tracking determination means for determining whether the subject's eyes are tracking the gaze point based on the relationship between the movement of the gaze point and the movement of the eyes.
[0003] Furthermore, Patent Document 2 discloses an eye-gaze input device that uses an image having information on gaze coordinate points as learning data, detects a face area from the learning data, and uses a trained model of the gaze area created by machine learning a convolutional neural network for the face area.
[0004] International Publication No. 2021 / 161380 Japanese Patent Application Laid-Open No. 2022-115480
[0005] An aspect of the present disclosure aims to simplify gaze estimation processing.
[0006] A gaze estimation device according to one aspect of the present disclosure includes a display unit including a display surface, a photographing unit that acquires a photographed image, and a processing unit that estimates the position of the gaze point of the estimated subject on the display surface based on the photographed image showing the face of the estimated subject gazing at the display surface and a learning model that takes the photographed image as an input and outputs the position of the gaze point of the estimated subject on the display surface, and outputs information regarding the position of the gaze point, wherein the processing unit includes a learning unit that generates the learning model by machine learning, and the learning unit generates the learning model based on a basic learning model that has been trained using as input the photographed image showing the face of a gazer other than the estimated subject gazing at the display surface, and that has been additionally trained using as input the photographed image showing the face of the estimated subject.
[0007] According to one aspect of the present disclosure, it is possible to simplify the gaze estimation process.
[0008] 13A is a schematic diagram showing the overall configuration of the gaze estimation device according to the first embodiment. FIG. 13B is a block diagram showing the hardware configuration of a processing unit included in the gaze estimation device according to the first embodiment. FIG. 13C is a block diagram showing the functional configuration of a processing unit included in the gaze estimation device according to the first embodiment. FIG. 13D is a diagram explaining a machine learning method used by the gaze estimation device according to the first embodiment. FIG. 13D is a diagram showing a learning data collection image displayed on a display screen of a display unit included in the gaze estimation device according to the first embodiment. FIG. 13C is a flowchart showing an operation of estimating the position of a gaze point by the gaze estimation device according to the first embodiment. FIG. 13D is a diagram showing an image captured by the gaze estimation device according to the first embodiment. FIG. 13D is a diagram showing a gaze target displayed on a display screen in a learning experiment. FIG. 13E is a diagram showing a face area extraction process by a face area extraction unit of the gaze estimation device according to the first embodiment. FIG. 13F is a diagram showing a face area calculation method used in the gaze estimation device according to the first embodiment. FIG. 13F is a diagram showing a gaze estimation error using a basic learning model used by the gaze estimation device according to the first embodiment. FIG. 13G is a diagram showing a gaze estimation error using a learning model used by the gaze estimation device according to the first embodiment. FIG. 13H is a diagram showing a rule explanation screen, which is an example of a data collection application screen in the gaze estimation device according to the first embodiment. FIG. 13H is a diagram showing a circle before contraction, which is an example of a data collection application screen in the gaze estimation device according to the first embodiment. FIG. 13I is a diagram showing a circle after contraction, which is an example of a data collection application screen in the gaze estimation device according to the first embodiment. Fig. 1 is a diagram showing a lattice point arrangement in a patient data collection target arrangement. Fig. 2 is a diagram showing an example of a peripheral weighted arrangement in a patient data collection target arrangement of the gaze estimation device according to the first embodiment. Fig. 3 is a block diagram showing the overall configuration of a gaze training system according to a second embodiment. Fig. 4 is a flowchart showing a gaze training operation by the gaze training system according to the second embodiment.
[0009] Hereinafter, embodiments for carrying out the present disclosure will be described in detail with reference to the drawings. However, the embodiments shown below are examples of a gaze estimation device and a gaze training system for realizing the technical concept of the present disclosure, and are not limited to the following. In addition, in each drawing, the same components are given the same reference numerals, and duplicate explanations will be omitted as appropriate.
[0010] [First embodiment] <Configuration of gaze estimation device according to first embodiment> (Overall configuration) The overall configuration of a gaze estimation device 100 according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a schematic diagram showing an example of the overall configuration of the gaze estimation device 100 according to the first embodiment.
[0011] The gaze estimation device 100 has a display unit 1 including a display surface 10, and a capturing unit 2 that acquires captured images. The gaze estimation device 100 also has a processing unit 3 that estimates the position of the gaze point of the estimation subject U on the display surface 10 based on an image captured by the capturing unit 2 that shows the face of the estimation subject U gazing at the display surface 10, and a learning model, and outputs information E related to the position of the gaze point. The learning model receives an image captured by the capturing unit 2 as an input, and outputs the position of the gaze point of the estimation subject U on the display surface 10.
[0012] The gaze estimation device 100 uses the processing unit 3 to estimate the position of the gaze point on the display surface 10 of the person to be estimated U who is positioned opposite the display surface 10 of the display unit 1, and outputs information E relating to the position of the gaze point of the person to be estimated U to an external device or equipment other than the processing unit 3. The gaze point means a point on the display surface 10 at which a person viewing the display surface 10 gazes. The position of the gaze point corresponds to the gaze direction of the person viewing the display surface 10. The gaze estimation device 100 can estimate the gaze of the person viewing the display surface 10 by estimating the position of the gaze point on the display surface 10 of the person viewing the display surface 10.
[0013] The information E regarding the position of the gaze point of the estimated subject U is, for example, coordinate information indicating the position of the gaze point of the estimated subject U. However, the information E regarding the position of the gaze point of the estimated subject U is not limited to coordinate information, and may be code information or an electrical signal indicating the position of the gaze point of the estimated subject U. The external device is an information processing device such as a PC (Personal Computer), an information terminal such as a smartphone or tablet, or an external server connected to be able to communicate via a network such as the Internet. The external device is a display device such as a display, a memory such as an HDD, etc. Note that the external device may be the display unit 1 or the imaging unit 2.
[0014] As a technology for performing gaze estimation, for example, Patent Document 1 discloses a gaze estimation system including a display control means for displaying a gaze point at which a subject is gazing so as to move in a predetermined movement manner, a detection means for detecting the subject's eye movement from an image of the subject, and a tracking determination means for determining whether the subject's eyes are tracking the gaze point based on the relationship between the movement of the gaze point and the eye movement. Furthermore, Patent Document 2 discloses an eye-gaze input device that uses an image having information on gaze coordinate points as learning data, detects a face area from the learning data, and uses a trained model of the gaze area created by machine learning a convolutional neural network for the face area.
[0015] However, the gaze estimation system described in Patent Document 1 estimates the gaze from the relationship between the movement of the gaze point and the movement of the eyes, which can make the gaze estimation process complicated. Furthermore, the gaze input device described in Patent Document 2 performs machine learning to output the gaze angle of the estimation target. Therefore, in order to estimate the position of the gaze point on the screen of the estimation target viewing a monitor screen using a learning model obtained by machine learning, a mechanism for estimating the positional relationship between the monitor and the face of the estimation target is required in addition to a learning model for estimating the gaze angle of the estimation target. This can make the gaze estimation process complicated.
[0016] The gaze estimation device 100 according to this embodiment estimates the position of the gaze point of the estimation target person U on the display surface 10 based on an image captured by the imaging unit 2, which captures the face of the estimation target person U gazing at the display surface 10, and outputs information E relating to the position of the gaze point. Because the gaze estimation device 100 does not output the gaze angle of the estimation target person U, it can estimate the position of the gaze point of the estimation target person U without using a learning model for estimating the position of the face of the estimation target person U. This allows the gaze estimation device 100 according to this embodiment to simplify the gaze estimation process.
[0017] Furthermore, the gaze estimation device 100 performs gaze estimation based on the image captured by the image capture unit 2, without using an expensive gaze tracking device. This allows a general-purpose camera such as a monocular RGB camera to be used as the image capture unit 2, reducing the introduction cost of the gaze estimation device 100 and allowing the gaze estimation device 100 to be used easily.
[0018] The display unit 1 displays still images or moving images on a display surface 10. The display unit 1 may be a liquid crystal display, an organic EL (Electro Luminescence) display, a CRT (Cathode Ray Tube) display, or the like.
[0019] The image capturing unit 2 captures captured images of an estimated subject U or the like positioned opposite the display surface 10 of the display unit 1. The captured images include still images and videos. The image capturing unit 2 can be a monocular RGB digital camera or video camera that outputs an RGB image as the captured image Im. The display unit 1 and the image capturing unit 2 may be integrated into a single device, or may be separate devices. The gaze estimation device 100 may be a camera-equipped display, smartphone, tablet, PC, notebook PC, or the like, in which the display unit 1 and the image capturing unit 2 are integrated.
[0020] (Configuration of processing unit 3) Next, the configuration of the processing unit 3 included in the gaze estimation device 100 according to the first embodiment will be described with reference to Figs. 2 to 5. Fig. 2 is a block diagram showing an example of the hardware configuration of the processing unit 3. Fig. 3 is a block diagram showing an example of the functional configuration of the processing unit 3. Fig. 4 is a diagram explaining an example of a machine learning method performed by the gaze estimation device 100. Fig. 5 is a diagram showing an example of a learning data collection image Ic displayed on the display surface 10 of the display unit 1.
[0021] 2, the processing unit 3 includes a central processing unit (CPU) 31, a read-only memory (ROM) 32, and a random access memory (RAM) 33. The processing unit 3 also includes a solid state drive (SSD) 34 and a communication interface (I / F) 35. These are connected to each other via a system bus B so as to be able to communicate with each other.
[0022] The CPU 31 executes control processing including various types of arithmetic processing. The ROM 32 stores programs used to drive the CPU 31, such as an IPL (Initial Program Loader). The RAM 33 is used as a work area for the CPU 31. The SSD 34 is a non-volatile storage device that stores various data, programs, etc. The data includes a facial image of the estimated subject U captured by the imaging unit 2. The SSD 34 may also be an HDD (Hard Disk Drive). The communication I / F 35 is an interface for communication between the processing unit 3 and devices or apparatuses other than the processing unit 3. The communication I / F 35 may also communicate with devices or apparatuses other than the processing unit 3 via a network or the like.
[0023] (Functional Configuration) As shown in FIG. 3, the processing unit 3 has an input unit 301, a face area extraction unit 302, a learning unit 303, an estimation unit 304, a notification unit 305, a collection image generation unit 306, and an output unit 307.
[0024] The functions of the input unit 301 and the output unit 307 are realized by the communication I / F 35, etc. Note that some of the functions of the input unit 301 and the output unit 307 may be realized by the CPU 31 executing processing defined in a program stored in the ROM 32, etc. The functions of the face area extraction unit 302, the estimation unit 304, the notification unit 305, and the collection image generation unit 306 are realized by the CPU 31 executing processing defined in a program stored in the ROM 32, etc. The function of the learning unit 303 is realized by the SSD 34 and the CPU 31 executing processing defined in a program stored in the ROM 32, etc.
[0025] Some of the functions of the processing unit 3 may be realized by equipment or devices other than the processing unit 3. Furthermore, some of the functions of the processing unit 3 may be realized by distributed processing between the processing unit 3 and equipment or devices other than the processing unit 3. Furthermore, some of the functions of the processing unit 3 may be realized by one or more processing circuits. This processing circuit is an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or the like designed to execute each of the above functions.
[0026] The input unit 301 controls communication with the imaging unit 2 to input an image Im captured by the imaging unit 2 from the imaging unit 2. The captured image Im is an image, etc., showing the face of the presumed subject U gazing at the display surface 10. However, the captured image Im is not limited to an image showing the face of the presumed subject U gazing at the display surface 10. The captured image Im may also be an image, etc., showing the face of a gazing person other than the presumed subject U gazing at the display surface 10.
[0027] The facial area extraction unit 302 extracts a facial area F of the presumed subject U gazing at the display surface 10 or a gazing person other than the presumed subject U from the image Im captured by the imaging unit 2 input via the input unit 301. The facial area extraction unit 302 passes the extracted facial area F to each of the learning unit 303 and the estimation unit 304.
[0028] The learning unit 303 receives the captured image Im as input and generates a learning model M by machine learning, outputting the position of the gaze point of the estimated subject U on the display surface 10. The learning unit 303 receives as input, for example, the face area F extracted from the captured image Im by the face area extraction unit 302 and the actual gaze coordinates on the display surface 10 of the estimated subject U or a gazer other than the estimated subject U, to generate the learning model M. The processing unit 3 having the learning unit 303 makes it possible to learn the learning model M even in an environment in which communication with an external device such as an external PC or an external server is not possible.
[0029] 4, the learning unit 303 can generate a learning model M by additionally learning using as input a captured image Im showing the face of the estimated subject U, based on a basic learning model Mb that has been learned using as input a captured image Im showing the face of a gazing person other than the estimated subject U gazing at the display surface 10. The additional learning can also be called fine tuning.
[0030] The gazers other than the estimated subject U are, for example, an unspecified number of gazers other than the estimated subject U. By additionally learning using a captured image Im containing the face of the estimated subject U as an input, the estimation accuracy of the position of the gaze point of the estimated subject U can be improved compared to when using a learning model trained using a captured image Im containing the faces of gazers other than the estimated subject U as an input. Furthermore, by using a basic learning model Mb trained using a captured image Im containing the faces of gazers other than the estimated subject U as an input, a large amount of learning can be performed while reducing the burden on the estimated subject U due to photography. This can improve the estimation accuracy of the position of the gaze point of the estimated subject U.
[0031] From the perspective of updating the learning model M using a captured image Im showing the face of the person to be estimated U as input, the learning unit 303 can also update the learning model M by using the basic learning model Mb as a base and performing additional learning using a captured image Im showing the face of the person to be estimated U as input.
[0032] The learning unit 303 does not necessarily need to use the face region F extracted by the face region extraction unit 302 for learning, as long as the captured image Im contains the face of the estimated subject U or a gazing person other than the estimated subject U. However, by using the face region F extracted by the face region extraction unit 302 for learning, the influence of the background region of the face in the captured image Im can be reduced, and the complexity of learning can be reduced.
[0033] 3 estimates the position of the gaze point of the estimated subject U on the display surface 10 based on a captured image Im showing the face of the estimated subject U gazing at the display surface 10 and a learning model M that receives the captured image Im as an input and outputs the position of the gaze point of the estimated subject U on the display surface 10. The estimation unit 304 outputs information E relating to the position of the gaze point of the estimated subject U on the display surface 10 to an external device or the like via the output unit 307.
[0034] The notification unit 305 notifies the estimation subject U that the face of the estimation subject U has shifted from a reference position based on the position of the face of the estimation subject U included in the captured image Im. Here, in order to improve the estimation accuracy of the gaze estimation device, it is preferable to fix the face of the estimation subject U to a certain extent so that the position of the face of the estimation subject U included in the captured image Im does not fluctuate. However, using a face fixing member that fixes the face by resting the chin or the like to prevent movement may impair the comfort of the estimation subject U. In the gaze estimation device 100, the notification unit 305 notifies the estimation subject U that the face of the estimation subject U has shifted from the reference position, thereby prompting the estimation subject U to correct the face position and stabilizing the face position without using a face fixing member or the like. This allows the estimation accuracy of the gaze estimation device 100 to be improved without impairing the comfort of the estimation subject U.
[0035] The notification unit 305 can notify that the face of the presumed target person U has shifted from the reference position by, for example, displaying text information St indicating that the face of the presumed target person U has shifted from the reference position on the display unit 1 or the like via the output unit 307. However, this is not limited to this, and the notification unit 305 may, for example, notify that the face of the presumed target person U has shifted from the reference position by voice using a speaker or the like.
[0036] The collection image generation unit 306 generates a learning data collection image Ic including a plurality of target images Co, each of whose display position changes in response to an instruction input. The learning data collection image Ic is used to efficiently acquire captured images Im and gaze coordinates of an estimated subject U gazing at various positions on the display surface 10 while reducing the burden on the estimated subject U when generating a learning model M. The instruction input is an input signal passed from the processing unit 3 or the like to the collection image generation unit 306 at a predetermined cycle, or an operation input by the operator of the gaze estimation device 100 received by the gaze estimation device 100 via an operation unit or the like of the gaze estimation device 100.
[0037] When generating the learning model M, the processing unit 3 can have the presumption target person U input the captured image Im showing the face of the presumption target person U in a game-like manner. Alternatively, the processing unit 3 can have the presumption target person U input the captured image Im showing the face of the presumption target person U in a game format.
[0038] For example, when generating the learning model M, the processing unit 3 displays the learning data collection image Ic generated from the learning data collection image Ic on the display surface 10 of the display unit 1 via the output unit 307, and causes the estimation subject U to gaze at a predetermined target image Tg among the multiple target images Co. The processing unit 3 inputs, to the learning unit 303, a captured image Im of the estimation subject U gazing at the target image Tg and the position of the target image Tg on the display surface 10.
[0039] The learning data collection image Ic illustrated in Fig. 5 includes a plurality of target images Co. The target images Co represent various objects such as apples, oranges, cakes, rice balls, and globes. In the example illustrated in Fig. 5, among the plurality of target images Co, the target image Co of an apple is predetermined as the target image Tg. The estimated subject U is also informed in advance that the target image Co of an apple is the target image Tg.
[0040] In the learning data collection image Ic, the display positions of the multiple target images Co change in response to instruction input. The estimation subject U searches for a target image Tg among the multiple target images Co, the display positions of which change in response to instruction input, in the learning data collection image Ic displayed on the display surface 10. When the estimation subject U finds the target image Tg, he or she gazes at the target image Tg and notifies the gaze estimation device 100, via an operation unit or the like of the gaze estimation device 100, that the target image Tg has been found. The gaze estimation device 100 inputs the center coordinates of the target image Tg as the actual gaze coordinates of the estimation subject U on the display surface 10, and can acquire a captured image Im in which the estimation subject U gazes at the target image Tg.
[0041] As described above, the processing unit 3 uses the learning data collection image Ic generated by the collection image generation unit 306 to have the estimated subject U gaze at the target image Tg and inputs the center coordinates of the target image Tg as gaze coordinates. Since the estimated subject U is made to find and gaze at the target image Tg, whose display position changes, in a game-like or game-like manner, it is possible to have the estimated subject U gaze at various positions on the display surface 10 while reducing the likelihood of the estimated subject U becoming bored. In this way, the gaze estimation device 100 can efficiently acquire the captured image Im and gaze coordinates of the estimated subject U gazing at various positions on the display surface 10 while reducing the burden on the estimated subject U.
[0042] <Operation of Gaze Estimation Device 100> An operation of estimating the position of a gaze point by the gaze estimation device will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of an operation of estimating the position of a gaze point by the gaze estimation device 100. The gaze estimation device 100 starts the operation of Fig. 6 when, for example, input of captured images Im from the imaging unit 2 is started as a starting condition. After starting the operation of Fig. 6, the gaze estimation device 100 repeatedly inputs captured images Im in accordance with the imaging cycle of the imaging unit 2 until input of captured images Im from the imaging unit 2 is stopped.
[0043] First, in step S11 , the gaze estimation device 100 detects, by the notification unit 305 , the position of the face of the estimation target person U included in the image Im captured by the imaging unit 2 and input via the input unit 301 .
[0044] Subsequently, in step S12, the gaze estimation device 100 determines, via the notification unit 305, whether or not the face of the estimation target person U is at the reference position.
[0045] If it is determined in step S12 that the face of the estimation target person U is displaced from the reference position (step S12, NO), the gaze estimation device 100 causes the notification unit 305 to display text information St indicating that the face of the estimation target person U has displaced from the reference position on the display unit 1 or the like via the output unit 307 in step S13. This notifies the user that the face of the estimation target person U has displaced from the reference position. Thereafter, the gaze estimation device 100 performs the operations from step S11 onwards again.
[0046] On the other hand, if it is determined in step S12 that the face of the presumed subject U is within the reference position (step S12, YES), in step S14, the gaze estimation device 100 causes the face area extraction unit 302 to extract a face area F of the presumed subject U gazing at the display surface 10 or a gazing person other than the presumed subject U, from the image Im captured by the imaging unit 2 input via the input unit 301. The face area extraction unit 302 passes the extracted face area F to the learning unit 303.
[0047] Next, in step S15, the gaze estimation device 100 estimates the position of the gaze point of the estimation target person U on the display surface 10 using the estimation unit 304 based on the face area F in the captured image Im and the learning model M.
[0048] Subsequently, in step S16, the gaze estimation device 100 outputs information E relating to the position of the gaze point of the estimation target person U estimated by the estimation unit 304 to an external device or the like via the output unit 307.
[0049] Next, in step S17, the gaze estimation device 100 determines whether to end the gaze estimation operation. For example, the gaze estimation device 100 determines to end the gaze estimation operation when an operation input indicating the end of the operation is received from the estimation target person U via an operation unit or the like.
[0050] If it is determined in step S17 that the operation should not be terminated (step S17, NO), the gaze estimation device 100 performs the operations from step S11 onwards again. On the other hand, if it is determined in step S17 that the operation should be terminated (step S17, YES), the gaze estimation device 100 stops inputting the captured image Im in step S18. Thereafter, the gaze estimation device 100 terminates its operation.
[0051] In this way, the gaze estimation device 100 can estimate the position of the gaze point of the estimation target person U on the display surface 10 and output information E relating to the position of the gaze point.
[0052] <Learning Experiment> An example of a learning experiment in the gaze estimation device 100 will now be described.
[0053] (Experimental Method) Table 1 shows the main conditions of the display unit 1 and the photographing unit 2 used in the learning experiment. Note that the display in Table 1 corresponds to the display unit 1, and the camera corresponds to the photographing unit 2.
[0054]
[0055] In the learning experiment, the positional relationship between the subject and the imaging unit 2 was measured before the start of imaging. Subject U0 was instructed to avoid extreme changes in facial position and facial expression, and was told that he or she could change the direction of his or her face and eyes. The subject in the learning experiment was either the estimated subject U or a gazer other than the estimated subject.
[0056] FIG. 7 is a diagram showing an example of a captured image Im captured by the gaze estimation device 100. The captured image Im shows the face of subject U0. The resolution of the camera used in the image capture unit 2 is 1280 x 960 (pixels). The image capture time by the image capture unit 2 was 20 to 50 minutes per person. More detailed information regarding the capture of the subjects is shown in Table 2.
[0057]
[0058] Fig. 8 is a diagram showing an example of a gaze point Gp displayed on the display surface 10 in a learning experiment. In the learning experiment, as shown in Fig. 8, a white circle gaze point Gp was displayed on the black display surface 10, and the subject was photographed by the photographing unit 2 while looking at the gaze point. The coordinates (x, y) of the gaze point Gp on the display surface 10 were also recorded. The unit of coordinates is pixels. The coordinates shown hereinafter will also be in pixels.
[0059] In the learning experiment, to confirm that the subject was looking at the gaze point Gp, the subject was asked to indicate the gaze point Gp using a cursor 11 displayed on the display surface 10. When the subject indicated the gaze point Gp with the cursor 11, the subject was photographed by the photographing unit 2, and the next gaze point Gp was displayed on the display surface 10.
[0060] 9 is a diagram showing an example of face region extraction processing by the face region extraction unit 302 of the gaze estimation device 100. In the face region extraction processing in the learning experiment, four face landmark coordinates calculated simultaneously with the face region F, which is a rectangular region, namely the coordinates of the four points of the right eye 81a, the left eye 81b, the right edge of the mouth 82a, and the left edge of the mouth 82b, were used. A calculation method was then devised and used for extracting a square face region that includes the eyes, the mouth, etc.
[0061] First, to define the face region, the centers of gravity of four points, the right eye 81a, the left eye 81b, the right edge of the mouth 82a, and the left edge of the mouth 82b, were determined as the center of region IX in FIG. 9 corresponding to the face region F. Then, the maximum difference (X max -X min ), the maximum difference in the Y direction (Y max -Y min The four points of the right eye 81a, left eye 81b, right edge of the mouth 82a and left edge of the mouth 82b are (X1, Y1), (X2, Y2), (X3, Y3) and (X4, Y4), and are summarized in the following formulas (1) to (6). The minimum value of the X coordinate is X min =MIN(X1, X2, X3, X4) ... Formula (1) Minimum value Y of the Y coordinate min =MIN(Y1, Y2, Y3, Y4) ... Formula (2) Maximum value of X coordinate X max =MAX(X1, X2, X3, X4) ... Equation (3) Maximum Y value of the Y coordinate max =MAX(Y1, Y2, Y3, Y4) ... Formula (4) Length of one side of face area F=1.8×MAX((X max -X min ), (Y max -Y min )) ... Equation (5) Center of face area F=((ΣXi) / 4, (ΣYi) / 4) ... Equation (6)
[0062] FIG. 10 is a diagram illustrating an example of a face region calculation method in the gaze estimation device 100. FIG. 10 is an enlarged view of region IX in FIG. 9. Size Fx indicates the size of face region F in the horizontal direction, and size Fy indicates the size of face region F in the vertical direction perpendicular to the horizontal direction. Face region center 80 corresponds to the center of face region F. The coefficient "1.8" was empirically determined so that both eyes of subject U0 are captured in the captured image Im while reducing the background. In the learning experiment, after extracting face region F, it was resized to 224 x 224 (pixels) to match the input size of learning model M.
[0063] The face area F and correct gaze coordinate values obtained as described above were used as learning inputs to train the learning model M. The learning model used was ResNet34 (Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun, "Deep Residual Learning for Image Recognition," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016) included in the PyTorch library (A. Paszke et al., "PyTorch: An imperative style high-performance deep learning library," NIPS'19: Proceedings of the 33rd International Conference on Neural Information Processing Systems, Curran Associates Inc, Red Hook, NY, USA, pp. 8026-8037, 8 December 2019).
[0064] The correct value to be learned was the coordinate (XX, YY) of the gaze point Gp on the display surface 10 normalized by the width and height of the display surface 10. The output obtained in the learning experiment was also a normalized value, and the value multiplied by the width (1920 pixels) and height (1080 pixels) of the screen was used to determine the estimated position of the gaze point.
[0065] In the learning experiment, the evaluation index was the gaze point (XX, YY) on the display surface 10 at the time of shooting and the estimated gaze position (X est , Y est ) and the error in Euclidean distance between the two points. The error in Euclidean distance is shown in Equation 7. This error in Euclidean distance will be referred to as the position error hereafter. est ) 2 + (Y-Y est ) 2} ... Formula (7)
[0066] 11 is a diagram showing an example of gaze estimation error using the basic learning model Mb by the gaze estimation device 100. For a display resolution of 1920 × 1080 pixels of the display unit 1, the average error in gaze estimation using the basic learning model Mb was 158.3 (pixels), the median was 121.8 (pixels), the minimum was 4.1 (pixels), and the maximum was 1354.6 (pixels).
[0067] 12 is a diagram showing an example of gaze estimation error using learning model M by gaze estimation device 100. For a display resolution of 1920 × 1080 (pixels) of display unit 1, the average error in gaze estimation using learning model M was 60.9 (pixels), the median was 46.6 (pixels), the minimum was 1.0 (pixels), and the maximum was 1376.0 (pixels).
[0068] In gaze estimation using the learning model M by the gaze estimation device 100, the average error was smaller for all subjects U0 than in gaze estimation using the basic learning model Mb, and the overall error was less than half.
[0069] <Evaluation Results of Gaze Estimation Accuracy Using the Gaze Estimation Device 100> Table 3 shows the evaluation results of gaze point error for multiple subjects and the average. To improve the gaze point error, the data collection application can be modified to encourage subjects to focus their gaze more on the target. Figures 13A, 13B, 13C, and 13D are diagrams showing an example of a data collection application screen in the gaze estimation device 100. Figure 13A is a diagram showing a rule explanation screen. Figure 13B is a diagram showing a circle before contraction. Figure 13C is a diagram showing a circle after contraction. Figure 13D is an enlarged view of the circle in Figure 13C. In the examples shown in Figures 13A, 13B, 13C, and 13D, the target changes over time. The subject can click when the target reaches a predetermined state. This allows the subject to focus their gaze more on the target.
[0070] Table 4 shows the evaluation results of the gaze estimation accuracy of the gaze estimation device 100. In Table 4, the results of the generic model show the average of five-fold cross-validation, and the results of the patient model show the average of the estimation results for four subjects. The generic model is a model trained on six subjects. The patient model is a model that was fine-tuned for each patient. The central 60% of the screen was defined as the center, and the area around the center was defined as the periphery. Equal estimation accuracy was obtained in the central and periphery.
[0071] FIG. 14A is a diagram showing the arrangement of grid points in the patient data collection target arrangement. FIG. 14B is a diagram showing an example of a peripherally weighted arrangement in the patient data collection target arrangement. In the gaze estimation device 100, for example, the patient data collection target arrangement can be the peripherally weighted arrangement shown in FIG. 14B, and the model input image can be the area of both eyes. Previous research has shown that the accuracy of patient models is lower in the peripheral areas than in the central area. In contrast, the gaze estimation device 100 can improve accuracy in the peripheral areas and reduce the difference from the central area by considering the effects of data expansion that shifts the patient data collection target arrangement, the model input image, and the trimming position, as shown in Table 4.
[0072] The gaze estimation device 100 can implement a learning method for a gaze estimation model that utilizes the dominant eye. Just as humans have a dominant hand, one of the two eyes also has a dominant role in vision. The dominant eye refers to the eye that is dominant in a human's two eyes. The dominant eye plays an important role in processing visual information and spatial recognition. The gaze estimation device 100 can implement a learning method for a gaze estimation model that utilizes the dominant eye and add information about the dominant eye to gaze estimation, thereby improving the accuracy of gaze estimation.
[0073] Second Embodiment Next, a description will be given of a gaze training system according to a second embodiment. Note that the same names and symbols as those in the already described embodiments indicate the same or similar members or configurations, and detailed descriptions thereof will be omitted as appropriate.
[0074] Fig. 15 is a block diagram showing an example of the overall configuration of a gaze training system 200 according to the second embodiment. The gaze training system 200 includes a gaze estimation device 100. The gaze training system 200 shown in Fig. 15 also includes a control unit 110 and an operation unit 120. The control unit 110 controls the overall operation of the gaze training system 200. The control unit 110 includes a CPU, ROM, RAM, SSD, a communication I / F, and the like. The operation unit 120 accepts operations by an operator, such as a trainer, who operates the gaze training system 200.
[0075] The gaze training system 200 trains the gaze behavior of the estimation target U using information E on the position of the gaze point output from the gaze estimation device 100. The gaze training system 200 can perform quantitative evaluations with the aim of enabling social anxiety disorder patients to make eye contact face-to-face through step-by-step training. Here, social anxiety disorder refers to a mental illness characterized by significant fear and anxiety in social situations where the patient may be gazed at by others. Eye contact is a behavior that plays an important role as a means of communication in social situations. In the gaze training system 200, a trainer uses the operation unit 120 to perform operations such as starting and temporarily pausing gaze training. The subject can receive gaze training by following instructions displayed on the training screen.
[0076] By including the gaze estimation device 100, the gaze training system 200 can simplify the gaze estimation process and can simplify processes such as determining whether eye contact has been made based on information E related to the position of the gaze point, which is the result of the gaze estimation process. The gaze training system 200 also has the advantage of being simple in configuration and easily portable. Furthermore, the gaze training system 200 has the advantage that an operator can start and stop gaze training with a simple operation, and no specialized knowledge is required.
[0077] Fig. 16 is a flowchart showing an example of the gaze training operation by the gaze training system 200. The gaze training system 200 starts the operation in Fig. 16 when a training start operation is received by an operator such as a trainer using the operation unit 120, as a start condition.
[0078] First, in step S21, the gaze training system 200 starts inputting the captured image Im from the imaging unit 2 by the gaze estimation device 100. Thereafter, the gaze estimation device 100 repeatedly inputs the captured image Im according to the imaging cycle of the imaging unit 2 until the input of the captured image Im from the imaging unit 2 is stopped.
[0079] Next, in step S22, the gaze training system 200 starts additional learning by inputting a captured image Im showing the face of the estimated subject U using the gaze estimation device 100. The estimated subject U corresponds to a trainee undergoing gaze training. After the start of the additional learning, the gaze estimation device 100 repeatedly inputs a captured image Im showing the face of the estimated subject U and performs additional learning.
[0080] Next, in step S23, the gaze training system 200 determines whether or not to end the additional learning using the control unit 110. For example, the control unit 110 determines to end the additional learning when 100 captured images Im showing the face of the estimation target person U have been input to the learning model M. Note that the number of captured images Im in the gaze training system 200, 100, is a suitable number for ensuring high gaze estimation accuracy while reducing the burden on the trainee.
[0081] If it is determined in step S23 that the additional learning should not be terminated (step S23, NO), the gaze training system 200 continues to input the captured image Im containing the face of the person to be estimated U and perform additional learning using the gaze estimation device 100, and repeats this process until it is determined in step S22 that the additional learning should be terminated.
[0082] On the other hand, if it is determined in step S23 that the additional learning is to be ended (step S23, YES), in step S24, the gaze training system 200 causes the control unit 110 to display an instruction image on the display unit 1 of the gaze estimation device 100. The instruction image includes a gaze point and means an image that instructs the trainee to gaze at the gaze point.
[0083] Subsequently, in step S25, the gaze training system 200 causes the control unit 110 to display a stimulating video on the display unit 1 of the gaze estimation device 100. The gaze training system 200 indicates a gaze point using the indication image displayed in step S24, and causes the trainee to gaze at the indicated gaze point when the stimulating video is played back in step S25.
[0084] Next, in step S26, the gaze training system 200 estimates the position of the gaze point of the trainee on the display surface 10 of the display unit 1 by the gaze estimation device 100. The gaze estimation device 100 outputs information E relating to the estimated position of the gaze point to the control unit 110.
[0085] Subsequently, in step S27, the control unit 110 of the gaze training system 200 determines whether or not the trainee has made eye contact with the gaze point indicated by the instruction image.
[0086] Subsequently, in step S28, the gaze training system 200 causes the control unit 110 to measure the cumulative time.
[0087] Subsequently, in step S29, the gaze training system 200 provides feedback of the determination result of step S27 to the trainee. The gaze training system 200 may provide feedback by displaying characters on the display surface 10 of the gaze estimation device 100 using the control unit 110, or may provide feedback by voice.
[0088] Next, in step S30, the gaze training system 200 determines whether or not to end the training through the control unit 110. For example, when one stage is defined as performing a series of actions from step S24 to step S29, the gaze training system 200 can determine that the training should end when 40 stages have been performed.
[0089] If it is determined in step S30 that the training should not be terminated (step S30, NO), the gaze training system 200 performs the operations from step S24 onwards again, and repeats this until it is determined in step S30 that the training should be terminated.
[0090] On the other hand, if it is determined in step S30 that the training is to be completed (YES in step S30), in step S31, the gaze training system 200 causes the control unit 110 to display the training results on the display surface 10 of the gaze estimation device 100. The training results include the determination result of whether or not the trainee was able to make eye contact in each stage, the accumulated training time, etc.
[0091] Subsequently, in step S32, the gaze training system 200 causes the control unit 110 to stop the input of the captured image Im by the imaging unit 2. After the input of the captured image Im has stopped, the gaze training system 200 ends its operation.
[0092] As described above, the gaze training system 200 can provide gaze training to the trainee.
[0093] Although the preferred embodiments have been described in detail above, the present invention is not limited to the above-described embodiments, and various modifications and substitutions can be made to the above-described embodiments without departing from the scope of the claims.
[0094] All ordinal numbers, quantitative numbers, and other figures used in the description of the embodiments are provided as examples to specifically explain the technology of the present disclosure, and the present disclosure is not limited to the illustrated figures. Furthermore, the connection relationships between components are provided as examples to specifically explain the technology of the present disclosure, and do not limit the connection relationships that realize the functions of the present disclosure.
[0095] The division of blocks in the functional block diagram is an example, and multiple blocks may be realized as a single block, one block may be divided into multiple blocks, or some functions may be moved to another block. Furthermore, the functions of multiple blocks with similar functions may be processed in parallel or time-shared by a single piece of hardware or software. Furthermore, some or all of the functions may be distributed across multiple computers.
[0096] The gaze estimation device according to the present disclosure can simplify the gaze estimation process and is therefore suitable for use in medical applications such as gaze training systems and systems for treating social anxiety disorder. It can also be used in input interface devices that use an operator's gaze as operational input. For example, an input interface device includes the gaze estimation device 100 according to the first embodiment and uses information E regarding the position of the gaze point output from the gaze estimation device 100 as operational input. Using information E regarding the position of the gaze point as operational input can simplify the configuration of the interface device. Furthermore, since operational input can be performed without using hands, a highly convenient interface device can be provided. However, an input interface device including the gaze estimation device 100 can also be used in conjunction with an input interface device that uses hands, such as a keyboard or mouse. An input interface device using the gaze estimation device can be used for operational input in information processing devices such as PCs, information terminals such as smartphones, or display devices such as displays or head-mounted displays. The input interface device may also be mounted on a moving object such as an automobile, train, ship, or airplane and used to operate the moving object. However, the gaze estimation device according to the present disclosure is not limited to the above-mentioned applications and can be used for a variety of applications.
[0097]
[0013] Aspects of the present disclosure are as follows, for example. <1> A gaze estimation device including: a display unit including a display surface; a capturing unit that captures captured images; and a processing unit that estimates a position of a gaze point of the estimation subject on the display surface based on the captured images showing the face of the estimation subject gazing at the display surface and a learning model that takes the captured images as input and outputs the position of the gaze point of the estimation subject on the display surface, and outputs information about the position of the gaze point, wherein the processing unit includes a learning unit that generates the learning model by machine learning, and the learning unit generates the learning model based on a basic learning model that has been trained using captured images showing the faces of gazers other than the estimation subject gazing at the display surface as input, and that has been additionally trained using captured images showing the face of the estimation subject as input. <2> The gaze estimation device according to <1>, further including a notification unit that notifies that the face of the estimation subject has shifted from a reference position, based on the position of the face of the estimation subject included in the captured images. <3> The gaze estimation device according to <1> or <2>, wherein the processing unit causes the estimated subject to input the captured image showing the face of the estimated subject in a game-like manner. <4> The gaze estimation device according to <1> or <2>, wherein the processing unit causes the estimated subject to input the captured image showing the face of the estimated subject in a game-like manner. <5> The gaze estimation device according to <3> or <4>, wherein, when generating the learning model, the processing unit displays a learning data collection image on the display surface, the image including a plurality of target images, each of whose display position changes in response to an instruction input, causes the estimated subject to gaze at a predetermined target image among the plurality of target images, and inputs the captured image of the estimated subject gazing at the target image and the position of the target image on the display surface to the learning unit. <6> A gaze training system includes the gaze estimation device according to any one of <1> to <5>, and trains the gaze behavior of the estimated subject using information about the position of the gaze point output from the gaze estimation device. <7> An input interface device having the gaze estimation device according to any one of claims 1 to 5, wherein information relating to a position of the gaze point output from the gaze estimation device is used as an operation input by the person to be estimated.
[0098] This application claims priority based on Japanese Patent Application No. 2024-026661 filed with the Japan Patent Office on February 26, 2024, and includes the entire contents of these Japanese patent applications.
[0099] REFERENCE SIGNS LIST 1 display unit 10 display surface 11 cursor 2 shooting unit 3 processing unit 31 CPU 32 ROM 33 RAM 34 SSD 35 communication I / F 301 input unit 302 face area extraction unit 303 learning unit 304 estimation unit 305 notification unit 306 collection image generation unit 307 output unit 80 face area center 81a right eye 81b left eye 82a right edge of mouth 82b left edge of mouth 100 gaze estimation device 110 control unit 120 operation unit 200 gaze training system B system bus Co target image E information on position of gaze point F face area Fx size in horizontal direction Fy size in vertical direction Gp gazed point Ic learning data collection image Im captured image M learning model Mb Basic learning model St Character information Tg Target image U Estimated subject U0 Subject
Claims
1. A gaze estimation device comprising: a display unit including a display surface; a photographing unit that acquires photographed images; and a processing unit that estimates the position of the gaze point of the estimated subject on the display surface based on the photographed image showing the face of the estimated subject gazing at the display surface and a learning model that takes the photographed image as input and outputs the position of the gaze point of the estimated subject on the display surface, and outputs information regarding the position of the gaze point, wherein the processing unit has a learning unit that generates the learning model by machine learning, and the learning unit generates the learning model based on a basic learning model that has been trained using as input the photographed image showing the face of a gazing subject other than the estimated subject gazing at the display surface, and has been additionally trained using as input the photographed image showing the face of the estimated subject.
2. The gaze estimation device according to claim 1, further comprising a notification unit that notifies that the face of the person to be estimated has shifted from a reference position based on the position of the face of the person to be estimated included in the captured image.
3. The gaze estimation device according to claim 1, wherein the processing unit causes the person to be estimated to input the captured image showing the face of the person to be estimated in a game-like manner.
4. The gaze estimation device according to claim 1, wherein the processing unit causes the person to be estimated to input the captured image showing the face of the person to be estimated in a game format.
5. The gaze estimation device of claim 4, wherein, when generating the learning model, the processing unit displays on the display surface a learning data collection image including a plurality of target images whose display positions change according to instruction input, has the estimated subject gaze at a predetermined target image among the plurality of target images, and inputs to the learning unit the captured image of the estimated subject gazing at the target image and the position of the target image on the display surface.
6. A gaze training system having the gaze estimation device according to any one of claims 1 to 5, which trains the gaze behavior of the person being estimated using information relating to the position of the gaze point output from the gaze estimation device.
7. An input interface device having the gaze estimation device according to any one of claims 1 to 5, wherein information relating to the position of the gaze point output from the gaze estimation device is used as an operational input by the person to be estimated.
Citation Information
Patent Citations
Line-of-sight detection device, and correction coefficient determination program
JP2012065781A
Information processor, information processing method, and computer-readable recording medium recorded with program
JP2017134558A
Subject recognition device
JP2020106552A
Program, information processing device, and terminal device
JP2021145785A
Line-of-sight input device and line-of-sight input method
JP2022115480A